ExamJuanReview. Prepare. Pass.

Engineering Mathematics · Lesson 15 of 28

Multivariable Calculus and Optimization

Push functions of several variables to their extremes the way the CELE tests it, from partial derivatives and the total differential for error estimation to locating and classifying critical points with the second-derivative test and solving side-constrained problems with Lagrange multipliers, every rule carried to a number.

15 min read · Super EaFree lesson

Optimization is where multivariable calculus earns its keep on the MSTE paper: the shortest distance, the least material, the largest area, the safest column, each is a function of several variables pushed to an extreme. This lesson builds the machinery in order, functions of several variables, partial derivatives, and the total differential that carries measurement error, then spends the bulk of its time on the payoff, locating and classifying critical points with the second-derivative test and handling side conditions with Lagrange multipliers. Every rule is carried all the way to a number so you can check your own habits against it.

Functions of several variables

A function of two variables, z = f(x, y), assigns one output to each input pair (x, y). Its graph is no longer a curve but a surface floating above the xy-plane, with the height z read straight up. A right circular cylinder's volume V = pi r^2 h, a beam's deflection as a function of load and span, and the cost of a footing as a function of width and depth are all functions of several variables. Two habits make them readable. First, freeze all but one input and you recover an ordinary single-variable function, a slice through the surface. Second, set the output to a constant, f(x, y) = c, and you get a level curve, the set of inputs that share one height, exactly like a contour line on a topographic map.

Worked example: For f(x, y) = 3x^2 y - 2x + y^3, the value at (1, 2) is 3(1^2)(2) - 2(1) + 2^3 = 6 - 2 + 8 = 12. The level curve through that point is the set of (x, y) with 3x^2 y - 2x + y^3 = 12, one contour on the surface at height 12.

Partial derivatives

A partial derivative measures the rate of change of f when only one input is allowed to move. To find partial f / partial x you differentiate f with respect to x while holding y (and any other variable) constant; for partial f / partial y you hold x constant. Nothing new is memorized: the power, product, quotient, and chain rules all still apply, and every frozen variable rides along as a constant.

Worked example (first partials): For f(x, y) = 3x^2 y - 2x + y^3, holding y constant gives f_x = 6xy - 2, and holding x constant gives f_y = 3x^2 + 3y^2. At the point (1, 2), f_x = 6(1)(2) - 2 = 12 - 2 = 10 and f_y = 3(1^2) + 3(2^2) = 3 + 12 = 15. Worked example (second partials): Differentiate f_x = 6xy - 2 again in x to get f_xx = 6y, and in y to get the mixed partial f_xy = 6x. Starting instead from f_y = 3x^2 + 3y^2 gives f_yx = 6x, the same result. This equality, f_xy = f_yx whenever the second partials are continuous, is Clairaut's theorem, and it is a fast self-check: if your two mixed partials disagree, one of them is wrong.

The two first partials collected into a vector form the gradient, grad f = (f_x, f_y). At (1, 2) above it is (10, 15). The gradient points in the direction of steepest ascent of the surface, its length is the maximum rate of increase, and it is always perpendicular to the level curves. Setting it to the zero vector is exactly the condition that locates the flat spots we optimize over, which is the whole point of the second half of this lesson.

The total differential and error estimation

Small changes in the inputs produce a small change in the output, and the total differential is the linear estimate of that change:

dz = f_x dx + f_y dy.

Read it as bookkeeping: each input contributes its own partial derivative times how far that input moved. This is precisely how a measured quantity carries its measurement error forward. If z = f(x, y) and x and y are known only to within dx and dy, then dz = f_x dx + f_y dy estimates the resulting uncertainty in z. Dividing through by the quantity turns absolute error into relative (percent) error, and for a product of powers that relative error separates into a clean weighted sum in which each variable's exponent is its weight.

Worked example: The power dissipated by a resistor is P = I^2 R, measured at I = 2.0 A and R = 100 ohm, so P = (2.0^2)(100) = 400 W. Suppose the current is uncertain by dI = 0.05 A and the resistance by dR = 2 ohm. The partials are partial P / partial I = 2 I R and partial P / partial R = I^2, so

dP = 2 I R dI + I^2 dR = 2(2.0)(100)(0.05) + (2.0^2)(2) = 20 + 8 = 28 W.

The relative form is cleaner still. Dividing dP by P = I^2 R gives dP / P = 2 dI / I + dR / R = 2(0.05 / 2.0) + (2 / 100) = 2(0.025) + 0.02 = 0.05 + 0.02 = 0.07, so the power is uncertain by about 7 percent. Notice the current enters with a factor of 2 because it appears squared in the formula, the single most useful pattern in error propagation.

Critical points

An unconstrained maximum or minimum of a smooth surface sits at a flat spot, a point where the tangent plane is horizontal. That happens exactly when both first partials vanish at the same time:

f_x = 0 and f_y = 0, that is, grad f = (0, 0).

Such a point is called a critical point. Solving the two equations together gives the candidate locations; there may be one, several, or none. A local maximum is a peak, a local minimum is a valley bottom, but a third possibility has no single-variable analog: a saddle point, where the surface rises in one direction and falls in the perpendicular one, so the tangent plane is still horizontal yet the point is neither a peak nor a valley.

At a local maximum the tangent plane is horizontal and grad f equals zero z x y local maximum grad f = 0 horizontal tangent plane at the peak
A local maximum of z = f(x, y) is a peak where the tangent plane lies flat, so both first partials are zero and grad f = (0, 0). A local minimum looks the same upside down, and a saddle point also has grad f = 0 but rises one way and falls the other.

Worked example: The dome f(x, y) = 4 - x^2 - y^2 has f_x = -2x and f_y = -2y, both zero only at (0, 0), so the single critical point is the origin, where f = 4. Since the surface falls away in every direction from there, that critical point is a local maximum with value 4, the top of the dome in the figure.

Classifying critical points: the second-derivative test

Finding grad f = 0 locates the flat spots but does not say which kind each is. The second-derivative test settles it using the three second partials assembled into the discriminant:

D = f_xx f_yy - (f_xy)^2.

Evaluate D and f_xx at the critical point and read off the verdict. A positive D means the surface curves the same way in the x and y directions, so it is a genuine peak or valley, and the sign of f_xx then says which. A negative D means the two directions curve oppositely, the signature of a saddle. When D = 0 the test is silent and you must look more closely.

Discriminant D = f_xx f_yy - (f_xy)^2 Sign of f_xx Conclusion at the critical point
D > 0 f_xx > 0 local minimum
D > 0 f_xx < 0 local maximum
D < 0 either sign saddle point
D = 0 either sign test inconclusive, examine the surface directly

Worked example: Classify every critical point of f(x, y) = x^3 + y^3 - 3xy. The first partials are f_x = 3x^2 - 3y and f_y = 3y^2 - 3x. Setting f_x = 0 gives y = x^2, and substituting into f_y = 0 gives 3x^4 - 3x = 0, that is x(x^3 - 1) = 0, so x = 0 or x = 1. The critical points are (0, 0) and (1, 1). The second partials are f_xx = 6x, f_yy = 6y, and f_xy = -3, so D = (6x)(6y) - (-3)^2 = 36xy - 9. At (0, 0), D = 0 - 9 = -9 < 0, so the origin is a saddle point. At (1, 1), D = 36 - 9 = 27 > 0 and f_xx = 6 > 0, so it is a local minimum, with value f(1, 1) = 1 + 1 - 3 = -1. One function, two critical points, two different verdicts, exactly what the test is for.

Constrained optimization by Lagrange multipliers

Most board optimization problems come with a side condition: the least fence for a fixed area, the largest area for a fixed perimeter, the point on a given line nearest the origin. You are no longer free to roam the whole surface; you must stay on the constraint curve g(x, y) = 0. The Lagrange multiplier method turns this into algebra. At a constrained optimum the level curve of f just touches the constraint curve, so the two curves are tangent there and their gradients point along the same line. Introducing a scalar lambda for that proportionality gives the working equations:

f_x = lambda g_x, f_y = lambda g_y, g(x, y) = 0.

That is three equations in the three unknowns x, y, and lambda. Solve them together, evaluate f at each candidate, and pick the largest for a maximum or the smallest for a minimum.

At a constrained optimum the level curve of f is tangent to the constraint and the gradients are parallel f = c g(x, y) = 0 P grad f grad g
The constraint line g = 0 is tangent to the optimal level curve f = c at P. Because a gradient is perpendicular to its own level curve, grad f and grad g both point along the common normal at P, so grad f = lambda grad g; the multiplier lambda is just the ratio of their lengths.

Worked example (a maximum): A rectangular sedimentation basin is to be built from 40 m of wall, so 2x + 2y = 40, and we want the largest floor area A = xy. Write the constraint as g = x + y - 20 = 0, with grad g = (1, 1) and grad A = (y, x). The Lagrange equations y = lambda and x = lambda force x = y, and the constraint x + y = 20 then gives x = y = 10. The maximum area is A = (10)(10) = 100 m^2. The square shape is no accident: for a fixed perimeter the rectangle of greatest area is always a square.

Worked example (a minimum): Find the point on the line x + 2y = 5 nearest the origin, that is, minimize the squared distance f = x^2 + y^2 subject to g = x + 2y - 5 = 0. Here grad f = (2x, 2y) and grad g = (1, 2), so 2x = lambda and 2y = 2 lambda, which gives y = 2x. Substituting into x + 2y = 5 yields x + 4x = 5, so x = 1 and y = 2. The minimum squared distance is 1^2 + 2^2 = 5, and the shortest distance itself is sqrt(5) = 2.236 m, which matches the standard formula |c| / sqrt(a^2 + b^2) = 5 / sqrt(1 + 4) = 5 / sqrt(5) = sqrt(5).

A surface picture makes the difference from the free case plain. Unconstrained, you seek the top of the whole hill. Constrained, you must walk only along the path that the constraint traces across the hill, and the constrained maximum is the highest point of that path, which is generally lower than, and to one side of, the free summit.

The constrained maximum is the highest point of the path the constraint traces on the surface free summit constrained max constraint g(x, y) = 0
The hollow dot at the top is the free maximum of f. Restricted to the path that g = 0 traces over the surface, the best you can reach is the filled dot, the highest point of that path; its footprint on the base plane sits on the constraint curve directly below.

The multiplier lambda is not just scaffolding. If the constraint is written g(x, y) = c, then lambda equals the rate of change of the optimal value of f with respect to c, the amount the best answer improves when you relax the constraint by one unit. Engineers read it as a shadow price: how much more area one more meter of wall would buy, for instance.

Unconstrained problem Constrained problem (one side condition g = 0)
What you solve f_x = 0 and f_y = 0 f_x = lambda g_x, f_y = lambda g_y, g = 0
Unknowns x, y x, y, lambda
Geometry at the optimum horizontal tangent plane, grad f = 0 level curve of f tangent to g = 0, grad f = lambda grad g
How to classify second-derivative test with D compare f at every candidate, plus any endpoints

Exam-day strategy

  • To take a partial derivative, cover every variable except the one you are differentiating and treat the rest as plain constants; the constant multiple just comes along for the ride.
  • Use f_xy = f_yx as a free check on any second-partial problem: compute the mixed partial both ways, and if they disagree you have an arithmetic error to hunt down.
  • For error propagation, prefer the relative form dz / z, where a product of powers becomes a weighted sum such as 2 dI / I + dR / R; the exponent of each variable is exactly its weight, so a squared term contributes double.
  • Locate critical points by solving f_x = 0 and f_y = 0 together, then classify each one with D = f_xx f_yy - (f_xy)^2 before trusting it: D < 0 is always a saddle, and a horizontal tangent plane alone never proves you have a maximum.
  • On D > 0, let the sign of f_xx break the tie, positive for a valley (minimum) and negative for a peak (maximum); when D = 0 the test is useless, so reason from the surface itself.
  • For a constrained problem, write the three Lagrange equations f_x = lambda g_x, f_y = lambda g_y, g = 0 first, then eliminate lambda by dividing the first two equations, which usually collapses to a simple relation like x = y before you ever touch the constraint.
  • When a problem asks for an absolute maximum or minimum on a closed, bounded region, check the interior critical points and the boundary separately, since the extreme value can hide on an edge or corner where grad f is not zero.

Marking it done updates your Exam-Ready progress.

Lesson quiz

Check you actually have it

20 items on this lesson alone, randomized each try, with the reasoning on every answer.

Multivariable Calculus and Optimization: quick check

Item 01 / 20 · Score 0

Lagrange multipliers

Minimize f(x, y) = x^2 + y^2 subject to x + 2y = 5. What is the minimum value of f?

This whole first section is free

Read every lesson in Engineering Mathematics and take its quizzes free. The full CELE reviewer unlocks the other 5 subjects, all section tests, and the timed mock exams — one payment, lifetime access, ₱399.

Unlock the full reviewer