Tag: math

  • Chain Rule: An In-Depth Analysis

    Chain Rule: An In-Depth Analysis

    The Chain Rule is one of the 4 or 5 main derivative rules that allow you to break down a function into smaller functions in order to find that function’s derivative. As with the other 4 or 5 main derivative rules, the Chain Rule is only applicable under certain conditions. Unlike the other derivative rules, it can be a little difficult to understand and recognize when you can actually use it. If you notice, the phrasing is “can” use it. Remember, don’t think of these derivative rules as things you HAVE TO USE to solve a problem, but rather as tools that can be useful for solving a problem that you may or may not use depending on the details of the problem and what you feel is the best way to approach it.

    When Can You Use the Chain Rule?

    You “can” use the chain rule (but may not “have to”) when the function you are tasked with taking the derivative of “can be thought of” as (again, don’t think of it as “it is this”; think of it as “it can be thought of as this” since some functions can be thought of as being constructed multiple ways) a function with a function inside of it. What does this mean? Well remember, a function is just some rule that maps inputs (the totality of all possible inputs for a function is called that function’s Domain; this input variable is typically going to be x, but it doesn’t have to be) to outputs (the totality of all possible outputs of a function is called that function’s Range; this output variable is typically going to be y, but it doesn’t have to be). In a calculus class, the Domain of a function will usually consist of a continuous interval of input values (examples of this would be something like All Real Numbers, or x can be any real number from 2 to 100), and the Range of the function will usually consist of a continuous or piece-wise continuous interval of output values. This relationship is usually written as a mathematical expression:

    f(x)=x2−2x+1f(x)=x^{2}-2x+1

    This mathematical relationship is what maps the inputs to the outputs. If you want to know what a number in the domain produces in the function’s range, or set of output values, all you need to do is “plug it in” (I prefer to use the word “substitute”, but to each, their own) to the mathematical expression that defines the function. As the mathematical equation specifies, f(x) is equal to that. This mathematical expression is like a machine that transforms the input number (or expression, as we’ll see later) into an output number (or expression, as we’ll see later). For example, using the previously defined function, let’s determine what the function value is for the input value of 2:

    f(2)=(2)2−2(2)+1=1f(2)=(2)^{2}-2(2)+1=1

    Notice the notation here. f(2) means “the function value when the input is 2”. If you notice, all we did here was replace or substitute all of the original input with the new input. Think of the original function with x as a blank template just showing you the steps you do for any value of x. Don’t like the variable x? Experiment around and put different placeholders. Whatever floats your boat. It’s just a variable. For example, another way to write the original function is:

    f()=()2−2()+1f(\hspace{0.5cm})=(\hspace{0.5cm})^{2}-2(\hspace{0.5cm})+1

    Or

    f(whatever)=(whatever)2−2(whatever)+1f(\text{whatever})=(\text{whatever})^{2}-2(\text{whatever})+1

    These all mean the same thing, because they all are telling you to do the same thing. This is an important thing to remember when you are taking a calculus class or any other advanced math class. Things that don’t look the exact same thing written down or described may be the exact same thing as far as we’re concerned because they are mathematically equivalent. In other words, they mean the exact same thing mathematically. So be on the lookout for and practice being able to recognize functions or theorems or problems that look different but are really mathematically equivalent to one another. So nothing too difficult yet, right? Where does the chain rule come into all this?

    Now we will take it one step further. Remember, the chain rule can be used to differentiate (take the derivative of) a function that can be thought of as smaller functions that have been put inside one another (these are called composite functions). Let’s look at how these are formed. To form a composite function, you don’t put a single number into a function like we did before, you put in an entire other function. For example, let’s take the original function we had before and introduce a new one as well:

    g(x)=1xg(x)=\frac{1}{x}

    Let’s find f(g(x)). What does this even mean? It looks pretty weird. Well, if you know how to find f(2), or f(3), or f( whatever ), you know how to find f(g(x)). Just replace the original variable in f(x) with g(x), just like the symbols f(g(x)) indicates and just like we did before with all the other substitutions into the original function:

    f()=()2−2()+1f(\hspace{0.5cm})=(\hspace{0.5cm})^{2}-2(\hspace{0.5cm})+1

    Therefore:

    f(g(x))=(g(x))2−2(g(x))+1f(g(x))=(g(x))^{2}-2(g(x))+1

    And:

    f(g(x))=f(1x)=(1x)2−2(1x)+1f(g(x))=f\left(\frac{1}{x}\right)=\left(\frac{1}{x}\right)^{2}-2\left(\frac{1}{x}\right)+1

    This is just what some people would consider a slightly more complicated version of what we’ve already done, but someone who can see through the clutter and really cut to the heart of what these ideas mean would see it as no different than substituting in a single number. So to get to the burning question of when you can use the chain rule, the answer is whenever you have a function that you can visualize as a composite function (a function within a function). This new function that we created f(g(x)) is an example of a composite function. But this is where it gets a little more difficult. When using the chain rule, you need to recognize when a big function is composed of two smaller functions where one has been put inside the other. This is where some students get tripped up because this can require a little bit of imagination. You have to take on the role of a detective investigating a crime scene. The detective wasn’t there when the crime took place, so he/she didn’t observe exactly what happened. Instead, they need to see the results of what happened and determine what series of events could have caused that end result. We are in this same situation when determining if the chain rule can be used. However, we have one big advantage over the situation the detective finds themself in. The detective can never really know with 100% certainty if their reconstruction of events is correct. We can be 100% certain of whether our conclusion of how a function is a composite function is correct or not by testing it. So in a chain rule type problem, you’ll be given this:

    h(x)=(1x)2−2(1x)+1h(x)=\left(\frac{1}{x}\right)^{2}-2\left(\frac{1}{x}\right)+1

    And you’ll need to be able to look at it and determine two different things:

    1. That this is potentially composed of two simpler functions where one has been put inside of another
    2. What those two functions are exactly (one inner function, and one outer function)

    In the example we just created, if we think of it as ourselves being the detective trying to recreate the series of events that created the crime scene or the function h(x), it’s pretty easy because we were also the criminal that created the crime scene. Remember, we ourselves put g(x) = 1/x inside of f(x) = x^2 – 2x + 1. But if we hadn’t created the function h(x) ourselves, what clues are there that this can be thought of as a composite function and what those functions could be? You can focus on grouping symbols like parentheses and brackets. So in this function, we see the same function inside a bunch of parentheses pairs. Let’s think of that as our inside function. Now here’s where we get to actually test our theory. If the inside function is indeed 1/x like we think, then we should be able to determine a function that 1/x can be put inside of to create h(x). Here’s kind of a full-proof way to do that. Let’s assign 1/x to a new variable:

    u=1xu=\frac{1}{x}

    Now we can just replace 1/x with u in the function h(x) to determine what the outer function is:

    h(x)=f(1x)=(1x)2−2(1x)+1h(x)=f\left(\frac{1}{x}\right)=\left(\frac{1}{x}\right)^{2}-2\left(\frac{1}{x}\right)+1
    f(u)=(u)2−2(u)+1f(u)=(u)^{2}-2(u)+1

    So here we have our answer. h(x) can be thought of as f(u) = (u)^2 – 2(u) + 1 with g(x) = 1/x put inside of it. Before we continue to finally go through how to use the chain rule and use it in this example, let’s double check our analysis so far. If we’re wrong about what the inner and outer functions are at this point, we will most likely get the final answer wrong no matter how well we actually use the Chain Rule. So how do we test that we are correct so far? Well, let’s remind ourselves of what we are claiming at this point. We are claiming that if you take g(x) and substitute into f(u), the resulting function would be h(x). To determine if that is true or not, let’s do it. Let’s substitute g(x) into f(u) and see what we get:

    f(g(x))=(1x)2−2(1x)+1f\left(g(x)\right)=\left(\frac{1}{x}\right)^{2}-2\left(\frac{1}{x}\right)+1

    Is that mathematically equivalent to h(x)? Yes, they are the exact same. Therefore, our current claim is correct. h(x) can be thought of as f(u) with g(x) inside of it. Now that all that is out of the way, we need to do the easy part: just use the chain rule. Here’s what the chain rule says as a mathematical formula:

    ddxf(g(x))=f′(g(x))⋅g′(x)\frac{d}{dx}f\left(g(x)\right)=f^{\prime}(g(x))\cdot g^{\prime}(x)

    Oh no, we’re looking at another complicated formula. Let’s break down the meaning of this formula a little bit at a time, and you’ll notice that it’s not quite as complicated as it looks. In plain English, I would describe this formula as saying:

    If you are taking the derivative of a function that can be thought of as a function within a function, you can do this by: 1) identifying the inner and outer functions, 2) taking both of their derivatives, and 3) writing your answer as the derivative of the outer function, containing the inner function, times the derivative of the inner function.

    Let’s finally take the derivative of h(x) using the chain rule. Keep in mind, we’ve already done the hard part of analyzing the function and seeing it as a function with a function inside of it, determining what those two smaller functions are, and checking to see if our analysis was correct. Now we just take both of the derivatives of these functions (maybe using other derivative rules) and put the result together as the chain rule formula says:

    f(u)=(u)2−2(u)+1f(u)=(u)^{2}-2(u)+1

    Therefore:

    f′(u)=2u−2f\prime (u)=2u-2

    And:

    g(x)=1xg(x) = \frac{1}{x}

    Therefore:

    g′(x)=−1x2g\prime(x) = -\frac{1}{x^2}

    That’s it. That’s all we need. Now we just put the pieces together according to the formula:

    h′(x)=ddxf(g(x))=f′(g(x))⋅g′(x)h\prime(x)=\frac{d}{dx}f\left(g(x)\right)=f\prime(g(x))\cdot g\prime(x)

    Put g(x) inside of f’(u) just like the formula says:

    f′(g(x))=2(1x)−2f\prime\left(g(x)\right)=2\left(\frac{1}{x}\right)-2

    Multiply that by g’(x) just like the formula says:

    h′(x)=f′(g(x))⋅g′(x)=(2(1x)−2)⋅(−1x2)h\prime(x)=f\prime\left(g(x)\right)\cdot g\prime(x)=\left(2\left(\frac{1}{x}\right)-2\right)\cdot \left(-\frac{1}{x^2}\right)

    And we’re finished. We can simplify this if we want, but mathematically, we’re done. So let’s take what we’ve learned and create a set of steps you can use to evaluate a derivative using the chain rule:

    1. Determine if you can see the function as a composite function, that is, a function with a function inside of it. This takes some creativity. How do you get good at this? Practice. Practice putting functions inside of functions and looking for patterns of common forms that appear. How do you become good at writing? Maybe start by reading. Read a lot and start to analyze what makes good writing good writing.
    2. If you can see the function as a function inside of a function, great. You just need to determine what those functions are. As we discussed before, if you can identify what you think the inside function is, you can replace every instance of that function with some variable (u, for instance) to tease out what the outer function is. If you don’t see the whole function as a function within a function, does that mean the chain rule can’t be applied? Well, literally yes. If you can’t see the function as an inner function and an outer function, you can’t continue. But that doesn’t necessarily mean it can’t be thought of as an inner function and an outer function. Maybe you’re just not looking at it in the best way.
    3. Once you have the two functions, the inner and the outer, you just take both of their derivatives. Keep in mind that to take both of these derivatives, you may need to use other derivative rules including the chain rule again. Once you have those derivatives, you have everything you need. You just need to put the pieces together. How you put them together is as follows:
    f′(g(x))⋅g′(x)f\prime\left(g(x)\right)\cdot g\prime(x)

    Now finally, to get back to a point I made near the beginning of the article: don’t think of the chain rule as “has to be” applied, but rather “can be” applied. In the example that we’ve been looking at, I wouldn’t have chosen to take its derivative using the chain rule normally. There is another way to look at it that makes it an easier problem. I’m not going to think of the function h(x) as

    h(x)=(1x)2−2(1x)+1h(x)=\left(\frac{1}{x}\right)^{2}-2\left(\frac{1}{x}\right)+1

    I’m going to think of it as something that is mathematically equivalent:

    h(x)=1x2−2x+1=x−2−2x−1+1h(x)=\frac{1}{x^{2}}-\frac{2}{x}+1=x^{-2}-2x^{-1}+1

    Here I just multiplied everything out. Now I see something that makes this problem easier. It’s just a bunch of constants and powers of x multiplied and added together. We can use the sum and difference rule, constant multiple rule, and power rule only. No Chain Rule required.

    h′(x)=−2x−3+2x−2h\prime(x)=-2x^{-3}+2x^{-2}

    But wait. That doesn’t look like what we had when we solved this problem using the chain rule. What went wrong? Nothing. Remember, we don’t care if things look the same, necessarily. We care if they are mathematically equivalent. And our two solutions are, in this case.

    h′(x)=(2(1x)−2)⋅(−1x2)=−2x3+2x2h\prime(x)=\left(2\left(\frac{1}{x}\right)-2\right)\cdot \left(-\frac{1}{x^2}\right)=-\frac{2}{x^3}+\frac{2}{x^2}

    They’re the exact same. If they weren’t, something would have gone wrong. Either we performed the derivative rules incorrectly, or we applied them when they didn’t actually apply. So remember, just because you can do something doesn’t mean you should. Be on the lookout for solving problems the simplest, easiest way possible. That’s how you’ll want to do it on a test. For homeworks and practice, it actually may be beneficial to challenge yourself to solve problems in unnecessarily complicated ways. That can further help you understand the inner workings of how these ideas and theorems work and go together, help you practice your mathematical creativity of seeing complicated things where you would normally see simple things and vice versa, and it will definitely help you appreciate the more efficient ways of doing things. Clean an entire kitchen with just a toothbrush just one time, and you’ll really learn to appreciate the more efficient tools for the job.