Home/Lessons/Lesson 4

The First Law of Thermodynamics

Work, heat, internal energy, and the first law.

Internal energyFirst lawPressure-volume workHeat transferState functionExact differentialQuasi-static processAdiabatic processExtensivity

The previous lesson introduced thermodynamic systems, equilibrium states and state functions. We now turn to a central question: how can the energy of a system be incorporated into this macroscopic description?

Some courses treat this question too quickly: the first law is presented as a mere statement of energy conservation, and the energy balance formula follows immediately, with no further justification. We shall meet these formulas in Sections 3 and 4.3, and the reader in a hurry may indeed turn to them directly (see also the short summary in Section 7).

The construction leading to these formulas is, however, a subtle one. First, the first law asserts more than energy conservation, which (nowadays) seems obvious from a microscopic point of view. We shall show that this law in fact bridges the logical gap between the existence of an energy UmicroU_{\mathrm{micro}} defined mechanically on a space of some 6N6N coordinates, and that of a state function UthU_{\mathrm{th}} depending on only a handful of macroscopic variables.

We shall then show that the additivity of energy is a further assumption, needed to pass from the conservation of energy of an isolated system to an energy balance at the boundary of a closed system. It is this last step that finally makes the epistemic and logical status of heat precise: heat is defined as whatever transfer of energy remains in that balance once the macroscopic work done on a closed system has been accounted for, Q  =def  ΔUWQ \equiv \Delta U - W.

In this lesson we first consider closed systems: no matter crosses their boundary. The reason is simple: in an open system, matter itself carries energy across the boundary and complicates the energy balance. Open systems will be treated later in the course.

1. The microscopic internal energy

Consider a classical, non-quantum and non-relativistic system in an inertial frame1. Consider NN point particles of masses mim_i, positions ri\vec r_i and velocities vi\vec v_i. At any instant, its microscopic state is completely determined by

Note 1 : Those cases require a suitably adapted treatment, beyond the scope of this course.
(r1,...,rN,p1,...,pN)Γ,withpi=mivi,\boxed{ (\vec r_1,...,\vec r_N, \vec p_1,...,\vec p_N) \in\Gamma, \, \mathrm{with} \,\,\, \vec p_i=m_i\vec v_i, }

where Γ\Gamma denotes the classical phase space of the system, of dimension dimΓ=6N\dim\Gamma=6N in this simple model. We assume that all internal forces, for instance the attractive or repulsive forces between molecules, are conservative. A total potential energy may then be associated with them, Epint(r1,...,rN),withFiint=riEpintE_p^{\mathrm{int}}(\vec r_1,...,\vec r_N), \, \mathrm{with} \,\,\, \vec F_i^{\,\mathrm{int}} =-\nabla_{\vec r_i}E_p^{\mathrm{int}}. The system may also be subject to external conservative forces, its weight for instance. In the same way, an external potential energy Epext(r1,...,rN)E_p^{\mathrm{ext}}(\vec r_1,...,\vec r_N) is associated with them.

The system need not be at rest in the chosen frame. At every instant one may define its center of mass and its velocity V\vec V, which may itself depend on time. Introducing the total mass of the system and the velocities relative to the center of mass,

M=i=1Nmi,V=1Mi=1Nmivi,ui=viV,M=\sum_{i=1}^{N} m_i, \qquad \vec V=\frac{1}{M}\sum_{i=1}^{N} m_i\vec v_i, \qquad \vec u_i=\vec v_i-\vec V,

the total kinetic energy may then be written as

Ec=12i=1Nmivi2=12i=1Nmi(V+ui)2.E_c=\frac12 \sum_{i=1}^{N}m_i \vec {v_i}^2 = \frac12 \sum_{i=1}^{N}m_i (\vec V + \vec u_i)^2.

Expanding the square, one finds that the cross terms cancel exactly, since imiui=0\sum_i m_i\vec u_i=\vec 0. This result is known as König's theorem in classical mechanics:

Ec=12MV2+12i=1Nmiui2+Vi=1NmiuiimiviMV=0E_c =\frac12MV^2 +\frac12 \sum_{i=1}^{N}m_i \vec{u_i}^2 +\vec V\mathbin{\cdot} \underbrace{\sum_{i=1}^{N} m_i\vec u_i}_{\sum_i m_i\vec v_i-M\vec V=\vec0}

Hence:

Ec=Ecmacro+Ecmicro\boxed{E_c = E_c^{\mathrm{macro}} + E_c^{\mathrm{micro}}}

(1)

The total kinetic energy therefore always splits2 into the macroscopic kinetic energy of the center of mass, and that of the residual relative motions, the “disordered” ones, which make up what is called thermal agitation.

Note 2 : The translation of the center of mass does not necessarily exhaust the macroscopic kinetic energy. The system may also undergo an ordered global rotation. In that case one may write the decomposition vi=V+ω×(riR)+ui,\vec v_i = \vec V +\vec\omega\times(\vec r_i-\vec R) +\vec u_i', where R\vec R is the position of the center of mass and where ui\vec u_i' now denotes the residual motion left after the global translation and rotation have been subtracted. Choosing ω\vec\omega so that this residual motion carries zero total angular momentum about the center of mass, one obtains Ec=12MV2+12ωICMω+12i=1Nmiui2,E_c = \frac12 MV^2 + \frac12\vec\omega\cdot \mathbf I_{\mathrm{CM}}\vec\omega + \frac12\sum_{i=1}^{N}m_i {u_i'}^2, where ICM\mathbf I_{\mathrm{CM}} is the inertia tensor about the center of mass. The separation between macroscopic and microscopic kinetic energy therefore remains possible.

Collecting the various terms, the total mechanical energy finally reads

Etot=Ecmacro+Epext+[Ecmicro+Epint(r1,...,rN)+Eintautres]internal energy.E_{\mathrm{tot}} = E_c^{\mathrm{macro}} +E_p^{\mathrm{ext}} +\underbrace{\left[ E_c^{\mathrm{micro}} +E_p^{\mathrm{int}}(\vec r_1,...,\vec r_N) +E_{\mathrm{int}}^{\mathrm{autres}} \right]}_{\text{internal energy}}.

Remark: the term EintautresE_{\mathrm{int}}^{\mathrm{autres}} collects the contributions of the internal degrees of freedom that this minimal point-particle model does not describe, for instance the rotational or vibrational degrees of freedom of the molecules. If these are modeled explicitly, the phase space Γ\Gamma must be enlarged accordingly.

Definition 1 (Microscopic internal energy)
We shall write Umicro=Ecmicro+Epint+EintautresU_{\mathrm{micro}}= E_c^{\mathrm{micro}} +E_p^{\mathrm{int}} +E_{\mathrm{int}}^{\mathrm{autres}} for the microscopic internal energy. It is a function defined on phase space, Umicro:ΓR,U_{\mathrm{micro}}:\Gamma\longrightarrow\mathbb R,

which therefore depends, in general, on a considerable number of variables, and which satisfies:

Etot=Ecmacro+Epext+Umicro.E_{\mathrm{tot}} =E_c^{\mathrm{macro}}+E_p^{\mathrm{ext}}+U_{\mathrm{micro}}.

(2)

In what follows we shall consider systems that are macroscopically at rest, free of global rotation, and whose external potential energy does not change. Up to a choice of origin, we may therefore write

Etot=Umicro.E_{\mathrm{tot}}=U_{\mathrm{micro}}.

We now introduce the thermodynamic energy state function, provisionally written UthU_{\mathrm{th}}, and explain why it must not be identified a priori with UmicroU_{\mathrm{micro}}.

2. The thermodynamic internal energy

Take as our system a saucepan of water placed on a hotplate. We remain within the simple framework described above: the saucepan is macroscopically at rest and its external potential energy does not change. We may therefore write Etot=UmicroE_{\mathrm{tot}}=U_{\mathrm{micro}} up to a constant.

When the water is heated, no macroscopic kinetic energy appears and the saucepan does not move. Granting the strict conservation of energy in all its forms, the energy supplied to it must therefore end up in its microscopic internal energy:

ΔEtot=ΔUmicro.\Delta E_{\mathrm{tot}}=\Delta U_{\mathrm{micro}}.

Thermodynamics seeks to describe exactly the same change. We should therefore like to construct a quantity UthU_{\mathrm{th}} such that

ΔUth=ΔUmicro.\Delta U_{\mathrm{th}}=\Delta U_{\mathrm{micro}}.

The solution looks obvious. Is it not enough to identify the thermodynamic internal energy with the microscopic internal energy?

Such an identification would in fact be mathematically incorrect, since the two functions do not have the same domain of definition. UmicroU_{\mathrm{micro}} is a function defined on phase space and depends on at least 6N6N variables, whereas UthU_{\mathrm{th}} depends only on a small number dd of independent variables describing the equilibrium state X=(x1,...,xd)EX=(x^1,...,x^d)\in\mathcal E (recall that the xix^i are macroscopic state variables, such as temperature, pressure or volume).

The two functions must nevertheless be related in such a way that Uth(X)U_{\mathrm{th}}(X) represents, at the macroscopic scale, the internal energy corresponding to the equilibrium state XX. How this dimensional reduction works must be explained. The question is a deep one and goes beyond the scope of this introductory course. A microscopic description is needed if one wishes to derive this passage between the two scales. Subsection 6.2, which lies outside the syllabus, outlines how statistical physics solves the problem.

From a strictly thermodynamic point of view, however, no justification of this microscopic identification is possible, and the existence of UthU_{\mathrm{th}} as a function of a few macroscopic variables only must be postulated. This will be the role of the first law of thermodynamics, which therefore asserts not only the conservation of energy, but also, and above all, that this energy can be represented at equilibrium by a state function defined on the macroscopic space E\mathcal E.

Once the dd independent state variables xix^i have been chosen, say (TT, VV, NN) for an ideal gas, the function U(T,V,N)U(T,V,N) defines a scalar field on the state space; in the enlarged space of coordinates U,T,V,NU,T,V,N, its graph is a hypersurface, called the equilibrium surface, whose role in the rest of this course will be crucial, cf. Figure 1.

The energy hypersurface defined over the space of equilibrium states.
Figure 1. The energy hypersurface defined over the space of equilibrium states.

Let us finally note an immediate corollary, which illustrates the difference between UthU_{\mathrm{th}} and UmicroU_{\mathrm{micro}} rather well: in the equilibrium thermodynamics developed here, the thermodynamic internal energy is in general not defined outside equilibrium states. During a violent process undergone by a gas, for instance, it is possible, and indeed common, that no single temperature or pressure can be assigned to it any longer. In that case the thermodynamic internal energy UthU_{\mathrm{th}} cannot be assigned to it directly either, since it is defined only on the space of equilibrium states E\mathcal E. The difficulty is not that the energy has ceased to exist, for its microscopic energy remains perfectly well defined at every instant, but simply that the system is no longer represented by a point XEX\in\mathcal E.

All this being said, from the next section onwards the mechanical function UmicroU_{\mathrm{micro}} will no longer appear directly: we shall simply write UU for the thermodynamic internal energy UthU_{\mathrm{th}}, understood to be defined on the space of equilibrium states only.

Remark 1 (Near-equilibrium extensions)
Let us mention that, for systems sufficiently close to equilibrium, the thermodynamic description can be extended by enlarging the state space. For instance, in the example already met of a metal bar placed between two thermal reservoirs, one may introduce a stationary local temperature field T(x)T(x), rather than a single temperature, and define a local density of internal energy. The total energy then becomes a functional of these fields. We shall not develop this so-called near-equilibrium thermodynamics here.

3. The first law of thermodynamics

The previous section has done much of the groundwork; we can now state:

First law of thermodynamics for a closed system
For every closed thermodynamic system there exists a state function U:ER,U:\mathcal E\longrightarrow\mathbb R,

called the internal energy, determined only up to an additive constant, at least twice differentiable, such that, for every equilibrium state XEX\in\mathcal E,

Etot=Ecmacro+Epext+U(X).E_{\mathrm{tot}} = E_c^{\mathrm{macro}} +E_p^{\mathrm{ext}} +U(X).
(3)

The first law also asserts the conservation of energy: for an isolated system,

ΔEtot=0.\Delta E_{\mathrm{tot}}=0.

The regularity of the function UU is likewise a non-trivial assumption. In practice, UU may be taken to be C\mathcal C^\infty. Since UU is determined only up to a constant, because of the potential energies involved in equality 3, only its variations carry physical meaning. For simplicity, we shall restrict ourselves in what follows to processes in which the macroscopic kinetic and potential energies do not change. We then have ΔEtot=ΔU\Delta E_{\mathrm{tot}}=\Delta U.

Consider now a closed system SS, which may exchange energy, but not matter, with its surroundings. The global system

S=S+ExtS'=S+\mathrm{Ext}

can always be embedded in an isolated system. The first law then imposes

ΔES=0.\Delta E_{S'}=0.

Under the further assumption that energy is additive, that is, that their interaction energy is negligible, we have ES=ES+EExtE_{S'}=E_S+E_{\mathrm{Ext}}, so that

ΔES=ΔEExt.\Delta E_S=-\Delta E_{\mathrm{Ext}}.

The assumption of additivity thus turns the conservation of energy into a balance between subsystems: whatever energy is gained by the system is lost by its surroundings, and conversely. It remains to specify how this energy can be transferred across the boundary of the system. We shall distinguish two modes of transfer: work, written WW, which has an independent definition inherited from mechanics, and heat, written QQ.

The energy balance will then take the form

ΔES=ΔU=Q+W\boxed{ \Delta E_S = \Delta U = Q+W }
(4)

where WW denotes the work received by the system and QQ the heat received by it.

Note that the assumption that energy is additive is not a trivial one. At an elementary level, it may simply be granted for the usual thermodynamic systems. It is not always valid, however. Long-range interactions in particular can defeat both the additivity and the extensivity of energy at once, see Section 6.4.

It now remains to define these two modes of transfer. Before that, it is essential to note a sign convention, an entirely standard one, that will be used throughout this course.

Key point (The banker's sign convention)
Throughout this book, a transfer of energy is counted positively when it is received by the system, and negatively when it is supplied by the system. Thus Q>0, W>0Q>0,\ W>0

correspond to energy received by the system, whereas

Q<0, W<0Q<0,\ W<0

correspond to energy given up to the surroundings.

4. Work, heat and the energy balance

4.1. Work

One of the advantages of work is that it has a definition independent of thermodynamics, inherited from mechanics. When an external force Fext\vec F_{\mathrm{ext}} acts on a point of the boundary of the system and that point is displaced by drd\vec r, the elementary work supplied by this force, and therefore received by the system, is

δW=Fextdr.\delta W = \vec F_{\mathrm{ext}}\cdot d\vec r.

For an extended system, the contributions of all the external forces doing work on its boundary must be summed.

The case that will serve us most often in practice is that of a fluid contained in a vessel with one movable wall, which leads us to evaluate the work of the pressure forces. Consider a gas enclosed in a cylinder by a flat piston of cross-section S\mathcal S, cf. Figure 2. The xx axis points outward from the gas. When the piston is displaced by an amount dxdx, the volume changes by

dV=Sdx.dV=\mathcal S\,dx.

Let PextP_{\mathrm{ext}} denote the external pressure exerted on the boundary of the gas. The force received by the gas is directed inward and is

Fext=PextSex.\vec F_{\mathrm{ext}} = -P_{\mathrm{ext}}\mathcal S\,\vec e_x.

The elementary work received by the gas is therefore

δWpression=Fext(dxex)=PextSdx,\delta W_{\mathrm{pression}} = \vec F_{\mathrm{ext}}\cdot(dx\,\vec e_x) = -P_{\mathrm{ext}}\mathcal S\,dx,

that is

δWpression=PextdV\boxed{ \delta W_{\mathrm{pression}}=-P_{\mathrm{ext}} dV }
(5)

The banker's rule makes the sign immediate to check. In a compression, dV<0dV<0, hence δW>0\delta W>0: the gas receives work. In an expansion, dV>0dV>0, hence δW<0\delta W<0: the gas supplies work to its surroundings.

Note that the pressure appearing in this formula is indeed the pressure external to the system, which in general has no reason to equal the pressure of the gas itself. This is particularly important to keep in mind for violent, out-of-equilibrium processes, where the internal pressure is in general not even defined.

Work received by a gas when a piston of cross-section S is displaced. In an expansion, dV>0 while the force exerted by the external pressure opposes the displacement: the work received is therefore negative.
Figure 2. Work received by a gas when a piston of cross-section S\mathcal S is displaced. In an expansion, dV>0dV>0 while the force exerted by the external pressure opposes the displacement: the work received is therefore negative.

Integrating, we obtain the following formula for a finite process ABA \to B during which the external pressure is known:

Wpression[AB]=ABPextdV\boxed{ W_{\mathrm{pression}}[A \to B] = -\int_{A}^B P_{\mathrm{ext}} dV }
(6)

This integral shows that, in order to compute the total work, the value of PextP_{\mathrm{ext}} must be known all along the process. This has a major consequence: the work can in general not be determined from the initial and final states AA and BB alone: its expression depends on the process followed.

Three particular cases are worth pointing out here.

  1. If PextP_{\mathrm{ext}} is constant, then Wpression[AB]=PextΔVW_{\mathrm{pression}}[A\to B] = -P_{\mathrm{ext}}\Delta V.
  2. Along an isochoric process, dV=0dV=0, and the work of the pressure forces vanishes: δWpression=0\delta W_{\mathrm{pression}}=0 and Wpression[AB]=0W_{\mathrm{pression}}[A\to B]=0.
  3. In an expansion into a vacuum, Pext=0P_{\mathrm{ext}}=0, so the work of the pressure forces vanishes as well.

Remark 2 (Other forms of work)
The work of the pressure forces is only one example among others. For a wire being stretched, if LL denotes its length and FextF_{\mathrm{ext}} the tensile force received, δW=FextdL.\delta W=F_{\mathrm{ext}}\,dL.

For an interface whose area AA is increased, the surface tension γ\gamma leads, under the usual conditions, to work of the form

δW=γdA.\delta W=\gamma\,dA.

Energy may also be transferred without any visible mechanical displacement. For instance, moving a charge dqdq through a potential difference can give rise to electrical work of the form

δWeˊlec=Vextdq,\delta W_{\mathrm{élec}}=V_{\mathrm{ext}}\,dq,

where VextV_{\mathrm{ext}} is the external electric potential, often written Φext\Phi_{\mathrm{ext}} in this context so as not to confuse it with the volume. Here too the sign of dqdq is chosen in accordance with the banker's convention.

Subsection 6.3 will make the general form of the external work terms precise.

4.2. Heat

Work cannot be the only mode of energy transfer, since heating a gas held in a rigid vessel raises its temperature and its energy without any external work being done. Energy has therefore crossed the boundary of the system without having been transferred as work. This second mode of energy transfer is called heat, or heat transfer.

In the construction adopted here (see Subsection 6.1 for an equivalent alternative construction), for a physical process ABA \to B connecting two equilibrium states AA and BB, once all the work terms received by the system have been identified, the heat received is defined by

Q[AB]  =def  U(B)U(A)W[AB].\boxed{ Q[A \to B] \equiv U(B)-U(A)-W[A \to B]. }
(7)

so that we recover the relation announced above,

ΔU=Q+W\boxed{ \Delta U=Q+W }

Since the work depends on the path of the process ABA \to B, while the energy change ΔU=U(B)U(A)\Delta U = U(B)-U(A) does not, the amount of heat exchanged must necessarily depend on it as well.

In thermodynamic language, one says that the internal energy UU is a state function (by postulate), whereas the work WW and the heat QQ are not. Recall what this means: the fact that WW and QQ cannot be functions of the state of the system at a given instant means that the system does not possess “an amount of work or of heat”; rather, WW and QQ are transfers of energy at its boundary during some process ABA \to B, exactly what we were after, in view of the various experiments described in Lesson 2.

Let us note that in a cyclic process AAA \to A the system returns to its initial state, so that

ΔUcycle=0.\Delta U_{\mathrm{cycle}}=0.

The first law then imposes

Qcycle+Wcycle=0.\boxed{Q_{\mathrm{cycle}}+W_{\mathrm{cycle}}=0.}

A cyclic machine can therefore not supply work indefinitely without receiving an equal amount of energy from its surroundings. This is, historically, what is called the impossibility of perpetual motion of the first kind.

4.3. Differential form of the first law

The relation

ΔU=Q+W\Delta U=Q+W

connects two equilibrium states AA and BB. It always holds, in the sense that it does not require the intermediate states themselves to be representable by points of the space of equilibrium states E\mathcal E. If the process is a violent one, the system may temporarily leave the equilibrium surface defined over E\mathcal E, while U(A)U(A) and U(B)U(B) remain perfectly well defined.

One special case must now be spelled out. Suppose the process is quasi-static. At every instant the system is then, by definition, in an equilibrium state. Such a process can be represented by a path

γE\gamma\subset\mathcal E

made of infinitesimally close equilibrium states, which never leaves the equilibrium surface. Since UU is a state function that is differentiable on E\mathcal E, its change between two infinitesimally close states is an exact differential, written dUdU.

We then write δQ\delta Q and δW\delta W for the corresponding elementary transfers of heat and work. In this case the first law takes the form

dU=δQ+δW\boxed{ dU=\delta Q+\delta W }
(8)

The difference in notation between dUd U on the one hand and δW\delta W or δQ\delta Q on the other is not merely cosmetic. It indicates that the former is an exact differential while the latter two are not, which is the mathematical counterpart of the fact that the energy change does not depend on the path followed, whereas QQ and WW do. The next section reviews, for the newcomer, the mathematics needed to grasp this crucial point.

In the particular case where the only work is that of the pressure forces, and where PextP_{\mathrm{ext}} may be identified with the pressure PP of the system, this relation becomes

dU=δQPdV.dU=\delta Q-P\,dV.
Key point (Two formulations of the first law)
The integrated form ΔU=Q+W\Delta U=Q+W

connects two equilibrium states and remains usable even when the intermediate process is not quasi-static. The differential form

dU=δQ+δWdU=\delta Q+\delta W

is valid only for a quasi-static process, which can be represented by a path in the space of equilibrium states.

5. Exact and inexact differential forms

This section gathers the minimum of mathematics needed to give a precise meaning to the distinction between dUdU on the one hand, and δQ\delta Q and δW\delta W on the other. Readers already familiar with it may skip it.

5.1. The differential of a function

Let ff be a function of the variables x1,...,xdx^1,...,x^d, assumed differentiable. Its differential is the expression

df=i=1dfxidxi,df=\sum_{i=1}^{d}\frac{\partial f}{\partial x^i}\,dx^i,

which measures the change of ff when one moves from the point XX to the neighboring point X+dXX+dX. The essential point is the following: if one follows a path γ\gamma from a point AA to a point BB, then summing all these elementary changes gives

γdf=f(B)f(A).\int_\gamma df=f(B)-f(A).

In particular, along a closed path we have:

df=0.\oint df=0.

5.2. Differential forms

Consider now an expression of the same kind, which must on no account be confused with the differential of a function. Let Ai(x1,...,xd)A_i(x^1,...,x^d) be arbitrary functions. We define ω\omega by

ω=i=1dAidxi.\omega=\sum_{i=1}^{d}A_i\,dx^i.

Such an object is called a differential form. Once the functions AiA_i are given, integrating it along a path γ\gamma raises no difficulty. But nothing guarantees that there exists a function ff of which ω\omega is the differential, that is, such that Ai=f/xiA_i=\partial f/\partial x^i for every ii.

  • If such a function ff exists, the form is said to be exact, and one writes ω=df\omega = d f. The integral γω=γdf=f(B)f(A)\int_\gamma\omega = \int_\gamma df = f(B)-f(A) then depends only on the endpoints.
  • Otherwise the form is said to be inexact, and it is then written δω\delta\omega rather than ω\omega, as a reminder that it is the differential of no function. (Mathematicians would simply write the form as ω\omega.)

5.3. Schwarz's criterion

A simple test for deciding whether a differential form is exact is Schwarz's criterion. For simplicity we work with two variables (the generalization is immediate), with

ω=A(x,y)dx+B(x,y)dy.\omega=A(x,y)\,dx+B(x,y)\,dy.

If ω\omega were exact, we would have A=f/xA=\partial f/\partial x and B=f/yB=\partial f/\partial y, hence

Ay=2fyx=2fxy=Bx,\frac{\partial A}{\partial y} =\frac{\partial^2 f}{\partial y\,\partial x} =\frac{\partial^2 f}{\partial x\,\partial y} =\frac{\partial B}{\partial x},

since the order of differentiation is immaterial for a twice continuously differentiable function. We thus have a convenient test:

AyBx    ω is not exact.\boxed{ \frac{\partial A}{\partial y}\neq\frac{\partial B}{\partial x} \;\Longrightarrow\; \omega\ \text{is not exact.} }

Up to subtleties that will not arise in thermodynamics, the converse holds as well.

5.4. Example: the work of the pressure forces

When the elementary transfers can be expressed in terms of the state variables along the quasi-static path, δQ\delta Q and δW\delta W become differential forms on E\mathcal E, and one can check whether they are exact or not.

By way of example, take nn moles of an ideal gas undergoing a quasi-static process, with Pext=PP_{\mathrm{ext}} = P. We then have δW=PdV\delta W=-P\,dV. Since P=nRT/VP=nRT/V, the elementary work reads

δW=PdV=nRTVdV=0×dTnRTVdV,that isA=0,B=nRTV,\delta W=-P dV = -\frac{nRT}{V}\,dV = 0\times dT-\frac{nRT}{V}\,dV, \qquad\text{that is}\qquad A=0, \quad B=-\frac{nRT}{V},

identifying x=Tx=T and y=Vy=V. This is indeed a differential form defined on the state space. Schwarz's criterion gives

AV=0,BT=nRV0.\frac{\partial A}{\partial V}=0, \qquad \frac{\partial B}{\partial T}=-\frac{nR}{V}\neq0.

The two cross derivatives are not equal: the elementary work is indeed not an exact differential. There is therefore, in this case, no state function W(T,V)W(T,V) of which δW\delta W would be the change, and the work received between two states does depend on the path followed. The same reasoning applies to δQ=dUδW\delta Q=dU-\delta W: since dUdU is exact and δW\delta W is not, their difference cannot be exact either.

Key point (why δ\delta and not dd)
dUdU is an exact differential: its change depends only on the initial and final states, and its integral over a cycle vanishes. δQ\delta Q and δW\delta W are inexact forms: they depend on the path followed, and there exists neither a function QQ nor a function WW of the state of the system such that δW=dW\delta W = d W and δQ=dQ\delta Q = dQ.

6. Going further

This section gathers the more advanced considerations announced earlier. It may be skipped on a first reading.

6.1. Two possible constructions of the first law

We have chosen here the following logical order: the existence of the state function UU is postulated by the first law, work is defined independently by mechanics, and heat is then defined by the balance

Q=ΔUW.Q=\Delta U-W.

A remarkable fact follows. If the process is adiabatic, then ΔU=W\Delta U = W, which forces the adiabatic work WadW_{\mathrm{ad}} to be, in its turn, always independent of the path followed.

This points to another possible construction of the internal energy: an operational one, hence independent of any microscopic consideration, and as it happens closer to the historical route. One begins by characterizing adiabatic processes without introducing the quantity QQ beforehand, so as to avoid any circularity in the reasoning.

The idea is a clever one: what is defined is not the transfer itself, but the apparatus. A wall is said to be adiabatic when the state of the system it encloses can be changed only by moving the external mechanical coordinates, that is, the piston, a stirrer or an electric current. The test is then direct: one holds these coordinates fixed and alters the surroundings arbitrarily, by plunging the enclosure into an ice bath, say, or bringing it close to a flame. If no state variable of the system changes, then the wall is adiabatic.

This characterization involves only states and mechanical displacements, never a transfer of energy: it therefore precedes any notion of heat. In practice, one approaches it by good insulation, or by working fast compared with the thermal relaxation time.

The work received, for its part, is measured in a purely mechanical or electrical way. This is the whole empirical content of Joule's experiments, carried out between 1843 and 1850 and repeated since with ever greater precision: a mass mm falling through a height hh supplies mghmgh, with no heat transfer involved. All the arrangements he devised and built in an adiabatic enclosure in the above sense (paddle wheel, heating resistor or compression of a gas) led him to the same conclusion: the same work supplied produces the same change of state, which for him constituted “the mechanical equivalent of heat”, as explained in Lesson 2. Joule was thus the first to show (an experimental indication of the fact) that adiabatic work does not depend on the path followed.

It was then Carathéodory, in 1909 [2], and above all Born, in 1921, who proposed to elevate this experimental fact into a postulate. Granting it, it becomes easy to construct a state function UU by setting

U(B)U(A)=Wad(AB),U(B)-U(A)=W_{\mathrm{ad}}(A\to B),

and then to define heat by difference for general processes, as we have done.

This construction is adopted in many textbooks, notably those of Pippard [3] and Callen [4]. It is equivalent to ours: postulating UU and then deducing the path-independence of adiabatic work is equivalent to postulating the path-independence of that work and then constructing UU. A detailed account of this conceptual evolution, from Joule to Carathéodory and Born, will be found in Rosenberg [5].

6.2. Micro- and macrostates

If one does not follow the operational route described above, then Section 2 has brought a fundamental question to light: how can an energy defined on a space of 6N6N coordinates be reduced to a function of only a few macroscopic variables? Statistical physics constructs this passage between the two scales. Its outline is as follows.

In statistical physics, a point XX of the space of equilibrium states E\mathcal E is called a macrostate. To one and the same macrostate there generally corresponds a gigantic number of microstates γΓ\gamma\in\Gamma compatible with the same macroscopic constraints.

One then shows that the energies of these microstates are distributed extremely narrowly about their mean value as NN\to\infty, the so-called thermodynamic limit. In other words, the relative fluctuations of the energy tend to zero:

σ(Umicro)Umicro0.\frac{\sigma(U_{\mathrm{micro}})} {\langle U_{\mathrm{micro}}\rangle}\longrightarrow0.

One may then set Uth(X)=UmicroXU_{\mathrm{th}}(X)= \langle U_{\mathrm{micro}} \rangle_X, the average being taken over all microstates compatible with the macrostate. It is the concentration just mentioned that gives this definition its meaning. At every instant the system occupies a single microstate, not the average: only because nearly all of them carry the same energy can one speak of the energy of the macrostate, and can two identical preparations yield the same measurement.

One may picture the situation by imagining that to a macrostate there corresponds a very large collection of microstates that are microscopically different but macroscopically indistinguishable. This picture resembles that of an equivalence class on phase space. The precise construction will, however, be slightly different, and will involve a probability distribution on phase space.

The following figure summarizes this procedure.

One and the same macrostate X is compatible with an immense number of microstates. On our scale their energies are extremely concentrated about a single macroscopic value, which is what makes it possible to define U_ th(X).
Figure 3. One and the same macrostate XX is compatible with an immense number of microstates. On our scale their energies are extremely concentrated about a single macroscopic value, which is what makes it possible to define Uth(X)U_{\mathrm{th}}(X).

6.3. Generalized work terms

The various forms of work met above share a common structure. We have seen the forms

δW=PextdV\delta W=-P_{\mathrm{ext}}\,dV

for pressure work,

δW=FextdL\delta W=F_{\mathrm{ext}}\,dL

for the stretching of a wire, and again

δW=γextdA\delta W=\gamma_{\mathrm{ext}}\,dA

for surface work. Note that in every case the quantity being differentiated is extensive, while its prefactor is intensive. Generalizing, we shall write any mechanical work term in the form

δW=iνiextdxi,\boxed{ \delta W=\sum_i \nu_i^{\mathrm{ext}}\,dx^i, }

where νiext\nu_i^{\mathrm{ext}} denotes the external generalized force (generally intensive) conjugate to the coordinate xix^i (generally extensive), with the sign matching our convention.

When the external generalized forces can be identified with the corresponding thermodynamic forces of the system (for example Pext=PP_{\mathrm{ext}}=P), the first law may then be written in the form

dU=δQ+iνidxi.\boxed{ dU=\delta Q+\sum_i\nu_i\,dx^i.}

6.4. Extensivity of the internal energy

In the previous lesson we presented the internal energy as an extensive quantity: multiplying the size of a homogeneous system by a factor λ\lambda, at fixed intensive variables, multiplies its energy by λ\lambda,

UλU.U\longrightarrow\lambda U.

This property is not postulated by the first law. It is a further assumption, which we shall often make, but which does not always hold. One indeed expects the presence of long-range forces between the constituents of the system, gravitational ones in particular, to ruin its extensivity.

Let us make this precise. Consider NN constituents distributed at constant density in a space of dimension DD, and suppose that their interaction potential energy behaves as

Ep(r)1rα.E_p(r)\sim\frac{1}{r^\alpha}.

At fixed density ρ=N/V\rho=N/V, the typical linear size of the system grows as

LN1/D.L\sim N^{1/D}.

Let us now evaluate the total interaction energy in three steps.

Counting the neighbors.

Fix one constituent and ask how many others lie at a distance between rr and r+drr+dr. This is the density times the volume of the corresponding shell:

dn(r)=ρSDrD1dr,dn(r)=\rho\,S_D\,r^{D-1}\,dr,

where SDS_D denotes the area of the unit sphere in dimension DD, that is 4π4\pi in three dimensions.

Summing over distances.

Each of these neighbors contributes Ep(r)rαE_p(r)\sim r^{-\alpha}. The interaction energy of a single constituent with all the others is therefore

uρaLrD11rαdr=ρaLrD1αdr.u\sim\rho\int_a^L r^{D-1}\,\frac{1}{r^{\alpha}}\,dr =\rho\int_a^L r^{D-1-\alpha}\,dr .

The lower bound aa is the minimum distance of approach, below which the 1/rα1/r^\alpha law ceases to hold and which prevents the integral from diverging as r0r\to0. The upper bound LL is the size of the system: there is no neighbor beyond it.

Summing over constituents.

We multiply by the number of constituents (dividing by two so as not to count each pair twice):

UintN2uNρaLrD1αdr,U_{\mathrm{int}}\sim\frac{N}{2}\,u \sim N\rho \int_a^L r^{D-1-\alpha}\,dr,

where the geometrical factor SD/2S_D/2 has been absorbed into the order of magnitude. Let us now look at the value of the integral:

aLrD1αdr=[rDαDα]aL(αD).\int_a^L r^{D-1-\alpha}\,dr =\left[\frac{r^{D-\alpha}}{D-\alpha}\right]_a^L \qquad(\alpha\neq D).

If α>D\alpha>D, the exponent is negative: the integral converges as LL\to\infty and is dominated by its lower bound, hence of the order of the constant aDα/(αD)a^{D-\alpha}/(\alpha-D). If α<D\alpha<D, it is on the contrary dominated by its upper bound and is of the order of LDα=N1α/DL^{D-\alpha}=N^{1-\alpha/D}. Finally, in the case α=D\alpha=D, the antiderivative is a logarithm and the integral equals ln(L/a)=1DlnN\ln(L/a)=\tfrac1D\ln N. Substituting, we thus obtain

Uintcste×{N,α>D,NlnN,α=D,N2α/D,α<D.\begin{aligned} U_{\mathrm{int}} \sim \text{cste} \times \begin{cases} N, & \alpha>D,\\[1mm] N\ln N, & \alpha=D,\\[1mm] N^{\,2-\alpha/D}, & \alpha<D. \end{cases} \end{aligned}

Physically, the interpretation is clear. When α>D\alpha>D, the forces are short-ranged: the decay of EpE_p outweighs the growth in the number of neighbors. Each constituent feels only its immediate neighborhood, so that its energy does not depend on the size of the system, and the total is proportional to NN, which is extensive.

When αD\alpha\leq D, the opposite happens: the distant neighbors, far more numerous, win out, each constituent feels the system as a whole, and its own energy grows with the size of that system. The total then grows faster than NN, and extensivity is lost. Sufficiently short-ranged interactions therefore lead naturally to an extensive energy, whereas long-range interactions can destroy that property.

In three-dimensional Newtonian gravity, for instance,

Ep(r)1r,D=3,α=1,E_p(r)\sim-\frac1r, \qquad D=3, \qquad \alpha=1,

and the above estimate gives, at fixed density,

UgravN5/3.U_{\mathrm{grav}}\sim-N^{5/3}.

The minus sign, which the order-of-magnitude argument does not by itself provide, reflects the attractive character of gravity.

The extensivity of energy will be used extensively in the rest of this book. Gravity is of course universal, but it is a very weak force: when studying two volumes of gas brought into contact, it can be neglected in practice. On the other hand, the internal forces within the gas, of van der Waals type, decay very fast (as 1/r61/r^6), so that extensivity is in that case very nearly exact.

For the description of so-called self-gravitating systems (stars, galaxies, ...), on the other hand, gravity is of course the essential ingredient, and the whole thermodynamic analysis must be taken up again from the start, since the extensivity of energy is necessarily lost. As a result, the thermodynamics of such systems is almost a subject of its own, and displays unexpected behavior: a star that radiates energy away, for instance, heats up instead of cooling down. We shall return to this in the advanced part of the book.

7. Summary

Let us summarize what this lesson has set up, and what will be used throughout what follows.

  • The first law postulates the existence of a state function U:ERU:\mathcal E\longrightarrow\mathbb R, the internal energy, and asserts the conservation of the energy of an isolated system.
  • In the rest of the book we shall also assume it to be extensive, and sufficiently continuous and differentiable in all its variables.
  • Supplemented by the additivity of energy, the first law allows conservation to be expressed as an energy balance between a closed system and its surroundings. Since work is defined independently by mechanics, heat is then defined as the remaining transfer, which leads to the useful formula ΔU=Q+W\Delta U=Q+W.
  • For a quasi-static process, this energy balance takes the differential form dU=δQ+δWdU=\delta Q+\delta W, in which only dUdU is an exact differential.

The analysis of open systems will be discussed elsewhere in the course.

8. References

On the conceptual evolution of the first law, from Joule to Carathéodory and Born, see Rosenberg [5]. For a classical presentation, see Pippard [3] and Callen [4].

  1. J. P. Joule, “On the Mechanical Equivalent of Heat,” Phil. Trans. R. Soc. Lond. 140, 61—82 (1850)
  2. C. Carathéodory, “Untersuchungen über die Grundlagen der Thermodynamik,” Math. Ann. 67, 355—386 (1909)
  3. A. B. Pippard, Elements of Classical Thermodynamics for Advanced Students of Physics, Cambridge University Press (1957)
  4. H. B. Callen, Thermodynamics and an Introduction to Thermostatistics, 2nd ed., Wiley (1985)
  5. R. M. Rosenberg, “From Joule to Caratheodory and Born: A Conceptual Evolution of the First Law of Thermodynamics,” J. Chem. Educ. 87, 691—693 (2010)
  6. A. M. Steane, “First Law, internal energy,” chap. 7 in Thermodynamics: A Complete Undergraduate Course, Oxford University Press (2017)
  7. E. A. Gislason and N. C. Craig, “Cementing the foundations of thermodynamics: Comparison of system-based and surroundings-based definitions of work and heat,” J. Chem. Thermodynamics 37, 954—966 (2005)