ISBN: 0-471-05998-6

Text
                    PURE  AND  APPLIED  MATHEMATICS
 A  Wiley-lnterscience  Series  of  Texts,  Monographs,  and  Tracts
Founded  by  RICHARD  COURANT
 Editors:  LIPMAN  BERS,  PETER  HILTON,  HARRY  HOCHSTADT
 ARTIN—Geometric  Algebra
ASH^Information  Theory
AUBIN—Applied  Abstract  Analysis
AUBIN—Applied  Functional  Analysis
AUBIN—Applied  Nonlinear  Analysis
 BEN-ISRAEL,  BEN-TAL,  ZLOBEC—Optimality  in  Nonlinear  Programming
CARTER—Simple  Groups  of  Lie  Type
CLARK—Mathematical  Bioeconomics
 COLTON  and  KRESS—Integral  Equation  Methods  in  Scattering  Theory
CURTIS  and  REINER—Representation  Theory  of  Finite  Groups  and  Associative  Algebras
CURTIS  and  REINER—Methods  of  Representation  Theory:  With  Applications  to  Finite
Groups  and  Orders,  Vol.  I
 EHRENPREIS—Fourier  Analysis  in  Several  Complex  Variables
 FISHER—Function  Theory  on  Planar  Domains:  A  Second  Course  in  Complex  Analysis
 FRIEDMAN—Differential  Games
 FRIEDMAN—Variational  Principles  And  Free-Boundary  Problems
GRIFFITHS  and  HARRIS—Principles  of  Algebraic  Geometry
HANNA—Fourier  Series  and  Integrals  of  Boundary  Value  Problems
HENRICI—Applied  and  Computational  Complex  Analysis,  Volume  1
Volume  2
HILLE—Ordinary  D
HILTON  and  WU-i
HOCHSTADT-The
HOCHSTADT-Inte
HSIUNG-A  Firet  C
KELLY  and  WEISS-
KOBAYASHI  and  N<
 KRANTZ—Functior
KUIPERS  and  NIEE
LALLEMENT-Sen
LAMB—Elements  of


LAY—Convex Sets and Their Applications LINZ—Theoretical Numerical Analysis: An Introduction to Advanced Techniques LOVELOCK and RUND—Tensors, Differential Forms, and Variational Principles MARTIN—Nonlinear Operators and Differential Equations in Banach Spaces MELZAK—Companion to Concrete Mathematics MELZAK—Invitation to Geometry NAYFEH—Perturbation Methods NAYFEH and MOOK—Nonlinear Oscillations ODEN and REDDY—An Introduction to the Mathematical Theory of Finite Elements PASSMAN—The Algebraic Structure of Group Rings i^ETRICH—Inverse Semigroups PRENTER—Splines and Variational Methods RIBENBOIM—Algebraic Numbers RICHTMYER and MORTON—Difference Methods for Initial-Value Problems, 2nd Edition RIVLIN—The Chebyshev Polynomials ROCKAFELLAR—Network Flows and Monotropic Optimization RUDIN—Fourier Analysis on Groups SAMELSON—An Introduction to Linear Algebra SCHUMAKER—Spline Functions: Basic Theory SHAPIRO—Introduction to the Theory of Numbers SIEGEL—Topics in Complex Function Theory Volume 1—Elliptic Functions and Uniformization Theory Volume 2—Automorphic Functions and Abelian Integrals Volume 3—Abelian Functions and Modular Functions of Several Variables STAKGOLD—Green’s Functions and Boundary Value Problems STOKER—Differential Geometry STOKER—Nonlinear Vibrations in Mechanical and Electrical Systems STOKER-Water Waves TURAyN—On A New Method of Analysis and Its Applications WHITHAM—Linear and Nonlinear Waves WOUK—A Course of Applied Functional Analysis ZAUDERER—Partial Differential Equations of Applied Mathematics
APPLIED NONLINEAR ANALYSIS
APPLIED NONLINEAR ANALYSIS JEAN-PIERRE AUBIN and IVAR EKELAND A WILEY-INTERSCIENCE PUBLICATION JOHN WILEY & SONS New York Chichester Brisbane Toronto Singapore
Copyright© 1984 by John Wiley & Sons, Inc. All rights reserved. Published simultaneously in Canada. Reproduction or translation of any part of this work beyond that permitted by Section 107 or 108 of the 1976 United States Copyright Act without the permission of the copyright owner is unlawful. Requests for permission or further information should be addressed to the Permissions Department, John Wiley & Sons, Inc. Library of Congress Cataloging in Publication Data: Aubin, Jean Pierre. Applied nonlinear analysis. (Pure and applied mathematics) “A Wiley-Interscience publication.” Bibliography: p. Includes index. 1. Nonlinear functional analysis. I. Ekeland, I. (Ivar), 1944- II. Title. III. Series: Pure and applied mathematics (John Wiley & Sons) QA321.5.A93 1984 515.7 83-26011 ISBN 0-471-05998-6 Printed in the United States of America 10 987654321
This book is dedicated to LAURENT SCHWARTZ
PREFACE For a long time now, functional analysis has been linear. Since the beginning of this century, thanks mainly to David Hilbert and Stefan Banach, the theory of infinite-dimensional vector spaces has become the appropriate framework for studying linear equations. The tremendous success of this approach is illus¬ trated by Laurent Schwartz’s theory of distributions, the gateway to partial differential equations. The methods of linear analysis are now fixed by tradition in a standard package, which can be found in numerous textbooks, and belong to the common background of all mathematicians. On the other hand, the same period also saw the growth of fixed-point theory, from Brouwer’s finite-dimensional result to the Leray-Schauder theory in Banach spaces, with its application to nonlinear partial differential equations. In more recent years, the interest in nonlinear problems and methods has dramatically increased. Scholars in many fields, from partial differential equa¬ tions to economic theory, have developed methods of coping with nonlinear problems, such as solving/(a:)=0 with x an infinite-dimensional variable or even F(.x) aO wtih Fa set-valued map. In addition, the success of linear programming has set a pattern for the whole of optimization theory to follow: Find necessary conditions for optimality, discover dual formulations of the problem, and use them to evolve efficient numerical methods for computing the solution. This program has been com¬ pleted for convex optimization and is now in the process of being carried over to some nonconvex problems. On the way, new tools have been developed for analyzing nonsmooth functions, which are constantly encountered in optimi¬ zation theory. Nonlinear analysis now stands on its own; it is no longer a subsidiary of linear analysis, but has its own methods and its own applications. Better still, it is now possible to introduce students to this very active field without going through the whole of linear analysis as a preliminary. The ideas in nonlinear analysis are simple, their proofs direct, and their applications clear. No more prerequisites are needed than the elementary theory of Hilbert spaces; indeed, many of the results are most interesting in Euclidian spaces. In order to remain at an introductory level, we have chosen not to delve into technical difficulties or sophisticated results that are not of current use. App¬ lications are given as soon as possible, and the theoretical parts are written with Vll
VIH PREFACE that purpose in mind. We feel that the added benefit from the availability of applications more than compensates for scattering topics. Several themes run throughout the book. The first one is the resolution of nonlinear equations f(x)=0 and inclusions F{x) aO. Chapter 6 gives a unified treatment of these problems in close relationship with game theory and mathematical economics, which provide insight and applications. In the particular case when the equation to be solved is V'(x)=0 or dV(x)B 0, where F is a real-valued function and d denotes a generalized gradient, we are dealing with a variational problem. This is the second theme of the book. In Chapter 2, these problems are studied geometrically by means of the inverse function theorem. The framework is finite dimensional or becomes so after a suitable reduction. In Chapter 4, we investigate the special case when V is convex: Any solution of dV(x) bO must then minimize K, and we enter the realm of convex optimization. We give, and sometimes improve, the classical results of convex analysis and duality theory. In Chapter 5, we state a general variational principle, which is applicable to infinite-dimensional nonconvex situations, and give several applications to nonconvex variational problems. The tools developed in Chapter 5 are applied in Chapter 8 to a specific example of great practical importance: finding periodic solutions to differential equations of Hamiltonian type, in other words, describ¬ ing periodic oscillations of conservative nonlinear systems. Another theme is nonsmooth analysis, which is developed for the purposes of optimization and stability theory in dynamical systems. This is done in Chapter 7: Two generalized derivatives are introduced and compared and the calculus rules extended accordingly up to the inverse function theorem, even for set¬ valued maps. Of course, this is but a glimpse into the vast and expanding field of nonlinear analysis. Much has been left unsaid—degree theory, inverse function theorems of Nash-Moser type, Morse theory, Liusternik-Schnirelman theory, obstacle and free boundary problems in partial differential equations, problems from differential geometry or nonlinear elasticity—and the list is depressingly long. Ours is but a personal selection, strongly motivated by our background in opti¬ mization theory. It ranges from very smooth functions to nonsmooth ones, from convex variational problems to nonconvex ones, from economics to mechanics. It is our hope that this will be enough to lead students into the field, to make them feel how lively and exciting it is, and to stimulate them to learn more. Jean-Pierre Aubin IVAR EkELAND Paris, France February 1984
CONTENTS Chapter 1. Background Notes 1. Set-Valued Maps, 1 2. Complete Metric Spaces, 5 3. Banach Spaces, 9 4. DifTerentiable Functionals, 16 5. Support Functions and Barrier Cones of Convex Subsets, 26 Chapter 2. Smooth Analysis 33 1. Iterative Procedures for Inverting a Map, 34 2. Milnor’s Proof of Brouwer’s Fixed Point Theorem, 3. Local Study of the Equation /(a:)=0, 46 4. Birth and Death of Critical Points, 58 5. Further Degeneracies: Bifurcation, 70 6. Transversality Theory, 83 7. Proof of the Transversality Theorem and Applications, 41 91 Chapter 3. Set-Valued Maps 1. Upper and Lower Semicontinuity of Set-Valued Maps, 2. Maps with Closed Convex Values, 121 3. Maps with Closed Convex Graphs, 130 4. Eigenvalues of positive Maps with Closed Convex Graphs, 146 Chapter 4. Convex Analysis and Optimization 1. Tangent and Normal Cones to Convex Subsets, 166 2. Derivatives and Codifferentials of Set-valued Maps with Convex Graphs, 177 3. Epiderivatives and Subdifferentials of Convex Functions, 186 4. Conjugate Functions, 200 103 107 159 IX
X CONTENTS 5. The Subdiiferential of the Marginal Function and Lagrange Multipliers, 214 6. Convex Optimization Problems, 220 7. Regularity of Solutions to Convex Optimization Problems, 225 8. Lagrangians and Hamiltonians, 231 Chapter 5. A General Variational Principle 1. Walking on Complete Metric Spaces, 239 2. Fixed Points of Nonexpansive Maps, 248 3. The 6-Variational Principle, 254 4. Applications to Convex Optimization, 261 5. Condition (C) of Palais and Smale, 269 6. Generic Differentiability, 279 7. Perturbed Optimization Problems, 285 239 Chapter 6. Solving Inclusions 295 1. Main Concepts of Game Theory, 299 2. Two-Person Zero-Sum Games: The Minimax Theorem, 312 3. The Ky Fan Inequality, 325 4. Existence of Zeros of Set-Valued Maps, 336 5. Walras Equilibria and Price Decentralization, 354 6. Monotone Maps, 363 7. Maximal Monotone Maps, 379 8. Existence and Uniqueness of Solutions to Differential Inclusions, 396 Chapter 7. Nonsmooth Analysis 401 1. 2. 3. 5. 6. Contingent and Tangent Cones, 405 Contingent Derivatives and Derivatives of a Set-Valued Map, 411 Epicontingent Derivatives and Epiderivatives of Real- Valued Functions, 418 Generalized Second Derivatives of Real-Valued Functions, 428 The Inverse Function Theorem for Set-Valued Maps, 429 Calculus of Contingent and Tangent Cones, Derivatives and Epiderivatives, 439
Chapter 8. Hamiltonian Systems CONTENTS xi 451 1. 2. 3. 4. 5. Comments Bibliography Anthor Index Subject Index The Least Action Principle, 452 A Dual Action Principle, 456 Nonresonant Problems, 464 Resonant Problems, 469 Transresonant Problems, 475 487 495 517
APPLIED NONLINEAR ANALYSIS
CHAPTER 1 Background Notes In this preliminary chapter, we group together some basic definitions, notations, and facts that will be used later in the book. In the first section, we define set-valued maps and fix the vocabulary employed. Baire’s theorem and some of its fundamental consequences are recalled in the second section. Examples of lower semicontinuous functionals on Banach spaces of functions are provided in the third section. The fourth section deals with different classical notions of differentiability for functionals and presents several examples. These background notes end with the presentation of support functions of closed convex sets and a quite useful closed image theorem. 1. SET-VALUED MAPS We shall not escape using set-valued maps in this book, for the good reason that we deal with them in a very natural way, as shown in the list of examples mentioned in the first section. Nor do we choose to regard a set-valued map from a set A"to a set У as a single-valued map from X to the set ^(Y) of the sub¬ sets of У: By doing so, we would lose a lot of information, since, in most cases, the structures on ^{Y) are much poorer than the original structures on У Let X and У be two sets. A set-valued map, or a correspondence F from X to is a map that associates with any xe Xdi subset F(x) of Y, called the image or the value of F at x We say that a set-valued map is proper if there exists at least an element xe X such that F{x)^0, that is, if F is not the constant map 0. In this case, we say that the subset (1) Dom {F):={xeX\F(x)^0} is the domain of F. Actually, a set-valued map Fis characterized by its graph, the subset of A' x y defined by (2) Graph (F):={(x, y)\y e F(x)}
2 CH. 1, SEC. 1 BACKGROUND NOTES Indeed, if G is a nonempty subset of the product space X x X it is the graph of the set-valued map F defined by (3) у 6 F{x) if and only if (jc, y)^G The domain of F is the projection on X of Graph (F) and the image of F, the subset of Y defined by (4) Im (f):= U U Fix) xeX xeDom(F) is the projection on Y of Graph (F). The inverse F“ * of f is the set-valued map Ffrom У to Л'defined by (5) if and only if € F(a:) or, equivalently, (6) xeF~ Hy) if and only if (x, y) e Graph (F) Therefore, we obtain the formulas (7) Dorn (F" ‘)=Im (F), Im (F" ‘)=Dorn (F) and (8) Graph (F -1)={(y, ;c) 6 У X A'l {x, y) e Graph (F)} We shall say that a set-valued map F from to У is strict if Dom (F)=X, that is, if the images F(x) are nonempty for all xe X. When A'is a nonempty subset and when F is a strict set-valued map from to it may be useful to “extend it” to the set-valued map Fk from X to У defined by (9) , , {F{x) when x & X when X i F whose domain Dom (Fk) is K. When F is a set-valued map from A' to y and F<= A", we denote by F|k its restriction to K. Let (^) be a property of a subset (for instance, closed, convex, monotone, maximal monotone, etc.). As a general rule, we shall say that a set-valued map F satisfies the property ( ^) if the graph of F satisfies this property. For instance, we shall speak of a closed, convex, monotone, maximal mono¬ tone map, which is a set-valued map whose graph is closed, convex, monotone, maximal monotone, and so on.
CH. 1, SEC. 1 SET-VALUED MAPS 3 If the images of a set-valued map F are closed, convex, bounded, compact, and so on, we say that F is a closed-valued map, convex-valued map, bounded¬ valued map, compact-valued map, and so on. When ★ denotes an operation on the subsets, we use the same notation for the operation on set-valued maps, which is defined by (10) Fi ★ F2 : F2(x) We define in that way Fi n F2, Fi u F2, Fi\F2, and Fi -h F2 (in vector spaces). Similarly, if a is a map from the subsets of Y to the subsets of Y, we define (11) a{F): x:-^a(F(x)) For instance F:=cl(F): x^F{x\ Int F: x->Int F(x), co(F): x-^co(F(x)), cb(jp): x:->cb F(.x), and so on. Let us mention the following elementary properties: (12) Examples i. F[K,kjK2) = F{K,)kjF{K2) ii. F(K^^ n K2)^F(K^)r\F(K2) iii. F{X\K)^F{x)\F{K) IV. K^czK2^F{K^)c:F{K2) a. The first natural instance where set-valued maps occur is the inverse/" ^ of a single-valued map from Xto Y We always can define/"^ as a set-valued map whose domain is Im(/), which is strict when / is surjective and single-valued when/ is injective. This map plays an important role when we study equations f{x)=y and are interested in the behavior of the set of solutions f~^(y)^sy ranges over Y b. We shall associate with a function V from a set A" to Ru{ + oo} the set¬ valued map (13) |F(x)-I-IR+ when F(x)<+00 '^'^1 0 when F(x)=+ 00 The domain of V+ is the subset of elements x satisfying F(x:)< +00, and the graph of V + is the epigraph of the function V. We say that V is proper if V + is proper and that Dom V:= Dom V+ is the domain of V (14) {.!* [11. c. Another instance of set-valued maps is associated with a family of functions V proper<=>{ € X\ V{x)< +00} Dom V:={xeX\ F(x)<+oo}
4 CH. 1, SEC. 1 BACKGROUND NOTES /(•, m) from XtoY when u ranges over a set U of parameters. In this case, we set (15) F{x) = {f{x,u)Uv Control theory provides examples of such maps, called parametrized maps. d. We shall associate with a lower semicontinuous convex function V from a Hilbert space A" to IRu{ + oo} its subdifferential dV(x\ defined by (16) dV(x):=^p 6 X*\(p, x)- V(x)=max [<p, y)- K(j)]| It is a closed convex subset of X'^, which may be empty. It “generalizes” the concept of gradient in the sense that if V has a gradient VF(jc) 6 at Xy then dV(x)={VV(x)}. The set-valued map x-^dV(x) will play a crucial role in many applications described in this book. Another important fact is that the inverse of 3K is also the subdifferential of a lower semicontinuous convex function K*, called the conjugate of K defined on the dual X* of A" by (17) F*(p)=sup [_{p,y)-V{y)'\ yex More generally, when K is a locally Lipschitz function from an open subset ^ of a Hilbert space A" to (R, we shall define its generalized gradient^ dV(xX which is a bounded closed convex subset that reduces to {VK(x)} whenever V is con¬ tinuously differentiable, and which was introduced by Clarke, e. Let ly be a function from A" x y to IR. We consider the minimization prob¬ lems (18) "iyeX, V(y)= \ni W(x,y) xe X The function V is called the marginal (or performance or value) function. Let G(y):={xeX\ W(Xyy)=V{y)] be the subset of solutions to the minimization problem V(y). One of the main purposes of optimization theory is to study the set-valued map G (continuity and differentiability in a suitable sense, and so on). We shall call it the marginal map. It is no wonder that game theory and mathematical economics use set-valued maps in a natural way.
CH. 1, SEC. 2 COMPLETE METRIC SPACES 5 2. COMPLETE METRIC SPACES A distance on a set A" is a real-valued function on A" x A" that satisfies the following for all X, y, and z€ X\ (1) (2) (3) d(x, y)>0 d(x, y)=0<^x=y d{x, z)^d(Xy y) + d(y, z) The third property is the triangle inequality. A set X, together with a dis¬ tance d, is a metric space. On such a space, we can define convergent sequences and Cauchy sequences. A sequence « 6 N, in A" is convergent if there exists some x e A" such that for every 6 > 0, some N eN can be found so large that (4) Wn^N, d(x„,x)^e The point X is uniquely defined and called the limit of the sequence x„. A sequence « 6 f^, in A" is Cauchy if for every £>0, some N eN can be found so large that (5) ^n^N, Vm^ TV, £ Any convergent sequence is Cauchy. The converse is not true without further assumptions on the sequence or on the space. For instance, if a Cauchy sequence contains a convergent subsequence, then it is itself convergent. Assumptions about the space are much more fruitful; we are led to the following definition. DEFINITION 1 A metric space (X, d) is complete if every Cauchy sequence is convergent. A Most familiar metric spaces, and certainly all those we shall deal with in this book, are complete. This includes (R, finite-dimensional vector spaces like R”, function spaces like IS or and all closed subsets thereof. It is clear from the definition that if (X, d) is complete and Fc: X is closed, then (F, d) is complete. Of course, the most important property of complete metric spaces lies in definition 1 itself: there is practically no way to prove that a sequence x„, n e N,is convergent, except by proving that it is Cauchy and lies in a complete metric space. The major difficulty, of course, is to prove completeness: It may be easy for us, knowing that U is complete, but it certainly was a major achieve¬ ment of nineteenth-century mathematics to build up U from the set of all rationals (noncomplete) in such a way that it was complete.
6 CH. 1, SEC. 2 BACKGROUND NOTES However, working from the definition, we obtain other properties of complete metric spaces. First an easy one; recall that the diameter of a subset in a metric space is the upper bound of all distances measured within that subset (6) F<= X, diam F=sup {d{x, y)\x e F,y e F] PROPOSITION 2 Let {Xy d) be a metric space and F„, neN, a decreasing sequence of closed subsets whose diameters go to zero (7) F„ciF„+i and diam Then there is a single point x belonging to all of the F„ (8) n ■f’«=W W=1 Proof Choose one point x„ in each of the We claim the sequence a:„, neN, thus obtained is Cauchy. Indeed, let e>0 be given and choose JVso large that: diam Fn^8 If both n and m are greater than N, then both x„ and x^ belong to Fn, since the sequence is decreasing, and we get the Cauchy property ^n>N, ^m>N, d(Xn, Xni)^diam Fn^s Since the space {X, d) is complete, the Cauchy sequence x„ has to converge to some X. Any one subset Fk contains ic, since it is closed and contains the sequence x„ except for the starting terms xu ..., Xk-i- Since x belongs to all of the it belongs to their intersection, as in equation (8). Let X be another point with the same property. Since both x and 3c belong to every Fn, we must have V«, c/(3c, x)^diam Since the right-hand side goes to zero, we get d(x, 3c)=0; hence, 3c=x, and uniqueness is proved. ■ Example Equation (7), particularly the fact that the diameters must go to zero, is essential. Let us give two counterexamples, one in finite dimension and the other infinite. Take X=U with the usual distance, and set F„:= [«, + oo]. This is a decreas¬
CH. 1, SEC. 2 COMPLETE METRIC SPACES 7 ing sequence of closed subsets with empty intersection (note that diam Fn = H- oo). Take the Hilbert space of square-summable sequences, x=((^„ i e N), with ^ (^?<oo. Set This is a decreasing sequence of closed subsets with empty intersection (note that diam F„ = l). One step further away from the definition, we find a very important result: the Baire theorem. ■ THEOREM 3 Let (X, d) be a complete metric space and i7„, neN, a sequence of open sets, each of w^hich is dense in X, Then so is their intersection (9) X=clf] Un M=1 Proof Call this intersection G. We want to prove that Gr\Bo^0, where Bo is any prescribed open ball in X. Since Ui is dense and Bq is open, UinBo is nonempty. Since Ui is open, so is UinBo. It follows that UinBo contains some closed ball Bu the diameter of which can be taken less than 1. We now proceed by induction. Assume that we have found a closed ball (10) c ^0 n Pi Uii k=l with diam B„^ 1/«. We then call B„ the interior of B„: It is an open ball. Since Un+i is open and dense, the intersection Un+i(^B„ is open and nonempty. Thus, we are able to pick another closed ball «+1 5„+i«=5„ni/„+,c:5on n Uk k=l 1 diam B„+i< n + l We now apply theorem 2 to the sequence B„, neN and obtain a point x that belongs to all of the B„. It follows from equation (10) that X eBo and x e U„, Hence, xeGnBo, and the result is proved. for all n
8 CH. 1, SEC. 2 BACKGROUND NOTES An alternative version is THEOREM 4 Let {X, d) be a complete metric space and F„, neN, a sequence of closed sets. If their union covers X, then one of them has nonempty interior (11) (J F„=X ^ 3n: F„f0 n=i Proof, Set Un = X\F„, The i/„ are open sets with empty intersection. If they were all dense in X, so would their intersection be by the Baire theorem. Since the empty set cannot be dense, it follows that one of the t/„ at least, say Un, is not dense in X, This means that there is an open ball B in X that does not intersect Un. Bn Un=0^B^Fn It follows that the interior of Fn contains B. ■ In both these theorems, the fact that we are dealing with sequences, that is, countable families, is very important. For instance, with each real number x associate the open set Ux-={y e^\y-hx] which is obviously dense in IR. Now G\=nUx for a: eg is dense (it is just the set of irrational numbers), whereas G\=nUx forxelR is empty! The reason the Baire theorem applies in one case and not in the other is because the first intersection is countable (recall that the set Q of rational numbers is countable), whereas the second is not. Note also that not all dense subsets can be obtained as countable intersec¬ tions of open dense sets. In fact, we introduce DEFINITIONS Let {X, d) be a complete metric space, A subset G^X is residual if it contains a countable intersection of open dense subsets, A It follows from the definition that the intersection of two residual subsets, or a countable number of residual subsets, is still residual. Indeed
CH. 1, SEC. 3 BANACH SPACES 9 GinG2=>(f] C/«Vfn n \ n / \ k J {n,k) ni?« = n(n U„,,)=f] i/,a n n \ k J (n,k) where n and k range over N and («, k) over x H which is countable. The Baire theorem states that in a complete metric space, any residual subset is dense. For instance, two residual subsets have a nonempty intersection: It is residual, hence dense, hence nonempty. In the same way, any countable family of residual subsets has a nonempty intersection. In the real line R, for instance, the set G:=R\6 of irrational numbers is residual, whereas Q itself is not (although it is dense!). For if G and Q were both residual, they would have to intersect, which is not the case. We see that residual subsets behave much better than ordinary dense subsets: Two dense subsets need not intersect (e.g., Q and R\2), whereas two residual subsets always do. In this way, we can think of residual subsets G as being “full” subsets of A", as R\ (2 is a full subset of U. The following definition emphasizes this and has become very popular in various areas of mathematics. DEFINITION 6 A property P(x\ where x runs through a complete metric space {X, d) is called generic if the set of points G^X where it holds true is residual A For instance, the property that a real number be irrational is generic. The interesting thing is that if two properties Pi(x) and P2(x) are generic, then so is the property “Pi(x) and Pii^V' For that matter, if the properties P„(x), for neN, are all generic, then so is the property that all P„(x) hold simultaneously. It happens very often in mathematics, and it will happen to us, that to prove that there is one point x e X such that jP(x) holds, we have to prove that P{x) is generic. In other words, we find a large number of points where P(x) is satisfied, just because we need one. 3. BANACH SPACES A norm on a vector space A" is a function x-^||x|l from A" to R that satisfies the following for all x,yin X and all 2 e IR: (1) (2) (3) (4) IWI^o IW|=o<^IW|=o 11^^11=WIWI
10 CH. 1, SEC. 3 BACKGROUND NOTES It follows that d{Xy j^)=||a:—;;|| is a distance on X. Unless otherwise specified, X will be endowed with this distance, which turns it into a metric space: It is the so-called norm topology on X. If the space (X, d) is complete, it will be called a Banach space. In the normed linear space X, the (closed) unit ball B is defined by (5) ^:={xe AIIWKl} The corresponding sphere S will be the set of points in at a distance 1 from the origin. The closed ball with center xe X and radius p=0 is the subset x+pB. The open ball, denoted by x+pÉ, is obtained by removing the boundary x+pB={y e X\ ||x->'||=ep} x+pÉ={yeX\ ||x-;^||<p} It should be noted that balls in a normed linear space are never compact unless the space is finite dimensional (a theorem of Riesz). Finite-dimensional spaces are just for some suitable «, albeit perhaps with some non-Euclidian norm, and all closed balls will certainly be compact. Now this cannot happen in infinite-dimensional spaces, which puts us on the lookout for another, weaker, topology. If A" is a normed linear space, denote by X* the set of all continuous linear functionals p on X. It is a normed linear space, and we call it the (topological) dual of X\ its norm is given by (6) lbll*=sup{|<j:,p)|||WI<l}. The weak topology (as opposed to the norm topology) on X will be defined as follows: X converges weakly to 3c in A' if for every fixed p in X* (7) (p,x)-^(p,x) Of course, if the point x converges to x in the norm topology, then it converges in the weak topology. The converse is not true, except again in the finite-dimen¬ sional case. For most practical cases, that is, for U{Q) and IF*’^(Q), with 1 < oo, the weak topology is metrizable, so that we can be content with considering sequences The main usefulness of weak topologies lies in the following result. THEOREM 1 Assume X is reflexive, that is, X=(X^)’^. Then all norm-closed balls in X are weakly compact. A
CH. 1, SEC. 3 BANACH SPACES 11 The most important example of reflexive spaces are Hilbert spaces, such as the space N) of square-summable sequences or the space I?(Q) of square- integrable functions over some subset ilczR” or the Sobolev spaces //*(Q) and But there are other examples, such as the spaces N) and L^(0), for 1 </? < oo, the duals of which are and Z?(Q), p~^ = \, or the Sobolev spaces l^^’^(Q) and always with 1 </? < oo and A: < oo. Now that we have indeed found some compact subsets in some infinite¬ dimensional Banach spaces, we shall try to exploit this for minimization pur¬ poses. For this, we need yet another notion, lower semicontinuity of functionals. The following definition is stated in terms of a topological space, that is, it can apply to either the norm topology or the weak topology of a Banach space. Note also that we shall allow + oo as a possible value for the functions we consider. This device will be very useful in minimization problems. DEFINITION 2 Let X be a topological space. A function U: A"-^[Ru{ + oo} is lower semicon- tinuous {sometimes shortened to l.s.c.) if for each point xe X, we have (8) liminf U{x)>U{5c) This means that when x goes to x in X, all cluster points of U{x) in R u {+ oo} are to be above U(x). In other words, jumps are allowed [inequality (8) may be strict], but they must occur downward. The most useful characterization of such functions is given by the following classical result. PROPOSITION 3 Let X be a topological space, U: u{ + oo} a function, epi U its epigraph epi U = {{x, a)eX X U\a^U{x)} Then U is lower semicontinuous if and only if epi U is a closed subset of XxR. A Of course, a function U: X-^Uu {- oo} will be called upper semicontinuous if for each point jc e A", we have lim sup i/(x)^ U{x) In other words, U is upper semicontinuous if (— i/) is lower semicontinuous. If a finite function U is both upper and lower semicontinuous, it is continuous. Lower semicontinuity is a much weaker property than continuity, and it will be much easier to check in practical situations. On the other hand, it is also all that is required for the purposes of minimization, as the following result shows.
12 CH. 1, SEC. 3 BACKGROUND NOTES PROPOSITION 4 Let the topological space X be compact. Then any lower semicontinuous function LTiX-^Rul + oo} attains its minimum 3x: U{x)^U{x) A Let us now go back to Banach spaces. We have stressed the fact that there are two possible topologies on such a space, the norm (strong) topology or the weak one. We speak accordingly of strongly or weakly lower semicontinuous func¬ tions, and they are not the same. Indeed, we note the following important result. PROPOSITIONS Let X be a Banach. Any weakly l.s.c. function is strongly l.s.c. The converse is true in the convex case: Any convex function that is strongly l.s.c. is also weakly l.s.c. A Proof Let U: X^Ru {+ oo} be some function and consider its epigraph. Any weakly closed subset of X-^R is also strongly closed, so by proposition 3, if U is weakly l.s.c., it also is strongly l.s.c. Conversely, by the Hahn-Banach theorem, any strongly closed subset of A" XIR that is convex is also weakly closed. The result follows by proposition 3. This leads to the following consequence, which is about the only way to find minimizers in infinite-dimensional problems. Note that it is stated in terms of the norm topology, although its proof requires the weak topology. PROPOSITION 6 Let X be a reflexive Banach space and U: A"->(Ru{ + oo} a convex and lower semicontinuous function. Let K be a bounded subset of A", convex and closed. Then U attains its minimum on K. 3xe K: U(x)^ U(x\ for all xe K A Proof. Switch to the weak topology. The function U is still lower semi¬ continuous by proposition 5. The subset K is still closed by the Hahn-Banach theorem (see Section 5). Being bounded, K is contained in some closed ball, which, by theorem 1, is weakly compact. Now any closed subset of a compact set is compact, so K is (weakly) compact. By proposition 4, U attains its mini¬ mum on K. ■ Of course, convexity is a very stringent assumption. In Chapters 2 and 5, we shall introduce ways of doing without it and still obtaining some information about the existence of minimizers. A typical way of constructing functionals on function spaces is by using integrals. We conclude by studying such functionals on I}.
CH. 1, SEC. 3 BANACH SPACES 13 Example 1 Let Q be an open subset of IR", and let u: QxIR''->IRu{h-oo} be a borelian function such that a. b. Vcu 6 Q, y) is lower semicontinuous y{o),y)eQxU\ u(o),y)^0 Define a functional U on I}{Q) by V;cGl?(i2), i/(x)= I u{(o, x{co))do) Jn The right-hand side is always well defined, since the function co-^u{co, x{(o)) is measurable and nonnegative: The integral is either finite or +oo. It follows that U is well-defined, with values in R u {+ oo}. (This could be the case even if u itself were restricted to values in U.) We claim that U: L^(Q)->(Ru{ + oo} is lower semicontinuous. A Indeed, let ;c„, /7 e N, be a sequence converging to some x in L^(ii). We first extract a subsequence x„' such that lim inf U{Xn)= lim U{Xn') n-* 00 n'-^OO and then we extract from Xn' a second subsequence that converges almost everywhere x„»(co) jc(co) in IR ” From assumption (a) it follows that for almost every co in Q lim inf m(co, Xn'ico))^u{(o, ic(co)) Integrating both sides yields 1 lim inf m(co, x„"{u)))dco^ u{o), x((o))do) Jn w"-^oo Jq Since the integrands are nonnegative, we can apply Fatou’s lemma lim inf u{(o, Xn"{(o))d(D^ lim ini u{co, Xn' {co))d(o n"-^oo Jii Jq m"->oO
14 CH. 1, SEC. 3 BACKGROUND NOTES Adding the two last inequalities lim inf OO Xn»{oj))d(o^ \ u(o), x{oji))da> a Jn But the integral on the left-hand side is V{Xn>), and the right-hand side is U{x\ Finally, we obtain the desired result lim inf U{x^= lim U{xn-)> U{x) ■ Example 2 Let Q be a borelian subset of R", and let u:ilx{+ oo} be a borelian function such that for some a e I}{Q) and ceU (^) Vco, j->w(co,;;) is lower semicontinuous (b) V(co, y\ u(co, y)^-a(co)- cy ^ Then the functional U: L^(Q)->Ru{h-oo}, defined as before, is lower semi¬ continuous. A Indeed, consider the functional V on defined by V(x)= [u{(o, x(o)))-^a{co)+cx^{o))']d(o Jn = U(x)-\- [ a{co)do) + c\\x\\l Jn The integrand, inside the brackets, is nonnegative, so V is lower semicon¬ tinuous by example 1. But V differs from [/ by a constant [the integral of (a)] plus a continuous function (the square of the norm in i}\ So U itself must be lower semicontinuous. Example 3 Let Q be a borelian subset of R”, and let w: Q x R^'-^R satisfy the following: (a) (b) (c) V;;6R^ co-*u{o),y) is measurable Vo) 6 Q, y-^u(co, y) is continuous V (co,;;), |m(co, y)\ ^ a{co) + cy^ for some a e L?(Q) and c e R. Then the functional U: is continuous. A
CH. 1, SEC. 3 BANACH SPACES 15 Indeed, we break inequality (c) into two parts u(o3, y)^— a(co)—cy^ — u{u>, y)'^— a(co)— The first one with example 2 tells us that the functional U is lower semi- continuous with values in Rvj{-l-oo}, and the second one tells us that the functional — 1/ is also lower semicontinuous with values in Ru{-l-oo}. The result follows immediately. ■ Example 4: A Minimization Problem mthout Solutions Let fi be the interval (0, 1) in R, and consider the functional U: Ho{0, !)-► R u {-f- oo} defined by U{x) with x:=dxldco. Recall that the functions x in Ho(0, 1) are all continuous and vanish on the boundary x(0)=0=x(l) We claim that there is no point jc that minimizes C7 on Z =if ¿(0,1). In other words, the problem of minimizing U on X has no solution VxeZ, U{x)>iniU A To see this, we note that the infimum of U is zero. On the one hand, U is clearly nonnegative, that is, inf U>0. A particular sequence x„ can be built in X such that i/(x„)-»’0. We define it by cutting up the interval (0, 1) into 2n equal intervals, /n=((k—1)/2«, k/2n), and setting xJo>)= -(-1 if to 6 4 with k odd x„(co)= — 1 if o) € 4 with k even This, together with x:„(0)=0, defines a function x„eHo(0, 1). It is easily checked that |x„(co)| < 1/2« and x„{o>)^ = 1 for almost every to in (0,1). Substituting this into the definition of U, we have U(x„)={2n)-^-*0 Finally, inf U=0 as claimed.
16 CH. 1, SEC. 4 BACKGROUND NOTES On the other hand, there could never be an x with U{5c) actually zero. Indeed, the integrand is the sum of two squares and could not be zero without each of them being zero. The condition U{x)=0 breaks down into two pointwise conditions x^^(co)=0 for almost every co dx — (co)= +1 for almost every co d(o These conditions are clearly incompatible: There is no x minimizing {/. ■ Note that the functional U itself is lower semicontinuous on //¿, by example 1. Changing (l—x^)^ to (1 — makes it continuous, and even but the corresponding minimization problem will still have no solution (argu¬ ment unchanged). In other words, there is nothing pathological about the func¬ tion U that causes this situation to occur. 4. DIFFERENTIABLE FUNCTIONALS In this section, X will be a normed linear space and X* its dual. Let C/ be a real-valued function defined on an open subset, a: some point in X, and p some continuous linear functional. There are several possible meanings to the state¬ ment that p is the derivative of U at x We list four of them: (a) Gateaux Differentiability. For any ve X lim - \^U(x-\-hv)— U{x)— {p, hv)^=0 h-*0 n (b) Frechet Differentiability. 1 ïüiï [^^(•^+1')- U{x)~ (p, u>]=0 v-^0 (c) Strict Differentiability. lim [[/(;;+t>)- U(y)~ (p, iJ>]=0 izi Any one of these definitions will give a single possible value for p e X*,
CH. 1, SEC. 4 DIFFERENTIABLE FUNCTIONALS 17 henceforth denoted by i/(x) or V U(x), to show dependence on the functional U and the point x, and called the gradient, or the derivative, of U at x. We give a fourth definition (d) The Property. U is Gateaux differentiable on a neighborhood of x in X, and U'{y)-* V{x) in X* when_v-»x in X. LEMMA 1 (d) ^ (c) => (b) => (a), with the same p. ^ Proof. The last two, (c) => (b) and (b) ^ (a), are obvious. As for the first one, assume (d)is satisfied, choose e>0 so small that U is Gateaux differentiable on the ball with center x and radius e, and pick any y and v with Hj —x||<6/3 and l|a||<e/3. Consider the function </> = [0, 1]^IR defined by <l){t)=Uiy + tv)-U(y) It is derivable, and we can use the mean value theorem 30 6 [0,1]: cf>{l)-m=<l>'(0) U{y+v)~ U(y)=(U'{y+ev), v) When y-^x and v-*0, we have [t/(;^ + a)- U(y)-(p, v}}\\v\\-^ = (Lny + &v)- «-‘>-^0 which is exactly property (c). I Some comments on these definitions may be useful. Gateaux differentiability is the weakest notion—and therefore the easiest to check ; for purposes of mini¬ mization, it is often enough. For instance, it is easily seen that if a Gateaux- differentiable function U attains its minimum over X at some point x, then its derivative at x is zero (1) i/(x) = inf U:=^U'ix)=0eX* On the other hand, for matters relying on the inverse function theorem. Gateaux differentiability is not enough: We require strict differentiability at least. To understand Gateaux differentiability, we write it as follows. (2) lim j [ U(x + hv) — U{x)\ = ( U(x)y v) /1-0+ h The left-hand side is the directional derivative toward y, and Gateaux differ¬
18 CH. 1, SEC. 4 BACKGROUND NOTES entiability simply expresses that it is a continuous linear functional of v. On the other hand, the function U need not even be continuous at x! The reader will check that the function i/: R^-»IR defined by U{x, y)=(x^+y^) for X 0 t/(0,7)=0 is Gateaux differentiable, but not continuous, at the origin [with t/'(0)=0]. Frêchet differentiability, on the other hand, implies continuity. It simply means that the first-order Taylor expansion is valid (3) U{x+a) = U(x)-I- (U'{x), v)+£(a)||a|| with e(a)^0 when ||a||^0. Strict differentiability will mean that the same expan¬ sion is going to be valid in a uniform way for points y close to x (4) U(y + v)= U(y)-1- <[/(x), v)+e(v, ;;)||a|| with s(a, y)-^0 when v-^0 uniformly for all y in a neighbourhood of x. As for the C ^-property, it means that the derivative U'(x) depends continuously on the point X, that is, [/ is continuous as a (nonlinear) map from X to X*. Observe that the definition of Frêchet differentiability involves a priori knowledge of the gradient! Therefore, to compute a Frêchet derivative, we must begin by computing the limit [equation (2)] of the differential quotients by means of the usual calculus of functions of one variable, then check whether the dependence on v is linear and continuous, and, finally, verify whether proper¬ ty (b) holds true. Even the requirement of Gateaux differentiability is too stringent, for the limit of the differential quotients may not exist, and if it does exist, it may be either nonlinear or discontinuous. We shall relax the concept of limit of differ¬ ential quotients in several ways to obtain even weaker notions of directional derivatives, retaining enough properties, however, to be useful. We shall investigate these properties for convex functions in Chapter 4 and for more general functions in Chapter 7. Now for some examples. Example 1 Let Q be an open subset of R", with the usual Lebesgue measure da>. Let there be given a function u: fixR^-^R, borelian with respect to both variables (co, x). Assume that
CH. 1, SEC. 4 DIFFERENTIABLE FUNCTIONALS 19 (5) for any fixed cü e Q, the function x) is C* over R*“ (6) there is some a e L^(£î, R) and some constant b such that I u'y{(o, y)| < a(co)+¿>1 y I for ail y e R (7) I \u(a>,0)\dco<co Ja Then the functional U: I3(CÏ, R*)-» R given by U(x)= u(co, x{a>))da> Ja is well defined and Gateaux differentiable. À To see this, we must make a computation. Let x and v be given in R'^). For each fixed co e £2, consider the derivative (8) ^ u{<o, x(m)+=Uy((o, x{(o)+fi;(a)))* u(cü) (9) Using estimate (6), we obtain d dt u{a>, x(co) + ii;((a)) < b(a>)l(a(m)+6|;c((ü)+it)(cu)l) < |u(m)|(a(cü)+6|x(co)| + ô|ü((ü)|) provided 1. We note that the right-hand side is an integrable function of CO (10) ^«(co, x(co)-|-ti;(cu)) ^g{(o), withi?eU(£2, The consequences are twofold. First, for any i; e L\ we have, setting x=0 in the preceding formula, m(o), d(o)))=u(co, 0) + I ^ m(o), tv{(o))dt |«(ft), y(0)))| \u{CO, 0)1 -I-g(co) I U(v)\^ j |m(cu, 0)|dco-f- j g{(o)d(o <00
20 CH. 1, SEC. 4 BACKGROUND NOTES So U is well defined. We now take any x el} and write the directional derivative toward vel3 ^U(x+tv) I u(o), x{co)+tv((o))d(a Thanks to the estimate in inequality (10), we can use the Lebesgue theorem on the differentiation of integrals with respect to a parameter si"*“’ x(œ)+tv{œ))dœ x{(a)+tv{(o))do) (11) Using equation (8), we finally obtain d dt U{x+tv) — I Wji(iO, ( = 0 Jn x{o)))v{(o)d(o Note that because of condition (6), the function a)-+My(to, x((o)) is square integrable, just as x itself «;°A:€L^(fi, R") so that the right-hand side of equation (11) is indeed a continuous linear functional on X=Ú{íl, R**). We write it as (12) (13) — U{x+tv) = (u'y ° X, v) t = 0 U'(x) = Uy° X e A"* Example 2 Same assumptions as example 1. We now claim that U has the property. We have already shown that U is Gateaux differentiable and that its deriva¬ tive U is given by equation (13). We now have to show that U is a continuous map from X to X*, that is, from L^(i2, into itself. This is a consequence of a theorem by KrasnosePskii, which we now state as theorem 2. THEOREM 2 Let X and Y be two Banach spaces, Q a Bor el subset of R” and u: Q x 7 a mapping, measurable with respect to co in Q, continuous with respect to x in X. For all X in U{il, X), let 0(x) be the {measurable) function from Qto Y defined by (14) 0(x) : 0)-> u{œ, x(o))) If^ maps Ii{Q, X) into I3(Q, Y), then it is continuous {l^p,q<co). k
CH. 1, SEC. 4 DIFFERENTIABLE FUNCTIONALS 21 Proof. We shall prove that <f> is continuous at the origin 0; continuity at any point X follows by replacing 4>(x) with (f>{x—x). Let x„,« e N, be any sequence of functions in if(Q, X) such that ||jc„—3c||p-»0. We are going to show that there is a subsequence ^ € N, such that (f)(0). This is enough to ensure continuity at the origin. We choose the sub¬ sequence x„^ to satisfy It is well known that this implies that almost everywhere. It follows that (15) «(<w. x„^((o))^u(o), 0) a.e. In other words, <f>(x„i)-*<l>(0) almost everywhere. To show convergence in 13, we have to use Lebesgue’s theorem. For this, we note that for almost every cu in i2, the sequence u((o, x„^(o))) converges to u(o), 0) in 1^ and, hence, there must be some k(co) e N such that the distance from m(o), x„^(co)) to u(o), 0) is maximum. We choose the k(co) in such a way that the function defined by x(cu)=x:fc,„)(a)) is measurable (here, we need a measurable selection theorem). By definition, we have (16) \\u(o), x„f(o))-tAp), 0)11 < ||m(o), xio)))-ui(o, 0)11 a.e. We claim that x belongs to If, indeed [ ll.x(a))||'’i/a)< [ Sup ||x„t(co)||'’i/cu Jn Jn k < [ E \\x„^i(oWdco Jil k k k By the assumption on <f>, it follows that <^(3c) has to belong to 13. The right- hand side of inequality (16) thus belongs to 13. Putting (15) and (16) together and using Lebesgue’s theorem, we obtain the desired result <!>{x„^)-*<t>i0) in I3(Q,Y) m This applies immediately to the map x-*u'y°x in example 1. Indeed, because of estimate (6), this map sends L^(Q, R *‘) into itself IKo^||2<||fl + Z,|^| ||2<|M|2 + 6|W|2
22 CH. 1, SEC. 4 BACKGROUND NOTES It follows that this map is continuous. Because of equation (13), the function U has the C‘ property. ■ Example 3 Again take a function u satisfying condition (5), (6), and (7) in example 1, with k=n. This time, we are going to define a functional V on the Sobolev space HhiCl, R). We set V(x) = u{co, Vx(co))d(o It is easily seen that V is finite everywhere and has the C* property over Hq. Equation (8) for the derivative here becomes dt F(x+it)) = u'u{(o, t = 0 Jn Vx{(o))Vv{o))d(o The right-hand side is certainly a continuous linear functional of v. We want to write it as (V'(x), v), with V'{x) e ^ This is done by integrating by parts V{x+tv) dt n ^ Z T— tt'yioi, yx(03))d(j3 n i = i 00),• There is no boundary term because of the condition ve Ho (and not H^). The sum on the right-hand side is the divergence operator. We finally write E'(x)=-div(«;,° Vx) ■ Example 4 We can now combine examples 1, 2, and 3 to get the usual functionals of the calculus of variations. For instance, let «i and «2 be two functions satisfying conditions (5), (6), and (7), with ki =k and kz =nk, and consider the functional {/: defined by t/(x) = [«1 (O), X(C0)) + M2(0), Vx(0)))]i/0) Jn This is a functional, and its derivative U(x) is the distribution in H~ ‘ which is defined as Mi(o), x(co))-div «2(0), V«(0))) If U attains its minimum over Hq at some point x, then the function x
CH. 1, SEC. 4 DIFFERENTIABLE FUNCTIONALS 23 satisfies the Euler-Lagrange equation u\{0), x(co))-div M2(co, Vm(co))=0 in For instance, if x minimizes Dirichlet’s integral (with <p given in iJ) U{x)= £ [<0(coWa))+l(Vx(ü)))2]i/cü then X solves Laplace’s equation (;()(io)=div Vx(a))= j] ^(co) 1 = 1 doJi <P=Au Of course, in most cases the growth conditions on ui will be too restrictive. They can be considerably relaxed by using the Sobolev embedding theorems and inequalities. For instance, if « = 1 and fi is a bounded domain, the norm is stronger than the C° norm, and it will be sufficient that t/'i(co, y) be bounded for (cu, y)eiixB, whenever 5<= is bounded. ■ These various definitions can be extended to maps between Banach spaces; there are two different ways of doing this. Given a map (f> : X-* Y, a point xe X and a map A e^^{X, T), we shall say that (a)' A is the Gateaux derivative of (/> at x if lim - \\(t)(x-\-hy)-(l)ix)-hAy\\=0 h-^O n (b)' A is the Fréchet derivative of (/> at x if Jim ||<^(x+y)-</>(x)-^y|| =0 (a)" A is the weak Gateaux derivative of </> at x if Vp e Y*, lim - (p, 4>{x+hy)-(l>{x)-hAy}=0 h-^0 n
24 CH. 1, SEC. 4 BACKGROUND NOTES (b)" A is the weak Fréchet derivative of </> at a: if V/> e y* lim TT^ (p, <f>(x+y)-<t>(x)-Ay)=0 Clearly, (by =>(b)"=>(a)", and (b)'=>(a)'=>(a)" with the same henceforth denoted by 0'(x), and called the derivative of </> at x. This relationship enables us to define—again, in several different ways—the second derivative of a function 1/ : IR as the derivative of U': X'^). We single out the most and the least restrictive definitions to single out two important classes. A function U: X-> R has the property if U : X-^ X'^ is Fréchet differenti¬ able everywhere and U\x) € S£{X, X'^) depends continuously on X. A function U\ A"-^R is twice weakly differentiable Gateaux if U : X-^X'*^ is weakly Gateaux differentiable everywhere. Example 5 Let Q be an open subset on R”, with the usual Lebesgue measure rfco. Let there be given a function u: QxR*^->R, borelian with respect to both variables (co, y). Assume that (17) for any fixed co 6 £2, the function y-^u{co, y) is over R*^ (18) there is some constant c such that \uyy{(o^ y)\ ^ c all (io, j;) € £2 X R^ (19) I Iw(co, 0)|dA:<00 and | \uy{co,0)\d(o<co h h Then the functional U: L^(£2, R*)-^R given by U(x) = u(co, JCl x{(o))dœ is twice weakly differentiable Gateaux. ■ It is a simple matter to check that the growth condition in example 1 is satisfied, so that by example 2, 17 is a functional, and lf(x)=Uy^ X el} For any z e Û, we have
CH. 1, SEC. 4 DIFFERENTIABLE FUNCTIONALS 25 ( U(:»;), z) = (z(co), u'y(io, x((o)))dco Ja Again, the integrand satisfies all the assumptions in example 1, so Wm \ {If {x+hy)-U(x), z) = f (z((o), x(oi))y{o3))dw /»->0 n Jci So [/ is Gateaux differentiable everywhere, and (20) (z, U"x)y)= [ (z{o}), Uyyico, x{ù})y{(û))dœ Jn Example 6 Same assumptions as example 5. We now claim that U cannot have the property unless u is precisely quadratic, that is, u{(o, y)=Uy, with A{œ) e (R^) ■ This fact was noticed by Skrypniak, who used it to point out a few gaps in the original work of Palais and Smale, which fortunately turned out to be of little consequence. However, it shows how prudent one must be in making differ¬ entiability assumptions in infinite-dimensional spaces. What we have to prove is that if Uyy(œ, • ) is not a constant matrix for almost every CO 6 Q, then the map U” : l3) cannot be continuous. This map is defined by the integral in equation (20), from which it follows that U" will be continuous if and only if the map (¡){x) = Uyy^ X is continuous from 1} to Now, it follows from condition (18) on Uyy that </> maps 1} into L°°. Unfortun¬ ately, we have the case q = oo where KrasnosePskii’s theorem does not hold (see example 2): We cannot conclude that (¡) is continuous. To avoid technicalities, we assume from now on that Uyy is continuous in (o), y) jointly. Assume that for some coq ^ Uyy(o)o, • ) is not constant, say Wyy(Cl)o, ^l)^Uyy{0)Qy (^2) Let xo be any continuous function such that Xo(coo) = ^i. With any 8>0, we associate the function ( if|co-cuo|<e ^ Vo(<^) if|cu —cool^e
26 CH. 1, SEC. 5 BACKGROUND NOTES For small enough £, we shall have u;y(co, xJ(Oji))^Uyy{(£>, Xoico)) if |co-cool <e When £-^0, we have Xe-^Xo in However, Uyy^ Xe does not converge to Uyy^Xo in since the set of points where they differ always has positive measure. So (/) is not continuous, and U is not C^. ■ 5. SUPPORT FUNCTIONS AND BARRIER CONES OF CONVEX SUBSETS Linear and convex analysis are based to a great extent on the Hahn-Banach Theorem, which, as is well known, has many different—but equivalent— formulations. We shall use the formulation that links the geometrical and analytical approaches, that is, the formulation that allows a characterization of closed convex subsets in terms of convex functions. With this tool at hand, we can pass at will from the sometimes cumbersome handling of closed convex sets to the more flexible and traditional handling of convex functions. We shall prove that any nonempty closed convex subset of a Banach space X is characterized by its support function defined on the dual X'^ of X by V/7 6 X*, <Tk(/?): = sup {p, x) € ] - 00, + QO] xgK Indeed, the Hahn-Banach Theorem states that K={x e X\>/p e X*, {p, x)<(7k(p)} In this formula, we can restrict p to range in the convex cone h[K):={peX^\G^{p)<^cx,] which is called the barrier cone of K. It measures the “boundedness” of K\ The larger the barrier cone is, the smaller is K at infinity. The negative polar cone of the barrier cone is the recession cone of K, since the following formula holds true ^xeK b(K)~=f]HK-x) A>0 But the importance of the role played by the barrier cone b( A3 of K lies in the following result: Let Xand Y be Banach spaces, A a continuous linear operator from Xto Y, and K a weakly closed subset of Y*.
CH. 1, SEC. 5 SUPPORT FUNCTIONS AND BARRIER CONES OF CONVEX SUBSETS 27 If zero belongs to the interior of the set Im A-\-h{K)y then A*^{K) is strongly closed in X'^, This “closed image” theorem will obviously be a precious tool. We conclude Section 5 with a “calculus” of support functions and barrier cones, which allows us to “compute” support functions and characterize barrier cones of images, inverse images, sums, products, and intersections of closed convex sets. DEFINITION 1 Let K be a nonempty subset of a Hausdorff locally convex space X. The function Ok’. X*-^Rkj{-\-co} defined by (1) V;; e X*, at^py.=^a{K, /7):=sup {p, x) xeK is called the support function of K. Its domain (2) b(A):= {p e X*\<Tk{p)< + <»} is called the barrier cone of K. It is readily seen that (3) and (4) ffjc is a proper lower semicontinuous convex positively homogeneous function that is equal to o^k), the support function of the closed convex hull of K b{K) is a convex cone (not necessarily closed). Before giving examples and formulas, we state and prove the version of the Hahn-Banach theorem that motivates the introduction of support functions. THEOREM 2 (HAHN-BANACH) Let a subset of a Hausdoiff locally convex vector space X. The closed convex hull co{K) of K is given by (5) co(A0={x eX\^pe X*, {p, x)<o^(p)} Proof Set k.= {x € XlVp e X*, (p, x)^<Tk(p)}, which is a^closed convex subset containing K. If ^(K)^K, there exists some point x€K that does not belong to co(A[). The separation theorem implies the existence of po e X* such that ffKiPo) < (Po> x), which is a contradiction to the fact that x belongs to K. ■
28 CH. 1, SEC. 5 BACKGROUND NOTES We now introduce the negative polar cone of the barrier cone. It is the cone b(A)" defined by b(A0~:={i^eAlVp6f>(A), PROPOSITIONS Let Kbe a closed convex subset. Then for every Xq e K, (6) b(A)-=n HK-Xo) X>0 Proof. Take any point Xq in K and let us set Lo'^=r\x>o HK—Xq). a. Take xeLo. For any A>0, there exists yxsK such that x=^yx — Xo). Hence, {p, x}==X((p, yx)-{p, (a ^o))- By letting X converge to 0, we find that (/?, x:)^0 whenever p 6 h{K). Hence, x e h{K)~. b. Conversely, let x6b(A3~ and A>0. Since xjX belongs to b(A3~j for all p eh(K), we have Consequently, since K is closed and convex, theorem 2 implies that x/A + xq belongs to K, that is, xsLq. ■ DEFINITION 4 The negative polar cone h(K)~ of the barrier cone is called the recession cone of K. A The first use of the barrier cone that we mention is the closed image theorem, which illustrates the merits of this concept, THEOREM 5 (CLOSED IMAGE) Let X and Y be Banach spaces, A a continuous linear operator from X to Y, and K^Y* a weakly closed subset of 7*. Assume that (7) 0 elnt(lm /H-b(A^)) Let A* e ^(Y*, X*) denote the transpose of A. Then (8) is strongly closed in X'*^ More generally. A* is proper in the sense that i. A'^ maps weakly closed subsets of K to closed subsets of X'*^. ii. For all strongly compact M <= X*, the set A*~ ^(M) is weakly compact.
CH. 1, SEC. 5 SUPPORT FUNCTIONS AND BARRIER CONES OF CONVEX SUBSETS 29 Proof. Let US consider a sequence of elements q„e K such that A*q„ converges strongly to some p in X*. Let y > 0 be such that by assumption (7), yB is contained in Im A + h{K). Then for all y ^Y, there exist points x e X and z 6 b(A!) such that (yl\\y\\)y=Ax-\-z. Therefore, {qmy}=\q, — (Ax+z))=— i(A*q„, x) + (q„, z}) y / y (10) <— (IWI \\A*q„\\ + aK,{z))< + <Xi y because the sequence of elements A*qn is bounded. It follows that the sequence of elements q„ is also weakly bounded and thus relatively weakly compact in Y*, Therefore, a generalized subsequence of elements q„> e K converges to some q that belongs to K, because K is weakly closed. Hence, A*qn' converges weakly to A^q. Since A'^qn' converges strongly to /?, we deduce thatp=y4*^. ■ Since we used the weak compactness of weakly bounded subsets, we have to be careful of the dual statement. THEOREM 6 Let X be a reflexive Banach space, Y a Banach space, A a continuous linear operator from X to Y, and K a weakly closed subset of X, Assume that (11) 0eInt(Im.4* + b(A:)) Then (12) A{K) is strongly closed in Y. k Examples and Elementary Properties We shall agree to set We observe that 1+00 when/7:^0 (13) (^x(p) = 0 when/7=0 More generally, let Pc A" be a cone of X. Then the barrier cone of P is the negative polar cone P~ oi P (14) b(/')=p-:={peX*K(/^)<0}
30 CH. 1, SEC. 5 BACKGROUND NOTES and the support function of P is defined by iO when/) 6P" (15) Ok{p) = 4-00 when/?iP“ As a consequence, we deduce from theorem 2 the following important result. PROPOSITION 7 (THE BIPOLAR LEMMA) If P is a closed convex cone of X, then (16) p={p-)- k When M is a vector subspace of X, the preceding formulas become (17) b{Ai)=M^ is the orthogonal subspace (annihilator) of M and if A/ is a closed subspace of X, then (18) Support functions of points {x}<=-M are the weakly continuous linear func¬ tionals on X* because (19) <^{x){p)={p,x) It is obvious that if B denotes the unit ball of X, then (20) o^b(p)=IIpII* is the dual norm of X*. Let /if be a closed convex subset of X; then (a) 0 €/if if and only if (b) K is symmetric if and only if <7^ is even. (c) K is weakly bounded if and only if is finite everywhere, that is, if and only if b(/C)=Z*. In Chapter 4, we shall prove the converse statement in theorem 8. THEOREM 8 Let a: X*-* Ryj co} be a proper positively homogeneous lower semicontinuous convex function. It is the support of the closed convex subset K defined by (21) /i::={x e Z|V/) 6 AT* </),x>«T(p)}
CH. 1, SEC. 5 SUPPORT FUNCTIONS AND BARRIER CONES OF CONVEX SUBSETS 31 Formulas for Support Functions and Barrier Cones (22) If Kc^L, then b(L)c:b(A) and n (23) If A::= n Ki where Ki <= Xi(i=I,, n), then 1=1 b(A)=n b(A:i), and <Tk(;>)= S M/'i) i = 1 i = 1 (24) If A::=co ( u a:.-\ then b(l^= f) b(A:i) and M/?)=sup \i6 / / ie I » e / (25) UAc=.^iX,Y), then b(:i(^))=^*-ib(/C), and In particular, (26) l(Ki,K2<=X, then b(^i +A:2)=b(A:i)r>b(A:2) and + K2(p)=®’ki (p)+Klip) (27) If P is a convex cone, then b( +P)=b( .K) ^ ^ ~ and / ^ \(^k(p) iipeP~ '’***'> = 1 + 00 MptP- (28) Ifxo6^,then (^K + xoiP) = <^Kip)+ <P. ^o> (29) If 6 ^{X, Y),LciX and Mc= y are closed and convex, and if 0 6lnt(.4(L)-M) then b(L n/4”^(M))=b(L)+.4*b(M), and for all/> 6b(X), there exists q 6 b(M) such that <^i,nA-i(M)(p)=i^r.(p-^*i)+‘^M®= ‘"f, \pdp-A*q)-YaMiq)'\ qe Y In particular, (30) liAe ^(X, Y), if M is a closed convex subset of Y, and if 0 e Int (Im A — M), then h(A-\M))=A*h{M)
32 CH. 1, SEC. 5 BACKGROUND NOTES and for all p 6 b(A there exists q e b(M) such that A*q=p and inf Mi) A q=^p (31) If Ki and K2 are two closed convex subsets of X such that 0 6 Int (/(fi -/sTj) then b(/i:ini^2)=b(^i)+b(/:2) and for all/j e b( ATi n K2), there existspi e b( andp2 e b{K2) such that p = Pi+Pi and <^K,nK2(p) = <^KriPi) + <^K2(P2)= inf {(^KtiPl) + <^K2(P2)) P = Pl +P2 Remark Formulas (22)-(28) are straightforward. Formula (29) is not obvious at all: It follows from Chapter 5. Formulas (30) and (31) are obvious consequences of formula (29).
CHAPTER 2 Smooth Analysis The inverse function theorem and the (closely related) implicit function theorem are two of the very few general methods available for studying nonlinear prob¬ lems. The importance of these theorems in modern analysis, whether finite or infinite dimensional, can hardly be overrated. This explains why the inverse function theorem appears at both the beginning and the end of this book. In Chapter 7, the theorem is proved under very weak hypothesis for set-valued, nonsmooth maps, whereas in the present chapter, we prove it with strong differentiability assumptions. The reader may wonder why we give two proofs, and two statements, of the inverse function theorem, since the later statement encompasses the first one. The reason is twofold. We do not believe that the more general statements are the clearer ones, and we prefer to give the theorem, explain it, and use it, in the differentiable case, before turning to more general, and perhaps less common, cases. Moreover, in the differentiable case, it is possible to give a constructive proof, that is, to devise an iterative procedure that actually converges toward a solution of the nonlinear problem under consideration. In other words, the solution is not only shown to exist, but can also be computed. The structure of this chapter now becomes clear. We state the inverse func¬ tion theorem and deduce the implicit function theorem, fully aware that better statements will be available later on. We emphasize the constructive aspects of the proof by giving two different algorithms to compute the solution. The second one (Newton’s method) is much more efficient than the first, classical, one and has recently received a great deal of attention. Later sections are devoted to applications. The depth of the inverse function theorem is revealed in that we are able to derive from it two very deep finite¬ dimensional results: Brouwer’s fixed-point theorem and Morse’s lemma on the singularities of a function. We then find we have all the tools needed to investigate equations/(A, x)=0, depending on a parameter A. In Section 4, our equation will be 0^(A, x)=0, and x will be a finite-dimensional variable, so that, in fact, we are investigating the critical points of a real-valued function 0(A, •) depending on a parameter A: We shall show that they appear (or disappear) in pairs. In Section 5, x will be infinite dimensional, and the equation /(A, x)=0 will have the trivial solution x=0 for all values of A: We look for another set of solutions, “bifurcating” from 33
34 CH. 2, SEC. 1 SMOOTH analysis the trivial one at some value of L Because of the practical importance of the problem, we also investigate the stability of the solutions we find, trivial or not. We conclude by stating and proving Thom’s transversality theorem, prob¬ ably one of the most brilliant achievements of modern mathematics. It is an everyday tool in topology and geometry and is making its way into analysis: We cannot do it complete justice, but we do provide a full proof and immediate applications. This enables us to end the chapter as we began it, by exploring Newton’s method. 1. ITERATIVE PROCEDURES FOR INVERTING A MAP We denote by X and Y Banach spaces and by ^ an open subset of X. Let /: y be a (nonlinear) map. We are interested in solving the equation (1) f(x)=y Assume that equation (1) has been solved for some particular case, that is, values (xo, jo) € x y have been found such that (2) f(xo)=yo The inverse function theorem then gives us conditions under which equation (1) can be uniquely solved in x for all y, provided we consider only values of x and y close to xq and yo- In its classical version, the theorem reads as follows. INVERSE FUNCTION THEOREM Assume is a mapping. let XqB^^ and yo^Y be given such that (2) (3) f{xo)=yo f'{xo) Y) has an inversef'{xo)~ ^ e^{Y, X) Then there is an open neighborhood Лof Xq such that f{J/')='f is an open neighborhood of у о and the map f: has a C^ inverse The derivative is given by (4) [/.^Ч(у) = [/'°/.^Чу)] \ for all у 6 Before proceeding with the proof, let us spell out what the theorem says. Note first that the derivativef'{xo) is required to be invertible. In the case where X is finite dimensional, this is tantamount to requiring/'(xo) to be injective or surjective. This means that Y has the same dimension as X and that in any system of coordinates for X and Y, the jacobian of f at xq does not vanish. Conditions (2) and (3) now read
(5) (6) CH. 2, SEC. 1 ITERATIVE PROCEDURES FOR INVERTING A MAP 35 .Vo = (/l(^o), •••./n(^o)) -(^o) dxi ^f” /V ) |^(^o) dxn ^{Xo) dx„ ^0 The conclusion of the inverse function theorem is that for any ye'f', there is a unique x 6 solving the equation f{x)=y (there might be many more out¬ side), and it depends smoothly (C‘) on x. But this is not all. As pointed out in the introduction, there are also iterative procedures to compute x from y; here is the simplest one (7) (8) (start from xo [run x„+1 :=x„-H/'(xo)‘‘(y-/(x„)) Here is another one, known as Newtorfs method istart from xo ^run Xn + \ .—Xn'\~f (^m) iy These procedures are illustrated in the following figure; we have taken X=R=Xy=0, and f{x)={x^-l)/4. Figure 1.
36 СН. 2, SEC. 1 SMOOTH ANALYSIS Simple method Xq=2,X„+i=X„-f (jc„)/6 = 1.708333333 X2 = 1.542266469 л:з = 1.431082585 X4 = 1.350630361 X5 = 1.289637732 Newton’s method •^0 2, Xfx + 1 = Xfi (.Xn)/3X|| ATI = 1.416666667 д:2 = 1.11053441 X3 = 1.010636768 лг4=1.000111557 д:5 = 1.000000012 Both methods are seen to converge to 1, the second one much faster than the first. Let us confirm these results. LEMMA 1 Choose any к € ]0, 1[. For any a 6 ]0, 1[, there is some e>0 such that whenever Цл:—Xo||<fi, we have (9) IU-/'(^o)“y'WNA: 11/'(^о)“‘(/(д^)-/(л:о)-/'(д:о)(А:-л:о))||<а||д:-д:о|| ^ Proof. Consider the map from to SF{X, AO defined by (11) ^^/-/'(хо)-У'(х) This map is continuous, since/ is a map and it is zero for л:=л:о. Inequality (9) then follows by continuity. Since/ is Frechet differentiable at xq, for any )3>0 some e>0 can be found such that Цл:—Xol|<8 implies (12) ll/(^)-/(.^o)-/'(xo)(;c-;co)||<j8||x-;col| Taking P=а/||/'(л:о)“'ll yields inequality (10). ■ LEMMA 2 i/=(l -а)е/||/'(л:о)‘Ч|. For any у such that lb-/(xo)||<i?, the map (13) 9'(л:)=л:+/'(а:о)'‘(;;-/(л:)) is a contraction of the closed ball Xo + sB into itself A Proof We have /(д:)=/—/'(дго)“'/'(д:). It follows from lemma 1 that ||/(a:)||^A: on ло + еД so that g is к Lipschitz on that ball, with A:<1. All we have to prove is that g sends xo + гВ into itself.
CH. 2, SEC. 1 ITERATIVE PROCEDURES FOR INVERTING A MAP 37 Take any x such that ||x—xo|| <e. Let us check that Xo|| <«• We have (14) g(x) -xo=x-xo +f'(xo) ~' -f{x)) =f'{XoY '^(f{Xo)+f'{Xo){x-X(i)~f{x))-\-f'(X(iY\y-f(Xo)) Hence, by inequality (10) (15) llfif(:>c)-A:olKa||.x-;iCo|| + (l-a)e<e ■ LEMMA 3 For any y such that ||y—/(xo)||<>7, the sequence x„ generated by algorithm (7) converges geometrically to the one and only solution off{x) =y in the ball Xq + sB (16) ||a:„—x|| < constant-k" Proof. The first part, and estimate (16), follow immediately from Banach’s contraction principle applied to the contraction g and the closed ball xo + ^B. LEMMA 4 Take any two points y andy within distance q off{xo). Let x and x solve equations f(x)=y andf(x)=y in the ball Xq+vB. We then have (17) Proof. Set (18) 9(x)=x+f'{xo)-^{y-f(x)) (19) 0(^)=^+/'(xo)'‘(y-/(;c)) We have g(x)=x and g{x)=x. It follows that (20) Il^-^ll = ll0(^)-0(.x)|| <\\gix)-gix)\\+\\m-9(x)\\ Using equations (18) and (19), we get (21) llgf(-x)-0(i)||<||/'(xo)“‘|| lly-yll
38 CH. 2, SEC. 1 SMOOTH ANALYSIS Lemma 2 tells us that g is k Lipschitz, so that (22) ll0(^)-0(^)ll<A:||x-.x|| Substituting this into inequality (20) and remembering that 0 <it < 1, we get the desired result. ■ This almost proves the inverse function theorem. If we define (23) r:={y\\\y-f{xo)\\<ri} (24) Ji:=f-Hr)n{x\\\x-xo\\<e} and we denote by/r the restriction of/ to Ji, lemma 3 tells us that/^ is bijective. It is known to be continuous, and lemma 4 tells us that//^ also is continuous (even Lipschitz). All that remains to be proved is that it is To begin with, note that/'(A:) is invertible for all x in Indeed, because of inequality (9), the series (25) I {i-f'{xorYix)T=s(x) n = 0 is convergent in S^(X, X). It is easily checked that (26) S{x)-{I-f'(xo)-rix))Six)==I (27) S(x)-5(x)(/-/'(;co)- Y(x))=I In other words, (28) f'ixo)- Y(x)S(x)=I=S(x)f'{xo)- Y(x) So f'(xy ‘ is well defined and equal to iS(x)/'(xo)” *. Now fix y in ‘f, and take some other point y in f. Set x=fY(y) and x=fY(y)- Since/ is Frechet differentiable at 3c, we have (29) fix)-fix) =f'ix)ix-x)+||x - 3c||e(x) with s(x)^0 in Y when x-^x in X. We rewrite this as (30) y-y=f'WYiy)-fYiy))+s°f:ir\y)\\fYiy)-f:r'^{y)\\ We note that e °fYiy)^Q when y-^y and that//* is Lipschitz. It follows
CH. 2, SEC. 1 ITERATIVE PROCEDURES FOR INVERTING A MAP 39 that the last term can be rewritten as tj{y)\\y—y\\, with r]{y)-^0 'fih^ny^y (31) y-y-'i(y)\\y-y\\=f'ixif?'^{y)-fjr^{y)) (32) fjr^iy)-fjr^iy)=f'{x) '^iy-y)-f\x) ^ri(y)\\y-y\\ The latter relationship means precisely that f Frechet differentiable at y, with derivative (33) U^^li^)=f'ix) ^=U'°fj^^iy)'\ ‘ The last term depends continuously on y, since/^Ms continuous and / is on and the inverse function theorem is completely proved. ■ We detailed the proof in this way because we wanted to show that the neighborhoods of Xo and 'V of yo =f(xo) could be evaluated. This is import¬ ant for some applications. Of course, we want 'f' and Jf to be as large as pos¬ sible, and formulas (23) and (24) tell us that 6 and y\ should then be chosen as large as possible. Looking up lemmas 1 and 2, we find that k should be chosen close to 1 and that a expresses a trade-off between 6 and rj. Using Newton’s method (8) instead of the preceding one will result in much faster convergence, albeit possibly on smaller domains ^ and 'f'. However, Newton’s method requires the function/ to be at least C^. To be precise, LEMMA 5 Assume f is C^, and pick any ó with 0<5<1. Then there exists some a>0 and 6 > 0 such that for all x and x belonging to xq + sB, we have (34) ll/'(^)-H/(x)-/(i)-/Mx-x))||^a||x-x||2 (35) l|/-/'(x)-y'(^)ll^¿ Set y=(x{i—k) ^ and ^=£[2||/'(a:o) ^|| max (1, 7)]"L For any y with lb“/(-^o)ll'^^) the sequence generated by Newton's procedure (8) converges to a solution X of the equation f\x)=y and satisfies (36) Proof. Since /'(xo) is invertible and /' is a continuous map, f\x) will be invertible for all x in some neighborhood of Xq. The existence of 8 such that inequalities (35) and (34) are satisfied follows from the fact that / is C^. Assume for the time being that Xi, xj,..., x„ all belong to Xo + e5. Calculating
40 CH. 2, SEC. 1 SMOOTH ANALYSIS x„+i by Newton’s procedure, we have (37) +1 - x„=x„ -x„-i +f'(x„) ~Hy-f (a:,,))-f'{x„ _ i) “' (j -/(x„ _ i)) =/'(■^»1 -1) “ -1) +f'{x„ -1 )(a:„ -x„-i)-f (x„)) + if'ix„)~^~f'(x„.,)-^)iy-fix„)) =f'(x„-i)~^(f{x„.i)+f'(x„-t )(x„ -x„-i)-f (x„)) + _ 1) - YixMXn + 1 - x„) Using estimates (35) and (34) with jc=x„-1 and a:=a:„ yields (38) 11 a:,. +1 - a:„| I < a| I- a;„ -11 p + A:| I a:„ +1 - x„| I Finally, (39) I|a:„+1 - a:„|| ^ ,1 From this, we easily derive relationship (36). Of course, all these calculations hold true only if the successive points x„ belong to xo+eB. We check this by induction. We have (40) We claim ||x„ —Xo||^e(l — 2 ”). It is true for « = 1. If it is true for it is true for («-fl) (41) 2” ^8(1-2-') + !^) ^8(1-2-'-^) The result follows. It should be noted that the very rapid convergence we obtain by Newton’s methods exacts its price, since it requires us to invert f'(Xn) at every step. It has therefore been attempted to improve the procedure by doing away with operator inversions while retaining the quadratic convergence. Here is one such scheme (42) Xn+i= Xn + Aniy-f {Xn)) (43) A„+i=An + A,,(I-fiXn+i)A„) and Aq = I Here An e X) can be thought of as an approximation off'{xn)~^ that
CH. 2, SEC. 2 milnor’s proof of Brouwer’s fixed point theorem 41 is close enough to give quadratic convergence. Such schemes, however, do not in principle require thatf\x) be invertible: Proceeding in that direction, we are able to prove very sophisticated “hard” inverse function theorems, which happily fall beyond the scope of this book. Another remark about Newton’s method. We can imagine a continuous procedure instead of a discrete one, thus transforming the induction formula (8) into a differential equation (44) J=f\x)-\y-f{x)) This can be rewritten as (45) f'ix)^^=y-f{x) Any solution x(t) of this equation, starting at a point xo where f'(xo) is invertible, has the property that (46) ^ [y-fixity)] =-[y-f (40)] This integrates to (47) y-fixit))=iy-fixo))e' as long as the solution 40 is defined. Now certainly, iff'ixo) is invertible and ;; is close enough to/(xo), the solution x(t) is going to stay in some neighborhood of xo where f'(x) is invertible, so that x{t) will be defined for all i>0, and f{x(t)) -^y exponentially when r->oo. But the really interesting things occur when we start at a point xq where f{xo) and y are far apart. Then, either the solution of the differential equation (45) runs into a point x wheref'{x) is not invertible, or it converges to a solution X off(x)=y. This situation lends itself to further analysis, and we shall come back to it in later sections. ■ 2. MILNOR’S PROOF OF BROUWER’S FIXED POINT THEOREM Nothing seems more appropriate to illustrate the depth of the inverse function theorem than to derive Brouwer’s fixed point theorem from it. The proof we give is due to John Milnor (1978). As a matter of fact, he derives a stronger geometrical statement: There is no way to “comb” an even-dimensional sphere.
42 CH. 2, SEC. 2 SMOOTH anlyysis THEOREM 1 Let be the unit sphere in ^ Then any continuous tangent vector field on must have a zero. ^ Note that this is obviously false for and (less obviously) for all odd¬ dimensional spheres. The way dimensionality comes into Milnor’s proof is perhaps its most pleasing feature. We do not assume anything about dimension¬ ality to begin with: We simply work on an m-dimensional sphere S^. We start by restricting ourselves to vector fields. Indeed, let ^ be a continuous tangent vector field. Using, for instance, the Stone-Weierstrass theorem, we can find for each 8>0 some vector field ^ such that ||<i —in the uniform norm. Projecting on the tangent space to ST at X, we get a tangent vector field ^[(x) that is still and still satisfies ||(^ — ^e|| < 6. Now if the theorem has been proved for the case, then for each 6>0, there will be a point XeSST where (^¿(A:e)=0. Letting £-^0 and using the compactness of ST, we obtain a cluster point x where (^(3c)=0, as desired. We now proceed with the proof. Assume is a nonvanishing vector field. We shall prove that m is odd. Since for all x, we can associate with every t eR a. map (/),: ■s/l + ?sr, as follows: -1 Since ^{x) is tangent to S” at x, we have ||<^,(x)|| = VT+7^, so that </>, does send the unit sphere 5™ into the sphere + Certainly, (/►, is C‘ for all t. LEMMA 2 There is some e>0 such that (j>t is a diffeomorphism for all lil<e. A Proof. Define a map \j/,: by il/,{x)={i+t^) ^'^<l>,{x) Clearly, (f>, is a diffeomorphism from S™ to VT+f%“ if and only if [¡/, is a diffeomorphism from 5” to S”. We shall work with ij/, instead of </>, so that the range does not change with t. Take any point xeST', and choose a local coordinate system (0i, . . ., 0J for the sphere, valid in some neighborhood of x. Since i/^o is the identity, so that ^o(x)=I, and ij/'t depends continuously on (f, x\ there is some e>0 and some neighborhood ^ oix such that whenever |f| < e and x e we have \\i-mr^nx)w^\ II m-'mx)-nx)-m)ix- .\\x-x\\
CH. 2, SEC. 2 milnor’s proof of Brouwer’s fixed point theorem 43 From the results of the preceding section, there will be an open neighborhood f of x=\l/o(x), and for every |i| <e, a map <o,: f such that i/',»o),=/y. If we carry out this construction for all jc in 5™, we define two coverings of the sphere by open subsets. By compactness, there will be some eo>0 and finite subcoverings , ^„) and {'f i,..., iT„) with the property that for all |f|<6o, there is a map cuf: such that ij/, — Since {'f i,..., 'f„) is a covering, it follows that is surjective for |f|<6o. We claim there is an Sj > 0 such that ij/, is one to one for |i| < «i. If it were not so, there would be a sequence f„->0 and points y„^y'„ such that fl/t„(yn)=^t„{yn)- By compactness, these sequences would have cluster points j and /, with tl/o{y)= >l/oiy'), so that y=y'. For n large enough, y„ and y'„ would belong to some common *, and this would contradict the existence of (o,^. We have proved that, for |t| <min(so, Si), the map ij/, is bijective. Its inverse must be CO? on hence the result. ■ We now extend (f>, to the unit ball by setting (1) to:=||x||<^,(x||x||-‘) forllxNl The map defined in this way is clearly a diffeo- morphism for |i| < s. We use it to compute the volume of + ^ Namely, we use the well-known formula for changing variables in a multiple integral. Vol (Vl + i^B"+‘)= [ |Det ^',{x)\dx jBm+1 There is no need to compute precisely the right-hand side. Note only that Det 0i(x) does not vanish, so that it has the sign of Det (l>oix\ which is positive. Note also that all derivatives are taken with respect to Xy so that all the terms of the determinant are affine functions of t and the final result may be a poly¬ nomial in t Vol (Vr+7^B’"+i)=polynomial (f) Now this is where we have a contradiction, for the left-hand side is homog- eneous of degree (m +1): + Vol (5'”polynomial (i) This obviously implies that (mH-1) is even, hence m is odd, as stated. We now derive a similar result for vector fields on the unit ball, which will now hold true in all dimensions.
44 CH. 2, SEC. 2 SMOOTH analysis THEOREM 3 Let B"-»/?" be a continuous vector field on the n-dimensional unit ball. Assume that on the boundary S‘‘~^ of B", this vector field points outward (2) Then ^ has a zero |lx|| = l=><:ic, i(x))^0 (3) 3xeB" suchthat i(3c)=0 Proof. We study successively the cases when n is even and odd. a. Assume first that n is even. Extend to a vector field |; 2B"-yR” by settinjg setting i{x):=i{x) if|W|<l f(x):=(|W|-l)x+(2-|W|K(ji^) ifl<|W|<2 Clearly, ^ is continuous and coincides with ^ on B". For ||x|| > 1, that is, for all X outside B", we have m, x>=(|lx|| - i)\\xr + (2- lIxIDIIxll ( i XIWI-1)IWI^>0 So J does not vanish on 2B"\B". On the boundary 25"'“^ of 2B", we have ^{x)=x. We now identify 2B" with the southern hemisphere of the sphere 25" and extend ^ by symmetry to a continuous tangent vector field J on 25". This is best explained by Figure 1; the continuity across the equator is ensured by the fact that ^ is normal to the boundary. According to theorem 1, since n is even, I must have a zero on 25". By symmetry, there must be a zero in the southern hemisphere, so that ^ has a zero on 2B". Since ^ does not vanish on 2B"\B", this zero must belong to B", and the result is proved. b. Assume now that n is odd. Then B" can be identified with a subset of B"'^\ and any continuous vector field B^-^R" can be extended to a continuous vector field B""''' ‘ as follows:
CH. 2, SEC. 2 milnor’s proof of Brouwer’s fixed point theorem 45 =|(Xi, . . . , X„+i) X jB”={(x„...,x„+i)65"+‘|x„+i=0} • • • » Xji + i) —...» Xn), ¿IXfi+i) Here a is a positive number, to be determined presently. For any point (xi,..., x„+1) on the boundary of B"''' we have M + 1 __ n ^ • • • > -^H+l) ^ ••• 9 ^n) 1 i = 1 1=1 M / n = E x„)+a (1 - E i=l \ 1=1 Now if the vector field ^ on F' points outward on the boundary, we will be able to choose a>0 so large that the vector field ^ on also points outward on the boundary. Since n is odd,« +1 is even, and so ^ must have a zero: ^(xiy..., x„+1)=0. Looking up the definition of we see this means that 1 = 0 and ^(xi,..., 3c„)=0, as desired. ■ N = >'?=4| We have represented a diffeomorphism of 2B" onto the southern hemisphere of 25", taking the center of 2B" to the south pole of 25" and the boundary to the equator. The vector field ^ on 2F‘ is transformed into a tangent vector field on the northern hemisphere by the formula i(yiy . . • , yn, yn+l)=- l(yi9 • • . , ym -yn+l) for 1 ^ I ^ n
46 CH. 2, SEC. 3 SMOOTH ANALYSIS Of course, if a continuous vector field on 5" points inward on the boundary, (x, (J(x))<0 for ||x|| = 1, then it also has a zero, because — i* will point outward and (J has the same zeroes as — <^. As an immediate corollary, we have Brouwer’s theorem. THEOREM 4 (BROUWER) Any continuous map of the unit ball B" into itself has a fixed point. A Proof Let f:ff'-^ff'b&a. continuous map. Then ^(х)=/(л:)—x is a con¬ tinuous vector field on B" that points inward on the boundary. By theorem 3, there is some point x where 0=^(x) =/(x)—x. ■ Of course, all these results extend to a wider class of objects. We shall later see that theorem 3 extends to closed convex subsets of infinite-dimensional vector spaces and continuous compact maps i. We note now that Brouwer’s theorem extends to any topological space X that is homeomorphic to some n-dimensional ball B". Indeed, let g: X^B' be this homeomorphism, and let/: X^Xbe a continuous map. Then g°f °g~^ is a continuous map of В into itself, which, therefore, has a fixed point y= 9°f (j')- Setting X=gf “ ‘ (j) € AT, we get X =/ (x), so that x is a fixed point of /• Among such spaces X homeomorphic to B", we single out for later use the «-dimensional simplex И+1 ^ x, = 1 and X(>0 for all i i=l COROLLARY 5 Any continuous map of into itself has a fixed point. 3. LOCAL STUDY OF THE EQUATION /fxj=0 Let X and Y be Banach spaces and/: Y a smooth map. We are interested in the subset E of X defined by the equation/(x)=0 (1) £:=/-‘(0)={x|/(x)=0} More precisely, we wish to know if the subset E is in some sense “smooth.” Of course, if/ is continuous, then E is closed. But this is meager information, since a closed subset can be very bad indeed, such as a Cantor subset of the real line or even worse. Perhaps we can do better if/ is smoother, (7, for instance, with 1 ? The following result gives us the answer, and it is negative.
CH. 2, SEC. 3 LOCAL STUDY OF THE EQUATION/(x) = 0 47 THEOREM 1 Let E be any closed subset ofR". Then there is some C® function f: that E=f-\a). ► R such k Proof. If F is closed, its complement = is open. It is a standard property of R" that any open subset is the reunion of a countable family of open balls. There is, therefore, a sequence B*, A: 6 N, of balls in R", the center of B* being Xk and its radius p^, such that n= U B, kef^ (2) Define for each k&C°° function/*: R”-+R by i. /kW=exp [-(p*-|lx-x*||)”^] »• /tW=0 ifxiB* lixeBk It is easily checked thatis and vanishes with all its derivatives outside Denote by p = {pu • • • element of by \p\ the sum Y!i = iPi DJ the partial derivative For each keM, choose Sk>0 so small that (3) Now set Vx e B*, IpI \D%ix)\ < (e*2'‘) * /W= E £kfk{x) k = 0 We claim this defines a C® function/: R"->-R. It will obviously be non¬ negative everywhere and zero only if all theyi(x) are zero; that is, if x does not belong to U“=o B*=Q=R''\B. The result then follows. Now to prove our claim. Take any p 6 N", and consider the series of partial derivatives (4) Z X ej)%(x)+ S 6*Z)%(x) fc = 0 fc<ipi fc^ipi The first term on the right-hand side is just a finite sum. As for the second, we see that \SkD%{x)\ =0 \fxiBk and is less than 2"^ if x e by the definition of £fc, so that VxeR", E 8*|£»%(x)|< Z 2-*<00 fc^ipi k>\p\ The series in equation (4) is uniformly convergent for any choice oi p e N”. It follows that/ is C®, as announced. ■
48 сн. 2, SEC. 3 SMOOTH analysis Therefore, if we want/'^(0) to be smooth, it is not enough for/ itself to be smooth; something more is needed. This is precisely where the inverse function theorem comes in. PROPOSITION 2 Let f'.X-^YbeaC*' map, 1, between Banach spaces. Let x be some point in f~^{0)=:E. Assume that f\x) is surjective and that there exists a continuous projector Tit from X onto Ker f\x). Then there are neighborhoods 'V of x in X and ^ of the origin in X, and a C diffeomorphism p of ^ onto 'f' such that p{0)=x and (5) /-i(0)nir=p(Ker/'(x)n^ Proof We set N\=(I-n)X. We start by bringing x to the origin: 3c=0. Any xe X can be written x=Xi +X2, with xi eN and xi e Ker/'(3c), and this decomposition is unique. We first apply the implicit function theorem to the map il/: Nx Ker/'(3c)-^ Y defined by il/(xu X2)=f{xi-\-X2y Here, X2 is considered a parameter, while xi is the true variable. The derivative iAxi(0,0) is the restriction to N of/'(0), which is an isomorphism between N and Y by the open mapping theorem. It follows that for suitably small Xi and X2, there is a unique C solution xi =g(x2) of the equation il/{xu X2)=0 with g(0)=0. In other words, p: W^Ker/'(0) satisfies the identity f{g{x2) + x2)=0 near 0 We now apply the inverse function theorem, this time to the map p: Xi~^X2-^Xi +g(x2) + X2, defined on a suitably small neighborhood of the origin in X. Clearly p'(0, 0) is invertible, so that p" Ms well defined and C in a neigh¬ borhood of the origin in X. Setting xi =0, we get p(0, X2)=g(x2)+X2, so that / o p(0, X2) is identically 0 e T The converse follows from the uniqueness in the inverse function theorem, and the result is proved. ■ It may not seem so at first glance, but proposition 2 gives us a lot of informa¬ tion about the subset E near x. For instance, the following configurations in Figure 1 are excluded. The only allowable configuration near 3c is the following: The map p“ ^: 'f'“straightens out” the subset E near 3c, turning it into a linear subspace. If this can be done for any choice of 3c in E, we say that £* is a submanifold of X. Let us give a formal definition (setting p~^=il/). tA projector is a linear map n such that n^=n. In Hilbert spaces, and finite-dimensional spaces, every closed linear subspace is the range of some continuous projector. This is no longer true in general Banach spaces.
CH. 2, SEC. 3 LOCAL STUDY OF THE EQUATION f{x) = 0 49 Figure 1. (a) Cluster, (b) self-intersection, and (c) cusp. Ker f (x) DEFINITION 3 A subset E of X is a C submanifold, if for any xeE, there is a neighborhood ^ of X in X, a neighborhood 'f' of 0 in X, a continuous projector n from X onto L, and a C diffeomorphism such that Lnr = ^(En^) If dim L=p<oo for all x, we say that E has dimension p. If dim (I—n)X= q<oo for all x, we say that X has codimension q, A A closed submanifold is the formalization of one’s intuitive idea of a smooth surface, without self-intersections and cusps. The requirement that the sub¬ manifold be closed helps to eliminate other singularities; for instance, the following spiral (center excluded) is a manifold, but it is not closed (Fig. 2). The finite dimensional case, dim X=n<oo, is particularly interesting. In this case, for any x can describe il/(x) g ^ by its n coordinates (<^i,..., ^„) in some basis of X, which gives us a set of n equations U Eisp dimensional and modeled on the subspace L, we usually choose the p first basis vectors of Xin L, so that is described by \htq=(n—p) equa¬ tions i/^p+i(x)= • • • =il/„(x)=0. We then refer to as a local chart of {E, X) around X. Since is a diffeomorphism, the linear functionals l^i^n, are linearly
50 CH. 2, SEC. 3 SMOOTH analysis Figure 2. A nonclosed submanifold. independent on X, so that the q equations (x, \l/'i{x))=0 for p-\-1 define a /?-dimensional linear subspace of X. We set, indifferently, Te{x):=TxE:={x e X\<x, il/'i(x)> =0,p + i^i^n} and call it the tangent space to E at x. It is independent of the local chart ij/ that has been chosen for {E, X) around x. We can also directly define the position of x in ^ by the values of the l^i^n. We then refer to • • • > ^n) as a local, curvilinear, coordinate system for X near x, and En^ \s defined by the equations = • • * =^p=0. They are now linear, but the coordinate system is not! This finite-dimensional framework enables us to greatly extend proposition 2. We start with a definition. DEFINITION 4 Let beaC*‘ map, r ^ 1, and aC‘‘ submanifold of codimension q. Let X be a point in IR”. We say that g is transversal to M at x if either g{x) does not belong to M or g(x) belongs to M, and for some local chart ij/ of (M, IR"*) around g(x), the (q, n) matrix {({d/dxi)\l/j<^ g{x))\ l^i^n, m-q + \^j^m, has rank q. We say that g is transversal to M if it is transversal to M at every point x € IR". It may be easier to understand when g is not transversal to M at x. This means that g{x) belongs to M and the {q, n) matrix (((ô/ôx,#j ® ^(x))) has rank ^q— 1. In other words, the corresponding linear map from IR” to IR^ is not sur¬ jective. Geometrically, this means that there is a vector rj that is neither in the tangent space T-^M nor the image by the tangent map g\x) of some vector 6 [R” nor a linear combination of both. This property does not depend on the local chart ij/ that has been chosen for (M, 1R"‘) around x.
CH. 2, SEC. 3 LOCAL STUDY OF THE EQUATION/(л:) = 0 51 THEOREM 5 Let g: be a C map, r> 1, and McR"* a C submanifold of codimension q. Assume g is transversal to M. Then g~^{M) is a C submanifold of W with the same codimension q. A Proof We are going to construct around every point x eg~^{M) a local chart for (g~^(M), R”) of a special kind. Set N=g~^{M) and g{x)=y for the sake of simplicity. Choose a local chart ij/ of (M, R"*) around j;. In some neighborhood ^ of y, we have yeM<^il/p+i{y):=- - =il/„,{y)=0 Choose a global linear coordinate system . . . , for IR", denoting as usual by Tti the ith projection ^i=7ii{x), for allxelR" By assumption, the matrix ({{dldxi)7tj ^ g{x))), l^i^n,p-\-i^j^m, has rank q=m—p. This means that the q rows of that matrix are linearly independent. Introducing the transpose of g'{x) e R"”), we recognize in the yth row, the n components of the linear functional g'(x)*il/'j(y). It follows that we can choose (n — q) of the Ui, say tcj, ..., n„-q, in such a way that the n linear functionals (TCi,..., 7t„_„ g'(x)*il/'p+ i(y),..., g'{x)*il/'„{y)) are linearly independent. We recognize the n components of the derivative at X of the nonlinear map from R” to itself (tIi, . . . , 7C„_„ 4fp+i°g ^m°g) Using the inverse function theorem, we see that this map is locally invertible near X. Moreover, in some appropriate neighborhood of x, the set N:=g~^(M) is described by the q equations i^p+i ° gM= • • * ° g(x)=0. We have thus found a local chart for {N, R”) around x and thereby proved that TV is a sub¬ manifold of codimension ^ in R”. ■ Of course, if M is a closed submanifold, so is g~^{M), since g is continuous. As a particular case of theorem 5, we obtain the finite-dimensional version of proposition 2.
52 сн. 2, SEC. 3 SMOOTH analysis COROLLARY 6 Let be a C map, r'^l. Assume that for any x ef~^{G) = \N,the deriva¬ tive f{x) e if(lR”, U^) is onto. Then is either empty or a closed C sub¬ manifold of codimension m. A Proof This is an easy consequence of proposition 2, because in finite dimension, any closed subspace is the range of a continuous projector. It can also be proved from theorem 5 that the origin {0} is a closed submanifold of codimension m in (R"*, the tangent space being the linear subspace {0}. Trans- versality then means precisely that f{x) is onto whenever/(x)=0. The assump¬ tion in theorem 5 is thus satisfied and so is the conclusion. ■ The transversal case being disposed of, we now turn to nontransversal cases. We wish to see what N looks like near a point xeN where f\x) is not surjective. There are, of course, an infinite number of possibilities according to the way /' degenerates at 3c. We limit ourselves to investigating the simplest possible degeneracy. From now on, we take У = IR, which means that the set N is de¬ scribed by a single equation. DEFINITION 7 Let X be a Hilbert space and f: X-^U a C function, r^2. Any point x where f'(x)=0 is called a critical point. A critical point x is nondegenerate if its Hessian is nondegenerate as a quadratic form {r\x)y,z)=0forallz<^y=0 If all the critical points of f are nondegenerate, we shall say that f is a Morse function. A THEOREM 8 (Morse Lemma) Let X be a nondegenerate critical point of f Then there exist neighborhoods 'V of X and ^ of the origin and a C~^ diffeomorphism p of ^ onto 'f' such that p(0)=x and '^y e f{x)+\{f"{x)y, j) =/ ° p{y) к We will provide a proof for r> 2 only. PROPOSITION 9 (Hadamard Lemma) Let X and Y be Banach spaces,/: X-* Y a (7 map, r^l, and x a point where f{x)=Q. Then there exists a C~ ^ map u: X-^SC{X, У) with u{x)=f'{x) such that fix)=u{x){x-x) к Proof. Let t 6 [0,1] be a real variable. We have
CH. 2, SEC. 3 LOCAL STUDY OF THE EQUATION/(a:) = 0 53 'ixeX, fix) = x)'\dt = 1 <f'[x + tix—x)\x—x>dt = < I f'\x+t(x—x)~\dt,. , x—x> Calling the integral m(a:) gives the desired result. ■ Proof of Theorem 8. The proof of the Morse lemma now runs as follows. Start withf :X-*R and apply Hadamard’s lemma tof —f (x). We get a C “' map u: X-* X such that u(x)=0 and fix)=(m(a:), x-x)+f (x) Apply Hadamard’s lemma once more, this time to u. We get a C map w: X^SeiX, X) such that Vx € X, m(x) = w(x)(x —x) Substituting this into the preceding equation, we obtain Vx 6 A', /(x)=(w(x)(x—x), (x—x)) +f (x) Denote by w(x) the symmetric part of w(x) wW=i[w(x)+w(x)*] Clearly w: X^^iX, X) is still C' and Vx e A', fix)=(w(x)(x—x), (x—x)) +/(x) Comparing this with the Taylor expansion of/ at x and bearing in mind that /'{x)=0 and w is continuous at x, we obtain w(x)=2/"(x) which is invertible, since the Hessian is nondegenerate. We now introduce the operator t)(x)=(w(x)“‘w(x))‘^^ defined by i;(x)=/+ f c„[w(x)-‘w(x)-/]" 11=1 where the coefficients c„ are defined by Vl+f = l+Zr=i for the real
54 CH. 2, SEC. 3 SMOOTH analysis variable t. This series has radius of convergence 1, so i;(x) is well defined for ||vv(x)”‘h'(3c)-71|<1, that is, for x in some neighborhood of x and satisfies the relationship d(xMx)= iv(x)“ ‘ vv(x) It is clear from the series expansion that v(x) will be close to /, and hence invertible, when ||x:|| is sufficiently small. Moreover, since ivix) is self-adjoint, we have [ vv(x) “ ‘ iv(x)] * = w(x) w(x) “ ‘ so that [ vv(x) ~ ^ w(x)] * vv(x)=w(x)=vv(x)[ w(x) ~ ^ w(x)] Summing up the series expansion for v(x) gives d(x)* vv(x)=w(x)i;(x) so that i;(x)* w(x)d(x)=w(x)i;(x)t)(x) = w(x)vv(x)" ^ w(x) = w(x) In other words, the linear change of variables described by invertible operator i)(x) will always bring the variable quadratic form described by w(x) to the constant form vvix). We take advantage of this by the change of variables The right-hand side has d(x)”* as its tangent map at x=x. By the inverse function theorem, it is invertible near x=3c andy=0, and its inverse is precisely the map p{y) we are looking for. Indeed, we have fix)=(w(x)(x - x), (x - x)) -l-/(x) =(vv(x)p(x)y, i;(x)y)+/(3c) =(y(x)* vv(x)y(x)j, y)+f (3c) =(vv(%, y)+f{x) =hif"ix)y,y)+f{x) m
СН. 2, SEC. 3 LOCAL STUDY OF THE EQUATION/(x) = 0 55 This concludes the proof. The Morse lemma Is particularly striking in the finite-dimensional case: COROLLARY 10 (FINITE-DIMENSIONAL MORSE LEMMA) Assume f: R"->IR is a C map, r>2, and x a point where (6) . df OXi u. Det dxidxj' (^) M=o Then there exists a local, curvilinear, coordinate system (ii, ... , ^„) for R" near X and an integer k, with 0^k<^n, such that fix)=f{x)- Z Z if i = 1 for all X in some neighborhood of x. A DEFINITION 11 The integer k is called the index of the nondegenerate critical point x, A Corollary 10 follows immediately from theorem 8 by finding a suitable (linear) basis for IR”, so that the quadratic form with matrix 1 2 dxidxj In the old basis now reads ((±¿0)) (recall that the Kronecker symbol Sij is 0 if i^j and 1 if i=y). By standard results in linear algebra, this is always possible if the quadratic form is nondegenerate and the number of k times — 1 occurs does not depend on the way this diagonalization is performed. Note that in the coordinate system the function/ becomes pre¬ cisely quadratic: It coincides with its second-order Taylor expansion (there are no first-order forms). This very simple expression for/ is achieved at the expense of curving the coordinate system, which is no longer linear. The index of a nondegenerate critical point is a geometric notion, that is, it is invariant by diffeomorphisms (in particular, change of coordinates) of the base space. The index determines the behavior off near x, and, hence, the shape of N=f~^(0) near 3c. For instance, if k=0, then/ achieves a local minimum at 3c. If k=n, then/ achieves a local maximum at 3c. In the one-dimensional case, « = 1, these are the only possibilities for nondegenerate critical points. When «=2, a third possib-
Figure 3. Two-dimensional critical points; (a) index 0. (b) index 1, and (c) index 2. ility appears, namely, 1. We then have/(x)=/(x)+(^f - in an appropriate coordinate system near x\ We say that x is a saddle point (Fig. 3). Still in the two-dimensional case, we now have a complete picture oiE=f~'- (0) near a nondegenerate critical point xeE (so that/'(3c)=0). By the Morse lemma, there is a neighborhood ^ofx and a local coordinate system (^i, of valid in such that (k=0 case) 3c, with coordinates (0,0), is the only point of E in (k=2 case) 3c, with coordinates (0, 0), is the only point of E in ‘W. (k=l case) all points x in £ with coordinates (^i, ^2) such that —{2=0 belong to ‘W. The last case is the most interesting. The equation iJi — <^2=0 splits into two straight lines ^i = ±^2, which intersect at the origin. Therefore, N has two smooth branches that intersect at 3c. The situation is shown in the following figure.
CH. 2. SEC. 3 LOCAL STUDY OF THE EQUATION/(x) = 0 57 It follows from this analysis that all nondegenerate critical points are isolated. This is a general result, independent of the dimension; indeed, going back to theorem 8, it is clear that (/ ^ p){y)=p\yTf ^ p{y)=r\x)y^Q when y^O, so f\x)4^0 for X close to x but distinct from it. To conclude with a specific example, let us consider the subset Nx of defined by the equation {{Xi + lf++ xl)=x When X varies, the Nx are the level sets of the function f{xu X2) = [(Xi + 1)^ +x^][(x, - 1)^ +xi] Let us look for critical points by solving the equations ^ (xi, ^2) = 0 = ^ (Xi, X2) We obtain 2(xi + l)[(xi - \f +xi]+2(x, - l)[(xi + lf+xi\ =0 2X2 [((x 1 -1 )^+xi)+((^2 + 1 )^+xi)] = 0 The second equation yields Xx =0, and substituting it into the first one, we get Xi =0, — 1, or 1. So we have three critical points, the origin 0 and the points y4(-l, 0) and 5(1, 0). The Hessian at the origin turns out to be ^) = — 4xi + 4x2 The origin, then, is a nondegenerate critical point of index 1. The points A and B turn out to be nondegenerate critical points with index 0. The shape of the level sets Nx now is easy to understand: They will all be one-dimensional closed k< 1 k> 1
58 CH. 2, SEC. 4 SMOOTH analysis submanifolds, that is, smooth curves, except for those values of A such that Nx contains A, B, or 0. This occurs when A =0 (Nq is reduced to the two points A and B) and A=1 (Ni has a double point at the origin, all other points being regular). Geometrically, Nx is the set of points M in the plane such that MA • MB=y/k. The Ni is known as Bernoulli’s lemniscate. 4. BIRTH AND DEATH OF CRITICAL POINTS Say we have two C® Morse functions,/o and/i, on R"; then all critical points of /o are nondegenerate and so are all critical points of/i. Now connect/o and/i by a one-parameter family of smooth functions. By this, we mean a C® map: /: [0,1] X R''->IR such that (1) /(0, x) =foix), for all a: e IR" f{i, x)=fi(x), for all a: €l From now on, we shall write either /(A, a:) or fx{x), so that fx is the map a:^(A, x). Since /o and /i were chosen at random, they presumably have a different number of critical points. Therefore, the number of critical points of the function fx must vary when A ranges from 0 to 1. This means that for special values of A, critical points must appear or disappear. We wish to investigate the manner in which this happens. . We begin by two simple results from propositions 1 and 2. PROPOSITION 1 Assume thatfor some value A of the parameter,fx is a Morsefunction. Then for any compact subset K of R", there is some e>0 such that whenever |A —A|<6, the function fx has no degenerate critical point on K. A Proof Assume otherwise. Then there is a sequence A*-vA and a sequence Xx in K such that xx is a degenerate critical point offxx ^^(A*, a:ji)=0 for all A: G xu^=0 By compactness, there is a subsequence, still denoted by Xk, that converges to some X in K. Taking limits in these equations, we see that 'x is a degenerate critical point offii in K, which contradicts the assumption that/j is a Morse function ■
CH. 2, SEC. 4 BIRTH AND DEATH OF CRITICAL POINTS 59 PROPOSITION 2 Assume that for some value X of the parameter, fx is a Morse function. Let K be any compact subset ofW'. Then there is some 6>0, some neighborhood ^ of K, and a finite number N of maps — A + l^i^N, such that the ^j{X) are precisely the critical points of f in A Proof Since ^ is a Morse function, all its critical points are isolated, so it can have only a finite number in any compact set. It follows that there is some open neighborhood ^ oi K such that ^ is compact andy^ has no critical point in^\X. Let X be a critical point oif^ in K. It is known to be nondegenerate, so that det This means that we can apply the implicit function theorem to the n equa¬ tions f-{X,x)=0, OXi near (X, x): For (2, x) close enough to (I, x), the only solution to this system is given by x = ^(A), with ^ a smooth map. Let Xi,...,Xn be all the critical points offj in K. For each of them, we proceed as above. We end up with an e > 0, disjoint open neighborhoods of the xj, and maps ^j: ]A —s, A + such that whenever |A —A|<e and xe the equation/'(2, x)=0 implies that x = ^j{X) for some j. The proof is now complete, except that we might have/'(2, x)=0 with xi We claim this never happens when x e ^ and |A —X| is sufficiently small. Otherwise, there would be sequences x„ in ^ and with x„ a critical point of fx^ not belonging to Letting «-►00 and using the compactness of we obtain in the limit a critical point X of/j in ^ not belonging to any and, hence, different from all the Xj, But this is impossible, since the xj accounted for all the critical points off ■ Propositions 1 and 2 already tell us a lot. If X is a value of the parameter for whichis a Morse function, the onjy way the number of critical points of fx can change when X crosses the value X is for critical points to “come in from’’ or “run away to” infinity. A simple instance of this occurs with the function /(X, x)=Xx^ — X (here, « = 1 and X=0). If there is a bounded set K that contains all critical points offx, for 1 and if/o and/i have different numbers of critical points, then proposition 2 tells us that the fx, for 0<X< 1, cannot all be Morse functions. The values of X for which fx has degenerate critical points are precisely those for which critical points appear or disappear (at finite distance). By proposition 2, they constitute a closed subset of [0, 1].
60 CH. 2, SEC. 4 SMOOTH analysis Let us investigate this phenomenon more closely. We define the indicatrix of the family fx> DEFINITION 3 Let f:UxW-^Ubea function. The indicatrix of f is the subset 1(f) of IR X R" defined by (2) I(f)=\(K x)eUxW^ dxi (A, a:)=0, l^i^n We wish to study 1(f) under reasonable assumptions on f As we have just seen, the assumption that the critical points of fx are always nondegenerate would not be reasonable, since it would exclude precisely those families we find most interesting. So we weaken it as follows: Consider the determinants (3) and / \ dXdxi (4) A(x, A) = Det dxidxj ay dXdx„ dS dS dd ^ dxi dxn dX I Our standing assumption throughout this section is given in assumption A. ASSUMPTION A There is no point (A, x)elRx[R" that satisfies simultaneously the following equations: (5) 1. ii. iii. (A, x)=0 l^i^n dxi d(X, x)=0 A(A, a:)=0 Note that this is a system of («4-2) equations with (« +1) unknowns, so that it does seem reasonable to assume that it has no solution. A rigorous argument along this line, showing that assumption A holds for almost all / in C^, is possible and relies on transversality theory (see Section 7).
CH. 2, SEC. 4 BIRTH AND DEATH OF CRITICAL POINTS 61 LEMMA 4 Assume A. Then I{f) is either empty or a C°° one-dimensional closed submanifold ofUxU". k Proof. All we have to do, according to corollary 6 in the preceding section, is prove that the n x (« +1) matrix has rank n for all (x, x) e I{f). If d{X, x) ^ 0, the first n columns already have rank n. If x)=0, then A{X, x):^0 by assumption A. Computing the determinant A{X, x) from its last line, with d{X, x)=0, we obtain (6) i=l OXi So one of the cofactors C,(A, x) must be nonzero, say Ci(A, But we recognize in Ci(A, x) the determinant formed by the last columns of the matrix M{X, x). So M{X^ x) has rank n again. ■ We shall say that a vector (fi, y) eU xW is vertical if its first coordinate vanishes: /¿=0. LEMMAS Assume A and Then the plaints (X, 3c) g /(/) where the tangent is vertical are precisely those points where S{X, x)=0. Any such point has a neighborhood ^ such that 1(f) lies on one side only of the vertical hyperplane through (X, 3c). A _ Proof Assume first that ¿(X, 3c):?^:0. We claim that the tangent to 1(f) at (A, 3c) is nonvertical. Indeed, going back to the proof of lemma 4, this is the case when the first n columns of M(A, x) have rank n. By the implicit function theorem, this means that the n equations x)=0, i^i^n OXi can be solved in terms of X: There is a neighborhood ^ of (X, 3c) such that 1(f) n % is the graph of some smooth map x=p(X), So the tangent to 1(f) at (X, x) carries the vector (1, p'(X)\ which is certainly nonvertical.
62 CH. 2, SEC. 4 SMOOTH analysis Now assume that 5(A, x)=0. By assumption A, we have A(A, 3c):ji:0. Suppose we have the case where the last n columns of M(A, x) have rank n. By the implicit function theorem again, this means that the n equations dxi (2, x)=0, can be solved in terms of Xi. We shall now express {x2,..., x„, 2) in terms of a new parameter u, related to Xi by dx\ „ ^ =Ci = Det au ay ay ay dxidx2 dxidxn dx\dX ay ay ay dx„dx2 dx„dx„ dx„dX All the quantities involved are to be considered as functions of xi through (2, x), and since C((2, x)^i=0, this change of variable is well defined near xi =xi, u—0. We now compute dX/du and d^XIdu^. To do so, we have to differentiate the basic equations, (^/5x,)(2, x)=0, 1 < i < «, with respect to Xi (7) ^ gy dxj ay dx , Lj J.. 'a.. 2 1 J.. j dxidxj dxi dxidX dxi We rewrite this as (8) y gy dXj ^ gy dX ^ d^f 1 dxidxj dx\ dxidX dxi dxidxi This is a nonhomogeneous system of n linear equations in n unknowns dx2ldxi,.. ., dxjdxi, dX/dx^. Using Cramer’s rules to solve this system gives dXj Cj , dX 5 and dxi Cl (9) and hence, (10) Turning now to the second derivative, we have dx\ Cl dX dX dxi ^ du dxi du
CH. 2, SEC. 4 BIRTH AND DEATH OF CRITICAL POINTS 63 (P'X dS _ dd dxi_^ dS du^ du dxi du ‘ dxi ’’ 1 dxj dxi dX dxi\ r_^ y Cl ^ ±^1 ‘ j-fi Q dxj Cl dxj =A by definition of the C, Now set M=0 to get the point (A, x). We find «• ^(0)=A(I,x)^0 The result follows immediately. We now can visualize the indicatrix in the following figure: Here, the function fx has one critical point for X<Xi, two for Ai<A<A2, four for two for ^3<A<>l4, and none for X>X^. When 2. increases through Ai, a critical point comes from infinity. When A increases through A2, two critical points are born together at X2- When A increases through A3 or A4, two critical points kill each other at X3 or x^. The fact that critical points are
64 CH. 2, SEC. 4 SMOOTH analysis born or die in pairs depends only on the concavity d^XIdu^ not vanishing at the points where the tangent is vertical, where dX!du=0. We state this fact as a general result in proposition 6. PROPOSITION 6 Assume A. Also assume that in the interval [Aq, all critical points of fx are contained in some bounded subset K ofW\ Then there is a finite number N{X) of them, and there is a finite set Lcz[Ao, Xf\ of values of X such that fx is not a Morse function. N(X) changes by an even number when X crosses a value in L and is constant on all open subintervals determined by L. It follows that if and fx^ are Morse functions, then (12) TV(2o) = M^i)mod2 The fact that parity is preserved can be very important. For instance, if / has an odd number of critical points, so will fx^, which proves that fx^ has at least one critical point; that is, the equationfxfx)=Q has a solution. If we proceed as we did in the beginning, choosing Morse functions/o and /i at random, we shall always be able to connect them by a path satisfying A, even if/o has an even and /i an odd number of critical points. What happens then is that an odd number of critical points are thrown away at infinity or drawn in from infinity. We now turn to another useful tool. DEFINITION 7 The graphic of f is the subset G{f) of defined by (13) Gif) :=|a f(X, x)) dxi (A, x)=0 for ¿=1,...,« G{f) is the image of 1(f) in the plane by the map y: (A, x)^(A, /(A, x)) LEMMA 8 Let (A, x) be a point in 1(f) where <5(A, x)^0. Then there is a neighborhood of (A, x) such that the image of 1(f)by y is a C" one-dimensional subman fold of the plane through (A, /(A, x)), with nonvertical tangent at this point. A Proof Since ¿(A, x)^^0, lemma 4 tells us there is a neighborhood ^ of (A, x) such that I(f)r\‘2i is the graph of some C°° map Ah>x(A) near A. So the image of 1(f) by y is the graph of the C” map A i-»/(A, x(A)). Hence the result. Setting c=/(A, x), we obtain in this way a portion of curve through (A, c) entirely contained within G(/). Now there might well be other points x', x",... such that /(A, x')=c=/(A, x")= . . . , and, hence, other smooth branches of
CH. 2, SEC. 4 BIRTH AND DEATH OF CRITICAL POINTS 65 G{f) intersecting at (X, c). Because of this, G(f) is not a one-dimensional sub¬ manifold of the plane; but there is also another problem, as stated in lemma 9. LEMMA 9 Assume A. Let (A, ic) be a point in 1(f) where d(X, x)=0, and setf(X, x)=c. Then G(f) has a cusp at (A, c). The half-tangent at this point is nonvertical and lies inside the cusp. A Proof We use the parameter u introduced in lemma 5. By differentiating and remembering that df/dxt=0 along 1(f), we obtain (14) III. ^(0)=afti)=o du So the curve u^(X(u\ f{X{u\ x(u))) does have a singular point for w=0. To show that it is a cusp, we have to show that d^f/du^ and d^Xjdu^ do not vanish simultaneously at w=0. Differentiating once more, we obtain du 2(0)=A(A, 3c)^0 (see lemma 5) du^ dX^\du) dXdu^ =^(X,mlx) The direction of the half-tangent to the cusp is given by {») and its first coordinate is nonzero. To check the last assertion that the half-tangent lies inside the cusp, we have to prove that the following determinant does not vanish du^
66 CH. 2, SEC. 4 SMOOTH ANALYSIS The computation finds this determinant to be equal to and the bracketed term is ^ gy dxj ^ d^f dX /=1 dXdxi du dX^ du If this were zero, multiplying by duldx^ we would obtain y gy dx, ay dX j=i dXdxj dxi dX^ dxi Remembering the n defining equations, gy dxj ^ gy dX j = 1 dxidxj dxi dxidX dxi we find that dxjdxi = ly dxjdxu • • • , dxjdxu dX/dxi solve a homogeneous system of linear equations, whose matrix is dxidxj d^f dXdxi - - SH - _ / Some algebraic manipulations show that this matrix is invertible if ¿(1, jc)=0 and A{X, x)^0. So all the unknowns should be zero, including dxJdxu a contradiction. ■ We now can draw the graphic off and observe the life of critical points from another point of view. Let us, for instance, observe the same family fx as in the preceding figure. Here/(A, x(yl))-^oo as X^X^ H- and x{X) goes to infinity on the vertical branch of I{X\ but this might very well not be the case; for example,/(A, x(A)) could converge to a finite limit. The graph tells us that when two critical points are born together, corre¬ sponding critical values separate very slowly. We conclude our investigation by asking what the function/^ looks like near a point X such that (A, x) e 1(f). If ¿(A, then 3c is a nondegenerate critical point oifi, and the Morse lemma provides the answer: There is a neighborhood
CH. 2, SEC. 4 BIRTH AND DEATH OF CRITICAL POINTS 67 ^ of X, and a local (curvilinear) coordinate system a:=(^i, ... of R", valid in ‘if, such that x=(0,...,0) /R a)=/(I, x)±|f± •••Vxe^Hf If ¿(X, x)=0, then X is a degenerate critical point offj, and the Morse lemma fails. However, the assumption A(A, ic)^0 restricts the possible degeneracies, so that we shall still be able to put/^ in a canonical form near x. LEMMA 10 Assume A. Let (A, x) € /(/) be a point where ¿(1, ic)=0. Then there exists a linear coordinate system x={^i,..., of U” such that (15) Vx, /(X.x)=/(X,x)+(^i-fi)2±--- +(^„_,-?„_i)2 i^j^k where x={lu ...,Q and a„„„{^)f0. k Proof. Let (^x,..., ^„) be a linear coordinate system that diagonalizes the quadratic form f"i(x). We then have 1 = 1
68 CH. 2, SEC. 4 SMOOTH analysis Therefore, one at least of the must be zero; say it is the last one. By assumption A and lemma 5, we know that (i/(5/fliw)(0)=A(A, x)4^0. Com¬ puting this derivative with d^fjd^l=Q gives dd dx\ d^f f-T d^f No term on the right should vanish. It follows that fl i ^¡cl Cl. • .., ln)f0 a i ^(A,Ci,. ...c„)=o a i ^(A,Ci,. .., ln)f0 We can assume the d^fld^} are actually ± 1. We now use Hadamard’s lemma (proposition 9 in the preceding section) three successive times, to prove that (16) f{l x)-f(X, x)-ij"i,xtx-x\ x-x)=Y.aM^i-m}-l)\L-U ■ LEMMA 11 Assume A. Let (A, x) e /(/) be a point where 5{X, x)=0. Then there exists a neigh¬ borhood ® of X and a local (curvilinear) coordinate system x=(rji,, rjn) of W valid in ^ such that x=(0,..., 0) = for \^i^n-i /(I, x)=f{X,x)±ril± - ±t]^-i+ri^+ £ bijMmm ^ i^j<k ifn Proof. Set (17) = for / = (18) I ^(^.-l,)l L ^ if n tinmAS) J The determinant of the dr\Jd^j at | is a„„„(|)'^^, which is nonzero. By the
CH. 2, SEC. 4 BIRTH AND DEATH OF CRITICAL POINTS 69 inverse function theorem, the map •-►i/ is a local diffeomorphism near so that (^ij • • •, ^n) is a local coordinate system near 3c. Replacing the by the rji in lemma 10, we obtain /(X, x)±t]j ± • • • ±>/2_ !+/?(»/) m= E au,mi-mj-mu-^,) (19) =a„„„mn-l? + I ifn + E =ri^n- E ijfn ijfn Here the djk are coefficients of monomials with degree zero or one in — when expanding rjf^ in equation (18). Replacing the (^, by their values in terms of the rji, we obtain the desired result. ■ PROPOSITION 12 Assume A. Let (I, x) be a point in 1(f) where <5(1, x)=0. Then there exists a neigh¬ borhood ^ of X and a local (curvilinear) coordinate system (Ci,..., Cw) of R", valid in such that for all x in ^ /(I,x)=/(I,x)±C?±-* - ±C,?-1+C,i A Proof Consider the expression for /(I, x) obtained from lemma 11. All the monomials in the remainder have joint degree two or three in the combined variables (f/i,..., i). We shall give a procedure for absorbing them into the squares rjf. Let us, for instance, eliminate all remainder terms containing rj^. We first list them. Q(n)=biiMrii+'Z biikMiifik kti + E bijk(ri)iurijri^ 1 <j^k
70 CH. 2, SEC. 5 SMOOTH analysis We then have ±'/i + (2(>/)=>/i Z +'?! E bijkVijflk L kfl J Kj^n Bjri) T Bjtff 2A(ri)j 4A(nf We define a new set of local coordinates (Ci,..., C«) by Ci=\A(r,r'^ and Ci=»ii for i>2 By the inverse function theorem, this is possible since ^(0)= +1. We have B(rjf ±ni + Q{ri)=±i:l- 4A(rif Here B(riy 4A(t])- is a polynomial m(rj2,..., rjn) with variable coefficients. All its monomials have degree ofat least two jointly in(fj2,---,^n-i)-Turningtothevariables(Ci,..., C„) preserves these properties, so that we obtain fil x)=fil x)±a± • • • ±a-,+a + P(Cu Hz,..., U The new remainder P is a polynomial in (Ca,..., in) with variable coefficients. All its monomials have degree of at least two jointly in (^2, • • • , i„-i). This enables us to iterate the procedure, thus eliminating the variables Ci, • ■ ■ ■> Cn from the remainder and reaching the desired result. ■ If« = 1, we have/(A, x)=f (I, ic)+(x—3c)^ in appropriate coordinates, so that X is an inflexion point. In higher dimensions, the points x described in proposi¬ tion 12 will be called generalized inflexion points. 5. FURTHER DEGENERACIES: BIFURCATION In some situations modeled by the equationjT(A, x)=0, we may observe several solutions in X for parameter values close to A: a trivial solution, xo(A)=0,_say, and bifurcating solutions xi(A),..., x„(A), which coincide with Xo when A=A, so
CH. 2, SEC. 5 FURTHER DEGENERACIES! BIFURCATION 71 that Xi(X)=0 for all i but move away from it when 2:^1. We wish to gain some mathematical insight into such situations. We begin by a formal definition. We are given a map f:AxX^Y between Banach spaces, with k>l, such that VAgA, M0)=0 DEFINITION 1 We say that X is a bifurcation value of the equation f{X, x)=0 if every neighborhood of (A, 0) in AxX contains nonzero solutions. A In other words, there is a sequence in A and a sequence x„-^0 in X such that/(A„, x„)=0 and x„^0 for all n. Of course, bifurcation values are interesting only if they can be reached through nonbifurcation values of the parameter. For the zero map /(2, x)=0, all values of the parameter are trivially bifurcation values. A few facts are already clear. First,/x(2, 0) cannot be invertible. Otherwise, by the implicit function theorem, for X close to X and small x, there would be a unique solution of /(A, a:)=0, namely, x=0. Neither could lm\_fx{X, 0)] = Im[/'(A, 0)] be the whole space Y. Otherwise, assuming that Ker f'{X, 0) has a closed complementing subspace, we would apply proposition 3.2. The equation f {X, x)=0 would define in some neighborhood ^ x'f' o{ (A, 0) a closed submani¬ fold M containing^ X {0} and some additional points Thus, x {0}, so that ^ X {0} has empty interior in M, which means precisely that all Xe^ are bifurcation values. Remembering that is a neighborhood of I, we find ourselves in the situation we have described as uninteresting. So the only nontrivial possibility left is to look for higher degeneracy of / at (1,0); that is, to make/x(X, 0) nonsurjective. For the problem to be tractable at all, we make the assumption that/i(2, 0) is a Fredholm operator. This means that (1) Ker fx(X, 0) has finite dimension d. (2) Im[/x(X, 0)] is closed and has finite codimension r. In this case, the equation /(2, x)=0 in A x 7 can be reduced to a finite¬ dimensional system, namely, r equations in (¿/-hi) unknowns, by the so-called Lyapounov-Schmidt procedure. To see this, write A"=A"o©Ker /^(I, 0) and y = Im[/^2, 0)]©7o and denote by tc: Y^Yq the projection. The equation /(2, x)=0 splits into two parts (3) (4) 7T/(;i,;c) = 0 in To (/- n)f{X, x)=0 in Im[/;(X, 0)] The restriction off'x{X, 0) to Xq is an isomorphism onto lm[f'x(x, 0)]. But
72 CH. 2, SEC. 5 SMOOTH analysis it is also the derivative at 0 of the map ato), defined on A'o- Applying the inverse function theorem to the map Xo^(I-n)f{X,Xi+Xo) depending continuously on the parameters (A, xi) e A X Ker f'AK 0) we define on some neighborhood ^ of (X, 0) a C‘ map (A, xi)^Xo(A, a: i) such that (5) xx +Xo(X, xx))=0 V(A, Xx)e<^ The only condition left to satisfy then is (6) nf {X,xi+ Xo(A, Xx))=0 Equation (6) is called tlw bifurcation equation, and it is a system of r equations in the space A x Ker /'x(A, 0). We have thus reduced the problem to a finite¬ dimensional one: Find the zeros of a function gi: near some degenerate zero (7) [solve ^(A, a:)=0 near (A, 0) [with g(I, 0)=0 and g'(X, 0)=0 This is a formidable task! As a matter of fact, it is one purpose of singularity theory to provide partial answers to this question. We have already met this problem in Section 4 and have investigated the simplest case, namely, a single equation (r = l) near a nondegenerate critical point, Det g"{X, 0)^0, when g is C" with p^2. The Morse lemma tells us that there will be some neighborhood of (A, 0) in A X Ker f'xiK 0) such that the set of solutions for g{X, a:)=0 in ^ is C’~^ diffeomorphic to some cone C: d + I ^ +A?=0 in i = 1 the vertex 0 of this cone being carried to (1,0). Composing with the map (A, Xi)-^ (>1, Xi +Xo(^, x)) and assuming / itself to be we see that the set of solutions of /(A, x)=0 in A X A" near (1,0) will also be diffeomorphic to this cone C near its vertex. By the way, since the set of solutions contains a piece of A x {0}, the cone C must be real, so that ^"(A, 0) cannot be positive or negative definite. The simplest nontrivial case is already very useful;
CH. 2, SEC. 5 FURTHER DEGENERACIES: BIFURCATION 73 THEOREM 2 Take A=R andletf: R x X-^ Y be a function, p^2, with/(2, a:)=0for all x. Assume that (8) (9) 0) is one dimensional y = Im[/;(1,0)] ®f"Ul 0) Ker/Ul, 0) Then A is an isolated bifurcation value. Precisely, there exist £ > 0 and maps x: (—£, e)-^ XandX. {—s, £)->R such that (10) i. A(0)=A, ;c(0)=0, il. x'(0) € Ker f'fX, 0), y(0) f 0 iii. / (A(5), x(i)) = 0 for56(-£, fi) Moreover, there is some neighborhood ^ of (A, 0) in A x X such that whenever (A, x)&^ solves f(X, x)=0, either x=0 or A=A(j) and x=x(i)for some s. A Proof We use the Liapounov-Schmidt procedure with d= 1 =r. We claim that the function g(f Q=nf(X, (^+a:o(A, ^)) has a nondegenerate critical point at (A, 0). Indeed, computing the second derivatives yields the following 2x2 matrix 9 ,"a "A(^,oy nfUlOl The upper diagonal terjn is zero, since/(A^0)=0 identically. By assumption, there is some 6 Ker/i(A, 0) such that nfx^iX, 0)jci ^^0, so that the nondiagonal terms are nonzero. It follows that the determinant is strictly negative. By the Morse lemma, there is a C'’“^ diffeomorphism A=A(s, t) and i=^{s, t), defined for |i| <e and |f| <£, such that 0(A(i, f), ^{s, t))=s^-t^ It follows that the set ^=0 decomposes into s = t and s=—t inj(.y, r) coordin¬ ates. In coordinates, this gives two curves, 5-s), (^(^, ^)) and s-^ (A(5, —s), (^(^, — ^)). We know that one of these curves must be the straightjine (^=0. So the other one yields the nontrivial soliUions of (^)=0 near (I, 0). Note that both curves intersect transversally at (X, 0). Taking the image of the nontrivial curve by the map i+ we obtain a curve ^->(2(^), x(i)), with (.1(0), a:(0)) = (2, 0) and x'(0):^0, which
74 CH. 2, SEC. 5 SMOOTH analysis describes all nontrivial solutions of /(A, x)=0 in some neighborhood ofthe origin. Everything is now proved, except the fact that x'(0) belongs to Ker/i(A, 0). We simply differentiate the identity/(A(i), x(j))=0 at 5=0 0=/'(A, 0)A'(0)+/i(A, 0)x'(0)=/;(A, 0)x'(0) ■ A few examples might be useful. Set X=R^ = Y and consider the map y=f(X, x) given by yi=Axi+X2 y2=Xx2-x\ We can easily check that X2yi — Xiy2=xi+X2 for all A, so f(x, A)=0 has x=0 as its only solution. So A=0 is not a bifurcation value, and yet f'JO, 0) is singular. The point is that it vanishes completely, so that Im[/*(0, 0)] has co¬ dimension 2, and theorem 1 is not applicable. ■ As another example, set X=R"=Y, and let g: R"- with gf(0)=0. Define f: Rx R”-^R” by /(A, x)=g(x)-Xx ►7?" be a given function. We have /*(A, 0)=^'(0)—A/, which will degenerate whenever A is an eigen¬ value of g'{0). Let A be such an eigenvalue: Is it a bifurcation value? Theorem 1 gives us sufficient conditions. Taking into account /L=L we see that /i(A, 0) should have corank 1 and its kernel should not be contained in its range: (g'{0)—^lfx=0 if and only if (gf'(0)—^?)3f=0. In other words, A should be a simple eigenvalue. By theorem 1, any simple eigenvalue of g'{0) is a bifurcation value off (A, x)=0. ■ These bifurcation problems take on a special flavor if we consider the differ¬ ential equation (11) dx Jt =/(A, x) associated with a map f'Rx R"^R". If /(A, 0)=0 for all A, the origin is an equilibrium point, that is, the equation has the constant solution x{t)=0, all i € In a physical situation, such equilibria can be observed only if they are asymptotically stable. This means that for any neighborhood ^ of the origin, a smaller neighborhood 'f' can be found such that for all starting points xq in 'f', the solution of the initial-value problem dxldt=f{X, ,x), x(0)=xo, stays in ^ for all i>0 and converges to zero when i-^oo. It is known that this will certainly be the case if the eigenvalues off'x(X, x) have
CH. 2, SEC. 5 FURTHER DEGENERACIES! BIFURCATION 75 negative real part. On the other hand, it will certainly not be the case if one of the eigenvalues has positive real part. The case when some eigenvalues have zero real parts, all othersJ?eing negative, requires further analysis. In the case when A is a bifurcation value for/(A, x)=0, new equilibrium points appear for values of A close to A. It is generally the case that the trivial equilibrium x=0 loses its stability property when crossing the bifurcation value: If it was stable for A<A it becomes unstable for A> A; and if it was unstable, it becomes stable. It is our purpose now to explain this law of experience, within the limited framework of theorem 1. In view of differential equation (11), we shall limit ourselves to the case where X=B"=Y Say A is a bifurcation value for/(A, x)=0, so that/i(A, 0) has zero as an eigenvalue. We wish to know what becomes of this eigenvalue for A^A: Assuming the (n-l) other eigenvalues to have negative real parts it is the sign of this remaining one that will determine the stability of the systenu The first question to ask is whether the eigenvalue zero for /x(A, 0) can be continued into an eigenvalue /i(A) for/i(A, 0). LEMMA 3 Let A : be a linear operator, having Po as a simple eigenvalue and Xo as an associated eigenvector. Then there is a neighborhood ^ of A in i?"), and C® maps p: ^-^R andx: such that (12) i. p{A)=po and p(B) is a simple eigenvalue of В ii. x(A)=Xo and x(B) is an eigenvector for p{B) The map p is unique, and so is x if ш add the condition that x{B)-xq belongs to a fixed linear subspace complementing Ker (A —PqI) in A Proof Let = Ker (A-PqI)®Xi.^q want to solve the equation {B-pI)[xo + Xi)=0 with peR and Xy e X^. To do this, we apply the implicit function theorem to the map {p, Xx)^{B-pi)(xo-^Xi) depending smoothly on the parameter В € .5^(Л”, /?”). The derivative of this map at {po, 0) for the value A of the param¬ eter {p,Xi)\^-pxo + (A-pol)xi This is a linear map from RxXi into R\ Since po is a simple eigenvalue, R^ splits as the direct sum of Ker (A-Pol) and Im (A-poI). It follows that the preceding linear map is invertible. By the implicit function theorem, there will be neighborhoods tT of ^ in Se{R\ i?”) and tT of (po, 0) in x X\ and a mapping (p, x): such that all solutions in 'f' xiT of the equation (B-pl){xo-i-Xi)=0 can be written Sisp=p(B), xi =Xi{B). Thus p{B) is an eigen¬
76 CH. 2, SEC. 5 SMOOTH analysis value for B and ato+a:i(5) an associated eigenvector. Note that fi(A)=0 and xi(>l)=0. Since fio is a simple eigenvalue of A, we have {A — fiofjXi ^0 for all Xi e Xi, in particular for ||xi|| = l. By compactness, it follows that {B—n{B)I)xi^d for all 6 A'l with llxill = 1 and for all B in some smaller neighborhood of A. It follows that n(B) must be a simple eigenvalue of B. ■ In the situation in theorem l^^we know that the set of solutions off{k, a:)=0 in some neighborhood of (A, 0) consists of two curves intersecting non- tangentially at (A, 0), namely, (—e, s) 9iH^(i+A, 0) (-e, e) 95H^(A(i), 4s)) We know that/x(A, 0)=0. Assume that zero is a simple eigenvalue for/i(A, 0), and apply the preceding lemma. We first choose Xo=(i/x/</s)(0) and take A'l <=/?" so that; Q^xo 6 Ker/;(A, 0) and i?"=Ker/;(A, O)0Ari We then get C'’" ^ maps x?) and (/r*, x{) from (—6, e) into RxXi such that (13) i. (/',(s + A, 0) - !x\smxo+x?(s))=0 ii. (/'x(A(4 x{s))-(iHs)I)ixo+x\is))=0 We have /i®(0)=0=/i^(0), and we wish to know something about /t°(s) and jU^(s) for s^^O, at least their sign. Fortunately, we can do so without actually computing the eigenvalues of/'x(A, x)—an enormous task—because the deriva¬ tives {dpL°lds)(fS) and {dii^lds)i^) are related to each other and to dXjdsif)). PROPOSITION 4 Assumptions as in theorem 1. Assume moreover X=R'‘ and the eigenvalue zero for /'x(A, 0) is simple. Then (14) ^(0)i0 and ~{0)+^{0)^{0)=0 ds ds ds More generally, there is some rj>0 such that whenever |^| < r\, we have {d?^!ds)(s) =0 if and only if fi^(s)=0. Along any sequence such that [dXlds)[s,) never vanishes, we have (15) Sn I ds ds
CH. 2, SEC. 5 FURTHER DEGENERACIES: BIFURCATION 77 Proof. Let us first work on the branch of trivial solutions. Differentiating equation (13), i. with respect to s, we obtain (16) f'UK 0)^0 -/'x(A, 0) ^ (0)=-^ {0)xo as as Now we cannot have zero on the right-hand side, otherwisefxx(^f 0)xo would fall within the range of/'*(2, 0) in R", in contradiction to the assumptions of theorem 1. The first relation is proved. Consider the branch of nontrivial solutions. We start by differentiating the identity /(2(i), x(5))=0 and obtain: (17) AiHs), x(j)) ^ (5)-f/'x(2(5), x(j)) (j) = 0 For i=0, both terms vanish, as we have seen in the proof of theorem 1. This induces us to look at the higher order terms. Differentiating the preceding identity at 5=0, and remembering that (dx/ds){0)=xo, we have (18) 2f’UK 0) ( ^ (0), XoV/i^(^. 0)(^o, Xo)+f'x{l 0) ^ (0)=0 ds^ We have another identity to differentiate, namely (13, ii), which gives (19) /'i,(I,0)(^^(0),xo^-h/".(X,0)(xo,Xo)- ^(0)xo-h/;(I,0)^(0)=0 Substracting the second equation from the first yields (20) f'Ul 0) (0), (0)xoH-/',(X, 0) (0)- ^ (0)^=0 We now recall equation (16). Multiplying it by {d^/ds){0\ and substracting from the preceding one to eliminate/L(^, 0), we have /A\ //Ay (21)/',(2, 0) - (0)^ (0) + ^(0)- ^ (0) + (0) + - (0)(0) Xo=0 ds ds ds ds dfJi^ d^ ds ds ds Now, since zero is a simple eigenvalue off'x{K 0) and the vector xq belongs to its kernel, it cannot belong to its range. It follows that each of the two terms in the foregoing sum must be zero; hence, the second equation.
78 CH. 2, SEC. 5 SMOOTH analysis To treat the case when {dn^/ds){0)=0, we begin by writing for |j| small enough (22) A:(j)=iA:o+>’i(5). with yi(i)eA"i and ^(0)=0 ds Indeed, we already know that (dx/ds){0)=xo, which means that we can write dd> x(i)=<^(s)xo+yi(i). with — (0)=1 and ;^i(s)6A'i Further identifications give <f>{0)=0 and yi(0)=0=(<a[yi/(/s)(0). Setting t=<l>(s), and applying the inverse function theorem to </►: R^R, the preceding identity becomes jc(f)=ixo+>'i ° for |f| small enough Renaming the independent variable, we obtain equation (22). We start from the identity (23) /',(2(4 x(4 ^ (i)+/',(A(4 (i)-;ci(i))+/iHi)(^o+;ci(5))=0 We write the Taylor expansions in s /24X i *• /'a(%). ^(■s))=iCL(^, 0)xo+4fi(i) [H. /;(A(4x(4=/'*(A,0)+ir2(i) Here, ri and r2 are continuous functions of s. Substituting into the preceding identity, we have (25) ^ 0)xo+liHs)xo(i)f i(i)^o+m‘(j)xi(j) + 'SI-2(.y) (i)-xi(i)^= 0) (5)_;c}(5)^ Now,_(i/yi/</5Xi)-x{(i) belongs to AT, for small ls|. Since Xj complements Ker /';,(2, 0) in R”, the restriction to A:, of the operator/';,(2, 0) has a left- inverse. It follows that there are constants ki &ndk2 suchthatfor |i| small enough ^^1 l/^i(5)l + dX + ■^^2 + dX^^
CH. 2, SEC. 5 FURTHER DEGENERACIES: BIFURCATION 79 We subsume everything under a single constant ks (26) 1 + dX ds is) We now recall, as in the preceding case, that f'Ul G)xo=f'Al 0) (0)+-£r (0)xo Since zero is a simple eigenvalue off'^iX, 0), its kernel and its range are com¬ plementary subspaces of R”, so that we may take Xi =Im/*(A, 0). With this choice of Xu relationship (25) projects onto Ker f'x{X, 0) as ^ (•y) ^ (0)+/iHi)=i ^ (j)ri(i)xo + i-2(j) On the right, we have wriUen the first component of the bracketed term, once R" has been split as Ker f'x(X, 0)-|-^i and Xo has been chosen as a base for the one-dimensional subspace Kerf'x(X, 0). Using estimate (26), this yields dX , ^ dfi° ,, , j — (5) ^ (0)-|-/il(5) ds ds 5 ^ lkl('^)-^oll +*3lk2(-s)ll ( l/il(5)l + Since ri and /*2 are continuous at the origin, we can find some constant such that for |^1 small enough (27) dX ' ds is) Everything follows from this inequality and the fact that (dfJ,^/ds)(0)^0. If //H5)=0 and (dX/ds)(s)=/=0, we have ds (0):^/:4|S| and hence \s\^k^^ ^iO) Similarly, we see that whenever |i|<A:4*, we cannot have {dX/ds){s)=^0 and /r*(5)^0.
80 CH. 2, SEC. 5 SMOOTH analysis Finally, if there were some sequence s„^0 such that s„(d^/ds)(s„){dfi°/ds)(0) and n^{s„) have the same sign, we would have Sn ^ (5„) ^ (0) + /r‘(i«) ds ds dX ds du^ + lA<‘(5n)l dfi° ,:r 1( ~T (‘5’n) ds + lAi‘(5,.) in clear contradiction to inequality (27). This proves that s„{dXlds)(,s,)(dii^ld)'.)(X) and have opposite signs, and it is easily seen that their quotient must converge to—1. ■ We now look at the practical implications of proposition 3 for equation (11). The first conclusion, {dfi°/ds)(0)^0, together with /i°(0)=0, tells us that the trivial equilibrium a:=0 cannot remain stable jor remain unstable) when X crosses its bifurcation value A (remember 2=s+A). Assume, for instance, it was asymptotically stable for A<A. Then ix°{X) must have been negative for A<A and becomes positive for A>A, so that x=0 becomes an unstable equilibrium. Let us turn our attention to the branch of nontrivial equilibria 5->(A(5), x(s)). If (</A/ife)(0)>0, relationship (14) shows that (dn^/ds){0) and (i//£°/ifc)(0) have opposite signs. Recalling that and coalesce at zero for A=A, we see that stability is transferred from the trivial to the nontrivial equilibrium when A crosses its_ bifurcation value. If, for instance, x=0 was asymptotically stable when A<A, then a:(A) was not unstable but will become so for A>A, whereas jc=0 becomes unstable. The same conclusions hold if {dX/ds){0)<0, as the following Figure shows. The situation is somewhat more complicated when (dX/ds)(0)=0. We then appeal to the more refined equation (15), which implies that Hi{s) and A(i) are going to have opposite signs near j=0, provided (d^i°ldX){X)>0 (they will have the same sign if (c?/i°/ifA)(A)<0). The outcome is that whenever there are other
CH. 2, SEC. 5 FURTHER DEGENERACIES! BIFURCATION 81 equilibria for a certain value of X beside x=0, then all nontrivial equilibria are unstable if x=0 is stable, and they are all stable if x=0 is unstable. The follow¬ ing figure shows various possibilities. At this point, the reader might think we have just about exhausted stability analysis for the differential equation (28) ±=f(X,y)eR” near the trivial equilibrium x=0. This is not so. What we have done is to in¬ vestigate (partly) the case when an eigenvalue of f'x(X, 0) vanishes. We have proved that in general, that is, except for further degeneracies, when this hap¬ pens, the trivial equilibrium becomes unstable, and there is a bifurcating branch of nontrivial equilibria that stability is transferred to. But there is another way for the trivial equilibrium to become unstable: This can happen whenever the real part of an eigenvalue vanishes (and notjhe eigenvalue itself). We should, therefore, investigate values X such that fx(K 0) has two purely imaginary eigenvalues. This will not prevent f'x(X, 0) from being invertible, so that the implicit function theorem will apply and no new equi¬ librium will appear to relieve the trivial one, x=0, once it has been destabilized. Thus, any physical system represented by equation (11), which is at rest at x=0 for X<X, cannot stay at rest (even at some other equilibriurn) when X crosses the value X. If we expect its behavior to be continuous across X, we must find a stable solution branching off continuously from the equilibrium x=0 when X=X. The answer is simple and due to Hopf: periodic motions. A function x: R IR ” is periodic if there is some number T>0 such that x(t + T)=x{t) for all r; it can then be considered a function on RITZ. The number T>0 is the period. Take a function x: R/Z^R*"; the function y(t)=x(t/T) then is T periodic, and is a solution of equation (11) if and only if x satisfies the equation dx It = Tf(X, x)
82 CH. 2, SEC. 5 SMOOTH analysis Set X=C^(RIZ\ R") and Y = C°{R/Z; R"), both Banach spaces, and A = R^. Define a map 0: A x y by (l>iT,X;x)=^-Tf(^.,x) Equation (11) for y{t)=x(t/T) becomes <f>{T, X; jc)=0 for x, with T ajid X as parameters. Assume the two imaginary eigenvalues ico and — ia> off'x{X, 0) are simple, and let n{X) and /i(A) be the corresponding (by lemma 2) complex eigen¬ values of f'x{X, 0), with n{X)=i(o. The Hopf bifurcation theorem states that if suitable nondegeneracy conditions are met, namely, for no integer k^±l iskio) an eigenvalue of/'*(1,0) then {2nco~^, I) is an isolated bifurcation value for 0. In other words, bifurcation from the trivial equilibria does not occur within the space of equilibria (i.e., constant solutions), but within the larger space of periodic solutions. We refer to Crandall and Rabinowitz for the proof and detailed analysis, contenting ourselves with depicting the two possible evolutions of a stable trivial equi¬ librium for «=2. X<X X>X
CH. 2, SEC. 6 TRANSVERSALITY THEORY 83 6. TRANSVERSALITY THEORY The preceding sections have studied the set of solutions of/(x)=0, or/(2, a:)=0, under various sets of assumptions on the derivatives. Even though we have limited ourselves to the very simplest cases, the variety of situations is already considerable. The complexity increases very quickly if we consider higher degeneracy in the derivatives, and we soon reach a stage where classification is impossible. So the question arises: Which singularities is it really necessary to study? Let us forget about pathological situations and concern ourselves with singul¬ arities that are most likely to occur in the class of problems we are investigating. We owe to René Thom a precise formulation of this idea as well as the mathematical tools to implement it in practical situations. The whole subject is referred to as transversality theory and has many ramifications besides the one we are giving. DEFINITION 1 Let X be a complete metric space and P{x) a statement about points x in X. We say that P{x) is a generic property if the set of points where it holds true contains a dense G à subset of X. k Recall that a subset is defined as the intersection of a countable family of open subsets. In other words, P{x) is generic if 00 {xeX\ P{x) is true} = Q„ 11=1 with each an open and dense subset of X. Let Pn{x\ ne f^, be a countable family of generic properties. Then the property Л Pnix) = Pi{x) and P2W and и = 1 and P„{x) and is generic also. This follows immediately from Baire’s theorem and would fail miserably if we had defined a generic property to hold on a dense subset of X (the inter¬ section of two dense subsets can be empty). In an earlier section, we defined transversality of a map/: to a sub¬ manifold MdR”'. Let us recall definition 3.4, while extending it to functions defined on Banach manifolds. DEFINITION 2 Let N be a C submanifold of some Banach space and f: a C map. Let
84 CH. 2, SEC. 6 SMOOTH analysis X be a point in N. We say that f is transversal to a submanifold M of U”' at x if (1) [either f(x) does not belong to M jor/(x) belongs to M and R'" = T^fT^N + 7>(*)M We now turn our attention to families of smooth maps depending on a parameter u, which belongs to a separable Banach space V. To be precise, let Ü be an open subset of R" and /; Fxfi-^R'’ be a C map, 1. Let M be a C® submanifold of R*" with codimension q. We look for the following property of the parameter value u P(m)={/(«, •) is transversal to M} We state Thom’s famous transversality theorem, whose proof is deferred to Section 7. THEOREM 3 Assume r>max(l, «—^+1) andf : V xQ-^R*" is transversal to M. Then P(u) is a generic property on V k The transversality of/ : K x ii->R*” to Mmeans that whenever/(m, x)e M, any vector j 6 R'’ can be written (maybe in several ways) as (2) У =/!<(«> x)u +f'x{u, x)x+z, with z 6 TyM In contrast, property Р{й) means that whenever/(ii, x)=y e M, any vector j 6 R*" can be expressed as y=f'^{ti, x)x+z, with z 6 TyM without contribution from variations in the parameter. The fact that V may be chosen infinite dimensional allows us to pick the function itself as parameter, which can be done as follows. Define C«,(D; R") to be the space of all C functions/: Ü-+R'’ such that/ and all its partial deriva¬ tives up to order r are uniformly bounded over Q. It is a Banach space for the natural norm. Set V = C„{Q.', R*") and define a map Ф: V xR"->R'’ (called the evaluation map) by Ф(/ x)=f(x) LEMMA 4 _ The evaluation map is C, and Ф>(/ x) is surjective for all {f, x). A
CH. 2, SEC. 6 TRANSVERSALITY THEORY 85 Proof. Note that <I) is continuous and linear with respect to f and C with respect to x. Jhe fact that it is C in (/, x) follows immediately. We have, by linearity, <D}(/, x)f =f(x\ and the map/-^f{x) from C« to is obviously onto. So is certainly going to be transversal to any submanifold M of As an easy application, we give corollary 5. COROLLARY 5 Let M be a submanifold of with codimension q. Then the property (3) P{f)=[f : is transversal to M) is generic in C«(Q; (R^), for all r^max(l, « +1 -^). A Just apply Thom’s theorem to <I) : K x R^. We might want to have generic- ity in (7(0; R^) instead of (7«,(0; R^); in that case, use lemma 6. LEMMA 6 Let 0 be an open subset ofW^ and O^, k eN, a sequence of open subsets such that (4) i. Q= Q Qj fc=l ii* VÆ, Çljç c c ^ j Assume a property P(/) is generic in C*(Qt; R'’) for every k. Then it is generic in C (ii; R"). A Proof. Note that/ In ^ belongs to (7„(Qj,; R'’) wheneverf belongs to C(Q; R More generally, iff belongs to (7„(0j-; R"), then/|q^ belongs to (7„(i2n; R'’) for all Call (f>k the map/-^f\a^ from (7(Q; R**) to R'’), and the map from R'’) to (7„(Qfc; R'’). Fork^j, the diagram C{Sl; R'’) is commutative <t>k=4>jk° 4>i The topology of (7(Q; RP) is defined to be the weakest topology that makes
86 CH. 2, SEC. 6 SMOOTH ANALYSIS all maps continuous. It does not make C(Q; a Banach space, although it does make it a complete metric space, so that Baire’s theorem holds. Now let Gfc be the subset of C«,(£ifc; R^) where property P{J) holds. By assump¬ tion, we have n n = l where each {/*,„ is an open dense subset of M'’). Each 0k” then is an open subset of (7(Q; R'’), and n n <l>u\VKn) Jt=l n=l is a countable intersection of open subsets of (7(i2; R'’); that is, a Gi subset. We claim that each 0k” Hi/k,«) is dense in C(n; R'’). To see this, pick any open subset U of Cifl; R'’). By definition, there is some i and some open subset Ui of CTJQ,-; R") such that U=<!>r\Ui). If i=k, Uiri Uk,„^0 since Uk,„ is dense in (7«,(Q(; R*"), and so U meets (¡>k^{Uk,„). If i<k, we write U=^k\(l>ki^{Ui)), and ^ki^(Ui), being an open subset of (7„,{£2k; R'’)» has to meet 14,m so that again Un4>k\Uk„)^0 If i >k, pick some/ € Ui and some 6 > 0 so small that ||/ - 0|| < e in ; R'’) will imply that g e U. Now consider /|nk=0ik(/)- Since Uk,„ is dense in C^(Qk: R*”), there must be some 0 6 C/k,„ such that ||0,k(/)-0|| <e/2inC|o(Qk; f*'’)- We now extend ¿f to a map g € C„(iii; R'’) in such a way that ||/—0|| <£. This will always be possible provided the open sets Qk are sufficiently regular —if, for instance, they are all finite unions of balls, which can always be managed. We then have geUi and (l>tk(g)=g e Uk,„, so that Uin<t>a^{Uk,n)^0, and so 0rHUi)n<l>rH0fk'(Uk.„)) =Un4>k ^iUk,„) =h0 So each (f>k^{Uk,n) is open and dense, and by Baire’s theorem, G is a dense Gs subset. ■ Corollary 5 then leads to Corollary 7. COROLLARY? Let M be a C°" submanifold of with codimension q. Then the property
CH. 2, SEC. 6 TRANSVERSALITY THEORY 87 (5) P(f) = [f : is transversal to M] is generic in C'(Q; IR^)for ail r^max(l, n-q-\-l). ^ It follows, for instance, that if any particular map we are dealing with is not transversal to M, it can be made so by an arbitrarily small C perturbation. For further applications, we use refinements of the evaluation map O of lemma 4. The results in propositions 8 and 9 are typical. PROPOSITION 8 The property (6) p(/)={/ : is a Morse function) is generic in C(Q; R) for all r>2. A Proof By lemma 6, it will be enough to prove that P{f) is generic in Set V = (7,^(0; R) and consider the map ^ :V x Q-^R" defined by (7) This map is linear inf, and C' ^ in jc, so that it is jointly C' We have '¥'0,x)f=f'{x)eU’' This shows that x): is surjective, so '¥ will be transversal to any submanifold M of R". Take Af={0}, a submanifold of dimension zero. Apply Thom’s theorem to »P and M. It follows that the property (8) P'(/)={«/'(/ •) is transversal to M} is generic in C^(ii; R). This means that whenever 'Vif, x) € M, the tangent map 'P* is surjective. In other words, whenever /'(x)=0, the matrix/"(x) is non¬ degenerate. Therefore, properties P{f) and fif) are the same, and the result is proved. ii Very often, we can guess the correct result by simple counting arguments. For instance, iff is not a Morse function, the system f'{x)=0 and Det/ (x)=0 has a solution. But this is a system of (n-l-1) equations for n unknowns (xi,..., x«)
88 CH. 2, SEC. 6 SMOOTH analysis and so should have no solution in general. Such heuristics, although incorrect, can give us foresight. For instance, let us ask whether we can arrange for all critical values to be distinct. The equality of two critical values, corresponding to critical points Xi and X2, requires that the system f(Xl)=f{X2) OXi I»'*'“» has a solution. We see that there are (2«+1) equations for 2n unknowns, so it would seem to be an exceptional situation. Proposition 9 confirms this observa¬ tion. PROPOSITION 9 Let us say that a function f: is excellent if it is a Morse function on Q and all its critical values are distinct. The property (9) P(f) = {/ is excellent] is generic in 07(0; R), all r'^^. A Proof We shall show that the property (10) 6(/)={aH critical values off are distinct} is generic in 07(0; R). By proposition 8, the property of being a Morse function is also generic, and by proposition 2, the conjunction of both, namely, P{f\ will be generic. By Lemma 6, it will be enough to prove that Q(f) is generic in C«(0; R) = L Introduce the space (11) fi = {(xi,X2)ei2xi2|xi^X2} It is an open subset of R Now define a map V xQ->IR^ xR^" (12) 'P(/ Xi, X2) = (/(Xi), fix2\ f'(Xi)> f\X2)) It is linear inf, and (7 in (xi, X2), so that it is globally (7, and 'V'fif, Xu X2)f={f(x,\ f{x2), /'(Xi), f'(X2))
CH. 2, SEC. 6 TRANSVERSALITY THEORY 89 So is surjective, and 'P will be transversal to any submanifold M of xU^*\ Let us choose M={(a, a, 0, 0)|aelR} By Thom’s theorem, the property = •, *) Is transversal to M} is generic. But •, •) Is a map from Q, an open subset of into Saying that this map is transversal to M means that whenever T^(/, xu ^2) eM, any vector in can be expressed as the sum of a vector in ('Pij, 'Px2)(0^" x and a vector in M. But M is one dimensional and ('Pi,» x IR") is at most 2«-dimensional, so they cannot make up IR So 'Pi/, Xu ^i) cannot belong to M, which means precisely that properties Q(f) and Q(f) coincide. ■ We now investigate one-parameter families of smooth functions/(/I, x) with XeU and xelR'*, as in Section 4. Because of the supplementary variable A, such families may contain non-Morse functions in a stable way; that is, all neighboring families will contain non-Morse functions, possibly for different values of the parameter >1. What we shall show is that the one-parameter families we investigated in Section 4 were typical; that is, the assumption we made is, in fact, a generic property. PROPOSITION 10 Recall assumption A of Section 4: There is no point (>l, x) e IR xQ such that the df/dxiiX, jc), l^i^n, 3(2., x) and A(/ x) vanish simultaneously. The property (13) p(f)z=[f satisfies assumption A} is generic in C‘((R x O; U)for all r'^2. k Proof Again, it is sufficient to prove that statement (13) is generic in CJR xQ;IR)=K Now consider the set M of symmetric nxn matrices; it is an n(n-\-l)/2- dimensional vector space, and so isomorphic to with p=n(n+l)/2. Define a map ^P: VxU xQ-^[R"xM ^(f x) = (f'x(K x\ fUK x)) It is linear inf and C " ^ in x, so it is jointly C " ^ and ^f(fJ^,x)f = (f'Alx),fUlx))
90 CH. 2, SEC. 6 SMOOTH analysis so that is surjective. It follows that 'F will be transversal to any submanifold of IR” X M. Consider in M the subset N of all singular matrices. It is not a submanifold, but it is a finite union of disjoint submanifolds, all of which have codimension ^ 1. The component of codimension one consists of all symmetric nxn matrices with rank (« — !); call it Nq. Now set F = {0} X X M. It is a finite union of submanifolds F,j, (« + !), the codimension of Fij being i. We have UjF,j = {0} x No- Applying Thom’s theorem to each of the Fij we see that the property % *) is transversal to all the Fij} IS generic. That'PC/; image •) is transversal to Fij means that whenever T^(/, X,x)=z 6 Fij, the TOl,x),'P;(/,X,3i))(RxR^«) and the tangent space T^Fij span the whole space. Since the former has dimen¬ sion less than or equal to ^« + 1, which implies that the codimension of Fij must be 1. So 'Pi/ •, *)=z misses the Fij altogether when i>« +1. It can only meet {0} x No- So F{f) implies thatfxxiK x) has rank ^ — 1) for all X and x; note here that we cannot prove it always has rank n, which would mean /(1, •) is a Morse function throughout, and that would be false in general. To go farther, we have to use the fact that ^(/, •, •) is transversal to {0} x A^o- The subset of Af is described by the equation Det m=0, and all points of No are regular for the map Det: M^U. Near any point mo e No, there is a local coordinate system (<^i = Det m, ^2, • • •, ip) for M. So {0} x A^o is defined by the equations = • • • =y„=0, =0in R" x M, and the transversality of^(/ •, •) = to {0} X No means that the tangent map dxidxj d Det f' dxi 5y dx:dX gy dx/fdX d Det/" is onto whenever '¥{f, X, x) e No- In other words, A(A, x)/0 whenever/*(A, x)=0 andi(A,x)=0. ■ We should picture the space C^(i2; K) as being partitioned into open cells by a network of submanifolds of codimension 1,2, and higher, similar to a beehive.
CH. 2, SEC. 7 PROOF OF THE TRANSVERSALITY THEOREM AND APPLICATIONS 91 The open cells consist of Morse functions; if a function is non-Morse, it will lie on the boundary between two or more cells, and a slight perturbation will send it into a cell, that is, make it a Morse function. A family/(^ a:) of functions depending on the real parameter A is really a map •) of IR into C'(Q; R), that is, a path in C‘(Q; R). If the points A=0 and A = 1 lie in two different cells, the path will have to cross some boundary between them at some A between zero and one. This situation cannot be altered by slightly changing the path. What can be arranged, though, is for the path to avoid boundaries of codimension 2 or higher, and cross only codimension-1 boun¬ daries. For this reason, unavoidable singularities are called codimension-1 singularities. Thanks to proposition 10, and to the analysis we carried out in Section 4, we can now state that the singularities of codimension 1 in C‘(f2; U) have the form f{x)=f(0)±xl± ■ ■ ■ ±xl-i+xl in an appropriate local coordinate system. 7. PROOF OF THE TRANSVERSALITY THEOREM AND APPLICATIONS In the preceding section, we discussed Thom’s transversality theorem. We now turn to its proof, following the approach of Stephen Smale, who managed to tie up transversality theory with a classical result in nonlinear analysis, the Sard- Brown theorem. The results of this section, which are mostly due to Smale and range from the very theoretical to the very practical, will show how intimate the connection is. From the technical viewpoint, in this last section, we abandon (linear) Banach spaces in favor of (nonlinear) Banach manifolds. We shall define a Banach manifold to be a submanifold of some Banach space (see definition 3.3). Although this is not the classical definition, it is equivalent to it when the dimension is finite and covers all infinite-dimensional situations arising from functional analysis. 1. The theorem of Brown, Sard and Smale Henceforward, we shall consider C maps, 1, from a Banach manifold Af to a Banach manifold Y. We recall that x e AT is a critical point off if the tangent map Txf: TxX-^Tf(x)Y is not surjective. We say that € T is a critical value of/ if it is the image of some critical point. Conversely, a point xe Xis regular if it is not critical, and j; 6 7 is a regular value for the mapping/ iff~^{y) contains no critical point. Note that iff~^{y) is empty, y is considered to be a regular value. When X and Y have finite dimension, the tangent map can be identified with
92 сн. 2, SEC. 7 SMOOTH analysis the Jacobian matrix in some suitable local coordinates. Then л: is a critical point for/ iff (1) rank Txf <dim Y and is a regular value for/ iff (2) /W=;;=>rank Txf = d\m Y THEOREM 1 Assume /: X^Y is a C map, l^r^oo, between finite-dimensional separable manifolds, with (3) r>dim A"—dim У Then the set of critical values for f is negligible, A By a negligible subset of У we mean a subset N that is negligible in all local coordinate systems. To be precise, whenever 0 is a local chart of some open subset U of Y onto an open subset of when ф{11 r\N) has Lebesgue measure zero in This theorem was first proved by Brown for the case r = oo and then by Sard for finite r, the proof being significantly more difficult. Condition (3) is automatically satisfied if r = oo or dim y>dim A"—1. Condition (3) requires that the dimension of the target not be too low, namely, dim У > dim X—r. Counterexamples are known where this requirement is not heeded. The Sard-Brown theorem can actually replace Thom’s transversality theorem in some simple situations, as shown in corollary 2. COROLLARY 2 Let H be a linear subspace of with codimension k, let X an n-dimensional manifold, andf : i/->R'^ a C map, r ^ 1. Associate with any a eU^ the map fa'. xh*f{x)-a Assume r>n-k. Then, for almost all a eU^, the map fa is transversal to H. A Proof Let be a complementary subspace to H, so that U^ = F@H, and 7c: R^^Fthe associated projection. Saying that/, is transversal to H means precisely that zero is a regular value of n ^fa, that is, n{a) is a regular value of 7c°/ Let Rc:FbQ the set of regular values for n^f By theorem 1, F\R has measure zero, and тг" is the set of a g R^ such thatfa is transversal to H. ■ The general situation, where/ depends on the parameters in a more complic¬ ated way, will not yield to this simple-minded approach. This situation moti-
CH. 2, SEC. 7 PROOF OF THE TRANSVERSALITY THEOREM AND APPLICATIONS 93 vated Smale to find an infinite-dimensional version of theorem 1 that could be used directly in the function space. This version requires a few definitions first. Let X and Y be Banach spaces and L: X-^ Y a continuous linear map. We say that L is a Fredholm operator if both (4) i. KerL = L ^(0) is finite dimensional. ii. Coker L = YlL{X) is finite dimensional. The index of L is the number (5) dim(Ker L)—codim(L(A")) e Z It follows from condition (4) that L{X), the range of L, is closed. If L is a Fredholm operator with index i and is a compact operator, then L-h К is a Fredholm operator of index i. For instance, all operators A-\- K, where A is an isomorphism and К is compact, are Fredholm with index zero. Now let X and У be Banach manifolds and/: X-^ У a map. We shall say that/ is a Fredholm map if for all xe X, the linear map Txf: ТхХ-^Тдх)У is a Fredholm operator. If X is connected (for instance, if A" is a Banach space), the index of Txf will not depend on the particular choice of the point л: in A" and is referred to as the index of f Note that if X and У are finite dimensional, then any map /: A^-> У is Fredholm with index dim X — dim У THEOREM 3 (SMALE) Letf: X^Ybe a C Fredholm map between separable Banach manifolds Xand Y, with l^r^oo. Assume Y is complete, X connected, and (6) r> index (/) Then the set of regular values for f contains a dense Gs subset ofY A In the finite-dimensional case, theorem 3 reduces to the original Sard- Brown theorem, with one major difference: The property ;; is a regular value for / is stated to be generic instead of true almost everywhere. These two statements are not equivalent: We can easily find a dense Gs subset of [0, 1] that is neg¬ ligible (so its complement has full measure and cannot contain a dense G¿).t However, the genericity statement is the only one to make sense in infinite¬ dimensional spaces, where there is no analogue to the Lebesgue measure. JLet be the set of all rationals in [0,1]. For any £>0, define to be the union of all openw intervals (p„ - c2 “p„ + £2 ""), so that meas( Ue)^ 2e. Then n ^> o t/e is a dense Gs subset with measure zero.
94 CH. 2, SEC. 7 SMOOTH analysis We prove Smale’s theorem in two steps. LEMMA 4 The set of critical points for f is closed in X. Proof We shall prove that its complement, the set of regular points, is open. Recall that a point x e X is regular if T^f: TxX^Tff^xiY is surjective. Set K= Ker Txf which is finite dimensional, since 7i/ is a Fredholm operator, and let n: TxX-^K be a continuous projection. By the open mapping theorem, the map f'(x)0 from TxX to Kx Tf^x)Y is an isomorphism; call it i. Set y=f{x) to simplify notations. Let p: KxTyY-^TyY be the projection. With any map u: TxX-^ KxTyY we associate p^u. In this way we define a continuous linear map L from ^{TxXy KX TyY) into SP(TxXy TyY). It is clearly surjective, so it is open, using the open mapping theorem again. We have just seen that Txf =p°U with i e the set of isomorphisms in SP{TxX, KxTyY). So Txf eL{^). Now L(®) is open since ® is open, and consists only of surjective maps (they can all be written sisp^Uy with u an isomorphism). Hence, the result. ■ LEMMA 5 The restriction of f to a suitably small neighborhood of any point maps closed subsets onto closed subsets. k Proof Take any point xsX. Set AT=Ker Txf and R = Txf X. Let n: TxX-^K and p: TyY-^R be continuous projections. By the open mapping theorem, the map f'{x)^) is an isomorphism of TxX onto KxR. This is the tangent map to the nonlinear mapping p °/((^)) at =x By the inverse function theorem, we can use it as a local coordinate system for Xnear ^=x. In these new coordinates, the mapp°f now reads and the map/ itself /: (<^i,. • •, U >?),•••, >?)> f?) Here dim K and /?=codim R, so that «-/? = index(/). Now let [/ be a closed bounded neighborhood of x where this local chart is valid and x^ a sequence in U such that f{x^) converges to some y eY Reading off the coordinates x^={^\ rj^), we see that p °/(<^^ rj^)=rj^ converges to p{y). As for the they stay within n{U)y a bounded set in the finite-dimen-
CH. 2, SEC. 7 PROOF OF THE TRANSVERSALITY THEOREM AND APPLICATIONS 95 sional space K, so a suitable subsequence will converge. The corresponding subsequence of converges also. ■ The proof of Smale’s theorem is as follows Choose local charts as in lemma 5 and call (j> the map with components ..., </>p). Then x=(^,rj)e U will be a critical point of/ if and only if Os a critical point of (j)(% rj\ a map from IR^. And y ef{U) will be a critical value if and only if y—p(y) is a critical value of (/>(•, rj), where rj=p{y). By Theorem 1 for every rj sR, the set of critical values of (/>(•, rj) has measure zero. It follows that it has empty interior. In other words, for each rj eR, the set oiy ep~^{rj) that are critical values of/ over U has empty interior. Then the set Cu of critical values of/ over U has empty interior. By lemmas 4 and 5, the set Cu is also closed. We now find a countable family of U covering the whole space X. The set of critical values for/ over X is the union of the Cu, and its complement is a dense Gs by Baire’s theorem, since Y is a complete metric space. 2. Proof of the Transversality Theorem We shall restate the transversality theorem in a nonlinear framework. Now Q is a connected separable «-dimensional manifold (instead of an open subset of (R”), Y a separable /7-dimensional manifold (instead of IR'"), and V a separable Banach space. THEOREM 6 Let f'. V X Y be a C map, 1 oo, and M a submanifold of Y, Assume that f is transversal to M and that (7) r > dim Q - codim M We denote by fu the map co-^f(u, co). Then the property (8) P[u) = {fu:ii-^ Y is transversal to M} is generic in V. ^ Let us introduce the set E=f~\M) and the map tc: V, the restriction to E of the first projection (w, co)-^w. Since/ is transversal to M, the set £" is a closed submanifold of F x Q, and the map n is C. The geometric situation is shown in the following figure.
96 CH. 2, SEC. 7 SMOOTH analysis \uu W2, M3} is the apparent contour of E. It is the set of critical values for n. We shall denote by x=(w, (o) the points of E. LEMMA 7 u is a regular value for n if and only iffu is transversal to M on Q. Proof Ify;,(n) does not meet M, then/„ is transversal to M, and there is no point a: in E:=f~^(M) with n{x)=u, so w is a regular value. Now assumefjfl) n M ^0, and let x=(w, co) e E. We claim that x is a regular point for n if and only /„ is transversal to M at jc. The result will then follow. Set M=0 for the sake of convenience. Assume first that/0 is transversal to M at x. Proceeding as in theorem 3.5, we can find local coordinates ..., 1/^^) in Y around j;=/(0, x), such that M is locally defined by k equations iAi(>^)= * * ‘ =^kiy)=^i and corresponding local coordinates (<^i,..., lAi °/o, • • • j lAk °/0) in Q around co. By the implicit function theorem, the equations iAi°/(w, co)= • • • = °/(w, co)=0 can be solved near (0, a>) in terms of (¡>1 and u. In other words, the </>,, l^i^n—k, and 71: E-^V constitute a local coordinate system for Enear x, which implies x is a regular point for ti. Conversely, assume that x=(0, co) is a regular point for n in E. This means that the map T^n: T^E-^T^V is onto. Now T^E is closed linear subspace of TqV X Tafl, and Tx7i is the restriction to TxE of the first projection. It follows that (9) Tx{VxSi) = TxE-\-T^n Here we have identified Tx(V x Q) with TqV x and Tafl with 7i({0} xii); the sum on the right is not direct, that is, we are not claiming that TxEn ={0}*
CH. 2, SEC. 7 PROOF OF THE TRANSVERSALITY THEOREM AND APPLICATIONS 97 Applying Txf to both sides TxfTxiV X Q)=TxfTxE+ TxfTjl Since/ is transversal to M, we also have TxfTx{Vxa)+TyM=TyY, withj=/(0,co) Replacing the first term by its value from the preceding equation, TxfTxE+ TxfTJ^ + TyM= TyY The first term on the left is contained in the third, TxfTxEaTyM, since f(E)=M. The second term is easily identified with EofoTJ^. Finally, we have TxfoTJ^ + TyM=TyY This means preeisely that/0 is transversal to M at x, as desired. ■ LEMMA 8 n is a Fredholm map, with constant index n—k. A Proof. As we just pointed out, TxE is a closed linear subspace of x Tofl, with codimension k, and is the restriction to TxE of the first projection. We introduce spaces N1, L, K, F, and N2 such that N1 = Ker Txit = TxEn Tefl TxE=Ni®L To F X r„Q = TxE®K (dim К=k) N2=KnTafl K=N2®F Clearly, N1 ®Nz = Tjl. We have ToF X Tjl=L®Ni ®N2®F=L®TJ[1®F So ТхП{ТоУ X Tji)=Txn(L)®Txn{F) It follows that Txn{F), which is isomorphic to F itself, is a supplementary sub¬ space to Txn(L)=Txn{TxE).
98 CH. 2, SEC. 7 SMOOTH analysis codim T*7t(T*£)=dim F=^-dim N2 dim (T*7i)"‘(0)=dim iVi=«-dim N2 So the index of Tin is n —k. ■ The proof of the transversality theorem now follows from theorem 3 applied to n: V, noting that dim Q—dim Y=n—k 3. Newton’s method revisited We conclude this chapter by applying the latest results we obtained to the first problem we started with, namely, the problem of solving the equation (10) f(x)=0 in some domain 5 of K". For the sake of convenience, we shall take B to be the unit ball, with boundary S. The function/ maps a bounded open subset QcR" containing B into R". At this point, we are looking for a priori conditions that will ensure that equation (10) has a solution and for a practical procedure to solve it numerically. Smale popularized a method that achieves both at one stroke. Smale’s method Choose some point XqbS. If /(xo)^O, construct the straight line D=Uf(xo), and the set C=Bnf-\D) k Clearly, C is not empty, since it contains Xo, and C contains any zero off in B, since 0 6 Z). The merit of considering C lies in the following two properties, (11) and (13), which we shall require of/ (11) / is transversal to D and/ '(2)) to S This tells us that/ ^(Z)) is a closed submanifold of Q with codimension (« — 1), that is, a smooth curve. For all x 6 define l(x:) e R by (12) f{x)=X(x)f{xo) By the implicit function theorem, we can use A as a local coordinate for in the neighborhood of any point 3c where/'(3c) is invertible, that is, Det/'(.x)^0. On the other hand, when Det /'(3c)=0, one of the projections in
CH. 2, SEC? PROOF OF THE TRANSVERSALITY THEOREM AND APPLICATIONS 99 IR”, x-^Xi, say, will serve as a local coordinate near x for/“*(£)), since the latter is a one-dimensional submanifold. Our second property (13) follows (all deriva¬ tives taken at Xi=Xi with the preceding definition of i). (13) fix)eD, Det/'(x)=0=>^^0, Det/'(;c)^0 axi axi Differentiating the equation f(x)=Xf(xo) at x, we have (14) Since Det/'(x)=0, the mapf'{x) is not surjective. We know from assumption (11) that/(:^o) and the image of f'{x) span the whole space R", so f(xo) cannot belong to the image of/'(3c). Equation (14) then implies that ^-0 <fe, at 3c Assumption (13) now tells us that dX/dxi and Det f'{x) change signs simul¬ taneously : The critical point off cutsf~^{D) into arcs on which X{x) is monotone, alternatively increasing and decreasing on consecutive arcs. We can now prove existence results constructively. PROPOSITION 9 Assume (11) and (13). Assumef{x)=xfor all x eS. Thenf has a zero in the interior ofB. k Proof The set/ " ^(/)) is a closed submanifold of Q. By assumption (13), the set of critical points of/ that belong to D is discrete: If Det/'(3c)=0 and 3c € f~^{D\ neighboring points of D have all Det f\x)4^Q, By compactness, it follows that there can be only finitely many critical points of/ on/"^(Z))n5. Now investigate the connected components of f~^(D)r\B containing xq. It is an arc contained in B, starting at xq transversally to 5; call it Let Xu... yXf^ be the (finitely many) critical points of f(x) on Delete the portion between xq and from ^ and consider the remaining arc, originating at xiq. On this arc, there are no more critical points of/(x), so X is monotone, and we can use it as a global coordinate. Using the compactness of 5, we see that X is bounded, and since it is monotone, it converges to a limit Xy and the corre¬ sponding point a: of ^ converges to a limit 3c. Since ^ is closed, it must contain 3c. So ^ ends at 3c. If 3c were an interior point of By this would contradict the fact that/" ^(D) is a submanifold. So 3c belongs to S. We cannot have 3c=a:o, or xq would be a self-intersection point of/"^(D), contradicting again the fact that it is a submanifold. So x^xoy and yet f{x)=
100 CH. 2, SEC. 7 SMOOTH analysis ^f(.Xo)\ since X belongs to S, this reduces to x=Xxq. The only possibility is x=—Xo, so A= —1. Since we started with A = 1 at xo, we must have A=0 somewhere on m This proposition can easily be extended in two different directions. COROLLARY 10 Assume (11) and (13). Assume f does not vanish on S, and consider the map f: S-^S defined by (15) /W=/(x)||/(x)|| -1 If there is an odd number d of points x such that f{x)=f{xo), then f has a zero in B. A Proof This contains the preceding result as a particular case:/(x)=/(x:)=A: in S, and xo is the only antecedent off(xo\ so d= 1. The proof is essentially the same. Take Xo and let xi be the endpoint of the corresponding arc, so x^fx^ and/(x:i)=A/(xo). If A<0, we are done, since A must change sign between xo and xi. If A>0, we have/(xi)=/(xo). Since d is odd, and hence fl, we can take another starting point X2, withy(x2)=j^(xo), and get a corresponding endpoint X3, with/(x3)=-^(x2). If A<0, we are done; if A>0, we can continue since XfA. Since d is odd, it will be impossible to appariate all the antecedents of/(xq), so we shall eventually find/(x„)^^x„+i) for some suitable starting point x„. ■ COROLLARY 11 Assume (11) and (13). Assume the scalar product (x, /(x)) does not vanish on S. Then f has a zero in B. A Proof Say (x, fix)) is always positive on S. This enables us to extend / to the ball A=25, with boundary 2=25, in such a way that the extension g: A-»R" is the identity on 2 and never vanishes on A\B. This is done by finding a smooth function [1,2]->[0,1] such that </>(l)=0 and (¡>(2)= 1 and [0W=/(^) iflWKi \д{х)=0(||x||)x+(1 - 4>i\\x\\))f () if 1 < iixll < 2 Clearly, g(x)=x if ||x||=2, and (^(x), x)>0 for l<||x||<2, so ^(x)^0 in that region. We then apply Smale’s method to g on A, starting from a suitable point on 2. The corresponding curve may break when crossing 2 (it remains con¬ tinuous, but may lose differentiability), but the argument in proposition 9 otherwise works to give a point x e A where ^(x)=0. Since g does not vanish outside B, we have xeB and/(x)=^(x)=0. ■
CH. 2, SEC. 7 PROOF OF THE TRANSVERSALITY THEOREM AND APPLICATIONS 101 At this point, the reader may wonder about the role of properties (11) and (13). As a matter of fact, they are generic in C„(Q), all r>2. For (11), it follows immediately from Thom’s transversality theorem. For (13) (once Xq^S is prescribed), it is also true but more intricate. So corollaries (10) and (11) extend by density to all of (7„(i2): Assumptions (11) and (13) can be dropped and so can the differentiability assumption. Cor¬ ollary (11) extends to any continuous/ with {f (x), x) 0 on the boundary, while corollary (10) reqmres an appropriate definition of d, which is now called the degree of the map/: S^S. But at this point, we are more interested in the computational problem, and this is where (11) and (13) are handy. We are supposed to find the curve f{x)= ^f(xo)- We do this by starting at xq and integrating numerically the correspond¬ ing differential equation /w|=/w or (16) dx , ^=/'W-y(xo) which will break down when the curve approaches a critical point 3c, where f (x) is singular. We should then take another coordinate along the curve xu say. Equation (16) is then replaced by (17) y/ X dX which can be solved for dx/dxi and dX!dx\. Once the critical value is crossed, we can resume using equation (16). All this is done automatically if we replace equation (16) by the system (18) ^=Cof/'(x)-/(xo) dX ^ Here, Cof f'(x) is the matrix of cofactors of/'(x:) /'(x)- Cof /'(x)=Det f\x)‘ I Cof/ (x)*/(xo) never vanishes along the curveotherwise/would not be transversal to Z). So the trajectory of the first equation is the whole of/ " ‘ (Z)).
102 CH. 2, SEC. 7 SMOOTH analysis The second equation reminds us that the points where Det /'(x)=0 separate f~\D) into arcs where X is monotone, alternatively increasing and decreasing. Finally, note that with/(jc)=2/'(xo), the original equation (16) can be written dx This is a continuous version of the discrete algorithm (18) of Chapter 2 for solving/(A:)=0 by Newton’s method Xn+l-X„=f'{x„)-^f{x„) which is why Smale also refers to it as Newton’s method. We shall later give another set of assumptions ensuring both the existence of a zero for f and the convergence of Newton’s method in a more general setting.
CHAPTER 3 Set- Valued Maps When X and Y are Hausdorff topological spaces, we face the problem of de¬ fining continuity of set-valued maps. In the case of single-valued maps/ from A' to continuous functions are characterized by two equivalent properties (a) For any neighborhood X(f{xo)) of /(xo), there exists a neighborhood J^(xo) of Xo such that/(J^(xo))= .A^(/(xo)). (b) For any generalized sequence of elements x„ converging to xo, the se¬ quence /(X;,) converges to/(xo). These two properties can be adapted to the case of strict set-valued maps from X to Y; they become (A) For any neighborhood A^(Fl(xo)) of F(xo), there exists a neighborhood Ji{xo) of Xo such that Fl(A^(xo))«= A^(F(xo)). (B) For any generalized sequence of elements converging to xq and for any there exists a sequence of elements 3;;, e F{x^) that converges to yo. In the case of set-valued maps, these two properties are no longer equivalent. We call upper semicontinuous maps those that satisfy property (A), lower semi- continuous maps those that satisfy property B, and continuous maps the ones that satisfy both properties (A) and (B). In the first section of Chapter 3 we review elementary properties of semi¬ continuous maps and present a list of examples. 1. If IF is a function from A" x y to i?, we set IFc:3;->IFc(;^):=IF(x,>^) and study the upper semicontinuity of the map epigraph (IFJ 2. Iff maps X X C/ to X we give sufficient conditions for the set-valued map F defined by Hxy.= {f(x,u)\uB U] to be upper or lower semicontinuous. 103
104 СН. 3, SET-VALUED MAPS 3. If/ is a map from KxY to Z, U and T are set-valued maps from К to Yand Z, respectively, and we prove that the set-valued map C defined by C{x):={y 6 U{x)\f{x, y) 6 T{x)} is lower semicontinuous under a convenient set of assumptions. 4. We consider a map T sending elements л: e A" to closed convex cones T{x) of We prove that T is lower semicontinuous if and only if the graph of the map x-^T{x)~ is closed. 5. Finally, we investigate the continuity properties of the marginal function V of a family of maximization problems V{y):= sup W{x,y) xeG(y) depending on a parameter as well as the upper semicontinuity of the marginal map associating with the parameter у the set of maximizers M{y):={x e (7(y)|F();)= W{x, In Section 2, we single out an important class of set-valued maps, namely, maps with closed convex values. Such maps Ffrom A" to a Banach space У can be characterized by their support functions a{F{x),p):= sup {p,y) yeF{x) since the Hahn-Banach separation theorem tells us that F{x) = {ye Y\ip e У^ (/7, y) ^ g{F{x\ p)} These support functions are very easy to manipulate. For instance, we shall observe that the upper semicontinuity of F implies the upper semicontinuity of the functions x^a{F{x\p) when p ranges over У^. So, we shall select maps en¬ joying the latter property, which we call upper hemicontinuous maps. A theorem due to Castaing states that any upper hemicontinuous map with compact convex values is, conversely, upper semicontinuous. Upper hemicontinuous maps with closed convex values enjoy many fixed point and surjectivity properties, which are presented in Chapter 6, Section 4. We devote the third section to studying maps with convex graphs as well as maps whose graphs are cones (called processes). Convex processes whose graphs are convex cones are the set-valued analogues of linear operators and share some of their properties. Convex processes will be used for defining derivatives of set-valued maps (see Chapter 4, Section 2 and Chapter 7, Section 7). They also enjoy spectral properties (eigenvalues and eigenvectors), which will be studied in Section 4.
CH. 3, SET-VALUED MAPS 105 When /4 is a (continuous) linear operator and P<=^X and Q^Y are (closed) convex sets, the set-valued map F defined by m-= Ax—Q whenxeP 0 when x^P provides an example of a (closed) convex map. The inverse of this map is defined by ^yeY, F ^(y) = {x e P\Ax eQ-\-y) These subsets F~^(y) provide the main class of subsets on which we minimize or maximize functions in optimization theory. When P and Q are (closed) convex cones, the map F is a (closed) convex process. We can adapt to closed convex maps the Banach open mapping and closed graph theorems, and we shall prove that such maps are lower semicontinuous on the interior of their domain. We have an even stronger result: If Xq e Int Dom Fand;^o ^ ^->^o) are chosen, there exists y > 0 such that "ixexo + yB, SyelmF d{y, i^.x))^- d(Xy F"^(y))(l + lb-;^oll) y This theorem implies that any closed convex process from a Banach space X to a Banach space Y whose domain is the whole space X is Lipschitz: There exists y > 0 such that Vxi, X2 G X, F(X2)^F{Xi)-\--\\Xi-X2\\ y We also deduce that if F is closed, convex, and locally bounded (for every X e Int Dom the image of some neighborhood is bounded), then F is locally Lipschitz on the interior of its domain. (When F is the inverse of a surjective continuous linear operator A, this is the Banach open mapping principle.) We shall use this theorem in many crucial instances; for example, for proving the nontrivial formulas of convex analysis. We then define the transpose F* of a closed convex process F, which genera¬ lizes the usual transpose of continuous linear operators. If .4 is a continuous linear operator, we prove the expected formulas (FA)* = A* F*, (AF)* = F*A* and (Fi+F2)^ = f? + n
106 CH. 3, SET-VALUED MAPS Contrary to the case of a continuous linear operator, these formulas are not always valid (nor obvious). Actually, we shall even define the transpose of any closed convex map. Indeed, properties of the transpose are used for proving “closedness” theorems, such as the one stating that the sum of two closed convex maps is still closed. We prove in the last section of this chapter that several spectral properties of positive matrices can be extended to positive set-valued maps with closed convex graph. Let us review the theorems we plan to generalize. PERRON-FROBENIUS THEOREM Let G be a positive matrix, g{ > 0 for all i and j. i. It has a positive eigenvalue S that is larger than or equal to the absolute value of any other eigenvalue of G. ii. It is the only eigenvalue of G to which there corresponds a nonzero non¬ negative eigenvector. iii. p — G is invertible, and {p — G)~^ is positive if and only if p> 5. k M-MATRICES Let H:={h{) be a matrix satisfying (★) Vi ^7, hi^O Then the following conditions are equivalent: n i. Vi = l,Sq^>0suchthat ^ hjq^>0 ii. H is invertible and H~ ^ is positive. k Matrices H satisfying condition (★), and either one of the equivalent conditions i or ii, are called M matrices. VON NEUMANN-KEMENY THEOREM Consider two matrices F:={f{) and G:=(g{) from R" to R’” satisfying '• Vi,y, gi^O {G is nonnegative.) n ii- Vi = l m, Y, 9i>0 j=i n iii. Wj=l,...,n, Z fi>0 i=l
CH. 3, SECT. 1 UPPER AND LOWER SEMICONTINUITY OF SET-VALUED MAPS 107 Then there exist d>0,xeR\,x^0 andp e satisfying i. SFx^Gx ii. 5F'^p>G*p iii. 5 (p, Fx) = {p, Gx) Furthermore Jor all p> 6 and for all y 6 Int ), there exists x e such that pFx—Gx^y Von Neumann devised the following economic interpretation. The economy is assumed to have n production processes that produce and consume m goods. The entry /\ denotes the quantity of good i consumed by process j when oper¬ ating at unit intensity, and entry g\ denotes the corresponding quantity pro¬ duced. We assume constant returns to scale, so that the pair of matrices (F, G) completely describes the production possibilities of the economy. The operation of the economy is further specified by a vector x e R, representing the inten¬ sities at which the n production processes are operated, and hy p e F"**, repre¬ senting the price systems on the commodity space R. A triple (x,p, S) satisfying condition (★★) is called an equilibrium. A 5 such that dFx^Gx can be regarded as a growth rate, whereas a number 6 such that dF^p^G'^p can be regarded as an interest rate. This theorem states the existence of intensities and prices for which both rates coincide. In Section 4, we not only prove these theorems, but actually deduce them from analogous statements for set-valued maps with closed convex graphs. We prove the existence of solutions xeR\ such that dF{x)eG(x)-R\ when F is a convex operator and G a positive set-valued map with closed convex graph from which we deduce equilibrium theorems of the von Neumann type. When jR”=F"', we add more specific requirements that imply a generalization of the Perron-Frobenius theorem. We finally define M convex processes and prove that in some sense, they map the positive cone R\ onto itself. 1. UPPER AND LOWER SEMICONTINUITY OF SET-VALUED MAPS Let F be a set-valued map from a Hausdorff topological space X to another Y.
108 CH. 3, SEC. 1 SET-VALUED MAPS DEFINITION I fVe say that F is upper semicontinuous {in short, u.s.c.) at xqsX if for any neigh¬ borhood Ji{F(xo)) of F(xo), there exists a neighborhood X(xo) of Xo such that (1) fxeXixo), F{x)<= X{F{xo)) We say that F is upper semicontinuous if F is u.s.c. at every point xeX. k Remark We recall that when a subset is a compact subset of a metric space X, the subsets (2) B(K,riy.= {yeX\d{K,yHrj} form a fundamental basis of neighborhoods of K, in the sense that any neigh¬ borhood of K contains B{K, rj) for some >/>0. Therefore, if F is a compact-valued map F from a metric space X to a metric space Y, F is upper semicontinuous at xo € A' if and only if for all e>0, there exists rj>0 such that (3) Vx e B{xo, ri), F(x) <= B{F{xo), e) Note that if F(xo) is not compact, this property may be true even if F is not upper semicontinuous at Xo: Consider, for instance, the set-valued map from R to defined by F{^):={{x, y)\x=^}. Property (3) obviously holds when s=rj. Moreover, F is not upper semicontinuous in the usual sense: Indeed, the subset X'-={{x, 7)1 l7l<l/|x|} is a neighborhood of F(0), but for every xfO, F(x) is not contained in X- ■ In the case of single-valued maps, as we have remarked, continuity can also be characterized in terms of generalized sequences. This second point of view leads to definition 2. DEFINITION 2 We say that F is lower semicontinuous {in short is.c.) at x° eX if for any 7° 6 F(x°) and any neighborhood X{y^) of 7°, there exists a neighborhood ./V(x°) of X® such that (4) Vx e M^(x°), F(x)n X{y^)^0 We say that F is lower semicontinuous if it is lower semicontinuous at every x°eX. k Definition 2 could be phrased as follows: Given any generalized sequence X;, converging to x° and any 7° € Ftx°), there exists a generalized sequence
CH. 3, SEC. 1 UPPER AND LOWER SEMICONTINUITY OF SET-VALUED MAPS 109 e F{x^) that converges to y^. When X and Y are metric, this last characteriza¬ tion holds true with usual (i.e., countable) sequences. Examples The mapping F+, from R into its subsets, defined by i^+(0) = [—1, +1] and /’+(a:)={0}, is U.S.C., while it is not l.s.c. The mapping F_, defined by F_(0) = {0}, F_(a:) = [— 1, +1], x^O, is I.S.C., while it is not u.s.c. I +1 -1 This map is u.s.c. and not l.s.c. I ill I to-- 1 This map is l.s.c. and not u.s.c. DEFINITION 3 A set-valued map F from X to Y is said to be continuous atXoeX if it is both u.s.c. and l.s.c. at Xq. It is said to be continuous if it is continuous at every point xe X. An important class of continuous set-valued maps is provided by Lipschitz maps. DEFINITION 4 Let F be a strict set-valued map from a metric space X to a metric space Y. We say that F is Lipschitz around Xq e X if there exist a neighborhood Ж{хо) and a constant c>0 {the Lipschitz constant) such that (5) Vx, j 6 J\i{xo), Hx) <= B{F(y), cd{x, j)) It is locally Lipschitz if it is Lipschitz around every Xq e X. It is Lipschitz if there exists c>Q such that (6) V;c, yeX, F{x) c B{F\y\ cd{x,;;)) A Note that definition 4 is symmetric in x and y. We shall use a weaker nonsymmetric statement in definition 5. DEFINITION 5 We shall say that a strict set-valued map from a metric space X to a metric space Y is upper locally Lipschitz at Xq e X if there exist a neighborhood J/'ixo) of Xo
110 CH. 3, SEC. 1 SET-VALUED MAPS and a constant oO such that (7) Vx e F(x)<^B{F(xo\ cd{xo, x)) k We begin by stating that the composition of two upper semicontinuous maps is upper semicontinuous. PROPOSITION 6 Let F and G be two set-valued maps from X to Y andfrom Y to Z, respectively. Define GF by (8) GF{x)-= U G(y) У e F(x) If F and G are upper semicontinuous^ so is GF. A PROPOSITION 7 Let F be an upper semicontinuous set-valued map from X to Y with closed values. Then F is closed {i.e., its graph is closed). A Proof. Let us consider a sequence of elements yp) of the graph of F that converges to some {Xyy)eXxY. Since Fis upper semicontinuous, we can associate to any closed neighborhood Jf(F{x)) an index po such that ^p>Poy yp € Jf(F{x)). Hence, y belongs to every neighborhood of i^x). ■ There is a partial converse to this result. THEOREM 8 Let F and G be two set-valued maps from XtoY such thatWx e X, F{x)nG{x)fi0. We suppose that (9) i. F is upper semicontinuous at Xq. ii. i^Xo) is compact. Hi. G is closed. Then the set-valued map Fr\G\ x-> /T(x)n G(x) is upper semicontinuous at xo- A Proof. Let Ж:= Ж(Дхо)пС(хо)) be an open neighborhood of Яхо)пС(хо). We have to find a neighborhood Ж(хо) such that Vx 6 Л{хо\ F{x) n G(x) <= Л If Л"=>Р{хо\ this follows from the upper semicontinuity of F. If not, then we introduce the subset K:^F{xo)\y
CH. 3, SEC. 1 UPPER AND LOWER SEMICONTINUITY OF SET-VALUED MAPS 111 which is compact (since F(xo) is compact). Let P:=graph (G), which is closed. For any ;; 6 a:, we have y i G(xo) and thus (xo, y) i P. Since P is closed, there exist open neighborhoods Jiy(xo) and Ji{y) such that Pr\{Jf j,(xo) x Therefore, (10) Vx € ЖДхо), G(x)n Since K is compact, it can be covered by n neighborhoods Ji(yi). The union ^(yi) is a neighborhood of K and Ji'o Ji a neighborhood of f(xo). Since F is upper semicontinuous at xo, there exists a neighborhood -^o(^o) of Xo such that (П) Vx e Jioixo), F{x) Я We set J^(xo):= J^o(A:o)nn"=i ^y.-i^o)- Hence, when xe Jf{x^, properties (10) and (11) imply that (12) i. F{x)<=--y^\jX lii. G(x)n M=0 Therefore, F (x) n G(x) <= Ж when x e Ж(хо). ■ COROLLARY 9 Lex G be a closed set-valued map from X to a compact space Y. Then G is upper semicontinuous. ^ Proof We take F to be defined by F(x):= У for all x e X, and we apply theorem 8. И Corollary 9 provides a very useful tool for proving that a given set-valued map is upper semicontinuous. Note, though, that the assumption of the compact¬ ness of Y is essential. Consider the map F from R into the subsets of R^, defined by f(<^)={(x, ;^): >'=<^x}. Then F has closed graph: Assume that (x„, y„)-^(x*, y*), (x„, y„) e F(Q Then y„ = (J„x„, and passing to the limit, y* = i*x*, that is, (x*, y*) e F{^*). However, consider, for instance, ^°=0. For no e>0 and for no (^^0 is f (^)<= F{0)-yeB. ■ When is a complete metric space and Y a compact metric space, a map with closed graph is not only upper semicontinuous but also almost continuous.
112 СН. 3, SEC. 1 SET-VALUED MAPS THEOREM 10 Let X be a complete metric space, Y a compact metric space, and F a strict set¬ valued map with closed graph. The subset of points at which F is continuous is residual. к Proof. Since У is a compact metric space (and thus separable), there exists a countable family of open subsets Un such that for every open set ^ of £ and for every xeU, there exists i/„ such that xe U. We set К„:={хеХ\Р{х)пи„ф0} These sets are closed, because the graph of F is closed and the subsets Un are compact. We claim that the set of points x at which F is not lower semicon- tinuous is contained in the union M:=(J,7=i of the boundaries дКп of the closed subsets Indeed, to say that F is lower semicontinuous at x amounts to saying that whenever Kn contains x, then x belongs to the interior of Kn- Hence, if F is not lower semicontinuous at x, there exists Kn containing x such that x does not belong to the interior of Kn, that is, such that x belongs to the boundary дКп of Kn. Since Kn is closed, the interior of дКп is empty. Hence, Baire’s theorem implies that the interior of Ur=i is also empty. Therefore, the interior of the set of points of discontinuity of F is empty. ■ Remark We shall see in Section 3 that when X. and Y are Banach spaces and the graph of F is closed and convex, then F is lower semicontinuous on the interior of its domain. Finally, we have the following result in proposition 11. PROPOSITION 11 Let F be a strict upper semicontinuous map with compact values from a compact space X to Y. Then F(X) is compact. к We proceed now with several examples of upper or lower semicontinuous maps. Example 12 Let X and У be Hilbert spaces and A e^(X, Y). To say that the set-valued map F:=y4 ~ Ms lower semicontinuous amounts to saying that there exists a constant c>0 such that (13) '^yeY, d(0,A '(j))=^c||j1|
CH. 3, SEC. 1 UPPER AND LOWER SEMICONTINUITY OF SET-VALUED MAPS 113 The fundamental Banach isomorphism theorem (also called the open mapping principle) states that when ^ is a continuous surjective linear operator from a Banach space A" to a Banach space Y, the set-valued map ^4 “ Ms lower semicontinuous. There is no wonder that extensions of this theorem play an important role: Banach’s theorem will be extended to maps with closed convex graph. ■ Example 13 Let X be a Hilbert space, V a proper function from Xto/iu{ + oo} and V+ be the set-valued map defined by V+(x):=K(-^) + i?+ when K(-^)<+oo and y+(x):=0 when L(x)=+oo It is a well-known fact that V is lower semicontinuous if and only if y + has a closed graph. ■ It will be useful to use the upper semicontinuity of the map that associates to x the epigraph of a function depending on this parameter .v. PROPOSITION 14 Let X and Y be Hausdorff topological spaces and W\ XxY-^R a lower semi¬ continuous function. We denote by the function y-^Wx(y):=W{x, y) and by Ep Wx its epigraph. If for all xe X, y-^ V{x, y) is continuous, and if Y is compact, then (14) ^-►Ep Wx is upper semicontinuous Proof. Let ^ be an open neighborhood of Ep Wx^. Take y eY, and fix > 0 such that the pair (y, W{xq, y) - 2&y) belongs to ^. Then the set =^{(z,X)\W[xo. z)-A<г,} is open and nonempty. Let «^/(y) be the projection of ^(y) onto Y, which is open. Since W is lower semicontinuous, there exist neighborhoods Ji{y)^ ^^iy) and such that (15) Vx 6 ..f,(A-o), Vz e Jiiyl W(xo, yH W(x, z)+2e, Since y is compact, it can be covered by n neighborhoods Ji(yd- We set ^o(A-o):=n"=i Therefore, for all x € ^o(ao) and any e y there exists j;,-such that j e J^(y,);soby (15), iy(xo,>’i)< W(x,y) + 2ey.. In other words, for all 6 y and W{x, y), there exists j, e Y such that the pair (y„ A) satisfies yd-eyi<2-, that is, belongs to the set which is contained in Hence, Ep for all x € J^o(ao)- ®
114 CH. 3, SEC. 1 SET-VALUED MAPS Remark When K is a convex function from an open subset ^ of a Hilbert space X to R, we shall prove that the set-valued map x-^dV(x) from ^ to A"* associating with X its subdifferential dV(x) is upper semicontinuous when X* is supplied with the weak * topology (t(X*, X), This result remains true when V is locally Lipschitz and dV(x) is the generalized gradient. ■ Control theory provides set-valued maps of the following type. We consider three sets X, Y, and U and a map/ from X x U to Y We associate with it the set-valued map F from X to Tdefined by (16) := {/(x, m)|m € {/} We say that F is “parametrized.” PROPOSITION 15 Assume that X and Y are Hausdorff topological spaces. a. If we suppose that (17) Vw e U, x^f{x, u) is continuous then F is lower semicontinuous. b. If we suppose that (18) i. U is a compact topological space ii. / is continuous from Xx U to Y then F is upper semicontinuous. A Proof. a. The first statement is obvious. b. Let Ji be an open neighborhood of F(xq)- For any ueU, Jf is a neighbor¬ hood of/(xo, u). The continuity of/ implies that there exist open neigh¬ borhoods *^,i(xo) and Ji(u) such that /(x, v) belongs to Jf whenever x belongs to c^m(xo) and v to Ji{u). We can cover the compact set Uhyn such open sets Jf{ui). Hence ^{xoY= n ^Ui(Xo) 1 = 1 is still a neighborhood, and if v is chosen in i/, hence, in some set Ji{ui\ then f{x, v) belongs to Ji. Therefore, F(x) is contained in Ji whenever ranges over ^(xo)- ■
CH. 3, SEC. 1 UPPER AND LOWER SEMICONTINUITY OF SET-VALUED MAPS 115 The following theorem allows the construction of lower semicontinuous maps. THEOREM 16 Let К be a Hausdorff topological space, Y and Z Banach spaces,/: К x Y-^Z a continuous map that is affine with respect to the second variable. Let U and T be lower semicontinuous maps from К to Y and Z, respectively, with convex values. We suppose that U is locally bounded (U is bounded on some neighborhood of each point). We assume that there exists “у > 0 such that (19) Va: e K, Vz € Z, satisfying ||z|| 3y e U{x) such that f(x,y) + zeT{x) Then the set-valued map C defined by (20) C(x) := {y e U(x)\f{x, y) e T(x)} is lower semicontinuous {with convex values). k Proof. Let y e C(;co) and 6 > 0. We have to prove that there exists a neigh¬ borhood ^(xo) of xo 6 such that (21) Vx6W(xo), 6 C(x)n(jVo + efi) Let ^o(^o) be the neighborhood on which U is bounded, d equal to sup {diam (i/(x:))|x: € J^o(-^o)} which is finite, and e<2d. We set cc:=ye/(2d — e). Since/ is continuous and U and T are lower semicontinuous, we can find a neighborhood Mxo) *= J^oixo) and ri < e/2 such that Vx 6 ^(xo), i. If\\yo-y\\<>1, then ||/(x,>')-/(xo,jo)ll<^ (22) ■ ii. yeU{x)+riB HI. f{xo,y)eT{x)+'^B Hence, (23) Vx 6 J/'ixo), ^yx e C/(x) such that f(x, y^) e T{x) + aB
116 CH. 3, SEC. 1 SET-VALUED MAPS Assumption (19) can be written (24) '^xeK, yB<=T{x)-f{x, U{x)) Let us set 0:=y/(a+y)<l and multiply inclusion (23) by 6. By noticing that 0a=(1 — 0)y, we obtain (25) df(x,y,)eeT(x)+(l-dyyB We multiply inclusion (24) by (1 — 0) and obtain (1 -0)yBc:(l -0)T(x)-(1 -d)f{x, U(x)) by using this result in (25), we obtain (26) Of(x, y,) e dT{x)-Kl -0)T(x)-(1 -0)/(;c, U{x)) Since T(x) is convex and/ is affine with respect toy, we have proved the existence of yi 6 U{x) such that that is, such that It remains to see that Of iy, + (1 - 0)y 1) 6 T(x) Ух-=0ух + {1—0)У1 €C(x) llyo-y*ll<llyo-y*ll+(l-0)lly*-J^ill<| + ^^lb*-yiH| + ^^<5 = e So, C is lower semicontinuous. ■ It is much easier to get upper semicontinuous maps in this way. PROPOSITION 17 Let К be a Hausdorff topological space, Yand Z Banach spaces,/: К x У-^Z a continuous single-valued map. Let U and T be set-valued maps from К to Y and X, respectively, both with closed graph. Then take set-valued map C: K^Y defined by C(x)\= {y 6 U(x)\f(x, y) 6 T(x)} has a closed graph.
СН. 3, SEC. 1 UPPER AND LOWER EMICONTINUITY OF SET-VALUED MAPS 117 Proof. The proof is an easy exercise. ■ We recall that given a cone % we denote by T" its negative polar cone (27) T-:={peX^\\/ve%{p.v)^ PROPOSITION 18 LetX be a finite dimensional space and T(*) be a set-valued map associating with any xeK a closed convex cone T(x). Let N{*) be the set-valued map X-* N(x):=T(x)~. The following conditions are equivalent: (28) (29) The set-valued map 7V(*) has a closed graph. The set-valued map T(*) is lower semicontinuous. Proof, a. Let us assume that T(*) is lower semicontinuous. Let (x„, p,) e graph (M*)) be a sequence converging to {x, p). To prove that p e N{x\ let us choose any v eT{x) and check that (/?, v) ^ 0. Since T( •) is lower semicontinuous, i; = lim„_^ Vn where Vn € T(x„). Since (/?„, v„)^0 (forp„ e N{Xn)\ we deduce that (/?, v) = l\m„^^(p„, v„}^0. Hence, (x,p) belongs to graph (N(*)X and thus TV has a closed graph. b. Let us assume that the graph of TV(*) is closed; let x„ converge to x and V e T{x). We shall prove that the projection Vn:=nT(x„){v) € T(Xn) converges to v. Indeed, we know that Pn'='ttNix„)iv) e N(xn) and that \\Pn\\^\\v\\ Since X is a finite dimensional space, some subsequence (again denoted by) Pn converges top eX. Since the graph of A^(*) is closed, we deduce thatp e N{x). Hence, the subsequence Vn = v—pn converges to v—p. Since (p„, i;„)=0, we deduce that (/?, v—p)=0, that is, — (/?, v)^Q. Thus/7=0 and ¿; = lim„_^ i;„. ■ Let X and Y be two sets, G a set-valued map from Y to X, and W a real¬ valued function defined on X xY. We consider the family of maximization problems (30) V(y):= sup W(x,y) X e G(y)
118 CH. 3, SEC. 1 SET-VALUED MAPS that depend on the parameter y. The function V is called the marginal function and the set-valued map M associating to the parameter y s Y the set (31) M(y):= {x € G{y)\ V(y)= W(x, y)} of solutions to the maximization problem V(y) is called the marginal map. By the stability of this family of optimization problems, we usually mean the study of various continuity properties of the marginal function V and the mar¬ ginal map M. We begin with proposition 19. PROPOSITION 19 Suppose that {\. W is lower semicontinuous on X xY (32) { [ii. G is lower semicontinuous at yo- Then the marginal function V is lower semicontinuous at yo, ^ Proof Fix 8>0, and choose xq eG(vo) such that V(yo)—^l2^W(xo, yo\ Since W is lower semicontinuous, there exist neighborhoods Ji{xo) and J^i(yo) such that (33) Vx € jYixoX ^y e J^dyol mxo, yo)-^^ y) Since G is lower semicontinuous, there exists a neighborhood ^2(3^0) such that (34) Vj; e J^2iyo\ G{y)n Ji(xo)^0 So we take «^Oo)-= *^i(Vo)^ ^2(yo) For any y e know that there exists .x 6 G(y) such that W{x, y)> W{xo, yo)-^^ V(yo)-2^ Hence V is lower semicontinuous at yo. We single out a useful consequence in corollary 20.
CH. 3, SEC. 1 UPPER AND LOWER SEMICONTINUITY OF SET-VALUED MAPS 119 COROLLARY 20 Let Y be a metric space. If G: X-*Y is lower semicontinuous, then the function (35) (a:, 6 A" X y-^d{G(x), y) is upper semicontinuous. A We consider now the case when IFand G are upper semicontinuous. PROPOSITION 21 Suppose that (36) i. W is upper semicontinuous on X xY. (il. G{yo) is compact, and G is upper semicontinuous at yo. Then the marginal function V is upper semicontinuous at yo. A Proof. We have to prove that for any e>0, there exists a neighborhood .^(yo) of yo such that (37) Vy 6 jY(yo), F(v)< l^(yo)+s Since W is upper semicontinuous, we can associate with any x eX open neigh¬ borhoods Ji{x) of a: and jY,,(yo} such that (38) \/x'6ji(x), yyejY:,(yo), Vy(A',y)<iy(x,yo)+e Since G(yo) is compact, it can be covered by n neighborhoods Ji(xi). Therefore, (39) j/':z= Q Ji{xi) is a neighborhood of G{yo) 1=1 The set-valued G being upper semicontinuous, there exists a neighborhood Jioiyo) such that (40) '^y^Jfoiyo), G(y)<=J^ We consider the following neighborhood of yo .A^(yo):= *^o(yo)^ n •^*i(yo) i = 1 When y belongs to Jf{yo) and x belongs to (?(y), then x belongs to Ji and thus
120 CH. 3, SEC. 1 SET-VALUED MAPS to some Hence, since y belongs to J^x^iyoX we deduce from (38) that W(x, y) ^ W(xi, >^o) + e ^ V{yo) H- e Hence, by taking the supremum when x ranges over G{y\ we obtain inequality (37). ■ We point out the following often used consequence in corollary 22. COROLLARY 22 Let X be a Hausdorff topological space, Y a compact space, and W an upper semicontinuous function from X y^Y to R. Then the marginal function V defined on Y by (41) V{y):=sup W(x,y) yeY is upper semicontinuous, A When the function VFand the set-valued map G are continuous, so is the mar¬ ginal function V, thanks to the preceding theorems. We now investigate the behavior of the sets of maximizers Miy):={xeG(y)\V{y)=W{x,y)} PROPOSITION 23 Suppose that (42) i. W is continuous on X xY ii. G is continuous with compact values. Then the marginal function V is continuous, and the marginal set-valued map M is upper semicontinuous. ^ Proof We note that M{y) = G(y)r\ K{y) where K{yY=[xeX\V(y)=W(x,y)} Since V and W are continuous functions, the graph of К is closed. The subsets M{y) are obviously nonempty. Since G is upper semicontinuous, theorem 8 implies that M is also upper semicontinuous. ■ Finally, we consider the case when IFand G are both Lipschitz.
CH. 3, SEC. 2 MAPS WITH CLOSED CONVEX VALUES 121 PROPOSITION 24 Suppose that (43) i. W is Lipschitz on X xY with Lipschitz constant (. (ii. G is Lipschitz with Lipschitz constant c. Then the marginal function V is Lipschitz with Lipschitz constant ^(c +1). A Proof. We take yt e Y and we choose Xi 6 G{yi) satisfying W(x\, ;^i)+e. By assumption (43), there exists Xa 6 G^yi) such that Iki -ATall <d{xi, G(y2))+^ <c||yi-y2ll+e Since V{y2)> W{x2, ja)» we deduce that V(yt)~ V{y2H mxi, yi)~ W(X2, y2)+s <tf(|lxi -Xall +lbi -J2ID+« ^^(c+l)||>'i—ya||+(^ +1)8 Since e is arbitrary, it follows that f"(:»^i)-F(ya)<%+l)llyi-j2ll ■ 2. MAPS WITH CLOSED CONVEX VALUES Let X be a Hausdorff topological space and Y a Hausdorff locally convex vector space. We associate with a set-valued map F from AT to V its support function defined by (1) fxeX, fpeY*, a{F(x),p}:= sup (p,y) y e F(x) DEFINITION 1 We say that F is upper hemicontinuous at Xo 6 Dom F if for any p e the function x-^(j(F{x\p) is upper hemicontinuous at Xq, The map F is said to be upper hemicontinuous if it is upper hemicontinuous at any point of its domain. k Upper semicontinuous maps are upper hemicontinuous.
122 СН. 3, SEC. 2 SET-VALUED MAPS PROPOSITION 2 We supply Y with the weak topology <t(Y Any set-valued map from X to Y upper semicontinuous at Xq is upper hemicontinuous at Xq. ^ Proof Let/7^0 belong to У* and г>0 be fixed. The subset Bp{£)-={y ^Y\{p, is a neighborhood of 0 for the weak topology. Since F is upper semicontinuous at xo, there exists a neighborhood J/'ixo) of Xo such that Vx 6 Ji{xo), F(x) c F{xo)+Bp{e) By taking the support functions, we obtain the inequality Vx e Ji{xol <^(F{x), p) < o[F(xo), p)+e which expresses the upper semicontinuity of the function x-*(r(F(x),p) at Xo- Ш Remark The converse is not true. For example, consider the map F from R to defined by F(x):={iyuy2)\y2>{iFx)yi} Then F is upper hemicontinuous at x=0. Indeed, fixр={риР2У EitherP2>0, and there is nothing to prove, because <t(F(0), p)= +oo, or p2<0. In the latter case, an easy calculation shows that 0-(f(x), p) = -^piip2(i +^)] ‘ is continuous at x=0. However, it is easy to see that F is not upper semicon¬ tinuous at x=0. ■ Actually, the converse becomes true if we make the additional assumption that the images F{x) of F are convex and weakly compact. This theorem, due to Castaing, is proved in the last part of Section 2. We now mention some useful elementary facts. PROPOSITION 3 A finite sum and a finite product of upper hemicontinuous set-valued maps is upper hemicontinuous. A
CH. 3, sec.2 maps with closed convex values 123 PROPOSITION 4 If X is compact and F is upper hemicontinuous with bounded values, then F{X) is bounded in Y. A Proof Since the images of F are bounded, the functions x-^(t(F{x\ p) are finite. Since X is compact and F is upper hemicontinuous, then the function (¡> defined on Y* by (2) V/7 € Y'^, Ф(рУ= sup o(F{x), /?)< + 00 xeX is a finite positively homogeneous lower semicontinuous convex function: It is the support function of a bounded closed convex subset KoiY. ■ PROPOSITION 5 The graph of any upper hemicontinuous map with closed convex values is closed in X xY (when Y is supplied with the weak topology). A Proof Let us consider a generalized sequence of elements {x^, y^) of graph (F) converging to (x, y) in X xY Since for all p e У*, {p, yn)^o{F{x^i), p) and x^o{F{x\ p) is upper semicontinuous, we deduce by taking the limit that {p,y)=lim <P. Уц) ^ lim sup p)^p) Hence, у eco{F{x)=F{x) The following theorem plays an important role in the theory of differential inclusions as well as in other domains. THEOREM 6 (CONVERGENCE THEOREM) Let F be an upper hemicontinuous set-valued map with closed convex values from a Banach space X to a Banach space Y. Let Q. be a measured space. We consider two sequences of functions x„(•) belonging to X) andу„{•) 6 L^(Q, Y)(p, q>l) satisfying (3) for almost all co g Q, for every there exists an integer N:=N{o), s) such that V«^N, d{{x„((o), y„(co)), graph {F))^s
124 CH. 3, SEC. 2 SET-VALUED MAPS If we assume that (4) i. Xfc(') converges strongly to x{*) in X). ii. ) converges weakly to y(^) in Z?(Q, Y). then we conclude that (5) for almost all co 6 Q, y{co) e F{x{(o)) A Proof a. We recall that for convex subsets of a Banach space (here, I?(Q, Y)\ the strong closure coincides with the weak closure (for the weakened topology o{I3(Q, y), I3*{Q, y^)) with 1/^4-= We use this fact. For any integer «, the weak limit j(*) belongs to the weak closure of the convex hull ^o{ym(')}m>n of the subsets Hence, we can choose functions (6) (where the coefficients a” are nonnegative, equal to zero except for a finite number, and Zr=n=1) belonging to this convex hull such that n In other words, the sequence z„(*) converges toy{') strongly in i?(Q, y). b. Since x„{ -) converges strongly to x( •) in if(i2, X) and z„( •) converges strongly to >'(•) in i?(Q, y), there exist subsequences (again denoted by) x„(') and z„(*) that converge to x(*) and y{-) for almost all cu € Q. Let us choose cu e fi such that x„(cu) converges to x{o}), z„(co) converges to y{co), and assumption (3) holds true. Let us fix e>0 andp e V*. Since F is upper hemicontinuous, there exists rj e]0, e/2||p||*[ such that (7) I|m - 4<w)ll < 2t]=^a{F(u), p)< a{F(x{(o), p)+- Assumption (3) implies that we can associate to t\ an integer N such that .g, iVw ^ A^, 3 (m„„ u,„) e graph (f) such that Since x„,(cw) converges to x(co), there exists such that l|A:m(iw)-x(ft))||<i/ when
CH. 3, SEC. 2 MAPS WITH CLOSED CONVEX VALUES 125 Hence, (8) implies that ^o(F(u,„, p)+ti\\p\\i, and (7) implies that <t{F(u„), p) < <7{F{x{co)), p) 4- 2 because ||m„—x(f)||^2i;. Therefore, since >/<e/2||p||,^ we obtain (9) 'im^Nu {p,yJ(o))^o(F(x{oi),p)+z Let n be larger that Ni. By multiplying inequalities (9) by for all and adding them together, we deduce that (10) >/n>Ni, (p, z„(co)}<<r(f(л:(й)),p) + e By taking the limit when n-*oo and e-»0, we obtain Vp 6 y*, (p, p((o)> < ff(F(x(co), p) that is, y{a>) 6 со f (x(cu))=F{co). ■ As a first consequence, we deduce the following corollary. COROLLARY 7 Let F be an upper hemicontinuous map from a Banach space X to a Banach space Y andO. a measured space. We supply lF{f).,X) with the strong topology and У) with the weak topology. Then the graph of the map from X) to E{Sl, Y) defined by (11) Щх(’)):={у 6 i?(Q, y)| for almost all o) 6Q, y{co) e F(x(o)))} is closed in И(П, X) x i?(Q, У). ▲ Proof. We take a sequence of elements (x„(*), Уп(‘)) belonging to graph (#) satisfying (4). Since property (3) holds true, we deduce that the limit (x(-), y(*)) belongs to graph ■ COROLLARY 8 We consider two Banach spaces X and Y, a compact convex subset Kof Y and a
126 CH. 3, SEC. 2 SET-VALUED MAPS lov^er semicontinuous function W\ X x K-^R satisfying (12) Vx e X, y-^ W{x, y) is convex and continuous Let us consider sequences of functions x„ e L^(Q, X\ynS Z?(Q, K) and w„ e I3{Q) such that (13) We assume that i. Xn converges strongly to x in Z?(fi, X). ii. y„ converges weakly to y in I?(Q, 7). [iii. w„ converges weakly to v in Z?(Q). J/or almost all co e Q^for every e > 0, there exists an integer such that V7(x„(co), j^„(co))^ w„(co)+e for all n>N Then (15) for almost all co g Q, W(x((o\ ;;(co))^ o(co) Proof We apply theorem 6 to the set-valued map F from Z to 7 x defined by F(x):=Ep(Wi) ■ A slight variant of the proof of corollary 8 by means of the convergence theorem and proposition 14 yields the following useful proposition. PROPOSITION 9 We consider two Banach spaces X and 7, a compact subset KofY and a nonnega- tive lower semicontinuous function W\X x satisfying (16) Vx G X, y-^ W{x, y) is convex and continuous We define the following functional W on X) x 7) (17) W{x{03), y{(o))d^i((o) e [0, +oo] This functional is lower semicontinuous when X) is supplied with the strong topology and Z?(0, 7) is supplied with the weak topology. A
CH. 3, SEC. 2 MAPS WITH CLOSED CONVEX VALUES 127 Proof. We take a sequence of functions x„(-) converging strongly to x(*) in if(Q, X) and a sequence of functions ) converging weakly to y(-) in i?(Q, 7), then consider t):=lim inf„^„ y„(*))- If f = ‘he theorem is true. If not, there exist subsequences (again denoted by) x„(‘) and y„(*) such that (18) V«, (f(x„(-),7»(-))^i^ + 1 Introduce the sequence of functions z„(*) defined by (6) that converges strongly to y(*) in i?(Q, y). Therefore, there exist subsequences (again denoted by) x„(’) and z„(*) such that for almost all cu e Q, x„(co) converges to x{co) and z„{co) converges to y(co). Let Wx denote the function y->- Wx{y)'.= W(x, y). Proposition 14 implies that the map x->Ep (W^) is upper semicontinuous. Therefore, for every i/>0, there exists N:=N(rj, x(co)) such that fn>N, {y„{o}), W(x„{(o\ y„(co)) € Ep (Wx)+ri(B x B) where BxB denotes the unit ball in 7 x Since the epigraph of Wx is convex by assumption (16), we obtain by setting (19) v^(o):= ^ a"W(x„(o)),y„(w)) that (20) fn>N, (z„(o)), t)„(co)) 6 Ep (17*) + ;;(ß X B) We introduce (21) ü(io):=Iim inf t)„(cu) n-*oo We can associate to e anij < e such that 17(x:(cü),y(co))<I7(x(a)), z)+8 when I|z-;;(cü)||^2/; Let Ni>N he chosen such that ||z„(co)-;;(co)||<>j when n^N, By (20) there exist z„ ez„(o>)+>;5 such that ■ j v /> 17(x(oj), z„) < v„(a))+ri^ v„((o)+e Finally, by (21), there exists n^Ni such that t)„(o))<t,(ft))+£ Therefore, W(x(o}), Mo>))<t)(co)+2e
128 CH. 3, SEC. 2 SET-VALUED MAPS By letting e converge to zero, we have proved that (22) for almost all co e Q, W{x((o), ;;(io))< lim inf v„(a>) 00 We integrate this inequality and apply Fatou’s lemma, which is possible because the function W is nonnegative. We obtain i Wm \n{ v„{(jo)dti{co) Jii n->00 < lim inf v„{co)dn{(o) n-*oo Jci But, we observe that by (18) and (19), 1'’”' io))dn(co)= Y, air ),>'».('))<t^+- 1и=и и SO that 4m inf |Ял:и(-),Ы-)) ■ fJ-»00 We now prove the announced partial converse of proposition 2. THEOREM 10 Let F be a strict upper hemicontinuous map from X to Y, If F{xo) is convex and weakly compact, F is upper semicontinuous at Xq. ^ We now need lemma 11. LEMMA 11 We posit the assumptions of theorem 10. Let ^ be a weakly open set containing F{x^). Then there is a finite set of pairs (pi, s,), such that = {y\{pi, y)^o{F{x\ Pi) + Ei \ i = \, is contained in' Proof of Theorem 10. Let be a neighborhood of F{xo)- Since F(xo) is weakly compact, there exists a finite set of pairs (p,-, Ej) such that iP" is contained in ‘W by lemma 11. So, it remains to note that since F is upper hemicontinuous
CH. 3, SEC. 2 MAPS WITH CLOSED CONVEX VALUES 129 at Xo, there exists a neighborhood Ji(xo) of xo such that for all x € J/'ixo), f(x)«= #■, and, consequently, f(x)<='2f. B Proof of Lemma 11. We proceed in two steps. a. Case when Y=R". We can replace by F(xo)+5, since F(xo) is com¬ pact. Set K to be the complement of in the closure of F{x°) + 2B: is a compact set. Fix any p e Y*, ||p|| = 1 and any e e ]0, 1 [. The set K{p, e)\= {<p, F) < ff(F(x®), p)+s}nK is (compact and) nonempty: Consider a point e F(x”) such that (p, =inf {{p, y)\y e F(x°)} and add a vector (—2z), where z is such that (p, z) = ||z|| = 1. Then the distance of Pm — 2z from F (x°) is exactly 2; hence, it belongs to K. Assume that lemma Ills false. Then the family {K{p, e)} has the finite intersection property, that is, no matter which finite set of pairs (p„ e,) we choose, the set f]K(p¡,ed^0 i Since K is compact, the intersection over all the K(p, e) must be nonempty. Let ^ be in this intersection. By a basic separation argument, there are p, e such that (p, 0>(T(F{x°),p)+e. Hence, i /C(p, s), a contradiction. This proves the lemma. b. General case. For every y e F(xo), let (F, E)y be a finite set of pairs (p„ e,) such that y-I- A^(F, E)yC.‘^, where N(P, F)/={z| Kpi, z)| <6,- for all i} Since F(xo) is compact, let {y¡-\-N(,P, E)¡] be a finite subcover of F(xo). [We set N(P, E)j :=N{P, E)yj.'\ Consider the finite set of all functionals {p,j} for all points y¡, and write Y=M+N, where M is a finite dimensional space and N the intersection of all the kernels of these functionals. Since M is a finite dimen¬ sional space, the result is true on M. Set II to be the projection onto M. For each j, n(yj+N(P, E)j) is open, and the union over; covers n(F(x®)). On M, consider the bounded closed convex set riF(x®), its open neighbor¬ hood UjIl(yj-l-iV(F, E)j), and let a^{TlF{x% p) be its support function. Then there is a finite set of pairs (P*^, E) with pf^ e M*, such that the set (23) {yeM\{pf‘,y)<a^(nF(x%pf‘)+e¡} is contained in [j {yjF N{P, E)j).
130 CH. 3, SEC. 3 SET-VALUED MAPS Extend each />“ to a pt e Y* by (Ph y} = (Ph m+n):= (pf^, m) Then it is easy to verify that a{F{x%pd=<T^(nFix%pf*) We claim that the set '={y\(Pi> y) «^(P(x°),pd+er. 1 = 1,..., m} is contained in Let y belong to it, and consider Ily. For every /, (pf^, n_p) = (pi,y}<a(F(x°),Pi)+Si=<T"(rif (x”),p!^)+£,• By (23), there exists some j* such that Ily e n(jj* 4- N(P, E)j*), that is. Fly =yj* -I- Ilu with V 6 N{P, E)j*. Fix (pif, e,) in (P, E)f. Note that since N is contained in the null space of pip for every vector (pip, i) = {Pij*y IlO- Hence, (PiA y) = (PiA ny) = (PiA Hy/+Hu) < (pij*, yj*}+ej This means that y eyf + NiP, E)fcz<^ ■ 3. MAPS WITH CLOSED CONVEX GRAPHS We observe that the graph of a set-valued map F from a Banach space X to a Banach space Y is convex if and only if for any convex combination of elements x,- € AT, we have E A,F(x,)cf Set-valued maps satisfying the opposite inclusion S A,F(x,)=>F X ?iiXi i=l \i=l for all convex combinations were introduced by Ioffe, who called them fans.
CH. 3, SEC.3 MAPS WITH CLOSED CONVEX GRAPHS 131 Let F be a map with convex graph; then its images F{x\ its domain, and its range are convex subsets. The inverse of a map with convex graph has a convex graph. We observe that a convex map is closed if and only if it is weakly closed, because in Banach spaces, closed convex subsets and weakly closed convex subsets do coincide. Examples of convex maps are provided by convex functions V from X to i?u{+oo}: The function V is convex if and only if the set-valued map V+ defined by y fK(x)-h/?+ when K(x)<-f-oo ‘ [0 when K(x)= +00 is convex (because the graph of V+ is the epigraph of V). More generally, if P is a closed convex cone of a Banach space Y, we associate with a single-valued map from its domain D(A) to Y the set-valued map A+ defined by . . {A{x)-\-P whenxe/)(^) ’“W whenx^Z)(^) and we observe that the following conditions are equivalent: i. The map A+ is convex. n Í ” \ ii. For all convex combination, ^ ^íA(xí) e A ( ^ ) + P We say that such maps A are P-convex operators. We shall pay special attention to the convex processes, which are set-valued maps whose graphs are convex cones, that is, the set-valued maps satisfying the properties i. Vxi,X2 6X, F{Xi)-\-F{X2)^F(Xi-\-X2) ii. VA >0, Vx 6 X, F (Ax)=AF (x) A closed convex process is a set-valued map whose graph is a closed convex cone. Convex processes are set-valued analogues of linear operators, since a strict single-valued map F is a convex process if and only if F is linear. Closed convex set-valued maps from Banach spaces to Banach spaces are lower semicontinuous on the interior of their domains. This important result is an extension of the closed graph theorem for continuous linear operators. Since it is customary to prove the closed graph theorem from the Banach open mapping principle, we state the Robinson-Ursescu theorem in the following way in theorem 1.
132 СН. 3, SEC. 3 SET-VALUED MAPS THEOREM 1 (CLOSED GRAPH) Let X and Y be Banach spaces^ F a proper closed convex map from X to Y, whose range has a nonempty interior. Let уо 6 Int (Im F) and xoeF~ be chosen. There exists у >0 such that (1) Vx 6 Dom F, V;; ey^FyB, d(x, F~^{y))^-d{y, F(x))(l +||x-Xo||) A У In particular, F~ Ms lower semicontinuous on Int F( By taking x=Xo, we have (2) "^уеуо + уВ, d(xo,F 4y))^~ d(y, F{xo)) COROLLARY 2 Let X and Y be Banach spaces, F a proper closed convex map from X to Y, whose range has a nonempty interior. Assume furthermore that F~^ is locally bounded. Then F~^ is locally Lipschitz on Int Im F. к When F is a closed convex process, theorem 1 yields corollary 3. COROLLARY 3 Let F be a closed convex process from a Banach space X to a Banach space Y If (3) Im F = У {i.e., F is surjective) then F " ^ is a Lipschitz: There exists у >0 such that (4) '^УиУ2 € У, Р-НУ2)<=Р-^{У,) + -\\У,-У2\\В У Proof. Since the image of a convex process is a convex cone, then 0 belongs to the interior of this image if and only if it coincides with the whole space Y. We take xq=x=0 and yo=0 in inequality (2). We obtain 1 1 fyeyB, d{0,p-\y))^-d(y,Pm^- Since P * is positively homogeneous, this inequality applied to y(;^/||;^||) implies that (5) 4y&Y, d{0,F-^iy)H~ У
CH. 3, SEC. 3 MAPS WITH CLOSED CONVEX GRAPHS 133 Consider now yi and 72 in Yand X2 e F~'^(y2)- For all 6>0, we can choose XeeF~^{yi—y2) such that ||Xe||<</(0, f + ll>'l-3'2ll+e Hence, Xie'=X2+Xt, belongs to F"‘(>'2+.Fi-J'2) = F“Hj'i) (because f is a convex process) and consequently, X2=xu—belongs to Ibi —i'211+e^. F Hyi)+(-|bi-l'2ll + e)B foralls>0 By letting e converge to 0, we deduce that inclusion (4) holds true. Let us apply theorem 1 to the map F, defined by Ax—M whenxeL (6) when x^L COROLLARY 4 Let L^X and MczY be nonempty closed convex subsets of Banach spaces X and Y, and let A belong to S£"{X, 7). We set (7) ^yeY, F~Hy)'={xeL\AxeM+y} Let us assume that (8) lni{A(L)-M)^0 Then for all yo e Int [A{L) — M) and xq sF~ exists y>0 such that (9) VxeL, Wyeyo + yB, ¿/(x, +||x-Xo||) A y Proof The graph of this map is obviously closed and convex, its image is equal to A{L)-M, its inverse is defined by (7), and d(y, F{x))=dM(Ax-y). These remarks made, corollary 4 follows from theorem 1. ■ When L and M are cones, corollary 3 implies the following statement in corollary 5.
134 СН. 3, SEC. 3 SET-VALUED MAPS COROLLARY 5 Let PcX and Q<=Y be closed convex cones of Banach spaces X and Y, and let A be a continuous linear operator from X to Y. We assume that (10) Y=A(P)+Q Let be the map definedfrom Y to X by (11) Vjey; F-^{y):={xeP\AxeQ+y} Then F~^ is Lipschitz: There exists y>0 such that (12) '^yuyi^Y, F 4y2)<=F ‘(.Pi)+-Ibi-72II^ У By taking P:=X and Q := {0}, we obtain the Banach open mapping principle stated in corollary 6. COROLLARY 6 Let A be a surjective continuous linear operator from a Banach space X to a Banach space Y, Then the inverse is a Lipschitz set^valued map from Y to X, Another consequence of theorem 1 is the following property of lower semi- continuous convex functions stated in corollary 7. COROLLARY 7 Any proper lower semicontinuous, convex function V from a Banach space X io {+ oo} is locally Lipschitz on the interior of its domain. A Proof We recall that Baire’s theorem implies that V is continuous on the interior of its domain. We apply theorem 1 to the inverse f “ ^ V; ^ of the map y + : X^R defined by У+(д^):=К(л:) + Л+when л: 6 Dom К and \+(x)=0 otherwise Let Xq 6 Int Dom V = lni Im F and yo = V{xo)eF ‘(xo) be given. Since V is continuous at jcq, there exists Vo < V such that | V{x) — V(xo)\ < 1 when xexo+joB. Let Xi and X2 belong to Xo -I-yoB. Assume that F(x:i)> Vixz); then, we apply theorem 1, which states that since Xi belongs to Xq + yoBcxo+yB
CH. 3, SEC. 3 MAPS WITH CLOSED CONVEX GRAPHS 135 and V(x2) belongs to Dom F, we have \V{X,)-V{X2)\ = V{X,)-V{X2)= inf \X-V(X2)\ AeF-‘{x,) = d(F-\x,\ K(X2))<- d{Xy, F{V{X2)m+\V{X2)-F(Xo)|) y The proof of theorem 1 is rather involved. We begin by deducing it from a more concise statement, given in proposition 8. PROPOSITION 8 We posit the assumptions of theorem 1. Let e Int Im(f) and XqbF ” ^(>'o) be given and B denote the unit ball of X. Then after setting K.= Dom F, we have (13) yo eint f(A:n(xo + B)) We shall decompose the proof of proposition 8 into three lemmas. But first, we deduce theorem 1 from proposition 8. Proof of Theorem 1 from Proposition 8. Let x e L be fixed. By proposition 8, there exists y>0 such that (14) PoF XyB^ F{Kr^ (xo + B)) Let;;e>;o+yfibe given. If x belongs to F \y), then ¿?(x, f ”^(y))=0 and the conclusion is satisfied. If not, for all s>0, there exists zeF(x) such that lb-z||<i/(;^, f(x))(l+e) Since y + yBcF{Kr^{xo + B)\ we obtain (15) y(y-z) Ib-^ll eF(/i:n(xo + 5))- y Let us set lb —z||/(y + ||>’—z||), which belongs to ]0, I[. We can write inclu¬ sion (15) in the form (16) {i-X)(y-z)eXF{Kr\ (xo + B)) - Xy
136 CH. 3, SEC. 3 SET-VALUED MAPS Since (1 — A)z 6 (1 — A)F (x) and the graph of F is convex, we obtain (17) :peAF(/:n(xo-l-5))+(l-A)f(x)cf(A(^n(xo-l-5))-l-(l-A)A) Then there exists Xi € A!^n(A:o-l-F) such that, by setting Xy\=Xxi+{\—X)x, we have Xy e f ' Hj)- Furthermore, we observe that l|x.-x|| = ||Axi+(l-A)x-x|| =A||xi -x|| A(||x-Xo|| +11^1 - a:o||)< A(||x-Xo|| +1) because Xi exQ-\-B. Now and thus \\y-z\\ ^d{y,F(x)t\+e) y + lb-z||" y (18) j/ 17-1/ M^ii F(x))(l+s) di,x, F : (Ik-xoll + l) By letting 6 converge to zero, we obtain (19) d(x, F - i(y))<i d{y, f (x))(||x-Xo|| +1) The proof of proposition 8 is analogous to the proof of the open mapping theorem. We use lemma 9. LEMMA 9 Let T be a subset of a Banach space Y satisfying (20) 1 °° ^ 2-<‘T<=T ^ k = 0 If zero belongs to the interior of the closure ofT, it actually belongs to the interior of T A Proof By assumption, there exists y>0 such that 2yB^ f Hence, for every /:> 1, we have 2-2~^yB^2~^T. Let y e yB. Then there exists VoeT such that 2y-Voe2-2-^yBc:2-^f since 2y ef.
CH. 3, SEC. 3 MAPS WITH CLOSED CONVEX GRAPHS 137 Let us assume that we have constructed a sequence of elements Vk^T (O^k^n-l) such that (21) 11-1 2y- ^ 2-\<=2’2-"yB k = 0 Since 2’2 "yB<=2 "7; we can find e T such that n- 1 2y- Y, 2-'‘Pfc-2“"i;„e2-<"+i>v5<=2"<"+‘>r k = 0 In this way, we have constructed a sequence of elements v^eT such that 1 oo 1 00 y=^ Z 2-Ve- £ 2-'=T By assumption,;; belongs to T. So, we have proved yB^T. ■ We apply this lemma to the subset (22) T:=F{Kn{xo + B))-yo Lemma 10 proves that zero belongs to the interior of the closure of T and lemma 11 states that this subset T satisfies property (20). Hence, we shall con¬ clude that zero belongs to the interior T, that is, the conclusion of proposition 8. LEMMA 10 We posit the assumptions of proposition 8. Then zero belongs to the interior of the closure of the subset T:=F{Kr\{xQ-\- B))—yo. A Proof We set Kn'.=Kri{xo-\-nB\ Hence, T=F{Ki)—yo. We remark that A^ = (Jr=i thus F(iC)=(J*^i F(/C„). We note also that 1 ) Xo + — Kn ^ K\ nJ n The graph of F being convex, we deduce that (23) Since 0 G Int (F(K)—yo\ there exists y>0 such that yBciF{K)-yo=0 {F{K„)-yo) i--]F{xo)+-FiK„)<=F{K,) n) n
138 CH. 3, SEC. 3 SET-VALUED MAPS But Jo e F(xo), and (23) implies that 1- that is, F(A:„)-jo<=«(F(A^i)-Jo)=nF Therefore, yF<=U"=i nT. Baire’s theorem inches that the interior of some subset n r is nonempty. Then there exists Xo e T and 5 > 0 such that xo -I- ¿5 <= T Since — yxo/||A:o|| belongs to yB, and thus to the union of the nVs, there exists n such that — yxo/«||A:o|| belongs to T. Let -l:=y/(y-l-n||xo||) e]0,1[. Hence, XbB=Xxo-(\-X)^^¡^+XдBcXT+{\-X)TclT «Ikoll because f is convex. We have shown that zero belongs to the interior of T It remains to check that T satisfies property (20). LEMMA 11 We posit the assumptions of proposition 8. Then the subset T:=F{Kr\ (xq + B))—yo satisfies the property \ £ 2-*TcT A ^ k = 0 Proof. We takejo=0 for the sake of simplicity. Letj*belongt02^”=o2”*‘T We set «„:=l/E*=o 2'* so that a„->i and yn-=ct„tl=o'^~\ where v^eT so that y„ converges to j*. By definition of T, we can find e /Cn(xo -I- B) such that Vk € F{uk). Since the graph of F is convex, we deduce that (24) y„ea„Y, 2 '‘F(Mfc)cF(a„ ^ 2 \ k=0 \ k=0 Let us consider the sequence of elements x„:=(x„Y,l=o Since Uk exo+B, we deduce that it is a Cauchy sequence, which converges to some x^ for X is complete. Since Kn{xo-\-B) is convex, x„ belongs to Kn{xo-\-B). This set also being closed, x^ belongs to K n {xq + B). Inclusion (24) says that the sequence of elements (x,,, y„) belongs to the graph of F. Since it is closed, it follows that {x^, y^) e graph F, that is, y^ belongs to F(x*)^F{Kn{xo-\-B)) = :T. ■ We can adapt the concept of transpose to set-valued maps.
СН. 3, SEC. 3 MAPS WITH CLOSED CONVEX GRAPHS 139 DEFINITION 12 Let F be a set-valued map from a Banach space X to a Banach space Y. We associate with F the convex process F* from У* to X* defined in the following way (25) i/j e F*(q) if and only if ip, —g) belongs to the barrier cone [b(graph (f)) of graph (f) We say that F* is the transpose of F. k In other words, (26) p e F*iq) if and only if sup sup {{p, x) — (g, j))< + oo xeX ye Fix) When F is a process, then its transpose is the closed convex process defined by {p € F'^(q) if and only if (/?, —q) belongs to the negative ^ [polar cone (graph {F))~ of graph F or, equivalently, (28) p € F*(q) if and only if Vx € X,;; g F(x), (p, x) ^{q,y) since the barrier cone of a cone is its negative polar cone. When G is a set-valued map from 7^ to we define the convex process G* from X to y by (29) y 6 G*(x) if and only if (x, -y)e b(graph (G)) [instead of requiring that ( — x,y) belongs to the barrier cone of graph (G)]. Remark When F is a continuous linear operator from XioY, its transpose as a continuous linear operator coincides with its transpose as a closed convex process. This is why the minus sign appears in the definition of F*. It could have appeared in front of X instead of it is a matter of convenience. ■ We now list formulas allowing the characterization of transposes. The transpose of F“ ^ the inverse of F, is given by (30) (F-^np)=^(F*)-\-p) PROPOSITION 13 Let X, X cmd Z be Banach spaces, F a set-valued mapfrom X to У, and В eS£(Y, Z).
140 CH. 3, SEC. 3 SET-VALUED MAPS Then (31) {BF)* = F*B* A Proof. Indeed, the graph of BF is equal to (1 x B) graph (F). By formula (25) of Section 5, Chapter 1 b((l xB) graph (F))=(l x5)*”‘b(graph (F)) so that (p, —q) belongs to b(graph {BF)) if and only if {p, —B*q) belongs to b(graph (F)); that is, ifp belongs to F*{B*q). ■ When A 6 if(ATo, 2f), the set-valued map is proper if and only if the inter¬ section of Im A and Dom F is nonempty, that is, if and only if zero belongs to Im A — Dom F. We shall make a stronger assumption in proposition 14. PROPOSITION 14 Let Xq, X, Y be Banach spaces, F a closed convex map from X to Y, and A 6 SnXo, X). If (32) then (33) 0 6 Int (Im A — Dom F) {FA)*=A*F* Proof Indeed, graph {FA)={Axl)~^ graph (F) and, consequently, graph (F.4)*=b((^ X1)”^ graph (F)). We apply formula (30) of Section 5, Chapter 1. For that purpose, we have to check that condition (34) (0,0) € Int (Im {A X l)-(-graph (F)) holds true. It follows from assumption (32). Indeed, let y>0 be such that yBczlmA — Dom F. Let (x, y) belong to the ball of radius y > 0 in x L Since X can be written x=Axo — xi, where xo belongs to Xq and jci belongs to the domain of F, then we can write {x,y)={Axo,yo)—{xi,yi) wheree F(xi) and jo=F+.Vi- Hence, {x, y) 6 Im (/4 X1)- graph (F) Therefore, formula (18) of Section 5, Chapter 1 implies equality (35) graph ((F^l)*)=b(graph {FA))={A* x l)b(graph (F))
CH. 3, SEC. 3 MAPS WITH CLOSED CONVEX GRAPHS 141 and thus, (36) r e (FA)*(q) if and only if there exists p e F*(q) such that y—A*p ■ COROLLARY 15 Let V, X, Y, and Z be Banach spaces, F a closed convex map from X to Y, let A belong to £F(V, X), and BtoSe{Y, Z). If (37) then (38) 0 e Int(Im A — Dom F) (BFA)*=A*F*B* COROLLARY 16 Let X, Y, and Z be Banach spaces, A e ^(X, Y), F a closed convex map from X to Z, and G a closed convex process from YtoZ. If (39) 0 e Int(/1 Dom F— Dom G) then (40) (F+GA)*=F*+A*G* A Proof Let us set H(x):=F{x)+G{Ax), B{y, z):=y+z and (i'xG)(x, j):=f(x)xG(y) and {I xA)(x)=(x, Ax) Then (41) H:=B{FxG){lxA) Since Dom (F X G)=Dom F x Dom G assumption (39) implies that (42) 0 6 Int(Im(l xA)— Dom(F x G))
142 CH. 3, SEC. 3 SET-VALUED MAPS Indeed, if («, p) 6 X x i; then there exist s> 0, e Dom F and y e Dorn G such that e{Au—v)— —Ax+y By setting z=x + eM, we see that s(u, v)=(z — x, Az—y) e(l xA)z—Dom (f xG) Hence, assumption (42) holds true. Therefore, the preceding corollaries imply that ^*=(1 xAf(FxGfB^, Since (f xG)*(^)=F*(p)xG*(p) and (1 xA)* = 1 + A*', we deduce that COROLLARY 17 Let F be a closed convex map from X to Y and K<^ X be a closed convex subset We assume that (43) 0Glnt(A:-DomF) Then the transpose of the restriction F|k F to K is defined by (44) {F\Knq) = F*{q) + b{K) A Proof We apply corollary 16 with ^4 = identity and G defined by G(x)=0 when X € K\ G{x)=0 when x$ K, whose domain is Ky whose graph is x (0), and whose transpose G* is the constant map defined by G*(q)=^b(K), I PROPOSITION 18 Let Fi and F2 be two closed convex maps from X to Y. If (45) then (46) 0 6 Int(graph(Fi)- graph(F2)) (Fi oF2)*(i)=Ff(^i)-l-Ff(^2) where q=qI +q2 Proof Since the graph of FioF2 is the intersection of the graphs of Fi and F2, assumption (45) and formula (29) in Section 5, Chapter 1 imply that 6(graph(F 1) n graph(F2))=¿>(graph(Fi)) -H ¿(graph(F2)) Therefore, the graph of (Fi oF2)* is equal to the sum of the graphs of Ft and Ft ■
СН. 3, SEC. 3 MAPS WITH CLOSED CONVEX GRAPHS 143 The formula (Im ЛУ = Ксг satisfied by continuous linear operators can be adapted to a set-valued map and yields the same surjectivity conditions. PROPOSITION 19 Let F be a set-valued map from X to Y. Then (47) b(ImF)=-F*"40) Proof It is obvious that an element q belongs to the barrier cone of Im F if and only if the pair (0, q) belongs to the barrier cone of the graph of F; that is, if (0, —q) belongs to the graph of F*. ■ We deduce the following interesting result in proposition 20. PROPOSITION 20 a. Let F be a proper closed convex map from a Banach space X to a Banach space Y. Then the image of F is dense if and only if (48) F*-H0)={0} b. Let F be a closed convex processfrom X to Y Then F is surjective if and only if the image of F is closed and F^ “ ^(0) = {0}. A Proof a. Since the image of F is convex, then Im F is equal to Y if and only if its barrier cone is equal to {0}, that is, if and only if F^"^(0) = {0} by proposition 19. b. Since the image of F is a closed convex cone, then Im F=(Im F)~~ = (b(ImF))- = T ■ COROLLARY 21 Let F be a set-valued map from X to Y and let K^X be a closed convex subset such that (49) Then 0 6 Int(Dom F — K) (50) b{F(K))=-F*-\-b[K)) Proof. We apply proposition 9 to the map F\k, since its image is equal to F(K). Then q belongs to the barrier cone of F{K) if and only if zero belongs to (f Ik)*(-?)> which is equal to F*(-q)+b(K) by corollary 17. This amounts to saying that — q belongs to F* ” *(—b{K)). ■
144 CH. 3, SEC. 3 SET-VALUED MAPS As a consequence, we obtain the following extension of Farkas’s lemma to closed convex processes. COROLLARY 22 Let F be a closed convex process from X to Y and let K be a closed convex cone of X such that (51) Then (52) If F{K) is closed, then (53) DomF-K = X F(K)- = -F*-\-K-) F{K)=^(F^-H-K-)r We shall use proposition 20 to prove an extension of the Lax-Milgram theorem to closed convex processes. DEFINITION 23 We shall say that a set-valued map F from X to X'^ is X elliptic if (54) Г 3 c > 0 such that for any two (л:', У) e graph(F) Ix^ —x^)'^c\\x^ LEMMA 24 The image of an X-elliptic map F with closed graph is closed, and its inverse is single valued and Lipschitz with constant c~^. A Proof The fact that F~^ is single valued follows from (54) by taking y=y^ =/ an Xu X2 in F~ ^0). Inequality (54) also implies that c\\F Hyi)-F Hy2W^c\\yi-y2\\\\F Hyi)-F Hy2)\\ To prove that Im(f) is closed, let us consider a Cauchy sequence of elements Pn G Im F. Let us take x„ in F " ^ (/?„). Since F is X elliptic, we deduce that and, therefore, that the sequence of elements x„ is a Cauchy sequence. Then the sequence of elements e graph (F) converges to some (x,p), which belongs to the graph of F, since the latter is closed. Hence,/? belongs to Im(F). We have proved that Im(F) is complete and, thus, closed. ■
СН. 3, SEC. 3 MAPS WITH CLOSED CONVEX GRAPHS 145 We deduce a surjectivity criterion analogous to the Lax-Milgram theorem on A"-elliptic continuous linear operators. PROPOSITION 25 Let F be an X-elliptic closed convex process from X to X*. If (Dom F)~ dm F and if the domain of F is closed, then F is surjective, and its inverse is a single¬ valued Lipschitz map from X* to X. к Proof By assumption, — F^" ^(0)=(lm F)" (by proposition 19) is contained in (Dom F)~~ = Dom F, since the domain of F is closed. Let us pick xoeF*~ ^(0), and choose у о e F( — Xo). Since (0, — Xq) belongs to graph (F)“, we deduce that (0, Xo) — (xo, уо)^0. Since F is a X-elliptic process, we deduce that Iл:оIP = I -^0 -0|P^ -Xo -0, -0) = - (xo, Уо)^0 Hence, Xo =0. Therefore, F* " ^(0) is equal to {0}. Since Im F is closed by lemma 24, proposition 20 implies that F is surjective. ■ COROLLARY 26 Any X-elliptic, closed convex process F vahóse domain is X is surjective and F ~ ^ is a single-valued Lipschitz map from X to X"^. к COROLLARY 27 (LAX-MILGRAM COROLLARY) Any X-elliptic continuous linear operator from X to X* is an isomorphism. к It is clear that if F is a closed convex set-valued map from X to У and if A e £F(Xo, X), then FA is still a closed convex set-valued map from Xq to X. If F 6 SF{X, Z), the set-valued map BF is convex, but not necessarily closed. We denote by BF the set-valued map whose graph is the closure of the graph ofFF. PROPOSITION 28 Let X, Y, and Z be reflexive Banach spaces, F a closed convex map from X to Y, and В e S£(Y, Z) a continuous linear operator. If (55) 0 G Int(Im B^ - Dom F^) then the graph of BF is closed and convex. к Proof. We have to prove that the graph of BF is closed. Since graph {BF) = (1 xB) graph (F), we can apply theorem 1.5.5. We easily check that assumption (55) implies that (56) (0, 0) G Int(l X 5)* + b(graph (F))
146 CH. 3, SEC. 4 SET-VALUED MAPS Since graph (F) is convex and closed, it is weakly closed. Hence, theorem 1.5.5 implies that (1x5) graph (f )=graph(5f) is closed. ■ COROLLARY 29 Let V, X, Y, and Z be reflexive Banach spaces, F a closed convex map from X to Y, let A belong to ^(V, X), and В to ^(Y, Z). If (57) Об lnt(Im B*- Dorn F*) then BFA is a closed convex map from X to V. A COROLLARY 30 Let Fi and F2 be two closed convex set-valued mapsfrom a reflexive Banach space X to a reflexive Banach space Y. If (58) 0 6 Int(Dom f f — Dom f |) Then the set-valued map Fi -I- f 2 has a closed convex graph. A Proof Let A e ^(X, XxX) be the map defined by^ у1х:=(д:, x) and BeSF{YxY,Y) the^nap defined by B(yi, y2):=yi +У2- Let F be the set-valued map defined by F(xi, X2)=Fi(xi)xF2(x2}. Then, Fi+F2=BFA. Since ~F*{pu P2)=F*iPi)^PtiP2) and thus, Dom F* = Dom_Ft xDom Ff, we see tlwt assumption (58) implies that 0 6 Int(lm B* - Dom F*). Hence, Fi -I- F2 = BFA is closed and convex by corollary 29. ■ 4. EIGENVALUES OF POSITIVE MAPS WITH CLOSED CONVEX GRAPHS We set (1) S":=-^x6 5"+ (=1 We introduce the following items: a single-valued map / from S" to K" whose components (2) and (3) fi are lower semicontinuous and convex a strict upper semicontinuous convex map G with compact images from 2" to /?+
CH. 3, SEC. 4 EIGANVALUES OF POSITIVE MAPS 147 We posit the following positivity conditions: (4) Vp 6 Z“, 3 X e Z" such that a{G(x), p) > 0 and (5) 3p € Z™ such that V;c e Z", {p, f(x)) > 0 Then we can define the positive number 5 by (6) 1 . f ^:= sup inf 0 pezm xezn (T{G(x), f) PROPOSITION 1 We posit assumptions (2), (3), (4), and (5). Then there exists a solution xel/" to the inclusion (7) 5f{x)eG(x)-Rl and there exists p e Z"' such that the minimax property holds (8) 1 _ (p, fix)) _ (p, f{x)) _ ip,f(x)) S a(G(x),p) xe in pGim (TiG{x),p) Furthermore, when a: e Z” andp>0 satisfy the inclusion (9) pf{x)eG{x)-R\ then p is not larger than S. A Remark We can say that when a pair (p, x) satisfies (9), p is an eigenvalue and x is an eigenvector of G(*)~R\ with respect tof So proposition 1 states the existence of a positive eigenvalue that is the largest of the nonnegative eigenvalues. ■ Proof a. We set F+(x):=/(x) + R+. Since the functions/ are convex, the set-valued map F+ is convex, and, consequently, the set-valued map G—6F+ is also convex. Hence (G-5F+)(Z") = Im(G-5F+) is a convex subset. The continuity properties (2) and (3) imply that (10) (G—¿F+)(Z”) is a closed convex subset.
148 CH. 3, SEC. 4 SET-VALUED MAPS Indeed, let e G{Xk)—SF+{Xk) belong to (G —that converge to some z in For all /? e we deduce that m {p,Zk)^<T{G{xu),p)~d Y. PifiiXk) i = 1 Since S" is compact, a subsequence of elements x*- converges to some x 61". The upper semicontinuity of x-*ff(G(x),p)-S Y Pifiix) « = 1 implies that (p, z) = lim (p, Zk'} fc'-> 00 lim sup (T{G{xk ),p)~5 Y PifiiXk) k'-*ao ' • ' 1 = 1 <:a{G{x),p)-d X PiMx)=a{G{x)-Sf(x),p) i= 1 Since this inequality holds true for all p eRI, we deduce that z e G(x)-Sf (x)-R”l= G(x)-df+(x) b. Next, we prove that Oe(G-SF+X^") If not, we can apply the separation theorem. There exists po e R”' and e > 0 such that for all x e Z", <t(G(x), Po)-<5(Po. fi.x)} + (t{-R"l, p^)^-s This implies that po^Rl and that a{—R"l, po)=0. Since po^O, we have Po:=Po/Er= 1 Po e S’" and we deduce that for all x 6 E" i^{Po,f{x)) £ ^{po,f{x)) e 6'' ff(G(x), Po) S(^{G{x), Po)cr(G(x), Po) dc where c:=supx€ i« /?o)- Therefore, we obtain the contradiction . f <Po. /(^)) £ ^ 1 £ ¿"*eI"<T(G(x),Po) 6c" d 6c
CH. 3, SEC. 4 EIGENVALUES OF POSITIVE MAPS 149 c. Hence, there exists 3c € S" such that OeG(x)-Sf{x)-R"i Consequently, for all p e Z’", we have 0 < iT(G(x), p)-S(p, fix)) and thus, (P^ /W) ^ 1 sup — r^-r pe2m(7(G(x),/?) 0 This implies the minimax equality fix)} ^ . {p, fix)) ^ peint (T(G{x\ p) X e Z»» p e Z»« <^(G(x), p) Since p^{p, f{x))!o(G(x\p) is upper semicontinuous and T” is compact, there exists p G Z"' such that 1 _ <P. fix)) _ {p> fix)} 6 (t{G{x\ p) ;c 6 z« (^{G(x\ p) c. Let us assume that there exist jc € Z” and ^>0 satisfying pf(x) e G(x)—R\. Thus, for all p G Z'”, we would have P{p. f{x))^(j(G{x\p) -= inf sup — r< sup ■ -aT— (5 X 6 ZN p e zm <y{G(x\ p) peim (T(G(xX p) P thatis,■ Remark We can replace Z” by any convex compact subset and replace assumption (3) on G by the weaker assumptions (11) i. The graph of G is convex. ii. The images of G are closed and bounded above, iii. V/7 G Z"‘, x-^(t{G{x\ p) is upper semicontinuous. Proposition 1 implies the existence of a solution x to the inclusion QeG(x)-df(x)-Rri
150 CH. 3, SEC. 4 SET-VALUED MAPS Now, we shall prove that when n>5, for every y e Int R1, there exist a negative number P and a solution x € Z" to the inclusion (12) pyeG(x)-nf{x)-R”, PROPOSITION 2 We posit assumptions (2), (3), (4), and (5) of proposition 1. We associate with any p>d andy 6 Int (/?+) the negative number (13) B:= inf sup 7 r- y) (j(Gix)-pf(x),p) pe Zm xe tn Then there exists x 6 Z”, a solution to inclusion (12). A Proof, a. We begin by checking that )?<0. Indeed, since ^>5, then -<^= inf for some^ 6 Z"* fx d xei»(^iG{xlp) Hence o(G{x)-pf(x\p) te E" iP^y) b. Next, we prove that (14) py&iG-pF^^iy) If not, we can apply the separation theorem. There exist po ^ K” and e>0 such that, for all x e Z", ct(G(x), po)-M{po,/(x))+(t(-R7,Po)^P(po,y)-e We deduce that po e /?+ and sincepo fO, that Po-=Pol E /'oeZ" i = 1 Then, o(G{x)-pfix),po)^„ e sup jt: c -p. r X€lW \P^yy) which is a contradiction.
CH. 3, SEC. 4 EIGENVALUES OF POSITIVE MAPS 151 c. It follows that there exists x e Z" such that liyeG(x)-fif{x)-R\ and that for all p ® S'”, MG(x)-pf{x),p) SO that ß^- (p^y) sup inf 7 r <P xer>ipeZ>« \P>y} Remark We have proved that the minimax property (15) O i„r <^iG{x)-pf(x),p) jj — svip ml y . xe in pe im (/7, y) holds true. Furthermore, since Z" is compact, we deduce that there exists p € Z” such that (16) that is, such that (17) O <^iG{x)-pf{x),p) (T{G{x)-ftf{x),p) P — sup / "^ \ / ^ \ xeX” \P> y) \P, y) I*- (.«• VxeZ", c{G(x)-pf(x)-ßyjHO a(G{x)-pf(x)-ßy, p)=0 We consider the case when the convex map G is defined by G{x):=g{x)— where ^ is a single-valued concave operator. PROPOSITION 3 (KY FAN) Let f and g be single-valued maps from S" to satisfying (18) !• The components fi are convex and lawyer semicontinuous, ii. The components gi are non negative, concave, and upper semicontinuous. Hi. 3p € Z"* such that Vx € Z", {p, f (x)) > 0 . iv. 3 jc e Z” such that gi(x)>0 for i = l,..,,m
152 CH. 3, SEC. 4 SET-VALUED MAPS Then there exist 5 >0, x € Z" and pel."' such that (19) i. df(x)<:g{x) ii. VxeZ", {g{x)-df{x),p)^0 in. S(p,f{x)) = {p,g{x)) Furthermore, for all p> 5 and for all y e Int{R"l), there exist /3o<0 and x el" such that (20) Poy^g{x)-pf(x) When g and / are linear operators from R" to we obtain the following solution to a problem posed by von Neumann. PROPOSITION 4 Let f :=(fi) andg'-=(g{) be two matrices satisfying (21) i. gi>0 forallUj (g is nonnegative) ii. Vi = l, ...,w, Y g{>0 j=i m I Hi. y/=l, ...,n, Z fi>0 1 = 1 Then there exist x eZ”,p e'L”' and5>0 such that (22) Sfx^gx bf*p>g*p • b(p, fx) = (p,gx) Furthermore,for all p>3 andfor all e Int (/?7), there exists x e R"+ such that (23) pfx-gx^y A When the dimensions n and m are equal, a boundary condition on/ implies that the solutions 3c e Z" and 3c e Z" to the inclusions (24) dfix)eg{x)-R\ and Py eg{x)-pf{x)-R\ are actually solutions to the equations (25) df{x)=g{x) and Py=g(x)-pf{x)
СН. 3, SEC. 4 EIGENVALUES OF POSITIVE MAPS 153 THEOREM 5 Letf be a single-valued map from Z" to R" satisfying the properties {I. The components f of f are convex and lower semicontinuous. ii. 3^6lnt(^'V) such that VxeZ", <^,/(x)>>0 ii. when Xi=0, then f^x)^0 {boundary condition) and let g be a singlevalued map from Г" to R” satisfying m\ I *’ components gi of g are concave and upper semicontinuous. ' ^ lii. VxeZ", = 0i(^)>O a. Let 5>0 be defined by (28) 1 f.— sup inf {p, fix)) pe (P,gix)) Then there exist x € Int R"+ andp e Int R"+ such that (29) i. 3f{x)=g{x) Ii. Vx 6 S", (gf(x)-¿/(x),p)^0 b. Let p>5 and ^ e Int {R"+) be given and P<0 defined by (p, gix)-pf(x)) (30) )3:= inf sup pe in xG in (p,y) Then there exists a solution x e Int R\ to the equation (31) Py=gix)-pf(x) A Proof, a. We denote by e^ theyth element of the canonical basis: e]=\ and e{=0 when k j. The boundary condition (26, iii) on/ implies that (32) VA:^y, Me^)^0 and joined with the positivity condition (26, ii) on f it implies that (33) y/=l,...,«, fy)>0 because there exists p e Z"nlnt (/?+) such that pne^)+ Z P%ie^)=(p, fie^))>0 kfj
154 CH. 3, SEC. 4 SET-VALUED MAPS The conclusions of theorem 5 actually follow from these two properties (32) and (33). b. By proposition 3, there exist 3c e X” and p 6 X" satisfying (34) i. Sf(x)^g{x) ii. Vx 6 Z”, (p, g{x) - df (x)} ^ 0 iii. {p,g{5c)-df{x))=0 We check first that p belongs to Int By taking x:=e^ in inequality (34, ii) and using positivity assumption (27, ii), we deduce that kfj Then inequalities (32) and (33) imply that p^ is strictly positive. Let z=g{x)-6f(x) belong to R\ by (34, i). Then property (34, iii) implies that (p, z) =0. Since p belongs to Int (Л+), then z is equal to zero. Therefore, for all i = 1,..., «, Sfi{x)=^gi{x), and since gi{x) is positive,/(x) is also positive for all i = l,. . . , n. We thus deduce from the boundary condition (26, iii) that x,>0 for all / = 1,..., n. c. The proof of the second statement of theorem 5, which is analogous to the proof of the first one, is left as an exercise. ■ The identity mapping obviously satisfies assumption (26) required on the map f. In this case, we state the following corollary on eigenvalues of positive concave operators. COROLLARY 6 Let g be a single-valued map from Z” to Int R\ whose components are concave and upper semicontinuous. a. Let 5>0be defined by (35) 1 . f {p>x) sup inf ЕП * 6 Z" {p, g(x)) There exist x € Int R\ andp e Int R”+ such that (36) b. Let p>3 andy 6 Int. Then there exist j?<0 and.x 6 Int R\ satisfying 5x=g(x) ii. Vx e X", (p(x)—5x,p)<0 (37) Py=g(x)-px A When (7 is a positive matrix, we obtain the Perron-Frobenius theorem.
CH. 3, SEC. 4 EIGENVALUES OF POSITIVE MAPS 155 THEOREM 7 Let G be a positive matrix, a. Then G has a positive eigenvalue S and a corresponding eigenvector x with positive components, b. d is the only eigenvalue of G for which there corresponds an eigenvector xer\ c. <5 is larger than or equal to the absolute value of any other eigenvalue of G, d. The matrix p — G is invertible and {p — G)~^ is positive if and only if p>d, k Proof, a. By corollary 6, we know that there exist 5 > 0, ic 6 Int R\ and p e Int R\ such that 5x = G(x) and G'^p — Sp^O, Actually, G*p — dp=0 because (3c, G*p — Sp) = (p, G*x — dx) =0 and because the components of 3c are strictly positive. b. Let X e Z” and p satisfy px = G{x), We deduce that i^(p,x) = (p, G{x)} = {G*p, x) =5{p, x) Since (p, x) is strictly positive because x belongs to Z" and p belongs to Int R"+, we can divide by (p, x) and observe that p=5. Hence, the second statement is proved. c. Let X be an eigenvalue of G and let z e R” be an associated eigenvalue. Equalities (38) imply that (39) '^z, = Y gUj j=i (j = \M lz.l=s Z gi\^j\ j=i Let \z\ denote the vector of components \zj\. Then \X\ \z\ belongs to G\z\ — R\ and thus \X\^d, d. We know that when p>3, the matrix p — G is invertible because S is the largest eigenvalue of G, We know that for all;; e Int R\, the solution {p — G)~ ^y belongs to Int R\ by the second statement of corollary 6. This implies that {p — G)~^ is positive. Conversely, assume that p — G is invertible and {p — G)~^ is positive. We cannot have the inequality p^S, because we would deduce that px^Sx=Gx and thus that — x = {p — G)~^(Gx — px) eR\ since p — G is invertible, (p — G) ^ is positive, and Gx — px is a positive vector. Hence,//>5. ■
156 CH. 3, SEC. 4 SET-VALUED MAPS We now recall the definition of a M-matrix M = {m{): It is a matrix that satis- V/^7, mi^O fies (40) and (41) Vx 6 Z", 3 ^ e Z" such that Mx) > 0 Let 6>max, = i,„„„ m\. Then it is clear that condition (40) is equivalent to (42) Vx e Z”, bx e Mx -I- R\ We shall extend the concept of M matrix to the concept of M-convex operator in the following way. DEFINITION 8 Let H be a convex processfrom R\ to R*\ We say that H is a M-convex process if it satisfies both properties (43) and (44): (43) and ^b eR such that Vx e /?+, bxe H(x) + R\ (44) Vx 6 Z”, 3^ 6 Int R\ such that inf 3^) >0 A yeH(x) PROPOSITION 9 Let H be a strict upper semicontinuous, convex process with compact images from R\ to R”. If H is a M-convex process, then (45) e Int , 3 X e R\ such that y e H(x) + R\ Proof Let p be larger than the number b involved in assumption (43). We introduce the map G from Z” to R!' defined by (46) Vx e Z", G(x) :=px—H(x) and we take F to be the identity mapping. Therefore, we can use proposition 1: There exist 5 > 0 and x e Z" satisfying (47) ¿X 6 G(x) — R\ =px — H{x) — R+ We make use of assumption (44): Let q e Int R\ such that infyeH(x) (q, y}>0.
СН. 3, SEC. 4 EIGENVALUES OF POSITIVE MAPS 157 Since {ц — д)х belongs to H{x)+R\ by (44), we deduce that inf (q,y)>0 yeH{x) then fi — 5>0 because (q, x) is strictly positive. Now, we apply proposition 2. We can associate to any ;; e Int elements jc € Z” and P<0 such that Py e G(x) -iix-R\ = - H(x) - R\ We divide this inclusion by — p, and we set x:=x/( — P). Then;; e H{x)-\-R\. ■ We consider now the case when H{x)=h{x) + R\ where his a. positively homogeneous concave operator. DEFINITION 10 Let h be a single-valued map from R\ to y^^hose components satisfy (48) Vi = l,..., hi is convex, lower semicontinuous, and positively homogeneous We say that it is a M-operator if (49) eR such that Vx € R\, bxi^hix) and (50) Vx 6 3 ^ e Z” such that (q, h(x)} >0 A THEOREM 11 Let h be a single-valued map from R\ to R!^ satisfying properties (48), (49), and (50) . Then h maps Int R\ onto itself (51) V;;eInt/?+, Зд:е1т Л+ such that hx=y A When A is a matrix, we obtain the characterization of M matrices. THEOREM 12 Let h be a matrix satisfying (40). Then the following statements are equivalent: a. his a M matrix. b. h is invertible, and h~^ is positive. c. h* is invertible, and h~^ is positive. A
158 CH. 3, SEC. 4 SET-VALUED MAPS Proof. The implication a=i>b follows from the Perron-Frobenius theorem. The implication b=>c is obvious. It remains to check that c implies a. Let p e Int R\ be the solution to h*p = \ where H is the vector of components 1. Then, for all x € Z", {p, hx) = (h*p, x)=Y. Xi=i i=l Therefore, property (41) is satisfied and A is a M matrix.
CHAPTER 4 Convex Analysis and Optimization The main objective of this chapter is the study of convex minimization problems (*) W{y)\= inf V{x,y) xeX depending on a parameter y when the function V is convex. Besides sufficient conditions implying the existence of solutions Xy to these minimization problems, we are looking for equations or inclusions that char¬ acterize those solutions (variational principles), and we are studying differenti¬ able properties of the “marginal function” W defined by (*) and the set of mini- mizers Xy of V{x, y) with respect to the parameter y. Characterization of solutions to minimization problems as solutions to equations (or inclusions) is a very old problem, since Fermat discovered the famous rule if X minimizes C/, then i/'(3c)=0 for algebraic functions in 1637, and later in 1684, Leibniz extended it to differ¬ entiable functions. This “Fermat rule” is still the object of recent works, when the function U is no longer differentiable. Indeed, usual differentiability is not stable for the pointwise supremum, so that in the framework of optimization and game theory, we very naturally meet nondifferentiable functions. Before considering the most general case in Chapter 7, we shall restrict our attention to the point- wise suprema of affine continuous functions, which are convex, lower semi- continuous functions (and, as we shall see, which make up the whole class of convex, lower semicontinuous functions). Consider the simplest example: We minimize the convex function x-^\xl It achieves its minimum at x=0, and we are unable to write the Fermat rule, because this function is not differentiable at this point. But x-*\x\ is the supre¬ mum of the affine functions x-^ax-\-b with ae\_ — l, +1] and 6^0. When x is negative, there is only one such affine function passing through (x, \x\\ that is, its tangent x-^ —x, whose slope is — 1: It is the gradient of the function at this 159
160 CH. 4 CONVEX ANALYSIS AND OPTIMIZATION point. When X is positive, there is still a smaller affine function passing through (x, |a:|); that is, its tangent x-^x, whose slope is 1. When a:=0, there is no tangent but “subtangents”; that is, the smaller affine functions x^ax, a e [-1, +1], passing through (0, 0). The revolutionary idea was to suggest taking the set [—1, +1] of all the gradients of these affine functions—called subgradients— as a candidate for replacing the missing concept of gradient; this set is called the subdifferential of x-^\x\ at zero. The Fermat rule still holds true in this case, because zero belongs to the subdifferential [—1, +1]. This idea works in the general case; the price to pay to hold the Fermat rule true for convex minimiza¬ tion problems was to accept dealing with “set-valued gradients”—so to speak— of convex functions. The adaptation of the Fermat rule and a decent “subdifferentiable calculus” made convex analysis more and more indispensable not only for studying convex programs (convex minimization problems in finite dimensional spaces) but also for problems in calculus of variations and optimal control. The ideas of Euler, Lagrange, and Hamilton can be adapted to nonsmooth problems as long as they are convex. Convex analysis also deals with set-valued maps with closed convex graph and convex sets. It is customary to begin the presentation of convex analysis with the study of convex function and then deduce the properties of normal cones to subsets, and so on. We shall follow the inverse route: In the first section, we begin with the study of tangent and normal cones to convex subsets, then give some examples and devise a set of formulas allowing the characterization of tangent and normal cones. In the second section, we adapt to the case of set¬ valued maps the ancient geometrical concept of the derivative of a real-valued function, whose graph is the tangent to the graph of the function. When F is a set-valued map from Z to 7 with a closed convex graph and (xoy yo) belongs to the graph of F, we define the derivative of F at (xo, ;^o) as the closed convex process DF{xo, >^o) from X to 7, whose graph is the tangent cone to the graph of F at (xo, yo)- It possesses most of the virtues we expect from a derivative. Its transpose DF(xq, a closed convex process from 7* to A"*, called the codifferential of F at {xq, yo), naturally plays an important role. When F is a single-valued map from Z to Fu {+ oo} and we are interested in only minimization properties where the order relation of R plays a crucial role, we associate with V the set-valued map V+ from X to R defined by \4x):=V(x) + R+ if F(x)< +00, V+(x):=j2T if F(x)= + oo whose graph is the epigraph of V. Hence, V+ has a closed convex graph if and only if Fis lower semicontinuous and convex. In section 3, we observe that the graph of the derivative Z)V+(a:, V{x)) of the set-valued map V+ at {x, V{x)) is the epigraph of a function we shall denote by Z)+ V(x), called the epiderivative, defined by D+V{x){v) = \im inf V{x^hu)-V(x) /l-0^
CH. 4 CONVEX ANALYSIS AND OPTIMIZATION 161 We also observe that the transpose D\+(x, K(x))*, a closed convex process from Rto X*, satisfies [0 when X < 0 where dV(x):=D\4x, V(x)ni) is the subdifferential of the function V at x. In summary, we present a unified treatment of convex analysis; starting with the definition and properties of tangent cones to convex sets, deducing the definition and the properties of derivatives to set-valued maps with convex graphs, which are convex processes, and then the definition and properties of the epiderivative of a convex function. But there is more to that. In the fourth section, where we consider the specific class of perturbations of a convex minimization problem V/7 6 X*, W(p)\= inf \V{x)-(p, x>] xeX we observe that the function V* defined by K*(p):= — W(p) is also a convex, lower semicontinuous function on the dual of X, which is called the conjugate function: K*07):=sup[</?, x)-K(x)] xeX We define the biconjugate function K** by supJ</7, x)-K»] peX* which is also convex and lower semicontinuous. The condition F=F**, which is necessary for V to be convex and lower semicontinuous, happens to be also sufficient. This shows that the correspond¬ ence is a one to one correspondence between convex, lower semi¬ continuous functions on X and its dual. This duality result will play an important role. We also observe that this result implies that any lower semicontinuous, convex function F can be obtained as the supremum of the continuous affine functions (/?, x) — V*{p) as p ranges over X*. It is easy to check that the subdifferential dV(x):= {p e X^\{p, x) = F(x)+ F»} is the set of gradients p of the smaller affine functions (/?, x) — F*(p) passing
162 CH. 4 CONVEX ANALYSIS AND OPTIMIZATION through {x, V{x)). The symmetry of this formula implies at once that p e dV(x)^x e dV*{p) that is, the set-valued mapp-^dV^(p) is the inverse of the map x-*dV(x) dV*=(dV)-^ This formula shows that the conjugate function plays the same role in the convex framework as the Legendre transform in the framework of smooth analysis; actually, they coincide when V is both convex and smooth. Now, if we return to the framework of the minimization problem v:= inf V{x) xeX when F is a convex, lower semicontinuous function, we observe that (a) inf V{x)=-V*(0) xeX (b) X minimizes V on edV{x) (c) The set of minimizers of V is dV*{0). Statement (b) is the Fermat rule, and statement (c) allows us to characterize the set of minimizers as the subdifferential of V* at zero. Therefore, conditions on the conjugate function F* (implying that F* is subdifferentiable at zero) are sufficient conditions for the existence of a minimizer of F This is the essence of the duality theory for the convex minimization problems. In Section 5, we consider a family of optimization problems W{y):= inf V(x,y) xeX where F is a proper lower semicontinuous, convex function from X xY to /?u{ + oo}. We associate with these families of problems the function /i* defined on X* xY by h*(p, y):= sup [_{p, y) - V{x,;;)] xeX which is convex with respect to p and concave with respect to y. Assume that X is a solution to the minimization problem W{y). We shall prove that the follow¬ ing statements are equivalent. (a) qedWiy) (b) {0,q)^dV{x,y) (c) 3c € j3) and q^dy{-h*my)
CH. 4 CONVEX ANALYSIS AND OPTIMIZATION 163 The elements q given by one of these equivalent statements are called Lagrange multipliers of the minimization problems. Formula (c) shows how the pairs (x, q) of solutions and Lagrange multipliers evolve with respect to the parameter y: They depend on the regularity of the set-valued map y^d^*{0,y)xdy{-h*)(0,y) We shall consider the particular case when the minimization problems are of the form W(y):= inf V(x) xe F~Hy) where U: X-^/?u{ + oo} is a proper lower semicontinuous function and F is a set-valued map from X to Y with a closed convex graph. Let 3c 6 F " be a solution to the problem W{y). Then we shall prove that the condition implies the formula 0 e Int (Dom F— Dom U) d W{y) =D(F~^ )(y, xYd U(x) between the subdifferentials dW and 5 i7 and the codifferential of the set-valued map F defining the constraints. We consider in the sixth section more specific problems of the form W{y):= inf iU(x)~ (p, x) + V(Ax-\-y)'\ xeX where U: ^'-►Ful + oo} and V: y-^Fu{-foo} are proper lower semicon¬ tinuous, convex functions and ^ is a continuous linear operator from X to K We shall associate to it its “dual problem” W%):= inf \_U*i-A*q+p)+V*(q)-(q,y}'\ qeY* We shall prove that the condition p G Int (A* Dom V* -f- Dom U*) on the dual problem implies that there exist solutions x to the problem W(y) and that the condition y 6 Int (Dom V—A Dom U) on the initial problem implies the existence of solutions q to the dual problem
164 CH. 4 CONVEX ANALYSIS AND OPTIMIZATION (called Lagrange multipliers of the initial problem). When both assumptions are satisfied, we prove that the set of solutions to W{y) is the subdifferential dW®{p) and that the set of solutions to the dual problem W®{p) is the subdif¬ ferential dW{y): A very interesting phenomenon for economists of the marginal school. The solutions x and q to the minimization problems W{y) and W®(p) are solutions to the system of inclusions pedU{x)+A*q y e -Ax-ydV*(q) (abstract Hamiltonian system) By eliminating q in these inclusions, we observe that the solutions x to W{y) are solutions to the inclusions pedU(x)-yA*dV(Ax+y) (abstract Euler-Lagrange equation) and by eliminating x, we note that the solutions ^ to W*(p) are solutions to the inclusion y edV*{q)—AdU*(p—A*q) (dual Euler-Lagrange equation) We shall define in Chapter 7 (on nonsmooth analysis) the concepts of genera¬ lized second derivatives, derivatives of set-valued maps and devise inverse function theorems for set-valued maps. Applied to our minimization problems, they imply that if the matrix of closed convex processes ¡d^U A* \ \-y4 d^V*) from AT X y* to X* X y is surjective, then the solutions (ic, q) of the problems W(y) and W®{p) depend in a Lipschitz manner on the parameters y and p, and the marginal variations dx, 5q of the solutions depend on the marginal varia¬ tions of the parameters through the formula (d^U A* VU^P In the last section, we consider minimization problems of the form v.= inf L{x, Ax) xsX where L is a proper lower semlcontinuous function from Xxyto/iul-l-oo} and A belongs to if (X, y). We associate to it its dual problem ’*:= 'mî^IÎ{—A*q,q) qey
CH. 4 CONVEX ANALYSIS AND OPTIMIZATION 165 and the nonnegative function d defined on X x 7* by s4(x, q)-.=L{x, Ax) + I^{-A*q, q) We observe that xeX is a solution to v and q eY* is a solution to v* and t>+t;*=0 if and only if d(x, q)=0. We associate to the function L, which plays the role of a Lagrangian in the calculus of variations, the function H defined on x 7* by H{x, q):= sup \_{q, y) - L(x, y)] y^Y which plays the role of an Hamiltonian, which is concave with respect to x and convex with respect to q. We shall observe that the solutions (3c, q) of the mini¬ mization problems v and i;* are the solutions to the Hamiltonian inclusion -A^edJ(-H){x, q) and Axed^H(x, q) and that the solutions 3c to the minimization problem v are solutions to the Euler-Lagrange inclusion 0 G (1 ©yi*)5L(3c, Ax) where 1 ®A* is the continuous linear operator from x y to X* defined by (l©^*)(p, q):=p + A*q If we assume that the Hamiltonian is convex and, thus, the Lagrangian is concave with respect to x and convex with respect to y, we can still write the Hamiltonian inclusion in the form (A% Ax)edH{x,q) and the Euler-Lagrange inclusion in the form 0 e —dx{ — L){x, Ax)-\-A*dyL(x, Ax) Since the Lagrangian is no longer convex, there is no hope of solving the mini¬ mization problem V. But we shall prove that solutions (3c, q) to the Hamiltonian inclusion are the solutions to the equation i,q)=0 where ^ is the nonnegative function defined on X xY* by ^(x, q):={H(x, q)~ (q, Ax)) + (H*{A% Ax)-(A% x>)
166 CH. 4, SEC. 1 CONVEX ANALYSIS AND OPTIMIZATION We observe that the solutions (3c, q) of the minimization problems w:= inf q)- {q. Ax)) (x,q)GXx.Y and w*:= inf (//*(/1*^, Ax)-(A% x}) (x,q)eXxY* are also solutions to the Hamiltonian inclusion. The problem w plays the role of the /east action principle, whereas problem w* can be said to be the dual least action principle» We shall apply these ideas to solving problems in calculus of variations in the last chapter. 1. TANGENT AND NORMAL CONES TO CONVEX SUBSETS Let Z be a normed space. DEFINITION 1 Let K^X be a convex subset andxeK. We denote by (1) Sk(x):= U /i>0 U the cone spanned by K — x and by (2) Tk(x):=ci(^UJ(^-^)) its closure. Tic{x) is called the tangent cone to K at x. (3) VveSKix), 3h>0 such that Vfe[0,/;], x + tveK
CH. 4, SEC. 1 TANGENT AND NORMAL CONES TO CONVEX SUBSETS 167 Indeed x + ii;=^l- ^^x+^(x+hv) is a convex combination of elements belonging to the convex set K. We can interpret this remark by saying that when v belongs to SkW, the beginning of the “half-curve” starting from x defined by Vie[0,<j){t)=x+tv is contained in K. Hence, the cone 5k(x) is the set of directions v such that the half-lines x+R+v intersect K. This subsumes the idea lying behind the concept of tangent. We point out that v belongs to the tangent cone 7i(x) if and only if (4) Ve>0, 3uev+eB, 3k>0 such that x+kueK or, equivalently, there exist sequences of vectors u„eX converging to U (5) and of numbers h„>0 such that x+h„u„ e K for all n. Condition (3) implies that if condition (4) is satisfied for some A:>0, it is also satisfied for all h<k. It follows that (6) n U (1(K-x)+sb) £>0 a>0 he ]0,a] / We shall see in Chapter 7 on nonsmooth analysis that the right-hand side of this formula defines the contingent cone to K at x. We observe that (7) and that (8) Vx€K, Tk{x) = Tk(x) if X 6 Int (K), then 5k(x)=T«(x)=X since K—x contains a neighborhood of the origin. It is also clear that Tk(x) is contained in the closed vector subspace M{K)=c\[{aK—PK}a,pe r] spanned by K. We note that (9) X <= X -t- 5k(x) <= X -(- Tk{x)
168 CH. 4, SEC. 1 CONVEX ANALYSIS AND OPTIMIZATION PROPOSITION 2 The cones Sk{x) and Tk(x) are convex. A Proof. Indeed, if Vi and V2 belong to Sk{v\ then jc+/?,y, eK for / = 1, 2; let /?=min (hu hi)- Then by (3), x + hvi e K for / = 1, 2. Hence, x-hh{oiVi-\-{l—(x)v2)eK when a 6 [0,1] Since the closure of a convex cone is still a convex cone, we have proved the proposition. ■ A normal to a smooth submanifold at a point x is any vector orthogonal to the tangent space. In the case of tangent cones to convex subsets, the concept of an orthogonal subspace to the tangent space has to be replaced by the negative polar cone to the tangent cone, which will be called the normal cone. The introduction of these two concepts, tangent and normal cones, allows the use of duality relations. DEFINITION 3 Let K be a nonempty convex subset of X. The normal cone Nk{x) to K at x is the negative polar cone to the tangent cone. A PROPOSITION 4 The normal cone N^^x) is equal to (10) Nk{xY={p&X* such that {p,x)=m&\{{p,y)\y eK):=aKip)]. A Note that Tk(x)=Nk(x)~, since Tk(x) is a closed convex cone. Proof. If/» 6 Tn{xy, then (p, y—x)^0 for all y e K, since v=y—x 6 T^ix) when yeK. Conversely, let p satisfy (p, x)=(Tii{p) and v = \im„^„X„(y„ — x) e Tk{x), where A„^0 and y„eK. Hence, (p, y)<0, since (p, A„(y„—x))=A„(/>, y„—x:)<0 for all n>0. ■ PROPOSITION 5 Assume that X is a Hilbert space. Let be the projection of best approximation onto a closed convex subset K. Then n^^{x)=x+ Nk{x) and v belong to Tk{x) if and only if (y-x,v)<0 fyen^Hx)
CH. 4, SEC. 1 TANGENT AND NORMAL CONES TO CONVEX SUBSETS 169 PROPOSITION 6 Let K be a closed subset of a normed space x. Then (11) Nk(x) has a closed graph. If X is finite dimensional, then (12) Tk(x) is lower semicontinuous. k Proof a. Let (x„, p„) be a sequence of elements of the graph of A^(*) converging to {x, p). For all yeK,v/Q have (p„, y)^(p„, Hence, letting w-^oo, we deduce that (p, y)^{p, x). Thus, (x,p) belongs to the graph of N{*). b. The second statement is equivalent to the first by proposition 3.1.18. ■ Remark We can prove that Tk(‘) is lower semicontinuous when A" is a Hilbert space. PROPOSITION 7 If K has a nonempty interior, then for all x e K, the interior of the tangent cone is nonempty and spanned by\n\K — x (13) Int Шх)= и l(Int К-х) /1>0 ^ Furthermore, the set-valued map x-^lnt Tk{x) has an open graph.
170 CH. 4, SEC.l CONVEX ANALYSIS AND OPTIMIZATION Proof, a. The cone (J(i>o K —x) is open, being a union of open subsets. Hence, it is contained in Int Tk{x). Since Int TK(x)=Int 5k(x), it suffices to prove that any v € Int Sic(x) belongs to some (l/A)(Int K — x). Let >j>0 be such that v+rjBcS,c{x). If x+u belongs to Int K, the proof is finished. If not, let xo belong to Int K and let us set Vo'=Xo — x. Hence, —(>//||uoll)fo belongs to 5k(x), and, consequently, there exists /i>0 such that x+/i{t)-(>//||i;oll)t^o) belongs to K. Let us set c^-=hr]/(hri + \\vo\\). We observe that x+(l -a)/it)=axo + (l -a)| x+h ( v- jj—Vq Since Xo belongs to Int K and x+/i(a—(>//||yo||)uo) belongs to K and since a belongs to ]0, 1[, then x + (l — a)/»; also belongs to the interior of K. This proves that v belongs to 1 {i-<x)h (IntK-x) b. Let uoelnt Tk(xo). Therefore, i;o 6(l//io)(Int K-xo) for some ho>0; hence, there exists a > 0 such that Xo+/iot)o + e-8=Xo+/io(i>o+7-fi )<=Int K V ho Take X e Xo + e/2B and v evo + e/2hoB. Then x+hov exo+Aot>o + e^<=Int K and, therefore, v e Int Tk-(x). Hence, the graph of x->Int Tk(x) is open. We now list a few examples. PROPOSITION 8 Let B be the unit ball of a Hilbert space and x belong to B. Then (14) I*’ xelnt5, and Tb{x)={x}~ if |W| = 1 ^ [ii. 1Vb(x) = {0} if xeintfi, and Nb{x)=R^x if ||x|| = l A Proof We take ||x|| = 1. Then/? 6 Nk(x) if and only if 1|/?||,^=sup,,e b {p, y) = <p, x). By the Cauchy-Schwarz inequality, this is equivalent \o p=Xx with A>0. By polarity, we deduce the formula for the tangent cone. ■
CH. 4, SEC. 1 TANGENT AND NORMAL CONES TO CONVEX SUBSETS 171 PROPOSITION 9 Let KcX be a closed convex cone. Then Nk{x) = K~ r\{xYy and thus (15) V e Tk(x) if and only if (p, i;)^0 for all p eK~ satisfying {p, x) =0 If K is a closed subspace, then Tic(x) = K and Nk{x) = K^. A Proof It is clear that K~ n{x}'^ is contained in Nk{x). Conversely, if peNK(x), then {p, x)=max3,e/c (a y)> Since X is a cone, we deduce that {p, x)=0 p e K~. ■ PROPOSITION 10 Let A 6 SP{X, Y) andK=A~^{y)be an affine subspace. Then if Ax =y. (16) TA-Hy)ix) = ^^r A Proof a. If i;€Ker A, then v-{-x eA~^(y)=K, and thus v = v-i-x—x belongs to Tk{x). b. Conversely, if i;=lim„_^ v„ e Tk(x), where v„=X„{x„—x) with x„eK and 2„>0, then Vn 6 Ker A and thus v e Ker A. PROPOSITION 11 Let and Then (17) and Z •^«•=1 i=l I{x):={i = l,..., A2|A:i=0} V e Tru^ (x) if and only if Vi^O for all i e /(x) (18) veTzn(x) if and only if Vi^O for all iel(x) and Z ^ /= 1 Proof, a. If K = R\, the first statement follows from proposition 9. Indeed, lip e satisfies Z Pi^i= Z Pi^i=^ i=l i^Hx)
172 CH. 4, SEC. 1 CONVEX ANALYSIS AND OPTIMIZATION then pi=0 whenever i $ 7(x); hence, v e Tnn^{x) if Y,ienx)PiVi<:0 for all peR"+, that is, if and only if (17) holds. b. Let V satisfy if i e /(x)and Xi'=i ^1=0. If ^¡=0 for all i i 7(x), then u=0. If not, let X= min 7^>0 ¡tl(x) I I’ll VifO Therefore, x+At) e S" since x,+At)i=A(t)f^0 if i 6 7(x) X/ + At),- > X,-—A| t),-| > X( — Xf=0 if t i 7(x) and Hence, f] (x,-+At),)=1+0=1 t) 6 y (2"—x) € ri«i(x) c. If t)=A(y—x), where y e 2" and A>0, then we deduce that Vi=A(ji - X,)=Ay^ 0 when i e 7(x) and Ya=i Vi=0. Therefore (J A(2"-x)cji) 6 TRi^.(x) ^ D,-=0 A>0 (. 1 = 1 Since the latter subset is closed, we deduce that t Vi=0 i = l Ti«(x)=<t) 6 Tr»^(x) The concept of tangent and normal cones is useful only if enough formulas allow us to characterize (or compute) tangent cones. For instance, we would like to know tangent cones to products, intersections, direct or inverse images by linear operators of sets whose tangent cones are known. We begin by stating obvious properties. All the subsets involved are naturally assumed to be convex.
CH. 4, SEC. 1 TANGENT AND NORMAL CONES TO CONVEX SUBSETS 173 PROPOSITION 12 a. If X € K <= L, then (19) Tk(x)<=Tl{x) and b. Let K:=f]i^j Ki and let J{x)={i\xt i Int K,} Then (20) 7k(a:)c TKi(x) i e J{x) PROPOSITION 13 Let X:=n"=i Ki andx={xi,..., x„) e K. Then (21) r^(x)= fl Tk,(x,) and Nic(x)=f\ NK^ixd i =1 1 = 1 Proof It is obvious that Tj5(3c)cn^=i ^«.(^f)- Conversely, let vieTK,(Xi) for i = 1 n. Then there exist sequences of elements vf converging to y, and of hi>0 such that Xi+hfvf 6 Ki for all i'=l,. . ., n. We set /i'‘:=min, = i „hf>0. Since the subsets Ki are convex, x'+hit^ e Ki for all i, that is, x+hv'‘eK. Hence, V e 7g(3c). We deduce the formula on normal cones by polarity. ■ PROPOSITION 14 Let A € SC(X, y) and K<=X. Then fxeK, TAiK)(Ax)=c\(ATK(x)) NMK){x)=A*-^{Ndx)) (22) and (23) Proof Since {p. Ax) =max (p. Ay) =max(A*p, y) = {A*p, x) yeK yeK we obtain the formula for the normal cones and deduce it by polarity for tangent cones. B COROLLARY 15 Let K and L be two closed convex subsets, x e K andy eL. Then (24) TK+dx+y)=d(TK{x)+TL(y)) and NK^dx+y)=NK,(x)r\Ndy) A
174 CH. 4, SEC. 1 CONVEX ANALYSIS AND OPTIMIZATION THEOREM 16 Let A e S^(X, Y) be a continuous linear operator and let LczX and M<=Y be closed convex sets. We set (25) K:={xeL\AxeM}=LnA-HM) Assume that K ^0, that is, 0 e A(L)—M, and choose x in K. The inclusions TK{x)^Tdx)nA-\T,,(Ax)) and Ndx)+A*NM{Ax)<^Ndx) are always true. If we assume that (26) Oe\nt(A(L)-M) then the equalities (27) TK{x)=TUx)nA-\Wx)) and Nk{x) = Nl{x)^ A^Nm(Ax) hold true. ^ Proof, a. The first inclusion is obvious. The equality follows from the closed graph theorem (see theorem 3.3.1.). Let xq^K and Vq eTtixojn A ~ ^ Tm{Axq). There exist sequences of elements v„eX and € 7 converging to Vo and Avo, respectively such that for all n, xo+hnVneL and Axo + hlu„ eM. We set h„:=mm{hn, h„, 1)>0. Since L and Mare convex, we deduce that (28) for all n, Xn-=Xo-\-h„Vn eL dfnAyn\=Axo+hnUn eM The theorem is proved if u„=Av„ for an infinite subset of indices. If not, we apply corollary 3.3.4. to the set-valued map F defined from L to 7 by F{xy.=Ax—M We take j^o=0 and Xq e F~^(0) = K. By assumption (26), F{L)= Int (A{L)—K). Hence, there exists y>0 such that Wy eyo + yB Vx€L, i/(x, f-H;^))<^i/(J, f(x))(||xo-x|| + I) We take y=0 and x=xo+h„Vo. So ||xo —x||=/i„||poII. ^^(0. F{xo + h„Vo))= d(Axo+h„Avo, M)^d(Axo+h„u„, M)+h,\\Avo — u,\\=hJ\Avo-u,\\. Therefore, di,Xo+h„Vo, ^ ^(0^ T(;co+/i«t^o))(ll^--^oll + l) hn yh„ =^^||^Po-M„||(/!„||Poll + l)
CH. 4, SEC. 1 TANGENT AND NORMAL CONES TO CONVEX SUBSETS 175 Hence, since u„ converges to Avq, we deduce that . , d(xo+h„Vo, K) inf k„>0 This means that Vo e Tk(xo). b. By polarity, we deduce that h„ -=0 (29) Nk(x)=cl (Ndx)+A * Nm{Ax)) Assumption (26) also implies that Ndx)+A*Nm{Ax) is closed. For that purpose, let + be a sequence converging tor, wherepn e Ndx)&ndq„ e Nm{Ax). We prove that Vu € i; sup (q„, v}<+co n Indeed, there exist 2>0,yeL, and zeM such that v=^z—Ay) by assumption (26). Hence, (q„, v) =X(qn, z-Ay) =H(q„, z) - (A*q„, =H(Pn, y} + <?m z) - <r„, y)) < ({Pn, x) + {q„, Ax) - (r„, y)) (since p„ e Ndx) and q„ 6 Nm{Ax)) =A(r„, x-y)<2||r„|| ||x-j||< +00 since the converging sequence r„ is bounded. Therefore, the sequence of elements q„ is bounded by the uniform boundedness theorem and thus relatively compact in the weak-* topology. Some subsequence (again denoted by) q„ converges to ^ € Nm(Ax). Hence, p„=r„~A*q„ converges top=r—A*q e NJix), and thus r=p + A*q belongs to Ndx)+A*Nm{x). ■ COROLLARY 17 If Lis a closed subset of X and if A e I£(X, 7), then for any y 6 Int A(L) and X eLnA~'^(y), we have (30) TtnA-'(rt(.x)=Tdx)nKer A Nlha-'(rt(A:)=Ndx) + Im ^* COROLLARY 18 Let PcX be a closed convex cone and po^P~ such that the subset K={xe P/(po, x) = -1} not empty. Then (31) veTdx)if andonlyif (pq,v)=0 and (p,v)^0 for all p eP~ sat ikying {p, x) =0
176 CH. 4, SEC. 1 CONVEX ANALYSIS AND OPTIMIZATION and (32) TV*(x)=Ker;7o + (P-n{;c}") A By taking X=Y and A to be the identity, we obtain corollary 19. COROLLARY 19 Let K and L be two closed convex subsets of X. If (33) then (34) Oeint {K-D fxeKnL, Tk ni,(x)=Tk(x) n Tl{x) The equality between the tangent cone of a finite intersection and the intersection of tangent cones holds true under the following assumptions. PROPOSITION 20 Let K:= j Ki be the intersection of n closed convex subsets: we assume that (35) 3y>0 such that Vu,-eyB (i = l,...,«), f] iKi — vi)f0 1=1 then (36) 'ixeK, mx)=f]T4x) i= 1 I m K, -K, Tk, n A2<0)={o} K, n a:2={o} Oe lnt(ifi -K^) Tk^ (0) ={* Ijc, < 0 } T/c^ (0) = {acI», > 0,X2> o} Tk, (0) n Tk2(0)={;« Ia;, =0,at2>0} Z
CH. 4, sec.2 derivatives and codifferentials of set-valued maps 177 Proof. Let D be the closed vector subset of X" of constant sequences 3c =(a:, X,..., a:). Then K is identified withPnfI"= i Xi. Assumption (35) implies that (37) OelntJ^n Since Tp{x)=D, proposition 13 and corollary 19 imply that This proves that Tk{x) is equal to the intersection of the tangent cones Tk,(x). 2. DERIVATIVES AND CODIFFERENTIALS OF SET-VALUED MAPS WITH CONVEX GRAPHS We now propose to develop a differential calculus for convex set-valued maps. We proceed as in elementary calculus, when the derivatives of real-valued functions are defined from the tangents to the graph. Let f be a convex set¬ valued map from a Banach space X to a Banach space Y. We fix a point (xq, To) in the graph of F.
178 СН. 4, SEC. 2 CONVEX ANALYSIS AND OPTIMIZATION The tangent cone Tg,ф^F){xo, З'о) to the graph of F at (xo, j'o) is a closed convex cone of X T We regard this cone as the graph of a closed convex process, denoted by DF{xo, уо) and call it the derivative of F at xo e and jo ^ F(xo). DEFINITION 1 The derivative DF(xq, Jo) of a set-valued map with convex graph at (xo, Fo) € graph(F) is the closed convex process whose graph is the tangent cone to the graph of F at (xo, j'o)- shall say that the transpose DF(xq, Fo)* of DF{xo, Fo) is the codifferential of the convex set-valued map F at (xo, In other words, (1) i. veDF (xo, j'o)(m)»(m, v) 6 Tgraph(f)(xo, Уо) ii. p 6 DF(xo, Fo)*(i)o(p, -q)e Wgraph(F)(xo, Уо) We begin by pointing out the following obvious (but useful) formulas. .2) i '• D{F - ‘)(yo, Xo) =DF{xo, yo)~' ^ [ii. D(F -i)(yo, Xo)*(p)= -{DF{xo, yo)*)~ H~p)- Al^. since the normal cone is the negative polar cone of the tangent cone, we have the following equivalent definitions. (3) i. peDF{xo,yo)*{g) ii. fxeX, 'fyeF{x), (q, уо-у)^{р, Xq-x) Hi. \/ueX,fveDF(xo,yo){u), (p,u)<(q,v)
CH. 4, sec.2 derivatives and codifferentials of set-valued maps 179 We point out the following monotonicity property. (4) If />, € DF (xi, yi)*igd (i = 1, 2), then <91-^2. yi -У2)^(Р1 -Рг, -Хг) which follows obviously from (3, ii). We observe that (5) i. Dom DF (xo, J^o) Tdomf)(^o) ii. Im DF (xo, З'о) «= 7J„(f>(J'o) Example Let Фа be the set-valued map from to У defined by (6) Фк{х)'=0 when xeK and Фк(х)=0 when x^K Then (7) Vx6/C, Dфк(x,0)=фт^^,^ A Indeed, graph Dфкix, 0)=Tg„ph^«(A:, 0)=Tr,<(o)(x, 0)=Tk(x)x{0}= graph (ФткмУ • These remarks having been made, we now prove an analytical characteriza¬ tion of the derivative to a convex set-valued map that captures the idea of a derivative as a suitable limit of “differential quotients.” PROPOSITION 2 Let (xo, >^o) belong to the graph of a set-valued map F with convex graph from X to Y. Then Vq e DF{xq, y^(uo) if and only if (8) 1- C p Ai F{xo + hu)-yo\ lim mf mf dl Vq, 7 1=0 M->Mo h>0 \ h Proof. Indeed, (mq, vq) e 7¡raph(f)(A:o, yo) if and only if for all e, >0, £2>0. there exist «£, eX and e Y satisfying and ||i;Ej|<e2 and /io>0 such that (xo -t- h(uo + Ue,), yo+Mvo + ^ S^nph (F ) for all h e ]0, ho\_. Hence f(xo+A(uo + Me,))-yo Po 6 Vc
180 CH. 4, SEC. 2 CONVEX ANALYSIS AND OPTIMIZATION and, consequently, F(xo + h{uo + Ue,))-yo inf d\vo, 0<h< ho The convexity of the graph of F implies that for all y e F(x), when then F{x + h2u)—y F{x+hiu)—y hi Indeed, for all у e F{x), we can write ^ F(x+h2v)+(^l - ^^y<=F (I (x+/i2u)+(l - ^yyF{x+hiv) This amounts to saying that the function 6->'d{v,{F{x + 6u)-y)/d) is increasing. So, we can write ,04 1- jI F{xo+hu)-yo\ . . ,( F{xo+hu)-yQ (9) hm i/ uo. 7 =inf</ uo,—^ л-»о+ V h ) h>o \ h Therefore, (10) inf d I Vo, 0<h<ho = inf if ( Vo, h>0 = Urn d\ Vo, F(,Xo+h{uo + Ut,))-yo h F(xo+h{uo+Uc,))-yo h Fixo+hiuo + Uet))-yo We have proved that (11) inf iiM-uoli h inf d (Vo, li>0 \ F(xo+hu)-yo <62 By letting 6i and 62 converge to zero, we obtain formula (8). ■ PROPOSITION 3 Let Xo and x belong to the domain of a set-valued map F with convex graph. For any уо 6 F{xo), we have (12) F{x)-yo^=-DF{xo, Уо)(л:-Хо)
CH. 4, SEC. 2 DERIVATIVES AND CODIFFERENTIALS OF SET-VALUED MAPS 181 Proof. Indeed, for any /i € ]0,1 [ and any y e F(x), we have (1 - h)yo + hyc F(xo+Mx - Xq)) Hence, F(xo+Mx-Xo))-yo y-yo e This implies that y-yo 6 DF(xq, yoXx-Xq). ■ Let P<= y be a closed convex cone of Y defining a preorder. We say that XoeK achieves the minimum of set-valued map F: K-^Y at yoS F(xo) if (13) fxeK, F(x)^yo + P Proposition 3 allows us to characterize the minimum of a set-valued map (variational principle). PROPOSITION 4 Let F be a set-valued map from К to Y with convex graph. Then Xq^K achieves the minimum of F on К at уо e F (xo) if and only if one of the equivalent conditions (14) hold true. Vmo e , DF (xo, Уо)(мо)=P 4qeP*, OeDFixo, To)*(4') Proof, a. Let д: e K. By proposition 3, property (12) implies that F{x)f=.yo+DF{xo, Уо)(^--х:о)<=Уо + Р
182 СН. 4, SEC. 2 CONVEX ANALYSIS AND OPTIMIZATION Hence, Xo achieves the minimum of F at yoe F{xo). b. Let Xo achieve the minimum of F at j;o ^ F(xo) and let vq eDF(xo, yo)iMo\ For all 8>0, there exist ueuo-^^B and h>0 such that (15) ,.efiii±M:2i+es=f+.s by (14, ii). Since P is closed, we deduce that vq belongs to P by letting e converge to zero. c. We prove that (14, i) implies (14, ¡1). Indeed, for all и e Dom DF (xo, Уо), V €DF(xo,yo)(M)<=i'.9 6i’^,wehave(0,M)-(^,a><0.Hence,0 еВР(хо,Уо)*(я)- d. Conversely, assume that 0 e DF(xq, уо)*(4') for all qeP^. Let vq belong to DF{xo, >'o)(mo)- Then we have (0, mo) — {q, i>o) <0 for all qeP^, that is, vo e P. Ш PROPOSITION 5 Let X and Y be Banach spaces, LcX and M<=Y closed convex subsets, and A 6 if(X, У) a continuous linear operator. We define the set-valued map F on X by (16) F{x):= Ax—M whenxeL [0 when xtL For any X eL and у e Ax—M, we have (17) and (18) DF(x, Au - Тм(Ах-у) when и e Tl(x) when и i Tiix) DF(x, A*q+ Ni{x) when q € Nm{Ax-y) when q i Хм{Ах-у) Proof, a. Let v belong to DF{x, y){u). Then there exist sequences of elements u„ converging to u, v„ converging to v, and h„>0 such that y+h„v„ belongs to F{x+h„u„) for all n. This means that x+h„u„ belongs to L for all n (and, thus, that u belongs to Tdx)) and Ax—y + h„{Au„—v„) belongs to M for all n. Therefore, belongs to Tu(Ax—y). b. Conversely, let u belong to Tl(x) and let v belong to Au- Tm(Ax-y). There exist sequences of elements u„ converging to u, w„ converging to Au-v, and h},, hl>0 such that x+h],u„ belongs to L and Ax-y-\-hlw„ belongs to M for all «>0. We set A„:=min(/ii, hi) and v„:=Au„—w„, which converges to v. We observe that y+h„v„ belongs to F{x+h„u„) and, consequently, v belongs to DF{x, y)(u).
СН. 4, SEC. 2 DERIVATIVES AND CODIFFERENTIALS OF SET-VALUED MAPS 183 c. By definition, p belongs to DF{x, y)*{q) if and only if sup sup [_{p,u) + {—q, Au—w)^ “^ ^L(^) we T\f{Ax -y) = sup {p — A*q,u}+ sup (q,w}^0 tie Tl(x) we T\t(Ax-y) This amounts to saying that p — A*q belongs to Ni{x) and that q belongs to Nu(Ax-y). ■ We now study calculus for derivatives and codifferentials of set-valued maps with convex graph. Let us start with the chain rule. THEOREM 6 Let F be a set-valued map from X to Y with closed convex graph and let A belong to X). We assume that (19) 0 6 Int(Im/1 —Dom T) Then the following chain rule formulas hold: (20) i. D{FA)(zo,yo)=L>F(Azo,yo)A ii. D{FA){zo, yo)*=A*DF{Azo, уоГ Proof We know that the graph ^ of G=FA, which is closed and convex, is equal to (Axl)~'^, where ^ is the graph of F. The assumption Oe Int(Im/4 —DomF) obviously implies that in X x zero belongs to Int(Im(y4 x 1) — ^). So by theorem 1.16, we know that T*r(zo, ;^o)=(^ X 1)” ‘ T^(Azo, Jo) and that iV»(zo. Jo)=(^* X l)N^(Azo, Jo)- This implies formulas (20, i and И). ■ Note that assumption (19) is always satisfied when A is surjective or when Im /Inlnt Dom Ff0. Let us consider a set-valued map F with convex graph from to У and В 6 if(y Z). The graph of the set-valued map BF from A" to Z is still convex. PROPOSITION 7 Let F be a set-valued map from X to Y with convex graph and let В belong to SF{Y, Z). Then (21) i. D{BF )(xo, Bjo)=BDF{xo, Jo) ii. D(BF )(xo, Bjo)* = DF (xo, Jo)*B*
184 СН. 4, SEC. 2 CONVEX ANALYSIS AND OPTIMIZATION Proof. We note that the graph ^ of G:=BF is (1 xB)^, where ^ is the graph of F. We deduce from proposition 1.14 that T^(xo, Вуо) = d((l XB)T^(xo, 3^o)) and N^{xo, Вуо) = {{1 xB)*~^уо). So formulas (21, i and ii) ensue. ■ The two preceding results are summarized in theorem 8. THEOREM 8 Let U, X, Y, and Z be Banach spaces, F a map with closed convex graph from X to Y. Let A belong to SF(IJ, Y) and В to Z); let UoeU and уо e F(Auq) be fixed. If (22) then 0 e Int(Im A — Dom F), As a consequence, we obtain theorem 9. D(BFA)[uq, Byo)—BDF{Auo, уо)А D(BFA)(uo, Вуо)* =A*DF(Auo, Уо)*В* THEOREM 9 Let X, Y, and Z be Banach spaces, F a set-valued map from X to Z,G a set-valued map from Y to Z, and let A belong to SF(X, У). We assume that the graphs of F and G are closed and convex and that (24) 0 e Int Dom F- Dom G) Let X 6 Dom FnA~^ Dom G, ye F(x), and z e G(Ax) be given. Then, (25) i. D{F + GA)(x, y+z)~DF{x, y)-i-DG(Ax, z)A ii. D{F+GA)*{x, y+z)=DF (x, y)*+A *DG(Ax, z)* Proof We observe that F-\-GA can be written as BGA where a. ~A: X eX-^{x, Ax) eX x Y b. F:(x,y)eXxY-^F(x)xG(y)<=ZxZ c. 5:(zi, Z2)6ZxZ->Zi+Z2 We also observe that assun^tion (24) implies that 0 6 lnt(lm A - Dom F). It then suffices to check that DF(x, y, zi, zf)=DF{x, Zj) xDG{y, Z2). ■ COROLLARY 10 Let F and G be two set-valued maps with closed convex graphs from a Banach space X to a Banach space Y. Assume that (26) 0 e Int(Dom F — Dom G)
CH. 4, SEC. 2 DERIVATIVES AND CODIFFERENTIALS OF SET-VALUED MAPS 185 and take x 6 Dom F n Dom G, y e F(x), and z e G(x). T/ie«, (27) i. D(F-\-G){x,y-\-z)=DF (x, y)+DG{x, z) ii. D{F-\-G)(x,y-yz)* =DF(x, y)*+DG{x, z)* We can compute the derivative of the restriction of a set-valued map with a closed convex graph. We note that F\k = F+(I>k, where is defined by (6). COROLLARY 11 Let F be a set-valued map from X to Y with closed convex graph and let F + фк be its restriction to the closed convex subset KczX. If (28) then for all xeK and у e F{x), (29) Oeint (Dom F-K) (.«• Ik)(^, y)*{d)=DF{x, yVi^ + N^x) to Tk(x). к = F + G. ■ COROLLARY 12 Let F be a set-valued map from X to Y with closed convex graph and let P<=Y be a closed convex cone. We assume that the subsets F+{x):=F{x)+P are closed for all X. Then for all x e Dom F and for all у eF{x) (30) i. DF+{x,y){u)=DF{x,y)(u)+P 10 if я ^ P [DF-i-(x, yY is the restriction of DF{x, yY to A Proof We apply corollary 10 with G defined by G(x):=P for all xeX. ■ We now compute the derivative of an intersection. PROPOSITION 13 Let us consider n set-valued maps Fifrom X to Y with closed convex graph. We assume that (31) and that л:о e Pi Int Dom F,- 1 = 1 (32) 3 у > 0 such that V(w,-, y.) e y{B x B\ fj (F,(x + щ) -1;,) 1 = 1
186 CH. 4, SEC. 3 CONVEX ANALYSIS AND OPTIMIZATION Then for any x 6 j Dom F,- and any y € j F,(x), we have (33) D ^ n (-^> y)= n DFi(x, y) Proof. The graph of f]"=i F, being the intersection of the graphs of the set-valued maps F„ we use proposition 1.20. We observe that assumptions (31) and (32) imply that there exists ¿>0 such that (34) Vi) e 5{B X B), n (graph (F,)-(«,-, Vi))f0 i = l Indeed, we take xo in the intersection of the domains F, and 5^y such that + Dom F,. Then we choose y^ in ni'=i (f^i(^o+«,)—u.), which exists by assumption (32). Hence, (xo, Fo) belongs to the intersection of the (graph F, —(i/f, Vi)). Therefore, proposition 13 follows from proposition 1.20. 3. EPIDERIVATIVES AND SUBDIFFERENTIALS OF CONVEX FUNCTIONS Let F be a real-valued function defined on a convex subset K of X. We extend it to X by setting (1) F(x):=oo whenxiK (we say that K = Dom F) We associate to F the set-valued map from A' to F defined by (2) V+(x):=F(x)-l-F+ if x€F,V+(x)=;2r if x$K whose graph is the epigraph of V. The latter is convex if and only if the function F is convex. Therefore, we can define the derivative £)V+(x, F(x)) of V+ at (x, F(x)), which is a closed convex process from X to R. Since DV+(x, F(x))(m) is either R, or a half-line [a, oo[ or {-I- oo}, this set is characterized by its lower bound. DEFINITION 1 The epiderivative of V at x is (3) D+F(x)(M):=inf{t)|u 6£)V+(x, F(x))(m)} eR We can check that (4) inf
CH. 4, SEC.3 CONVEX FUNCTIONS 187 We note that the convexity of ^implies that (5) V{xo-\-hu)-V(xo) V(xo-\-hu)-V(xo) /1-^0+ h h>o h The transpose of the derivative D\+{x^ V(x)) is a convex process from R to A"*, whose domain is contained in /?+. Then for any q'^0, we have (6) D\4x, V(x)nq)=qD\4x, V(x)ni) This brings us to definition 2. DEFINITION 2 Let V be a convex function form X to Rkj{-{-00} and x belong to Dom V. The subset 5F(x):=Z) V+(x, K(x))*(l) of the dual X* of X is called the subdifferential of Vat X. k We immediately give an analytical characterization of the subdifferential in proposition 3. PROPOSITION 3 Let V: X-*^Ru{-\-co} be a convex function. The following statements are equivalent: (a) pQ€dV(xo) (b) Vx e X, V{xq)-F(x)^ (/?o, Xq-x) (c) Vw 6 X, (/;o, u) V(xo)(u) A When K is a proper convex function, we obtain the following characteriza¬ tion of a minimizer of V, PROPOSITION 4 Let V be a convex function from X io ] —00, +00] whose domain is nonempty. a. For any Xq, xeK.we have (7) V(x)- K(xo)^Z)+ F(xoXx-Xo) b. xo minimizes V on K if and only if (8) Vw e Dom D+ F(xo), 0^Z)+ K(xo)(w) or, equivalently^ (9) 0g5F(xo) a
188 CH. 4, SEC. 3 CONVEX ANALYSIS AND OPTIMIZATION Theorem 2.6 implies the following property of the epiderivatives and sub¬ differentials of lower semicontinuous convex functions, which plays an import¬ ant role in optimization theory. THEOREM 5 Let X and Y be Banach spaces, V: X^Ru{-\-(X)} and W: + oo}proper lower semicontinuous, convex functions, and let A belong to S^{X, 7). We assume that (10) Then, (11) 0 € Int(^ Dom V— Dom W) fi. [ii. d(V (F-h WA)(x)=^D^ V(x)^D^ W(Ax)A d(V-h WA)(x)=dV(x) + A*dW(Ax) Consequently, x is a minimizer of x^ V{x)-\- W{Ax) if and only if x is a solution to the inclusion (12) OedVix)+A*dW{Ax) that is, if and only if there exists q eY* such that (13) -A*qedV(x), qedW(Ax) A It is useful to list several consequences of this theorem. COROLLARY 6 a. Let A e if (A", Y) be a continuous linear operator from a Banach space X to a Banach space Y and let W: Y^Ru{ + (X)} be a proper lower semicontinuous convex function. If (14) then (15) 0 e Int(Im A — Dom W) i. D^{WA)(x)=D^W{Ax)A ii. d(WA)(x) = A'^dW{Ax) b. Let Vand W be proper lower semicontinuous, convex functions from a Banach space X to Ru{ + (X)}. If (16) 0 6 Int (Dom V— Dom W)
CH. 4, SEC. 3 CONVEX FUNCTIONS 189 theft (17) i. D+(V+ W)(x)=D+ V(x)+D+ W(x) ii. 3( F+ H^)(x)=d V{x)+d W{x) c. Let V be a proper lower semicontinuous, convex function from a Banach space A" to/?u{ + co} and let KcX be a closed convex subset. If (18) then (19) 0 elnt(K —Dom V) i. Z)+(F|K)(x)=Z)+F(x)|rKW «• a(FU)(x)=aF(x) + iVK(x) Consequently, x eK minimizes V on K if and only i^O e 5K(x)+ Njfx) d. Let Vbea proper lower semicontinuous, convexfunction from a Banach space X to R'o{-\-<x>], let A belong to S£fX, y), and M be a closed convex subset of K If (20) then (21) 0 g \ni{M—A Dom V) -1 (M))(-x)=Z) + F(x)| X -1 tm(Ax) ii- d{V\A-HM^)=dV(x) + A*NM{Ax) Consequently, xeA ^(M) minimizes V on A ‘(Af) if and only if 0 edV{x)+ A*Nm{Ax). a More generally, we introduce (22) a proper lower semicontinuous, convex function L:Xxy^/?u{ + oo} to which we associate the lower semicontinuous, convex function V defined by (23) \/xeX,V{x):=L(x,Ax) where AeSi'(X,Y) We denote by A®-1 e S^(X x i; y) and 1 ©.4* 6 if(X* x Y*, X*) the con¬ tinuous linear operators defined by (24) {A®-l)(x,y)=Ax-y, {l®A*){p, q)=p + A*q
190 CH. 4, SEC. 3 CONVEX ANALYSIS AND OPTIMIZATION We observe that (25) Dom V 4^0 if and only if 0 e (^ © -1) Dorn L because x € Dom V if and only if there exists (w, v) e Dom L such that x=u and Ax = v, that is, if and only ifO = Au — v 6(AQ—1) Dom L. THEOREM 7 Let L satisfy (22) and V be defined by (23). If y^e posit the assumption (26) 0 € Int ((^ © -1) Dom L) then (27) dV(x) = {\ ®A*)dL*(x, Ax) Consequently, x eX is a minimizer ofVif and only if (28) Oe{i®A*)dL{x,Ax) that is, if and only if there exists q eY* such that (29) {-A"^q,q)e dL{x, Ax) A Proof Let 1 (g)/l be the map: x-^(x. Ax) from X to X xY whose transpose is 1 ®A*. The function V defined by (23) is L(1 ®A). We point out that assump¬ tion (26) implies that (0, 0)elnt (Im(l®^)—Dom L). Indeed, let (u, v)€X. By assumption (26), there exist A>0 and (x, ;;) e Dom L such that A(v — Au)= Ax—y. Set z:=Au + x. Then Az=Av-\-y and A(u, v)=(z, Az)—(x, y) 6 Im (1 ®^)— Dom L Then formula (27) follows from formula (15, ii), and the second equivalent statements follow from b of proposition 7. ■ PROPOSITION 8 Let us consider n convex lower semicontinuousfunctions Vifrom X to ^ — co, + oo]. Let V be defined by (30) and (31) V(x)\= max Vi(x) 1=1 n J{x)\=[i = l,..., n\Vi{x)=V(x)]
CH. 4, SEC. 3 CONVEX FUNCTIONS 191 Let us assume that the n functions Vi are continuous at x. Then (32) i. D+V(x)(y) = max D + l^(x)(p) i 6 J(x) ii. 3F(x)=co( (J dVi{x) \i« J(x) Proof. We observe that V(x)=i Vi+(x), hence, our result follows from proposition 2.13. We know that x belongs to the intersection of the interiors of the domains of the maps Vj+. Assumption (2.32) is obviously satisfied. Then by proposition 2.13, Z)V+(x, V(x))= П V{x)) i=l When i belongs to J(x\ then D\i^(x, V(x)\u)=D\,^(x, Vi(x)){u) = [p^Vi{x\u\ oo[ When i does not belong to J(x), then {x, V{x)) belongs to the interior of the graph of V,+, so that Z) V,+(x, 7(л:))(м) = Л for all и eX. Then D\+{x,V(x))(u) = П D\i+(x,V(x)){u) i e J{x) SO that formulas (32) ensue. ■ Let us point out the monotonicity property of the subdifferential of a convex function. PROPOSITION 9 Let V be a proper convex function from A"io/?u{ + oo}. Then the set-valued map X € X-^dV(x) e is monotone (33) V(x,p), f (y,q)egraph (dV), (p-q,x-y)>0 Proof Indeed, since pedV{x) implies that F(x)—F(y)<(p, x—y) and since q edV(y) implies that V{y)— F(x)< {q, y—x), we obtain formula (33) by adding these two inequalities. ■ We proceed by computing the subdifferential of convex functions of the norm in a Banach space. PROPOSITION 10 Let X be a Banach space. Then the norm is subdifferentiable and (34) Vx€X,d||-||(x);={p6X*Kp,x)=|W| and M| = l}
192 сн. 4, SEC. 3 CONVEX ANALYSIS AND OPTIMIZATION ^ a> 1, then ^хеХ,д(^^\\‘Гух)=ИГЩИ\){х) = {peX*\ (35) , = \\x\\ and <p,x) = ||x||*} A Proof, a. The Hahn-Banach extension theorem implies that we can associate with any xqeX a continuous linear functional po e X* such that </>o, a:o)=1|xo|| and ||polU = l Indeed, we extend the continuous linear functional defined on Rxo by p(axo)'-= “lUoll. whose norm is equal to one, by a continuous linear form po on X whose norm llpolU is still equal to one. Then (po, x-xo) =^||/>olUlNI - (po, Xo>=l|x|| - lixoll This implies thatpo e 5(|| • ||)(xo). Conversely, ifpo e 5(|| • ||)(xo), then ||xoll - ||x|| < (po, xo-x). By taking x=Axo, this implies that (l-2)«po,xo)-||xoll)^0 By successively taking A > 1 and A < 1, we deduce that (po, Xo)=|lxo||. Therefore, (po, x) < ||x|| for all X € AT and (po, xo) =xq. This implies that ||polU = 1- b. The second statement follows obviously from the first. ■ Actually, any lower semicontinuous, convex function on a Hilbert space is subdifferentiable on a dense subset of its domain. This is also true in Banach spaces, as we shall see in theorem 4.4.3 in the next chapter. THEOREM 11 Let X be a Hilbert space and V a proper lower semicontinuous, convex function from X to Ли{ + оо}. a. The domain Dom of the set-valued map dV(-) is dense in Dom V. b. For any X eX and A>0, there exists a unique solution Xx to the inclusion (36) X BXx+^dV{xx) Proof We associate to each positive A the minimization problem (37) Vx(.x):= inf yeX ^(y)+^lb-x|p
CH. 4, SEC. 3 CONVEX FUNCTIONS 193 As we will see later on (theorem 4.2), a proper lower semicontinuous, convex function is bounded below by an affine function: There existp eX* and a eR such that WyeX, V{y)> (p, y) + a. Therefore, ^'(y)+¿ i\\y-x+W-^^\\pr)-a+(p, x) so that Vx(x)^ —('^/2)||/j||^+a+ {p, x) is a finite number, a. We begin by proving that there exists xx such that (38) Vx{x)=V(xx)+^ \\xx-x\\^ To this end, we consider a minimizing sequence of elements y", which satisfy F(>'")+(l/2A)||y—Vx(x)+{l/n). It is a Cauchy sequence because ll/-/‘lP=2||/-x|p + 2||/‘-x|p-4 — x\\ <42 ^l+l+2K,(x)- F(/)- F(y’")j + 82 I^a(x)J <42 [j+^+2F <42 f-+—I (for V is convex) \« mj Hence, y" converges to an element xx- Since V is lower semicontinuous, it follows that f"(^A)+:il|.x:A-x||^<lim inf K(/)+ Hm i II/-^IP< Ia(:>c) ZA ;i->oo n-*ao ZA b. Let xa be a solution to the minimization problem (37). Then by taking y=xx+0{xx—z), we deduce that F(x>i)+ilkA-;c||^<(l-0)F(xA)+0nz) 22 +-4 (II^A - + 20 (xa -x,Xx-z) + O^Wxx - z|P) ZA After simplification and division by 0>O, we obtain F(xa)-F(z)<j (xx-x, Xa-Z>+ 0||XA-r||
194 CH. 4, SEC. 3 CONVEX ANALYSIS AND OPTIMIZATION By letting 0^0, we have proved that (39) j(x-xx)edV(xx) that is, V is subdifferentiable at xx, a solution to inclusion (36). c. We now prove that the domain of the map 5K(*) is dense in the domain of V. Let X belong to Dorn V and p belong to the domain of its conjugate function V*, which is not empty (see theorem 4.2 below). Since (40) and since we deduce that ^||xa-x|P + K(x,)=K,(x)^F(x) - 1^(xa)^-a-<p, xa> ^ ||xa - x|p < F(x) - a - (p, x) + (p, x-xx) IIa: - xaIP + F(x) - a-<p, x>+A||/> P (for ab^a^/4X+b^X). Hence, when A converges to zero, ||x-XAP<4A(K(x)-a-(p,x>+A|pP)->0 Since xx belongs to the domain of d F( •), we have proved the required statement. Remark By taking 7=x in the definition of Vx, we see that F(xa)<L(x) for all A>0. Since Xx converges to x 6 Dom V and since V is lower semicontinuous, we deduce that (41) VxeDomF F(x)= lim Vx{x) Remark Proposition 9 states that 5F(*) is a monotone map, and theorem 11 states that 1+A5F is surjective. By Minty’s theorem, we know that the subdifferential map 5F(') is a maximal monotone set-valued map; properties of maximal mono-
CH. 4, SEC. 3 CONVEX FUNCTIONS 195 tone maps are studied in Chapter 6, Section 7. We know in particular that (42) (1 Ms a single-valued nonexpansive map defined on the whole space X which is called the resolvent Jx'={^ ‘ of 3K The map defined by (43) A;^{x) :=j (x-Jxx) is Lipschitz with constant 1//: It is called the Yosida approximation of In our case, Ax(x\ the Yosida approximation of dV, coincides with the Frechet derivative of the convex function Ki. ■ PROPOSITION 12 The Yosida approximation Vx of a proper lower semicontinuous, convex function V: X->/?u{ + ooj is a differentiable function defined on X, which converges pointwise to V. k Proof We already observed that Vx does converge pointwise to K; the convexity of Vx is obvious. We remark that Jxx is the solution xa to inclusion (36), and we set yx'.=^Jxy- We shall prove that Axx:=j (x-Xx) = VVx(x) On one hand, Vx{x)-Vx{y)= V{xx)-+ ^ \\xx-x\\^-^ \\yx-y\\^ + ^\\xx-x\\^-^\\yx-yr +l\\xx-x\\\\yx-y\\ ^\j{x-Xx), x-x
196 CH. 4, SEC. 3 CONVEX ANALYSIS AND OPTIMIZATION Hence, (1/A)(x-xa) e On the other hand, since (1/A)(y—>»a) 6 5Fa(t), we obtain Vx{x)-FaO)>( ^ (y-yx), x-y ^{\(,x-xx\ x-jV(t (x-Xa)- \ iy-yx), x-y ^{\(x-xx),x-y}- j{x-xx)-\{y-yx) j(x-Xa), x-W- jllx-jlP because Axx-={\!^){x—xx) is Lipschitz with constant 1/1. So we have proved that Vx{x)- Vxiy)-<(1/A)(x-Xa), x-y) This proves that AxX = ’VVx(x). iix-7ii <0 Before proving that a continuous convex function is subdifferentiable, we show that convex continuous functions are locally Lipschitz. PROPOSITION 13 Let V\ X^R\j { + (Xi} be a proper convex function. The following conditions are equivalent: (44) {ri; V is bounded above on an open subset {contained in Dom V). V is locally Lipschitz on Int Dom V. Proof a. It is clear that (44, ii) implies (44, i). b. Let us assume that V is bounded by a on a ball xo+i/5<=Dom V. We associate with any x e X the element y-= xo-(l-0)x ^ ^ \\x-Xo\\ where 0= <1 O' ||x-Xo|| + ii Hence, lb—^oll =>7 and, consequently, ^(y)^^ The convexity of V implies that F(xo) = V{6y + (1 - d)x)^6a + (1 - 0)V{x) Thus, ^'(^o)< V(x)+j^ («- F(xo))= K(x)+^^^^ ||x-Xo|| 1 —(/ rj
CH. 4, SEC. 3 CONVEX FUNCTIONS 197 Now, take xexo + t]B and y—(x—{\ —6)xo)l6, where 0 = Hxo —x|l/f}< 1. Thus, \\y-Xo\\=ri, and, consequently, F(y)<a. The convexity of V implies that F(x) = V{ey + (1 - e)xoH 0a + {l- 9)V{xo)=0{a - F(xo)) + V(x) =^^^^^\\x-Xo\\ + V{x) Therefore, when x e xo + t]B, then (45) |F(x)-F(xo)|<^ c. Let us prove that F is Lipschitz on the ball Xo+pB, where ^ e]0, >/[. Let Xu X2 belong to xo+PB. We choose an integer n^||xi-X2||/(f;—/5). For j=0, . . . , n, we introduce the elements yj=xi+{j/n'^X2 — xi). Then yo=Xu yn=X2, \\yj-i-i-yj\\=\\xi-X2\\/n^ti-P, ||xi-X2||=I]”:^ Ibj+i-yjll, and for y=0,.. .,n,yj belongs to Xo + pB. Hence by (45), a- F(;^j)<2(a— F(xq)) and since F is bounded above by a on yj+{ri — P)B, we deduce from (45) (where xo is replaced by yj) that IV{yj^,)~ F(j;,)llb;+1-yM Consequently, in^i)- F(x2)|^'I \V{yj^i)~ V(yj)\ j=0 ^2(a-F(xo))y ,11 2(a-F(xo)),, d. Finally, we shall prove that for any xi e Int(Dom F), F is bounded above on a neighborhood of Xi and thus by the preceding result F is Lipschitz on a neighborhood of Xi. Since there exists y such that Xi +yBc Dom V, then i ^ / V X\—^Xo ^ X2-.=Xo-y-—^(x,-2Co)=^;—€ Dom F 1-A 1-A since
198 CH. 4, SEC. 3 CONVEX ANALYSIS AND OPTIMIZATION Let y €Xi +Xrj. Then the element z:=j {y-\-^Xo-Xi)=j {y-{l-X)x2) satisfies \\z—Xo\\=(i/À)\\y — Xi\\^ri. Consequently, V{z)^a, and the convexity of V implies that K(y)=F(2z4-(l-A)x2)^AK(z)+(l-A)Ffe)^Aa + (l-A)F(x2) = :ô Hence, V is bounded above on a neighborhood of Xi. ■ We deduce the following important consequences. COROLLARY 14 If the interior of the domain of a convex function V\ {+ oo} nonempty, then V is locally Lipschitz on Int(Dom V). k Proof Let B{xo, rj) be a ball with center xq and radius rj contained in Dom V. We can then find n points x,- e B{xoy rf) such that the vectors Xi — Xo are linearly independant. The subset S of convex combinations ^”=o where A,->0 for all U is open and contained in Dom K Consequently, since K(^”=o the convex function V is bounded above on S. The statement follows from proposition 13. ■ COROLLARY 15 If the interior of the domain of a lower semicontinuous, convex function V from a Banach space to Ru{ + oo} is nonempty, then V is locally Lipschitz on Int Dom V Proof By a consequence of Baire’s theorem, V, a lower semicontinuous, real¬ valued function defined on the open set Int(Dom V), is bounded above on a nonempty open subset. The statement follows from proposition 13. ■ Continuous convex functions are epidifferentiable and subdifferentiable on the interior of their domain. THEOREM 16 Let Xo belong to the interior of the domain of a lower semicontinuous, convex function V: -foo}. Then (46) VweA", D+V(xq)(u)= lim is finite /i->o+ a
CH. 4, SEC. 3 CONVEX FUNCTIONS 199 There exists rj>0 such that (47) F(xo)-l^(^o-w)<ß+l^(^o)(M)<l^(^o + M)-l^(jio) In addition, for some constant oOwe have (48) I'■ ^ [ii. (x, u) € Int Dorn V xX-^D+ V{x){u) is upper semicontinuous. A Proof. Let f/>0 such that the ball Xo-\-y\B is contained in Dorn V. Since xo = (-^0 ^ (-^0 “ we deduce from the convexity of V that -00 < n Since Ä->[F(xo+/i«)- F(xo)]//i is increasing, it follows that n /i-»0+ n h>0 We also know that there exists f/ > 0 such that V is Lipschitz on the ball xo + r\B. If oO denotes the Lipschitz constant, then r. I// V ^^^(^o^-M-F(xo)^ II II D + F(xo)(m) ^c| I m| I Finally, Z)+F(x)(w) being the infimum of the continuous functions (x, w)-^ [K(x+/?«)— V(x)'\lh on Int Dom F xX, it is upper semicontinuous. ■ We now translate these results into terms of the subdifferential. THEOREM 17 Let Xq belong to the interior of the domain of a lower semicontinuous, convex function F: —00, H-oo]. Then (49) dV{xo) is a nonempty, bounded, closed, convex subset of (and thus weakly compact)
200 CH. 4, SEC. 4 CONVEX ANALYSIS AND OPTIMIZATION and (50) X 6 Int Dom V^dV(x) is upper hemicontinuous on Int Dom V. Furthermore, (51) D+V{xo)(u) = (T{dV{xo\ u) that is, D+ V{x)(*) is the support function of dV{x). If dV{xo) contains exactly one element, this element is the gradient VK(xo) of V at Xq. A Proof Since dV(xo) = {p e X*|Vw e X, {p, u)^D^ V{xo){u)} and Z)+7(x)(*) is convex, lower semicontinuous, and positively homogeneous, it is the support function of a nonempty, closed convex, which is 3F(xo). Since o-(K(xo), u)=D^V(xo){u)^c\\u\\=o{cB^, u) we deduce that dV(xo)^cB^, that is, it is bounded. By theorem 16, dV(*) is upper semicontinuous from Int Dom V to X*. If dV{xo) = {po] contains exactly one point, then M->D+ K(xo)(«)=ff({;>o}, «)= {po, u) is a continuous linear functional. Hence,/?o=VK(xo) is the gradient of V at xq. 4. CON JUGATE FUNCTIONS We can introduce the conjugate function K* of a proper function V: X^R\j {4- oo} in many ways. We choose the following point of view. Instead of studying simply the minimization problem: inf^ex y{x), we study the whole family of perturbed problems -K»:= inf {V{x)^{p,x)) xeX when the function V is perturbed by the continuous linear functionals x^ (p, x). (They constitute the class of simplest perturbations that can be considered.) This defines a function /?-^F*(/?) on the dual X* of X, which is obviously convex and lower semicontinuous, since K* is the pointwise supremum of the continuous affine functions p-^(p, x) — V{x). Observe that ifF:X->/^u{ + oo} has a nonempty domain, then V* maps X* into /^u{ + oo}, because if xo belongs to Dom V, then for all /? 6 X*, we have - 00 < (/?, Xo) - T(xo)^sup {(p, x) - K(x)) = :K*(/?)
CH. 4, SEC. 4 CONJUGATE FUNCTIONS 201 DEFINITION 1 Let V be a proper function from X to /?u{ + oo}. The function from to Ru{ + oo} defined by (1) V/7 e X*, F*(/?):= sup ((/?, x) - V(x)) xeX is called the conjugate function of V. Similarly, if W is a proper function from X"^ to Ru{-\-(X)},w;e define W^from X to Rkj{-\-co} by (2) W*(x):= sup «/?, x)-W(x)) peX* and consequently, the biconjugate K** of V is defined by K**:=(K*)*. ^ We observe that both K* and K** are lower semicontinuous, convex functions and that (3) Va:, V**(x)^ sup i(p, x)-{{p, a:) - K(a:)))^ V{x) peX* The crucial point is that the conjugacy operation is a one to one cor¬ respondence between proper lower semicontinuous, convex functions defined on X and X*, respectively, which allows us to use the convenient interchange properties of duality: We shall have each time the choice of working with either a function F or a function F* according to the properties needed to solve the problem at hand. THEOREM 2 A proper function F: X->/^u{ + oo} is convex and lower semicontinuous if and only if F= F**. In this case, the domain of F* is nonempty. A Proof Since F is convex and lower semicontinuous, Ep{Vy.= {(x,X)eX xRIV{x)-X^Qi} is a closed convex subset oi X xR. a. We assume a<V{x). Since the pair (x:, a) does not belong to Ep(V), there exists a continuous linear functional (p, -a)eX*xR that strictly separates (x, a) from Ep{v). Then there exists e>0 such that (4) V_v6Dom K VA^O, (p,j')—aF(j)—aA<(p,x:)—aa—e By taking the supremum when A > 0, we deduce that a ^ 0 and (5) VpeDomK (p,y)-<xV{y)^(p, x)-ixa-e
202 CH. 4, SEC. 4 CONVEX analysis and optimization b. When a>0, we can divide by a and set p=p!a. Inequality (5) becomes (6) VyeDomK (^..y>-I^(FK<Ax)-a-- and thus by taking the supremum when y ranges over Dom V, (7) x)-a- - The first consequence is that Dom K* 4^0, that is, 7* is proper. The second consequence is that whena< F(a:), then a< {p, x)— V*(p)^ K**(x). So by letting a converge to V{x), we deduce that V{x)= F**(x). c. We observe that a>0 whenever x e Dom V; indeed, we take y=x in in¬ equality (5) and find that 0L{a— — г, a>0. When x $ Dom K both situa¬ tions a >0 and a=0 can occur. It remains to study the latter case. d. When a=0 and x i Dom V, inequality (5) implies that (8) V;; eDom V, (p, y — x) + s^0 Let p e Dom K*, which exists by (b) and (c). By multiplying (8) by «>0 and adding to it inequality (p, y) — V(y)— V*{p\ we obtain (9) Vj^eDom K (p-^npyy}-n(pyx) + ne-V*{p)-V{y)<:0 which implies that by taking the supremum with respect to j; € Dom K -{-np)—n (p, x)+ns— V*(p) ^ 0 By adding and substracting (p^ x), we deduce that ns + (p, x) - (p-^np, x) - V*ip + npH V*^{x) By letting «-►00, we conclude that F**(x)=oo. ■ The very definition of the conjugate function V* implies that for all x eX, p eX*, inequalities (10) <p, x)^V{xHV^p) known as Fenchel inequalities, hold true. It is quite natural to distinguish the pairs (x,p) such that inequality (10) is actually an equality. PROPOSITION 3 Let V: X->jRu{ + oo} be a proper convex function and x 6 Dom V be fixed.
СН. 4, SEC. 4 CONJUGATE FUNCTIONS 203 The following statements are equivalent: (a) pGdV(x) (b) {p,x) = V(x)+V*{p) A In other words, dV(x) is the set of gradients p of the affine functions x-* (p, x) - V*{p) that pass through {x, F(x)) and are below V. Proof Indeed, pedV{x)if and only if by proposition 3.3, VyeA', <;?, J>-K(y)<<p,x)-F(jc) that is, if and only if F*(p):=sup [(p, y) - V(y))=(p, x) - V{x) ■ yeX When V is lower semicontinuous, we deduce the following reciprocity formula in theorem 4. THEOREM 4 Let V be a proper lower semicontinuous, convex function from X to Ru{-hoo}, Then the inverse of the set-valued map dV{*): X-^X* is dV*, (11) pedV{x) if and only if лге5К*(/?)
204 сн. 4, SEC. 4 CONVEX analysis and optimization Proof. It follows obviously from proposition 3 applied to K* and equality As a consequence, we obtain proposition 5. PROPOSITION 5 Let V be a proper lower semicontinuous, convex function. Then is the set of minimizers of the perturbedfunction V(y)— (p, y). ^ We now proceed with some examples of conjugate functions. PROPOSITION 6 The support function of a subset К is the conjugate of its indicator. Conversely^ if К is closed and convex, Фк=<^к- ^ Proof. Indeed (Tk(/>):=sup {p, x)=sup {(p, х)-фк(х))=:фЦр) хек When К is closed and convex, {¡/¡c is convex and lower semicontinuous, and thus Фк = Фк*=<^% ■ In particular, we mention this important example: The conjugate function of the continuous linear functional p e X* is the indicator of the singleton {p} Р*=Ф{р) because p can be regarded as the support function of the singleton {p}: (7(p}(x) = (p, x) for M X eX ■ PROPOSITION 7 A proper lower semicontinuous, positively homogeneous convex function cr: {+ oo} is the support function of the subset K:={xeX\ypeX*, {p,x)^(r{p)} k Proof. Since a is positively homogeneous, we observe that (t*(x)=0 when xeK and (t*(x)= + oo when x^K. Therefore, <7* is the indicator of K, and o-=(T** is its support function. ■ We now compute the conjugate functions of the functions x->(I/a)||xH®.
CH. 4, SEC. 4 CONJUGATE FUNCTIONS 205 PROPOSITION 8 a. The conjugatefunction of x-> ||x|| is the indicator of the unit ball B* of the dual (12) !»={' 0 ifWph^i + 00 if\\p\\*>i b. Let 4>: R->^Rkj{ + oo} bea proper lower semicontinuous, even convex function and (t>* its conjugate. Then (13) 0(11-ID*=«^*(11-lU) c. In particular, if a > 1, then by setting <x* =a/(a -1), we obtain (14) (‘iHl'T-i"-"* * Proof a. When a = 1, we observe that || x| = Ob-^x) is the support function of the unit ball B* of the dual. Hence, || • |U=1^»'indication of B*. b. We check that 0(||-||)*(;»):=sup [<p,x)-0(IWD] xgX =sup sup [(/?, x) —0(IUID] A^O xeX 11X11 = A = sup [A||/>|U-0(A)] A>0 =sup [AIIpII,^ - 0(A)] (because 0 is even) XeR =4>*{\\p\U) c. Holder’s inequality states that 1 1 . for all a, 6/?+, The equality holds true when a°‘=b“’. If we set 0(A)=(l/a)|A|®, this shows that <j>*(p)=(l/a*)| Therefore, the last statement follows from the second. ■ We now present elementary properties of conjugate functions.
206 CH 4 SEC. 4 CONVEX ANALYSIS AND OPTIMIZATION PROPOSITION 9 If V^W,thenW*^V*. Let us set W(x)= V(x-Xo)+ \Po, x)+a. Then (15) W*(p) = ^*(P-Po)+{p,Xo)-{a. + {po,Xo)) A Proof. The first statement is obvious. We verify the second: sup [_{p, x)- IT(x)]=sup {{p-Po, x)- K(x-Xo)]-a xeX =sup i(p-po> X-Xo)- F(x-Xo)] + <P, Xo)-(a + </>o, Xq)) xeX = V*{p -Po) + (P> JCo) - (a + (po, Xo» ■ PROPOSITION 10 Let U: X X y-+/?u{ + °o} be a proper function and A e 7). Let V(y)\= infx^x U{x,y-Ax). Then for all qeY*, (16) Proof Indeed V*{q}= U*(A*q, q) K*(^):=sup l(q,y>" yeY =sup sup [(i'. T-x) - U(x, y- Ax)~\ xe XyeY =sup sup l(A*q, x) + (q, z) - C/(x, z)] x€ X ze Y = U*(A*q, q) ■ In particular, we obtain the following formulas. a. Let U-. X^/?u{ + oo} be a proper function and A e S£{X, 7). Set F(y):= m[Ax=y U{x). Then (17) V*{q)= U*{A*q) b. Let Ux and U2 be two proper functions from X to {+ 00}. Set V{y)\= inf,6 X (+ Viiy- ^))- Then (18) V*(q)=V*{q)+Ui{q) c Let U- X-^Ru{ + oo}, W: Y-^Ryj{ + co} be proper functions and A e ^(X, 7). Set K(y):=inf,.x my~Ax)). Then (19) y*{q)= V*{A*q)+W*{,q)
CH. 4, SEC. 4 CONJUGATE FUNCTIONS 207 We now turn our attention to computing the conjugate functions of functions of the form VA, V1 + V2, V+WA and, more generally, functions of the form x-^L{x, Ax), where L is a proper lower semicontinuous, convex function (20) , [from X y to /?u{ + oo}. We introduce the lower semicontinuous, convex function defined by (21) V{x):=L(x, Ax) whose domain is nonempty if and only if 0 € ((.^ ©— l)Dom L). Let L* denote the conjugate function of L defined on A"* x Y* by (22) I^(p> &•= sup [ (p, x) + {q, y) - L{x, :v)] (x,}») eX y.Y Hence, the domain of the function q-^I^(p — A*q, q) is nonempty if and only if p belongs to (1 ©yi*)Dom L*. THEOREM 11 Let L satisfy assumption (20) and V be defined by (21). We posit assumption (23) 0€lnt((^©-l)DomL) Then if po G (1 ©y4*)Dom L*, there exists q eY* such that (24) V*(po) = L*(po - A % q) = inf L*(po ” A *q, q) k qeY* Before proving this theorem, we mention explicitely the following particular cases. COROLLARY 12 a. Let W: Y-^Ru{ + oo} be a proper lower semicontinuous, convex function andAe^iX, Y). If (25) 0Glnt(Imy4-Dom W) then for all p e A* Dom W*, there exists qeY* satisfying A*q=p and (26) {WA)*(p)=W*{q)= inf W*{q) Aq = p b. Let Wii X-^Ru{-i-oo} and 1^2- +cx)} be two proper lower semi¬ continuous, convex functions such that (27) 0 6 Int(Dom Wi — Dom W2)
208 CH. 4, SEC. 4 CONVEX analysis and optimization Then for all p € Dom Wt + Dom W*, there exists q eY* such that (28) m + W2)*{p)= WUq)+ Wi(p-q)= inf {W\{.q)+ Wt(p-q)) qeX* c. Let U: A'-»'i?u{ + oo} and W: Y-^Rkj{ + co} be two proper lower semi- continuous, convex functions and A e Jf(X, Y) satisfying (29) 0 6 Int(/4 Dom U- Dom W) For all p € Dom U*-\-A* Dom W*, there exists q eY* such that (30) (U-yWA)*{p)=U*(p-A*q)+W*{q)= \rd^{V*{p-A*q)+W*{q)) A qe Y* First proof of Theorem 11 (self contained). Let po belong to (1 ©y4*)Dom li. a. We introduce the map ip from Dom If <zX* x Y* to Rx X* defined by (31) \l/{p,q):={Il>‘(P+Po,g)>P+A*^) and we consider the set (32) iA(Dom L*-(po,0))-i-R^.x{0}<=RxX* It is easy to check that this subset is convex. We shall deduce from assumption (23) that it is closed. Indeed, let us consider a sequence of elements (v„, r„) belong¬ ing to this set and converging to (p#, r*) in /? x *. There exist elements/>„ e * and q„ € Y* such that v„>L*(p„ +po, r„ =p„ + A*q„ and there exists a ball of radius y > 0 contained in (^4 ® — 1 )Dom L by assump¬ tion (23). Therefore, for all z e K there exists (x, y) e Dom L such that yz/||z|| = y-Ax. Consequently, 7t4 < z) = (q„, y) - x) = {qm y) + {P»’X)-{r„,x) ^ L*(/J„ +P0, q«) + L(x, y) - (rm x) - (po, x) <U„-|-L(x, y)- {rm x) - (po, x) Since the sequences v„ and (r„, x) are convergent and thus bounded, we have proved that (33) 'izeY, sup (q„,z}<+CO n^O
CH. 4, SEC. 4 CONJUGATE FUNCTIONS 209 The boundedness theorem implies that the sequence of elements q„ e y* lies in a weakly relatively compact subset. Therefore, a subsequence q'„ converges weakly to some q^ in Y*, and, consequently, q'„=r'„-A*q'„ converges weakly top^.=r^-A*q^f. Since the function L* is lower semicontinuous for the weak topology of X* X y*—as a pointwise supremum of the weakly continuous aihne functions ipy (py y'i ~ we deduce that ^iP*+Po> inf L*(p^+po> iii)^ I™ v'„ = v^ n-><30 "■*<» Thus we have proved that v^^L*(p^+Po,9*)> r^=P* + A*q*, that is, (v^, rj belongs to iA(DomL*-(po, 0)) + /?+ x{0} b. Now we shall prove that (34) (V*(Po), 0) 6 i/^(Dom L*-(po, 0))+/?+ x {0} Indeed, this statement implies the theorem because there exist q e Y* such that (35) V*{po)>L*(po-A*q,q) But inequalities (po, x) = {po-A*q,x} + (q,Ax}y^If(po~A*q, q) + L{x, Ax) imply that V*(po)^inf^evPf‘(po — '^*9y 9)- Hence, equation (24) follows from (34). Assume that inclusion (34) is false. Since the subset <//(Dom L*-(po, 0)) + R+ x(0} is convex and closed, the separation theorem implies the existence of a pair (a, —x)eRxX and e>0 such that aF*(po)=S inf (oiI^{p+po,q)+{p + A*q,-x})+inf aO-B (p,q)eX'xY* 9>0 We deduce that a is not negative and infe>o «0=0. Further, sincepo belongs to
210 CH. 4, SEC. 4 CONVEX ANALYSIS AND OPTIMIZATION (l©y4*)Dom L*, there exists (p, ^)6Dom L* — (/?o, 0) such that p-\-A^q=0. This implies that a>0, because if a=0, we would deduce from the preceding inequality that 0^ — e. Dividing by a>0 and setting x:=x/oc, rj=e/oL, we would obtain V*{po) < inf [L*(p +po, q) - (p +Po, x) - {q. Ax)'] + {po, x)-q p,q = - L(x, Ax) + (poy x}-rj ^V^{po)-rj a contradiction. Hence, statement (34) holds true, and our theorem ensues. ■ Second proof of Theorem 11. a. We temporarily set W(p):=\niqeY*L*(p-\-A*q,q). Proposition 10 implies that (36) W*{x)=L{x, Ax)=V(x) b. If we prove that the convex function W is also lower semicontinuous, then theorem 2 implies that V^(p)=W^\p) = W{p\ that is, the second statement of the theorem holds true. Since W{p)=m{qeY* U(q) where U{q):=If{p — A*q, q\ it suffices to prove that i/isinfcompact. Indeed, let K:={^ 6 T^IU{q)^ W^0^)+1}, which is then a weakly compact set. Then IP(/?)=minqgx L*(p —^4*^, q\ Since L* is lower semicontinuous, corollary 3.22 implies that W is lower semicontin¬ uous. c. Therefore, the second statement of theorem 3 follows from the following proposition. ■ PROPOSITION 13 // 0 e Int((y4 © — l)Dom L), then the functions q-^l^{p — A* q, q) are inf compact. Proof We begin by estimating the conjugate function of U. If Zo-=yo — Axo € —(y4© —l)Dom L then i/*(zo)< L(xo, yo)~ (p, Xo) Indeed, U*{zo)= sup l(q, Zo)-I^{p-A*q, q)] qeY' = sup inf inf [(q,Zo)-(p-A*q, x)-{q,y) + L(x, 3^)] qeY* xeX yeY
CH. 4, SEC. 4 CONJUGATE FUNCTIONS 211 < sup [{q, Zo + Axo—yo)~iP' a:o) + L(a:o,;'o)] «ey’ =L(xo, yo)-{p,Xo) Therefore, assumption (23) implies that there exists 7>0 such that for all z 6 Y, there exists (x,y) € Dom L satisfying vz/||z|| =y—Ax and thus, l/*(yz/||z||) <C + 00. Let Kji:={ql U(q)^Ji} be a level set of U. Hence, for all z e i; yz llzll sup VT737< sup i/(i)+i^*(if5r)<+°0 «eKA \ l|2||/ 4®XA Then Ka is weakly bounded and thus weakly relatively compact. ■ Remark Formula (1.5.29) on support functions is a consequence of formula (30). Indeed, if K:={x e LjAx e M}, then its indicator can be written i/'k(x)=i/'l(x)+i/'m(Ax) Therefore, if Oeint (y4L-M)=Int(^ Dom ^¿-Dom i/'m), for all p^b(K), there exists ^ 6 Y* such that (37) OK{p)=<^L(p-A*q)+<TM{q)= inf i<rL(p-A*q)+(TM(q)) ■ leY' It will be convenient to investigate further partial conjugate functions of bi¬ convex functions Let X and Y be two Banach spaces and (38) V: X xy->Ru{ + oo}a proper lower semicontinuous, convex function We introduce the partial conjugate U defined on X x y* by (39) U(x, q):=sup [{q, y) - V{x, y)] yeY which is equal to —co on K:={x e Xl'^y e Y, V(x, >^)= + oo}. This is obviously a function satisfying (40) Vx 6 X, q-^ U(x, q) is lower semicontinuous and convex, e y*, X-* U(x, q) is concave. Therefore, by theorem 2, we can write (41) V(x, y)= sup [(q, y) - U(x, q)'\ ««r*
212 CH. 4, SEC. 4 CONVEX ANALYSIS AND OPTIMIZATION and (42) V*(p, i)=sup [(p, x)+U{x, q)] xeX because y*{p> (l)=sup [<p, x) + sup {{q, y) - V{x, y))] xgX yey If we assume that (43) U{x, q) is upper semicontinuous then we can also write that (44) U{x,q)=m[^\y*{p,q)-{p,x)-] P G X because by (40, ii) and theorem 2, x->^ — U{x, q) is the conjugate function of p^V*{p,q). Remark Conversely, starting with a concave-convex function U, defined on X-*Y*, we can define a “biconvex” function F on Z x 7 by (41). Assumption (43) implies that V is lower semicontinuous. The next result will be often used in the following proofs. ■ THEOREM 14 We posit assumption (38). The two following conditions are equivalent: (45) and (46) ip, q)edV(x, y) pedxi- U)(x, q) and yedqU(x, q) where dx(— U) denotes the subdifferential of the convex function x-^ — U(x, q) and dqU the subdifferential of the convex function q^ U(x, q). A Proof a. We begin by proving that (46) implies (45). We can write (46) in the form Uix,q)-\-V{x,y)={q,y)
сн. 4, SEC. 4 CONJUGATE FUNCTIONS 213 and V*(p,q)~U{x,q)={p,x) By adding these equalities, we obtain (47) V*(p, q) + V{x, y)={p, x) + {q, y) that is, (p, q) edV(x, y). b. Conversely, assume (45), that is, (47). Since we always have {q, y)^ V(x, >')+ U(x, q), we obtain V*(p, q)- U{x, q)^ (p, x), which by (42) implies that ped^{-U){x, y). By using equality V*(p, q) = (p, x)+U{x, q), we obtain V (jc, y)+U(x,q)={q,y),ih&X\s,y i/(x, q). ■ We set (48) dxU{x, q):= -d,(- C/)(x, q) and call it the superdifferential of the concave function x-* U{x, p). We say that (xo, qo)eX X y* is a saddle point of the function defined by (49) (x, q)^ U{x, q) - (po, x) - (q, yo) if for all (x, q)&X X Y*, . qo)- (Po, x) - (qo, yo> < U(xo, qo)~ {po, Xo) - (qo, yo) (50) |t/(x,, < U{xo, q)~ (po, Xo) - (q, yo) The following characterization in proposition 15 is obvious. PROPOSITION 15 We posit assumption (38). The following conditions are equivalent: (51) and (52) Poe- дус{ - U)(xo, qo) and у о e t/(xo, qo) (^o> qo) is a saddle point of the function defined by (49) In particular, (x, q) is a saddle point of U if and only if (0,0) 6 dxU(x,p) x dq U{x,p). к PROPOSITION 16 We posit assumption (38). The set-valued map (x, q)^dx{— U)(x, q) x dqU{x, q) is monotone, к
214 CH. 4, SEC. 5 CONVEX ANALYSIS AND OPTIMIZATION Proof. This follows obviously from theorem 14. Indeed, let {pi, yt) e l/)(xi, gi) X dpU(Xi, q) (i=1, 2). Then {{Puyi)-{P2, yi), (^1. qi)-{x2, q2)) = {Pi -P2, xi - X2) + {qi -q2,yi - J2> = (iPl> gi)-(.P2y ^2), (^1, X2)-{yuy2)}>0 because (x, y)^dV{x, y) is monotone. ■ 5. THE SUBDIFFERENTIAL OF THE MARGINAL FUNCTION AND LAGRANGE MULTIPLIERS We consider a family of minimization problems (1) where (2) W(y):= inf V{x,y) xeX F is a proper lower semicontinuous, convex function from X y to i?u{ + oo}. Fix y and assume that x eX achieves the minimum of x-*V(x, y). We can characterize the subdifferential of the marginal function IF at y in the following way. PROPOSITION 1 Let y eY and x eX such that W{y)=V{x, y). Then the following conditions are equivalent: (3) q^dWiy) and (4) {%q)BdV{x,y) k Proof We observe that VF*(^)=sup [{q, y) - inf V{x, y)] ye y xeX =sup sup [<0, x) + (q, y) - V(x, v)l = F*(0, q) xe X ye y
CH. 4, SEC. 5 THE SUBDIFFERENTIAL OF THE MARGINAL FUNCTION 215 Hence, q e dW(y) if and only if 9)' <0, x) + {q, y) - V{x, y) = {q, y) - W{y) = W*{q) = F*(0, q) that is, if and only if (0, q) e dV(x, y). ® We consider a more specific problem. Let i i. F be a set-valued map from A" to F with a closed convex graph (5) i ¡1. t/ be a proper lower semlcontinuous, convex function from I XxF toFu{-l-oo} We set f U(x, y) if X 6 F " ^(y) (6) n^.F):=C/(^.F)+«/'g™ph,f-,(^.F)=|^oo ifxiF-‘(y) so that the minimization problem (1) can be written in the form (7) W{y)\= inf U{x,y) xeF~^(y) THEOREM 2 Let us assume that (5) holds true. We posit (8) 0 6 Int(Graph (F)— Dom U) LetyeY and xeF ^(y) be such that W{y)= U{x, y). Thefollowing statements are equivalent'. (9) and q^dWly) (10) 3(p, r) edi7(jc, y) such that qeD(F '^Xy, x)*{p)+r A Proof. By proposition 1, qedWiy) if and only if (0, q)&dV{x, y)= 5( C/-I-(/'g„ph(F)(^. f))- Assumption (8) and corollary 4.12. imply that (0, q)e d U(x, y) + Wgraph(f) (^, y)- Then there exists (p,f)ed U(x, y) such that {-p,q-rie A^graph(f)(3c, y), that is, such that —peDF{x, y)*{r—q). We use the inversion formula r-qeDF{x,y)* \-p)=-D{F~ ^){y, x)*{p) ■
216 сн. 4, SEC. 5 CONVEX ANALYSIS AND OPTIMIZATION We mention the following corollary. COROLLARY 3 Let U: X-*Ryj{ + oo} be a proper lower semicontinuous, convex function and F a closed convex set-valued map from X to Y. We assume that (11) 0 €lnt(Dom Dom i/) Let W: Y^Rkj{-\-oo} be the marginal function defined by (12) Wly)= inf U{x) xeF~^{y) Let 3c 6 f ”^(v) achieve the minimum of U on F~^iy). The subdifferential of the marginal function W is equal to (13) dW{y)=D(F-^)(y,x)*dU{x) A Many minimization problems are set in the following form: (14) W(y):= inf U(x,y) xe L Axe M + y where LcX, M^Y mq closed convex subsets and A e Y) is a continuous linear operator. COROLLARY 4 Let X eL satisfy Ax e M-\-y and let W{y)= U{x, y) be a solution to the minimiza¬ tion problem (14). If we assume that (15) (0, y) 6 Int((l xA)L—{0} X M— Dom U) then the following statements are equivalent: (16) qedWiy) and (17) 3(p, r)edU(x, y) such that qer— NM(Ax—y) and A*q е^ + Л*г + Л^ь(^) (18) Proof We apply theorem 2 when F is the set-valued map defined by Ax—M when xeL F{x):= 0 when X it
CH. 4, sec.5 the subdifferential of the marginal function 217 Assumptions (15) imply assumption (8) of theorem 2. We recall that (19) DF{x, /i H- Nl{x) when q e Nm{Ax—y) {0 when q $ NM{Ax-y) by proposition 2.5. Therefore, (20) D(F~^)iy, x)*{p)=-NM{Ax-y)nA*~\p+NL(x)) ■ Remark Since the marginal function W defined by (1) is convex, it is sufficient to prove that it is subdifferentiable at y to establish the existence of q. The Robinson- Ursescu theorem allows us to find sufficient conditions for W to be sub¬ differentiable. ■ PROPOSITION 5 Let us assume assumptions (6) and (7). Assume moreover that (21) y e Int Im F Let X eF ^{y)be a solution to W(y)= U(Xy y). If we assume that (22) U is continuous at (x, y) then the marginal function W is continuous at y and thus subdifferentiable on the interior of its domain. k Proof The Robinson-Ursescu theorem (see theorem 3.3.1.) states that the set-valued map F“ Ms lower semicontinuous on the interior of Im(F). Therefore, proposition 3.2.19 implies that the marginal function W is upper semicontinuous at y and, thus, W is bounded above on a neighborhood of y. Therefore, the interior of the domain of W is nonempty. Hence, W is continuous (and thus subdifferentiable) on Int Dom W. ■ We now add supplementary characterizations of the subdifferential dW(y) of the marginal function involving the partial conjugates (23) and (24) h(x, ^):=sup l(q, y) - V(x, y)} yeY h*{p, >^):=sup [{p, x) - V{x, jf)] xeX = sup inf [(p,x}-(g,y)+h(x,q)'] xeX qeY*
218 CH. 4, SEC. 5 CONVEX ANALYSIS AND OPTIMIZATION Then theorem 4.14 and proposition 4.16 imply at once the following character¬ izations of the subdifferential of W. PROPOSITION 6 Let y e Y and x eX such that W{y)= V{x, y). The following statements are equi- valent : (25) qedWiy) (26) (0, q) edV(x, ÿ) (27) X e d,ji*{0, y) and q edy(—h*)(0, j) (28) Oedf- h){x, q) and ÿ e d^h(x, q) Statement (27) is quite important, since it gives a way of obtaining both the optimal solution x and the subgradients of W in terms of the perturbation y. The regularity of this set-valued map is thus embodied in knowledge about the function h*. It is traditional to let the function h play an important role through the function €y defined on X xY* by (29) q)=(q,y}~h(x, q) called the Lagrangian of the minimization problem W{y). ■ PROPOSITION 7 a. The conjugate function W* of the marginal function W defined by (1) is equal to (30) W^*(^) = sup h{x, q)= F*(0, q) xeX b. The marginal function W is lower semicontinuous if and only iffor all ye Y, (31) lF(v)=sup miey{x,q)=M sup^)-(x, ^) ▲ qer’xeX xeX ye¥* Proof We observe that formula (23) implies Wiy)= inf sup [_{q,y)-h{x, q)'\ xgX qeY* = inf sup €y{x, q) xeX qeY*
CH. 4, SEC. 5 THE SUBDIFFERENTIAL OF THE MARGINAL FUNCTION 219 On the other hand, W^%)=sup [_{q,y)~ inf xeX =sup sup [(q, y) - V(x, >')]=sup h{x, q) xeXyeY xeX =sup sup [(0, x) + {q, y) - V{x, ;^)] = F*(0, q) xe XyeY and 1^**0;):= sup [(q,y}-W*{qj] qeY* = sup inf [{q, y)—h(x, q)~\ = sup inf €y{x, q) qeV xeX qey xeX Since the function W is convex, theorem 4.2. states that it is lower semicon- tinuous if and only if W=W*, that is, if and only if formula (31) holds true for all j s y ■ When the function x-*€^x, q):= {q, y) - h{x, q) is simpler to minimize than the function x->^V{x, y), it can be useful to replace the problem W{y) by a problem of the form (32) W^y{q):= inf €yix, q) xeX We note that formula (30) implies (33) Wf{q) = (q,y)-v*{0, q)=(q, y)-W*{q) and formulas (31) and (32) imply (34) WqeY*, W^,%)<W'(y) Hence, we distinguish the elements qeY* (if any) for which (35) W(y)=W*{q):=inUy{x,q) by calling them Lagrange multipliers of the minimization problem (1).
220 CH. 4, SEC. 6 CONVEX ANALYSIS AND OPTIMIZATION Remark The problem of finding Lagrange multipliers is called the dual problem of the minimization problem W(y). When the function W is lower semicontinuous, this amounts to maximizing the function Wy^{q). Hence, in this case, we can solve the minimization problem W{y) by (a) First, finding q that maximizes the function q-^Wf(q). (b) Second, minimizing the function x^^y{x, q). The existence of Lagrange multipliers can be proved in the framework of minimax inequalities. For the time being, we emphasize an important—and simple—property of Lagrange multipliers. ■ PROPOSITION 8 The set of Lagrange multipliers of the minimization problem W(y) coincides with the subdifferential dW{y) of the marginal function at y, A Proof Formula (33) implies that {q, y) - W*(q)=M£y{x, q)= Wf{q) Hence, q&dWiy) if and only if {q, y) = W*{q)+W{y)\ that is, if and only if Wy%)=W{y). Remark Theorem 3.17 implies that there exist Lagrange multipliers whenj^ e Int Dom W. When y is a Hilbert space, the set of elements y for which there exist Lagrange multipliers is dense in Dom W by theorem 3.11. 6. CONVEX OPTIMIZATION PROBLEMS We consider (1) i. two Banach spaces X and Y ii. two proper lower semicontinuous, convex functions (7: Z^Rul + oo} and F:y^Ru{ + oo} ill. a continuous linear operator A e V) iv. two parameters/? eTf* ande y We shall study the class of minimization problems (2) WO):= inf [t/W- {p, x) + ViAx+yi] xeX
CH. 4, SEC. 6 CONVEX OPTIMIZATION PROBLEMS 221 to which we associate the dual problems (3) W®{p)-.= inf \U*(-A*q+p)+V*(q)-{q,y)} qey* We observe that the Fenchel inequalities imply that (4) p^X*, W{y)+W®(p)^(> Note that W{y)< +oo if and only if (5) y 6 Dom V —A Dom U and W*[p)< + 00 if and only if (6) p eA* Dom V* + Dom U* Inequality (4) and conditions (5) and (6) imply that infima W(y) and W®(p) are finite. Existence of solutions to minimization problems W{y) and W®{p) and the main properties of their solutions follows from assumptions slightly stronger than conditions (5) and (6), which we assume to be satisfied. THEOREM 1 We posit assumptions (1). a. Assume that X is reflexive and that (7) pe\ni(A* Dom F* + Dom U*) Then there exists a solution x to the problem W{y) and equality W{y)-\- W®(p)=0 holds true. b. Assume that (8) y 6lnt(Dom V—A Dom U) Then there exists a solution q to the problem W®{p) and equality fF(y)+ holds true. c. Assume that both assumptions (7) and (8) are satisfied. Then x is a solution of W{y) and q is a solution to W®(p) if and only if they are solutions to the system of inclusions (9) pedU{x)+A*q y € -Ax+dV*{q)
222 CH. 4, SEC. 6 CONVEX ANALYSIS AND OPTIMIZATION d. Assume that both assumptions (7) and (8) hold true. The following conditions are equivalent: (10) i. X eX is a solution to W{y). il. p edU{x)+A*dV{Ax+y) iii. xedW®{p) e. Assume that X is reflexive and that both assumptions (7) and (8) are satisfied. The following conditions are satisfied. (11) i. q eY* is a solution to W®{p). ¡i. y€dV*(q)-AdU*(-A*q+p) iii. q^dWiy) Proof, a. The two statements follow from part c of corollary 4.12 to theorem 4.11. For instance, the second statement is obtained by taking in corollary 4.12 the function W to be defined by W(z)=V(Az+y); assumption (8) implies assumption (4.29), and W®{q)=V*{q)— (q, 3^). b. Equality W{y) + W®(p)=0 implies that solutions x and q to the minimiza¬ tion problems Wly) and W®(p) satisfy (t/(3c)-l- U*{-A*q+p)-(p-A% x))+{V(Ax+y)+ V*{q)-{q, ^x-i-;^»=0 Since each term on the left-hand side of these equations is nonnegative, they are both equal to zero, which means that inclusions (9) hold true. Conversely, (9) implies that lF(y)-)- W®(p)=0, so that 3c and q are solutions to the minimization problems. c. Theorem 3.5 shows that when X is reflexive, assumption (7) implies that 3c is a solution to W(y) if and only if inclusion (10, ii) holds true and assumption (8) implies that ^ is a solution to lT®(p) if and only if inclusion (11, ii) is satisfied. d. Proposition 5.1 states that q belongs to dW(y) if and only if (0, q) belongs to the subdifferential at (3c, y) of the function (x, ;^)^ t/(x)- {p, x) + V{Ax+y)=L(x, Ax+y) where L{x, z):= U(x)— (p, x) + V{z). Assumption (7) implies that 0eInt(DomL —Im(l x(yl©l)) where (1 x (1 ©^4)) denotes the continuous linear operator defined by (1x(>1©1))(x,;^):=(x,z1x-H;^) Since its transpose is equal to (1 ©^4*) x 1, we obtain that q belongs to dW{y) if and only if there exist r €dU(x) and s edV(Ax+y) such that 0=r-p-\-A*s
СН. 4, SEC. 6 CONVEX OPTIMIZATION PROBLEMS 223 and q=s, that is, if and only if Ax+у edV*(q) and x edU*(p —A* q).By elimin¬ ating X in these inclusions, we find that q belongs to dW(y) if and only if q solves inclusion (11, ii). We use the same arguments for proving the equivalence between statements (10, ii) and (10, iii). ■ COROLLARY 2 Let X and Y be Hilbert spaces, A 6 У) and U\ X^R\j{ + <x>} be a proper lower semicontinuous, convex function. We denote by J e^(Y, У*) the duality map from Y onto Y*. We consider the minimization problem (12) u:=mf There exists a Lagrange multiplier and the following conditions are equivalent: (13) X eX minimizes x-^ on X. (14) xsX is a solution to 0 eA"^JAx-]rdU{x\ (15) The solutions qeT^ to 0 eq—JAdU*{—A*q) are the Lagrange multipliers, (16) xeX is related to a Lagrange multiplier q by the relation q = JAx and —A*q edU{x). If we assume that (17) 0 e Int(Dom f/* + Im ^4*) then such solutions x eX and q eY* do exist. A Proof We take/?=0, ;;=0, and V defined by V(y)=^^\\y\\^, whose domain is K Then dV(y) = Jy. V^(q)=M\l and dV^{q) = J~\ ■ COROLLARY 3 Let X and Y be Banach spaces, K<=-Y be a closed convex subset, U:X-^Ru{+^) a proper lower semicontinuous, convex function, and A e SP{X, 7). We consider the following minimization problem: (18) Let us assume that (19) v:= inf U{x) Axe к 0€lnt(/l Dom U—K)
224 CH. 4, SEC. 6 CONVEX ANALYSIS AND OPTIMIZATION Then there exists a Lagrange multiplier q, and the following conditions are equivalent: (20) xeK minimizes U(x) under the constraint Ax e K. (21) X e X is a solution to 0 e3C/(x) + /l*A^/c(^3c). (22) The solutions q eY* to 0 e daK(q) — Ad — A"^q) are the Lagrange multipliers. (23) X is related to a Lagrange multiplier q by the relation q € Nk{Ax) and OedU(x)-i-A*q Moreover, if we assume that X is reflexive and that (24) 0 e Int(Dom i/* + ^ ^b(K)) [where b(K) is the barrier cone of K\ then such solutions x eX and q eY* do exist. ^ Proof. We take V(y) = ^K(yl Then #ic(y)=^ic(p), V*iq)=(^Ki^X and Dorn V*=b{K). ■ These kinds of results, involving Lagrange multipliers, are known in mathe¬ matical economics as decentralization principles. Specifically, let us consider (25) n Banach spaces X, n proper lower semicontinuous, convex functions i7,:A"-^i?u{ + oo} n continuous linear operators Ai e S^{Xi, Y) The problem (26) y:= inf ( t I ^ \i =1 \l == 1 ^ is decentralizable in the following sense: Lagrange multipliers exist provided that (27) 0 6 Int X Dorn t/,- - Dorn V If 96 y* is such a Lagrange multiplier, any optimal solution xeX satisfies (28) V/ = 1,..., n, Xi minimizes x,—> i/i(x,)- {q, ^iX,)
CH. 4, SEC. 7 CONVEX OPTIMIZATION PROBLEMS 225 Optimal solutions do exist when the Banach spaces Z, are reflexive and when (29) VxiSXi, there exist A > 0 and ^eDomF* such that for all I = 1,, n, J.Xi — A*qi e Dom [/,• A Once the Lagrange multiplier q is known, problem (26) is split into n problems (28) of smaller dimension. 7. REGULARITY OF SOLUTIONS TO CONVEX OPTIMIZATION PROBLEMS We introduce (1) We take (2) i. two finite dimensional spaces X and Y Ü. a linear operator Ao from X to V iii. two proper lower semicontinuous, convex functions Í7: X-+/?u{ + oo} and L: y->7?u{ + co} i. yo 6lnt(Dom V—Ao Dom Í7) ii. po € Int(^o Dom V* + Dom t/*) We recall that the solutions (xo, qo)^Xx Y* of the optimization problem i. U(xo)+V{AoXo+yo)~ (Po> Xq) =min (i/(x)+ V{Aox-\-yo)~ (po, x:)) xeX H. U*(—Aoqo+Po)+^*(^o)—(^o>yo} (3) =min {U*{-Atq+po)+V*{q)- (q, yo)) qe ¥• Hi. C/(xo)+ U*{—A^qo+Po)+ y{AoXo+yo)+ l^*(io) = (po, Xo) + {qo, yo) are the solutions to the system of inclusions (4) i. poedU{xo)+Atqo ii. yo e -AoXo+dV*(qo) (See theorem 6.1.) We shall study the behavior of the solutions (xo, qo) to this system with respect to the parameters po, Jo. and Ao-
226 CH. 4, SEC. 7 CONVEX ANALYSIS AND OPTIMIZATION For that purpose, let us denote by F“ ^(/?, y, A) the subset of solutions (x, q) to the problem (5) i. p €d[/(x) + A*q ii. ye-Ax + dV*(q) We use the definition of the generalized second derivative of a proper convex function introduced in dealing with nonsmooth analysis (definition 7.4.1). DEFINITION 1 Let U: 2i-^Fu{ + oo} be a proper convex function. Assume that U is subdiffer- entiable at xq and let po^^ be a subgradient of U at Xq. We shall say that the derivative of the set-valued map x-^dU(x) [/(xo, po) •= Cd f/(xo, po) is the second derivative of U at (xq, po). A Then U(xo, po) is a monotone closed convex process from X to X*. We set d(y4, F):=sup inf l|x-;;|| xe A ye В We now make precise what we mean by Lipschitz behavior. DEFINITION 2 Let F be a proper set-valued map from X to Y and let (xq, ;^o) belong to the graph of F. We say that F is pseudo Lipschitz around (xo, j^o) if there exists a neighbor¬ hood iV' of Xo, two neighborhoods Ш and Y of уо, ^ iT, and a constant ^ > 0 such that i. Vx6Ti^, F{x)n^^0 H. Vxb X2 e IT, d(F(xi)n^, F(x2)n Г-X2II A We can now state the regularity theorem. THEOREM 3 We posit assumptions (1) and (2). Let (xq, ^0) be a solution to problem (3). We assume that the monotone closed convex process from X to Y to itself defined by id^U{xo,po-A%qo) V -Ao S^ViAxo+ycqoY is surjective.
CH. 4, SEC. 7 CONVEX OPTIMIZATION PROBLEMS 227 Then (6) F~^ is pseudo Lipschitz around (po, yo, Aq, Xq, ^o)- Furthermore, the derivative of F~^ is defined as follows: (7) (dx, dq) eCF~^ (po, yo, Aq ', Xq, qoK^P, ^y, ^A) if and only if d^U(xo,Po-A%qo) AS (8) /Sx\ (d U) A A '^i^p—^A*‘qo\ -Ao d^V{Axo+yo,<lo)~^) \Sy+SA'Xo/^ This theorem is a consequence of the inverse function theorem 7.5.2. applied to the map F defined by (5). For simplicity, we set G(x):=5 U{x) and H{x)=dV(y), so that H~^{q)=dV*{q). Let F be the map from Xxyto^fxYx L{x, y) defined by (9) if and only if (10) {p, y, A) 6 F(x, q) f i. peG(x)+A*q |ii. y e — Ax+H~ Hi) We shall characterize the derivative of F in terms of derivatives of the set-valued maps G and H (or H~ H, respectively, in Section 7.5. Our theorem follows then from theorem 7.5.2. and lemma 7.5.9., which we state here. LEMMA 4 Let Xo, io a solution to the system of inclusions (11) i. poeG(xo)+ASqo ii. yoe -Axo + H~^{qo) The following conditions are equivalent: (12) (13) (bp, by, bA) € CF(xo, io> yo, Ao)(bx, bq) i. bp—bA*'qo & CG(xo,Po~ Aoio)(^x)+A%bq ii. by + bA‘Xo e -AoSx+CH~‘(^o.yo+Axo)(bq)
228 CH. 4, SEC. 7 CONVEX analysis and optimization Example 5 We consider a minimization problem with equality constraints, defined by (14) We take (15) i. two finite dimensional spaces X and Y ¡1. a linear operator Ao from X to 7 iii. a lower semicontinuous, convex function U from X to R i. yo e -Int {Ao Dorn U) ii. 0 6 Int(Im A§ + Dom U*) Let Xo be a solution to the minimization problem i. Axo=-yo (16) ii. U{xo)= min U{x) = - >10 and qo the associated Lagrange multiplier. Assume that (17) i. ^0 is surjective. ii. U is twice continuously differentiable at Xo and U{xo) is positive definite. Then F ' is pseudo Lipschitz around (0, yo, Ao, Xo, qo)- We set (18) i. J{xo)={Ao'V^U{xo)-'At)-^ ii. Ao = U{xo)~ '/4gy(xo), which is a right inverse of Ao ^ iii. q*®-Ae£F{x,y)-*A*qeX* iv. X®: A e SF{x, Y)-yAx e Y The derivative of the map f " Ms given by the formula '{i-AoAo)'^^U{xo) ^ -Ao -qt® J{Xo) .Vo ® (/loM* bp by bA Proof. We apply theorem 3 to the case when V is the indicator of {0}. Then d^V*{qo, AoXo+yo)=^^y*i<io, 0) is the constant map equal to zero. So inclusion (8) can be written 'bx\ (V^U{xo) Al\^(bp-bA*qo (bx\JV^U{xo) At\-^f> \bq) ^ -^0 0^ V which can be inverted explicitly. by-YbAqo
CH. 4, SEC. 7 CONVEX OPTIMIZATION PROBLEMS 229 Example 6 We consider the items defined by (14). We set Y:=Ji" and we take (19) i. yo e —Ao Dom U—R"+ ii. 0 e Int {A^R\ + Dom U*) Let xo be a solution to the minimization problem (20) i. ¡1. U{xo)= min U{x) < 0 and let ^o^O be an associated Lagrange multiplier. We denote by /1 the set of indexes such that (.^o^o+;^o)i=0. We posit assump¬ tion (17) and (21) V/e/i, ^O(>0 Then F is pseudo Lipschitz around (0, yo, Ao, Xo, <7o)- We write (22) i. R"=R'‘ X R’\ where /2 = {i = L ..., n|i /,} ii. q={q\q\ A=(A\A^) in. Ji(xo)=(AoV^U(xo)~'Ao')~^ iv. =V^i/(xo)“‘/4oVi(xo) Then the derivative of the set-valued map F ' at (0, Ao) is given by the inclusion written symbolically l{\-Ah^Ah)V^U{xo)-^ -Ah^ (^0^)* -/i(a:o) 4® ^ 0 0 0 Proof. We apply theorem 1 to the case when V is the indicator of the cone -/?+.Then is the indicator of/?+ and dF* = is the normal cone to 7?+. We take —yo eInt(/lo Dom V-\-R!\)=Ao Dom IJ+R\ (which is the Slater condition). Then (xo, qo) is a solution to the inclusion (23) i. Q=U{xo)-\-A%qo ii. yo€-AoXo + Nr»^ {qo)
230 CH. 4, SEC. 7 CONVEX ANALYSIS AND OPTIMIZATION The latter condition implies that (24) <^o. ^oJCo+Jo>=0 Since AoXo+yo e —/?+, we deduce that (25) if {A oXo + ;^o)i < 0, then =0 Then ql is equal to zero, and we assume that > 0 for all i e /i. By corollary 7.2.12., an element 5y of CNRn^iqo, AoXo+yo){^q) is defined by i. Forie/i, ¿y, =0 and is arbitrary. ii. For i € ¡2, dyi 6 R and Sqi is equal to zero. Let us write R"=R’' x R‘^ and q={q'^, q^). The domain of 3^F*(^o> ^^oXo+J'o) is R‘‘ X {0} and dV*{qo, AoXo+yo\dqu 0) = {0} xR'^ Hence, the matrix of second derivatives can be written symbolically 5p] /v^C/(xo) Ah' 5x dyi 6 -Ah 0 0 \ -Ah 0 ^0, Then it is surjective if and only if the matrix of linear operators iV^U(xo) Ah' V -Ah 0 from X xR}' to itself is surjective. This is the case by assumption (17). We can even invert the preceding inclusion explicitely and obtain the formula for the derivative of the map f " ■ Example 7 Still considering the items defined by (14), we introduce a closed convex subset Pof K We take (26) yo 6lnt(/*—/4o Dom U) 0eInt(^§6(P)+Dom U*) where b{P):= {^|sup,6i>(^, y) < + oo} is the barrier cone of P. Let xo be a solu-
CH. 4, SEC. 8 LAGRANGIANS AND HAMILTONIANS 231 tion to the minimization problem (27) i. Axo e P-yo [ii. l/(xo)= min U(x) Axe P-yo and go an associated Lagrange multiplier. We posit assumption (17) and the following assumption on P: (28) absolution to - J(xo)q € Cnp(Axo + + go)(y + (1 - A^o)k) Then the conclusion of theorem 3 holds true. A Proof. It is sufficient to check the surjectivity of V^U(xo) At — Ao CNp(Axo-hyof Po) ^ This method, using the inverse function theorem for set-valued maps, can be used to treat more general convex minimization problems. 8. LAGRANGIANS AND HAMILTONIANS The convex minimization problems W(y) studied in the two preceding sections are particular cases of minimization problems of the form (1) v:= inf L(x, Ax) xeX where X and Y are Banach spaces, A a continuous linear operator from X to and L a proper lower semicontinuous, convex function from Xxyto/?u{ + oo}. Because of their formal analogy with problems arising in the calculus of varia¬ tions, we may call the function L a Lagrangian. In the calculus of variations, as we shall see, spaces X and Y are spaces of functions or distributions and A is a differential operator. The following results, which use the transpose A* of Ay cannot be applied, because several differential operators cannot be trans¬ posed explicitly (depending on the spaces on which they are defined). These operators possess only explicit “formal transposes” related to /i by a Green formula instead of the usual relation {A*qy x) = (q. Ax). Still, it is worth begin¬ ning with the simpler case when A is an abstract continuous linear operator, because it embodies all the main ideas that we shall adapt in the case of calculus of variations (see Chapter 8).
232 CH. 4, SEC. 8 CONVEX analysis and optimization This being said, we associate with problem (1) its dual problem (2) y*:= inf L*{-A*q,q) qeY* We observe that the function defined on X xY* by (3) q):=L(x, Ax) + If(-A*q, q) is always nonnegative, because L(x, Ax)-\-I^(-A'^q, q)^{-A'^q, x)-\-{q, Ax)=0 Hence, (4) v-\-v*= inf j!^{x,q)>0 (x,q)eX xY* We say that an element ^ g 7* is a Lagrange multiplier if 11, , + „*=0 ^ [ii. V* = L?{ — A*q, q) THEOREM 1 Let us assume that L is a proper lower semicontinuous, convex function from X xY to Ru{-\-oo} satisfying (6) 0 6 Int((yi 0 — l)Dom L) Then there exists a Lagrange multiplier q of the minimization problem (1), and the two following conditions are equivalent: (7) and (8) xeX minimizes x->L(x, Ax) X eX is a solution to the inclusion 0 g (1 ®A"^)dL{x, Ax) Proof a. Let V be the function defined by V{x):=L{x, Ax). Hence, v = m{xexVM= - L*(0). By part (b) of theorem 4.11., there exists ^ g 7* such that K*(0) = L*( —.4*^, ^)= inf L*( —= qe Y* This means that ^ is a Lagrange multiplier.
CH. 4, SEC. 8 LAGRANGIANS AND HAMILTONIANS 233 b. We know that xeX minimizes K on X if and only if OedV(x) = (1 ©^*)3L(3c, Ax\ by theorem 3.7. ■ We now mention a list of useful—and simple—equivalent statements. We introduce the Hamiltonian of the problem y, which is the function H defined on X X y* by (9) H(x, q}:=sup [{q, y) - L{x, j^)] PROPOSITION 2 The following statements are equivalent: (10) xeX minimizes x-^L(x, Ax) on X, and qeY* is a Lagrange multiplier (11) sé{x,q)=Q(= min sé(x,q)) (x,q)eX xY* (12) i-A*q,q)edL(x,Ax) (13) A"^q e dxH(x, q\ and Ax e dqH(x, q) (14) X is a solution to the inclusion 0 e (1 ©/4*)3L(x, Ax) (15) q is a solution to the inclusion 0 g ( — © 1 )5L*( — A q) k Proof The implications (10)=»(11)=>(12)=>(10) are obvious. Equivalence between statements (12) and (13) follows from theorem 4.14. Statement (12) is equivalent to (14) by eliminating q and equivalent to (15) by eliminating x in the relation (x, Ax) g ôL*( - .4*^, q\ ■ Remark When L is Gateaux differentiable at [x, Ax\ the inclusion 0 g (1 ®A*)dL(x, Ax) becomes the equation (16) V;cT(x, Ax)-hA*VyL(x, Ax)=0 If A=d/dt (formally), so that Ax=x, we recognize the Euler-Lagrange equation from the calculus of variations. For this reason, we shall call the inclusion 0 G (1 ®A*)dL(x, Ax) the Euler-Lagrange inclusion. When the Hamil¬ tonian H is Gateaux differentiable at {x, q), inclusion (13) can be written (17) A*q=VxH{x,q) and Ax = Vq{5c,q) We recognize the Hamilton equations', therefore, relations (18) A^^qe dxH(x, q) and Ax g dqH(x, q)
234 CH. 4, SEC. 8 CONVEX analysis and optimization will be called the Hamilton inclusions. When the dual function L* is Gateaux differentiable at { — q\ inclusion (15) becomes (19) AS7pL^(-A% q)-S7,L*i-A% q)=0 We shall call the relation (20) Oe{-A®i)dL*i-A%q) the dual Euler-Lagrange inclusions. ■ We observe that a Lagrange multiplier of the dual problem y* is a solution to the minimization problem v. By applying the preceding theorem 1 to the dual problem [where X is replaced by Y*, Y by X*, Ahy A*, Lhy {q, /?)->L*(p, q\ etc.], we obtain THEOREM 3 Let us assume that X is a reflexive Banach space and that (21) 0 6lnt((l©yl*)DomL*) Then v = v*, and there exists x eX such that (22) v=L{x, Ax)=mm L(x, Ax) A xeX When both assumptions are satisfied, we obtain the existence of both an optimal solution and a Lagrange multiplier. THEOREM 4 Let us assume that X is a reflexive Banach space and that (23) i. 0 e Int((/1 © — 1 )Dom L) ii. 0 € Int((l ©/l*)Dom L*) Then there exist xeX andp eY* satisfying the equivalent conditions (10)-(15). Remark We have used the terminology Lagrange multiplier in two different instances. However, the terminology is consistent when we regard the minimization problem V defined by (1) as the particular case v = W{^) of the family of minimiza¬ tion problems (24) W{y):= inf L(x, Ax-Vy) xeX
CH. 4, SEC. 8 LAGRANGIANS AND HAMILTONIANS 235 So if we set V(x, y):=L{x, Ax+y), we observe that h(x, q):=sup i(q, y) - V(x, y)]=H{x, q)-(q, Ax) yer and that ^y{x, q):= (q, y) - U(x, q)=(q,y+Ax)- h(x, q) There is, at this point, a slight terminology ambiguity concerning the word Lagrangian, which designates both the function L(x, Ax-\-y) we want to mini¬ mize (terminology coming from the calculus of variations) and the function (terminology used in optimization). The functional of the dual problem is then equal to Wy*‘{q):= inf fyix, q)=(q,y)-L!’{-A*q, q) xeX Hence, for y=0, ^ is a Lagrange multiplier in the sense that W{0) = Wo^{g) if and only if v= — L*(—>4*^, g) that is, if and only if ^ is a Lagrange multiplier according to definition (5). ■ We now proceed with the case of convex Hamiltonians. Theorem 4 states that when the Lagrangian L is convex and the assumptions (23) hold true, the solutions X to the minimization problem (1) V = inf L{x, Ax) xeX are the solutions to the Euler-Lagrange inclusion (14) and, by duality, the solu¬ tions q to the dual minimization problem (2) v*= inf Lf{ — A*q,q) qeY* are related to the solutions x to the Euler-Lagrange inclusion (14) through the formula (3c, A5c) e dL*(—A*q, q). We can still use such variational principles (allowing us to solve inclusions by solving minimization problems) for solving the Euler-Lagrange inclusion when L is no longer convex but concave-convex. In this case, the Hamiltonian
236 CH. 4, SEC. 8 CONVEX ANALYSIS AND OPTIMIZATION H is a convex function. Therefore, from now on, we assume that (25) iH: X X y*^/îu{ + oo} is a proper lower (semicontinuous, convex function to which we associate the Lagrangian L: x (note that the value — oo is allowed) defined by (26) L(x, y):= sup [(q, y) - H(x, q)'\ qeY* Such a Lagrangian L is concave with respect to x, convex and lower semicontinuous with respect to y. If L is assumed to be upper semicontinuous with respect to ;c, then the function H is its Hamiltonian in the sense that (27) H(x, 9)=sup [{q, y)-L(x, q)] yeY In this case, the Euler-Lagrange inclusion takes the form (28) 0 6 — 5x( - L){x, Ax)+A*dyL(x, Ax) and the Hamiltonian inclusion can be written (29) (A*q, Ax) e dH{5c, q) When the Hamiltonian is convex and lower semicontinuous, these equations are equivalent, because theorem 4.14 implies that the Hamiltonian inclusions (29) are equivalent to the pair (30) q e dyL(x, Ax) and A*q e dx(—L)(x, Ax) Eliminating q from (30) yields the Euler-Lagrange inclusion (28). Note that the function ^ defined on X x y* by (31) ^(x, q):=H(x, q)-(q, Ax) + H*{A*q, Ax)-(A*q, x) is always nonnegative because H(x, q) + H*(A*q, Ax)> (A*q, x) + (q, Ax) We obtain a result analogous to proposition 2. PROPOSITION 5 The following statements are equivalent: (32) , ^)=0(= min âS(x,q)) (JC,«)€X xT*
(33) (34) CH. 4, SEC. 8 LAGRANGIANS AND HAMILTONIANS 237 X is a solution to the Euler-Lagrange inclusion 0 e —dx(—L)(x, Ax)-\-A*dyL{Xy Ax) The pair (x, g) is a solution to the Hamiltonian inclusion {A*q, Ax) e dH(x, q). A Proof. Conditions (32) and (34) are obviously equivalent, and we already mentioned that statements (33) and (34) are equivalent. ■ We shall now prove that we can solve these equations by solving either the abstract least action principle (35) w:= inf (H{x, q) — (q, Ax}) (x,q)eXxY* or the dual least action principle (36) w*inf (H*{A"^q, Ax)— (A*q, x)) {x,q)eXGY* THEOREM 6 We assume that the Hamiltonian is proper, convex, and lower semicontinuous. Any minimizer {x, q) of the least action principle (35) solves the Hamiltonian inclusion (34). If we assume that (37) 0 e Int(Im(y4* x .4)- Dom //*) one of the minimizers to the dual least action principle (36) also solves this Hamil¬ tonian inclusion. k Proof, a. It is easy to check that if (3c, q) is a solution to the minimization problem (35), then 'i{x,q)BX xY*, (A *q, x) + {q, Ax}^D+H{x, q)[x, q) Therefore, {A*q, Ax) belongs to dH{x, q). b. Let (x, q) be a solution to the dual least action principle. We easily deduce that, V(x, q)eXx Y*, (A*q, x) + {q. Ax)^D^H*{A*q, Ax){A*q, Ax) We apply corollary 3.6 and obtain (A*x, Aq)e(A xA*)dH*(A*q, Ax)
238 CH. 4, SEC. 8 CONVEX analysis and optimization Then there exists (3c, q) such that {x, q) edH*(A*q, Ax\ A*q = A*q, Ax = Ax that is, such that (x, q) e dH*(A*q, Ax) Since (3c, q) is also a solution to the minimization problem (36), the last statement ensues. ■ Unfortunately, neither the least action principle nor its dual are convex minimization problems, because we add to the convex functions I/(x, q) and Ax) the bilinear function — {q. Ax), But the least action principles do retain enough structure to allow us to solve the Hamiltonian inclusions (34) even when the Hamiltonian is convex. Since H and //* have dual properties, we shall be able to choose among the least action principle and its dual the one that has a solution. In Chapter 8, these ideas will be developed, and methods will be given for solving the Hamiltonian inclusions by finding critical points (not necessarily minimizers) of the dual least action principle.
CHAPTER 5 A General Variational Principle This chapter introduces Ekeland’s s-variational principle (corollary 3.2), which is an important tool for nonlinear analysis, as we shall see in subsequent chapters. We have tried to motivate its use as much as possible, by providing an appealing proof and giving early and powerful applications. Our proof relies on the notion of dissipative dynamical system. Along the way, we prove Caristi’s fixed-point theorem and Baillon’s non¬ linear mean ergodic theorem. Among applications, we give a strengthened version of the Ambrosetti- Rabinowitz “mountain pass” theorem. This is the simplest (and most powerful) method known for finding critical points that are not local maxima or minima. Its use is illustrated in the last chapter on Hamiltonian systems, as will the global inverse function theorem we give in example 5.8. The last sections give applications to the geometry of Banach spaces and to optimization theory. These are very active fields of research, and many related questions are still open. 1. WALKING ON COMPLETE METRIC SPACES Let us start somewhere in a complete metric space X and walk around. At each point xeX, the rules of the game specify the set F{x) of points that can be reached from x in one step. Starting from, say, xq e X, we reach xi e F{xo) at the first step, then X2 6 F{xi) at the second step, then хз € Т(хг) at the third step, and so on. The question is whether we shall end up somewhere. We formally introduce definition 1. DEFINITION 1 Let X be a complete metric space. A dynamical system on X is a set-valued map X with F{x)^0 for all x. Any infinite sequence х^=Хо, Xi,..., x„,... nt I F:X such that (1) x„+ieF{x„), for all n 239
240 CH. 5, SEC. 1 A GENERAL VARIATIONAL PRINCIPLE is called a motion starting at Xq. The set (2) is the trajectory of this motion. The set ^x)= u{C(x^)\xo =x} is the forward cone of the point x in X. A Here the integer n must be understood to denote successive times, or steps; it runs from zero (initial time) to infinity. In general, starting from any point x in X, several motions are possible unless F is single valued. For instance, if F(x) contains two different points, say ye F(x) and zeF(x) with yfz, then we can start two different motions from x, say yj^ and z^, to wit (3) (4) yo=x, yi =y, y2 e F(y\ and so on, inductively zo=x, zi =z, Z2 e F(z), and so on, inductively A trajectory is the set of points through which an individual motion will run. Note that knowing the trajectory does not give us full information about the motion, since we are not told when an individual point will be reached. The for¬ ward cone of a point x is the set of all points in X that can be reached in a finite number of steps if we start at x It has the obvious properties (5) i. X e ^(x) (reflexivity) ii. \_y 6 ^(x) and z e ^(y)] => z e ^(x) (transitivity) Indeed, if€ ^{x\ there is a motion xt and an integer n such that xq =x and Xn=y. Since z e ^(y\ there is a motion y^ and an integer k such that yo=y and yk=z. Setting Zp—Xp for p^n and Zp=yp-„-i for p'^n-\-1, we obtain a sequence zt such that ^0 =-^0» ^k + n+ 1 — and ZpeF{zp-i) for all p So Zt is a motion, and z e ^(x). We are interested in finding motions that converge. Let us start with a simple criterion. It is stressed that all the arguments rely heavily on the fact that the metric space X is complete. PROPOSITION 2 Assume there exists a nonnegative function U: X->IRu{4-oo} such that (6) Vx e X, ^y e F{x): U{y)-\-dix, y)^ U(x)
CH. 5, SEC. 1 WALKING ON COMPLETE METRIC SPACES 241 Then for all points x where U{x)< +oo, there is a motion xt starting at x that converges to a limit point x: (7) Xq = x and Xn-^x If the graph of F is closed, then x is a fixed point. ^ Proof Construct by induction, using assumption (6) at each step, a motion xt such that (8) (9) d{xn, +1)^ i/(xJ - U(Xn +1), for all n This implies that U{xt) is a decreasing sequence of real numbers. Since all its terms are positive, this sequence has to converge. Adding up both sides of inequality (9) from A2 = A^to/t = M—1, we have (10) M- 1 X d{Xn,Xn+])^U{Xi4)-V(XM) n = N Using the triangle inequality, this yields (11) d{xN,XMHU{xN)-U(xM) Since the sequence C/(x„) converges, the right-hand side can be made smaller than any prescribed e>0 by choosing N<Mlarge enough. This proves that xt is a Cauchy sequence, and since the space is complete, it converges. Assume now that the graph of F is closed. We know x„+iS F{x„X since jct is a motion. It follows that the pair (;>c„, Xn+i) belongs to the graph of F. Since (x„, x„+i) converges to (x, x), the latter pair must also belong to the graph of F. Hence, the result. ■ Of course, when F is single-valued, F(x) = {f(x)} for all x, and the statement of proposition 2 can be simplified: PROPOSITION 3 Let f: X-^X be single-valued. Assume there is a positive function U: X-^Ru {+ 00} such that (12) Vx 6 X, U{f(x)) + d{x, f{x)H U(x) Then for all points x where U{x)< -f oo, the sequence x^ defined recursively by (13) Xo=x, x„ +1 =/ (^m) =r\xo)
242 CH. 5, SEC. 1 A GENERAL VARIATIONAL PRINCIPLE converges to some limit x. If f has a closed graph (for instance, if f is everywhere finite and continuous), then x is a fixed point (14) f(x)=x No proof is needed, since this is just a transposition of proposition 2. Criterion (6) is easy to understand. Think of the space X as the map of a mountain range, with U{x) as the altitude of point x. Then inequality (9) implies that U(x„+i)^ U(Xn)y which means going down. Moreover, we have U(x„)-U(x„^,) d{Xn, Xn+i) which means losing height at some fixed rate. It is intuitively clear that we cannot go on like this indefinitely (unless, of course, we fall into a bottomless pit; this is why we have to assume С/(л:)^0 everywhere). We have to end up somewhere. Criterion (6) is useful in so far as we can construct a function U. When F has nice continuity properties, we know even the smallest function U satisfying property (6). PROPOSITION 4 Let us associate with a strict set-valued map F from X to X the nonnegative func¬ tion Uf from X to U + u{ + co} defined by (15) WxeX, Uf(x):= inf ^ d(x„,x„+i) {x^/xo = x} n-0 If F is upper semicontinuous with compact values, then Up satisfies property (6) (16) 'ix eX, 3_v 6 F(x): Uf(x)> UF(y)+d(x, y) Moreover, if U is any other function satisfying property (6), we have (17) VxeX, Uf(x)^U(x) A Proof Inequality (17) is obvious. Let us associate with any e > 0 the function i/e defined by (18) t/e(A:):=inf< Y d(x„, x„+i)\xo=x and x, (n = 0 and the function Uo defined by „ +1 e B(F(x„), e)| t/o(A^):=4m C/£(x)=sup i/t(x) £-»0 £>0
CH. 5, SEC. 1 WALKING ON COMPLETE METRIC SPACES 243 These properties are well defined, for if then f/^, < Uq^ Up- Next, for each £ > 0, select Xe 6 B(F{x), e) such that (19) Ut(xc)+d{x, Ut(x)+e Since F{x) is compact, we can select a sequence x* e B{F{x), i/k) converging to some X e F{x). Also, since F is upper semicontinuous, for any 5 > 0 there exists ko>i/S such that for all k>ko. В ■■B{F{x), d) Thus for k^ко, (20) Us{x)^d(x,Xk)+Uiiu(Xk) Combining (19) and (20), we have for k>ko, Ub(x)-d{x, Xfc)+rf(x, Xk)^ Uiik{x)+l/k Letting A: go to 00 in this inequality, we obtain U3{x) + d{x, x)^ Uoix) and finally, since d is arbitrary, we find that for some x eF{x) (21) C/o(-i) + d{x, x) ^ Uo{x) Hence, Uo satisfies property (6) and thus, by the first part of the proposition, Uq, Since Uo^ Up, we deduce that Uo=Up. ■ DEFINITION 5 A set-valued map F is called a contraction if it is Lipschitz with constant Я e ]0,1 [. It is nonexpansive if X = l. к THEOREM 6 A (set-valued) contraction from a complete metric space X to itself has afixed point. Proof Since F is Lipschitz—and thus upper semicontinuous—with com¬ pact values, the function Up defined by (15) satisfies property (6). The function Up is also finite. To see this, we select a motion such that d(Xn+u Xn)=d(Xn, F(xn))
244 CH. 5, SEC. 1 A GENERAL VARIATIONAL PRINCIPLE which is possible because the values of F are compact. Since F is Lipschitz, we observe that 00 oo j Uf{x)^ X! d(Xn+i, x„)<: X ^"d(xo, ;ci)=-—T d(xo, Xi)< + oo n = 0 n = 0 1—A Our theorem follows from proposition 2. ■ When F is simple valued, this is the well-known Banach contraction principle. THEOREM 7 Any single-valued contraction from a complete metric space to itself has a unique fixed point. k Proof Uniqueness is a special feature here. Assume there are two fixed points X and y, and apply the Lipschitz condition. We have (22) d{x, j)=i/(/(x), /(y))<2i/(x, y) which is absurd since ^ < 1; therefore, x=y. ■ We now go back to the general, set-valued case, and we strengthen condition (6) to bring it closer to the single-valued case. DEFINITION 8 A dynamical system F: X-^X is called dissipative with respect to a function U: A"-^Ru{ + oo}, positive and not identically + oo, j/’ (23) Vx e X, "^yeF (x), U{y)+d(x, y) ^ U{x) In our mountaineering analogy, this means that there is no way to go but down. The terminology comes from physics: We can think of x as describing the state of a physical system, with U{x) its energy. Condition (23) tells us that whichever way the system evolves, it is going to lose energy at some fixed rate. It is then natural to think that the system will evolve toward a stable equilibrium. It is this idea that is expressed in the following proposition. PROPOSITION 9 Assume F is dissipative. Then for any x where U{x)< H-oo, there is a motion xt and a point x such that (24) Xo=x and {x}= Q neN Proof. For any y where U{y)< + oo, define (25) V(y)=m.{{U{z)\zB^y)}
CH. 5, SEC. 1 WALKING ON COMPLETE METRIC SPACES 245 Since U is positive and ^(y) nonempty, V(y) is some real number. Now take any xeX and ;; € ^{x). By definition of the forward cone ^(x), there is a se¬ quence X t and an integer k such that xq=x and Xk =y. Writing assumption (23) at each step «=0 to n=k — l and summing up, we have (26) Vj; e ^{xl d{x, yH U(x)- U(y) By definition of K this yields (27) Vj; e ^(x), d(x, y)^ U(x)— V{x) and, hence, for any xeX (28) diameter ^(x) < 2( U{x) - V{x)) We now consider a sequence (not a motion) in X defined inductively as follows. Start at yo=x, and pick y„+i e ^(y„) such that (29) U(y„^,HV{y,H2-^^ This is always possible by definition of V. By relation (5), we have (30) ^{yn^i)^^iyn) This implies K();„+i)^ K(y„). Let us estimate V{y„+i). We have by relation (29) (31) K(;;,^i)^^0;„ + i)^L0;„)+2-"^F0;„^i) + 2-” Hence, C/(;^M+i)““ 1^(Vm+i)^2"”. Setting x=y„+i in formula (28), we obtain (32) diameter ^(yn +1) ^ 2 ^ ” It follows that ^(y„), ne N,is 3. nested (decreasing) sequence of closed sets whose diameters go to zero. Since Z is a complete metric space, their inter¬ section is a singleton (33) ^xeX: f) ^{y„)={x} neN Since yn-n^ ^{ynX there is a motion xf and an integer k„ such that xS=7„ and xZ„=7„+i. Using motion x? to go from yo=x to 3^1 and then motion xt to go from yi to y2, and so on inductively, define a motion xt that will eventually carry us through all the points y„. Clearly, for any k, we can find n and m such that (34) Xk e ^{y„) and y„ 6 ^Xk)
246 CH. 5, SEC. 1 A GENERAL VARIATIONAL PRINCIPLE Hence, we have (35) and the desired result (24) follows from (33). ■ We have proved that the diameter of ^(x„) goes to zero, which conveys the idea of stability: For any e > 0, there is an integer N such that if a motion x't branches off from xt at a later time than N, it will stay within a distance e of x. We derive from proposition 9 several results pertaining to the existence of points X where F(x) = {x}. Such points will be called invariant. They are dead ends where the system is stuck. This is stronger than the fixed-point property, X 6 F{x), except, of course, in the single-valued case. COROLLARY 10 Assume F is dissipative and lower semicontinuous. Then there is an invariant point X where (36) F(x)={x} Proof. Start from a point x where U(x)< +oo, and define x as in proposition 9. Take any z in F{x). Since F is lower semicontinuous, a sequence z„ e F{x„) can be found converging to z. We have (37) Zk 6 ^{x„), for a\\k>n and, hence, letting ^->^00 (38) z 6 ^{x„\ for all n So z belongs to the Intersection of all ^{x„). By formula (24), we have z=x I COROLLARY 11 Assume F is dissipative and all forward cones are closed. Then whenever U{x) is finite, the forward cone of x contains an invariant point x. (39) U{x)< +oo=i>3x 6 ^(x):F{x)={x} Proof Start from a point x where U{x) < +00, and define x as in proposition 9. Since all the ^{x„) are closed, condition (24) becomes (40) {x}= '^(x„) neN
CH. 5, SEC. 1 WALKING ON COMPLETE METRIC SPACES 247 Since X 6 we have F{x)ci<^{x„) for all n. Hence, (41) F(x)= n ^ix„) neN The result follows immediately from (40) and (41). ■ The question naturally arises, how do we construct dissipative dynamical systems; there is a standard way to do this. PROPOSITION 12 Let U: X-^Ru{-\-co} be a positive function, not identically -f- oo. The dynamical system G defined by (42) G{x) = [y\U(yHd{x,y)^ U{x)} is dissipative with respect to U and has the property that (43) G{x)=^^(x\ for all X k Proof G(x) is not empty, since it certainly contains x itself. Thus formula (42) does define a dynamical system, and it is obviously dissipative with respect to U. Indeed, G could even be defined as the largest dynamical system that is dissipative with respect to U. Relation (43) is obvious when t/(x)=+oo, both sets being equal to X. Now assume U{x) is finite. We claim that G(y)<=G(x) whenever 6 G(x). Indeed, if z e G(yX we have (44) (45) i/(y) + rf(x,;;)^ [/(x) U(z)-i-d(y, z)^ U(y) Since [/(x) is finite, so are U(y) and U(z). Adding them up and using the tri¬ angle inequality, we obtain (46) U(z)-i-d(x, z)^ U(x) which means that z 6 G{x), as desired. It follows by inducation that ^(x)cG(x). Since the converse inclusion holds by definition, both sets coincide. ■ A dynamical system F is dissipative with respect to U if and only if F{x) c G(x) for all X Of course, any point x that is G invariant will also be F invariant (47)
248 cH. 5, SEC. 2 a general variational principle This simple remark leads us to a remarkable result. PROPOSITION 13 Let F be a dynamical system, dissipative with respect to some lower semicon- tinuous positive function + oo. Then F has an invariant point x (48) F(x)={x} Proof We associate with U a dynamical system G as in proposition 9. As previously noted, it is enough to show that G has an invariant point. Since U is lower semicontinuous, the set G(x) defined by (42) is closed and so is the forward cone ^(x) because of equation (43). Applying corollary 11, we see that G has an invariant point. A This last result is particularly striking, because it places no continuity require¬ ment on the mapping F itself. Let us rephrase it in the single-valued case to obtain Caristi’s original statement, which was the starting point of this investiga¬ tion. THEOREM 14 Let f: X^X be a {single-valued) map and U: X- tinuous function such that a {finite) lower semicon- (49) d{x, f(x))^ U(x)-U(f{x)\ for allxeX Then f has a fixed point x (50) f{x)=x k Note that/ is not required to be continuous and there may be several fixed points (take f to be the identity and U any positive constant). Note also that the proof of corollary 11 does not provide us with an iterative procedure for finding x. Indeed, if we choose some starting point x, there is no guarantee that the motion xt defined by x„=/"(x) will converge to anything, let alone to x. Finally, note that a strictly positive constant k can be introduced on the left- hand sides of formulas (6) and (49) to read kd{x, y\ without altering subsequent statements. Indeed, kd is again a distance on X, equivalent to the original one d, and we can use either one indifferently. 2. FIXED POINTS OF NONEXPANSIVE MAPS The Banach contraction theorem implies the existence of an unique fixed point of a contraction/ from a complete metric space K to itself.
CH. 5, SEC. 2 FIXED POINTS OF NONEXPANSIVE MAPS 249 When X is a closed, convex bounded subset of a Hilbert space, we can relax the contraction assumption and replace it by the assumption that/ is non- expansive. THEOREM 1 Let К be a nonempty, closed, convex bounded subset К of a Hilbert space and let f be a nonexpansive map from К to K. Then f has a fixed point. Furthermore, the subset of fixed points is closed and convex. A Proof. We shall write the proof in three steps by constructing a sequence of elements Xt^K satisfying (1) lim ||д:,-/(хО||=0 then by showing that the weak cluster points of such a sequence are fixed points of the nonexpansive map and finally, by proving that the set of fixed points is closed and convex. a. Let xo^K and i> 1. We define the single-valued map f from K to K by /,(x):=y xo + ^ fix) Since ll/rW-/,(y)||<( 1- y)ll/W-/(v)||<( the maps/ are contractions. Then there exist fixed points x, e X of the maps f ,=yXo + ^l- Therefore, 1 ^<-/WII=- lko-/W||<-sup {||x|| \xeK} t" t since K is bounded and/(x,) € K. Hence, lim ||x,-/(x,)||=0 i-» 00 b. Since Xt belongs to K, which is weakly relatively compact, there exists a generalized subsequence of elements x,- that converges weakly to some xsK.
250 CH. 5, SEC. 2 A GENERAL VARIATIONAL PRINCIPLE We set xa :=(1 — 2)Jc + 2/’(x), where A e ]0,1 [. Since/ is nonexpansive, we deduce that <xa -/(xa) - (x,- -/(x,.)), a:a - X,.) > I |xa - x,-| p -1 |/(xa) -f{x,)\111Xa - x,.| | ^ 0 By letting X,' converge to x, this inequality implies that (xa-/(xa)-0, xa-x)^0 because x,--/(x, ) converges strongly to zero. By dividing by A > 0, we obtain <(1 -A)x+A/(x)-/((1 -A)x + A/(x)), fix)-x)^0 By letting A converge to zero, we deduce that ||/(x)—x||^<0, that is, x=/(x) Hence, there exists at least a fixed point off. c. The subset of fixed points is obviously closed, because/ is continuous. Let xo and Xi be two fixed points off, and let us prove that xa:=(1 — A)xo+Axi is also a fixed point for A 6 [0,1]. We have ||/(xa)-XoII = ||/(xa)-/(xo)||<I1xa-XoII=A||xo-Xi|| and ||/(xa)-xi|| = ||/(xa)-/(xi)||<1|xa-Xi||=(1-A)||xo-xi1| It follows that ||/(xa)-Xo|| + ||/(xa)-Xi||<||xo-Xi||^||/(xa)-xo|| + ||/(xa)-Xi|| that is ||/(xa)-Xo||=A||xo-Xi|| and ||/(xa)-Xi||=(1-A)||xo-Xi|| Since Af is a Hilbert space, we deduce that /(xa)=(1 -A)xo+Axi =Xa that is, xa is a fixed point. ■ It is not necessarily true that the sequences of elements /"(xo) converge to some fixed point when / is nonexpansive. But we shall prove that the Cesaro
CH. 5, SEC. 2 FIXED POINTS OF NONEXPANSIVE MAPS 251 means (2) >'r:=4 S f‘(xo) i i=l converge weakly to a fixed point. For that purpose, it is convenient to introduce the concept and the properties of the asymptotic center of bounded sequences of a Hilbert space. It always exists, is unique, and coincides with the weak limit whenever the latter exists. The sequence Xt is bounded. We associate with it the function (/> defined on Xhy (3) 0(y):=lim sup ||x:i->^p= inf sup s>o r^s The function </) is nonnegative, locally Lipschitz, strictly convex, and satisfies (4) lim (j)(y)=oo WyU^OO DEFINITION 2 The unique point that minimizes ^ on X is called the asymptotic center of the bounded sequence of elements Xf ^ PROPOSITION 3 The asymptotic center of a bounded sequence of elements XteX belongs to the closed convex hull of its weak cluster points. k Proof Let a: 00 be the asymptotic center of the bounded sequence of elements Xt and let ;;oo be the projection of Xoo onto the closed convex hull C of the weak cluster points of the sequence {xt}. It is nonempty, because any bounded sequence is weakly relatively compact. By definition of (¡), there exists a sub¬ sequence of elements Xt> such that 4>{yJ=\m \\x„-y, t'-*O0 Hence, lim sup ||x,.-x«,||^<lim sup \\x,-xj\^ = 4>{xj i'-»00 t~* oo We can also find a subsequence Xr that converges to some z eC. Therefore, lim sup - -^oc. .y00 -) = (I'oo - JCco. Jo, - z> <0 because y^ is the projection of x^ to C.
252 CH. 5, SEC. 2 a general variational principle From the identity lk,-A:„|p=||A:,-3^JP + lb„-JC„|P + 2{A:,-7<„, we deduce that ^{xj>4>{yj+\\x„-yj\'^^<i){yj From the uniqueness of the minimum of (/>, it follows that x^=y^ eC. The second part of the proposition is an immediate consequence of the first. ■ We now detail the preceding result. PROPOSITION 4 Let us consider a bounded sequence of elements Xj. We denote by C the closed convex hull of its weak cluster points and by N the subset defined by (5) N\=<y eX\(l>{y)=l\m ||x, .-.i'lp} which we call the ''attractor'' of the sequence. If NnC^0, this intersection reduces to the asymptotic center x^ (6) NnC = {xJ Proof Let;; e Nr\ C; we shall prove that ;;=Xoo, and, for that purpose, we check that 0(z)^ </>00 for all z e X and;; is the unique minimum of <l>. Let w be any weak cluster point of {xt}; then w is the weak limit of a sub¬ sequence {xf} of {x}. Passing to the limit as t'-^co in the identity \\Xf-z\\^ = \\Xf-y\\^ + \\y-z\\'^+2{Xf-y,y-z) we obtain, since ;; belongs to the attractor N, (p(z)^4>(y)+\\y-z\\^ + 2(w-y, y-z) Since this inequality holds for all weak cluster points, it also holds for all w € C. In particular, we can take w =y for y eC. Hence, <f){z)'^(l){y)+\\y-z\\^><l>{y) ■ As a corollary, we obtain the following sufficient condition for weak con¬ vergence.
CH. 5, SEC. 2 FIXED POINTS OF NONEXPANSIVE MAPS 253 COROLLARY 5 If the weak cluster points of a bounded sequence of elements Xt belong to its attractor N, then this sequence converges weakly to its asymptotic center. A PROPOSITION 6 Let us consider two bounded sequences of elements {xr} and Let N be the attractor and C the closed convex hull of the weak cluster points of {xt}. If the weak cluster points of {yt} belong to Cr\N, then yt converges weakly to the asymptotic center of the sequences {xr}. A Proof Since the weak cluster points of {yt} belong to CnN, this set is nonempty and by proposition 4, reduces to x^o, the asymptotic center of {xt}. Since the sequence is relatively compact and since it has a unique cluster point Xoo, then converges weakly to x^ as i->oo. ■ We are now ready to prove that Cesaro means converges to a fixed point. THEOREM 7 Letf be a nonexpansive map from a closed, convex bounded subset К of a Hilbert space to itself For all initial point Xq e K, the sequence of elements ут'= (V^)Zr=i/Vo)» T e N converges weakly to the asymptotic center of the sequence f\xo\ t e I^, which is a fixed point off A Proof We consider the sequence of Cesaro means i t=i It is bounded, since/ is nonexpansive. For any y&K, write \\гы-т+т-уг =ll/W-/(y)ll"+2</Vo)-/(v), m-y)Hfiy)-yf By adding these inequalities from i = 1 to Г and dividing by T > 0, we obtain ^ E тхо)-уГ=^ E \\nxo)-f{yW + \\f{y)-yr+2(yT-f(y\f{y)-y} ^ t=l Л t=l We now take;^;=;^7-. We deduce that Y E \\Г(хо)-УтГ=^ E \\Пхо)-/(Ут)Г-\\/(ут)-УтГ ^ t=l 1 r=l
254 CH. 5, SEC. 3 A GENERAL VARIATIONAL PRINCIPLE Which we rewrite as Since/ is nonexpansive, mxo)-f{yT)\W-Hxo)-yT\\ and all terms in these sums cancel except the first and last, yielding: <ymax{|W||xeK} So, Iim7'-,„||_p7'—/(yT-)ll =0. On the other hand, part (b) of the proof of theorem 1 implies that any weak cluster point of the sequence of elements yr is a fixed point of f. But any fixed point 3c of/ belongs to the set N associated to the sequence x,:=f‘{xo), because the inequality (7) ||/'(x:o)- x||=||/'(xo)-/(x)|l < II/' H^o)-^11 shows that the sequence of real numbers ||/‘(xo)—3c|| is nondecreasing and bounded below and thus convergent. So the weak cluster points of belong to the attractor N. They also belong to the closed convex hull C of the weak cluster points of {xr}. Hence, proposition 6 implies that yj converges weakly to the asymptotic center of the sequence of elements/'(xo), which is a fixed point of / ■ 3. THE 8-VARIATIONAL principle We now address ourselves to the problem of minimizing a lower semicontinuous function i7 on a complete metric space X. Of course, if the space X is compact, there will always be a minimizer, that is, some point x where (1) U(x)^ U{x), for aWx&X Such will not be the case when no compactness assumption is made. In other words, the infimum of the set of real numbers {i/(x)|x: € X}, denoted by inf U, need not be attained (it may even be — oo). However, by the very definition of an
CH. 5, SEC. 3 THE £-VARIAT10NAL PRINCIPLE 255 infimum, there will be a sequence x t in AT such that (2) i/(x„)-^inf U when «^oo Such sequences are called minimizing. When inf U is finite, this implies that for any 6>0, there is some such that (3) Uix)>U{x,)-e, for all X 6 X We shall show that even in the general, noncompact case, there is more to be said. Provided that inf U is finite, we shall associate with any e > 0 points that satisfy, in addition to (3), other conditions, to be interpreted later. THEOREM 1 Let X be a complete metric space, C/:A'-»Ru{+oo}a proper, nonnegative, and lower semicontinuous function, and Xo in Dom U. Then there exists y in Dom U such that (4) i. i/Cv)+d{xo, j) < U(xo) ii. fx4^y, t/(y) < i/(x)+d(x, y) COROLLARY 2 («-VARIATIONAL PRINCIPLE) We posit the assumptions of theorem 1. Let there be given e > 0 and x^eX such that (5) i/(x,)<£-l-inf [/ Then for any k>0, there is some point y^eX such that (6) i/(Ve)< U(x,) (7) (8) k U{x)> U{yc)-ked{x, yX forallx^y^ Proof of Theorem 1. As in proposition 1.12, we associate with C/thedynamic system defined by (9) C?(x)={y|i/(y) + i/(x,y)^[/(x)} Since i/is lower semicontinuous, G{x) is a closed subset of X, which coincides with the forward cone ^(x) by proposition 1.12. Since i/(xo) is finite, it follows
256 CH. 5, SEC. 3 A GENERAL VARIATIONAL PRINCIPLE from corollary 1.11 that its forward cone contains an invariant pointy ye^(xo) and G{y)={y} Recall that <^{xo)=Gixo)- The first condition then becomes U{y)+^Xo, U{xo) We now write that y is G invariant [/(x)+d(y,A:)<C/(y) -» y=x This is exactly relation (4, ii). The theorem is proved. ■ Proof of Corollary 2. Replace distance </by k&d, the function Uhy U— inf U, and the point xq by some Xe satisfying (5). Denote by y^ a solution provided by theorem 1. Relations (6) and (7) then follow from combining (4, i) and (5). Relation (8) follows from (4, ii). ■ This chapter views condition (8) in several different lights. For the time being, we choose to treat it as a variational principle. In other words, we relate it to the theory of necessary conditions for local minima. From now on, X will be a Banach space, so that </(x,y)=||x—y||. Transcribing theorem 1 in this new notation would be a loss of time and space, but it might be useful to draw a picture (see Fig. 1). The set ‘¡^:= {(x, a)|a+^e||x|| < 0} is a cone of revolution in x /?, with vertical axis and vertex at (0,0), oriented downward. Its angle co is defined by tan (u=A:e. Translating this cone by a vector (3c, a) in X x R will bring its vertex to (3c, a).
CH. 5, SEC.3 THE a-VARIATIONAL PRINCIPLE 257 The translate of ^ can be defined directly as follows: (10) 5) = {(x, a)|(a —¿)-l-A:6||x —x||^0} It is now geometrically clear what condition (8) means: The cone ^+(ye, U(ys)) lies entirely under the graph of the function U \n X xR and does not touch it except at its vertex. The smaller we make ke, the flatter is the cone ^ and the closer ^ + (ye, U(ye)) to the horizontal hyperplane through (y^, Uiye)l The limiting position ke=0 is not permitted by the statement of theorem 1. Writing ke=0 in equation (8), leads to the recognition of ye as a unique global minimizer for U over X. Such a strong minimizer need not exist under the very weak assumptions that X be complete and U lower semicontinuous and bounded from below. If in some particular case, it can be directly shown that U has a unique global minimizer yo over Xy then of course this point yo will satisfy condition (8) with ks=0. The interest of corollary 2 lies in the fact that for any 6 > 0, there will always be some point Xe satisfying condition (5) and, hence, some point ye satisfying conditions (6), (7), and (8). The last two are somehow complementary, and the choice of A:>0 allows us to strike a balance between them according to which application we have in view. If k is small, the cone ^ will be flat, and the point ye will be close to being a global minimizer for U. On the other hand, the right- hand side of inequality (7) will be large, so that we shall have little information on the whereabouts of ye. Conversely, if k is chosen large, ye will be located close to our initial point Xe, but the cone ^ will be acute, and inequality (8) will give us little information. The two most important ways to choose k are k = l, which means we are losing interest in condition (7), and k = e~^^^, which means we are keeping ^ flat and ye close to Xg (having our cake and eating it). In these two statements, we make X a complete metric space again. COROLLARY 3 Let X be a complete metric space and i/:X^Ru{+oo}aproper lower semi¬ continuous function bounded below. Then for any 6>0, there exists some point ye where (11) (12) i/(ye)^a-i-inf U U(x) > Uiy,) - ed{x, ye), for all xfye COROLLARY 4 Let X be a complete metric space and U'. >IR u {+co} a lower semicontinuous function bounded below. Let £>0 and Xe^X be such that (13) f/(Xe)<£ + inf U
258 CH. 5, SEC. 3 A GENERAL VARIATIONAL PRINCIPLE Then there exists some point yc where (14) U(ye)^ U(x^) (15) d(Xe,yeHy/ê (16) U(x)^U{ys)-\fed{x,yc), for all y ^Xt k We now go back to Banach spaces and figure 1. Imagine U to be differentiable at y^. The tangent hyperplane to U at this point will obviously lie above the cone ^+{ye, U{yf). The flatter the cone % the more horizontal To be precise, if X* is the slope of J'H’, we must have ||A:*||#<tan co=ke. Let us state and prove this result. PROPOSITION 5 Assume X is a Banach space and U is finite and Gateaux differentiable at y^. Then condition (8) of theorem 1 implies that (17) Proof The left-hand side denotes the norm of U\y^ as a continuous linear functional on X. Taking any unit vector y&X, ||>'|| = 1, and setting x=yc + ty, with t>0, in formula (8), we have (18) 1 \_U{yt + ty)- U{yf\> -ke, for all i >0 Letting i^O, we obtain by the definition of l/iy^) as a Gateaux derivative (19) (U'(ye),y)^-ke, forall^eJf with ||j Since — y is also a unit vector, we have (20) —(G'(ye),y)^—ke, for all 7 € with | From both inequalities, it follows that (21) Kl7'(>'£),>')|<^e, forall>'6 2i with ||> This is precisely formula (17). 1 = 1 1 = 1 1 = 1 We now have the interpretation we were looking for; condition (8) as a vari¬ ational principle. If A:e=0, we have V'{y^=0, the familiar first-order necessary condition for a local minimum. But it has already been pointed out, and indeed will presently be illustrated by an example, that the assumptions of corollary 2
CH. 5, SEC. 3 THE 6-VARIATIONAL PRINCIPLE 259 by themselves do not warrant that such a point exist. However, for any ke > 0, we are assured that there is some point ye where condition (8), and hence (17), is satisfied. In other words, if C/'(ye) can not actually be made zero, it can at least be made arbitrarily small. This point of view is illustrated by the following results. COROLLARY 6 Assume X is a Banach space, U\ X^U is lower semicontinuous, Gateaux differ¬ entiable, and bounded below. Let a > 0 and Xe^X be given, with (22) i/(;cj^8 + inf U Then there exists some point ye where (23) UiyeH UiXe) (24) (25) l|t/'(:F.)IU<Vi A COROLLARY? Assume X is a Banach space, U'. X-*U is lower semicontinuous. Gateaux differ¬ entiable, and boundedfrom below. Then there is a sequencey„ such that when n-*oo (26) (27) t/(7„)->inf U U'iy„)-*0 in X* The latter result is particularly interesting. It asserts the existence of minimizing sequences of a particular kind: Not only do they minimize the function V, but they also satisfy the first-order necessary conditions, up to any desired approx¬ imation. To relation (26), which defines all minimizing sequences, we have added condition (27), and our result is that both can be satisfied at once. We illustrate these results with a few examples. Example 8 Let X be a Banach space, and U: X-^R a lower semicontinuous, Gateaux- differentiable function. Assume that for some constant a>0 and c, we have (28) U(x)'^a\\x\\ + c, for all x Then U'{X) is a dense subset of the ball aB* in X*. If condition (28) is strengthened to (29) U{x)><t>{\\x\\), for all X
260 CH. 5, SEC. 3 A GENERAL VARIATIONAL PRINCIPLE where [0, +co)-*R is continuous and i” +<* when t-^co, then U'{X) is dense in A"* for its norm topology. ^ To see this, note first that condition (29) implies that for all a>0, some c will be found such that inequality (28) holds. So U'{X) will be dense in all the balls around the origin in A* and, hence, in A* itself. Now, if condition (28) holds, take any p eX* such that ||p||*<a. In other words, the point p belongs to aB*. We are going to show that there is a sequence p„ in (/'(A) that converges to p, which is easy. Simply define a new function F: A->/îby (30) V{x)=U{x)—{p, x), for all X. Clearly, V is lower semicontinuous and Gateaux differentiable. Since (31) V(x)> i/(x)-||p|Ullx||>(a-||plU)l|x|| + c the function V is bounded from below by c. Applying corollary 7, we obtain a sequence y„ in A such that (32) V'{y„)=U'{y„)-p^0inX* Setting p„ = U'{y„), we obtain the desired result. ■ Example 9 We now wish to present a practical situation where there is no exact minimizes Let A be ITo’‘^(0, 1), the Sobolev space of all continuous functions x on the interval (0,1) such that its distributional derivative x=î/x/î/( belongs to L‘^(0,1) and x(0)=x(l)=0. Let U: X^R be defined by (33) l/(x) = I {(x(i)^ -1 )^ + x(t)^} i/i This is the situation of example 3.4 in chapter 1. We have seen that there is no minimizer. Corollary 7 will provide us with a substitute. Set A(f)=| x(i)ifc and write the function U as follows: (34) i/(x)= ix{t)^-lfdt- f* X(t)x{t)dt Jo Jo
CH. 5, SEC. 4 APPLICATIONS TO CONVEX OPTIMIZATION 261 It is now apparent that U is Frechet differentiable and its derivative is the linear map on Wo'^iO, 1) defined by (35) y} = £ (4Mx^-l)-2X)ydt Nowy can be any function in ii(0,1) such that [ y(t)dt=0 Jo It follows that (36) \\U 'WIU=sup|| i4x{x^-l)-2X)ydt\yet,^ ydt=( =inf (||4x(x^-l)-2Z-a|U4/31 a e R} Corollary 7 now yields the following; There is a sequence y„ 6 IFo’‘*(0,1) and a sequence € R such that (37) (38) {(yl-'^f+yl]dt^(J l|4y„(j„"-l)-2y„-a„|L4/3^0 4. APPLICATIONS TO CONVEX OPTIMIZATION Let us consider a lower semicontinuous, convex proper function U: X-»Ru {+00}, where is a Banach space. We have proved in theorem 4.3.11 that in the particular case when X is a Hilbert space, U is subdifferentiable on a dense subset of its domain. We now extend this result to Banach spaces. Let V be bounded below and xo any point in X. The s-variational principle states that there exists a point x^&X such that (1) \\xo-xJi\^U{xo)-U{x,) (2) Wx=f=Xc, i/(x:e)< C/(x)+£||x:-X£|| Since Xc minimizes the function C/(x) + £||jic-x:e||
262 CH. 5, SEC. 4 A GENERAL VARIATIONAL PRINCIPLE the subdifferential calculus (corollary 4.3.6, in particular) implies that (3) 0€dU(x,)^-eB Theorem 1 follows immediately. THEOREM 1 Let U: X-^Uu{+oo} be proper, lower semicontinuous, convex, and bounded below. Choose Xq in Dom U and e>0. Then there exists x^ in Dom U and /?£ € 5 U{Xe) such that (4) llXo-Xell^ C/(Xo)- t/(Xc) (5) IIaIU<« ^ Changing the norm of X from ||x|| to /:e||x||, we obtain a more precise result, COROLLARY 2 Let U: A'->Ru{+oo} be proper, lower semicontinuous, convex, and bounded below. Let e>0 and ;co € Dom U be given such that (6) i/(xo)<e+inf i/ Then, for any k>0, there is some point Xc e Dom U and some pc edU{xc) such that (7) Uixo) (8) 11^0 (9) ||pe||*<ke ▲ We deduce theorem 3, which is due to Br^nsted and Rockafellar. THEOREM 3 Let X be a Banach space and C/:Ji-+IRu{+oo}a convex, lower semicontinuous function. Then the set of points where U is subdifferentiable is dense in Dom U. More precisely, for any xeX where U{x)< +oo, there is a sequence Xk,keN such that (10) . Xk-*X . l/(Xk)-> U{x) . dU(Xk)f0, for all k
CH. 5, SEC. 4 APPLICATIONS TO CONVEX OPTIMIZATION 263 Proof. If U=-\-<x>, Dom U is empty and there is nothing left to prove. If +00, there is somepeX* and a 6 R such that (11) V(x) = U{x) — (p, x) — a>0, for all X Let X be any point where U (and hence V) is finite. Apply corollary 2 to K with e= L(x)- inf V (not necessarily small) and k any integer. We obtain some point Xk^X and pk 6 dV(xk). (12) (13) (14) L(x,)<F(x) This obviously implies that U is also subdifferentiable at xu, with dU(xk)= p+dV(Xk). Conditions (10, i and il) have thus been proved, and we are left with condition (10, ii). Expressing V in terms of U in equality (11), we obtain U(xk) < U{x) —{p,x—Xk), for all k Letting k-^co, we have lim sup i/(xii)< U{x). On the other hand, since U is lower semicontinuous, we have lim inf i/(x*)> U(x) So U(xk)~* U{x), and the theorem is proved. ■ We now apply Corollary 2 to convex optimization problems of the form treated in Section 4.8 in situations where there exists a Lagrange multiplier, (given by theorem 4.8.1) but no optimal solution. PROPOSITION 4 Let X and Y be Banach spaces, L: X x y^i?u{+co} a lower semicontinuous, convexfunction, and A: X-^Y a continuous linear operator. Consider the optimiza¬ tion problem (15) \n[L(x,Ax) X Assume that 0 6 Int((/1 ® - /)Dom L), so that there exists a Lagrange multiplier (16) — L^{—A*g,q)= ln( L{x, Ax) xeX
264 CH. 5, SEC. 4 A GENERAL VARIATIONAL PRINCIPLE Let x„, ne N, be a minimizing sequence in problem (15). Then there is a sequence (x„, y„, P„, q„) such that (17) (18) (19) (20) (21) Ikn-^IU-^o ll/>«+^*9lU-*o ip„, q„) 6 dL{x„, y„) Proof. Consider the function IJ on X xY defined by (22) U{x, y)=L{x, y) + (x, A*q) - (y, q) +1*{-A*q, q) This function is always positive. By assumption (16), its infimum is zero, and the sequence (x„, Ax„) is minimizing. Setting U(x„, Ax„)=s„->'0 and applying corollary 2 to the point (x„, Ax„), with >Je„ = l/k, we obtain points {x„, y„)eXxY and (p'n, q'„)eX* X y*, where (23) (24) (25) ||xn and iPtti qtd ^ ^ LJ{Xny ^n) llPnlU^Ve» and Setting /?„ = - y4 -f pn and ^ + q'n, these relations yield the desired result. If there is an optimal solution 3c to problem (15), proposition 4 will hold with (Xm YmPm Ax, —A*q, q) for all n. More generally, in the absence of any optimal solution, it may well happen that relations (17)-(21) imply some kind of convergence of the sequence (x„, yn) to some limit (3c, Ax), which will be regarded as a “weak,” or “generalized,” solution to problem (15). This kind of result was first observed by Temam in the particular case of Plateau’s problem. We now describe this example. Example 5 Let Q be a bounded, open subset of and let xq be some function in the Sobolev space py^’^(Q). We want to minimize (26)
CH. 5, SEC. 4 APPLICATIONS TO CONVEX OPTIMIZATION 265 subject to (27) This is a simple version of Plateau’s problem: Integral (26) is the area of the hypersurface S in described in parametric form by (28) 5={(cu, x(cu))| cu € Q} and condition (27) means that x and xo coincide on the boundary dil of Q. In other words, we are seeking a hypersurface S (if any), admitting a representation of type (28), with prescribed boundary 55={(co, xo(cu))|a> e 5i2} and having minimum area. Soap bubbles apparently solve this problem, but mathematicians do not. The set xo + lTo’^(n) is a closed affine subspace of fT*’‘(Q), Integral (26) defines a function U(x), which is positive and continuous by Krasnoselskii’s theorem. It is also Frechet differentiable, and its derivative at x is the continuous linear map on ITo’*(i2) defined by (29) fFj’* ay- -1/2 y dx dy i = i da>i dcOi dœ By definition this functional is an element of W which we can write as where the derivatives are to be understood in the sense of distributions. Finding a minimizer for integral (26) under contraint (27) therefore means finding a weak solution for the boundary-value problem ,?i ôa>i [aco, ( do)J J (32) x-xo 6 IFè'Uiî) Unfortunately, if no further assumptions are made on the shape of Q, this problem need not have a solution. A sufficient condition for a solution x to exist for all xq is that the domain Q be convex (more generally, that the mean curvature of the boundary 5Q be positive). This is very surprising at first glance, since we could always build a wire to the prescription {(co, xo(co))lco e 5Q}, dip it in soapy water, and note the shape S of the bubble. The point is that even though such a “physical” solution will exist in general, it will not admit a parametric representation of the form (28)—except, of course, when Q is convex.
266 CH. 5, SEC. 4 A GENERAL VARIATIONAL PRINCIPLE In the general case, when equations (31) and (32) have no solution, we apply the e-variational principle to get a substitute. Set z=x—Xo, so that problems (26) and (27) becomes (33) (34) inf U{z+Xo) Z € It is well known (Poincar6’s inequality) that ||grad x||ti and are equivalent norms on The following inequality then holds for some A:>0 and c; (35) U(z+Xo)>k\\z\\-c We have the situation in Example 3.8: There exists in the haWkB* of a dense subset such that for all T ^ if, the equation t/'(z+Xo)=T has a solution. Taking into account the fact that U is strictly convex, this means that the problem (36) X-Xo € has a unique solution x for all T in if. We now apply proposition 4, taking Ax= Y=n(ilf L(x,y)= dx dx\ d(Oi ’■■-’dcoj 1 + Y + 00 otherwise HQ) Starting from a minimizing sequence x„ €Xo + we find a sequence (x„, y„, p„, q„) such that (37) (38) (39) x„-x„^0 in IT‘’Hii) y„-Ax„-^0 inL‘(il)'' q„-q-*0 inL”(Q)'''
CH. 5, SEC. 4 APPLICATIONS TO CONVEX OPTIMIZATION 267 (40) p„ + A*q->’0 in [W'i ‘{Q)]* r 1~ 1/2 (41) 9n=l';.|^l+ E {y'„f^ , l<j<N, (42) (Pn, x) =0, for all X 6 W'J *(fi) A simple computation shows that (43) Ii{-A*q,q) = {-A*q,Xo)- S if ||^(co)||^l almost everywhere and L!^{—A*q, q)=-\-co otherwise. Since q is a Lagrange multiplier, it minimizes (43), and it follows that (44) ll^(co)|| [N “11/2 < 1 a.e. Using very refined a priori estimates, due to Ladyjenskaia and Uraltseva, Temam has associated with any open set whose closure is contained in fl, a constant c(‘^)>0 such that (45) Vcoe'^, ||^(a))||<l- Condition (39) them implies that ||9„(<o)|| is also bounded by l—c{^) on Relation (41) can be reversed (46) 1^1 - Z ^U")^ j 1/2 to show that the sequence yn converges uniformly on ^ toward the function 1-1/2 (47) y{o})=q{(o) 1^1 - X It is known that the operator A has closed range, so that y=Ax, where X 6 iy*’‘(i2) is defined up to an additive constant (48) (49) -1/2
268 CH. 5, SEC. 4 A GENERAL VARIATIONAL PRINCIPLE Conditions (40) and (42) imply that {A*q, a:) =0 for all x € which means that in the sense of distributions, (50) div^=0 in0'(Q) Substituting (49) into (50), we have This is the so-called equation of minimal hypersurfaces, which any optimal solution of Plateau’s problem has to satisfy, together with the boundary condi¬ tion (52) x-XoeWh'\Q) As noted before, there may well be no solution to boundary-value problem (51) and (52) and, hence, no minimizer. What we do show is that there always exist minimizing sequences x„ and a solution x to equation (51) such that x„ converges uniformly to x on all compact subsets of Q. The function x, however, does not necessarily satisfy boundary condition (52) (if it does, it is a minimizer). It is a “weak” or “generalized” solution of Plateau’s problem; a more detailed analysis, carried out by Lichnewski, shows that x{a>)=xo(o)) at all points co of the boundary dCl where the mean curvature is non-negative. A typical situa¬ tion is illustrated in Figure 1, with TV = 1, Q=(— 1,0) u (0,1) and xo( — 1)=Xo(l)= 1, x:o(0)=0. The generalized solution is 3c(co) = l; it is shown in the following figure as well as a minimizing sequence x„ converging to it.
CH. 5, SEC. 5 CONDITION (c) OF PALAIS AND SMALE 269 5. CONDITION (C) OF PALAIS AND SMALE In the preceding section, we have shown the existence of minimizing sequences of a particular type. The question naturally arises whether these sequences converge, thereby proving the existence of an actual minimizer. Some kind of compactness assumption is needed, the weakest being condition (C) of Palais and Smale. Historically, condition (C) was first stated for functions on Banach mani¬ folds, where it serves as a substitute for convexity. The differentiability require¬ ment has since been lowered to ^ (locally Lipschitz derivative) and even for certain questions. It is within this nonlinear framework that condition (C) has proved the most useful. We shall, however, deal with linear spaces only, because Banach manifolds are outside the scope of this book. The proofs we give will, nevertheless, carry over to the general nonlinear case, since they rely on the 6-variational principle and use only the differentiability structure of the underlying space. DEFINITION 1 Let X be a Banach space and U: a Gdteaux-differentiable function. We say that U satisfies condition (C) on a subspace Q<^X if whenever there is a sequence x„, n e N, in ii, with (1) I i/(.x:„)| ^ constant (2) U'{x„)^0 inX* then in the closure of the set [xn\n eiV}, there is some point x where U'{x)=0. Note how carefully this condition is worded: It is not stated that there is a convergent subsequence in the sequence x„. Indeed, \{ X = R and t/=0, then condition (C) is satisfied. For instance, taking the sequence x„=n, we have f/(x„)=0= U'{xn) for all n, There will be no convergent subsequence, but we can take any of the for 3c. Condition (C) enables us to use Morse theory or Liusternik-Schnirelman theory, that is, to find lower bounds for the number of critical points of U by topological methods (if, for instance, some group acts on X and leaves U invariant). Since our aim is much more modest, we are content with a weakened form of condition (C), which we call (weak C). DEFINITION 2 Let X be a Banach space and U\ X-^R a Gdteaux-differentiable function. We say that U satisfies condition {weak C) on a subspace O.C.X if whenever there is a
270 CH. 5, SEC. 5 A GENERAL VARIATIONAL PRINCIPLE sequence x„, n € N, in Q with (3) I U{x„)\ ^ constant (4) U(x„) ^ 0, for all n (5) t/'W-O inX^ then there is some point xeX such that (6) lim inf U{Xt)^ i7(x)<lim sup t/(x„) (7) C/'(3c)=0 A PROPOSITION 3 If U is continuous and satisfies condition (C) on Q, it satisfies condition {weak C) on O.. If X is reflexive and U is convex, lower semicontinuous and U{x)^ -foo when ||a:||->oo, then U satisfies condition {weak C) on X. A Proof The first statement is trivial. Its converse, that (weak C) implies (C), is false in general. We prove the second statement. Take a sequence satisfying (3), (4), and (5). Then take a subsequence such that lim {/(>;„)=lim sup i/(x„). It follows from the assumptions on X and U that contains a weakly convergent subsequence its limit being called x. Since U is weakly lower semicontinuous, we have i7(x)<lim C/(zJ=lim sup Fix X in X. Since U is convex, it satisfies inequalities {U\zn\ x-Zn)^-U{zn)^ U{x) Taking the limit, we obtain, because of (5), lim sup {7(xJ = lim i/(z„)^ U{x) Since X is any point in X, everything follows at once. First C/(x)^ i/(x), so that X is a minimizer and i/(x)=0. Then lim sup t/(xn) = inf U, so that t/(xj actually converges to U{x). ■ Our first existence result concerns minima. PROPOSITION 4 Let U: X-^U be Gateaux-differentiable, lower semicontinuous, and bounded from below. Assume the restriction of U' to straight lines is continuous and U
CH. 5, SEC. 5 CONDITION (c) OF PALAIS AND SMALE 271 satisfies condition (weak C) on X. Then U attains its minimum on X (8) 3xeX: U(x) = \nf U A Proof. By corollary 3.7, there is a sequence neN, in X such that U and U\yt)-^Q. There are now two cases to consider: Either we can extract a subsequence (denoted by x„) such that U\Xf)4^0 for all or i/'(j;„)=0 for all but a finite number of n. In the former case, by condition (weak C), there is some point xeX such that t/'(x)=0 and f/(i)^lim i/(x„) = inf U So 3c is a minimizer, and the result is proved. In the latter case, set S={x\U'(x)=0}. liS=X, then U is constant, and any point is a minimizer. USfX, there is some point z where U'(z)^0. For fixed n, consider the line segment t-^tz + (l — t)x„, 0 ^ i ^ 1, from to z, and set i = inf {i|iz + (l — By assumption, the restriction to this line segment of U is continuous. It follows that we can find ti^t^t2 with xi=iiz + (l-ii)x„ eS xi=t2Z + (i-t2)x„^S \U(xj)-U{x},)\^n-^ wvixm^n-^ Since iiwe have fz+(l — for all f<fi. It follows from the mean value theorem that U{x!,)= U{x„) We have found, for each ne N, a. point xl such that U'{x^)^0 and Letting «->•00, we have [/(x^)->inf U and U'ix^)-*0. We are back in the preceding case. ■ If we drop the continuity assumption on U', our argument will give us a critical point i/'(x)=0, without telling us whether it minimizes U. Note, however, that this will follow automatically when U is convex; as is well known, the result will hold in this case without any differentiability assumption at all on U.
272 CH. 5, SEC. 5 a general variational principle A judicious use of condition (weak C) will enable us to find critical points of various types, not only local minima or maxima, but saddle points as well. There are now many theorems of this kind, the simplest and perhaps the most beautiful being the following, a strengthened version of a result originally due to Ambrosetti and Rabinowitz. THEOREM 5 Let X be a Banach space and U\ X^U a continuous and Gdteaux-differentiable function. Assume that If: X^X’*' is strong-to-weak"^ continuous and (9) (10) (11) 3a>0: m(a):=inf {U{x)\ ||x:|| =a}> i/(0) 3z € A': ||z|| > a and U(z)< m(a) U satisfies condition (weak C) on {x\ U{x)'^m((x)} Then there is a point xsX where (12) U(x)>m(ot) and U'(x)=0 A Proof A path from zero to z is a continuous map c: [0,1]-^X with c(0)=0 and c(l)=z. Denote by ^ the set of all paths from zero to z, endowed with the distance C2)=max {||ci(i)-C2(t)ll It is a complete metric space. We define a function /: as follows: /(c)=max {C/(c(t)) 10 < t < 1} The function I is lower semicontinuous, since it can be written /(c)=sup, /,(c), each function /,(c)= C/(c(i)) continuous. Note also that it is bounded below by m(a). Indeed, since c continuously joins zero to z, it must cross the ball of radius a somewhere: There is some f« e [0,1] where ||c(ta)|| =a and, hence, I{c)> t/(c(f„))>m(a) It follows that we can apply corollary 3.2: For any e>0, there is some path Ce where /(CcXinf {/(c)|c e <^} + e l(c)^I(Cc)-ed{c, c^) Now let y be any continuous map from [0, 1] to X such that y(0)=0 and
CH. 5, SEC. 5 CONDITION (C) OF PALAIS AND SMALE 273 y(l)=0. For any /» e R, we have the inequality /(ce+hy)~ I(c,) ^ - ed{c,+hy, c,) which becomes h-^\_I(c,+hy)- /(ce)]^ -fi max ||y(i)|| t On the other hand, we can write I(ce + hy)- I(ce)=m?i\ U{ceit)-^hy{t))-ma\ U{Cs{t)) t t =max {t/(Ce)+/i<i/'fe), r>}+o(/»)-max C7(ce) r ^ with h~^o{h)-^0 when fi-^0 (since [0, 1] is compact). Set U(ce)=f and (U'{ce\ y}=9- Our assumptions will imply that/and g are continuous maps from [0, 1] to X. Define a function : C®([0,1])--^IR O)(0)=max |<^(i)| t This function is convex and continuous, hence, subdifferentiable everywhere. The dual of C°([0,1]) is the space of Radon measures n on [0,1], and the sub¬ differential of d) is given by =|/^>ol j dn = 1 and supp jUczM(0) where M(0) = {i|0(i)=<!)((/))} We now see that -6 max l|y(i)||^lim [/(c^ + Ay)-/(cJ]ä ^ t h-*0 = lim W^hg)-<t>{f)-]h-^ h-^O =max{(g,fi)\ned^{f)} =max I j < t/(ce), y)dn \ n e dd>(/)} On both sides take the infimum over the set of y e C°([0, 1]; X) such that ||y||<l, r(0)=0=y(l)
—E<max 274 сн. 5, SEC. 5 a general variational principle We obtain -e< inf max < (U'(cc), y)dn ‘ Y M iJ ■ The set дФ(/) is weak* compact, so that we can use the inf sup theorem 5.2.7 nax inf I [ (U'(Ce), y)dn‘ Ц 7 U : ^í6 5Ф(/),IMI<l y(0)=0=)-(l) цедФ{/), r(0)=0=r(l) =max = -min {II í/'(Ce(í))|U|í e M{U0Ce)} So there is some tj such that i/(Ce(te)) = maX U(Ceit)) t l|i/Ve(te))ll*<e Now set 6=«“* and Ce{te)=x„. We have proved that m(a)< [/(xj^inf 7+e and U'{x„)^0. If U'(x„)=0 for some n, we are done. If not, we apply condition (weak C) to find some point x where i/'(^)=0 and i/(x)>lim inf i/(jc„)>m(a) ■ It was quite easy to picture the underlying geometric situation. Imagine the graph of U as the shape of a mountain range lying over X. Then the origin lies in a closed valley, and it is known that there is some point outside this valley with lower altitude than the surrounding mountains. Clearly then, there must be a mountain pass out of the valley, and this is exactly what our proof is looking for. The nontrivial critical value in theorem 5 is given by V = inf max i/(c(i)) ce^ 0<i<1 The idea of finding a critical value by an inf max formula of this kind can be extended to other situations, yielding different kinds of critical points. THEOREM 6 Let X be a Banach space and U\ X-^U a continuous and Gdteaux-differentiable function. Assume that U'\ AT-->Ar* is strong-to-weak* continuous and X splits into a direct sum X = Xq®X^, with (13) Xq is finite dimensional
(14) (15) (16) CH. 5, SEC. 5 CONDITION (c) OF PALAIS AND SMALE 275 3R>0: [xo eXo and ||xo|| = 7?]=> {/(xo, 0)<0 x„ 6Z„=>C/(0, xJ^O Usatisfies condition {weak C) on {x| C/(x)^0} Then U has a critical point: (17) Proof. Set 3x € 2Í: i/'(x)=0 and t/(x)>0 = S Xq I ||Xo||</?} Sq = {xo 6 Xo I ||xo|| =7?} <^={</>6C°(Bo, XJ|(/)(5o)=0} We endow ^ with the uniform topology, which turns it into a complete metric space, and we define a function I: as follows: 7(0)=max { U{xo, </>(xo))|xo 6 5q} The supremum on the right-hand side is attained, since Bq is finite dimen¬ sional and hence compact. It cannot be attained on the boundary Sq because of conditions (14) and (15) Xo €So=> f/(xo, <^(xo))= f/(xo, 0)< i/(0, m) The function I is lower semicontinuous and bounded below by zero. It follows that we can apply corollary 2.2. For any e>0, there is some such that I{<l>)>I{(l>t)-e\\(l>-(l>c\\ Now pick any y in For any /i 6IR, we have the inequality I{<t>.+hy)-I{<t>.)^-eh\\y\\ Arguing as in the proof of theorem 5, with C”(5o) replacing C°([0, 1]). we find some xo € Bo such that U{xo, </>c(xo))=max {t/(xo, 0c(xo))lxo 6 Bq} I|í7'(a:o, </>e(A:o))IU<« The result follows from condition (weak C). ■
276 CH. 5, SEC. 5 A GENERAL VARIATIONAL PRINCIPLE Example 7 Let X be a reflexive Banach space and A : X-^X* a compact linear operator, with A*=A. Let L: be convex and C‘. Assume that (18) 3A:o>0, 3co 6lR: L(x:)^A:olkl|-Co VxreX (19) ^ki<2, 3ci eR: <F'(x:),x)<^iK(jc)+ci VxeX The latter condition limits the growth of V at infinity; if ci =0, it can be shown to be equivalent to the condition that V{Xx)<X'‘^V{x) for all a: e F and 2>1. Set U(x)=^(Ax,x) + V{x) We claim that the function U satisfies condition (weak C) on X. A To prove this, we take two constants a and b and a sequence a:„ e X such that for all n 6 W V{x„)=]^ (Ax„, x„)+V(x„)^b U'(x„)=Ax„+V'(x„)=p„-yO in X* Substituting the latter relation into the first, we obtain (Pn-y'M, X„) + V(x„)^b LFsing the assumptions on V, this yields ^ (Pm a:„) + ^1- F(a„)-Ci <p„, x„) + i 1- j(/:olW|-Co)-Ci Since p„-»0, it follows that the sequence x„ is bounded: ||a:„||<constant. Since X is reflexive, there is a subsequence shortened to Xk, such that (20) x/,->-x weakly in X
CH. 5, SEC. 5 CONDITION (c) OF PALAIS AND SMALE 277 Since is a compact operator, Axk converges to Ax strongly in X. It follows that: (21) V'(Xk)=Pk-Axk->Ax Since V is convex and continuous, it is weakly lower semicontinuous, so that F(3c)^lim inf K(xfc). Taking limits in the inequalities V{x)>V{xk)+ (y'{Xk\ x—Xk)y we obtain V{x)^V(x)-\-(Ax, x—x), so that — Ax=V'(x). This means that t/'(jc)=0. Starting from V(x)^V{xk)+(V'(xk), x — Xk), and letting k-^co, we have V{x)>lim sup Vixk). We have just proved the converse inequality, so that K(3c)=lim V{xk). Since A is compact, the quadratic term (Ax, x) is weakly continuous on bounded subsets. Finally i/(3c) = lim U{Xk\ and condition (weak C) is satisfied. ■ Example 8 We retain the assumptions and notations of the preceding example. We add the following: (22) 3a>0: inf {U{x) \ ||x|| =a} > F(0) (23) 3z6X:||z||>a and i/(z)^F(0) Then there is some point x eX such that (24) x^O and Ax+V'{x)=0 A This is the situation in theorem 5. Conditions (9) and (10) are assumed and condition (11) has been proved. We shall use this result in Chapter 8. Example 9 Let X be a Hilbert space and F : X^U SiC^ functional. Assume F is twice weakly differentiable Gateaux and there is a constant k>0 such that (25) VxeAT ||F"(x)*i''(^)||>A:||f'(x)|l Then F has a critical point on X (26) 3xeX: F'{x)=0 A Note that the assumption is certainly satisfied if F has the property and F"(x)e SF{X, X) is invertible, with F"(x)”‘ bounded uniformly (27) ||F"(x)-‘||<A:-‘ VxeX
278 CH. 5, SEC. 5 a general variational principle To prove our result, assume that there is no critical point, so F'{x)=^0 for all x. We shall derive a contradiction. LEMMA 10 The real function x-^\\F'{x)\\ is continuous and Gateaux differentiable everywhere, k Proof Consider the map <P{x)=\\F'{x)\\^={F'{x), F'ix)) It is obviously continuous. Check that it is Gateaux differentiable. y [<p(x+iy)-<i()(x)]=i liF'ix+ty), F'ix+ty))-{F'(x), F'W)] =-(F'(x + ty)-F'{x),F'(x)) Now let i^O. By assumption, F' is weakly differentiable Gateaux, so 7 (F'(x+ij^)-f'(x), F'{x))-*{F"{x)F'ix), F'(x)) Restrict t to a sequence Since the sequence ti^^\_F'ix-\-tny)—F'{x)'] converges weakly in X, it is bounded (Banach-Steinhaus theorem). Since F' is continuous, F'{x-\-tny)—F'{x) converges strongly to zero in X. So the second term on the right converges to zero, and \im^ [<p{x+ty)-(p{x)']=2(F"{x)y, F'{x)) i->0 t So q> is Gateaux differentiable. Since l|F'WII=>/^W and <p never vanishes, the function x-* with derivative is also Gateaux differentiable. ^ V ' wnxW V’ ^ ’ \\FM\)
CH. 5, SEC.6 GENERIC DIFFERENTIABILITY 279 We now apply corollary 3.6 to the function |1F'||. We obtain a sequence x„ along which the derivatives goes to zero F"(x ^ -»0 in F* contradicting assumption (25). Hence, the result. ■ 6. GENERIC DIFFERENTIABILITY From now on, U will be any lower semicontinuous function, but X will be a special kind of Banach space. DEFINITION 1 A Banach space X is smooth if there exists a continuous function such that (1) (2) (3) <I>(x) ^ 0 for all X eX D = {.x|<I)(a:) > 0} is bounded and nonempty. <I> is Frechet differentiable on D. Combining a translation with a homothety, we can assume that <D(0)>0 and D is as small as need be. In the sequel, we shall use the function 'P = I/O. It is well defined and lower semicontinuous as a mapping from X to IR u {+oo} and Frechet differentiable on D. PROPOSITION 2 Any Banach space X on which there is an equivalent norm, Frechet differentiable on X \{0} is smooth. If X is reflexive, or if X"^ is separable, then X is smooth. The spaces and are not nor is any space that contains one of them. A Proof. Denote by ||x|| this norm on X. Pick any C® function a: R-^[0, oo) with a(l)>0 and a(i)=0 for i^l/2 and t'^2. Set <D(x):=a(||x||). The second part of the proposition follows from standard renorming theorems due to Kadec, Klee, and Asplund and John and Zizler, to be found in Diestel [1975] Chapter 4, Sections 4 and 9, respectively. ■ On smooth Banach spaces, all lower semicontinuous functions enjoy proper¬ ties that are related to differentiability; namely, those that follow.
280 CH. 5, SEC. 6 A GENERAL VARIATIONAL PRINCIPLE DEFINITION 3 Let e^O be given. We say that a continuous linear functional p eX* is e supporting to U at the point xeX if U(x) < +co and there is some g>0 such that (4) \\x-y\\^ri=>U(y)> f/(x)+<p, j-x)-£||a:-;^|1 The set of all such p is called the s support of U at x and denoted by ScU(x). If it is nonempty, we say that U is e supported at x. A The following properties are easy consequences: (5) i. Se U(x) 4= 0=^ U is lower semicontinuous at x. ii. SeU(x) is a convex subset of X*. I Hi. 5,i/(x)+5«KWc5e+.(i/+F)(x) iv. a^e=>5.C/(x)=>S'ei/(x) The relationship with Frechet differentiability is given in the following result. PROPOSITION 4 The function U is Frechet differentiable at x if and only iffor every e>0, both U and — U are e supported at x. We then have (6) fl SMX) n -5e(-t/)(x)={C/'(x)} £ > 0 £>0 Proof The only if part follows immediately from the definitions. The con¬ verse is less obvious. Set £=«" ‘ and assume U and — i7 are £ supported at x. We find for every n some >/„ > 0 and two continuous linear functionals p„ and q„ such that for l|x-j||<f?„ we have: fi. U{y)>U{x)+{p„,y-x)-n-^\\x-y\\ |ii. -i/(>^)^-t/(x)-|-<^m;^-X>-«“‘||x-j|| Write the first inequality for some m>n, and add it to the second. We obtain ||x-;;|| < </»m+in. - x) 2n ~ 1 l|x-;^|| and hence, (8) Vk e fm^n, ||/>m + ^nlU<2« '
CH. 5, SEC.6 GENERIC DIFFERENTIABILITY 281 Similarly, (9) Vh e - 1 This proves that both p„ and q„ are Cauchy sequences. Since X* is complete, they converge top and q, respectively. Taking limits in the preceding inequalities, we see thatp + ^=0. We claim that p is the Fr^chet derivative of U at x To see this, take any 6>0. Choose and take >;= min (>/„, e„). By the preceding inequalities, we have for ||j — <p„, >»-x)-«■ ^||x-;^||< U(y)- U{x)< {q,„ Letting m-*co, and taking limits in Wpm + qnW* and ||/>„+^mlU. we obtain (10) lip+^„11=^2«“* and ||i+p„||<2«"‘ Substituting this into the preceding inequality, with 65= 3«” ‘, we have for all y such that ||x—p|| <>/, (p,p-x)-s||x->'||< t/(y)- t/(x)<<^,>'-x> + e||x-7l| which is precisely the definition of Frechet differentiability. ■ We now state the main result in theorem 5. THEOREM 5 Let X be a smooth Banach space and (7:Ar^lRu{+oo}a lower semicontinuous function on X. For any s>0, the set of points x where U is e supported is dense in Dom U. A Proof We are given a point Xo in Dom U (so that i/(x)< +oo), a neighbor¬ hood iP" of the origin in X, and we seek a point x e iii^-i-xo where U is locally 6 supported. Since U is lower semicontinuous, we can find a smaller neighborhood of the origin such that U is bounded from below on "iC-i-xo 3m: 'fxeU^+xo, U{x)>m Take a function as in definition 1, assuming <l)(0)=0 and and define a lower semicontinuous function 'F:2f^Ru{ + oo}by 'P(^)= 1 O(x-xo)
282 CH. 5, SEC. 6 A GENERAL VARIATIONAL PRINCIPLE Set F=U + '¥. The function F is lower semicontinuous and bounded from below on the Banach space X. Applying the 6-variational principle with k we obtain some point x where '^xeX, F(x)>F{x)- £||x-x|| This implies that zero is locally e/2 supporting to F at x. Letp=<I)'(jc). so that —p is locally e/2 supporting to — <I) at 3c. It follows that 0—/> is locally e support¬ ing to F—<I>= [/at 3c. Moreover, SO that jc e Dorn F = Dorn i/nDorn'PcDorn Ur\('f'-hxo) ■ In the sequel, we shall prove a similar result for a=0, by slightly strengthening the assumption on X (see proposition 7.7). We apply the preceding result to the special case of convex functions. We wish to prove the following result. THEOREM 6 Let X be a smooth Banach space and U: 2i^lRu{4-oo} a convex, lower semi¬ continuous function. Assume U is finite {and hence continuous) on some open convex subset il<=X. Then U is Frechet differentiable at all points x of some resi¬ dual subset of Si. ^ Let us begin by noting that since U is continuous at every point x 6 fi, it is also subdifferentiable: dU{x)fi0- This gives us indications on S,U(x) and Sei- t/)(x). LEMMA 7 Let qbea subgradient of U at x. Then q is zero supporting to U at x. Moreover, if p is £ supporting to — Uat x, then ^ Proof. We have by definition fyeX, U(y)>U{x)-\-(q,y-x) So q is zero supporting to U. If F is « supporting to - U, we have for some q>0 \\y-x\\<ri=^-U{y)^ - U(x)-\-(p,y-x)-£\\y-x\\
CH. 5, SEC.6 GENERIC DIFFERENTIABILITY 283 Adding, we obtain + for \\y-x\\<t] Hence the result. ■ Using proposition 4, we see that all we have to do is to find a dense Gg subset ^ of fi such that for sd\xe A and e>0, the function -Vise supported at X We proceed to do this by defining a sequence A„ of open dense subsets of Q, whose intersection will be A. For each n>l, define A„ to be the set of all points x € ft such that for some (5 > 0, we have ||.x-a:i||<^ and /’i |i<? Il•^-^2ll<¿ and P2 e5i/n(—U)te)J ^ ^ « It is clear from the definition that A„ is an open set. LEMMA 8 Any point X e O where — U is l/n supported belongs to A„. A Proof. Take a: eQ and p e5i/„(- t/)(x). By definition, there is some t]>0 such that \\y-x\\<t\^-U{y)'^-U{x)+{p,y-x)-n-^\\y-x\\ On the other hand, dU(x) is not empty, which means that there is some qeX* with Vy, Gly) ^ U(x) +{q,y-x) By lemma 7, we have ||/?+4r|| \ and we rewrite the inequality as follows: Vy, U(y)^U{x)-{p,y-x)-n~'^\\y-x\\ Finally, we have proved that \\x-y\\<qMU{y)- UW-</),y-x>|<«"*||y-A:|| Now choose a:' 6 fi such that \\x'-x|| <q/2. The function Uis still continuous, and hence subdifferentiable, at x'- Taking any q edU(x), we have Vy, U{y)>U(x')+q',y-x') For any y such that l|x'-/ll “^e both preceding inequalities.
284 CH. 5, SEC. 6 A GENERAL VARIATIONAL PRINCIPLE This leads us to {q', [C/(;;)- i/(x)] + [t/(x)- t/(x')] Restricting ourselves to vectors y with ||x' — x|| < l|x'—j'll <>?/2, we obtain the inequality which implies {q'+P,y-x')<'in ^||;^-x'|| Now ifp' belongs to 5i/„(— t/)(x'), we have + by lemma 7 and hence, If x" eQ is another point where ||x"—x||<>7/2, and — 1/ is 1/« supported, we also have ||;7-/?"||<4n“^ for allp" eSi/„(- V)(x"). Hence, ||p'-/7"||<8«”*, which means x belongs to A„. ■ By theorem 5, the set A„ is dense in ii. Thus we obtain a sequence A„ of open and dense subsets of Q. Define A=^^A„ fi=i It is a Gi subset of i2 and dense by Baire’s theorem. We claim it has the desired property, given in lemma 9. LEMMA 9 The function U is Frechet differentiable at each point of A. k Proof Take x € ^. All we have to show is that - G is 1/« supported at x for all n: Since U is zero supported, differentiability will follow from proposition 4. By definition, for each n there is some d„>0 such that the set i^n={5,/„(-G)(>')|lb-xl|<5„} has diameter less than 8/«. Letting «-»oo and assuming the sequence 5„ to be decreasing, we see that the closures K„ build up a nested sequence of closed subsets whose diameters go to zero. Since X* is complete, their intersection is a singleton 3peX*: n ^x = {p} n=l
CH. 5, SEC. 7 PERTURBED OPTIMIZATION PROBLEMS 285 This implies that for any n, for all yeil with ||A:-y||<5„, and all p e5i/„(- U){y), we have Take any qedUiy)-, by lemma 7, we know that which yields U(yHU{x)-(q,x-y} < U(x)+(p, x-y)+n~mx-y\\ < U{x)+{p, x-y) + 10/j”i||x-y|| By theorem 5, the set of y with 5i/„(- V){y)4^0 is dense in fi. Since U is continuous, the preceding inequality extends to all y such that ||x—y||<5„, which means precisely that — [/ is 10«”^ supported by p at x. The result now follows. ■ 7. PERTURBED OPTIMIZATION PROBLEMS Let t/: X^IRu{+oo} be a lower semicontinuous function (not necessarily convex). Assume it is bounded from below inf U>-CO so that the conjugate function £/*: u {+oo} is proper and 0 e Dom U*. By definition, U* is convex; if Z* happens to be smooth, the function (/* will be Frechet differentiable almost everywhere on Dom U* by theorem 6.6. What does it mean for U itself? From an extensive analysis by Asplund, we extract the relevant information, given in lemma 1. LEMMA 1 Assume U* is Frechet differentiable at peX* and that the derivative is an element xof X. Then for any sequence x„eX such that U{x„)-{p, x„)^ - U*(p) we have ||x„-x||^0 A Proof Define y*: [0, +oo)->[0, +oo)u{+co} by y*(f):=sup {l/*(i)- U*{p)-{x,q-p)\\\q-p\U^t} {p is fixed). It is easily checked that y* is convex and lower semicontinuous.
286 CH. 5, SEC. 7 A GENERAL VARIATIONAL PRINCIPLE Since U* is Frechet differentiable at p, we have when i-+0. The conjugate function of y* is given by y(i)=sup {st-y*{t)\t>0} It is another convex, lower semicontinuous function on [0, +oo), and y(i)>0 for all s^O. We rewrite the inequality U*(qH U*{p)+ (x, g-p) + y*{\\q-p\U) for all qeX* as follows: U*{q+pH U*{p)+ (x, i>+y*(||ilU) for all qeX* and take convex conjugates. We obtain U**{y)-(p,y}>-U*{p)+y{\\y-x\\) for ally-6 X Since U**^ U, it follows that U{y)-{p. y)>- U*(pHy{\\y-x\\) for all;; 6 X If is a sequence such that U{Xn)-{py Xn) converges to — U*{p\ the preceding inequality yields ydlx«—x||)->0, which implies ||x„ —x||-^0 as desired. We can now combine this with theorem 6.6 to obtain results for closed sets and lower semicontinuous functionals. We first give a definition. DEFINITION 2 Let A be a subset of some Banach space X and let x be some point in A. We say that X is a strongly exposed point of A if there exists some /7 e X * such that x„ 6 A for all n e N, and {p,x„)^'m{ (p,x) =^X„-^X Any p with this property is said to expose x. A PROPOSITION 3 Assume X* is a smooth Banach space and A<=X is closed and bounded. Then there is a residual subset GofX* such that all p eG expose some x e A. A Proof Take U(x) = '^a{x\ the indicator function of A. It is lower semi¬ continuous, since A is closed, and [/* is everywhere finite, and hence contin¬ uous, since A is bounded. Now apply theorem 6.6 and lemma 1. ■
CH. 5, SEC. 7 PERTURBED OPTIMIZATION PROBLEMS 287 PROPOSITION 4 Assume X* is a smooth Banach space and A^X is closed, bounded, and convex. Then A is the closed convex hull of its strongly exposed points. A Proof. Let Ec/i be the set of exposed points, and consider co E. Obviously coEc/i. If coE:^/l, there is some point x e A, x $c6E. Separating x from coE by the Hahn-Banach theorem, we obtain some e A" * and a € (R such that (p, x)<(x< (p, y) for all;; 6 CO E Since A is bounded, ||;;|| for all ;; e ^ say, choosing another q eX* will perturb these inequalities into (q, x) = {p, x) + (q-p, x) =a' (q,y)>a-\\p-q\\m=a" forall;^eco£ For ||^-;?|| small enough, a'<a", so q will still separate x from co£. By proposition 3, there will be some q separating x from co £ and exposing some point jc in /1. So X 6 £, and <^,x>=inf {q,y)^{q,x) ye A But this contradicts the fact that q separates x from co£. Hence the result. ■ We now turn to functions U on X, with the aim of refining theorem 6.5. If X is a smooth Banach space and O: .Y-» IR satisfies the requirements of definition 6.1, let us agree to call functions of type <I>“^ all functions 'P; X->Rvj{+oo}, which can be written as 'P(x) :=c + <p, x> + aO ■ ‘ (m(x - Xq)) for some c6lR,p6Jli*, a6M,m€lR, Xo6.if- THEOREM 5 Let X be a Banach space and t/: X-+R u {+oo} a lower semicontinuousfunction on X. Assume X and X * are smooth. Then there is a dense subset D of Dom U such that whenever x 6 Z), there is a function 'P of type O “ * with (1) i. i/(x)^'P(x) forallxeX ii. i/(x)='P(x) Proof Define £ = i7+ T as in the proof of Theorem 6.5. Now consider the epigraph of £. It is a closed subset of but it is not bounded.
288 CH. 5, SEC. 7 a general variational principle However, it is easy to see that it has strongly exposed points. Indeed, consider its indicator function iI/ep(F)(x, a); we have a{Ep{F), (p, - b)) = sup {(p, x) - bF(x)\x 6 X} <00 provided 6>0 Applying lemma 1 and theorem 6.6, we see that the set of {p, — b) that expose some point (x, a) of Ep{F) is a dense G& in ]0, oo[ x X*. Let us choose {p, —b), with ¿>0 and the corresponding {x, a). It is easy to see that a=f(jc), so that (2) i. (p, x}—bF{x)> {p, x)—bF{x) ii. F(x)'^F(x)+{x—x,p/b) for all x: 6 X Now write U=F—'V to obtain (3) U{x)>{x—x,plb)+U{x)F^{x)—^{x) for all x: 6 A" The right side is a function of type O" *, which agrees with U at x=x. Hence the result. ■ COROLLARY 6 Let X be any point in D. Then there is some p 6 X* such that for any e > 0, there is some t]>0 with (4) ||x:—3c||<>/=> f/(x:)> l/(x)+ (p, x:—x)—e|lx—x|| Proof. Any function of type d> ‘ is Fréchet differentiable on its domain. The result then follows from theorem 5, with p=<I)'(x) ■ We now turn our attention and efforts to perturbed optimization problems. We are given two Banach spaces X (parameter space) and V (state space), an open subset Q c X, and a function F: Q X F-^IRu{+oo} such that (5) F is lower semicontinuous on Q x F (6) F{x, v) finite^F(•, i;) is Gateaux differentiable at x. We define {/(x)=inf {F(x, t))|i; 6 V}
CH. 5, SEC. 7 PERTURBED OPTIMIZATION PROBLEMS 289 If this infimum is finite, we ask ourselves whether it is attained. In other words, with every x 6 X, we associate the optimization problem infF(x,i;) veV A solution of (is a point v eV where F{x, v)= U{x) It is a well known fact from optimization theory that the existence of a solu¬ tion for {^x) is closely related to the subdifferentiability of C/ at x (uniqueness being related to differentiability). Let us give a precise statement. THEOREM 7 Assume that at some point 5c eX, there is an open neighborhood Jf of 5c and a function 'V : Ji-t-U such that (7) (8) (9) 'F is 'P(x)=i7(x) 'P(x)< V(x) for all xe Ji Then for any minimizing sequence of v„ of (10) F{x,v„)^U(x) there is a sequence x„^x in Ji such that ^ i «• F(x„, v„)-^U(x) in IR ^ ^ lii. f;(x„,p„)^4"(x)mX* ▲ Proof Define e„=f(x, p„)-'P(x), so that e„^0 and e„->0. For fixed «, define a function G„ on ^ by G„(x)=F(x, «^-^(x) This is a non-negative Gateaux-differentiable, lower semicontinuous function on We can assume n to be so large that the closed ball B„ with center x and radius 2^„ is contained in Ji and apply the e variational principle to G on B„, starting from the point x. We obtain a point x„ such that
290 CH. 5, SEC. 7 A GENERAL VARIATIONAL PRINCIPLE l|GUxn)IU = ||f;(x„, t^„)-4"(x„)|U<V^ Letting «->0, we obtain the desired result. COROLLARY 8 Assume moreover that the family Fi(-, t;„),« 6 N, is equicontinuous at x (12) Va>0, 3rj>0: ||;c-x||^>y=>||F;(A:, Vn)-F'x{x, t^„)IU<e V« Then, Fi(x, !;„)-►'P'Cx). Proof Write WF'xix, i;„)-4^'(^)|U^||FUa:„, t;„)-Fi(x, i^JU + l|Fi(x„, t^„)-4^'(i)IU The second term on the right goes to zero by condition (11, ii) of theorem 7, and the first term also goes to zero because x„-^x and the family F'x{^, v„) is equicontinuous. ■ Certainly corollary 8 is more transparent than previous statements: It tells us that minimizing sequences v„ at x have to be such that F^x, Vn) converges strongly in X*. This is a considerable amount of additional information and can be used in many cases to prove that the sequence v„ itself converges. For the sequel, recall that a map between metric space is proper if any sequence whose image converges has a convergent subsequence. PROPOSITION 9 Assume that for some point x ef2, there is a function 'P: X-^U, which is in some neighborhood of x and satisfies (13) ^(3c)= i7(x) and 4^(x)^ U(x) for all xsO. Assume moreover that there is some e>0 such that the family {Fx(*, v)\F{x, pX i/(i)+e} is equicontinuous at x. Assume finally that the map Fi(x, •)from V to X* is proper. Then problem (^x) has at least one solution x. If in addition U is Gateaux differ¬ entiable at X and the map Fi(x, •) is one to one, the solution is unique. к
CH. 5, SEC. 7 PERTURBED OPTIMIZATION PROBLEMS 291 Proof. Any sequence Vn such that Fi(3c, v„) converges must have convergent subsequences, Vnf^^v, say. If we take a minimizing sequence for we obtain a minimizer for v, since F is lower semicontinuous. For the last part, just note that we must have ^'(ic)= i/'(3c), and the equation v)= U'ix) defines v uniquely. For instance, if U is on some open set, we can take U=^ on that set. Only rarely do we come across such regular behavior of U. On the other hand, it is often the case that U is continuous, which motivates the following proposi¬ tion. PROPOSITION 10 Assume that X is reflexive and there is an open subset on which U is continuous. Assume moreover that whenever we have sequences Xn in Q and v„ in V such that the sequence (x„, F{xn, Vn\ F'x{x„, v„)) converges in X xU xX"^, then the sequence v„ must be compact. Then there is a dense subset such that whenever xeD, the problem {^x) has a solution. Proof. Since X is reflexive, then X and X* are smooth, and we can apply theorem 6.5. Now norms that are Frechet differentiable away from the origin are in fact (see Diestel, Chapter 11, Section 2). It follows that function of type ^ are CMf we build from a differentiable norm. So the conditions in theorem 7 are fulfilled at all points of a dense subset; we conclude as in proposition 9. If we slightly strengthen the assumptions, we obtain uniqueness in a strong way. PROPOSITION 11 Assume that X is reflexive and there is an open subset 0.ciX on which U is con¬ tinuous. Assume moreover that whenever we have sequences x„ in Q and v„ in V such that {Xfi, F(Xfl, Vfi\ Fl^n)) converges in X xU x X*, then v„ converges in V. Then there is a dense subset such that whenever x eD, the problem J has a unique solution v, and all minimizing sequences of (converge to v. k
292 CH. 5, SEC. 7 a general variational principle Proof, Existence at all points of D follows from proposition 10. Take xeD\ if had two solutions v\ and ^2, we could construct a minimizing sequence by alternatively setting v„ = vi and v„=V2- By theorem 7, we would get a sequence x„-^x with Fx{x„, v„) converging in X*. But then v„ itself would converge, which is absurd, unless vi =V2. So there is only one minimizer. If v„ now is any minimizing sequence, the Fx(x„, p„) will converge, so v„ itself will converge, and the limit must be the minimizer. ■ If we throw in equicontinuity, we can strengthen the result still further. In the following proposition, we keep the notations and assumptions of proposition 11. PROPOSITION 12 Assume in addition that, for every x eQ, there is some a>0 such that the family {F'xi', y)|F(x, v)^ C/(x)+a} is equicontinuous at x. Then the set D of existence and uniqueness contains a dense Gs of SI. A Proof Density is the previous result; to obtain a dense we have to use Baire’s theorem. To do this, take any s>0. We shall say that x has the property Pg, or simply Pfx), if all minimizing sequences of {^x) are eventually confined within some ball of radius <6. 3r<e: [F(x, v„)-^ U(x)]=>[3N: ||f;(:)c, v„)-F'x(x, v, m/IlH« V«, m^N] Set i2e = {A:lP£(x:)}. Clearly so 0^ is dense. On the other hand, Ps(x) is easily seen to be equivalent to: 3r<s, 35>0: F(x, v)— U(x)^5 and F{x,w)-U(xH5 =>\\Fx{x, v)-F'x(x, w)|U^r In this form, it is clear from the equicontinuity that Qg is open. The set of existence and uniqueness then certainly contains n iii/.. n e which is a dense Gs by Baire’s theorem. ■ Note that in none of these propositions have our assumptions been strong enough to ensure that any individual problem has a solution. We are thus
CH. 5, SEC. 7 PERTURBED OPTIMIZATION PROBLEMS 293 in a position where we know that for most, or at least many, values of the param¬ eter X the problem can be solved, without being able to pinpoint any of them. Let us conclude with an example. Let ^ be a closed subset of a Banach space X, iQtxeX be given. We seek the point in A closest to x if there is any. The problem is set up as follows: inf (llx-J'll + lAxi^) yey So V = Xy F(x, y) = \\x—yW + \l/A(y\ and U{x) obviously is the distance from X to A U{x)=d{x, A) It is known to be continuous everywhere. The function F itself is lower semi- continuous. If X is reflexive and the norm is chosen to be Frechet differentiable away from the origin, F is Frechet differentiable at all points x^y, with deriva¬ tive F'Ax, y) = '-j(x-y) Set i2=X \^. If X eQ and F{x, y)< +co, we must have ye A, so x^y and F{’,y) is differentiable at x; therefore, assumption (6) is satisfied. Suppose we have sequences x„ in Q and y„ in X such that x„-*x in Q Ik»-J»ll + ^Aiy»)-^d{x, A) x„-y„ j {x„-y„)=j jkn->'»ll inX* Now j sends the unit sphere S<=X into the unit sphere S*<=X*, and is characterized by (j{z), z) = 1, all z € S (see Diestel Chapter 2). Since X is reflexive, its unit ball is weakly compact, so/must be surjective. The derivative j*: S*^S of the dual norm is characterized by (/*(9)> ^ t ^ It follows that j*=j ~ and j is a homeomorphism. Thus we have x„-y lk»-.y»ll in S Since x„-*x i A and y„ stays in A, it follows that y„ converges to some limit ye A. All the assumptions of proposition 11 are satisfied, and we obtain a well- known result, stated in proposition 13.
294 cH. 5, SEC. 7 a general variational principle PROPOSITION 13 Assume X is reflexive and the norms of X and X * are Frechet differentiable off the origin. Then for any closed subset Aof X, there is in X \A a dense set D ofpoints X with a unique projection on A: All sequences yn e A such that \\x—yn\\-^d{x. A) converge to the same limit in A. k If we strengthen the assumptions slightly, we obtain a much better result. Recall that X * is uniformly convex if for every 6 > 0, there is some <5 > 0 such that (14) peS*, p eS^ and p+p PROPOSITION 14 Assume in addition that X* is uniformly convex. Then the subset Z)cQ contains a dense Gs. k Proof. We have to check the additional assumption of proposition 12, namely, that the family , y\ iox ye, is equicontinuous on Q. Fix X e SI and 6 > 0, and set Since AT* is uniformly convex, there is some ¿>0 such that ||/>j,+;>IU>2 —5 implies II/?),—Take x' in S with ||x—x'||<5, and set F'xix', y)=p'y Then 11/^)’“hPj^*^\^Py~^Pyy -^)l 5= I{py, x) + {p'y, X') + (p'y, x-x')\ ^2-S thereby forcing II/?),—/?J,IU<fi independently of y, as desired.
CHAPTER 6 Solving Inclusions This rather long chapter deals with one central theme: solving “inclusions”; that is, when F is a set-valued map from XioY, finding x 6 Dom (F), a solution to y eF{x) when ;; is given in Y which offers a wide array of applications. Curiously enough, mathematical problems arising from game theory moti¬ vate important theorems in nonlinear analysis. Relations between game theory and nonlinear analysis are so deep that a detour through game theory does, in fact, save time before presenting nonlinear analysis. Moreover, a deeper insight is gained by building up intuition. After a short presentation of two-person games in Section 1, we prove in the second section the following theorem. LOP SIDED MINIMAX THEOREM Let M be a compact convex subset, N a convex subset, andf:Mx N^R satisfy i. Vj; 6 N, x-^f{x, y) is convex and lower semicontinuous, ii. Vx e M, y-^f {x, y) is concave. Then there exists xeMsuch that sup/(x,>;)=sup inf f(x,y) y^N yeN xeM Observe that this equality implies the minimax equality inf sup/(j>c,;;)=sup inf f{x,y) xe M ye N yeN xe M We shall propose a proof where the roles of convexity assumptions, on one hand, and topological assumptions on the other are well separated. We then present a more sophisticated version of the minimax theorem, where the compactness assumptions are dramatically relaxed. In the third section, we drop the convexity assumption of / with respect to x. However, we still prove a 295
296 CH. 6, SEC. 1 SOLVING INCLUSIONS minimax equality involving continuous maps between the strategy sets of the two players. THEOREM Let M be a compact space, N a convex subset, andf:Mx N^R satisfy i. € N, x-^f(x, y) is lower semicontinuous. ii. VxeM, y^f{x, y) is concave. Then there exists xeM such that sup/(x,j)= sup inf/(x, C(x))= inf sup/(£>(>;), j) A yeN xeM De^{NM)yeN This is a less known but very useful equality. It is equivalent to the Ky Fan inequality, which plays a crucial role in proving in a very easy way the existence theorems in this chapter. KY FAN INEQUALITY Let K be a compact convex subset and (¡> \ KxK-^R satisfy i. 6 K, x^(l)ix, y) is lower semicontinuous. ii. "^xeK, y^4>(x, y) is concave. iii. "^yeK, (l>{yyy)^0 Then there exists x eK satisfying sup (¡>{x, yeK We proceed by generalizing this inequality to the case where we replace the lower semicontinuity of (/> with respect to x by the monotonicity condition Vx, yeK, (j)(x, y)^-(!>{y, x)^0 and the lower semicontinuity of (/> with respect to x for only a very strong topology, called the finite topology. Ky Fan’s inequality is equivalent to the Brouwer fixed point theorem. Its analytical formulation—unlike the pleasant geometrical form of the Brouwer theorem—makes it more operational for proving most of the results in nonlinear analysis. In the fourth section, we in¬ vestigate the problem of finding a zero of a set-valued map F, solution to the inclusion 3 3c e K such that 0 6 F (3c)
CH. 6, SEC. 1 SOLVING INCLUSIONS 297 or a fixed point of F, solution to the inclusion K such that e F(x^) The common feature is that we shall derive all these results from the Ky Fan inequality, and, therefore, in the final analysis, from the Brouwer fixed point theorem. All the set-valued maps investigated in this section are upper hemicontinuous maps with closed convex values from a compact subset K of X to a topological vector space Y. The main theorem of this section states that when F: K-^X=Y satisfies the tangential condition WxeK, F{x)nTK{x)^0 then i. there exists xeK,a, solution to 0 € F(x). ii. "iy eK, 3 jc € X, a solution io y ex—F{x). We single out several generalizations and consequences. Among them we obtain the Kakutani fixed point theorem and other fixed point theorems. We follow Poincare’s continuation method for carrying the preceding theorem by homotopy and thus proving an adaptation of the Leray-Schauder theorem to set-valued maps. We shall apply these results to prove the existence of a Walras equilibrium of an exchange economy as well as another concept of equilibrium, which avoids some of the shortcomings of the Walras equilibrium (Section 5). There are, however, examples of set-valued maps that are not upper semi- continuous; for instance, the subdifferential x-^d U{x) of a lower semicontinuous, convex function. Unfortunately, existence theorems in the fourth section do not apply to them. Let us consider an example: X is a weakly compact convex subset of a Hilbert space X, and [/ is a lower semicontinuous, convex function, related to X by the condition 0 e Int(Dom U — X). We know that there exists an element 3c 6 X that minimizes U on X, that is a solution to the inclusion (*) 0 edU{x) + Nk{x) Therefore, there must be a property of x-^d U(x) other than upper hemicontinu- ity that allows us to solve such an inclusion. This property is monotonicity. A set-valued map A from X to X is monotone if and only if W{x,p), graph (^), (p-q,x-y)>0. It happens that this algebraic property balances some of the continuity requirements we made in the preceding sections.
298 CH. 6, SEC. 1 SOLVING INCLUSIONS Inclusion (*) is a particular case of a class of problems of the form (*♦) f sA(x)^dV{x) where .4 is a monotone map with weakly compact convex values, which we assume to be only finitely upper semicontinuous and where K is a lower semi- continuous, convex function. We shall prove that Int(Dom F*4-y4 Dom V)cz\m{A + dV)c:Dom V* + A Dom V and how more restrictive monotonicity assumptions imply that A-\-dV is surjective. Subdifferentials of lower semicontinuous, convex functions are actually maximal monotone, in the sense that there is no strict monotone extension. Minty’s theorem characterizes maximal monotone maps: They are monotone maps such that l-\-A\s surjective. They enjoy many properties. In particular, they can be approximated by Lipschitz single-valued maps Ax called the Yosida approximations of A, which are systematically used when dealing with maximal monotone maps. We study surjectivity properties of maximal monotone maps and prove that strongly coercive maximal monotone maps are surjective. We also provide conditions under which two maximal monotone set-valued maps A and B satisfy the properties i. Int(Im (A + B)) = Int(Im (A) -h Im(^)) ii. cl(Im (A + B))=cl(Im(yi)+Im (5)) We deduce the existence of solutions to inclusions y 6x-hABx when A and B are maximal monotone maps. We conclude this section with the most important property of maximal monotone maps related to the existence of solutions to differential inclusions x'(t) € —A(x(t)); x(0)=xo is given in Dom(^). We prove that there exists a unique solution to the initial value problem, which, amazingly enough, is “lazy.” For almost all i, -x'(t)=m(A(t)), the velocity with the smallest norm. Furthermore, if we denote by Txo e ^(0, oo; X) the solution of the preceding initial value problem, we observe that sup ||T(xo)(f)- <lko-^ill i>0
СН. 6, SEC. 1 MAIN CONCEPTS OF GAME THEORY 299 1. MAIN CONCEPTS OF GAME THEORY Let us consider two players, Mike and Nancy. Let M and N be their strategy sets. Our purpose is to single out pairs {x, y)eMxN by various optimization methods. This being said, how can we devise these methods? In this chapter, we suggest deriving them from decision theory, which means providing decision¬ makers mechanisms for selecting elements (called decisions) in given subsets (called decision sets). History shows that parlor games provided mathematicians since Blaise Pascal with various problems, so such terms as players (instead of decision makers) and strategies (instead of decisions) were coined quite early, and tradition maintains them. The status of games as an application of mathe¬ matical theory is due to von Neumann, who proposed the general framework of conflict and cooperation. An elementary mechanism that allows Mike and Nancy to choose their respective strategies is obtained by giving them decision rules. DEFINITION 1 A decision rule for Mike is a set-valued map Cm from N to M. It assigns to each strategy у e N played by Nancy a strategy x e См{у) that can be implemented by Mike when he knows that Nancy plays y. A Similarly, a decision rule for Nancy is a set-valued map from M io N associating with each strategy a: 6 M a strategy ;; 6 Cn{x) played by Nancy constrained by the choice xeM made by Mike. N M Figure 1. Case of a game with no consistent bistrategies.
300 CH. 6, SEC. 1 SOLVING INCLUSIONS Decision rules Cm and Cn being given to Mike and Nancy, we are naturally led to single out pairs of strategies—called bistrategies—that are consistent in the sense that (1) a: e CM(y) and ;; e Cn{x) DEFINITION 2 A pair of strategies—or bistrategy—{x, y) satisfying property (1) is said to be a pair of consistent strategies—or a consistent bistrategy—for decision rules Cm and Cn of Mike and Nancy. ^ iWX N M = Mike's strategy set The relevance of such concepts depends on the choice of the decision rules. A trivial example is given by constant decision rules. Indeed, we can identify a strategy X 6 M of Mike with the constant decision rule ;; e N-^x € M, which describes stubborn behavior by Mike, who plays x whatever the strategy played by Nancy is. So, when Mike and Nancy, respectively, play constant decision rules X and y, the associated consistent bistrategy is the pair (x, y). The set of consistent bistrategies may be empty or very large, depending on the properties of Cn and Cm- The problem of findin^pairs of consistent strategies amounts to a fixed point problem. We denote by C the set-valued map from MxN io itself defined by (2) C(x,>')=Cm(3')xCw(x)
CH. 6, SEC. 1 MAIN CONCEPTS OF GAME THEORY 301 Inclusions (1) can obviously be written (3) (x, y) e C(x, y) The search for fixed point being a central theme of this book, we shall describe various methods for devising sufficient conditions for the existence of consistent bistrategies. So, decisions rules for Mike and Nancy provide a selection mech¬ anism yielding consistent bistrategies. If it is not sharp enough, in the sense that the set of consistent bistrategies is too large, we may need a further mech¬ anism selecting bistrategies among the consistent ones and so on. Before going further, a first question arises: where do the decision rules come from? How can we construct them? Gaines in Normal (or Strategic) Form: Noncooperative Equilibria The traditional way of presenting game theory is to posit that each player classifies pairs of strategies through a real-valued function. We can think of such a function as a map that associates to each pair of strategies {x, y) its cost, measured by a real number. Since the concept of cost involves the notion of money, which is quite difficult to master in economics, we prefer to call it a loss. So, a player uses a loss function f:Mx N^R to define the preference preorder on M X N sis follows: (4) (xu yi) is preferred to (x2, yi) if and only iffixi, yi)^ f{x2, yil Whatever the relevance of this assumption is, we posit from now on that Mike and Nancy select their strategies according to given loss functions/^: M x N^R and ffi \ Mx N^R, respectively. We set fix, yY=ifMix, y\ fNix, y)) e R^ DEFINITION 3 A game in normal (or strategy) form is defined by a map f from M xN- called the biloss map. R\ к There is a natural way to associate decision rules with a game described by loss functions. Indeed, let/m be Mike’s loss function. If he has the opportunity of knowing a strategy yeN played by Nancy, he will be tempted to choose a strategy X that minimizes his loss given Nancy's choice: In other words, he will choose a strategy in the subset Cuiy) defined by (5) См(у):= {x 6 М\/м{х, у) = min/м(х, 7)}
302 CH. 6, SEC. 1 SOLVING INCLUSIONS This set-valued map Cm defines a decision rule for Mike. Similarly, Nancy may associate to her loss function ffi the decision rule defined by (6) Vx 6 M, Cn(x):= {e M|/jy(x, j5)=min/wix, y)} yeN DEFINITION 4 The decision rules Cm cind Cn associated to the lossfunctions/м andby formulas (5) and (6) are called canonical decision rules. Л consistent bistrategy (i, y) for the optimal decision rules is called a noncooperative equilibrium of the game A Therefore, a pair (3c, y) is a noncooperative equilibrium if (7) хеСм(у) and yeCN{x) Formulas (8) and (9) provide an equivalent definition. PROPOSITION 5 A pair of strategies {x, y) is a noncooperative equilibrium if and only if (8) and (9) fM{x,y)=rmnf{x,y) xeM ffiix,y) = mirifn{x,y) ye M So, a noncooperative equilibrium is a situation where each player optimizes his or her own criterion, assuming that the partner’s choice is fixed. In other words, this is a situation of individual stability. Pareto Optima and Conservative Strategies Can we accept the concept of a noncooperative equilibrium as the only reason¬ able concept of solution? Not necessarily; in particular, not when we assume that the players can communicate and exchange informations. When they do so, they may see that there exist other pairs of strategies that satisfy (10) /м(х, у) < /м{х, у) and /д,(х, у) < /ц{х, у) that is, yield both Mike and Nancy (strictly) smaller losses than those assigned by the noncooperative equilibrium. This, when it happens, reveals a lack of collective stability, because both players can find better strategies.
CH. 6, SEC. 1 MAIN CONCEPTS OF GAME THEORY 303 DEFINITION 6 A pair of strategies t*) is said to be Pareto optimal if there is no other pair (x,y)BMxN such thatfuix, < fu(x^, andfdx, y) < f^x^, A Does there exist noncooperative equilibrium that are not Pareto optimal? Unfortunately, there are many examples of this situation. No general theorem asserting the existence of noncooperative equilibria that are also Pareto optimal is known to the authors. The preceding diagram represents the subset f{MxN)cR” of bilosses yielded by all the pairs of strategies. We also represent by a thick line the bi¬ losses of Pareto optima. We see that selecting Pareto optima is not a sharp selection procedure. Assume, for instance, that there exists a pair (3c, j;) that achieves the minimum of Mike’s loss function (11) /m(3c, y) = inf /m(x, y) = -MM xeM yeN It is clear that (x, y) is a Pareto minimum. For Nancy to agree to this situation would probably mean that her only aim in life is to please Mike. Also, any pair (ic, y) of strategies that minimizeson M x N is a. Pareto optimum (12) fdx, y)= inf fN(x, t)=:«n xeM yeN
304 CH. 6, SEC. 1 SOLVING INCLUSIONS We note that if a pair (3c, y) minimizes both andon M x Ny then it is the best candidate for a concept of solution because /m(^, y) = «M' /n(^, y) = O^N However, such a situation is naturally quite exceptional. This is why we call the vector (13) <Xn) the shadow minimum (or virtual minimum) of the game. We also note intuitively that neither the pair (3c, j;) [defined by (11)] nor the pair (3c, j)) [defined by (12)] is a realistic choice in the framework of game theory: If it is reasonable for players to agree to choose pairs of strategies that do not yield both players smaller losses, it is not obvious that one of them will agree to give the other the entire benefit (see the preceding figure). Actually, cooperative game theory provides mech¬ anisms of selecting among Pareto optima. Conservative Strategies and Values The case when Nancy’s behavior is to please Mike without taking into account her own interest leads to strategies (;c, y) defined by (11). Assume that Nancy exhibits the opposite behavior. Her only aim is to hurt Mike, and Mike knows it. (Actually, whenever Nancy behaves kindly, it suffices for Mike to believe that she is nasty.) So, he assigns to each strategy x 6 M the worst loss/m(^) (read / sharp sub M) defined by (14) ye N and by doing so, looks for strategies x*eM that minimize / m over M (15) f ^ix*)= ini fiSix) xe M We say that x* is a conservative strategy for Mike. We set (16) inf sup/m(jc,3/)= inf /¿¡(x) xe M ye N xeM and call it Mike’s threat value. Indeed, he can always reject a pair of strategies (Xy y) that satisfies (17) fM{x,y)>v^ since any conservative strategy x*eM yields Mike a lossfu(x *, y) (strictly)
CH. 6, SEC. 1 MAIN CONCEPTS OF GAME THEORY 305 smaller than y)> So without any agreement with his opponent, he can always threaten to implement his conservative strategies. Similarly, Nancy’s threat value is defined by (18) vn:= inf sup Mx,y)= inf //(j) ye N xe M ye M We say that the vector (19) is the threat vector. So, pairs of strategies worth considering are those satisfying /(x, y)^v^. The Duopoly We present the basic example of duopoly, where both players are producers. The loss functions are the cost functions, and they depend on the production of the two players. This game and the concept of noncooperative equilibrium was introduced by the French economist Cournot in 1838. We suppose that Mike and Nancy produce a given good, assumed to be homogeneous. We denote by x €/?+ and yeR+ the quantities of this good produced by Mike and Nancy, respectively. We assume that the price (20) p(x,y):=ot-p{x+y) is an affine function of the total production x+y(<x>0y P>0) and cost functions c and d of each producer are also affine functions of the individual productions (21) c{x)=yx-^S, The net cost for Mike is d(y)=yy-\-S, y>0, 0^0 y — OL fM{x, yY=yx-\-6-p{x, y)x=px[x-\-y-h—^ 1 + 5 P and the net cost for Nancy is /n(x, y)=yy + d-p(x-^y)y=Py(^x+y+^-j^ + S We do not change the game by setting j? = 1 and 5 =0. So by setting w=y — a, the duopoly can be regarded as a two-person game, where (22) M:=N:=[0,u]
306 CH. 6, SEC. 1 SOLVING INCLUSIONS and the loss functions are (23) '• fM{x,yy.=x{x-\-y-u) »'• fN{x,y):=y{x+y-u) The biloss map is defined by (24) J{x, y)={x{x+y-u), y(x+y-u)) which maps the upper triangle (25) r+ := {(x, y) € [0, m] x [0, u]\x+y^u} onto the square S+ := [0, m^] x [0, m^], the diagonal (26) To := {(x, y) € [0, m] X [0, u\\x+y=u} onto {0}, and the lower triangle (27) T_ :={(x, y) 6 [0, m] x [0, m]|x+j^m} onto the triangle S-:=|(/,0)€|^-^,oJx|^-^,O (28) We observe that the subset f+g>- (29) is mapped onto the subset (x, j) € [0, m] X [0, m]|x+;^=2 7(T)={(y;0)e Therefore, (30) the subset P defined by (29) is the set of Pareto strategies, The bistrategy (31) u u
CH. 6, SEC. 1 MAIN CONCEPTS OF GAME THEORY 307 is Pareto optimal; if the producers agree to cooperate, a reasonable compromise is the pair (xp, yp). X It is clear that (32) fu{x):= sup x(x+y-u)=x^ achieves its minimum at x=0 and symetrically, / *{y) =y^ achieves its minimum at y=0. Therefore, the conservative bistrategy is (0, 0), that is, both players produce nothing. The threat vector v* is equal to (0,0). We note that the shadow minimum is equal to {—u^/4, —u^l4). The concept of noncooperative equilibrium in the case of the duopoly was introduced by Cournot in 1838. Let y be Nancy’s production. Then Mike will implement a production x that minimizes over [0, m] his net cost function x->x(x+y—m); the minimum is achieved at the point (33) and is equal to (34) x=CMiy)=Uu-y) fhiy) = inf [x(x +y-u)'\ = - (u-yf So, the map Cm defined by (33) is Mike's canonical decision rule, and similarly, the map Cn defined by (35) CN(x)=Utt-x)
308 CH. 6, SEC. 1 SOLVING INCLUSIONS is Nancy's canonical decision rule. Therefore, the noncooperative equilibrium of the duopoly is the fixed point of the map (x, y)-^{CM(y), Cn{x)), that is, the point (36) ^c=3-yc=3 which yields to each player a cost equal to — u^/9. We note that the noncoopera¬ tive equilibrium is not Pareto optimal. /3 Mike's strategy set Actually, we observe that the noncooperative equilibrium can be reached by implementing the following algorithm^ When Nancy produces y2,,-i at the odd period 2« —1, Mike produj:es X2n=CM(y2n-i) at the even period 2«, and then Nancy produces y2n+i =C^(x2„) at the odd period 2n +1, and so on, the players taking turn in responding to each other. Since the sequences (x2n) and (y2n+i) are, respectively, the even and odd subsequences of the sequence of elements Zk defined by 2zk+i+Zk = u multiplying each equation by ( -1)*^ *2* and summing them, we obtain M 1+2-" — -- _ + (_l)«+i2 71 + 1 - W - 1 , 2 1+2- So z„ and, consequently, both X2n and y2n+i converges to m/3.
CH. 6, SEC. 1 MAIN CONCEPTS OF GAME THEORY 309 We can associate another game with the duopoly, where instead of choosing strategies, the producers choose decision rules. Let us consider Mike’s point of view. He may decide to play an affine decision rule of the form (37) CM{yy-=a{u—y) where ae[0,1] This means that he produces nothing whenever,Nancy produces the maximal production u and he decides to produce au whenever Nancy produces nothing. When Nancy implements an affine decision rule of the form (38) C'’N{x):=b{u—x) where ¿6 [0,1] the associated consistent bistrategy is ffl(l — b)u b(\ — a)ii\ ' \-ab ' l-ab which yields the following costs: a{l-a)(\-bfu^ (40) i. guia, b):- - ii. gl^(a, b)\= - {\-abf 6(l-6)(l-a)V {\-abf Therefore, we have a new game, whose strategies are the slopes of the decision rules. For instance, if Nancy plays slope b, Mike will play the slope a^Guib) that minimizes a-*gM(a, b). We find (41) a=aM{b):= 1 2-6 For instance, if Nancy plays her canonical decision rule C^, which corre¬ sponds to the slope 6=1/2, and if Mike knows (or guesses) her move, he will play the slope <t(1/2)=2/3. The associated decision rule C*/^ is called Mike’s Stackelberg decision rule. The consistent strategies for the decision rules CUf and C}/^ = Cn are equal to (42) xs=^, Js=4 They form the so-called Stackelberg equilibrium for Mike. The associated costs are given by (43) fM(xs,ys)=- 8’ fnixs, 3's)= ~
310 CH. 6, SEC. 1 SOLVING INCLUSIONS By implementing a Stackelberg equilibrium, Mike is then in a better situation than in the original noncooperative equilibrium ( —wVS instead of — wV9). Mike’s advantage lies in the fact that he knows that Nancy will play her canonical decision rule. What would happen if Nancy followed the same reasoning? She would also implement her Stackelberg decision rule Consistent bistrategies for the Stackelberg decision rules are equal to (44) 2u yu=- 2u They form the so-called Stackelberg disequilibrium.
CH. 6, SEC. 1 MAIN CONCEPTS OF GAME THEORY 311 The producer’s cost will equal —2u^l25, and both Nancy and Mike are in a worse situation than in the noncooperative equilibrium case. For both players, the Stackelberg decision rules yields a smaller loss than the canonical decision rule, but both players would be better off both choosing the canonical decision rule than both choosing the Stackelberg decision rule. TABLE 1 ( u^\ ( 2«^ 2m^' Stackelberg ~ 1^, The noncooperative equilibrium (a, 5) for the game played on decision rules is the solution to the problem (45) a=<r(b) and b=a(a) which is 5=1, ¿»=1. The associated consistent bistrategies {x, y) are those defined by equation x-\-y=u. We note that for such strategies, the costs are equal to zero. This non- cooperative equilibrium can be achieved by implementing the following algorithm. Nancy begins by playing the slope 1/2 of her canonical decision rule, then Mike implements his Stackelberg decision rule 2/3=<r(l/2), then Nancy implements the slope ct(2/3)=<t^(1/2), and so on. It is easy to show that the se¬ quence of slopes is equal to 1 — 1/« since fr-lL—!--1--L V nj 2-l + l/n n+1 and they obviously converge to the slope one. Therefore, Mike will implement the slopes 1 —1/3, 1 —1/5,..., 1 —1/(2«+1) during the successive even period, while Nancy implements the slopes 1/2,1 -1/4,1 -1/6,..., 1 -1/2« during the odd periods. During the even periods, the consistent strategies will be X2n = r yin = (2«-2)m 2(2«-1)
312 CH. 6, SEC. 2 SOLVING INCLUSIONS and during the odd periods, the consistent strategies will be ■^2/1+1 =~ (2n—\)u 4n ’ 3^2«+1 —^ U They converge to the bistrategy (w/2, u/2). The analysis of this simple game emphasizes the conceptual difficulties of game theory. The bistrategies we described—the conservative solution (0,0), the Pareto minimum (w/4, m/4), the noncooperative equilibrium (w/3, m/3), the Stackelberg equilibria (m/4, u/2) and {u/2, m/4), the Stackelberg disequilibrium (2m/5, 2m/5), and the bistrategy {u/2, m/2) just mentioned—each have their own interest. This variety of reasonable concepts indicates the need for richer structures. We shall now turn our attention to the mathematical concepts of the theory. 2. TWO-PERSON ZERO-SUM GAMES: THE MINIMAX THEOREM We now consider the important class of two-person games that satisfy the condition (1) Vx6M, yeN, /m{a:, j)+/w(x,3;)=0 So, Nancy’s loss is Mike’s gain and vice versa. We observe that every pair of strategies is Pareto optimal, so that this concept is no longer interesting: Indeed, /(Mx N) is contained in the second bisector. (2) So if we set /m(^, y)-=f{x, y), /n(x, y):T= -f(x, y) (3) /’^(x):=sup/(.x, j), f'^:= inf sup/"(x, y) xeMyeN
CH. 6, SEC. 2 TWO-PERSON ZERO-SUM GAMES: THE MINIMAX THEOREM 313 (read/-sharp and a-sharp) and (4) /'’(j):= inf/(x, j), a'’:=sup inf/(x,>') x€ M ye N xe M (read/ flat and v flat) and (5) M*:={x*eM\f*{x*)=v*}, N^:={y” e N\f’iy‘’)=v'>} we have (6) fUx)=f*{x), fM(y)=-f\yl vti=v*, vH=-v'’ We observe that AT** and N'’ are subsets of Nancy and Mike’s conservative strategies, respectively. Since /'’(y)^ f*(x) for all X 6 M, 6 N, we deduce that (7) or equivalently, (8) v*\=:{v*, — a'’) lies above the second bisector. We say that the interval (a'*“, a'’) is the duality gap of / The set of pairs of stra¬ tegies (x, y) satisfying fix, y)^v’^ is (9) K:= {(x, y) e M X N\v'’^/(x, y)^v*} It contains the set M* x if’. There are situations where a*" is strictly less than v*. Example Consider the finite game M={1, 2}, W={1, 2, 3}, where/ is described by the following matrix. Nancy selects columns 1 2 3 1 -6 2 2 4 -5 -4 Mike selects rows
314 CH. 6, SEC. 2 SOLVING INCLUSIONS The entries of this matrix represent Mike’s losses. So Mike’s biggest losses are 2 and 4, respectively, and thus Mike’s conservative strategy is the first row and v^=2. Nancy’s least gains are, respectively, —6, —5, and —4 and thus her conservative strategy is the third column and v^=—4. The pairs (1, 2), (1, 3), and (2, 3) belong to the subset K, Let us try to play that game for ourselves. First let Mike implement its conservative strategy (first row). He expects Nancy to choose the second column. But the conservative strategy for Nancy is the third column, and she expects Mike to choose the second row. But if Mike is informed of this choice (or guesses it), then he would do better selecting his second row (with a loss of — 4) instead of the first one. Similarly, if Mike chooses his conservative strategy, then Nancy would do better playing her second row (with a gain of two) instead of the third. This “wheels-within-wheels” situation illustrates the lack of noncooperative equilibrium; indeed, the canonical deci¬ sion rule C ^:=Cm for Mike is defined by C’^(1)={1}, c*{2)={2], C*(i)={2} and the canonical decision rule C'’.=C*for Nancy, by Q(l)={2}, C‘(2)={1} So the decision rule C:=(C^ C*) has no fixed point, as can be directly checked. The absence of noncooperative equilibria when v^<v* is actually a general fact: The following result shows that its existence requires very stringent conditions. PROPOSITION 1 The following conditions are equivalent: . {x, y) is a noncooperative equilibrium. (10) j ii. W{x,y)eMxN, [iii. v^=v^ andxeM^ andy e are conservative strategies. A Proof. The equivalence between properties (10, i) and (10, ii) is obvious, as well as implication (10, ii)=»(10, iii). The converse is easy: Let v denote the common value v^=v\ x belong to M and y e h& conservative strategies. Then v=f\y)^fix, y)<f *{x)=v and inequality (10, ii) ensues. ■ When v^=v^, this common value is called the value of the game, and a non- cooperative equilibrium is called a saddle point. Indeed, the graph of such a function / then looks like a saddle.
CH. 6, SEC. 2 TWO-PERSON ZERO-SUM GAMES: THE MINIMAX THEOREM 315 There are examples where saddle points do exist. Example Consider the finite game M={1, 2}, 2, 3}, where/ is described by the following matrix. 2 1 3 M 1 -2 -1 -4 2 1 0 -6 We observe that t; = — 1 and the pair of conservative strategies (1, 2) is a non- cooperative equilibrium. ■ To find sufficient conditions for the equality we shall introduce another value, (t;-natural) falling within the duality gap and prove successively that =v^ and = v^. Let if denote the set offinite subsets K of N. We set (11) (12) (read V natural) VK^:=ini sup fix, y) xeM yeK i;^:= sup i;jf= sup inf sup f(x,y) Ke Se Kg y xe M ye K
316 CH. 6, SEC. 2 SOLVING INCLUSIONS Since each point 6 Mean be identified with the subset {we observe that vfy)=/‘’(j) and thus Also since, u'’ = sup sup vt = V^ yeN Key sup f(x, sup fix, y), ye K yeN we deduce that Vk^v^ and thus Hence (13) We now use convexity assumptions to prove that THEOREM 2 Let us assume that M and N are convex subsets of vector spaces and that I *• Vj; € Ny x^f{Xy y) is convex. [ii. Vx € M, y^f {Xy y) is concave. Thenv^ = v^. A Proof. We set Z":={Ae/?+, ^"^^¿¿ = 1}. We associate with any K\= [yu ..., the map Fk from M to PT defined by (15) and we set (16) i’K(^):=(/U, yi),fix, y„)) wjc:=sup inf <A,Fk(x)> A e Zm X € M We shall prove successively that (17) i. Fk(M)-{-R\ is a convex subset (lemma 3). ii. VK 6 v^^wk (lemma 4) Hi. VK e 5^, Wi^^v^ (lemma 5) Consequently, inequalities (18) i;^:—sup i;^^sup wk^v^^v^ Key hold true and prove our theorem.
CH. 6, SEC. 2 TWO-PERSON ZERO-SUM GAMES! THE MINIMAX THEOREM 317 LEMMA 3 If M is convex and in the functions x-^f{x, y) are convex^ then the subset Fk{M) H- R\ is convex. ^ We set Z” := {A 6 R\ ^ = 1}, which is convex and compact, and (19) Wx!=sup inf {KFk{x)) AeZ« x€ M LEMMA 4 If M is convex and thefunctions x-^f (x, y) are convex, then for allfinite subsets K, we have (20) V^^Wk Proof Let6>0, H:=(l,1). We claim that (21) {wK + ¿)'^eFK{M) + R\ Assume the contrary, we can use the separation theorem in finite dimensional spaces, because Fk(M)-\-R\ is convex by lemma 3. There exists X^R!*, such that ¿ ^i{^K + fi) = (^9 (wk -f e) 1) ^ inf Fk{x)) + inf (X, u) veFK(N) + R% M€K?|. i=l Therefore, infueniiK u) is bounded below; consequently, XeR\ and infM€j?i^ {K m)=0. Since X^Q, then ^¿>0. We set X=Xj^.^^ A/ eZ” and obtain w/c + e^ inf (X, Fk{x))^ sup inf {X,Fk(x))=Wk xeM AeI”xeM which is impossible. Hence, there exist x^eM and u^^R\ such that (wk + s)H =Fx(x£)+«e. By the very definition (15) of F^, we deduce that Therefore, (22) Vi= 1, f{Xt, Wk + £ max f{x„y¡HwK+e Í = 1 n By letting 6 converge to zero, we deduce our lemma.
318 CH. 6, SEC. 2 SOLVING INCLUSIONS LEMMA 5 Let N be convex and the functions y^f(x, y) concave. Then for allfinite subset K, we have Wk^v^. ^ Proof We associate with any A e L” the point yT-=YH= i belongs to N (by convexity). The concavity of the functions y-^f{x, y) implies that 'ix 6 M, Ya= 1 ytHfix, yx)- Consequently, inf X J.X inf /(x,yA)=^sup inf f{x,y)-.=v'’ xe M i = I xeM ye N xe M Hence, by taking the supremum on Z”, we obtain w^^v^. THEOREM 6 Let us assume that (23) i. 3 Jo e N such that x^f(x, jo) is inf compact. ii. Vj G N, x-*f(x, y) is lower semicontinuous. Then v^ =v^, and there exists x eK such that (24) sup/(x,y)=t)^ A ye N Proof. We introduce the subsets Sy defined by (25) Sy-.={xeM\f{x,yHv‘'} They have thefinite intersection property: For each finite subset K = {yo,...,y„} <= Af containing yo, we have f]yeicSyf0. Indeed, the function f* defined by fK.ix)=ra&XyeKf{x, y) is lower semicontinuous [assumption (23, ¡¡)] and inf-compact [assumption (23, i)] and yo e K. So, it achieves its minimum at a point that belongs to Furthermore, the subsets Sy are closed [assumption (23, il)] and Sy^ is compact [assumption (23, i)]. Hence, the intersection nonempty. Any element X e riysN satisfies (26) sup f(x, y)<:v‘’ yeN Therefore, Since the other inequality holds, the theorem ensues. tWe recall that a function/ is inf-compact if for all XeRy the lower level sets {.jc|/(x)^A} are relatively compact.
CH. 6, SEC. 2 TWO-PERSON ZERO-SUM GAMES; THE MINIMAX THEOREM 319 Putting theorems 2 and 6 together, we obtain the lopsided minimax theorem. THEOREM 7 (LOPSIDED MINIMAX) Let M and N be convex subsets of vector spaces, M being supplied with a topology. We assume that (27) i. fy e N, y) is convex and lower semicontinuous. ii. ^yo e N such that x-^f{x, jo) is inf compact. and (28) Vx 6 M, y-^f{x, y) is concave Then f has a value (v^=v*), and there exists xeM such that supye n fix, y)=v . We deduce von Neumann’s minimax theorem. THEOREM 8 (MINIMAX) Let M and N be convex subsets of vector spaces, supplied with topologies. We assume that fy 6 N, x-*fix, y) is convex and lower semicontinuous. (29) and {li. Then there exists a saddle point {x,y)eM^N. ii. 3jo e N such that x-*f{x, )^o) is inf compact i. Vx e M, y-*f{x, y) is concave and upper semicontinuous. Sxo^M such that y-^f[x,yo) is sup compact. We state a corollary to theorem 6 that uses the conjugate functions /* from X* to ]-oo, -boo] defined by (31) We set (32) /*(p):=sup i(p,x)-fix,y)'] xeM Dornff :={p eX*\ftip)< +<»}.
320 CH. 6, SEC. 2 SOLVING INCLUSIONS COROLLARY 9 We assume that M is a weakly closed subset of a reflexive Banach space X. We posit the following assumptions: (33) {'• [ii. e N, x-*f{x, y) is lower semicontinuous. 3yo e N such that 0 € Int(Dom/J,) {for the strong topology of the dual 2i*) Then =v*, and there exists x eK such that supy e n /(3c, y)=v*. A Proof It suffices to prove that assumption (33) implies that the function x-^f{x, yo) is inf compact. There exists rj>0 such that fy^cDom /*. Let Sx:={x e M\f{Xy be a lower level set of x->/(x, j^o)- We prove that it is bounded. Indeed, for all /? e X *, we deduce from definition (31) offy{p) that »/<yo) + fUnpl\\p\\) and thus sup {p, a+fUw/\\p\\))< +00 The uniform boundedness theorem implies that 5a is bounded and thus relatively compact in X. Since/ is lower semicontinuous, it is actually weakly compact in X and therefore in M. We then apply theorem 6. ■ Remark Let X and X* be two paired vector spaces. Corollary 9 remains true when X is supplied with the weak topology o(X, X *) and X * with the Mackey topology t(X*, X), since the neighborhoods for the Mackey topology are the polar sub¬ sets of weakly compact subsets of X. ■ The compactness assumption we made in Corollary 9 happens to be too strong for many problems. We shall relax it when M is a subset of a Banach space. We consider two Banach spaces X and Y and a function / from X xY toR:=[-oo, -f-oo]. Weset (34) and M:={x 6 6 Y,f{x, y)< -l-oo} (35) ^:={y€y|Vx6X,/(x,y)>-oo}
CH. 6, SEC. 2 TWO-PERSON ZERO-SUM GAMES: THE MINIMAX THEOREM 321 We shall say that MxN is the domain of f. We assume that M and N are non¬ empty and set (36) / is the restriction of/ io MxN Thus, / maps MxN io R. The following theorem and its proof will be very useful in many applications. For that purpose, we replace the covering ^ of TV by finite subsets by a covering sd of TV, which we can choose as a parameter, satisfying (37) when K and L belong to sd, then KuL belong to sd. We associate with such a convering the number (38) u^(j!^):= sup inf sup/(x,>') Keii xeM yeK We point out that when jd c (39) and that where ^ denotes the covering of N by all the subsets K of N. THEOREM 10 Let X be a reflexive Banach space and f be a function from X xY to R whose domain M xN is nonempty. (40) Let si be a countable covering of N. We assume that 0 eint ( U Dorn /*) \ye AT / (41) and that (42) Then there exists xeM such that (for the strong topology) (43) Vj; e N, x-^f(x, y) is lower semicontinuous on X for the weak topology. sup f(x,y)=v‘^(sd)=v* yeN
322 CH. 6, SEC. 2 SOLVING INCLUSIONS Among all the corollaries we can form from these results, we state only the following very useful theorem. THEOREM 11 (RELAXED LOPSIDED MINIMAX THEOREM) Let X and Y be reflexive Banach spaces and f be a function from X xY to R whose domain M xN is nonempty. We assume that (44) that (45) that (46) Wy e N, x-*f{Xy y) is convex, lower semicontinuous Vx e Af, y-^f{x, y) is concave, upper semicontinuous 0 e Int ( (J Dom ff I {for the strong topology) \yeN / Then f has a value v:=v^ = v^, and there exists x eM such that supy e n f(x,y)=v. Proof We set f\y):= inf^ e y) and observe that 6 y|/'’(;;)>-00} is the domain of/^ We introduce the following countable covering si oi N by the subsets (47) K,:={yeN\\\y\\<p and f‘’(y)^-p} Since /'’ is concave and upper semicontinuous by (45), the subsets Kp are closed, convex, and bounded. Since Y is reflexive, they are weakly compact. We introduce (48) u^(j)/):=sup inf sup fix, y) p>l xe M yG Kp Theorem 10 implies that there exists x e M such that sup fix,y)=v‘’isl)=v^ yeN Now we can apply the lopsided minimax theorem 7 to —f, which yields inf sup/(x, j)= sup inf f(x,y) xeMyeKp yeKpxeM Indeed, Kp is convex compact, M is convex, y-*f(x, y) is concave and upper
CH. 6, SEC. 2 TWO-PERSON ZERO-SUM GAMES: THE MINIMAX THEOREM 323 semicontinuous for all xbM, and x-*f{x, y) is convex for all y e Kp. Therefore, t)^(j^)=sup sup inf /(x, ;^)=sup inf f{x,y)=v'’ p>l yeKpxeM ye N xe M (or Kp. m It remains to prove theorem 10. For that purpose, we need lemma 12. LEMMA 12 Assume that is finite. Then there exists a generalized sequence of elements Xpof M satisfying (49) fKesd, 3/iK such that lim sup I sup/(xp, \yeK ) When si is a countable covering, this sequence is a usual one (i.e., countable). A Proof. By the very definition of v\si), inequalities (50) inf sup /(x, y)<U^(ji/) xeM yeK hold true for all subsets K e si. Therefore, we can associate with any « e an element xk,„ e M such that (51) sup f(XK,m y)^v‘'isi) + 1 yeK we set Ji\=s4 x N, preordered by the relation (52) (Ki, «i)<(X2, «2) if and only if Ki<=K2 and /Ji<«2 which is filtered (or directed): Any pair ((/Ci, «1), (Kz, «2)) has an upper bound (Ki u K2, max («1, «2))- It is countable whenever si is countable. Therefore, the map {K, n)e ^^xk.„ 6 Mis a generalized sequence. By taking L^K and m^n, we obtain sup f{XL.m, I'Xsup f{XL,m, y)^v''{si) + ^^v‘'(si) + 2 yeK ye L ^ “ Hence, if we fix Xo e j;/ and K = Kq, we obtain m sup ( sup/(xL,m,.y))<u^(^)+- (L,m)>(K,n) \yeKo / ^
324 CH. 6, SEC. 2 SOLVING INCLUSIONS Consequently, limsup [sup/(xjc,„,y)] (K,n)>(Ko,l) yeKo = inf sup [sup■ (K.n) WKo. 1) »(K,n) y 6 Ko Proof of Theorem 10. LetiV be the union of an increasing sequence of sub¬ sets K„ and si the family of the We set (53) t>^(j!/):=sup inf sup f{x,y) n>0 x€ Af ye K„ By lemma 1, there exists a countable sequence of elements x„eM satisfying (54) fyeN, ^n(y) suchthat lim sup/(x„,;'Xu^(j!/) By assumption (41), there exists f/>0 such that (55) U Dorn /* yeN So for any peX*, there exists yp e N such that fyp(yip/\\p\\)< +oo. By (54), we deduce that there exists np>n{yp) such that (56) Therefore, /(x„,y)<i;^(y)-l-l fn^rip, { x„)^ f{x„, yp) +fy \\p\ (heI ’VIHl = ‘^V)+1+/? <-foo f—V Since the sequence is countable, we deduce that sup (p, X„)< +00 Therefore, the uniform boundness theorem implies that the sequence x„ is bounded and thus weakly relatively compact. A subsequence x„> converges weakly to some xeX, and the lower semicontinuity of / with respect to x implies that
(57) CH. 6, SEC. 3 THE KY FAN INEQUALITY 325 V;; e N, f{x, 3;)^ lim inf 3;)^ ) Then X belongs to M, and the theorem ensues. 3. THE KY FAN INEQUALITY Let M and N be the strategy sets of Mike and Nancy and/ be Nancy’s gain function. She can use it to assign to each decision rule C^: M->7Va gain defined by (1) f\Crd'= inf fix, C^ix)) xeM This represents the worst gain she can expect using decision rule Cn and assum¬ ing that Mike’s behavior is noncooperative. Note that this definition is con¬ sistent with the definition of the worst gain yielded by a strategy j;, regarded as a constant decision rule x-^y f\y):= inf f(x,y)= inf fix,y{x)) xe M xe M Consequently, if is a set of decision rules containing N, (2) v^:=sup inf /(x:,3^)< sup /\Cn)^ mi f(x, y):=v^ yeNxeM xeMyeN PROPOSITION 1 If Nancy is allowed to use all possible decision rules, then (3) sup f\Ctd=v* CnsnM Proof. Indeed, we can associate with every £>0 and every x&M& strategy C|i(x) 6 N such that u’*‘^sup f(x,y)<^f(x, Cn(x)) + £ ye N From this, we deduce that inf sup/(x,;^)</'’(aW)+e< sup /'’(Q)+e xeM ye N Cj^eN^ Since this inequality holds for every £>0, we obtain the inequality supcpieN>^ f\CN), which, together with inequaity (2), proves the proposition. ■
326 CH. 6, SEC. 3 SOLVING INCLUSIONS We show that with additional hypotheses, equality (3) remains true when we require Nancy to use only continuous decision rules (in order to describe a stable behavior for her). THEOREM 2 Let M be a topological space, N a convex subset of a topological vector space, and f a function from M xN to R. Let us suppose that (4) and that (5) i. ^yo e N such that x-^f(x, is inf compact ii. \fy e N, x-^f{x, y) is lower semicontinuous Vx € M, y-^f{x, y) is concave If ^(M, N) denotes the set of continuous mapsfrom M to N, thefollowing equality holds true: (6) sup inf f{x,Cfi{x)) = V^ xeM Proof We already know from (2) that (7). sup inf /{x, Ca^(a:)) ^ ^ xeM Thus, we must establish the opposite inequality. First of all, we can associate with every 6 >p a map C% from M to N, not necessarily continuous, that satisfies inequality (3). Moreover, since the functions x^f(x, y) are lower semicontin¬ uous, there exist open neighborhoods such that (8) yx'eJiix), f(x,C%ix)Hfix',C%{x))+e We introduce the subset Mo'.= {xeM\ fix, yo) which is compact, since x^f(x, ;^o) lower semicontinuous and inf compact. Therefore, it can be covered by n neighborhoods := ^(xi). Let Jio:=M \Mq, which is an open set. Then Let {pi},=o „ be a continuous partition of unit subordinate to this finite convering. We introduce the function Cn from Mio N defined by (9) CJ,(x):=poW>'o+ E Pi(x)C%ixi)
CH. 6, SEC. 3 THE KY FAN INEQUALITY 327 This map is continuous. Furthermore, since the function y) is concave andp,(A:)^0 for i=0,...,« and X”=o 1» we obtain (10) /(x, CMx))^po(^)/(:v, >^o)+ Ё Pi(x)f{x, C^(Xf)) i=l Now, if po(^)>0, then xejf^ and consequently, /(x, yo)'^v*>v*-г, If /?,(x)>0, then X 6 and consequently /(x, C^(x,))^y^(xf, C^(x,)) —e [inequalities (7) and (8)]. Since Z”=o /^«W = we deduce from (10) that Hence, /(x, C%{x))> Y, Pi(x)(v^-&) = v*-e i = 0 v^-e^f^iChix))^ sup inf /(x, C^(x)) Cj^e<if{M,N) xeM Letting 6 converge to zero, we conclude the proof. ■ We shall prove an equivalent formulation of the Brouwer fixed point theorem, known as the Ky Fan inequality. Since we have proved the Brouwer fixed point theorem for compact convex subsets of RP, we begin by proving the Ky Fan inequality for compact convex subsets of R!\ We shall extend it to any compact convex subset (see theorem 5). LEMMA 3 Let K be a compact convex subset of and let (f> be a real-valuedfunction defined on KxK satisfying (11) i. 6 K, x-m/>(x, y) is lower semicontinuous ii. Vx e K, у-^ф{х, у) is concave Then there exists x e К satisfying (12) sup ф{х, ;;)^ sup ф{у, у) к ye к ye К Proof. Since the functions х-*ф(х, д') are lower semicontinuous and the functionsу-*ф{х, у) are concave, theorem 2 and theorem 2.6 imply the existence of 3c 6 К satisfying u'''=sup j)= sup inf C(x)) ye к Ce'«(K,K) хек
328 сн. 6, SEC. 3 SOLVING INCLUSIONS Since к is compact and convex and C is a continuous map, there exists a fixed point Xc e К of C. Hence, inf </>(x, С(х))^ф{Хсу С(Хс))=ф{Хсу Xc)^sup ф{у, у), хек уек Lemma 3 is proved. ■ Remark Conversely, assume that Ky Fan’s inequality holds true. We associate with any C e K) the function ф defined on X x by (13) Ф{х, у) = <C(x) - X, у-x) This function satisfies the assumptions of Ky Fan’s inequality. Hence, there exists xeK such that supyeK(C(3c) —x, x)^0. By taking j; = C(x) 6X, we get ||C(x)-xlP^0, that is, x is a fixed point of C. This proves that the Ky Fan inequality in finite dimensional space is equiva¬ lent to the Brouwer fixed point theorem. ■ We proceed to prove the other inequality mentioned earlier. Let M and N be the strategy sets of Mike and Nancy and let f be Mike’s loss function. Mike assigns to each decision rule Cm* N-^M the worst loss (14) / *(Cm)=sup fiCuiy), y) yeN If ^M is a set of decision rules containing the set M of constant decision rules, we have (15) inf THEOREM 4 Let M be a topological space, N a convex subset of a topological vector space, and f a function from M x Nto R. Let us suppose that (16) and that (17) f i. 3;;o e N such that x-^/(x, ^o) is inf compact |ii. Wy 6 N, x-^/(x, y) is lower semicontinuous Vx 6 M, y-^f (x, y) is concave If ^{N, M) denotes the set of continuous maps from N to M, there exists xeM such that (18) sup /(x, y) = inf sup /(См(Л у) yeN Cf^e<g(N,M) y^ M
CH. 6, SEC. 3 THE KY FAN INEQUALITY 329 Proof. Since inequality infcj^e«(N,M)/’*‘(CM)<t^’*‘ holds true, we have to prove the existence oix&M such that for all continuous maps Cm from N to M, we have the following inequality: (19) sup f(x, yH sup fiCuiy), y) ye N ye N By theorem 2.6., we know that there exists xeM such that (20) sup/(x, j)=u^:= sup inf sup/(x, j) yeN Ke xe M ye K where ^ is the family of finite subsets of N, Since N is convex, co(K)c and Cm(co{K))<=M. We set K\={yu • • • , y,,} and 2'':={A A, = l}. We can write n inf max f(x,yi)= inf sup Z xeM xeMAeL"i=l n < inf sup Z fix, yd xeCMicoiK)) A el" 1=1 = inf sup Z ^if I CmI Z Pjyj > Pi /i6l"AeZ"i=l \ \j=l We set (pin, >^):= E V It maps S” X Z" to Ry is lower semicontinuous with respect to n (because Cm is continuous from N to M), and is affine with respect to X. Hence, Ky Fan’s inequality in finite dimensional spaces (lemma 3) implies that inf sup (f>{^y sup 0(A, X) /ieZ"AeZ" A el" Since the functions y^f(x, y) are concave, we obtain (piK == S ^if (cm( E hyj )> y> Thus, we have proved that for all K 6 and for all Cm e ^{N, M), (21) inf sup/(x, jX/'^(Cm) xeM yeK
330 CH. 6, SEC. 3 SOLVING INCLUSIONS Hence, (22) inf and our theorem ensues. As a consequence, we obtain the Ky Fan inequality for compact convex subsets of any topological vector space. THEOREM 5 (KY FAN) Let K be compact convex in a topological vector space and <j>\ KxK-^R be a function satisfying (23) i. Vj; 6 K, y) is lower semicontinuous. ii. VxeX, y^(j){x,y) is concave. Then there exists x eK satisfying (24) sup <l>{x,yXs\xp (¡>{yyy) yeK yeK Proof We take M=K, N=K, and f{Xy y)=(l>{Xy y). Since the identity is a continuous map from K to K, we deduce that there exists xeK such that ^ sup 0(x, y) = inf sup (f>{Ciy\ y) ^ sup (l)(yy y) ■ yeK Ce<&(K,K) yeK yeK Let us consider theorems 2 and 4, stating that (25) sup inf f{XyCN{x))= inf supfiC^iyly) Cffe^iM^N) xeM Cme ^(N.Af) yeN In both theorems, the topology on N can be chosen as a parameter ; also, the stronger the topology on N, the larger the set ^{N, M) the smaller the set ^(M, N)y and consequently, the stronger equality (25). We can design a topology on Ny which is not (necessarily) a vector space topology, that is stronger than any vector space topology and for which theorems 2 and 4 hold true. Let be a convex subset of a vector space Y. We associate with any finite set K = {yu ... yy»} of N the affine map Pk from 2” to N defined by (26) VAeS”, ^k(A)= E i= 1
CH. 6, SEC. 3 THE KY FAN INEQUALITY 331 DEFINITION 6 The finite topology on a convex subset N is the strongest topology for which the maps Pk cire continuous when K ranges over the family if offinite subsets of N. So, a map C from N, supplied with the finite topology, to a topological space M is continuous if and only if (27) ViC 6 the maps CPk from S" to M are continuous Also, any map C from a topological space M to TV of the form (28) C{x) :=P,$(x)= £ Pi(x)yi 1=1 where ^ is a continuous map from M to Z”, is continuous from M to TV supplied with the finite topology. PROPOSITION 7 The finite topology on a convex subset N of Y is stronger than any vector space topology. Any affine map C from N to a vector space X is continuous when both TV and X are supplied with the finite topology. к Proof, a. Suppose that Y is supplied with a vector space topology. Let us take C:=7 to be the canonical injection from TV to К Since for all Keif, the map IpK- ^ obviously continuous, we deduce that I is continuous, that is, the finite topology is stronger than the restriction of the vector space topology to TV. b. Let C be any affine map from TV to X. We have to prove that for any the map СРк: A 6S"^C/Sk(A)=C t t Шуд=Рс(ю^ (A) is continuous. But CPk=Pc{K) is indeed continuous by the very definition of the finite topology on X. ■ THEOREM 8 Let M be a topological space, TV a convex subset supplied with the finite topology, andf:Mx N->R a function satisfying (29) f i. Зуо e TV such that x-^f{x, уо) is inf compact. |U. У у € TV, x^f{x, y) is lower semicontinuous.
332 CH. 6, SEC. 3 SOLVING INCLUSIONS and (30) Vx e M, y-^f {x, y) is concave. Then there exists xeM such that (31) sup/(x,;;)= sup inf /(x, C^(x)) yeN xgM = inf sup /(Cm(j^), y) yeN Proof. Indeed, in the proof of theorem 2, the map C% defined by (9) can be written PkP> where K:={yo, C^Xi),..., Cn{xj))SLndp{x)=ipo{x\pi{x\.., ,p„{x)l which is continuous from M to JV supplied with the finite topology. In the proof of theorem 4, we needed the continuity of Cm only to infer that for all X e Z", X) is lower semicontinuous. But we can write 0(w, X) = ^”=1 ^if{CMPK(p\ yi)’'> therefore, ju-^0(ju, X) is lower semicontinuous whenever Cm is continuous for the finite topology on N. ■ We now prove an extension of the Ky Fan inequality in which the assumption of lower semicontinuity with respect to x is relaxed. THEOREM 9 (KY FAN’S INEQUALITY FOR MONOTONE FUNCTIONS) Let KciX be a convex subset of a topological vector space and (p: KxK-^R be a function satisfying (32) I ** ^ x-^<l>(x, y) is lower semicontinuous for the finite topology. [ii. Vx e K, y-^<l>(x, y) is concave and upper semicontinuous. We also assume that (33) 3yQ 6 K such that x-^cpix, yo) is inf compact and that (j) is monotone in the sense that (34) i. ^yeK, (l>(y,yH0 ii. "^x, ye K, (j){x, y) + (piy^ x)^Q Then there exists xeK such that (35) Proof. Since we have the inequality (36) sup (l>{Xyy)^0 yeK v^ ^ sup inf max </>(x, y) {yi,...*yn)^ ^ X^co ye co{yi y«}
CH. 6, SEC. 3 THE KY FAN INEQUALITY 333 and since the Ky Fan inequality in finite dimensional spaces implies that (37) inf sup <l>(x, sup (¡>{y, 0 x€co{yi y«} we deduce that (38) Lemma 2.12 implies the existence of a generalized sequence of elements x^eK satisfying (39) G K, such that lim sup By the compactness assumption (33), we infer that the subsequence remains in a compact subset and consequently, a subsequence (again denoted by) con¬ verges to some Jc 6 K. We shall prove that x solves Inequality (35). If not. (40) there existsy eK such that 0<<l>{x,y) Since the function t-^(f>{x+t{y-x), y) is lower semicontinuous by assumption (31, i), we deduce from (39) that there exists i 6 ]0,1[ such that (41) 0<4>{x + t(y-x),y) We prove now that (42) 0^<t>(x + t(y - x\ x) Indeed, the monotonicity (33, ii) of </> implies that by setting z=x-l-r(y-.x), O^lim sup {<i>{xn, z)+0(z, Xf,)) <lim sup (f>(x^, z)-l-lim sup </>(z, X;,) but lim sup;,>^(,,)<^(x^, z)<i;^<0 by (39) and lim sup^>^,^) <^(z, x^)^<f>{z, x) because y->(/>(x, y) is upper semicontinuous [(32, ii)], hence, (42) holds true. The concavity of the function z-^(/>(x-i-t(y — x), z) and inequalities (41) and (42) imply that (43) 0 < 0(3c + liy - x), 3c + t{y - x)) which contradicts assumption (34, i). ■
334 CH. 6, SEC. 3 SOLVING INCLUSIONS As a first application of Ky Fan’s inequality, we prove the existence of non- cooperative equilibria in «-person games. Let us consider n players. When i is a player, we denote by i = {y e {1,.. .,n}\j^i} the set of the« -1 other players. Each player must choose a strategy Xi in a strategy set M,- according to a given mechanism. An example of such a mechanism occurs when each player chooses a strategy through a deci¬ sion rule Ci that constrains his or her choice when the « — 1 other players j € i have implemented their strategy xje ... jfiMj. For simplicity, we set (44) Vi, xt={xj}j,t, M:=Y\Mi jfi i=l A decision rule for the ith player is a set-valued map C, from M¡ to M,. DEFINITION 10 A multistrategy x={xu . .., x„) e M is consistent with respect to decision rules Ci if and only if (45) If we set Vi = 1,..., «, Xf e Ci{Xi) V3c€M, C{x)=t\ CiiXi) i= 1 the subset of consistent multistrategies is the subset of fixed points x of the set¬ valued map C X e C(3c) This problem can be solved by various available fixed point theorems. We con¬ sider here the particular case where decision rules are canonical decision rules constructed from loss functions as follows. We describe the game by (46) « loss functions/ from Mio R that associate to each player i and to every multistrategy x e M a loss/(x) e R. From the point of view of player i, we can write a multistrategy x in the form {xi, xt) e Mi X Mf, since each player can choose Xi e Mi but has no control over the choice of xf. So, we set (47) fi{x):= fiixi, Xi) We associate with xt e M\ the subset (48) C,(x?):={jc, 6M,|/(jc„ xf)= inf fiy^xf)] yi e Mi
СН. 6, SEC. 3 THE KY FAN INEQUALITY 335 of the strategies Xi that minimizes this loss function x¡) when the other players implement xt. DEFINITION 11 A n-person game in normal {or strategic) form is defined by n loss functions f defined on M. The maps Ci from М/ to Mi defined by (7) are called canonical decision rules. The multistrategies x that are consistent with these canonical decision rules are called noncooperative equilibria {or Nash equilibria) of the game. We set (49) Vx, yeM, ф{х, y):= Z х^-Дуь xf)) i=l PROPOSITION 12 The following statements are equivalent: (50) i. X eM is a noncooperative equilibrium. ii. Vi = 1, 2,, И, fi(xi, Xt)= min Дуи д:?) yieMi Hi. X eM satisfies sup 0(x, j?)=0. A yeM Proof, a. Conditions (50, I and ii) are obviously equivalent by the very definition of canonical decision rules. —+ b. We prove that (50, i) implies (50, iii). Let jt:=(ji,..., € M be given. Since for all i = 1,..., n, fi{xi, X?)-/((>'., .vi)<0, by adding these inequalities we obtain (j){%y)^0. c. We prove that (50, iii) implies (50, ii). Let i be fixed. We choose y such that yj=Xj for all j e i and y,- e A/,. Therefore, /i(Xi, x?)-/i(y.-, xr)=0(x, y)<0 ■ We then prove the fundamental theorem of noncooperative game theory. THEOREM 13 (NASH) Suppose that for all i = i,n. (51) i. the strategy set Mj is convex and compact. ii. the loss function fi is continuous, andfor every xi e Mu X?) is convex. Then there exists a noncooperative equilibrium. A Proof. The set M:=Y\l=i is convex and compact and the function $ defined on M x M by (49) satisfies (52) J i. Vy € M, х-*ф{х, у) is lower semicontinuous. |ii. WxeM, у-*ф{х,у)18сопсл\е.
336 CH. 6, SEC. 4 SOLVING INCLUSIONS Hence, Ky Fan’s inequality implies the existence of 3c e M satisfying (53) sup sup J;)=0. yeM yeM By proposition 12,3c is a noncooperative equilibrium. ■ 4. EXISTENCE OF ZEROS OF SET-VALUED MAPS In a straightforward way, Ky Fan’s inequality gives an array of useful sufficient conditions implying the existence of solutions to inclusions of the form xeK and 0 e F{x) + P [or F{x)n where P is a closed convex cone. The case where P={0} yields results on the existence of zeros of F. Many problems arising in economics and game theory fall in this class. Namely, let 7 and Y* be two paired vector spaces supplied with their weak topologies and F<= Y* a closed convex cone and F“ c Y be its nega¬ tive polar cone. THEOREM 1 Let К be a compact metric space and F : K->Y* a strict set-valued map. We assume that (1) (2) i. F is P-upper hemicontinuous in the sense that eP x-^(t{F{x\ y) is upper semicontinuous. ii. Vx e K, F (л:) 4- P is closed and convex. there exists a finitely continuous map C from P~ to К such that ^y eP~, (T(F{C{y)\ y)>0. Then there exists x e K such that 0 e F{x)-\-P. A Proof. We introduce the function (f) defined on KxP~ by (3) (t>(x,y):=-(T{F{xly) This function is lower semicontinuous with respect to x thanks to the assump¬ tion (1, i), and concave with respect to j;. Hence, theorem 3.8 implies the existence of 3c 6 K such that, C being finitely continuous. (4) sup ф{х,у)^ sup ф(С{уХу) yeP~ yeP~
СН. 6, SEC. 4 EXISTENCE OF ZEROS OF SET-VALUED MAPS 337 By assumption (2), ф(С{у)у for all;; e P". Hence, we have proved that (5) 4yeP~, 0^a{F{x\y) Since a{P~,y)=0 when>» € P~ and (т{Р~,у)= +oo when;; $ P", we deduce that (6) V;; 6 0^а(Р{х) + Руу) The subset F(x) + P being closed and convex by assumption (1, ii), it follows that 0 e P(x) + P. В The problem, then, amounts to finding finitely continuous maps from p~ io К satisfying assumption (2). When they do exist, we can think of taking a retrac¬ tion to Ky a subset of У: A retraction from У to a subset X c у is a map r from У to jFC such that r{y) =y for all у eK. For instance, if К is a closed convex subset of a Hilbert space, the projection r :=7i¡^ is a continuous retraction. So by taking P={0} and C a retraction onto Ky we obtain the following corollary. COROLLARY 2 Let К be a compact subset of Y that has a finitely continuous retraction r. Let F be a strict upper hemicontinuous map from К to У* with closed convex images satisfying (7) Vj^ey, of{F{r(y)),y)>0 Then there exists x e К such that 0 e F{x). A For instance, when К is the ball of radius a, we obtain the following conse¬ quence, the set-valued extension of theorem 3 of Chapter 2, Section 2. COROLLARY 3 Let Y be a finite dimensional space and F an upper semicontinuous map from the ball of radius a to the closed convex subsets of Y satisfying (8) Vxey ||x||=a, o{F(x),x)>0 Then there exists a solution x to (9) OeFiy) к Proof. Indeed, take r(y)=y when 1|;^|| and r{y)=ayl\\y\\ when ||;^|| ^a, so that (8) implies (7). ■
338 CH. 6, SEC. 4 SOLVING INCLUSIONS The following theorem has many applications to economics, as we shall see in Section 5.5. We consider the simplex C":=|x€/i"+ Z ^¡ = 1 i=l THEOREM 4 Let F be a strict upper hemicontinuous map from Z” to with compact convex values. If (10) Vx€l”, a(F{x\x)^0 then there exists 3c 6 Z” such that 0 € F(3c)—/?+. A Proof We take Y = Y^ = R\ P=-R\, P'=R\, X = Z", and C to be the map defined on К by C(y)\=ylY!^= i Assumptions of theorem 1 are obviously satisfied, and the existence of a solution x € Z” to 0 e F(x)—R\ ensues. ■ Remark More generally, we can assume that (11) and take (12) /¡C is a weakly compact subset of У* such that 0^ К P=K because we can prove that P~ = is spanned by K. It can also be proved that P" is spanned by a weakly compact subset that does not contain the origin if and only if the interior of P (for the Mackey topology) is nonempty. ■ The following generalization, although somewhat technical, may be very useful. THEOREM 5 We posit assumptions (i) and (ii) of theorem 1. We assume that (13) Ve> 0, there exists a finitely continuous map from P to К such that V;; € P“, (y(F{Ce(y)\ у)^ Then there exists xeK such that 0 e P(x)+P. A Proof The proof is the same as the proof of theorem 1, where we replace (4) by the following consequence of theorem 3.8: sup (l>{x,y)^ in{ sup (l>iCe{y\y) ye P~ £>0yeP~
CH. 6, SEC. 4 EXISTENCE OF ZEROS OF SET-VALUED MAPS 339 Assumption (10) implies that inf sup inf e=0. ■ £>0 Ye P~ Remark We can also relax the assumption that K is compact. For instance, it is suf¬ ficient to assume that there exists yo e P~ such that (14) Ko:={xeK\a(F(x),yo)>0} is compact Remark We can imagine applying an analogous proof for upper hemicontinuous maps F with closed convex values from a compact convex subset Kc=X to X* by applying theorem 3.8 to the function (f> defined by (15) (f>ix,p):=-ff{Fix),p) We then need the existence of a finitely continuous map C from X* to X such that (16) Vp€X*, <t(F(Qp),p))>0 One natural example of such a map C would be a selection of the subdiffer¬ ential d(Txi—p) of the support function of K (which is continuous—hence subdifferentiable—because K is compact). Since a finitely continuous selection does not necessarily exist, we have to devise another strategy for proving that assumption (16) for da^ implies the existence of zeros of F. We recall that x belongs to d(Tx{—p) if and only if —p belongs to the normal cone Nx(x) to K at X. Hence, condition (16) for can be written (17) Vx 6 X, Vp 6 Nk{x), <7(F(x), -p)>0 This condition—which will play an important role—is called the normal condition. By using the tangent cone 7i(x) to X at x, which is the negative polar cone of the normal cone, we obtain a dual version of (17), which we call the tangential condition. We investigate their properties in a more general frame¬ work. ■ Let X and Y be two Hausdorff, locally convex vector spaces, A belong to if(X, y), X<=X a closed convex subset, and F: K-*Ya strict set-valued map. DEFINITION 6 We shall say that F satisfies the tangential condition with respect to A if (18) Vx6X, F{x)ncl{ATx{x))f0
340 СН. 6, SEC. 4 SOLVING INCLUSIONS and the normal condition wiiA respect to A if (19) VxeX, \/peA* ^Nk{x\ then (j{F(x\—p)>0 A When X=Y and Л = 1, we omit mentioning with respect to A. LEMMA 7 The tangential condition (18) implies the normal condition (19). The converse is true when the images F{x) of F are convex and compact. A Proof a. Let x€ К and v e F{x) n c\(A Tk(x)) be chosen. Hence, v=lim„_ Au„ where u„ belongs to Tk(x). Let p such that >4*/? belongs to Njdx). Then, (t(F(x), -p)>(-p,v} = \im (-p,Au„)=\\m (~A*p,u„}^0 И-> 00 П-* ao because (A*p, u„)^0 for all u„ e Tk(x)=Nk{x)~. b. Conversely, assume that F(x) is convex and compact and that 0$F{x) — cI{ATk{x)), that is, F{x)ncl{ATK(x))^0. The Hahn-Banach separation theorem implies the existence of p eY* and e>0 such that <t{F{x\ —p)^ inf (—/7, Av) — s<0 ve Tk(x) Since Tk(x) is a cone, this implies that A*p e Tk{x)~ = Nk{x) and that a(F{x\ —p) ^ — e<0. So, the normal condition (19) is contradicted. ■ The calculus on tangent cones to convex subsets allows us to check the tangential condition in many instances of closed convex subsets. We point out the following important, although obvious, remark. PROPOSITION 8 If two set-valued maps Fi and F2 satisfy the (dual) tangential conditiony so does the set-valued map oi^Fi Ч-агТз for au OC2>0. A In particular, we use the fact that when F satisfies the tangential condition, so does the map x->^F(x) — Ax-hy when у is given in A(K). Remark This is a first property of stability of the tangential condition and, consequently, of its consequences, under perturbations of F. We mention another property of stability under small perturbations: PROPOSITION 9 Let К be a convex subset with nonempty interior and let F be a set-valued map from KtoX with compact graph satisfying the strong internal tangential condition (20) Vx e X, F(x)<=Int Tk(x)
СН. 6, SEC. 4 EXISTENCE OF ZEROS OF SET-VALUED MAPS 341 Then there exists a > 0 such that any set-valued map G close to F in the sense that (21) graph Gcgraph x B) also satisfies the strong internal tangential condition. A Proof. Condition (20) implies that the graph of F, which is compact, is contained in the graph of Int which is open. Then there exists a such that condition (21) holds true. ■ We now prove our basic result—the existence of a zero of F that belongs to К when the tangential condition holds true. THEOREM 10 Let X and Y be two Hausdorff locally convex vector spaces^ A 6 SP{X, У) a continuous linear operator, K<=-X compact convex, and F an upper hemicon- tinuous map from К to Y with nonempty closed convex images. We posit the following tangential condition: (22) Then, (23) ЧхеК, Т(х)пс\(АТк(х))ф0 j i. there exists a zero xeK of F:0 eF{x) |ii. Vj; € A{K), 3xeK such that Ax—y 6 F(x) Remark We can regard the second condition as a perturbation result (24) if there exists a solution xeKto the linear equation y=Ax, then there exists a solution xeKto the perturbed inclusion y eAx—F{x) which also amounts to saying that (25) {A-F)~^nK IS a strict set-valued map from A{K) to K. Ш When X = Y and A = l, we obtain the following consequence, stated in theorem 11. THEOREM 11 Let X be a Hausdorff, locally convex vector space, K<^X a compact convex subset, and F a strict upper hemicontinuous map from К to X with closed convex images satisfying the tangential condition (26) Vx e K, F(x)n Тк{х)ф0
342 CH. 6, SEC. 4 SOLVING INCLUSIONS Then, (27) i. there exists x eK, solution to 0 e F(x). ii. eK, 3xeK, solution to y ex —F(x) Remark Haddad’s viability theorem states that under the assumptions of theorem 11, for all xq e K, the differential inclusion (28) x'(t) e F (x(i)), .x(O)=xq has a solution x(.) that is viable in the sense that (29) Vi^O, x{t)eK Actually, the tangential condition (26) is also necessary: It holds true whenever problems (28-29) has a solution for all xqsK [see Aubin-Cellina [1984], Chapter 4]. ■ Proof of Theorem 10. a. The second conclusion follows from the first: Since the set-valued map G defined by G{x):=F{x)-\-y — Ax is the sum of the two maps F and y —A that satisy the tangential condition, then G satisfies it and, consequently, has a zero x that is a solution to the inclusion Ax-ye F{x). b. We denote by (t{F {x\ q):=supv e f(x) {q, v) the support function of the closed convex subset F(x). To prove the existence of a zero, we assume the contrary: Vx 6 0 i F(x) and derive a contradiction. Since the subsets F{x) are closed and convex, the separation theorem implies that Vx eK, 3p eY* such that <r(F(x), —p)<0 We set (30) A,:={xeK\a(F{x),-p)<0} So, the statement that no zero exists takes the form K<= U Ap peY* c. Since F is upper hemicontinuous, the subsets Ap are open. Hence, the com¬ pact subsets K can be covered by n open subsets Ap;. Let {a,} for i = 1,...,« be a continuous partition of unity associated to this covering. We introduce the function (¡) defined KxKhy (31) Hx, ;^) = - E aiix){A*pi, x-y) 1=1
CH. 6, SEC. 4 EXISTENCE OF ZEROS OF SET-VALUED MAPS 343 It is continuous with respect to x, affine with respect to;^, and satisfies (f>(y,y)=0 for all y eK. So the assumptions of the Ky Fan inequality (theorem 3.5) hold and, con¬ sequently, there exists x* e K such that (32) 'iyeK, Ф{х^,у)=(-А*р^,х^-у)^0 where we set p*=X"=i In other words, A*p^ belongs to Nic(x^,). The tangential condition (18) implies that the normal condition (19) holds true (by lemma 7). Therefore, (33) (j(F(x^,), -p^)^0 d. The latter inequality is impossible: Let / be the set of indices i such that «i(A:*)>0. It is nonempty since aj(x*)=l. If i e /, then x# e Ap^ and thus, ff(F(x*), —Pi) <0. Therefore ^(F(x,j(), P)(t)—^(F(x,j(), ^ ^i(x,i<)pi)^ ^ ^i(x,[{)^(F(x,j(), Pi) i€ I i e I (by the convexity of support functions). We have proved that —pj<0, which is the contradiction we were looking for. ■ Remark Theorem 10 remains true even if we assume only that F satisfies the normal condition (19). ■ Remark We can even let the continuous linear operator A depend on x as in theorem 12. THEOREM 12 We introduce (34) i. K, a compact convex subset of X ii. F, an upper hemicontinuous map from К to Y with closed convex values lii. A: Y) a continuous map associating with each xeK a continuous linear operator from X to Y We posit that the tangential condition, (35) Vx € K, F{x)nc\{A{x)TK{x))^0
344 СН. 6, SEC. 4 SOLVING INCLUSIONS holds. Then (36) i. There exists a zero xeK of F: 0 6 F{x). ii. Wy e K, there exists xe K satisfying A{x){x—y) e F{x). Proof a. The second statement follows from the first applied to the map G defined by G{x):=F{x)+A{x)(y — x). b. The proof of the first statement is the same as the proof of theorem 10, where the function (f) defined by (31) is replaced by the function 0 defined by (37) ф{х,у)=- x; ai{x)(pi, A(x)(x-y)) i=l We deduce now from theorem 11 the well known Kakutani fixed point theorem as well as some of its extensions. THEOREM 13 (KAKUTANI) Let К be a compact convex subset and G an upper semicontinuous map from К to К with compact convex images. Then there exists a fixed point x^eK of G. A Proof. We set F(x):=G{x)—x<^K — x<^Tk(x). Hence, F(.) is an upper hemicontinuous map from К to X that satisfies the tangential condition. By theorem 11, there is a zero x^ e K, of F, which is a fixed point of G. ■ The same proof implies the following statement. We say that G is inward if (38) Vx € X, G(x)n(x+ Тк{х))ф0 THEOREM 14 Let G be an upper hemicontinuous map from a compact convex subset K^X to the closed convex subsets of X. If G is inward, then it has a fixed point x^ e K. We also mention the following result. We say that G is outward if (39) Vx e X, G(x)n(x- Тк(х))ф0 THEOREM 15 Let G be an upper hemicontinuous map from a compact convex subset K<=^X to the closed convex subsets of X. If G Is outward, then (40) [ii. It has a fixed point x^ 6 X. KCG(K) for allyeK, 3xeK such that y 6 G(x)] A
СН. 6, SEC. 4 EXISTENCE OF ZEROS OF SET-VALUED MAPS 345 Proof. It follows from theorem 11 applied to the map F defined on К by F(x):=x — G(x). Then a zero 3c of F is a fixed point of G, and a solution 3c to у ex —F{x) is a. solution x to у e G{x). ■ We now mention a convenient sufficient condition implying that a set¬ valued map G is outward. We recall that (41) док{р):={х e K\(p, х}=<Тк{р)} is the support zone of К at p and that x 6 da^ip) if and only ifp e Nk{x). PROPOSITION 16 Let К be a closed convex subset of X and let G be a set-valued map from К to К satisfying (42) Vp 6 X*, Va: 6 д(Тк{р\ Gix)ndajc(p)f0 Then G is outward. In particular, this is the case of a map G satisfying (43) V/7 G X *, G(dGj,(p))<zdG^,{p) к Proof Let xedoj({p) and yoeG{x)ndoj({p) be fixed. Hence, v=yo-x belongs to G(x)-x and -Tk{x) because, УреМк(х), (p, уо)-(р, x) = We can relax the assumption that К is compact by replacing it with a coer¬ civeness assumption. THEOREM 17 Let К be a closed convex subset of a finite dimensional space and let F be a strict upper hemicontinuous map with compact convex images satisfying the coercive¬ ness assumption (44) lim (t{F{x\x)<0 llxll-oo xeK We posit the tangential condition, (45) VxgX, F{x)nTK{x)f0 Then there exists a zero xeKof F. If we posit the stronger coerciveness assump¬ tion.
346 СН. 6, SEC. 4 SOLVING INCLUSIONS (46) a{F{x), x) hm —П—П— = —00 llxll-»oo xeK then for all y ^ K, there exists a solution xb Kto the inclusion y ex- F{x\ A Proof Coerciveness assumption (44) implies that there exist e>0 and ^>0 such that sup (t(F{xX x)^—e<0 WxW^a xeK This implies that F{x)cTaB{xX By taking the number a large enough so that KnaB=^0, we know that TKnaBix)=TK{x)nTaB{x). So the tangential condi¬ tion (45) implies that (47) "^xeKnaB, F(x) n Tk пав(х) Ф0 It suffices to apply theorem 11. To prove the second part of the theorem, we replace F by the map G defined by G{x)=F{x)-\-y—Xy which satisfies the coer¬ civeness assumption (44) whenever F satisfies the stronger coerciveness assump¬ tion (46). ■ We can deduce the Leray-Schauder theorem on the existence of stationary points from theorem 11 by Poincare’s continuation method. We take X = RP and X to be a compact convex subset with a nonempty interior, so that the boundary dK = Knílnt К of К is distinct from K. THEOREM 18 (LERAY-SCHAUDER) Consider a compact convex subset KciR^ with nonempty interior and a strict upper hemicontinuous set-valued map F from [0,l^x К to with closed convex values. Suppose that for Я=0, the set-valued map x-^F(0, л:) satisfies the tangential condition, (48) Vx 6 дК, F(0, x)n Шх)ф0 and (49) Vx e дК, УЯ 6 [0,1 [, Oí Т(Я, x) Then there exists xeK such that 0 eF{l,x). A Proof We suppose that the conclusion is false and derive a contradiction. We set N:=dK, a closed subset of К and introduce the subset (50) M:={xeK \ ЭЯ € [0,1] satisfying 0 e F(A, x)}
CH. 6, SEC. 4 EXISTENCE OF ZEROS OF SET-VALUED MAPS 347 The subset M is nonempty because it contains a solution 3c € X to the inclusion 0 € F(0,3c), which exists by theorem 11 [thanks to assumption (48)]. It is closed, since the graph of F is closed. The intersection Mn N is empty: If x e and if X € [0, 1[, assumption (49) implies that x$M. We then introduce a continuous function (/> from K to [0,1] that is equal to zero on N and to one on M; for instance. ф{х):=- d{x, N) d(x, M)+d(x, Л0 We define the set-valued map G on К by (51) в(х):=Г{ф{х\ x) Then G is clearly upper hemicontinuous with nonempty closed convex values. It coincides with F(0, •) on N=dK and, consequently, satisfies the assumptions of theorem 21. Hence, there exists a solution xeK to 0 eG(3c)=F((/>(3c), 3c). But this implies that 3c 6 M and, therefore, (/>(3c) = 1, so that 0 e F(l, 3c). This is a contradiction. ■ The following consequence in corollary 19 is very useful. COROLLARY 19 Let Kbea compact convex subset of F” mth a nonempty interior and let G and H be two strict upper hemicontinuous maps from K to F” with closed convex values. We assume that (52) Vjc 6 dK, G{x)n Tk{x)^0 and (53) Vx e dK, V/i ^ 0, 0 i G(x)+pH{x) Then there exists a solution xeK to the inclusion 0 eH(x). A Proof We apply the preceding theorem with F(A, x):=(l — X)G{x)-\- XH{x). Condition (49) can be written in form (53). It follows that 7/=F(l, •) has a zero in K. ■ COROLLARY 20 Let Xq belong to the interior of a compact convex subset and let H be a strict upper hemicontinuous map from К to F” with closed convex values. We
348 CH. 6, SEC. 4 SOLVING INCLUSIONS suppose that (54) Vx e dK, V/i > 0, xo ^ x+pH{x) Then there exists a solution xe KtoOe H(x). ^ Proof. We apply corollary 19 with G(x)=x — xo. ■ We proceed by giving a method for selecting a fixed point of a set-valued map F. THEOREM 21 We assume that (55) and (56) К is a compact convex subset of a Hilbert space X F is an upper hemicontinuous set-valued map from К to К with nonempty closed convex values. We consider a function f:Kx K-^R satisfying (57) i. ^y e K, x->/(x, y) is lower semicontinuous. ii. Vx € K, y-^f (x, y) is concave. iii. f(y,y)^Q Finally, suppose that the set-valued map F and the function f are related by the property (58) {x e К such that sup /(x, >^)^0} is closed. у e F{x) Then there exists a solution x e K to the ''quasi-variational inequalities'" (59) {'• ^ ’ [ii. sup /(x,;;K0 A yeF(x) Remark Assumptions (55) and (56) are those of Kakutani’s fixed point theorem, which imply conclusion (59, i). Assumptions (55) and (57) are those of the Ky Fan inequality, and assumption (58) is a consistency hypothesis between / and F.
CH. 6, SEC. 4 EXISTENCE OF ZEROS OF SET-VALUED MAPS 349 Proof. We argue by contradiction: If the conclusion is false, then for all xeK, either x^F{x) or a(x:):=sup,eF(*) fix, y) is strictly positive. Saying that X 6 F(x) implies that there exists p 6 X* such that (p, x) — cr(F(x), p)>0. We set (60) i. Ao:={x 6X|a(x)>0} ii. Ap:={x e/CKp, x)-<T(f(x),p)>0) Then the negation of the conclusion can be expressed in the form (61) KczAou U Ap peX* Assumptions (56) and (57, i) imply that the sets Aq and Ap are open. Since K is compact, it follows that there exist/?i,... such that (62) KcAou U Ap, i=l and there exists a continuous partition of unity {a,} for i =0,...,« associated to this covering. We then introduce the function 4>: KxK-^R defined by (63) <i)ix,yy.=aoix)fix,y)-\- Yj aiix){pi,x-y) i=l This function (j) is lower semicontinuous with respect to x, concave with respect to and satisfies (¡)(y, >^)=^0 for all y e K, thanks to assumption (57, iii). The Ky Fan inequality (see theorem 3.5) implies the existence of x eK satisfying (64) sup (¡){x, y)^0 yeK We are going to contradict this inequality by proving that there exists y eK such that (65) We take (f){x, ;;)>0 (66) ye F(x) arbitrary when (x(x)^O and satisfying fix, y)'^OL{x)/2 when a(3c)>0 (which is possible) Since [ai] for i=0,. . ., A2 is a partition of unity, then a,(x)>0 for at least one
350 CH. 6, SEC. 4 SOLVING INCLUSIONS index 1=0,Inequality (65) will then follow from the following statements: (67) i. ao{5c) > 0 implies that/(3c, j^) > 0 ii. a,(x)>0 implies that (/?,, 3c—j;) >0 Let us verify these statements. If ao(.x)>0, then 3ceAo and, consequently, a(3c)>0. Hence, /(3c, j;)^a(3c)/2>0. If ai{x)>0 for then 3c e Ap. and, con¬ sequently (pi, x}>(7{F{x),Pi)>(pi,y} because y e f (3c). Hence, (p„ x - ) > 0. ■ Since the function a defined by oi(x):=supyeF(x) f{Xy y) is lower semicon- tinuous when F and / are lower semicontinuous, we obtain the following corollary because assumption (58) is satisfied. COROLLARY 22 We assume that К is convex and compact, F : K-^K is a strict continuous set¬ valued map with closed convex values, and f is a lower semicontinuous function defined on К xK, concave with respect to у and satisfying sup yex /(>^, >^)^0. There exists a solution xeK to the quasi-variational inequalities (59). ▲ We continue the study of л-person noncooperative games begun earlier. A game is^defined by n strategy sets Mi and n loss functions / defined on the product М:=П”=1 of the strategy sets, to which we have associated the canonical decision rules С,- defined by Ci(xf):= {xi 6 Mi I /(xi, xf) = inf fiyi, X?)} У16 Mf This time, we add a concept of feasibility to the game through the n set-valued maps Fi restricting the strategies of the ith player to the subset Fi(x\)^Mi when all the other players have chosen their strategies Xj e In such a game (often called a metagame), the other players influence player j (a) Indirectly, by restricting у’s feasible strategies to Ffx]). (b) Directly, by affectingy’s loss function/•. The conjunction of those two effects lead to feasible decision rules defined by (68) Ci(xi):={xieF,(xf)|/i(xi, x?)= inf fiyi,xi)] yieFi(xi) The multistrategies xeM that are consistent with these decision rules are called
CH. 6, SEC. 4 EXISTENCE OF ZEROS OF SET-VALUED MAPS 351 social equilibria of the metagame. So, they are defined by the conditions (69) i. Vi=1,..., Й, xi€ Fiixi) (feasibility) ii. Vi = fi{xi,xi)^ inf Myi,xi) yi€Fi(xD THEOREM 23 Suppose that for all i = \,,.. (70) i. The strategy set Mi is convex and compact. ii. The loss function f is continuous andfor every xt e Mf, yi ^МУь xt) is convex. iii. The feasible map F,-: Mt-^Mi is continuous with closed convex images. Then there exists a social equilibrium x, that is, a multistrategy x satisfying condi¬ tions (69). A Proof We set (71) and (72) K:=Y\Mi, ВД:=ПВД 1=1 i=l fix, y)'.= Ё ifiiXi, xi)-fiiyi, Xi)) i=l Then the set K is convex and compact, and the set-valued map F is continuous with closed convex images. The function / is continuous and concave with respect to y and satisfies supyg^/O?, j?)=0. Thus corollary 22 implies the existence of a solution to the quasi-variational inequalities 3ceF(3c) and sup f{x,y}^0 y e Fix) The first statement implies that xi e Fi{xt) for all i = 1,..., n. The second state¬ ment implies that for all / = 1, . . . , « and for all yi e Mi,f(xi, xi)^f(yi, xt). It suffices to take y such that yi=yi and yj=Xj for all j^i. Hence y e F{x) and fix, y)=fiixi, xi)-f^yi, Xf)+ X ifAxj, Xj)-fjiXj, Xj))=fiXi, xt)-fi(yi, Xi)^0 jf‘ m In 1929, Knaster, Kuratowski, and Mazurkiewicz provided a new proof of Brouwer’s fixed point theorem based on lemma 26. This theorem was genera¬
352 CH. 6, SEC. 4 SOLVING INCLUSIONS lized by Shapley in view of applications to cooperative game theory. It is a direct consequence of theorem 11. Let I.^:=<xeR\ i=l be the (« — 1) simplex. We associate with any nonempty subset T of the set A^:={1,the face (73) {;c 6 S’lXi^O for all i 6 T} Let ^(N) denote the family of nonempty subsets T of N and let Cr be the characteristic function of the subset T [defined by Ct-(0=1 when i e T and Cr(0=0 when i i T], THEOREM 24 (KKMS) Let {F7-}t6 »(N) ^ family of closed {possibly empty) subsets of £" satisfying the property (74) VT€^(7V), U i'll Then there exist nonnegative scalars m(T) such that (75) i. Cn= E m{T)CT Te ^{N) ii. n Ft^0 m(T)>0 We observe that condition (75, i) can be written as (76) fieN, Y. m{T)=l ie T DEFINITION 25 A family ^ of subsets T e ^(N) such that (77) CN = Y,m{T)CT, m(T)>0 for all T Si T is called a balanced family. 4 Hence, theorem 24 can be restated as saying that assumption (74) implies the existence of a balanced family ^ such that H Ft^0- This is a generaliz¬ ation of the famous “Three Poles lemma” or “KKM lemma”.
CH. 6, SEC. 4 EXISTENCE OF ZEROS OF SET-VALUED MAPS 353 LEMMA 26 (KKM) Let us consider n closed subsets Fi of S" satisfying (78) Then (79) VxgZ”, xe (J Fi {i\xi>0} n i=l Proof of Lemma 26. We apply theorem 24 with Fs:=F, when T = {i}, i = l, . . . , and Fj'=0 when card(F)^2. Then assumption (78) implies the corresponding assumption (74). Then condition (75, i) shows that m{T)=m{{i}) is equal to one for all i so that (75, ii) is reduced to (79). ■ Proof of Theorem 24. We apply theorem 11 to the set-valued map G from Z” to R” defined by card(A^) {card(r)}i-j.9;t The values of G are obviously convex and compact, and G is clearly upper semi- continuous. It satisfies the tangential condition, VxgZ”, G{x)nTMx)^0 Indeed, let T be the subset of indexes i such that x/>0. By assumption (74), there exists a nonempty subset RczT such that x belongs to Fr. Hence, y'= Cn Cr card(A^) card(F) belongs to G(x). It also belongs to Tzn(x) because Y!i = i if i does not belong to T and, consequently, does not belong to R, so that;;, = l/card(A^)^ 0. Therefore, theorem 11 implies the existence of ic e Z” such that 0 e G(3c), that is, such that (80) Cn= Yj aTT^—xe f] Ft B card(T) A(T)>o Remark We can prove Brouwer’s fixed point theorem from the KKM lemma. Indeed, let/ be a continuous map from the simplex Z” to itself. We introduce the sub-
354 CH. 6, SEC. 5 SOLVING INCLUSIONS sets Fi defined by Fi:={xel.''\xi^fi(x)} which are closed because/ is continuous. They satisfy assumption (78), other¬ wise there would be x g L”, which belongs to Q{fUi>o} ^ other words, we would have/(x)>Xi whenever Xf>0. Since both x and f(x) belong to L", we obtain the contradiction 1 =X^=i^ jc, = l. Therefore, there exists X that belongs to the intersection of the subsets F„ that is, that satisfies Vi = 1,Xi>fi(x) We cannot have the strict inequality Xi>fix) for at least one U because such an inequality would imply that 1 > 1. Hence, 3c, =/(3c) for all / = 1,■ Remark Ky Fan’s inequality and, more generally, theorem 3.4 can also be deduced from the KKM lemma. Therefore, Brouwer’s fixed point theorem, Ky Fan’s inequality, theorem 11, Kakutani’s fixed point theorem, the KKM lemma and KKMS theorem are all equivalent, as indicated by the diagram on the opposite page. 5. WALRAS EQUILIBRIA AND PRICE DECENTRALIZATION We apply the preceding theorems to find a possible explanation for the role of price systems in decentralizing the behavior of different consumers; that is, knowledge about the price system allows each consumer to make a choice without knowing the global state of the economy and, in particular, without (necessarily) knowing the choice of other consumers. There is no doubt that Adam Smith is at the origin of what we now call decentralization, that is, the ability of a complex system moved by different actions in pursuit of different objectives to achieve an allocation of scarce resources. He introduced this mysterious and quite paradoxical property in a poetic way. Let us quote the famous passage from The Wealth of Nations, published in 1776, two centuries ago. Every individual endeavours to employ his capital so that its produce may be of greatest value. He generally neither intends to promote the public interest, nor knows how much he is promoting it. He intends only his own security, only his own gain. And he is in this led by an invisible hand to promote an end which was no part of his intention. By pursuing his own interest, he frequently thus promotes that of society more effectually than when he really intends to promote it. Adam Smith did not provide a careful statement of what the invisible hand manipulates nor, a fortiori, a rigourous argument for its existence. We had to
355
356 CH. 6, SEC. 5 SOLVING INCLUSIONS wait a century for Leon Walras to recognize that price systems are the elements on which the invisible hand acts and that actions by various agents are guided by those price systems, providing enough information to all the agents to guarantee the consistency of their actions with the scarcity of available com¬ modities. Walras presented in 1874 the general equilibrium concept as a solution to a system of nonlinear equations in “Elements d’économie politique pure.” An equal number of equations and unknowns led him and his followers quite optimistically to assume that a solution does exist. But it required one more century for Arrow, Debreu, Gale, Nikaido, and others to provide rigorous statements and proofs of the existence of an equilibrium. In order to represent an exchange economy, we begin by introducing ( types of commodities that are endowed with an “unit,” so that we can speak of x units of a commodity. A commodity may involve not only its physical proper¬ ties, but also the place where it is available, the date when it is available, and, in the case of uncertainty about the future, the elementary event that will be realized (for instance, 100 kilograms of bread that will be available in New York in 32 days if there is a truckers’ strike). Services may also be included as long as units are perfectly well defined. So, a commodity bundle is a vector x e R\ which describes the quantity xh of each commodity A = 1,..., ^. The description of an exchange economy begins with (1) the subset oi available commodities and continues with the specification of n consumers / = 1,...,«. A first definition of a consumer i begins with (2) the consumption set Li^zR^ which is interpreted as the set of commodity bundles he or she needs. If x € L,-, then Xh is I’s demand of commodity h when Xh>0, and \xh\ is the i’s supply of commodity h when Xh<0. The question now arises, can consumers share an available commodity? We define an allocation 3c e {R^Y as n commodity bundles x, e Li such that their sum ^ Xi is available. We denote by (3) K:=-^3c€n Li i=l Z XiSM the set of allocations. The next question that arises is whether we can devise mechanisms that provide allocations. Indeed, when the number of consumers is large, finding an allocation is difficult, because it requires possessing a large amount of informa¬ tions. Can we find a way of summarizing enough information to allow each consumer to choose his or her consumption in a decentralized manner?
CH. 6, SEC. 5 WALRAS EQUILIBRIA AND PRICE DECENTRALIZATION 357 This is possible by using price systems. A price system is a linear functional p that associates with any commodity xeR^ its value (/?, x) e R. We denote by l.^:={p eRi\Y,i=iPi = i} the price simplex. We next regard the support function (4) r(p):=sup {p,y} yeM as the gross income, which is the maximum value of the available commodity bundles for the price system p. The solution proposed by Walras and his followers consists in letting price systems play a crucial role by summarizing enough information about the economic system for the n consumers. A consumer is defined as an automaton associating to every price p and every income r (in monetary units) his or her demand r) e Li, which is the commodity bundle that the consumer buys when the price system is p and income is r. In other words, a consumer i is char¬ acterized by the demand function 1,^ xR-^Li, which describes the behavior of consumers. We observe that it is possible to devise other mathematical descriptions—models—of the same behavior, and we shall do so. We also mention that neoclassical economists assume that demand functions are derived by maximizing a utility function in compliance with the first sentence in Adam Smith’s passage. But this is by no means necessary. Anyway, assume for the first time that consumers are just demand functions 5i{., .), given independently of the set M (which cannot be known by the con¬ sumers). We need another assumption to define the Walrasian mechanism. If p is the price system, we assume that the gross income r{p) is allocated among consum¬ ers in incomes rip) (5) Z rip)=r(p) i = 1 We insist on the fact that the model does not provide this allocation of income but assumes that it is given. In summary, the mechanism we are about to describe depends on (a) the description of each consumer i by its demand function ¿¿(.,.) (b) an allocation K/^)=X”= i ^i(p) of the gross income Hence, when p is the price on the market, the income of consumer i is rip) and demand is dip, rip)). The mechanism works if and only if demand balances supply, that is, if and only if there exists a price p such that (6) Z fiiP)) ^ M 1=1
358 CH. 6, SEC. 5 SOLVING INCLUSIONS DEFINITION 1 A price system p (^) Isolds true is called a Walras equilibrium; we say that ''it clears the market'' A Observe that this mechanism is decentralized: The choice of the ith con¬ sumer depends only on the price system p (through its income function rO and does not require knowledge of the other consumers’ choices. This illustrates better Adam Smith’s quotation, because by choosing a Walras equilibriump^ the invisible hand promotes an end, Yl= i which was not part of the consumers’ intentions. The task now is to solve problem (6), which requires us to choose among all the sufficient conditions that can be devised those that have an economic interpretation. It is remarkable that a budgetary constraint on consumers’ behavior, known as the Walras law, provides such a sufficient condition. The Walras law forbids consumers to spend more than their incomes, that is, (7) Vi = l, (p,^i{p,r)}^r Actually, this law can be less rigorous; it is sufficient to assume that (8) E '■()>< E '■i i=l i=l This latter law, the collective Walras law, allows financial transactions among consumers. We insist on the fact that the Walras laws (7) or (8) do not involve the set M of available resources. We shall prove that Walras law (8) implies the existence of a Walras equi¬ librium, that is, allows Adam Smith’s invisible hand to provide a price system summarizing information on the state of this economy—the subset M and behavior of each consumer described by the demand functions ¿f, and their share r/(.) of the gross income. After pioneer work by Wald, von Neumann, Kakutani, and so on, started in the 1930s, the first proof of the existence of a Walras equilibrium was due to Arrow and Debreu in 1954. Further work on this problem was due to McKenzie, Nikaido, Uzawa, and many others. THEOREM 2 Let us assume that (9) M=Mq — R+ is closed and convex, where Mq is compact and (10) the demandfunctions are continuous and satisfy the collective Walras law (8)
CH. 6, SEC. 5 WALRAS EQUILIBRIA AND PRICE DECENTRALIZATION 359 and (11) the income functions r,- are continuous. Then there exists a Walras equilibriump A Proof It is consequence of theorem 4.4 applied to the set-valued map F defined by (12) F{p):=Mo- 'Z Hp,n(p)) i=l which is obviously upper hemicontinuous with compact convex values and satisfies (13) a{F(p),p) = rip)- Z (Py^-iP’ri{p)))>r(p)- Z ri(p)>0 i=l /=1 Hence, there exists a solution ^ 6 2^ to the inclusion 0 6 F(p)-Ri =M- t Hp, riip)) m i = l Remark We can relax assumption (9) (see for instance Aubin, 1979b, theorem 8.2.2, p. 248). Theorem 1 can also be extended to the case of set-valued demand maps. We are going to propose another model that keeps the essential ideas under¬ lying Adam Smith and Léon Walras’s proposals but can take into account the dynamic nature of the behavior of each agent by its instantaneous demandfunc¬ tion di'. which sets the variation in consumer’s i demand when the price is p and consumption is x. We assume that (14) i. M = Mo — R + is closed and convex where Mo is compact. ii. Vi = 1,...,«, L, is closed, convex, and bounded below. ¡ii. 0 6lnt(Xi=i L,-M) We posit the following assumptions on the instantaneous demand functions dr. i i. Vi = l,..., n, the function di: is continuous. [ii. VxeLj, di(x,p)eTL^(x)
360 CH. 6, SEC. 5 SOLVING INCLUSIONS and (16) Vx e Li, p-^di{x, p) is affine. We still have a concept of equilibrium associated to this mechanism. DEFINITION 3 An economic equilibrium is a sequence (x^ ,Xn,p) of n consumptions 5c,- and a price system p such that (17) Vi = l,...,«, N/ i = 1,..., n, di(Xi,p)=0 Xi 6 Li, p In order to keep all the good features of the Walras model, we must check that there are sufficient conditions with an economic interpretation. This is still the case, since we shall prove that equilibria (17) do exist if the instantaneous demand functions di satisfy the instantaneous Walras law (18) V/7 e Vx, € L„ (p, diixi, p)) <0 This is a financial rule that requires that for each price, the value of the rate of change of each consumer is not positive, that is, each consumer does not spend more than he or she earns on an instantaneous exchange of goods. As in the Walras model, the instantaneous Walras law does not involve the subset M of available supplies. More generally, we shall assume that the instantaneous collective Walras law (19) Vx 6 n U '^p e /= 1 </>, Yj di(Xhp))<^0 i = 1 holds true. THEOREM 4 We posit assumptions (14) on the economy and assumptions (15) and (16) on the instantaneous demand functions di. We assume that the instantaneous collective Walras law (19) holds true. Then there exists an equilibrium {x, p)eKx'L^ Isatirfying (17)']. A Proof. It follows from theorem 4.21 applied to the set-valued map D from K<={Ry to (Py defined by (20) D(x):={d{x,p)}pexf, where d(x, p):=idi{xup\ ■■■, d„{x„, p)) € {R^f
CH. 6, SEC. 5 WALRAS EQUILIBRIA AND PRICE DECENTRALIZATION 361 Obviously, assumptions (15) and (16) imply that D is upper hemicontinuous with compact convex images. Furthermore, we deduce from assumptions (14, i and ii) that K is convex compact and from assumptions (14, ill) that the tangent cone to at x is equal to (21) 7K(x):=|t)6 n Tn(x) 1 = 1 \i = l It remains to check that (22) VxeK, F{x)nTK{x)^0 or, equivalently, thanks to assumption (15, ii) and formula (21), that (23) Vx e K, 3pe'L^ such that ^ di(xi,p) e Tm ^ Otherwise, there would exist x eK such that Tm Z Z di(Xi, ^)| ^^=0 The separation theorem implies that in this case, there would exist peTM ( Z XiJ =-^M ( Z ) and e>0 such that (24) inf (p, Z di(Xi,q)/^e \ i=l Since Nm(Ya=i Xi)<^Ri because M=Mo — R+, we deduce a contradiction to the instantaneous collective Walras law (19) by taking q‘-=p/Yh=i Ph- Inequality (24) implies that Z {q, diiXi, q))>e>0 . ■ i = 1 Walras defined not only the concept of equilibrium but proposed a process known as Walras’s tâtonnement {tâtonnement means tentative process, trial and error, etc.—literally, groping, feeling one’s way in the dark). Indeed, a Walras
362 CH. 6, SEC. 5 SOLVING INCLUSIONS equilibrium is an equilibrium for the “excess demand map” E defined by (25) V/7, E(p):=Y, 0i{p,n)p))-M i=l So, the idea was to associate the dynamic system (26) p'{t) e E{p(t))\ p(0)=po and study under which conditions we can prove that p{t) converges to a Walras equilibrium p, solution to the inclusion 0 e E(p). We observe that if p(t) is a price supplied by the Walras tâtonnement process (26) and p{t) is not a Walras equilibrium, it cannot be implemented, because the associated total demand nipit)) is not necessarily available. Hence, this model forbids consumers to transact as long as the price p{t) is not an equilibrium. It is as if there were a superauctioneer calling prices and receiving transactions offers from consumers. If the offers do not match, he or she calls another set of price according to rule (26) but does not allow trans¬ actions to take place as long as the offers are not consistent. Tâtonnement does not provide a model of how prices are actually evolving. The fundamental nature of the Walras world is static, while we live in a dynamic¬ al environment where no equilibria have been observed. By means of the second mechanism, which describes consumers i as auto¬ mata associating to each price system p and consumption bundle Xi their rate of change we can regard an equilibrium (jci,..., x„,p) as an equilibrium for the dynamical system (27) Vi = 1,..., x'i(i)=diixiit), pit)) ; x,(0)=X? When the price p{t) evolves, so do consumptions A:,(i) according to the differen¬ tial equations (27). So, a viability problem arises: Does there exist a price functionp{t) such that the sum ^f(^) consumptions remains available? In other words, do the trajectories Xf(.) of the n coupled differential equations satisfy the viability condition (28) Vi>0, 1 = 1 We observe that this mechanism shares with the Walras model the decen¬ tralization property: The price system p{t) summarizes enough information about the economic system to allow each consumer to change her or his own consumption independently of other consumers and ignore the set M of avail¬ able supplies.
CH. 6, SEC.6 MONTONE MAPS 363 It is possible to prove that under the assumption of theorem 4, we can associate to each initial allocation xo a price function p(t) such that trajec¬ tories in (27) satisfying viability condition (28) do exist (see Aubin and Cellina, [1984], Chapter 5). Furthermore, we can show that the price system evolves as a feedback control: It depends on time through the state of the system, in the sense that there exists a set-valued map C from the set of allocations to the set of prices such that (29) for almost all t>0, p{t) e ..., x„(t)) In this framework, we see Adam Smith’s invisible hand (which, more to the point, should be called invisible brain) setting prices as functions of allocations for the purpose of promoting consumers to respect scarcity constraints. It is in this sense that we may regard such a dynamical system as a regulation mechanism. So the feedback condition (29) involves the price system, but does not influence its variation. This is the second important difference between dynamical systems such as (26) that require, so to speak, Adam Smith’s invisible hand to actively, as in mechanics, set prices to adjust demand to supply. In this model, only consumers are supposed to have the ability to act dynamically according to differential equation (27), and prices “follow” consumption according to relation (29). ■ 6. MONOTONE MAPS Let K be a closed convex subset of a Hilbert space X (identified with its dual). The problem arises how to transform a set-valued map A : K-^X that does not satisfy the tangential condition. (1) WxeK, -A{x)nTK{x)f0 to another set-valued map F that satisfies the condition. The simplest idea is to project the images of A(x) onto the tangent cone Tk(x\ that is, to define F by (2) F(x):=7i7-xw(-^W) which, by construction, satisfies the strong tangential condition. The zeros of A (or — A) are zeros of F. Unfortunately, we cannot apply theorem 4.21 because this map F inherits neither the upper hemicontinuity of A nor the convexity of the images of A. But we observe that the zeros of a set-valued map F are the zeros of the map m(F): x-^m{F(x)) associating to x the elements of F(x) with the smallest norm. Since Nk(x\ the normal cone to K at x, is the negative polar cone
364 CH. 6, SEC. 6 SOLVING INCLUSIONS of the tangent cone Ta{x), we know that the orthogonal projections satisfy (see Fig. 1). This implies the following lemma. LEMMA 1 (4) (See Fig. 2). m{^TK(x)i - = - m(A{x) + Nk{x)) Proof. Set A:=A(x), T:=Tic{x) and N:=Nk{x). The lemma follows from the equality inf l|jtr(->')||= inf Il-_v-7tft,(-j)||= inf inf ||->'-z||= inf Hull ■ yeA yeA yeAzeN veA + N So, the zeros of the set-valued map x-^7iTj^^x)i—Mx)) are the zeros of the set- values map x-^A{x)-\- Nk(x\ which is simpler to handle (see Fig. 3). This justifies the following definition. DEFINITION 2 Let K be a closed convex subset of X. A ''variational inequality'' for A on K is an inclusion of the form (5) 1. xeK ii. 0 e.4(ic)+A^k(-5c)
CH. 6, SEC. 6 MONOTONE MAPS 365 Example Optimization with constraints provides examples of variational inequalities, which we saw when U was a lower semicontinuous, convex function such that (6) 0 6lnt(Dom U-K)
366 CH. 6, SEC. 6 SOLVING INCLUSIONS because the elements x eK achieving the minimum of (/ on are solutions to the inclusion (7) OedU{x) + NK(x) that is, a solution to the variational inequality with A :=dU. ■ Remark Variational inequalities is an expression that was coined because in the single¬ values case, (5) can be written (8) ii. xeK |ii. Vy e K, (A{x),x-y)<:0 We observe that when K is a cone, this system becomes (9) i. xeK ii. Axe -K~ Hi. (Ax, x) =0 ■ Remark Both x->nт¡^^x)A{x) and x->^/4(jc)-I-Wk(x) coincide with A(x) when x belongs to the interior of K. Hence, any solution to the variational inequalities that belongs to the interior of K is a zero of A. When — A satisfies the strong tangential condition. (10) VxeK, -A(x)<=Tk(x) any solution X to the variational inequalities (5) is a zero of A. Indeed, there exists veA{5c) such that 0=ij-l-7twKef)(5), and, consequently, 7tjvK(S)(t))=-ti belong to Tk(x. This implies that i5=0 and, thus, that x is a zero of A. ■ We recall that Nk{x), the normal cone to K at x, is the subdifferential of the indicator ij/ic- Therefore, variational inequalities are particular cases of inclusions of the form (11) feA{x)+dV{x) when V: —oo, -f-oo] is a proper lower semicontinuous, convex function and ^ is a set-valued map from the Hilbert space X to itself. We assume once and for all that (12) i. Dom V <= Dom A ii. Vx 6 Dom A, A{x) is convex and weakly compact.
CH. 6, SEC. 6 MONOTONE MAPS 367 We associate to the function V, the map A, and an element/ e X the function (f> defined on Dom V by (13) We observe that (14) <t>(y):=V(y)+ inf iV*{f-u)-(f-u,y)) u 6 A{y) VjeDomK <f>{y)^0 since, for all ueA{y\ V{y)+V*{f—u)—(f — u, y)>0, thanks to the Fenchel inequality. We can also characterize the set-valued map A by the function y defined on Dom A X Dom A by (15) y{x, y):= inf (p,x-y) = -a{A{x), y-x) peA(x) PROPOSITION 3 We posit assumption (12). The following problems are equivalent: (16) i i. 33c € Dom Vsuch that f e Ax+dV{x) ii. 3/6 Dom V such that f ep+AdV*{p) iii. 3x6 Dom V such that Vy 6 Dom V, y(x, y)-(f x-y) + V(x)-F(y)<0 Iv. 3x e Dom V such that (j)(x)=0 {= min <l>{y)) ^ y e Dom V Proof a. Let 3c be a solution to (16, i); then there exists p^dV(x) such thatf—p 6 Ax<=-AdV*{p). Conversely, let/ be a solution to (16, ii). Then there exists 3c 6dV*(p) such that f ep +Ax. Since p 6 dV{x), then/ 6 dV(x)+Ax. _ b. Let 3c be a solution to (16, i). There exists ü e A(x) such that/ sdV(x)+u, that is, such that Vj 6 Dom V, (m, x-y)-{f, x-j') + F(x)- F(j)<0 By taking the infimum on A{x), we deduce inequality (16, iii). c. Inequality (16, iii) can be written sup inf [F(3c)-F(y)-(/-M, y e Dom Vue i4(x) Since Dom V is convex, A{x) is convex weakly compact, the lopsided minimax
368 CH. 6, SEC. 6 SOLVING INCLUSIONS theorem 2.7 implies that the left-hand side of this inequality is equal to inf_ sup \V(x)-V{y)-{f-u,x-y)'] u e A(x) y e Dorn V = inf_ \_V{x)-\-V*{f-u)- if-u,x)'] = <t>{x) u e A{x) Hence, (f)(x)<0, and since we have already observed that 0(x)>O, we conclude that (¡){x)=0. d. Let X 6 Dom V satisfy <^(jc)=0. Since A{x) is weakly compact and F* is weakly lower semlcontinuous, there exists iJ 6 A(x) such that (¡){x)= V{x)+ K*(/-w)“ </-w, x)=0 This is equivalent to saying thatf—uedV{x\ that is, x solves (16, i). ■ The equivalence between (16, i and iv) allows us to interpret the solutions to problem (11) as solutions to a minimization problem (minimization of the functional (¡)) and provides a variational principle. The equivalence between (16, i and Hi) allows us to solve problem (11) (and, in particular, variational inequalities) by applying minimax inequalities from Section 3 to the function (/> defined by (17) <t>(x, y):=y{x, j)- if, x-y) -I- V{x)~ V(y) We observe that (18) i. Vx, y-^(l>{x, y) is concave. ii. V;;, <l){y,y)=0 If we want to apply Ky Fan’s inequality (theorem 3.5), we have to verify that (19) V>^, x^(f>{x, y) is lower semicontinuous. To satisfy this assumption, we must require A to be upper semicontinuous from X supplied with the weak topology to X supplied with the strong topology. This is not reasonable (in the case of infinite dimensional spaces), since this requirement is not met by the subdifferential x^dU(x) of lower semicontinuous convex functions. Still, when K is weakly compact, there exists an element X eK achieving the minimum of U on X, that is, a solution to the variational inequalities 0 edU{x) + N¡c{x)
CH. 6, SEC. 6 MONOTONE MAPS 369 But the subdifferential 3 i/ is a monotone map, and, consequently, the functions y and ф are monotone in the sense that y(x, y) + y{y, x)^0 (ф{х, у) + ф{у, x)>0) Therefore, we shall apply theorem 3.9 (Ky Fan’s inequality for monotone functions) for solving variational inequalities and problem (11) when A \s a monotone map. DEFINITION 4 We shall say that a set-valued map A from X to X is monotone if its graph is monotone in the sense that (20) V(.x,/?)egraph (^), V(j;, ^) egraph (^), (p-q,x-y)>0. к The terminology comes from the monotone maps from Rio R; for instance, a nondecreasing map F from Rio R Is monotone. Example The map Ifjc<0,/W=x lfx=0, f{x)=0. Ifx>0, f{x)=x+\. is a monotone map as well as the map 1Гд:<0, f{x) = x. Ifjc=0, /(x) = [0, 1]. Ifx>0,/(jc)=jc+l. More generally, if/ is a nondecreasing map from R to /?, the map x-^A(x):= ifix-Xfix-^-flnR is monotone. We give several other examples of monotone maps.
370 CH. 6, SEC. 6 SOLVING INCLUSIONS Let / be a nondecreasing function from an interval Dom /<=i? to R. Let Q 6 7?" be open and X :=l3(Q). We define the set-valued map A from X to X by (21) for almost all m e Q, ))(co);=/(x(co)) It is clear that A is monotone. Let V: be a proper lower semicontinuous, convex function. Then the set-valued map x^dV{x) is monotone. Let l/:Xxy^/?u{—oo}u{-l-oo}bea proper function satisfying (22) i. Vj e У x-^ U{x, y) is concave and upper semicontinuous. (¡i. '^xeX, y-»C/(x,j) is convex and lower semicontinuous. Then the set-valued map (23) (x,y)^d*(- Щх,y)y.dyU(x,y)<zXxY is monotone. DEFINITION 5 We say that a proper set-valued map F from X to X is tionexpansive if (24) V(p, x), (q, y) 6 graph (F), \\p-q\\ \\x-y\\ PROPOSITION 6 If F is a proper nonexpansive set-valued map from X to X, then A:=l—F is monotone. A Proof Let (x, p) and {y, q) belong to the graph of F. Then {x-p-(y-q\ x-y)=\\x-y\Ÿ--{p-q, x-y) >\\x-y\?-\\p-q\\\\x-y\\^Q ■ The converse is not necessarily true. ■ If К is a closed convex set, the projector is a monotone map, since <Як(л:)-Як(у), х-у)^Цяк(л:)-Ях(у)|Р>0 ■ Since monotonicity is a property bearing on the graph of A, then (25) A is monotone if and only if ^ ' is monotone. If A and В are monotone and X>0, p>0, then XA+pB is monotone. If A is monotone, then cd{A): x->cô(v4(x)) is also monotone. If we supply X xX with
СН. 6, SEC. 6 MONOTONE MAPS 371 the product of the strong topology and the weak topology, we can see that the closure of a monotone graph is still monotone. We begin by mentioning the following characterization in proposition 7. PROPOSITION 7 A set-valued map A from X to X is monotone if and only if (26) V2>0, V(a:,p), V(>', ^) € graph (^), ||x-j^||<||x-y+A(p-^)|| A Proof We compute Wx-y+X(p-q)\\'^ = \\x-y\\^-\-X^\\p-q\\^-\-2X{p-q,x-y) Hence, if A is monotone, inequality (26) ensues; conversely, if (26) holds true, we deduce that X^\\p-q\?' + 2X{p-q, x-y}>0 We obtain monotonicity by dividing by Я and letting Я converge to zero. ■ Remark We observe that condition (26) does not involve the scalar product; therefore, it can be used on Banach spaces. In this case, maps A satisfying (26) are called accretive maps. ■ We associate to the map A its resolvent (27) Л;=(1+А^)-1 An important consequence of property (26) is given in proposition 8. PROPOSITION 8 The resolvent Jxof a monotone map A is a single-valued nonexpansive map from \m{l-\-XA)to X. A Proof Let xeJx{u) and yeJxiv). We can write uex-\-XA{x) and vey +XA(y). Property (26) implies that ^ ,(u—x V—y\ r) = ||w-i;|| This shows that Jx is nonexpansive and by taking u = v, that Jx contains a unique point. ■
372 CH. 6, SEC. 6 SOLVING INCLUSIONS We shall characterize monotone maps A such that \m(I-\-XA) = X for >i>0; they are the maximal monotone maps, studied in the next section. There are maps that enjoy stronger monotonicity properties, such as maps satisfying (28) or (29) V(x,p), V(;;, i) egraph(^), {p-q, x-y)^c\\x-y\\ V(x,p), V(j, 9) € graph(^), {p-q,x-y)-^c\\x-y\\^ (with oO, a> 1). We shall describe the monotonicity of /i by a proper non¬ negative lower semicontinuous, convex function p: X-^R+u{ +oo}. DEFINITION 9 Let P: X-^/^+u{+cx)} be a proper lower semicontinuous, convex function. We say that A is p-monotone if (30) V(x,p), (7, i) 6 graph(^), (p-q,x-y)>P{x-y) A For instance, we can always take (31) i. P(z):=0 (and thus P*=il/{Oh P* = {0}) ii. P{z):=\\z\\ {and thus p*=il/B,^om P*=B) iii. p{z):=- Wzll'^iandthus j?*=— || ||“*, -+—=1, Dom P* = X*\ a \ a ct,,, y In the following theorem, we shall measure the degree of monotonicity of A by the size of the domain of /3*: the larger Dom /?*, the more monotone A. ■ We now solve the problem (32) feA{x)+dV{x) when /1 is a monotone set-valued map with weakly compact convex values. We recall that we have introduced the function y, defined by (33) yix,y)-= inf (p,x-y) peA(x) We observe that the function y is monotone in the sense that ii. V;;eDom(^), r(j',>')<0 [il. Vx, y € Dom (A), y(x, y)+y(y, x)>0
СН. 6, SEC. 6 MONOTONE MAPS 373 Indeed, у(л:, j) + r(j, x)= inf {p-q,x-y)>0 P e A{x) q e A(y) Also, y-^y{x, y) is obviously concave and upper semicontinuous. We shall need the weakest continuity property we can think of. DEFINITION 10 We say that a set-valued map from a convex subset M<^X to X* is finitely upper semicontinuous if it is upper semicontinuous from M supplied with the finite topology to X supplied with the weak topology. A LEMMA 11 If A is finitely upper semicontinuous from M to X, then the functions x-^y{x, y) are lower semicontinuous for the finite topology. A Proof. Let К := {xi,..., be a finite subset of M, and set Pk'. X eZ"->^x(A):= ^ Я,х,- i=l We have to prove that X^y{Pj^{X\ y) is continuous. Let Ж '.= \p 6 X* max Kp, be a neighborhood of zero for the weak topology. Since A is finitely upper semicontinuous, there exists rj>0 such ||A —implies that A(PkW)c: A{Pk(^o))-{- Jf. Also, Y\ can be chosen so small that sup (/?, PkW—Pk{^o))\^-;^ pGAip^iXo)) ^ Hence, for any p eA{Pn(^)), there exists po eA{Pni^o)) such that p—poejf, and thus, (p—po, j) <e/2 by the very definition of On the other hand, {po, Pk{^o)-PkW)^^/'2- We deduce that yiPki^o), yH (po, PK(^o)-y) = (p> PKW—y} + {po—p, PKi^)—y)+(poy Pki^o)—PkW) ^ (py PkW~y) +6
374 CH. 6, SEC. 7 SOLVING INCLUSIONS By taking the infimum when p ranges over we deduce that P-^11 ^^=>y(pKi^ol yHyiPAh y) + e We shall now solve inclusion (32) when (35) i. K is a proper lower semicontinuous convex function. ii. ^ is a monotone finitely upper semicontinuous map with weakly compact convex values. Hi. Dom V c Dom A We observe that a necessary condition for the existence of a solution to this problem is that (36) / e Dom V* + A Dom V Indeed, proposition 3 implies the existence of a solution to the equivalent problem f ep + AdV^ip); we deduce that p € Dom 5F*c:Dom K* and 5K*(^)c:Im 5K* = Dom SKcDom 7 and thus / e Dom + ^ Dom V. We shall prove that this condition is almost sufficient. THEOREM 12 We posit assumptions (35). Assume moreover that A is p monotone. Then there exists a solution x to the inclusion (32) when (37) / e Int(Dom K* + Dom V + Dom )?*) Remark The size of Dom jS* balances the interiority condition in assumption (37), as the following corollary shows. ■ COROLLARY 13 We posit assumptions (35). a. If A is monotone, then (38) Int(Dom + Dom F)cIm(yl + 5F)c:Dom V^-^A Dom V b. If there exists c> 0 such that (j, ^) €graph(^), {p-q,x-y)'^c\\x-y\\ then (39) \m(A +dV)=Dom V* + A Dom V
CH. 6, SEC. 6 MONOTONE MAPS 375 c. If there exists c> 0 and a > 1 such that Vx,/7), (y, egraph(^), {p-q. x-y)^c\\x-yT then (40) Dorn V^-yA Dorn V = \m{A + dV) = X A We also state the consequence of this theorem for variational inequalities. We recall that (41) b{K)\={p&X I <Tk(/>):=sup (p,x)+co} xeK is the barrier cone of K, whose size measures the lack of boundedness of K. COROLLARY 14 Let A be a strict monotone finitely upper semicontinuous map from a closed convex set K^X to the weakly compact convex subsets of X. Assume that A is P monotone. If (42) 0 6 \пЩК) + A(K) + Dorn p*) then there exists a solution x eK to the variational inequality 0 e >4(х)Н- 7Vk(x).A Assumption (42) shows how the lack of boundedness of К is compensated by the degree of monotonicity of A. We point out that (42) is satisfied when one of the following instance is satisfied: i. К is bounded {b{K) = X). ^ ii. ^ is surjective (A{K) = X). 1 Hi. /1 satisfies (28) (Domj5* = ^) and A{K)n—b{K)^0. iv. A satisfies (29) (Dom P* = X). COROLLARY 15 We posit assumptions (37) of theorem 12. Then the map A-bdV is surjective. Proof of Theorem 12. We set K„\={x 6 Dom l/|K(jc)^«^and \\x\\^n}. The subsets K„ are weakly compact and convex and Dom V = [jn=i a. We set (44) Ф{х, j'):=y(x, </, x-y) + V{x)-V{y)
376 CH. 6, SEC. 6 SOLVING INCLUSIONS This function satisfies (45) and (46) i. e Dom K x-^(l)ixy y) is lower semicontinuous for the finite topology. ii. Vx e Dom K y^^ixy y) is concave and upper semicontinuous (for the weak topology). fi. Vj: [ii, Vx, cDomK y € Dom K <l>(x, j)+4>{y,x)>0 Since K„ is weakly compact and convex, Ky Fan’s inequality for monotone functions (see theorem 3.9) implies that for all 1, there exists x„ e K„ solution to (47) 'iyeK„, (¡)(x„,yX0 b. We shall now use assumption (37) to prove that x„ remains in a weakly compact subset of X. For that purpose, thanks to the uniform boundedness theorem, it is sufficient to prove that (48) 'ipeX*, 3«(p) such that sup (/?, x„)<+oo n^n{p) By assumption (37), there exist >j>0, re Dom q e Dom V*, y e Dom V, u 6 A(y) such that (49) We choose n(p) to be the smallest n such that y e K„. By taking the duality product with x„, we obtain A (p, x„) = <r, x„-y) + (q, x„) + (u, x„-y) - (f, x„-y) + (r+u-f,y) M We use Fenchel’s inequalities <r, x„-y)^ ^(Xn-y)+P*{r) and (q, x„>< V{Xn)+ V*(q). We obtain (50) (P, x„)^(u, x„-y) + V{x„)~ V{y)-</, x„-y) +P{Xn-y)+P*(r)+ V*(q) + F(y) + <r+M -/, j )
CH. 6, SEC. 6 MONOTONE MAPS 377 Since A is P monotone, we deduce that (51) y(x„,y)-{u,x„-y}= inf (p-u,x„-y)^ P{x„-y) P e A(Xn) Therefore, inequality (50) becomes {p, X„)^(y{x„, y)-(f, x„-y) + V{x„)~ V{y)) + P*{r) + V*{q)-y V{y)+ (r + u-f,y) Consequently, for all n>n(p), we deduce from (47) that (52) {p, mr)+ V*(qn V{y)+ {r + u-f,y)) The right-hand side is finite, because r 6 Dom /5*, q e Dom V*, and y e Dom V. Hence, the sequence is bounded and thus weakly relatively compact. c. Therefore, a subsequence of elements converges weakly to some xeX. Since V is lower semicontinuous, we deduce from the monotonicity of A and the variational inequalities (47) that F(.xXlim inf V{x„) n <lim inf [_{V(y)+ (f, x„-y)+y{y, x„))-y(y, x„)-y{x„, y)] n <lim sup {V(y)-\-{f, x„-y)+y(y, X,.)] n < y(y)+ if, x-y)-yy{y, x) Therefore, x 6 Dom V and (53) “ V^eDomK 0^<^(y, x) d. We deduce from properties (45) and (46) that (54) Vz 6 Dom K </>(x, z)<0 as in the proof of theorem 3.9. Indeed, if the conclusion is false, there_would exist z e Dom V such that 0 < <^(x, z) and by (45,1), there would exist t e ]0,1 [ such that (55) 0< (j){x + t{z—x), z)
378 CH. 6, SEC. 7 SOLVING INCLUSIONS By taking y=5c-\-t{z—x), inequality (53) implies that 0^(f){x-\-t{z — x)y x) Hence, the concavity of (f) with respect to the second variable yields (56) 0 < (t){x -f- t{z — x),x+t{z—x)) a contradiction of (46, i). Then proposition 3 implies that the solution x of (54) is a solution to problem (32). ■ Remark We notice that the property (57) A-\-dV IS surjective provides a decomposition akin to the one furnished by the projection theorem onto closed vector subspaces or convex cones. Indeed, any/ can be written (58) i. / eAx-Vp H. <ax) = K(x)+F*(p) since the latter equation expresses that p e dV{5c). Actually, when A is the identity and V is the indicator of a closed cone P, V* is the indicator of the negative polar cone P~ and (58) becomes (59) i. f=x+p, xeP, peP ii. (p, 3c)=0 7. MAXIMAL MONOTONE MAPS We continue to assume that X is a Hilbert space identified with its dual. DEFINITION 1 A monotone set-valued map A is maximal if there is no other monotone set-valued map A whose graph strictly contains the graph of A. A We begin by pointing out the following: A set-valued map is maximal mono¬ tone if and only if its inverse ^4 " Ms maximal monotone. Also, the graph of any monotone set-valued map is contained in the graph of a maximal monotone set-valued map by Zorn’s lemma, because the union of an increasing family of graphs of a monotone set-valued map is the graph of a set-valued monotone map. Actually, we will use the following equivalent ana¬ lytical definition of a maximal monotone set-valued map.
СН. 6, SEC. 7 MAXIMAL MONOTONE MAPS 379 PROPOSITION 2 A necessary and sufficient condition for a set-valued map A to be maximal mono¬ tone is that the property (1) V(j;, y)6graph(/i), {u-v,x-y)>Q be equivalent to (2) и eA(x) ^ This provides a useful and manageable way of recognizing that и belongs to A(x). PROPOSITION 3 Let A be maximal monotone. a. Its images A(x) are closed and convex. b. Its graph is weakly-strongly closed in the sense that if converges to x, and if u„€ A{xn) converges weakly to u, then и e A(x). A Proof, a. By the preceding proposition, A{x) is the intersection of the closed half-spaces {u e X\{u — v, x—y)X)] when {y^ v) ranges over the graph of A. Hence, A(x) is closed and convex. b. Let x„ converge to л: and let u„ e converge weakly to u. Let us choose {y, v) in the graph of A. Inequalities (u„-v,x„-y)>0 imply, by going to the limit, inequalities (u—v, x—y)'^0 Hence, и e A(pc) by proposition 2. ■ PROPSOITION4 Any finitely continuous monotone single-valued map A from X to X is maximal monotone. A Proof. Let л: eX and и eX such that (3) {u —A{y\ x—y)>Q forall;;€X To show that A is maximal monotone, we have to check that u = A(x) (proposi¬ tion 2). For that purpose, we take y=x—X{z — x) where Я e]0, 1[ and z eX. Inequality (3) becomes (4) {u — A(x—^(z — x)\ z — x)^0 for гЛ\ z eX
380 СН. 6, SEC. 7 SOLVING INCLUSIONS By letting A converge to zero and by the continuity of A for the finite topology, we deduce that (u—A(xX z—x)>0ior all z eX, that is, и=A{xX ■ The following theorem provides a very important characterization of maximal monotone maps. THEOREM 5 (MINTY) A monotone map A is maximal if and only if 1+A is surjective. A Before proving Minty’s theorem, we use it to provide examples of maximal monotone maps. PROPOSITION 6 a. Let V be a proper lower semicontinuous, convex function from a Hilbert space X to i?u{+oo}. Then the subdifferential map dV: x e X-^dV(x)<=^X is maximal monotone. b. Let U be a proper function from ХхУго/^и{—oo} such that (5) eYy x-^ U{x, y) is concave and upper semicontinuous. Vx e X, y-^ U{Xy y) is convex and lower semicontinuous. Then the set-valued map (x, y)eX xY-^dx{- U){x, y)xdyU(Xy y) is maximal monotone. A Proof. The first statement follows from theorem 3.3.11, which states that \-\-dV is surjective, and the second statement follows from theorem 3.4.14, stating that the graph oidU is, up to a canonical isomorphism, the graph of the subdifferential 5 К of the convex function defined on X x У* by V(x, q) = sup [ {q, y) - U{x, y)] Ш уеУ PROPOSITION 7 Let A be a monotone {respectively, maximal monotone) set-valued map from X to X. Let si be the set-valued map from L^iO, T,X)tol3{0, T,X) defined by (si x){t):= A{x{t)) a.e. Then si is a monotone (respectively, maximal monotone) set-valued map. A Proof We know that si is monotone. If /i is maximal monotone, then Ji={\+A)~^ is Lipschitz from X to X. Hence, if j(*) is given in L^(0, T, X), the function x(*) defined by x(f)=^j(f) a.e.
CH. 6, SEC. 7 MAXIMAL MONOTONE MAPS 381 is obviously measurable. Also, ||x(f)-yi(0)||<||j(f)|| and thus x(') belongs to l}(0, T; X) and is obviously a solution of (1+^)W-)). ■ Corollary 6.15 also shows that the sum A+dV ofa monotone finitely upper semicontinuous map with weakly compact images and the subdifferential of a lower semicontinuous, convex function is maximal monotone when Dom V c Dom A. Proof of Minty's Theorem, a. Assume that 1 + A is surjective. Let x eX and ueX satisfy (6) V(;;, v) egraph(^), (u-v, x-y)^0 By Proposition 2, we have to prove that u e A(^). Since 1 +^4 is surjective, we can choose y in (6) to be a solution yo of the inclusion u+x eyo+A{yo). Let Vo e A(yo) such that u+x=yo + Vo- Then ll■x-3'oГ = (^->'0, because A is monotone. Hence, x=yo, and thus m=Uq 6 A(yo)=A(x). b. Assume that A is maximal monotone. Let y eX. We have to prove that there exists x such that y ex+A(x). It is sufficient to choose y=0, since this amounts to replacing A by x-* —y+A(x), which is also maximal monotone. By proposition 2, we must prove that there exists x satisfying (7) V(>^, u)6graph(A), (-x-v,x-y)^0 For convenience, set <f>(x; (y, v)):= (x+v, x-y), or, equivalently, <p(x; (y, t))):=||x||^ + (x, v-y)-(v, y) We have to prove that there exists x such that (8) V(y, v) 6 graph(A), (pix; {y, u))<0 Fix (jO) Vo) to be any point in the graph of A. If a solution to (8) exists, it certainly belongs to (9) L:={x 6 Dom(/l)l(/)(x; (j'o. i^o))<0}
382 CH. 6, SEC. 7 SOLVING INCLUSIONS The map {yo, Vq)) is quadratic; hence, the set {xl<^(x; {yo, Vo))^0} is convex, closed, and bounded, hence, weakly compact. Its intersection L with Dom A need not be so, but so is cd(L). Also, the maps {y, i;)) are convex and (strongly) continuous. Therefore, their lower sections are convex and closed, hence, weakly closed. It follows that they are weakly lower semi- continuous. We use theorem 2.6: Set 9" to be the family of all finite subsets K:= {{yu ViX ..., {yny v„)} of graph(y4), then there exists x e co(L) such that sup (/>(x; (y, v)) (y,v)e graph (X) < sup inf max (f>{x; (yi, d,)) Ke y xeco{L) i= (10) < sup inf max (¡){x; (7,-, t;,)) Ke y xeco{yi,..,,y„) i=l n n < sup inf sup X {yj, Vj)) Ke y xeco(yi,...,y„) j=i < sup inf sup ¡x) Ke y Aein neZn <i>KiK M):= s iyj, vj)), m--= E ^iyi j=i i=i where we set (11) This function is continuous with respect to ^ and n i,k = 1 = E ^<'‘^í'‘(P^^í\Уk-Уi)+ E f^‘t^'‘(0hyk-yi} i,fc=l i,k = l The first term is zero for reasons of symmetry, while the second can be written 1 ” 1 ” - ^ fi‘iJ‘(vhyk-yi}+x E ^i,k=l ^ j,i=i 1 ” - X ^í‘^^■'‘(vi-Vk,Уk-Уi)<0 ^ i,k = l since A is monotone. Hence, assumptions of the lopsided minimax theorem are satisfied (see theorem 2.7). Therefore, for all K, inf sup </>k(A,/i)^0 A e Z” /1 e Z”
CH. 6, SEC. 7 MAXIMAL MONOTONE MAPS 383 and thus by (10) inequality (8) is satisfied for 3c. We have shown that 3c is a solu¬ tion of 0 e 3c+y4(3c). ■ Suppose that A is maximal monotone. We show that A can be approximated in some sense by single-valued maps Ax that are also maximal monotone. These maps, called Yosida approximations, play an important role. THEOREM 8 Let A be maximal monotone. Then for all A>0, (12) the résolvant Л=(1 -\-^A) ^ is a nonexpansive single-valued map from X to X and the map Ax '-={1 — /л)Д satisfies (13) 'VjceX, Ax{x)eA(Jxx). [ii. Ax is Lipschitz with constant 1/A and maximal monotone. Let m(A{x)) denote the element of A(x) with smallest norm. We also have (14) VxeDom(y4), ||>4A(x)-m(^(x))|p<||m(^(x))|p-||/4A(x)||^ andfor all x e Dom(/4), (15) JxX converges to x. ii. AxX converges to m{A(x)). DEFINITION 9 The maps Ax are called the Yosida approximations of A. Proof a. Let xt (i=1, 2) be solutions to the inclusions (16) }>i 6 Xi+^A{xi) (i = 1, 2) So yi=Xi + Xvi when vi € A(xi). We obtain Ibl -l'2lP = ll^l-X2 + HVi-V2)\\^ = 11^1 -•^2lP+‘^^l|fi -i^2ll^ + 2A(ui -1)2, --x:2> >\\Xl-X2\\^+^^\\Vi-V2\\^ Hence, (17) i. ||Xi-X2ll<||ji-;)2ll ii. ||t)i-t)2ll<(lM)lbl-V2l|
384 CH. 6, SEC. 7 SOLVING INCLUSIONS By taking yi =y2, (17, i) proves the uniqueness of the solution. We note that (18) Xi = Jxyi and Vi=Ax(yd Hence, inequalities (17) prove that Jx and Ax are Lipschitz with constants 1 and 1/A, respectively. b. By the very definitions of Jx and Ax, we have (19) Ax(y)=j (y- Jxiy)) 6 A(Jxy) for ally eX Therefore, since yi = Jxiyi)+Ax(yd, we obtain (Ax(yi)-Ax(y2), yi -yi) = {Axiyi)-Ax{y2), Jx{yi)~ Jxiyx)) + A\Ax{yi)-AxiyiW >M\Ax{yi)-Axiy2)^\^>0 Hence, Ax is monotone (and by proposition 4, maximal monotone). c. Let X 6 Dom A. We compute \\Ax(.x)-m(A{x)W- = Ma(x)|P + ||m(^(x))||^-2(Ax(x), m(A(x))} = - |U;i(x)|P + ||m(^(x))||^-2<^,(x), m{A{x))-Ax(x)} But since A is monotone, m{A{x)) 6 A{x), and Axix) € A{Jxx\ we obtain (Axix), m(A{x))~ Ax{x)) =j (x-Jx{x), m{A{x))-Ax(x))>0 Therefore, we have proved inequality (20) IMa(x) - m(A(x))\\^ < ||m(yl(x))|p - |Ma(x)|P d. Then when x 6 Dom A, we have lk-/A(x)||=A||^,(x)||<2||m(^(x))|| Therefore, Jx(x) converges to x when X converges to zero. e. Since =x-XAa(x) and Aa(x) € A(Jxx), we see that j=/1 a(x) is a solution to the equation y e A{x — Xy). Conversely, any such solution y is Ax(x). Indeed, set z=x—Xy; this equation becomes x ez+XA{z). Hence, z = Jx(x) and y=(l/X)(x-Jx(x))=Ax(x) f. This remark implies that A^+x(x)=(AMx)
CH. 6, SEC. 7 MAXIMAL MONOTONE MAPS 385 Indeed, y=A^+x(x) is a solution to the equation y eA(x—Xy—fiy)-, then y e A^(x—Xy). By again applying the preceding remark to the Yosida approxima¬ tion which is maximal monotone, we deduce that y={A^)x{x). g. Now we use inequality (20), replacing A by Af,. Since m{A^{x))=A^{x), we obtain Then the sequence ||y<^(x:)||^ is an increasing sequence of real numbers bounded above by |lm(y4(j))||^. Hence, it converges to some real number a. This implies that lim ||.4^+;i(x)-.4p(x)|p<a-a=0 Hence, A;((x) satisfies the Cauchy criterion and converges to some element v in X, Since Ax{x) e A{Jx{x)) and the graph of A is closed, we deduce that v 6 A(x). Also, ||i;|| = lim |M,W||^|W^W)|| Since A{x) is closed and convex, the projection of zero to A{x) is unique and, consequently, v=m(A{x)). Therefore, Ax{x) converges to m{A{x)) for all X € Dom(y4). ■ We now provide a handy sufficient analytical condition for an open subset to be included in the image of a maximal monotone map. This condition motivates the introduction of a property that is satisfied, for instance, by the subdifferentials of convex functions. When two monotone maps satisfy this property and the sum is maximal monotone, then we obtain the following inclusion: Int(Im A-\-lm 5)c:Im ^ + Im B which allows us to solve inclusions of the form y e Ax-\-Bx as well as inclusions of the form ;; ex-\-ABx We recall solving such problems in Section 6, where B was the subdifferential of a lower semicontinuous convex function. We begin with the fundamental result stated in theorem 10.
386 СН. 6, SEC. 7 SOLVING INCLUSIONS THEOREM 10 Let A be a maximal monotone map and K<=-X a subset satining (21) VueK, 3yeX and ceR inf (u—v,x—y)>c (x,m) e graph(.4) Then (22) co(K)<=Im^ and Int co(K)c:Int(Im/1) Proof, a. First, we check that we can replaced by co(K) in property (21). Indeed, let v=Y!!= j be a convex combination of elements a,- e K. By assump¬ tion (21), there exist yi and Ci e R such that for all {x, u) e graph(y4), {U-Vi, X-yi)>Ci Let us set j ^¡yi. We deduce that Hence, (23) (u, x) - (v, x) - (u, y> ^ E Uci - (vi, ;^i>) i=l inf (u-v,x-y)> Y - (vi,;^i>)-b(v,y) (x.M)e graph(^) f=i b. We now assume that K is convex and prove that 1C <= Im Indeed, since A is maximal monotone, we can associate with any e>0 and any u e K the solu¬ tion Xc to the inclusion v eexc + A{xe). Let y eX and c eR such that property (21) holds true. By taking (x« a —exj) 6 graph(/4), we obtain (-ex„ Xe-y)>c and thus — c Hence, \fexe remains in a bounded set rB. Consequently, v eA{x^)-\-\fe-Jlxt<=' \m A-¥\ferBfor all e>0 which shows that v eIm A. c. We still assume that K is convex and prove that Int Kdnt Im A. Let u € Int iC and y >0 such that u-l-yBc:K^. By assumption, for all zbX, v+yzl\\z\\ belongs to K, and there existy ^X and c €/? such that (24) V(x, m) 6 graph(^). u-v-
CH. 6, SEC. 7 MAXIMAL MONOTONE MAPS 387 Let Xs be the solution of the inclusion v e sxe+A{xe). By taking (xg, v—sxe) € graph(>4), we obtain yz and thus r IWI^+/ï7 <z, IMI^+n^ {z,y)-c Therefore, for all zeX, (25) sup (z, Xg)<4-00 The uniform boundedness theorem implies that Xg remains in a weakly compact subset. Then a subsequence Xg' converges weakly to some x^. Since v — s'xe> e A(xe') and v—s'xe> converges strongly to d, proposition 3 implies that v e A{x^)^ Im A. ■ We deduce the following interesting consequence. PROPOSITION 11 If A is a maximal monotone map, then (26) i. Dom A and Int Dom A are convex. ii. Im A and Int Im A are convex. Proof. We take K = lm A. Since A is monotone, property (21) is obviously satisfied. Then co(Im A) dm A and Int co Im A^lni Im A by theorem 10. This implies that Im A and Int Im A are convex. Since /1 " Ms maximal monotone and Dom A =Im ^4“^ we deduce that Dom A and Int Dom A are convex. ■ We also deduce a surjectivity criterion. THEOREM 12 Let Abe a maximal monotone map. We assume that either (27) or that (28) Dom A is bounded <. irifw e A(x) -X) i A • * 1 ' \ lim = -foo (A IS strongly coercive.) llxll->oo X Then A is surjective.
388 CH. 6, SEC. 7 SOLVING INCLUSIONS Proof. We use theorem 10 with K = X. First, we replace A by A defined by A(x):=A(x + xo) where xq € Dom A, so that 0 e Dom A. The map A is still maximal monotone, its domain is bounded when Dom A is bounded, and A is strongly coercive when A is strongly coercive. Let veX. We check property (21) when;;=0. a. Assume that Dom A^rB. We take z e /1(0). Then when (x, u) e graph(/l), we have (29) <«-1;, x-O) = z, x-0> H- <z, x) ^ <2, x) ^ - ||z|| ||x|| ^ - r||z|| Hence, property (21) is satisfied with j;=0 and c= “H|z||. b. Assume that A is coercive. There exists r>0 such that, for all x^e Dom A, ||x||^r, for all u6A(x), (u, x)^||i;|| ||x||. Then when (x, m) egraph(^), ||x||^r, we obtain (30) {u-v,x-Q)>\\v\\ ||x||-<t;, x>^0 By taking into account inequality (29) when ||x||^r, we see that property (21) is satisfied whenj;=0 and c = — r||z||. c. In both cases, we have X = lnt co(X)czIm A, that is, A is surjective or, equivalently, A is surjective. ■ We now turn our attention to the surjectivity properties of the sum of two maximal monotone maps. For that purpose, we have to provide conditions under which the sum is still a maximal monotone map and to assume that one of these maps satisfies another property, which is satisfied, for instance, by subdifferentials of proper lower semicontinuous, convex functions. We begin by giving a sufficient condition for the sum of two maximal mono¬ tone maps to be maximal monotone. The sum of two monotone maps is mono¬ tone, whenever it is proper, that is, whenever 0 e Dom A — Dom B. We shall prove that a stronger condition, 0 6 Int(Dom A — Dom B) implies that A-\-B is maximal monotone. We have already encountered similar conditions in particular cases. When A=dV and B=dW 2irQ subdifferentials of lower semicontinuous, convex functions V and W, we proved that the weaker condition 0 6 Int(Dom V— Dom W) implies that dV -\-dW=d(V -}-W) and that thus dV-^dW is maximal monotone.
СН. 6, SEC. 7 MAXIMAL MONOTONE MAPS 389 THEOREM 13 Let A and В be two maximal monotone mapsfrom X to X. If (31) ОбInt(Dom A — Dom B) then AВ is a maximal monotone map. A Proof Since A^-B \s monotone, we must prove that 1 + ^4 -f 5 is surjective. For that, we shall approximate one of the maps, B, for instance, by its Yosida approximation 5я, which is maximal monotone and Lipschitz. We shall prove that A-\-Bx '\s maximal monotone and thus there exists a solution хд € Y to the inclusion (32) у exx + Axx + BxXx In the second step, we shall deduce from assumption (31) that (33) sup ||^я^я||< +00 A>0 In the third and last step, we shall deduce that the solutions xx of (32) do con¬ verge to the solution x of (34) у ex-\-Ax^-Bx a. There exists a solution xx to (32). We have to prove that A-\-Bx '\^ maximal monotone, that is, 1 -\-p(A-\-B}) is surjective for some /^>0. If;^ 6 Y is given, we have to find xx, solution to (35) Xx = {i-\-pA) \y-pBxXx) so that Xx is a fixed point of the map x->(l +A)~^{y — pBxx). Since (1 -\-pA)~^ is Lipschitz with constant 1 and x-^y — pBxXx is Lipschitz with constant /i/A, then {l-\-pA)~^{y—pBx{^)) is a contradiction when Therefore, there exists a fixed point xx of (35) and, consequently, a solution to (32). b. Let us prove estimate (33). We begin by proving that (36) sup ||хя||< +00 A>0 We choose zeDom .4 —Dom B. Since y-xx^(A-\-Bj)(xx) by (32) and m{Az)+Bxz € (A-\-Bx)(z\ we deduce that (y-xx-m{Az)-BxZ, Xx-z}={y-{z+m{Az)+Bxz), xx-z)-\\xx-z\\^>0
390 CH. 6, SEC. 7 SOLVING INCLUSIONS Hence, \\xx-z\\^\\y-{z + m(Az)+Bxz)\\ Since ||5A(z)||<|lm(5z)ll by theorem 8, solutions xx to (32) remain bounded (37) ||;ca|| <2||z|| + II jII + ||m(^z)|| + ||m(Bz)|| By assumption (31), there exists y>0 such that yBfx Dom A — Dom B We can associate to every p&X elements u € Dom A and v e Dom B such that ypl\\p\\=v-u- Then (38) {p, BxXx) = (v, BxXx) - <M, BxXx) The monotonicity of Bx implies that {BxXx-Bxv, xx-v)^Q and thus {BxXx, v)< {BxXx, Xx) - {Bxv, Xx~v)^ {BxXx, xx) + ||m(Bt;)|| ||xa-n|| Hence, (39) jP {p, BxXx) < {BxXx, XA-M) + ||m(Bn)|| ||xa-u|| Since Xx is a solution to inclusion (32), we can write that BxXx=y — xx — zx when zx e A(xx)- So, inequality (39) becomes y IP {p, BxXx)^{y-Xx-zx, XA-M) + ||m(Bu)|| HxA-t^ll But the monotonicity of A implies that {zx-m{Au), xx-u)>0 and thus - {zx, xx-u)< {m{Au), xx-u)^ l|m(.4M)|| Hxa-m||
CH. 6, SEC. 7 MAXIMAL MONOTONE MAPS 391 Finally, (p, (1|xa-m||(||>'-x,i|| + ||m(^M)|l)+ ||m(5p)|| ||xa-p||)< +oo Consequently, for allpeX, supA>o (p> BxXx) < +oo. The uniform boundedness theorem implies that BxXx remains bounded. c. We prove now that xx converges to a solution x to inclusion (34). For that purpose, we prove that the sequence xx satisfies the Cauchy criterion. We set ^3.'=y—Xx-BxXx 6 A{xx) and Zf,:=y—x^~B^x,, e A(x^) Since 0=xx+Zj+BxXx—x^—z^ — B^Zf„ by taking the scalar product with xx—x^, we obtain 0 = \\Xx -X^\\^ + (zx-Z^, Xx-X^)+ (BxXx - Bf,Xn, xx - x„) ^Wxx-x^W^ + iBxXx-B^x^, Xx-x„) (because A is monotone). Now, we set Jx.={\^B)~ ^. By definition of the Yosida approximation, Xx Xft — ^BxXx pBfiXft "i" JxXx JttX^ We recall that BxXx e B{JxXx) and B^x,, e B(J^tX^. Since B is monotone, we deduce 0>\\xx-x^\\^ + (BxXx-B^x^, ^.Bxxx-pB^x„) and, consequently, \\xx-< IIBaXa- B^X,\\ \\XBxXx-pB^^x^W Since ||Ba.xa||<c and by (33), we obtain (40) ||xa-< 2c^(A + p) converges to zero. Hence, Xx converges strongly to some x*. Since BxXx is bounded, a subsequence Bx Xx- converges weakly to some v^, and thus Zx=yx-Xx-BxXx converges weakly to some =y—x^ — u*. Proposition 3 implies that €y4(x„,). Also, xx~JxXx=^BxXx implies that ||;c^_7^xa||<A||BaXa||=SAc converges to zero, so that JxXx converges strongly
392 СН. 6, SEC. 7 SOLVING INCLUSIONS to л:*. Inclusions Ддосд- € B(Jx-Xx ) imply that a* e B(x^\ because В is maximal monotone. Hence, y=x^ + u^ + v^ex^ + A(x^) + B(x^,) Ш We now introduce a property that we need for studying the image of the sum of two maximal monotone maps. DEFINITION 14 Let A be a monotone map from X to X. It satisfies the L property if (41) fVw6lm/l, V^eDom/l, ^c:=c{y,w) such that I inf (x,ii) e graph(/i) inf {u—w,x—y)'^c. A e graphic) We begin by giving two examples of monotone maps satisfying the L property. PROPOSITION 15 The subdifferential of a proper convex function satisfies the L property. A Proof Let u 6 dV{xX v e dV{y\ and w e 5F(z)c: Im 5F be chosen. We deduce from the inequalities V{x)- V(yH (uy x-y), V{y)-L(z)^ (vyy-z) and V(z)— (w, z—x) that (42) Then (m, x-y)-\'{vyy — z)-}-{w, x — x)^0 <w-w, x-y)^{w-v,y-z) The L property is satisfied by the constant c:=(w — Vyy — z) that depends only on y and w. ■ Remark Actually, any monotone map satisfying inequality (42) satisfies the L property. Another example is provided by the following proposition. PROPOSITION 16 Let Abe a monotone map such that either (43) Dom A is bounded
CH. 6, SEC. 7 MAXIMAL MONOTONE MAPS 393 or (44) • I. € /1(X) {w, x} I. hm ^ =+00 llxll-oo ||x|| {A is strongly coercive). I .. sup (A is bounded). (x, m) egraph(i4) Then A satisfies the L property. A Proof. Let y 6 Dom A and w e Im /4 be fixed. a. If Dom A czrB is bounded, we take v e A(y); the monotonicity of A implies that for any (x, u) e graph(.4), (45) {u-w,x-y)>{u-v, x-;^> + <u-w, x-;^>^0-l|i;-w||(||j|| + r) So the L property holds true with c:= — ||y—w||(||y|| + r) b. Assume that A is strongly coercive and bounded. There exists r>0 so large that for all ||x;||>r, there exists ueA{x) such that (m, x)>(||w|| + c||j||)||x|| and ||m||<c||x||. On the other hand, (w, x)> — ||w|| ||x|| and — (m, y)> — ||m|1 \\y\\> -dWIlWl.Then <M-w, JT-y)>(||w|| + c|WI)||x|| -||w|l ||x|| -c||x|| \\y\\ + <w, y) = <w, y) This inequality, together with inequality (45), implies the L property with c=min«w,7>,-l|t)-w||(||y|| + r)) ■ The introduction of the L property is justified by the two following results. THEOREM 17 Let A and B be two monotone maps satisfying (46) Then i. Dom A <= Dom B ii. A-yB is maximal monotone. iii. B satisfies the L property. (47) Int Im(y4 + 5)=lnt(Im(.4)+Im B) ii. \m(A + B)=\m A+ \m B Proof. We know that lm(A + 5) c Im /4 + Im 5. We apply theorem 10 to the maximal monotone map A + B and to K :=Im /1 + Im B. For that purpose, we
394 CH. 6, SEC. 7 SOLVING INCLUSIONS have to check that (48) >jv elm By 'iyeXy 3c gsuch that inf (u — v,x—y)>c (x,u)e graph(/l + B) holds true. Take v=Vi where € Im ^4 and V2 e Im B, Ltiy e/l"^(i;i)cDom Dom By Since A is monotone, for all x g Dom A and ui g A{x), we have (49) (Ui-VuX-y)^O Since B satisfies the L property, there exists ceR such that for all x € Dom A < Dom B and «2 ^ B{x), (50) (u2-V2,X-y)^C By adding these two inequalities, we obtain (48), and thus, Im ^ + Im 5<= lm{A + B) and Int(Im ^ + Im 5)<= Im /1 + Im B thanks to theorem 10. ■ THEOREM 18 Let A and B be two monotone maps satisfying the L property. If A-y Bis maximal monotone, then (47) i. Int lm{A + B)=Int(Im A-y\mB) ii. Im(/l + jB)=Im/1 + Im B Proof We proceed in a manner analogous to the proof of theorem 17. We apply theorem 10 for the maximal monotone map A + B with K = \m A-y\m B. We take yo e Dom A n Dom B and u=i;i + V2, where vi elm A and V2 e Im B. Since A and B satisfy the L property, there exist constants C\ and C2 such that i i. Vx € Dom r> Dom B, 'fu\eA{x), {u\ — Vx,x—yo)'?'C\ |ii. Vx 6 Dom A n Dom B, V1/2 £ B{x), («2 - V2, x-yo)> C2 By adding these two inequalities, we find that {v;: 6lm..4 + ImB, ByQeX, 3c:=Ci+C2 e/? such that M)€graph(/1 + B), (u-v,x-yo)>c Hence, by theorem 10, we obtain the inclusions Im ^ + Im B<^\m(A + B) Int(Im y4 + Im B) <= lm{A + B) from which our theorem ensues.
CH. 6, SEC. 7 MAXIMAL MONOTONE MAPS 395 We deduce a result on the surjectivity of maps of the form THEOREM 19 Let A and B be two maximal monotone maps satisfying (53) i. 0 e Int(Dom A — lmB) ii. ;; € Int(Im A + Dom B) Then either the assumptions (54) A satisfies the L property and Im Dom A or the assumption (55) B satisfies the L property and Im ^4 <= Dom B-\-y imply the existence of a solution x to the inclusion (56) yex-\-ABx A Proof a. We assume (54). Inclusion (56) can be written in the form y eB~^u-\-Au, where xeB~^u, We have Dom B~^=lm 5<=Dom A, Assump¬ tion (53, i) can be written 0 € Int(Dom —Dom B~^) and theorem 13 implies that A + B~^ is maximal monotone. Then theorem 16 states that Int(Im y4-hDom ^) = Int(Im yl + Im B~^)^\m{A-\-B~^) so that assumption (53, ii) implies that there exists a solution uXoy eAu-VB~^u. Hence, there exists xeB~ ^(u\ which is a solution to the inclusion (56). b. We assume (55). Inclusion (56) can be written in the form (57) QeBx-\-A{x\ where y4(x):= Then ^ is a maximal monotone map, Dom A =y — Im Ay and Im ^4 = — Dom A We have Dom ^cDom B, Assumption (53, ii) implies that OeInt(Dom A — Dom B)y and theorem 13 implies that A-\-B \^ maximal monotone. Theorem 17 states that Int(Im Dom y4) = Int(Im ^-i-Im A)^\m(A^-B) Hence, condition (53, i) implies that there exists a solution x to inclusion (57), that is, to inclusion (56). ■
396 сн. 6, SEC. 8 SOLVING INCLUSIONS 8. EXISTENCE AND UNIQUENESS OF SOLUTIONS TO DIFFERENTIAL INCLUSIONS THEOREM 1 Let A be a maximal monotone set-valued map from X to X. Consider the initial value problem for the differential inclusion (1) xe—A{x\ x(0)=xo when Xo is given in Dom A. Then there exists a unique solution x(*) defined on [0, 00 ], which is the slow solution. Let m{K) denote the projection of zero onto a closed convex set K; a slow solution is a solution to the differential equation (2) for almost all i ^ 0, x'{t) = — m(A{x(t))) Moreover^ (3) t^\W{t)\\ is nonincreasing. Let ) andy{^) be solutions starting at Xq andyo- Then (4) Vt>o, IWi)-j(f)ll<lko-;^oll Finally, (5) Vi ^ 0, x'{t) = lim —and x'{ •) is continuous from the right. h-^0+ h Proof. The proof is based on defining approximate solutions as solutions to ordinary differential equations where A is replaced by the Yosida approxima¬ tions of A. To show the convergence of a sequence of approximate solutions, we shall prove that they are a Cauchy sequence. Properties (3) and (4) have been pointed out explicitely in the statement of the theorem; property (3) would give a bound on ||y4(x(i))|| if we knew we had a solution x(-), yields the same bound on the set of approximate solutions. Property (4), which shows that the solution is unique, also implies that the distance between two approxi¬ mate solutions is small, that is, that they are a Cauchy sequence. We shall divide the proof into the following steps: Step a. We prove (4) [i.e., that the map xo-^x(') is nonexpansive] and derive uniqueness. Step b. We prove that i->||V(i)|| is nonincreasing. Step c. We consider the solutions xa(') to x'x(t) A xxx{t); xx(0) = X
CH. 6, SEC. 8 EXISTENCE AND UNIQUENESS OF SOLUTIONS 397 where is the Yosida approximation of A and prove that xa(') is a Cauchy sequence that converges to some x(*) in ‘^(0, oo, X). Step d. We check that x(i) e Dom A for all t>0. Step e. We prove that x'(f) 6 —Ax{t) for almost all t>0. Step f. We show that t-ym{A(x{t))) is continuous from the right. Step g. We conclude by establishing that x'(t)= —m{A{x{t))) [i.e., that x(*) is the slow solution] and x'(‘)=dldt x(t) is the derivative from the right. a. Assume that x(*) and ) are solutions of the initial value problems I i. x'{t)e-Ax{t), x(0)=xo 1«. y'it)e-Ay{t\ j(0)=yo Therefore, since A is monotone, ii;c(i)-7(i)ip=<x'(i)-y(f), x{t)-m<o and by integrating from zero to i, we deduce that supllxW-yOll^lIxo-joll t>0 In particular, when xq =yo> this implies that the solution of (1), if any, is unique. b. Consider now 7(t)=x(t + /i) where h>0 and x(*) is a solution of (6, i). It satisfies y'{t) e - Ay(t), y(0)=x(h) Therefore, (3) implies that (7) ||x(i+/i)-x(i)||<|W/i)-x(0)|| By dividing by/i>0 and letting h tend to zero, we deduce that ||x'(i)||<||x'(0)|| foralli^O c. We consider solutions xx to the approximate problems (8) x'x(t)=-Axx(t), xa(0)=Xo Since Ax is a Lipschitz map from X into X (see theorem 7.8), it has a unique, continuously differentiable solution xa( •) defined on [0, oo [. We shall prove that x,i(*) is a Cauchy sequence in <^(0, oo; Y) and check that xa(’) converges to the solution to the differential inclusion (1).
398 CH. 6, SEC. 8 SOLVING INCLUSIONS Since Ax is also monotone, we deduce from (3) [with x(*) replaced by xa(*)] that (9) IU-iXa(î)|| =iui(f)|| ^\\xm\ = WAxXoW < l|m^(xo)|| We compute i||xA(i)-:>C;,(i)||^ =a(t) l|;c,(f)-x,(i)l|^ = II ^ \\xxi^)-x,(rrdr = - (AxXx{‘t)-Ap,x^(t), xx{t)-x^(x)>dx We now use the relation l — Jx=Ax, the fact that Axx eAJx{x), and the mono¬ tonicity of A a(t)= I (AxXxi'^)-A^x^(r), XAxXxi'^)-dA^x^{t))dT - j (Axxx(r)-A^x^(x), JxXx{t)-J^x^{'^))d-c < - j (AxXx('^)-A^x^(t), XAxXx{T:)-I^A^x^(x))dr = [ X(AxXx{x), A^x^{x))dr+ f ti(Af,x/r), AxXx(x))dx Jo Jo -i X\\AxXx{x)\\^-\-n\\A^x^{x)\\^dx We note that X{AxXx{x), IM^x^(t)||<A||/Î;iXa{t)||^ -I-^\\A^x^(x)\\^ and in the same way KA^X^ix), ^AXA(T)><^|M^Xp(T)|p-f^ \\AxXx(xW Therefore, using this and (9), we obtain <^^l|m(^(^o))ll
CH. 6, SEC. 8 EXISTENCE AND UNIQUENESS OF SOLUTIONS 399 Hence, xa(*) is a Cauchy sequence in <^(0, oo; X) and thus converges to some continuous function x(*) uniformly over compact intervals. d. The inequality l|.XA(i) - /AA:A(i)|| = A||^AA:A(f)|| < A||w(^(xo))|| yields that JxXxi’) converges to jc(*) uniformly. Also, since ||yiAA:A(i)||<||/n(/i(xo))||, there exists a subsequence Ax„xx„{t) that converges weakly in X to some a(i). But Ax„xx„{t) belongs to A(Jx„xx„it))', there¬ fore, we deduce from proposition 7.3 that v(t) e A(x(t)). In particular this implies that x(t) 6 Dom A for all t^O. e. We note that x'x(t) remains in a bounded set of L°°(0, oo; X) and thus a sub¬ sequence x'x„(') converges weakly to some function, which is equal almost everywhere to x'(*). So, for any T>0, x’x„(') converges weakly to x'(*) in l}(0, T ; X), and x„(-) converges strongly to x(*) in I?(0, T; x). Proposition 7.7 implies that the map x( • )^( •)=-^ W •)) is maximal monotone in L^(0, T;X) ; we deduce that x'(-) e -Ax{-) in l3{0, T;X). Hence, x'{t)eAx{t) almost every¬ where. f. Let t > to- Since ||x'A(i)||=|MA;CA(f)||^||m(/l(x(io)))|| we deduce that IbWNlim inf||x'A(i)||<||m(^Wto)))ll A->0 and the solutions xa(') are uniformly Lipschitz. Hence, x(*) is also Lipschitz. Moreover, ||m(/4(x(i)))||=^||t)(i)||<||m(/4(x(fo)))||. This proves that f-> ||m(/i(x(i)))|| is nonincreasing. g. Let us check that t->m(^(x(i))) is continuous from the right. Let converge to io+; then x(t„) converges to x(to), and since ||m(/l(x(i„)))|| < ||t>(f„)ll<||m(ylx(io))||, we deduce that some subsequence (again denoted) m(/4(x(i„))) converges weakly to some y in X. Proposition 7.3 implies that y e A(x(to)). But II <lim inf ||m(^(x(t„)))|| <||m(^{x(fo)))|| tn-^oo Hence, y=m(A(x(to))X and m(A(x(to)) is the weak limit of the sequence m(A(x(t„))). Since ||m(^(x(io)))|| = lim ||m(^(x(i„)))|| we deduce that m(A(x(t„))) converges strongly to »i(/l(x(io))) when Hence, t-^m(A(x(t))) is continuous from the right.
400 CH, 6, SEC. 8 SOLVING INCLUSIONS h. Let N be the subset of [0, oo[ where neither x(-) is differentiable nor x'(i) i -A(x(t)). Let to i N. We deduce from (7) (where zero is replaced by to) that \\x{t+h)- x(i)| I < I |x(i 0+/i) - x(to)| I But \\x{to+h)-x{to)\\^ i*° ||x'(t)||</t< [ ||m(^(x(to)))||dT</i||m(^(x(io)))|| Jto ^0 (By (9)) Hence, \\x'(to)\\= lim fi-»0+ 4fo + A)-4io) ;||m(^(x(to)))ll Since x'((o)6 -Ax(to), we deduce that x'(io)=-m(^(4io)))- Integrating from to to to+h, we deduce that x(to-t-h)—x(to) 2 ” ^ 1 m(A(x(r)))cix Since m(/4(x(-))) is continuous from the right, we deduce that dt x(to)= -mA(x(to))
CHAPTER 7 Nonsmooth Analysis Nonlinear analysis must provide sufficient conditions for solving inclusions (*) ysF(x) when F is a set-valued map from a Banach space AT to a Banach space Y. Our principal objective in this chapter is to prove an inverse function theorem for set-valued maps allowing us to say that when xq is a solution to yoeFixo) then there exist neighborhoods U of Xq and V of j^o such that inclusion (*) has solutions in U whenever y ranges over K Furthermore, as in the smooth case, we require that the set of solutions F~^{y)nU of (*) depends in a Lipschitz manner on the data;;. Since the inverse function theorem for usual smooth maps plays such an important role in solving many problems of pure and applied analysis, we can expect an adaptation of the inverse function theorem to be very useful, more than just a generalization that had to be made. Actually, it can be used for solving convex minimization problems and proving the Lipschitz behavior of its solutions when the natural parameters vary. Economists claim that this problem is of utmost importance in their field (marginal theory). We shall take our inspiration from the smooth situation where the sufficient condition is very simply stated: The derivative at Xq must be surjective. The question arises, can we define derivatives of set-valued maps such that the sur¬ jectivity of the derivative at {xq, yo) is sufficient for solving the surjectivity of F around yo^ The answer to this question is one purpose for this chapter. We now explain how we shall proceed to define derivatives of set-valued maps. We adopt the very first strategy, apparently suggested by Fermat, which defines the graph of the derivative to a smooth function as the tangent to the graph of this function. Therefore, we postpone questions about derivatives until after having tackled the matter of tangent spaces to subsets K of a Banach space X. They do not exist when K is no longer a smooth manifold. However, it is known in convex analysis that we can define in a natural way “tangent cones” to convex sets which retain enough properties of tangent spaces to be quite useful. This is r
402 CH. 7, SEC. 1 NONSMOOTH ANALYSIS enough, because most of the set-valued maps we shall meet have nonconvex graphs. When K is neither smooth nor convex, there are many ways of defining tangent cones, each one being as “natural” as the other. We shall retain only two concepts among the many candidates: the contingent cone and the tangent cone. Namely, they are defined in the following way: Let xq belong to K, The contingent cone, defined by 7ic(A:o):= n H U (^7; + e>0 a>0 /ie]0,a] / was introduced by Bouligand in the early thirties, and the tangent cone defined by Ci:(-^o)’= n U n e> 0 a,p >0 he ]0,a] jc e Bjc(xo,P) ^{K-x)+eB was introduced by F. H. Clarke in 1975. We see at once that the tangent cone Ck(x) is contained in the contingent cone Tjcix). They are both closed, and the tangent cone is always convex. We can say that they form a kind of “dipole,” in the sense that the tangent cone Ck(xo) is the Kuratowski lim inf of the contingent cones Tjc{x) when x-^xq (when the space X is finite dimensional). So, several properties of the tangent cone Ck{xo) at xo “diffuse” to generally weaker properties of the contingent cones Tk(x) in a neighborhood of xq. This dipole collapses to the usual tangent cone of K or the tangent space of K when K is convex and a smooth manifold, respec¬ tively. We shall see in the sixth section that the contingent and tangent cones enjoy dual properties. These are some of the reasons for studying both contingent and tangent cones and treating them as a pair rather than individuals. Let Fbe a set-valued map from A" to 7 and (xq, yo) belong to its graph. We define the contingent derivative DF{xo, j^o) as the closed process from X to y whose graph is the contingent cone to the graph of F Vo eDF{xo, 7o)(wo)^(mo, Vq) e Tgraph(F)(xo, yo) and the derivative CF(xo, yo) as the closed convex process from X io Y whose graph is the tangent cone to the graph of F Vq G CF{X0i yo)iPo)^iM09 1^0) ^ ^graph(f)(-^0j Yo) They may be different. Consider, for instance, the Lipschitz single-valued map n from IR to R defined by 7i(x)=0 when x^O, n(x)=x when x>0
CH. 7, SEC. 1 NONSMOOTH ANALYSIS 403 Then the contingent derivative of n at (0,0) is defined by Dn{0, 0)(w)=0 when w ^ 0, Dn(0, 0)(w)=u when w ^ 0 and the derivative of n at (0, 0) is defined by Cn{O,O)iu)=0 when Cti(0, 0)(0)=0 This example shows that the price to pay for having a closed convex process as a derivative is sometimes too high. These definitions provide intrinsic definitions of derivatives of single-valued maps defined on subsets K that may have an empty interior, as well as formulas for computing them when they are restrictions to of a smooth map. When Fis continuously differentiable on an open neighborhood of K, then the contingent derivatives and derivatives of the restriction jF]/c of F to K are the restrictions of the Jacobian VFof Fto the contingent and tangent cones, respectively ii. C(A1k)(xo, F(xo)) = VF(a:o)|c/^(;c^) Also, these concepts of derivatives allow us to compute the inverse of the derivative of a map, in particular, the inverse of the Jacobian of a single-valued map, because we infer immediately from the definitions that i. DF(xo, yo)~ ^ =D{F~ ^){yo, Xq) ii. CF(xo, >^o)" ^ = C(F" ^o) Since the derivative is a closed convex process, it is useful to distinguish its transpose CF(xo, yo)*, a closed convex process from Y* to A"*: We shall call it the codifferential of F at (jcq, j^o)- For real-valued functions, we can take into account the order relation, which is used in optimization problems (or in the theory of Lyapunov functions, i.e., functions that decrease along the trajectories of a dynamical system).
404 СН. 7, SEC. 1 NONSMOOTH ANALYSIS We associate with a proper function V from A" to (R u {H-oo} the set-valued map V+ defined by V+(a:):=F(jc)-1-IR+ when К(л:)<+сх), У+(x):=0 when V(x)= +00. We observe that there are numbers D+ V(xq)(uo) and C+ K(xo)(wo) such that i. £)V+(xo, ^^(a:o))(wo)=^+^Cvo)(wo) + IR + H. CV+{xo, V(xo))(wo) = C+V(xoKwo) + ^ + where - 00 ^ Z) + К (xo)(wo) ^C+V (xo)(wo) ^ + oo We shall say that the functions D+ F(xo)(*) and C+ K(xo)(*) from A" to {—oo} u (Ru{ +00} are the epicontingent derivative and epiderivative of the function V. In other words, the epigraphs of D+V{xo)(*) and C+V{xo)i*) are the con¬ tingent and tangent cones at the epigraph of V at (л:о, У{хо)У When the derivative C+ Vixo){^) is a proper function from A" to (R u {+oo}, it is convex, positively homogeneous, and lower semicontinuous. This is, then, the support function of the closed convex subset dV{xo):={p e Х*\Уи e X, (p, V{xo){u)} We shall call this subset the generalized gradient introduced by F. H. Clarke in 1975. Indeed, the terminology is justified by the fact that when V is con¬ tinuously differentiable at Xo, then 5K(xo)={VK(xo)}. We also observe that when V is convex, the generalized gradient coincides with the subdifferential dV(xo) of convex analysis. It is then natural to consider the derivatives of the set-valued map x-^dV{x) as candidates for the role of second derivatives. Let po belong to dV(xo)', the derivative CdV(xo,po)'=d^y (-^0, Po) is a closed convex process from X to X*, which is monotone when V is convex. These tangent cones and derivatives enjoy enough properties to make a decent calculus. But the main justification for including this study here is their use in the inverse function theorem. When X and Y are finite dimensional, it has a very simple formulation. Let F be a set-valued map with a closed graph and let (xo, Уо) belong to the graph of F. Assume that the derivative CF{xq, уо) of F at (xo, д^о) ^ surjective. Then F ^ is pseudo Lipschitz around (xo, yo) in the sense that there exist a neighborhood W of yo, two neighborhoods U and V of Xo, i7<= K ctnd a constant
CH. 7, SEC. 1 CONTINGENT AND TANGENT CONES 405 ^ > 0 such that i. 'iye W, F \y)r\U^0 «. ^yu yi e W, d(F~^(y,)n U, F-Hy2)nV))^€\\y,-y2\\ where d(A, B):=sup^ e a infy e b d{x, y). It is itself a consequence of a more general inverse function theorem, valid in infinite dimensional spaces and involving surjectivity properties of the con¬ tingent derivative of F not only at (xo, yo), but at all neighboring points. We conclude this chapter with a section devoted to the calculus of tangent cones, derivatives of set-valued maps, and epiderivatives of real-valued functions. 1. CONTINGENT AND TANGENT CONES o Let X be a nonempty subset of a Banach space X. We denote by sB and sB the ball (respectively, open ball) of center zero and radius e > 0. We set Bidxo, e) := Xn(xo + £B), and the symbol x^xo denotes the convergence of x to a:o in K. DEFINITION 1 We say that the subset (1) M:=n n U £>0 «>0 0</i<a / is the contingent cone to K at x, ^ In other words, v e Tk(x) if and only if (2) Vs>0, Va>0, ^uev-ysB, 3/ie]0, a] such that x + e K or, equivalently, v e Tk(x) if and only if there exist sequences of strictly positive numbers h„ and elements u„eX satisfying (3) !• lil^w-»oo^w — ii. lim„_«,/i„=0. iii. V« ^ 0, X 4- h„u„ e K. We characterize the contingent cone by using the distance function ¿k(*) to K defined by i//^(x):=inf{||x->^|| \ y^K} (4) V e Tk{x) if and only if lim inf =0 h^o+ h
406 CH. 7, SEC. 1 NONSMOOTH ANALYSIS It is quite obvious that the contingent cone is a closed cone, which is trivial when X belongs to the interior of K; (5) when X e lni{K\ then Tk{x)=X. For all xeX,sNQ have Tx(x) = X. We set Тф(х)\=0. It is convenient to introduce the definition of the lim inf of a family of subsets F(u). DEHNITION 2 Let U be a metric space, uq belong to U, and F be a set-valued map from U to X. We set (6) M-+M0 e>0 tj>0 ue B{uo,tj) We observe that when the images of F are closed, (7) lim inf F{u)<=-F{uo) U-*UO and F is lower semicontinuous at uo if and only if (8) F{uo)=lim inf F(u) U-^UO It is useful to note that v belongs to lim inf„^^ F(u) if and only if (9) Ve>0, 3f;>0 such that sup d{v, F(u))^e u 6 B{uo,n) DEFINITION 3 We say that the subset (10) C*(xo):=liminfJ(X-x)=n U 0 (UK~x)+eB /i-»0+ n e>0 a,P>0 xe BxixotOi) \n x^xo he]0,p] is the tangent cone to K at Xq. In other words, v e Ck{xo) if and only if iVs>0, 3a>0, 3p>0 such that Vx e 5|c(,xo, a) |v/i6]0, j5], 3uev-\-eB satisfying x-\-hueK or, equivalently, if and only if
CH. 7, SEC. 1 CONTINGENT AND TANGENT CONES 407 (12) for all sequences of elements e A", > 0 converging to xo and zero, there exists a sequence of elements e AT converging to v such that Xn + hnUn belongs to K for all n. It is also characterized in the following way: (13) V e Cx(xo) if and only if lim dic(x+hv) =0 h^0+ We observe that when xelni{K\ then Ck{x) = X. For all xeX, we have Cx{x)=X. We shall set C^f,{x):=0. Tangent cones enjoy a very attractive property. PROPOSITION 4 The tangent cone Ck(xo) to K at xq is closed and convex. A Proof. Let and belong to Ck(a:o). We take any sequence of elements {Xny h„)eKx ]0, cx)[ converging to (xo, 0). There exists a sequence of elements vjt converging to such that the elements y„:=Xn+hnv!t belong to K for all n. Since y„ converges to xq, there exists a sequence of elements v„ converging to such that y„ + h„Vn —Xn + hn(vl H- vl) belongs to K for all n. Since H- vl converges to we deduce that belongs to Ck(xo). Hence, the tangent cone is convex. ■ We note that Cx(^o)^Tx(xo)<=cl [ (J -(K —x:o)^ \h>on ) PROPOSITION 5 If K is a convex subset, these three cones coincide (14) Ck(a:o)=7k(xo)—cl (J -(K — Xo) h>o n Proof. We have to prove that any Wo eel (V^)(^~-^o) belongs to Ck(xo). Let 6 > 0 be fixed; there exist y eK and p>0 such that uo - ^ {s/2)B. Let us take (x:=Pe/2, x in Bxixoy a), and h e ]0, P~\. We set Then x + hu = \ i--jx-\-py
408 CH. 7, SEC. 1 NONSMOOTH ANALYSIS belongs to K, because both x and y belong to K and h/P ^ 1. Also, y-Xo II Ii^ll^-^oll , I|m-MoII=^ ;;—+ P Uo- p a e ^—l--=г P 2 Hence, uo belongs to CkÍXq)^ These two cones may be different. Consider, for instance, the set K from which is the graph of the map n from IR defined by Then 7r(x)=0 when n{x)=x when x^O if X < 0, C¡c{Xy 0)=T¡c{x, 0)=R X {0} ifx=0, Ck(0,0)={0,0}, TKÍ0y0)={-U^x{0})u{uyu}ueR^ if X>0, Ck(x, x) = TKÍXy x) = {Uy u}ueU The tangent cone to K at (0, 0) is convex but trivial, whereas the contingent cone to K at (0, 0) is nonconvex but quite large. We also observe that when iC is a smooth manifold (of class C^), then both the tangent cone and the contingent cone coincide with the usual tangent vector space to K at X of differential geometry. The contingent and tangent cones are related by the following interesting relation. PROPOSITION 6 Assume that X is finite dimensional. Then (15) Vxo e Ky Cjt(xo)<= lim inf Tk(x) Proof. By definition of the tangent cone, we have Q(xo)=n U U n n £ a>0 P>0 X€ Bk{xo,(x) he]0,p] \n J Let e and a be fixed. It is clear that U n n Q(K-x)+efi)<= nun (\(K-x)+bB p>0 xeBK{xo,<x) h€]0,P] \n J xeBxixo.a) P>0 he]0,p]
CH. 7, SEC. 1 CONTINGENT AND TANGENT CONES 409 Since X is finite dimensional, we observe that any v in (J n — ) belongs to Tk{x)-h /?>0 /ie]0,/?] / Indeed, there exist j? and elements xh such that .Xh-x V e- + sB forh^P A subsequence of {xh—x)/h converges to some w in Tk(x). Hence, CK(^o)<=n U n (rKW+6fi)=lim inf TxW £>0 a>0 xeBKixoA) This inclusion is actually an equality. ■ THEOREM 7 Let K be a nonempty weakly closed subset of a Hilbert space. The following inclusions hold true: (16) lim inf Tx(x) c lim inf (co T^íx)) c Ck(xq) IVhen X is finite-dimensional, equalities hold true. A Then the set-valued map is lower semicontinuous at Xo if and only if the contingent cone to K at Xo coincides with the tangent cone to K at Xq. ■ The proof ensues from the following lemmas. LEMMA 8 Let K<^X be a weakly closed subset. We denote by 7Tk(x) the nonempty subset of elements xbK such that ||x—j^|| =dx(y). We obtain the following: (17) fy$K, VxeTtK(y), VuecorK(x), then {y-x, v)^0 A Proof. Let X e n^iy) and v 6 Tk(x). We deduce from the inequalities 11 y - x| I - i/ji(x-I-/rt))=- i/jcU+A«) < 11T --X - MI that Ilm ll>-^11 -11/ -irJl*^Ita Inf =0 \\y-x\\ h h
410 CH. 7, SEC. 1 NONSMOOTH ANALYSIS {oxy^x, since m->||m|| is differentiable at u^Q. So u)<0 for all v eTx(x) and, consequently, for all y e co Tk(x). ■ LEMMA 9 For any y eX, we have (18) lim inf ^ {dK.(y-\-hvf -dK.{yf)<dK,{y)d(v, co TK(nK(y))) A Proof, Let us take ;c in 7tj^(y). We observe that ^idK{y+hvf-dK(y?H^i\\y+hv-x\\^-\\y-x\\^) because —a:||. Therefore, lim inf (dK(y+hvf-dK(y)y^(y-x, v) h-,o+ 2/z and for all w € co Tk(x), we deduce from lemma 8 that lim inf (dK(y+hv)^-dic(y)^)^ (y-x, t)-w)<||j>-x|| ||y-w|| =i/K(j)||y-w|| h-*0+ Lemma 9 ensues by taking the inffmum when w ranges over coTjcW and x over TtKij)- B LEMMA 10 Let us consider the Lipschitz function f defined by f(t):=^dK(x + tvfi- For almost all t^O, we have (19) f'{t)^dK{x+tv)d(v, CO 7’K(rtK(x+iy))) Proof of Theorem 7. Let Vq belong to lim inf*.,,^ c^k(^)- Then, for all 8>0, there exists >j>0 such that for all x e Bk(xq, h)> Vo € co Tk(x)+£B. Now if x belongs to Bic(xo, «) and ie]0, )?[, then nic{x + tvo)^BK(xo, h) whenever 2a + /S||yoll<»?. This happens, for instance, when a:=ril4 and P:=tl/‘2\\vo\l By setting f{t):=^dK{x + tvo)^, we deduce from lemma 10 that CO Tic{nK{x + tVo)))^edK{x + tVo)^et\M\ because ifK(^ + iyo)<i||i^oll
у 2 CH. 7, SEC. 2 CONTINGENT DERIVATIVES AND DERIVATIVES 411 Therefore, for all x 6 Bk(xo, a) and h 6 ]0, /?], ^ d^ix + hvo)^ =m-m= ll’ /'(t)rff^elkoll ■ and, consequently, Urn *<^+'"’■>,0 x^xo h h-io+ This implies that Vq belongs to the tangent cone Cx(xo). Then, we obtain lim inf Tx(x)c:lim inf ^Tk(x)<=^Ck(xq) When X is finite-dimensional, proposition 6 implies that these three cones are equal. ■ The tangent cone C/^(xq) being a closed convex cone, it is equal to Ck{xo)~ ”, its negative bipolar cone. Since this duality relation is quite useful, we introduce the following definition. DEFINITION 11 We shall say that the negative polar cone (20) Nk(a:o) ^ = Ck(xq) to the tangent cone to K at Xq is the normal cone to K at Xq. A 2. CONTINGENT DERIVATIVES AND DERIVATIVES OF A SET-VALUED MAP We adapt the intuitive definition of a derivative of a function in terms of the tangent to its graph to the case of a set-valued map. Let F be a proper set-valued map from X to Y and let (xo, уо) belong to graph(F). We denote by DF{xq, уо) the set-valued map from X to Y whose graph is the contingent cone Tgraph(F)(^o> j^o) to the graph of F at (xq, ^o)- In other words, (1) Vo e DF(xo, yo){uo) if and only if («о, i^o) e 7’graph(f)(^Vo, уо) We observe that vq belongs to DF(xoy yo){uo) if and only if Jthere exist sequences «„-►wq and Vn-^Vo [such that Vn e (F(xo + for all n
412 СН. 7, SEC. 2 NONSMOOTH ANALYSIS DEFINITION 1 We shall say that the set-valued map DF{xq, уo) from X to Y is the “contingent derivative” of F at (xo, ;^o) e graph(F). ▲ It is a “process,” that is, a positively homogeneous set-valued map (since its graph is a cone) with closed graph. We now give an analytical characterization of DF(xq, уо), which justifies that the preceding definition is a reasonable candidate for translating the idea of a derivative as a (suitable) limit of differential quotients Vo belongs to DF(xo, ;Vo)(mo) if and only if F{xo+hu)-yo^ (3) lim inf d\vo, /i->0+ =0 When f is a single-valued map, we set (4) DF{xo):=DF(xo, F(xo)) since >>o=F(a:o). This formula shows that in this case, Vq belongs to Df (xo)(mo) if and only if 11 f (a:o -I- Am) - T (xo) - /luo 11 (5) lim inf- «-►MO -=0 If F is C^ then DF(a:o)(mo)=VF(xo)mo. When the graph of F is convex, we ob¬ serve that Mo belongs to DF(xo, To)(mo) if and only if (6) lim inf I inf if ( Vo, u-*uo \h>0 F{xo+hu)-yo =0 PROPOSITION 2 Assume that F is Lipschitz on a neighborhood of Xq {belonging to Int Dom f). Then Vo belongs to DF{xo, Jo)(mo) if and only if (7) lim inf d ii->0+ ^1^0, F{xoFhuo)-yo =0 Furthermore, if the dimension of Y is finite, then (8) DomDF(xo,yo) = X Proof a. The first statement follows from the fact that (9) F(xo-\-hu)-yo<=^F(xo-^huo)-yo-\-^h\\u-Uo\\B when both h and ||w —«oil are small.
СН. 7, SEC. 2 CONTINGENT DERIVATIVES AND DERIVATIVES 413 b. Let Wo belong to X. Then for all A>0 small enough, (10) уо eF(xo)^F(xo+huo)-^^h\\uo\\B Hence, there exists vn e F{xo-\-huo) such that {vh—yo)/h belongs to ^||wo||^ which is compact. A subsequence {vh„—yo)thn converges to some uq, which belongs to DF (xq, yo){uo). ■ We point out that (11) Vxo6X, \/yoeF(xo), DF(xo,yo)~^=D{F~^){yo,Xo) Indeed, to say that (wo, vq) eTgrap^nixo, уо) amounts to saying that (vq, uq) e Tgraph(f“ 0(3^0» -^o)* Contingent derivatives allow us to differentiate restrictions of a map or a set-valued map to a subset. PROPOSITION 3 Let F be a single-valued map from an open subset O. of X to Y of class and let К be a nonempty subset of Q containing Xq. Then (12) [ 0 if Uo^Tk{xo) Proof If F is a single-valued map at xq and wo belongs to T^ixoX there exist sequences and w„->Wo such that Xo-^hnUn belongs to K. Since F\k(xo + h„u„) = F (xo + hnU„) = F (xq) + h„{VF{xo)u„ + 0(/i„)) we deduce that the elements u„:=VF(xo)w„ + 0(/i„) converge to VF(xo)wo and belong to {F\k{xo + h„u„)~F\K{xo))/h„. Therefore, DF\k(xo, F(xo))(wo)=VF(xo)wo ■ We follow the same procedure in defining the derivative of a set-valued map from X to Y. Let (xo, yo) belong to the graph of F. We denote by CF(xq, yo) the closed convex process from X to 7 whose graph is the tangent cone C graph(F)(-^o, To) to the graph of F at (xq, To)- Briefly, (13) Vo eCF (xo, To)(wo) if and only if (wo, t^o) e C graph(F)(xo, To) DEFINITION 4 We shall say that the closed convex process CF(xo, To) from X to Y is the deriva¬ tive of F at Xo^ Dom F and yo 6 F(xo). A
414 CH. 7, SEC. 2 NONSMOOTH ANALYSIS We observe that vq belongs to CF{xo, >'o)(mo) if and only if (14)!'^®'’ 3a, ^>0 such that V(o:,;;) e fig„ph(f)(xo, «). ^ (V/i 6 ]0,/S], 3ueuo + £iB, 1)6 1)0 + 625 such that v e(F{x+hu)-y)lh or, equivalently, if and only if (15) For all sequences of elements {x„, y„, h„) e graph(F) x ]0, oo[ converging to (xo, yo, 0), there exist sequences of elements u„ converging to Mo and v„ converging to Vq such that yn+^nV„ eF(x„+h„u„) for all n > 0. The analytical formula Involving “differential quotients” is quite compli¬ cated. It is simpler when F is locally Lipschitz; we begin with it. PROPOSITION 5 Assume that F is Lipschitz on a neighborhood of an element Xo e Int Dom F. Then Do belongs to CF{xo, Fo)(mo) \f ond only if (16) x-^xo,h-*0'^ V h Remark We observe that the domain of the derivative of a Lipschitz function is not necessarily the whole space, while the domain of the contingent derivative is the whole space when the dimension of Y is finite. Take, for instance, the map n associating toxeU, ti(x):=0 if and n{x)=x if x>0. We saw that Cn{0, 0)(m) =0 when u^O and Ctc(0, 0)(0)=0, whereas Z)7c(0, 0){u)=n{u) for all m € R. ■ For the analytical formula in the general case, we need the following defini¬ tion. DEFINITION 6 Let U and V be metric spaces and ф be a function from UxV to U. We set (17) lim sup inf ф{и, i;):=sup inf sup inf ф{и, v) к u-*uo v-*vo «>0 tj>0 t4€B(uo,fi) veB(vo,e) PROPOSITION 7 Let F be a proper set-valued map from X to Y and let (xo, Уо) belong to graph(F). Then Vo belongs to the derivative CF{xo, Fo)(mo) if ond only if (18) limsup graph (f) h-^0 +
CH. 7, SEC. 2 CONTINGENT DERIVATIVES AND DERIVATIVES 415 Proof of Propositions 5 and 7. Formula (14) can be written sup inf sup inf if I Do. Ei>0 (x,P>0 (Jc,y) e Bgraph(F)(xo,yo;a) Mo + £iB \ This proves proposition 7. When F is Lipscitz around xq, ,„f M€Uo + £iB \ ^ / V ^ / and the formulas become . ^ F(x + huo)-y\ ^ . mf sup d I Vo, 7 1=0 ■ a,/?>0 (x,y)e Bgraph(f)(xo,yo,a) \ ^ / he]0,p] When F is single-valued, we set (19) CF{xoy.= CF{xo,F{xo)) If F is continuously differentiable at xq, we have (20) CF{xo) = yF{xo) Naturally, the formula for derivatives of inverses is obvious (21) 'f{xo, уо) e graph(F), CF{xo, Fo)" ‘ = C(F” ‘)(j'o, лго) PROPOSITION 8 Let F be a single-valued map from an open subset Q of X to Y, continuously differentiable at Xq 6 П and let К be a nonempty subset of X containing дго- Then (22) ГЧГ1 1 \ i^P(xo)uo if иоеСк(хо) Cf|x(^:o)Mo=i ^ ^^/4 10 if Uo t Cx(a:o) Proof Let {x„, h„)eKx ]0, co[ converge to (xo, 0) in K x IR + . If mq belongs to Cx(xo), there exists a sequence of elements u„ converging to Uo such that x„+h„u„ belongs to K for all n. Then F |x(x„ 4- h„u„)=F(x„+h„u„)=F (x„)+/i„(VF (x„)«„ + 0(h„)) Since F is continuously differentiable, the sequence of elements v„ :=WF {x„)u„ +o{h„) converges to VF(x:o)mo. and we have F\idx„)+h„v„=F\icix-i-h„u„) for all n. m
416 CH. 7, SEC. 2 NONSMOOTH ANALYSIS Since the derivative CF (xoy yo) is a closed convex process, it is equal to its bi¬ transpose CF{xoyyo)**- This suggests that we introduce the following definition. DEFINITION 9 We shall say that the transpose CF{xo, >^o)* of the derivative of F at (xo, yo) € graph(F) is the codifferential of F at (xq, >^o)- ^ It is a closed convex process from У* to X* defined by (23) Po € CF(xoy Уо)*(^о) if and only if Vm e X, Vi; 6 CF{xo, Jo)(m), (Po, u) - (qo, v)^0 We mention an example of derivatives of a set-valued map that we shall use later. PROPOSITION 10 Let X and Y be Banach spaces, A a continuously differentiable operator from an open subset SI of X to Y, and LaSl, M^Y closed subsets of X and Y, respec¬ tively. Let F be the set-valued map from X to Y defined by (24) F(x)-.= A(x)—M when xeL 0 when x^L Let (xo, yo) belong to the graph of F.The following conditions are equivalent (25) Vo 6 CF{xo, Уо)(мо)- Mo e Cfixo) and Vo e VA(xo)uo - Cm(Axo - уо)- Proof a. Let us prove that (i) implies (ii). We take sequences (x„, z„, h„) e LxMx]0, 00[ converging to (xo, Ахо~Уо, 0)- Then ;^„:=.4(x„)-z„ converges to Уо, and by (i), there exist sequences u„ and v„ converging to uo and vo such that x„-i-h„u„eL and A{x„+h„u„) e M-\-y„-i-h„v„ for all n. This implies that mq belongs to Cl(xo) and that V/4(xo)mq-Uo belongs to См(Ахо~уо) because w„:=A{x„-\-h„u„)-A{x„)-v„ converges to V/1(xo)mo-uo and z„ + h„w„ belongs to M for all n. b. Conversely, let us show that (i) follows from (ii). We take a sequence (л:«, Ут h„) 6 graph(f) x ]0, oo[ converging to (xo, уо, 0). There exists a sequence M„ converging to Mo such that x„-I-h„u„ belongs to L, and since Ax„-y„ converges to Ахо—Уо in M, there exists a sequence of elements w„ converging to VA(xo)uo — Vo and satisfying Ax„—y„-t-h„w„ e M for all n. Then the sequence of elements u„:={[y4(x„-l-/i„M„)—^x„]/A„}-w„ converges to Vo and satisfies y„-\-h„v„e F(x„ -I- h„u„) for all n. ■ PROPOSITION 11 Let К be a closed convex subset of a Hilbert space X and let po belong to the normal cone Nk{xo). Let Nk denote the set-valued map x->^Nk(x) and the
СН. 7, SEC. 2 CONTINGENT DERIVATIVES AND DERIVATIVES 417 Lipschitz single-valued map associating to x its best approximation 7iic(x) 6 К by elements of K.T hen the two following statements are equivalent: (26) i. qo e CNk(xo> Po){uo)- ii. Mo 6 Слк{хо + Po){uo + qo)- The same result holds when the derivative is replaced by the contingent derivative. k Proof We recall that p belongs to the normal cone Nk{x) if and only if X:=n^X+p). a. Assume that go belongs to CA^x(^o. Po)(mo)- Let us consider a sequence of elements (, /i„) 6 X X ]0, 00 [ converging to (xo+po, 0) We set x„:=7Cx(y„), which converges to Xo=n,fxQ+po), and p„'.=y„ — x„, which converges to p^. Then there exist sequences of elements u„ and q„ con¬ verging to Mo and qo such that p„+h,^„ belongs to NK.{x„+h„u^ for all «; that is, such that TtK{yn)+h„u„=%K{yn + h„{q„-iru„)) for all n Hence, Mo belongs to C;ik(a:o+/Jo)(«o+^o)- b. Conversely, assume that mo belongs to Cn^ixo + PoKuo+io)- Let (x„,p„, h„) e graph TVk x]0, 00 [ converge to (xo, po, 0)- Since x„+p„ converges to xo+po. there exist sequences of elements u„ and w„ converging to uq and uo+qo such that x„+h„u„=UgiXn + Pn)+h„u„=7tK(x„ +p„ + h„w„) for all « Then q„:=w„—u„ converges to uo and we deduce that P«+h„q„ 6 NK{x„+h„u„) for all n Hence, qo belongs to CNk{xo, />o)(«o)- ■ COROLLARY 12 Let us consider the set-valued map associating to xeR\ the normal cone Nri(x) to R+ at X. Let belong to ЛГд»[^(х®). Then q^ belongs to CWr![^(x°,p®)(m®) if and only if (27) {0} if Xi>0 (and thus pi=0) 0 if x?=0, and UifO IR if x?=0, p?<0 and M(=0 {0} if x?=0, p?=0 and Mi=0
418 CH. 7, SEC. 2 NONSMOOTH ANALYSIS Proof. We observe that A!„) = (7t(Xi), . . . , n{x„)) where jr(x)=0 when and n{x)=x when x>0. Since CJt(jc)(M)=0 when x<0, u when x>0 and Cn(O){u)=0 when u^O and C7t(0)(0)=0, we obtain corollary 12. ■ 3. EPICONTINGENT DERIVATIVES AND EPIDERIVATIVES OF REAL¬ VALUED FUNCTIONS We can use the concept of contingent derivatives and derivatives for single¬ valued maps V from Dom FcX to R. We obtain, for instance, (1) Vo eDV(xXuo)^lim inf u-*uo V(x+/iu)~ V(x) — Vo = 0 In many problems, such as minimization problems, the order relation plays an important role. This is why we associate with a proper function F: X-»IR u {+00} the set-valued map V+ defined by V+(x)=F(x)-l-R+ when F(x)< +00 and V+(x)=0 when V{x)— +00. Its domain is the domain of V, and its graph is the epigraph of V. We consider its contingent derivative D\+(x, F(x)), whose images are closed half lines. Therefore, for all «0 € X, D\+(x, F(x)(mo) is either R or a half line [uo, oo[ or empty. We set (2) £>+F(A:)(M):=inf {i>|u 6£)V+(x, F(x))(m)} It is equal to —00 ifZ)V+(x, F(x))=R, to Vo ifDV+(x, F(x))(«)=[uo, 00 [, and to +00 \{D\+{x, V(x)Xu)=0. DEFINITION 1 We shall say that Z)+ F(x)(m) is the epicontingent derivative of V at x in the direc¬ tion u. A We begin by computing epicontingent derivatives. PROPOSITION 2 If V is a proper function from X to7?w{+oo}, then (3) D+F(xo)(«o)=lim inf li-»0+ U-*UO V(xo+hu)-V{xo) The function u-^D+ F(xo)(m) is positively homogeneous and lower semicontinuous when D+V(xo)(m)> — 00 for all ueX. k
CH. 7, SEC. 3 EPICONTINGENT DERIVATIVES 419 Proof. Indeed, let vobDY+{xq, K(xo))(mo); then Vai>0, 62>0, Va>0, there exist « e Mq + £2^ and h<a. such that .. ^YAxo+hu)-V{xo) . „ Vq g \-SiB This implies that V(xo-^hu)-V{xo) Voi Therefore, i„r fi <a llu-uoll <62 ^ h h-*0+ U-*UO For the time being, let us set . ,V{xo-\-hu)-V{xo) a:=lim ini ; h^O+ U-^UO Thus, we have proved that a^D+ l^(xo)(«o)- On the other hand, we know that for any M>a, . t . f V{xo+hu)-V{xo) sup inf inf — —<M a>0 /i<a ^ i>0 that for all a,d>0, there exist h<a and ueuo + ^B such that V{xo + hu)- V{xo) Hence, M 6 [V+(xo+/im)- F(xo)]//*, which proves that a e£>V+(xo, T(xo))(«o)- Since it is smaller than all the other ones, we infer that a =£> + F(xo)(«o)- ■ If V is C‘ at jco, then (4) Vmo e X, Z)+ F(xo)(mo)=<VF(xo), uf) If F is convex, then ■. , F(xo+/im)-F(xo)' (5) Vmo 6 X, Z)+ F(xo)(«o)=lim inf ( inf «-♦MO \h>0 We deduce from propositions 2.2 and 2.3 the following statements.
420 CH. 7, SEC. 3 NONSMOOTH ANALYSIS PROPOSITION 3 Let us assume that V is Lipschitz on a neighborhood q/’xo 6 Int Dom V. Then (6) w n T/^ V ^ 1. • y(xo-^huo)-V{xo) vuoeX, D+V(xo)(uo) = lim inf /1-0+ n and the epicontingent derivative is finite. A PROPOSITION 4 Let V be a proper function from X to Rkj{ +00} and K a subset of X. Let V\k denote the restriction of V to K {in the sense that V\fc{x) equals V{x) when x eK, 00 when X ^ K). Then (7) VxoeK, Vt;o eT’xixo), Z)+K(xo){moX^+I^U(^o)(mo) If F is C* at xo, we have +00 if Uo^Tk(xo) a We state the obvious property of the epicontingent derivative at a minimizer. PROPOSITION 5 Let V be a proper function from a Banach space X to Rul+ooj.TjTxe Dom V minimizes V on X, then (9) ViteX, 0<D+F(x)(m) A More generally, the e-variational principle can take the following form. THEOREM 6 Let V be a proper lower semicontinuous function bounded below from a Banach space X to Ru{+oo} and let xo belong to Dom V. Then for any s>0, there exists Xc € Dom Vsatisfying (10) i- F(Xe) + 6||Xe-Xo||<F(Xo). ii. V«6X, 0<Z)+F(xe)(M)+ellMl|. Proof By theorem 5.3.1 (the e-variational principle), there exists Xc e Dom V satisfying (10, i) and F(Xe)=min;,ex[f'(jc)+fi||x-X£||]. Let u € Dom D+ F(xj). Then for any i/>0, ¿>0, a>0, there exist h^a and v&u+dB such that V(x,+hv)- F(Xe) h <Z)+F(x£)(m)-I->/
CH. 7, SEC. 3 EPICONTINGENT DERIVATIVES 421 Theorem 5.3.1 implies h Therefore, we infer that 0^-0+ (Xe)(m)+e||m|| + e5 + By letting d and ri converge to zero, we obtain the desired inequality. ■ In the same way, we define epiderivatives of functions V from X to R u {+oo}. Since images of the derivative CV+{xo, K(xo)) are either Rora half line [uq, oo[ or empty, we set (11) C+ F(xo)(Mo):=inf{u|t; € CV^xo, K(;co))(i/o)} It is equal to -oo when CV+{xo, V{xo)) = R, to Vq when CV+{xo, K(xo))(wo)= [i;o, 00[, and to +oo when CV+(xo, V{xo)){uo)=0^ DEFINITION 7 We shall say that C+ K(xo)(wo) is the epiderivative of V at xq in the direction uq- The epigraph of u^C+V{xq){u) is a closed convex cone, because it is the graph of the set-valued map w->CF+(a:o, V{xq))[u\ which is a closed convex process. We deduce at once the following important property. PROPOSITION 8 The epiderivative u^C+V(xq){u) is a positively homogeneous lower semicon- tinuous, convex function when C+V(xo)(u)> —oo for all ueX. A It is easy to check that the codifferential of V+ at (xq, V{xo)) is a closed convex process from IR to X*, defined by its values CV+(xq, K(xo))*(—1) and CV+ixo, 1^(a:o))*(1). We observe that CV+(xo, l^(:^o))*(- l)=j3'and the support function of CV+{xo, l^Uo))*(l) equals C+V(xq){-) when it is not empty. ■ DEFINITION 9 We say that the closed convex subset of X* defined by (12) dV(xoY=CV^{xo. K(xo))*(l)={/7eX*|VweX, {p,u)^C^V{xo)(u)] is the generalized gradient of V at Xq. ^ It is empty whenever there exists a direction uq for which C+ F(xo)(wo)= -oo.
422 CH. 7, SEC. 3 NONSMOOTH ANALYSIS If V is continuously differentiable at xo, then VwoeX, C+F(xo)(mo)= (VK(xo), Mo) and, consequently, 5K(xo)={VF(xo)} This justifies the term generalized gradient. When V is convex, it coincides with the subdifferential of V at xq. PROPOSITION 10 Let us assume that V is Lipschitz on a neighborhood of Xq^ Int Dom V. Then (13) Vmo e X, C+ T(xo)(mo)=lim sup X-^XQ /i->0+ V(x+huo)~ V(x) h and the epiderivative is finite. Furthermore, the following properties hold true'. i. Vmo 6 X, (x, u)-^ C + V(x)(u) is upper semicontinuous at (xo, Mq). (14) •! ii. u->C+V(xo){u) is Lipschitz. ill. C+(- T)(xo)(m)=C+ K(xo)(-m). In terms of generalized gradients, these properties become (15) i. x-^dV{x) is upper hemicontinuous at Xq. ii. dV{xo) is {closed convex and) bounded. Hi. d( - K)(xo) = - d T(xo). Proof. Since V is Lipschitz on a neighborhood of Xo 6 Int Dom V, there exist ao>0 and ^>0 such that for any a, P, tj satisfying a+j?(||Moll + 'i)<“o> we have: Vx €Xo + aj5, V/i €]0, j8], Vm euo + >lB, (16) V(x + hu)-V{x) <<i’(llMoll + >?) a. Let Vo belong to the derivative CF+(xo, V{xo)) of V+ at (xo, T(xo))- This means that for all e, >? > 0, there exist a, /? > 0 such that for all x e Xo+t^P, V/i 6 ]0, P~\, there exists ueuo+flP such that
Hence, CH. 7, SEC. 3 EPICONTINGENT DERIVATIVES 423 ^V{x-^hu)-V{x) ^V{x-^huo)-V(x) ^ Vo^ -j e—trj (because V is Lipschitz around Xo). Consequently, Vo^lim sup X-^Xq h-^0* V{x+huo)— V(x) and thus C+F(xo)(Mo)>lim sup X->Xo /i->0 + V(x+huo)- V{x) h Conversely, let us set u:=lim sup X-^Xo /1-^0 + V(x+huo)— V(x) h which is finite by inequality (16). Then we can associate to any s>0 constants a, P>0 such that a + e> V(x+huo)— V{x) This implies that a belongs to CT+(xo, V(xo)). Hence, formula (13) ensues, b. The upper semicontinuity of (.x, m)-»C+F(x)(m) at (xo, uo) follows at once from formula (13). Also, inequality (16) implies that (17) C+F(xo)(«)<^||m|| and thus that u-^C+V{xo){u) is Lipschitz. To prove (14, iii), we observe that - V{x + huo)-{-y(x)) V({x+huo)+h{-uo))- V{x+huo) Since x+huo is in a neighborhood of Xo when x is a neighborhood of Xo and h is small, we deduce that C+(- F)(xo)(mo)=C+ F(xo)(-Mo)
424 CH. 7, SEC. 3 NONSMOOTH ANALYSIS c. Since C+F(xo)(’) is proper, it is the support function of 5 F(a:o). ■ Remark More generally, we can prove the following formula for epiderivatives of arbi¬ trary functions. For that purpose, it is expedient to use the notation (18) (x, X)lxoC^^>V(x), and definition 2.6 of lim sup inf. and A->F(xo) PROPOSITION 11 Let Xo belong to the domain of a function V from X to R u {+°o}. T hen (19) C+F(xo)(«o)='int sup inf (x,A)lxi) M-+MO /1-^0 + V(x-\-hu)—X The proof is left as an exercise. When V is lower semicontinuous at xo, formula (19) becomes (20) C+ F(xo)(«o) = lim sup inf X-^Xo M-+U0 Vix)^V(xo) h-^0+ V(x^hu)-V{x) h It may be useful to use another concept of derivative, easier to manipulate than the epiderivative. DEFINITION 12 Let V be a proper function from X to Uu{+co} and let Xq belong to Dom K We set (21) 5+K(xo)(mo):=1™ sup V(x-^hu)-X (x,A)lxo tt->U0 We shall say that B+V{xo)(uo) is the strict epiderivative of V at Xq in the direction of Uo and V is strictly epidifferentiable at xq if the function m-> B+ V(xo)(u) is a proper function from X to Uu{+^}- ^ We always have (22) Vw e X, D+ V(xo){uH V(xo)iuHB^ V{xo){u) Clearly, a function V that is Lipschitz around Xo is strictly epidifferentiable at Xo-
CH. 7, SEC. 3 EPICONTINGENT DERIVATIVES 425 The introduction of this concept is justified by the following result, stated in proposition 13. A PROPOSITION 13 Let us assume that the function V is strictly epidifferentiable at Xq 6 Dom V. Then (23) and Dom 5+F(A:o)=Int Dom C+V{xo) (24) V«o e Dom C + V(xo), C+ F(a;o)(mo) = lim inf B+ T(xo)(mo) Furthermore, for any Uq 6 Int Dom C+ V{xo), (25) i. (;c, m)->C+ F(x)(m) is upper semicontinuous at (jcq, Mo) (ii. u-*C+V (a:o)(m) is continuous at Uq. If we assume that Dom 5+ V(xq)=X, then (26) d(-V){xo)=-dV(xo) A Proof a. Let uo belong to the domain of 5+ V(xol Equation (21) implies at once that Dom B^V{xq) is open and (x, m)->5+F(x)(m) is upper semicon¬ tinuous at (Xo, Uq) b. Formula (21) implies that (27) B+V (xo)(mo -I- Ml) < 5+ F (xo)(mo) -I- C + F(xo)(mi ) We deduce that any u interior to the domain of C+ F(xo) belongs to the domain of 5+F(xo). For that purpose, take uq 6 Dom B+V(xq) and A>0 such that u—Xuq belongs to the domain of C+F(xo). Then inequality (27) implies that B+ F(xo)(m)< C + F(xo)(m - Xuq) + A5+ F (xo)(mq) < +oo that is, M belongs to the domain of fi+F(xo). Hence, the domain of 5+F(xo) coincides with the interior of the domain of C+F(xo). Inequality (27) implies also that the epigraph of 5+ F(xo) is dense in the epigraph of C+ F(xo). Conse¬ quently, lim inf B+ F(xo)(m)^ C+ F(xo)(mo) U~-*UQ Since u^C+ F(xo)(m) is lower semicontinuous, equality (24) ensues. Furthermore, by letting A go to zero in the preceding inequality, we obtain (28) Vm e Dom B+ F(xo), 5+ F(xq)(m)=C+ F(xq)(m)
426 CH. 7, SEC. 3 NONSMOOTH ANALYSIS Inequality (24) implies that (29) dV{xo)={pBX*\\/u^X, {p,u)^B^V(xoM} Hence, property (26) follows from (30) Vm 6 Dorn 5+ K(xo), 5+ F(xo)(-«)=£+(- F)(xo)(m) To prove this statement, we set vo:=B+V(xoX-Uoy, for all e>0, there exist «0, Po, »/o>0 such that, for all e Dorn Fn(xo+aoB), h e ]0, ]8], u 6 mo+»?oB. \y{y-hu)-F(j)]/(j^i;o + e- Let us take a e ]0, ao], P e ]0, Po], and rj e ]0, /?o] such that a+/5(1|moII+»?) < «o- Hence, for all xeDom Fn(xo+«B), ^eV(xo)+ocB, satisfying A>-V(x), h € ]0, /5], u € Uo+y\B, we have, by setting ;^:=x:+Am, - V{x+hu)~X ^ V(x)- V{x+hu) _ V{y-hu)- V(y)^ . ^ . — 'i'.' i>o+e h h h because y belongs to Dom Fn(x:o+aoB). This implies that B+(- F)(a:o)(mo)<Uo:=B+F(a:o)(-Mo) By exchanging the roles of V and — K we have proved equality (30). ■ PROPOSITION 14 Let P be a closed convex cone. Assume that V is nonincreasing with respect to P in the sense that F(x+_y)< F(x) for all y eP. Let xq belong to Dom V. Then (31) Vmo6P, C+F(;co)(mo)<0 and (32) dV{xo)<zp- If Int P4^0, then (33) VMo€lntP, B+F(xo)(mo)<0 a Proof. Indeed, for any xexo+aB, 2eF(xo)+aB, A^V(x), Ae]0, j8[. Mo e P, we have V(x+huo)-A^V(x+huo)- h '' h '' because F is nonincreasing. Hence, C+F(a:o)(mo)<0. If mo belongs to the interior of P, there exists >/o > 0 such that mq + rjoB^P. Hence, for all m 6 mo +
we would have CH. 7, SEC. 3 EPICONTINGENT DERIVATIVES 427 V(x-^hu)-X^V{x+hu)- and thus B+ F(xo)(mo)<0. ■ We deduced the property of epicontingent derivatives and epiderivatives from the properties of contingent cones and tangent cones. Conversely, we can derive properties of contingent cones and tangent cones from those of epi¬ contingent derivatives and epiderivatives, because we remark that when xo belongs to a subset K, (34) D+ ^k(xo, 0)=i/'tko.b) and C+iAk(xo, 0)= where ipt denotes the indicator of the subset L. We also mention the following useful properties. PROPOSITION 15 a. If p&X* satisfies (p, Xo)=maXy«K (p, y), then p belongs to the normal cone Nk{xo). b. Assume now that X is a Hilbert space. If y f K and x e Jt^(y) is a projection of y to K, then y—x belongs to the normal cone Nk(x). ^ Proof a. If p eX* satisfies {p, Xo)=maXyeK(A j)> then xo^X mini¬ mizes on K the linear functional x^{p, x), and thus 0 6 5(—pIk)(.xo)= —p + Nnixo) by propositions 4 and 5. b. Since the function V\ x-»||y—x||=:K(x:) is continuously differentiable at all X ^y and x 6 Ti^y) minimizes V on K, we deduce that 0 6d(FU)(x)<=VK(x)-HWx(x)= x-y + Nk{x) Hence, y—xe Nk{x). PROPOSITION 16 Let V be a proper upper semicontinuous function from X to set (35) K:={xeX\V{x)^c} Let Xo gK satisfy V{xo)=c. Then (36) Tdx)<={v 6 X|Z)+ K(x)(u)<0} If we assume that (37) 3Mo 6X such that C+V(xo)(mo)<0
428 CH. 7, SEC. 4 nonsmooth analysis then inclusion (38) {u eX\C^ V(xo)(u)^0}czCk(xo) holds true. Proof. We first check that if uq satisfies C+K(xo)(mo)<0, then to Ck{xq). Let us set Vq:= — C+V{xo){uo)>0. For all a6]0, i;o[, a>0 and jS>0 such that for all xexo + o^B, h e]0, jS[, there exists such that V(x + hu)^V{x)-{-h{—Vo-^e). Hence, for all xeB^ixo, a), there exists ueuo-\-sB such that V{x + hu)^ V{xo\ that is, x + hueK. uo belongs to Ck{xo). Now, if u satisfies C+F(xo)(w)^0, then for all X e]0, 1[, ma*=(1 satisfies C+F(x:o)(ma)<0 by convexity, and thus ux e Ck{uoI Hence, that u belongs to Ck{uo) by letting A converge to zero. Uo belongs there exist и euo-\-sB A€]o, pi Therefore, — A)w + Xuq we deduce 4. GENERALIZED SECOND DERIVATIVES OF REAL-VALUED FUNCTIONS Let F be a proper function from X to/?u{+oo}. We consider fhe set-valued map dV from X to X* associating to each xqgX the generalized gradient of V at xq. Therefore, if (xq, po) belongs to the graph of 5K the derivative C(5F)(xo, Po) of dV at (xo,po) plays the role of a second derivative of V. DEFINITION 1 We shall say that the derivative (1) d^V{xo,po)-C{dV){xo,Po) of the map dV at (xq, Po) ^ graph (dV) is the generalized second derivative of V at {xo, Pol T herefore, d^'V(xo, Po) is a closed convex process from X to X*. к It is clear that when V is twice continuously differentiable at xo, then po = W(xo) and d^V{xoi Po) coincides with the Hessian V^F(xo), mapping X to X*. PROPOSITION 2 Let V be a proper lower semicontinuous, convex function from X to /^u{+oo} and F* its conjugate function. Then d^V(xo, po) is a monotone closed convex process and (2) S^V*iPo,Xo)={d^V{xo,por'
CH. 7, SEC. 5 THE INVERSE FUNCTION THEOREM FOR SET-VALUED MAPS 429 Proof. Let {u\ q')(i = \, 2) be two pairs of the graph of d^V(xo, Po\ Let /t„ converge to 0*^. Then we know that there exist sequences of elements wl, and Vn converging to M* and t;‘ such that {xo^Kun, po-\-hnq't) belong to the graph of 5L for i = 1,2. Since the graph of 5 F is monotone, we deduce that hi (gl - ql, ul-ul) = {po+h„ql ~(po + h„ql), xq+h„ul - (xq+h„ul)} > 0 Hence, 5^F(xo, Po) is monotone. Inequality (2) is straightforward, since 3F* is the inverse of 5F 5. THE INVERSE FUNCTION THEOREM FOR SET-VALUED MAPS We denote by pB and pBP the closed and open balls of radius p, respectively. We set dl(y4, ^):=sup inf ||x—j;|| xe A yeB Note that B)=^ means that A is contained in the closure of B. We shall extend the usual inverse function theorem for continuously dif¬ ferentiable single-valued maps to the case of set-valued maps. We need the following definition. DEFINITION 1 Let F be a proper set-valued map from X to Y and let (л:о, у о) belong to the graph of F. We say that F is pseudo Lipschitz around (xo, Уо) if there exist a neighbor¬ hood W of Xo, two neighborhoods U and V of у q, and a constant £>Q such that (1) i. УхбЖ F{x)nU^0. ii. Va:i, ХгеЩ d(F(xi)n U, F(x2)n V)^€\\xi-Хг\\. We note that if F is single-valued on W^ it is pseudo Lipschitz if and only if it is Lipschitz. THEOREM 2 Let F be a proper set-valued map with closed graph from X to Y and let (xo, J^o) belong to graph (F). We assume that (2) i. Both X and Y are finite dimensional. ii. The derivative CF(xo, yo) of F at (xq, j^o) is surjective (i.e., Im CF(xo, ;^o)= L). Then F ^ is pseudo Lipschitz around (;ио> ^o)-
430 CH. 7, SEC. 5 NONSMOOTH ANALYSIS We start with the following lemma. LEMMA 3 Let us assume that the spaces X and Y are finite dimensional. Let (xq, yo) belong to the graph of F. We assume that (3) the derivative CF{xq, yo) maps X onto Y. Then, for all a>0, there exist constants oO and rj>0 such that for all (x, y) e graph (F) satisfying lk-^oll+llj-;'oll<>/ andfor all veY, there exist ueX and weY satisfying (4) t) eZ)F(x, j)(m)+w, ||m||^c||i;||, and ||w||<a||i;|| A Proof Since CF{xo, j'o) is a closed convex process, proposition 3.3.8 with Xo=0 and =0 implies the existence of y > 0 such that (5) yB<=CF(xo, yom Let us introduce the subset (6) Xr.=(BxyB)ngraph CF{xo, Fo) Since the spaces X and Y are finite dimensional, the subset Jf is a compact subset of the tangent cone C^(xo, yo) to the graph of F, at (xo, jo). which is the lim inf of the contingent cones F^(x, y) to the graph of F at points (x, y) con¬ verging to (xo, Fo)- Hence, we can associate to every a > 0 a positive number ri such that for all («o, Vo) 6 , (x, y) 6 B^(xq, yo, rj), we have (uo, vo) e T^(x, y) + a(BxB). Now, take t) in L Then Uo:=yt)/|lD|l belongs to yB and by (5), there exists Uo sB such that (uo, Vq) belongs to Jf. Then, for all (x, y) e B^(xo, yoh), there exist Ua € xB and u* e xB such that (uo — u„ Vo — t)«) € T^{x, y), that is, such that Do e Z)F (X, f)(Mo - M«) -I- Da We set M:=(||D||/y)(Mo —Ma) and w=(||D||/y)Da. Then ,l+a„ „ veDF(x,y){u)+w, ||m||<- and Theorem 2 then becomes a consequence of the following general Inverse function theorem, valid in all Banach spaces.
CH. 7, SEC. 5 THE INVERSE FUNCTION THEOREM FOR SET-VALUED MAPS 431 THEOREM 4 Let F be a proper set-valued map from a Banach space X to a Banach space Y with closed graph. Let {xq, Уо) € graph(F) be fixed. We assume that there exist constants a e [0, 1[, y/>0 and c>0 such that for all (x, y) e graph(F) satisfying + for all V eY, there exist ueX and w eY such that (7) Let us set (8) |i. . 1«. 11« eDF{x,yXu)+w. /||=^c|lt)ll and ■“3(l+a + c)’ and Ff47):=f’"4>')n^Xo+^ji^r5^ Then F" ^ is pseudo Lipschitz around (xo, ;^o)- Namely, i. ^уеуо + гВ, Ро^(у)ф0, (9) jib V^i, Ьг&Уо + гВ, d(Fo c+2a '' 1-a ll>’i-F2ll Proof. Let yi and y^ belong to the open ball yQ+rB. Assume for the time being that there exists Xi satisfying (10) eFo Hji):=F”‘(7i)n(jco+t?r5), where^:=^^!^ 1 —a (This is possible when we take yi=y^ and xi=a:o!) We associate with any P € ]ll3^i —yiW 2r[ the number 61= \\У1-У2\\+^Р that satisfies (П) 3|bi-:t^2ll^^ ^ l-g 2»/ '' 1 + c+a We apply theorem 1 of Chapter 5, Section 3, to the continuous function V defined on the graph of F by V{x, >'):=||3'2—3^||.
432 CH. 7, SEC. 5 NONSMOOTH ANALYSIS Since it is complete, there exists (3c, y) e graph(F) such that (12) I i. ||j5->'2ll + e(ll^-Xill + llF-;^i||)<||ji->»2ll- 111. V(x,;;)€graph(F), llF-j^2lNllF-F2ll+e(IU-3c|| + |b-j;||). Inequality (12, i) implies that Therefore, ll^-^ill + llF-Fill<^|lFi-j2Ny 2ri ll^-^oll + llF-Foll<y+l|A:o-Xi|| + ||_vo-Fill ^2f/ /c+2a \ 2ri l+cc + c Irj rj Consequently, we can use property (7) with v:=y2~y- There exist u and w satisfying (13) •• F2 - F e £>F (x, 7)(m)+w II. ||MNc||y2-j'll and l<a|lF2-Fll By the very definition of the contingent derivative DF{x, y), we can associate to any 5 >0 elements h e ]0, ¿], us e SB, and 6 SB such that the pair (x, y) defined by x=x+hu-\-hus, y=y+h{y2—y)—hw—hvs belongs to the graph of F. Using this pair in inequality (12, ii), we obtain IIf2 - jll <(i -/i)lly2 - i'll+/i|IHI+a«(IIm||+11^2-y\\ +1 + /l((l+£)||l^ill+«ll»<'’ll^ We divide this inequality hy h>0 and let 5 converge to zero. We get llF2-Fll=i(«(c + l)+a(l+e))llF2-i'll Since £ < (1 — a)/(c +1 — a), we infer that j2 = and thus that 3c is a solution to the inclusion >>2 6 F{x). By setting ;^2 =y >n inequality (12, i), we obtain -lV,-F2ll=^P<2^r
CH. 7, SEC. 5 THE INVERSE FUNCTION THEOREM FOR SET-VALUED MAPS 433 Therefore, x belongs to F~^{y2)r^{xi+2frB)<=Fi^(y2), and thus d{xu i'r‘(>'2))<ll^-JCill<0 -3^211 By letting p converge to || —yill we deduce that (14) d(xu Fi\y2)H^\\yi-y2\\ We can always take (xu yi)-={xo, yol We thus have proved (15) "^y2 eyo-\-rB, X2eFoHy2)’=F~\y2)r\(^Xo-\-^^^rJ^ (because \\y2—yo\\<r instead of 2r). In other words, the set-valued map Fq ^ has nonempty images when ;; ranges over the open ball yo + rB. Inequality (14) implies that d{FoHyilFi\yi)):= sup d{xuFi\y2)) ^^T^Wyi-yiW ■ As a first consequence, we obtain the usual Liusternik theorem. COROLLARY 5 Let f be a continuously differentiable map from an open subset Q. of a Banach space X to a Banach space Y. Assume that for Xo € П, (16) y/’(xo) is surjective. Then there exist neighborhoods U and V of Xo, t/c: K cind W of /(xo) such that for ally eW, there exists a solution xeU to the equation f(x)=y. Furthermore, (17) V;;i, ;;2 eW, d(f~ Hyi)n U,f~\y2)n V)^€\\y, Proof Let iC be a closed neighborhood of xq contained in Q. We apply theorem 4 to the restriction F of / to K. Since V/'(xo) is surjective, there exists a constant oO such that for all t; e there exists a solution u of the equation Vf(xo)u=v satisfying ||m||^c||i;||. Leta>0 be given and rj>0 such that liy/^(-^)~ y/^(->^o)ll when X 6 ^(xo, ^;)<= Int K. Then the assumptions of theorem 4 are satisfied, because v=Vf{x)u-\-w where ||M||^c||y|| and w:=(V/(x)-V/(xo))y is such that ||w||^al|i;|| ■ By taking ?/ = 00, we obtain the following corollary.
434 СН. 7, SEC. 5 NONSMOOTH ANALYSIS COROLLARY 6 (NORMAL SOLVABILITY) Let F be a proper closed set-valued map from a Banach space X to a Banach space Y. We assume that there exists a constant c> 0 such that for all (x, y) € graph(F), for all V eY, there exists ueX satisfying v eDF{x, y){u) and Hwll^dlj^ll- Then F maps X onto Y, and its inverse F~^ is a Lipschitz set-valued map with Lipschitz constant equal to c. A Let US also mention the following consequence of the proof of theorem 4. COROLLARY 7 Let F be a proper closed set-valued map from a Banach space X to a Banach space Y, Assume that there exists a constant oO such that iV(x,y)€graph(F), 3m 6 A' satisfying j)(m) and ||м||<с||:и|| Then the set of zeros of F is nonempty and (19) Vx 6 Dom(f), d{x, F-H0)Hcd{0, F{x)) Remark When F=/|k is the restriction to a closed subset K of a continuously differ¬ entiable single-valued map f assumption (18) becomes (20) VxeK, 3m€Tk(x) such that -f(x)=Vf{x)u and ||m|I<c|I/WII We then deduce that there exists a solution 3c e K to the equation/(3c)=0 and (21) Vx€K, d(x,/-i(0))<c|W| In the book by Aubin and Cellina [1983], it is shown that assumption (20) implies the existence of a trajectory of the implicit differential equation (22) satisfying (23) and (24) |V/(x(f))x'(i)=-/(x(f)) [x(0)=xo given in К Vi^O, x(t) belongs to К d{x{t),f-\0))^e-‘4\xm\
CH. 7, SEC. 5 THE INVERSE FUNCTION THEOREM FOR SET-VALUED MAPS 435 We observe that the differential equation (22) is the continuous version of the Newton method and inequality (24) implies the convergence of the Newton method. ■ Remark When the graph of F is compact (i.e., when the domain of F is compact and F is upper semicontinuous with compact values), we need only to assume that (25) V{x, y) e graph(F), 3ueX satisfying —yeDF(x, y){u) for deducing that F~^(0) is nonempty. Indeed, we minimize on the graph of F the function (x,>’)->||;^|| and denote by (3c, y) 6graph(F) a minimizes We proceed as in the proof of theorem 4 with г=0. ■ The use of set-valued maps abolishes the formal distinction between the inverse function theorem and the implicit function theorem. Let X, Y, and Z be three Banach spaces and let G be a set-valued map from X X y to Z. The implicit function theorem deals with the behavior of the map that associates to any {y, z)eY xZ the set of solutions a: to the inclusion z e G{Xy y). This amounts to studying the inverse of the set-valued map F from X to y X Z defined by (26) {y, z) 6 F{x)<^z e G{x, y) Since the graphs of the set-valued maps F and G coincide as subsets of X xY xZ, there are close relations between the derivatives of F and G at (^o> yoi Zo)y because the graphs of these two derivatives coincide with the tangent cone to the graph of G at (xo, yo> ^o)- Then we can state the implicit function theorem. THEOREM 8 Let G be a proper set-valued map with closed graph from X xY to Z and let (^o> Уо) zo) belong to the graph of G. We assume that (27) I ^ finite dimensional. [ii. Vy, vveTxZ, 3ueX such that w e CG{xo, уо, zo)(w, i^). Then (28) F ^ is pseudo Lipschitz around (xo, (уо» ^o))- In the case when G is a continuously differentiable function, we obtain the following useful corollary.
436 CH. 7, SEC. 5 NONSMOOTH ANALYSIS COROLLARY 9 Let g be a function from an open neighborhood of (xo, Jo) X xY to Z satis¬ fying (29) Vygf(xo, Jo) is surjective from Y to Z Then there exist neighborhoods Vq and V^of yo' Vo<^Vu neighborhoods Uof Xo and W of Zo and a constant oO such that (30) Vxet/, fzsW, 3j 6 Fo such that g{x,y)=z and, if we set F” ^(x, z):= {j 6 Y\g{x, y)=z}, (31) Vxi, X2 e K zu zieW, d{F~ ^xi, zi)n Vo, F~ ‘(X2, Z2)n Fi)</(||xi -X2II + ||zi -Z2II) A Proof It is analogous to the proof of Corollary 5 and follows from Theorem 4 applied to the set-valued map F from Y to X xZdefined by (32) F(j):= (x, z)eKx Z\g{x, j)=z} when j e L 0 when J € L where K and L are closed neighborhoods of zo and jo on which ^ is C‘. The graph of F is closed and assumption (7) is satisfied: let (u,w)eXxZbe chosen and define t) 6 7 as a solution to the equation (33) satisfying (34) Vy0(Xo, jo)u = W - Vxgixo, Jo)« |<c||w-V,0(xo, jo)«|| thanks to the Banach open mapping principle. We set w:=Vg{x, y){u, v) ~^9ixo, Jo)(w, i^). We see that (35) with (36) and, a being given, (37) 11(0, w)||^||V0(x, y)-Vg{xo, jo)||(||«|| + ||w||)<a(||«|| -H|w||) (m, w) eZ)F(j, (x, z))(u)-M0, w) I max (1, IIV;t0(xo, Jo)||)(||«|| + ||w||
CH. 7, SEC. 5 THE INVERSE FUNCTION THEOREM FOR SET-VALUED MAPS 437 provided that {x, y) remains in a small neighborhood of (xq, yoX Hence theorem 4 implies that F“Ms pseudo-Lipschitz around (xq, (yo, zo)\ which is what the conclusion of corollary 9 states. ■ This having been said, it is not always obvious how to obtain “nice” formulas for the derivatives. For instance, let X and Y be Banach spaces, A a continuous linear operator from X io Y, and G\ and H\ Y-^Y* set-valued maps. We consider the set of solutions x eX to the inclusion (38) p e G{x)-\-A*H{Ax-\-y) where p is given in X *. The first idea is to apply the inverse function theorem to the set-valued map E from X to X* xY defined by (39) E{x)={p,y)\peG{x)+A*H{Ax+y)} Unfortunately, there is no nice expression for the tangent cone to the graph of£inA^xX*x]K But we can introduce an auxiliary variable q eY* and write inclusion (38) as the equivalent inclusion (40) i. peG(x)+A*q. ¡1. y e—Ax + H~^{q). The set of solutions (x, q) to this problem is denoted by F *(p, y. A), where F is the set-valued map from X xY* to X* xY x SF{X, Y) defined by (41) ip, y, A) 6 F{x, q) if and only if (40) holds We shall characterize the derivative of F in terms of the derivatives of the set-valued maps G and H{or FI~^\ respectively. LEMMA 10 Let xo, qo be a solution to the system of inclusions (42) i. po ^G{xo) +Aoqo. ii. yQ& -Axo + H~\qo). The following conditions are equivalent: (43) (dp, dy, dA)eCF(xo, qo i £o. yo> ^o)(5x, dq) (44) I*’ bp-dA*'qo eCGixo,Po-Mdo){bx)+A^dq. |ii. dy + dA’Xo e —Aobx+CH~^(qo, yo+Axo)(dq).
438 СН. 7, SEC. 5 NONSMOOTH ANALYSIS Proof, a. We prove that (43) implies (44). We choose sequences {x„, q„, p'„, y'„, h„) converging to (xo, qo, Po~A *qo, Jo + Axo, 0). By setting A„:=Ao, Pn =p'n+Atq„, and we see that (x„, q„,Pmym A„, h„) converges to (a:o. 9o. Po, Jo. ^o> 0). Therefore by (43), there exist sequences of elements dx„, Sqm Spm dy„, and dA„ converging to ¿x, 5q, dp, dy, and dA such that •• p'«+h„{dp„ - Atdq„-dA*q„+h„dA*dq„) e G{x„+h„dx„). ii. y'„ + h„(dy„+Aodx„+dA„-x„ + h„dA„dx„) eH~\q„+h„dq„). Hence, the system of inclusions (44) holds true. b. Conversely, let us consider the system of inclusions (44) and prove (43). We choose sequences (x„, q„, j„, p„, A„, h„) converging to {xo, qo, Jo, Po, Aq, 0). Then we know that there exist sequences of elements dx„, dq„, u„, and v„ converg¬ ing to dx, dq, dp — dA*qo — Atdq, and dy+dAxo + Aodx, respectively. We set (45) Hence, i. dA„:=dA, which converges to dA. ii. dp„:=u„+dA*‘q„ + A¡jidq„+h„dA*dq„, which converges to dp. iii. dy„:=v„—dAx„—A„dx„ — h„dAx„, which converges to ¿x. {Pn-^hnSp„, j„-l-Mjn, A„ + h,^A) € F{x„+h„dx,„ q„+hdq„) and, consequently, inclusion (43) holds true. ■ COROLLARY 11 Let X and Y befinite dimensional andletG.X^X* and H: Y-^Y* be set-valued with a closed graph. Let {po, jo, Ao) belong to X* xYx £P(X, T). Assume that there exists a solution (xo, qo) to the system of inclusions (46) i. poeG(xo)+A^qo ii. Jo 6 - AqXo + H~ ^{qo) If the matrix of closed convex processes iCG(xo,Po-A^qo) A^§ \ -.4o CH{AoXo+yo,qo)~^ is surjective, then there exist neighborhoods U and V of (xo, qo), U<=V and W of iPo, Jo, Ao) such that the set-valued map (47) (p, j. A) e W^F-\p, y,A)nU has nonempty values and is pseudo Lipschitz. Furthermore, the derivative of F~^
СН. 7, SEC. 6 CALCULUS OF CONTINGENT AND TANGENT CONES 439 is given by 5;ic\ / CG(xo, Po -A^go) 5q) \ -Ao Vdp-5A*'qo 6. CALCULUS OF CONTINGENT AND TANGENT CONES, DERIVATIVES, AND EPIDERIVATIVES The applications of nonsmooth analysis to nonlinear analysis to which we have devoted the preceding section motivate the development of a calculus of contingent and tangent cones, contingent derivatives and derivatives of set¬ valued maps, and epicontingent derivatives and epiderivatives of real-valued functions. In the tables at the end of this book, we summarize this calculus, adding the formulas of convex analysis for the sake of comparison. PROPOSITION 1 a. Let K<^L^X be two nonempty subsets. Then (1) WxoeK, TKixo)<=Tdxo) b. Let Kibe the union of subsets Ki. Then (2) VxogK, (J TKi{xo)^TK{xo) i€l If I={1,... yn} is finite, equality holds true. c. Let K := Hi e / intersection of subsets Ki. Then (3) ^XoeK, Tk(xo)<= n ie I d. Let ^^=11*6/ ^ finite product of subsets Ki. Then n (4) \/xo'.={xo)ieI ^ Ky Tk(xo)^Y\ TiCiixo) and CK(xo)=f~[ Cjj'.(.x:o) i=l i€l PROPOSITION 2 Let X and Y be Banach spaces, A a map from an open subset Q. of X to Y, and KczQa subset of X. Then (5) Vxo 6 X, ^A{xo)Tk(xo)<^Ta(K)(M^o))
440 СН. 7, SEC. 6 NONSMOOTH ANALYSIS Proof. Let Vo belong to 7]c(xo); there exist sequences and v„-^Vo such that xo+h„v„ belongs to К for all n. Then the sequence of elements u„-.={A(xo+h„v^-A{xo))lh„ converges to A(xo)vo and A(xo)+ftnU„ belongs to A(K) for all n. Hence, A(xo)vo belongs to Ta(K)(Axq). ■ In particular, ifAe Sf(X, Y), we obtain the formula (6) VxeK, АШх)<=Та(ю(А(х)) We now study the contingent cone to the preimage of a set by a smooth map. PROPOSITION 3 a. Let X and Y be two Banach spaces, L<=AT and MczY two subsets, and A a C* map from an open neighborhood of L to Y. We set (7) Then (8) K:={x6L\A(x)eM}=LnA-\M) VxeK, Гк(х)<=Т£,(х)пУ/1(х) ^Тм(А{х)) b. Let X and Y be finite dimensional spaces, A a map from an open subset Q.cX to Y, and L<=fí and M<=Y closed subsets of X and Y, respectively. We assume that there exists Xo ^Ln A ~ ^(M) such that (9) V/i(xo)Cb(xo) - CmÍAxo) = Y Then (10) Сь(хо)г>У.4(хо) ^Cm{Axo)^Ck{xo) ^ Proof, a. By proposition 1, Tk(x)c=7l(x), because К=L. By proposition 2 VA{x)Tк{x)cTA^к){Ax)<= Tm{Ax) because A{K)^M. Hence, Tk{x)<=VA(x)~'^Tm{Ax), and, consequently, formula (8) holds true. b. We introduce the set-valued map F from X to У defined by (11) F{x):=A{x)-M when xeL, F{x):=0 when xiL We observe that F~^(0)=K. We shall prove that there exists a neighborhood Uo of xo in L such that (12) Vx e Uo, d(x, F " ^)^l'dM(Ax)
CH. 7, SEC.6 CALCULUS OF CONTINGENT AND TANGENT CONES 441 Indeed, we take ;^o=0 and Xo eF“'(0). The inverse function theorem implies that F“ Ms pseudo Lipschitz around (0, xo). Then there exist a neighborhood U of xo, a ball of radius /• in T and a constant F > 0 such that 'iyerh, VxeF~^{y)r\U, if(x, F"'(0))<^|IfII We can choose U so small that ||/l(x)-.4(xo)||<r when x ranges over U. Any xeLnU belongs to F~ ‘(^(x)-7tM(^W)), and M(x) - n^iAixM ll^(x)- ^(xo)ll < r Therefore, we know that for all xe Uo'-=LnU, d{x, F-i(0))<c|M(x)-M^(x))-01| =dMiA{x)) c. Let Mo belong to CL(xo)nV/l(xo)''‘CM(^A:o). There exist a>0 and P>0 such that x+huo belongs to Uo'=LnU when ||x—Xoll<a and h^p. Since F~^{0)=Lr\A~^{M), we deduce from (12) d(x+huo, F~^{0)) dM(A(x+huo)) h ^ du(Ax+AV.4(xo)mo) \\A(x+huo)-A{x) - АУу1(хо)мо|| < he ; The first term on the right-hand side converge to zero, because VA(xq)uo belongs to Cm(Axo) and the second converges also to zero, because A is con¬ tinuously differentiable. Hence, mq belongs to Cf-i(0)(mo). ■ COROLLARY 4 Let X and Y be finite dimensional spaces, A a continuously differentiable map from X to Y, and M a closed subset of Y let Axq belong to M. If (13) then (14) Im VA(xo)-Cm{Axo)= Y VA(xo) Cm{Axo)'^Ca-í{M){xo) COROLLARY 5 Let L and M be two nonempty closed subsets of a finite dimensional space X and let Xo belong to LnM. If (15) Cl{xo)—Cjvf(xo)=X
442 CH. 7, SEC. 6 NONSMCX)TH ANALYSIS then (16) ^ ^ nA/('^o) ^ COROLLARY 6 L^i Kf(i = l, n) be n nonempty closed subsets of X and let xq belong to Ki. We posit the following assumption: (17) Then (18) Vt)i, ...,v„bX, f) (CKi{xo)-Vi)=h0 i=l П Ск,(хо)<=Сп;_.«, (л:о) i=l Proof. Let DdX" denote the closed vector space of constant sequences 3c:=(x, ., x). Then K is identified with K,-. We observe that Cd(x)=D and C n;., k, (5)= nlгCк ((jc). Assumption (17) implies that (19) Co{x)-Cx\:.,kXx)=D-Y[C4x)=X” i=l Therefore, corollary 5 implies that (20) fl Ск((л:)с:Сп;.,к, (x) 1= 1 that is, inclusion (18). We shall derive a calculus of contingent derivatives and derivatives of set¬ valued maps from the properties of the contingent and tangent cones. PROPOSITION 7 a. Let F be a set-valued map from X to Y and let В be a map from an open neighborhood Q. of \m F^Y to Z, Then (21) Vwo e X, S/B(yQ)-DF{xo, yo)(uo)^D(BF)(xo, Byo)(uo) V (22) F is Lipschitz around Xo e Int Dom F with compact values and dim У < +oo, then (23) Vmo 6 X, VB{yo)'DF(xo, yo)iuo)=D{BF)(xo, Byo){uo)
СН. 7, SEC. 6 CALCULUS OF CONTINGENT AND TANGENT CONES 443 b. Let F be a set-valued map from X to Y and let A be a 0 map from Xq to X. Then if Axo belongs to Dom F, (24) УмобХо, D(fy4)(xo,j'o)(Mo)=^i'M^o,3'o)(V^(xo)(Mo)) If we assume that either (25a) F is Lipschitz around Axq or (25b) V/4(xo) is surjective and dim Xq < +oo then (26) fuoeXo, D(FA)(xo,yoKuo)=I>FiAxo,yo)(yA{xo)uo) ^ Proof a. Let (1 x B) be the map {x, y)eXx Ci-^{x, B{y)) € У X Z The graph of the set-valued map G:=BF is related to the graph of F by the relation graph(G)=(l x5)graph(f). By proposition 2, we know that (1 X Vfi(yo))Tgraph(F)(xo, Уо) is Contained in Tgraph(o(A^o. Руо)- This implies formula (21). b. Let Wo belong to D{BF){xo, Bj;o)(mo)- There exist sequences h„^0, u„-*uo, and w„-»wo such that h„w„ belongs to B(F(xo+h„u^)~В(уо) for all n. Hence, there exists v„ e (F (xo -h h„u„)—yo)/h„ such that h„w„=В{уо+h„v„)—B( jo)- Since F is Lipschitz around xq, v„ belongs to F(x:o)—Уо + Лп1|мп||Д which is contained in a compact set, because the values of F are compact and the dimension of Y is finite. Hence, a subsequence (again denoted by) v„ converges to some vq, which belongs to DF{xo, уо)(мо)> and thus the sequence of elements w„ converges to ^B{yo)-Vo=wo. Therefore, D{BF){xo, Byo)(uo)f^'^B{yo)DF{xo, Jo)(mo) c. Let Axl: XqX Y-*X x У be the map defined by (A x 1) (x:q, y)={Axo, y). The graph of the map G — FA is related to the graph of F by the formula graph G=(/l X1)“ ^ graph F. By corollary 4, we know that Tgraph(C)(^0, Уо) = (У^(^о) X 1)" ‘ Tg„ph(F)(/lXo, Уо) which implies formula (24). d. To prove (26), let us pick wo in DF{Axo, Fo)(V^(a:o)mo)- There exist se¬ quences h„-^0 + , v„^VA(xo)uo and w„-»wo such that h„w„ belongs to
444 CH. 7, SEC. 6 NONSMOOTH ANALYSIS F(/4(a:o)+/i„u„)—>>o- Assume first that F is Lipschitz around xoi we deduce that F (xo)+h„v„)) <= F{A(xo+h„uo))+f\\A(xo+h„uo)—A{xo)—h„v„\ \ B Hence, there exists a sequence of elements w'„ converging to w© such that hnw'„e F{A{xo+h„Uo))—yo) for all n. Thus wq belongs to D(FAXxo, yoXwo)- Assume now that V/1(a:o) is surjective. Then corollary 5.5 implies the existence of a constant F>0 such that for n large enough, there exists a solution x„ to the equation y4(x„)=/l(xo) + /i„y„ satisfying ||x„—xol|<^/i„lli;„ll. Since dim Xo< +oo, we deduce that a subsequence of elements u„:={x„—xo)/h„ converges to some element uo- Since >^0+h„w„ e FA{xo+h„u„) for all n then Wo 6 D(FA){xo, >'o)(mo)- ■ We now investigate the chain rule formulas. PROPOSITION 8 Let Xq, X, and Y be three finite-dimensional spaces, F a set-valued map with closed graph from X to Y, and A a continuously differentiable map from Xoto X. Let Xo belong to A~' Dom F andуо eF(xo)- Wfe assume that (27) Then (28) Im V/<(xo)- Dom CF{Axo, Уо)=Х i. fuoeXo, C{FAXxo,yo){»o)^CF{Axo,yoXyMxo)uo) ii. fqoeY*, QFAXxo, yo)*ido)*='^Mxo)*CF{Axo, Уо)*Ш Proof Since gtaph(FA)={A x 1)“‘ graph(F), we apply the second part of proposition 3, which states that (V/l(xo) x l)”‘Cgraph(F)('^^o. To) is contained in Cgraph(f.4)(^o. To); that is formula (28), provided that property (29) Im(V.4(Xo) X 1) (7graph(F)(-^-^0> To) X X T is satisfied. But this property follows from assumption (27). Assumption (27) also implies that {CF(Axo, To)-V^(xo))* = V^(xo)*CF(.4xo, To)* thanks to proposition 3.3.14.
СН. 7, SEC. 6 CALCULUS OF CONTINGENT AND TANGENT CONES 445 PROPOSITION 9 Let F be a proper set-valued map from X to Y, К a subset of X, and let Xq belong to К n Dom F. Then (30) D(F\k)(xo, yo){uo)^F)F(xo, ;^о)1тк(*й)(ио) If X and Y are finite dimensional, the graph of F is closed, К is closed, and (31) Ck(xo) ~ Dom CF(xo, 3^0) = X then (32) CF{xo, yo)\cK(xo)bto)^C{F\K){xoy yo){uo) and for all qo € У*, (33) C{F\k){xo, yo)4qo)^CF(xo, yo)4qo)F Nk{xo) a Proof We observe that graph(F|x)=graph(F)n(K x У). Then "^graph(F K)(^0i Уо)^ TgrгphiF){Xo, Уо)^^{Тк{Хо) X У) from which we deduce formula (30). We observe that assumption (31) implies that ^ graph(F) (-^Oj Уо) C[c{Xo) X Y = X X Y Therefore, corollary 5 implies that Cgraph(F)(^0,Уо)(^{Ск{хо) X Y)c^C graph(F I K)ixo, Jo) from which we deduce formula (32). Assumption (31) allows us to deduce formula (33) from formula (32). ■ PROPOSITION 10 a. Let us consider n set-valued maps Fi from X to Y. (34) ,i=l Vho 6 X, D ( (J f i) (xo, Jo)(mo)= U F>Fi{xo, Jo)(mo) i= 1 b. Let us consider n set-valued maps Fi with closed graph from a finite dimen¬ sional space X to a finite dimensional space Y Let {xq, yo) belong to the inter-
446 CH. 7, SEC. 6 NONSMOOTH ANALYSIS section of the graphs of Ft. Assume that (35) Then (36) V(m„ Vi)&X X y(i = l,... ,/j), 3(mo, Do) 6X X у such that Do € CFi(Xo, Уо)(Мо + «f) - D( for i = l, . . . ,n Vmo e X, C ( П Fi I (xo, Jo)(mo) = П CFi{xo, Уо)(мо) .»= 1 i = 1 Proof. We note that graph(uF,)= ugraph(F,) and graph(nF/)= n graph (Fj) and apply proposition 3 and corollary 6, respectively. ■ We deduce at once a calculus of epicontingent derivatives and epiderivatives of real-valued functions V from the calculus of contingent derivatives and derivatives of the set-valued map K+ defined by V+{xY=V(x)^-R+ when V{x)< +00 and F+(x):=j0'when F(A:j={+oo}. PROPOSITION 11 a. Let V be a proper function from Y to Ru{+oo} and let A be a map from X to Y If Axo belongs to the domain of К then (37) Vwo e V{Axo){^A{xo)uoHD^{VA){xo){uo) b. Let X and Y be finite-dimensional spaces, A a map from X to Y, and V : y-)-[Ru{+oo} a proper lower semicontinuous function. Let xq belong to Л "^(Dom V). If we assume that (38) Im VA(xo)-C+ V{Axo)= Y then (39) Vwo e X, C+ V(AxoWA{xo)uo)> С^УЛШЫ and (40) d{VA){xo) ^ ^A{xo)*d V (Axq) We observe that assumption (38) is satisfied when either V is Lipschitz around A{xo), because in this case Dom C+V(Axo)=Y, or V^(xo) is surjective. Proposition 9 implies the following formula for epiderivatives of restrictions.
CH. 7, SEC. 6 CALCULUS OF CONTINGENT AND TANGENT CONES 447 PROPOSITION 12 Let X be a finite-dimensional space, V a proper lower semicontinuous function from X to R\j{ +oo}, and K<=X a closed subset. Let xo belong to Kn Dom V. We assume that (41) DoniC+K(xo)—CKi^o) — X Then (42) V« e Ck(xo), C+(K|k)(xo)(m)=S C+ F(xo)(m) and (43) a(KU)(xo)«=5F(xo)+iVK(^o) In particular, these formulas hold true when V is Lipschitz around xq. A PROPOSITION 13 Let V be a proper function from the product X xY of two Banach spaces to Ru{+oo}. We set (44) W(yy.= inf V{x,y) xeX If XySX minimizes x-* V(x, y) on X, then (45) '^veY, D+ W^(j')(i>)< inf D+V(xy, yXu, v) A ueX Proof Let u and v belong to X and Y, respectively. Then the following inequalities hold true: W(y+hv)-W{y)^ V{Xy+hu,y+hv)-V{xy, y) h " h This implies that for all {uo, Vo) in x 1^ D+ W^(y)(t)o)<-0+ V(Xy, y)(uo, uo) and, consequently, inequality (45). ■ Remark When t)=0, we obtain proposition 3.5 as a consequence. The analogous state¬ ment holds true for the supremum of a family of functions. Let (46) t/(_);):=sup V{x, y)= V{xy, y) xeX
448 CH. 7, SEC. 6 NONSMOOTH ANALYSIS Then (47) Vt; e y, D + C/( ^ sup Z) + F{xy, y){u, v) ■ ueX We now provide a formula on the epiderivative of a supremum of a finite number of functions. PROPOSITION 14 Let us consider n proper lower semicontinuous functions from a finite dimensional space X to IRu{h-oo}. We set i/(x):=maxi=i„„,„ii(x). We assume that n (48) Xo belongs to P) Int Dom Vi « = 1 Let us set J(xo):= {i = 1,..., Vixo) = i/(A:o)}. Assume also that (49) iVi e /(xo), Vw,- 6 X, there exists uq [such that uo - Ui e Dom C + K(^o) for all i e J{xo) then (50) and (51) Vwo e X, C+ U{xo){uo)^ max C+ K(^o)(wo) i e J(xo) 5f/(xo)=co U dVi(xo) i e J{xo) Proof. Assumption (48) implies that when i i /(xo), then {xo, U{xo)) belongs to the interior of EpVi, so that the tangent cone Cspivoixo, t/(^o)) is equal to X xR. Assumption (49) implies that V(«i, A;) e X X /?(/=1,..., rt), there exists (52) {(mo>Ao) such that (mo-Mi, Ao-A,) belongs to CEp^Vi)(xo> U{xo)) for all /=1,...,« Indeed, uo is given by property (49), and we take Ao:= max {C+V^xoKuo-Ui)+^i) i € J{Xo) Therefore, we are allowed to use corollary 6, since EpU=f]i=o Then n C£p(K,)(Xo, f/(^o))= n (^Ep(Vi){X09 UiXo))^CEpu{X09 U(Xo)) i 6 J(Xo) * = 1
CH. 7, SEC. 6 CALCULUS OF CONTINGENT AND TANGENT CONES 449 This inclusion implies inequality (50), from which we deduce inclusion (51), because the support function of a union is the supremum of the support func¬ tions. ■ We mention the following corollary. COROLLARY 15 Assume that the n proper lower semicontinuous functions are Lipschitz around Xo- Then (53) and (54) fuoeX, C+ U{xo)iuo)^ max C + 1^(a:o)(mo) i e J(xo) dt/(xo)«=co (J dViixo) »€ J(Xo) We now turn our attention to the behavior of epicontingent derivatives and epiderivatives of the sum of two functions. This time, we need the concept of strict epidifferentiability. PROPOSITION 16 Let V and W be proper functions from X to K u {-foo}. Then (55) D+K(;co)(«o)+Z)+fT(xo)(Mo)</)+(K + fT)(xo)(«o) Assume that W is strictly epidifferentiable at Xq and (56) Then (57) and (58) Dom C+T(xo)nDom B+W{xo)f0 C+(F-H fF)(xo)(«o)< V{xo){uo)+ C+ 1T(xo)(mo) d(V+ W)(xo)^dV{xo)+dW{xo) Remark Assumption (56) is satisfied when W is Lipschitz around xq. Proof The formula for epicontingent derivatives is obvious. We observe that for all uo belonging to Dom C+V{xo)<^ Dom 5+ W{xo), we have (59) C+(F + IF)(xo)(«o)< C+ V{xo){uo)+B^ fF(^o)(«o)
450 CH. 7, SEC. 6 NONSMOOTH ANALYSIS Now, let McDom C+F(xo)nDom C+W{xo) and Ae]0, 1[ be fixed. Since Dom B+W^(xo)=Int Dom C+W(xo), we deduce that (1-A)m+A«o belongs to Dom C+ F(xo)n Dom W(xo), so that formula (59) implies that C+{V + WtxoW Amo) < (1 - A)C + F (xo)(mo)+AC + F(xo)(mo)+B + if (x:o)(( 1 - A)m+Amq) By formula (3-27), we have B + lF(xo)((l - A)m -I- Amo) ^ (1 - A)C+ 1F(x:o)(m) -t- AB + iF(xo)(Mo) By letting A converge to 0, we deduce formula (57), which implies formula (58), because the support function of a sum is the sum of support functions. B Remark Observe that assumption (56) implies that (60) Dom C+F(xo)—Dom C+W(xo)=X In the same way, without using the inverse function theorem, we can prove that the assumption (61) Dom C+ F(xo)n V^(xo)” ^ Dom B+ W{xo)^0 implies that (62) C+(F -i- IF^)(xo)(mo)< C+ F(xo)(mo)+ C+ W(Axo)(S A(xo)uo) and that (63) d(V+ WA)ixo) c d F(xo) -I- VA(xo)*d IF(^Xo) without assuming that the spaces are finite dimensional. B
CHAPTER 8 Hamiltonian Systems We now apply the abstract methods of earlier chapters to a specific example of great practical importance: finding periodic solutions to certain differential equations. The differential equations we are investigating are of Hamiltonian type, that is, they can be written (1) dqi dH dPi dH. . for a suitable choice of the variables {q,p) 6 R^" and the function H\ [0, T] x R2"->R. In the particular case where the equations are autonomous, that is, if is a function of q and p only, the physical or mechanical systems they represent are conservative (2) -H{q(t\p{t))=0 The quantity H{q, p), which is preserved throughout the motion, is usually understood as the energy of the system. Since it does not dissipate, the system cannot come to rest, and its motion can be very complicated as ±oo. This is why it is so interesting to find periodic solutions: They are the only ones whose behavior can be completely described. Of course, this problem has been studied for a very long time, and giving a survey of the field in anything shorter than a full-length book is out of the question. What we do in this last chapter is to present three selected results, chosen to illustrate the abstract methods described in earlier chapters, par¬ ticularly Chapter 5. It may come as a surprise that abstract variational methods are helpful in solving problems in ordinary differential equations. This is the case here 451
452 CH. 8, SEC. 1 HAMILTONIAN SYSTEMS because Hamiltonian systems have a special, variational, structure: It has been known since the time of Fermat that the trajectories of equation (1) minimize a certain path integral. This is known as the least action principle, actually a misnomer, since nowadays we know that the trajectories of equation (1) may correspond to critical points other than minima. A precise formulation is given in Section 1. It turns out, however, that the least action principle is not the appropriate tool for our purpose: The associated functional on the path space is too com¬ plicated. For instance, in the cases we consider, it will be unbounded from above and from below. So we introduce another path integral that is much better behaved than the least action principle but yields the same critical paths. This new variational principle, which is in some sense dual to the least action principle, is explained in Section 2. In later sections, we apply the abstract results of Chapters 2 and 5 to obtain existence theorems for periodic solutions. We use the inverse function theorem in Section 3 and the Ambrosetti- Rabinowitz theorem in Section 5. In Section 4, minimization is enough. 1. THE LEAST ACTION PRINCIPLE We start from a function H: U referred to as the Hamiltonian. The integer n is called the number of degrees of freedom. It is assumed that the Hamiltonian is time periodic (1) 3T>0:H(t + T,x)=H(t, x) V(i, x) We are given a linear operator J: such that (2) -J=J* = J-^ Note that 7^4- /=0. We can always pick a basis of where the matrix of J will be (3) ;) where 0, 1, and -1 denote nxn diagonal matrices with 0, 1, and -1, respec¬ tively on the diagonal. In such a basis, the antisymmetric bilinear form {Jx,y) is written (4) {^JXy y)— ^ iyi^i + n -XiJ^i + m) i=l
CH. 8, SEC. 1 THE LEAST ACTION PRINCIPLE 453 We are Interested in the boundary-value problem (5) dt x(Q)=x{T) Clearly, since H itself is T periodic in time, any solution x to (5) will be extended as a T-periodic solution over U. Let us first dispel any notion the reader might entertain that just because the right-hand side is T-periodic, the equation x = JHx{t, x) will always have a T-periodic solution. LEMMA 1 Set ♦ " ft)- " H{t,x)= Yj -7rixf + Xi + „)+ Y fi(t)Xi + n i=l ^ 1=1 with OiSU and fi {2n/(o) periodic, l^i^n. If co eZOk for some k, and Jq//c(0^xp (ift)fcr)^/i:?^0, then problem (5) has no solution. In all other cases, it has at least one solution. A Proof. Let us write the equations i. jc,=ft),x/-n+yi(i) IL Xi-\-n— O^Xi 1 which are equivalent to !• Xi + 0)( Xi — fiif) 11. Xi -f- M — 0)iXi l^i^n In Other words, we are dealing with n uncoupled harmonic oscillators. The result then follows from the standard theory of second-order equations with constant coefficients. If co i Zco,- for all i, then there is precisely one In/co- periodic solution (the so-called particular solution). If ft) € Zco^, we have to expand the right-hand side f in Fourier series. If the coefficient of Qxp(icoi^t) does not vanish, we are in the so-called resonant case, and we know that all solutions must contain the nonperiodic term t exp(/ft)/^i). Hence, the result. ■ In the case of autonomous systems, that is, when the Hamiltonian H does not depend on time, there is a further complication. The differential equation
454 CH. 8, SEC. 1 HAMILTONIAN SYSTEMS may have constant solutions. To be precise, any point xo e where is called an equilibrium, and x{t)=xo is a solution of the equation. Such constant solutions are T periodic for all T. The question then is whether there are other kinds of periodic solutions. A classical method for approaching these problems is by means of the least action principle. It is generally attributed to Maupertuis, who was the first to formulate it in modern language, but mathematicians in the seventeenth century, particularly Fermat, were well aware of it. We shall state the least action principle in the language of the calculus of variations. Recall that an extremal of the integral (6) I Lit, ^(t), x{t))dt is a solution X of the corresponding Euler-Lagrange equation ddL dL , . ^ PROPOSITION 2 (LEAST ACTION PRINCIPLE) The solutions of Hamilton's equation (8) ~ = JHUt,x) are precisely the extremals of the integral (9) j [i( Jx, x) + H{t, x)\dt A Proof Write the Euler-Lagrange equation \jt Jx-^H':,(t, x) Since 7* = / " ^ = — y, this is precisely Hamilton’s equation ■ X is an extremal if the first variation of the integral at x, among all smooth curves satisfying appropriate boundary conditions, is zero. In other words, X is a critical point of the integral on a appropriate subspace of Note that X need not minimize the integral, and in most situations in physics, it does not. It would be more appropriate to speak of a “stationary action” principle. The least action principle must be tailored to suit the specific problem we
CH. 8, SEC. 1 THE LEAST ACTION PRINCIPLE 455 are working with. This means taking into account the boundary conditions and choosing the function space we want to work with, preferably a Hilbert space or a reflexive Banach space. In the case at hand, problem (5), we shall work in the Sobolev space T: IR^"), with 1 <a<oo. Recall that dt (10) T;IR2")=|x6L“ Denote by Wp\f the subspace consisting of T-periodic curves W^pi;*(0, T; e H^‘-«|x(0)=x(T)} The action functional on given by integral (9), is the sum of two terms 1 (12) <I)i(x)=2 j (Jx, x)dt (13) <D2(x)=jJ//(f,x(t))ifi The first term, ^i(x), is well defined, because all functions in are con¬ tinuous. It is easy to see that the canonical injection T;R^") is continuous, so is a continuous quadratic form on and hence, in¬ definitely differentiable. Recall that the Hamiltonian H(t, x) is assumed at this stage to be C^. The map x-> x(t))dt from C° to R then is also. Composing it on the left with the canonical injection we see that O2 is a map. The Frechet derivative at X e W^’“ is given by (see Chapter 1) ^'2{x)y= (//i(i, x(t)), y(t))dt We now state the least action principle in this setting PROPOSITION 3 Define the action functional O on by
456 CH. 8, SEC. 2 HAMILTONIAN SYSTEMS Then <I> is C^, and O'(jc)=0 if and only if x solves problem (5). Proof Let us write i>'(x)=0. For all y e W^r*. we have x)+U Jx, y)+x), y)~\dt=0 Integrating the second term by parts and taking into account the fact that X and y are T periodic, we obtain iy, x)+ Jx)dt=Q, for all y e Now x(i)) is a continuous function, and Jx belongs to If, so their sum is in L“. On the other hand, is dense in the dual if of L®, so we have {y, H'x{t,x)+Jx)dt=Q, forall^^eL^ a ^+/5 ‘ = 1 and hence. H'x{t,x)+Jx=Q in L® So x(t)=JH'x{t, x(f)) almost everywhere. Since H'xisC^ in the (t, x) variables, this implies that x is a classical solution to the differential equation x = JH'x(t, x). Since X 6 W^;®, we also have x(0)=x(T). ■ 2. A DUAL ACTION PRINCIPLE This section is concerned with nonconvex duality. We first prove abstract results and then apply them to the special case of Hamiltonian problems. We begin with a very simple setting. Let X be a reflexive Banach space, Q a continuous quadratic form on X, and F:A'^Ru{ + oo}a convex l.s.c. function. We are interested in the function <I):A'->IRu{ + oo} defined by (1) 0=Q + f It is essential to note that unless Q is positive, an assumption we do not make, <I) is not a convex function. We shall say that ti is a critical point of <I) if Q'{v)+dF{v) bO
CH. 8, SEC. 2 A DUAL ACTION PRINCIPLE 457 If F happens to be Gâteaux differentiable at p, then so is <I), and the definition becomes (l)'(i;)=0. In the general case, Q'{v)FdF(v) is clearly the generalized gradient of the function <I> at t; (see Chapter 7, Section 3), and the definition can be written 0 6 5<l)(p). It follows from Chapter 7, Section 6 that if v minimizes <I> locally, then p is a critical point. It is also true that all local maximizers are critical points. It is well known that there is a unique self-transposed operator A: given by (2) (Au, v) = [0(m + d)- Q(u)- 0(f)] (3) Q(v)=i(Av, v) Thus we have (4) A=A* (5) Q'(v)=AveX* We now turn to our duality result. THEOREM 1 Consider the two functionals <I> and ^ on V (6) (7) ^(v)=j(Av, v) + F(v) ^¥(v)=^{Av, v)-\-F*(-Av) If visa critical point of O, it is also a critical point of If v is a critical point of 'F and if (8) 0€lnt(^(2f)-hDomF*) then there is some w e Ker A such that u = v—w is a critical point of A Proof Let i; be a critical point of 0. By definition 0 6Av + dF(v) This can be written — Av edF(v) Using the Legendre reciprocity formula (theorem 4.4.4), we obtain V edF*( — Av)
458 cH. 8, SEC. 2 Hamiltonian systems Applying A to both sides, AveAdF*(-Av) Set G{v)=F'*'(-Av). We have AdF*(-Av)<=-dG(v) Hence, Av-i-dG(v) 3 0, and i; is a critical point of T. Let i; be a critical point of Condition (8) implies that (see corollary 4.3.6) AdF*(-Av)=-dG(v) so that the equation Av-\-dG{v) aO can be written A\_v — dF*{ — Av)'] 90 This means that there is some w eV such that v — dF*( — Av) Bw and Aw=0 So w e Ker A, and u = v — w satisfies ii edF*(—Av) Using the Legendre reciprocity formula (theorem 4.4.4) — Av edF(ii) But Av=A(v — w)=Au since Aw=0. Hence, Au + dF{u) bO and M is a critical point of d>. ■ We shall now apply this result to the least action principle for Hamiltonian systems. We first recast the least action principle to fit this framework. Set X:=I?(0, T; R^”), with l<a<oo. Its dual X* is lS(0, T; R^”), with a“^ ^ = 1. We introduce the closed subspace (9) U=1^ 6 L“(0, T; R^")| y{t)dt =o| The condition on y can be rewritten (z,y)=0 for all constant functions z.
CH. 8, SEC. 2 A DUAL ACTION PRINCIPLE 459 The constant functions z form a subspace of if that we identify with So, (10) L«o=(lR2n)x R2"=(U)^ For each y 6 LJ, define n>> to be the primitive of y with zero mean d -r (11) The map (12) — Tly=,y and J njy(f)ifi=0 41 x{t)dt is an isomorphism from onto LI x R^". Its inverse is the map (13) {y,Qi^Tiy+^ Replacing by L© x R^» ^ve reformulate the action integral as follows: (*T (14) [ [i(/v, Tly + ^)y Hit, Jo = \hiJy,m+Hit,uy+mt since ;; e which leads to proposition 2. PROPOSITION 2 Assume that H{U x) is a continuous function of (i, x\ convex with respect to x for every fixed t and (15) (j>{\x\HH{u xHil/(\x\) for all it, x) with s ^(l>is)^ +00 and s ^^(s) bounded when ^->oo. Then iy, 0 ^ critical point of the functional (16) <I>(y, 0=^ j\jy, ny)dt+ jjf/(f, ny + m on Lo xK^” if and only if x=II j;+^ solves the boundary value problem fx 6 x(i)) a.e. ^ ^ W0)=x(r)
460 CH. 8, SEC. 2 HAMILTONIAN SYSTEMS Proof. First note that (I)=Oi+<I)2, with quadratic and <I>2 convex. We have ^i(y, 0=^ {Jy, (-JTly,y)dt On account of the boundary condition, it is easily seen that J: Lo~*lJ is self-transposed. So (—70,0) is the self-transposed operator associated with 4>i on Lo On the other hand, the integral (18) =i> x{t))dt defines a convex and finite function on L?. The growth assumption on H implies that it is bounded from below. Using Fatou’s lemma, as in Chapter 1, we see that I is lower semicontinuous. Since it is finite everywhere, it is also continuous and, hence, subdifferentiable everywhere. We have (19) dl(x) = {z 6 L“|z(i) e dxH{u A:(i)) a.e.} The proof of this formula is rather technical and uses the measurable selection theorem (see Ekeland-Temam [1972] lemma 10.4.1). We are interested in the function ^2 = / ° ^ on LJ x with A{y, = Here, A sends Lq x into L^, so /1* will send L® into the dual of L% x R^", which is L§ X R^". An integration by parts gives A*{z) f z(t)dt) where Ilz denotes the primitive of z- 1/Til with mean zero. Since I is continuous, we have d<t>2iy, ^)=A*I{A{y, 0) = |(-riz, I z(f)i/f)|z 6 5/(n7 + ^) Let us now write that {y, is a critical point of 0+d^2(y, ^30
CH. 8, SEC. 2 A DUAL ACTION PRINCIPLE 461 This means that there is some 2 6 L® such that z{t)ed,H{t, Uyit)+0 — 711^ — 112=0 in 0 + f z{t)dt = Jo 0 in Differentiating the second equation gives 7;;+ 2=0, and substituting it into the first, we obtain y(t)eJ8M n>^(f)+^) [%(f)7f = 7 [' Jo Jo 2(f)i/f =0 Setting x(t) = ri7(i)+so that x =y, we obtain precisely problem (17). ■ We now wish to recast this equation into the form of theorem 1. The first step consists in eliminating ^ from the functional d>(j, <^), thereby obtaining a function ^(y) defined on L% only. COROLLARY 3 Take H{t, x) as in the preceding equation, and consider the functionals (20) ^(j)=^ I (Jy,Uy)dt-\-G(Yly), yeLo (21) G(A:)=min f H(t, x{t)-\-^)dt, xelS Jo Then G is a convex, continuous function on LP, and y is a critical point of^ on LI if and only if there is some e R such that x = Oy + (^ solves problem (17). ▲ Proof LetG(x):=min{/(x + (J), (^eR^"}. First, we observe that the infimum in (7(x) is achieved. This follows from the fact that the function H(t,x(t)+m=i{x+i) Jo is continuous on R" and goes to +co when |<i|-^oo.
462 CH. 8, SEC. 2 HAMILTONIAN SYSTEMS We compute dG{x). Set V{x, <^)= I(x-\-^)- Since G(x)= min V(x, V(x, then proposition 4.6.1 implies that gedG(x) if and only if (^,0) belongs to dV(x, i). Subdifferential calculus shows that dV(x,0=ly, [ Jo )yedHx + i) qedG(x) o q edl{x+^)&nd [ q{t)dt=0 Jo So, It follows that r is a critical point of O if and only if 0 e JTly+115(7(11;^) since n * = - IT This means that there is some ^ 6 and q eH such that 0 6 JTly + Tlq G(Tly)=my + ^) [ ^(r)i/i=0 Jo qBdI(Uy + ^) In other words, q 6 L©, JTly = —Tlq, and by formula (19), q{t)edMt,^y{t)+^) a.e. Setting x(r)=riy(f)+<i, this can be written X e JdxHit, jc(i)) a.e. and a:(0)=x(T). This is the desired result. ■ We now apply theorem 1 to obtain a dual version of the least action principle in Hamiltonian mechanics.
CH. 8, SEC. 2 A DUAL ACTION PRINCIPLE 463 THEOREM 4 Assume that H{t, a:) is a continuous function of (i, x), convex with respect to x for every fixed i, and (22) <^(||x||)<//(i, x)<«/r(|x|) for all (i, x) with s + <X) and s ^^{s) bounded when j-»oo. Define a functional 'P on L%by (23) ny)+H*{t, -Jy)-\dt where H*(t, -) is the conjugate of H{t, •) with respect to x. Then y is a critical point of ^ if and only if there is some such that x=Hy-\-^ solves the boundary value problem (24) x=JH'x{t, x) x(0)=x(T) Proof Apply theorem 1 to T, with X=Lo, A = —JTl, and (25) F(y)=j^H*(t,-Jy)dt Let us compute F*. We first extend the function f to a function F defined on the whole of L“ by the same formula (26) F(y)=j'^H*(t,-Jy)dt So F is the restriction of f to LI Since L% is the kernel of the map e:y-^ily(t)dt, we can write, using the indicator function of the set {0}: F(y) = F{yH^{o}(0y) We apply corollary 4.4.12. We have to check that 0 belongs to the interior of 0(Dom F). By condition (22), x) is bounded on every bounded subset of UxR^\ It follows that F(0 is finite for all constant functions <^. Since the restriction of 9 to the constant functions is the multiplication ^ 6(Dom F) is the whole space Consequently, (27) F*iq)= niin F*(q-^0
464 CH. 8, SEC. 3 HAMILTONIAN SYSTEMS because the transpose 6*: is the map associating the constant function 0)=^ in iS with 6 Define A::Z?->IRu{+00} by (28) We have (29) K*(x) Jo x)dt Hence, since f(>^)=iC(—/>») and J is an isomorphism, (30) F*{x)=K*i-Jx) (31) '*(q)= min [ ieR2»Jo Hit, -Jx+m Using the definition of G, formula (21), we obtain (32) F*iq)=G(-Jx) The dual formula «h associated with 'P by formula (7) now reads (33) <l>(y)=^ j + =- r 2 Jo {Jy, ny)dt-¥Giny) This is just what we called <I)(j) in corollary 3. Its critical points correspond to solutions of (17), and the result is proved. ■ In the following sections, we shall find critical points of T by three different methods: the inverse function theorem, global minimization, and the Ambrosetti-Rabinowitz theorem. 3. NONRESONANT PROBLEMS We shall prove the following result stated in theorem 1.
CH. 8, SEC. 3 NONRESONANT PROBLEMS 465 THEOREM 1 Let H e R). Assume that two real numbers a and P can be found with (1) 2tc 0 2tc y w<a^/?<y (w + 1) for some integer m'^Q (2) a/^ff;Ax)^P/ for all X € Then for all f € L'(0, T; R^"), the boundary value problem (3) has one solution at least. x=JH'^x)+f{t) x{0)=x(T) Inequality (2) must be understood in the sense of symmetric matrices: /li </<2 if the eigenvalues of (A2—A1) are nonnegative. Together with (1), it is a nonresonance condition ■ We begin by casting the problem into Hamiltonian form. Set (4) and denote by F the antiderivative of /—m with zero mean value d (5) —F(t)=f{t)-m and F{t)dt=0 dt Jo Set (6) z{t)=x{t)- Fit) Note that z(T)=z(0) if and only if x(T)=x(0) We introduce the Hamiltonian (7) K(t,z)=H(z+F(t))-(Jm,z+F(t)) Problem (3) is equivalent to the following: (8) z(t) 6 JK'ft, z(t)) a.e. z(0)=z(T)
466 CH. 8, SEC. 3 HAMILTONIAN SYSTEMS The function F is continuous, and inequality (2) sets growth conditions on H i44^+cc'\\x\\+o^’’^H{xHmxV+P’\\4+P" for adequate constants a', a", p\ and p". So the new Hamiltonian K(t, x) satisfies all the assumptions of theorem 4 with a=2=p. Solving problem (3) is therefore equivalent to finding critical points of the functional (9) \\uJy,ny)+K*(t, -Jy)-]dt Jo on the space Lq. Here K*(t, •) is the conjugate of K(t, •). LEMMA 2 For every fixed t, the function K*{t, •) is on and (10) y)=K'Uu zY' xY ‘ where y=K'i(t, z), z = Kf{t, y), x=z+F(t) in We also have (11) ^^KK*;{t,yH^-I foralliuy) Proof. By definition of the conjugate function, we have (12) X*(i, y)=yz-K{t, z), with y = K',(t, z) For fixed t, we can apply the inverse function theorem to y and z, since the derivative KUU z) = H'Uz + F{t)) is invertible by assumption (2). We have ^ dzi' ^yj =KAu zY with obvious notations. By writing z in terms of y in the right side of equation (12), it follows im¬ mediately that K*(f, •) is C^. Differentiating the Legendre reciprocity formula, (13) /Cr(i,K;(i,z))=z yields equations (10), from which (11) follows. ■
CH. 8, SEC. 3 NONRESONANT PROBLEMS 467 LEMMA 3 The function ^ is and twice weakly differentiable Gateaux (14) i^"iy)zu Z2)= \ [^{Jzi,Uzz)+(K*y(t,-Jy)Zi,Z2)'\dt Jo Proof. The first term in 'P is quadratic, continuous, and, hence, C". The second term satisfies all the assumptions of example 5 in Section 1.4. Hence, the result. ■ Now remember that - 711 is a compact self-adjoint operator on Ll, so it has a sequence A„-»0 of real eigenvalues, and there is an orthonormal basis of eigenvectors. An easy computation from the equation Xy=-JY\y gives (15) kel, kfO All these eigenvalues have multiplicity 2n, and the corresponding eigenspaces are (16) Ek = \ >'(f)=exp( -Tkn^] K eR^" They are of course orthogonal, and we have the Hilbertian sum (17) (18) U = 0 £fc = £'0£", with kfO k=-l /-(m + 1) \ / 00 £'= 0 £* and £"= 0 £J0 0£, k=-m \k = -ao / \k=l Any z 6 Lo can be written z=z'-)-z" with z e £' and z" 6 £". By lemma 2, we have ')dt (19) /c=-n,A:27r J (20) ('P"(x)z", z")= Z k>l 1 T kin ")dt k^-m-I Here is where we use assumptions (1) and (2), substituting inequality (11) into the preceding equations, we obtain
468 CH. 8, SEC. 3 HAMILTONIAN SYSTEMS /^1,2 ^ '^Ikn OC k= -m Zk (-T . 1\„ , (21) 2nm a.) (T"(x)z", z")> S ^ i 2^^ Jo P )dt /c< -m- 1 “I. I (j- P k>l *< -m- 1 \P 2(m + l)Ti Z* (22) .1- P 2{m-\-\)n ."II 2 Set a={T/2nm—il(x) and i?=(!/)?—T/2(m+l)7r). By assumption, ^>0 and 6>0, and we have (23) Vz'eE', (24) Vz"g£", It follows that ^"(x) is nondegenerate. LEMMA 4 There is some k>0 such that (25) Vz6£, |'P"(x)zl|^itizl| A Proi^ Assume otherwise. Then there is a sequence z„=z'„ + z|| such that: (26) ||z„p = ||z;,P + |zi'P = l (27) 'P"(x)z„^0 in L§
CH. 8, SEC. 4 RESONANT PROBLEMS 469 We also have (z'„, 'i"'(x)z„)=(z’„, 'i"’(x)z'„)+(z'„, 'I'"(x)z:) (28) <-a||z;p + (z;,T"(x)z;') (z:, '¥"(x)z„)=(z", '¥"(x)z'„)+(z:, 'f"(x)z'') (29) Mzl'V'\x)z'„)+b\\z';,f Letting «->00, this yields (30) lim inf a||zi,|| ^ <lim inf(z;„ '¥'\x)zl) (31) lim sup(z;.', '¥"{x)z'„) < lim sup - ft || zj,' P But (z'„, T"(x)z;')=(z", 'i"'{x)z'„), so that lim(z;, 'P"(x)z")=0 Inequalities (30) and (31) now give us lim sup a||z;,P<0 lim sup ¿)||z;;p<o which is impossible, since ||z'„P + ||z^'P = l. Hence the result. ■ We now have the situation in example 8, section 5.5, so 'P has at least one critical point y, and by theorem 2.4, there is some 6 such that x=Tly + i solves problem (3). 4. RESONANT PROBLEMS We shall prove the following result of Clarke and Ekeland. THEOREM 1 Let H 6 C^(R^", R)- Assume that it is convex and (1) 0(11 a:|| ) ^ //(x)^ il/{\\x\\) for all x with s~^(l)(s)-^ 4-00 ands~^il/{s)-^k/2>0 when s^co. Assume kT<2n. Then for all f e L^(0, T; the boundary value problem
470 CH. 8, SEC. 4 HAMILTONIAN SYSTEMS (2) has one solution at least. x=JH'^(x)+f{t) x(0)=x(T) Note that if we take w=0 in theorem 3.1, we have 0<a^P<2nT ^ and (xI^H'xx{x)^^I. Setting a=0 is not allowed in theorem 3.1, but the situation we obtain falls within the range of theorem 1. In other words, theorem 1 allows us to “touch” the m=0 resonance. The proof is quite straightforward: We shall prove that the functional 4^ of theorem 2.4 has a global minimum on Lq. The point where it is attained is the critical point we were looking for. We introduce the same Hamiltonian K{t, x) as in the preceding section and the dual action function T defined by (3) ikJy, ny)+K*(t, - Jy)-]dt We begin with some estimates. LEMMA 1 in^Njl Proof. Expand y in Fourier series ^ / 2iknt\ ks7L Note 3^0 =0 since y e Lq. Now integrate termwise We have „ y. T (2iknt\
CH. 8, SEC. 4 RESONANT PROBLEMS 471 LEMMA 2 For any choice ofk'>k, there is some constant c such that (4) K*(t, x)^^\\x\\^-c for all if, x) 2k Proof. Recall that m, z)=H{z + Fit)) -iJm,z+Fit)) Since Fit) is continuous, it is bounded on [0, T] by some constant M, and we have sup Kit, z)^i;i\\z\\+M)-II ym||(||z|| -M) t So, k lim sup ||z|| sup K{u z)^- Ilzll-^OO t 2 This means that for any choice of k'>k, there will be some constant c such that K(u z)^y Taking conjugates, we obtain the desired inequality. ■ LEMMA 3 'F attains its minimum on Lq. A Proof. Since /: < 2tcT" S we can choose k' so that k<k'<2nT-^ It follows that jjiJy, w+ - w dt ^ - 4k
472 cH. 8, SEC. 4 Hamiltonian systems Now let y„ be a minimizing sequence for T The sequence is bounded from above by some constant c'. Substituting this into the preceding inequality -cT 2\k‘ Since k! <2nT S the coefficient on the left is positive, from which it follows that (6) We now extract from the bounded sequence y„ a weakly convergent sub¬ sequence, which we still denote by yn- Let;; be its weak limit yn-^y weakly The second term in ^ is convex and continuous (see Section 2) and, hence, weakly lower semicontinuous lim inf I X*(i, -Jy„)dt I K*(i, -Jy)dt The first term is weakly continuous. To be precise, we have Uy„^Ily Jyn)-{^y, Jy) = {^yn-^y. Jyn)-^(^y. Jyn-Jy) The first term on the right converges to zero because ||i;^r,|| stays bounded (Banach-Steinhaus theorem), and the second term converges to zero, since \jyn-Jy)-^^- Finally, lim "V(yn)^^{y) So y must be a minimizes Theorem 1 now follows from theorem 2.4.
CH. 8, SEC. 4 RESONANT PROBLEMS 473 We now turn to autonomous problems, that is, we set / =0. Of course, theorem 1 still applies, but the solution we obtain might be the trivial one (as pointed out in Section 1). To exclude this solution, we need one more assumption onH. THEOREM 4 Let H e U) be convex and H{x)^H{P)=0. Set (7) (8) 1 K lim inf -2 min{//(A:)| ||jc|| =r}=— »•-►0 2 1 k lim sup^max{7/(x)| \\x\\ =R}=^ R-^ao ^ ^ Assume 0<k<K^oo. Then for all T in the open interval {2nK ^ 2nk ^), the equation x=m{x) has at least one periodic solution with smallest period T. A If a: is T periodic, it is also 2T periodic, 3T periodic, and so on. The solution that theorem 4 gives us will not be T/2 periodic or T/k for any integer k>\. In particular, it cannot be constant. To prove theorem 4, we first apply theorem 1. We have found a minimizer for T, and x=^Ylyf-^ is a T periodic solution. Assume that x is actually T/k periodic for some integer/:> 1. Then so is j;. Set yk We have 6 Lq. We claim that 'P(ji()< which contradicts the fact that y minimizes T and concludes the proof. Indeed, = f [HTly„ Jy,)+H*{- Jy)-]dt Jo = 1^5 (n j kds dt
474 CH. 8, SEC. 4 Hamiltonian systems = ^ (n^a), jy(t))+H*{ - jy(t))dt Jo 2 =ytTO + (l-A:) \\*i-Jy)dt Jo Now H{x)~>H{Qi)=Q, so H*{z)^H*{Qi)=Q. Hence, (9) At this point in the proof, we need an intermediate result. LEMMA 5 'P(j)=minT<0 Proof. Clearly, *P(j)<'P(0)=0, but this is not sharp enough. Fix i with ||ii|| fO. With every i>0 associate the path ( 2nJt\ ^ y,{t)=sexp\^- We have eL§ and | =i||<^||r‘^^ Moreover, -2nJt\ , T i {Jy„ TlyM exp j ^ y exp ( K | di — InJi 2 2n " " Now it follows from our assumptions that lim sup4{^*WI ll^ll='-}=^ r-o 2/C Pick some K' with 2nT~^<K'<K, and choose |j| so small that H*{ - Jys)dt < ^ II JysW' dt
CH. 8, SEC. 5 TRANSRESONANT PROBLEMS 475 We have '¥(ys)= ^y^)dt+ [ H*(-JyMt Hence, '*'(>'*)< 0 as desired. ■ Since T(>^)<0, we have k'i’(y)<'¥{y) whenever /:> 1 and hence the desired contradiction that concludes the proof of theorem 4 (10) COROLLARY 6 Assume H is convex and for some a with 1 < « < 2, (11) (12) liminf r‘“min/i'(x)| ¡A:|=r}>0 r^O limsup/? “max.H(x)| l|x||=/?}<+00 R-* ao Then for all T > 0, the equation x=JH'x(x) has aperiodic solution with minimal period T ^ 5. TRANSRESONANT PROBLEMS We shall now deal with the case when H"(x) ranges from zero to +oo. This is a more difficult situation to handle, and the nonautonomous case, for instance, is not fully understood. We shall be content to give a simple result. THEOREM 1 Let H 6 C(IR^”, (R). Assume that (1) H is strictly convex (2) H(x)^H(0)=0 Assume moreover that for some jS > 2, w£ have (3) {x,H'{x))>pH(x)
476 CH. 8, SEC. 5 HAMILTONAIN SYSTEMS Then for any T>0, the equation (4) x = JH'{x) has at least one nonconstant T-periodic solution. A Assumption (3) can be put in a more readily accessible form (5) foralU>l, xeU^" Since /?>2, we obtain (note X>\ and not A>0), (6) ||jc|| +00 when ||x||->oo (7) when |l.x||-»0 We shall first prove the theorem under the added assumption that for some constant k>Q (8) lIxP, for all a: 6 R2" The general case will follow later. LEMMA 2 Set min {/7(a:)| ||a:|| =1}. It is strictly positive, and vie have (9) Vx, H(x)^j{\\xY-l) (10) 1 |x|l<l=>||f/'(x)|<^(^''2''-c^)||xr-‘ Proof The first inequality follows immediately from assumption (3). For the second, start with the convexity inequality (H\x\z-x)^H(z)-H{x) Take the supremum over all z such that ||z—x|| = ||x||. Using assumption (3), we obtain which is the desired result.
CH. 8, SEC. 5 TRANSRESONANT PROBLEMS 477 LEMMA 3 //* is everywhere finite and C^. k Proof. Taking conjugate functions on both sides of inequality (9), we obtain (11) I where a is the conjugate exponent of /3, so a“ * 4-/?" ^ = 1. The property follows from the fact that H" is positive definite, as in lemma 3.3.1. ■ LEMMA 4 H* satisfies all the following: (12) (13) (14) ccH*{y)My, H*’(y)) 1 ||;;||>l=i>||//*'(>^)l|<-(2V-“-^-“) II«-1 Proof. The first inequality follows from (8) by taking conjugates of both sides. The second follows from assumption (8) and the Legendre formula H*{y)={x, H'(x))-H{x) =OL-\H'(x),x) =«-~\y,H*'(y)) with x=H*'{y) and j=if'(x). The third inequality proceeds from (11) and (12) as in lemma 1. ■ We can now proceed to the heart of the proof. By theorem 2.4, we shall be seeking a critical point of the dual action functional on Lq (15) 'P(3')=|^ i^{Jy,Yly)+H*{-Jy)-\dt The origin ;;=0 is an obvious critical point. It does not interest us, since the corresponding solution A: = ri;;4-^ is constant, that is, the equilibrium. We are
478 CH. 8, SEC. 5 Hamiltonian systems looking for another critical point, which we shall find by applying the Ambrosetti-Rabinowitz theorem 5.5.5. LEMMA 5 Constants y>0 and r>0 can be found such that (16) \\y\\=r=>'¥{y)>y (17) 0<|;;l|<r=i>4'(:(;)>'P(0)=0 A Proof. Clearly, ^(0)=0. By condition (12), we have Ù0 ^ a Using Cauchy-Schwarz on the first term, we obtain invL ^iy>^\\y\\l-h\\ Now n sends Lq into which injects into and J^l|« for some constant b. Hence, Since a < 2, the first term outweighs the other near the origin, and the result follows. ■ LEMMA 6 T here is some y eU v^here ;^) < 0. A Proof Define ys as in lemma 4.5, so that On the other hand, by inequality (11), we have [\*(-jy,)dt^i^\\y:\%+^T Jo « Finally, nys)^Ç UVts^- ^ ll^ll v+^ T
CH. 8, SEC. 5 TRANSRESONANT PROBLEMS 479 Since a <2, we have 'P(>'s)<0 for large s. Hence, the result. ■ We now have the situation in examples 6 and 7 in Section 5.5, so that there exists at least one nonzero critical point y and the result is proved. We can also locate this solution in the phase space by giving an upper bound for its energy level h. We first relate h to the critical value v='I'fj;) we have just found, starting from (18) [h{Jy, Uy)+H*(-Jmdt Jo and recalling that x=TIy+ie dH*{ - Jy), since is a critical point. Hence, V= ^\{Jy, Uy)H-Jy, + № = - \hiJx, X)+Hix)\dt = [\mx),x)-H{x)-\dt (19) I H{x)dt But H(x{t))=h along the trajectory. Hence, v^h(li/2 — l)T. We then estimate v; remember that it is found by applying the Ambrosetti- Rabinowitz theorem 3.5.5, which defines v as follows; (20) u=inf max iA(y(i)) ye r 0<5< 1 where F is the set of all paths y: [0, l]->Lo such that 'y(0)=0 and y(l)=y, the latter being a fixed point where il/{y)<0. Taking the path y: s-^ys described in lemma 5 gives max 1 <max 04s^i I a " " 4n ) P ^^-2a/(2-a)j2(l-a)/(2-ap^^a/(2-a)^l _ j (21) - ^+jT
480 CH. 8, SEC. 5 HAMILTONIAN SYSTEMS Finally, we obtain an a priori estimate for the energy level (22) 2/(i-2,^-W-2)(2„)/./U>-2) + This allows us to prove theorem 1 for the general case, that is, when the additional assumption (8) is not satisfied. Indeed, note that the constant k in (8) plays no role in the energy estimate we have just found. Given any H in U), satisfying conditions (1), (2), and (3) but not (8), we compute the right-hand side of (22) and call it ho [recall that is the minimum of H(x) over the unit sphere |a:| = 1]. With any constant r>0, associate the convex function Gr(x)=sup{(A:,y)-/f*(>')| ||y|| K,={xW\H'(x)\\^r} (23) to the set (24) and the numbers (25) Mr=vci2i\{H{x)\x e Kj, mr=mm{H{x)\x i K,} We have G,.(a:) = H{x\ for all x e K,. G^(A:)<r||x|| for all x e Define (j)/. [0, oo)-^IR as follows: (26) \(j),{t)=at^- if b if t'^nir the constants >0 and ¿>0 being adjusted so that 0, is at i=m, and hence everywhere. Since P>1 and m,.> 1 (for r large enough), 0,. will be convex and increasing. Now consider the Hamiltonian (27) H,{x)=MG,(x)) It is and strictly convex. We have (28) H(x) H(x)=Hy(x)
CH. 8, SEC. 5 TRANSRESONANT PROBLEMS 481 (29) H{x)'^nir^ Hr{x)^a{r\x\ +Mrf — b It follows from condition (3) or (5) that (30) ||x||^l=^//(A:)<||x||^max{/f(A:)| ||x|| = l} Conditions (28), (29), and (30) together imply that for any r, we can choose some kr large enough so that (31) 1 So condition (8) will hold for H,. The proof of theorem 1 for H now runs as follows. Choose r>0 so large that ntr > ho- Apply theorem 1 to and find a T-periodic solution of x=JH,{x) with energy level It follows that, for all t Hr(x(t))^ho<:mr By (27) and (28), this implies that HMt))=mx(t)) So in fact we have solved the equation x = JH'{x) as desired. ■
G О О а D Un 0 Он Он D СЛ 1 сл С О и g Ö0 с (2 i сл 'S il .ÍN «л § > сН о 'S 8 1« & S сл 75 ё О S W X й ^ в" § я о (L> Й 8 н—• Й 0> ÖJ) Й Й о и 0> Й о и Й 4> W) Й cd Н Й о Он о 3 b V/ S b и 2 и 2 и ш b с Он г Sv b ^Z)V ОТ ICJI II "»'S; -5 ч ^ "cd ^*5 b b '-' Л J!^ 2 2 ‘“ä ÏÏ b II II I 3 ^ <s Й ^ Ш 3 iS ^ Д 3 c 2 II и 3 и 0 1 о 3^ « nÍS у с líí II i< b “= Wn >3 b ‘■CÏ = CT и í< и 'Ел 'Сл II >- 3 и b ш 3 -I p 3 II ^2 Ш -- is 3 и 2 > 6 о & и В ^ о fi сл ^ 2 cd тз :: b, Ü^oíS. " kJ 3 î3? 'ï' Ol сл nJ * ^ I b ^ H I к.) Ш fi ^ Í't: ^ Ш I Й 2 "ÍI is о m b II 1-? Л ^ ' I II r ß 5 Ш C ^ Ш V is о ^ 3 S' I 3 > 3 >h'' II f ^ X I 3 t> и M ^ is > в, £ ^ K- .S2 2 ^ >< 482
§■ Ü Ш cd О i Си cd s •s > I ü CO cd cd > U û
i Q Ш o X c3 T 4- 3 Î c o ü G D Un ’S 'л > 13 <ü «-I s. o Wn Oh Cd cd > 'П <u -О ‘Бн ш I s 484
485
Comments CHAPTER 2: SMOOTH ANALYSIS Section 1 The inverse function theorem is so classical that we don’t even know who started it. In modern times, the need has arisen for an inverse function theorem that would cover situations where the underlying space is or C" (which are not Banach spaces) and the linear inverse loses derivatives (sends onto say). The answer is the Nash-Moser inverse function theorem, where Newton’s method, because of its very rapid convergence, plays an essential role in the proof. It falls outside the scope of this book; see Moser (1973) and Nirenberg (1974) for an exposition; Hamilton (1982) for a survey; Ray (1983) for an origi¬ nal proof; and Arnold (1978) for applications to classical mechanics (the Kolmogorov-Arnold-Moser theorem). Section 2 Milnor’s proof is in Milnor (1978). Other proofs use either a combinatorial lemma of Knaster-Kuratowski-Mazurkiewicz type (1926) or homology theory under more or less clever disguises (Brouwer, 1910). Section 3 Theorem 1 is due to Whitney. Our proof of the infinite-dimensional Morse lemma (theorem 8) follows Lang. In finite dimensions, the Morse lemma is but the first step of singularity theory, a good introduction to the subject being Brocker’s book. For a proof of the Morse lemma in the case on a Hilbert space, see Cambini (1973). Section 4 This follows Cerf’s exposition in F4=0. 487
488 COMMENTS Section 5 The results in this section are taken from Crandall and Rabinowitz (1971,1973) and so are the proofs. Nirenberg’s exposition (1974) follows the same pattern. Theorem 1 will also be found in Prodi and Ambrosetti (1973). See the books by Chow and Mallet-Paret (1982), Hassard et al. (1981), and looss and Joseph (1980) for further information on bifurcation theory. Section 6 and 7 Transversality theory is a creation of René Thom. See Abraham and Robbin (1967) for a more complete exposition, and a proof of Sard’s theorem. We have followed this excellent reference in our presentation. Smale’s method was given in (1976c). Here we have chosen different boundary conditions. See Milnor (1975) for the impact of the Sard-Brown theorem on topology. Smale’s method itself came out as a smooth version of Scarfs al¬ gorithm (1967b) for computing fixed points. Continuation methods such as Smale’s are very popular nowadays because they are simple, flexible, and often very efficient in solving numerically a large variety of nonlinear problems see Robinson (1980) and the references therein; Garcia and Gould (1978, 1980), and Lasry and Siconolfi (1983). CHAPTER 3: SET-VALUED MAPS Definitions of continuity for set-valued maps and most results of the first section are now classical. For more details, we refer to the forthcoming book by Rockafellar and Wets. Some related material appears in the book by Berge (1954). The extension to set-valued maps with closed convex graphs of the closed graph theorem and the open mapping principle was found by Robinson (1976a) and Ursescu (1975). The concept of transpose of closed convex processes was used by Rockafellar (1967b) for applications to economic models (1974b). (See also Makarov and Rubinov, 1970,1973). The economic model studied in the fourth section is due to von Neumann (1937) and has been thoroughly studied since (see Gale, 1956; Nikaido, 1968, etc.). The extension to the nonlinear case is due to Ky Fan (1958). Frobenius’s theorem goes back to 1908. This section is taken from Aubin (1978c). The book by Castaing and Valadier (1971) gives a thorough study of measurable set¬ valued maps. A study of continuous selections of set-valued maps can be found in the first chapter of the book by Aubin and Cellina (1983). CHAPTER 4: CONVEX ANALYSIS AND OPTIMIZATION Convex analysis and optimization are by now very established. It started with the fundamental work of Fenchel (1949,1951), followed by the works of Moreau
COMMENTS 489 and Rockafellar, and summarized in Moreau (1967) and Rockafellar (1970a, 1974a„ 1976). We refer to these books (among others) for further comments. We mention only that the duality theory for constrained minimization problems was first established in the special case of linear programming, thanks to the pionneering work of Dantzig (1963) and Kuhn and Tucker (1951), following earlier ideas of J. von Neumann (unpublished). The concepts of derivatives and co-differentials was introduced in Aubin (1981) and Pchenitchny (1980). The computation of tangent cones to Lr\A~^ {M) is due to Aubin (1979c). Section 3.7 on the regularity of the set of solutions and Lagrange multipliers of convex minimization problems is taken from Aubin (1982a). CHAPTER 5: A GENERAL VARIATIONAL PRINCIPLE Section 1 The most important results of this section are theorems 7 (the Banach contract¬ ion principle) and 14 (the Caristi fixed-point theorem). The first one is classical and has already been used in Chapter 1 to prove the inverse function theorem. The second one is recent (Caristi, 1976; Kirk and Caristi, 1975) and came as a surprise to experts: it was the first fixed-point theorem that did not require the self-map to be continous. Since then, this vein has been exploited, with further interesting results, by Kirk, Ray, and their students. The connection between Caristi’s fixed-point theorem and Ekeland’s e- variational principle was first noted by F. Browder. The approach we use here, through dissipative dynamical systems, is due to J. P. Aubin and J. Siegel (1980). Our exposition follows theirs. Section 2 Theorem 1 is due to F. Browder (1965a). See also related work by Browder (1965b), Edelstein (1963, 1965), and Kirk (1965). Theorem 7 is a nonlinear ver¬ sion of the mean ergodic theorem (von Neumann, 1932). It is due to Bâillon (1978). The proof is due to Pazy (1979). Section 3 The e-variational principle (theorem 1) is due to Ekeland (1972, 1973, 1974). Many consequences, including some we describe in this book, are given in the survey paper Ekeland (1979a). Example 9 is classical in the Russian literature on the calculus of variations (see the excellent book by Ioffe and Tihomirov, 1974). Much more can be said about this problem and related ones : it would take us into relaxation theory, which is described in the books of Ioffe and Tihomirov and Ekeland and Temam.
490 COMMENTS Section 4 Theorem 3 is due to Br^ndsted and Rockafellar. Proposition 4 is an abstract version of Temam’s results for the Plateau problem (see Temam, 1971 or Chapter 5 of the book by Ekeland and Temam), which we describe in the remain¬ der of the section. Section 5 Palais and Smale introduced condition (C) in a series of papers where they extended Morse theory and the Liusternik-Schnirelman theory to infinite¬ dimensional manifolds. The condition (weak C) appears here for the first time. Proposition 4 and theorem 5 strengthen existing results. Theorem 5, with the assuption that U be C^, is a celebrated result of Ambrosetti and Rabinowitz (1973). We found a simpler proof, relying on the s-variational principle instead of the so-called deformation lemma, and Brézis showed us how to extend it to the case when U is not . Example 8 is found in Ekeland (1979a). Section 6 The results of this section are due to Ekeland and Lebourg (1976) (see also Ekeland, 1979a). Theorem 6 solves half of the so-called Asplund conjecture (see Asplund, 1968) for the first steps in this direction, and the book by Diestel and Uhl for connections with the Radon-Nikodym property). The other half remains open (if every continuous convex function on a Banach space is Fréchet differentiable on a dense G^, must the space have an equivalent Fréchet differentiable norm?), as do all related questions when we replace Fréchet differentiability by Gâteaux differentiability. Section 7 Proposition 12 is the main result of Ekeland and Lebourg’s paper (1976). The approach we use here, however, is different, and due to Lebourg. Proposition 13 is due to Edelstein (1968). Similar results are known when one seeks to maxi¬ mize the distance to a given point in a closed subset (instead of minimizing it). They will be found in Edelstein (1966) and also follow from the results of this section. CHAPTER 6: SOLVING INCLUSIONS The main concepts of game theory, which we present in the first section, were defined by early economists. Cournot, for instance, introduced the duopoly
COMMENTS 491 model in his book (1838). Even if we don’t give everyone his due, we should at least quote the very influential books of Walras (1874) and Pareto (1909). Their work, however, never achieved great levels of formal sophistication. The first one to formulate a truly mathematical theory was, as in so many other instances, J. von Neumann (1944). His very early minimax theorem (1928), which he deduced from Brouwer’s fixed point theorem (1910), has set standards of rigor for the whole field and has received many extensions. Ky Fan’s inequality (theorem 3.5) was proved in 1972. The extension to mono¬ tone functions (theorem 3.9) is due to Brezis, Nirenberg, and Stampacchia (1972). The existence of noncooperative equilibria in game theory is due to Nash (1950a). The first fixed-point theorem for set-valued maps (with compact convex values) was proved by Kakutani (1941) for the needs of game theory. The particular needs of mathematical economics led to successive extensions of Kakutani’s theorem, culminating in the Gale-Nikaido-Debreu result (theorem 4.4). The inner trend of pure mathematics also led to fixed-point theorems for set-valued maps, such as theorem 4.13, which is essentially due to F. Browder (1968) and Ky Fan (1972), and theorem 4. 15, which is due to Rogalski (1972). Haddad and Lasry (1983) have also given fixed-point theorems for maps with nonconvex values. Leray and Schauder proved their famous theorem in their 1934 paper, and applied it to nonlinear partial differential equations. The extension to set-valued maps follows a method given in Granas (1976). Theorem 4.21 on quasi- variational inequalities originates in Arrow and Debreu (1954) and was extended by Joly and Mosco (1974). The KKM lemma was found by Knaster, Kuratowski, and Mazurkiewicz (1926) to give a combinatorial proof of Brouwer’s fixed-point theorem. It was generalized by Shapley (1973) (theorem 4.24), who used his result to provide a proof that the core of a balanced game is nonempty, a fact first stated by Scarf (1967a). The proof we give of the KKMS lemma is due to Ichiishi (1981a, b), and further extended by Ky Fan. The concept of Walras equilibrium was introduced by Walras (1874) and since then many partial proofs were offered, until the basic one due to Arrow and Debreu (1954). The alternative modelization of an equilibrium was pro¬ posed by Aubin (see Aubin and Cellina, 1983). Monotone maps were introduced by Zarantonello (1960). For a detailed account of monotone maps and variational inequalities, see Brezis (1968,1973), J. L. Lions (1969), and F. Browder (1976). Maximal monotone maps were introduced and characterized by Minty (1965). The proof of the theorem on the sum of two maximal monotone maps is due to Attouch (1981), extending a result due to Brezis, Crandall, and Pazy (1970) and Rockafellar (1974a). We refer to the paper by Brezis and Haraux (1976) for the study of the range of the sum of two maximal monotone maps, and applications to Hammerstein
492 COMMENTS equations (see also Brezis and Browder, 1975 and F. Browder, 1975). The exist¬ ence and uniqueness of a solution to a differential inclusion for maximal mono¬ tone maps is due to Crandall and Pazy (1969). The literature on fixed-point theory and monotone maps is quite large. Further references are listed in the bibliography. CHAPTER 7: NONSMOOTH ANALYSIS Nonsmooth analysis started at the end of the 1960s when the need to extend the successful subdifferential calculus to nonconvex and nonsmooth functions or to use convenient “tangent cones” for expressing the necessary conditions became pressing. Let us quote, among many other works, Dubovitskii and Miljutin (1971), Ioffe and Tihomirov (1972), Laurent (1972), and Neustadt (1976). The concept of generalized gradient and normal cone introduced by Clarke (1975) gave a new impetus in the field and was at the origin of a considerable amount of work. Other attempts for defining other concepts of generalized gradients were made by Russian mathematicians (see, e.g., Demianov and Vassiliev, 1981; Demianov and Rubinov, 1983, and Pchenitchny, 1980. The part of this chapter dealing with generalized gradient and normal cones is based on the works of Clarke (1975,1976d, 1977b, 1981a) regrouped in Clarke (1983), and the works of Rockafellar (1979a, b, c, 1980). The importance of the role of the Bouligand tangent cone (Bouligand, 1930, 1932) in viability theory for differential inclusion was recognized by Haddad (1981a), following many papers (Brezis, 1970; Clarke, 1975; Crandall, 1970; Gautier and Penot, 1979; Ladde and Lakshmikantham, 1974; Redheffer, 1972; Yorke, 1967,1969). See the book by Aubin and Cellina (1983) for further comments. The fact that the tangent cone is the Kuratowski liminf of the contingent cone was discovered by Cornet (1981a), Penot (1981), and Rockafellar and Wets (unpublished). Many works were denoted to tangent cones and derivatives (Auslender, 1978a, b; Crouzeix, 1977, for quasiconvex functions; Gauvin, 1979; Gollan, 1981; Halkin, 1976; Hiriart-Urruty, 1978, 1979a, b, c; Hiriart-Urruty and Thibault, 1980; Hogan, 1973; Ioffe, 1981a; Janin, 1982; Lebourg, 1975, 1979; Lemarechal, 1975, Lempio-Maurer, 1980; Penot, 1974,1978a, b, c; Shi Shu Chung, 1980; Thibault, 1979; Warga, 1976, 1978a). In Frankowska (to appear), one can find an intermediate cone lying between Ck{x) and Tk(x), namely, Pk{x) : = liminf j(K-x) h^o h which coincides with the tangent space of a differentiable manifold. She also introduces the asymptotic cone of a nonconvex cone T, defined by T”: = {v 67’|v+r<=7’}
COMMENTS 493 which is a convex subcone of T, The asymptotic cone Pk{x) of which is a closed convex cone larger than Ck{x), plays quite an important role. The associated generalized gradient d^V{x), the set of p such that (p, — 1) y(x))~, is smaller than the generalized gradient, and is reduced to the usual gradient when V is only Frechet differentiable (instead of being of class as for Clarke’s generalized gradient). Epi-contingent derivatives are quite useful for the theory of Hamilton- Jacobi equations (see Aubin, 1981 and Aubin and Cellina, 1983) and are related to the concept of generalized solutions introduced in Crandall and Lions (1981) and P.-L. Lions (1981a, 1982). Many concepts of generalized derivatives of vector-valued maps have been proposed and studied. Let us mention the fans, introduced by Ioffe (1979,1982). See also Aubin (1982b), Ioffe (1981c), and the papers of Clarke (1976e), Hiriart- Urruty (to appear), Kutateladze (1977), McLeod (1965), Sweester (1977), and Thibault(1982), The concept of contingent derivative of set-valued map was introduced in Aubin (1981) and the concept of derivative in Aubin (1982a). Other concepts of derivatives of set-valued maps were proposed by Banks and Jacobs (1970), de Blasi (1976), Boudourides and Shinas (1981), Gautier (unpublished), Mirica (1980), Nurminski (1978), Petcherskaja (1980), Shinas and Boudourides (1981), and Spingarn (1981). By using the asymptotic tangent cone Pk(x) she introduced in 1983, H. Frankowska has defined asymptotic derivatives and asymptotic co-differentials of set-valued maps and has used them for proving necessary conditions for optimal trajectories of differential inclusions. The inverse function theorem is taken from Aubin (1982a). See also the papers of F. Clarke (1976c), Halkin (1976), Ioffe (1981c), and Warga (1978). Corollary 6 subsumes many earlier results of “normal solvability theory”: see the papers by F. Browder (1976) and Kirk (1975). CHAPTER 8: HAMILTONIAN SYSTEM There is a vast literature on the subject of periodic solutions for Hamiltonian systems, beginning with the classical work of Poincaré, “Les méthodes nouvelles de la mécanique céleste” (1892-1899). We refer to the book by Moser (1973) for a survey and a bibliography, and to Jorna (1978) and Lichtenberg and Lieberman (1983) for a practicioner’s view on the subject. Most of this literature deals with systems depending on a small parameter, by the various methods of perturbation theory. The interest in global, topolo¬ gical methods was rekindled by the Rabinowitz paper (1978). See the surveys by Berestycki (1983) and Ambrosetti (1983) for a bibliography of recent develop¬ ments. The results in Section 2 can be traced to F. Clarke (1978, 1980a). Theorem 2.1 is found in Ekeland and Lasry (1980a, 1983). Theorem 2.4 is found in Clarke
494 COMMENTS and Ekeland (1978, 1980). There is also related work by Aubin and Ekeland (1980) and Brezis, Coron and Nirenberg (1980). Theorem 3.1 (the nonfesonant case) is well known: see, for instance, Mahwin (1976) and the book by Fucik (1983). The proof we give is of course new. Theorem 4.1 (the resonant case) is due to Clarke and Ekeland (1978, 1980). See also Clarke and Ekeland (1982), Ekeland (1981a, 1981b), and Willem (to appear) for a study of nonautonomous systems by this method. Theorem 5.1 (the trans-resonant case) is a particular case of the paper by Rabinowitz (1978), where convexity is replaced by a weaker assumption. The use of the Ambrosetti-Rabinowitz theorem in this context was initiated by Ekeland (1979b) and extended by Brezis, Coron, and Nirenberg (1980) to a nonlinear wave equation. See Ambrosetti and Mancini (1981) for a different proof, and progress on the question of minimality: does the solution found in theorem 5.1 have minimal period T? The nonautonomous trans-resonant case has been solved by Bahri and Berestycki (1981, 1983), who have found infinitely many periodic solutions. The question of finding periodic solutions with prescribed energy (instead of prescribed period) has also given rise to interesting developments; see Weinstein (1973), Moser (1976), Ekeland and Lasry (1980), and Ekeland (1984). Finally, note that direct variational methods, without the help of the duality theory of Section 2, have also proved quite useful; see Rabinowitz (1978), Bend and Rabinowitz (1979), and Bahri and Berestycki (1981,1983).
Bibliography Abraham, R., and Robbin, J. (1967). Transversal mappings and flows. Benjamin, New York. Amann, H., and Zehnder, E. T. (1980). Nontrivial solutions for a class of nonresonance problems and application to nonlinear differential equations. Ann. Sc. N. Sup. Pisa., 7, 539-603. Ambrosetti, A., and Mancini, G. (1981). Solutions of minimal period for a class of convex Hamilton¬ ian systems. Math. Ann., 255,405-421. Ambrosetti, A., and Mancini, G. (1981 ). On a theorem of Ekeland and Lasry concerning the number of periodic Hamiltonian trajectories. J. Diff. Eq., 43,1-6. Ambrosetti, A., and Prodi, G. (1973). On the inversion of some differential mappings with singu¬ larities between Banach spaces. Ann. Mat. Рига Appl., 93, 291-247. Ambrosetti, A., and Rabinowitz, P. (1973). Dual variational methods in critical point theory and applications. J. Funct. Anal., 14,349-381. Antosiewicz, H. A., and Cellina, A. (1975). Continuous selections and differential relations. J. Diff. Eq., 19, 386-398. Antosiewicz, H. A., and Cellina, A. (1977). Continuous extension of multifunctions. Ann. Polonisi Мал, 34,107-111. Arnold, V. (1974). Méthodes mathématiques de la mécanique classique (French translation) MIR, Moscow, 1976. Arnold, V. (1978). Chapitres Supplémentaires de la Théorie des Équations Différentielles Ordinaires. (French translation). MIR, Moscow, 1980. Arrow, K. J., and Debreu, G. (1954). Existence of an equillibrium for a competitive economy. Econometrica, 22,265-290. Arrow, K. J., and Hahn, F. M. (1971). General Competitive Analysis. Holden-Day, San Francisco. Artstein, Z. (1974). On the calculus of closed set-valued functions. Indiana Univ. Math. 7., 24,433- 441. Asplund, E. (1966). Farthest points in reflexive locally uniformly rotund Banach spaces. Isr. J. Mar/i.,4,213-216. Asplund, E. (1968). Fréchet differentiability of convex functions./1стМа/Л., 121,31-47. Asplund, E., and Rockafellar, R. T. (1969). Gradients of convex functions. Trans. Am. Math. Soc., 139,443-467. Attouch, H. (1979). Famille d’opérateurs maximaux monotones et mesurabilité. Ann. Mat. Рига Appl.,m, 35-111. Attouch, H. (1981). On the maximality of the sum of two maximal monotone operators. Nonlinear Anal. TAM, 5,143-147. Attouch, H. (To apppear). Variational Convergence for Functions and Operators, Research Notes in Mathematics. Pitman, London. 495
496 BIBLIOGRAPHY Attouch, H., and Damlamian, A. (1972). On multivalued evolution equations in Hilbert spaces. Isr. J. Math., 12, 373-390. Attouch, H., and Damlamian, A. (1975). Problèmes d’évolution dans les Hilbert et applications. J. Math. Pure Appl., 54,53-74. Attouch, H., and Wets, R. (1983). A convergence for bivariate functions aimed at the convergence of saddle values. In Mathematical Theory of Optimization, P. Cecconi and T. Zolezzi (Eds.), Springer-Verlag, Berlin. Attouch, H., and Wets, R. (1983). Convergence de points min/sup et de points fixes. CRAS, 296, 657-660. Attouch, H., and Wets, R. (1983). A convergence theory for saddle functions. Trans. Am. Math. Soc., in press. Aubin, J. P. (1963). Un théorème de compacité. CRAS, 265,5042-5045. Aubin, J. P. (1970). Abstract boundary-value operators and their adjoints. Rend. Sem. Padova, 43,1-33. Aubin, J. P. (1972). Théorème du minimax pour une classe de fonctions. CRAS, 274,455-458. Aubin, J. P. (1974). Règles de décision optimales en théorie des jeux à deux personnes. CRAS, 279, 173-176. Aubin, J. P. (1977). Evolution monotone d’allocations de biens disponibles. CRAS, 285, 293-296. Aubin, J. R., (1977). Applied Abstract Analysis. Wiley-Interscience, New York. Aubin, J. P. (1978). Propriété de Perron-Frobenius pour des correspondances. CRAS, 286,911-914. Aubin, J. P. (1978). Analyse fonctionnelle nonlinéaire et applications à l’équilibre économique. Ann. Sc. Math. Quebec, 2,5-47. Aubin, J. P. (1978). Gradients généralisés de Clarke. Ann. Sc. Math. Québec, 2,197-252. Aubin, J. P. (1979). Cônes tangents à un sous-ensemble convexe fermé. Ann. Sc. Math. Québec, 3, 63-80. Aubin, J. P. (1979). Mathematical Methods of Game and Economic Theory. North-Holland, Amsterdam. Aubin, J. P. (1979). Applied functional analysis. Wiley-Interscience, New York. Aubin, J. P. (1980). Further properties of Lagrange multipliers in nonsmooth optimization. Appl. Math.Opt.,6,19-9^. Aubin, J. P. (1981). Contingent derivatives of set-valued maps and existence of solutions to nonlinear inclusions and differential inclusions. In Advances in Mathematics Supplementary Studies, L. Nachbin (Ed.), Academic Press, New York, pp. 160-232. Aubin, J. P. (1982). Ioffe’s fans and generalized derivatives of vector-valued maps. In Convex Analysis and Optimization, J. P. Aubin and R. Vinter (Ed.), Pitman, London. Aubin, J. P. (1982). Comportement lipschitzien des solutions de problèmes de minimisation convexes. CRAS, 295,235-238. Aubin, J. P., and Cellina, A. (1984). Differential Inclusions. Springer-Verlag, Berlin. Aubin, J. P., and Clarke, F. M. (1977). Monotone invariant solutions to differential inclusions. J. London Math. Soc., 16,357-366. Aubin J. P., and Clarke, F. M. (1977). Multiplicateurs de Lagrange en optimisation nonconvexe et applications. J. London Math. Soc., 16,357-366. Aubin J. P., and Clarke F. M. (1979). Shadow prices and duality for a class of optimal control problems. SIAMJ. Opt. Control, 17,567-586. Aubin, J. P., and Cornet, B. (1976). Règles de décision en théorie des jeux et théorèmes de point fixe. 283,11-14. Aubin, J. P., Ekeland, I. (1976). Estimation of the duality gap in nonconvex optimization. Math. Op.Res.,l,\-\4. Aubin, J. P., Ekeland, I. (1980). Second order evolution equation with convex Hamiltonian. Can. Math.Bull.,23,%\-94.
BIBLIOGRAPHY 497 Aubin J. P., Siegel, J. (1980). Fixed points and stationary points of dissipative multivalued maps. Proc. Am, Math. Soc., 78,391-398. Aubin, J. P., Vinter, R. (1982). Convex Analysis and Optimization. Pitman, Boston. Aumann, R. J. (1965). Integrals of set-valued functions. J. Math. Anal. Appl., 12,1-12. Auslender, A. (1976). Optimisation : Méthodes Numériques. Masson, Paris. Auslender, A. (1977). Minimisation sans contraintes de fonctions localement lipschitziennes. CRAS, 284,959-961. Auslender, A. (1978). Stabilité différentiable en programmation nonconvexe nondifférentiable. C. R. Acad. Sci. Paris, 286,575-577. Auslender, A. (1978). Differentiable stability in nonconvex and nondifferentiable programming. In Mathematical Programming Study, No. 10, P. Huard (Ed.). Averbuh, V. I., and Smolyanov, O. G. (1968). Different definitions of derivatives in linear topo¬ logical spaces. Usp. Mat. Nauk, 23,67-116. Bahri, A., and Berestycki, H. (1981). A perturbation method in critical point theory. Trans. Am. Math.Soc.,261,\-Z2. Bahri, A., and Berestycki, H. (to appear). Existence of forced oscillations for some nonlinear differential equations. Comm. Pure. Appl. Math. Bahri, A., Berestycki, H. (to appear). Forced vibrations of superquadratic Hamiltonian systems. Acta Mathem. Bâillon, J. B. (1978). Générateurs de semi-groupes dans les espaces de Banach uniformément lisses. J. Funct. Anal., 29,199-212. Bâillon, J. B., and Brézis, H. (1976). Une remarque sur le comportement asymptotique des semi- groupes nonlinéaires. Houston J. Math., 2,5-7. Bâillon, J. B., and Haddad, G. (1977). Quelques propriétés des opérateurs angle-bornés et n-cycli- quement monotones. Isr. J. Math., 26,137-150. Baiocchi, C., and Capelo, A. (1978). Disequazioni variazionali e quasivariazionali. Applicazioni a problemi di frontiera libera. Pitagora Editrice, Bologna. Banks, H. T., and Jacobs, M. Q. (1970). A differential calculus for multi-functions. J. Math. An. Appl.,29,246-212. Barbu, V. (1976). Nonlinear Semi-Groups and Differential Equations in Banach Spaces. Noordhoff, Leiden, Netherlands. Barbu, V. (1981). Necessary conditions for nonconvex distributed control problems. J. Math. Anal. /1/7/?/., 80,566-597. Barbu, V., and Cellina, A. (1970). On the surjectivity of multivalued dissipative mappings. Boll. Un. MflL/to/., 3,817-826. Beer, G. (to appear). Oh functions that approximate relations. Proc. Am. Math. Soc. Begle, E. (1950). A fixed-point theorem. Ann. Math., 51,544-550. Belluce, L. P., and Kirk, W. A. (1969). Some fixed point theorems in metric and Banach spaces. Can. Math. Bull., 127,481^89. Belluce, L. P., and Kirk, W. A. (1969). Fixed point theorems for families of contraction mappings. Proc. Am. Math. Soc., 29,141-146. Benci, V. ( 1981 ). A geometrical index for the group S^ and some applications to the study of periodic solutions to ordinary differential equations. Comm. Pure Appl. Math., 34,393-432. Benci, V., and Rabinowitz, P. (1979). Critical point theorems for indefinite functionals. Inv. Math., 52,241-273. Bensoussan, A. (1982). Stochastic Control by Functional Analysis Methods. North-Holland, Amsterdam. Bensoussan, A., and Lions, J. L. (1978). Applications des inéquations variationnelles et controle stochastique. Dunod, Paris.
498 BIBLIOGRAPHY Bensoussan, A., and Lions, J. L. (1982). Contrôle Impulsionnel et Inéquations Quasi-Variationnelles. Dunod, Paris. Berestycki, H. (1983). Solutions périodiques de systèmes Hamiltoniens, Séminaire Bourbaki No. 603. Berestycki, H., Lasry, J. M., Mancini, G., and Ruf, B. (to appear). Existence of multiple periodic solutions on star-shaped Hamiltonian surfaces. Comm. Pure Appl. Math. Berge, С. (1959). Espaces Topologiques et Fonctions Multivoques. Dunod, Paris. Bishop, E., and Phelps, R. (1963). The support functional of a convex set. Proc. Symp. Pure Math., 7,27-35. de Blasi, F. S. (1976). On differentiability of multifunctions. Рас. J. Math., 66,67-81. Bliss, G. A. (1946). Lectures on the Calculus of Variations. University of Chicago Press, Chicago. Bondareva, O. N. (1962). Theory of the core in w-person games, (in Russian). Vestn. Leningrad, 13, 141-142. Bony, J. M. (1969). Principe du maximum, inégalité de Harnack, et unicité du problèm de Cauchy pour des operateurs elliptiques dégériéiés Ann. Inst. Fourier, 19,277-304. Borwein, J. (1978). Weak tangent cones and optimization in Banach spaces. SIAM J. Cont. Opt., 16, 512-522. Boudourides, M., and Shinas, J. (1981). The mean value theorem for multifunctions. Bull. Math Ломт., 25,128-141. Bouligand, G. (1930). Sur les surface dépourvues de points hyperlimites. Ann. Soc. Polon. Math., 9, 32-41. Bouligand, G. (1932). Introduction à la géométrie Infinitésimale Directe. Gauthier-Villars, Paris. Bourbaki, N. (1981). Espaces Vectoriels Topologiques. Masson, Paris, Chaps. 1-5. Bourbaki, N. (1982). Variétés Différentielles et Analytiques. Masson, Paris, Fasc. 1-5. Bourbaki, N. (1982). Topologie Générale. Masson, Paris, Chap. 1-4. Brézis, H. (1968). Equations et inéquations nonlinéaires dans les espaces vectoriels en dualité. Ann. Inst. Fourier, 18,115-175. Brézis, H. (1970). On a characterization of flow invariant sets. Comm. Pure Appl. Math., 23,261-263. Brézis, H. (1973). Opérateurs Maximaux Monotones et Semi-Groupes de Contractions dans les Espaces de Hilbert. North-Holland, Amsterdam. Brézis, H. (1983). Analyse Fonctionnelle et Applications. Masson, Paris. Brézis, H. (1983). Periodic solutions of nonlinear vibrating strings and duality principles. Bull. Am. Math.Soc.%,m^26. Brézis, H., Coron, J. M., and Nirenberg, L. (1980). Free vibrations for a nonlinear wave equations and a theorem of P. Rabinovitz. Comm. PureApp. Math., 33,667-689. Brézis, H., and Nirenberg, L. (1978). Characterization of the range of some nonlinear operators and applications to boundary valve problems Ann. Scuola Norm. Sup. Pisa, 5,225-326. Brézis, H., and Browder, F. (1974). Some new results about Hammerstein equations. Bull. Am. Math.Soc.,9IS, 567-572. Brézis, H., and Browder, F. (1975). Nonlinear integral equations and systems of Hammerstein type. Adv.Math.,\%,\\5-\Al. Brézis, H., and Browder, F. (1976). Nonlinear egodic theorems. Bull. Am. Math. Soc., 82,959-961. Brézis, H., and Browder, F. (1976). A general principle on ordered sets in nonlinear functional analysis. Adv. Math., 21,355-369. Brézis, H., and Browder, F. (1977). Remarks on nonlinear ergodic theory. Adv. Math., 25,165-177. Brézis, H., Crandall, M., and Pazy, A. (1970). Perturbations of nonlinear maximal monotone sets in Banach spaces. Comm. Pure Appl. Math., 23,123-144. Brézis, H., Crandall, M., and Pazy, A. (1976). Un principe variationnel associé à certaines équations paraboliques. CRAS, 282,971-974.
BIBLIOGRAPHY 499 Brézis H., and Haraux, A. (1976). Image d’une somme d’opérateurs monotones et applications. Isr.J.Math.,l\\6S-m. Brézis, H., and Lions, P. L. (1978). Produits infinis de résolvantes. Isr. J. Math,^ 29,329-345. Brézis, H., Nirenberg, L., and Stampacchia, G. (1972). A remark on Ky Fan’s minimax principle. Boll. Un. Mat. liai., 6,293-300. Bröcker, T. (1972). Differenzierbare Abbildungen Der Regensburger Trichter. (English translation: Cambridge University Press, Cambridge, U.K.). Brondsted, A., and Rockafellar, R. T. (1965). On the subdiflferentiability of convex functions. Proc. Am. Math. Soc., 16,605-611. Brouwer, L. (1910). Uber eindeutige stetige Transformationen von Flächen in sich. Math. Ann., 67, 176-180. Brouwer, L. (1911). Beweis der Invarianz der Dimensionzahl. Math. Ann., 70. Brouwer, L. E. J. (1912). Uber Abbildungen von Manningfaltigkeiten. Math. Ann., 71, 97-115. Browder, F. (1965). Fixed point theorems for non-compact mappings in Hilbert spaces. Proc. Natl. Acas. Sei. USA, 53,1272-1276. Browder, F. (1965). Non expansive nonlinear operators in a Banach space. Proc. Natl. Acad. Sei. (75/4,54,1041-1044. Browder, F. (1968). The fixed point theory of multivalued mappings in topological vector spaces. Math.Ann.,\ll,m-30\. Browder, F. (1976). Nonlinear operators and nonlinear equations of evolution in Banach spaces. Am. Math. Soc. Proc. Symp. Pure Math., 18,2. Browder, F. (1976). Normal solvability for nonlinear mappings into Banach spaces. Bull. Am. Math. 5oc., 79,328-350. Browder, F. (1983). Fixed point theory and nonlinear problems. Bull. Am. Math. Soc., 9, (1), 1-39. Bruckner, A. J., Leonard, J. L. (1966). ‘Derivatives’ in the Slaught memorial papers. Am. Math. Mo«., 73,24-56. Cambini, A. (1973). Sul lemma di Morse. Boll. Un. Mat. Ital., 1,87-93. Carathéodory, C. (1967). Calculus of Variations and Partial Differential Equations of the First Order. Holden-Day, San Francisco. Caristi, J. (1976). Fixed point theorems for mappings satisfying inwardness conditions. Trans. Am. Math.Soc.,l\5,2A\-25\. Castaing, C, Valadier, M. (1977). Convex analysis and measurable multifunctions. Lecture Notes in Mathematics., Volume 580. Springer-Verlag, Berlin. Castro, A., and Lazer, A. (1979). Critical point theory and the number of solutions of a nonlinear Dirichlet prob. Ann. Mat. Pura. Appl., 120,113-137. Céa, J. (1971). Optimisation: Théorie et Algorithmes. Dunod, Paris. Céa, J. Glowinski, R., and Nedelec, J. C. (1971). Minimisation de fonctionelles non différentiables. Lecture Notes in Mathematics, 22%, Morris (Ed.), Springer-Verlag, Berlin. Cellina, A. (1969). Approximation of set valued functions and fixed point theorems. Ann. Mat. Pura Appl.,S2,17-24. Cellina, A. (1969). A theorem on the approximation of compact multivalued mappings. Rend. Accad. Naz. Lincei, 47,434-440. Cellina, A. (1969). Multivalued Functions and Multivalued Flows. University of Maryland Tech. Note BN 615. Cellina, A. (1970). A further result on the approximation of set valued mappings. Rend. Accad. Naz. Lincei, 48,230-234. Cellina, A. (1970). Multivalued differential equations and ordinary differential equations. SIAM J.Appl.Math.,\%,533-53%.
500 BIBLIOGRAPHY Cellina, A. (1971). The role of approximation in the theory of multivalued mappings. In Differential Games and Related Topics, H. W. Kuhn and G. P. Szego (Eds.), North-Holland, Amsterdam. Cellina, A. (1971). On mappings defined by differential equations Zesz. Nauk. Uniw. Jagiellon. Pr. Ma/., 15,17-19. Cellina, A. (1976). A selection theorem. Rend. Sem. Univ. Padova, 55,143-149. Cellina, A. (1980). On the differential inclusion x’e(—1, +1). Rend. Acc. Naz. Lincei, 69, 1-6. Cellina, A., and Lasota, A. (1969). A new approach to the definition of topological degree for multivalued mappings. Rend. Accad. Naz. Lincei, 47,434-440. Cellina, A., and Marchi, M. V. (1982). Nonconvex perturbations of maximal monotone differential inclusions. Cerf, J. (1966). Г"^ = 0. Lecture Notes in Mathematics, Volume 25. Springer-Verlag, Berlin. Chaney, R. W. (1982). Second order sufficiency conditions for non differentiable programming problems. SIAM J. Contré Opt., 20,20-33. Chow, S. Mallet-Paret, J. (1982). Methods of bifurcation theory. Springer. Clark, C. W., Clarke, F., and Munro, G. R. (1979). The optimal exploitation of renewable resource stocks. Economics, 47,25-47. Clarke, F. H. (1975). The Euler-Lagrangedifferential inclusion. J. Diff. Eq., 19,80-90. Clarke, F. H. (1975). Generalized gradients and applications. Trans. Am. Math. Soc., 205,247-262. Clarke, F. H. (1976). On the inverse function theorem. Рас. J. Math., 64,97-102. Clarke, F. H. (1976). The maximum principle under minimal hypotheses. SIAM J. Control Opt., 14, 1078-1091. Clarke, F. H. (1976). The generalized problem of Bolze. SIAM J. Control Opt., 14,682-699. Clarke, F. H. (1976). A new approach to Lagrange multipliers. Math. Op. Res., 1,165-174. Clarke, F. H. (1976). Optimal solutions to differential inclusions. J. Opt. Th. Appl., 19, 469-478. Clarke, F. H. (1977). External arcs and extended Hamiltonian systems. Trans. Am. Math. Soc., 231, 349-367. Clarke, F. H. (1977). Multiple integrals of Lipschitz functions in the calculus of variations. Proc. Am. Math. Soc., 64,260-264. Clarke, F. H.(1978). Pointwise contraction criteria for the existence of fixed points. Bull. Can. Math. Soc.,l\,l-\\. Clarke, F. H. (1978). Nonsmooth analysis and optimization. Proceedings of the International Con¬ gress of Mathematicians, Helsinki, 1978. Clarke, F. H. (1978). Solutions périodiques des équations Hamiltoniennes. CRAS, 287, 951-952. Clarke, F. H.(1979). Optimal control and the true Hamiltonian. SIAM Rev.,l\, 157-166. Clarke, F. H. (1980). The Erdmann condition and Hamiltonian inclusions in optimal control and calculus of variations Can. J. Math., 32,494-509. Clarke, F. H. (1981). Generalized gradients of Lipschitz functionals. Adv. Math.,40,52-67. Clarke, F. H. (1981). Periodic solutions of Hamiltonian inclusions. J. Diff. Eq., 40,1-6. Clarke, F. H. (1982). On Hamiltonian fiows and symplectic transformations. SIAM J. Control Opt.,20,355-359. Clarke, F. H. (1983). Optimization and nonsmooth analysis. Wiley-Interscience, New York. Clarke F. H., and Ekeland, I. (1978). Solutions périodiques, de période donnée, des équations Hamiltoniennes. CRAS, 287, 1031-1015. Clarke, F. H., and Ekeland, I. (1980). Hamiltonian trajectories having prescribed minimal period. Comm. Pure Appl. Math., 33,103-116. Clarke, F. H., and Ekeland, I. (1982). Nonlinear oscillations and boundary-value problems for Hamiltonian systems. Arc. Ration. Mech. Anal., 78,315-333.
BIBLIOGRAPHY 501 Conley, C., and Zehnder, E. (1983). The Birkhoff-Lewis fixed point theorem and a conjecture of V. I. Arnold. /. Math,, 73,33-49. Conley, C., and Zehnder, E. (to appear). Morse type index theory for flows and periodic solutions for Hamiltonian equations. Comm. PureAppl. Math. Cornet, B. (1977). Accessibilité des optima de Pareto par des processus monotones. CRAS, Ser. A, 641-644. Cornet, B. (1977). An abstract theorem for planning procedures. Lee. Notes Econ. Math. Syst., 144, 53-59. Cornet, B. (1981). Regularity properties of tangent and normal cones. Cah. Math. Décision, Univ. Paris Dauphine, 81-30. Cornet, B. (1981). Contributions à la théorie mathématique des mécanismes dynamiques d’allocation des ressources. Thèse de doctorat d’Etat, Université de Paris-Dauphine. Cornet, B., and Lasry, J. M. (1976). Un théorème de surjectivité pour une procédure de planification. C/?/15,282,1375-1378. Costa, D., and Willem, M. (1983). Multiple critical points of invariant functionals and applications. MRC Technical summary report 2532. Cournot, A. A. (1838). Recherches sur les Principes Mathéorie des Richesses. H. Rivière and Co., Paris. Crandall, M. G. (1970). Differential equations on convex sets. J. Math. Soc. J., 22,443-455. Crandall, M. G. (1972). A generalization of Peano’s existence theorem and flow invariance. Proc. Am.Math.Soc.,36,\5\-\55. Crandall, M. G. (1977). An introduction to constructive aspects of bifurcation and the implicit funct. th. Applications of bifurcation theory, P. Rabinowitz (Ed.), pp. 1-35. Academic Press, New York. Crandall, M. G. (1979). Nonlinear Evolution Equations. Academic Press, New York. Crandall, M. G., and Lions, P. L. (1981). Conditions d’unicité pour les solutions généralisées des equationsd’Hamilton-Jacobi. CRAS, 292,183-186. Crandall, M. G., and Pazy, A. (1969). Semi-groups of nonlinear contractions and dissipative sets. J. Funct. Anal.376^18. Crandall, M. G., and Pazy, A. (1970). On accretive sets in Banach spaces. J. Funct. Anal., 5,204-217. Crandall, M. G., and Rabinowitz, P. (1971). Bifurcation from simple eigenvalues. J. Funct. Anal., 8, 321-340. Crandall, M. G., and Rabinowitz, P. (1973). Bifurcation, perturbation of simple eigenvalues and linearized stability. Arch. Rat. Mech. An., 52,161-180. Cronin, J. (1964). Fixed points and topological degree in nonlinear analysis. Am. Math. Soc. Math. Surveys,\\. Crouzeix, J. P. (1977). Contribution à l’étude des fonctions quasi-convexes. Ph D Thesis, Université de Clermont. Crouzeiz, J. P. (1980). On second order conditions for quasi-convexity. Math. Prog., 18, 349-352. Danes, J. (1979). Two fixed point theorems in topological and metric spaces. Bull. Aust. Math. Soc., 14,259-265. Dantzig, G. B. (1963). Linear Programming and Extensions. Princeton University Press. Debreu, G. (1959). Theory of Value. Wiley, New York. Demianov, V. F. (1982). Non Smooth Analysis (in Russian). Leningrad University Press. Demianov, V. F., and Malosemov, V. N. (1972). Introduction to Minimax (in Russian). Nauka, Moscow. Demianov, V. F., and Rubinov, A. M, (1980). On quasidifferentiable functionals (in Russjaiflr^ Dokl. Acad. Nauk SSSR Sov. Math. Dokl., 21,14-17.
502 BIBLIOGRAPHY Demianov, V. F., and Vassiliev, L. V. (1981). Non differentiable optimization (in Russian). Nauka, Moscow. Desolneux-Moulis, N. (1979). Orbites périodiques des systèmes Hamiltoniens autonomes. Séminaires Bourbaki, 32, (552). Diestel, J. (1975). Geometry of Banach spaces. In Lecture Notes in Mathematics, Vol. 485, Springer- Verlag, Berlin. Diestel, J. (1977). Vector measures. Am. Math. Soc., Math Surveys, 15. Dolecki, Z., Salinetti, G., and Wets, R. (1983). Convergence of functions: equi-semicontinuity. Trans. Am. Math. Soc., 276,409-429. Downing, D., and Kirk, W. A. (1977). A generalization of Caristi’s theorem and application to nonlinear mapping theory. Рас. J. Math., 69,339-347. Downing, D., and Kirk, W. A. (1977). Fixed point theorems for set-valued mappings in metric and Banach spaces. Math. Jpn., 22,99-112. Downing, D., and Ray, W. (to appear). Renorming and the theory of phi-accretive set-valued mapp¬ ings. Рас. J. Math. Dubovicki, A. I., and Milyutin, A. M. (1971). Necessary Conditions for a Weak Extremum in the General Problem of Optimal Control {m Russian). Nauka, Moscow. Dugundji, J. (1951). An extension of Tietze’s theorem. Рас. J. Math., 1,353-367. Dugundji, J. (1954). Topology. Allyn and Bacon, Boston. Dunford, N., and Schwartz, J. T. (1958). Linear Operators. Wiley-Interscience, New York. Du vaut. G., and Lions, J. L. (1972). Les Inéquations en Mécanique et en Physique. Dunod, Paris. Eaves, B. C. (1972). Homotopies for computation of fixed points. Math. Program., 3,1-22. Edelstein, H. (1963). A theorem on fixed points under isometries. Am. Math. Mon., 70, 298-300. Edelstein, H. (1965). On nonexpansive mappings of Banach spaces. Proc. Cambridge Phil. Soc., 60, 439-447. Edelstein, H. (1966). Farthest points of set uniformity convex Banach spaces. Isr. J. Math., 4,171- 176. Edelstein, H. (1968). On nearest points of sets in uniformly convex Banach spaces. J. London Math. 5oc.,43,375-377. Edelstein, H. (1972). The construction of an asymptotic center with a fixed point property. Bull. Am. Math. Soc., 78,206-208. Edelstein, H. (1974). Fixed point theorems in uniformly convex spaces. Proc. Am. Math. Soc., 44, 369-374. Edelstein, H. (1975). On some aspects of fixed point theory in Banach spaces. In The Geometry of Metric and Linear Spaces, Lecture Notes, Vol. 490, Springer-Verlag, Berlin. Eggleston, H. C. (1958). Convexity. Cambridge Tracts in Mathematics, Volume 47. Cambridge University Press, London. Eilenberg, S., and Montgomery, D. (1966). Fixed point theorem for multivalued transformations. Am.J.Math.,S%,2\4-222. Ekeland, I. (1972). Remarques sur les problèmes variationnels 1. CRAS, 275,1057-1059. Ekeland, I. (1973). Remarques sur les problèmes variationnels 2. CRAS, 276,1347-1348. Ekeland, I. (1974). On the variational principle. J. Math. Anal. Appl., 47,324-353. Ekeland, I. (1979). Periodic solutions of Hamilton’s equations and a theorem of P. Rabinowitz. J.Diff.Eq.,34,523-534. Ekeland, I. (1979). Nonconvex minimization problems. Bull. Am. Math. Soc., 1,443-474. Ekeland, I. (1981). Oscillations forcées de systèmes Hamiltoniens nonlinéaires. Bull. SMF, 109,297- 330. Ekeland, I. (1981). Forced oscillations for nonlinear Hamiltonian systems. In Advances in Mat- hematics., L. Nachbin (Ed.), Academic Press, New York.
BIBLIOGRAPHY 503 Ekeland, I. (1982). Dualité et stabilité des systèmes Hamiltoniens. CRAS, 294,673-676. Ekeland, I. (1983). A perturbation theory near convex Hamiltonian systems. J. Diff, Eq.y 50,407-440. Ekeland, I. (1984). Une théorie de Morse pour les systèmes Hamiltoniens convexes. Am. IHP, Nonlinear Anal. 1,19-78. Ekeland, I., and Lasry, J. M. (1980). Problèmes variationnels non convexes en dualité. CRASy 291,493^95. Ekeland, I., and Lasry, J. M. (1980). On the number of closed trajectories for a Hamiltonian flow on a convex energy surface. Ann. Math.y 112,283-319. Ekeland, I., and Lebourg, G. (1976). Generic Frechet-differentiability and perturbed optimization problems in Banach spaces. Trans. Am. Math. Soc. 224,193-216. Ekeland, L, and Temam, R. (1976). Convex Analysis and Variational Problems. Elsevier North- Holland, Amsterdam. Fadell, E. (1978). Recent results in the fixed point theory of continuous maps. Bull. Am. Math. Soc.y 76,10-29. Fadell, E., and Rabinowitz, P. (1978). Generalized cohomological index theories for group actions with an application to bifurcation questions for Hamiltonian systems Inv. Math.y 45,139-174. Fan, Ky (1952). Fixed point and minimax theorems in locally convex topological linear spaces. Proc. Natl. Acad. Sci. US Ay 38, 121-126. Fan, Ky. (1953). Minimax theorems. Proc. Natl. Acad. Sci. US Ay 39,42-47. Fan, Ky. ( 1956). On systems of linear inequalities. Ann. Math. Stud. 38,99-156. Fan, Ky. (1957). Existence theorems and extreme solutions for inequalities concerning convex functions or linear translations. Math. Z., 68,205-216. Fan, Ky. (1958). On the equilibrium value of a system of convex and concave functions. Math. Z., 70,271-280. Fan, Ky. (1961). A generalization of Tychonoffs fixed point theorem. Math. Ann.y 142, 305-310. Fan, Ky. (1963). On the Krein-Milman theorem. Proc. Symp. Pure Math. Am. Math. Soc. 7, 211- 219. Fan, Ky. (1964). Sur un théorème minimax. C.R. Acad. Sci.y 259,3925-3928. Fan, Ky. (1965). A generalization of the Alaoglu-Bourbaki theorem and its applications. Math. Z., 88,48-60. Fan, Ky. (1966). Applications of a theorem concerning sets with convex sections. Math. Ann.y 163,189-203. Fan, Ky. (1970). A conbinatorial property of pseudomanifolds and covering properties of simplexes. J. Math. Anal. Appl.y 31,68-80. Fan, Ky. (1972). A minimax inequalities and applications. In InequalitieSy Volume 3, Shisha (Ed.). Academic Press, New York, pp. 103-113. Fan, Ky. (1979). Fixed point theorems and related theorems for non-compact convex sets. In Game Theory and Related Topics y North-Holland, Amsterdam, pp. 151-156. Fenchel, W. (1949). On conjugate convex functions. Can. J. Math.y 1, 73-77. Fenchel, W. (1951). Convex Cones, Sets and Functions, Mimeographed Lecture Notes. Princeton University. Fenchel, W. (1952). A remark on convex sets and polarity. Medd. Lunds Univ. Mat. Sem. (Supple¬ ment Band), 82-89. Fenchel, W. (1956). Uber konvexe Functionen mit vorgeschriebenen Niveaumannigfaltigkeiten. Mfli/i.Z., 63,496-506. Figueiredo, D. (1967). Topics in nonlinear analysis. Lecture Notes 48, University of Maryland. Frankowska, H. (1983). Inclusions adjointes associées aux trajectoires minimales d’inclusions diff. CRASy (in press).
504 BIBLIOGRAPHY Frankowska, H. (to appear). First order necessary conditions for nonsmooth variational control problems. SIAM. J. Opt. Frankowska, H. (to appear). On the single-valuedness of Hamilton-Jacobi operators. J. Nonlinear Anal. TAM. Frobenius, G. (1908). UberMatrizenauspositivenelementen. Sitzungsber Pr. Akad. 471-476. Fucik, S. (1983). Nonlinear Problems. Reidel, Dordrecht, Netherlands. Furi, M., Martelli, M., and Vignoli, A. (1970). On minimum problems for families of functionals. Ann. Mat. Рига Appl.,%6,181-187. Galé, D. (1956). The closed linear model of production. Ann. Math. Stud., 38,285-303. Garcia, C. B., and Gould, F. J. (1978). A theorem on homotopy paths. Math. Op. Res., 3,282-289. Garcia, C. B. and Gould, F. J. (1980). Relations between several path following algorithms and local and global Newton methods. SIAM Rev., 22, 263-274. Gautier, S. (1973). Différentiabillitédes multiapplications. Publ. Math. Pau. Gautier, S., and Penot, J. P. (1973). Fermés invariants par un système dynamique. CRAS, 276, 1457-1460. Gauvin, J. (1979). The generalized gradient of a marginal function in mathematical programming. Math. Opt. Res., 4,458-463. Gollan, B. (1981). Higher order necessary conditions for an abstract optimization problem. Math. Program. Study, 14,69-76. Gollan, B. (to appear). A general perturbation theory for abstract optimization problems. J. Opt. Th.Appl. Granas, A. (1962). Sur la multiplication cohomotopique dans les espaces de Banach, CRAS, 254, 56-57. Granas, A. (1976). Sur la méthode de continuité de Poincaré. CRAS, 282,983-986. Haddad, G. (1981). Monotone viable trajectories for functional differential inclusions. /. Diff. Eq., 42,1-24. Haddad, G. (1981). Monotone trajectories of differential inclusions with memory. Isr. J. Math., 39, 83-100. Haddad, G. (1981). Topological properties of the set of solutions for functional differential inclus¬ ions. Nonlinear Anal. Theory, Meth. AppL, 5,1349-1366. Haddad, G., and Lasry, J. M. (to appear). Periodic solutions of functional differential inclusions and fixed points of (j-selectionable correspondences. J. Math. Anal. Appl. Hale, J. (1977). Theory of Functional Differential Equations. Springer-Verlag, Berlin. Halkin, H. (1972). Extremal properties of biconvex contingent equations. In Ordinary Different¬ iable Equations (NRL-MRC Conference), Academic Press, New York. Halkin, H. (1976). Interior mapping theorem with set-valued derivative. J. Anal. Math., 30,200-207. Halkin, H. (1976). Mathematical programming without differentiability. In Calculus of Variations and Control Theory, D. L. Russell (Ed.), Academic Press, New York, 279-288. Halpern, B. R. (1968). A general fixed point theorem. Proc. Symp. Nonlinear Funct. Anal. Amer. Math. Soc. Halpern, B. R., and Berginan, G. M. (1968). A fixed point theorem for inward and outward maps. Trans. Am. Math. Soc., 130,353-358. Hamilton, R. (1982). The inverse function theorem of Nash and Moser. Bull. Am. Math. Soc., 7, 1-64. Hassard, B., Kazarinoff, N., and Wan, Y. (1981). Theory and applications of Hopf bifurcation. London Mathematical Society Lecture Notes 41, Cambridge University Press, London. Hildenbrand, W. (1974). Core and Equilibria of a Large Economy. Princeton University Press, Princeton.
BIBLIOGRAPHY 505 Himmelberg, C. J. (1972). Fixed points of compact multifunctions. Indiana Univ. Math. 7., 22, 719-729. Himmelberg, C. J. (1975). Measurable relations. Fund. Math., 87,53-12. Himmelberg, C. J., Jacobs M. Q., and Van Vleck, F. S. (1969). Measurable multifunctions, selectors and Filippov’s implicit function lemma. J. Math. Anal. Appi, 25,276-284. Himmelberg, C. J., and Van Vleck, F. S. (1971). Selection and implicit function theorems for multi¬ functions with Souslin graph. Bull. Acad. Pol. Sc., 19, 911-916. Himmelberg, C. J., Van Vleck, F. S. (1972). Lipschitzian generalized differential equations. Rend. Sent. Mat. Padova, 4S, 159-169. Himmelberg, C. J., and Van Vleck, F. S. (1972). Fixed points of semi-condensing multifunctions. Boll. Un. Mat. Ital.,5,187-194. Hiriart-Urruty, J. B. (1978). Gradients généralisés de fonctions marginales. SIAM J. Cont. Opt., 16, 301-316. Hiriart-Urruty, J. B. (1979). Refinements of necessary optimality conditions in nondifferntiable programming 1. Appl. Math. Opt., 63-82. Hiriart-Urruty, J. B. (1979). New concepts in nondifferentiable programming. Bull. Soc. Math. F. Mém., 60,57-85. Hiriart-Urruty, J. B. (1979). Tangent cones, generalized gradients and mathematical programming in Banach spaces. Math. Oper. Res., 4, 79-97. Hiriart-Urruty, J. B. (1980). Mean value theorems in nonsmooth analysis. Numer. Funct. Anal. O/?/., 2(1), 1-30. Hiriart-Urruty, J. B. (1981). Optimality conditions for discrete nonlinear norm-approximation problems. In Optimization and Optimal Control, Springer-Verlag, Berlin, pp. 29-41. Hiriart-Urruty, J. B. (to appear). Refinements of necessary optimality conditions in nondifferent¬ iable programming 2. Math. Programming Study. Hiriart-Urruty, J. B., and Thibault, L. (1980). Existence et caractérisation de différentielles générali¬ sées. 290,1091-1094. Hirsch, M. W. (1976). Differentiable Topology. Springer-Verlag, Berlin. Hirsch, M. W., and Smale, S. (1974). Differential Equations, Dynamical Systems and Linear Algebra. Academie Press, New York. Hirsch, M. W., and Smale, S. (1978). On algorithms for solving/(x)=0. Not. Am. Math. Soc., 25, 501-544. Hogan, W. (1973). Directional derivative for extremal value functions with applications to the completely convex case. Oper. Res., 21, 188-209. Holmes, R. D. (1976). Fixed point for local radial contractions. Proc Sympos Fixed Point Theory and Its Appl, Dalhousie University Academic Press, New York, pp. 79-89. Hopf, H. (1928). Eine Verallgemeinerung der Euler-Poincaréschen Formel. Nachr. Ges. Wiss. Goettingen, 127-136. Hopf, H. (1929). Uberdie algebraische Anzahl Fixpunkten. Math. Z., 29,493-524. Huard, P. (1975). Optimization algorithms and point to set maps. Math Program., 8, 308-331. Ichiishi, T. (1981). A social coalitional equilibrium existence lemma. Econometrica, 49. Ichiishi, T. (1981). On the Knaster-Kuratowski-Mazurkiewicz-Shapley theorem. J. Math. Anal. Appl. Ioffe, A. D. (1976). An existence theorem for a general Bolza problem. SIAM J. Cont. Opt., 14, 458-466. Ioffe, A. D. (1977). On lower semicontinuity of integral functionals 1. SIAM J. Cont. Opt., 15, 521-538.
506 BIBLIOGRAPHY loflfe, A. D. (1978). Survey of measurable selection theorems: Russian literature supplement. SIAM. J.Cont.Opt.,\6,m-m. Ioffe, A. D. (1979). Différentielles généralisées d'applications localement lipschitziennies d’un espace de Banach dans un autre. CRAS, 289,637-639. Ioffe, A. D. (1981). A new proof of the equivalence of the Hahn-Banach extension. Proc. Am. Math. Soc.,%1,385-390. Ioffe, A. D. (1981). Nonsmooth analysis : differential calculus of nondifferentiable mappings. Trans. Am. Math. Soc., 266,1-56. Ioffe, A. D. (1981). Sous-différentielles approchées de fonctions numériques. CRAS, 292,675-678. Ioffe, A. D. (1982). Nonsmooth analysis and the theory of fans. Convex Analysis and Optimization, J. P. Aubin and R. Vinter (Eds.), Pitman, Boston, 93-117. Ioffe, A. D. (to appear). Approximate subdifferential of nonconvex functions. Trans. Am. Math. Soc. Ioffe, A. D., and Levin, V. L. (1972). Subdifferentials of convex functions. Trans. Moscow Math. Soc.,26,\-12. Ioffe, A. D., and Tihomirov, V. M. (1974). The Theory of Extremal Problems. Nauka, Moscow (English translation, North-Holland, Amsterdam, 1979). loss, G., and Joseph, D. (1980). Elementary Stability and Bifurcation Theory. Springer-Verlag, Berlin Istratescu, V. I. (1981). Fixed point theory. Reidel, Dordrecht. Itoh, S., and Takahashi, W. (1977). Single valued mappings, multivalued mappings and fixed-point theorems. J. Math. Anal. Appl., 59,514-521. James, R. C. (1964). Weakly compact sets. Trans. Amer. Math. Soc., 113,129-140. Janin, R. (1982). Sur des multiapplications qui sont des gradients généralisés. CRAS, 294,115-117. Jorna, S. (1978). Topics in nonlinear dynamics. Conference Proceedings, 46. American Institute of Physics, New York. Kakutani, S. (1941). A generalization of Brouwer’s fixed point theorem. Duke Math. J., 8,457-459. Kato, T. (1967). Nonlinear semi-groups and evolution equations. J. Math. Soc. J., 19, 508-520. Kirk, W. A. (1965). A fixed point theorem for mappings which do not increase distance. Am. Math. Mo«., 72,1004-1006. Kirk, W. A. (1976). Caristi’s fixed point theorem and metric convexity. Colloq. Math., 81-86. Kirk, W. A., and Caristi, J. (1975). Mapping theorems in metric and Banach spaces. Bull. Acad. Polon.Sci.,23,391-394. Knaster, B., Kuratowski, C., and Mazurkiewicz, C. (1926). Ein Beweis des Fixpunktsatzes fur n dimensionale Simplexe. Fund. Math., 14,132-137. Krasnoselski, M. A. (1963). Topological Methods in the Theory of Nonlinear Integral Equations. Pergamon Press, New York. Krasnoselski, M. A. (1964). Positive Solutions of Operator Equations. Noordhoff, Groningen. Krasnoselski, M. A., and Rutickii, Y. B. (1961). Convex Functions and Orlicz Spaces. Noordhoff, Groningen. Krasnoselski, M. A., Zabreiko, P. P. (1975). Geometrical methods of nonlinear analysis (in Russian). Nauka, Moscow. Kuhn, H. W., and Tucker, A. W. (1951). Nonlinear Programming. Proceedings of the 2d Berkeley Symposium on Mathematical Statistics and Probability, University of California Press, Berkeley, pp.481-492. Kuratowski, K. (1958). Topologie, Vols. 1 and 2 4th. Ed. corrected. Panstowowe Wyd Nauk, Warszawa. (Academic Press, New York, 1966. Kutateladze, S. S. (1977). Subdifferentials of convex operators. J. Math. Sibirsk, 13, 1057-1064. Ladde, G., and Lakshmikantham, V. (1974). On flow invariant sets. Рас. J. Math., 51, 215-220.
BIBLIOGRAPHY 507 Landesman, E., and Lazer, A. (1970). Nonlinear perturbations of linear elliptic boundary value problems at resonance. J. Math. Mech., 19,609-623. Lang, S. (1962). Introduction to differentiable Manifolds. Wiley-Interscience, New York. Lasota, A., and Olech, C. (1966). On the closedness of the set of trajectories of a control system. Bull. Acad. Pol. Sc., 14,615-621. Lasota, A., and Olech, C. (1968). On Cesari’s semicontinuity condition for set valued mappings. Bull. Acad. Pol. Sc., 16,711-716. Lasota, A., and Opial, Z. (1965). An application for the Kakutani-Ky Fan theorem in the theory of ordinary differential equations. Bull. Acad. Pol. Sc., 13,781-786. Lasota, A., Opial, Z. (1968). Fixed point theorems for multivalued mappings and optimal control problems. Bull? Acad. Pol. Sc., 16, 645-649. Laurent, P. J. (1972). Approximation et Optimisation. Hermann, Paris. Lazer, A. (1972). Application of a lemma on bilinear forms to nonlinear oscillations. Proc. Am. Math.Soc.,33,89-94. Lazer, A., and Sanchez, P. (1969). On periodically perturbed conservative systems. Michigan Math. J., 16, 193-200. Leach, E. B. (1961). A note on inverse function theorems. Proc. Am. Math. Soc., 12,694-697. Lebourg, G. (1975). Valeur moyenne pour un gradient généralisé. CRAS, 281, 795-797. Lebourg, G. (1978). Problèmes d'optimalisation dépendant d’un paramètre à valeurs dans in espace de Banach. Séminaire d’Analyse Fonctionnelle, 1978-1979, Ecole Polytechnique, 11. Lebourg, G. (1979). Generic differentiability of Lipschitzian functions. Trans. Am. Math. Soc., 256, 125-144. Leibniz, G. (1684). Nova methoda pro maximis et minimis. Lemaréchal, C. (1975). An extension of Davidon methods to nondifferentiable problems. Math programming study. Nondifferentiability Optimization, North-Holland, Amsterdam, 95-109. Lempis, F., and Maurer, H. (1980). Differential stability in infinite dimensional nonlinear pro¬ gramming. Appl. Math. Opt., 6,139-152. Leray, J. (1946). Sur les équations et les transformations. J. Math. Pure Appl., 24,201-248. Leray, J. (1952). La théorie des points fixes et ses applications en analyse. Proc. Int. Cong. Math. (Cambridge 1950) Am. Math. Soc., 2,202-208. Leray, J. (1959). Théorie des points fixes, indice total et nombre de Lefschetz. Bull. Soc. Math. Fr., 87,221-233. Leray, J., and Lions, J. L. (1965). Quelques résultats de Visik sur les problèmes elliptiques nonlinéaires. Bull. Soc. Math. Fr., 93,97-107. Leray, J., Schauder, J. (1934). Topologie et équations fonctionnelles. Ann. Sci. ENS, 51, 45-78. Levitin, E. S. (1974). On differential properties in the optimum value in parametric problems of mathematices programming. Dokl. Acad. Nauk SSSR, 215, (Sov. Math. Dokl. 1974,15, 603- 608). Lichtenberg, A., and Liebermann, M. (1983) Regular and stochastic motion, Springer-Verlag, Heidelberg. Lions, J. L. (1968). Contrôle Optimal de Systèmes Gouvernés par des Équations aux Dérivées Partielles. Dunod, Paris. Lions, J. L. (1969). Quelques Méthodes de Résolution de Problèmes Nonlinéaires. Dunod, Paris. Lions, J. L. (1976). Sur Quelques Questions d*Analyse, de Mécanique et de Contrôle Optimal. Presses de l’Université de Montréal, Montréal. Lions, P. L. (1977). Approximation de point fixe de contractions. CRAS, 284,1357-1359. Lions, P. L. (1978). Produits infinis de résolvantes. Isr. J. Math., 29,329-345.
508 BIBLIOGRAPHY Lions, P. L. (1978). Une méthode itérative de résolution d’une équation variationnelle. Isr. J. Math., 31,204-208. Lions, P. L. (1980). Résolution des équations de Hamilton-Jacobi-Bellman pour des opérateurs uniformément elliptiques. 290,1049-1052. Lions, P. L. (1981). Minimization problems in LL J. Funct. Anal., 41,236-215. Lions, P. L. (1981). Solutions généralisées des équations de Hamilton-Jacobi du 1er ordre. CFA S, 292,953-956. Lions, P. L. (1982). Generalized Solutions of Hamilton-Jacobi Equations. Pitman, Boston. Lions, J. L., and Magenes, E. (1968). Problèmes aux Limites Non Homogènes. (3 vol.) Dunod- Gauthier-Villars, Paris. Lions, J. L., and Stampacchia, G. (1965). Inéquations variationnelles noncoercives. CRAS, 261, 25-27. Lions, J. L., and Stampacchia, G. (1967). Variational inequalities. Comm. Pure Appl. Math., 20, 493-519. Lipschitz, R. (1877). Lehrbuchder Analyse, Bonn. Liusternik, L. A. (1947). The Topology of the Calculus of Variations in the Large. (English trans. Vol. 16, American Mathematical Society, 1966). Liusternik, L. A., and Schnirelman, L. G. (1930). Méthodes Topologiques dans les Problèmes Variationnels. Hermann, Paris. Makarov, V. L., and Rubinov, A. M. (1970). Superlinear point to set mappings of economic dynamics (in Russian). Uspehi Mat. Nauk, 25,125-169. Makarov, V. L., and Rubinov, A. M. (1973). Mathematical theory of economical dynamics and equilibria (in Russian). Nauka, Moscow. Mangasarian, O. L. (1966). Sufficient conditions for the optimal control of nonlinear systems. SIAMJ. Contr. Opt.,4,139-152. Mangasarian, O. L. (1969). Nonlinear Programming. McGraw-Hill, New York. Mangasarian, O. L., and Fromovitz, S. (1967). The Fritz John necessary optimality conditions in presence of equality and inequality constraints. J. Math. Anal. Appl., 17,37-47. Markin, J. T. (1973). Continuous dependence of fixed point sets. Proc. Am. Math. Soc., 38, 545-547. Martelli, M., and Vignoli, A. (1974). On differentiability of multivalued maps. Boll. Un. Mat. Stat., 10,701-712. Martin, R. M. (1973). Differential equations on closed subsets of a Banach space. Trans. Am. Math. Soc.,\19,399-4\4. Martin, R. M. (1976). Nonlinear Operators and Differential Equations in Banach Spaces. Wiley- Interscience, New York. Maschler, M., and Peleg, B. (1976). Stable sets and stable points of set-valued dynamical systems. SiamJ. Cont. Opt., 14,985-995. Maurer, H. (1979). First order sensitivity of the optional value function in mathematical programm¬ ing and optimal control. Proc. Symp. Math. Prog, with Data Perturbations. Washington, D.C. Maurer, H. (1979). Differential stability in optimal control problems. Appl. Math. Opt., 5,283-295. McCormick, G. P. (1975). Optimality criteria in nonlinear programming. In R. W. Cottle (ed.). Proceedings of the Symposium on Applied Mathematics, Vol. 9. Me Kenzie, L. (1959). On the existence of general equilibrium for a competitive market. Econome- trica,27,54-11. McLeod, R. M. (1965). Mean value theorems for vector-valued functions. Proc. Edinburgh Math. Soc. 14(2), 197-209. MeShane, E. J. (1939). On multipliers for Lagrange problems. Am. J. Math., 61,809-819. Michael, E. (1956). Continuous selections 1. Ann. Math., 63, 361-381.
BIBLIOGRAPHY 509 Michael, E. (1956). Continuous selections 2. Am. Math., 64,562-580. Michael, E. (1957). Continuous selections 3. Ann. Math., 65,375-390. Michael, E. (1959). A theorem on semicontinuous set valued functions. Duke Math. J., 26,647-651. Michel, P. (1974). Problèmes d’optimisation définispar des fonctions qui sont sommes de fonctions convexes et dérivables. J. Math. Pure. AppL, 53,321-330. Milnor, J. (1963). Morse Theory. Princeton University Press. Milnor, J. (1965). Topology from the Differentiable Viewpoint. University of Virginia, Charlottesville. Milnor, J. (1978). Analytic proof of the hairy ball theorem and the Brouwer fixed point theorem. Am. Math. Mon., 85,521-524. Minty, G. (1961). On the maximal domain of a monotone function. Mich. Math. J., 8, 135-137. Minty, G. (1962). Monotone (nonlinear) operators in a Hilbert space. Duke Math. J., 29, 341-348. Minty, G. (1964). On the monotonicity of the gradient of a convex function. Рас. J. Math., 14, 243-247. Minty, G. (1965). A theorem on maximal monotone sets in Hilbert space. J. Math. Anal. Appl., 11, 434-439. Minty, G. (1967). On the generalizing of the direct method of the calculus of variations. Bull. Am. Math.Soc.,l\l\5-m. Minty, G. (1974). A finite-dimensional tool-theorem in monotone operator theory. Adv. Math., 12, 1-7. Mirica, S. (1980). A note on the generalized differentiability of mappings. Trans. Am. Math. Soc., 4,567-575. Moreau, J. J. (1965). Proximité et dualité dans un espace Hilbertien. Bull. Soc. Math. Fr., 93, 273-299. Moreau, J. J. (1967). Fonctionnelles Convexes. Collège de France, Paris. Moreau, J. J. (1974). La convexité en statistique. Analyse Convexe et ses Applications, J. P. Aubin (Ed.), Springer-Verlag, Berlin. Moreau, J. J. (1978). Un cas de convergence des itérés d’une contraction d’un espace Hilbertien. 286,143-144. Mosco, U. (1969). Convergence of convex sets and of solutions of variational inequalities. Adv. Math.,3,5\0-5S5. Mosco, U. (1970). Perturbations of variational inequalities. Proc. Symp. Pure Math., Am. Math. 5oc., 18,182-194. Mosco, U. (1976). Implicit Variational Problems and Quasi-Variational Inequalities. Lecture Notes 543, Springer-Verlag, Berlin. Moser, J. (1966). A rapidly converging iteration method and nonlinear partial differential equations. Ann. Scuola Norm. Pisa, 265-315,499-535. Moser, J. (1973). Stable and Random motion in dynamical systems. Princeton University Press, Princeton. Moser, J. (1976). Periodic orbits near an equilibrium and a theorem by A. Weinstein. Comm. PAM, 29,727-747. Moulin, H. (1980). Théorie des Jeux pour Г Économie et la Politique. Hermann, Paris. Nadler, S. B. (1969). Multivalued contraction mappings. Рас. J. Math.,31^, 475^88. Nash, J. (1950). Equilibrium points in n-person games. Proc. Natl. Acad. Sei. USA, 36, 48-49. Nash, J. (1950). The bargaining problem. Econ., 18,128-160. Nash, J. (1956). The imbedding problem for Riemannian manifolds. Ann. Math., 63,20-63. von Neumann, J. (1929). Zur allgemeinen Theoriedes Masses. Fma/î/. Math., 13,73-116. von Neumann, J. (1932). Zur Operatorenmethode in der klassischen Mechanik. Ann. Math., 33, 587-643.
510 BIBLIOGRAPHY von Neumann, (1937). Uber ein Ökonomische Gleichungssystem und eine Verallgemeinerung des Brouwersschen Fixpunktsatzes. Ergebnisse eines Math. Collo.,%, 73-83. von Neumann, J., and Morgenstern, O. (1944). Theory of Games and Economic Behaviour. Princeton University Press, Princeton. Neustadt, L. (1976). Optimization. Princeton University Press. Nikaido, H. (1968). Convex Structure and Economic Theory. Academic Press, New York. Nirenberg. L. (1972). An abstract form of the Cauchy-Kowalewski theorem. J. Diff. Geom., 6, 561-576. Nirenberg, L. (1974). Topics in Nonlinear Functional Analysis. Lecture Notes, New York University, New York. Nirenberg, L. (1981). Variational and topological methods in nonlinear problems. Bull, Am. Math. 5oc., 4,267-301. Numinskii, E. A. (1978). On differentiability of multifunctions. Kibernetika, 6,46-48. Olech, C. (1965). A note concerning extremal points of a convex set. Bull. Acad. Pol. Sc., 13,347-351. Olech, C. (1965). A note concerning set-valued measurable functions. Bull. Acad. Pol. Sc., 13, 317- 321. Olech, C. (1967). Lexicographical order, range of integral and bang-bang principle. Mathematical Theory of Control, Academic Press, New York, 35-45. Olech, C. (1968). Approximation of set-valued functions by continuous functions. Colloq. Math., 19,285-293. Olech, C. (1969). Existence theorems for optimal control problems involving multiple integrals. J. Diff.Eq.,b,5\2-S2(i. Olech, C. (1975). Existence of solutions of nonconvex orientor fields. Boll. Un. Mat. Ital., 11(4), 189-197. Opial, Z. (1967). Weak convergence of the sequence of successive approximations for nonexpansive mappings. Bull. Am. Math. Soc.,13,591-597. Palais, R. S. (1963). Morse theory on Hilbert manifolds. Topology, 2,299-340. Palais, R. S. (1966). Ljusternik-Schnirelman theory on Banach manifolds. Topology, 5, 115-132. Palais, R. S. (1970). Critical point theory and the minimax principle. Proc. Symp. Pure Math. Am. Math.Soc.,\S,\%5-2\2. Palais, R. S., and Smale, S. (1964). A generalized Morse theory. Bull. Am. Math. Soc., 70,165-171. Panagiotopoulos, P. D. (1980). Superpotentials in the sense of Clarke and in the sense of Warga and applications. ZAMM, 62. Panagiotopoulos, P. D. (to appear). Optimal control and parametrized identification of structures with convex or nonconvex strain energy density. Solid Mech. Archives. Papageorgiou, N. S. (to appear). Nonsmooth analysis on partially ordered spaces. Part 2: nonconvex case, Clarke’s theory. Рас. J. Math. Pareto, V. (1909). ManuelcTEconomie Politique. Girard & Briere, Lausanne. Pascali, D., and Sburlan, S. (1979). Nonlinear Mappings of Monotone Type. Noordhoff, Leiden, Netherlands. Pazy, A. (1977). On the asymptotic behaviour of iterates of nonexpensive mappings in Hibert spaces. Isr.J.Math.,16,\91-20^. Pazy, A. (1978). On the asymptotic behaviour of semigroups of nonlinear contractions in Hilbert spaces. J. Funct. Anal., 21,293-307. Pazy, A. (1979). Remarks on nonlinear egodic theory in Hilbert spaces. Nonlinear Anal. TMA, 3, 863-871. Pazy, A. (1983). Semigroups of Linear Operators and Applications. Springer-Verlag, Berlin. Pchenitchny, B. W. (1971). Necessary Conditions for an Extremum. Dekker, New York.
BIBLIOGRAPHY 511 Pchenitchny, В. W. (1980). Convex Analysis and Extremal Problems (in Russian). Nauka, Moscow. Peano,G.(1892). Sur la définition de la dérivée. Mathesis, 2,12-14. Penot, J. P. (1973). Calcul différentiel dans les espaces vectoriels topologiques. Stud. Math., M, 1-23. Penot, J. P. (1974). Sous-différentiels de fonctions numériques nonconvexes. Cr. Acad. Sei. Paris, 278,1553-1555. Penot, J. P. (1978). The use of generalized subdifferential calculus in optimization theory. In Methods of Operations Research, Vol. pp. 31,495-511. Athenäum, Berlin. Penot, J. P. (1978). Utilisation de sous-différentiels généralisés en optimisation. First meeting AFCET-SMF on Applied Mathematics Palaiseau, AFCET-SMF, Paris. Penot, J. P. (1978). Calcul sous-différentiel et optimisation. J. Funct. Anal., 27(2), 248-276. Penot, J. P. (1979). Fixed point theory without convexity. Bull. Soc. Math. Fr. Mém., 60, 129-152. Penot, J. P. (1981). A characterization of tangential regularity. Nonlinear Anal. TAM, 5, 625-643. Petcherskaja, N. A. (1980). On differentiability of set-valued maps (in Russian). Vestnik Leningrad University, 115-117. Petryshyn, W. V. (1968). Fixed point theorems for P-compact, semi-contractive and accretive operators not defined on all of a Banach spaces J. Math. Anal. Appl., 23,336-354. Petryshyn, W. V. (1970). Nonlinear equations involving noncompact operators. Nonlinear Func¬ tional Analysis, Proceeding of Symposia in Pure Mathematics, 10, Vol. 1, pp. 206-233. Picard, E. (1890). Mémoire sur la théorie des équalions aux dérivées partielles et la méthode des approximation successives. J. Math.y 6,145-210, Poincaré, H. (1899). Les Méthodes Nouvelles de la Mécanique Céleste. Gauthier-Villars, Paris. Prodi, G., and Ambrosetti, A. (1973). Analisi Nonlineare. Scuola Normale Superiore, Pisa. Rabinowitz, P. (1974). Variational methods for nonlinear eigenvalue problems. CIME, Edizioni Cremonese, 141-195. Rabinowitz, P. (1977). A variational methods for finding periodic solutions of differential equations. Proceedings Symposium on Nonlinear Evolution Equations, Academic Press, New York, pp. 225-251. Rabinowitz, P. (1979). Periodic solutions of a Hamiltonian system on a prescribed energy surface. J.Diff.Eq.,33,336-352. Rabinowitz, P. (1980). Subharmonic solutions of Hamiltonian systems. Comm. PAM, 33,609-633. Rabinowitz, P. (1982). Periodic solutions of Hamiltonian systems: a survey. SIAM J. Math. An., 13,343-352. Rabinowitz, P. (to appear). Periodic solutions of large norm of Hamiltonian systems. J. Dijf. Eq. Rabinowitz, P. H. (1978). Periodic solutions of Hamiltonian systems. Comm. Pure Appl. Math., 31, 157-184. Ray, W. (1983). A rapidly convergent iteration process and Gâteaux differentiable operators. J. Math. An. Appl., in press. Ray, W., and Cramer, W. (1981). Solvability of nonlinear operator equations. Рас. J. Math., 95, 37-50. Ray, W., and Walker, A. (1982). Mapping theorems for Gâteaux differentiable and accretive operators. J. Nonlinear Ann. TMA,6,423-433. Redheffer, R. M. (1972). The theorems of Bony and Brézis on flow invariant set. Am. Math. Mon., 79,790-797. Redheffer, R. M., and Walter, W. (1974). A diff ineq for the distance function in normed linear spaces. Math. Ann., 211,299-314. Reich, S. (1977). Nonlinear evolution equations and nonlinear ergodic theorems. Nonlinear Anal. ГЛ/у4, 1,319-330. Reich, S. (1978). Approximate selections, best approximations, fixed points and invariant sets. J. Math. Anal. Appl., 62,104-113.
512 BIBLIOGRAPHY Reich, S. (1980). Convergence and approximation of nonlinear semigroups. J. Math. Anal. Appl., 76,77-83. Roberts, A. W., and Varberg. (1974). Another proof that convex functions are locally Lipschitz. Am. Math. Mon., 81,1014-1016. Robinson, S. (1976). Regularity and stability for convex multivalued functions. Math. Op. Res., 1, 130-143. Robinson, S. (1976). Stability theory for systems of ineq Part 2: differentiable nonlinear systems. SIAMJ. Num. Anal., 13,497-513. Robinson, S. (1979). Generalized equations and their solutions. Part 1. Math. Prog. Stud., 10,128- 141. Robinson, S. (1980). Generalized equations and their solutions. Part 2: application to nonlinear programming. Cahiers de Math, de la Décison, 8006, Université de Paris-Dauphine. Robinson, S. (to appear). Some continuity properties of polyhedral multifunctions. Math. Prog. Studies. Robinson, S. (Ed.), (1980). Analysis and computation of fixed points. Academic Press, New York. Rockafellar, R. T. (1966). Characterization of the subdifferentials of convex functions. Pacific J. Л/агЛ., 17,497-510. Rockafellar, R. T. (1967). Monotone processes of convex and concave type. Mem. of AMS, 77. Rockafellar, R. T. (1967). Conjugates and Legendre transforms of convex functions. Can. J. Math., 19,200-205. Rockafellar, R. T. (1968). Duality in nonlinear programming. Lectures in Applied Math., Vol. 2, Math. Soc., pp. 401-422. Rockafellar, R. T. (1968). Integrals which are convex functionals. Рас. J. Math., 24, 525-540. Rockafellar, R. T. (1969). Measurable dependence of convex sets and functions oparameters. J. Math. Anal. Appl., 28,4-25. Rockafellar, R. T. (1970). Convex Analysis. Vol. 28 of Princeton Math. Series, Princeton Univ. Press 1970. Rockafellar, R. T. (1970). Generalized Hamiltonian equations for convex problems of Lagrange. Рлс.У.Ма/Л.,33,411^28. Rockafellar, R. T. (1970). Confugate convex functions in optimal control and the calculus of variations. J. Math. Anal. Appl., 332,174-222. Rockafellar, R. T. (1971). Integrals which are convex functionals 2. Рас. J. Math., 39, 429-469. Rockafellar, R. T. (1971). Existence and duality theorems for convex problems of Bolza. Trans. Am. Math.Soc.,\S9,\^(i. Rockafellar, R. T. (1972). State constraints in convex problems of Bolza. SIAMJ. Control, 10,691- 715. Rockafellar, R. T. (1973). Optimal arcs and the minimum value function in problem of Lagrange. Trans. Am. Math. Soc., 180,53-83. Rockafellar, R. T. (1974). Confugate duality and optimization. No. 16 in Conference Board of Math. Sci. Series, SIAM Publications. Rockafellar, R. T. (1974). Convex algebra and duality in dynamic models of production. Mathe¬ matical Models in Econmics, Los (Ed.), North-Holland Amsterdam. Rockafellar, R. T. (1975). Existence theorems for general control problems of Bolza and Lagrange. Adv.Math.,\5,3\2-333. Rockafellar, R. T. (1976). Lagrange multipliers in optimization. SIAM-AMS Proc Vol. 9 (R. W. Cottle and C. E. Lemke(Eds.), pp. 145-168. Rockafellar, R. T. (1976). Integrals functionals, normal integrands and measurable selections. Nonli Op and the Calc of Var (L. Waelbroeck (Eds.), Lecture Notes in Mathematics 543, S-V, 157-207.
BIBLIOGRAPHY 513 Rockafellar, R. T. (1979). Clarke’s tangent cones and the boundaries of closed sets in Rn. Nonlinear Anal. Th. Meth. Appl.,3y 145-154. Rockafellar, R. T. (1979). The theory of subgradients and its applications to problems of opti¬ mization Convex and nonconvex functions, Helderman Verlag, W. Berlin, 1981. Rockafellar, R. T. (1979). La théorie des sous-gradients et ses applications à l’optimisation. Fonct Conv. et Nonconv., Col. Chaire Aisenstadt, Presses de l’Univ. de Montréal, 130. Rockafellar, R. T. (1979). Directionally Lipschitzian functions and subdiflferntial calculus. Proc. London Math. Soc., 39,331-335. Rockafellar, R. T. (1980). Generalized directional derivatives and subgradients of nonconvex functions. Can. J. Math., 32, 157-180. Rockafellar, R. T. (1982). Proximal subgradients, marginal values, and augmented Lagrangians. Math.Op.Res.,6,421-431. Rockafellar, R. T. (1982). Lagrange multipliers and subderivatives of optimal value funct in nonlin¬ ear prog. Math. Program. Stud., 17,28-66. Rockafellar, R. T. (1982). Favourable classes of Lipschitz continuous functions in subgradient optimization. Nondifferential Optimization, E. Nurminski (Ed.), Pergamon Press, New York. Rockafellar, R. T. (1982). Augmented Lagrangians ans marginal values in parametric optimization problems. Generalized Lagralgian Methods in Optimization, A. Wierzbicki (Ed.), Pergamon Press, New York. Rockafellar, R. T. (to appear). Directional differentiability of the optimal value function in a non¬ linear programming problem. Math. Program. Stud. Rockafellar, R. T. (to appear). Marginal values and second-order conditions for optimality. Math. Program. Studies. Rockafellar, R. T., and Wets, R. (1976). Stochastic convex programming. SIAM J. Contr. Opt., 14, 574-589. Rogalski, M. (1972). Surjectivité d’applications multivoques dans les convexes compacts. Bull. Sc. Mat. (1970). Some remarks on vector fields in Hilbert spaces. Proc. Symp. Pure Math., Am. Ma//2.5oc., 18(1), 251-269. Rouche, N., and Machwin, J. (1973). Equations Différentielles Ordinaires. Masson, Paris. Rubinov, A. M. (1968). Dual models of production. Dokl. Akad. Nauk SSSR, 180,795-798. Rubinov, A. M. (1969). Effective trajectories of a dynamical model of production. Dokl. Akad. NaukSSSR,\U{6). Rubinov, A. M. (1969). Point-to-set mappings defined on a cone. Opt. Planirovanis, 14, 96-114. Rubinov, A. M. (1970). Sublinear functionals defined on a cone. Sibirsk. Mat. Z., 11, 429-441. Sadovski, V. N. (1971). Application of topological methods. Theory of periodic solutions of non- differentiable operational equations of neutral, type., Dokl. Akad. Nauk SSSR, 200,1037-1041. Salinetti, G., and Wets, R. (1977). On the relation between two types of convergence for convex functions. J. Math. Anal. Appl., 60,211-226. Salinetti, G., and Wets, R. (1979). On the convergence of sequences of convex sets in finite dimension 57/1M Rev., 21,18-34. Sard, A. (1942). The measure of critical values of differentiable maps. Bull. Am. Math. Soc., 48, 883-890. Scarf, H. (1967). The approximation of fixed points of continuous mappings. SIAM J. Appl. Math., 15,1328-1343. Scarf, H. (1967). The core of a n-person game. Econ., 35, 50-69. Scarf, H. (1971). On the existence of a cooperative solution for a general class of n-person game. J.Econ.Th.,l,\(i9-m. Scarf, H., and Hansen, P. (1973). The Computation of Economic Equilibria. Yale University Press, New Haven, CT.
514 BIBLIOGRAPHY Schauder, J. (1930). Der Fixpunktsatz. Studia Math.,1,171-180. Schinas, J., and Boudourides, M. (1981). Higher order differentiability of multifunctions. Nonlinear 5,509-516. Schwartz, J. T. (1960). On Nash’s implicit function theorem. Comm. Pure Appl. Math., 13,509-530. Schwartz, J. T. (1964). Generalizing the Liusternik-Schnirelman theory of critical points. Comm. Pure Appl. Math., 307-315. Schwartz, J. T. (1969). Nonlinear Functional Analysis. Courant Institute (NYU) (preprint). Schwartz, L. (1966). Théorie des Distributions. Hermann, Paris. Schwartz, L. (1970). Topologie Générale et Analyse Fonctionnelle. Herman, Paris. Shapley, L. S. (1969). On the core of an economic system with externalities. Am. Econ. Rev., 59,678- 684. Shapley, L. S. (1971). Cores of convex games. Int. J. Game Theory, 1,12-26. Shapley, L. S. (1973). On balanced games without side payments. T. C. Hu and S. M. Robinson (Eds.), Mathematical Progress, Academic Press, New York, 261-290. Shapley, L. S. (1975). An example of a slow-converging core. Int. Econ. Rev., 16,345-361. Shapley, L. S. (1976). Noncooperative general exchange. Theory and Measurement of Economic Externalities, Academic Press, New York. Shi, Shu-Chung. (1980). Remarques sur le gradient généralisé. CRAS, 291,443-446. Siegel, J. (1977). A new proof of Caristi’s fixed point theorem. Proc. Am. Math. Soc., 66, 54-56. Smale, S. (1964). Morse theory and a nonlinear generalization of Dirichlet problem. Ann. Math., 80,382-396. Smale, S. (1965). An infinite dimensional version of Sard’s theorem. Ann. J. Math., 87, 861-866. Smale, S. (1976). Exchange processes with price adjustment. J. Math. Econ., 3,211-216. Smale, S. (1976). A convergent process of price adjustment and global Newton method. J. Math. Econ.,3,107-120. Smale, S. (1976). Dynamics in general equilibrium theory. Am. Econ. Rev., 66,288-294. Smart, D. R. (1974). Fixed point theorems. Cambridge University Press. Spanier, E. (1966). Algebraic Topology, McGraw Hill, New York. Sperner, E. (1928). Neuer Beweis fur die Invarianz der dimensionzahl und des Gebietes. Ab. Math. Sem. Univ. Hamburg, 6,265-272. Spingarn, J. E. (1981). Submonotone subdifferential of Lipschitz functions. Trans. Am. Math. Soc., 264,77-89. Spingarn, J. E., and Rockafellar, R. T. (1979). The generic nature of optimality conditions in non¬ linear programming. Math. Opt. Res., 4,425-430. Stampacchia, G. (1964). Formes bilinéaires coercives sur les ensembles convexes. CRAS, 258,4413- 4416. Stampacchia, G. (1970). Regularity of solutions of some variational inequalities. Proc. Symp. Pure 18,271-280. Strauss, W. A. (1970). Further applications of monotone methods to partial differential equations. Proc. Symp. Pure Math., 18,282-288. Strodiot, J. J., and Hien Nguyen, V. (1979). Caractérisation des solutions optimales en pro¬ grammation nondifférentiable. CPyl5',288,1075-1078. Sweetser, T. H. (1977). A minimal set-valued strong derivative for vector-valued Lipschitz functions. /. Opt. Th. Appl., 23,549-562. Tartar, L. (1974). Sur l’étude directe d’équations nonlinéaires intervenant en théorie du contrôle optimal. J. Funct. Anal., 17,1-45. Tartar, L. (1974). Inéquations quasi-variationnelles abstraites. CRAS,21%, 1193-1196. Tartar, L. (1975). Equations with order preserving properties. Math. Res. Center Univ. Wisconsin.
BIBLIOGRAPHY 515 Temam, R. (1970). Solutions généralisées de certains problèmes de calcul des variations. CRAS, Paris, 271,1116-1119. Temam, R. (1970). Remarques sur la dualité en calcul des variations. CRAS, Paris, 270, 754-757. Temam, R. (1971). Solutions généralisées de certains problèmes de calcul des variations. Arch. Rat. Mech. Anal.,44,121-156. Temam, R. (1975). A nonlinear eigenvalue problem : the shape of equilibrium of a confined plasma. Arch. Rat. Mech. Anal.,60,5\-13. Thibault, L. (1979). Cônes tangents et épi-différentiels de fonctions vectorielles. Travaux Sem. Anal. Convexe, 9(2), Exp. No. 13. Thibault, L. (1982). Subdifferentials of nonconvex vector-valued functions. J. Math. Anal. Appl., 86,319-344. Thom, R. (1972). Stabilité Structurelle et Morphogénèse. Benjamin, New York. Toland, J. F. (1973). Asymptotic linearity and nonlinear eigenvalues problems. Q. J. Math., 24, 241-250. Toland, J. F. (1978). Duality in nonconvex optimization. J. Math. Anal. Appl., 66, 399-415. Toland, J. F. (1979). A duality principle for nonconvex optimization and the calculus of variations. Arch. Rat. Mech. Anal., 71,41-61. Toland, J. F. (1979). Stability of heavy rotating chains. J. Diff. Eq., 32,15-31. Trêves, F. (1967). Topological Vector Spaces, Distributions and Kernels. Academic Press, New York. Trêves, F. (1976). Basic Linear Partial Differential Equations. Academic Press, New York. Tuy, Hoang (1964). Concave programming under linear constraints. Doki Acad. Nauk. SSSRy 159, 32-35. Tuy, Hoang (1980). Solving equations Об f(x) under general boundary conditions. Numerical Solu¬ tions of Highly Nonlinear Prob, W. Forster (Ed.), Elsevier North-Holland, Amsterdam, 271-296. Tuy, Hoang (1981). On variable dimension algorithms and algorithms using primitive sets. Math. Oper. Forsch. Stat. Ser. Optimization, 12, 361-381. Ursescu, C. (1973). Sur une généralisation de la méthode de différentiabilité. Ac. Naz. Lincei, 8, 199-204. Ursescu, C. (1975). Multifunctions with closed convex graph. Czech. Math. J., 25,438-441. Uzawa, H. (1962). Competitive equilibrium and fixed point theorems. Econ. Status Q., 8, 59-62. Vainberg, M. M. (1956). Variational methods for the study of nonlinear operators (English trans.). Holden-Day, San Francisco, 1964. Vainberg, M. M., and Trenogin, V. A. (1962). The methods of Lyapounov and Schmidt in the theory of nonlinear equations and their development. Uspeki Mat. Nauk, 17,13-75. Valadier, M. (1970). Contribution à Г Analyse Convexe. Thesis, Paris. Valadier, M. (1972). Sous-différentiabilité de fonctions convexes à valeurs dans un espace vectoriel ordonné. Math. Scand., 30,65-74. Valentine, F. A. (1960). Convex Sets. McGraw-Hill, New York. Wagner, D. M. (1977). Survey of measurable selection theorems. SIAMJ. Contr. Opt., 15,859-903. Walras, L. (1874). Eléments <TEconomie Politique Pure. Corbaz, Lausanne. Warga, J. (1962). Relaxed variational problems. J. Math. Anal. Appl., 4,111-128. Warga, J. (1972). Optimal Control of Differential and Functional Equations. Academic Press, New York. Warga, J. (1976). Derivate containers, inverse functions and controllability. Calc of Var and Control Theory, D. L. Russel (Ed.), Academic Press, New York, 33-46. Warga, J. (1978). An implicit function theorem without differentiability. Proc. Am. Math. Soc., 69, 65-69. Warga, J. (1978). Controllability and a multiplier rule for nondifferentiable optimization prob. SIAMJ. Com. Opt., 16,803-812.
516 BIBLIOGRAPHY Weinstein, A. (1973). Normal modes for nonlinear Hamiltonian systems. Im. Math., 20, 47-57. Weinstein, A. (1978). Periodic orbits for convex Hamiltonian systems. Ann. Math., 108, 507-518. Willem, M. (1981). Density of the range of potential operators. Proc. Am. Math. Soc., AMS, 83, 341-344. Willem, M. (1982). Remarks on the dual least action principle. Zeit.fur Anal, und Anw., 91-95. Willem, M. (to appear). Subharmonic oscillations of convex Hamiltonian systems. Nonlinear Anal., TMA. Wolfe, P., and Balinski, M. (1975). Nondifferentiable optimization. Math. Program Studies, North- Holland, Amsterdam. Yorke, J. A. (1967). Invariance for ordinary differential equations. Math. Syst. Theory, 1, 353-372. Yorke, J. A. (1969). Invariance for contingent equations. Lecture Notes in Operations Research and Mathematical Economics, Vol. 12, Springer-Verlag, Berlin, 379-381. Yorke, J. A. (1970). Differential inequalities and non-Lipschitz scalar functions. Math. Syst. Theory, 4,140-153. Yorke, J. A. (1971). Another proof of the Liapounov convexity theorem. SIAM J. Control Opt., 9,351-353. Yoshizawa, T. (1966). Stability Theory by Liapounov's Second Method. The Mathematical Society of Japan, Tokyo. Yoshizawa, T. (1975). Stability Theory and the Existence of Periodic Solutions and Almost Periodic Solutions. Springer-Verlag, Berlin. Yosida, K. (1974). Functional Analysis. Springer-Verlag, Berlin. Young, L. C. (1969). Lectures on the Calculus of Variations and Optimal Control Theory. Saunders, Philadelphia. Zarantonello, E. (1960). Solving functional equations by contraction averaging. Math. Res. Center Rep.n. 160, University of Wisconsin, Madison. Zarantonello, E. (1964). The closure of the numerical range contains the spectrum. Bull. Am. Math. 5oc., 70,781-787. Zarantonello, E. (1971). Contributions to Nonlinear Functional Analysis. Academic Press, New York. Zarantonello, E. (1973). The product of commuting projections is a projection. Proc. Am. Math . Soc.,3%, 591-594. Zlobec, C., and Massam, H. (1974). Various definitions of the derivatives in mathematical program¬ ming. Math. Prog., 1,144-161. (Addendam, 14,108-111,1978). Zowe, J. (1975). A duality theorem for a conv prog theorem in ordered complete vector latticex. J. Math. Anal. Appl., 50,273-287.
Index Ambrosetti- Rabinowitz theorem, 272 Asymptotic center, 251 Attractor (of a sequence), 252 Baire’s theorem, 7 Barrier Cone, 26, 143 Bifurcation value, 71. See also Hopf bifurcation Bipolar lemma, 30 Brouwer’s fixed point theorem, 42, 46, 354 Caristi’s theorem (fixed point), 248 Closed graph theorem, 28 Closed image theorem, 110, 112, 117, 132 Codifferential, 178, 184, 215 Condition (C) of Palais and Smale, 269, 270 Conjugate function, 201, 207, 212, 221, 225, 232, 236 Conservative strategy, 300, 334 Conservative value, 304, 307 Consistent multistrategy, 300, 334 Contingent cone, 405, 409 Contingent derivative, 430, 492 Contraction, 243 Convergence theorem, 123 Crandall - Rabinowitz theorem, 73, 76 Critical point, 52 Decision rules, 299 canonical, 302, 307 Derivative: of a function, 446, 447, 448, 449 of functional, 16 of a map, 413, 430, 439, 444 of a map with convex graph, 178, 184 Differentiable functionals, 16 Differentiable maps, 23 Dissipative system, 244, 247 Domain; of a function, 3 of a set valued map, 1 Dual least action principle, 237, 456 Dynamical system, 239, 247 Epicontingent derivative, 418, 446, 449 Epiderivative, 421 of a convex function, 186, 188, 190, 198 Epigraph, 3 Equilibrium, 336, 337, 339, 341, 343, 346 economic, 360 Ky Fan, 151 Nash, 302, 335 noncooperative, 302, 335 social, 351 Stackelberg, 310 Walras, 358 Ergodic theorem, 253 Farkas’s lemma, 144 Fenchel inequality, 202 Finite topology, 331, 373 Fixed point, 42, 46, 241, 244, 248, 249, 344 Fr6chet differentiability: of a function, 16 of a map, 23 Generalized second derivative, 226, 428 Generic differentiability, 282 Generic property, 9, 83, 85 Graph of a map, 2 Hadamard lemma, 52 Hahn - Banach theorem, 27 Hamiltonian, 233, 236, 237 Hamilton inclusion, 233 Hopf bifurcation, 81 Implicit function theorem, 435, 436 Index of nondegenerate critical point, 55 Indicatrix, 60 Inverse function theorem, 34, 429, 431, 433, 435 517
518 INDEX Kakutani fixed point theorem, 344, 348 Knaster - Kuratowski - Mazurkiewicz (KKM) lemma, 353 Krasnosel’skii theorem, 20 Ky-Fan: equilibrium, 151 fixed point theorem, 344 inequality, 296, 327, 330, 332 Lagrange multiplier, 219, 220, 222, 232, 233, 235 Lagrangian, 232, 233, 235, 236 of a minimization problem, 218, 235 Lax-Milgram theorem, 145 Least action principle, 237, 454, 456 Leray-Shauder theorem, 346 Lipschitz function, 196 Lipschitz maps, 121, 134, 414 Liustemik theorem, 433 Lower semicontinuous function, 11 Lower semicontinuous map, 108 Lyapunov - Schmidt procedure, 71 Manifold, 49 Marginal function, 4, 118, 119, 120, 121, 214, 215, 218 Maximal monotone map, 194, 378, 380, 386, 389, 395, 396 Maximum theorem, 118, 119, 120, 121 M-convex processes, 156 Minimax theorem, 318, 320, 321, 326, 327 lop-sided, 295, 319, 322 M-matrices, 106, 157 Monotone function, 332 Monotone map, 369, 371, 372, 374 Morse function, 52, 87 Morse lemma, 52, 55 Nash equilibrium, 302 Nash theorem, 335 noncooperative equilibrium, 302, 335 nonexpansive maps, 249, 370 Normal cone: to convex subsets, 168, 169, 174 normal solvability, 434 Palais Smale, condition (C) of, 269, 270 Pareto optimum, 302, 306 Perron- Frobenius theorem, 106, 153, 155 Positive maps, 147, 150 proper set valued map, 1 of a function, 3 Subdifferential, 187, 188, 190, 198, 203, 212, 215,218, 262, 421,446, 448 Support function, 26, 27, 30, 31, 204 €-Supporting functional, 280 Tangent cone: to convex subsets, 166, 169, 174 to subsets, 406, 409, 440 Thom’s transversality theorem, 84 Transpose of a set valued map, 139, 140, 141, 178, 215, 216 Transversality, 50, 83, 84, 87, 89, 95 Upper hemicontinuous map, 121 Upper semicontinuous function, 11 Upper semicontinuous map, 108, 373 Variational inequality, 181, 364, 367, 368, 375 Variational principle, 181, 368 €-Variational principle, 255, 258, 262 von Neumann equilibrium, 152 von Neumann - Kemany theorem, 106 Walras equilibrium, 358, 359 Yosida approximation: of a function, 193 of a subdifferential, 195 Yosida approximation map, 383