Selection (1) Texto

METHODS AND SETTINGS FOR SENSITIVITY ANALYSIS
As we already know that for our model all cross-product terms with i = j are zer
o, this can be reduced to the convenient
in which uncertainty is divided among themes , each theme comprising an indicator a
nd its weight. It is easy to imagine similar applications. For example, one coul
d divide uncertainty among observational data, estimation, model assumptions, mod
el resolution and so on.
1.2.16 Further Methods
So far we have discussed the following tools for sensitivity analysis:
derivatives and sigma-normalized derivatives;
regression coefficients (standardized);
variance-based measures;
scatterplots.
We have shown the equivalence of sigma-normalized coefficients S
sensitivity indices Sifor linear models, as well as how Siis a model-free extens
ion of the variance decomposition scheme to models of unknown linearity. We have
discussed how nonadditive models can be treated in the variance-based sensitivi
ty framework. We have also indicated that scatterplots are a powerful tool for se
nsitivity analysis and shown how Sican be interpreted in relation to the existen
ce of shape in an Xiversus Yscatterplot.
At a greater level of detail (Ratto et al., 2007) one can use modern regression t
ools (such as state-space filtering methods) to interpolate points in the scatte
rplots, producing very reliable EYXi=xi*curves. The curves can then be used for
sensitivity analysis. Their shape is more evident than that of dense scatterplot
s (compare Figure 1.7 with Figure 1.8). Furthermore, one can derive the first-or
der sensitivity indices directly from those curves, so that an efficient way to
estimate Siis to use state-space regression on the scatterplots and then take th
e variances of these.
In general, for a model of unknown linearity, monotonicity and additivity, varia
nce-based measures constitute a good means of tackling settings such as factor f
ixing and factor prioritization. We shall discuss one further setting before the
end of this chapter, but let us first consider whether there are alternatives t
o the use of variance-based methods for the settings so far described.
Why might we need an alternative? The main problem with variancebased measures is
computational cost. Estimating the sensitivity coefficients takes many model ru
ns (see Chapter 4). Accelerating the computation of sensitivity indices of all o
rders or even simply of the SiSTicouple
is the most intensely researched topic i
n sensitivity analysis (see the filtering approach just mentioned). It can reaso
nably be expected that the estimation of these measures will become more efficie
nt over time.
At the same time, and if only for screening purposes, it would be useful to have
methods to find approximate sensitivity information at lower sample sizes. One
such method is the Elementary Effect Test.
1.2.17 Elementary Effect Test
The Elementary Effect Test is simply an average of derivatives over the space of
factors. The method is very simple. Consider a model with kindependent input fa
ctors Xii=12k, which varies across plevels. The input space is the discretized p
-level grid . For a given value of X, the elementary effect of the ith input fac
tor is defined as
YX1X2Xi-1Xi+Xk-YX1X2Xk
EEi=(1.48)
where p is the number of levels,
X = is any selected value in such that the transformed point
ei is still in for each index i = 1 and ei is a vector of zeros but with a unit
as its ith component. Then the absolute values of theEEi, computed at r differen
t grid points for each factor, are averaged
r
*1 j
i=EEi(1.49) rj=1
and the factors ranked according to the obtained mean U.
In order to compute efficiently, a well-chosen strategy is needed for moving fro
m one effect to the next, so that the input space is explored with a minimum of
points (see Chapter 3). Leaving aside computational issues for the moment, is a
useful measure for the following reasons:
1.It is semi-quantitative
the factors are ranked on an interval
scale;
2.It is numerically efficient;
3.It is very good for factor fixing
it is indeed a good proxy for
STi;
4.It can be applied to sets of factors.
METHODS AND SETTINGS FOR SENSITIVITY ANALYSIS
Due to its semi-quantitative nature the *can be considered a screening method, e
specially useful for investigating models with many (from a few dozen to 100) un
certain factors. It can also be used before applying a variance-based measure to
prune the number of factors to be considered. As far as point (3) above is conc
erned, *is rather resilient against type II errors, i.e. if a factor is deemed n
oninfluential by *it is unlikely to be identified as influential by another meas
ure.
1.2.18 Monte Carlo Filtering
While *is a method of tackling factor fixing at lower sample size, the next meth
od we present is linked to an altogether different setting for sensitivity analy
sis. We call this factor mapping and it relates to situations in which we are espe
cially concerned with a particular portion of the distribution of output Y. For
example, we are often interested in Ybeing above or below a given threshold. If
Ywere a dose of contaminant, we might be interested in how much (how often) a th
reshold level for this contaminant is being exceeded. Or Ycould be a loss (e.g.
financial) and we might be interested in how often a maximum admissible loss is
being exceeded. In these settings we tend to divide the realization of Yinto good
and bad . This leads to Monte Carlo filtering (MCF, see Saltelli et al., 2004, pp.
151 191 for a review). In MCF one runs a Monte Carlo experiment producing realizat
ions of the output of interest corresponding to different sampled points in the
input factor space, as for variance-based or regression analysis. Having done th
is, one filters the realizations, e.g. elements of the Y-vector. This may entail c
omparing them with some sort of evidence or for plausibility (e.g. one may have
good reason to reject all negative values of Y). Or one might simply compare Yag
ainst thresholds, as just mentioned. This will divide the vector Yinto two subse
ts: that of the well-behaved realizations and that of the misbehaving ones. The sa
me will apply to the (marginal) distributions of each of the input factors. Note
that in this context one is not interested in the variance of Yas much as in th
at part of the distribution of Ythat matters for example, the lower-end tail of
the distribution may be irrelevant compared to the upper-end tail or vice versa,
depending on the problem. Thus the analysis is not concerned with which factor
drives the variance of Yas much as with which factor produces realizations of Yi
n the forbidden zone. Clearly, if a factor has been judged noninfluential by eit
her *or STi, it will be unlikely to show up in an MCF. Steps for MCF are as foll
ows:
A simulation is classified as either B, for behavioural, or B, for nonbehavioural
(Figure 1.11).
Thus a set of binary elements is defined, allowing for the identification of two
subsets for each Xi: one containing a number nof elements denoted
(X|B) B
(X|B)
B
XY
Figure 1.11 Mapping behavioural and nonbehavioural realizations with Monte Carlo
filtering
XBand a complementary set XBcontaining the remaining n=N-nsimulations (Figure 1.
11).
A statistical test can be performed for each factor independently, analy
sing the maximum distance between the cumulative distributions
of the XBand XBsets (Figure 1.12).
If the two sets are visually and statistically18 different, then Xiis an influen
tial factor in the factor mapping setting.
1
Prior
Fm(Xi|B)
0.8
Fn(Xi|B)
0.6
F(Xi )
0.4
dn,n
0.2
0
0 0.2 0.4 0.6 0.8 1
X
Figure 1.12 Distinguishing between the two sets using a test statistic
18 Smirnov two-sample test (two-sided version) is used in Figure 1.12 (see Salte
lli et al., 2004, pp. 38 39).
1.3 NONINDEPENDENT INPUT FACTORS
Throughout this introductory chapter we have systematically assumed that input f
actors are independent of one another. The main reason for this assumption is of
a very practical nature: dependent input samples are more laborious to generate
(although methods are available for this; see Saltelli et al., 2000) and, even
worse, the sample size needed to compute sensitivity measures for nonindependent
samples is much higher than in the case of uncorrelated samples.19 For this reas
on we advise the analyst to work on uncorrelated samples as much as possible, e.
g. by treating dependencies as explicit relationships with a noise term.20 Note t
hat when working with the MCF just described a dependency structure is generated
by the filtering itself. The filtered factors will probably correlate with one
another even if they were independent in the original unfiltered sample. This co
uld be a useful strategy to circumvent the use of correlated samples in sensitivi
ty analysis. Still there might be very particular instances where the use of cor
related factors is unavoidable. A case could occur with the parametric bootstrap
approach described in Figure 1.3. After the estimation step the factors will in
general be correlated with one another, and if a sample is to be drawn from thes
e, it will have to respect the correlation structure.
Another special instance when one has to take factors dependence into considerati
on is when analyst A tries to demonstrate the falsity of an uncertainty analysis
produced by analyst B. In such an adversarial context, A needs to show that B s an
alysis is wrong (e.g. nonconservative) even when taking due consideration of the
covariance of the input factors as explicitly or implicitly framed by B.
1.4
POSSIBLE PITFALLS FOR A SENSITIVITY ANALYSIS
19 Dependence and correlation are not synonymous. Correlation implies dependence
, while the opposite is not true. Dependencies are nevertheless described via co
rrelations for practical numerical computations. 20 Instead of entering X1 and X
2 as correlated factors one can enter X1 and X3, with X3 being a factor describi
ng noise, and model X2 as a function of X1 and X3.
As mentioned when discussing the need for settings, a sensitivity analysis can f
ail if its underlying purpose is left undefined; diverse statistical tests and m
easures may be thrown at a problem, producing a range of different factor rankin
gs but leaving the researcher none the wiser as to which ranking to believe or p
rivilege. Another potential danger is to present sensi-tivity measures for too m
any output variables Y. Although exploring the sensitivity of several model outp
uts is sound practice for testing the quality of the model, it is better, when p
resenting the results of the sensitivity analysis, to focus on the key inference
suggested by the model, rather than to confuse the reader with arrays of sensiti
vity indices relating to intermediate output variables. Piecewise sensitivity an
alysis, such as when investigating one model compartment at a time, can lead to
type II errors if interactions among factors of different compartments are negle
cted. It is also worth noting that, once a model-based analysis has been produce
d, most modellers will not willingly submit it to a revision via sensitivity ana
lysis by a third party.
This anticipation of criticism by sensitivity analysis is also one of the 10 com
mandments of applied econometrics according to Peter Kennedy:
Thou shall confess in the presence of sensitivity. Corollary: Thou shall anticip
ate criticism [ ] When reporting a sensitivity analysis, researchers should explain
fully their specification search so that the readers can judge for themselves h
ow the results may have been affected. This is basically an honesty is the best p
olicy approach, advocated by Leamer, (1978, p. vi) (Kennedy, 2007).
To avoid this pitfall, an analyst should implement uncertainty and sensi-tivity
analyses routinely, both in the process of modelling and in the oper-ational use
of the model to produce useful inferences.
Finally the danger of type III error should be kept in mind. Framing error can o
ccur commonly. If a sensitivity analysis is jointly implemented by the owner of
the problem (which may coincide with the modeller) and a practitioner (who could
again be a modeller or a statistician or a practitioner of sensitivity analysis
), it is important to avoid the former asking for just some technical help from th
e latter upon a predefined framing of the problem. Most often than not the pract
itioner will challenge the framing before anything else.
We Sgs for sensitivity analysis, such as:
1.5 CONCLUDING REMARKS
factor prioritization, linked to Si;
1.
factor mapping, linked to MCF;
metamodelling (hints). factorfixing,shownTior*;
linkedto
have justdifferentsettin
2.
The authors have found these settings useful in a number of applications. This d
oes not mean that other settings cannot be defined and usefully applied.
3.
We have discussed fitness for purpose as a key element of model quality. If the
purpose is well defined, the output of interest will also be well identified. In
the context of a controversy, this is where attention will be focused and where
sensitivity analysis should be concentrated.

4.
As discussed, a few factors often account for most of the variation. Advantage s
hould be taken of this feature to simplify the results of a sensitivity analysis
. Group sensitivities are also useful for presenting results in a concise fashio
n.
5.
Assuming models to be true is always dangerous. An uncertainty/sensitivity analys
is is always more convincing when uncertainty has been propagated through more t
han just one model. Using a parsimonious data-driven and a less parsimonious lawdriven model for the same application can be especially effective and compelling
.
6.
When communicating scientific results transparency is an asset. As the assumptio
ns of a parsimonious model are more easily assessed, sensitivity
The reader will find in this and the following chapters didactic examples.
ed sensitivity measures) can be computed analytically. In most practical
at instances the model under analysis or development will be a computational one
, without a closed analytic formula.
Typically, models will involve differential equations or optimization algor thepurposefamiliarizationby a withsensitivitymeasures. Most of the
analysisshouldofbe followedsimplification.
fo
model
rithms involving numerical solutions. For this reason the best available practice
s for numerical computations will be presented in the following chapters.
exercises will be basedon models whose output(andpossiblytheassoci
For the Elementary Effects Test, we shall offer numerical procedures developed by
Campolongo et al. (1999b, 2000, 2007). For the variance-based measures we shall
present the Monte Carlo based design developed by Saltelli (2002) as well as th
e Random Balance Designs based on Fourier Amplitude Sensitivity Test (FAST-RBD,
Tarantola et al., 2006, see Chapter 4). All these methods are based on true poin
ts in the space of the input factors, i.e. on actual computations of the model a
t these points. An important and powerful class of methods will be presented in
Chapter 5; such techniques are based on metamodelling, e.g. on estimates of the m
odel at untried points. Metamodelling allows for a great reduction in the cost o
f the analysis and becomes in fact the only option when the model is expensive t
o run, e.g. when a single simulation of the model takes tens of minutes or hours
or more. The drawback is that metamodelling tools such as those developed by Ra
tto et al. (2007) are less straightforward to encode than plain Monte Carlo. Whe
re possible, pointers will be given to available software.
1.6 EXERCISES
1. Prove that
VY=EY2-E2Y
2.
Prove that for an additive model of two independent variables X1 and X2, fixing
one variable can only decrease the variance of the model.
3.
Why in *are absolute differences used rather than simple differences?
4.
If the variance of Yas results from an uncertainty analysis is too large, and th
e objective is to reduce it, sensitivity analysis can be used to suggest how man
y and which factors should be better determined. Is this a new setting? Would yo
u be inclined to fix factors with a larger first-order term or rather those with
a larger total effect term?
5.
Suppose X1 and X2 are uniform variates on the interval [0, 1]. What is the mean?
What is the variance? What is the mean of X1 +X2? What is the variance of X1 +X
2?
6.
Compute Sianalytically for model (1.3, 1.4) with the following values: r=2=12and
n =21.
7.
Write a model (an analytic function and the factor distribution functions) in wh
ich fixing an uncertain factor increases the variance.
8.
What would have been the result of using zero-centred distributions for the
Equation (1.27)?
s in
1.7 ANSWERS
1. Given a function Y=fXwhere X =X1X2Xkand X ~pX
where pXis the joint distribution of X withpXdX =1, then the function mean can b
e defined as
EY=fXpXdX
and its variance as
VarY=fX-EY2pXdX
=f2XpXdX +E2Y-2 EYfXpXdX
=EY2+E2Y-2E2Y
=EY2-E2Y
Using this formula it can easily be proven that VarY=VarY+c, with can arbitrary
constant. This result is used in Monte Carlo-based variance (and conditional var
iance) computation by rescaling all values of Ysubtracting EY. This is done beca
use the numerical error in the variance estimate increases with the value of Y.
2. We can write the additive model of two variables X1 and X2 as Y=f1 X1+f2 X2
, where f1 is only a function of X1 and f2 is only a function of X2.
Y2
Recalling that the variance of Ycan be written as VY=EE2 Y, where Estands for the expectation value, and applying it to Y
we obtain
VY=Ef12 +f22 +2f1f2-E2 f1 +f2
Given that Ef1f2=Ef1Ef2for independent variables, then the above can be reduced
to
(
VY=Ef12 +Ef22 -E2 f1-E2 f2
which can be rewritten as
VY=Vf1+Vf2
which proves that fixing either X1 or X2 can only reduce the variance of Y.
3.
Modulus incremental ratios
are used in order to avoid positive and negative
values cancelling each other out when calculating the average.
4.
It is a new setting. In Saltelli
et al. (2004) we called it the variance
cutting setting, when the objective of sensitivity analysis is the reduction of
the output variance to a lower level by fixing the smallest number of input fact
ors. This setting can be considered as relevant in, for example, risk assessment
studies. Fixing the factors with the highest total effect term increases our ch
ances of fixing, besides the first-order terms, some interaction terms possibly
enclosed in the totals, thus maximizing our chances of reducing the variance (se
e Saltelli et al., 2004).
5.
Both X1 and X2 are uniformly distributed in 01, i.e.
pX1=pX2=U01
This means that pXiis 1 for Xi.01and zero otherwise. Thus
1
12
x1
EX1=EX2=pxxdx==x=0 202
Further:
1 1 2 1 1 VarX1=VarX2=pxx-dx=x2 -x+dxx=02x=04
x3 x21 1111
=-+x1 =-+=
3 24032412 EX1 +X2=EX1+EX2=1
as the variables are separable in the integral. Given that X1 +X2 is an additive
model (see Exercise 1) it is also true that 1
VarX1 +X2=VarX1+VarX2=
6
The same result is obtained integrating explicitly
. 1 . 1 VarX1 +X2=pxx1 +x2 -12dx1dx2x1=0 x2=0
6. Note that the model (1.3, 1.4) is linear and additive. Further, its probabilit
y density function can be written as the product of the factors marginal distribu
tions (independent factors). Writing the model for r=2 we have
YZ1Z2=1Z1 +2Z2
with
1 -Z12
Z1 ~N01or equivalently pZ1=veZ122Z21
and a similar equation for pZ2. Note that by definition the distributions are no
rmalized, i.e. the integral of each pZiover its own variable Z1 is 1, so that th
e mean of Ycan be reduced to
EY=1 Z1pZ1dZ1 +2 Z2pZ2dZ2
-f 22
These integrals are of the type xe-xdx, whose primitive -e-x/2 vanishes at the e
xtremes of integration, so that EY=0. Given that the model is additive, the vari
ance will be
VY=VZ1 +VZ2 =V1Z1+V2Z2
For either Z1 or Z2 it will be
VZi=ViZi=2VZi
i
We write
VZi=EZi2-E2Zi=EZi2
and
1 2 -zi/22
EZi2=vzie2 Zidzi2Zi
The tabled form is
v
+
2-at2
tedt=
-04
which gives with an easy transformation
EZi2=Z2 i
so that
22
=
VZiiZi
and
VY=212 +222
Z1 Z2
and
22
iZi
=
SZiVY
Inserting the values =12and f =21we obtain VY=8 and SZ1 =SZ2 =21.
The result above can be obtained by explicitly applying the formula for Sito our
model:
VEYZi
SZi=
VY
which entails computing first EYZi=zi*. Applying this to our model Y=1Z1 +2Z2 we
obtain, for example, for factor Z1:
+***
EYZ1 =z1=pz1pz21z1 +2z2dz1dz2 =1z1
the variance over z*1 of 1z*1
is, as before, equal to 12Z21 and
Hence VZ1
2 iZ2 i
SZi=
212 +222
Z1 Z2
7. We consider the model Y=X1 X2, with the factors identically distributed as
X1X2 ~N0
Based on the previous exercise it is easy to see that
EY=EX1X2=EX1EX2=0
so that
D
VY=EX12X22 =EX12 EX22 =4
If X2 is fixed to a generic value x2*, then
EX1x2*=x2*EX1=0
as in a previous exercise, and
**22 2
VYX2 =x2=VX1x2=EX12 x2*=x2*EX12 =x2*2
It is easy to see that VX1x2*becomes bigger than VYwhenever the modulus of x2 *i
s bigger than . Further, from the relation
VYX2 =x2*=x2*22
one gets
EVYX2=4 =VY
Given that
EVYX2+VEYX2=VY
it must be that
VEYX2=0
i.e. the first-order sensitivity index is null for both X1 and X2. These results
are illustrated in the two figures which follow.
Figure 1.13 shows a plot of VX~2 YX2 =x2*, i.e. VX1 YX2 =x2*at different values
of x2 *for =1. The horizontal line is the unconditional variance of Y. The ordin
ate is zero for x2 *=0, and becomes higher than VYfor x2 *~1.
Figure 1.14 shows a scatterplot of Yversus x1 *(the same shape would appear for
x2*). It is clear from the plot that whatever the value of

Selection (1) Texto

Cargado por

Información del documento

Derechos de autor

Formatos disponibles

Compartir este documento

Compartir o incrustar documentos

Opciones para compartir

¿Le pareció útil este documento?

¿Este contenido es inapropiado?

Copyright:

Formatos disponibles

Selection (1) Texto

Cargado por

Copyright:

Formatos disponibles

METHODS AND SETTINGS FOR SENSITIVITY ANALYSIS

sensitivity analysis should be concentrated.

También podría gustarte