Tapering (mathematics): Difference between revisions

From formulasearchengine
Jump to navigation Jump to search
en>Troglodyto
Added a few more examples in matrix notation.. i'll try to put up a diagram or two of what a linear and nonlinear taper might look like
 
en>LilHelpa
m disambiguate "scaling" with Scaling (geometry)
Line 1: Line 1:
<br><br>Everybody has a will need for a multi tool Guys, girls and even your children will obtain themselves in a situation where they will say, "I wish I had a multi tool suitable now". Every single automobile, motorcycle, ATV, snowmobile and boat should really have one particular on board. You will often discover a use for 1 of these, specifically in an [http://www.aixpan.com/?q=node/434270 emergency].<br><br>I carry my multitool on my belt and really do not like when the device is shifting inside the case or pouch. Anyhow, the setup on the Milwaukee M18 Multi-Tool is sweet. Just flip the top rated lever 180°, twist out the retaining bolt by hand, and swap out your blade. Now twist the bolt back in by hand, flip the lever down, and fire that undesirable boy up. It's swift and effortless, the blade stays securely in place, and it does not matter Where your Allen wrench is. I LIKE this tool-free of charge blade changing! A die cast pinching tool for holding or manipulating materials. The pliers are pretty strong and perform pretty well for their size. Nonetheless, they are not spring-loaded so when you go to use them to clamp on one thing, you have to manually open them back up with your fingers.<br><br>A further [http://Www.flashbored.com/profile/781601/taoyxj.html variety] of tool could be emerging also. This is the hinged-lid tool. If you beloved this posting and you would like to obtain extra info relating to [http://www.Thebestpocketknifereviews.com/best-fillet-knife-electric-fish-knives-reviews/ fish fillet Knife Set] kindly visit our internet site. These tools may possibly be exciting for these who do not will need pliers. Kutmaster has such a tool, and Gerber just came out with the Multi-Lite. There are several tools including scissors, and the lid has an LED light in it. Leatherman - - note the Leatherman tools now have a tool adaptor that can be utilised to drive common 1/four" bits (hex, torx, etc.). This can be used with the PST, PST II, and Super Tool. A precise die-cast pinching tool for holding or manipulating materials, that automatically spring back to the open position when not being employed. Spring-loaded tools are beneficial for reducing hand fatigue that can happen when manually adjusting the grip on the tool with one particular hand. Serrated Knife<br><br>The Genesis GMT15A Oscillating Tool is a paltry $35 at Amazon. It essentially has a quite sharp look about it as effectively. Who knows what you happen to be going to get at the price, but it might be worth skipping a night out with the fellas in order to locate out. It helps to take a closer look at the particular tools that will make this a a single-quit item. 1 of the most crucial characteristics is having robust, reputable locks on the knife blades. A knife blade that can unexpectedly close when working beneath pressure is dangerous. The very best portion was the cost which is nevertheless quite fantastic for a compressed air tool with wrenches, but what is the point if I have to obtain a further wrench set to do the job? So, what are you waiting for? Your mission to uncover the ideal multitool is over!<br><br>One particular of the most significant characteristics is the castle nut wrench. The castle nut is what secures your stock tube to your lower, some thing you have to deal with when you're building up a reduced, or want to alter various buttstock elements (for instance, installing a stock that has an integrated tube). The castle nut wrench on the DPMS Multi-Tool stands out mainly due to the fact of its strength. If you are trying to remove a tube on a factory rifle and it really is most likely staked, you will want a nicely-made tool do the job. The DPMS will manage it with ease.<br><br>The Milwaukee® M18 Cordless Multi-Tool Kit cuts up to 50% more rapidly and delivers up to 2X extra cuts per charge than the competitors. With finest-in-class speed and efficiency, the cordless Multi-Tool presents the versatility to total problematic jobsite applications. Excellent for remodelers, flooring contractors, maintenance repair technicians and electricians, the cordless Multi-Tool cuts, grinds, sands and scrapes at odd angles and in locations difficult to function. It really is powered by the M18 REDLITHIUM Battery to deliver up to 40% more runtime. Use it to make flush<br><br>The Crunch multitool is a tiny diverse from the rest. It really is constructed around locking pliers (vise-grips) which are massive enough to grab a 1-inch pipe. The Crunch has been about for more than ten years now and still going sturdy due to its one of a kind feature set. It is roughly about the very same size as the common models like the Wave – with a closed length of 4 inches. 1 of the most utilised tools is the blade and as with all knife blades it really is all about the steel. You can study my entire guide to knife blade steels but in general 420HC is very good, 154CM is much better and S30V is the bee's knees. The superior the steel, the longer it will hold its sharp edge. Style and ergonomics Produced of 100% stainless steel, the Wave packs a lot of tools into a smaller All Locking Blades and Tools<br><br>There was one particular time I didn't tighten the Tulio multitool all the way when I fixed a flat and it came loose flopping all over the place but it by no means dropped nor did any pieces wiggle out of it even with it banging about for a handful of moments. What a helluva tool! Why don't additional persons believe of clever items like these!? This knife is pure great. When I will under no circumstances put it to it really is intended use, it is nonetheless good to have a fantastic sharp knife with you at times. Plus, it appears truly poor-ass. Conclusion Yes  No  Report from got turkey wrote 13 weeks 1 day ago this is a great multitool my dad and i each have one light and come with a handy cace A knife crafted from 420HC, a high-carbon stainless steel that is corrosion resistant and can be quickly maintained.
{{tone|date=December 2009}}
'''Sliced inverse regression (SIR)''' is a tool for [[dimension reduction]] in the field of [[multivariate statistics]].
 
In [[statistics]], [[regression analysis]] is a popular way of studying the relationship between a response variable ''y'' and its explanatory variable <math>\underline{x}</math>, which is a ''p''-dimensional vector. There are several approaches which come under the term of regression. For example parametric methods include multiple linear regression; non-parametric techniques include [[local smoothing]].
 
With high-dimensional data (as ''p'' grows), the number of observations needed to use local smoothing methods escalates exponentially.  Reducing the number of dimensions makes the operation computable. [[Dimension reduction]] aims to show only the most important directions of the data. SIR uses the inverse regression curve, <math>E(\underline{x}\,|\,y)</math> to perform a weighted principal component analysis, with which one identifies the effective dimension reducing directions.
 
This article first introduces the reader to the subject of dimension reduction and how it is performed using the model here. There is then a short review on inverse regression, which later brings these pieces together.
 
 
==Model==
 
Given a response variable <math>\,Y</math> and a (random) vector <math>X \in \R^p</math> of explanatory variables, '''SIR''' is based on the model
 
<math>Y=f(\beta_1^\top X,\ldots,\beta_k^\top X,\varepsilon)\quad\quad\quad\quad\quad(1)</math>
 
where <math>\beta_1,\ldots,\beta_k</math> are unknown projection vectors. <math>\,k</math> is an unknown number (the dimensionality  of the space we try to reduce our data to) and, of course, as we want to reduce dimension, smaller than <math>\,p</math>. <math>\;f</math> is an unknown function on <math>\R^{k+1}</math>, as it only depends on <math>\,k</math> arguments, and <math>\varepsilon</math> is the error with <math>E[\varepsilon|X]=0</math> and finite variance <math> \sigma^2 </math>. The model describes an ideal solution, where <math>\,Y</math> depends on <math>X \in \R^p</math> only through a <math>\,k</math> dimensional subspace. I.e. one can reduce to dimension of the explanatory variable from <math>\,p</math> to a smaller number <math>\,k</math> without losing any information.
 
An equivalent version of <math>\,(1)</math> is: the conditional distribution of <math> \,Y </math> given <math>\, X </math> depends on <math>\, X </math> only through the <math>\,k</math> dimensional random vector <math>(\beta_1^\top X,\ldots,\beta_k^\top X)</math>. This perfectly reduced vector can be seen as informative as the original <math> \,X </math> in explaining <math>\, Y </math>.
 
The unknown <math>\,\beta_i's</math> are called the ''effective dimension reducing directions'' (EDR-directions). The space that is spanned by these vectors is denoted the ''effective dimension reducing space'' (EDR-space).
 
==Relevant linear algebra background==
 
 
To be able to visualize the model, note a short review on vector spaces:
 
For the definition of a vector space and some further properties I will refer to the article [http://teachwiki.wiwi.hu-berlin.de/index.php/Basic_Linear_Algebra_and_Gram-Schmidt_Orthogonalization|Basic Linear Algebra and Gram-Schmidt Orthogonalization] or any textbook in linear algebra and mention only the most important facts for understanding the model.
 
As the EDR-space is a <math>\,k</math>dimensional subspace, we need to know what a subspace is. A subspace of <math>\R^n</math> is defined as a subset <math>U \in \R^n</math>, if it holds that
 
: <math>\underline{a},\underline{b} \in U \Rightarrow \underline{a}+\underline{b} \in U</math>
 
: <math>\underline{a} \in U, \lambda \in \R \Rightarrow \lambda \underline{a} \in U</math>
 
Given <math>\underline{a}_1,\ldots,\underline{a}_r \in \R^n</math>, then <math>V:=L(\underline{a}_1,\ldots,\underline{a}_r)</math>, the set of all linear combinations of these vectors, is called a linear subspace and is therefore a vector space.  One says, the vectors <math>\underline{a}_1,\ldots,\underline{a}_r</math> span <math>\,V</math>. But the vectors that span a space <math>\,V</math> are not unique. This leads us to the concept of a basis and the dimension of a vector space:
 
A set <math>B=\{\underline{b}_1,\ldots,\underline{b}_r\}</math> of linear independent vectors of a vector space <math>\,V</math> is called ''basis'' of <math>\,V</math>, if it holds that
 
: <math>V:=L(\underline{b}_1,\ldots,\underline{b}_r)</math>
 
The dimension of <math>\,V (\in \R^n)</math> is equal to the maximum number of linearly independent vectors in <math>\,V</math>. A set of <math>\,n</math> linear independent vectors of <math>\R^n</math> set up a basis of <math>\R^n</math>. The dimension of a vector space is unique, as the basis itself is not. Several bases can span the same space.
Of course also dependent vectors span a space, but the linear combinations of the latter can give only rise to the set of vectors lying on a straight line. As we are searching for a <math>\,k</math>dimensional subspace, we are interested in finding <math>\,k</math> linearly independent vectors that span the <math>\,k</math>dimensional subspace we want to project our data on.
 
==Curse of dimensionality==
 
The reason why we want to reduce the dimension of the data is due to the "[[curse of dimensionality]]" and of course, for graphical purposes. The curse of dimensionality is due to rapid increase in volume adding more dimensions to a (mathematical) space. For example, consider 100 observations from support <math>[0,1]</math>, which cover the interval quite well, and compare it to 100 observations from the corresponding <math>10</math> dimensional unit hypersquare, which are isolated points in a vast empty space. It is easy to draw inferences about the underlying properties of the data in the first case, whereas in the latter, it is not. For more information about the curse of dimensionality, see [[Curse of dimensionality]].
 
==Inverse regression==
 
Computing the inverse regression curve (IR) means instead of looking for
*<math>\,E[Y|X=x]</math>, which is a curve in <math>\R^p</math>
 
we calculate
 
*<math>\,E[X|Y=y]</math>, which is also a curve in <math>\R^p</math>, but consisting of <math>\,p</math> one dimensional regressions.
 
The center of the inverse regression curve is located at <math>\,E[E[X|Y]]=E[X]</math>. Therefore, the centered inverse regression curve is
 
*<math>\,E[X|Y=y]-E[X]</math>
 
which is a <math>\,p</math> dimensional curve in <math>\R^p</math>. In what follows we will consider this centered inverse regression curve and we will see that it lies on a <math>\,k</math>dimensional subspace spanned by <math>\,\Sigma_{xx}\beta_i\,'s</math>.
 
But before seeing that this holds true, we will have a look at how the inverse regression curve is computed within the SIR-Algorithm, which will be introduced in detail later. What comes is the "sliced" part of SIR. We estimate the inverse regression curve by dividing the range of <math>\,Y</math> into <math>\,H</math> nonoverlapping intervals (slices), to afterwards compute the sample means <math>\,\hat{m}_h</math> of each slice. '''These sample means are used as a crude estimate of the IR-curve''', denoted as <math>\,m(y)</math>. There are several ways to define the slices, either in a way that in each slice are equally much observations, or we define a fixed range for each slice, so that we then get different proportions of the <math>\,y_i\,'s</math> that fall into each slice.
 
==Inverse regression versus dimension reduction==
 
As mentioned a second before, the centered inverse regression curve lies on a <math>\,k</math>dimensional subspace spanned by <math>\,\Sigma_{xx}\beta_i\,'s</math> (and therefore also the crude estimate we compute). This is the connection between our Model and Inverse Regression. We shall see that this is true, with only one condition on the design distribution that must hold. This condition is, that:
 
: <math>\forall\,\underline{b} \in\R^p:\,E[b^\top X|\beta_1^\top X=\beta_1^\top x,\ldots,\beta_k^\top X=\beta_k^\top x)
==c_0+\sum_{i==1}^{k} c_i\beta_i^\top x</math>
 
I.e. the conditional expectation is linear in <math>\beta_1 X,\ldots,\beta_k X</math>, that is, for some constants <math>c_0,\ldots,c_K</math>. This condition is satisfied when the distribution of <math>\,X</math> is elliptically symmetric (e.g. the normal distribution). This seems to be a pretty strong requirement. It could help, for example, to closer examine the distribution of the data, so that outliers can be removed or clusters can be separated before analysis
 
Given this condition and <math>\,(1)</math>, it is indeed true that the centered inverse regression curve <math>\,E[X|Y=y]-E[X]</math> is contained in the linear subspace spanned by <math>\,\Sigma_{xx}\beta_k(k=1,\ldots,K)</math>, where <math>\,\Sigma_{xx}=Cov(X)</math>. The proof is provided by Duan and Li in ''Journal of the American Statistical Association'' (1991).
 
==Estimation of the EDR-directions==
 
After having had a look at all the theoretical properties, our aim is now to estimate the EDR-directions. For that purpose, we conduct a (weighted) principal component analysis for the sample means <math>\,\hat{m}_h\,'s</math>, after having standardized <math>\,X</math> to <math>\,Z=\Sigma_{xx}^{-1/2}\{X-E(X)\}</math>. Corresponding to the theorem above, the IR-curve <math>\,m_1(y)=E[Z|Y=y]</math> lies in the space spanned by <math>\,(\eta_1,\ldots,\eta_k)</math>, where <math>\,\eta_i=\Sigma^{1/2}_{xx} \beta_i</math>. (Due to the terminology introduced before, the <math>\,\eta_i\,'s</math> are called the ''standardized effective dimension reducing directions''.) As a consequence, the covariance matrix <math>\,cov[E[Z|Y]]</math> is degenerate  in any direction orthogonal to the <math>\,\eta_i\,'s</math>. Therefore, the eigenvectors <math>\,\eta_k (k=1,\ldots,K)</math> associated with the <math>\,K</math> largest eigenvalues are the standardized EDR-directions.
 
Back to PCA. That is, we calculate the estimate for <math>\,Cov\{m_1(y)\}</math>:
 
: <math>\hat{V}=n^{-1}\sum_{i=1}^S n_s \bar{z}_s \bar{z}_s^\top</math>
 
and identify the eigenvalues <math>\hat{\lambda}_i</math> and the eigenvectors <math>\hat{\eta}_i</math> of <math>\hat{V}</math>, which are the standardized EDR-directions. (For more details about that see next section: Algorithm.) Remember that the main idea of PC transformation is to find the most informative projections that maximize variance!
 
Note that in some situations SIR does not find the EDR-directions. One can overcome this difficulty by considering the conditional covariance <math>\,Cov(X|Y)</math>. The principle remains the same as before, but one investigates the IR-curve with the conditional covariance instead of the conditional expectation. For further details and an example where SIR fails, see Härdle and Simar (2003).
 
==Algorithm==
 
The algorithm to estimate the EDR-directions via SIR is as follows. It is taken from the textbook ''Applied Multivariate Statistical Analysis'' (Härdle and Simar 2003)
 
'''1.''' Let <math>\,\Sigma_{xx}</math> be the covariance matrix of <math>\,X</math>. Standardize <math>\,X</math> to
 
: <math>\,Z=\Sigma_{xx}^{-1/2}\{X-E(X)\}</math>
 
(We can therefore rewrite <math>\,(1)</math> as
 
: <math>Y=f(\eta_1^\top Z,\ldots,\eta_k^\top Z,\varepsilon)</math>
 
where <math>\,\eta_k=\beta_k\Sigma_{xx}^{1/2}\quad\forall\; k</math>
For the standardized variable Z it holds that <math>\,E[Z]=0</math> and  <math>\,Cov(Z)=I</math>.)
 
'''2.''' Divide the range of <math>\,y_i</math> into <math>\,S</math> nonoverlapping slices <math>\,H_s(s=1,\ldots,S).\; n_s</math> is the number of observations within each slice and <math>\,I_{H_s}</math> the indicator function for this slice:
: <math>n_s=\sum_{i=1}^n I_{H_s}(y_i)</math>
 
'''3.''' Compute the mean of <math>\,z_i</math> over all slices, which is a crude estimate <math>\,\hat{m}_1</math> of the inverse regression curve <math>\,m_1</math>:
 
: <math>\,\bar{z}_s=n_s^{-1}\sum_{i=1}^n z_i I_{H_s}(y_i)</math>
 
'''4.''' Calculate the estimate for <math>\,Cov\{m_1(y)\}</math>:
 
: <math>\,\hat{V}=n^{-1}\sum_{i=1}^S n_s \bar{z}_s \bar{z}_s^\top</math>
 
'''5.''' Identify the eigenvalues <math>\,\hat{\lambda}_i</math> and the eigenvectors <math>\,\hat{\eta}_i</math> of <math>\,\hat{V}</math>, which are the standardized EDR-directions.
 
'''6.''' Transform the standardized EDR-directions back to the original scale. The estimates for the EDR-directions are given by:
 
: <math>\,\hat{\beta}_i=\hat{\Sigma}_{xx}^{-1/2}\hat{\eta}_i</math>
 
(which are not necessarily orthogonal)
 
For examples, see the book by Härdle and Simar (2003).
 
==See also==
*[[Curse of dimensionality]]
==References==
 
*Li, K-C. (1991) "Sliced Inverse Regression for Dimension Reduction", [[Journal of the American Statistical Association]], 86, 316&ndash;327 [http://www.jstor.org/stable/2290563 Jstor]
 
*Cook, R.D. and Sanford Weisberg, S. (1991) "Sliced Inverse Regression for Dimension Reduction: Comment", [[Journal of the American Statistical Association]], 86, 328&ndash;332 [http://www.jstor.org/stable/2290564 Jstor]
 
*Härdle, W. and Simar, L. (2003) ''Applied Multivariate Statistical Analysis'', Springer Verlag. ISBN 3-540-03079-4
 
*Kurzfassung zur Vorlesung Mathematik II im Sommersemester 2005, A. Brandt
 
==External links==
*[http://www.hmwu.idv.tw/hmwu/index.php/publication/80 Sliced Inverse Regression: Literature Collection]
 
[[Category:Regression analysis]]
[[Category:Dimension reduction]]

Revision as of 20:33, 22 January 2013

I am Keisha from Menzingen. I love to play Lute. Other hobbies are Insect collecting.

Stop by my weblog Hostgator Discount Coupon Sliced inverse regression (SIR) is a tool for dimension reduction in the field of multivariate statistics.

In statistics, regression analysis is a popular way of studying the relationship between a response variable y and its explanatory variable x_, which is a p-dimensional vector. There are several approaches which come under the term of regression. For example parametric methods include multiple linear regression; non-parametric techniques include local smoothing.

With high-dimensional data (as p grows), the number of observations needed to use local smoothing methods escalates exponentially. Reducing the number of dimensions makes the operation computable. Dimension reduction aims to show only the most important directions of the data. SIR uses the inverse regression curve, E(x_|y) to perform a weighted principal component analysis, with which one identifies the effective dimension reducing directions.

This article first introduces the reader to the subject of dimension reduction and how it is performed using the model here. There is then a short review on inverse regression, which later brings these pieces together.


Model

Given a response variable Y and a (random) vector Xp of explanatory variables, SIR is based on the model

Y=f(β1X,,βkX,ε)(1)

where β1,,βk are unknown projection vectors. k is an unknown number (the dimensionality of the space we try to reduce our data to) and, of course, as we want to reduce dimension, smaller than p. f is an unknown function on k+1, as it only depends on k arguments, and ε is the error with E[ε|X]=0 and finite variance σ2. The model describes an ideal solution, where Y depends on Xp only through a k dimensional subspace. I.e. one can reduce to dimension of the explanatory variable from p to a smaller number k without losing any information.

An equivalent version of (1) is: the conditional distribution of Y given X depends on X only through the k dimensional random vector (β1X,,βkX). This perfectly reduced vector can be seen as informative as the original X in explaining Y.

The unknown βis are called the effective dimension reducing directions (EDR-directions). The space that is spanned by these vectors is denoted the effective dimension reducing space (EDR-space).

Relevant linear algebra background

To be able to visualize the model, note a short review on vector spaces:

For the definition of a vector space and some further properties I will refer to the article Linear Algebra and Gram-Schmidt Orthogonalization or any textbook in linear algebra and mention only the most important facts for understanding the model.

As the EDR-space is a kdimensional subspace, we need to know what a subspace is. A subspace of n is defined as a subset Un, if it holds that

a_,b_Ua_+b_U
a_U,λλa_U

Given a_1,,a_rn, then V:=L(a_1,,a_r), the set of all linear combinations of these vectors, is called a linear subspace and is therefore a vector space. One says, the vectors a_1,,a_r span V. But the vectors that span a space V are not unique. This leads us to the concept of a basis and the dimension of a vector space:

A set B={b_1,,b_r} of linear independent vectors of a vector space V is called basis of V, if it holds that

V:=L(b_1,,b_r)

The dimension of V(n) is equal to the maximum number of linearly independent vectors in V. A set of n linear independent vectors of n set up a basis of n. The dimension of a vector space is unique, as the basis itself is not. Several bases can span the same space. Of course also dependent vectors span a space, but the linear combinations of the latter can give only rise to the set of vectors lying on a straight line. As we are searching for a kdimensional subspace, we are interested in finding k linearly independent vectors that span the kdimensional subspace we want to project our data on.

Curse of dimensionality

The reason why we want to reduce the dimension of the data is due to the "curse of dimensionality" and of course, for graphical purposes. The curse of dimensionality is due to rapid increase in volume adding more dimensions to a (mathematical) space. For example, consider 100 observations from support [0,1], which cover the interval quite well, and compare it to 100 observations from the corresponding 10 dimensional unit hypersquare, which are isolated points in a vast empty space. It is easy to draw inferences about the underlying properties of the data in the first case, whereas in the latter, it is not. For more information about the curse of dimensionality, see Curse of dimensionality.

Inverse regression

Computing the inverse regression curve (IR) means instead of looking for

we calculate

  • E[X|Y=y], which is also a curve in p, but consisting of p one dimensional regressions.

The center of the inverse regression curve is located at E[E[X|Y]]=E[X]. Therefore, the centered inverse regression curve is

which is a p dimensional curve in p. In what follows we will consider this centered inverse regression curve and we will see that it lies on a kdimensional subspace spanned by Σxxβis.

But before seeing that this holds true, we will have a look at how the inverse regression curve is computed within the SIR-Algorithm, which will be introduced in detail later. What comes is the "sliced" part of SIR. We estimate the inverse regression curve by dividing the range of Y into H nonoverlapping intervals (slices), to afterwards compute the sample means m̂h of each slice. These sample means are used as a crude estimate of the IR-curve, denoted as m(y). There are several ways to define the slices, either in a way that in each slice are equally much observations, or we define a fixed range for each slice, so that we then get different proportions of the yis that fall into each slice.

Inverse regression versus dimension reduction

As mentioned a second before, the centered inverse regression curve lies on a kdimensional subspace spanned by Σxxβis (and therefore also the crude estimate we compute). This is the connection between our Model and Inverse Regression. We shall see that this is true, with only one condition on the design distribution that must hold. This condition is, that:

b_p:E[bX|β1X=β1x,,βkX=βkx)==c0+i==1kciβix

I.e. the conditional expectation is linear in β1X,,βkX, that is, for some constants c0,,cK. This condition is satisfied when the distribution of X is elliptically symmetric (e.g. the normal distribution). This seems to be a pretty strong requirement. It could help, for example, to closer examine the distribution of the data, so that outliers can be removed or clusters can be separated before analysis

Given this condition and (1), it is indeed true that the centered inverse regression curve E[X|Y=y]E[X] is contained in the linear subspace spanned by Σxxβk(k=1,,K), where Σxx=Cov(X). The proof is provided by Duan and Li in Journal of the American Statistical Association (1991).

Estimation of the EDR-directions

After having had a look at all the theoretical properties, our aim is now to estimate the EDR-directions. For that purpose, we conduct a (weighted) principal component analysis for the sample means m̂hs, after having standardized X to Z=Σxx1/2{XE(X)}. Corresponding to the theorem above, the IR-curve m1(y)=E[Z|Y=y] lies in the space spanned by (η1,,ηk), where ηi=Σxx1/2βi. (Due to the terminology introduced before, the ηis are called the standardized effective dimension reducing directions.) As a consequence, the covariance matrix cov[E[Z|Y]] is degenerate in any direction orthogonal to the ηis. Therefore, the eigenvectors ηk(k=1,,K) associated with the K largest eigenvalues are the standardized EDR-directions.

Back to PCA. That is, we calculate the estimate for Cov{m1(y)}:

V̂=n1i=1Snsz¯sz¯s

and identify the eigenvalues λ̂i and the eigenvectors η̂i of V̂, which are the standardized EDR-directions. (For more details about that see next section: Algorithm.) Remember that the main idea of PC transformation is to find the most informative projections that maximize variance!

Note that in some situations SIR does not find the EDR-directions. One can overcome this difficulty by considering the conditional covariance Cov(X|Y). The principle remains the same as before, but one investigates the IR-curve with the conditional covariance instead of the conditional expectation. For further details and an example where SIR fails, see Härdle and Simar (2003).

Algorithm

The algorithm to estimate the EDR-directions via SIR is as follows. It is taken from the textbook Applied Multivariate Statistical Analysis (Härdle and Simar 2003)

1. Let Σxx be the covariance matrix of X. Standardize X to

Z=Σxx1/2{XE(X)}

(We can therefore rewrite (1) as

Y=f(η1Z,,ηkZ,ε)

where ηk=βkΣxx1/2k For the standardized variable Z it holds that E[Z]=0 and Cov(Z)=I.)

2. Divide the range of yi into S nonoverlapping slices Hs(s=1,,S).ns is the number of observations within each slice and IHs the indicator function for this slice:

ns=i=1nIHs(yi)

3. Compute the mean of zi over all slices, which is a crude estimate m̂1 of the inverse regression curve m1:

z¯s=ns1i=1nziIHs(yi)

4. Calculate the estimate for Cov{m1(y)}:

V̂=n1i=1Snsz¯sz¯s

5. Identify the eigenvalues λ̂i and the eigenvectors η̂i of V̂, which are the standardized EDR-directions.

6. Transform the standardized EDR-directions back to the original scale. The estimates for the EDR-directions are given by:

β̂i=Σ̂xx1/2η̂i

(which are not necessarily orthogonal)

For examples, see the book by Härdle and Simar (2003).

See also

References

  • Härdle, W. and Simar, L. (2003) Applied Multivariate Statistical Analysis, Springer Verlag. ISBN 3-540-03079-4
  • Kurzfassung zur Vorlesung Mathematik II im Sommersemester 2005, A. Brandt