Website Worth

Total Pageviews

Showing posts with label Biostatistics. Show all posts
Showing posts with label Biostatistics. Show all posts

Thursday

Reliability after Factor Analysis: Composite Reliability Calculator

Measurement is a serious topic in social science. Validity and reliability should always be the central interest of social scientists who are in charge of developing data collection tools and analysis plans.

This tool allows you to calculate composite reliability based on standardized factor loadings and error variance

http://www.thestatisticalmind.com/calculators/comprel/composite_reliability.htm

Please send thanks to the author.

Tuesday

Composite Reliability Calculator

Measurement is a serious topic in social science. Validity and reliability should always be the central interest of social scientists who are in charge of developing data collection tools and analysis plans.

This tool allows you to calculate composite reliability based on standardized factor loadings and error variance

http://www.thestatisticalmind.com/calculators/comprel/composite_reliability.htm

Please send thanks to the author.

Wednesday

Survey Errors and Survey Costs

Survey Errors and Survey Costs

Copyright © 1989 John Wiley & Sons, Inc. All rights reserved.

Survey Errors and Survey Costs

Author(s): Robert M. Groves

Published Online: 28 JAN 2005

Print ISBN: 9780471611714

Online ISBN: 9780471725275

DOI: 10.1002/0471725277

Book Series: Wiley Series in Probability and Statistics

About this Book

The Wiley-Interscience Paperback Series consists of selected books that have been made more accessible to consumers in an effort to increase global appeal and general circulation. With these new unabridged softcover volumes, Wiley hopes to extend the lives of these works by making them available to future generations of statisticians, mathematicians, and scientists.

"Survey Errors and Survey Costs is a well-written, well-presented, and highly readable text that should be on every error-conscious statistician's bookshelf. Any courses that cover the theory and design of surveys should certainly have Survey Errors and Survey Costs on their reading lists."
-Phil Edwards
MEL, Aston University Science Park, UK
Review in The Statistician, Vol. 40, No. 3, 1991

"This volume is an extremely valuable contribution to survey methodology. It has many virtues: First, it provides a framework in which survey errors can be segregated by sources. Second, Groves has skillfully synthesized existing knowledge, bringing together in an easily accessible form empirical knowledge from a variety of sources. Third, he has managed to integrate into a common framework the contributions of several disciplines. For example, the work of psychometricians and cognitive psychologists is made relevant to the research of econometricians as well as the field experience of sociologists. Finally, but not least, Groves has managed to present all this in a style that is accessible to a wide variety of readers ranging from survey specialists to policymakers."
-Peter H. Rossi
University of Massachusetts at Amherst
Review in Journal of Official Statistics, January 1991

More about this book summary

Table of contents

Select All

    1. You have free access to this content

      Frontmatter (pages i–xxiii)

    2. You have full text access to this content

      An Introduction to Survey Errors (pages 1–47)

    3. You have full text access to this content

      An Introduction to Survey Costs (pages 49–80)

    4. You have full text access to this content

      Costs and Errors of Covering the Population (pages 81–132)

    5. You have full text access to this content

      Nonresponse in Sample Surveys (pages 133–183)

    6. You have full text access to this content

      Probing the Causes of Nonresponse and Efforts to Reduce Nonresponse (pages 185–238)

    7. You have full text access to this content

      Costs and Errors Arising from Sampling (pages 239–294)

    8. You have full text access to this content

      Empirical Estimation of Survey Measurement Error (pages 295–355)

    9. You have full text access to this content

      The Interviewer as a Source of Survey Measurement Error (pages 357–406)

    10. You have full text access to this content

      The Respondent as a Source of Survey Measurement Error (pages 407–448)

    11. You have full text access to this content

      Measurement Errors Associated with the Questionnaire (pages 449–499)

    12. You have full text access to this content

      Response Effects of the Mode of Data Collection (pages 501–552)

    13. You have free access to this content

      References (pages 553–579)

    14. You have free access to this content

      Index (pages 581–590)

Monday

10 data analysis techniques that social researchers must know

1. Exploratory factor analysis

2. Confirmatory factor analysis

3. Adapting a scale

4. Thematic analysis

5. Logistic regression

6. Case studies

7. Longitudinal analysis

8. Multi-level analysis

9. Ordinal regression

10. Social network analysis

11. Discourse analysis

12. Grounded theory

Wednesday

Sample size calculation: comparing two proportions

We can use the following formula for the sample size n:

n = (Zα/2+Zβ)2 * (p1(1-p1)+p2(1-p2)) / (p1-p2)2,

where Zα/2 is the critical value of the Normal distribution at α/2 (e.g. for a confidence level of 95%, α is 0.05 and the critical value is 1.96), Zβ is the critical value of the Normal distribution at β (e.g. for a power of 80%, β is 0.2 and the critical value is 0.84) and p1 and p2 are the expected sample proportions of the two groups.

Friday

When should we use ‘non-parametric’ techniques?

We only use parametric techniques (like t test, z test) when we are certain about the distribution of the variable of interest.

When we don’t know its distribution, it is safer to use non-parametric tests. These tests have no assumptions about distribution of the dependent variable.

In fact, it has been argued, quite sharply, that in all social sciences, we should use non-parametric, rather than parametric, tests.

Principles of good measurement

 

1 – Mutually exclusive attributes: every response is clearly different from others

2 – Exhaustive attribute: every response has a place to go

3 – Uni-dimensionality : only one construct is measured

Sunday

Marginal effect

obin2-mod

SEM Measurement Model with AMOS 18 (example 2)

1) Unconstrained Model
image
2) Constrained factor loading
image
3) Constrained structural covariance
image
4) Constrained measurement residuals
image
5) Model comparison

Assuming model Unconstrained to be correct:
Model DF CMIN P NFI
Delta-1
IFI
Delta-2
RFI
rho-1
TLI
rho2
Measurement weights 4 2.632 .621 .011 .012 -.010 -.012
Structural covariances 7 3.306 .855 .014 .015 -.023 -.026
Measurement residuals 13 11.212 .593 .048 .051 -.011 -.012
Assuming model Measurement weights to be correct:
Model DF CMIN P NFI
Delta-1
IFI
Delta-2
RFI
rho-1
TLI
rho2
Structural covariances 3 .674 .879 .003 .003 -.012 -.014
Measurement residuals 9 8.581 .477 .037 .040 -.001 -.001
Assuming model Structural covariances to be correct:
Model DF CMIN P NFI
Delta-1
IFI
Delta-2
RFI
rho-1
TLI
rho2
Measurement residuals 6 7.906 .245 .034 .037 .012 .013

Friday

A measurement model, cross-group validation, SEM, AMOS 18

fourmodels

First, what is 'constraint': A model of this type relates to in-variance. We want to test if factor loadings, variances of factors, covariances between factors, and variances of the 'error' terms remain the same ACROSS groups - or, if they change, the change is INSIGNIFICANT. We do that by artificially 'keep' constant one of the elements listed above (in the least restricted model, keep constant factor loadings only) or all of them (in the fully restricted model). The computer will do it for you. In AMOS, you should choose 'unstandardized estimates' if you want to see which element is kept constant.

Then we compare estimates of the less restricted model with estimates of the more restricted model, and hope that the change would be insignificant. Again, the computer will do it.

A good cross-group measurement model should have four components like above: Top left: No constraints (two models are run at the same time, chi-square difference is computed at the end). Top right: Constrained factor loadings (the least restricted model - this is minimum requirement). Bottom Left: Constrained structural covariances (the more restricted model). Bottom Right: Constrained measurement residuals (the fully restricted model).

Perfect: All models have good fit
Very good: Except the last model, other three have good fit
Good: The model constraining for factor loadings has good fit.
Poor: The model constraining for factor loadings has poor fit. In this case, we conclude that the groups that are compared do not share the same factors. Two separate measurements (one for men, another for women, for instance) should be used.

The above model has perfect fit. How can I say that?
  • All 4 models have RMSEA smaller than 0.06, GFI close to 1, p value is smaller 5%
  • Compared with the less constrained model, the more constrained model has a good fit and the 'gap' between 'less constrained' and 'more constrained' is insignificant. In AMOS, look for this at 'Model Comparison' tab.
Assuming model Unconstrained to be correct:
Model DF CMIN P NFI
Delta-1
IFI
Delta-2
RFI
rho-1
TLI
rho2
Measurement weights 4 2.632 .621 .011 .012 -.010 -.012
Structural covariances 7 3.306 .855 .014 .015 -.023 -.026
Measurement residuals 13 11.212 .593 .048 .051 -.011 -.012

  • The model constraining for factor loadings is  insignificantly different from the unconstrained model.

Thursday

A normal distribution: an univariate approach

 

If a variable is normally distributed, most of its values centre around the mean, within two standard deviations (up to the blue areas, at both sides). If value is more than two SD away from the mean, it is potentially an outlier. Remove it. If there are two many values like that, we should not consider the variable normally distributed.

In order to judge if a variable is ‘normally distributed’

  • calculate the mean and SD : SD = square root[sum((xi – mean)2)/mean]
  • convert each value into z score: (xi – mean)/SD
  • compare each z score with 1.96*SD
  • if the result is greater than 1.96*SD, that value is an outlier.

image

Figure Caption in SEM

This is the text to use:

chi squared = \cmin
df = \df
p = \p
RMR = \rmr
GFI = \gfi
RMSEA = \rmsea
image
image
This is the complete list:
\cmin is a "text macro", a code that Amos fills in with the minimum value of the discrepancy function, C (see Appendix B), once the minimum value is known. Similarly, \df is a text macro that Amos fills in with the number of degrees of freedom for testing the model and \p is a text macro that Amos fills in with the "p value" for testing the null hypothesis that the model is correct. Here is a list of text macros.
\agfi Adjusted goodness of fit index (AGFI)
\aic Akaike information criterion (AIC)
\bcc Browne-Cudeck criterion (BCC)
\bic Bayes information criterion (BIC)
\caic Consistent AIC (CAIC)
\cfi Comparative fit index (CFI)
\cmin Minimum value of the discrepancy function C in Appendix B
\cmindf Minimum value of the discrepancy function divided by degrees of freedom
\datafilename The name of the data file. \longdatafilename displays the fully qualified path name of the data file.
\datatablename The name of the data table (for those file formats that allow a single file to contain multiple data tables, such as Excel workbooks.)
\date Today's date in short format. \longdate displays today's date in long format. The displayed date is made current whenever the path diagram is read from a file, saved or printed.
\df Degrees of freedom
\ecvi Expected cross-validation index (ECVI)
\ecvihi Upper bound of 90% confidence interval on ECVI
\ecvilo Lower bound of 90% confidence interval on ECVI
\f0 Estimated population discrepancy (F0)
\f0hi Upper bound of 90% confidence interval on F0
\f0lo Lower bound of 90% confidence interval on F0
\filename Name of the current AMW file. Use \longfilename to display the complete path to the current AMW file.
\fmin Minimum value of discrepancy function F in Appendix B
\format Format name (See Formats tab.)
\gfi Goodness of fit index (GFI)
\group Group name (See Manage groups.)
\hfive Hoelter's critical N for =.05
\hone Hoelter's critical N for =.01
\ifi Incremental fit index (IFI)
\longdatafilenameThe fully qualified path name of the data file. \datafilename displays the data file name without the path.
\longdate Today's date in long format. \date display's today's date in short format. The displayed date is made current whenever the path diagram is read from a file, saved or printed.
\longfilename Fully qualified path name of the current AMW file. Use \filename to display the file name without the path.
\longtime The time in long format. \time displays the time in short format. The displayed time is made current whenever the path diagram is read from a file, saved or printed.
\mecvi Modified ECVI (MECVI)
\model Model name (See Manage models.)
\ncp Estimate of non-centrality parameter (NCP)
\ncphi Upper bound of 90% confidence interval on NCP
\ncplo Lower bound of 90% confidence interval on NCP
\nfi Normed fit index (NFI)
\npar Number of distinct parameters
\p "p value" associated with discrepancy function (test of perfect fit)
\pcfi Parsimonious comparative fit index (PCFI)
\pclose "p value" for testing the null hypothesis of close fit (RMSEA < .05)
\pgfi Parsimonious goodness of fit index (PGFI)
\pnfi Parsimonious normed fit index (PNFI)
\pratio Parsimony ratio
\rfi Relative fit index
\rmr Root mean square residual
\rmsea Root mean square error of approximation (RMSEA)
\rmseahi Upper bound of 90% confidence interval on RMSEA
\rmsealo Lower bound of 90% confidence interval on RMSEA
\time The time in short format. \longtime displays the time in long format. The displayed time is made current whenever the path diagram is read from a file, saved or printed.
\tli Tucker-Lewis index (TLI)

















































Understanding, checking collinearity

 

When you build a model to predict something, you would like to choose the ‘right’ predictor. There should be no redundancies among the set of predictors. Or, each predictor should explain some part of the variance of the dependence value that others don’t. In other words, predictions  should not overlap.

If you happen to use two predictors that correlate with each other significantly, they together can only explain a small amount of variance. Collinearity is a problem here. In stata a collinearity diagnosis can be obtained by typing vif  after regression


    Variable |       VIF       1/VIF 
-------------+----------------------
          d3 |      1.27    0.784681
          d2 |      1.25    0.800108
          d4 |      1.09    0.916960
-------------+----------------------
    Mean VIF |     1.20

If any individual score is larger than 10, that is the indication of collinearity. (None in this example)

If the mean of VIF is substantially larger than 1 –collinearity may exist. (Not very far from 1 in this example)

What to do?

Delete the variable with greater ViF. 

Using principle component analysis in Stata (PCA) to combine two or more variable to make an composite variable.

Monday

SEM in Stata 12 and SEM in AMOS 16: which one is better?

In Stata 12, it is much more difficult to draw a SEM diagram. You should use syntax instead.

The advantage: it allows to take into account effects of research design (survey, for example), estimation is usually better.

In AMOS 16, it is much easier to draw a SEM diagram. It is also much easier to interpret outputs because you can get everything (goodness of fit statistics) in one go. However, if your data need to be adjusted for research design, AMOS does not have any function for it.

Conclusion: if you know that research design is not a several problem, use AMOS, otherwise, use STATA 12.

Export graphs in Stata

Instead of doing it manually (copy and paste), we can do it much quicker: if you want to save a histogram graph:

1) hist d1, freq
2) graph export d1.eps, replace
3) graphout d1.eps using example.rtf
4) Open file example.rtf, you will see this:


The only problem is that it is not clear as we use Excel. 
Always use Excel if the graph is simple. If it is more complicated, use Stata's graphic functions.

Friday

Checking for autocorrelation /serial correlation


What is autocorrelation?
1) So you have an equation to predict values of a dependent variable Y….
2) The predicted values (Y’i)  are only ‘close’ the the real values (Yi)….
3) The difference between predicted and real values are called by different names: residual (the left over), or error term (the part that equation cannot predict)
4) There are two important assumptions related to the residual/error term: – they should not correlate with each other, and their variance should be constant.
5) If the first assumption is violated, or the residuals (between different waves of data, or different groups) correlate with each other, we have autocorrelation. This is particularly so in time series data analysis (called ‘serial correlation’)
6) If the second assumption is violated, we have the phenomenon called heteroskedasticity
How to check for them?
How to know if the first assumption is violated?
We can set the data as panel data, then use Wooldridge test as following in Stata
xtserial DEP (list of predictor)  ------> if p value smaller than 0.05, assumption is violated. If I have d1 as DEP and age is the only predictor,
xtserial d1 age
Wooldridge test for autocorrelation in panel data
H0: no first-order autocorrelation
    F(  1,       1) =      0.059
           Prob>| F| =      0.8486
-->This indicates there is no autocorrelation
Another way is to do this: set the data as times series, then
reg DEP (list of predictor)
dwstat
How to know if heterokesdasticity exist?
reg DEP var list
estat hett
Breusch-Pagan / Cook-Weisberg test for heteroskedasticity
         Ho: Constant variance
         Variables: fitted values of d3
         chi2(1)      =     0.73
         Prob >|chi2|  =   0.3942
-->This indicates no heterokesdascity
What do to?
What to do if you find autocorrelation? Remove its effect by:
prais d1 age, corc
Iteration 0:  rho = 0.0000
Iteration 1:  rho = 0.0795
Iteration 2:  rho = 0.0808
Iteration 3:  rho = 0.0808
Iteration 4:  rho = 0.0808
Cochrane-Orcutt AR(1) regression -- iterated estimates
      Source |       SS       df       MS              Number of obs =     127
-------------+------------------------------           F(  1,   125) =    4.85
       Model |  14.3667947     1  14.3667947           Prob >| F|      =  0.0295
    Residual |  370.582352   125  2.96465882           R-squared     =  0.0373
-------------+------------------------------           Adj R-squared =  0.0296
       Total |  384.949147   126  3.05515196           Root MSE      =  1.7218
------------------------------------------------------------------------------
          d1 |      Coef.   Std. Err.      t    P>|t|     [95% Conf. Interval]
-------------+----------------------------------------------------------------
         age |   .0306289   .0139136     2.20   0.030     .0030922    .0581656
       _cons |   3.001655   .5507311     5.45   0.000     1.911689     4.09162
-------------+----------------------------------------------------------------
         rho |   .0808052
------------------------------------------------------------------------------
Durbin-Watson statistic (original)    1.838574
Durbin-Watson statistic (transformed) 2.000578
What do do if you find heteroskedasticity?
Heteroskedasticity can distort  estimations. In stata, we remove its effect by using robust regression
reg Y X, robust
 reg d1 age, robust
Linear regression                                      Number of obs =     128
                                                       F(  1,   126) =    5.55
                                                       Prob>| F|      =  0.0201
                                                       R-squared     =  0.0409
                                                       Root MSE      =   1.722
------------------------------------------------------------------------------
             |               Robust
          d1 |      Coef.   Std. Err.      t    P>|t|     [95% Conf. Interval]
-------------+----------------------------------------------------------------
         age |   .0322989   .0137162     2.35   0.020      .005155    .0594428
       _cons |   2.945031   .5509205     5.35   0.000     1.854776    4.035286
------------------------------------------------------------------------------

Tuesday

Non-parametric tests

1) Test for normality: Kolmogorov-Smirnov one-sample test

2) Comparing two related samples: Wilcoxon Signed Ranks Test

3) Comparing two unrelated samples: Mann-Whitney U-Test (both variables are categorical, ranksum2 in Stata)

4) Comparing more than two related samples: The Friedman Test

5) Comparing more than two unrelated samples: the Kruska-Wallis H-Test (kwallis2 in Stata)

6) Comparing variables of ordinal or dichotomous scales: Spearman Rank-Order (both variables are ordinal, spearman in Stata), Point-Biserial (pbis in Stata), Biserial (if one is ordinal, one is continuous) Correlations

7) Tests for nominal scale data: Chi-square and Fisher Exact Tests

Monday

Checking for multivariate normality

 

When we run a multiple regression, we assume that the predictors in our model are normally distributed. Not each of them! But the combination of them is normally distributed. That is called ‘multivariate normality’.

We, again, need to check for that. In the Z table, we note that data points that lie below –3 standard deviations or above +3 standard deviations would mean ‘outliers’ because they are too far from the mean (3 sd). Conventionally, we can use these points (-3: +3) to say which observations are ‘out of the range’ when we consider multivariate normality.

zdemo

 

What we need to do now is to combine all predictors and see which points are outliers. This task is extremely difficult, unless we use a computer program called biplot.

 

Graph MV

We now see that observations 8, 265, 266 are potentially outliers, though they are not out of the range. We may delete them if our sample size is large enough, or keep them since these outliers are not severe.

Basic nonparametric tests in Stata

There are many variables that do not follow the rule of normal distributions. We may call them ‘qualitative’, ‘discrete’ or ‘categorical’ variables. What to do ? Consider the following cases
1) One continuous variable to be compared across 2, 3 groups or more? And you want to do something like ANOVA ?
If alcohol is an index, if it is a continuous variable following normal distribution, and you would like to compare alcohol across 4 groups (numbered from 1 to 4 in variable ‘ex’):
anova alcohol i.ex
pwcompare i.ex
--------------------------------------------------------------
             |                                 Unadjusted
             |   Contrast   Std. Err.     [95% Conf. Interval]
-------------+------------------------------------------------
          ex |
    2 vs 1  |    7.84005   .6146453       6.63063     9.04947
     3 vs 1  |   1.063768   .7986802     -.5077717    2.635308
     4 vs 1  |   11.31003   .7748954      9.785287    12.83477
     3 vs 2  |  -6.776282   .8104726     -8.371025   -5.181539
    4 vs 2  |   3.469976   .7870442      1.921332     5.01862     4 vs 3  |   10.24626   .9378378      8.400902    12.09161--------------------------------------------------------------
Yellow lines indicate significant differences. But if it is clearly that, as an index, alcohol does not follow normal distribution? We should think about rank.
kwallis2 alcohol, by(ex)
(where alcohol is a continuous variable, not normally distributed)
One-way analysis of variance by ranks (Kruskal-Wallis Test)
ex       Obs   RankSum  RankMean
--------------------------------
  1      115   9896.00     86.05
  2      104  22142.00    212.90
  3       45   4775.00    106.11
  4       49  12328.00    251.59
Chi-squared (uncorrected for ties) =   178.123 with    3 d.f. (p = 0.00010)
Chi-squared (corrected for ties)   =   191.758 with    3 d.f. (p = 0.00010)
Multiple comparisons between groups
-----------------------------------
(Adjusted p-value for significance is 0.004167)
Ho: alcohol(ex==1) = alcohol(ex==2)
    RankMeans difference =    126.85  Critical value =     32.31
    Prob = 0.000000 (S)
Ho: alcohol(ex==1) = alcohol(ex==3)
    RankMeans difference =     20.06  Critical value =     41.98
    Prob = 0.103737 (NS)
Ho: alcohol(ex==1) = alcohol(ex==4)
    RankMeans difference =    165.54  Critical value =     40.73
    Prob = 0.000000 (S)
Ho: alcohol(ex==2) = alcohol(ex==3)
    RankMeans difference =    106.79  Critical value =     42.60
    Prob = 0.000000 (S)
Ho: alcohol(ex==2) = alcohol(ex==4)
    RankMeans difference =     38.69  Critical value =     41.37
    Prob = 0.006809 (NS)
Ho: alcohol(ex==3) = alcohol(ex==4)
    RankMeans difference =    145.48  Critical value =     49.30
    Prob = 0.000000 (S)

2) Compare two groups?
Two-sample Wilcoxon rank-sum (Mann-Whitney) test
ranksum2 bin, by(sex)
Two-sample Wilcoxon rank-sum (Mann-Whitney) test
         sex |      obs    rank sum    expected
-------------+---------------------------------
      Female |       39      2518.5      3607.5
        Male |      145     14501.5     13412.5
-------------+---------------------------------
    combined |      184       17020       17020
unadjusted variance    87181.25
adjustment for ties   -25532.48
                     ----------
adjusted variance      61648.77
Ho: bin(sex==Female) = bin(sex==Male)
             z =  -4.386
    Prob > |z| =   0.0000
U/mn         .69257294
3) Association between two categorical variables?
spearman bin sex, stats(rho p)
Number of obs =     184
Spearman's rho =       0.3242
Test of Ho: bin and sex are independent
    Prob > |t| =       0.0000
4) Regression with many variables?
logistic bin sex ethnic

Thursday

All possible models

This interesting command allows an exploration of all possible models

In STATA only.

allpossible modelcmd depvar varlist [if exp] [in range] [weight] , {
               eclass(statistics_list) | rclass(progname statistics_list) } [ detail
               npmax(#) format(format) cellwidth(#) modelcmd_options]