T tests in censored normal distribution

T-tests assume normal distributed random variates. We experience designs in which data take on a particular value and are otherwise normal distributed on a half axis. It might be the case in experiments in which a regimen is only applied due to symptoms or the participant's request. In such case a censored normal distribution assumption will increase the strength of the experiment, i.e. statistical power will be increased and assumptions of an underlying normal distribution will be easier to justify.
The suggested test is based on the chi-square likelihood ratio test. The example is two independent groups of identical independently distributed samples and the null hypothesis is identical distributions. The density function splits into single probability, given as the integral of a normal density from negative infinity to the lower bound and the density of the normal distribution for values larger than the lower bound.
The expression involves the distribution function for a Gaussian variate due to normalization. Likelihood ratio value can be calculated in SAS using PROC LIFEREG. P-value is then calculated with a lookup of a chisq quantile value.

%macro tobittest(var_name,group_name,datafile);
/*https://support.sas.com/documentation/cdl/en/statug/63033/HTML/default/viewer.htm#statug_lifereg_sect035.htm;*/
data analysedata;
set &datafile;
lower=&var_name.;
if missing(&var_name.) then delete;
run;
proc lifereg data=analysedata outest=OUTEST(keep=_scale_);
class &group_name.;
model (lower, &var_name.) = &group_name. / d=normal;
probplot ppout npintervals=simul;
inset;
output out=OUT xbeta=Xbeta;
ods output ParameterEstimates=PE;
ods select ParameterEstimates FitStatistics;
run;
data predict;
drop lambda _scale_ _prob_;
set out;
if _n_ eq 1 then set outest;
lambda = pdf('NORMAL',Xbeta/_scale_)
/ cdf('NORMAL',Xbeta/_scale_);
Predict = cdf('NORMAL', Xbeta/_scale_)
* (Xbeta + _scale_*lambda);
res=(&var_name.-predict)/_scale_;
label Xbeta='MEAN OF UNCENSORED VARIABLE'
Predict = 'MEAN OF CENSORED VARIABLE';
run;
title1 "&var_name.";
title2 "Tobit regression";
proc univariate data=predict(where=(NOT(missing(&var_name.))));
var res;
qqplot /normal(mu=est sigma=est);
ods select qqplot;
run;
title2 "Test for differences between groups";
proc print data=PE(where=(Parameter EQ '&group_name.'));
run;
title2 "Summary statistics across groups";
proc tabulate data=analysedata;
class &group_name.;
var &var_name.;
table &group_name. all,&var_name.*(min p25 median p75 max);
run;
data zero;
set analysedata;
zero=0;
if &var_name. EQ 0 then zero=1;
run;
title2 "Zero values (censored values)";
proc tabulate data=zero;
class &group_name. zero;
table &group_name. all,zero all;
run;
title1;
%mend;

ods pdf;
%tobittest(varnameOnFile,groupvarOnFile,datanameFile);

In R the package censReg by Arne Henningsen should perform a similar calculation, however I have found lack of convergence and thus recommend using SAS.

Alder/korrekt århundrede udfra cpr nummer

De fleste, der arbejder med registre eller databaser, står ofte med problemstillingen, at alder er uoplyst, medens cpr-nummer er kendt. Hvordan regner man den ud? Følgende regel er gældende: Hvis syvende ciffer er 0, 1, 2 eller 3 er man født i det 20. århunderede (1900-tallet) Ligeledes, hvis syvende ciffer er 4 eller 9, og årstallet (femte og sjette ciffer) er større end eller lig 37. Endelig er man født i det 19. århundrede (1800-tallet) hvis syvende ciffer er 5, 6, 7 eller 8 og årstallet er større end eller lig 58. Nedenfor finder du eksempel i SAS kode: En lille makro, der udover fødselsdato også udregner køn samt den præcise alder givet datovariabel. Kilde: Opbygning af CPR nummeret, cpr.dk proc format library=work; value gender 0="Female" 1="Male" ; run; %macro agefromCPR(cpr,datevar=inddto,birthvar=birth,agevar=age); dy_temp=input(substrn(&cpr,1,2),2.); mt_temp=input(substrn(&cpr,3,2),2.); yr_temp=input(substrn(&cpr,5,2),

Comorbidity indexes in SQL

Generating Elixhauser comorbidity index from Danish National Health Register as relational database. ( ICD 10 Coding  in SAS) A lookup-table based version of Charlson comorbidity index I made in SQL. A similar approach can be applied to Elixhauser. SELECT V_CPR, MAX(EI1)+MAX(EI2)+MAX(EI3)+MAX(EI4)+MAX(EI5)+ MAX(EI6)+MAX(EI7)+MAX(EI8)+MAX(EI9)+MAX(EI10)+ MAX(EI11)+MAX(EI12)+MAX(EI13)+MAX(EI14)+MAX(EI15)+ MAX(EI16)+MAX(EI17)+MAX(EI18)+MAX(EI19)+MAX(EI20)+ MAX(EI21)+MAX(EI22)+MAX(EI23)+MAX(EI24)+MAX(EI25)+ MAX(EI26)+MAX(EI27)+MAX(EI28)+MAX(EI29)+MAX(EI30)+MAX(EI31) AS Elixhauser FROM (SELECT V_CPR, -- Congestive Heart Failure CASE WHEN DIAG LIKE 'DI099%' OR DIAG LIKE 'DI110%' OR DIAG LIKE 'DI130%' OR DIAG LIKE 'DI132%' OR DIAG LIKE 'DI255%' OR DIAG LIKE 'DI420%' OR DIAG LIKE 'DI425%' OR DIAG LIKE 'DI426%' OR DIAG LIKE 'DI427%' OR DIAG LIKE 'DI428%' OR DIAG LIKE 'DI429%' OR D