{
  "version": "1.0",
  "built": "2026-06-29",
  "source": "Institute of Actuaries of India (IAI) CS1 past papers",
  "train_questions": 243,
  "topic_distribution": {
    "distributions": 70,
    "inference": 82,
    "regression_glm": 48,
    "bayes_credibility": 33,
    "data_analysis": 10
  },
  "subject_distribution": {
    "CS1A": 194,
    "CS1B": 49
  },
  "test_papers": [
    {
      "session": "2025-02",
      "subject": "CS1A",
      "total_marks": 100,
      "question_count": 50,
      "source_qp": "raw/CS1A_2025-02_QP.pdf",
      "source_sol": "raw/CS1A_2025-02_SOL.pdf"
    },
    {
      "session": "2025-05",
      "subject": "CS1A",
      "total_marks": 100,
      "question_count": 19,
      "source_qp": "raw/CS1A_2025-05_QP.pdf",
      "source_sol": "raw/CS1A_2025-05_SOL.pdf"
    },
    {
      "session": "2025-11",
      "subject": "CS1A",
      "total_marks": 100,
      "question_count": 12,
      "source_qp": "raw/CS1A_2025-11_QP.pdf",
      "source_sol": "raw/CS1A_2025-11_SOL.pdf"
    },
    {
      "session": "2025-02",
      "subject": "CS1B",
      "total_marks": 100,
      "question_count": 4,
      "source_qp": "raw/CS1B_2025-02_QP.pdf",
      "source_sol": "raw/CS1B_2025-02_SOL.pdf"
    },
    {
      "session": "2025-05",
      "subject": "CS1B",
      "total_marks": 100,
      "question_count": 4,
      "source_qp": "raw/CS1B_2025-05_QP.pdf",
      "source_sol": "raw/CS1B_2025-05_SOL.pdf"
    },
    {
      "session": "2025-11",
      "subject": "CS1B",
      "total_marks": 100,
      "question_count": 4,
      "source_qp": "raw/CS1B_2025-11_QP.pdf",
      "source_sol": "raw/CS1B_2025-11_SOL.pdf"
    }
  ],
  "questions": [
    {
      "q_num": 1,
      "marks": 8,
      "topic": "distributions",
      "subtopics": [],
      "stem": "",
      "parts": [
        {
          "label": "i",
          "marks": 3,
          "text": "In a game, a player pulls three cards at random from a full deck of 52 cards, and earns as\n                  many points as the number of red cards among the three. Assume 2 people each play this\n                  game once with their own decks, and let X be the sum of their combined points. Derive\n                  the moment generating function of X.                                                          (5)\n\n                                                                                    12                     2\n         ii)      Hence prove that he mean of total points earned by the players is (52) ( (26\n                                                                                            3\n                                                                                              ) + (26\n                                                                                                   1\n                                                                                                     )(26\n                                                                                                       2\n                                                                                                         ))     (3)\n                                                                                    3",
          "topic": null
        }
      ],
      "solution": "i) Probability distribution of points (P) scored is:\n                                              1  26 26\n                                    =          ((   ) ( ) for P = 0)\n                                          (52)3\n                                                  0    3\n\n                                          1     26 26\n                                              ((   ) ( ) for P = 1)\n                                          3\n                                                 1    2\n\n                                          1     26 26\n                                              ((   ) ( ) for P = 2)\n                                          3\n                                                 2    1\n\n                                          1     26 26\n                                              ((   ) ( ) for P = 3)\n                                          3\n                                                 3    0\n\nMGF of this function is\n            1   26   26         0\n                                         26 26             26 26           26 26\n      =         (( ) ( ) × e        +(      ) ( ) × e1t + ( ) ( ) × e2t + ( ) ( ) × e3t )\n           (52\n            3\n              )   0   3                   1    2            2  1            3  0\n\nHence MGF of X where X=P1+P2 (subscript 1 and 2 refer to the player 1 and player 2\nrespectively) can be stated as :\n\n𝐸(𝑒 𝑡𝑋 ) = 𝐸(𝑒 𝑡 ∑ 𝑃𝑖 ) = ∏ 𝐸(𝑒 𝑡𝑃𝑖 )\nAs the draws are independent,\n                           2\n∏ 𝐸(𝑒 𝑡𝑃𝑖 ) = (𝐸(𝑒 𝑡𝑃 ))\n\n      1\n= ((52) ((26\n          3\n            ) + (26\n                 1\n                   )(26\n                     2\n                       ) × e1t + (26\n                                  2\n                                    )(26\n                                      1\n                                        ) × e2t + (26\n                                                   3\n                                                     )×\n       3\n       2\n 3t\ne ))                                                                                        [1]\n\n                           𝑑(𝐸(𝑒 𝑡𝑋 ))\nii) Using MGF, 𝐸(𝑋) =                    𝑓𝑜𝑟 𝑡 = 0                                          [1]\n                               𝑑𝑡\n                                                                               1\n                2 26    26 26           26 26           26\n      𝐸(𝑋) = 52 (( ) + ( ) ( ) × e1t + ( ) ( ) × e2t + ( ) × e3t )\n            ( )    3     1  2            2  1            3\n                3\n                         26 26             26 26              26\n                     × (( ) ( ) × e1t + 2 ( ) ( ) × e2t + 3. ( ) × e3t ) 𝑓𝑜𝑟 𝑡 = 0\n                          1  2              2  1               3\n\n                                                                                   Page 1 of 11\n\fIAI                                                                                               CS1A-0619\n\n                      2                                                                     12             2\nHence 𝐸(𝑋) = (52) (2 (26\n                      3\n                         ) + 2 (26\n                                1\n                                  )(26\n                                    2\n                                      )) × (3(26\n                                              2\n                                                 )(26\n                                                   1\n                                                     ) + 3. (26\n                                                             3\n                                                               )) = (52) ( (26\n                                                                            3\n                                                                               ) + (26\n                                                                                    1\n                                                                                      )(26\n                                                                                        2\n                                                                                          ))\n                      3                                                                      3\n\nPart (i) was not answered correctly by majority of the candidates and hence they could not\nattempt the Part (ii). Students who could identify the PDF of the event scored highly on the\nquestion.",
      "has_math": false,
      "session": "2019-06",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2019-06_QP.pdf",
      "source_sol": "raw/CS1A_2019-06_SOL.pdf"
    },
    {
      "q_num": 2,
      "marks": 15,
      "topic": "inference",
      "subtopics": [],
      "stem": "A study into the average claim (in Rs. ‘000) per health insurance policy was performed for the\n      claims incurred in public and private hospitals. Data for some cities is given below:\n\n                        City 1   City 2    City 3    City 4   City 5    City 6   City 7    City 8    City 9\n              Public     24        45       29        33        20       40       26.5       25       27.5\n              Private    30       54.5      30        40       28.5      36       30.5      30.5      35.5",
      "parts": [
        {
          "label": "i",
          "marks": 6,
          "text": "Determine the sample mean and sample variance of average claim size in both the type\n                  of hospitals.                                                                                 (6)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 4,
          "text": "State the primary condition that needs to be true for testing equal mean and verify\n                  whether that condition is satisfied in the above example (You may assume that the\n                  samples come from a normal population).                                                       (4)",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 5,
          "text": "Test whether the treatments in private hospitals result in higher claim size at 95% level.    (5)",
          "topic": null
        }
      ],
      "solution": "i)\n                                            ∑ 𝑋𝑃𝑢𝑏𝑙𝑖𝑐 = 270                                             [0.5]\n                                         (𝑋̅𝑃𝑢𝑏𝑙𝑖𝑐 ) = 30                                               [0.5]\n                                              ∑ 𝑋𝑃𝑟𝑖𝑣𝑎𝑡𝑒 = 315.5                                        [0.5]\n                                               (𝑋̅𝑃𝑟𝑖𝑣𝑎𝑡𝑒 ) = 35.06                                     [0.5]\n                                                   2\n                                              ∑ 𝑋𝑃𝑢𝑏𝑙𝑖𝑐   = 8614.5                                        [1]\n             (subscript 1 refers public hospitals)\n                                       𝑆12 = (8614.5 − 9 × 302 )/8 = 64.31                                [1]\n                                                   2\n                                              ∑ 𝑋𝑃𝑟𝑖𝑣𝑎𝑡𝑒   = 11599.25                                     [1]\n                                       2\n                                     𝑆2 = (11599.25 − 9 × 35.062 )/8 = 67.42                              [1]\n\n       ii)       Two sided t-test can be applied in case the samples come from populations with\n                             equal variances.\n                 We are testing 𝐻0 ∶ 𝜎12 = 𝜎22 𝑣𝑠 𝐻1 ∶ 𝜎12 ≠ 𝜎22\n                                            𝑆 2 ⁄𝜎2\n                          Test statistic is 𝑆12 ⁄𝜎12 ~𝐹𝑛1 −1,𝑛2−1\n                                             2   2\n                                                                     64.31\n                                           Value of statistic is 67.42 = 0.9542\n             𝐹8,8 values at 5% levels are 0.2256 and 4.433 . Since the value of the test statistic is\n             between the above values, we have insufficient evidence to reject the hypothesis and\n             conclude that the population variances are equal.\n\n      iii)      We are testing 𝐻0 ∶ 𝜇1 = 𝜇2 𝑣𝑠 𝐻1 ∶ 𝜇1 < 𝜇2\n                                                           (𝑋̅2 −𝑋̅1 )−(𝜇2 −𝜇1 )\n                                       Test statistic is                           ~𝑡𝑛1 +𝑛2 −2\n                                                             2 (1⁄𝑛 +1⁄𝑛 )\n                                                           √𝑆𝑃     1    2\n\n                                                                                                 Page 2 of 11\n\fIAI                                                                                              CS1A-0619\n                          𝑆 2 (𝑛1 −1)+𝑆22 (𝑛2 −1)\n           Where 𝑆𝑃2 = 1        (𝑛1 +𝑛2 −2)\n           Using the values in section (I) above,\n                                                            64.31×8+67.42×8\n                                                    𝑆𝑃2 =          16\n                                                                            = 65.86\n           value of test statistic is\n                                                                 (35.06−30)−0\n                                                                                 = 1.32\n                                                               √65.86(1⁄9+1⁄9)\n\n                                                                     𝑃(𝑡16 > 1.32) = 20.5%\n\n           This is higher than 95% hence we do not have sufficient evidence to reject the\n           hypothesis and hence conclude that the cost of claims in private hospitals is similar to\n           that in public hospital\n\nPart (i) and (iii) of this question was generally well answered by most of the candidates. Many\ncandidates could not attempt the part (ii) correctly.",
      "has_math": false,
      "session": "2019-06",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2019-06_QP.pdf",
      "source_sol": "raw/CS1A_2019-06_SOL.pdf"
    },
    {
      "q_num": 3,
      "marks": 9,
      "topic": "distributions",
      "subtopics": [],
      "stem": "An academician proposes that the joint distribution of a number of weeks of study leave (X)\n      and the proportion of questions answered correctly (Y) is given by:\n\n                         9       1\n         𝑓𝑋𝑌 (𝑥, 𝑦) =      𝑥𝑦 2 + 𝑓𝑜𝑟 0 ≤ 𝑥 ≤ 2, 0 ≤ 𝑦 ≤ 1\n                        10       5",
      "parts": [
        {
          "label": "i",
          "marks": 3,
          "text": "Determine the Marginal distributions of X and Y.                                              (3)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 4,
          "text": "A student is selected at random, what would be the expected number of weeks of study\n                  leave and expected proportion of questions answered correctly.                                (4)",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 2,
          "text": "Compute the covariance of X and Y                                                             (2)",
          "topic": null
        }
      ],
      "solution": "i)\n\n                                                                                   [0.5 Marks for each step]\n\n                                                                                   [0.5 Marks for each step]\n\n                                                                                               Page 3 of 11\n\fIAI                                                                              CS1A-0619\n\n       ii)\n\niii)\n\nThis question was well answered by most of the candidates. Some candidates lost marks due\nto incorrect computation or evaluation of the integrals but the concepts were correctly\napplied by majority of the candidates.\n\n                                                                               Page 4 of 11\n\fIAI                                                                                   CS1A-0619",
      "has_math": true,
      "session": "2019-06",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2019-06_QP.pdf",
      "source_sol": "raw/CS1A_2019-06_SOL.pdf"
    },
    {
      "q_num": 4,
      "marks": 17,
      "topic": "regression_glm",
      "subtopics": [],
      "stem": "You are an actuarial analyst working at a life insurance company in India. Your actuary has\n      asked you to analyze the claims data to aid her in setting pricing assumptions for a new product.\n      You have obtained the following information from the claims department:\n                   Age band          Number of claims per 10,000 policies\n                     0-10                           112\n                    11-20                           122\n                    21-30                           133\n                    31-40                           187\n                    41-50                           258\n                    51-60                           400\n                    61-70                           522\n\n         The actuary has asked you to perform a regression analysis on this data to identify the\n         relationship between age and number of claims.",
      "parts": [
        {
          "label": "i",
          "marks": 6,
          "text": "Perform a linear regression of the claims as a function of the age                              (6)\n\n         Note: You can assume that the rate of claims in a particular age band is constant and hence\n         perform a regression on the middle age of each age band.",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 2,
          "text": "Calculate the R-squared for the regression model                                                (2)",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 3,
          "text": "Calculate a two-sided 95% confidence interval for the slope parameter                           (3)",
          "topic": null
        },
        {
          "label": "iv",
          "marks": 3,
          "text": "Carry out an F test to determine whether the slope parameter is zero                            (3)",
          "topic": null
        },
        {
          "label": "v",
          "marks": 3,
          "text": "Given below is the scatter plot of residuals against the values of the independent variable.\n                Please Comment on the plot.                                                                     (3)",
          "topic": null
        }
      ],
      "solution": "i) The slope and intercept parameters can be derived as the expected value of β and α in the\nfollowing equation\n\n        𝑦 = 𝛼 + 𝛽𝑥+∈                                                                   [0.5]\n        Where, ∈ are the error terms that are assumed to be identical independently\n        distributed normal random variables.\n        Under linear regression, α and β can be found by minimizing the squared errors -\n        distance between the observed and predicted values of y.\n\n        The mathematical expressions of the expected values of α and β are\n                                         ∑𝑛𝑖=1(𝑥𝑖 − 𝑥̅ ) (𝑦𝑖 − 𝑦̅)\n                                    ̂\n                                   𝛽=\n                                             ∑𝑛𝑖=1(𝑥𝑖 − 𝑥̅ )2\n\n                                                        𝛼̂ = 𝑦̅ − 𝛽̂ ∗ 𝑥̅                     [0.5]\n        Using the data given,\n        Slope = 6.825, intercept = 8.84                                    [3]\n        Alternative answer : intercept: 7.65, slope 6.78\n        *********************************************************************\n\n        Splitting the total sum of squares into regression and residual sum of squares:\n        SS Total = SS Regression + SS Residual\n        Using standard notations; SS Total = S yy; SS Regression = S2XY/ S XX\nii) R – squared = SS Regression / SS Total = 130425.8 / 149597.4 = 0.8718\n        alternate answer : 86.86%                                                               [2]\n\n                                  ̂2\n                                  𝜎           3834.34\niii) Standard error of 𝛽 = √            = √             = 1.170216                            [1.5]\n                                  𝑆𝑥𝑥            2800\n\n        The two sided 95% confidence interval for β = 𝛽̂ ± 𝑡0.025,5 ∗ 𝑠𝑒(𝛽)\n        i.e. 6.825 ± 2.571 ∗ 1.170216 = (3.8164, 9.8336)                                      [1.5]\n        alternate answer: 6.78 +/- 2.571*1.1782\n\niv) ANOVA Table\n\n          Source of variation           Degrees of freedom                  SS       MSS\n\n               Regression                           1                   130425.8   130425.8\n\n                 Residual                           5                   19171.68   3834.336\n\n                   Total                            6                   149597.4\n\n                                                                                    Page 5 of 11\n\fIAI                                                                                 CS1A-0619\n\n       F-test:\n       H0: β = 0\n       F-statistic = 130425.8/ 3834.336 = 34.01521 on (1,5) degrees of freedom               [1]\n       Critical value of F(1,5) = 10.01                                                    [0.5]\n       Since, F-statistic is greater than the critical value,\n       so H0 is rejected at the 2.5%% level.                                               [0.5]\n\n       Hence, the slope parameter is statistically significantly different from zero.        [1]\n\n       alternate answer:\n\n                                                                        Significance\n                             df        SS           MS           F             F\n            Regression          1 129951.4 129951.4 33.07313                 0.00223\n            Residual            5 19646.07 3929.213\n            Total               6 149597.4\n       Note for markers: At all levels of significance (1%, 2.5%, 5%), critical values for F\n       distribution will be lower than the test statistic value. Marks should be provided for\n       any level of significance used by the student.                                        [3]\n\nv) The residual plot is a U-shaped graph and the residuals are observed to follow a pattern [1]\n       The non-random pattern in the residuals indicates that the deterministic portion of\n       the regression model is not capturing some explanatory information. The\n       possibilities could include:\n       1. A missing variable\n       2. A missing of higher order term of a variable in the model to explain the\n           polynomial trend in residuals                                                 [2]\n        From, the above graph it looks likely that including a higher order term of the\n        independent variable should be able to resolve this problem.                       [3]\nQuestion (except part (v) ) was largely well answered. Computational errors were made by\nsignificant number of candidates. In part (v) most of the candidates identified that the\ndistributaion residuals does not seem to be normal but could not provide any additional\ncomments beyond that.",
      "has_math": false,
      "session": "2019-06",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2019-06_QP.pdf",
      "source_sol": "raw/CS1A_2019-06_SOL.pdf"
    },
    {
      "q_num": 5,
      "marks": 7,
      "topic": "inference",
      "subtopics": [],
      "stem": "If 𝜃̂ is an estimator of parameter 𝜃, answer the following:",
      "parts": [
        {
          "label": "i",
          "marks": 1,
          "text": "Define unbiased estimator                                                                       (1)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 1,
          "text": "Define ‘bias’.                                                                                  (1)",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 1,
          "text": "Define Mean Square Error (MSE) of this estimator 𝜃̂                                             (1)\n\n         There exist another estimator 𝜃̃ of the same parameter 𝜃, such that 𝜃̂ has no bias but higher\n         MSE than 𝜃̃ while 𝜃̃ has a positive bias.",
          "topic": null
        },
        {
          "label": "iv",
          "marks": 1,
          "text": "State giving reason, which estimator is ‘efficient’?                                            (1)",
          "topic": null
        },
        {
          "label": "v",
          "marks": 1,
          "text": "When would either of the two estimators be termed as consistent?                                (1)",
          "topic": null
        },
        {
          "label": "vi",
          "marks": 2,
          "text": "Outline (in one sentence each) any two methods of estimating 𝜃                                  (2)",
          "topic": null
        }
      ],
      "solution": "(i) 𝜃̂ said to be unbiased when 𝐸(𝜃̂) = 𝜃                                                    [1]\n(ii) measure of the ‘bias’ is given by 𝐸(𝜃̂) − 𝜃                                             [1]\n                                                                 2\n(iii) Mean Square Error (MSE) of this estimator 𝜃̂ = (𝐸(𝜃̂) − 𝜃)                           [1]\n      ̃\n(iv) 𝜃 is efficient as an estimator with lower MSE is said to be more efficient than one with\nhigher MSE.                                                                                [1]\n (v) An estimator is termed as consistent if MSE converges to 0 as the sample size tends to ∞\n\n(vi) 𝜃 can be estimated using:                         [mention any 2 methods, 1 mark each]\n                                                                                   Page 6 of 11\n\fIAI                                                                                    CS1A-0619\n\na. Method of moments: the population moments are equated to the sample moments to\nestimate the parameters.\n\nb. Maximum likelihood method: A maximum likelihood function 𝐿(𝜃) = ∏𝑛𝑖=1 𝑓(𝑥𝑖 ; 𝜃) is\n                                                                                         𝑑𝐿(𝜃)\ngenerated. A maximum likelihood estimate of the parameter is given by solution to                =0\n                                                                                          𝑑𝜃\n\nc. Bootstrap method: This is computer intensive method that allows us to avoid making\nassumption about the sampling distribution by forming an empirical sampling distribution\nwhich is possible due to re-sampling based on the available sample.\nThis was a bookwork question and was well answered. In part (vi) some students provided\nonly the names of the methods without the accompanying narration.",
      "has_math": true,
      "session": "2019-06",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2019-06_QP.pdf",
      "source_sol": "raw/CS1A_2019-06_SOL.pdf"
    },
    {
      "q_num": 6,
      "marks": 4,
      "topic": "regression_glm",
      "subtopics": [],
      "stem": "Define",
      "parts": [
        {
          "label": "i",
          "marks": 2,
          "text": "Pearson residuals                                                                             (2)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 2,
          "text": "Deviance residuals                                                                            (2)\n\n          Clearly describe all notations used.",
          "topic": null
        }
      ],
      "solution": "The residuals are based on differences between the observed responses, y, and the\n       fitted responses, 𝜇̂ .\n       Pearson residuals\n                                         ̂\n                                       𝑦−𝜇\n       𝑃𝑒𝑎𝑟𝑠𝑜𝑛 𝑟𝑒𝑠𝑖𝑑𝑢𝑎𝑙 =                                                                        [2]\n                                          ̂)\n                                     √𝑉𝑎𝑟(𝜇\n       Deviance residuals\n\n                                    𝐷𝑒𝑣𝑖𝑎𝑛𝑐𝑒 𝑟𝑒𝑠𝑖𝑑𝑢𝑎𝑙 = 𝑠𝑖𝑔𝑛(𝑦 − 𝜇̂ ) ∗ 𝑑𝑖\n       𝑤ℎ𝑒𝑟𝑒 𝑑𝑖 𝑖𝑠 𝑡ℎ𝑒 𝑐𝑜𝑛𝑡𝑟𝑖𝑏𝑢𝑡𝑖𝑜𝑛 𝑜𝑓 𝑦 𝑡𝑜 𝑡ℎ𝑒 𝑠𝑐𝑎𝑙𝑒𝑑 𝑑𝑒𝑣𝑖𝑎𝑛𝑐𝑒 (∑ 𝑑𝑖2 )                         [2]\nThis was a bookwork question and was generally well answered. Some students missed on\nproviding the meaning of the terms used.",
      "has_math": false,
      "session": "2019-06",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2019-06_QP.pdf",
      "source_sol": "raw/CS1A_2019-06_SOL.pdf"
    },
    {
      "q_num": 7,
      "marks": 4,
      "topic": "bayes_credibility",
      "subtopics": [],
      "stem": "State and derive the Bayes’ theorem                                                                     [4]",
      "parts": [],
      "solution": "If B1, B2, B3,…,Bk constitute a partition of a sample space S and P(Bi) ≠ 0 for i=1,2,3,….k,\nthen for any event A in S such that P(A) ≠ 0:\n                     𝑃(𝐴|𝐵𝑟 )𝑃(𝐵𝑟 )\n       𝑃(𝐵𝑟 |𝐴) =                     𝑤ℎ𝑒𝑟𝑒 𝑃(𝐴) = ∑𝑘𝑖=1 𝑃(𝐴|𝐵𝑖 ) ∗ 𝑃(𝐵𝑖 ) 𝑓𝑜𝑟 𝑟 = 1, 2, 3, … . 𝑘\n                         𝑃(𝐴)\n\n       Derivation:\n       P (A∩B) = P (A) P (B|A)\n                                          𝑃(𝐴∩𝐵)\n       On rearranging: 𝑃(𝐵|𝐴) =              𝑃(𝐴)\n\n       However, 𝑃(𝐴 ∩ 𝐵) = 𝑃(𝐵 ∩ 𝐴) = 𝑃(𝐵)𝑃(𝐴|𝐵)\n       Now, replacing B by Br, we have:\n                     𝑃(𝐵𝑟 ∩𝐴)       𝑃(𝐵𝑟 )𝑃(𝐴|𝐵𝑟 )\n       𝑃(𝐵𝑟 |𝐴) =               =\n                      𝑃(𝐴)              𝑃(𝐴)\n\n                                                                                      Page 7 of 11\n\fIAI                                                                                   CS1A-0619\n\n         And from the law of total probability:\n         𝑃(𝐴) = ∑𝑖 𝑃(𝐴|𝐵𝑖 ) ∗ 𝑃(𝐵𝑖 )                                                         [2.5]\nMany students failed to provide the proof. Majority of the attempts were limited to statement\nof the Bayes’ Theorem",
      "has_math": false,
      "session": "2019-06",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2019-06_QP.pdf",
      "source_sol": "raw/CS1A_2019-06_SOL.pdf"
    },
    {
      "q_num": 8,
      "marks": 13,
      "topic": "distributions",
      "subtopics": [
        "inference"
      ],
      "stem": "",
      "parts": [
        {
          "label": "i",
          "marks": 3,
          "text": "Wickets taken by a cricket team ‘A’ follows Poisson process with rate of 1 wicket per\n                 100 balls bowled. How many wickets will the team take with 95% confidence after\n                 bowling 500 balls?                                                                            (3)\n\n          ii)    A cricket team ‘B’ has batsmen for the last wicket that can score 1 run per ball with a\n                 probability of 40% and 0 runs with probability of 60%. The team B (playing against the\n                 above team A and having only 1 wicket in hand) needs to score 26 runs in 50 balls.\n                 Team A wins if it takes the 1 wicket in these 50 balls and team B wins in case it scores\n                 the required runs in 50 balls. Determine which team has a higher probability of win. State\n                 any assumptions you make.                                                                     (7)\n\n          iii)   Hence determine the probability that team B will bat for at least 30 balls                    (3)",
          "topic": null
        }
      ],
      "solution": "i) Wickets taken per 500 balls follow Poi(5) distribution. As the number of trials (balls) is very\nhigh and poisson parameter >= 5, we can use normal approximation to Poison Distribution.\n      Thus the wickets take approximately follow N(5,5).                                    [0.5]\n      Hence we need ‘X’ such that:\n               𝑋−5\n      𝑃 (𝑍 >        ) = 0.95                                                                [0.5]\n               √5\n\n      Critical value at 95% confidence is 1.65                                               [0.5]\n      Thus\n                    𝑋−5\n      𝑃 (1.65 >           ) = 0.95\n                     √5\n\n      Hence 𝑋 < 8.68                                                                         [0.5]\n     As number of wickets can only take whole values, we need to truncate the number to\n     lower whole number. Hence the team takes upto 8 wickets at 95 % confidence level.\nii) For team B the runs in 50 ball will follow Bin(50,0.4)                           [0.5]\nThe mean and variance for this Binomial distribution are 50 × 0.4 = 20 𝑎𝑛𝑑 50 × 0.4 ×\n(1 − 0.4) = 12 respectively                                                       [1]\nFor large number of trials and probability of success is close to 0.5 (or np>10), normal\napproximation can be applied & thus the runs per 50 balls follows approximately N(20,12) [1]\n                                                                           26−20\nProbability of team B scoring 26 or more runs in 50 balls is thus 𝑃 (𝑍 >           ) = 4.16% [1]\n                                                                            √12\n\nThe Poisson rate of taking wickets (by team A) is 1 per 100 balls i.e. 0.01 per ball. Hence, the\nwickets taken by team A in 50 balls has rate = 0.01 x 50 i.e it follows Poi(0.5) process.     [1]\n                                         0.50\nProbability of not taking any wicket is 1 ∗ 𝑒 −0.5 = 60.65%                                   [1]\n\nHence probability of winning is 1-0.6065=39.34%                                               [1]\nThus Team A has higher probability of winning.                                               [0.5]\n\n                                                                                     Page 8 of 11\n\fIAI                                                                                   CS1A-0619\n\niii) Probability that team B bats for 30 balls = (Probability of waiting time (in terms of number\nof balls)> 30) x (Probability of A not scoring 26 runs in 30 balls)                          [0.5]\nThe waiting time has Exp(0.01) distribution, hence P(T>30)=exp(-0.01*30) = 74.08%              [1]\n                                                26−30×0.4\nProbability of A not scoring 30 runs = 𝑃 (𝑍 <               ) = 97.41%                        [1]\n                                                √30×0.4×0.6\n\nHence there is a 74.08% x 97.41% = 72.16% chance that team B will bat for 30 balls.        [0.5]\n\nMost of the students struggled with this question. Not applying CLT, not rounding –off the\nnumber of wickets were the common mistakes in part (i). In part (ii) most of the students\ncomputed probability of scoring ‘exactly’ 26 runs and used that in the answer. Only a handful\nof candidates attempted part (iii)",
      "has_math": false,
      "session": "2019-06",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2019-06_QP.pdf",
      "source_sol": "raw/CS1A_2019-06_SOL.pdf"
    },
    {
      "q_num": 9,
      "marks": 11,
      "topic": "bayes_credibility",
      "subtopics": [],
      "stem": "The annual distribution of claims arising from a portfolio of motor insurance policies follows\n       a Poisson distribution with mean μ. The prior distribution for μ has a gamma distribution with\n       parameters α = 4 and λ =7.\n\n          Claim figures over the last n years are x1, x2, x3….xn.",
      "parts": [
        {
          "label": "i",
          "marks": 3,
          "text": "Show that the posterior distribution is gamma and determine its parameters                    (3)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 3,
          "text": "Given that n = 10, ∑10\n                                     𝑖=1 𝑥𝑖 = 15 determine the Bayesian estimate for μ under\n\n                 a) squared error loss                                                                         (2)\n                 b) ‘all-or-nothing’ loss                                                                      (3)\n                 c) absolute error loss                                                                        (3)",
          "topic": null
        }
      ],
      "solution": "i) Prior distribution of μ is Gamma (4,7)\n                                                    74\n                                     𝑓𝑝𝑟𝑖𝑜𝑟 (𝜇) = Г(4) ∗ 𝜇 3 ∗ 𝑒 −7∗𝜇\n\n       Thus 𝑓𝑝𝑟𝑖𝑜𝑟 (𝜇) 𝑖𝑠 𝑝𝑟𝑜𝑝𝑜𝑟𝑡𝑖𝑜𝑛𝑎𝑙 𝑡𝑜 𝜇 3 ∗ 𝑒 −7𝜇                                          [1]\n\n       The likelihood is the product of the Poisson probabilities:\n                                         𝜇 𝑥1 −𝜇 𝜇 𝑥2 −𝜇    𝜇 𝑥𝑛 −𝜇\n                                𝐿(𝜇) =        𝑒 ∗      𝑒 ∗…      𝑒\n                                         𝑥1 !     𝑥2 !      𝑥𝑛 !\n       The likelihood function is proportional to\n\n       𝐿(𝜇)𝑖𝑠 𝑝𝑟𝑜𝑝𝑜𝑟𝑡𝑖𝑜𝑛𝑎𝑙 𝑡𝑜 𝜇 ∑ 𝑥𝑖 ∗ 𝑒 −𝑛𝜇                                                  [1]\n\n       So, 𝑓𝑝𝑜𝑠𝑡𝑒𝑟𝑖𝑜𝑟 (𝜇) 𝑖𝑠 𝑝𝑟𝑜𝑝𝑜𝑟𝑡𝑖𝑜𝑛𝑎𝑙 𝑡𝑜 𝜇 3+∑ 𝑥𝑖 ∗ 𝑒 −7𝜇−𝑛𝜇\n\n       The posterior distribution of μ thus takes the form of a Gamma (4 + ∑ 𝑥𝑖 ,7+n).        [1]\nii) (a) Squared error loss\n       When n=10 and ∑ 𝑥𝑖 = 15, the posterior distribution of μ is Gamma (19, 17).\n       The Bayesian estimate of μ under squared error loss is the mean of the posterior\n       distribution.                                                                    [1]\n       Bayesian estimate = mean of posterior distribution = mean of Gamma (19, 17) =\n       19/17 = 1.1176                                                                         [1]\n  (b) All-or-nothing loss\n       The Bayesian estimate under all-or-nothing loss is the mode of the posterior\n       distribution.                                                                          [1]\n       To find the mode we need to differentiate the PDF and equate it to zero.\n                                                                                    Page 9 of 11\n\fIAI                                                                                  CS1A-0619\n\n       𝑓𝑝𝑜𝑠𝑡𝑒𝑟𝑖𝑜𝑟 (𝜇) = 𝑘 ∗ 𝜇18 ∗ 𝑒 −17𝜇 𝑤ℎ𝑒𝑟𝑒 𝑘 𝑖𝑠 𝑎 𝑐𝑜𝑛𝑠𝑡𝑎𝑛𝑡                               [1]\n\n       Taking logs and differentiating:\n                                          𝑑 𝑙𝑛𝑝𝑜𝑠𝑡 𝜇 18\n                                                    =   − 17\n                                             𝑑𝜇       𝜇\n       Equating the derivative to zero will give us the value of μ which maximizes the PDF\n       and thus will give us the mode of the distribution.\n       μ = 18/17                                                                           [0.5]\n       Differentiating again gives us (-18/μ^2) which is less than zero. This is a check that\n       the prior step gives us the maxima.                                                  [0.5]\n       So, the Bayesian estimate of all-or-nothing loss is 18/17\n  (c) Absolute error loss\n\n       The Bayesian estimate under absolute error loss is the median of the posterior\n       distribution.                                                                          [1]\n\n       The posterior distribution follows Gamma (19, 17). Let X denote the posterior\n       distribution. Hence X ~ Gamma (19, 17). Then 2*17X ~ Chi squared (2*19).              [1]\n\n       The median of the posterior distribution is the value of M such that P (X<M) = 0.5\n       Equivalently, P ( ϰ238 < 34M) = 0.5\n       From the tables we can see that the 50th percentile of ϰ238 is 37.34:\n       Hence, M = 37.34/34 = 1.098\n       So, the Bayesian estimate under absolute error loss is 1.098                       [1]\n\nPart (i) was answered nicely answered by the well prepared candidates. Students made\nmistakes in Part (ii) by equting the mean / median / mode to loss measures other than those\ngiven in the solution.",
      "has_math": true,
      "session": "2019-06",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2019-06_QP.pdf",
      "source_sol": "raw/CS1A_2019-06_SOL.pdf"
    },
    {
      "q_num": 10,
      "marks": 12,
      "topic": "bayes_credibility",
      "subtopics": [],
      "stem": "A life insurance company sells Group Term Insurance Policies to different private companies.\n       Each company will have a different degree of risk with the difference arising because of a\n       difference in the nature of business, the proportion of blue and white collared employees etc.\n\n          The Life Insurance Company has collated claims data across 2 such group policies. Claim\n          payments (in INR ‘000) and the number of employees covered in each company are as given\n          below:\n\n           Claim Payments: INR ‘000         Year1    Year2     Year3    Year4\n                  Company A                 4084     4387      4550     3456\n                  Company B                 4634     3203      2073     4485\n\n       Number of employees      Year1    Year2    Year3     Year4\n          Company A              121      119      120       110\n          Company B              150      135      122       145\n\n      You are an actuarial analyst in the life insurance company. You have been asked to compute\n      the expected claims payouts for both the companies over the coming year assuming that the\n      number of employees covered will be 135 and 155 for company A and company B.\n\n      Analyze the data using Empirical Bayes Credibility theory – EBCT model 2 and calculate the\n      expected claims payout for companies A and B.                                                [12]\n\n      Note: A group term insurance policy is an insurance product sold to corporates under which\n      all employees of the company are provided life insurance.",
      "parts": [],
      "solution": "Let 𝑌𝑖𝑗 and 𝑃𝑖𝑗 be the claim amounts and number of employees covered for company i and\nyear j respectively.\n               𝑌\nDenote 𝑋𝑖𝑗 = 𝑃𝑖𝑗 ; N=2, n = 4\n                   𝑖𝑗\n\n𝑃̅𝐴 = ∑𝑗 ̅̅̅̅\n         𝑃𝐴𝑗 = 121 + 119 + 120 + 110 = 470\n\n𝑃̅𝐵 = ∑𝑗 ̅̅̅̅\n         𝑃𝐵𝑗 = 150 + 135 + 122 + 145 = 552\n\n𝑃̅ = ̅̅̅\n     𝑃𝐴 + ̅̅̅\n          𝑃𝐵 = 1022\n\n                                                                                  Page 10 of 11\n\fIAI                                                                                 CS1A-0619\n       1                         470                   552\n𝑃∗ = 7 ∗ [470 ∗ (1 − 1022) + 552 ∗ (1 − 1022)] = 72.53                                      [1]\n\nTable for claims per unit employees, 𝑋𝑖𝑗 :\n\n                                 Year1         Year2   Year3   Year4\n           Company\n                                 33.75         36.87   37.92   31.42\n              A\n           Company\n                                 30.89         23.73   16.99   30.93\n              B\n\nUsing the formulae from the tables,\n̅̅̅\n𝑋 𝐴 = 35.0745\n̅̅\n𝑋̅̅\n  𝐵 = 26.0779\n ̅\n𝑋 = 30.2074\n\n𝐸[𝑚(𝜃)] = 30.2074                                                                          [2]\n\n                             2            4\n            1     1                         1\n𝐸[𝑠 2 (𝜃)] = ∗ ∑              ̅̅̅̅\n                     ∗ ∑ 𝑃𝑖𝑗 (𝑋      ̅ 2\n                                𝑖𝑗 − 𝑋𝑖 ) =   ∗ (1011.036 + 5904.065) = 3457.551\n            𝑁    𝑛−1                        2\n                         𝑖=1             𝑗=1\n\n                         1       1\n𝑉𝑎𝑟[𝑚(𝜃)] = 72.53 ∗ (7 ∗ 41214.23 − 3457.551) = 33.5061                                    [2]\n\nPutting the above derived values in the formulae:\n           470\n𝑧𝐴 =          3457.551   = 0.81997                                                        [1.5]\n       470+\n              33.5061\n\n           552\n𝑧𝐵 =          3457.551   = 0.842501                                                       [1.5]\n       552+\n              33.5061\nUsing credibility theory, credibility premium per unit risk volume is given by:\n\n                          𝐶𝑜𝑚𝑝𝑎𝑛𝑦 𝐴: 𝑍𝐴 ∗ 𝑋̅𝐴 + (1 − 𝑍𝐴 ) ∗ 𝐸[𝑚(𝜃)] = 34.1843\n                         𝐶𝑜𝑚𝑝𝑎𝑛𝑦 𝐵: 𝑍𝐵 ∗ 𝑋̅𝐵 + (1 − 𝑍𝐵 ) ∗ 𝐸[𝑚(𝜃)] = 26.72829\n\nThe EBCT claim amounts for the coming year for the two states are:\nCompany A: 4614.88; Company B: 4142.886                                                    [2]\n\nThis question was attempted by majority of the students. Coputational errors were made by\nsome while the rest scored highly in this question.\n\n                                                   ***************\n\n                                                                                  Page 11 of 11",
      "has_math": false,
      "session": "2019-06",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2019-06_QP.pdf",
      "source_sol": "raw/CS1A_2019-06_SOL.pdf"
    },
    {
      "q_num": 1,
      "marks": 6,
      "topic": "distributions",
      "subtopics": [],
      "stem": "A discrete random variable X has the following probability density function:\n\n        𝑃(𝑋 = 𝑥) = 𝑝(1 − p)x−1 where 0 < 𝑝 < 1 and x>0",
      "parts": [
        {
          "label": "i",
          "marks": 4,
          "text": "Derive the Moment Generating Function (MGF) and the Cumulant Generating Function\n           (CGF) for the above distribution.                                                               (4)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 2,
          "text": "Using either the MGF or the CGF determine E(X).                                                (2)",
          "topic": null
        }
      ],
      "solution": "i. MGF of X:\n\n𝑀 (𝑡) = 𝐸[𝑒 ] = ∑              𝑒 𝑃(𝑋 = 𝑥) = ∑           𝑒 𝑝(1 − p)                                          [1]\n\n𝑀 (𝑡) = 𝑝e + 𝑝(1 − p)e + 𝑝(1 − p) e + ⋯                                                                     [1]\n\nThis is infinite Geometric Series with summation given as:\n\n           𝑝e\n𝑀 (𝑡) =         (1 − (1 − 𝑝)e )                                                                             [1]\n\nCGF is hence given by:\n\n                                𝑝e\n𝐶 (𝑡) = ln(𝑀 (𝑡)) = 𝑙𝑛               (1 − (1 − 𝑝)e )                                                        [1]\n\nii. Determining E(X)\n\n Alternative solution 1: 𝑀′ 𝑡 𝑎𝑡 𝑡 = 0                    Alternative solution 2: 𝐶′ 𝑡 𝑎𝑡 𝑡 = 0\n\n        𝑀 (𝑡)                                                                      𝑝e\n                                                          𝐶 (𝑡) = ln(𝑀 𝑡) = 𝑙𝑛          (1 − (1 − 𝑝)e )\n          𝑝e\n        =      (1 − (1 − 𝑝)e )\n          𝑝e . (1 − 𝑝)e                                      𝐶 (𝑡) = ln(𝑝) + 𝑡 − ln(1 − (1 − 𝑝)e )\n        +\n                         (1 − (1 − 𝑝)e )\n                                                                        (−1)(1 − 𝑝)e\n                                                          𝐶 (𝑡) = 1 −                    (1 − (1 − 𝑝)e )\n At t=0,\n                𝑝            𝑝. (1 − 𝑝)\n      𝑀 (𝑡) =        (𝑝) +                      = 1/𝑝     At t=0\n                                          (𝑝)\n                                                                                 1−𝑝\n                                                                   𝐶 (𝑡) = 1 +          𝑝 = 1/𝑝",
      "has_math": true,
      "session": "2019-11",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2019-11_QP.pdf",
      "source_sol": "raw/CS1A_2019-11_SOL.pdf"
    },
    {
      "q_num": 2,
      "marks": 9,
      "topic": "inference",
      "subtopics": [],
      "stem": "",
      "parts": [
        {
          "label": "i",
          "marks": 4,
          "text": "Define the variable 𝑡𝑘 used in the t-test for sampling distribution of sample mean\n           describing all the symbols used.                                                                (2)\n\n        ii) State the mean and variance of 𝑡𝑘 for k>2.                                                     (1)\n\n        iii) A sample of 10 numbers from normal population has sample mean and sample variance\n             as 50 and 48.667 respectively.\n\n        Determine the confidence interval for the population mean at 99% confidence level-\n\n             a) Using the t-test tables                                                                    (2)\n             b) Assuming a Normal distribution with parameters as the results of part (ii) above           (4)",
          "topic": null
        }
      ],
      "solution": "( , )\n  i)     Variable 𝑡 is defined as 𝑡 ⇛                                                                       [1]\n                                                  ⁄\n\n         where k denotes the degrees of freedom                                                           [0.5]\n         and the two random variables 𝑁(0,1) and 𝜒 are independent.                                       [0.5]\n\n  ii)   Mean and variance of 𝑡 for k>2 are 0                                                              [0.5]\n        and k/(k-2) respectively.                                                                         [0.5]\n\n iii)\n\n         a) We know that for a sample from a normal population,\n\n           /√\n                ~𝑡                                                                                        [0.5]\n\n         For the given confidence level 𝑡 = 3.25,                                                         [0.5]\n\n                                                                                                  Page 2 of 9\n\fIAI                                                                                                                                                 CS1A-1119\n\n                                                                                                                .                           .\n         Confidence interval for 𝜇 is thus 50 − 3.25 ×                                                                      , 50 + 3.25 ×\n\n        i.e. (42.83 , 57.17)                                                                                                                                  [1]\n\n        b)\n\n         From (ii) above, we know that 𝑡                                            ~𝑁 0,                       ~𝑁(0 ,           )                            [1]\n\n         i.e.\n\n             /√\n                  ~𝑁(0 ,          )                                                                                                                         [0.5]\n\n                      ~𝑁(0 ,1)                                                                                                                              [0.5]\n         /        /\n\n        Critical value for given level of confidence is 2.58                                                                                                [0.5]\n\n                                                                                                            .                                   .\n        Confidence interval for 𝜇 is thus 50 − 2.58 ×                                                                   × , 50 + 2.58 ×             ×\n\n        i.e. (43.54 , 56.45)                                                                                                                                  [1]",
      "has_math": true,
      "session": "2019-11",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2019-11_QP.pdf",
      "source_sol": "raw/CS1A_2019-11_SOL.pdf"
    },
    {
      "q_num": 3,
      "marks": 10,
      "topic": "bayes_credibility",
      "subtopics": [],
      "stem": "The number of claims X, on an insurance policy over a year follows a Poison distribution\n        with unknown parameter 𝜃. The number of claims observed in the previous n years are\n        𝑥1 , 𝑥2 … 𝑥𝑛\n\n        Prior distribution for 𝜃 has gamma distribution with parameters 𝛼 and 𝜆, as defined in the\n        actuarial tables.",
      "parts": [
        {
          "label": "i",
          "marks": 4,
          "text": "Derive the posterior distribution of 𝜃 given 𝑥1 , 𝑥2 … 𝑥𝑛                                       (4)\n\n                                                                                   𝛼+∑ 𝑥",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 3,
          "text": "Show that the Bayesian estimate of 𝜃 under quadratic loss is equal to 𝜆+𝑛 𝑖                    (3)",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 3,
          "text": "Show that the mean of the posterior distribution can be written in the form\n             𝑍 ∗ (𝑠𝑎𝑚𝑝𝑙𝑒 𝑚𝑒𝑎𝑛) + (1 − 𝑍) ∗ (𝑚𝑒𝑎𝑛 𝑜𝑓 𝑝𝑟𝑖𝑜𝑟 𝑑𝑖𝑠𝑡𝑟𝑖𝑏𝑢𝑡𝑖𝑜𝑛), defining Z as\n             appropriate.                                                                                  (3)",
          "topic": null
        }
      ],
      "solution": "i)              The likelihood is given by\n                  𝐿(𝜃) =          !\n                                      𝑒               ×           !\n                                                                      𝑒     …×              !\n                                                                                                𝑒                                                             [1]\n                  ∝ 𝜃∑ . 𝑒                                                                                                                                  [0.5]\n                  The prior distribution 𝜃~𝐺𝑎𝑚𝑚𝑎(𝛼 , 𝜆) is given by\n                  𝑓(𝜃) =       ( )\n                                      𝜃(                  )\n                                                              𝑒            ∝ 𝜃(             )\n                                                                                                𝑒                                                             [1]\n                  The posterior distribution is given by:\n                  𝑝𝑜𝑠𝑡𝑒𝑟𝑖𝑜𝑟 ∝ 𝑝𝑟𝑖𝑜𝑟 × 𝑙𝑖𝑘𝑒𝑙𝑖ℎ𝑜𝑜𝑑                                                                                                            [0.5]\n                       (      )                           ∑\n                  =𝜃              𝑒               ×𝜃                  .𝑒                                                                                    [0.5]\n                       ∑      (           )           (           )\n                  =𝜃                          𝑒                                                                                                             [0.5]\n\n ii)              Bayesian estimate of 𝜃 under quadratic loss is the mean of the posterior distribution.\n                  The posterior distribution in part (i) is of the form\n                  𝐺𝑎𝑚𝑚𝑎(𝛼 + ∑ 𝑥 , 𝜆 + 𝑛)                                                                                                                      [1]\n                                                                                                    (           ∑       )\n                  The mean of above distribution is given by                                            (           )\n                              (       ∑           )\n                  i.e. ⏞\n                       𝜃=         (           )\n                                                                            (       ∑       )\n iii)             From part (ii) we have ⏞\n                                         𝜃=                                     (       )\n                       ( )      (∑            )\n                  =(         )\n                               +(             )\n                                                      ∑\n                  =                   +                                                                                                                       [1]\n                                                                                                                                                    Page 3 of 9\n\fIAI                                                                                                    CS1A-1119\n              = (𝑚𝑒𝑎𝑛 𝑜𝑓 𝑝𝑟𝑖𝑜𝑟 𝑑𝑖𝑠𝑡𝑟𝑖𝑏𝑢𝑡𝑖𝑜𝑛) × (1 − 𝑍) + (𝑠𝑎𝑚𝑝𝑙𝑒 𝑚𝑒𝑎𝑛) × (𝑍)\n              Where 𝑍 =       ℎ𝑒𝑛𝑐𝑒 (1 − 𝑍) =                                                                 [1]",
      "has_math": true,
      "session": "2019-11",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2019-11_QP.pdf",
      "source_sol": "raw/CS1A_2019-11_SOL.pdf"
    },
    {
      "q_num": 4,
      "marks": 6,
      "topic": "bayes_credibility",
      "subtopics": [],
      "stem": "The investment department of a life insurance company feels that the chances of the economy\n        moving into a high level of financial stress over the next month are 80%.\n\n        Based on an alternate study of macro-economic variables, it has been found that the relative\n        position of credit spreads vis-à-vis their long-term historical average at the beginning of a\n        month is a leading indicator of the level of financial stress in the economy over the following\n        month. The studies indicate that the economy may face high levels of stress when there is a\n        spike in credit spreads.\n\n        Based on data gathered for the past 10 years, it has been observed that:\n\n             75% of the times when the economy ends up in high level of financial stress, it is preceded\n              by high credit spreads; and\n             40% of the times when the economy ends up in a low level of financial stress, it is\n              preceded by high credit spread.\n\n        Compute the following:",
      "parts": [
        {
          "label": "i",
          "marks": 1,
          "text": "Prior probability of the economy moving into a high level of financial stress.                    (1)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 2,
          "text": "Find the conditional probability of the credit spreads being high in the beginning of the\n            month given that the level of financial stress was high over the following month.                (2)",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 3,
          "text": "Calculate the posterior probability that the financial stress index will be high given that\n             the credit spreads are high in the beginning of the month.                                      (3)",
          "topic": null
        }
      ],
      "solution": "i)         The prior probability is the probability assessed by the investment department, 80%             [1]\n  ii)         The conditional probability = P (CS = H| FS = H) = 0.75                                         [2]\n iii)         The posterior probability can be computed as:\n\n                  P (FS = H| CS = H) = P(FS=H Ո CS=H)𝑃(𝐹𝑆 = 𝐻|𝐶𝑆 = 𝐻) =\n                                  (     )∗ (     |     )                              . ∗ .\n                   (    )∗   𝐶𝑆 = 𝐻 𝐹𝑆 = 𝐿       (     )∗ (      |       )\n                                                                             =   . ∗ .     . ∗ .\n                                                                                                   = 0.88     [3]",
      "has_math": false,
      "session": "2019-11",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2019-11_QP.pdf",
      "source_sol": "raw/CS1A_2019-11_SOL.pdf"
    },
    {
      "q_num": 5,
      "marks": 8,
      "topic": "bayes_credibility",
      "subtopics": [],
      "stem": "You are given a coin. Let p be the probability of getting heads if the coin is flipped. You have\n        not been told if coin is fair or biased. To start with you assume that p can take any value\n        uniformly over the range 0 and 1. Next you decide to perform an experiment. You start\n        flipping the coin until you get the head for the first time.",
      "parts": [
        {
          "label": "i",
          "marks": 4,
          "text": "Derive the posterior distribution of p if you get the head for the first time after m flips.      (4)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 4,
          "text": "Derive the posterior distribution of p if you get the head for the second time after a further\n            n flips.                                                                                         (4)",
          "topic": null
        }
      ],
      "solution": "i)    Note: Alternate solution is possible depending on the assumption whether you get the head\n        on mth flip or m+1st flip.\n\n        Let X be the number of times the coin needs to be flipped before getting heads for the first\n        time. Then X|p is a Type 1 Geometric distribution with parameter p.                        [1]\n\n        The prior distribution of p can be assumed to be uniform over the interval [0, 1]. f prior (p) = 1,\n        0<= p<=1                                                                                     [0.5]\n\n        Likelihood function: L (p) = P(X = m) = (1-p)m-1.p                                                  [0.5]\n\n        Alternate:\n        Likelihood function: L (p) = P(X = m) = (1-p)m.p\n        Posterior distribution of p is proportional to f prior(p).L(p)                                      [0.5]\n        Therefore posterior distribution of p is proportional to (1-p) m-1.p                                [0.5]\n\n        Alternate:\n        Therefore posterior distribution of p is proportional to (1-p) m.p\n        Posterior distribution of p is Beta (2, m)                                                            [1]\n\n        Alternate:\n        Posterior distribution of p is Beta (2,m+1)                                                           [1]\n\n ii)    Again, X|p is a Type 1 Geometric distribution with parameter p\n        With an added observation the likelihood function changes to:\n        L (p) = (1-p) m+n-2.p2\n        Alternate: L (p) = (1-p) m+n.p2\n\n        The prior distribution of p can be assumed to be uniform over the interval [0, 1]. f prior (p) = 1,\n        0<= p<=1                                                                                     [0.5]\n\n        Posterior distribution of p is proportional to f prior(p).L(p)                                       [0.5]\n\n                                                                                                      Page 4 of 9\n\fIAI                                                                                         CS1A-1119\n        Therefore posterior distribution of p is proportional to (1-p) m+n-2.p2\n\n        Alternate: Therefore posterior distribution of p is proportional to (1-p) m+n.p2\n\n        Posterior distribution of p is Beta (3, m+n-1)\n\n        Alternate: Posterior distribution of p is Beta (3, m+n+1)                                      [1]",
      "has_math": false,
      "session": "2019-11",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2019-11_QP.pdf",
      "source_sol": "raw/CS1A_2019-11_SOL.pdf"
    },
    {
      "q_num": 6,
      "marks": 12,
      "topic": "inference",
      "subtopics": [
        "distributions"
      ],
      "stem": "A statistics student was observing the number of passengers sitting (other than the driver)\n        inside five-seat capacity cars entering the city centre in City A and City B respectively over\n        a particular period and summarised his observations as below:\n\n            Number of seats occupied apart from driver in the\n                                                                   0       1        2        3       4\n            observed car\n            Number of cars (City A)                                70     120      201       80     29\n            Number of cars (City B)                                40     100      160      170     70\n\n        It is proposed that a Binomial model with parameters n and p can be fitted for city A (where\n        p is the probability that a seat is occupied with 0 < p < 1).",
      "parts": [
        {
          "label": "i",
          "marks": 5,
          "text": "Determine the maximum likelihood estimate of p for City A.                                        (5)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 1,
          "text": "𝜃1 and 𝜃2 are the proportions of cars plying with less than 2 passengers in cities A and B\n            respectively.\n\n              a) Estimate 𝜃1 𝑎𝑛𝑑 𝜃2 using the sample data provided.                                          (2)\n\n               b) Calculate the symmetrical 95% confidence interval for the difference in proportion\n                  of cars plying with less than 2 passengers in both the cities.\n\n             c) Comment on your answer in part b.                                                          (1)",
          "topic": null
        }
      ],
      "solution": "i)    The likelihood function is given by\n\n          𝐿(𝑝) = 𝐶[(1 − 𝑝) ] . [𝑝(1 − 𝑝) ]           . [𝑝 (1 − 𝑝) ]   . [𝑝 (1 − 𝑝) ] . [(𝑝) ]\n\n        𝐿(𝑝) = 𝐶[𝑝             (1 − 𝑝)    ]                                                            [1]\n\n       Taking logs and differentiating w.r.t p:\n               d             d                                        878 1122\n                  ln 𝐿(𝑝) =    (ln(𝐶) + 878 ln(𝑝) + 1122 ln(1 − 𝑝)) =    −\n               dp           dp                                         𝑝   1−𝑝\n\n       Equating to zero\n\n              =          𝑜𝑟 878 = 𝑝(1122 + 878)𝑜𝑟 𝑝 = 0.439                                            [1]\n\n       Checking for maximum:\n\n              ln 𝐿(𝑝) = −           −(    )\n                                              < 0 ⇒ 𝑀𝑎𝑥𝑖𝑚𝑢𝑚                                            [1]\n\n ii)\n        a) 𝜃 =                            = 0.3800                                                     [1]\n\n              And 𝜃 =                          = 0.2593                                                [1]\n\n        b) For samples from independent binomial distributions, we know that\n\n                                          𝜃 −𝜃    − (𝜃 − 𝜃 )\n                                                                 ~𝑁(0,1)\n                                         𝜃 (1 − 𝜃 ) 𝜃 (1 − 𝜃 )\n                                             𝑛     +    𝑛\n\n        Substituting values of 𝑝 𝑎𝑛𝑑 𝑝 in the above equation,\n        ( .        ) (     )\n                               ~𝑁(0,1)                                                                 [1]\n               .\n\n        Critical value for given confidence level is 1.96.                                           [0.5]\n                                                                                            Page 5 of 9\n\fIAI                                                                                         CS1A-1119\n        Hence the confidence interval is calculated as:\n\n        (0.1207 − 1.96 × 0.02876 , 0.1207 + 1.96 × 0.02876 )                                       [0.5]\n\n        i.e. (0.0643, 0.1771)                                                                        [1]\n\n        c) As the interval does not contain zero, there is significant difference in the proportion of\n           cars plying with less than 2 passengers in both the cities.                               [1]",
      "has_math": true,
      "session": "2019-11",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2019-11_QP.pdf",
      "source_sol": "raw/CS1A_2019-11_SOL.pdf"
    },
    {
      "q_num": 7,
      "marks": 9,
      "topic": "inference",
      "subtopics": [],
      "stem": "Number of claims in a year on an insurance policy is believed to follow a Poisson distribution.\n        Claims on portfolio of 1000 such policies were observed for one year. It was suggested that\n        the value of Poisson parameter is 3. If the observed number of claims in that one year is less\n        than 3100 then the suggested value of 3 for the Poisson parameter is accepted else rejected.\n\n        You may use the result that probability distribution of summation of ‘n’ Poisson variables\n        with parameter 𝜇 is 𝑃𝑜𝑖(𝑛𝜇).",
      "parts": [
        {
          "label": "i",
          "marks": 4,
          "text": "Define Type I error and estimate it for the above case.                                         (4)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 1,
          "text": "Define Type II error.                                                                          (1)",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 2,
          "text": "Define power of a test and determine the power of test in terms of 𝜇 in above case.           (2)",
          "topic": null
        },
        {
          "label": "iv",
          "marks": 2,
          "text": "If the actual observed number of claims is 2900, determine the confidence interval for the\n            Poisson parameter at 99% confidence level.                                                     (2)",
          "topic": null
        }
      ],
      "solution": "i)    Type I error - Event of Rejecting the hypothesis when it is true                            [1]\n\n        Let X be the random variable denoting the total number of claims on the portfolio. X thus\n        follows Poi(nµ) i.e. Poi(3000) where µ is the Poisson parameter.\n\n        Null hypothesis 𝐻 is thus 𝑋~𝑃𝑜𝑖(3000)                                                      [0.5]\n\n        𝑃(𝑟𝑒𝑗𝑒𝑐𝑡 𝐻 𝑤ℎ𝑒𝑛 𝐻 𝑖𝑠 𝑡𝑟𝑢𝑒) = 𝑃(𝑋 > 3100 𝑤ℎ𝑒𝑛 𝑋~𝑃𝑜𝑖(3000))                                    [1]\n\n        Using normal approximation (as nλ is large enough)                                         [0.5]\n\n                                                𝑋~𝑁(3000,3000)\n\n        𝑃(𝑋 > 3100) = 𝑃 𝑍 >                       = 1 − 𝑃(𝑍 < 1.825) = 3.39%                         [1]\n                                        √\n\n  ii)   Type II error - Event of Accepting the hypothesis when it is false                           [1]\n iii)   Power of a test - Probability of Rejecting the hypothesis when it is false                   [1]\n\n        In terms of 𝜇 it is given by:\n\n        𝑃(𝑟𝑒𝑗𝑒𝑐𝑡 𝐻 𝑤ℎ𝑒𝑛 𝐻 𝑖𝑠 𝑓𝑎𝑙𝑠𝑒) = 𝑃(𝑋 > 3100 𝑤ℎ𝑒𝑛 𝑋~𝑃𝑜𝑖(𝑛𝜇)~𝑁(𝑛𝜇, 𝑛𝜇))\n\n        𝑃(𝑋 > 3100) = 1 − 𝑍 <                                                                      [0.5]\n                                            √\n\n        The value of power of test will depend on the value of parameter under alternate hypothesis.\n iv)    If 𝜇̂ is the estimator of 𝜇 (the poisson parameter), 𝜇̂ follows 𝑁(𝜇, 𝜇̂ /𝑛)                [0.5]\n\n        Hence         follows N(0,1)                                                               [0.5]\n                  /\n\n                                .\n        Or 𝑃 −2.5758 <                  < 2.5758 = 0.99                                            [0.5]\n                                . /\n\n        Hence the confidence interval for 𝜇 is (2.7613, 3.0387)                                    [0.5]\n\n                                                                                           Page 6 of 9\n\fIAI                                                                                          CS1A-1119",
      "has_math": true,
      "session": "2019-11",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2019-11_QP.pdf",
      "source_sol": "raw/CS1A_2019-11_SOL.pdf"
    },
    {
      "q_num": 8,
      "marks": 10,
      "topic": "inference",
      "subtopics": [],
      "stem": "There are 2 fair dice coloured yellow and black respectively which are rolled simultaneously.\n        Two random variables are defined as follows:\n\n            X: Number rolled on the yellow dice\n            Y: Number of times ‘3’ appears when both dice are rolled once together",
      "parts": [
        {
          "label": "i",
          "marks": 4,
          "text": "Determine the joint probability mass function (X, Y) by filling up the following table:\n\n                    X=1     X=2        X=3       X=4      X=5     X=6\n            Y=0\n            Y=1\n            Y=2                                                                                            (4)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 6,
          "text": "Compute the value var(X+Y)                                                                     (6)\n\n                                21           1            91            14\n        You are given 𝐸(𝑋) = 6 , 𝐸(𝑌) = 3 , (𝑋 2 ) = 6 , 𝐸(𝑌 2 ) = 36",
          "topic": null
        }
      ],
      "solution": "i)   X and Y can take whole values in the sample space (1,6) and (0,2) respectively. The joint mass\n       function of X and Y is as follows:\n\n                  X=1            X=2             X=3           X=4           X=5            X=6\n         Y=0      (1/6).(5/6)    (1/6).(5/6)     (1/6).0       (1/6).(5/6)   (1/6).(5/6)    (1/6).(5/6)\n         Y=1      (1/6).(1/6)    (1/6).(1/6)     (1/6).(5/6)   (1/6).(1/6)   (1/6).(1/6)    (1/6).(1/6)\n         Y=2      (1/6).0        (1/6).0         (1/6).(1/6)   (1/6).0       (1/6).0        (1/6).0\n        OR\n\n                 X=1             X=2            X=3            X=4           X=5            X=6\n         Y=0     5/36            5/36           0              5/36          5/36           5/36\n         Y=1     1/36            1/36           5/36           1/36          1/36           1/36\n         Y=2     0               0              1/36           0             0              0\n\n                                        (0.25 marks for each non zero values,1 marks for all zero values)\n\n ii)   𝑉𝑎𝑟(𝑋 + 𝑌) = 𝑣𝑎𝑟(𝑋) + 𝑣𝑎𝑟(𝑌) + 2𝑐𝑜𝑣(𝑋 + 𝑌)\n\n       𝑉𝑎𝑟(𝑋 + 𝑌) = 𝐸(𝑋 ) − 𝐸(𝑋) + 𝐸(𝑌 ) − 𝐸(𝑌) + 2 𝐸(𝑋𝑌) − 𝐸(𝑋)𝐸(𝑌)                                  [1]\n\n       𝐸(𝑋𝑌) is calculated by taking sum of all the values in following table. Each entry is\n       𝑥. 𝑦. 𝑃(𝑋 = 𝑥). 𝑃(𝑌 = 𝑦):\n\n                   X=1            X=2             X=3           X=4           X=5           X=6\n          Y=0      0              0               0             0             0             0\n          Y=1      1/36           2/36            15/36         4/36          5/36          6/36\n          Y=2      0              0               6/36          0             0             0\n\n       Hence 𝐸(𝑋𝑌) =                                                                                  [4]\n\n       𝑉𝑎𝑟(𝑋 + 𝑌) =       −         +      −      +2      −    .   = 3.027                            [1]",
      "has_math": true,
      "session": "2019-11",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2019-11_QP.pdf",
      "source_sol": "raw/CS1A_2019-11_SOL.pdf"
    },
    {
      "q_num": 9,
      "marks": 10,
      "topic": "regression_glm",
      "subtopics": [],
      "stem": "",
      "parts": [
        {
          "label": "i",
          "marks": 4,
          "text": "Consider a response variable Y which is related to the value of x by the following equation\n\n                                    𝑦𝑖 = 𝛼 + 𝛽. 𝑥𝑖 + 𝜀𝑖         𝑖 = 1,2,3 … , 𝑛\n\n        Here, 𝜀𝑖 are independent and identically distributed standard normal variables.\n\n        State the function which is to be minimized to arrive at the least squares estimates for 𝛼 & 𝛽.    (2)\n\n        ii) You have been given a sample set of size n of paired data (x, y) and it has been suggested\n            to use a model of the form:\n\n                                                   𝐸(𝑌𝑖 ) = 𝛾. 𝑒 𝑥𝑖\n\n        Derive the least squares estimate of 𝛾.                                                                (4)\n\n       iii) A weighted least squares regression is a regression where to arrive at the estimates of the\n            regression coefficients, the weighted sum of squared errors are minimized instead of a\n            simple sum of squared errors.\n\n        Now using the model prescribed in part (ii), derive the least squares estimate of 𝛾 by assigning\n        a weight to each of the error terms where the error term is inversely proportional to the value\n        of the independent variable 𝑥𝑖 . For instance, the weight applied to ith squared error term will\n            1\n        be 𝑥                                                                                                   (4)\n            𝑖",
          "topic": null
        }
      ],
      "solution": "i)   The least square estimates of the regression coefficients are the values of 𝛼 & 𝛽 for which:\n\n       𝑞= ∑        𝑒                                                                                  [1]\n        =∑        [𝑦 − (𝛼 + 𝛽. 𝑥 )]                                                                   [1]\n\n              is a minimum\n\n ii)   The least squares estimate of the regression coefficient 𝛾 is the value of 𝛾 for which:\n\n       𝑞= ∑       𝑒     = ∑     [𝑦 − 𝛾. 𝑒 ]                                                         [0.5]\n\n       is a minimum.\n\n                                                                                             Page 7 of 9\n\fIAI                                                                                        CS1A-1119\n\n        ∑    𝑒   = ∑        (𝑦 − 2. 𝛾𝑒 . 𝑦 + 𝛾 𝑒                        )                           [1]\n\n        In order to find the minimum value, differentiate with respect to 𝛾 and set it equal to zero:\n\n            = 2𝛾 ∑     𝑒        − 2. ∑                𝑦 .𝑒         =0                               [1]\n\n                            ∑      .\n        Therefore, 𝛾 = ∑                                                                            [1]\n\n             = 2∑       𝑒        > 0, 𝑡ℎ𝑒𝑟𝑒𝑓𝑜𝑟𝑒 𝑚𝑖𝑛𝑖𝑚𝑢𝑚                                           [0.5]\n\n iii)   The least squares estimate of the regression coefficient 𝛾 is the value of 𝛾 for which:\n\n                                                  [        .   ]\n        𝑞= ∑         𝑤 ∗𝑒       = ∑                                                                 [1]\n\n        is a minimum.\n\n                                           .      .\n        ∑    𝑒   = ∑                                                                                [1]\n\n        In order to find the minimum value, differentiate with respect to 𝛾 and set it equal to zero:\n\n                                                       .\n            = 2𝛾 ∑              − 2. ∑                         =0                                   [1]\n\n                            ∑          .\n        Therefore, 𝛾 =                                                                            [0.5]\n                            ∑\n\n             = 2∑       𝑒                                                                         [0.5]\n                                 𝑥 > 0, 𝑡ℎ𝑒𝑟𝑒𝑓𝑜𝑟𝑒 𝑚𝑖𝑛𝑖𝑚𝑢𝑚",
      "has_math": true,
      "session": "2019-11",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2019-11_QP.pdf",
      "source_sol": "raw/CS1A_2019-11_SOL.pdf"
    },
    {
      "q_num": 10,
      "marks": 20,
      "topic": "inference",
      "subtopics": [
        "distributions",
        "regression_glm"
      ],
      "stem": "You have been given the data set comprising of mortality rating and premium rates.\n\n         Mortality\n         Rating       25     50     75     100         135     180     215     245    300    365      400\n         (%)\n         Premium\n                     3.50   5.80   8.07    10.33      13.44    17.38   20.39   22.94 27.53   32.83    35.62\n         Rate\n\n        It has been specified that the premium rates can be expressed as a cubic function of mortality\n        rating. You have been given the task of deriving a simple formula for calculation of premium\n        rates at different mortality ratings.\n\n        Suppose, moving ahead in line with the proposed methodology, you decide to fit a simple\n        linear regression with premium rates being regressed on cubic function of mortality rating.\n        The transformed data is as below:\n\n         (Mortality\n         Rating)^3 0.02     0.13    0.42    1.00       2.46     5.83   9.94    14.71 27.00 48.63 64.00\n         : (X)\n         Premium\n                    3.50    5.80    8.07   10.33       13.44 17.38 20.39 22.94 27.53 32.83 35.62\n         Rate: (Y)\n\n        You are given the following summary statistics:\n\n        ∑ 𝑥 = 174.13, ∑ 𝑦 = 197.84,\n\n        ∑ 𝑥 2 = 7545.90, ∑ 𝑦 2 = 4747.45,\n\n        ∑(𝑥 − 𝑥̅ )(𝑦 − 𝑦̅) = 2176.84",
      "parts": [
        {
          "label": "i",
          "marks": 3,
          "text": "Derive the linear regression equation of premium rates Y on X.                                    (3)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 6,
          "text": "Perform a statistical test to investigate the hypothesis that there is no linear relationship\n          between X and Y.\n\n      State clearly all assumptions made.                                                                  (6)",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 2,
          "text": "Calculate the sample correlation coefficient.                                                   (2)",
          "topic": null
        },
        {
          "label": "iv",
          "marks": 7,
          "text": "Calculate the 95% confidence interval for the individual and mean responses\n          corresponding to x^3 =25.                                                                        (7)",
          "topic": null
        },
        {
          "label": "v",
          "marks": 2,
          "text": "Consider the residual plot of the fitted regression:\n\n                     1000\n\n                      500\n         Residuals\n\n                        0\n\n                      -500\n\n                     -1000\n                                                X\n\n      Comment on the fit of the model and any drawbacks of using the model rather than the full\n      premium rate table.                                                                                  (2)",
          "topic": null
        }
      ],
      "solution": "(∑ )\n  i)    𝑆    = ∑𝑥 −              = 4789.42                                                        [0.5]\n\n        𝑆    = 2176.84                                                                            [0.5]\n\n                        .\n        𝛽=       =      .\n                                = 0.4545                                                            [1]\n\n        𝛼 = 𝑦 − 𝛽 𝑥̅ = 17.895 − 0.4545 ∗ 15.83 = 10.79                                            [0.5]\n\n        The fitted regression equation is 𝑦 = 10.79 + 0.4545 ∗ 𝑥                                  [0.5]\n\n ii)    𝑆    = 4789.42                         from result of part i\n\n                        (∑ )\n        𝑆    = ∑𝑦 −              = 1189.21                                                        [0.5]\n\n        𝑆    = 2176.84                         from result of part i\n                                                                                           Page 8 of 9\n\fIAI                                                                                                 CS1A-1119\n                                                                                  .\n        𝜎 =                𝑆        −           = ∗ 1189.21 −                         .\n                                                                                          = 22.20           [2]\n\n                                                     .\n        𝑠. 𝑒. 𝛽 =                       =                    .\n                                                                     = 0.0681                               [1]\n\n                                    To test 𝐻 : 𝛽 = 0 𝑣 𝐻 : 𝛽 ≠ 0, 𝑡ℎ𝑒 𝑡𝑒𝑠𝑡 𝑠𝑡𝑎𝑡𝑖𝑠𝑡𝑖𝑐 𝑖𝑠\n\n                       .\n        . .\n                  =    .\n                                    = 6.674                                                                 [1]\n\n        Under the assumption that the errors of the regression are i.i.d 𝑁(0, 𝜎 ) random variables,\n        beta has a t distribution with n-2 degrees of freedom.                                [0.5]\n\n        Critical value for t-distribution with 9 degrees of freedom: t9,0.025 = 2.68.                     [0.5]\n\n        Since the critical value at 95% level of significance is less than the test statistic, there is\n        sufficient evidence to reject the null hypothesis. Hence, it cannot be concluded that there is\n        no statistically significant relationship between x and y.                                [0.5]\n\n iii)   Pearson’s correlation coefficient is computed as:                                                 [0.5]\n\n        S yy = 1189.21\n\n                                                 .\n        𝑟=                     =                                      = 0.91                              [1.5]\n                                    √       .   ∗                .\n\n iv)    The estimated value of y corresponding to 𝑥 = 25 𝑖𝑠 10.79 + 0.4545 ∗ 25 = 22.15                     [1]\n\n        𝜎 = 22.20                       … from earlier workings\n\n        The variance of the estimator of the mean response is given by\n              (       ̅)                                 .\n          +                𝜎 =                  +                .\n                                                                      ∗ 22.20 = 2.41                        [2]\n\n        The variance of the estimator of the individual response is given by\n\n                      (        ̅)\n        1+ +                         𝜎 = [1 + 0.1085] ∗ 22.20 = 24.61                                       [2]\n\n        Using t9 distribution, the 95% confidence intervals for mean and individual responses are:\n\n        22.20 ± 2.262 ∗sqrt(2.41) and 22.20 ± 2.262 ∗sqrt(24.61) = (18.7,25.7) & (11.0,33.4)                [2]\n\n  v)    The residual plot shows a definite pattern. Although the correlation coefficient is high, the\n        model does not seem to be appropriate.                                                     [1]\n        Using this model leads to underestimation of premium rates at low and high mortality ratings.\n                                                                       ***************\n                                                                                                    Page 9 of 9",
      "has_math": true,
      "session": "2019-11",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2019-11_QP.pdf",
      "source_sol": "raw/CS1A_2019-11_SOL.pdf"
    },
    {
      "q_num": 1,
      "marks": 8,
      "topic": "bayes_credibility",
      "subtopics": [
        "distributions"
      ],
      "stem": "The marks of a professional exam are normally distributed with unknown mean μ and\n         standard deviation σ = 20.\n\n         Prior belief is that μ follows Normal Distribution with mean 65 and standard deviation 17.",
      "parts": [
        {
          "label": "i",
          "marks": 2,
          "text": "Calculate the prior probability that μ is greater than 60                                (2)\n\n         Basis the sample of 150 scripts, mean marks are found to be 63.",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 2,
          "text": "Calculate the posterior probability that μ is greater than 60                              (4)\n\n         iii) Comment on your calculations from parts (i) and (ii).                                      (2)",
          "topic": null
        }
      ],
      "solution": "i)        P(μ>60) = P(z> (60-65)/17) = P(z>-0.2941)                                                                [1]\n                                                   = 61.57%                                                            [1]\n\n    ii)       From tables page 28, we know μ/X follows Normal (μ*, σ*^2)\n\n              μ* = (150*63/20^2 + 65/17^2)/(150/20^2+1/17^2)                                                           [1]\n                  = 63.02\n\n              σ*^2 = 1/(150/400+1/289) = 2.6423                                                                        [1]\n\n              So P(μ>60) = P(z> (60-63.02)/sqrt(2.6423)) = P(z>-1.86)                                                  [1]\n                                                                   = 96.86%                                            [1]\n\n    iii)      The probability has increased as there is higher certainty over the value of μ as a result of considering a\n              fairly large sample.\n\n              This is despite the fact that the mean belief about μ has fallen, which apriori might make a lower value\n              of μ more likely\n\n              The posterior distribution has thinner tails and hence lower volatility since the credibility increases\n              around the mean.                                                                                    [2]",
      "has_math": true,
      "session": "2020-11",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2020-11_QP.pdf",
      "source_sol": "raw/CS1A_2020-11_SOL.pdf"
    },
    {
      "q_num": 2,
      "marks": 7,
      "topic": "distributions",
      "subtopics": [],
      "stem": "",
      "parts": [
        {
          "label": "i",
          "marks": 2,
          "text": "Given a Weibull Distribution with probability distribution function: 𝐹 (𝑥) = 1 − e ( )\n               determine the equation that needs to be solved to simulate a random variable using\n               random number u from a U(0,1) interval.\n\n                a) 𝑥 = 1 − e ( )\n\n                b) 𝑥 =      −𝑙𝑛(1 − 𝑢)\n\n                c) 𝑥 = −𝑙𝑛 (√1 − 𝑢)\n\n                d) 𝑥 =      −ln(1 − 𝑢 )                                                                  (2)\n\n        ii)    Simulate a value at u=0.75 based on your answer in part (i)                               (1)\n\n        iii) Given a probability density function 𝑓(𝑥) = (1 + x)              for x > 0 determine the\n             corresponding probability distribution function\n\n                a) 𝐹 (𝑥) = 1 − (1 + 𝑥)\n                b) 𝐹 (𝑥) = (1 + 𝑥)\n                c) 𝐹 (𝑥) = 1 − (1 + 𝑥 + 𝑥 )\n                d) 𝐹(𝑥) = 1 − (1 + x)                                                                    (2)\n\n        iv) Hence simulate a random variable from the distribution function as per your answer in\n            part (iii) using random variable u=0.5 from U(0,1)                                           (2)",
          "topic": null
        }
      ],
      "solution": "i)        Let ‘u’ be the simulated value from the Uniform Distribution. Solving for ‘x’ in terms of ‘u’:\n\n                                                      𝑢 = 1−e ( )\n\n                  −(x) = ln(1 − 𝑢)\n\n                  x=     −ln(1 − 𝑢)\n\n                  Hence correct answer is ‘b’                                                                          [2]\n\n    ii)       for u= 0.75, x =   −ln(1 − 0.75) = 1.177                                                                 [1]\n\n    iii)      Deriving the probability distribution function:\n\n              𝐹(𝑥) = ∫ (1 + t) dt = [−(1 + t) ]\n\n              𝐹(𝑥) = 1 − (1 + x)\n\n              Correct answer is ‘a’                                                                                    [2]\n\n    iv)       Equating and Solving for ‘u’\n\n                                   𝑢 = 1 − (1 + x)\n\n           𝑥 = (1 − u)   −1                                                                                            [1]\n\n                                                                                                               Page 2 of 7\n\fIAI                                                                                                                   CS1A-1120\n             For u =0.5, 𝑥 = (1 − 0.5)        −1=1                                                                           [1]",
      "has_math": true,
      "session": "2020-11",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2020-11_QP.pdf",
      "source_sol": "raw/CS1A_2020-11_SOL.pdf"
    },
    {
      "q_num": 3,
      "marks": 6,
      "topic": "distributions",
      "subtopics": [
        "data_analysis"
      ],
      "stem": "Answer the following in case of a bivariate dataset with data points (𝑋 ) and (𝑌 ):",
      "parts": [
        {
          "label": "i",
          "marks": 1,
          "text": "Identify the correct definition of the Spearman rank correlation coefficient\n\n                                            ∑\n               a) 𝑟 = 1 −\n                                        (               )\n\n                                            ∑\n               b) 𝑟 = 1 −\n                                             (          )\n\n                                            ∑\n               c) 𝑟 = 1 −\n                                        (           )\n\n                                            ∑\n               d) 𝑟 = 1 −                                                                                        (1)\n                                        (               )",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 1,
          "text": "Define the symbols in your answer for part (i) and the condition that is necessary so that\n              the formula can be used to calculate the respective coefficients.                                  (1)",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 1,
          "text": "Define concordant pair of observation                                                               (1)",
          "topic": null
        },
        {
          "label": "iv",
          "marks": 1,
          "text": "Identify the correct definition of the Kendall rank correlation coefficient\n\n               a) 𝜏 =\n                            (           )\n\n               b) 𝜏 =\n                            (           )/\n\n               c) 𝜏 =\n\n                        (                   )/\n               d) 𝜏 =                                                                                            (1)\n                                (           )",
          "topic": null
        },
        {
          "label": "v",
          "marks": 2,
          "text": "Define the symbols in your answer for part (iv) and the condition that is necessary so\n              that the formula can be used to calculate the respective coefficients.                             (2)",
          "topic": null
        }
      ],
      "solution": "∑\n      i)        Spearman rank correlation coefficient 𝑟 is given by 𝑟 = 1 −       (       )\n                                                                                              correct answer is ‘d’           [1]\n\n      ii)       𝑑 = 𝑟𝑎𝑛𝑘 𝑜𝑓 (𝑋 ) − 𝑟𝑎𝑛𝑘 𝑜𝑓 (𝑌 )                                                                            [0.5]\n                Condition necessary: there should not be any ‘ties’ in the rank of variables                               [0.5]\n\n      iii)      A pair of observation (𝑋 , 𝑌 ); 𝑋 , 𝑌 𝑤ℎ𝑒𝑟𝑒 𝑖 ≠ 𝑗 is said to concordant if\n                  (𝑋 > 𝑋 ) 𝑎𝑛𝑑 (𝑌 > 𝑌 ) 𝑜𝑟 (𝑋 < 𝑋 ) 𝑎𝑛𝑑 (𝑌 < 𝑌 )                                                              [1]\n\n      iv)       Kendall rank correlation coefficient 𝜏 is given by 𝜏 =   (   )/\n                                                                                  correct answer is ‘b’                       [1]\n\n      v)        Where 𝑛 is the number of concordant pair in the data,                                                      [0.5]\n                𝑛 is the number of discordant pair in the data and                                                         [0.5]\n                𝑛 is the number of observations                                                                            [0.5]\n                Condition necessary: there should not be any ‘ties’ in the rank of variables                               [0.5]",
      "has_math": true,
      "session": "2020-11",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2020-11_QP.pdf",
      "source_sol": "raw/CS1A_2020-11_SOL.pdf"
    },
    {
      "q_num": 4,
      "marks": 13,
      "topic": "distributions",
      "subtopics": [
        "bayes_credibility"
      ],
      "stem": "Let p denote the proportion of policies where claim is registered. Prior belief about p is that\n         it follows Beta distribution with parameters 𝛼 and 𝛽. Underwriters estimate mean μ and\n         standard deviation as σ of p.",
      "parts": [
        {
          "label": "i",
          "marks": 3,
          "text": "Select how 𝛼 and 𝛽 would be represented in terms of μ and σ                                        (3)\n\n                                    (           )                       ( (       )        )(       )\n               a) 𝛼 =                                        and 𝛽 =\n                                (           )                         ( (     )       )(        )\n               b) 𝛼 =                                       and 𝛽 =\n                                (           )                         ( (     )       )(        )\n               c) 𝛼 =                                       and 𝛽 =\n                                    (           )                       ( (       )        )(       )\n               d) 𝛼 =                                        and 𝛽 =\n\n         A random sample of n policies is considered, and it is observed that d claims arise of them.",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 3,
          "text": "Derive the posterior distribution of p                                                             (3)",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 3,
          "text": "If the mean of posterior distribution is written in the form of a credibility estimate, select\n             the credibility factor, Z:\n\n                a) 𝑍 =\n                         (           )\n\n                b) 𝑍 = (             )\n                             (   )\n                c) 𝑍 = (             )\n\n                d) 𝑍 = (             )                                                                         (3)",
          "topic": null
        },
        {
          "label": "iv",
          "marks": 3,
          "text": "Select which of the below statement is true:\n\n                a) MLE of p is d/n ; Z increases with increasing σ\n                b) MLE of p is n/d ; Z increases with increasing σ\n                c) MLE of p is d/n ; Z decreases with increasing σ\n                d) MLE of p is n/d ; Z decreases with increasing σ                                             (3)",
          "topic": null
        },
        {
          "label": "v",
          "marks": 1,
          "text": "Comment on the chosen option regarding the dependency of credibility factor on σ\n               from part (iv)                                                                                  (1)",
          "topic": null
        }
      ],
      "solution": "i)        Option a                                                                                                      [3]\n\n      ii)       f(p/x) is proportional to f(x/p) * f(p)                                                                     [0.5]\n\n                = p^d * (1-p)^(n-d) * p^(alpha – 1) * (1-p)^(beta-1)\n\n                =p^(alpha+d-1) * (1-p)^(n-d+beta-1)                                                                         [1.5]\n\n                Which corresponds to the pdf of Beta distribution with parameters (alpha + d) and (beta + n - d)              [1]\n\n      iii)      Option d                                                                                                      [3]\n\n      iv)       Option a                                                                                                      [3]\n\n      v)        Higher σ means high volatility and less certainty in the prior estimate. Prior becomes less reliable and\n                leads to higher weight to observed data and hence higher credibility factor.                          [1]",
      "has_math": true,
      "session": "2020-11",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2020-11_QP.pdf",
      "source_sol": "raw/CS1A_2020-11_SOL.pdf"
    },
    {
      "q_num": 5,
      "marks": 13,
      "topic": "inference",
      "subtopics": [
        "distributions",
        "data_analysis"
      ],
      "stem": "A general insurance company is studying a portfolio of policies on which the company has\n         received more than 1 claims. The analyst has generated the following frequency table:\n\n              Number of claims           Frequency\n                    2                       230\n                    3                        54\n                   ≥4                         6\n\n         It is believed that the claims follow a Poisson distribution with parameter 𝜆 (for ease of\n         typing, you may write the parameter as ‘L’ in your solution)",
      "parts": [
        {
          "label": "i",
          "marks": 4,
          "text": "Show that the truncated probability function is given by:\n                                                            𝜆 𝑒\n                                       𝑃 (𝑋 = 𝑥 ) =\n                                                     𝑥! (1 − 𝑒 − 𝜆𝑒 )                                          (4)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 5,
          "text": "Prove that the relation between sample mean 𝑥̅ and 𝜆 is given by\n\n                                                            𝜆 (𝑛 − 1 ) 1 − 𝑒\n                                                     𝑥̅ =\n                                                            𝑛(1 − 𝑒 − 𝜆𝑒       )\n\n               (Hint: You can use method of maximum likelihood to find the answer)                             (5)",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 4,
          "text": "Using interpolation and guess value of 𝜆 as 0.6 and 0.7, calculate the estimate of 𝜆 upto\n             3 decimal places.                                                                                    (4)",
          "topic": null
        }
      ],
      "solution": "i)        Since only certain part of experience is given, we should use truncated Poisson distribution:\n                (one can use L to mean λ)\n                                      (   )\n                𝑃(𝑋 = 𝑥) = 𝑘                                                                                                  [1]\n                                     !\n                𝑤ℎ𝑒𝑟𝑒 𝑥 = 2,3,4, …\n                𝑎𝑛𝑑 𝑘 𝑖𝑠 𝑐𝑜𝑛𝑠𝑡𝑎𝑛𝑡 𝑠𝑢𝑐ℎ 𝑡ℎ𝑎𝑡 𝑡ℎ𝑒 𝑠𝑢𝑚 𝑜𝑓 𝑝𝑟𝑜𝑏𝑎𝑏𝑖𝑙𝑖𝑡𝑖𝑒𝑠 = 1                                                   [0.5]\n                We also have that 𝑃(𝑋 ≥ 2) = 1 − 𝑃(𝑋 = 0) − 𝑃(𝑋 = 1)\n                𝑃(𝑋 ≥ 2) = 1 − exp(−𝐿) − 𝐿. exp(−𝐿)                                                                           [1]\n                Since 𝑘 ∑   𝑃(𝑋 = 𝑥) = 1 ⇒ 𝑘(1 − exp(−𝐿) − 𝐿. exp(−𝐿)) = 1                                                    [1]\n                Or 𝑘 = 1/(1 − exp(−𝐿) − 𝐿. exp(−𝐿))                                                                         [0.5]\n\n                                                                                                                      Page 3 of 7\n\fIAI                                                                                                                                                                     CS1A-1120\n                       Hence the function is given by:\n                                                              (           )\n                       𝑃(𝑋 = 𝑥) =                            !\n                                                                              .(           (       )   .        (   ))\n\n        ii)            The maximum likelihood function is given by:\n                                                                                                 𝜆 𝑒                                              𝜆∑ 𝑒 (    )\n                                                         𝐿(𝜆) =                                                                = 𝑐𝑜𝑛𝑠𝑡𝑎𝑛𝑡 ×\n                                                                                          𝑥 ! (1 − 𝑒 − 𝜆𝑒                 )                   (1 − 𝑒 − 𝜆𝑒   )(    )\n\n                                             ⇒ log 𝐿(𝜆) = 𝑐𝑜𝑛𝑠𝑡𝑎𝑛𝑡 +                                                𝑥 log(𝜆) − 𝜆(𝑛 − 1) − (𝑛 − 1)log 1 − 𝑒       − 𝜆𝑒\n                       Differentiating wrt 𝜆\n                                                         ∑                                         (       )(                   )\n                           log 𝐿(𝜆) =                                − (𝑛 − 1) −                                                                                               [1]\n                                                                      ̅\n                       ⇒            log 𝐿(𝜆) =                            − (𝑛 − 1)(1 − 𝑒                           )/ 1 − 𝑒        − 𝜆𝑒                                       [1]\n                       Equating to zero,\n                                        ̅\n                       0=                   − (𝑛 − 1)(1 − 𝑒                           )/ 1 − 𝑒                  − 𝜆𝑒                                                         [0.5]\n                                ̅            (       )                                         (       )\n                       ⇒            =                                         ⇒ 𝑥̅ =                                                                                         [0.5]\n\n                                                                                           ×           ×        ×\n        iii)           From the given data, 𝑥̅ =                                                                     = 2.2275                                                   [1]\n                       With 𝜆 as 0.6, RHS of result in part (ii) is 2.2131                                                                                                      [1]\n                       With 𝜆 as 0.7, RHS of result in part (ii) is 2.2539                                                                                                      [1]\n\n                       Using linear interpolation,\n                                                 .               .\n                       𝜆 = 0.6 + .                               .\n                                                                                  (0.7 − 0.6) = 0.635                                                                           [1]",
      "has_math": true,
      "session": "2020-11",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2020-11_QP.pdf",
      "source_sol": "raw/CS1A_2020-11_SOL.pdf"
    },
    {
      "q_num": 6,
      "marks": 8,
      "topic": "inference",
      "subtopics": [],
      "stem": "The country of Actuaria has been affected by some novel virus species. The health\n         department has compiled the following data as regards the travel history and symptoms seen\n         in some of the people of Actuaria that were tested for the virus:\n\n                                    No Symptoms      Some Symptoms          Confirmed case\n                                        (NS)              (SS)                   (CC)\n              Travelled to place\n              outside Actuaria          86                   25                   80\n              (T)\n              No travel in recent\n              past                      130                 148                   31\n              (NT)\n\n         Test at 1% level if there is any association between travel history and the symptoms\n         indicating presence of virus in the body.                                                                [8]",
      "parts": [],
      "solution": "We need to perform 𝜒 contingency test with:\n\n𝐻 : There is no association between travel history and symptoms vs\n\n𝐻 : There is some association between travel history and symptoms                                                                                                               [1]\n\nTotal number of observations are 500.                                                                                                                                           [1]\n\nThe expected frequencies are:\n\n                                                                                           No Symptoms                              Some Symptoms       Confirmed case\n                                                                                           (NS)                                     (SS)                (CC)\n    Travelled to place outside Actuaria\n                                                                                           82.51                                    66.09               42.40\n    (T)\n    No travel in recent past\n                                                                                           133.49                                   106.91             68.60\n    (NT)\n\nThe test statistic is\n    (          )       (            .       )                (                    .   )\n∑                  =                             +⋯+                                       = 95.74                                                                              [2]\n                            .                                                 .\n\n                                                                                                                                                                        Page 4 of 7\n\fIAI                                                                                                       CS1A-1120\nThis is much higher than the critical value of 9.21 at 1% confidence. Hence there is enough evidence to conclude that\nthere is some association between the travel history and the virus symptoms.                                      [1]",
      "has_math": false,
      "session": "2020-11",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2020-11_QP.pdf",
      "source_sol": "raw/CS1A_2020-11_SOL.pdf"
    },
    {
      "q_num": 7,
      "marks": 9,
      "topic": "distributions",
      "subtopics": [],
      "stem": "",
      "parts": [
        {
          "label": "i",
          "marks": 2,
          "text": "Variance of random variable X in terms of probability generating function is given as:\n\n                 a) 𝑣𝑎𝑟 (𝑋) = 𝐺 (0) + 𝐺 (0) − 𝐺 (0)\n\n                 b) 𝑣𝑎𝑟 (𝑋) = 𝐺 (1) − 𝐺 (1) + 𝐺 (1)\n\n                 c) 𝑣𝑎𝑟 (𝑋) = 𝐺 (1) + 𝐺 (1) − 𝐺 (1)\n\n                 d) 𝑣𝑎𝑟 (𝑋) = 𝐺 (0) − 𝐺 (0) + 𝐺 (0)                                                               (1)\n\n        ii)     The infinite series ∶ 1 + 𝑡𝐸 (𝑋) +        𝐸 (𝑋 ) +       𝐸 (𝑋 ) + ⋯is expanded form of…\n                                                      !              !\n\n                 a) Probability Generating Function\n                 b) Moment Generating Function\n                 c) Cumulant Generating Function\n                 d) None of the above                                                                             (1)\n\n         A person is observing time taken for a bus to appear at the bus stop. It is believed that the\n         time period (denoted by random variable X) between the time at which two consecutive\n         buses arrive at the stop follows an exponential distribution with mean 𝜃.\n\n        iii) The following equation on further evaluation yields the moment generation function of\n             X.\n\n                a) 𝐸 (𝑒 ) =           − 𝑡 ∫ 𝑒𝑥𝑝(−𝑧) . 𝑑𝑧\n\n                b) 𝐸 (𝑒 ) = ∫ 𝑒𝑥𝑝(−𝑧) . 𝑑𝑧\n\n                c) 𝐸(𝑒 ) =        −𝑡       ∫ 𝑒𝑥𝑝(−𝑧) . 𝑑𝑧\n\n                d) 𝐸 (𝑒 ) =           −𝑡    ∫ 𝑒𝑥𝑝(−𝑧) . 𝑑𝑧                                                 (2)\n\n        iv) If random variable Y denotes the total observation time to count N buses, determine the\n            moment generating function of Y.                                                               (3)\n\n        v)     Hence identify the distribution followed by Y.                                              (2)",
          "topic": null
        }
      ],
      "solution": "i)        Correct answer is ‘c’                                                                               [1]\n    ii)       Correct answer is ‘b’                                                                               [1]\n\n    iii)      𝐸(𝑒 ) =          −𝑡      ∫ exp(−𝑧) . 𝑑𝑧\n              𝐸(𝑒 ) = (1 − 𝑡𝜃)\n              Correct answer is ‘d’                                                                               [2]\n\n    iv)       The total time 𝑦 to count N buses will be 𝑦 = 𝑥 + 𝑥 … + 𝑥                                           [1]\n              MGF of Y is given by 𝐸(𝑒 ) = 𝐸 𝑒 ∑ = ∏ 𝑒                                                            [1]\n              𝑀 (𝑡) = (1 − 𝑡𝜃)                                                                                    [1]\n\n    v)        The 𝑀 (𝑡) is of the form of MGF for 𝐺𝑎𝑚𝑚𝑎 (𝛼, 𝜆).\n              Hence the distribution in our case is 𝐺𝑎𝑚𝑚𝑎 𝑁,",
      "has_math": true,
      "session": "2020-11",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2020-11_QP.pdf",
      "source_sol": "raw/CS1A_2020-11_SOL.pdf"
    },
    {
      "q_num": 8,
      "marks": 16,
      "topic": "distributions",
      "subtopics": [
        "inference"
      ],
      "stem": "The theme park management wants to study the relationship between the number of visitors\n         and temperature of the day. In the sample collated for 25 low temperature weeks and 35\n         high temperature weeks over the span of 15 months, xi denotes the number of visitors in\n         week i. Given below is the summary of number of visitors:\n\n              Type of week      𝑋           𝑋\n\n               Cold week     10,250     4,718,190\n              Warm week      12,001     4,740,201",
      "parts": [
        {
          "label": "i",
          "marks": 5,
          "text": "Perform a test with the null hypothesis that the variance of number of visitors during\n               cold weeks is equal to the variance of number of visitors in warm weeks against the\n               alternative that the variance is higher in cold week when compared to warm week.\n\n               Significance level of 5% to be used.                                                        (5)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 4,
          "text": "Perform a test to check if there is any significant difference in the mean of the number\n               of visitors to theme park during cold weeks and warm weeks.\n\n               Significance level of 5% to be used.                                                        (4)\n\n         To obtain further insight into the relationship between temperature and number of visitors,\n         the management decides to carry out a linear regression analysis.\n\n         Let T denote the average temperature of the week and X denote the total number of weekly\n         visitors. The model suggested is E(X) = AT+B. Given below is the summary for temperature\n         ti and visitors xi for 78 weeks:\n\n         ∑ 𝑡 : 1,603\n         ∑ 𝑡 ∗ 𝑥 : 582,205\n         ∑ 𝑡 : 51,823",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 5,
          "text": "Estimate the correlation co-efficient ρ(T,X) between the temperature and visitors.\n             Comment on the same.                                                                           (5)",
          "topic": null
        },
        {
          "label": "iv",
          "marks": 2,
          "text": "Estimate the parameters A and B                                                                 (2)",
          "topic": null
        }
      ],
      "solution": "i)        Null hypothesis: There is no difference in variance of number of visitors in cold week and number of\n              visitors in a warm week\n\n              Test statistic: Scold^2/ Swarm^2 ~ F24,34                                                           [1]\n\n              Scold^2 = 1/24 * (4,718,190– 25*(10,250/25)^2)) = 21,487.08                                         [1]\n\n              Swarm^2 = 1/34 * (4,740,201 – 35*(12,001/35)^2)) = 18,389.10                                        [1]\n\n              The 95% quantile of F24,34 is 1.843 and observed value is\n\n              F = 21,487.08/18,389.10 = 1.17 < 1.83                                                               [1]\n\n              Hence there is no evidence that variance of visitors is higher in colder weeks                      [1]\n\n    ii)       Null hypothesis: There is no difference in mean number of visitors in cold and warm week\n              Assuming that the two population variances are equal, we have:\n              Sp^2 = (24 * 21,487.08 + 34 * 18,389.10)/58 = 19,671.02                                             [2]\n              t = (10,250/25 – 12,001/35) / sqrt(19,671.02 *(1/35+1/25)) = 1.8274                                 [1]\n\n              The 97.5% (two sided test) of the t58 is 2.009 and 2.000.\n\n              Hence there is no evidence suggesting that the mean number of visitors differ between cold and warm\n              weeks.                                                                                           [1]\n\n              Alternatively, students can use the z statistic considering that samples are large:\n\n                                                                                                         Page 5 of 7\n\fIAI                                                                                                      CS1A-1120\n             z = (10,250/25 – 12,001/35) / sqrt(21,487.08/25+18,389.10/35) = 1.8035\n\n             Since this is lower than 1.96, null hypothesis is accepted and same conclusion as t test.\n      iii)\n             Option 1 (considering number of weeks as 78):\n             Stx = 582,205 – 1/78 * 1,603 * (10,250+12,001) = 124,918.42                                         [1]\n\n             Sxx =(4,718,190 + 4,740,201)-(10,250+12,001)^2 /78 = 3,110,865.35                                   [1]\n\n             Stt = 51,823 - 1,603^2/78 = 18,879.29                                                               [1]\n\n             Rho(T,X) = 124,918.42/sqrt(3,110,865.35*18,879.29) = 0.5155                                         [1]\n\n             This shows a weak positive correlation between the number of visitors and temperature, i.e. number of\n             visitors increase with increase in temperature, however the correlation is weak                   [1]\n\n             Option 2 (considering number of weeks as 60):\n\n             Stx = 582,205 – 1/60 * 1,603 * (10,250+12,001) = -12,267.55                                         [1]\n\n             Sxx =(4,718,190 + 4,740,201) - (10,250+12,001)^2 /60 = 1,206,607.65                                 [1]\n\n             Stt = 51,823 - 1,603^2/60 = 8,996.18                                                                [1]\n\n             Rho(T,X) = -12267.55/sqrt(1,206,607.65 * 8,996.18) = -0.12                                          [1]\n\n             This shows a small negative correlation between the number of visitors and temperature, i.e. number of\n             visitors increase with drop in temperature, however the correlation is quite small to come to any\n             concrete conclusions.                                                                              [1]\n\n      iv)    Option 1 (considering number of weeks as 78):\n\n             A = (78 * 582,205 - 1,603 * (10,250+12,001)) / (78 * 51,823 – 1603^2) = 6.6167\n\n             Or from part iii),\n             A = 124,918.42 / 18,879.29 = 6.6167\n\n             B = ((10,250+12,001) – 6.6167*1,603)/78 = 149.29                                                    [2]\n\n             Option 2 (considering number of weeks as 60):\n\n             A = (60 * 582,205 - 1,603 * (10,250+12,001)) / (60 * 51,823 – 1603^2) = -1.3636\n\n             Or from part iii),\n             A = -12,267.55 / 8,996.18 = -1.3636\n\n             B = ((10,250+12,001) – (-1.3636) * 1,603)/60 = 407.29                                              [2]\n\n                                                                                                         Page 6 of 7\n\fIAI                                                                                                            CS1A-1120",
      "has_math": true,
      "session": "2020-11",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2020-11_QP.pdf",
      "source_sol": "raw/CS1A_2020-11_SOL.pdf"
    },
    {
      "q_num": 9,
      "marks": 7,
      "topic": "distributions",
      "subtopics": [],
      "stem": "It is known that for a shop of a particular category, the mean invoice amount towards sakes\n         is INR 2000 and standard deviation is INR 500. Stating clearly any assumptions that you\n         make, calculate the following:",
      "parts": [
        {
          "label": "i",
          "marks": 3,
          "text": "Probability that mean amount of next 10 invoices is less than INR 1700.                      (3)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 3,
          "text": "Probability that standard deviation of amount of next 10 invoices is less than INR 250.      (3)",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 1,
          "text": "Probability that both (i) and (ii) occur together. (i.e. mean amount of next 10 invoices is\n             less than INR 1700 and standard deviation of amount of these 10 invoices is less than\n             INR 250)                                                                                       (1)",
          "topic": null
        }
      ],
      "solution": "i)         Assumption : the sample is from a normal population.                                                  [0.5]\n               For a sample from normal distribution, we know that 𝑋~𝑁 𝜇,\n\n               𝑋~𝑁 2000,                                                                                              [0.5]\n\n               𝑃(𝑋 < 1700) = 𝑃 𝑍 <                                                                                     [1]\n                                              √\n\n               = 𝑃(𝑍 < −1.8974) = 1 − Φ(1.8974) = 1 − 0.9711 = 2.889 %                                                  [1]\n\n    ii)        Assumption : the sample is from a normal population.                                                  [0.5]\n                                                                      (    )\n               For a sample from normal distribution, we know that             ~𝜒                                     [0.5]\n                                    (     )\n               𝑃(𝑆 < 250 ) = 𝑃                < (10 − 1)        = 𝑃(𝜒 < 2.25) = 1.313 %                                 [2]\n\n    iii)       Assumption : the sample is from a normal population and hence the mean and variance are independent.\n\n            𝑃(𝑋 < 1700 ∩ 𝑆 < 250 ) = 𝑃(𝑋 < 1700). 𝑃(𝑆 < 250 ) = 0.02889 × 0.01313 = 0.038% [0.5]",
      "has_math": false,
      "session": "2020-11",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2020-11_QP.pdf",
      "source_sol": "raw/CS1A_2020-11_SOL.pdf"
    },
    {
      "q_num": 10,
      "marks": 13,
      "topic": "regression_glm",
      "subtopics": [],
      "stem": "",
      "parts": [
        {
          "label": "i",
          "marks": 6,
          "text": "Explain what is meant by a saturated model and discuss if such a model is useful in\n               practice.                                                                                    (2)\n\n        ii)    a) Define both Pearson and deviance residuals                                                (2)\n\n               b) Explain how these two types of residuals are different                                    (2)\n\n               c) State in which case they are the same                                                     (1)\n\n         It is believed that the claims arising out of Health Insurance primarily depends on two\n         variables – Age and BMI:\n\n         Age xi is the category 18 to 35, 36 to 45, 46 to 60 and above 60\n\n         BMI yi can be categorised as Under-weight, Normal, Over-weight and Obese depending on\n         the BMI values\n\n         The company wishes to model claim amounts using this data and is investigating models\n         which take into account the below variables:\n\n              Model    Choice of predictor     Scaled Deviance\n               A               1                    900\n               B              Age                   800\n               C           Age + BMI                770\n               D           Age * BMI                760\n\n        iii) Carry out the analysis of scaled deviance and recommend which model should be used\n             by the Insurance Company?                                                                      (6)",
          "topic": null
        }
      ],
      "solution": "i)         A saturated model has as many parameters as there are data points and would therefore be a perfect fit\n               to the data. It is not useful to predict and hence not used in practice.                          [2]\n    ii)\n               a) Pearson residuals are (y- μ)/sqrt(var(μ)) where μ is the fitted response estimator.                   [1]\n                  The deviance residuals are sign(y- μ)di where di is the contribution of the i-th to total deviances, sum\n                  of di ^2 is the scaled deviance.                                                                      [1]\n\n               b) The Pearson residuals tend to be skewed in non normal data while deviance residuals tend to be\n                  symmetric and hence normal assumption is more appropriate. For that reason, the latter is preferred\n                  in actuarial applications                                                                        [2]\n\n               c) For normally distributed data, the Pearson and deviance residuals are identical                   [1]\n    iii)       Using the information given, we can calculate the deviance differences and compare that with the\n               differences of the degrees of freedom for each of the model. If the decrease in the deviances is greater\n               than twice the difference in degrees of freedom this suggests an improvement.                        [1]\n\n           Model Choice of predictor      Degree of freedom Scaled Deviance Difference of scaled deviance\n            A            1                       16               900\n            B           Age                      13               800                    100\n            C        Age + BMI                   10               770                     30\n            D        Age * BMI                    1               760                     10\n           From the above table, it is seen that interaction model does not indicate significant improvement and hence\n           the recommended model is Age + BMI.                                                                      [1]\n                                                     ***********************\n\n                                                                                                               Page 7 of 7",
      "has_math": false,
      "session": "2020-11",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2020-11_QP.pdf",
      "source_sol": "raw/CS1A_2020-11_SOL.pdf"
    },
    {
      "q_num": 1,
      "marks": 7,
      "topic": "bayes_credibility",
      "subtopics": [],
      "stem": "",
      "parts": [
        {
          "label": "i",
          "marks": 1,
          "text": "A company hires candidates from 4 institutes. It selects 20% of candidates from Institute\n              A and B each. Balance candidates are selected from Institute C and D with equal share.\n              Probability that selected candidate will join the company is 99% for Institute A, 98% for\n              Institute B, 97% for Institute C and 96% for Institute D.\n\n         Calculate the probability in each of the following cases\n\n                 a Candidate who fails to join the company is from Institute A\n                 a Candidate who fails to join the company is from Institute B\n                 a Candidate who fails to join the company is from Institute C\n                 a Candidate who fails to join the company is from Institute D.                          (3)\n\n        ii)   Identify the incorrect statement/s, if any and correct them\n\n         a) Under squared error loss, the mean of the posterior distribution minimises the expected\n         loss function\n\n         b) Under absolute error loss, the mean of the posterior distribution minimises the expected\n         loss function\n\n         c) Classical approach treats unknown parameter as a random variable in statistical estimation\n         problems\n\n         d) In classical statistics, unknown parameter is a fixed quantity and probability can be\n         assigned to it                                                                                   (2)\n\n        iii) Calculate the Bayesian estimate of W under squared error loss when posterior distribution\n             of W follows Gamma(48,4).                                                                    (1)\n\n        iv) A student has correctly identified posterior distribution to be Ga(a,b) and correctly\n            calculated Bayesian estimate under squared error to be 'S'. Then, Bayesian estimate under\n            Absolute error loss and under All-or-nothing loss must be\n\n         a. Greater than 'S'\n         b. Lower than 'S'\n         c. equal to 'S'\n         d. none of the above                                                                             (1)",
          "topic": null
        }
      ],
      "solution": "i)\n\n       Institute Share Prob(fails to join)                    P(Institute i given fails to join)\n                       X             Y           Xi * Yi (Xi*Yi)/(Summation of Xi *Yi over A to D)\n             A        0.2          0.01          0.002                   7.41%\n             B        0.2          0.02          0.004                  14.81%\n             C        0.3          0.03          0.009                  33.33%\n             D        0.3          0.04          0.012                  44.44%\n                      1.0                        0.027                 100.00%\n\n       Correct Share for C and D (0.5 Mark)\n       Correct Probability for all 4 institutes (0.5 Mark)\n       Correct Formula (1 Mark)\n       Correct Calculation / Final Answer for each Institute (1 Mark)\n\n     ii) Only statement ‘a’ is True. All other statements are incorrect\n\n           Corrected version\n\n           b – Under absolute error loss, median of posterior distribution minimises the expected loss\n           function\n\n           c – Bayesian method assumes parameter to be a random variable\n\n           d – classical statistics assumes unknown parameter to be fixed and hence cannot assign\n           probability statements to it.\n\n     iii) Correct Formula and answer –\n\n           Mean of the Gamma is alpha /lambda = 48/4 =12                                                  [1]\n\n     iv)\n\n            Answer          Option b\n\n           As Ga(a,b) is positively skewed, mean > median > mode. (0.5 mark)\n\n           Bayesian estimate under squared error loss is mean equal to 'S'. (0.5 mark)\n           (calculations not expected as simple logic needs to be applied to solve this 1-mark\n           question)",
      "has_math": false,
      "session": "2021-03",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2021-03_QP.pdf",
      "source_sol": "raw/CS1A_2021-03_SOL.pdf"
    },
    {
      "q_num": 2,
      "marks": 9,
      "topic": "distributions",
      "subtopics": [],
      "stem": "",
      "parts": [
        {
          "label": "i",
          "marks": 2,
          "text": "What is the main purpose of performing Factor analysis? Comment on the original\n              variables and newly identified Principal components.                                        (3)\n\n        ii)   A student came up with 5 by 5 variance-covariance matrix of the Principal Components\n              (PC1, PC2, PC3, PC4, PC5) with these 5 diagonal entries: 0.456, 0.137, 0.080, 0.0165\n              and 0.012 respectively. Identify the percentage of the total variance explained by each\n              Principal Component.                                                                        (4)\n\n        iii) What can be concluded based on the results of part (ii).                                     (2)",
          "topic": null
        }
      ],
      "solution": "i)     Factor analysis / Principal Component Analysis is -\n\n       A method for reducing the dimensionality of data                                                 [0.5]\n\n                                                                                                   Page 2 of 9\n\f    IAI                                                                                              CS1A-0321\n\n           It seeks to identify key components necessary to model and understand data                    [0.5]\n\n           Original variables may be\n\n              correlated with each other                                                                [0.5]\n\n           While Newly identified principal components are chosen to be\n\n              uncorrelated                                                                              [0.5]\n              linear combinations of the original variables of the data                                 [0.5]\n              which maximise the variance                                                               [0.5]\n     ii)\n\n            Principal         Diagonal entry     PCi/ (Sum(PCi) over 1\n           Component               (PCi)                  to 5)\n               PC1                 0.456                 65.0%             of total variance explained by PC1\n               PC2                 0.137                 19.5%             of total variance explained by PC2\n               PC3                 0.08                  11.4%             of total variance explained by PC3\n               PC4                0.0165                 2.4%              of total variance explained by PC4\n               PC5                 0.012                 1.7%              of total variance explained by PC5\n                                  0.7015                100.0%\n\n    Correct formula (1 mark)\n    Sum of PCi (0.5 marks)\n    Correct calculation (2.5 Marks)\n\n     iii) As 1st 3 Principal components explain over 95% of total variance, dimensionality can be reduced\n          to 3 for this dataset                                                                       [1]\n\n         The 1st 3 Principal Components can then be used for building further classification or regression\n         modelling purpose                                                                             [1]",
      "has_math": false,
      "session": "2021-03",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2021-03_QP.pdf",
      "source_sol": "raw/CS1A_2021-03_SOL.pdf"
    },
    {
      "q_num": 3,
      "marks": 7,
      "topic": "inference",
      "subtopics": [],
      "stem": "",
      "parts": [
        {
          "label": "i",
          "marks": 2,
          "text": "Calculate the probability that the sample variance of a sample of 5 values from a normal\n                 distribution will be more than 3 times the sample variance of a sample of 17 values from\n                 an independent normal distribution with the same variance.\n\n         You are given that\n\n         F4,16 at 10% is 2.333\n\n         F4,16 at 5% is 3.007 and\n\n         F4,16 at 2.5% is 3.729.                                                                            (3)\n\n        ii)      In case of Kendall’s correlation, Identify the number of observations in the assignment,\n                 given that there are no ties and Nc = 87 and Nd = 123?\n\n         Note – Nc and Nd denote number of concordant and discordant pairs respectively.\n\n         a. 18\n         b. 19\n         c. 20\n         d. 21\n         e. 22                                                                                              (2)\n\n        iii) State the properties that can lead to data being classified as ‘big’ data.                     (2)",
          "topic": null
        }
      ],
      "solution": "i)\n\n      Let X denotes the sample with 5 values and Y denotes the sample with 17 values\n      Given that population variances are equal i.e. Sigma X = Sigma Y,                                   [0.5]\n      Therefore, P (Sx2 / Sy2 > 3) = P(F4,16 > 3)                                                           [1]\n      As upper 5% point of F 4,16 distribution is 3.007.                                                  [0.5]\n      So, required probability is just over 5%                                                              [1]\n\n    (As, F4,16 at 10% is 2.333 and F4,16 at 2.5% is 3.729.\n\n    Hence required number 3 is between 5% and 10%)\n\n    Deduct half mark if over 5% is not mentioned / incorrectly mentions under 5% or 5%\n     ii) Answer Option d\n\n           Nc + Nd = 87 +123 = 210                                                                       [0.5]\n\n           N (N-1)/ 2 = 210                                                                              [0.5]\n\n           N^2-N-420 =0\n\n                                                                                                   Page 3 of 9\n\fIAI                                                                                            CS1A-0321\n\n          N^2-21N+20N-420 = 0\n\n          N(N-21)+20(N-21) =0\n\n          (N+20) (N-21) =0\n\n          N=-20 / N=21                                                                             [0.5]\n          As N cannot be negative, N=21                                                            [0.5]\n     iii) Size of the dataset                                                                      [0.5]\n\n          Speed of arrival of the data                                                              [0.5]\n\n          Variety of different sources from which the data is drawn                                [0.5]\n\n          Reliability of the data elements might be difficult to ascertain                         [0.5]",
      "has_math": false,
      "session": "2021-03",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2021-03_QP.pdf",
      "source_sol": "raw/CS1A_2021-03_SOL.pdf"
    },
    {
      "q_num": 4,
      "marks": 8,
      "topic": "distributions",
      "subtopics": [],
      "stem": "",
      "parts": [
        {
          "label": "i",
          "marks": 2,
          "text": "If coefficient of variation is defined as\n\n         Coefficient of variation = Standard variation / mean\n\n         What will be the effect of an increase in variance on Coefficient of variation in case of\n         following distributions?\n\n         a) Chi Square Distribution\n         b) Poisson Distribution\n         c) Exponential Distribution                                                                        (6)\n\n        ii) Using inverse transform method, simulate 2 values from Exp(0.5) distribution using\n            random value 0.769 and 0.004 from U(0,1) distribution.                                          (2)",
          "topic": null
        }
      ],
      "solution": "i)      For Chi square –\n\nif degrees of freedom = n then mean =n and variance = 2n                                           [0.5]\n\nCoefficient of variation = Sqrt (2n)/ n = Sqrt(2/n)                                                [0.5]\n\nHence, as variance increases, coefficient of variation will reduce.                                   [1]\n\nFor Poisson -\n\nIf Poisson parameter is L then mean = variance = L                                                 [0.5]\n\nCoefficient of variation = Sqrt(L)/L = Sqrt(1/L)                                                   [0.5]\n\nHence, as variance L increases, coefficient of variation will reduce.                                [1]\n\nFor Exponential –\n\nIf exponential parameter is 1/L then mean =L and variance = L^2                                    [0.5]\n\nCoefficient of variation = Sqrt (L^2)/L =1 = constant value                                        [0.5]\n\nHence, Coefficient of variation will have no effect of increase (or any change) in variance           [1]\n\nii)       F(x) = u = 1-e^(-0.5x)                                                                   [0.5]\n\n1-u = e^(-0.5x)\n\nLN (1-u) = -0.5x\n\nX = -LN (1-u)/0.5                                                                                  [0.5]\n\nFor u = 0.769, x = -2*LN(1-0.769) =2.931 and                                                       [0.5]\n\nfor u= 0.004, x = -2*LN(1-0.004) = 0.008                                                           [0.5]",
      "has_math": false,
      "session": "2021-03",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2021-03_QP.pdf",
      "source_sol": "raw/CS1A_2021-03_SOL.pdf"
    },
    {
      "q_num": 5,
      "marks": 12,
      "topic": "distributions",
      "subtopics": [],
      "stem": "",
      "parts": [
        {
          "label": "i",
          "marks": 2,
          "text": "For the following joint distribution between X and Y, Calculate E[Y|X=2]\n\n                                 Y\n                                 1                   2                3\n                        1        0.1                 0.1              0.3\n             X          2        0.3                 0.2              0                                     (2)\n\n        ii)      If X and Y are independent standard normal variables then calculate the mean, variance\n                 and the distribution of 5X - 4Y.                                                           (2)\n\n        iii) If a random variable X follows Gamma (10,0.1) then P(X>50) is equivalent to finding\n             the probability using the certain chi square distribution. Identify the equivalent\n\n              probability calculation formula involving chi square distribution and calculate the\n              required probability using chi square tabulated value.\n              You are given that chi square tabulated value is 0.0318.                                          (3)\n\n        iv) Calculate the required probability as mentioned in Part (iii) above using normal\n            approximation.                                                                                      (3)\n\n        v)    Comment on the results obtained in part (iii) and (iv).                                           (2)",
          "topic": null
        }
      ],
      "solution": "i)        P(X=2) = 0.3 + 0.2 + 0 = 0.5\n\nRequired expectation is summation of y* P(Y=y| X=2)                                                   [1]\n\n                                                                                              Page 4 of 9\n\f  IAI                                                                                 CS1A-0321\n\n  = 1*0.3/0.5 + 2*0.2/0.5 +3*0/0.5\n\n  =3/5 + 4/5 = 7/5 =1.4                                                                      [1]\n\n ii) For 5X - 4Y,\n  Mean is 5*E(X) – 4*E(Y) = 5*0 -4*0 =0                                                   [0.5]\n  Variance is 5^2 * Var (X) + 4^2 * Var (Y) = 25*1 + 16*1 = 41                              [1]\n  Hence required distribution is N(0,41)                                                  [0.5]\n\niii) (2*Lambda *X) ~ Chi square distribution with (2 *alpha) degrees of freedom           [0.5]\n\n  P(X>50) = P(2*0.1*X > 2*0.1*50)\n\n  = P(Chi square >10)                                                                       [1]\n\n  with 2*10 =20 degrees of freedom                                                        [0.5]\n\n  Hence,\n\n  Required Chi square expression is\n\n  P(Chi square > 10) where chi square distribution will have 20 degrees of freedom\n\n  Required chi square probability is equal to 1-0.0318 = 0.9682                           [0.5]\n\n  Hence, probability of X greater than 50 is over 96.8%                                   [0.5]\n\niv) E(X) = alpha /lambda = 10/0.1 = 100                                                   [0.5]\n\n  Var(X) = alpha /lambda^2 = 10/0.1^2 = 1000                                              [0.5]\n\n  Hence using Central Limit theorem, using normal approximation                           [0.5]\n\n   (P X>50) = ~P(N(100,1000) >50)\n   ~P(Z>(50-100)/sqrt(1000))\n   ~P(Z(N(0,1)>-1.58114))\n   ~Z(N(0,1)<1.58114))\n\n  For correct equation as above                                                            [1]\n\n  From tables\n\n          x      phi(x)\n        1.58    0.94295\n        1.59    0.94408\n\n  Answers using interpolation are accepted though not expected.\n\n  There is over 94.3% probability that X is greater than 50\n\n  using normal approximation to underlying gamma distribution\n\n                                                                                     Page 5 of 9\n\f  IAI                                                                                          CS1A-0321\n\nv) Probability calculated using normal distributional assumption is lower as compared to the answer\n   obtained using chi square for the underlying Gamma distribution.                             [1]\n\n  As gamma distribution is positively skewed and has thick tail compared to normal and it will tend to\n  be more like normal only when alpha tends towards infinity. As in this case, value of alpha is only 10,\n  so normal approximation is not truly able to capture the correct thick tail found for Gamma.",
      "has_math": false,
      "session": "2021-03",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2021-03_QP.pdf",
      "source_sol": "raw/CS1A_2021-03_SOL.pdf"
    },
    {
      "q_num": 6,
      "marks": 7,
      "topic": "bayes_credibility",
      "subtopics": [],
      "stem": "",
      "parts": [
        {
          "label": "i",
          "marks": 3,
          "text": "Claim amounts x1,x2…xn are observed over n years. Claim amounts over 6 years equal\n              3600.\n\n         If, Total claim amounts per annum follow N(b, 1002) and Prior believes are that b is N(700,\n         702).\n\n         Identify mean, variance of posterior distribution of \n\n        ii)   Using the information from earlier part, calculate the posterior probability that b is greater\n              than 600 using normal-normal model.                                                               (2)\n\n        iii) Comment on the possible value of credibility factor Z and true mean based on the results\n             in part (i).\n             How would the value of Z change with change in standard deviation of prior and\n             likelihood?                                                                                        (3)",
          "topic": null
        }
      ],
      "solution": "i)     Prior mean (Mu0)=700 Var (sigma0)=70^2 = 4900\n\n  observed sample mean (x bar) over 6 (=n) years is 3600 /6 = 600 and Var of distribution\n  (sigma)=100^2 = 10000\n\n  Posterior Variance = 1/ [(n/sigma^2) + (1/ (Sigma0^2)]                                           [0.5]\n\n  = 1/ /[(6/10000)+(1/4900)] = 1/ 0.000804 =1243.655                                               [0.5]\n\n  posterior mean = [ (n* x bar) / sigma^2) + (mu0 /Sigma0^2)] / [(n/sigma^2) + (1/ (Sigma0^2)]\n\n  = [(6*600/10000) + (700/4900)] /[(6/10000)+(1/4900)] = (0.36 + 0.142857)/0.000804\n\n  =625.3807                                                                                        [0.5]\n\n  Hence posterior distribution of beta is N(625.3807, 1243.655) i.e (625.3807, 35.266^2)\n  ii) Required probability is\n\n   P(Z>600-625.3807/sqrt(1243.655))\n   ~P(Z<0.7197039)\n\n  For correct equation as above                                                                      [1]\n\n  From tables,\n\n    X    phi(x)\n   0.72 0.76424\n\n  Correct tabulated values                                                                           [1]\n  There is over 76% probability that beta is greater than 600\n\n  iii)\n\n                prior     likelihood posterior\n     mean       700           600       625.38\n       sd        70           100        35.27\n  As posterior mean is closer towards likelihood mean, Credibility factor is expected to be greater than\n  0.5.\n\n  If we use Z = 0.75, using simple weighted average, we will get\n\n  posterior mean = 0.75 (600) + 0.25(700)\n\n  = 450 + 175 = 625\n\n                                                                                             Page 6 of 9\n\fIAI                                                                                             CS1A-0321\n\nHence Z is expected to be close to 0.75                                                               [1]\n\nIf prior SD had been more than 70 then Credibility Factor Z would have increased (Higher the variance\nof prior, less reliable the prior belief would be)                                                [1]\n\nIf likelihood SD had been lower than 100 then Credibility Factor Z would have increased. (Lower the\nvariance of likelihood more reliable the data would be)                                         [1]",
      "has_math": false,
      "session": "2021-03",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2021-03_QP.pdf",
      "source_sol": "raw/CS1A_2021-03_SOL.pdf"
    },
    {
      "q_num": 7,
      "marks": 15,
      "topic": "bayes_credibility",
      "subtopics": [],
      "stem": "An insurance company has a sample of 500 motor insurance policies with claims. Say X\n         denotes the number of policies where the annual claim amount exceeded INR 10,000 and X\n         follows Binomial distribution with parameters (5000, p).",
      "parts": [
        {
          "label": "i",
          "marks": 4,
          "text": "Assuming annual total claim amounts per policy are independent and identically\n              distributed, estimate the unknown proportion “p” of claims higher than 10,000 per year\n              by deriving its maximum likelihood.                                                               (4)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 4,
          "text": "If there is a prior knowledge about p follows Uniform distribution with parameters (0, 1)\n            then derive the density of the Bayesian posterior distribution of p in terms of n and X.\n            Also state the name of the posterior distribution that p follows along with its parameters.         (4)\n\n         If 200 out of 500 policies with claims from the sample have a claim above INR 10,000 then",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 1,
          "text": "Estimate p using the MLE in (i).                                                                   (1)",
          "topic": null
        },
        {
          "label": "iv",
          "marks": 2,
          "text": "Estimate p using the Bayesian estimator under quadratic loss, based on the posterior\n            distribution derived in part (ii).                                                                  (2)",
          "topic": null
        },
        {
          "label": "v",
          "marks": 1,
          "text": "Comment on the difference between the values of p estimated in (iii) and (iv)                       (1)",
          "topic": null
        },
        {
          "label": "vi",
          "marks": 3,
          "text": "From part (iv), express the Bayesian estimator in the form of credibility factor and\n            determine the value of credibility factor.                                                          (3)",
          "topic": null
        }
      ],
      "solution": "i)      L(p) = constant * p^x*(1-p)^(n-x)                                                       [1]\n              logL(p) = constant + x logp+ (n-x)log(1-p)                                              [1]\n              Taking derivative w.r.t. p\n              logL’(p) = x/p – (n-x)/(1-p)                                                            [1]\n              Equating to 0\n              x(1-p)-(n-x)p = 0\n              x – xp – np + xp = 0\n              p = x/n                                                                                 [1]\n\n              OR (if instead of “n”, 5000 is substituted):\n\n              L(p) = constant * p^x*(1-p)^(5000-x)                                                    [1]\n              logL(p) = constant + x logp+ (5000-x)log(1-p)                                           [1]\n              Taking derivative w.r.t. p\n              logL’(p) = x/p – (5000-x)/(1-p)                                                         [1]\n              Equating to 0\n              x(1-p)-(5000-x)p = 0\n              x – xp – 5000p + xp = 0\n              p = x/5000\n\n      ii)     f(p) = 1/(1-0)\n              Let posterior distribution of p be denoted by P(p)\n\n              P(p) α L(p) * f(p)                                                                      [1]\n              P(p) α p^x*(1-p)^(n-x) * 1                                                              [1]\n              P(p) α p^(x+1-1)*(1-p)^(n-x+1-1)\n\n              Therefore, the posterior distribution is beta distribution with parameters x+1, n-x+1\n              OR if instead of “n”, 5000 is substituted, posterior distribution would be beta\n              distribution with parameters x+1, 5001-x+\n\n      iii)    p = 200/500 = 0.4                                                                       [1]\n\n      iv)     Under quadratic loss, the Bayesian estimator is the expectation of the posterior\n              distribution. In this case, p = (200+1)/ (200+1+500-200+1) = 201/502 = 0.4004 [2]\n\n      v)      The two estimates are almost equal, this is because the impact of prior distribution is very\n              limited and the Bayesian estimator is mainly determined by the actual data               [1]\n\n                                                                                              Page 7 of 9\n\fIAI                                                                                               CS1A-0321\n\n      vi)       Posterior mean can be written in credibility form as:\n\n                p = (x+1)/(n+2)                                                                         [1]\n                p = x/(n+2) + 1/(n+2)\n                = x/n * n/(n+2) + 2/(n+2) * (1/2)                                                       [1]\n                = E(X) * Z + E(p) * (1-Z)\n                Where E(X) = x/n and E(p) = ½ and Z = n/(n+2) = 500/502                                [1]",
      "has_math": false,
      "session": "2021-03",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2021-03_QP.pdf",
      "source_sol": "raw/CS1A_2021-03_SOL.pdf"
    },
    {
      "q_num": 8,
      "marks": 18,
      "topic": "inference",
      "subtopics": [
        "distributions"
      ],
      "stem": "The below data gives the marks scored and time spent on social media in hours per day:\n\n              Hours x           6        5.5       0.5       3.5       7.5        4.5    4    1.5     2       10\n              Marks y           65       45        75        78        45         80     49   91      67      49\n\n         For this data: ∑x = 45; ∑𝑥 = 277.5; ∑y = 644; ∑𝑦 =43,956; ∑xy = 2,602\n\n         Scatterplot for the above data is provided below:\n\n                                     Hours per day Vs. Marks\n                      100\n\n                       90\n\n                       80\n\n                       70\n              Marks\n\n                       60\n\n                       50\n\n                       40\n\n                       30\n                            0        2         4         6         8         10     12\n                                                   Hours spent",
      "parts": [
        {
          "label": "i",
          "marks": 1,
          "text": "Based on the above scatterplot, comment on the association between marks scored and\n                  hours spent per day on social media.                                                                   (1)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 4,
          "text": "Calculate the correlation co-efficient between the two variables and comment on the\n            same.                                                                                                        (4)",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 7,
          "text": "Investigate the hypothesis that there is a negative correlation using Fisher’s\n             transformation of the correlation co-efficient. You should clearly state the hypothesis of\n             your test and any assumption made for the test to be valid.                                                 (7)",
          "topic": null
        },
        {
          "label": "iv",
          "marks": 3,
          "text": "Fit a linear regression model to this data with marks being the response variable and\n            hours spent as the explanatory variable.                                                                     (3)",
          "topic": null
        },
        {
          "label": "v",
          "marks": 2,
          "text": "Calculate the coefficient of determination for this model and interpret the same.                      (2)",
          "topic": null
        },
        {
          "label": "vi",
          "marks": 1,
          "text": "Calculate the expected change in marks for every 1 additional hour spent on social media\n            basis the model fit in (iv).                                                                                 (1)",
          "topic": null
        }
      ],
      "solution": "i)    The scatter plot suggests an inverse relation between marks obtained and hours spent on\n            social media per day                                                                [1]\n\n      ii) Sxx = 277.5 – 45^2/10 = 75\n          Syy = 43,956 – 644^2/10 = 2482.4\n          Sxy = 2,602 – 644*45/10 = -296\n          r = Sxy/√(𝑆𝑥𝑥 ∗ 𝑆𝑦𝑦) = -0.686                                                                 [3]\n\n            -69% correlation co-efficient also implies a moderate negative linear relation between the two\n            variables as visible from the scatterplot.                                                  [1]\n\n      iii) Null hypothesis H0: ρ = 0 against H1: ρ < 0                                                  [1]\n\n            Need to assume that data come from a bivariate normal distribution.                         [1]\n\n            From page 25 of tables, r = 0.5 * ln(1-0.686/1.686) = -0.8404                               [1]\n\n            And under H0, this should be a value from the N(0, 1/7) distribution.                       [1]\n\n            Fisher’s standardized statistic = (-0.8404 – 0)/( 1/7) = -2.22                              [1]\n\n            This gives the p-value = P(z<-2.22) = 0.013 which is quite small and hence shows a strong\n            evidence to reject the null hypothesis with 95% confidence. We can conclude that marks\n            obtained and hours spent on social media are negatively correlated.                       [2]\n\n      iv) Beta = Sxy/Sxx = -296/75 = -3.9467\n          Alpha = mean of y – beta * mean of x = 644/10 + 3.9467 *45/10 = 82.16\n          Fitted line is y = 82.16 – 3.9476x\n\n      v) 𝑅 = -0.686^2 = 0.4706\n         This gives the proportion of total variation explained by the model.                           [2]\n\n      vi) For every additional hour spent on social media per day, the total marks reduce by 3.95 (~4\n          marks) basis the fitted equation.                                                         [1]",
      "has_math": true,
      "session": "2021-03",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2021-03_QP.pdf",
      "source_sol": "raw/CS1A_2021-03_SOL.pdf"
    },
    {
      "q_num": 9,
      "marks": 6,
      "topic": "distributions",
      "subtopics": [],
      "stem": "Let X1, X2, X3,…,Xn be a random sample of claim amounts which follow gamma distribution\n         with a parameter α = 5 and unknown parameter λ.",
      "parts": [
        {
          "label": "i",
          "marks": 3,
          "text": "Specify the distribution of ∑              𝑋𝑖 with parameters. Hence also state the distribution of\n                  2nλ𝑋.                                                                                                  (3)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 3,
          "text": "A random sample of 10 such claims has a mean of 100. Pick the 95% confidence interval\n                  for λ.\n\n          Option A: (0.03897, 0.06220)\n          Option B: (0.03711, 0.06480)\n          Option C: (0.01618, 0.03570)\n          Option D: (0.01738, 0.03375)                                                                          (3)",
          "topic": null
        }
      ],
      "solution": "i)        Sum of Xi follows Gamma distribution with parameters 5n, λ                            [1.5]\n\n                                                                                                Page 8 of 9\n\fIAI                                                                                               CS1A-0321\n\n              If Y ~ gamma (α, λ) then 2 λ Y ~ chi squared distribution with degree of freedom 2α\n              Hence 2nλ𝑋 follows chi squared distribution with df 10n.                            [1.5]\n\n      ii)     Option B is correct                                                                      [3]",
      "has_math": true,
      "session": "2021-03",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2021-03_QP.pdf",
      "source_sol": "raw/CS1A_2021-03_SOL.pdf"
    },
    {
      "q_num": 10,
      "marks": 5,
      "topic": "regression_glm",
      "subtopics": [
        "distributions"
      ],
      "stem": "An insurance company models claim numbers on health insurance portfolio using a Poisson\n          distribution and the mean depends on age and gender of policyholder.",
      "parts": [
        {
          "label": "i",
          "marks": 1,
          "text": "Specify the link function for fitting a generalized linear model for the mean of the Poisson\n                 distribution                                                                                   (1)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 2,
          "text": "State the linear predictor used for modelling the age and gender dependence as:\n                    a) Age + Gender                                                                             (2)\n\n                    b) Age + Gender + Age . Gender                                                              (2)",
          "topic": null
        }
      ],
      "solution": "i)      Link function is g(µ) = logµ                                                              [1]\n\n      ii)\n      a) The linear predictor is αi + βx where the intercept αi for i = 1, 2 depends on gender         [2]\n      b) The linear predictor is αi + βix where both parameters depend on the gender                   [2]",
      "has_math": false,
      "session": "2021-03",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2021-03_QP.pdf",
      "source_sol": "raw/CS1A_2021-03_SOL.pdf"
    },
    {
      "q_num": 11,
      "marks": 6,
      "topic": "inference",
      "subtopics": [
        "distributions"
      ],
      "stem": "In a survey by Industry, 20 organizations from each industry were questioned on their\n          attrition rate and a summary was prepared as below:\n\n               Industry                  Manufacturing IT / ITES Consulting\n               Average resignation %     27%           36%       30%\n               Sample standard deviation 5%            10%       8%\n\n          Perform a one-way analysis of variance test to test the hypothesis that the Industry has no\n          impact on resignation rate.                                                                           [6]",
      "parts": [],
      "solution": "H0: There is no difference among industries                                                            [1]\n\nH1: At least one industry differs significantly from the overall mean\n\nSSR = 19 (5^2 + 10^2 + 8^2) = 3591                                                                    [1.5]\n\nMean of resignation = (27+36+30)/3 = 31\n\nSSB = 20((27-31)^2 + (36-31)^2 + (30-31)^2)                                                           [1.5]\n\n= 840\n\nF2, 57 = (840/2)/(3591/57) = 6.667                                                                      [1]\n\nThe 1% point from F2,60 is 4.977 and since the test statistic is higher than this, the null hypothesis is\nrejected. We conclude that resignation rate is different across different industries.                    [1]\n\n                                 *********************************\n\n                                                                                                Page 9 of 9",
      "has_math": false,
      "session": "2021-03",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2021-03_QP.pdf",
      "source_sol": "raw/CS1A_2021-03_SOL.pdf"
    },
    {
      "q_num": 1,
      "marks": 7,
      "topic": "data_analysis",
      "subtopics": [],
      "stem": "",
      "parts": [
        {
          "label": "i",
          "marks": 1,
          "text": "Define three “key forms of data analysis” using examples.                                       (6)\n\n        ii) Which of the following is not a property for a data set to be classified as “big data”?\n\n           A. Speed of arrival of data\n           B. Size of the data set\n           C. Homogeneity within data set\n           D. Reliability of data elements may be difficult to ascertain",
          "topic": null
        }
      ],
      "solution": "i)    a) Descriptive analysis: It involves summarising the data or presenting it in a format which\n            highlights any patterns or trends i.e. producing summary statistics like measures of\n            central tendency and dispersion. It describes a data set rather than giving any specific\n            conclusions.                                                                                         (1)\n         Example (any one point should also fetch marks):\n\n                   Calculating mean and standard deviation of number of motor claims in a day\n                   Plotting graphs on the average rainfall every month to illustrate the months with\n                    heaviest rainfall                                                                            (1)\n         b) Inferential analysis: This involves estimating the summary parameters of a population\n            based on the sample data set under consideration and testing hypotheses.                             (1)\n         Example (any one point should also fetch marks):\n\n                   Any example of hypothesis testing (people visit malls more on weekend than on a\n                    weekday\n                   Rate of health claims in India is same as health claims made in Tier I cities                (1)\n         c) Predictive analysis: This extends the principle behind inferential analysis in order for the\n            user to analyse the past data and make predictions about the future event.                           (1)\n         Example (any one point should also fetch marks):\n              Predicting the number of lapses that will happen in the future years’ basis the\n                 number of lapses in the last 1 year\n              Forecasting the number of customers who would move to using electric vehicles                 (1)\n\n   ii)   C                                                                                                   (1)",
      "has_math": false,
      "session": "2021-09",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2021-09_QP.pdf",
      "source_sol": "raw/CS1A_2021-09_SOL.pdf"
    },
    {
      "q_num": 2,
      "marks": 5,
      "topic": "distributions",
      "subtopics": [],
      "stem": "The number of hospitalisations due to COVID19 have varied experience over months and\n        data corresponding to a group illustrates that one should expect:\n\n        Two hospitalisations in October\n        Three hospitalisations in November\n        One hospitalisation in December\n\n        Determine the probability that there will be less than five hospitalisations in the period\n        October to December given the hospitalisation in each month are independent and are\n        assumed to follow Poisson distribution.                                                            [5]",
      "parts": [],
      "solution": "Let Xi (i = 1,2,3) denote the number of hospitalisations in the month of October, November\n         and December respectively.                                                                         (0.5)\n         From the information provided, X1 ~ Poi (2), X2 ~ Poi (3) and X3 ~ Poi (1)                         (0.5)\n         Let the total hospitalisation over this period be denoted by X where X = X1 + X2 + X3                   (1)\n         Since all Xi’s are independent\n         X ~ Poi (2 + 3 + 1)                                                                                     (1)\n         Thus, P [X < 5] = ∑ (6x e-6 )/x! (summation over x = 0 to 4)                                            (1)\n                          = 0.0248 + 0.0149 + 0.0446 + 0.0892+ 0.1339\n                          = 0.2851                                                                               (1)",
      "has_math": false,
      "session": "2021-09",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2021-09_QP.pdf",
      "source_sol": "raw/CS1A_2021-09_SOL.pdf"
    },
    {
      "q_num": 3,
      "marks": 13,
      "topic": "distributions",
      "subtopics": [],
      "stem": "The amount of claims of a motor insurance company is modelled as an exponential random\n        variable with λ = 1.25 (in ‘000s). A data analyst is interested in assessing the probability of\n        Y exceeding 10 (INR 10,000 if represented in absolute amounts), wherein; Y is the total of\n        10 independent motor claim amounts.",
      "parts": [
        {
          "label": "i",
          "marks": 3,
          "text": "Show, using moment generating functions, that:\n\n           a) Y has gamma distribution and                                                                 (3)\n\n           b) 2.5Y has a χ2 20 distribution                                                                (3)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 2,
          "text": "Estimate the probability of Y > 10:\n\n           A. 0.7986\n           B. 0.2014\n           C. 0.2140\n           D. None of the above",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 4,
          "text": "Specify an approximate normal distribution for Y by applying the central limit theorem\n             and use this to calculate the approximate value of Y > 10.                                    (4)",
          "topic": null
        },
        {
          "label": "iv",
          "marks": 1,
          "text": "Comment on the use of this approximation and on the result.                                    (1)",
          "topic": null
        }
      ],
      "solution": "Page 2 of 9\n\f   IAI                                                                                           CS1A-0921\n    i)     a) Let Xi represent each motor claim amount for i = 1 to 10\n           Moment generating function for exponential distribution, Mx(t) = (1 – t/ λ)-1\n           Hence, for Y = ∑Xi (i = 1 to 10)\n           MY(t) = (MX(t)) ^ 10\n                 = (1 – t/ λ)-10                                                                           (0.5)\n\n           which is the moment generating function of gamma distribution with α = 10 and λ = 1.25          (0.5)\n\n           b) MGF of 2.5Y is E[e(2.5t)Y]                                                                   (0.5)\n           = MY[2.5t]                                                                                      (0.5)\n           = (1 – 2t) -10                                                                                  (0.5)\n           = (1 – t/ 0.5) -10                                                                              (0.5)\n           which is the moment generating function of gamma (10, 0.5)                                      (0.5)\n           i.e. χ2 20 distribution                                                                         (0.5)\n\n    ii)    B                                                                                                   (2)\n\n    iii)   From i.a. above, Y ~ gamma (10, 1.25)                                                           (0.5)\n           Therefore Y has mean 10/1.25 = 8 and variance = 10/(1.25)2 = 6.4                                  (1)\n           Applying central limit theorem Y ~ N (8,6.4)                                                    (0.5)\n           Thus, P[Y>10] = P [Z > (10 – 8)/(√6.4) = 0.791]                                                   (1)\n                         = 1 – 0.786                                                                       (0.5)\n                         = 0.214                                                                           (0.5)\n    iv)    n is not large enough for the central limit theorem to be used, but the approximation is still\n           close to the true probability                                                                     (1)",
      "has_math": true,
      "session": "2021-09",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2021-09_QP.pdf",
      "source_sol": "raw/CS1A_2021-09_SOL.pdf"
    },
    {
      "q_num": 4,
      "marks": 4,
      "topic": "distributions",
      "subtopics": [],
      "stem": "A teacher wanted to analyse if there has been any alteration in the marks of his nine students\n        after taking coaching classes. A test is conducted before students started the coaching classes\n        and after completing the coaching classes. The scores before and after the coaching classes\n        are given below:\n\n         Student #                       1 2 3 4 5 6 7 8 9\n         Before coaching classes         50 35 37 49 40 38 44 40 43\n         After coaching classes          47 34 33 54 42 17 42 31 44\n\n        Determine the Pearson’s correlation coefficient between the scores before and after the\n        coaching classes:\n\n        A. 0.8870\n        B. 0.6795\n        C. 0.7417\n        D. 0.5421\n        E. None of the above",
      "parts": [],
      "solution": "E                                                                                           [4 Marks]",
      "has_math": false,
      "session": "2021-09",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2021-09_QP.pdf",
      "source_sol": "raw/CS1A_2021-09_SOL.pdf"
    },
    {
      "q_num": 5,
      "marks": 5,
      "topic": "inference",
      "subtopics": [],
      "stem": "Let X and Y be random variables with joint density function as\n\n        f(x , y) = (1 / 27) * (2 x + y) and X and Y can assume values 0,1,2",
      "parts": [
        {
          "label": "i",
          "marks": 3,
          "text": "Calculate the marginal distribution for Y.                                                (3)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 2,
          "text": "Compute the conditional distribution of X =2 given Y = 1:\n\n        A. 0.4\n        B. 0.3333\n        C. 0.5556\n        D. 0.6667\n        E. None of the above                                                                         (2)",
          "topic": null
        }
      ],
      "solution": "i)    f(x,y) = (1/27)*(2x + y) where x = 0,1,2 and y = 0,1,2\n         Joint probability distribution of X, Y i.e f(x,y) is given by the table\n         f(x,y = 0,0) = 0\n         f(x,y = 0,1) = 1/27\n         f(x,y = 0,2) = 2/27\n         f(x,y = 1,0) = 2/27\n         f(x,y = 1,1) = 3/27\n\n                                                                                                 Page 3 of 9\n\f  IAI                                                                                            CS1A-0921\n          f(x,y = 1,2) = 4/27\n          f(x,y = 2,0) = 4/27\n          f(x,y = 2,1) = 5/27\n          f(x,y = 2,2) = 6/27\n\n          fY(0) = 6/27\n          fY(1) = 9/27\n          fY(2) = 12/27                                                                                    [3]\n\n   ii)    C                                                                                                (2)",
      "has_math": false,
      "session": "2021-09",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2021-09_QP.pdf",
      "source_sol": "raw/CS1A_2021-09_SOL.pdf"
    },
    {
      "q_num": 6,
      "marks": 20,
      "topic": "inference",
      "subtopics": [],
      "stem": "",
      "parts": [
        {
          "label": "i",
          "marks": 4,
          "text": "State the properties of a “good estimator” and “consistent estimator”.                    (2)\n\n        ii) In a particular small organisation, there are 5 Senior grade employees out of total 11\n            employees. You are told that ESOPs are granted to 3 employees.\n\n        Fisher’s test – Summary of data provided.\n\n         ESOP             Yes      No    Total\n         Senior Grade      A       C       5\n         Other             B       D       6\n         Total             3        8     11\n\n        Four possible ways to fill the grid above are represented by say, W1, W2, W3 and W4 having\n        probabilities of 0.0606, 0.1212, 0.3637 and 0.4545, respectively.\n\n        W1 to W4 are defined below: e.g., for W1, A=3, B=0, C=2 and D=6\n\n        W1      W2      W3      W4\n        3 2     0 5     2 3     1 4\n        0 6     3 3     1 5     2 4\n\n        a) Using the information provided above, calculate the p-value of finding two senior grade\n           employees having ESOPs.                                                                   (2)\n\n        b) Show that probability of W1 i.e., probability of finding 3 senior grade employees having\n           ESOPs is 0.0606.                                                                            (2)\n\n        iii) The number of claims per year arising from Group-A of policies has a Poisson\n             distribution with mean ‘mu’. The number of claims arising from Group-B of policies has\n             a Poisson distribution with mean 3 times ‘mu’.\n\n        A sample of 100 Group-A policies resulted in 12 claims in a year and a sample of 80 Group-\n        B policies resulted in 18 claims in a year. Determine the maximum likelihood estimate\n        (MLE) of ‘mu’ based on this information by selecting the option\n\n        A. 0.088\n        B. 0.120\n        C. 0.167\n        D. 0.225\n        E. 0.367\n\n        iv) A student has calculated MLE of the parameters of a lognormal distribution. Derive MLE\n            of the mean and variance of the lognormal distribution if ‘mu cap’=1 and ‘sigma square\n            cap’ = 0.5.                                                                                (3)\n\n        v) You are told that MLE of lambda is 0.1 based on sample of 20 values with total delay of\n           200 (resulting in X bar = 10).\n\n        MLE of 0.1 is obtained by setting first derivative of log likelihood with respect to lambda\n        equal to zero.\n\n        Second derivative of log likelihood with respect to lambda is negative and was found to be\n        -n/lambda2.\n\n        a) Calculate large-sample approximate variance (i.e. CRLB) of ‘lambda cap’.                    (2)\n\n        b) Further, calculate an approximate 95% confidence interval for ‘lambda’.                     (2)\n\n        c) Calculate an exact 95% confidence interval for ‘lambda’ using chi square result and         (4)\n           comment on the result in comparison to part (b) above.\n\n        You are told that\n\n        2*lambda*n*X bar ~ Chi square distribution with 2 times n degrees of freedom\n        24.43 and 59.34 are chi square 2.5% critical points",
          "topic": null
        }
      ],
      "solution": "i)    a) An estimator is said to be consistent when\n             mean square error tends to zero\n             as ‘n’ tends to infinity\n             where ‘n’ is sample size                                                                         (1)\n\n          b) A good estimator is one that\n              has small mean square error\n              is unbiased and\n              is consistent                                                                               (1)\n   ii)    a) probability of finding 2 senior grade employees having ESOPs is 0.3637 (i.e., W3)\n          W1 and W2 are more extreme scenarios than W3.\n          Hence, p-value of finding 2 senior grade employees is P(W1) + P(W2) + P(W3)\n          =0.0606+0.1212+0.3637 =0.5455\n          (Or p-value of W3 can be found as 1- P(W4) = 1-0.4545 =0.5455)                                       (2)\n\n          b) Required probability is 3C3 * 8C2 / 11C5\n          =1*(8*7/1*2) / ((11*10*9*8*7/1*2*3*4*5)\n          =4*7/(11*7*3*2)\n          =2/33\n          =0.0606                                                                                              (2)\n\n   iii)   Option A\n          (12+18)/ (100+3*80) = 30 / 340 = 0.088                                                               (3)\n\n   iv)    If ‘m’ is the mean of the lognormal distribution then by invariance property,\n          ‘m cap’ = e^(‘mu cap’ + ½ * ‘sigma square cap’)\n          =e(1.25) = 3.49\n\n          If ‘var’ is the variance of the lognormal distribution then by invariance property,\n          ‘var cap’ = e^(2*‘mu cap’ + ‘sigma square cap’) * (e^(‘sigma square cap’) -1)\n          =3.49^2 *(e(0.5)-1)\n\n                                                                                                 Page 4 of 9\n\f  IAI                                                                                                CS1A-0921\n             =12.1825*0.6487\n             =7.903                                                                                                (3)\n\n        v)\n             a) CRLB = -1/E[second derivative of log likelihood with respect to lambda]\n                = 1/E[n/lambda^2]\n                =lambda^2/n\n                =0.01/20\n                =0.0005                                                                                            (2)\n\n             b) ‘lambda cap’ ~ N(lambda, CRLB) approximately. Hence confidence interval is given by\n                (‘lambda cap’ -1.96 * sqrt(CRLB), lambda cap’ + 1.96 * sqrt (CRLB)\n                =(0.1-1.96*sqrt(0.0005), 0.1+1.96*sqrt(0.0005))\n                =(0.056173,0.143827)                                                                               (2)\n\n             c) Using, 2*lambda*n*X bar ~ Chi square distribution with 2*n degrees of freedom\n                 40*lambda*X bar ~ chi square distribution with 40 degrees of freedom\n                P(24.43<40*lambda*Xbar<59.34) =0.95\n                Hence 95% confidence interval for lambda is\n                (24.43/(40*10), 59.34/(40*10))\n                = (0.061075,0.14835)\n\n                Confidence interval using chi square result / exact result is narrower (i.e.\n                better)compared to result in part b.\n                Result in part b is impacted due to smaller sample size.\n                Larger sample could have resulted in better / narrower interval in part b                          (4)",
      "has_math": false,
      "session": "2021-09",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2021-09_QP.pdf",
      "source_sol": "raw/CS1A_2021-09_SOL.pdf"
    },
    {
      "q_num": 7,
      "marks": 15,
      "topic": "bayes_credibility",
      "subtopics": [
        "distributions"
      ],
      "stem": "",
      "parts": [
        {
          "label": "i",
          "marks": 4,
          "text": "The annual number of claims arising from a group of policies follows a Poisson\n             distribution with mean ‘mu’. The Prior distribution of ‘mu’ is Ga (alpha, lambda). In\n             previous 5 years, the number of claims arising from the group were 3. Determine the\n             posterior distribution of ‘mu’:\n\n        A. Ga (alpha+3, lambda+5)\n        B. Ga (alpha+5, lambda+3)\n        C. Ga (alpha, lambda+3)\n\n        D. Ga (alpha, lambda+5)\n        E. None of the above                                                                                  (2)\n\n        ii)       If the prior distribution of ‘mu’ is Chi square distribution with 10 degrees of freedom\n                  and likelihood is Poisson with ‘x’ claims in 12 months, then identify the posterior\n                  distribution of ‘mu’.                                                                       (1)\n\n        iii)      You are required to estimate ‘mu’ under all-or-nothing loss using posterior\n                  distribution as identified in part i.\n\n               Identify the correct option based on following statements-\n\n              Your friend (F1) thinks that ‘mu’ would have been higher if observations would have\n               been taken over longer time horizon irrespective of number of observed claims.\n              Friend (F2) thinks that ‘mu’ would have been higher if more claims would have been\n               observed over the same time horizon.\n\n        A. F1 is definitely correct\n        B. F2 is definitely correct\n        C. Both F1 and F2 are definitely correct\n        D. Both F1 and F2 are definitely wrong                                                                (1)\n\n        iv)       Posterior distribution of theta is Beta (5,15) distribution. Identify the correct option\n                  indicating Bayesian estimate under all-or-nothing loss function:\n\n        A. 5/15\n        B. 5/20\n        C. 4/19\n        D. 4/18\n        E. None of the above\n\n        v)        Total claim amount per annum on a particular insurance policy follows a normal\n                  distribution with unknown mean ‘theta’ and variance 1502. Prior beliefs about theta\n                  are described by normal distribution with mean 500 and variance 1002. Claim\n                  amounts x1, x2, …, xn are observed over n years.\n\n               a) State A, B, C and D for the posterior distribution of ‘theta’ which can be written in\n                  the form of Normal((A+B)/(C+D),1/(C+D)) using values provided above.                        (2)\n\n               b) Express mean of the posterior distribution of ‘theta’ in terms of credibility estimate\n                  by clearly defining Z, prior mean and MLE of ‘theta’.                                       (2)\n\n               c) Comment on change in value of Z, using Z as derived in part b above:\n                   if prior variance was 1502\n                   if likelihood variance was 1002                                                           (4)",
          "topic": null
        }
      ],
      "solution": "i)        Option A                                                                                              (2)\n\n   ii)       Chi square with 10 df can be written as Ga(5,0.5) distribution.\n             Hence, posterior distribution would be Ga(5+x, 1.5) using results from part 1                         (1)\n\n   iii)      Option B                                                                                              (1)\n\n   iv)       Option D                                                                                              (3)\n             (working is not required)\n             Bayesian estimate under all-or-nothing loss is mode of the distribution.\n             Differentiating the log of the posterior distribution and equating with zero, we get,\n             4/p-14/(1-p) =0\n             Hence, 4(1-p)-14p=0\n             4-18p=0. Hence, p=4/18\n\n                                                                                                     Page 5 of 9\n\f   IAI                                                                                             CS1A-0921\n   v)     a) Posterior distribution of theta can be written as, Normal((A+B)/(C+D),1/(C+D))\n          Where,\n          A = (n*x bar)/150^2\n          B =500/100^2\n          C =n/150^2\n          D=1/100^2                                                                                              (2)\n\n          b) Mean of the posterior distribution is (A+B) /(C+D)\n          This can be written in the forms of\n          [(A/(x bar)) / (C+ D) ]* (x bar) + ((B/500)/ (C+D)) * 500\n          i.e. (C/(C+D) )* (x bar) + (D/(C+D) * 500\n\n          i.e., Z* x bar + (1-Z) * 500 which is a credibility estimate\n\n          where Z = (n/150^2)/((n/150^2) + (1/100^2))\n          MLE of ‘theta’ = x bar\n          Prior mean = 500                                                                                       (2)\n\n          c) Impact on Z\n\n          If prior variance was 150^2 (instead of 100^2) –\n                 this will lead to reduction in denominator and Z will increase.                          (0.75)\n                 Increase in prior variance means that prior is less reliable and hence we need to rely\n                more on data and hence Z will increase.                                                    (1.25)\n\n          if likelihood variance was 1002 (instead of 150^2) –\n                 Z will increase with more increase in numerator compared to denominator.                 (0.75)\n\n                Reduction in variance of the observed data means data is more reliable and hence\n               more weight can be given to it and hence Z increases.",
      "has_math": false,
      "session": "2021-09",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2021-09_QP.pdf",
      "source_sol": "raw/CS1A_2021-09_SOL.pdf"
    },
    {
      "q_num": 8,
      "marks": 21,
      "topic": "regression_glm",
      "subtopics": [
        "inference"
      ],
      "stem": "You are an actuarial student working for a large life insurance company in India. Appointed\n        Actuary has asked you to analyse the claims data to understand the impact of age on number\n        of COVID claims. You have received the following information from claims department.\n\n         Age(X)     Number of COVID claims per 10,000 policies(Y)\n           5                           101\n           15                          120\n           25                          135\n           35                          186\n           45                          268\n           55                          540\n           65                          620\n\n        ∑(xi-x̅)2 =2,800      ∑(yi-𝑦)2 = 2,70,832      ∑((xi-x̅) (yi-𝑦)) =25,300\n\n        You have been asked to perform linear regression analysis on the data to identify the\n        relationship between age and number of COVID claims.",
      "parts": [
        {
          "label": "i",
          "marks": 3,
          "text": "Determine the fitted regression line with ‘no. of claims’ as the response and ‘age’ as the    (3)\n             explanatory variable.",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 3,
          "text": "Assuming the full normal model, calculate the estimate of the error variance 𝜎 and             (3)\n            obtain a 90% confidence interval for 𝜎 .",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 3,
          "text": "Calculate the proportion of variance explained by the model. Hence, comment on the            (3)\n             fit of the model.",
          "topic": null
        },
        {
          "label": "iv",
          "marks": 3,
          "text": "By considering the slope parameter, formally test whether the data is positively               (3)\n            correlated.",
          "topic": null
        },
        {
          "label": "v",
          "marks": 4,
          "text": "Calculate 95% confidence interval for mean predicted number of COVID claims                     (4)\n           corresponding to age of 60.",
          "topic": null
        },
        {
          "label": "vi",
          "marks": 3,
          "text": "Assess the fitness of the model by:\n\n               a) Completing the table of residuals (nearest rounded number)\n                    Age      5 15 25 35 45 55 65\n                    Residual 91   -56 -95  78                                                              (2)\n\n               b) Comment on appropriateness of linear model.                                              (3)",
          "topic": null
        }
      ],
      "solution": "i)    The sums of squares as given in the question\n         Sxx = ∑(xi-x̅)2 = 2, 800             Sxy = ∑((xi-x̅) (yi-𝑦̅) =25,300                               (0.5)\n              ∑ 𝑥𝑖               ∑ 𝑦𝑖                                                                         (1)\n          𝑥̅ = 𝑛 = 35 ,     𝑦̅ = 𝑛 = 281\n             𝑆𝑥𝑦   25,300\n          𝛽̂ = 𝑆 = 2,800 = 9.04                                                                               (1)\n               𝑥\n          𝛼̂ = 𝑦̅ - 𝛽̂ 𝑥̅ = 281 – 9.04 × 35 = -34.82                                                        (0.5)\n          Hence the fitted regression line of y on x is y = -34.82 + 9.04x\n\n   ii)    Syy = ∑(yi-𝑦̅)2 = 2,70,832\n\n                                                                                                   Page 6 of 9\n\fIAI                                                                                                   CS1A-0921\n                                  2\n                  1              𝑆𝑥𝑦        1                  25,3002                                          (1)\n       𝜎̂2 = 𝑛−2(Syy - 𝑆 ) = 5( 2,70,832 - 2,800 ) = 8,445.69\n                                  𝑥𝑥\n\n                  ̂2\n                 5𝜎\n       Now 𝜎2 ~ chi square 𝑥52 which gives a confidence interval for 𝜎 2 of:\n       5×8,445.69 5×8,445.69\n       (                   ,                    ) = (3,814.67, 36,880.74)                                       (2)\n               11.07              1.145\n\niii)\n       The proportion of the variability explained by the model is given by:\n                   2\n                  𝑆𝑥𝑦                  25,3002\n       R2 = 𝑆                  = 2800×2,70,832 = 84%                                                                (2)\n                 𝑥𝑥𝑆𝑦𝑦\n\n       84% of the variance is explained by the model , which indicate that the fit is fairly good. It\n       is still might be worthwhile to examine the residuals to double check that a linear model is\n       appropriate.\n\niv)    Testing:\n       H0 : 𝛽 = 0 vs H1 : 𝛽 > 0\n                   ̂− 𝛽\n                   𝛽\n       Now             2\n                                  ~t5\n                  ̂ ⁄\n                 √𝜎                                                                                                 [1]\n                     𝑆     𝑥𝑥\n\n       The observed value of test statistics is\n               9.04−0\n                                  = 5.20\n       √8445.63⁄2800                                                                                                [1]\n       This exceeds the 0.5% critical value of the t5 distribution of 4.032. So we have sufficient\n       evidence at the 0.5% level to reject H0 and the conclusion is that 𝛽 > 0 hence the data\n       are positively correlated                                                                                (1)\n\nv)     The variance of the distribution of the mean number of COVID claims corresponding to an\n       entry age of 60 is:\n           1   (𝑥0 − 𝑥̅ )2              1       (60− 35)2\n       [𝑛 +                    ] 𝜎̂2 = [7+                  ] × 8445.69 = 3,091.72\n                 𝑆𝑋𝑋                              2800\n\n       The predicted value of number of COVID claims corresponding to age 60 is\n       -34.82 + 9.04× 60 = 507.32\n\n       We have t5 distribution. Hence the 95% confidence interval is\n       507.32 ± 2.571 × √3091.72 = (364.37, 650.28)                                                             (2)\n\nvi)    a) The completed table of residuals are as follows\n        Age       5         15         25          35                                45     55   65\n        Residual  91        19         -56         -95                               -104   78   67                 (2)\n\n                                                                                                      Page 7 of 9\n\f  IAI                                                                                                 CS1A-0921\n          b) Clearly the trend of residual with progression of age is not pattern less. The residuals               (3)\n             are not independent of the age. This means that the linear model is missing something\n             and is not appropriate to these data",
      "has_math": true,
      "session": "2021-09",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2021-09_QP.pdf",
      "source_sol": "raw/CS1A_2021-09_SOL.pdf"
    },
    {
      "q_num": 9,
      "marks": 10,
      "topic": "regression_glm",
      "subtopics": [
        "distributions"
      ],
      "stem": "",
      "parts": [
        {
          "label": "i",
          "marks": 4,
          "text": "State three components of Generalised Linear Model (GLM).                                     (2)\n\n        ii) Show that exponential distribution with density function f(y) =          𝑒   , y>0 can be\n             written in the form of exponential family of distribution.                                    (1)\n\n        iii) Further, determine.\n\n               a) canonical link function                                                                  (1)\n\n               b) variance function                                                                        (1)\n\n               c) dispersion parameter                                                                     (1)\n\n      Medical claim amounts of a health insurance company are believed to have an exponential\n      distribution with mean µi\n\n                                           −𝑦\n                                  f(yi) = e   𝜇\n\n      The following data of medical insurance claims is available\n\n       Age                        30    35      40        45       50\n       Claim amount(Rs ‘000)      50    80      150       250      360\n\n      The insurer believes that a linear function of age affects the claim amount i.e. ɳ = α + βxi\n\n      wherein\n      xi is the age of policyholder i\n      ɳ is the claim amount of policyholder i\n      α and β are constants\n\n      iv)    Using the canonical link function, identify which one of the following equations\n             gives the correct maximum likelihood estimates for α and β based on the above\n             data.\n\n                A:\n                        +        +        +           +         – 890 = 0\n                        +        +        +           +     – 39550 = 0\n\n                B:\n                        +        +        +           +         – 890 = 0\n\n                        +        +        +           +     – 39550 = 0\n\n                C:\n                        +        +        +           +         – 890 = 0\n                        +        +        +           +     – 39550 = 0\n\n                D:\n                        +        +        +           +         – 890 = 0\n\n                        +        +        +           +         – 39550 = 0\n\n                E:\n                None of the above",
          "topic": null
        }
      ],
      "solution": "i)     a) A distribution of the response variable Y\n          b) A “linear predictor” ɳ\n          c) A ”link function” g                                                                                    (2)\n   ii)    The PDF of exponential distribution can be written as\n                       −𝑦\n                   1                     𝑦\n          f(y) = 𝜇 𝑒 𝜇 = exp{− 𝜇 − 𝑙𝑜𝑔𝜇}\n\n          Comparing the above with the standard PDF of exponential family of distribution\n          𝜃 = −1/𝜇, b(𝜃) =log𝜇 = -log(-𝜃) , ∅ = 1, a(∅) = ∅ and c(y, ∅) = 0                                         (1)\n\n   iii)\n                                                                      1\n          a) The canonical link function from part (ii) that 𝜃 = − 𝜇                                                (1)\n\n          b) The variance function is 𝑏 ′′ (𝜃). Differentiating b(𝜃) twice , 𝑏 ′′ (𝜃) = 1/𝜃 2 = 𝜇 2\n                So the variance function is 𝜇 2                                                                     (1)\n          c) The dispersion parameter or scale parameter is ∅ = 1                                                   (1)\n\n   iv)    The log of the likelihood function is\n                          𝑦\n          Log L(𝜇𝑖 ) = -∑ 𝜇𝑖 - ∑ log 𝜇𝑖\n                                    𝑖\n          The canonical link function for the exponential distribution is g(𝜇𝑖 ) = 1/𝜇𝑖 .\n          The canonical link function connects the mean response to the linear predictor , g(𝜇𝑖 ) =\n           ɳ𝑖\n          Hence we have\n          1\n               = α + βxi\n          𝜇𝑖\n          The log likelihood function in terms of α and β:\n          Log L(α, β) = ∑ 𝑦𝑖 (𝛼 + 𝛽𝑥𝑖 ) + ∑ log( α + βxi)\n          Differentiating the above equation with respect to 𝛼 𝑎𝑛𝑑 𝛽:\n          𝜕                                  1\n               log L(𝛼, 𝛽) = - ∑ 𝑦𝑖 + ∑ 𝛼+𝛽𝑥\n          𝜕𝛼                                     𝑖\n           𝜕                                     𝑥𝑖\n               log L(𝛼, 𝛽) = - ∑ 𝑥𝑖 𝑦𝑖 + ∑ 𝛼+𝛽𝑥\n          𝜕𝛽                                          𝑖\n          The equations satisfied by the MLEs of 𝛼 and 𝛽 are\n                            1\n          - ∑ 𝑦𝑖 + ∑ 𝛼̂+𝛽̂𝑥 = 0\n                                𝑖\n                           𝑥\n          - ∑ 𝑥𝑖 𝑦𝑖 + ∑ 𝛼̂+ 𝛽̂𝑖 𝑥       =0\n                                  𝑖\n\n                                                                                                      Page 8 of 9\n\fIAI                                                                                        CS1A-0921\n      Substituting in the given data values gives the following equations\n        1           1         1        1         1\n      ̂ +30𝛽̂   + 𝛼̂+35𝛽̂ + 𝛼̂+40𝛽̂ + 𝛼̂+45𝛽̂ + 𝛼̂+50𝛽̂ – 890 = 0\n      𝛼\n         30         35       40        45       50\n      ̂ +30𝛽̂   + 𝛼̂+35𝛽̂ + 𝛼̂+40𝛽̂ + 𝛼̂+45𝛽̂ + 𝛼̂+50𝛽̂ – 39550 = 0\n      𝛼\n      (The above is for information purpose only. Students are not expected to provide the above\n      derivation in the answer script)                                                                [4]\n      Correct answer is Option C\n\n                                             **********************\n\n                                                                                           Page 9 of 9",
      "has_math": true,
      "session": "2021-09",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2021-09_QP.pdf",
      "source_sol": "raw/CS1A_2021-09_SOL.pdf"
    },
    {
      "q_num": 1,
      "marks": 8,
      "topic": "distributions",
      "subtopics": [],
      "stem": "Assume for a health insurance policy the number of claims follows a Poisson process with\n        a rate of 0.3 per year.",
      "parts": [
        {
          "label": "i",
          "marks": 1,
          "text": "Determine the probability that no claim arises in the policy in one year.                       (1)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 2,
          "text": "For the policy determine the probability that, out of four consecutive years, there are one\n            or more claim(s) in two of the years and no claim in the remaining two years.                  (2)",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 2,
          "text": "Assuming that a claim has just occurred, determine the probability that more than three\n             year will elapse before occurrence of the next claim.                                         (2)",
          "topic": null
        },
        {
          "label": "iv",
          "marks": 1,
          "text": "For the health insurance policy, provide the steps to be followed in order to simulate an\n            observation of N number of claims, occurring in one year.                                      (1)",
          "topic": null
        },
        {
          "label": "v",
          "marks": 2,
          "text": "Using the following two random numbers (uniformly distributed between 0 and 1)\n           simulate the number of claims in one year for the above health insurance policy.\n\n           a) 0.7521\n           b) 0.6512                                                                                       (2)",
          "topic": null
        }
      ],
      "solution": "i)        The distribution of the number of claims, N, in one year, is Poi(0.3). Hence the probability of no\n             claims in one year is\n                      0.30\n             P(N=0) = 0! 𝑒 −0.3 = 0.740818\n   ii) The probability of one or more claims in one year is\n\n             P(N>=1) = 1-P(N=0) = 1 – 0.740818 = 0.259182\n\n             If X is the number of years with one or more claim, then\n\n             X ~ Bin( 4, 0.259182)\n\n             So we have:\n                 P(X=2) = 4C2 * 0.2591822 *(1-0.259182)2 = 0.221199\n\n   iii) The waiting time , T, in years follows Exp(0.3) distribution\n        P( T >3) = 1 – F[3] = 𝑒 −0.3∗3 = 0.40657\n\n   iv) To simulate a value from a discrete distribution, the following steps would be followed\n\n        1. Calculate DF, F(n)\n        2. For a random number r between 0 and 1, if F(n-1) <r <= F(n), then the simulated value is n\n\n   v) The CDF is\n\n   N                            0                             1                         2\n   P(N<=n)                      0.74                          0.96                      0.99\n\n   Since 0.74< 0.7521 <= 0.96, the first simulated value is 1. Since 0< 0.6521 < 0.74 the second\n   simulated value is 0",
      "has_math": false,
      "session": "2022-03",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2022-03_QP.pdf",
      "source_sol": "raw/CS1A_2022-03_SOL.pdf"
    },
    {
      "q_num": 2,
      "marks": 11,
      "topic": "inference",
      "subtopics": [
        "distributions"
      ],
      "stem": "Following measures are provided for 2 independent random samples.\n\n         Sample Sample size         ∑xi          ∑xi2\n           1       15               297          5970\n           2       15               297          6200\n\n        Sample Mean for both the samples = 19.8",
      "parts": [
        {
          "label": "i",
          "marks": 1,
          "text": "Calculate sample variance for both the samples.                                                 (1)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 2,
          "text": "Assuming that the samples are drawn from a normally distributed population\n\n           a) Carry out t-tests of the hypothesis H0: µ =18 vs H1: µ ≠ 18 at 98% level of confidence\n              for sample 1.                                                                                (2)\n\n           b) Calculate confidence interval for sample 2 to test the null hypothesis stated in sub-\n              part ii(a).                                                                                  (2)\n\n           c) Discuss the reasons for differences and similarities while comparing the test results\n              of Sample 1 and Sample 2.                                                                    (2)\n\n           d) A new 100 sized sample 3 is collected from the same population.\n              Explain how the width of the confidence interval of sample 3 will differ from\n              confidence interval of sample 2.                                                             (2)\n\n           e) It was observed that one of the claims in the sample 2 has an extremely large value\n              and can be considered as an outlier. The value is now replaced with a new randomly\n              selected one, which is not an outlier anymore.\n              For the updated sample, explain how the confidence intervals will differ from that of\n              sample 2 calculated in sub-part ii(b).                                                       (2)",
          "topic": null
        }
      ],
      "solution": "1\n        i)    Sample variance = 𝑛−1 (∑ 𝑥𝑖2 − 𝑛𝑥̅ 2 )\n\n              Sample                              Sample Variance\n              1                                   6.385\n              2                                   22.814\n\n  ii)\n                                𝑋̅−𝜇\n         a) Test Statistic :             ~ tn-1\n                               √𝑆 2 /𝑛\n         Sample 1\n         t = (19.8 -18) /sqrt(6.385/15) = 2.758\n         P-value = 2*P(t14 >2.68) = 0.0153 < .02\n         We reject null hypothesis and accept H1: µ ≠ 18 at 98% level of confidence                        (2)\n         b) Sample 2\n\n                                                                                                  Page 2 of 10\n\f    IAI                                                                                                  CS1A-0322\n\n          Confidence interval is given by\n                                                           𝑋̅ ∓ 𝑡.01,14 ∗ √𝑠 2 /𝑛\n\n          19.8 ∓ 2.6245 * sqrt(22.814/15) = (16.56 , 23.04)\n\n          Since 18 lies between the confidence interval, we cannot reject null hypothesis at 98% level of\n          confidence                                                                                 (2)\n\n          c) Sample 2 does not provide enough evidence to justify rejecting H0, despite having the same\n              size and mean to Sample 1.\n          The reason for the loss of significance is the much greater variation in the data in Sample 2 – the\n              variance is three times bigger than in Sample 1 (22.814 v 6.386)\n          – this greatly increases the standard error of estimation.                                      (2)\n\n          d) With the larger sample of 100 claims the standard error of the sample mean will be smaller,\n             giving a narrower confidence interval.                                                 (2)\n\n          e)    The replacement of the extreme value will lead to\n          -     [Bonus]Sample mean, which means that the interval will be shifted to the left.\n          -     Reduction of variance of the sample will also be smaller with replacement of outlier value,\n          -     which will again give a narrower interval.",
      "has_math": true,
      "session": "2022-03",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2022-03_QP.pdf",
      "source_sol": "raw/CS1A_2022-03_SOL.pdf"
    },
    {
      "q_num": 3,
      "marks": 7,
      "topic": "distributions",
      "subtopics": [],
      "stem": "The claim amounts X and Y (in units of INR 1000) for two different types of insurance\n        policy are modelled using a gamma distribution with parameters 𝛼 = 5, λ = 1/8 and 𝛼 = 3, λ\n        = 1/4 respectively i.e. X ~ Gamma(5,1/8), Y ~ Gamma(3,1/4). Assume that X and Y are\n        independent of each other.",
      "parts": [
        {
          "label": "i",
          "marks": 1,
          "text": "Identify which one of the following options describes the moment generating function\n           of X\n\n           A. (1-8t)^(-4)\n           B. (1-t/8)^(-4)\n           C. (1-8t)^(-5)\n           D. None of the above                                                                        (1)\n\n                                                     1   2",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 2,
          "text": "Use moment generating function to show 4 𝑋 ~𝜒10                                            (2)",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 2,
          "text": "Calculate the probability that the claim amount, X, exceeds INR 40,000.                   (2)",
          "topic": null
        },
        {
          "label": "iv",
          "marks": 2,
          "text": "An analyst argues that sum of X and Y must follow Gamma(8,3/8) i.e. X + Y ~\n            Gamma(8,3/8)\n\n           Comment on the analyst’s argument using moment generating function.                         (2)",
          "topic": null
        }
      ],
      "solution": "i)         Option C                                                                                         (1)\n\n                  X ~ Gamma(5,1/8)\n                                          𝑡\n                  From Tables, MX(t) = (1-𝜆 )^(-𝛼) = (1-8t)^(-5)\n\n                   1\n    ii)        Y = 4 𝑋 , MY(t) = E[ etY] = E[e(1/4)*(t*X)] = MX(t/4) = (1- 8(t/4))^(-5) = (1-2t)^(-5)\n\n    So the moment generating function of Y is (1-2t)^(-5)\n\n    By comparing this with the MGF of the Gamma distribution in the tables, we see that this is the\n    same as the MGF of Gamma(5, ½) distribution. Looking at the definition of the chi-square\n    distribution, we see that Gamma(5,1/2) is equivalent to chi-square distribution with 10 degrees of\n    freedom\n\n                                                                                                   2       1\n    By the uniqueness property of moment generating function , therefore we have shown that 4 𝑋 ~𝜒10\n\n    iii)       X ~ Gamma(5,1/8) where X is the claim amount in units of 1000 INR\n\n           2                2\n    2𝜆X ~ 𝜒2𝛼 i.e (1/4)X ~ 𝜒10\n\n                               2\n    P(X> 40) = (X/4 > 10) = P(𝜒10 >10) = 1 – 0.5595 = 0.4405                                                    (2)\n\n    (iv) Since X and Y are independent. The MGF of X + Y is given by the product of MGFs.\n      MX+Y(t) = E[etX] E[etY] = (1-8t)^(-5) * (1-4t)^(-3)\n      MZ(t) = (1-8t/3)^(-8)\n\n                                                                                                        Page 3 of 10\n\f   IAI                                                                                          CS1A-0322\n\n    MZ(t) ≠ MX+Y(t), and therefore X + Y does not have a Gamma( 8,3/8) distribution                    (2)",
      "has_math": true,
      "session": "2022-03",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2022-03_QP.pdf",
      "source_sol": "raw/CS1A_2022-03_SOL.pdf"
    },
    {
      "q_num": 4,
      "marks": 9,
      "topic": "bayes_credibility",
      "subtopics": [],
      "stem": "An analyst is working on renewals of health insurance policies. He has estimated following\n        aggregate claim statistics for four policies over last 3 years:\n\n         Policy Number                       3                               3\n\n                                            ∑ 𝑋𝑗                            ∑(𝑋𝑗 − 𝑋̅)\n                                            𝑗=1                             𝑗=1\n                1                           12,183                               12,504\n                2                           13,098                               12,718\n                3                           12,822                               12,432\n                4                           13,453                               12,242",
      "parts": [
        {
          "label": "i",
          "marks": 4,
          "text": "Using EBCT Model 1, compute the credibility premium of Policy Number 4 for the\n           upcoming renewal. Students may refer actuarial table for the formula.                       (4)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 3,
          "text": "The medical inflation in last few years is very high and increasing at a rate of 20% per\n            year on average. Describe the adjustment to be made in the above estimation (as\n            determined in sub part i) in order to incorporate the medical inflation. Also, state any\n            additional data that will be required.                                                     (3)",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 2,
          "text": "Explain, without performing any further calculation, what change is expected to the\n             credibility factor (determined in sub-part i), if Policy Number 1 is excluded.            (2)",
          "topic": null
        }
      ],
      "solution": "i) Mean and Variances of policies\n\n                Policy Number (i)                             𝑥̅𝑖                        𝑠𝑖2\n                        1                                   4,061                      6,252\n                        2                                   4,366                      6,359\n                        3                                   4,274                      6,216\n                        4                                   4,484                      6,121\n\n               E[m(θ)] = Overall mean = 4296.25\n               E[s2(θ)] = Mean of the variances = 6237\n\n                                  1                         1 1        1\n                   Var[m(θ)] = 3 ∑4𝑖=1(𝑥̅𝑖 − 𝑥̅ )2 − 3 [4 ∑4𝑖=1 2 ∑3𝑗=1(𝑥𝑖𝑗 − 𝑥̅𝑖 )]\n                      1\n                    = 3 [(4061 − 4296.25)2 + (4366 − 4296.25)2 + (4274 − 4296.25)2 + (4484 −\n                                      1\n                    4296.25)2 ] − 3 × 6237\n\n                  = 29,905\n\n                                                    3\n    The estimated credibility factor is Z =          6237    = 0.935\n                                               3+\n                                                    29905\n\n    Thus, the credibility premium for Policy Number 4 is\n    = 0.935 x 4,484 + 0.065 x 4296.25 = 4,471.8\n\n    ii) EBCT Model 2 can be used with risk volumes P(i,j) for each policy number for each year.\n    P(i,1) = 1 , P(i,2)=1/1.2 and P(i,3)=1/1.2^2\n\n    (Alternate solution: Adjust year-wise figures to bring it year 3 level. Multiply 1.2^2 to year 1 and 1.2\n    to year 2. And then apply the EBCT Model 1 to adjusted on-level figures.)\n\n    Data Required: Year wise data will be required since only totals are given.                         (3)\n\n    iii) Policy number 1 has a lowest mean. Removing it will reduce the variance of means, var[m(θ)].\n    The variance of policy number is similar to other policies and close to overall mean of variance.\n\n    Thus, 𝐸(𝑠 2 (𝜃)] will remain similar and won’t change much.\n    Hence, proportionately smaller var[m(θ)] will lead to reduction of credibility factor.",
      "has_math": true,
      "session": "2022-03",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2022-03_QP.pdf",
      "source_sol": "raw/CS1A_2022-03_SOL.pdf"
    },
    {
      "q_num": 5,
      "marks": 10,
      "topic": "distributions",
      "subtopics": [],
      "stem": "The joint probability density function of random variables X and Y is\n                               𝑦\n                          −(3𝑥+ )\n         𝑓(𝑥, 𝑦) = {𝑘𝑒         5    ,   𝑥 > 0, 𝑦 > 0\n                     0,                  𝑂𝑡ℎ𝑒𝑟𝑤𝑖𝑠𝑒",
      "parts": [
        {
          "label": "i",
          "marks": 3,
          "text": "Determine the value of k.                                                                        (3)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 2,
          "text": "Identify the right marginal density function of X and Y from the following expressions\n                                                𝑦\n            A. fX(x) = 1/3𝑒 −3𝑥 , fY(y) = 5*𝑒 −5\n                                                𝑦\n            B. fX(x) = 3𝑒 −3𝑥 , fY(y) = 1/5*𝑒 −5\n                               𝑥                  𝑦\n            C. fX(x) = 1/3𝑒 −3 , fY(y) = 1/5*𝑒 −5\n            D. fX(x) = 3𝑒 −3𝑥 , fY(y) = 5*𝑒 −5𝑦\n            E. None of the above                                                                             (2)",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 1,
          "text": "State whether X and Y are independent based on your answer in part (ii).                       (1)",
          "topic": null
        },
        {
          "label": "iv",
          "marks": 3,
          "text": "Determine the conditional density function f(x|X>5).                                            (3)",
          "topic": null
        },
        {
          "label": "v",
          "marks": 1,
          "text": "Identify which one of the following expressions is equal to the conditional expectation\n            of E[x|X>5]\n\n            A. ∫0∞ 15𝑡 𝑒−3𝑡dt + ∫0∞ 3 𝑒−3𝑡dt\n            B. ∫0∞ 3 𝑒−3𝑡dt + ∫0∞ 15𝑡 𝑒−3𝑡dt\n            C. ∫0∞ 15(𝑡 ^2)𝑒−3𝑡 dt + ∫0∞ 3 𝑒−3𝑡dt\n            D. ∫0∞ 3𝑡 𝑒−3𝑡dt + ∫0∞ 15 𝑒−3𝑡dt\n            E. None of the above                                                                             (1)",
          "topic": null
        }
      ],
      "solution": "i)    The integral over the domain\n           ∞\n         ∬0 𝑓(𝑥, 𝑦)𝑑𝑥𝑑𝑦 = 1\n\n           ∞                          ∞ −(3𝑥+𝑦)                   ∞        ∞   𝑦\n         ∬0 𝑓(𝑥, 𝑦)𝑑𝑥𝑑𝑦 = 𝑘 ∬0 𝑒              5   𝑑𝑥𝑑𝑦 = k∫0 𝑒 −3𝑥 𝑑𝑥 ∫0 𝑒 −5 dy\n\n                                                                                               Page 4 of 10\n\f   IAI                                                                                       CS1A-0322\n\n          ∞                𝑒 −3𝑥 ∞ 1\n         ∫0 𝑒 −3𝑥 𝑑𝑥 =-      3\n                                |0 = 3\n\n          ∞ −𝑦              −\n                             𝑦\n                                    ∞\n         ∫0 𝑒 5 dy = -5 𝑒 5 | 0 = 5\n\n                  1\n         Hence k*3*5 = 1 Hence k = 3/5\n\n   ii) Answer B                                                                                      (2)\n\n   The marginal distribution of X is\n                 ∞ −(3𝑥+𝑦)                              ∞   𝑦\n   fX(x) = 3/5 ∫0 𝑒          5      𝑑𝑦 = 3/5 * 𝑒 −3𝑥 *∫0 𝑒 −5 dy = 3𝑒 −3𝑥\n\n   The marginal distribution of Y is\n                 ∞ −(3𝑥+𝑦)                          𝑦\n                                                        ∞                   𝑦\n   fY(y) = 3/5 ∫0 𝑒             5   𝑑𝑥 = 3/5 *𝑒 −5 * ∫0 𝑒 −3𝑥 𝑑𝑥= 1/5*𝑒 −5\n\n   iii) From the results of point (ii), it is found that\n     𝑓(𝑥, 𝑦) = fX(x) * fY(y)\n\n   Hence joint probability function is the product of two marginal probability functions for all (x,y) in\n   the range of the variables hence X and Y are independent\n\n   iv) The conditional probability P(X<=x|X>5) is\n\n                      𝑃(𝑋≤𝑥,𝑋>5) 𝑃(5<𝑋≤𝑥) FX(x)−FX(5)\n   P(X<=x|X>5) =                =        =            , x>5\n                        𝑃(𝑋>5)    𝑃(𝑋>5)    𝑃(𝑋>5)\n\n   Therefore,\n                 fX(x)     3𝑒 −3𝑥\n   f(x|x>5) = 𝑃(𝑋>5) = 𝑒 −15 = 3𝑒 15−3𝑥 , x>5\n\n                       ∞                            ∞\n   Since P(X>5) = ∫5 3𝑒 −3𝑥 𝑑𝑥 = -𝑒 −3𝑥 | 5 = 𝑒 −15\n\n   v) Option D                                                                                       (1)\n\n   The conditional expectation\n                ∞                 ∞\n   E[X| X>5] = ∫5 𝑥 f(x|x>5)dx = ∫5 𝑥 3𝑒 15−3𝑥 dx\n\n   By taking t = x – 5\n                  ∞                             ∞               ∞\n   E[X| X>5] = ∫0 (𝑡 + 5) 3𝑒 −3𝑡 dt = ∫0 3𝑡 𝑒 −3𝑡 dt + ∫0 15 𝑒 −3𝑡 dt",
      "has_math": true,
      "session": "2022-03",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2022-03_QP.pdf",
      "source_sol": "raw/CS1A_2022-03_SOL.pdf"
    },
    {
      "q_num": 6,
      "marks": 5,
      "topic": "bayes_credibility",
      "subtopics": [],
      "stem": "Aggregate claim amounts of a health insurance portfolio follows a normal distribution with\n         mean μ where prior distribution of μ is N(μ0, σ2). For this distribution a model has been built\n         using n sized sample data with mean 𝑥̅ , variance s2 and credibility factor Z.\n\n         State the relationship (increasing, decreasing or nothing) of the credibility factor with\n             μ0,\n             σ2\n             and s2\n         and provide the reason for the same.                                                                [5]",
      "parts": [],
      "solution": "Credibility factor Z = n / (n + s2/ σ2 )\n   No requirement to specify credibility factor.\n\n   μ0 : No relationship\n\n                                                                                            Page 5 of 10\n\f   IAI                                                                                               CS1A-0322\n\n   1 Mark\n\n   σ2 : Z will be an increasing function of σ2 . Higher variability in prior means less relevance and high\n   spread making the prior less credible for estimation and hence, more credit to sample.\n\n   s2 : Z will be an decreasing function of s2 . High sample variance indicates high spread of sample\n   values over a wide range and hence, less reliable for estimation. Thus, with increase in sample\n   variance less credibility given to sample value.                                         [5 Marks]",
      "has_math": true,
      "session": "2022-03",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2022-03_QP.pdf",
      "source_sol": "raw/CS1A_2022-03_SOL.pdf"
    },
    {
      "q_num": 7,
      "marks": 10,
      "topic": "distributions",
      "subtopics": [],
      "stem": "Housing prices, Xi (in units of 1000) in region X follow a normal distribution with mean\n         500 and standard deviation 10 i.e. Xi ~ N(500, 10^2) for i=1,……..,10.\n         For another region Y the housing prices Yj (in units of 1000) are normally distributed with\n         mean 510 and standard deviation 5 i.e. Yj ~ N(510, 5^2) for j=1,……..,5. Assume that two\n         samples are independent of each other. Let 𝑋̅ and 𝑌̅ denotes the means of the two samples\n         and let 𝑆𝑋2 and 𝑆𝑌2 be the sample variance.",
      "parts": [
        {
          "label": "i",
          "marks": 3,
          "text": "Calculate the probability of “the region X sample mean is greater than the region Y\n           sample mean”.\n\n           A. 0.00561\n           B. 0.00489\n           C. 0.00494\n           D. 0.00572\n           E. None of the above                                                                          (3)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 2,
          "text": "Calculate the probability of sample variance for region X is greater than 100.               (2)",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 3,
          "text": "Calculate the approximate probability for “sample variance of region X is less than the\n             sample variance of region Y”.\n\n           A. 2.5%\n           B. 4.2%\n           C. 6%\n           D. 10%\n           E. None of the above                                                                          (3)",
          "topic": null
        },
        {
          "label": "iv",
          "marks": 2,
          "text": "Calculate the probability of “the region X sample standard deviation is more than 4 times\n            greater than the region Y sample standard deviation”.                                        (2)",
          "topic": null
        }
      ],
      "solution": "i)     Option C                                                                                          (3)\n\n   The distribution of sample means are\n\n   𝑋̅ ~ N(500, 10) 𝑌̅ ~ N(510, 5)\n\n   𝑌̅ - 𝑋̅ ~ N( 510-500, 5+10) ~ N(10,15)\n\n                                      0−10\n   P(𝑋̅ > 𝑌̅) = P(𝑌̅ - 𝑋̅ <0) = P(Z <     ) = P(Z <-2.58) = 1 – P(Z<= 2.58) =1 – 0.99506 = 0.00494\n                                         √15\n\n   ii) P(𝑆𝑋2 > 100) = 1 - P(𝑆𝑋2 <= 100)\n                           9𝑆 2\n                  = 1 – P(100𝑋 <=9)\n                  = 1 – P(𝜒92 <= 9)\n                  = 1- 0.5627 = 0.4373                                                                      (2)\n\n   iii) Option B                                                                                            (3)\n\n                                                        𝑆 2 /100\n   𝑆𝑋2 and 𝑆𝑌2 are independent and therefore 𝑆𝑋2 /25 ~ F9,4\n                                                          𝑌\n            2               𝑆2           𝑆 2 /100   25\n        P[ 𝑆𝑋   < 𝑆𝑌2 ] = P[ 𝑋2 < 1] = P[ 𝑋2      <    ] = P[F9,4 < .25]\n                            𝑆\n                            𝑌             𝑆 /25\n                                          𝑌         100\n\n   P[F9,4 < .25] = P[ F4,9 >1/.25] = P[F4,9 >4]\n\n   The value is between 2.5% and 5%, and interpolating, we find the probability is approximately 4.2%\n\n   iv) P(SX >4SY) = P(SX/ SY > 4) = P( 𝑆𝑋2 / 𝑆𝑌2 >16)\n\n                             𝑆2 / 𝑆2\n   P( 𝑆𝑋2 / 𝑆𝑌2 >16) = P( 𝑋 4 𝑌 > 4) = P( F9,4 > 4)\n\n   From Table the required approximate probability is 10%                                                   (2)\n\n                                                                                                 Page 6 of 10\n\f IAI                                                                                          CS1A-0322",
      "has_math": true,
      "session": "2022-03",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2022-03_QP.pdf",
      "source_sol": "raw/CS1A_2022-03_SOL.pdf"
    },
    {
      "q_num": 8,
      "marks": 10,
      "topic": "bayes_credibility",
      "subtopics": [],
      "stem": "",
      "parts": [
        {
          "label": "i",
          "marks": 3,
          "text": "Define conjugate prior distribution.                                                          (2)\n\n        ii) Assume that X1, X2,………, Xn are random samples from an exponential distribution\n            with parameter λ, where λ is a random variable.\n            Prove that the conjugate prior distribution for λ is a Gamma distribution.                   (3)\n\n        iii) If λ ~ Gamma (α, β),\n\n                               𝑛       𝑛𝛽\n           a) Deduce that E( 𝜆 ) = (𝛼−1) where n is the sample size.                                     (2)\n\n           b) Considering the distribution in part (ii), show that the posterior mean of n/λ can be\n              expressed as\n              Z * ∑𝑛𝑖 𝑥𝑖 + (1- Z ) * prior mean of n/λ                                                   (3)",
          "topic": null
        }
      ],
      "solution": "i) If, taking a sample from the distribution parameterised by λ, the posterior distribution of λ\n         belongs to the same family as the prior distribution then the prior is called a conjugate prior.\n    ii) Assume Prior distribution of λ is Gamma(α, 𝛽).\n          X is the sample taken from exponential distribution then posterior density satisfies:\n              f(λ|X) ∝ f(X|λ)f(λ)\n                                            𝛽 𝛼 𝜆𝛼−1 𝑒 −𝛽𝜆\n                  = [∏𝑛𝑖=1 𝜆𝑒 −𝜆𝑥𝑖 ] ×           𝛤(𝛼)\n\n                                            𝑛\n                  ∝ 𝜆𝛼+𝑛−1 𝑒 −𝜆(𝛽+∑𝑖=1 𝑥𝑖)\n                  ∝ pdf of Γ (α+n, β + ∑𝑛𝑖 𝑥𝑖 )\n\n              This means that the posterior distribution of also follows a Gamma distribution and\n              therefore the Gamma distribution satisfies the definition of a conjugate prior.\n\n       iii)\n              a) Given λ ~ Gamma (α, β)\n\n                  𝑛          ∞ 𝑛𝑓(𝜆)\n               𝐸 (𝜆 ) = ∫0      𝜆\n                                       𝑑𝜆\n                          ∞ 𝑛𝜆𝛼−1 𝑒 −𝛽𝜆\n                      = ∫0                  𝑑𝜆\n                              𝜆𝛤(𝛼)\n                         𝑛𝛽 ∞ 𝛽 𝛼−1 𝜆𝛼−2 𝑒 −𝛽𝜆\n                      = 𝛼−1 ∫0     𝛤(𝛼−1)\n                                               𝑑𝜆\n                         𝑛𝛽    ∞\n                      =\n                        𝛼−1 0\n                             ∫ 𝑝𝑑𝑓 𝑜𝑓 𝑔𝑎𝑚𝑚𝑎 (𝛼 − 1, 𝛽) 𝑑𝜆\n                        𝑛𝛽\n                      =𝛼−1 𝑋 1\n                        𝑛𝛽\n                      =\n                        𝛼−1\n\n              b) Using (ii), Posterior distribution of λ is Gamma (α+ n, β+ ∑𝑛𝑖 𝑥𝑖 ) .\n                     Using (iii a), Posterior mean of n/λ\n                                                        𝑛(𝛽+∑𝑛\n                                                             𝑖 𝑥𝑖 )\n                                                =   𝛼+𝑛−1\n                                                     𝑛𝛽        𝑛 ∑𝑛\n                                                                  𝑖 𝑥𝑖\n                                                =           +\n                                                  (α+ n −1)   (α+ n −1)\n                                                    𝛼−1       𝑛𝛽           𝑛\n                                                =           𝑋     +              𝑋 ∑𝑛𝑖 𝑥𝑖\n                                                  (α+ n −1)   𝛼−1      (α+ n −1)\n\n                                       𝑛\n                  Consider Z = (α+ n −1)\n                                 𝑛𝛽                 𝑛\n                  Using (iii a), 𝛼−1 is 𝐸 (𝜆 )\n                  Thus, posterior mean of n/λ = Z* ∑𝑛𝑖 𝑥𝑖 + (1- Z ) * prior mean of n/λ",
      "has_math": true,
      "session": "2022-03",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2022-03_QP.pdf",
      "source_sol": "raw/CS1A_2022-03_SOL.pdf"
    },
    {
      "q_num": 9,
      "marks": 30,
      "topic": "regression_glm",
      "subtopics": [
        "inference"
      ],
      "stem": "",
      "parts": [
        {
          "label": "i",
          "marks": 2,
          "text": "a) Poisson distribution (e-mu * muy ) / (y!) can be written as exp(y*log(mu)-mu-log y!), in\n           the form of a member of the exponential family.\n\n           Identify correct option indicating Variance function i.e. V(mu)\n\n           A. V(mu) = log(mu)\n           B. V(mu) = (1/mu) -1\n           C. V(mu) = (y/mu) -1\n           D. V(mu) = mu\n           E. V(mu) = 1/mu                                                                               (2)\n\n      b) Identify correct option indicating mean (mu) and variance as a function of ‘mu’ for a\n         particular distribution when written in the form of a member of the exponential family\n         having\n          b(theta) = -log(-theta)\n          theta = -1/mu\n          a(phi) = 1\n\n         A. mean = mu and variance = mu\n         B. mean = mu^2 and variance = mu^2\n         C. mean = mu and variance = mu^2\n         D. mean = mu^2 and variance = mu\n         E. None of the above                                                                         (2)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 4,
          "text": "a) A student modelling accidental hospitalisation claims found that interaction between\n          ‘age band’ and ‘gender’ is significant. Explain what this means (for the model predicting\n          accidental hospitalisation claims) with reference to beta parameter for gender.\n\n         (You are told that males are more prone to accidents than females and hence males\n         require more accidental hospitalisation compared to female across all age groups.)           (4)\n\n      b) A student is fitting GLM model to predict policy renewal rate for a group of policies.\n         Comment on the choice of link function that can be used with suitable expression.            (4)",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 3,
          "text": "Answer the following questions using ANOVA Table as provided\n\n       Source of variation           d.f.    Sum of Squares       Mean Sum of Squares\n       Regression                    1       A                    SSREG\n       Residual                      12      B                    C\n       Total                         13      5.38\n\n      a) Identify A, B and C to complete the above table and calculate F value when\n\n         SSREG =1.38 (for Model 1) and\n         SSREG =2.38 (for Model 2)                                                                    (4)\n\n      b) Comment on the results in part (a) above using tabulated values provided below, clearly\n         stating the hypothesis\n\n       F1,12 table values:\n              1%           2.50%            5%       10%\n             9.33           6.554         4.747     3.177                                             (3)\n\n      c) Comment on the suitability of Model 1 and Model 2 using R2.\n         Note - Students are required to provide calculation of R2 separately for Model 1 and\n         Model 2 clearly stating the formula.                                                         (3)",
          "topic": null
        },
        {
          "label": "iv",
          "marks": 4,
          "text": "a) A statistician is comparing models using R. He observes following results using R\n          software\n          Result 1\n          Analysis of Deviance Table\n\n      Model 1: y ~ 1\n      Model 2: y ~ x\n\n          Resid. Df Resid. Dev Df Deviance               F        Pr(>F)\n       1       11         1888\n       2       10        118.71      1     1769.3 149.05 2.48E-07 ***\n      ---\n      Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1\n\n      Result 2\n      Analysis of Deviance Table\n\n      Model 1: y ~ x\n      Model 2: y ~ x * region\n\n         Resid. Df Resid. Dev Df Deviance  F    Pr(>F)\n       1    10      118.706\n       2     8       31.289   2   87.417 11.175 0.00483 **\n\n              ---\n      Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1\n\n      Comment on the results and model you would prefer with reasons.                            (4)\n\n      b) You then further observe following result in R\n\n      Result 3\n\n      Model 1: y ~ x + region\n      Model 2: y ~ x * region\n\n         Resid. Df Resid. Dev Df Deviance   F   Pr(>F)\n       1     9      31.327\n       2     8      31.289    1   0.0374 0.0096 0.9245\n\n      Comment on the new result and model that you would prefer with reasons based on results\n      in part a and part b.                                                                      (4)",
          "topic": null
        }
      ],
      "solution": "i)\n a) Correct Option is Option D                                                                        (2)\n\n steps not required\n\n theta= log(mu) hence, mu = etheta\n E(Y) = b(theta) = mu = etheta\n\n                                                                                             Page 7 of 10\n\f  IAI                                                                                         CS1A-0322\n\n  b`(theta) = etheta = mu\n  V(mu) = b``(theta) = etheta = mu\n  Hence correct option is Option D\n\n  b) Correct option is Option C                                                                       (2)\n\n  steps not required\n\n  b(theta) = -log(-theta)\n  mean = E(Y)= b`(theta) = -1/theta\n  as theta=1/mu, b`(theta) = mu\n  variance function is V(mu) = b``(theta) = 1/theta2 = mu2\n  Hence, variance is V(mu) / a(phi) = mu2/1 = mu2\n  Hence, correct option is Option C i.e. mean=mu and variance = mu2\n\nii)\n  a) Interaction term means that effect of age band on accidental hospitalisation claims depends on\n     the gender of the insured and significant indicates that accidental hospitalisation claims are\n     better modelled with interaction term, when claims for any age group change with respect to\n     gender of the insured.\n\n        E.g. accidental hospitalisation claims are expected to be higher for males compared to females\n        for any age group.\n        This can be achieved in the model by having Beta_males > Beta_females across age groups. This\n        will result in higher expected accidental hospitalisation claims for males compared to females for\n        any age group (assuming only other parameter is for age band which is same for males and\n        females.                                                                                       (4)\n\nb) Policy renewal is binary event for a single policy. Hence, predicted policy renewal rate (say, mu) for\n   a group of policies can vary between 0 to 1.\n\n    If logit link function is used, then\n    eta = log(mu/(1-mu))\n    mu/(1-mu) = exp(eta)\n    mu = exp(eta) – mu*exp(eta)\n    mu(1+exp(eta)) = exp(eta)\n    mu = exp(eta)/ (1+exp(eta)) = 1/(1+exp(-eta)) = (1+exp(-eta))^-1\n\n    This is expected to result in the range of 0 to 1 for mu as required - renewal rate for a group of\n    policies.\n    Hence, logit link function can be used for renewal rate.                                       (4)\n\niii)\n   a)\n\n        Model 1  Model 2\n  SSREG    1.380    2.380             Given\n  A        1.380    2.380             as A/1 = SSREG\n  B        4.000    3.000             B=5.38-A\n  C        0.333    0.250             C= B/12\n  F        4.140    9.520             F=SSREG/C\n\n                                                                                             Page 8 of 10\n\f IAI                                                                                      CS1A-0322\n\n b)\n H0 = Beta parameter is zero (no linear relationship)\n H1 = Beta parameter is not equal to zero (linear relationship is present)\n\n For Model 1, F = 4.14, this is between 5% critical value of 4.747 and 10% critical value of 3.177\n Hence, we have sufficient evidence to reject null hypothesis that beta parameter is zero at 10% level\n (but cannot be rejected at 5% level) indicating linear relationship between response variable and\n predictor variable at 10% level\n\n For Model 2, F=9.52 is greater than critical value at 1% level as well.\n Hence, we have sufficient evidence to reject null hypothesis even at 1% level indicating linear\n relationship between response variable and predictor variable at 1% level\n\nc) R2 = SSREG /SSTOT\n\n For Model 1, R2 = 1.38/5.38 = 25.7% and\n For Model 2, R2 = 2.38/5.38 = 44.2%\n\n As % of variation explained by model 1 and Model 2 is low (based on low value of R2 of Model 1 and\n Model 2), we can conclude that none of the model is good fit to the data and hence, models are not\n suitable for prediction purpose (though there is linear relationship between response and predictor\n variable for Model 1 and Model 2 as indicated in part b)\n\n iv)\n a)\n Based on Result 1\n      Model using x as predictor is significant improvement over null model\n      As reduction in deviance (by 1769) is significantly more compared to 2 times the loss in\n        degrees of freedom by 1 when x is used as predictor over null model\n      Result is significant even at 0.000001 level as p value is smaller than that\n\n Based on Result 2\n     Model using interaction term between x and region is significant improvement over model\n        using just x as predictor\n     As reduction in deviance (by 87) is significantly more compared to 2 times the loss in degrees\n        of freedom by 2 when interaction between x and region is considered over model using just\n        x as predictor\n     Result is significant at 0.005 level as p value is smaller than 0.005\n\n b)\n Based on Result 3\n     Model using only main effect of x and region is not significantly different compared to model\n        using interaction between x and region\n     As reduction in deviance (by 0.037) is less than 2 times the loss in degrees of freedom by 1\n        when interaction between x and region is considered over model using just the main effect\n        of x and region\n     As p value (0.92) is much more than 90% and\n\n Comparing Result 2 and Result 3\n     Model using only main effects of x and region is significantly better than model using x as\n       predictor\n\n                                                                                         Page 9 of 10\n\fIAI                                                                                       CS1A-0322\n\n         as reduction in deviance is significantly more compared to 2 times the loss in degrees of\n          freedom by 1 when main effect of x and region are used as predictor over model using only x\n         and p value to be significant at 0.005 level\n\n         We prefer simpler model i.e. Model with only main effects of x and region over complex\n          model i.e. Model having interaction between x and region as additional complexity is not\n          justified by better prediction\n                                         *********************\n\n                                                                                       Page 10 of 10",
      "has_math": false,
      "session": "2022-03",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2022-03_QP.pdf",
      "source_sol": "raw/CS1A_2022-03_SOL.pdf"
    },
    {
      "q_num": 1,
      "marks": 7,
      "topic": "data_analysis",
      "subtopics": [],
      "stem": "",
      "parts": [
        {
          "label": "i",
          "marks": 2,
          "text": "List the key steps of a data analysis process.                                                        (4)\n\n        ii) Random variable X follows a distribution with an unknown parameter μ. A statement “The\n            probability of μ being between 10 and 50 is 95%” is valid in… (pick the correct option below)\n\n        A. Classical statistics only\n        B. Bayesian statistics only\n        C. Both Classical & Bayesian statistics\n        D. Neither Classical nor Bayesian statistics                                                             (1)\n\n        iii) Match items A-D with the corresponding items I-IV below.\n\n        A. Type I error\n        B. Type II error\n        C. Specificity\n        D. Power\n\n        I. True positive\n        II. True negative\n        III. False positive\n        IV. False negative                                                                                       (2)",
          "topic": null
        }
      ],
      "solution": "i)\n    1. Developing objectives to be met by the results of the data analysis\n    2. Identifying the data items required for the analysis\n    3. Collecting the data\n    4. Processing and formatting the data\n    5. Cleaning the data\n    6. Exploratory data analysis\n    7. Modelling\n    8. Communicating the results\n    9. Monitoring the process (updating the data & repeating the process if required)\n\nii) B                                                                                                     [1]\n\n       Explanation:\n       In classical statistics, μ is a fixed quantity, and therefore cannot have a probability distribution\n       associated with it.\n\n       However, in Bayesian statistics, μ is a random variable, and therefore statements about its probability\n       can be made\n\niii) A – III\n     B – IV\n     C – II\n     D–I",
      "has_math": true,
      "session": "2022-07",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2022-07_QP.pdf",
      "source_sol": "raw/CS1A_2022-07_SOL.pdf"
    },
    {
      "q_num": 2,
      "marks": 6,
      "topic": "distributions",
      "subtopics": [],
      "stem": "An Insurance company is analysing the claim data. Determine the probabilities of the following events.",
      "parts": [
        {
          "label": "i",
          "marks": 2,
          "text": "The number of claims reported in a year by 100 policyholders is less than 6.\n\n        Assume claims reporting from each policyholder follows Poisson distribution with mean 0.03 per year\n        independently of the other policyholder.                                                                 (2)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 2,
          "text": "The number of claims examined up to and including the fourth claim that exceeds £50,000 is less\n            than 7.\n\n        Assume the above follows negative binomial distribution with probability of a claim exceeding\n        £50,000 as 0.4 independent of any other claim.                                                           (2)",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 2,
          "text": "The number of deaths in the coming year amongst a group of 1000 policyholders is less than 10.\n\n        Assume each policyholder has a 0.015 probability of dying in the coming year independently of any\n        other policyholder.                                                                                      (2)",
          "topic": null
        }
      ],
      "solution": "i)         The number of claims incurred by each policyholder follows the poisson distribution with mean\n           0.03. Therefore X, the number of claims for the 100 policyholders follows the Poi(3), X ~ Poi(3).\n\n           Since the poisson distribution only takes integer value P(X<6) = P(X<=5)\n           Using the poisson cumulative probability tables gives 0.91608\n\nii)        Counting the numbers of trials up to and including the 4th success. This describes the variable (X)\n           is Type 1 negative binomial distribution with k= 4 and p = 0.4\n\n                     𝑥−1\n           P(X=x) = (   ) 0.44 0.6𝑥−4            x = 4,5,6,….\n                      3\n\n           So P(X <7) = P(X=4) + P(X=5) + P(X=6)\n\n                     3\n           P(X=4) = ( ) 0.44 = 0.0256\n                     3\n                                                     𝑥−1\n           Now using the iterative formula P(X=x) = 𝑥−4q P(X=x-1)\n\n                    4\n           P(X=5) = 1 ×0.6 × 0.0256 = 0.06144\n                    5\n           P(X=6) = 2 ×0.6 × 0.06144 = 0.09216\n\n           Hence, P(X <7) = 0.0256 + 0.06144 + 0.09216 = 0.1792\niii)       Here the variable(X) is binomial distribution with n = 1000 and p = 0.015\n           Since n is large and p is small, hence poisson approximation can be used\n\n                                                                                                  Page 2 of 9\n\fIAI                                                                                          CS1A-0722\n\n         Bin(1000,0.015) ~ Poi(15) (approximately)\n\n         Using the cumulative Poisson table gives\n\n         P(X <10) = P(X <=9) = 0.06985                                                               [2]",
      "has_math": false,
      "session": "2022-07",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2022-07_QP.pdf",
      "source_sol": "raw/CS1A_2022-07_SOL.pdf"
    },
    {
      "q_num": 3,
      "marks": 6,
      "topic": "distributions",
      "subtopics": [],
      "stem": "A Student Actuary is analysing the time taken between two consecutive claims in a health insurance\n        policy. It is believed that the time period (denoted by random variable X) between two consecutive\n        claims in a health insurance policy follows an exponential distribution with mean µ.",
      "parts": [
        {
          "label": "i",
          "marks": 2,
          "text": "Identify the correct expression for moment generating function of X.\n                      1        ∞\n        A. E(𝑒 𝑡𝑋 ) = (𝜇 − 𝑡) ∫0 𝑒 −𝑧 . dz\n                     1 1           ∞\n        B. E(𝑒 𝑡𝑋 ) = 𝜇 (𝜇 − 𝑡)−1 ∫0 𝑒 −𝑧 . dz\n\n                       1   ∞\n        C.   E(𝑒 𝑡𝑋 ) = 𝜇 ∫0 𝑒 −𝑧 . dz\n                       1 1        ∞\n        D. E(𝑒 𝑡𝑋 ) = 𝜇(𝜇 − 𝑡) ∫0 𝑒 −𝑧 . dz\n\n        E. None of the above",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 3,
          "text": "If random variable Y denotes sum of time periods of two consecutive claims of N policies,\n            determine the moment generating function of Y.                                                        (3)",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 1,
          "text": "Identify the distribution of Y.                                                                      (1)",
          "topic": null
        }
      ],
      "solution": "i)    Correct Answer (B)                                                                             [2]\n             1 1           ∞\nE(𝑒 𝑡𝑋 ) = 𝜇 (𝜇 − 𝑡)−1 ∫0 𝑒 −𝑧 . dz\nE(𝑒 𝑡𝑋 ) = (1 − 𝑡µ)−1\n\nii) Total time Y of time periods of N policies will be Y = 𝑋1 + 𝑋2 + ……. +𝑋𝑁\n\nMGF of Y is given by E(𝑒 𝑡𝑌 ) = E(𝑒 𝑡 ∑ 𝑋𝑖 ) = П1𝑁 𝐸(𝑒 𝑡𝑋 )\n\n𝑀𝑌 (𝑡) = (1 − 𝑡µ)−𝑁\n\niii) The 𝑀𝑌 (𝑡) is of the form of MGF for Gamma distribution\n                                        1\n     Thus, the distribution is Gamma(N, 𝜇)                                                            [1]",
      "has_math": true,
      "session": "2022-07",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2022-07_QP.pdf",
      "source_sol": "raw/CS1A_2022-07_SOL.pdf"
    },
    {
      "q_num": 4,
      "marks": 4,
      "topic": "inference",
      "subtopics": [],
      "stem": "Claim sizes on a fixed benefit health insurance policy are normally distributed about a mean of INR\n        900 and with a standard deviation of INR 100. Claim sizes on a indemnity based health insurance\n        policy are normally distributed about a mean INR 1,400 and with a standard deviation of INR 300. All\n        claim sizes are assumed to be independent and in units of INR 100.\n\n        To date, there have already been fixed benefit health insurance claims amounting to INR 900 and no\n        indemnity based health insurance claims. Assuming that there will be further 4 fixed benefit health\n        insurance claims and 3 indemnity based health insurance claims in next year, calculate the probability\n        that the total claim amount under the indemnity based health insurance claims exceeds the total claim\n        amount under fixed benefit health insurance claims.",
      "parts": [],
      "solution": "Let X be the amount of fixed benefit health insurance claims and Y the amount of indemnity\nbased health insurance claim.\n\nThen:\nX~ N(900, 1002) and Y ~ N(1400, 3002)\n\nWe require\n\nP((Y1+Y2 + Y3 ) > (X1+X2 + X3+ X4) + 900)\n= P((Y1+Y2 + Y3 ) - (X1+X2 + X3+ X4) >900)\n\nSo we need the distribution of (Y1+Y2 + Y3) - (X1+X2 + X3 + X4):\n\n(Y1+Y2 + Y3) - (X1+X2 + X3 + X4) ~ N( 3×1400 – 4×900, 3×3002+4×1002)\n\ni.e (Y1+Y2 + Y3) - (X1+X2 + X3 + X4) ~ N(600,310000)\n\nTherefore\n\nP((Y1+Y2 + Y3) - (X1+X2 + X3 + X4)>900)\n\n         900−600\n= P( Z > 310000 ) = P( Z > 0.54) = 1 – P( Z < 0.54) = 1- 0.70540 = 0.2946\n         √",
      "has_math": false,
      "session": "2022-07",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2022-07_QP.pdf",
      "source_sol": "raw/CS1A_2022-07_SOL.pdf"
    },
    {
      "q_num": 5,
      "marks": 21,
      "topic": "inference",
      "subtopics": [],
      "stem": "",
      "parts": [
        {
          "label": "i",
          "marks": 3,
          "text": "Briefly explain what is meant by Independent and Identically Distributed (IID) Variables.              (1)\n\n        A coin is tossed five times and the outcome is as follows: heads, heads, heads, heads, tails. Assume\n        that the coin tosses are IID and the probability of each coin toss being either heads or tails is θ and\n        (1-θ) respectively.\n\n        ii) Under the null hypothesis where θ = 0.5:\n\n        a) Calculate the probability of the outcome observed.                                                     (1)\n\n        b) Calculate the p-value of four or more heads and comment whether the null hypothesis can be\n           rejected at significance level of 5%.                                                                  (3)\n\n        iii) Considering θ as an unknown parameter:\n\n        a) Write down a formula for the likelihood of the outcome observed.                                       (1)\n\n        b) Which of the following is the Maximum Likelihood Estimate (MLE) of θ?\n\n             A. 0.5\n             B. 0.4\n             C. 0.8\n             D. 0.2\n             E. None of the above                                                                                 (3)\n\n        The prior distribution of θ was as follows:\n\n            θ           0.1           0.3           0.5          0.7            0.9\n          f(θ)          0.2           0.2           0.2          0.2            0.2\n\n      c) Calculate the prior expected value of θ.                                                                 (1)\n\n      d) Briefly explain what the prior distribution indicates about our knowledge of θ.                          (1)\n\n      e) Based on the outcome observed (4 heads, 1 tail), the posterior expected value of θ is… (pick the\n         correct option).\n\n          A. Less than the prior expected value\n          B. Equal to the prior expected value\n          C. Greater than the prior expected value\n          D. None of the above                                                                                    (2)\n\n      The coin is tossed 3 more times and the outcome is tails, tails, tails.\n\n      f) The posterior distribution is recalculated incorporating the 3 additional coin tosses as well. The\n         revised expected value of θ is also recalculated from this posterior distribution. This expected value\n         is… (pick the correct option)\n\n          A. Exactly 0.2\n          B. Greater than 0.2 but less than 0.5\n          C. Exactly 0.5\n          D. Greater than 0.5 but less than 0.8\n          E. Exactly 0.8                                                                                          (2)\n\n      g) The Bayesian estimator under quadratic loss is… (pick the correct option)\n\n          A. Mean of the prior distribution\n          B. Mean of the posterior distribution\n          C. Median of the prior distribution\n          D. Median of the posterior distribution\n          E. Mode of the posterior distribution                                                                   (1)\n\n      h) The Bayesian estimator under absolute loss is… (pick the correct option)\n\n          A. Mean of the prior distribution\n          B. Mean of the posterior distribution\n          C. Median of the prior distribution\n          D. Median of the posterior distribution\n          E. Mode of the posterior distribution                                                                   (1)\n\n      iv) So far, we have assumed that each coin toss is independent of the previous toss. It is now desired\n          to test this assumption.\n\n      a) State, with reason, whether the chi-squared test can be used for this purpose, considering the\n         sample of 8 coin tosses as above.                                                                        (1)\n\n      The coin is tossed 13 more times (21 tosses in total). A contingency table is prepared as below. The\n      columns shows the outcomes of the coin toss, and the rows shows the outcome of the immediately\n      preceding coin toss. (The first toss is omitted.)\n\n                     Heads         Tails        Total\n          Heads        7             3           10\n          Tails        3             7           10\n          Total       10            10           20\n\n        b) From the above table, test whether each coin toss is independent of the immediately preceding toss\n           and state your inference.                                                                                        (3)",
          "topic": null
        }
      ],
      "solution": "i) A group of random variables is said to be independent and identically distributed if the variables\n        are independent of each other and follow the same probability distribution\n    ii)\n\n                                                                                             Page 3 of 9\n\fIAI                                                                                               CS1A-0722\n             a) As there are 5 coin tosses and the probability of each coin toss being either heads or tails is\n                0.5, the probability of this exact outcome is 0.5^5 = 0.03125                              [1]\n\n             b) p-value is the probability of an observation at least as “extreme” as the actual observation.\n                Under the null hypothesis, the expected number of heads is 2.5, while the actual number of\n                heads is 4 (> 2.5). Thus, we need to calculate the probability of 4 or 5 heads.\n                Let the number of heads be X. Then X ~ Bin (5, 0.5)\n                Prob (X >= 4) = 1 – Prob (X <= 3) = 1 – 0.8125 (from tables) = 0.19 approx\n                As the p-value is 0.19 > 0.05, the null hypothesis cannot be rejected at 5% significance level\n      iii)\n             a) Likelihood can be calculated as:\n                                                           𝑛\n\n                                               𝐿(𝜃) = ∏ 𝑓(𝑥𝑖 ; 𝜃)\n                                                          𝑖=1\n                 which yields\n                                                       𝐿(𝜃) = 𝜃 4 (1 − 𝜃)\n\n             b) C                                                                                             [3]\n                Explanation:\n                 Differentiating the log likelihood,\n\n                                 𝜕             𝜕                         4  1\n                                   log 𝐿(𝜃) =    [4 log 𝜃 + log(1 − 𝜃)] = −\n                                𝜕𝜃            𝜕𝜃                         𝜃 1−𝜃\n\n                  Equating to 0,\n                                   4   1                       4\n                                     −   = 0 ⇒ 𝜃 = 4 − 4𝜃 ⇒ 𝜃 = = 0.8\n                                   𝜃 1−𝜃                       5\n\n                   Checking for maximum:\n                               𝜕2              𝜕 4   1        4     1\n                                 2\n                                   log 𝐿(𝜃) =    [ +     ]= − 2−\n                              𝜕𝜃              𝜕𝜃 𝜃 1 − 𝜃     𝜃   (1 − 𝜃)2\n\n                   Substituting θ = 0.8, this works out to -31.25, which is negative. Thus, θ = 0.8 represents the\n                   maximum.\n                   The MLE of θ is therefore 0.8.\n\n             c) Prior expected value of θ = 0.2*(0.1 + 0.3 + 0.5 + 0.7 + 0.9) = 0.5                           [1]\n\n             d) As the prior distribution is uniform across nearly the entire possible range of θ (0 to 1), it\n                indicates that we have no knowledge (or very little knowledge) about the value of θ.\n\n             e) C                                                                                       [2]\n             Explanation:\n             The posterior expected value would lie somewhere between the prior expected value and the\n             MLE (observed test statistic). Here, the prior EV is 0.5 and the MLE is 0.8.\n             Thus, the posterior EV would lie between 0.5 and 0.8 – i.e., it would be greater than 0.5.\n\n             f) C                                                                                           [2]\n             Explanation:\n             The posterior EV would lie between the prior EV and the MLE.\n             The prior EV is 0.5.\n             The MLE is simply the proportion of coin tosses that result in “heads” – in this case, also 0.5 (8\n             tosses, 4 heads, 4 tails).\n             Since both the prior EV and the MLE are 0.5, the posterior EV must also be exactly 0.5.\n\n                                                                                                      Page 4 of 9\n\fIAI                                                                                                CS1A-0722\n\n           g)       B                                                                                      [1]\n\n           h)       D                                                                                      [1]\n\niv)\n      a) No – For the chi-squared test, values less than 5 for any expected value are generally not\n         considered. If we try to form a contingency table (as is done in the next sub-question) based on 8\n         coin tosses, the expected value in each cell would be less than 5\n\n      b)   Number of degrees of freedom = (rows – 1 ) * (columns – 1) = 1 * 1 = 1\n           Expected values in each cell would be 5 [= row total * column total / table total]\n\n           Thus, the squared difference of the actual value in each cell with the expected value is (observed\n           value – 5)^2, i.e. 4, 4, 4, 4.\n           The χ2 statistic is therefore 4 * 4/5 = 3.2\n           For 1 df, the 5% value of χ2 is 3.841, which is higher than the figure of 3.2 calculated above.\n\n           Thus, there is insufficient evidence to reject the null hypothesis (i.e., that each coin toss is\n           independent of the preceding toss) at the 5% level.",
      "has_math": true,
      "session": "2022-07",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2022-07_QP.pdf",
      "source_sol": "raw/CS1A_2022-07_SOL.pdf"
    },
    {
      "q_num": 6,
      "marks": 5,
      "topic": "distributions",
      "subtopics": [],
      "stem": "The joint probability density function of random variables X and Y is:\n\n                         −(𝑥+3𝑦)\n          f(x,y) = {3𝑒             , 𝑥 > 0, 𝑦 > 0}\n                         0         𝑜𝑡ℎ𝑒𝑟𝑤𝑖𝑠𝑒",
      "parts": [
        {
          "label": "i",
          "marks": 1,
          "text": "Determine fY(y) the marginal density function of Y.                                                              (1)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 2,
          "text": "Determine the conditional density function f(y | Y>4).                                                          (2)",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 2,
          "text": "Identify which one of the following expressions is equal to the conditional expectation E[Y | Y>4]\n             :\n             ∞                 ∞\n        A. ∫0 3𝑒 −3𝑡 𝑑𝑡 + ∫0 12𝑒 −3𝑡 𝑑𝑡\n             ∞                 ∞\n        B. ∫0 3𝑒 −3𝑡 𝑑𝑡 + ∫0 12𝑡𝑒 −3𝑡 𝑑𝑡\n             ∞                  ∞\n        C. ∫0 3𝑡𝑒 −3𝑡 𝑑𝑡 + ∫0 12𝑒 −3𝑡 𝑑𝑡\n             ∞                  ∞\n        D. ∫0 3𝑡𝑒 −3𝑡 𝑑𝑡 + ∫0 12𝑡𝑒 −3𝑡 𝑑𝑡\n        E. None of the above                                                                                                (2)",
          "topic": null
        }
      ],
      "solution": "i) The marginal density is\n                ∞                         ∞\n      fY(y) = 3∫0 𝑒 −𝑥 𝑒 −3𝑦 𝑑𝑥 = 3𝑒 −3𝑦 ∫0 𝑒 −𝑥 𝑑𝑥 = 3𝑒 −3𝑦                                               [1]\n\n  ii) The conditional probability P( Y ≤ y | Y> 4) is FY(y)\n                                P( Y ≤ y ,Y> 4)   P( 4<𝑌 ≤ 𝑦)   P( 4<𝑌 ≤ 𝑦)   FY(y)−FY(4)\n           P( Y ≤ y | Y> 4) =                   =             =             =             ,y > 4\n                                   P( Y> 4)        P( Y> 4)      P( Y> 4)       P( Y> 4)\n           Therefore\n\n                           fY(y)         3𝑒 −3𝑦\n           f(y | Y>4) =              =          = 3𝑒 12−3𝑦 , y>4                                           [2]\n                          P( Y> 4)        𝑒 −12\n\niii) The correct option is (C)                                                                             [2]\n      The conditional expectation is given as\n                        ∞                           ∞\n      E[Y | Y>4] = ∫4 𝑦f(y | Y > 4)𝑑𝑦 = ∫4 3y𝑒 12−3𝑦 𝑑𝑦\n\n      By taking t = y-4,\n                        ∞                           ∞               ∞\n      E[Y | Y>4] = ∫0 3(t + 4)𝑒 −3𝑡 𝑑𝑡 = ∫0 3t𝑒 −3𝑡 𝑑𝑡 + ∫0 12𝑒 −3𝑡 𝑑𝑡",
      "has_math": true,
      "session": "2022-07",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2022-07_QP.pdf",
      "source_sol": "raw/CS1A_2022-07_SOL.pdf"
    },
    {
      "q_num": 7,
      "marks": 17,
      "topic": "regression_glm",
      "subtopics": [
        "distributions"
      ],
      "stem": "You are an actuarial analyst working at a life insurance company in India. The company is analysing\n        the force of mortality, µX, of a particular group of policyholders. The company believes that µX is\n        related to age, X, by the formulae:\n\n        µX = BCX\n        You are provided herewith the following summary results for 10 ages\n\n         Age,X                     30      32       34    36      38    40      42    44      46      48\n         Force of mortality,\n         µX(×10-4)             5.22        5.64    6.25   6.82   7.46   8.73   10.63 12.23 14.45 16.28\n         ln µX                 -7.56 -7.48 -7.38 -7.29           -7.20 -7.04 -6.85   -6.71   -6.54   -6.42\n\n        ∑ 𝑋𝑖 = 390 , ∑ 𝑋𝑖2 = 15,540 , ∑ ln µ𝑋𝑖 = -70.47 , ∑ (ln µ𝑋𝑖 )2 = 498.05 , ∑ 𝑋𝑖 ln µ𝑋𝑖 = -2,726.66\n\n        Your reporting manager has asked you to perform a linear regression analysis on this data to identify\n        the relationship between age and force of mortality.\n\n        It is decided to analyse the assumptions by using the linear regression model:\n\n        Yi = α + βXi + εi where εi ~ N(0, 𝜎 2 ) where Yi = ln µ𝑋𝑖 , α = ln B , β = ln C",
      "parts": [
        {
          "label": "i",
          "marks": 1,
          "text": "The following is a plot of a graph of ln µX against the age of the policyholder, X.\n\n             -6.2\n             -6.4                                                                             -6.42\n             -6.6                                                                     -6.54\n                                                                              -6.71\n             -6.8                                                     -6.85\n              -7                                              -7.04\n             -7.2                                      -7.2\n                                               -7.29\n             -7.4                      -7.38\n                               -7.48\n             -7.6      -7.56\n             -7.8\n\n           Comment on the suitability of the regression model.                                                       (1)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 3,
          "text": "Use the data to calculate least squares estimates of B and C in the original formula.                    (3)",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 2,
          "text": "Write down the value of Pearson’s correlation coefficient between the variables ln µ𝑋𝑖 and Xi\n             hence comment on the relationship between the variables after taking into consideration the value\n             of the slope parameter β.                                                                               (2)",
          "topic": null
        },
        {
          "label": "iv",
          "marks": 2,
          "text": "Calculate the coefficient of determination between ln µX and X. Hence comment on the fit of the\n            model to the data.                                                                                       (2)",
          "topic": null
        },
        {
          "label": "v",
          "marks": 3,
          "text": "Complete the table of residuals and use it to comment on the fit.\n\n             Age,X          30    32   34   36     38     40   42  44    46   48\n             Residual, 𝑒̂𝑖 0.079 0.028    -0.045 -0.087 -0.058    0.009 0.048                                        (3)",
          "topic": null
        },
        {
          "label": "vi",
          "marks": 4,
          "text": "Calculate a 95% confidence interval for the mean predicted response [ln µ45 ] and hence obtain a\n            95% confidence interval for the mean predicted value of µ45 .                                            (4)",
          "topic": null
        },
        {
          "label": "vii",
          "marks": 2,
          "text": "Comment on the width of a 95% confidence interval for the predicted mean response if X = 41, as\n             compared to the width of the interval in part (vi), without calculating the new interval.               (2)",
          "topic": null
        }
      ],
      "solution": "i) The graph appears to show an approximately linear relationship. However, it does appear to have a\n      slight curve and this would warrant closer inspection of the model to see if it is appropriate to the\n      data.                                                                                              [1]\n\nii) Least squares estimates:\n      Obtaining the estimates of α and β with Y = ln µX\n\n                                              390\n      SXX = ∑ 𝑋 2 - n𝑋̅ 2 = 15540 – 10( 10 )^2 = 330\n\n                                                                                                   Page 5 of 9\n\fIAI                                                                                           CS1A-0722\n                                         390 −70.47\n      SXY = ∑ 𝑋𝑌 - n𝑋̅𝑌̅ = -2726.66 – 10( 10 )( 10 ) = 21.67\n\n           SXY 21.67\n      β̂ = SXX = 330 = 0.0657\n\n                         −70.47                 390\n      α = 𝑌̅ - β̂𝑋̅ = ( 10 ) – 0.0657× 10 = -9.61\n      ̂\n      Therefore, we obtain\n      B = 𝑒 𝛼 = 0.000067\n      C = 𝑒 𝛽 = 1.07\n\n              𝑆𝑋𝑌\niii) r =                 = =.990645\n           √𝑆𝑋𝑋 𝑆𝑌𝑌\n\n    The correlation coefficient shows a strong positive relationship between the variables force of\n    mortality and age. The positive value of the regression slope parameter β̂ also suggest the positive\n    correlation between the variables.\niv) The coefficient of determination is given by\n              𝑆2             21.672\n      𝑅 2 = 𝑆 𝑋𝑌\n               𝑆\n                 = 330×1.45 = 98.14%\n              𝑋𝑋 𝑌𝑌\n\n      Where SYY= ∑ 𝑌 2 - n𝑌̅ 2 = 1.45\n\n      This says that 98.14% of the variation in the data can be explained by the model and hence indicates\n      an extremely good fit of the model\n\nv) The completed table of residuals using 𝑒̂𝑖 = yi - 𝑦̂𝑖 is:\n       Age,X              30    32    34     36     38     40     42    44    46    48\n       Residual, 𝑒̂𝑖      0.079 0.028 -0.004 -0.045 -0.087 -0.058 0.001 0.009 0.048 0.036\n\n     Age 34: -7.38 – (-9.61 + 0.0657×34) = -0.004\n     Age 42: -6.85 – (-9.61 + 0.0657×42) = 0.001\n     Age 48: -6.42 – (-9.61 + 0.0657×48) = -0.036\n\n The residuals should be pattern less when plotted against X, however it is clear to see that some pattern\n exists – this indicates that the linear model may not be a good fit.                                  [3]\n\nvi) The variance of mean predicted response is:\n\n 1      (𝑋0 −𝑋̅)2                  1       (45−39)2\n{𝑛 +      𝑆𝑋𝑋\n                  } 𝜎̂ 2 =      {10 +         330\n                                                    } × 0.0034 =   0.00071\n\n                     1            21.672\n     Where 𝜎̂ 2 = 8(1.45 - 330 ) = 0.0034\n\nThe estimate is Y = ln µ45 = -9.61+0.067×45 = -6.65\n\nUsing the t8 distributions , a 95% confidence interval for Y = ln µ45 is\n\n-6.65      ± 2.306√0.00071 = (-6.71, -6.59)\n\nThe corresponding 95% confidence interval for µ45 is (0.001219, 0.001374)\nvii) The width of the interval is only affected by the variance of the mean predicted response. Which\n     depends on the value of (𝑋0 − 𝑋̅)2 . This term will now be smaller as the new 𝑋0 = 41 value is closer\n     to 𝑋̅than 𝑋0 = 45. Therefore the interval will be narrower.                                       [2]\n                                                                                              Page 6 of 9\n\fIAI                                                                                              CS1A-0722",
      "has_math": true,
      "session": "2022-07",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2022-07_QP.pdf",
      "source_sol": "raw/CS1A_2022-07_SOL.pdf"
    },
    {
      "q_num": 8,
      "marks": 11,
      "topic": "distributions",
      "subtopics": [],
      "stem": "You have been investing in shares of unrelated industries X and Y for diversification. Your friend\n        opines that this does not achieve diversification, because share indices of the two industries are\n        strongly positively correlated with each other. To support this line of argument, he gives you the\n        following index values for two different dates over two different periods in time:\n\n         Period 1 Period 2\n         X Y X Y\n         10 10 10 10\n         30 20 30 15",
      "parts": [
        {
          "label": "i",
          "marks": 5,
          "text": "Calculate Pearson’s correlation coefficient between the two indices X and Y for each of the two\n           periods.                                                                                                  (5)\n\n        On digging deeper, it turns out that the bases of the indices differ between periods 1 and 2. To express\n        the period 2 figures in the period 1 base, the X figures need to be multiplied by 2 and the Y figures by\n        0.4.",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 4,
          "text": "Expressing all figures in the period 1 base, calculate the combined correlation of X and Y across\n             both the periods.                                                                                 (4)",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 2,
          "text": "Comment on whether your portfolio is diversified in view of your friend’s opinion based on the\n              results of (i) and (ii).                                                                              (2)",
          "topic": null
        }
      ],
      "solution": "i)\n     Period 1:\n     E(X) = (10+30)/2 = 20\n     SXX = (10-20)^2 + (30-20)^2 = (-10)^2 + (10)^2 = 200\n\n     E(Y) = (10+20)/2 = 15\n     SYY = (10-15)^2 + (20-15)^2 = (-5)^2 + 5^2 = 50\n\n     SXY = (-10 * -5) + (10 * 5) = 100\n\n     Correlation = SXY / √ (SXX * SXY) = 100 / √ (200 * 50) = 1\n\n     Period 2:\n     E(X) = (10+30)/2 = 20\n     SXX = 200, as above\n\n     E(Y) = (10+15)/2 = 12.5\n     SYY = (10-12.5)^2 + (15-12.5)^2 = (-2.5)^2 + 2.5^2 = 12.5\n\n     SXY = (-10 * -2.5) + (10 * 2.5) = 50\n     Correlation = 50 / √ (200 * 12.5) = 1\n\n     ii) In the period 1 base, the figures of X & Y are: (10, 10); (30, 20); (20, 4); (60, 6).\n\n     E(X) = (10 + 30 + 20 + 60) / 4 = 30\n     SXX = (-20)^2 + 0^ 2 + (-10)^2 + 30^2 = 1400\n\n     E(Y) = (10 + 20 + 4 + 6)/4 = 10\n     SYY = 0^2 + 10^2 + (-6)^2 + (-4)^2 = 152\n\n     SXY = (-20 * 0) + (0 * 10) + (-10 * -6) + (30 * -4) = -60\n     Correlation = -60 / √(1400*152) = -0.13\n\n     iii) In part (a), the correlation between X and Y was calculated separately for each period, and they\n        appeared to be perfectly positively correlated.\n     However, on combining the periods in part (b), X and Y turn out to be (weakly) negatively correlated.\n     Thus, while the friend’s assumption of strong positive correlation may be valid for some periods,\n     overall, the correlation between the two indices / industries X & Y appears to be very weak and\n     negative. As such, the portfolio is diversified.",
      "has_math": false,
      "session": "2022-07",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2022-07_QP.pdf",
      "source_sol": "raw/CS1A_2022-07_SOL.pdf"
    },
    {
      "q_num": 9,
      "marks": 5,
      "topic": "bayes_credibility",
      "subtopics": [],
      "stem": "An Indian life insurer has written a large book of policies which provides the Sum Assured on\n         diagnosis of major stage cancer.\n\n         The claim frequency per mille (number of claims per 1,000 policies) arising on this book over the past\n         3 years is as below:\n\n           Year   Claim frequency (per mille)\n          2019-20            16.4\n          2020-21            17.3\n          2021-22            16.7\n\n         The age composition and other aspects of the book have been relatively unchanged over this period.\n\n         The claim frequency per mille is modelled as following a Poisson distribution with an unknown\n         parameter λ.\n\n         The Pricing Actuary models λ as following a Gamma (A, B) distribution, with A = 15 and B = 1.\n\n         The Bayesian credibility factor is given by n/(n + B), where n is the number of years for which data is\n         available.",
      "parts": [
        {
          "label": "i",
          "marks": 3,
          "text": "Calculate the Bayesian credibility estimate for the number of claims per 1,000.                         (3)\n\n         The Pricing Actuary had chosen the prior distribution of Gamma (15, 1) based on inputs from a global\n         reinsurer, who expected the claim frequency to be 15 per mille.\n\n         The Appointed Actuary disagrees with the Pricing Actuary’s choice of Gamma parameters, as he\n         believes the reinsurer’s experience may not be very relevant to the Indian market. He therefore\n         suggests using Gamma (3, 0.2) as the prior distribution.",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 1,
          "text": "Briefly explain, by general reasoning, how the suggested parameters reflect the greater uncertainty\n             regarding the relevance of the reinsurer’s data.                                                       (1)",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 1,
          "text": "Calculate the revised Bayesian credibility estimate.                                                  (1)",
          "topic": null
        }
      ],
      "solution": "i)         Sample mean = (16.4 + 17.3 + 16.7) / 3 = 16.8\n           Prior mean = A/B = 15/1 = 15 (formula from Tables)\n           Credibility factor Z = 3/(3+1) = 0.75\n           Credibility estimate = Z * 16.8 + (1-Z) * 15 = 16.35\n\nii)        The variance of the gamma distribution is mean / B. Reducing the B parameter while\n           keeping the mean constant increases the variance, reflecting greater uncertainty.             [1]\n\n                                                                                                 Page 7 of 9\n\f IAI                                                                                               CS1A-0722\n  iii)   Revised credibility factor Z = 3/3.2 = 0.94 (approx.)\n         Revised credibility estimate = 0.94 * 16.8 + 0.06* 15 = 16.69                                    [1]",
      "has_math": true,
      "session": "2022-07",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2022-07_QP.pdf",
      "source_sol": "raw/CS1A_2022-07_SOL.pdf"
    },
    {
      "q_num": 10,
      "marks": 6,
      "topic": "inference",
      "subtopics": [],
      "stem": "An insurer has written n personal accident policies, which pay the Sum Assured in case of death of\n         insured due to accident.\n\n         The probability of a claim payout is assumed to be q. Each claim is IID.\n\n         The actual number of claims paid is x. Policy-wise data is available, showing exactly which policies\n         have turned into claims.",
      "parts": [
        {
          "label": "i",
          "marks": 1,
          "text": "Which of the following is the likelihood function for the situation described?\n\n              𝑛          𝑥 (1          𝑛−𝑥\n         A.       𝐶𝑥 𝑞          − 𝑞)\n         B. 𝑞 𝑥 (1 − 𝑞)𝑛−𝑥\n              𝑛\n         C.       𝐶𝑥 𝑞1−𝑥 (1 − 𝑞)𝑛−𝑥\n         D. 𝑞 𝑛−𝑥 (1 − 𝑞)𝑥\n              𝑛\n         E.       𝐶𝑥 𝑞 𝑛−𝑥 (1 − 𝑞)𝑥\n\n         Given that n = 10,000 and x = 3.",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 4,
          "text": "Using the normal approximation, calculate the 95% confidence interval for 𝑞̂ (the sample estimator\n             of q).                                                                                                (4)",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 1,
          "text": "Consider q = 0.2 per mille as the null hypothesis.\n              Based on the above, comment on the validity of the null hypothesis.                                  (1)",
          "topic": null
        }
      ],
      "solution": "i)      B\n         Explanation:\n         Likelihood is the probability of the exact outcome observed. As x policies have resulted\n         in a claim and 1-x have not, and the policies are independent, the probability is given by\n         the product: 𝑞 𝑥 (1 − 𝑞)𝑛−𝑥\n\n         The 𝑛𝐶𝑥 factor is not relevant here as for each policy, we know whether there has been\n         a claim or not.                                                                                      [1]\n\n ii)     Let X be the number of claims. Then X ~ Bin (n, q), with mean nq and variance nq(1-q).\n\n         By central limit theorem, 𝑞̂ = X/n approximately follows 𝑁(𝜇, 𝑆), where:\n\n         𝜇 = x/n = 3 / 10,000 = 0.3 per mille = 3 * 10^-4\n\n         S        = 𝑞̂ ∗ (1 − 𝑞̂)\n                  = 3 * 10^-4 * 0.9997 = 3 * 10^-4 (approx.) = 0.3 per mille\n\n                   𝑆                   3∗10−4\n         1.96 ∗ √ = 1.96 ∗ √                  = 0.34 per mille (approx.)\n                   𝑛                   10,000\n                                                             𝑆\n         The 95% confidence interval is 𝑞̂ ± 1.96 ∗ √𝑛\n         Plugging in the values calculated above, and noting that 𝑞̂ can’t be negative,\n         95% confidence interval: (0 per mille, 0.64 per mille)                                               [4]\n\n iii)    As 0.2 per mille falls within the 95% confidence interval, there is insufficient evidence to\n         reject the null hypothesis q at p = 5%.                                                              [1]",
      "has_math": true,
      "session": "2022-07",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2022-07_QP.pdf",
      "source_sol": "raw/CS1A_2022-07_SOL.pdf"
    },
    {
      "q_num": 11,
      "marks": 12,
      "topic": "regression_glm",
      "subtopics": [
        "distributions"
      ],
      "stem": "A random variable z has a binomial distribution with parameters n and µ and has the following density\n         function:\n\n         f(z) = (𝑛𝑧)𝜇 𝑧 (1 − 𝜇)(𝑛−𝑧) , where 0 < µ < 1\n\n                                                         𝑍",
      "parts": [
        {
          "label": "i",
          "marks": 4,
          "text": "Show that the distribution function of Y = 𝑛 can be written in the standard form of the exponential\n            family of distributions, stating the natural and scale parameters, 𝜃 and 𝜑, and the associated\n            functions of these parameters.                                                                         (4)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 3,
          "text": "Verify the mean and variance of the Binomial Distribution, using the expressions from part (i)\n             together with the properties of the exponential family of distributions.                              (3)\n\n         A researcher is investigating the number of students who pass in a particular examination. The\n         researcher believes that the number of students who pass follows binomial distribution.\n\n         He also believes that probability of passing, 𝜇, depends on the followings\n\n             The number of assignment, N, submitted by the student\n             The student’s mark in the mock exam S\n             Whether student attended tutorials or not (Yes/No)\n\n         The researcher specifies the following linear predictor, where 𝛼𝑖 , 𝛽1 and 𝛽2 are parameters to be\n         estimated\n\n         𝜂(𝜇) = 𝛼𝑖 + 𝛽1 𝑁 + 𝛽2 S\n\n         Where 𝛼𝑖 takes one value for those attending tutorials (𝛼𝑌 ) and a different value for those who do not\n         ( 𝛼𝑁 ).\n\n         The researcher then runs computer model that fits generalized linear model (using binomial canonical\n         link function) basis of data collected from 30 observation points.\n\n      Parameters:                   Estimate        Standard Error\n      Intercept, 𝛼𝑌                  -1.501         0.29190\n      Intercept, 𝛼𝑁                  -3.196         0.13401\n      𝛽1, no. of assignment           0.5459         0.08352\n      𝛽2, mark in mock exam          0.0251         0.00156",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 2,
          "text": "Explain, using the model output shown above, whether the variable “no. of assignment” is\n           significant or not.                                                                                (2)",
          "topic": null
        },
        {
          "label": "iv",
          "marks": 3,
          "text": "Estimate using the fitted model, the probability of passing for a student who attends tutorials,\n          submitted 4 assignments and scored 65 marks in the mock exam.                                       (3)",
          "topic": null
        }
      ],
      "solution": "i) The PF of Z is\n\n                       f(z) = (𝑛𝑧)𝜇 𝑧 (1 − 𝜇)(𝑛−𝑧)\n The PF function of Y can be obtained by replacing z with ny :\n\n                               𝑛\n                       f(y) = (𝑛𝑦 ) 𝜇 𝑛𝑦 (1 − 𝜇)(𝑛−𝑛𝑦)\n This can be written as :\n\n                        𝑛\n         f(y) = exp{ln (𝑛𝑦 ) + ny lnμ + n ln(1 − μ) − ny ln(1 − μ)}\n                                   𝜇                𝑛\n               = exp{𝑛𝑦 ln (1−𝜇) + nln(1 − μ) + ln (𝑛𝑦 )}\n                                𝜇\n                       𝑦 ln(      )+ ln(1−μ)\n                               1−𝜇                   𝑛\n               =exp{             1/𝑛\n                                               + ln (𝑛𝑦 )}\n Comparing this to the generalized form of exponential family of distributions:\n\n             𝜇                                         𝑒𝜃\n 𝜃 = ln (1−𝜇) . Rearranging this gives µ = 1+𝑒 𝜃\n\n                                                                                                   Page 8 of 9\n\fIAI                                                                                 CS1A-0722\n                              𝑒𝜃              1\nb(𝜃) = - ln(1 − μ) = - ln( 1-1+𝑒 𝜃 ) = - ln(1+𝑒 𝜃) = ln (1 + 𝑒 𝜃 )\n𝜑=𝑛,\n       1\na(𝜑) = 𝜑\n              𝑛          𝜑\nc(y, 𝜑) = ln (𝑛𝑦 ) = ln (𝜑𝑦 )]\n\nii) Using the properties of exponential distributions\n                   𝑑                             𝑒𝜃\nE(Y) = 𝑏 ∕(𝜃) = 𝑑𝜃(ln (1 + 𝑒 𝜃 )) =             1+𝑒 𝜃\n                                                             =µ\n\n                          𝑒 𝜃 (1+ 𝑒 𝜃 )− 𝑒 𝜃 𝑒 𝜃        𝑒𝜃\nV(Y) = a(𝜑) 𝑏 ∕∕ (𝜃) =             𝑛(1+𝑒 𝜃 )2\n                                                   = 𝑛(1+𝑒 𝜃 )2\n\n                               𝜇\nSubstituting 𝜃 = ln (1−𝜇)\n\n(\n              𝜇\n             1−𝜇           𝜇\nV(Y) =          𝜇 2    = 𝑛(1−𝜇) (1 − 𝜇)2 = 𝜇(1 − 𝜇)/n\n          𝑛(1+     )\n               1−𝜇\n\niii) Using the model output, we can see that\n\n𝛽1 > 2 × 𝑠𝑡𝑎𝑛𝑑𝑎𝑟𝑑 𝑒𝑟𝑟𝑜𝑟(𝛽)\n\ni.e 0.5459 > 2 × 0.08352 = 0.16704\n\nSince\n𝛽1 > 2 × 𝑠𝑡𝑎𝑛𝑑𝑎𝑟𝑑 𝑒𝑟𝑟𝑜𝑟(𝛽) , it can be concluded that the parameter 𝛽1 for the variable “no. of\nassignment” is significant in the model.\n\niv) Using binomial canonical link function,\n               𝜇\n𝜂(𝜇) = ln (1−𝜇) = 𝛼𝑖 + 𝛽1 𝑁 + 𝛽2 S\n\nSo for 𝛼𝑌 = - 1.501 , 𝛽1 = 0.5459, 𝛽2 = 0.0251 and N = 4, S = 65\n\n      𝜇\nln (1−𝜇) = - 1.501 + 0.5459 × 4 + 0.0251 × 65 = 2.3141\n\n𝜇 = ( 1 + 𝑒 −2.3141 )−1 = 91%\n\nHence probability of passing students in the given scenario is 91%\n\n                                                        *******************\n\n                                                                                    Page 9 of 9",
      "has_math": true,
      "session": "2022-07",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2022-07_QP.pdf",
      "source_sol": "raw/CS1A_2022-07_SOL.pdf"
    },
    {
      "q_num": 1,
      "marks": 3,
      "topic": "inference",
      "subtopics": [],
      "stem": "2           ̅ )2\n                                                                       ∑(Xi −X\n        To estimate the population variance σ2 the statistic Sn′ =                  is used rather than using\n                                                                           n\n        estimator Sn2 , where n is the sample size.\n                            2\n        Find the bias of Sn′ , when n = 13, mean is 5.24 and σ2 = 3.4224                                        [3]",
      "parts": [],
      "solution": "𝑆′2𝑛 =            =            𝑛\n                          𝑛\n                                     2\n                             (𝑛 − 1)𝑆𝑛⁄     (𝑛 − 1)𝐸[𝑆𝑛2 ]⁄ (𝑛 − 1)𝜎 2⁄\n              E(𝑆′2𝑛 ) = E [           𝑛] =                𝑛=          𝑛\n                                                   2\n                                   2     (13 − 1)𝜎  ⁄ = 3.1591\n              So for n = 13 , E(𝑆′13 )=              13\n                                2       2\n              Thus, bias = E(𝑆′13 ) − 𝜎 = 3.1591 – 3. 4224 = - 0.2633",
      "has_math": true,
      "session": "2022-12",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2022-12_QP.pdf",
      "source_sol": "raw/CS1A_2022-12_SOL.pdf"
    },
    {
      "q_num": 2,
      "marks": 4,
      "topic": "data_analysis",
      "subtopics": [],
      "stem": "Choose the correct option and provide a reason for your choice. No marks would be awarded\n        if reason is not provided.",
      "parts": [
        {
          "label": "i",
          "marks": 2,
          "text": "A multivariate model is fit with 10 explanatory variables and 100 observations. Due to\n           IT restrictions the model can’t be implemented. An Actuary decided to use Principal\n           Component Analysis (PCA) for reducing the dimensionality of the data set. How many\n           Components will be there after fitting PCA?\n\n        A. 3\n        B. 9\n        C. 10\n        D. 100                                                                                                  (2)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 2,
          "text": "Which of the following is not a Linear Predictor?\n\n        A. Y = α + β X2\n        B. Y = α + β2 X\n                       1\n        C. Y = α + β ( 𝑋 )\n        D. All of the above                                                                                     (2)",
          "topic": null
        }
      ],
      "solution": "i) Answer: C\n\n              PCA will not reduce the number of variables. So a 10 variable data will give 10 components. PCA\n              only transforms the data into uncorrelated linear combination of the variables. The components\n              are then selected to maximise variance. The initial PCAs (say PCA 1 and PCA 2) usually try to capture\n              maximum possible information.\n              ii) Answer: B\n\n              A linear predictor is linear in the parameters. It does not have to be linear in covariates as in case\n              of A) and C)",
      "has_math": true,
      "session": "2022-12",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2022-12_QP.pdf",
      "source_sol": "raw/CS1A_2022-12_SOL.pdf"
    },
    {
      "q_num": 3,
      "marks": 6,
      "topic": "distributions",
      "subtopics": [],
      "stem": "",
      "parts": [
        {
          "label": "i",
          "marks": 3,
          "text": "For a standard normal random variable Z, derive an expression for its Moment\n           Generating Function (MGF) using first principles.                                                    (3)\n\n        ii) Using the results obtained in part (i), prove that a normal variable X with mean µ and\n            variance δ2, is perfectly symmetrical about its mean i.e. the coefficient of skew-ness of\n            the normal variable X is equal to 0.\n\n        Hints:\n              a) Standard relationship between normal variable X and standard normal variable Z\n                 i.e. “Z = (X - µ) / δ” can be directly used without proof.\n               b) Use the fact that E (Zr) is the coefficient of the term tr / r! in the Taylor expansion.      (3)",
          "topic": null
        }
      ],
      "solution": "i)\n\n                                                                                 1 2\n                                                                         1\n              PDF of a standard normal distribution is:                       𝑒 −2 𝑥\n                                                                        √2𝜋\n\n              Mx(t) = E(etx)\n                                                  1 2\n                            ∞             1\n                       = ∫−∞ 𝑒 𝑡𝑥             𝑒 −2𝑥 𝑑𝑥\n                                √2𝜋\n                            1 ∞ 𝑡𝑥 −1𝑥 2\n                       =    ∫ 𝑒 𝑒 2 𝑑𝑥\n                         √2𝜋 −∞\n                          1   ∞ 𝑡𝑥+−1𝑥 2\n                       =    ∫   𝑒        2  𝑑𝑥\n                         √2𝜋 −∞\n                                  1    1       1\n                         1    ∞     𝑡 − 𝑡 +𝑡𝑥+− 𝑥 2\n                                     2    2\n                       =    ∫   𝑒 2    2       2    𝑑𝑥                           ………… Adding and subtracting ½ t2 in the\n                        √2𝜋 −∞\n              exponent\n                            1 2                    1 2            1 2\n                                      ∞   1\n                       = 𝑒 2𝑡 ∫−∞     𝑒 −2𝑡 +𝑡𝑥+−2𝑥 𝑑𝑥\n                                  √2𝜋\n                            1 2                    1    2          2\n                                      ∞   1\n                       = 𝑒 2𝑡 ∫−∞     𝑒 −2(𝑡 −2𝑡𝑥+𝑥 ) 𝑑𝑥\n                                  √2𝜋\n                            1 2                         1 𝑥−𝑡 2\n                                      ∞       1\n                       = 𝑒 2𝑡 ∫−∞ (1)√2𝜋 𝑒 −2( 1 ) 𝑑𝑥\n                            1 2\n                       = 𝑒 2𝑡 * 1\n                                                                                  ……. As the PDF is that of a normal\n                                                                                  distribution with mean t and sd 1 and it\n                                                                                  integrates to 1\n                             1 2\n                       = 𝑒 2𝑡                                                                                                        [3]\n\n              ii)\n\n              We have to prove that a normal variable X with mean µ and variance δ2 is symmetrical about its\n              mean.\n              It means that we have to prove that the coefficient of skew-ness of the normal variable X is equal\n              to 0.\n\n                                                                                                                     Page 2 of 14\n\f    IAI                                                                                                         CS1A-1222\n\n              Symbolically, we have to prove – E[(X - µ)3] = 0                            …………………………… (1)\n\n              We know that standard normal variable Z is a special case of a normal variable with µ = 0 and δ2\n              =1\n\n              We also know the relationship that Z = (X - µ) / δ\n\n              Rearranging, (X - µ) = Z * δ\n\n              Hence we have to prove that-\n              From (1)      E[(X - µ)3] = 0\n              i.e.                E(Z * δ)3 = 0\n              i.e.                δ3 * E(Z3) = 0                                         ……………………………… (2)\n\n                                                           1 2\n              Mx(t) for a standard normal variable = 𝑒 2𝑡\n\n                                                      𝑥𝑟\n              Using the Taylor expansion, ex = ∑∞\n                                                0     𝑟!\n\n                              1\n                             ( 𝑡 2 )𝑟\n              Mx(t) = ∑∞0\n                          2\n                            𝑟!\n                    = 1 + ½ t2 + (½ t2)2 / 2! + ……………………………..\n\n              It is given in the question that the E(Zr) is the coefficient of the term tr / r! in the Taylor expansion.\n\n              So,\n              E(Z) = 0                   …… since term t/1 is not there in the Taylor expansion\n              E(Z2) = 1                  …… coefficient of term t2 / 2 in the Taylor expansion\n              E(Z3) = 0                  …….. there is no term t3 / 6 in the Taylor expansion\n\n              Thus, E(Z3) = 0                                                        …………………………………. (3)\n\n              From (2) and (3),\n\n              L.H.S.\n              = δ3 * E(Z3)\n              = δ3 * 0\n              =0\n\n              R.H.S = 0\n\n              Hence we are able to prove that the coefficient of skew-ness for a normal variable X is 0 and hence\n              we infer that it is symmetrical about its mean. .",
      "has_math": true,
      "session": "2022-12",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2022-12_QP.pdf",
      "source_sol": "raw/CS1A_2022-12_SOL.pdf"
    },
    {
      "q_num": 4,
      "marks": 7,
      "topic": "distributions",
      "subtopics": [],
      "stem": "A health insurance company has recently launched a new one year health insurance product\n        which pays a fixed sum assured on the incidence of Heart, Cancer and Liver related ailments\n        in the next one year.\n\n        Even after a claim under one ailment, coverage continues for the other ailments. A fixed\n        sum assured would be paid on the incidence of the pre-defined ailment and no further claims\n        can then arise for that particular ailment.\n\n        It can be assumed that these three risks are independent.\n\n        Sum Assured for Heart, Cancer and Liver related ailments are INR 20 lakhs, INR 25 lakhs\n        and INR 15 lakhs respectively.\n\n        The company estimates the probabilities of claim arising in the next year to be 0.01 for Heart\n        related ailments, 0.02 for Cancer and 0.005 for Liver related ailments.",
      "parts": [
        {
          "label": "i",
          "marks": 2,
          "text": "Determine, for a single policy, using suitable Bernoulli variables, the mean and standard\n           deviation of the total claim amount to be paid over the next year.                              (2)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 3,
          "text": "You are informed that a claim has been reported under a policy. Given that there is a\n            claim under the policy, calculate the expected pay-out on this claim.                          (3)",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 2,
          "text": "Why does mean claim pay-out in part (i) differ from the expected claim pay-out in part\n             (ii)?                                                                                         (2)",
          "topic": null
        }
      ],
      "solution": "i)\n\n              Since only one claim is eligible for each of the ailments, claims from Heart, Cancer and Liver related\n              ailments can be modelled as three Bernoulli Variables (indicator variables). It is given in the\n              question that the three can be assumed to be independent.\n\n              H = Claims from Heart related ailments              H ~ Bernoulli (0.01)\n\n                                                                                                                  Page 3 of 14\n\fIAI                                                                                               CS1A-1222\n\n      C = Claims from Cancer related ailments              C ~ Bernoulli (0.02)\n      L = Claims from Liver related ailments               L ~ Bernoulli (0.005)\n\n      Let X be the claim amount to be paid out in the next year on a single policy\n\n      X = 20 × H + 25 × C + 15 × L\n\n      We have to find E(X) and s.d.(X)\n\n      E(X)    = 20 × E(H) + 25 × E(C) + 15 × E(L)\n              = 20 × (0.01) + 25 × (0.02) + 15 × (0.005) ……. E(A) = p for A ~ Bernoulli(p)\n              = 0.775 lakhs\n              = INR 77,500\n\n      Var(X) = 202 × Var(H) + 252 × Var(C) + 152 × Var(L)\n                                              …… Since H, C and L are independent, no co-variance terms\n\n              = 400 × (0.01)(1-0.01) + 625 × (0.02)(1-0.02) + 225 × (0.005)(1-0.005)\n                                                                …………. Var(A) = p(1-p) for A ~ Bernouli(p)\n              = 17.32938 lakhs\n\n      SD(X)   = (17.32938)1/2\n              = INR 4.1628 lakhs\n      ii)\n      Exactly one claim has occurred. We don’t know whether it is related to H, C or L.\n\n      P(exactly 1 claim)\n      = P(H) × (1-P(C)) × (1-P(L)) + (1-P(H)) × P(C) × (1-P(L)) + (1-P(H)) × (1-P(C)) × P(L)\n      = (0.01)(1-0.02)(1-0.005) + (1-0.01)(0.02)(1-0.005) + (1-0.01)(1-0.02)(0.005)\n      = 0.009751 + 0.019701 + 0.004851\n      = 0.034303\n\n      P(H | 1 claim has occurred) = 0.009751 / 0.034303 = 0.284261\n      P(C | 1 claim has occurred) = 0.019701 / 0.034303 = 0.574323\n      P(L | 1 claim has occurred) = 0.004851 / 0.034303 = 0.141416\n\n      These should total up to 1.\n\n      So, we have to find E(X | 1 claim has occurred)\n\n      E(X | 1 claim has occurred)\n      = 20 × P(H | 1 claim has occurred) + 25 × P(C | 1 claim has occurred) + 15 × P(L | 1 claim has\n      occurred)\n      = 20 × 0.284261 + 25 × 0.574323 + 15 × 0.141416\n      = INR 22.16453 lakhs\n                                     ………… Kindly note that even after taking account the condition\n                                     that one claim has occurred, H,C and L continue to be Bernoulli\n                                     variables and hence their mean will be equal to p …… although the\n                                     value of p has changed now.\n      iii)\n\n      There are three independent risks covered under this policy with relatively very small probability\n      of incidence of a claim in the next year.\n      The probability of no claim during the next one year = (1-0.01) (1-0.02) (1-0.005) = 0.965349\n\n                                                                                                    Page 4 of 14\n\f    IAI                                                                                                     CS1A-1222\n\n              Since in almost 96% of the cases, there will be no claim, the expected pay-out at the inception of\n              the policy is quite low (lower than1 lakh).\n              However, after one claim has occurred, we have actually experienced something which has a\n              possibility of 3.4% to occur. After its occurrence we are finding out the expected amount since we\n              don’t know whether it relates to H, C or L (otherwise there was no need of expectation, we could\n              directly infer it to be 20 lakhs, 25 lakhs or 15 lakhs).\n              Since something which was only 3.4% probable has actually occurred, there is a significant increase\n              in the expected claim pay-out from (i) to (ii).",
      "has_math": false,
      "session": "2022-12",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2022-12_QP.pdf",
      "source_sol": "raw/CS1A_2022-12_SOL.pdf"
    },
    {
      "q_num": 5,
      "marks": 7,
      "topic": "inference",
      "subtopics": [
        "distributions"
      ],
      "stem": "Out of the 85 tosses of a coin, 40 tosses turn out to be heads.",
      "parts": [
        {
          "label": "i",
          "marks": 2,
          "text": "Let N denote the total number of heads in 85 tosses, what is the most suitable distribution\n           of N? Estimate the mean and variance of N.                                                      (2)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 2,
          "text": "Find out the probability that N > 40 using approximate distribution.                           (2)\n\n        Let the distribution of N as specified in part (i)",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 1,
          "text": "Test the hypothesis that:\n              H0: probability of getting heads = 0.5 v. H1: probability of getting heads > 0.5 at the\n              significance level of 5% using the probability value calculated in part (ii) above.          (1)",
          "topic": null
        },
        {
          "label": "iv",
          "marks": 2,
          "text": "Find X where P(N > X) is less than the significance level of 5% leading to rejecting the\n            null hypothesis of the above test.                                                             (2)",
          "topic": null
        }
      ],
      "solution": "i)\n\n              Most suitable distribution for N is Binomial (85, p) where p is the probability of head.\n\n              Estimate of p is 40/85.\n\n              Mean of N is np =85*(40/85) = 40.\n\n              Variance is np(1-p) = 85*(40/85) * (1-40/85) = 21.1765.                                                         [2]\n              ii)\n\n              N approximately follows Normal with mean 85*(1/2) = 42.5 and variance 85*(1/2)*(1-(1/2)) =\n              21.25.\n\n              Using continuity correction,\n              PBin (N > 40) = PNor(N≥ 40.5)\n                                     p=½                                             p = 40/85\n               = PNor ( Z ≥ (40.5 -42.5)/(21.25^(1/2)))          = PNor ( Z ≥ (40.5 -40)/(21.1765^(1/2)))\n               = PNor (Z ≥ -0.43386) = 0.6678                    = PNor (Z ≥ 0.11) = 0.4562                                   [2]\n\n              iii)\n              Using the above probability, the P-value of the test is 0.6678 (or alternatively 0.4562). We do not\n              have sufficient evidence to reject Ho at 5% significance level.                                                 [1]\n\n              iv)\n\n              Null hypothesis to be rejected for P-value < 0.05,\n              Then PBin (N>=n) = 0.05\n                                                           p=½\n               P[((N-42.5)/21.25^0.5) >=((n-42.5)/21.25^0.5)] = 0.05\n               P(Z>=1.64485) = 0.05\n               Thus, n=1.64485*21.25^0.5+42.5 = 50.0824\n                                                        p = 40/85\n               P[((N-40)/21.1765^0.5) >=((n-40)/21.1765^0.5)] = 0.05\n               P(Z>=1.64485) = 0.05\n               Thus, n=1.64485*21.1765^0.5+40 = 47.56926                                                                      [2]",
      "has_math": false,
      "session": "2022-12",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2022-12_QP.pdf",
      "source_sol": "raw/CS1A_2022-12_SOL.pdf"
    },
    {
      "q_num": 6,
      "marks": 7,
      "topic": "inference",
      "subtopics": [],
      "stem": "A company offers an optional basic group term insurance policy to its employees, as well\n        as an accidental death benefit rider. To be covered under the accidental death benefit rider,\n        an employee needs to first opt for group term insurance policy.\n\n        Let X denote the proportion of employees who have opted for group term insurance cover.\n        Let Y denote the proportion of employees who have opted for accidental death benefit rider.\n\n        Let X and Y have the following joint density function f (x,y) on the region where both X\n        and Y are non-negative:\n\n                            f (x,y) = 2 (x + y)",
      "parts": [
        {
          "label": "i",
          "marks": 1,
          "text": "Clearly specify the bounds on values of X and Y for which the above joint density\n               function holds true.                                                                        (1)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 2,
          "text": "Determine the marginal density function of X.                                               (2)",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 4,
          "text": "Given that 10% of the employees opt for the group term insurance policy, calculate the\n             probability that less than 5% of the employees opt for the accidental death benefit rider.    (4)",
          "topic": null
        }
      ],
      "solution": "i)\n              We know that X and Y are proportions. Hence they need to lie between 0 and 1.\n\n              Hence, preliminary bounds are: 0 ≤ x, y ≤ 1\n\n              It is also given in the problem that you cannot opt for accidental death benefit rider unless you\n              have opted for the group term insurance policy. This further implies that y ≤ x.\n\n                                                                                                             Page 5 of 14\n\f    IAI                                                                                                  CS1A-1222\n\n              So, the final bounds are –\n              0 ≤ x ≤ 1 for X\n              0 ≤ y ≤ x for Y\n              ii)\n\n              We have to determine the marginal density function of X.\n                           𝑥\n              fx(X)     = ∫𝑦=0 𝑓(𝑥, 𝑦) 𝑑𝑦\n                           𝑥\n                        = ∫𝑦=0 2 (𝑥 + 𝑦) 𝑑𝑦\n                           𝑥             𝑥\n                        = ∫𝑦=0 2𝑥 𝑑𝑦 + ∫𝑦=0 2𝑦 𝑑𝑦\n                        = 2x (x – 0) + 2/2 (x2 – 02)\n                        = 2x2 + x2\n              fx(X)     = 3x2                                                                                              [2]\n\n              iii)\n\n              We are given that X = 0.10 and we have to calculate P(Y<0.05 | X = 0.10).\n\n                                             ℎ(𝑥=0.10,𝑦<0.05)\n              P(Y<0.05 | X = 0.10)       =      𝑓𝑥 (𝑥=0.10)\n\n                                               0.05\n              h(x=0.10, y<0.05)          = ∫𝑦=0 2(0.10 + 𝑦) 𝑑𝑦\n\n                                                                    1\n                                         = [0.2(0.05 − 0) + 2 ∗ 2 (0.052 − 02 )]\n                                         = (0.01 + 0.0025)\n                                         = 0.0125\n\n              fx(X=0.1)                  = 3(0.10)2\n                                         = 0.03\n\n              P(Y<0.05 | X = 0.10)       = 0.0125/0.03\n                                         = 0.4167\n\n              Hence the probability that less than 5% of the employees will opt for the accidental death benefit\n              rider, given that 10% of them have opted for the group term insurance policy is 0.4167.                      [4]",
      "has_math": false,
      "session": "2022-12",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2022-12_QP.pdf",
      "source_sol": "raw/CS1A_2022-12_SOL.pdf"
    },
    {
      "q_num": 7,
      "marks": 10,
      "topic": "inference",
      "subtopics": [
        "distributions"
      ],
      "stem": "Let the random variable X have the Poisson distribution with probability function:\n\n               𝑒 −𝜆 𝜆𝑥\n        f (x) =          , x = 0,1,2,...\n                  𝑥!\n\n                                             𝜆",
      "parts": [
        {
          "label": "i",
          "marks": 2,
          "text": "Show that P(X = k+1) = 𝑘+1 𝑃 (𝑋 = 𝑘) , k = 0,1,2,…                                             (2)\n\n        It is believed that the distribution of the number of claims which arise on insurance policies\n        of a certain class is Poisson.\n\n        A random sample of 1,000 policies is taken from all the policies in this class which have\n        been in force throughout the past year. The table shows the observed number of policies\n        with 0, 1, 2, 3, 4, 5, 6, 7 and 8 or more claims during the year:\n\n          No. of Claims(k)          0        1         2       3      4      5     6    7     8+\n          No. of policies(fk)       300      365       216     70     30     16    2    1     0\n\n        For these data the Maximum Likelihood Estimate (MLE) of the Poisson parameter λ is 𝜆̂\n        =1.186",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 3,
          "text": "Calculate the expected number of policies with 0, 1, 2, 3, 4, 5, 6, 7 and 8 or more claims\n            during the year under the Poisson model with parameter given by the MLE above, using\n            the recurrence formula of part (i) (or otherwise).                                            (3)",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 5,
          "text": "Perform an appropriate statistical test to investigate the assumption that the numbers of\n             claims arising from this particular class of policies follow a Poisson distribution.         (5)",
          "topic": null
        }
      ],
      "solution": "i)\n\n              P(X = k+1)\n              Substituting x by K+1\n              P(X=k+1) = 𝑒 −𝜆 𝜆𝑘+1 ⁄(𝑘 + 1)!\n              = 𝑒 −𝜆 (𝜆𝑘 ∗ 𝜆)⁄[(𝑘 + 1) ∗ 𝑘!]\n              = 𝑒 −𝜆 𝜆𝑘 ⁄𝑘! ∗ [𝜆/(𝑘 + 1)]\n                    𝜆\n              =𝑘+1 𝑃 (𝑋 = 𝑘) for k=0,1,2,3,…                                                                               [2]\n\n              ii)\n\n              As MLE of 𝜆̂ = 1.186,\n              P(X=0)=𝑒 −1.186 = 0.3054\n              Also, P(X=8+)=1-∑7𝑖=1 𝑃(𝑋 = 𝑖)\n               K              0        1               2        3       4     5      6        7        8+\n\n                                                                                                            Page 6 of 14\n\f    IAI                                                                                                  CS1A-1222\n\n               Probability\n               using MLE\n               and\n                𝜆\n               𝑘+1\n                   𝑃 (𝑋 =\n               𝑘)          0.3054 0.3623 0.2148 0.0849 0.0252 0.0060 0.0012 0.0002 0.0000\n               Expected\n               No       of\n               policies =\n               prob*1000 305.44 362.25 214.82 84.92 25.18 5.97       1.18   0.20   0.03\n              iii)\n              To perform Chi-square goodness of fit\n              combining 4 categories to obtain >5\n                                                               5&𝑚𝑜𝑟𝑒 (𝑓 − 𝑒 )2\n                                                                        𝑖   𝑖\n                                                     𝜒2 = ∑\n                                                               𝑖=0          𝑒𝑖\n\n               K                 0          1       2         3         4             5+\n\n               ei\n                                 305.4      362.3   214.8     84.9      25.2          7.4\n               fi                300        365     216       70        30            19\n\n                                     𝜒 2 = 0.10 + 0.02 + 0.01 + 2.61 + 0.91 + 18.18 = 21.84\n\n              Degrees of freedom = 6-1-1 = 4 due to MLE estimate\n                            2\n              From tables, 𝜒0.05,4 = 9.488\n              21.84>9.488\n              Thus, the number of claims does not come from a Poisson (1.186) distribution.",
      "has_math": true,
      "session": "2022-12",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2022-12_QP.pdf",
      "source_sol": "raw/CS1A_2022-12_SOL.pdf"
    },
    {
      "q_num": 8,
      "marks": 12,
      "topic": "distributions",
      "subtopics": [
        "data_analysis"
      ],
      "stem": "A new political party is investigating whether size of policemen impacts petty crimes.\n        Following data of 9 zones of a state has been collected:\n\n          Zones              A          B         C       D          E      F      G     H       I\n          Policemen         178        161       140     106        171    153    142   124     37\n          Cases             171        114        62      46        184    149     99    70     39",
      "parts": [
        {
          "label": "i",
          "marks": 4,
          "text": "Calculate Spearman’s and Kendall’s correlation coefficients.                                 (4)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 2,
          "text": "After investigation, party concluded that more crimes are done where more policemen\n            are deployed and suggesting reduction police force. Comment on their conclusion.              (2)",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 6,
          "text": "An expert suggested party to use Pearson’s correlation coefficients instead. Compute\n             Pearson’s correlation coefficient and test whether there is no correlation between\n             policemen and cases.\n\n             ∑ x = 1212 , ∑ y = 934 , ∑ x 2 = 178000 ∑ y 2 = 120476 and ∑ xy = 140790                     (6)",
          "topic": null
        }
      ],
      "solution": "i)\n\n              Spearman Rank correlation coefficient :\n\n                          Rank in low to high order and differences are :\n                   Zones            A        B       C        D       E           F         G   H    I\n                   Police           9        7       4        2       8           6         5    3   1\n                   Cases            8        6       3        2       9           7         5    4   1\n                   Differences      1        1       1        0      -1          -1         0   -1   0\n                   diff square      1        1       1        0       1           1         0    1   0\n\n                            6 𝑋 62\n              rs = 1 − 9 𝑋 (92 −1) = 0.95\n\n              Kendall Rank correlation coefficient :\n\n                        Arranging in order of Policemen rank\n                                                   Concordant Disconcordant\n                   Zones     Police     Cases\n                                                   Pairs      Pairs\n                   I                  1         1              8                      0\n                   D                  2         2              7                      0\n                   H                  3         4              5                      1\n                   C                  4         3              5                      0\n\n                                                                                                          Page 7 of 14\n\f    IAI                                                                                                   CS1A-1222\n\n                    G                 5          5                 4              0\n                    F                 6          7                 2              1\n                    B                 7          6                 2              0\n                    E                 8          9                 0              1\n                    A                 9          8                 0              0\n                    Total                                         33              3\n\n                                                                  33 − 3\n                                                           𝜏=            = 0.83\n                                                                  𝟑𝟑 + 𝟑\n              ii)\n\n              Both Spearman and Kendall rank correlation coefficient indicates a strong positive correlation\n              between policemen and cases. In other words, zones with more policemen have more cases.\n\n              However, correlation does not necessarily infer causation. Even though more cases are present\n              where more policemen are present, it doesn’t indicate cases will go down with reduction of\n              policemen.\n\n              It could be other way, i.e., more policemen are deployed where more crime is present. Or It could\n              be more policemen and crime depends upon the size of zones. Bigger zones have more crime and\n              more force.                                                                                                  [2]\n\n              iii)\n\n              Calculation of Sample correlation coefficient using Pearson’s Method:\n\n              X be Policemen and Y be Cases\n\n                        ∑ 𝑥 = 1212 , ∑ 𝑦 = 934 , ∑ 𝑥 2 = 178000 ∑ 𝑦 2 = 120476 ∑ 𝑥𝑦 = 140790\n\n                     Sxx = 14784.00\n                     Syy = 23547.56\n                     Sxy = 15011.33\n\n                            𝑆𝑥𝑦\n              𝑟=                      = 0.8045\n                      √𝑆𝑥𝑥 𝑆𝒚𝑦\n\n              Test whether is no correlation:\n\n                                                     𝐻0 : 𝜌 = 0         𝑣𝑠 𝐻1 : 𝜌 ≠ 0\n              Under H0 :\n                                                             𝑟√𝑛 − 2\n                                                                        ~ 𝑡𝑛−2\n                                                             √1 − 𝑟 2\n              Observed value of test statistic is\n                                                         0.8045√9 − 2\n                                                                       = 3.58\n                                                     √1 − 0.80452\n              Upper 0.5 % point of t distribution t7 distribution = 3.499 < 3.584 (observed value). Thus, we have\n              sufficient evidence to reject H0 at 1% level. Thereby, it indicates there is strong correlation\n              between policemen and cases.                                                                                 [6]",
      "has_math": true,
      "session": "2022-12",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2022-12_QP.pdf",
      "source_sol": "raw/CS1A_2022-12_SOL.pdf"
    },
    {
      "q_num": 9,
      "marks": 15,
      "topic": "distributions",
      "subtopics": [
        "bayes_credibility"
      ],
      "stem": "Five years ago, an insurance company began to issue insurance policies covering medical\n        expenses for dogs. The insurance company classifies dogs into three risk categories: large\n        pedigree (category 1), small pedigree (category 2) and non-pedigree (category 3).\n\n        The number of claims nij in the ith category in the jth year is assumed to have a Poisson\n        distribution with unknown parameter θi.\n\n         Data on the number of claims in each category over the last 5 years is set out as follows:\n\n          Category                                           Year                                     𝟓            𝟓\n\n                                        1         2           3            4               5         ∑ 𝒏𝒊𝒋        ∑ 𝒏𝟐𝒊𝒋\n                                                                                                     𝒋=𝟏          𝒋=𝟏\n                           1            28        41          47           54              62         232         11434\n                           2            35        48          55           57              65         260         14028\n                           3            26        29          20           39              31         145          4399\n\n         Prior beliefs about θ1 are given by a gamma distribution with mean 50 and variance 25.",
      "parts": [
        {
          "label": "i",
          "marks": 5,
          "text": "Find the Bayes estimate of θ1 under quadratic loss.                                                      (5)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 6,
          "text": "Calculate the expected claims for year 6 of each category under the assumptions of\n                           Empirical Bayes Credibility Theory Model 1.                                                              (6)",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 2,
          "text": "Explain the main differences between the approach in part (i) and that in part (ii).                                  (2)",
          "topic": null
        },
        {
          "label": "iv",
          "marks": 2,
          "text": "Explain why the assumption of a Poisson distribution with a constant parameter may\n                           not be appropriate and describe how each approach might be generalised.                                  (2)",
          "topic": null
        }
      ],
      "solution": "i)\n\n                                                                                                            Page 8 of 14\n\fIAI                                                                                            CS1A-1222\n\n      We need to find the parameters of the Gamma distribution, say α and λ\n      Then\n                                          𝐸(𝑋)       𝛼⁄             50\n                                                  = 𝛼 𝜆 = 𝜆=            =2\n                                        𝑉𝑎𝑟(𝑋)        ⁄𝜆2           25\n      And hence 𝛼 = 𝐸(𝑋) ∗ 𝜆 = 50 ∗ 2 = 100\n      The posterior distribution is given by:\n                                           𝑓(𝜃1 |𝑥) ∝ 𝑓(𝑥|𝜃1) ∗ 𝑓(𝜃1 )\n                                                        𝑛\n                                        ∝ (∏5𝑗=1 𝑒 −𝜃1 𝜃1 1𝑗 ) * 𝜃1𝛼−1 𝑒 −𝜆𝜃1\n                                                                  𝛼+∑5   𝑛1𝑗 −1\n                                         ∝ 𝑒 −(𝜆+5)𝜃1 𝜃1 𝑗=1\n      Which is the pdf of a gamma distribution with parameters\n                                                      5\n\n                                              𝛼 + ∑ 𝑛1𝑗 = 100 + 232 = 332\n                                                   𝑗=1\n      And 𝜆 + 5 = 7\n      Under quadratic loss the Bayes estimate is the mean of the posterior distribution. So we have an\n      estimate of 332/7 = 47.43                                                                                 [5]\n\n      ii)\n      We have\n                                                                232\n                                                          𝑛̅1 =\n                                                                 5\n                                                                260\n                                                          𝑛̅2 =\n                                                                 5\n                                                                145\n                                                          𝑛̅3 =\n                                                                 5\n                               46.4+52+29\n      This gives 𝑛̅ =               3\n                                          = 42.4667\n\n       5                           5          5\n                           2\n      ∑(𝑛1𝑗 − 𝑛̅1 ) = ∑ 𝑛1𝑗 − 2 ∑ 𝑛1𝑗 ∗ 𝑛̅1 + 5 ∗ 𝑛̅12\n      𝑗=1                         𝑗=1        𝑗=1\n      = 11,434 – 2*232 * 46.4 + 5*46.42\n      = 669.2\n\n      Similarly,\n\n       5                           5          5\n                           2\n      ∑(𝑛2𝑗 − 𝑛̅2 ) = ∑ 𝑛2𝑗 − 2 ∑ 𝑛2𝑗 ∗ 𝑛̅2 + 5 ∗ 𝑛̅22\n      𝑗=1                         𝑗=1        𝑗=1\n      = 14028 – 2* 260 * 52.0 + 5*522\n      = 508\n\n       5                           5          5\n                           2\n      ∑(𝑛3𝑗 − 𝑛̅3 ) = ∑ 𝑛3𝑗 − 2 ∑ 𝑛3𝑗 ∗ 𝑛̅3 + 5 ∗ 𝑛̅32\n      𝑗=1                         𝑗=1        𝑗=1\n      = 4399 – 2* 145 * 29.0 + 5*292\n      = 194\n\n      So\n\n                   1   1\n      E(s2(θ)) = 3 * 4 * (669.2 + 508 + 194) = 114.2667\n                       1                                                          1\n      Var(m(θ)) = 2 * ((46.2-42.4667)2+ (52-42.4667)2 + (29-42.4667)2) - 5 ∗ 114.2667\n               = 121\n\n                                                                                                 Page 9 of 14\n\f    IAI                                                                                                    CS1A-1222\n\n              So\n                            5\n              Z=         114.2667   = 0.8411\n                      5+\n                           121\n\n              So expected claims for next year are:\n              Cat 1 0.1589 × 42.4667 + 0.8411 × 46.4 = 45.78\n              Cat 2 0.1589 × 42.4667 + 0.8411 × 52 = 50.49\n              Cat 3 0.1589 × 42.4667 + 0.8411 × 29 = 31.14\n\n              iii)\n\n              The main differences are that:\n              • The approach under (i) makes use of prior information about the distribution of θ1 whereas the\n              approach in (ii) does not.\n              • The approach under (i) uses only the information from the first category to produce a posterior\n              estimate, whereas the approach under (ii) assumes that information from the other categories can\n              give some information about category 1.\n               • The approach under (i) makes precise distributional assumptions about the number of claims\n              (i.e. that they are Poisson distributed) whereas the approach under (ii) makes no such assumptions.           [2]\n\n              iv)\n\n              The insurance policies were newly introduced 5 years ago, and it is therefore likely that the volume\n              of policies written has increased (or at least not been constant) over time. The assumption that the\n              number of claims has a Poisson distribution with a fixed mean is therefore unlikely to be accurate,\n              as one would expect the mean number of claims to be proportional to the number of policies. Let\n              Pij be the number of policies in force for risk i in year j.\n              Then the models can be amended as follows: The approach in (i) can be taken assuming that that\n              the mean number of claims in the Poisson distribution is Pijθi . The approach in (ii) can be\n              generalised by using EBCT Model 2 which explicitly incorporates an adjustment for the volume of\n              risk                                                                                                          [2]",
      "has_math": true,
      "session": "2022-12",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2022-12_QP.pdf",
      "source_sol": "raw/CS1A_2022-12_SOL.pdf"
    },
    {
      "q_num": 10,
      "marks": 29,
      "topic": "regression_glm",
      "subtopics": [],
      "stem": "An actuarial trainee fit a simple linear model to detect bus cancellation charges (Y).\n\n         As an exploratory analysis, trainee determined that correlation coefficient between Y and X\n         is 0.623 as a first step.\n\n         The slope and intercept of the model are b = -5 and a = 10.\n\n         Further, after fitting the model, histogram of residuals (shown below) was prepared as per\n         manager’s request:\n\n                           60\n\n                           50\n\n                           40\n               Frequency\n\n                           30\n\n                           20\n\n                           10\n\n                               0\n                                   -0.4 -0.3 -0.2 -0.1   0     0.1   0.2       0.3   0.4       0.5   0.6    0.7\n                                                               Residual",
      "parts": [
        {
          "label": "i",
          "marks": 2,
          "text": "Plot graph showing relationship between Y and X for X = -2 to 2\n                           Note: Compute Y for each X (-2,-1, 0, 1, 2) and then plot a freehand graph.                              (2)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 2,
          "text": "Comment on the following:\n\n             a)   Validity of assumption of the linear model.                                             (2)\n\n             b) Possibility of error made by Actuarial Trainee.                                           (2)\n\n      Trainee fits a generalised linear model to predict number of free bus cancellations for prime\n      members. The following data for 20 prime members collected for 2 cities:\n\n       City I      2       2          0       0       1      0       0      0       1        0\n       City II     1       2          1       0       2      1       1      0       2        2\n\n      Trainee uses Poisson distribution to analyse bus cancellation charges and fits following\n      models:\n\n      Model 1 :        log 𝜇𝑖 = 𝛼\n                                                          1 𝑓𝑜𝑟 𝐶𝑖𝑡𝑦 𝐼\n      Model 2 :        log 𝜇𝑖 = 𝛼 + 𝛽 𝑥𝑖 𝑤ℎ𝑒𝑟𝑒 𝑥𝑖 = {\n                                                          0 𝑓𝑜𝑟 𝐶𝑖𝑡𝑦 𝐼𝐼",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 2,
          "text": "Show that Poisson distribution is a member of the exponential family of distributions.         (2)",
          "topic": null
        },
        {
          "label": "iv",
          "marks": 2,
          "text": "a) Calculate the maximum likelihood estimator for α and β under Model 1 and 2.                (5)\n\n            b) Compute the probability of 3 cancellations for City I and II under Model 2.                (2)",
          "topic": null
        },
        {
          "label": "v",
          "marks": 5,
          "text": "Compute Scaled deviance for Model 1 and 2.\n             Note: y * log y = 0 for y=0 can be assumed.                                                  (5)",
          "topic": null
        },
        {
          "label": "vi",
          "marks": 2,
          "text": "Suggest which model is better by using appropriate statistic                                 (2)\n\n      Trainee also tried another model.\n\n                                      𝛿 𝑓𝑜𝑟 𝐶𝑖𝑡𝑦 𝐼\n      Model 3:           log 𝜇𝑖 = {\n                                      𝛾 𝑓𝑜𝑟 𝐶𝑖𝑡𝑦 𝐼𝐼",
          "topic": null
        },
        {
          "label": "vii",
          "marks": 1,
          "text": "For Model 3, compute the following:\n\n            a) MLE for 𝛿 and 𝛾                                                                            (2)\n\n            b) AIC (Akaike’s Information Criterion)                                                       (1)",
          "topic": null
        },
        {
          "label": "viii",
          "marks": 2,
          "text": "Compare Model 2 and Model 3 and comment. Are there any similarities between the\n            models 2 and 3?                                                                               (2)",
          "topic": null
        },
        {
          "label": "ix",
          "marks": 2,
          "text": "For review, Plot of Pearson residual is used. Write down its disadvantages for checking\n             Poisson distribution.                                                                        (2)",
          "topic": null
        }
      ],
      "solution": "i)\n\n                  X                     -2     -1        0        1          2\n                  Y                     20     15       10        5          0\n\n                                                    Y = 10 -5x\n                                                       25\n\n                                                       20\n\n                                                       15\n                      Y\n\n                                                       10\n\n                                                        5\n\n                                                        0\n                       -3              -2      -1           0         1          2          3\n                                                            X\n\n                                                                                                           Page 10 of 14\n\fIAI                                                                                                  CS1A-1222\n\n      ii)\n       a)\n       Assumption of linear model does not seem valid in this case. The histogram shows a positively\n       skewed distribution for the residuals and suggests that the errors ~N(0,ϭ2) distribution does not fit\n       in this case.                                                                                                  [2]\n\n       b)\n       The plot in (i) suggests a negative relationship between Y and X with slope parameter = -5. However,\n       the correlation shows a positive relationship. Thus, it seems an error has been made in either fitting\n       the model or while computing correlation.                                                                      [2]\n\n      iii)\n\n       For exponential family, we have write in the form:\n\n                                                         𝑦𝜃 − 𝑏(𝜃)\n                                          𝑔(𝑦) = 𝑒𝑥𝑝 [             + 𝑐(𝑦, ∅)]\n                                                           𝑎(∅)\n       Poisson Distribution is:\n                                                           𝑒 −𝜇 𝜇 𝑦\n                                                  𝑓(𝑦) =\n                                                              𝑦!\n                                                       𝑦𝑙𝑜𝑔𝜇 − 𝜇\n                                          𝑓(𝑦) = 𝑒𝑥𝑝 [              − log 𝑦!]\n                                                           1\n       Where:\n\n                                  𝑏(𝜃) = 𝜇, 𝜃 = 𝑙𝑜𝑔𝜇 , 𝑎(∅) = 1 , 𝑐(𝑦, ∅) = − log 𝑦!                                  [2]\n\n      iv)\n       a)\n       Log Likelihood Function is:\n\n                                      𝐿𝑜𝑔 𝐿 = ∑ 𝑦𝑖 𝑙𝑜𝑔𝜇𝑖 − ∑ 𝜇𝑖 − ∑ log 𝑦𝑖 !\n\n       Model 1 log 𝜇𝑖 = 𝛼\n                              20                 20                             20\n                                             𝛼                              𝛼\n                   𝐿𝑜𝑔 𝐿 = 𝛼 ∑ 𝑦𝑖 − 20𝑒 − ∑ log 𝑦𝑖 ! = 18𝛼 − 20𝑒 − ∑ log 𝑦𝑖 !               − (∗)\n                             𝑖=1                 𝑖=1                            𝑖=1\n\n       Differentiating this with respect to α , and setting the result equal to 0, we get\n\n                                                                18\n                                    18 − 20𝑒 𝛼̂ = 0 → 𝛼̂ = log ( ) = −0.1054\n                                                                20\n\n                                                   1 𝑓𝑜𝑟 𝐶𝑖𝑡𝑦 𝐼\n       Model 2 :     log 𝜇𝑖 = 𝛼 + 𝛽 𝑥𝑖 𝑤ℎ𝑒𝑟𝑒 𝑥𝑖 = {\n                                                    0 𝑓𝑜𝑟 𝐶𝑖𝑡𝑦 𝐼𝐼\n\n                                                                                                      Page 11 of 14\n\fIAI                                                                                                         CS1A-1222\n\n                                        20             20                                     20\n                                                                    𝛼                𝛼+𝛽\n                        𝐿𝑜𝑔 𝐿 = 𝛼 ∑ 𝑦𝑖 + 𝛽 ∑ 𝑦𝑖 𝑥𝑖 − 10𝑒 − 10𝑒                              − ∑ log 𝑦𝑖 !\n                                        𝑖=1            𝑖=1                                    𝑖=1                           [5]\n                                                                                20\n\n                             = 18𝛼 + 6𝛽 − 10𝑒 𝛼 − 10𝑒 𝛼+𝛽 − ∑ log 𝑦𝑖 ! − (∗∗)\n                                                                            𝑖=1\n\n      Differentiating this in turn with respect to α and β and setting it equal to 0, we get\n                                                             ̂\n                                      18 − 10 𝑒 𝛼̂ − 10 𝑒 𝛼̂+𝛽1 = 0        (∗)\n                                                 ̂                      ̂\n                                    6 − 10 𝑒 𝛼̂+𝛽 = 0 → 6 = 10 𝑒 𝛼̂+𝛽 (∗∗)\n\n      Substituting this in (*), we get\n                                                                   12\n                                  18 − 10 𝑒 𝛼̂ − 6 = 0 → 𝛼̂ = log ( ) = 0.1823\n                                                                   10\n\n      Substituting 𝛼̂ 𝑖𝑛 (∗∗), 𝑤𝑒 𝑔𝑒𝑡\n                                                   ̂            6\n                                   6 = 10 𝑒 0.1823+𝛽 → 𝛽̂ = ln ( ) = −0.6932\n                                                                12\n\n      b)\n      City 1 : log 𝜇𝐶𝑖𝑡𝑦1 = 𝛼 + 𝛽 → 𝜇𝐶𝑖𝑡𝑦1 = 0.6\n      City 2 : log 𝜇𝐶𝑖𝑡𝑦2 = 𝛼 → 𝜇𝐶𝑖𝑡𝑦2 = 1.2\n                                         𝑒 −0.6 0.63\n      P(cancellation =3) for City 1 =                  = 0.0126\n                                             3!\n\n      P(cancellation =3) for City 1 = 0.0867                                                                                [2]\n\n      v)\n\n      Scaled Deviance = 2 (log Ls – Log Lm)\n      Where LogLs is the value of log likelihood function for the saturated model\n      And Log Lm is the value of log likelihood function for Model M\n      For Saturated model, 𝜇𝑖 = 𝑦𝑖, . This implies\n                             𝐿𝑜𝑔 𝐿 = ∑ 𝑦𝑖 𝑙𝑜𝑔𝑦𝑖 − ∑ 𝑦𝑖 − ∑ log 𝑦𝑖 ! = −13.84\n\n      Using iv.a. (*) and (**),\n                                                                                     20\n                                                                            ̂\n                                                                            𝛼\n                                   𝐿𝑜𝑔 𝐿(𝑚𝑜𝑑𝑒𝑙 1) = 18𝛼̂ − 20𝑒 − ∑ log 𝑦𝑖 !\n                                                                                     𝑖=1\n                                                                                       20\n\n                                  = 18 ∗ −0.1054 − 20 𝑒𝑥(−0.1054) − ∑ log 𝑦𝑖 !\n                                                                                      𝑖=1\n                                                            = −24.055\n\n      Similarly,\n                                                                                               20\n                                                                                         ̂\n                        𝐿𝑜𝑔 𝐿(𝑀𝑜𝑑𝑒𝑙 2) = 18𝛼̂ + 6𝛽̂ − 10𝑒 − 10𝑒         ̂\n                                                                        𝛼             ̂ +𝛽\n                                                                                      𝛼\n                                                                                             − ∑ log 𝑦𝑖 !\n                                                                                               𝑖=1\n\n      putting 𝛼̂ = 0.1823 𝑎𝑛𝑑 𝛽̂ = −0.6932 in above equations\n      logL (model2) = -23.036\n\n      Scaled deviance model 1 = 20.43\n      Scaled deviance model 2 = 18.39\n\n                                                                                                            Page 12 of 14\n\fIAI                                                                                              CS1A-1222\n\n      vi)\n\n      We can compare Model 1 and 2 by using chi-square distribution and scaled deviance difference.\n\n      Scaled deviance difference = 2[LogL(model 1) – LogL(model 1)) = 2.04\n\n      This follows chi-square distribution with 2-1 =1 degrees of freedom\n\n      At 5% level of confidence, value is 3.84.\n\n      No significant improvement. Thus, prefer Model 1                                                            [2]\n\n      vii)\n       a)\n                                    𝛿 𝑓𝑜𝑟 𝐶𝑖𝑡𝑦 𝐼\n      Model 3:         log 𝜇𝑖 = {\n                                    𝛾 𝑓𝑜𝑟 𝐶𝑖𝑡𝑦 𝐼𝐼\n\n      So Log Likelihood function is:\n\n                                         10         20                             20\n                                                                   𝛾           𝛿\n                          𝐿𝑜𝑔 𝐿 = 𝛿 ∑ 𝑦𝑖 + 𝛾 ∑ 𝑦𝑖 − 10𝑒 − 10𝑒 − ∑ log 𝑦𝑖 !\n                                        𝑖=1         𝑖=11                           𝑖=1\n                                                                         20\n\n                                     = 6𝛿 + 12𝛾 − 10𝑒 𝛾 − 10𝑒 𝛿 − ∑ log 𝑦𝑖 !\n                                                                         𝑖=1\n\n      Differentiating this , and setting the result equal to 0, we get\n\n                                         ̂                  6\n                                 6 − 10𝑒 𝛿 = 0 → 𝛿̂ = log ( ) = −0.5108\n                                                           10\n                                                             12\n                                 12 − 10𝑒 𝛾̂ = 0 → 𝛾̂ = log ( ) = 0.1823                                          [2]\n                                                             10\n\n      b)\n\n      Scaled Deviance = 18.39\n\n      AIC = -2 * LogL (model3) + 2 X number of paramters = 50.72                                                  [1]\n\n      viii)\n\n      Model 3 and Model 2 are essential the same but represented in different way.\n\n      Under Model 2, exp(α) represents mean of City II and exp(α + β) represents City I. In Model 3, α +\n      β is given as .\n\n      In other words, exp(β) under Model 2 is expressed as change in mean between City I and City II. In\n      Model 3, mean of the 2 cities are expressed separately.\n\n      Thus, it can be observed that Scaled Deviance for Model 2 and 3 are same.\n\n      ix)                                                                                                         [2]\n                                Pearson residuals is defined as:                                           [29 Marks]\n\n                                                                                                  Page 13 of 14\n\fIAI                                                                                         CS1A-1222\n\n                                                             𝑦𝑖 − 𝑦̂𝑖\n                                                               √𝑦̂𝑖\n\n                            As variance equals mean for poison distribution.\n\n      Pearson residual distribution is often skewed for non-normal distributed data. This makes the\n      interpretation of residual plots difficult.\n\n                                     ************************\n\n                                                                                             Page 14 of 14",
      "has_math": true,
      "session": "2022-12",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2022-12_QP.pdf",
      "source_sol": "raw/CS1A_2022-12_SOL.pdf"
    },
    {
      "q_num": 1,
      "marks": 10,
      "topic": "distributions",
      "subtopics": [],
      "stem": "",
      "parts": [
        {
          "label": "i",
          "marks": 1,
          "text": "Let M(t) be the Moment Generating Function (MGF) of a random variable Y. Given\n           below are four MGFs written in terms of M(t) of four different random variables.\n           Identify which one of the following is NOT a valid MGF.\n\n             A. M(t) * M(3t)\n             B. e-3t * M(0.5t)\n                 2\n             C. 3 * M(t)\n                     1\n             D. M(3 𝑡 )                                                                                         (1)\n\n         Let X be a two-parameter exponential random variable such that X ~ Exp (λ, a). It has the\n         following probability density function:\n\n                                         f(x) = λ e-λ (x-a), x ≥ a, where λ, a > 0\n\n                                                                                     𝑡 −1",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 3,
          "text": "Show that the moment generating function of X is given by: (1 − 𝜆)              * eat.              (3)",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 2,
          "text": "Using the result derived in part (ii), calculate E(X).                                             (2)",
          "topic": null
        },
        {
          "label": "iv",
          "marks": 1,
          "text": "Which of the following is an expression representing the distribution function Fx(X) of\n            random variable X?\n\n             A. e λa * (1 – e- λx)\n             B. e λa * (e- λa – e- λx)\n             C. e- λx * (1 – e- λa)\n             D. e- λa * (e λa – e- λx)                                                                          (1)",
          "topic": null
        },
        {
          "label": "v",
          "marks": 3,
          "text": "If λ = 0.10, a = 0.05, simulate two values from X using the distribution function as\n             determined in part (iv) using values 0.226 and 0.304 from U (0, 1).                               (3)",
          "topic": null
        }
      ],
      "solution": "i)        Correct Answer is Option C\n\n              It is a basic property of MGFs that for any random variable at t=0, the value of the MGF shall be\n              equal to 1. This is not true in case of option C which will return a value of 2/3, hence it is not\n              valid.                                                                                                   [1]\n\n      ii)     Mx(t) = E(etx) = ∫a∞ etx λ e-λ(x-a)dx                                                                    [1]\n\n                               =∫a∞ eλa λ e-x(λ-t)dx\n                               = eλa λ e-x(λ-t)/-(λ-t)|a∞                                                              [1]\n                               =λ/(λ-t) eat , Provided t<λ\n                               =(1-t/λ)-1 eat                                                                          [1]\n\n  iii)        Mx(t) =(1-t/λ)-1 eat\n              M’x(t) =(1-t/λ)-1 eat a+ (1-t/λ)-2 1/λ eat                                                               [1]\n              E(x) = M’x(0)=a+1/λ                                                                                      [1]\n\n  iv)         Correct Answer is Option B\n\n              The result can be obtained by integrating the pdf from a to x. It returns the following result:\n              - exp(λa) * [ exp(-λt) ] xa\n              = - exp(λa) * [ exp(-λx) - exp(-λa)]\n              = exp(λa) * [ exp(-λa) - exp(-λx) ]                                                                      [1]\n\n      v)      FX(x) = u = exp(λa) * [ exp(-λa) - exp(-λx) ]\n              u / exp (λa) - exp(-λa) = - exp(-λx)\n              exp (-λx) = exp(-λa) - u / exp (λa)                                                                      [1]\n\n              Taking logs to the base e on both sides and substituting values of λ and a,\n              -0.10 * x = ln ( exp(-0.10*0.05) – u / exp(0.10*0.05) )\n              x = - 10 * ln ( (1-u) * exp(-0.10*0.05) )\n              x = - 10 * ln ( (1-u) * 0.995012)                                                                        [1]\n\n              Using the two values given in the question,\n              @ u = 0.226, x = -10 * ln ( (1-0.226) * 0.995012) = 2.61183\n              @ u = 0.304, x = -10 * ln ( (1-0.304) * 0.995012) = 3.67406                                              [1]",
      "has_math": true,
      "session": "2023-05",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2023-05_QP.pdf",
      "source_sol": "raw/CS1A_2023-05_SOL.pdf"
    },
    {
      "q_num": 2,
      "marks": 10,
      "topic": "distributions",
      "subtopics": [],
      "stem": "A general insurance company is analyzing its portfolio of accidental insurance claims.\n         The random variables X and Y represents the accidental loss amounts and allocated\n         operating expenses respectively per policy.\n\n         Joint probability density function is given by:\n\n                                    f XY (x, y) = 3/106 e-(x/1000), 0 < 3y < x < ∞\n                                                =0                     otherwise\n\n         The industry has a practice of having ratio of 1:3(Y:X) between allocated operating\n         expenses and accidental loss amounts. However, the Chief Financial Officer of the company\n         feels that there is a comprehensive expenses management framework in the company and\n         hence this ratio has fallen below 1:4.",
      "parts": [
        {
          "label": "i",
          "marks": 3,
          "text": "Using the joint density function f XY (x, y) given above, calculate the probability that the\n             ratio has fallen below 1:4. You are given that current level of accidental loss amount per\n             policy is 1,000.\n\n             Hint: Calculate P (Y < X / 4)                                                                      (3)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 2,
          "text": "Derive an expression for the marginal density function of random variable Y.                   (2)",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 1,
          "text": "Based on the expression obtained in part (ii) identify the correct distribution and\n             parameters for random variable Y.\n\n              A. Y ~ Gamma(α=2, λ=1/1000)\n              B. Y ~ Exp(λ=3/1000)\n              C. Y ~ Gamma(α=2, λ=3/1000)\n              D. Y ~ χ2 with 3 degrees of freedom                                                          (1)\n\n         It is given that the marginal density function of random variable X i.e. fX(x) is given by the\n         following expression:\n\n                                       fX(x) = x/10002 e-x/1000    3y < x < ∞\n                                             =0                     otherwise",
          "topic": null
        },
        {
          "label": "iv",
          "marks": 1,
          "text": "Justify why random variables X and Y cannot be independent based on fX(x) as given\n               above and fY(y) as determined in part (ii).                                                 (1)",
          "topic": null
        },
        {
          "label": "v",
          "marks": 1,
          "text": "Identify the correct joint density function for which the random variables X and Y with\n               marginal density functions as above are considered to be independent?\n\n               A.    g XY (x, y) = 3x/109 e-(x+3y) / 1000    0 < 3y < x < ∞\n                                 =0                           otherwise\n               B.    g XY (x, y) = 3x/106 e-(x+3y) / 1000    0 < 3y < x < ∞\n                                 =0                           otherwise\n               C.    g XY (x, y) = 3xy/109 e-(x+3y) / 1000   0 < 3y < x < ∞\n                                 =0                           otherwise\n               D.    g XY (x, y) = 3y/109 e(x – 3y)/1000     0 < 3y < x < ∞\n                                 =0                           otherwise                                    (1)",
          "topic": null
        },
        {
          "label": "vi",
          "marks": 2,
          "text": "Suppose you decide to use the joint density function g XY (x, y) as determined in part\n               (v), determine the conditional expectation E (Y | X > 950).                                 (2)",
          "topic": null
        }
      ],
      "solution": "i)        Pr(Y < X / 4)\n              = Pr (Y < 250)\n              = Pr (3Y < X < ∞, 0 < Y < 250)                                                                         [0.5]\n\n              = ∫0250 (∫3y∞ 3/106 e-x/1000 dx) dy\n\n              = ∫0250 (3/106 e-x/1000/ (-1/1000))3y∞ dy                                                                [1]\n                                                                                                           Page 2 of 11\n\fIAI                                                                                                       CS1A-0523\n\n              =∫0250 3/103 (e-3y/1000-e-∞) dy\n\n              = 3/103 (e-3y/1000/(-3/1000))0250\n\n              = (e0 – e -3(250/1000))\n              = 1 – exp(-0.75)                                                                                         [1]\n              = 1 – 0.472\n              = 0.528\n              The probability that the ratio has in fact reduced to 1:4 is 0.528.                                     [0.5]\n                      ∞\n      ii)     f(y) =∫𝑥=3𝑦 3/10^6 e-x/1000 dx\n\n                   =3/106 (e-x/1000/(-1/1000))3y∞\n                   = 3/1000 e-3y/1000                                                                                  [2]\n\n  iii)        Correct Answer is Option B\n              From part (ii), we can observe that Y is an exponential random variable with the value of\n              parameter λ = 3/1000.                                                                                    [1]\n\n  iv)         For random variables X and Y to be independent, fX(x) * fY(y) = f XY (x, y) must be true.\n\n              From observation itself, we understand that since f XY (x, y) does not contain any term in y, the\n              above is not true and hence X and Y cannot be independent.                                               [1]\n\n      v)      Correct Answer is Option A\n              fX(x) * fY(y)\n              = x/10002 e-x/1000 * 3/1000 e-3y/1000\n              = 3x/109 * e-(x+3y)/1000\n              = g XY (x, y) as mentioned in Option A.                                                                  [1]\n\n  vi)         We have decided to use g XY (x, y) and we know that for this joint probability density\n              function X and Y are independent random variables.\n              Hence, conditional expectation E (Y | X > 950) is independent of X and hence it is equal to\n              E(Y).                                                                                                    [1]\n\n              From part (iii), we know that Y ~ Exp (3/1000).\n\n              Hence E (Y | X > 950) = E(Y) = 1000/3 = 333.33                                                         [1]",
      "has_math": true,
      "session": "2023-05",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2023-05_QP.pdf",
      "source_sol": "raw/CS1A_2023-05_SOL.pdf"
    },
    {
      "q_num": 3,
      "marks": 10,
      "topic": "regression_glm",
      "subtopics": [
        "distributions"
      ],
      "stem": "",
      "parts": [
        {
          "label": "i",
          "marks": 1,
          "text": "Which of the following is a genuine difference between a “general” linear model and\n               a “generalised” linear model?\n\n              A. The response variable is normally distributed in case of a “general” linear model\n                 whereas it can be non-normal in case of “generalised” linear models.\n              B. Link function can be used only for “generalised” linear models and cannot be used\n                 for a “general” linear model.\n              C. “Generalised” linear models can be non-linear in terms of the covariates whereas a\n                 “general” linear model has to be linear in terms of both parameters and covariates.\n              D. Interaction between various explanatory variables can be added in case of\n                 “generalised” linear models which is not possible in case of a “general” linear model.    (1)\n\n         Consider the discrete random variable Y with the following probability density function:\n\n         f(y,µ) = n C ny * µny * (1 - µ)n-ny            where y = 0, 1/n, 2/n, 3/n, ………, 1",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 4,
          "text": "Show that the above discrete distribution belongs to the exponential family of\n                distributions and specify different parameters of exponential family of distributions.     (4)\n\n         Let us define Z as a binomial variable such that Z = nY",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 3,
          "text": "Derive expressions for the mean and variance of the discrete distribution E(Z) and\n                Var(Z), using your answer in part (ii). Also check whether the expressions match with\n                mean and variance of Z ~ Bin(n,µ).                                                         (3)",
          "topic": null
        },
        {
          "label": "iv",
          "marks": 1,
          "text": "Is the random variable Y a Bernoulli variable with parameter µ?                            (1)",
          "topic": null
        },
        {
          "label": "v",
          "marks": 1,
          "text": "Choose the canonical link function to be used in this case from the following options.\n\n               A. g(µ) = log⁡(µ⁡/⁡(1⁡– ⁡µ) )\n               B. g(µ) = µ\n               C. g(µ) = log µ\n               D. g(µ) = 1 / µ                                                                             (1)",
          "topic": null
        }
      ],
      "solution": "i)        Correct Answer is Option A\n              All other options represent statements which are not true.                                               [1]\n\n      ii)     For exponential family, we have write in the form:\n\n                                                          𝑦𝜃 − 𝑏(𝜃)\n                                           𝑓(𝑦) = 𝑒𝑥𝑝 [             + 𝑐(𝑦, ∅)]\n                                                            𝑎(∅)                                                      [0.5]\n                                                                                                           Page 3 of 11\n\fIAI                                                                                                      CS1A-0523\n\n              The given discrete distribution is:\n              𝑓(𝑦)\n              = n C ny * µny * (1 - µ)n-ny\n              = exp ( n( y logµ + (1 – y) log(1 – µ)) + log (n C ny) )\n              = exp ( n( y log (µ / (1 – µ)) + log(1 – µ)) + log (n C ny) )\n\n              This is in the form of the exponential family as mentioned above.\n\n              Where:\n              𝜃 = log (µ / (1 – µ) )                                                                                 [0.5]\n              ∅=𝑛\n              a(∅) = 1 / ∅                                                                                           [0.5]\n              c(y, ∅) = log (n C ny)                                                                                 [0.5]\n              Using 𝜃 = log (µ / (1 – µ)), we can show that µ = 𝑒 𝜃 / (1 + 𝑒 𝜃 )\n              b(𝜃) = 𝑙𝑜𝑔 (1 + 𝑒 𝜃 )                                                                                   [1]\n\n   iii)       E(Y) = b’(𝜃) = 𝑒 𝜃 / (1 + 𝑒 𝜃 ) = µ                                                                    [0.5]\n                                     𝜃 2\n              V(µ) = 𝑒 𝜃 / (1 + 𝑒 ) = µ (1 − µ)                                                                      [0.5]\n\n              V(Y) = V(µ) ∗ a(∅) = µ (1 − µ) / n                                                                     [0.5]\n\n              E(Z) = n * E(Y) = n * µ                                                                                [0.5]\n\n              Var(Z) = n2 * V(Y) = n * µ (1 − µ)                                                                     [0.5]\n\n              This corresponds to the mean n * µ and variance n * µ (1 − µ) of Z ~ Bin(n, µ)                         [0.5]\n   iv)        Bernoulli variable say X with parameter µ would have the following probability distribution:\n              f(x) = µ ∗ (1 − µ)            𝑓𝑜𝑟 𝑥 = 0,1.\n              But the probability density function of Y is\n              f(y,µ) = n C ny * µny * (1 - µ)n-ny             where y = 0, 1/n, 2/n, 3/n, ………, 1\n              Hence Y is not Bernoulli. Although mean and variance of Y equal to the variance of a Bernoulli\n              distribution, it is not sufficient to conclude that Y is a Bernoulli variable.                          [1]\n\n      v)      Correct Answer is Option A\n              In case of binomial distribution, logit link function is the canonical link function.                  [1]",
      "has_math": false,
      "session": "2023-05",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2023-05_QP.pdf",
      "source_sol": "raw/CS1A_2023-05_SOL.pdf"
    },
    {
      "q_num": 4,
      "marks": 10,
      "topic": "inference",
      "subtopics": [],
      "stem": "A private tutor teaches actuarial subject CS2 in two batches – “Elite” and “Zenith”. Those\n         who have passed CS1 with more than 70 marks are a part of the “Elite” batch and others are\n         a part of the “Zenith” batch.\n\n         He recently conducted a mock test for CS2 and is trying to analyse their performance by\n         comparing the number of questions answered correctly. The mock question paper had a total\n         of 7 questions and the data relating to number of correctly attempted questions is presented\n         below:\n\n                      Batch               “Elite” “Zenith”\n             Total number of students        6       11\n          Total number of correct answers   33       37\n\n         He decides to use binomial distribution to model the number of questions correctly\n         attempted by each candidate and defines two variables XE and XZ for “Elite” and “Zenith”\n         batches respectively such that:\n\n                                               XE ~ Bin (n=7, pE)\n                                               XZ ~ Bin (n=7, pZ)",
      "parts": [
        {
          "label": "i",
          "marks": 1,
          "text": "Which of the following is NOT a valid method used to determine a point estimate for\n           the value of the unknown parameter using the information provided by above sample?\n\n               A. Method of Percentiles.\n               B. Non-Parametric Bootstrap Method.\n               C. Parametric Bootstrap Method.\n               D. None of the above.                                                                       (1)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 1,
          "text": "Determine the estimates p̂E and p̂Z and combined estimate p̂ using the method of\n            moments.                                                                                       (1)\n\n         The tutor suggests an alternative model wherein he decides to use the method of maximum\n         likelihood. He continues to use binomial distribution to model XE and XZ with n = 7, but for\n         Elite Batch he uses the parameter 2Ɵ and for Zenith Batch he uses the parameter Ɵ, where\n         Ɵ < 0.5.",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 2,
          "text": "Identify which one of the following corresponds to the log likelihood function of Ɵ\n                 given the observed data.\n\n               A. log L ∝ 33 In (2Ɵ) + 9 In (1-2Ɵ) + 37 In (Ɵ) + 40 In (1-Ɵ)\n               B. log L ∝ 9 In (2Ɵ) + 33 In (1-2Ɵ) + 40 In (Ɵ) + 37 In (1-Ɵ)\n               C. log L ∝ 33 In (Ɵ) + 9 In (1-Ɵ) + 37 In (2Ɵ) + 40 In (1-2Ɵ)\n               D. log L ∝ 9 In (Ɵ) + 33 In (1-Ɵ) + 40 In (2Ɵ) + 37 In (1-2Ɵ)                                        (2)",
          "topic": null
        },
        {
          "label": "iv",
          "marks": 4,
          "text": "Show using your answer in part (iii), that the maximum likelihood estimate for Ɵ is\n                 Ɵ̂ =0.412. You are NOT required to check that it is a maximum.                                     (4)",
          "topic": null
        },
        {
          "label": "v",
          "marks": 2,
          "text": "Without performing any formal test, by completing the below table, state which\n                 method gives a better estimate.\n\n                  Total number of correct answers                     “Elite” Batch          “Zenith” Batch\n                  Method of Moments Estimate (using                     _______                 _______\n                  combined estimate p̂)\n                  Method of Maximum Likelihood Estimate                    _______              _______\n                  Observed Values                                             33                   37               (2)",
          "topic": null
        }
      ],
      "solution": "i)        Correct Answer is Option D\n              All methods are valid methods to estimate the value of parameter based on information from a\n              sample.                                                                                                 [1]\n\n                        33                      37                    70\n      ii)     p̂E =             = 0.786, p̂Z = (11∗7) =0.481, p̂ = (17∗7) =0.588                                      [1]\n                      (6 ∗ 7)\n\n   iii)       Correct Answer is Option A\n\n              let ni be the total number of questions for the whole batch i,, Bi be the total number of correct\n              answer by Batch i.\n                      L(b;Ɵ) = (2Ɵ)BE (1-2Ɵ)7nE-BE (Ɵ)BZ (1-Ɵ)7nZ-BZ *constant                                        [2]\n\n                                                                                                           Page 4 of 11\n\fIAI                                                                                                           CS1A-0523\n\n                    l(b;Ɵ ) = InL(b;Ɵ)\n                      = 33In(2Ɵ)+(42-33) In(1-2Ɵ)+37In(Ɵ) +(77-37) In(1-Ɵ)+ constant\n                      = 33 In(2Ɵ) +9In(1-2Ɵ)+37 In(Ɵ) +40 In(1-Ɵ)+constant\n\n  iv)                     𝑑𝑙    66        18         37        40\n                               = 2Ɵ - 1−2Ɵ + Ɵ - 1−Ɵ\n                          𝑑Ɵ\n\n                                     70         18        40\n                                  = Ɵ − 1−2Ɵ - 1−Ɵ                                                                        [1.5]\n              Set equal to zero and solve\n                               70(1−2Ɵ)(1−Ɵ)−18Ɵ(1−Ɵ)−40Ɵ(1−2Ɵ)\n                                                                          =0\n                                               Ɵ(1−2Ɵ)(1−Ɵ)\n                         70-210 Ɵ-140 Ɵ -18 Ɵ+18 Ɵ -40 Ɵ+80 Ɵ2 =0\n                                                      2             2\n\n                         238 Ɵ2 -268 Ɵ +70 = 0\n                         Ɵ = 0.412 or 0.714\n                         As Ɵ<0.5, Ɵ̂ =0.412\n\n      v)       Total number of correct answers                          “Elite” Batch        “Zenith” Batch\n               Method of Moments Estimate                                42 * 0.588            77 * 0.588\n                                                                            = 24.7               = 45.3\n               Method of Maximum Likelihood Estimate                    42 * 2 * 0.412         77 * 0.412\n                                                                            = 34.6               =31.7\n               Observed Values                                                33                   37                      [1]\n\n              From the table it appears that MLE gives a better fit as predicted values are close to expected\n              values.                                                                                                    [1]",
      "has_math": true,
      "session": "2023-05",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2023-05_QP.pdf",
      "source_sol": "raw/CS1A_2023-05_SOL.pdf"
    },
    {
      "q_num": 5,
      "marks": 10,
      "topic": "inference",
      "subtopics": [],
      "stem": "A telecommunications company operating more than 1,000 circles in the country is planning\n         to perform Kaizen costing exercise to reduce its operational costs. Initially as a test run, the\n         exercise is being done over 10 circles in the country. The operational costs data for these 10\n         circles under the traditional costing approach and Kaizen costing approach has been\n         presented below:\n                                                                𝑛                    𝑛\n\n             Costing Approach         Sample Size n            ∑ 𝑥𝑖                  ∑ 𝑥𝑖2\n                                                               𝑖=1                   𝑖=1\n                 Traditional                 10                514.80            27804.64\n                   Kaizen                    10                401.40            18215.88\n\n         A statistical test is to be performed at the 5% level of significance, to determine whether\n         Kaizen actually leads to cost reduction i.e. for:\n\n         H0: There is no cost reduction versus H1: There is cost reduction",
      "parts": [
        {
          "label": "i",
          "marks": 1,
          "text": "Identify a suitable distribution for the test statistic.\n\n               A. t distribution\n               B. F distribution\n               C. Chi-square distribution\n               D. Normal distribution                                                                               (1)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 1,
          "text": "Show that the 95% two-sided confidence interval for difference in mean operating\n              costs is (-1.586, 24.266). What would you conclude in context of the null hypothesis?                 (4)\n\n         iii) Which one of the following is required to be assumed to perform the above mentioned\n              test?\n\n              A. Normality of the sample and equal population variances\n              B. Normality of the sample and equal sample variances\n              C. Normality of the population and equal sample variances\n              D. Normality of the population and equal population variances                                 (1)\n\n        The 95% confidence interval as determined in part (ii) was presented to the CEO of the\n        company and it was pitched to implement Kaizen costing over all the remaining circles in\n        the country. The CEO was not convinced with the width of the confidence interval and\n        suggested that the width should be reduced to below 20 in order to go ahead with full-blown\n        implementation.\n\n        iv)    How many more circles shall be at least covered under the test run in order to reduce\n               the width of the 95% confidence interval as determined in part (ii) as required by the\n               CEO?                                                                                         (3)\n\n        v)     At this reduced width, does your conclusion in part (ii) still hold true? Assume that the\n               difference between the sample means remains the same.                                        (1)",
          "topic": null
        }
      ],
      "solution": "i)        Correct Answer is Option A\n\n              Since the sample size is small and since the population variance is not known, t distribution\n              would be suitable to perform this test.                                                                      [1]\n\n      ii)     X̅A =51.48,    X̅B =40.14,\n              SA2 =1/9*(27804.64-10*51.48^2) = 144.7484\n              SB2 =1/9*(18215.88-10*40.14^2) = 233.7427                                                                   [1.5]\n\n              Pooled Variance:\n              S2p = 1/18*(9*144.7484+9*233.7427)\n              =189.2456                                                                                                    [1]\n                                                                                         1      1\n              95 % confidence interval is: (X̅A - X̅B) ± tnA+nB-2 * S2p* √𝑛𝐴 + 𝑛𝐵\n                                                                                                    2\n                                                     = (51.48-40.14)± 2.101 √189.2456 √10\n                                                     = (-1.58569, 24.26569)                                                [1]\n\n              Since 0 lies in the above confidence interval, we can conclude that at 5% level of significance,\n              there is insufficient evidence to reject the null hypothesis.                                               [0.5]\n\n                                                                                                               Page 5 of 11\n\fIAI                                                                                                    CS1A-0523\n\n  iii)        Correct Answer is Option D\n\n              Normality of the population data and equal population variances are the assumptions used for\n              conducting a t-test.                                                                                  [1]\n\n  iv)         Sample size:\n\n              The width of the confidence interval is\n                                 2     144.7484(𝑛−1)+233.7427(𝑛−1)\n              2 * t2.5%,2n-2√𝑛 √                  2𝑛−2\n\n                             38.9097\n              = t2.5%,2n-2\n                               √𝑛                                                                                   [1]\n              This should be less than 20, so using percentage points of the t distribution,\n              We have:\n              n = 15 => t2.5%,2n-2 =2.048\n                     => 38.9097 * 2.048/√15              = 20.57511 > 20                                            [1]\n              And n = 16 => t2.5%,2n-2 =2.042\n\n                                => 38.9097*2.042/√16 = 19.8634        < 20\n              Hence the minimum sample size required is 16. Hence, at least 6 additional circles must be\n              subject to a test run before going ahead with full blown implementation.                              [1]\n\n      v)      At this new width, confidence interval would be:\n              (11.34 – 19.8634/2, 11.34 + 19.8634/2)\n              = (1.408, 21.272)                                                                                    [0.5]\n              Since this confidence interval does not contain 0, we can conclude that there is sufficient\n              evidence to reject the null hypothesis.                                                            [0.5]",
      "has_math": true,
      "session": "2023-05",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2023-05_QP.pdf",
      "source_sol": "raw/CS1A_2023-05_SOL.pdf"
    },
    {
      "q_num": 6,
      "marks": 15,
      "topic": "distributions",
      "subtopics": [],
      "stem": "",
      "parts": [
        {
          "label": "i",
          "marks": 2,
          "text": "Generally, Kendal’s Tau tends to be lower than Spearman’s Rho. Consider n values\n               for two rank variables Rx and Ry which have the following pairs:\n\n               (1, n), (2, n – 1), (3, n – 2) ……………………………….. (n –2, 3), (n – 1, 2), (n, 1).\n\n               For the given values of the two rank variables, which of the following would be\n               TRUE?\n\n              A. Both Kendal’s Tau and Spearman’s Rho would be equal to +1\n              B. Both Kendal’s Tau and Spearman’s Rho would be equal to –1\n              C. Both Kendal’s Tau and Spearman’s Rho would be positive but Kendal’s Tau would\n                 be lower than Spearman’s Rho\n              D. Both Kendal’s Tau and Spearman’s Rho would be negative but Kendal’s Tau would\n                 be lower than Spearman’s Rho                                                               (2)\n\n        Inflation rates across ten different time periods for four developing economies were\n        analysed and the sample Kendall’s rank correlation coefficient between them was calculated\n        as given in the below table:\n\n                        Sample Kendall’s Rank Correlation Coefficient Matrix\n                            Zubrowka        Freedonia        Genovia                      Aldovia\n              Zubrowka         1.00            0.29             0.20                       0.42\n              Freedonia        0.29            1.00             0.11                       0.24\n               Genovia         0.20            0.11             1.00                       0.16\n               Aldovia         0.42            0.24             0.16                       1.00",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 4,
          "text": "Using normal approximation, test whether sample Kendall’s Rank correlation\n               coefficient supports the hypothesis that the inflation rates for Freedonia and Genovia\n               are positively correlated. Clearly state the null and alternate hypothesis, value of the\n               statistic and conclusion at 0.05% level of significance.\n\n             Hint: Variance = 2 (2n + 5) / 9n (n – 1)                                                (4)\n\n       While calculating Kendall’s Tau showing correlation between inflation rates for Freedonia\n       and Genovia, the concordant and discordant pairs were calculated as follows:\n\n                                                          Number of            Number of\n             Rank Freedonia        Rank Genovia\n                                                        Concordant Pairs    Discordant Pairs\n                    1                     ?                    4                    5\n                    2                     ?                    6                    2\n                    3                     ?                    7                    0\n                    4                     ?                    5                    1\n                    5                     ?                    2                    3\n                    6                     ?                    0                    4\n                    7                     ?                    0                    3\n                    8                     ?                    0                    2\n                    9                     ?                    1                    0\n                   10                     ?                    -                    -\n                                                              25                   20",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 5,
          "text": "Determine the values of “?” in the above table and hence calculate the sample rank\n              correlation coefficient between inflation rates for Freedonia and Genovia using\n              Spearman’s method.                                                                     (5)\n\n       Principal Component Analysis was carried out to reduce the dimensionality of the inflation\n       rates data-set in R using prcomp function using scale = TRUE and following extracts of\n       the R output and scree-plot have been obtained:\n\n                     Standard deviations (1, .., p=4):\n                     [1] 1.9583973 0.3047250 0.2437022 0.1114980\n\n                     Rotation (n x k) = (4 x 4):\n                                      PC1         PC2       PC3         PC4\n                     Zubrowka -0.4993881 0.01113661 -0.8530755 -0.15083017\n                     Freedonia -0.5015187 -0.48300910 0.3934253 -0.60033137\n                     Genovia   -0.4932231 0.81303126 0.3066223 -0.04115716\n                     Aldovia   -0.5057880 -0.32489745 0.1531718 0.78432047\n\n                     Importance of components:\n                                               PC1     PC2     PC3     PC4\n                     Standard deviation     1.9584 0.30472 0.24370 0.11150\n                     Proportion of Variance 0.9588 0.02321 0.01485 0.00311\n                    Cumulative Proportion 0.9588 0.98204 0.99689 1.00000",
          "topic": null
        },
        {
          "label": "iv",
          "marks": 3,
          "text": "Determine the principal component(s) to be retained using the following criteria:\n\n               a) Retaining those components which represent at least 90% of the total variance;\n               b) Scree Test;\n               c) Kaiser Test.                                                                               (3)",
          "topic": null
        },
        {
          "label": "v",
          "marks": 1,
          "text": "The correlation coefficient between the PC1 and PC2 as determined above is -\n\n               A. – 1\n               B. + 1\n               C. 0\n               D. None of the above.                                                                         (1)",
          "topic": null
        }
      ],
      "solution": "i)        Correct Answer is Option B\n              Since the data is perfectly monotonically decreasing, both the coefficients would be equal to -1\n              (perfect negative correlation).\n\n      ii)     H0: ρ = 0\n              H1: ρ > 0                                                                                            [0.5]\n\n              Under H0, the sampling distribution of Kendall’s rank correlation coefficient is approximately\n              normal with mean 0 and variance = 2(2n+5) / 9n(n-1)                                                  [0.5]\n\n              Variance = 2(20+5)/90(9) = 50/810 = 0.061728                                                         [0.5]\n              Observed value of the test statistic\n              = (rk – mean) / √𝑣𝑎𝑟𝑖𝑎𝑛𝑐𝑒\n              = (0.11 – 0) / √0.061728\n              = 0.4427                                                                                              [1]\n              This does not exceed the upper 0.05% point of the standard normal distribution (3.2905).             [0.5]\n              So we do not have sufficient evidence to reject the null hypothesis.                                 [0.5]\n                                                                                                         Page 6 of 11\n\fIAI                                                                                                       CS1A-0523\n\n              Hence, we can conclude that the inflation rates for Freedonia and Genovia are not positively\n              correlated.\n\n   iii)       The completed table with the rank values for Genovia is given below:\n\n                  Rank Freedonia           Rank Genovia              Number of                 Number of\n                                                                   Concordant Pairs         Discordant Pairs\n                         1                         6                      4                        5\n                         2                         3                      6                        2\n                         3                         1                      7                        0\n                         4                         4                      5                        1\n                         5                         8                      2                        3\n                         6                        10                      0                        4\n                         7                         9                      0                        3\n                         8                         7                      0                        2\n                         9                         2                      1                        0\n                         10                        5                      -                         -\n                                                                         25                        20\n\n              In case of Rank 1 (Freedonia), there are 4 concordant pairs and 5 discordant pairs. So, the rank\n              value here for Genovia is higher than 4 rank values (10, 9, 8, 7) but lower than 5 rank values (1,\n              2, 3, 4, 5). It must be 6.\n              In case of Rank 2 (Freedonia), there are 6 concordant pairs and 2 discordant pairs. So, the rank\n              value here for Genovia is higher than 6 rank values (10, 9, 8, 7, 5, 4) but lower than 2 rank values\n              (1, 2). It must be 3. Kindly note that 6 is already considered in the upper cell and it is not being\n              taken into consideration.\n              The above process can be continued till we get all rank values for Genovia.                             [2.5]\n\n                       6 ∑ 𝑑2\n              rs = 1 – 𝑛(𝑛2 −1)                                                                                       [0.5]\n\n              ∑ 𝑑 2 = -52 + -12 + 22 + 02 + -32 + -42 + -22 + 12 + 72 + 52 = 134                                       [1]\n\n              rs = 1 – 6*134 / 10(100 – 1) = 0.1879                                                                    [1]\n\n   iv)        If we want to retain those components which explain 90% of the total variance, PC1 should be\n              retained as it accounts for 95.88% of the total variance.                                                [1]\n\n              Based on the Scree Plot, the plot becomes flat from PC2 and onwards. Hence using Scree Test,\n              PC1 should be retained as variances level off after PC1.                                                 [1]\n\n              As per Kaiser’s Test only those PCs with variances greater than 1 should be retained (applicable\n              in case of scaled data). Since, only PC1 has variance greater than 1, only PC1 should be retained.       [1]\n\n      v)      Correct Answer is Option C\n              Principal components are un-correlated linear combinations of the variables of the original data\n              set. Hence correlation coefficient between PC1 and PC2 will be equal to 0.                            [1]",
      "has_math": false,
      "session": "2023-05",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2023-05_QP.pdf",
      "source_sol": "raw/CS1A_2023-05_SOL.pdf"
    },
    {
      "q_num": 7,
      "marks": 15,
      "topic": "bayes_credibility",
      "subtopics": [],
      "stem": "A Psephology Firm is conducting a survey to determine the proportion of people “p” who\n         will vote for XJP political party in a particular municipal ward for the local body elections.\n         The Chief Psephologist is an Actuary and he is reviewing the responses collected by his\n         team working on-field. A total of 200 responses have been collected. His prior beliefs about\n         “p” based on historical vote share of XJP are given by a uniform distribution on the interval\n         [0,1]. It turns out that after speaking to n1 respondents, the n1st respondent happens to be a\n         supporter of XJP. Others i.e. (n1 – 1) are non-supporters.",
      "parts": [
        {
          "label": "i",
          "marks": 1,
          "text": "If n is a random variable that represents the number of respondents to be covered till\n                one meets the first XJP supporter, then which of the following probability distributions\n                would be suitable to model n?\n\n               A. Hyper-geometric Distribution\n               B. Geometric Distribution\n               C. Bernoulli Distribution\n               D. Binomial Distribution.                                                                     (1)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 3,
          "text": "Specify the posterior distribution of “p” after the actuary finds the first supporter of\n                XJP political party.                                                                         (3)",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 4,
          "text": "State the prior mean of “p” and calculate the posterior probability that “p” is greater\n                than the prior mean using the posterior distribution determined in part (ii) and explain\n                what it signifies. You may assume that n1 = 3.                                               (4)\n\n         The second supporter is found after covering another n2 respondents, third supporter after\n         covering further n3 respondents and so on until the 50th supporter is found who happens to\n         be the 200th (last) respondent. In other words, n1 + n2 + n3 + ……….. + n50 = 200.",
          "topic": null
        },
        {
          "label": "iv",
          "marks": 3,
          "text": "Show that after conducting the above survey, the posterior distribution of “p” is a Beta\n                Distribution with parameters α = 51 and β = 151.                                             (3)",
          "topic": null
        },
        {
          "label": "v",
          "marks": 3,
          "text": "Determine the Bayesian estimate of “p” under “Squared Error Loss” and express it in\n                the form of a credibility estimate. Also determine the value of the credibility factor Z.    (3)",
          "topic": null
        },
        {
          "label": "vi",
          "marks": 1,
          "text": "Based on the value of the credibility factor calculated in part (v), which of the\n                following statements would be a reasonable inference for the survey?\n\n             A.    XJP is slated to face defeat in the local body elections in the ward under\n                   consideration.\n             B.    Under no circumstances, XJP would be able to achieve its historical vote share\n                   in the municipal ward in the coming local body elections.\n             C.    There is a need to increase the sample size of the survey in order to take a\n                   credible view on the vote share of XJP in the upcoming elections in the ward.\n             D.    It is virtually certain that the vote share of XJP will be reduced to almost a half\n                   of its historical vote share based on the survey results.                              (1)",
          "topic": null
        }
      ],
      "solution": "i)        Correct Answer is Option B\n\n                                                                                                            Page 7 of 11\n\fIAI                                                                                                    CS1A-0523\n\n            Since we are modelling N which represents the number of trials to be performed until the first\n            success occurs, the appropriate distribution would be geometric distribution.                            [1]\n\n      ii)   The prior distribution of “p” is uniform over the interval [0,1]\n            So f prior (p) =1         0≤p≤1                                                                        [0.5]\n\n            Sample contains only one observation n1. So the likelihood function of “p” is:\n            L(p) = P(N = n1) = (1 – p)(n1-1) * p\n            The above expansion is based on the fact that N | p ~ Geometric(p)                                       [1]\n\n            Combining the prior PDF and the likelihood function, we get,\n            f posterior (p) ∝ f prior (p) * L(p)\n            f posterior (p) ∝ (1 – p)(n1-1) * p\n            f posterior (p) ∝ (1 – p)(n1-1) * p(2-1)     for 0 ≤ p ≤ 1                                               [1]\n\n            So the posterior distribution of p is Beta (2, n1)                                                     [0.5]\n\n  iii)      Prior mean of p under U[0,1] is given by:\n            E(p) = (0+1)/2 = 0.50                                                                                  [0.5]\n            We have to calculate probability that P exceeds 0.50 using the posterior distribution of P.\n\n            P(P > 0.50)\n               1\n            = ∫0.50(1 − 𝑝)(𝑛1−1) ∗ p dp\n               1\n            =∫0.50(1 − 𝑝)(3−1) ∗ p dp\n               1\n            =∫0.50(1 − 𝑝)2 ∗ p dp\n                1\n            = ∫0.50(1 − 2p + 𝑝2 ) ∗ p dp\n               1\n            =∫0.50(p − 2𝑝2 + 𝑝3 ) dp\n            = (12 – 0.502) / 2 – 2/3 * (13 – 0.503) + ¼ * (14 – 0.504)\n            = 0.026042                                                                                             [2.5]\n            This signifies that there is only 2.6% chance that the vote share of XJP will cross 50% (which is\n            the prior expectation based on historical vote shares).                                                  [1]\n\n  iv)       Likelihood function now is given by:\n            L(p)\n            = P(N1 = n1) * P(N2 = n2) * P(N3 = n3) * ………………… * P(N50 = n50)\n            = (1 – p)(n1-1) * p * (1 – p)(n2-1) * p *(1 – p)(n3-1) * p * ………………….. * (1 – p)(n50-1) *\n            p\n            = (1 – p) (n1+n2+……+n50 – 50) * p 50\n            = (1 – p) (200 – 50) * p 50\n            = (1 – p)150 * p 50                                                                                    [1.5]\n\n            Hence, posterior distribution of p is given by:\n            f posterior (p) ∝ (1 – p)150 * p50\n            f posterior (p) ∝ (1 – p)(151-1) * p(51-1)    for 0 ≤ p ≤ 1                                              [1]\n\n            So the posterior distribution of p is Beta (51, 151)                                                   [0.5]\n\n                                                                                                          Page 8 of 11\n\fIAI                                                                                                       CS1A-0523\n\n      v)      Bayesian estimate of “p” under squared error loss is the mean of the posterior distribution\n              which is given by:\n\n              E(P | n) = α / (α + β) = 51 / (51 + 151) = 0.2525                                                         [1]\n\n              E(P | n) = Z * sample mean + (1 – Z) * prior mean\n              0.2525 = Z * (50/200) + (1 – Z) * 0.50\n              0.2525 = 0.25Z – 0.50 Z + 0.50\n              0.25Z = 0.2475\n              Z = 0.99\n              So, if we want to estimate the posterior mean as a credibility estimate,\n              E(P | n) = 0.99 * sample mean + 0.01 * prior mean                                                         [2]\n\n   vi)        Correct Answer is Option C\n              200 is a very small sample size to draw any concrete conclusion. Increasing the sample size\n              would be the ideal way forward.",
      "has_math": true,
      "session": "2023-05",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2023-05_QP.pdf",
      "source_sol": "raw/CS1A_2023-05_SOL.pdf"
    },
    {
      "q_num": 8,
      "marks": 20,
      "topic": "regression_glm",
      "subtopics": [],
      "stem": "You are working as the Chief Statistical Officer in the Ministry of Health Care, Government\n        of Actuaria. Your team has collected the following data relating to average heart rate (Y)\n        and average systolic blood pressure (X1) and average diastolic blood pressure (X2) for 10\n        major cities in Actuaria.\n\n                                  Average Heart          Average Systolic       Average Diastolic\n                                       Rate              Blood Pressure          Blood Pressure\n                  City\n                                 beats per minute            mm Hg                  mm Hg\n                                        (Y)                   (X1)                    (X2)\n             Orbit City                  94                   139                      84\n            Emerald City                 77                   125                      90\n             Shangri-La                  81                   126                      81\n           Tomorrow Land                 76                   129                      87\n             Kingsbury                   63                   112                      94\n              Rivendell                  76                   118                      78\n               Atlantis                  91                   153                      92\n             Cloud City                  73                   104                      70\n             Dark City                   88                   124                      72\n           Thugs Mansion                 84                   134                      79\n\n        The following multiple linear regression model was used to analyse the above data where\n        Y was the response variable and X1 and X2 were the explanatory variables:\n\n                                          y = α + β1x1 + β2x2 + e\n\n        The model was fitted in R and extracts from the R output for this model are given below:\n\n                    Call:\n                    lm(formula = Y ~ X1 + X2)\n\n                    Residuals:\n                        Min      1Q Median              3Q       Max\n                    -4.0127 -1.5872 -0.1759         1.7446    5.8372\n\n                    Coefficients:\n                                Estimate Std. Error t value Pr(>|t|)\n                    (Intercept) 47.56101   12.75027   3.730 0.007357 **\n                    X1            0.69237   0.08818   7.852 0.000103 ***\n                    X2          -0.66235    0.14896 -4.446 0.002985 **\n                    ---\n                    Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1\n\n                    Residual standard error: 3.322 on 7 degrees of freedom\n                    Multiple R-squared: 0.9005, Adjusted R-squared: ????\n                    F-statistic: 31.66 on 2 and 7 DF, p-value: 0.0003113",
      "parts": [
        {
          "label": "i",
          "marks": 1,
          "text": "Select the equation for the fitted multiple linear regression model using the extracts\n              given above from the following four options:\n\n             A. y = -4.0127 -1.5872 x1 -0.1759 x2 + e\n             B. y = 47.56101 + 0.69237 x1 – 0.66235 x2 + e\n             C. y = 12.75027 + 0.08818 x1 + 0.14896 x2 + e\n             D. y = 3.7300 + 7.8520 x1 – 4.4460 x2 + e                                                    (1)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 2,
          "text": "Calculate Adjusted R2 for the model.                                                        (2)",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 2,
          "text": "Calculate the predicted average heart rate for Emerald City and Dark City. Also\n              determine the value of residuals.                                                           (2)\n\n       Some of the eminent officials from the ministry who are medical doctors by profession are\n       of the opinion that instead of considering the levels of systolic and diastolic blood pressure,\n       “pulse pressure” i.e. the difference between systolic and diastolic blood pressures should be\n       considered as the explanatory variable for predicting the value of the average heart rate.\n\n       Let us define Average Pulse Pressure as Z where Z = X1 – X2.\n\n                                        𝑦̅ = 80.30⁡⁡⁡⁡⁡\n\n                                        𝑧̅ = 43.70\n\n                                        ∑(y − 𝑦̅)2 = 776.10⁡;\n\n                                        ∑(z − 𝑧̅)2 = 1472.10⁡;\n\n                                        ∑(y − 𝑦̅)(z − 𝑧̅) = 1013.90",
          "topic": null
        },
        {
          "label": "iv",
          "marks": 5,
          "text": "Using a bivariate linear regression model as given below, calculate the least square\n              estimates of λ and µ and re-calculate the predicted average heart rate for Emerald City\n              and Dark City clearly stating the value of residuals.\n                              y = λ + µz + e                                                              (5)",
          "topic": null
        },
        {
          "label": "v",
          "marks": 2,
          "text": "Calculate Adjusted R2 for the bivariate regression model.                                   (2)",
          "topic": null
        },
        {
          "label": "vi",
          "marks": 1,
          "text": "What is the expected reduction in pulse pressure which will ensure a decrease in the\n              heart rate by 2? Answer with reference to the bivariate regression model fitted in part\n              (iv).\n\n             A. 3.51\n             B. 0.52\n             C. 1.38\n             D. 2.90                                                                                      (1)\n\n       Your twenty-one year old son who is a freshly qualified statistics graduate has suggested\n       that instead of using Z as the explanatory variable, using deviations from its mean defined\n       as W = (Z – Z̅) would provide a better fit.\n\n      The improvised bivariate linear regression model using W as the explanatory variable is\n      given below:\n                                           y = δ + £w + e",
          "topic": null
        },
        {
          "label": "vii",
          "marks": 4,
          "text": "Show that the least square estimators of the parameters of the improvised bivariate\n           model are given by:\n\n           a)    £̂ = µ̂\n           b)    𝛿̂ = λ̂ + µ̂ z̅\n\n      where λ̂ and µ̂ are the parameter estimators of the model stated in part (iv).\n\n      Hint: Kindly note that w\n                             ̅ = 0.                                                                    (4)\n\n      Using the improvised bivariate linear regression model in part (vii), predicted average heart\n      rate for Emerald City and Dark City and associated residuals were recalculated as follows:\n\n             City                       ̂\n                                        𝒚                       e\n          Emerald City               74.3079                 2.6921\n           Dark City                 86.0166                 1.9834\n\n      Adjusted R2 for the improvised bivariate model is 88.72%.",
          "topic": null
        },
        {
          "label": "viii",
          "marks": 3,
          "text": "Compare the three models based the predictions for Emerald City and Dark City and\n            the values of Adjusted R2 and briefly state your observations.                             (3)",
          "topic": null
        }
      ],
      "solution": "i)        Correct Answer is Option B\n              Value of estimates for the intercept, X1 and X2 will be the values of α, β1 and β2 respectively.          [1]\n\n      ii)     Adjusted R2 = 1 – (n – 1) / (n – k – 1) * (1 – R2)                                                        [1]\n              We have n = 10 and k = 2 predictors and R2 = 0.9005                                                     [0.5]\n              Adjusted R2\n              = 1 – (10 – 1) / (10 – 2 – 1) * (1 – 0.9005)\n              = 87.21%                                                                                                [0.5]\n   iii)       𝑦̂ 𝐸𝑚𝑒𝑟𝑎𝑙𝑑 𝐶𝑖𝑡𝑦\n              = 47.5601 + 0.6924 (125) – 0.6624 (90)\n              = 74.4941                                                                                               [0.5]\n              𝑒̂ 𝐸𝑚𝑒𝑟𝑎𝑙𝑑 𝐶𝑖𝑡𝑦\n              = 77 – 74.4941\n              = 2.5059                                                                                                [0.5]\n              𝑦̂ 𝐷𝑎𝑟𝑘 𝐶𝑖𝑡𝑦\n              = 47.5601 + 0.6924 (124) – 0.6624 (72)\n              = 85.7249                                                                                               [0.5]\n              𝑒̂ 𝐷𝑎𝑟𝑘 𝐶𝑖𝑡𝑦\n              = 88 – 85.7249\n              = 2.2751                                                                                                [0.5]\n\n   iv)        Syy = 776.10\n              Szz = 1472.10\n              Syz = 1013.90                                                                                             [1]\n              µ̂\n              = Syz / Szz .                                                                                             [1]\n\n                                                                                                            Page 9 of 11\n\fIAI                                                                                               CS1A-0523\n\n           = 1013.90 / 1472.10\n           = 0.6887\n           ʎ̂\n           = 𝑦̅ – µ̂ * 𝑧̅\n           = 80.30 – 0.6887 * 43.70\n           = 50.2019                                                                                           [1]\n\n           𝑦̂ 𝐸𝑚𝑒𝑟𝑎𝑙𝑑 𝐶𝑖𝑡𝑦\n           = 50.2019 + 0.6887 (125-90)\n           = 74.3064                                                                                          [0.5]\n           𝑒̂ 𝐸𝑚𝑒𝑟𝑎𝑙𝑑 𝐶𝑖𝑡𝑦\n           = 77 – 74.3064\n           = 2.6936                                                                                           [0.5]\n           𝑦̂ 𝐷𝑎𝑟𝑘 𝐶𝑖𝑡𝑦\n           = 50.2019 + 0.6887 (124-72)\n           = 86.0143                                                                                          [0.5]\n           𝑒̂ 𝐷𝑎𝑟𝑘 𝐶𝑖𝑡𝑦\n           = 88 – 86.0143\n           = 1.9857                                                                                           [0.5]\n      v)   R2\n           = Sxz2 / (Sxx * Szz)\n           = (1013.90)2 / (776.10 * 1472.10)\n           = 89.9778%                                                                                          [1]\n\n           Adjusted R2\n           = 1 – (10 – 1) / (10 – 1 – 1) * (1 – 0.899778)\n           = 88.725%                                                                                           [1]\n  vi)      Correct Answer is Option D\n           2 = (0.6887) * zreduction\n           zreduction = 2/0.6887 = 2.90                                                                        [1]\n\n  vii)                          ̅ = 0.\n           a) We are given that W\n                       ̅)2 = ∑(w − 0)2 = ∑(z − 𝑧̅)2 = Szz\n           Sww = ∑(w − 𝑤                                                                                       [1]\n\n           Syw = ∑(y − 𝑦̅)(w − 𝑤\n                               ̅) = ∑(y − 𝑦̅)(w − 0) ∑(y − 𝑦̅)(z − 𝑧̅) = Syz                                   [1]\n\n           £̂ = Syw / Sww = Syz / Szz = µ̂                                                                     [1]\n\n           b) 𝛿̂ = 𝑦̅ – µ̂ * 𝑤\n                             ̅ = 𝑦̅ – µ̂ * 0 = 𝑦̅ = ʎ̂ + µ̂ * 𝑧̅                                               [1]\n\n viii)                                                                                   Improvised\n                                              Multiple Linear      Bivariate Linear\n                            City                                                      Bivariate Linear\n                                             Regression Model          Model\n                                                                                           Model\n                Emerald City (𝑦̂, 𝑒)           (74.49, 2.51)        (74.31, 2.69)       (74.31, 2.69)\n                 Dark City (𝑦̂, 𝑒)             (85.72, 2.28)         86.02, 1.99)       (86.02, 1.98)\n                  Adjusted R2                    87.21%                88.73%              88.72%\n\n                                                                                                   Page 10 of 11\n\fIAI                                                                                           CS1A-0523\n\n      In terms of the predicted responses and residuals, for Emerald City, the multiple linear\n      regression model appears to be a better fit. However for Dark City, the bivariate linear model\n      gives better results as compared to the multiple linear regression model.\n\n      However in terms of Adjusted R2 (which measures the variation of the predicted responses to\n      actual responses), Bivariate Linear Model appears to be a better fit.                               [1]\n\n      Improvised Bivariate Linear Model just employs a linear combination of the explanatory\n      variable of the Original Bivariate Linear Model and hence gives almost similar results like the\n      original model. Unlike the presumption made by your son, it is clear that the improvised model\n      does not provide a better fit as compared to the original model.                                    [1]\n\n                                       *****************\n\n                                                                                              Page 11 of 11",
      "has_math": true,
      "session": "2023-05",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2023-05_QP.pdf",
      "source_sol": "raw/CS1A_2023-05_SOL.pdf"
    },
    {
      "q_num": 1,
      "marks": 6,
      "topic": "distributions",
      "subtopics": [],
      "stem": "Let Mx(t) denote the moment generating function (MGF) and Cx(t) denote the cumulant\n        generating function (CGF) of a random variable X.",
      "parts": [
        {
          "label": "i",
          "marks": 1,
          "text": "Which of the following is TRUE for the relationship between Mx(t) and Cx(t)?\n\n               A. Cx(t) = e Mx(t)\n               B. Mx(t) = e Cx(t)\n               C. Cx(t) = log10 (Mx(t))\n               D. Mx(t) = log10 (Cx(t))                                                                   (1)\n\n        The series expansion formula of moment generating function (MGF) for a random variable\n        X is as follows:\n\n                                               t2      t3      t4\n                           MX (t) = 1 + tE(X) + E(X ) + E(X ) + E(X 4 ) + ⋯\n                                                   2       3\n                                               2!      3!      4!",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 3,
          "text": "Based on the above series expansion, show that the variance of random variable X is\n                 given by the following expression:\n\n                          var(X) = M ′′ X (0) − [M ′ X (0)]2                                              (3)",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 2,
          "text": "In terms of Cx(t), how will you represent variance of random variable X? Choose the\n                 correct option from those given below:\n\n                 A. var(X) = C′′ X (0) − [C′ X (0)]2\n                 B. var(X) = [C′ X (0)]2\n                 C. var(X) = C′′ X (0)\n                 D. var(X) = [C′ ′X (0)]2                                                                 (2)",
          "topic": null
        }
      ],
      "solution": "i)       Correct Answer is Option B\n\n                   We know that Cx(t) = ln (Mx(t)). Hence, Mx(t) = e Cx(t).                                                 [1]\n\n         ii)                            𝑡           𝑡          𝑡\n                   𝑀 (𝑡) = 1 + 𝑡𝐸(𝑋) +     𝐸(𝑋 ) +     𝐸(𝑋 ) + 𝐸(𝑋 ) + ⋯\n                                        2!          3!         4!\n                                   2𝑡            3𝑡           4𝑡\n                   𝑀 (𝑡) = 𝐸(𝑋) +       𝐸(𝑋 ) +       𝐸(𝑋 ) +     𝐸(𝑋 ) + ⋯\n                                    2!            3!           4!\n                                      𝑡           𝑡\n                   = 𝐸(𝑋) + 𝑡𝐸(𝑋 ) +      𝐸(𝑋 ) + 𝐸(𝑋 ) + ⋯\n                                      2!          3!\n                   𝑀 (0) = 𝐸(𝑋)\n\n                                                       𝑡\n                   𝑀     (𝑡) = 𝐸(𝑋 ) + 𝑡𝐸(𝑋 ) +           𝐸(𝑋 ) + ⋯\n                                                       2!\n                   𝑀 (0) = 𝐸(𝑋 )\n                   𝑣𝑎𝑟(𝑋) = 𝐸(𝑋 ) − [𝐸(𝑋)] = 𝑀                (0) − [𝑀 (0)]                                                 [3]\n\n         iii)      Correct Answer is Option C\n\n                   𝐶 (𝑡) = 𝑙𝑛 𝑀 (𝑡)\n                               1\n                   𝐶 (𝑡) =         𝑀 (𝑡)\n                             𝑀 (𝑡)\n                                   1                  1\n                   𝐶 (𝑡) = −            𝑀 (𝑡)𝑀 (𝑡) +       𝑀 (𝑡)\n                                [𝑀 (𝑡)]              𝑀 (𝑡)\n                     𝑀 (𝑡)       𝑀 (𝑡)\n                   =          −\n                      𝑀 (𝑡)       𝑀 (𝑡)\n                              𝑀 (0)      𝑀 (0)\n                   𝐶 (0) =            −        = 𝑀 (0) − [𝑀 (0)] = 𝐸(𝑋 ) − [𝐸(𝑋)]\n                              𝑀 (0)      𝑀 (0)\n                   𝐶 (0) = 𝑣𝑎𝑟(𝑋)                                                                                          [2]",
      "has_math": true,
      "session": "2023-11",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2023-11_QP.pdf",
      "source_sol": "raw/CS1A_2023-11_SOL.pdf"
    },
    {
      "q_num": 2,
      "marks": 8,
      "topic": "inference",
      "subtopics": [],
      "stem": "The Inspector General of Police wants to evaluate the effectiveness of the lie-detector test\n        carried out using a polygraph on those who are charged with the offence of first degree\n        murder at all police station lock-ups in the country.\n\n        The police authorities have prepared the following matrix for consideration of the\n        Inspector General based on the past records relating to 1,000 cases which have been\n        subsequently decided by the judicial authorities:\n\n                    Particulars             Persons found Innocent as    Persons found Guilty as\n                                            per the Lie-Detector Test   per the Lie-Detector Test\n\n                Persons acquitted by\n                                                       “Area A”                  “Area B”\n               Judicial Authorities as\n                      Innocent\n                                                        (✓– )                     ( +)\n               Persons found Guilty of\n                                                       “Area C”                  “Area D”\n                first degree murder by\n                 Judicial Authorities\n                                                        ( –)                     (✓ +)",
      "parts": [
        {
          "label": "i",
          "marks": 2,
          "text": "Which of the following is TRUE with reference to Type I Error (False Positives) and\n               Type II Error (False Negatives) based on the given matrix?\n\n               A. Type I Error = Area A and Type II Error = Area D\n               B. Type I Error = Area C and Type II Error = Area B\n               C. Type I Error = Area B and Type II Error = Area C\n               D. Type I Error = Area D and Type II Error = Area A\n\n        Following events have been defined in this context:\n\n          G     : A person charged with the offence is actually guilty of the offence\n          I     : A person charged with the offence is innocent\n          LG : A person charged with the offence is found to be guilty as per the lie-detector test\n          LI    : A person charged with the offence is found innocent as per the lie-detector test",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 2,
          "text": "Using the events as defined above, you are required to show the following:\n\n               a) Probability of committing Type I Error = P(I | LG);\n               b) Probability of committing Type II Error = P(G | LI).                                      (2)\n\n        Based on the historical records for 1,000 cases as collected by the police authorities,\n        following data has been obtained:\n\n                    A = 356; B = 111; C = 105; D = 428",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 1,
          "text": "Based on the above data, which of following is the probability that the Lie-Detector\n               Test correctly identifies the perpetration / non-perpetration of the crime?\n\n               A. 0.784\n               B. 0.461\n               C. 0.533\n               D. 0.216                                                                                     (1)",
          "topic": null
        },
        {
          "label": "iv",
          "marks": 3,
          "text": "Based on the historical data as given above, calculate the probabilities of Type I\n               Error and Type II Error as shown in part (ii).                                               (3)",
          "topic": null
        }
      ],
      "solution": "i)           Correct Answer is Option C\n\n                Type I error represents False Positives i.e. where actually innocent individuals (as decided by the\n                court later) are found to be guilty based on the lie-detector test. This is thus represented by Area B\n                in the matrix.\n\n                Type II error represents False Negatives i.e. where a person who is actually guilty of the crime (as\n                decided by the court later) is considered to be innocent based on the lie-detector test. This is thus\n                represented by Area C in the matrix.\n\n                Kindly note that here Positive means “being found guilty” and Negative means “being innocent”.              [2]\n\n   ii)          a) Probability of Type I Error\n                   = Probability (False Positive)\n                   = Probability (Person is Innocent but has been identified as guilty by the lie-detector)\n                   = P(I | LG)\n\n                b) Probability of Type II Error\n\n                                                                                                                 Page 2 of 11\n\f IAI                                                                                                         CS1A-1123\n\n              = Probability (False Negative)\n              = Probability (Person is Guilty but has been identified as innocent by the lie-detector)\n              = P(G | LI)\n\n  iii)    Correct Answer is Option A\n\n          Probability that the lie-detector correctly identifies the perpetration / non-perpetration of the crime\n          is given by (A+D) / (A + B + C + D)\n          = (356 + 428) / 1000\n          = 0.784                                                                                                      [1]\n\n  iv)     Probability of Type I Error\n          = P(I | LG)\n          = B / (B+D)\n          = 111/(111+428)\n          = 0.206\n\n          Probability of Type II Error\n          = P(G | LI)\n          = C / (A+C)\n          = 105 / (356+105)\n          = 0.228",
      "has_math": false,
      "session": "2023-11",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2023-11_QP.pdf",
      "source_sol": "raw/CS1A_2023-11_SOL.pdf"
    },
    {
      "q_num": 3,
      "marks": 10,
      "topic": "regression_glm",
      "subtopics": [],
      "stem": "",
      "parts": [
        {
          "label": "i",
          "marks": 4,
          "text": "Briefly explain the three components of a generalised linear model (GLM).                    (3)\n        An Actuary is using a GLM to model the remaining time until an actuarial aspirant\n        qualifies as an actuary. The covariates he uses are:\n           - Age\n           - Passes: List of papers already passed.\n                [This can be expressed as as a series of variables Pi, where i is a number between 1\n                and 13 (both inclusive), and each Pi is either 1 (the paper has been passed) or 0 (the\n                paper has not been passed).]\n\n               -    Experience\n               -    Duration: Time elapsed since clearing ACET\n\n        The model he fits is given below:\n\n                           Age + Passes + Experience + Duration + Experience . Duration\n\n        ii)        The number of parameters for the main effects (i.e., not interactions) under the above\n                   GLM is –\n\n                   A. 15\n                   B. 16\n                   C. 17\n                   D. 18                                                                                     (2)\n\n        iii)       The number of parameters relating to interaction terms in the above GLM is –\n\n                   A. 30\n                   B. 20\n                   C. 10\n                   D. 1                                                                                      (1)\n\n        The Actuary wants to simplify the model and check whether any of the covariates can be\n        removed. He has calculated scaled deviance of the model as 15. However, he has been\n        informed that the Akaike’s Information Criterion (AIC) could be a better metric to decide\n        on the model fit.\n\n        iv)        Calculate the AIC for the model based on the scaled deviance given above and the\n                   number of parameters determined in parts (ii) and (iii). You are given that the log-\n                   likelihood of the saturated model is 16.                                                  (4)",
          "topic": null
        }
      ],
      "solution": "i)     The three components of a GLM are:\n              (1) A distribution for the response variable – This belongs to the exponential family.\n              (2) A linear predictor η – This is a linear function of the covariates.\n              (3) A link function g – This connects the mean response to the linear predictor, g(μ) = η\n  ii)     Correct Answer is Option C\n\n          The number of parameters (standalone) for each factor (excluding passes) is:\n               Age – 2 (including intercept)\n               Experience – 2 (including intercept)\n               Duration – 2 (including intercept)\n\n          Exam passes: This is effectively 13 yes / no factors. Each of these 13 factors would contribute 2\n          parameters on standalones basis (including intercept).\n\n          The main effects are therefore going to contribute: 2 + 13*(2-1) + (2 – 1) + (2 – 1) = 17 parameters.        [2]\n\n  iii)    Correct Answer is Option D\n\n          The interactions will contribute (2-1) * (2-1) = 1 parameter.                                                [1]\n\n  iv)     Scaled Deviance = 2 (ln LS – ln LM)\n          15 = 2 (16 – ln LM)\n          ln LM = (32 – 15) / 2 = 8.5\n\n          AIC\n          = -2 * ln LM + 2 * number of parameters\n          = -2 *8.5 + 2 * (17+1)\n\n                                                                                                            Page 3 of 11\n\f IAI                                                                                                          CS1A-1123\n\n              = -17 + 2 * 18\n              = -17 + 36\n              = 19                                                                                                     [4]",
      "has_math": false,
      "session": "2023-11",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2023-11_QP.pdf",
      "source_sol": "raw/CS1A_2023-11_SOL.pdf"
    },
    {
      "q_num": 4,
      "marks": 12,
      "topic": "inference",
      "subtopics": [
        "distributions"
      ],
      "stem": "Space scientists from Planet Actuaria have been making attempts to understand the extent\n        of gravitational pull on various regions of a nearby Planet Numerica. For this purpose, an\n        object weighing 10 pounds on Actuaria has been sent through space vehicles on various\n        regions of Planet Numerica and the weight of this object at each of these regions is being\n        measured.\n        Let W be the weight (in pounds) of the object on Planet Numerica.\n        Let X be the ratio of the weight of the object on Planet Numerica to the weight of the\n        object on Planet Actuaria. In other words, X = W / 10.\n        X is assumed to follow a continuous uniform distribution within the interval [0, θ].\n        An attempt is being made to arrive at the estimate of θ (given θ > 0) in order to assess the\n        maximum gravitational pull on Planet Numerica.\n\n        A sample of 10 measurements for X has been obtained as given below:\n\n                               0.70, 0.55, 0.31, 0.40, 0.35, 0.62, 0.34, 0.77, 0.45, 0.64",
      "parts": [
        {
          "label": "i",
          "marks": 3,
          "text": "If θ̂MOM is the method of moments estimator of θ, obtain an estimate for θ̂MOM based\n             on the given sample.\n             ∑ x = 5.13\n             ∑ x 2 = 2.8761                                                                           (3)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 1,
          "text": "Which of the following is a correct expression for the likelihood function of θ\n             assuming a sample x = x1, x2, ……., x10 given 0 ≤ x1, x2, ……., x10 ≤ θ?\n                            1\n             A. L(θ) = θ10 for all xi in x\n                            1\n             B. L(θ) = θ10 if θ > max (xi) ;\n                     = 0 otherwise\n                            1\n             C. L(θ) = θ10 if θ > min (xi) ;\n                     = 0 otherwise\n                            1\n             D. L(θ) = θ10 if θ ≤ max (xi) ;\n                     = 0 otherwise                                                                    (1)",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 3,
          "text": "Obtain an equation to be solved for finding θ̂MLE i.e. the maximum likelihood\n             estimator of θ using the likelihood function chosen in part (ii). Also comment on why\n             method of differentiation doesn’t work in this case.                                     (3)\n\n      Assuming value of θ from 0.01 to 3.00, a graph was constructed plotting the value of θ on\n      x axis and the value of L(θ) on y-axis:\n\n                                              Graph of Likelihood Function\n                          Value of L(θ)\n\n                                          max(xi)          Value of θ",
          "topic": null
        },
        {
          "label": "iv",
          "marks": 2,
          "text": "Based on the given graph, show that max (xi) is θ̂MLE i.e. the maximum likelihood\n             estimator of θ.                                                                          (2)\n\n      Let us define a random variable Z = max (Xi) where Xis given i = 1,2, ….., 10 are\n      independent and identically distributed continuous uniform variables over interval [0, θ].\n\n      The probability density function of Z is given by the following expression:\n\n                                                           10 z9\n                                                    fZ(Z) = θ10 . for 0 ≤ z ≤ θ",
          "topic": null
        },
        {
          "label": "v",
          "marks": 2,
          "text": "Using the probability density function of Z, show that the bias of the estimator θ̂MLE\n               is – 11-1 * θ.\n               Hint: Bias (𝜃̂MLE) = E(𝜃̂MLE) – θ = E(Z) – θ.                                               (2)",
          "topic": null
        },
        {
          "label": "vi",
          "marks": 1,
          "text": "Which of the following is TRUE about the mean square error of the estimator θ̂MLE?\n               A. MSE (θ̂MLE) = Bias2(θ̂MLE)\n               B. MSE (θ̂MLE) = Variance (θ̂MLE)\n               C. MSE (θ̂MLE) = Variance (θ̂MLE) + Bias2(θ̂MLE)\n               D. MSE (θ̂MLE) = Variance (θ̂MLE) – Bias2(θ̂MLE)                                            (1)",
          "topic": null
        }
      ],
      "solution": "i)         Based on sample mean:\n              X_bar = ∑ 𝑥 / n = 5.13/10 = 0.513\n\n              Mean of population = (b+a)/2 = θ / 2\n\n              Equating population mean with sample mean,\n              X_bar = θ / 2\n              0.513 = θ / 2\n              θMOM = 0.513*2 = 1.026\n\n              Based on sample variance:\n              S2\n              = 1/(n-1) * (∑ 𝑥 – n * (∑ 𝑥 / n)2)\n              = 1/9 * (2.8761 – 10 * 0.5132)\n              = 1/9 * 0.24441\n              = 0.027157\n\n              Variance of the population\n              = (b – a)2 / 12\n              = θ2 / 12\n\n              Equating population variance with sample variance,\n              S2 = θ2 / 12\n              θMOM\n              = (0.027157*12)0.5\n              = 0.5709 …………… as θ > 0.                                                                                  [3]\n\n   ii)        Correct Answer is Option B\n\n              X is uniformly distributed over the interval [0, θ]. So, X can take values which lie between 0 and θ\n              and not beyond θ.\n\n              L(θ, x) = 1 / θ10 for all xi ≤ θ i.e. θ > max(xi)\n              L(θ, x) = 0       otherwise.                                                                              [1]\n\n  iii)        L(θ, x) = 1 / θ10\n              log L = - 10 * log θ\n\n              Differentiating both sides,\n              d/dθ (log L)\n              = d/dθ (-10 * log θ)\n              = -10/ θ\n\n              Required equation to be solved to get MLE is:\n              d/dθ (log L) = 0\n              -10/ θ = 0\n\n              Kindly note that when we try to equate it with 0, we will get that θ tends to infinity. So, we won’t\n              get a finite value which maximizes the likelihood function.\n\n                                                                                                             Page 4 of 11\n\f IAI                                                                                                              CS1A-1123\n\n              Similarly, when we take a second derivative to check maxima,\n\n              d2/dθ2 (LogL) = 10 / θ2 > --------- This is indicative of minima and not maxima\n              Hence, the method of differentiation does not work in this case.                                              [3]\n\n  iv)         L(θ, x) = 1 / θ10 for θ > max(xi)\n              L(θ, x) = 0       otherwise.\n\n              Hence, till the point of θ = max(xi) i.e. for all values of θ which are lower than max(xi), the value of\n              the likelihood function is equal to 0 as can be seen in the graph.\n\n              At θ = max(xi), the likelihood function sees a sudden spike and for all values of θ which are greater\n              than max(xi), the likelihood function keeps on declining.\n\n              So, it is evident from the graph that the likelihood function is maximized at θ = max(xi). So, θMLE =\n              max(xi)                                                                                                       [2]\n\n   v)         E(Z)\n              = integral ( 10 * z * z9 / θ10 dz)0θ\n              = 10 / θ10 * (z11 / 11) 0θ\n\n              = 10 / θ10 * (θ11 / 11)\n              = 10/11 * θ\n\n              Bias (θMLE)\n              = E(θMLE) – θ\n              = E(Z) – θ\n              =10/11 * θ – θ\n              = -1/11 * θ\n              = – 11-1 * θ                                                                                                  [2]\n\n  vi)         Correct Answer is Option C\n\n              MSE (𝜃MLE)\n              = E (𝜃 MLE – θ)2\n              = Var (𝜃MLE) + Bias2(𝜃MLE)                                                                                   [1]",
      "has_math": true,
      "session": "2023-11",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2023-11_QP.pdf",
      "source_sol": "raw/CS1A_2023-11_SOL.pdf"
    },
    {
      "q_num": 5,
      "marks": 15,
      "topic": "inference",
      "subtopics": [
        "distributions"
      ],
      "stem": "Country Money-Land has two leading stock exchanges MSE and LSE. Stock market\n        growth in the country is measured using two important indices:\n              1. MSE’s MIFTY which is an index comprising of 50 stocks from diverse sectors\n                 listed on MSE.\n              2. LSE’s LENSEX which is an index comprising of 30 large cap stocks listed on\n                 LSE.\n        There are no common stocks in MIFTY and LENSEX and hence annual returns from both\n        these indices are assumed to be independent of each other.\n\n        Column A represents various available data sets relating to MSE and NSE and Column B\n        represents the type of data.\n\n                      Column “A” – Data Sets                         Column “B” – Type of Data\n          1. LSE’s LENSEX returns over the last 10 years.             i. Truncated Data\n\n          2. Returns on stocks which were listed on MSE in the        ii. Longitudinal Data\n             middle of the year.\n          3. Closing value of LSE’s LENSEX on 31st March,            iii. Cross-sectional Data\n             2023.\n          4. Data of stocks on MSE during a truncated week           iv. Censored Data\n             (i.e. a week where there are less than 5 working\n             days in the week due to public holidays).",
      "parts": [
        {
          "label": "i",
          "marks": 2,
          "text": "Above pairs are incorrectly matched. Which one of the following options from those\n               given below represents correctly matched pairs?\n\n               A. 1 – iii, 2 – iv, 3 – i, 4 – ii\n               B. 1 – ii, 2 – i, 3 – iv, 4 – iii\n               C. 1 – i, 2 – iv, 3 – iii, 4 – ii\n               D. 1 – ii, 2 – iv, 3 – iii, 4 – i                                                           (2)\n\n        A leading firm in the stock market which undertakes technical analysis has fitted a normal\n        distribution on returns from both MSE’s MIFTY as well as LSE’s LENSEX.\n        Random variables X and Y represents the average annual returns on MSE’s MIFTY and\n        LSE’s LENSEX respectively. X ~ N(µ, σ2) and Y ~ N (α, β2).\n\n        Random samples (X1, X2, …….., X10) and (Y1, Y2, …….., Y10) comprising of average\n        annual returns on MIFTY and LENSEX respectively for the past 10 years have been\n                                    ̅ and Y\n        collected with sample means X     ̅ and sample variance S2x and S2y.",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 4,
          "text": "Show that P (S2x > σ2) is 0.41. Use χ2 result for sample variance.                           (4)",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 1,
          "text": "̅ independent of sample variance S2x?\n               Is the sample mean X                                                                         (1)",
          "topic": null
        },
        {
          "label": "iv",
          "marks": 3,
          "text": "̅ > µ | S2x > σ2) using the result for sampling distribution of X\n               Determine P (X                                                               ̅.              (3)\n\n        MSE’s MIFTY is considered to be a more reliable index for tracking stock market growth\n        due to its wider and more diverse composition. LSE’s LENSEX however, tends to show\n        higher annual returns coupled with higher volatility.\n\n        You are given that: µ = 17%, σ2 = 16%%, α = 19%, β2 = 24%%.",
          "topic": null
        },
        {
          "label": "v",
          "marks": 4,
          "text": "Show that P (Sx ≤ Sy) is greater than 10%. Use F-result for variance ratios.                 (4)",
          "topic": null
        },
        {
          "label": "vi",
          "marks": 1,
          "text": "Technical analysts in the firm have developed a 30 × 30 (scaled) variance /\n               covariance matrix for average annual returns of the 30 stocks which form a part of\n               LSE’s LENSEX using principal component analysis (PCA). What is the maximum\n               number of non-zero values that this matrix can contain?\n\n               A. 30\n               B. 900\n               C. 0\n               D. None of the above.                                                                        (1)",
          "topic": null
        }
      ],
      "solution": "i)         Correct Answer is Option D\n\n              Since 1) from Column A relates to returns over a period of time, it is longitudinal data (Option ii\n              from Column B).\n\n              Since 2) from Column A relates to stocks listed during the middle of the year (hence data prior to\n              their listing would not be publicly available). Hence, it is censored data (Option iv from Column B)\n\n              Since 3) from Column A relates to value of LENSEX at a point of time, it is cross-sectional data\n              (Option iii from Column B)\n\n              Since 4) from Column A relates to stocks on MSE during a truncated week (which has less than 5\n              working days due to presence of public holidays), it is truncated data (Option I from Column B)               [2]\n\n                                                                                                                 Page 5 of 11\n\fIAI                                                                                                    CS1A-1123\n\n ii)    Probability that the sample returns for MSE’s MIFTY are more volatile as compared to the\n        population\n        = P (S2x > σ2)\n        = P (S2x / σ2 > 1)\n        = P ( (10-1) * S2x / σ2 > 9)\n        = P (9 * S2x / σ2 > 9)\n\n        (9 * S2x / σ2) ~ χ29\n\n        Probability that the sample returns for MSE’s MIFTY are more volatile as compared to the\n        population\n        = P(χ29 > 9)\n        = 1 - P(χ29 ≤ 9)\n        = 1 – 0.5627 ……………….. taken from Actuarial Tables\n        = 0.4373 ~ 0.4 (as required)                                                                             [4]\n\n iii)   Sample mean X and sample variance S2x are independent on each other. This is true because the\n        distribution of the population is normal distribution. For non-normal distributions, it may not be\n        true.                                                                                                    [1]\n\n iv)    Conditional probability that the sample tends to overestimate the average annual return of MIFTY\n        as compared to the population\n\n        = P( X > µ | S2x > σ2)\n        = P (X > µ) ……………… X and S2 are independent on each other\n\n        = P (X – µ > 0)\n        = P ( (X – µ) / (σ / √n) > 0)\n        = P (Z > 0) …………………. as X ~ N(µ, σ2 / n)\n        = 1 – P(Z ≤ 0)\n        = 1 – 0.5 = 0.5                                                                                          [3]\n\n v)     Probability that the sample standard deviation of the returns of MIFTY is lower than the sample\n        standard deviation of the returns of LENSEX\n        = P(Sx ≤ Sy)\n        = P(S2x ≤ S2y)\n        = P(S2x / S2y ≤ 1)\n\n        = P((S2x / S2y) * 1.5 ≤ 1.5)\n        = P((S2x / 16) / ( S2y / 24)) ≤ 1.5)\n        = P((S2x / σ2) / ( S2y / β2)) ≤ 1.5)\n        = P (F 9, 9 ≤ 1.5)\n        = P (F 9, 9 > 2/3)\n\n        As per Actuarial Tables, P(F 9, 9 > 2.440) = 10%, P(F 9, 9 > 3.179) = 5%\n\n        Hence P(F 9, 9 > 2/3) > 10%                                                                              [4]\n\n vi)    Correct Answer is Option A\n\n        As mentioned in part (vi), principal components are un-correlated combinations of the original\n        random variables, so in this variance / covariance matrix as well, the value of covariances for any\n        two distinct PCs will be equal to 0. So, 870 values will be 0.\n\n                                                                                                      Page 6 of 11\n\f IAI                                                                                                             CS1A-1123\n\n              Only the 30 values in the diagonal will represent the variance of each PC. So, the maximum number\n              of non-zero values in the variance / covariance matrix will be 30.",
      "has_math": true,
      "session": "2023-11",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2023-11_QP.pdf",
      "source_sol": "raw/CS1A_2023-11_SOL.pdf"
    },
    {
      "q_num": 6,
      "marks": 15,
      "topic": "bayes_credibility",
      "subtopics": [],
      "stem": "Seismologists consider that the approximate time (in days) between the occurrence of two\n        earthquakes in a particular seismic zone can be modelled as a random variable X with an\n                                                                      1\n        exponential distribution, having the density function: f(x) = µ 𝑒 −𝑥/µ .\n\n        It is proposed to use the following prior distribution for µ:\n\n                                                   θα e−θ/μ\n                                            f(µ) = μα+1Γ(α)          µ>0\n\n        The mean of this distribution is: θ / (α – 1).",
      "parts": [
        {
          "label": "i",
          "marks": 2,
          "text": "Write down the likelihood function of µ, based on observations x1, ……………., xn\n               from an exponential distribution.                                                            (2)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 3,
          "text": "Determine the posterior probability density function of µ and state its parameters.          (3)",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 2,
          "text": "Based on the posterior distribution derived in part (ii), show that an expression for\n               the Bayesian estimate of µ under squared error loss is given by:\n\n                                                            θ+ ∑ x\n                                                         μ̂= n+α−1                                          (2)",
          "topic": null
        },
        {
          "label": "iv",
          "marks": 1,
          "text": "This is a case of conjugate priors. Which of the following is TRUE in respect of\n               conjugate priors?\n\n               A. If the prior distribution leads to a posterior distribution is exactly identical to the\n                  distribution of the prior, only then this prior is called the conjugate prior for this\n                  likelihood.\n\n               B. Conjugate distributions often make Bayesian calculations simpler.\n\n               C. If the distribution of the random variable X is identical with the posterior\n                  distribution, then it is considered to be a case of conjugate priors.\n\n               D. If the prior distribution of a parameter is uniform, then the posterior distribution\n                  of the parameter will always be a uniform distribution.                                    (1)",
          "topic": null
        },
        {
          "label": "v",
          "marks": 3,
          "text": "Show that the Bayesian estimate of µ can be written in the form of a credibility\n               estimate giving formula for the credibility factor.                                           (3)\n\n         You are given that the parameters of the prior distribution are θ = 40 and α = 1.5.\n\n         You are given the following summary statistics from the sample data relating to the past\n         100 earthquakes in the seismic zone:\n\n                                       n = 100, ∑ 𝑥 = 9,000, ∑ 𝑥 2 = 12,00,000",
          "topic": null
        },
        {
          "label": "vi",
          "marks": 3,
          "text": "Calculate the prior mean, the sample mean, the Bayesian estimate of µ and the value\n               of the credibility factor.                                                                    (3)",
          "topic": null
        },
        {
          "label": "vii",
          "marks": 1,
          "text": "Based on the calculations in part (vi), what can you infer about the estimation\n               exercise? Choose the right option from those given below:\n\n               A. Bayesian estimate of µ is closer to the mean of the sample data.\n               B. Bayesian estimate of µ is closer to the mean of the prior distribution.\n               C. Bayesian estimate of µ is an average of the sample mean and prior mean.\n               D. Bayesian estimate of µ is equal to the mean of the sample data.                            (1)",
          "topic": null
        }
      ],
      "solution": "i)         L(µ)\n              = 1/µ * e-x1/µ * …………… * 1/µ * e-xn/µ\n              = 𝑒 ∑ /µ / µn                                                                                                [2]\n   ii)        The posterior PDF is proportional to the prior PDF multiplied by the likelihood function.\n\n              fpost(µ) ∝ (e –θ/µ / µ α+1) * (𝑒 ∑ /µ / µn)\n\n              =𝑒 (       ∑ )/µ\n                                 / µ n+α –1\n\n              This has the same form as the prior distribution, but with different parameters. So we have the\n              same distribution, but with parameters:\n\n              α* = n + α\n              θ* = θ + ∑ x                                                                                                 [3]\n\n  iii)        Using the formula for mean given in the question,\n\n              E(µ | x)\n              = θ* / (α* - 1)\n                   ∑\n              =\n\n              This is the Bayesian estimate of µ under squared error loss.                                                 [2]\n\n  iv)         Correct Answer is Option B\n\n              Option A is incorrect as the prior and posterior distributions need not be exactly identical for\n              conjugate priors. Even if the prior and posterior distributions belong to the same family, even then\n              the two are considered as conjugate priors.\n\n              Option B is correct as conjugate priors make Bayesian calculations simpler.\n\n              Option C is incorrect as the distribution of random variable X need not be identical with the posterior\n              distribution, for conjugate priors.\n\n              Option D is incorrect, as a prior uniform distribution does not necessarily lead to a posterior uniform\n              distribution.                                                                                                [1]\n\n   v)         Splitting the formula for posterior mean into two parts, we see that:\n\n              E(µ | x)\n\n                   ∑\n              =\n\n                             ∑\n              =          +\n\n                             ∑\n              =          ×        +           ×\n\n                                                                                                                Page 7 of 11\n\f IAI                                                                                                           CS1A-1123\n\n                                                                                                                ∑\n              This is a weighted average of the maximum likelihood estimate of µ (which is the sample mean\n              ) and the mean of the prior distribution     . So it is a credibility estimate.\n\n              The credibility factor is:\n                                                          Z=                                                              [3]\n\n  vi)         Prior mean = 40 / (1.5 – 1) = 80\n\n              Sample Mean = 9000/100 = 90\n\n              Using the given figures, the Bayesian estimate of µ is:\n\n                    ∑\n              =\n              = (40 + 9000) / (100 + 1.5 – 1)\n              = 89.95\n\n              The value of the credibility factor is:\n\n              Z\n              =\n              = 100 / (100 + 1.5 – 1)\n              = 0.9950                                                                                                    [3]\n\n  vii)        Correct Answer is Option A\n\n              Since the value of the credibility factor is close to 1, the posterior estimate is closer to the sample\n              mean (direct data) as compared to the prior mean (collateral data). Sample mean in this case is 90\n              and prior mean is 80. It is evident that the posterior estimate 89.95 is closer to the sample mean.         [1]",
      "has_math": true,
      "session": "2023-11",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2023-11_QP.pdf",
      "source_sol": "raw/CS1A_2023-11_SOL.pdf"
    },
    {
      "q_num": 7,
      "marks": 16,
      "topic": "distributions",
      "subtopics": [],
      "stem": "The insurance regulator has announced a move to a new solvency regime to align with\n         international practice. Under the new regime, the capital requirement for each line of\n         business will be determined based on expected losses arising from a line of business.\n         Let Y be a random variable representing the expected losses arising from a line of\n         business. The capital requirement for that line of business is the value y such that P(Y≤y)\n         = 0.995.\n         The Appointed Actuary of a general insurance company is reading some examples of\n         capital requirement calculation. All examples are in lakhs of INR.\n         In the first example, the losses Y arising from a line of business are modelled as Y ~\n         Gamma(θ, 1/υ).\n\n         For a particular line of business, θ = 4 and υ = 25.",
      "parts": [
        {
          "label": "i",
          "marks": 2,
          "text": "Express the mean and variance of Y in terms of θ and υ and calculate the mean and\n             standard deviation of the losses from this business.                                        (2)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 1,
          "text": "What is the capital requirement for this line of business?\n             (The table of percentages for the distribution is at the end of this question.)             (1)\n\n      In the second example, frequency N and severity X are modelled separately. Here, N is a\n      random variable representing the number of claims, and X is a random variable\n      representing the cost of claim. Thus, Y is modelled as\n                                                        𝑁\n\n                                                 𝑌 = ∑ 𝑋𝑖\n                                                       𝑖=1\n      The Xi’s are assumed to be independent and identically distributed. N is assumed to be\n      independent of the Xi’s.",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 1,
          "text": "In the context of the above example, which of the following is necessary and\n             sufficient for random variables to be “independent and identically distributed”?\n\n             A. Both variables have identical distributions and are uncorrelated to each other.\n             B. Both variables belong to the same family of distributions and are not dependent\n                on each other.\n             C. Both variables have identical distributions and are not dependent on each other.\n             D. Both A and C                                                                             (1)",
          "topic": null
        },
        {
          "label": "iv",
          "marks": 1,
          "text": "Which type of random variables will be used to model N and X in the second\n             example? Choose the right option from those given below:\n\n             A. Discrete random variables would be used to model both N and X.\n             B. X would be modelled as a discrete random variable whereas N would be\n                modelled as a continuous random variable.\n             C. N would be modelled as a discrete random variable whereas X would be\n                modelled as a continuous random variable.\n             D. Continuous random variables would be used to model both N and X.                         (1)\n\n      The example further states that N ~ Poisson (μ) and X ~ Gamma (α, 1/β), where μ = 10, α\n      = 2/3 and β = 15.",
          "topic": null
        },
        {
          "label": "v",
          "marks": 10,
          "text": "Show that E(Y) = μαβ and var(Y) = μαβ2 (1+α). Using the values of μ, α and β given\n             above, calculate the mean and variance of Y.\n             Hint: Use the results E(Y) = E[E(Y|N)] and Var(Y) = E[Var(Y|N)] + var[E(Y|N)]               (7)\n\n      In the second example, capital requirement is calculated using Monte Carlo simulation.\n      Simulation is to be done using the following three steps:\n\n      Step 1: First, we need to simulate a value for number of claims ‘n’ given by random\n      variable N where N ~ Poisson (10).\n\n      Step 2: Using the simulated value ‘n’ generated above, ‘n’ values are simulated for claim\n      amounts ‘x’ represented by random variable X where X ~ Gamma (2/3, 1/15)\n\n        Step 3: The sum of these n claim amounts is the simulated value of Y.",
          "topic": null
        },
        {
          "label": "vi",
          "marks": 4,
          "text": "Follow the given steps to simulate a value of Y. Use the random variates and the\n              table of gamma probabilities as given below.\n              Random variates taken from a U(0,1) distribution: 0.19, 0.95, 0.70, 0.80, 0.20, 0.10,\n              0.20, 0.50, 0.60, 0.80\n              [The first random variate can be used to simulate n, and the remaining can be used to\n              simulate x values.]                                                                      (4)\n\n                  Table of percentages of Gamma distributions\n                P(Y≤y)               Gamma (4, 1/25)   Gamma (2/3, 1/15)\n                 0.1%                    10.71               0\n                 0.5%                    16.80              0.01\n                 1.0%                    20.58              0.01\n                 2.5%                    27.25              0.05\n                 5.0%                    34.16              0.14\n                10.0%                    43.62              0.41\n                20.0%                    57.42              1.21\n                30.0%                    69.09              2.32\n                40.0%                    80.28              3.77\n                50.0%                    91.80              5.65\n                60.0%                   104.38              8.11\n                70.0%                   119.06             11.48\n                80.0%                   137.88             16.46\n                90.0%                   167.02             25.39\n                95.0%                   193.84             34.64\n                97.5%                   219.18             44.10\n                99.0%                   251.13             56.82\n                99.5%                   274.44             66.56\n                99.9%                   326.56             89.43                                      [16]",
          "topic": null
        }
      ],
      "solution": "i)        Reading the formulae from the table,\n\n              E(Y) = α/λ = θυ\n\n              Var(Y) = α/ λ2 = θυ2\n\n              Mean = θυ = 4 * 25 = INR 100 lakh\n\n              Standard deviation = √θ𝜐 = INR 50 lakh                                                                      [2]\n\n   ii)        From the table, the 99.5th percentile value of Gamma(4,1/25) is INR 274.44 lakhs. This is the capital\n              requirement.                                                                                                [1]\n\n  iii)        Correct Answer is Option C\n\n              Independent and identically distributed needs the following two aspects –\n                  - Variables have identical distributions (belonging to the same family of distributions is not\n                     sufficient, the distributions have to be identical)\n                  - Variables are not dependent (uncorrelated variables are not sufficient, we know that when\n                     two variables are independent they are necessarily uncorrelated, but when two variables\n                     are uncorrelated they are not necessarily independent of each other)                                 [1]\n\n                                                                                                              Page 8 of 11\n\f IAI                                                                                                           CS1A-1123\n\n  iv)         Correct Answer is Option C\n\n              For N, a discrete distribution would be appropriate, as the number of claims would be a non-\n              negative whole number.\n\n              For X, a continuous distribution would be appropriate, as it is a monetary amount, which is a\n              continuous quantity.\n              (Although it may be argued that money is not infinitely subdivisible, modelling it as a discrete\n              quantity serves no purpose – e.g. if the 99.99th percentile of X is 1,000,000; modelling a range of 0\n              to 1,000,000 in steps of say 0.01 would get us no significant increase in accuracy.)                       [1]\n\n   v)         E(X) = αβ\n\n              E(Y) = E[E(Y|N)] = E[N E(X)] = E[N αβ] = E(N) αβ = μαβ\n\n              Var(X) = αβ2\n\n              Var(Y) = E[Var(Y|N)] + var[E(Y|N)]\n              E[Var (Y|N)] = E[N Var(X)] = E(N) αβ2 = μαβ2\n\n              Var[E(Y|N)] = Var[N E(X)] = var(N) α2β2 = μα2β2\n\n              Thus Var (Y) = μαβ2 + μα2β2 = μαβ2 (1+α)\n\n              E(Y) = μαβ = 10 * 2/3 * 15 lakh = INR 100 lakh\n\n              Var(Y) = E(Y) (1 + α) β = 100 lakh * 5/3 * 15 lakh = 25 * 1012                                             [7]\n\n  vi)         First, we need to simulate a value for N, which follows Poisson (10).\n              First random variate is 0.19.\n              From Tables, P(N = 6) < 0.19 < P(N=7). Thus the simulated value of N is 7.\n\n              We now need to simulate 7 claims.\n\n              The next 7 random variates and the corresponding values from the Gamma (2/3, 1/15) distribution\n              are:\n\n               Claim No Variate Corresponding X value\n                   1     0.95           34.64\n                   2     0.70           11.48\n                   3     0.80           16.46\n                   4     0.20           1.21\n                   5     0.10           0.41\n                   6     0.20           1.21\n                   7     0.50           5.65\n\n              Adding up the X values, the simulated value of Y is 71.06 lakh rupees.                                    [4]",
      "has_math": true,
      "session": "2023-11",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2023-11_QP.pdf",
      "source_sol": "raw/CS1A_2023-11_SOL.pdf"
    },
    {
      "q_num": 8,
      "marks": 18,
      "topic": "regression_glm",
      "subtopics": [],
      "stem": "A real estate analyst has collected data of the average floor price index (per-square-foot\n        apartment price) and the population density of various metropolitan areas in the country.\n        The data is given below:\n\n               Population Density                   Floor price index\n          (X) (In unit of 1000 persons)         (Y) (In unit of 1000 INR)\n                   per sq. km                           per sq. foot\n                        17                                  8.50\n                        12                                  6.90\n                        25                                 18.00\n                        10                                  4.00\n                         5                                  0.50\n        She decides to fit a linear model to the data, where population density is the explanatory\n        variable X and floor price index is the response variable Y.",
      "parts": [
        {
          "label": "i",
          "marks": 2,
          "text": "Calculate x̄ and ȳ.                                                                     (2)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 3,
          "text": "Calculate Sxx, Syy and Sxy.\n             You are given that:\n             ∑ 𝑥 2 = 1183.00\n             ∑ 𝑦 2 = 460.11\n             ∑ 𝑥𝑦 = 719.80                                                                            (3)",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 3,
          "text": "Determine the fitted regression line: 𝑦̂ = α + βx.                                       (3)",
          "topic": null
        },
        {
          "label": "iv",
          "marks": 4,
          "text": "Calculate a 99% confidence interval for the slope parameter β.                           (4)\n\n       The residuals 𝑦 − 𝑦̂ from the linear regression are as follows:\n                    x          Residual\n                    5            0.42\n                   10           -0.34\n                   12            0.85\n                   17           -1.81\n                   25            0.87",
          "topic": null
        },
        {
          "label": "v",
          "marks": 1,
          "text": "Which of the following is TRUE about the distribution of residuals? Choose the right\n             option from those given below:\n             A. The residual values don’t seem to have any relationship with x and seem to be\n                distributed approximately normally around the origin.\n             B. Residual values seem to be positively correlated with x and hence are not\n                normally distributed.\n             C. Residual values seem to be negatively correlated with x and hence are not\n                normally distributed.\n             D. Residual values are not independent of each other and hence are not normally\n                distributed.                                                                          (1)\n\n       The analyst obtains another data point to add to her data, so n = 6.\n       After adding this data point, Sxx = 250, Syy = 190, Sxy = 200.",
          "topic": null
        },
        {
          "label": "vi",
          "marks": 1,
          "text": "What proportion of variance in the floor price index is now explained by the model?\n             Choose the right option from those given below:\n             A. 84%\n             B. 72%\n             C. 42%\n             D. 93%                                                                                   (1)\n\n       The analyst adds a second explanatory variable to her regression and R2 becomes 87%.",
          "topic": null
        },
        {
          "label": "vii",
          "marks": 4,
          "text": "Calculate the adjusted R2 for both the one-variable model (considering 6 data points)\n             and the two-variable model, and comment on which is the better model.                    (4)",
          "topic": null
        }
      ],
      "solution": "i)         x̄ = 13.8\n              ȳ = 7.58\n\n                                                                                                              Page 9 of 11\n\fIAI                                                                                           CS1A-1123\n\n ii)\n        𝑆   =       𝑥 − 𝑛𝑥̅ = 230.8\n\n        𝑆    =       𝑦 − 𝑛𝑦 = 172.828\n\n        𝑆    =      𝑥 𝑦 − 𝑛𝑥𝑦          = 196.78                                                        [3]\n\n iii)       𝑆\n        𝛽=     = 0.85\n            𝑆\n        𝛼 = 𝑦 − 𝛽𝑥̅ = −4.186\n        So the fitted regression line is:\n        𝑦 = −4.186 + 0.85x .                                                                           [3]\n\n iv)    A 99% confidence interval for the slope parameter β is given by:\n\n                                                                 𝜎\n                                              𝛽 ± 𝑡      ; .\n                                                                 𝑠\n\n        From our data,\n\n        𝜎\n        = 1 / (n – 2) * (Syy – S2xy / Sxx)\n        = 1 / 3 * (172.83 –196.782 / 230.8)\n        = 1.6845\n\n        So the 99% confidence interval for the slope parameter β is given by:\n\n                           .\n        0.85 ± 5.841\n                               .\n        = 0.85 ± 0.499\n        = (0.351, 1.349)                                                                               [4]\n\n v)     Correct Answer is Option A\n\n        The residual values don’t seem to have any relationship with x and seem to be distributed\n        approximately normally around the origin. As such, the distribution is as expected.            [1]\n\n vi)    Correct Answer is Option A\n\n        Proportion of variance explained by the model is given by R2.\n        R2 = S2xy / (Sxx * Syy) = 2002 / (190*250) = 0.84                                              [1]\n\nvii)                                𝑛−1\n        𝐴𝑑𝑗𝑢𝑠𝑡𝑒𝑑 𝑅 = 1 −                 (1 − 𝑅 )\n                                   𝑛−𝑘−1\n\n        For one variable model,\n                                    6−1\n        𝐴𝑑𝑗𝑢𝑠𝑡𝑒𝑑 𝑅 = 1 −                  (1 − 0.84) = 0.80\n                                   6−1− 1\n\n        For two variable model,\n                                    6−1\n        𝐴𝑑𝑗𝑢𝑠𝑡𝑒𝑑 𝑅 = 1 −                  (1 − 0.87) = 0.78\n                                   6−2− 1\n\n                                                                                           Page 10 of 11\n\fIAI                                                                                                     CS1A-1123\n\n      As the adjusted R2 for the one variable model is higher, it can be considered the better model.           [4]\n                                       ************************\n\n                                                                                                  Page 11 of 11",
      "has_math": true,
      "session": "2023-11",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2023-11_QP.pdf",
      "source_sol": "raw/CS1A_2023-11_SOL.pdf"
    },
    {
      "q_num": 1,
      "marks": 2,
      "topic": "inference",
      "subtopics": [],
      "stem": "A renowned social media influencer Mr. Numero Uno has openly claimed that his new reel on\n        Xstagram will receive exactly 40 likes in a minute. Probability that this claim turns out to be\n        true is –\n\n        A. 0.193%\n        B. 6.295%\n        C. 8.115%\n        D. 0.075%                                                                                            [2]",
      "parts": [],
      "solution": "Correct Answer is Option D.\n\n              Let X denote the number of likes received in a minute.\n              Since, the number of likes received per second is 0.4, X ~ Poi(24).\n\n              P(X=40)\n              = 2440 / 40! * exp(-24)\n              = 0.075%                                                                                           [2]",
      "has_math": false,
      "session": "2024-05",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2024-05_QP.pdf",
      "source_sol": "raw/CS1A_2024-05_SOL.pdf"
    },
    {
      "q_num": 2,
      "marks": 2,
      "topic": "distributions",
      "subtopics": [],
      "stem": "The mean waiting time between two consecutive likes for a new post on Xstagram will be –\n\n        A. 4 seconds\n        B. 2.5 seconds\n        C. 5 seconds\n        D. 0.4 seconds                                                                                       [2]",
      "parts": [],
      "solution": "Correct Answer is Option B.\n\n              If the number of likes follow a Poisson process with mean of 0.4 likes per second, the waiting\n              time between two consecutive likes is exponentially distributed with parameter = 0.4.\n\n              Mean waiting time\n              = Mean(Exp(0.4))\n              = 1/0.4\n              = 2.5 seconds.                                                                                     [2]",
      "has_math": false,
      "session": "2024-05",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2024-05_QP.pdf",
      "source_sol": "raw/CS1A_2024-05_SOL.pdf"
    },
    {
      "q_num": 3,
      "marks": 2,
      "topic": "distributions",
      "subtopics": [],
      "stem": "A novice on social media, Mr. Numero Cero has already waited for 10 seconds for receiving a\n        like for his new post on Xstagram with no success. What is the probability that he will have to\n        wait for at least 10 more seconds for receiving his first like?\n\n        Hint: Use the memoryless property of an exponential distribution.\n\n        A. 0.005%\n        B. 0.034%\n        C. 1.832%\n        D. 2.732%                                                                                            [2]\n\n        Q.4 and Q.5 are based on the information presented below:\n        Two other social media platforms viz. Scapebook and Fritter are being compared by the\n        research firm in terms of the waiting time between two consecutive likes. Target audience on\n        both these platforms is different – Scapebook is used mostly by youngsters up to age 40 and\n        Fritter is generally used by the older population with ages 40 and above.\n        Waiting times are modelled using independent exponential variables X (Scapebook) and Y\n        (Fritter) with means of 1 second and 2 seconds respectively.",
      "parts": [],
      "solution": "Correct Answer is Option C.\n\n              Numero Cero has waited for 10 seconds already, and we has to wait for 10 more seconds.\n              Using the memoryless property of an exponential distribution, waiting for 10 more seconds\n              given that there is already a waiting of 10 seconds, is equivalent to waiting for 10 seconds.\n\n              P(W > 10)\n              = exp(-0.4*10)\n              = 1.832%                                                                                           [2]",
      "has_math": false,
      "session": "2024-05",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2024-05_QP.pdf",
      "source_sol": "raw/CS1A_2024-05_SOL.pdf"
    },
    {
      "q_num": 4,
      "marks": 2,
      "topic": "distributions",
      "subtopics": [],
      "stem": "Which of the following is the correct expression representing f (x, y) i.e. the joint distribution\n        of X and Y?\n\n        A. f (x, y) = exp(-x) × ½ exp(-½y) for ∞ > x, y > 0;\n                    =0                      otherwise\n        B. f (x, y) = 2 exp(-2x) × exp(-y) for ∞ > x, y > 0;\n                    =0                      otherwise\n\n        C. f (x, y) = exp(-x) × 2 exp(-2y) for ∞ > x, y > 0;\n                    =0                     otherwise\n        D. f (x, y) = ½ exp(-½x) × exp(-y) for ∞ > x, y > 0;                                                 [2]\n                    =0                      otherwise",
      "parts": [],
      "solution": "Correct Answer is Option A.\n\n              f(x) = exp(-x)\n              f(y) = ½ *exp(-½y)\n\n              As X and Y are independent variables,\n              f (x, y) = f(x) * f(y)\n\n              f (x, y) = exp(-x) × ½ exp(-½y) for ∞ > x, y > 0;\n                       =0                       otherwise                                                        [2]",
      "has_math": true,
      "session": "2024-05",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2024-05_QP.pdf",
      "source_sol": "raw/CS1A_2024-05_SOL.pdf"
    },
    {
      "q_num": 5,
      "marks": 2,
      "topic": "distributions",
      "subtopics": [],
      "stem": "Let the joint MGF M x, y (t, s) of the joint distribution of X and Y be defined as follows:\n\n                                           M x, y (t, s) = E [ exp(xt + ys) ].\n\n        Which of the following statements is TRUE in respect of the joint MGF of random variables X\n        and Y i.e. M x, y (t, s)?\n\n        A. M x, y (t, s) = Mx(t) + My(s)\n        B. M x, y (t, s) = Mx(t) × My(s)\n        C. M x, y (t, s) = Mx(t) – My(s)\n        D. None of the above                                                                                 [2]\n\n        Q.6 to Q.10 are based on the information presented below:\n\n        The number of cheques dishonoured y, in a batch of x cheques meant for clearing in a local\n        branch of a cooperative bank is modelled as a Poisson random variable with mean λx, where λ\n        is unknown.\n\n        Data is available from six independent batches of proposals as follows:\n\n         Batch No.          1               2             3             4         5            6\n            x:             50              75           105            150       40           60\n            y:              1              21            17             9         4            3\n\n             ∑ x = 480; ∑ x 2 = 46850; ∑ y = 55; ∑ y 2 = 837; ∑ xy = 5100",
      "parts": [],
      "solution": "Correct Answer is Option B.\n\n              Mx(t) = E[exp(tx)]\n              My(s) = E[exp(sy)]\n\n              M x, y (t, s)\n              = E [ exp(xt + ys) ]\n              = E[exp(xt)] * E[exp(ys)]\n              = Mx(t) * My(t)                                                                                    [2]",
      "has_math": true,
      "session": "2024-05",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2024-05_QP.pdf",
      "source_sol": "raw/CS1A_2024-05_SOL.pdf"
    },
    {
      "q_num": 6,
      "marks": 1,
      "topic": "inference",
      "subtopics": [],
      "stem": "The method of moments estimate λ̂MOM based on the given data is equal to –\n\n        A. 8.727\n        B. 0.786\n        C. 0.115\n        D. 1.271                                                                                             [1]",
      "parts": [],
      "solution": "Correct Answer is Option C.\n\n              λ̂MOM = ∑ 𝑦 / ∑ 𝑥 = 55 / 480 = 0.115                                                               [1]\n\n                                                                                                  Page 2 of 12\n\f   IAI                                                                                            CS1A-0524",
      "has_math": true,
      "session": "2024-05",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2024-05_QP.pdf",
      "source_sol": "raw/CS1A_2024-05_SOL.pdf"
    },
    {
      "q_num": 7,
      "marks": 3,
      "topic": "inference",
      "subtopics": [],
      "stem": "Let us define the least square estimator λ̂LSE . By definition, it is that value of λ for which\n        ∑(y − E(y))2 is minimized.\n\n        Which of the following represents the correct expression for the least square estimator of λ i.e.\n        λ̂LSE ?\n                   ∑ y2\n        A. λ̂LSE = ∑ x2\n                   ∑ xy\n        B. λ̂LSE = ∑ 2\n                     y\n                   ∑x\n        C. λ̂LSE = ∑ y\n                   ∑ xy\n        D. λ̂LSE = ∑ x2                                                                                      [3]",
      "parts": [],
      "solution": "Correct Answer is Option D.\n\n               Expression to be minimized\n               =∑(𝑦 − 𝐸(𝑦))2\n               =∑(𝑦 − ʎ𝑥)2\n\n               𝑑\n                  ∑(𝑦 − ʎ𝑥)2\n               𝑑ʎ\n\n               = -2∑ 𝑥𝑦 (1) + 2ʎ(∑ 𝑥 2 )\n\n               We equate this with 0 to find the value of least square estimator of ʎ\n\n               ʎLSE = ∑ 𝑥𝑦 / ∑ 𝑥 2\n\n               𝑑2\n                     ∑(𝑦 − ʎ𝑥)2 =2(∑ 𝑥 2 ) > 0 ------- this is indicative of minima\n               𝑑ʎ2                                                                                              [3]",
      "has_math": true,
      "session": "2024-05",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2024-05_QP.pdf",
      "source_sol": "raw/CS1A_2024-05_SOL.pdf"
    },
    {
      "q_num": 8,
      "marks": 2,
      "topic": "inference",
      "subtopics": [],
      "stem": "Which of the following is the likelihood function of λ?\n\n        A. L(λ) = e−λ ∑ x × ∏(λx)y × c\n        B. L(λ) = e−λ ∏ x × ∏(λx)y × c\n        C. L(λ) = e−λ ∑ x × ∑(λx)y × c\n        D. L(λ) = e−λ ∏ x × ∑(λx)y × c",
      "parts": [],
      "solution": "Correct Answer is Option A.\n\n               L(λ)\n               = 𝑒 −𝜆 ∑ 𝑥 × ∏(𝜆𝑥)𝑦 × 1/ ∏ 𝑦!\n               = 𝑒 −𝜆 ∑ 𝑥 × ∏(𝜆𝑥)𝑦 × constant                                                                   [2]",
      "has_math": true,
      "session": "2024-05",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2024-05_QP.pdf",
      "source_sol": "raw/CS1A_2024-05_SOL.pdf"
    },
    {
      "q_num": 9,
      "marks": 3,
      "topic": "inference",
      "subtopics": [],
      "stem": "Which of the following is TRUE regarding the maximum likelihood estimator λ̂MLE which\n        maximizes the likelihood function as selected in Q.8 –\n\n        A. λ̂MLE = λ̂LSE\n        B. λ̂MLE = λ̂MOM\n        C. λ̂MLE = λ̂LSE = λ̂MOM\n        D. None of the above                                                                               [3]",
      "parts": [],
      "solution": "Correct Answer is Option B.\n\n               log L(ʎ)\n               = −𝜆 ∑ 𝑥 + ∑ 𝑙𝑜𝑔ʎ y+ constant\n               = −𝜆 ∑ 𝑥 +𝑙𝑜𝑔ʎ ∑ 𝑦 + constant\n\n               d/dʎ (log L(ʎ))\n               = − ∑ 𝑥 + ∑ 𝑦 (1/ʎ)\n\n               Equating d/dʎ (log L(ʎ)) = 0\n               ∑ 𝑦 (1/ʎ) = ∑ 𝑥\n               ʎ = ∑𝑦/∑𝑥\n\n               d2/dʎ2 (log L(ʎ))\n               = ∑ 𝑦 (-1/ʎ2)\n               = − ∑ 𝑦 / ʎ2 -------------- this is indicative of maxima\n\n                               ∑𝑦                                      ∑𝑦\n               Thus, 𝜆̂𝑀𝐿𝐸 = ∑ 𝑥 . From part (i), we know that 𝜆̂𝑀𝑂𝑀 = ∑ 𝑥 but from part Q8, we know ʎLSE =\n               ∑ 𝑥𝑦 / ∑ 𝑥 2 . Hence 𝜆̂𝑀𝐿𝐸 = 𝜆̂𝑀𝑂𝑀                                                               [3]",
      "has_math": true,
      "session": "2024-05",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2024-05_QP.pdf",
      "source_sol": "raw/CS1A_2024-05_SOL.pdf"
    },
    {
      "q_num": 10,
      "marks": 1,
      "topic": "bayes_credibility",
      "subtopics": [],
      "stem": "The least square estimate λ̂LSE based on the given sample data is –\n\n        A. 0.0179\n        B. 6.0932\n        C. 8.7273\n        D. 0.1089                                                                                          [1]\n\n        Q.11 and Q.12 are based on the information presented below:\n\n        Following data tabulates the number of matches (in thousands) on a dating application Rumble\n        for the last five years for four different categories.\n                                           Number of matches in Year (j)\n                                    2019     2020     2021     2022      2023      Mean Variance\n                   Under 18          48       53       42       50        59        50.4   39.3\n          Category Adults            64       71       64       73        70        68.4   17.3\n             (i)   Divorced          85       54       76       65        90        74.0  215.5\n                   Seniors           44       52       69       55        71        58.2  132.7\n                                                       Variance of Means:          110.57\n\n        The credibility factors Zi for all categories are to be determined using the assumptions of EBCT\n        (Empirical Bayes Credibility Theory) Model 1.",
      "parts": [],
      "solution": "Correct Answer is Option D.\n\n               ʎLSE\n               = ∑ 𝑥𝑦 / ∑ 𝑥 2\n               = 5100/46850\n               = 0.1089                                                                                         [1]",
      "has_math": true,
      "session": "2024-05",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2024-05_QP.pdf",
      "source_sol": "raw/CS1A_2024-05_SOL.pdf"
    },
    {
      "q_num": 11,
      "marks": 3,
      "topic": "bayes_credibility",
      "subtopics": [
        "distributions"
      ],
      "stem": "Value of credibility factor ZA (for adults category) is –\n\n        A. 0.8176\n        B. 0.8485\n        C. 0.7812\n        D. 0.8169                                                                                          [3]",
      "parts": [],
      "solution": "Correct Answer is Option D.                                                                      [3]\n\n                                                                                                 Page 3 of 12\n\f  IAI                                                                                               CS1A-0524\n\n               E(s2(θ))\n               = Mean (Variances for individual categories)\n               = (39.3 + 17.3 + 215.5 + 132.7) / 4\n               = 101.20\n\n               Var(m(θ))\n               = Variance(Means for individual categories) – 1/n * E(s2(θ))\n               = 110.57 – 1/5 * 101.20\n               = 90.33\n\n               n=5\n\n               ZA\n               = n / (n + E(s2(θ)) / Var(m(θ)))\n               = 5 / (5 + 101.20 / 90.33)\n               = 0.8169",
      "has_math": false,
      "session": "2024-05",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2024-05_QP.pdf",
      "source_sol": "raw/CS1A_2024-05_SOL.pdf"
    },
    {
      "q_num": 12,
      "marks": 2,
      "topic": "bayes_credibility",
      "subtopics": [
        "distributions",
        "data_analysis"
      ],
      "stem": "Value of expected number of matches for adults category (in thousands) is –\n\n        A. 68.40\n        B. 67.37\n        C. 62.75\n        D. 63.81                                                                                           [2]\n\n        Q.13 to Q.15 are based on the information presented below:\n        A 4 × 4 correlation matrix using Karl Pearson’s Method has been constructed for the data\n        collected from the dating application Rumble tabulating the sample correlation coefficients:\n                                 Under 18           Adults            Divorced         Seniors\n            Under 18               1.00               0.63               0.12           0.13\n            Adults                 0.63               1.00              -0.51           0.02\n            Divorced               0.12              -0.51               1.00           0.33\n            Seniors                0.13               0.02               0.33           1.00\n        a",
      "parts": [],
      "solution": "Correct Answer is Option B.\n               Credibility Premium for adults category\n               = Z * (Category Mean) + (1 – Z) * (Overall Mean)\n               = 0.8169 * (68.40) + (1 – 0.8169) * (50.4 + 68.4 + 74.0 + 58.2)/4\n               = 67.3655\n               = 67.37                                                                                           [2]",
      "has_math": true,
      "session": "2024-05",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2024-05_QP.pdf",
      "source_sol": "raw/CS1A_2024-05_SOL.pdf"
    },
    {
      "q_num": 13,
      "marks": 1,
      "topic": "distributions",
      "subtopics": [],
      "stem": "Based on this table, which of the following statements is FALSE?\n\n        A. Teenagers(under 18) and adults tend to show similar dating behaviour.\n        B. Dating behaviour for seniors is hardly related to the dating behaviour for adults.\n        C. Adults and divorcees tend to show dissimilar dating behaviour.\n        D. Dating behaviour of seniors is more similar to teenagers (under 18) than to divorcees.            [1]",
      "parts": [],
      "solution": "Correct Answer is Option D.\n\n               Option A is true as teenagers and adults show a positive correlation of 0.63.\n\n               Option B is true as seniors and adults show a correlation of 0.02 which is very close to 0.\n\n               Option C is true as non-divorcee adults and divorcee adults have a negative correlation of -\n               0.51\n\n               Option D is false as the correlation between divorcees and seniors is 0.33. This is much\n               higher than 0.13 i.e. the correlation between seniors and teenagers.                              [1]",
      "has_math": false,
      "session": "2024-05",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2024-05_QP.pdf",
      "source_sol": "raw/CS1A_2024-05_SOL.pdf"
    },
    {
      "q_num": 14,
      "marks": 2,
      "topic": "inference",
      "subtopics": [],
      "stem": "It is decided to perform a statistical test to check whether the dating behaviour of teenagers\n       (under 18) and divorcees is uncorrelated with each other. Students’ t distribution is to be used\n       for performing this statistical test.\n\n        What is the value of the test statistic for the above test?\n\n        A. 0.27\n        B. 3.70\n        C. 0.21\n        D. 4.78                                                                                              [2]",
      "parts": [],
      "solution": "Correct Answer is Option C.\n               Value of t test statistic\n               = r √(𝑛 − 2) / (√1 − 𝑟 2 )\n               = 0.12 * √3 / (√1 − 0.122 )\n               = 0.209359\n               = 0.21                                                                                            [2]",
      "has_math": false,
      "session": "2024-05",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2024-05_QP.pdf",
      "source_sol": "raw/CS1A_2024-05_SOL.pdf"
    },
    {
      "q_num": 15,
      "marks": 2,
      "topic": "regression_glm",
      "subtopics": [
        "inference"
      ],
      "stem": "Which one of the following options correctly represents the p-value for this test and the\n       conclusion of this statistical test at 5% level of significance?\n\n                 p-value    Conclusion of the statistical test\n            A.   < 40%      Dating behaviour of teenagers (under 18) and divorcees is\n                            uncorrelated\n            B.   < 5%       Dating behaviour of teenagers (under 18) and divorcees is correlated\n            C.   > 40%      Dating behaviour of teenagers (under 18) and divorcees is\n                            uncorrelated\n            D.   < 1%       Dating behaviour of teenagers (under 18) and divorcees is correlated             [2]\n\n        Q.16 to Q.20 are based on the information presented below:\n\n        Your team member Mr. Left was working on a project which involved fitting a linear regression\n        model for testing effectiveness of new drug of a fertilizer manufacturing company. The\n        company is launching a new fertilizer DAP and is testing the impact of the fertilizer on the crop\n        yield based on trials done over 20 hectares of paddy crop. The model helps to predict the crop\n        yield(Y) of paddy crop per hectare based on amount of fertilizer(X) employed in the field per\n        hectare.\n\n        Mr. Left has resigned and the fitted linear regression model which was saved on his laptop\n        cannot be retrieved as his laptop has been formatted. You have to give a presentation to the\n        fertilizer manufacturing company tomorrow. On his writing desk you find a printout which\n        contains details of a sample of ten Xi values, corresponding fitted ̂\n                                                                            Yi values.\n\n                                 ̂i) from the printout are as follows:\n        The ten pairs of of (Xi, Y\n                    A = (10, 40); B = (35, 165); C = (20, 90); D = (80, 390); E = (90, 440);\n                    F = (10, 40); G = (95, 465); H = (45, 215); I = (30, 140); J = (60, 290).\n\n        You had called Mr. Left and he informed you that the model was fitted based on 100 values of\n        X and Y obtained from the trails conducted by the company. Mr. Left also informed you that\n        he remembers that the coefficient of determination i.e. R2 for this model is 70%.\n\n        Apart from the values of (Xi, Ŷi), the printout also contained information about residual values\n        ei for each pair A to J. Based on this, you have constructed the following scatter plot. Every\n        point shown in the plot is (X, Y) where Y = ̂  Y + e.\n\n        Also, the fitted regression line has been back-calculated based on the data available from the\n        printout and plotted in the graph as the red line.\n\n                                      ̂i) will you require to obtain the regression equation of Y",
      "parts": [],
      "solution": "Correct Answer is Option C.\n\n               Degrees of freedom for the t distribution = 5 – 2 = 3.\n\n               From Tables, P(t3 > 0.2767) = 40%.\n\n               p-value\n               = P(t3 > 0.21)\n               > 40%                                                                                             [2]\n\n                                                                                                  Page 4 of 12\n\f  IAI                                                                                                 CS1A-0524\n\n               As this p-value is much higher than 5%, we do not have strong evidence to reject the null\n               hypothesis and hence conclude that ρ = 0. So, dating behaviour of teenagers and divorcee\n               adults is uncorrelated.",
      "has_math": true,
      "session": "2024-05",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2024-05_QP.pdf",
      "source_sol": "raw/CS1A_2024-05_SOL.pdf"
    },
    {
      "q_num": 16,
      "marks": 2,
      "topic": "regression_glm",
      "subtopics": [],
      "stem": "How many minimum pairs of (Xi, Y\n       on X: Y = α + β × X?\n\n        A. 2\n        B. 3\n        C. 10\n        D. 1                                                                                                [2]",
      "parts": [],
      "solution": "Correct Answer is Option A.\n\n               Minimum of two pairs of (Xi, 𝑌̂i, ei) would be required to derive the fitted regression\n               equation of Y on X. This can be done by solving the two simultaneous equations in X and 𝑌̂\n               and thus the values of 𝛼̂ and 𝛽̂ can be found out.                                                   [2]",
      "has_math": true,
      "session": "2024-05",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2024-05_QP.pdf",
      "source_sol": "raw/CS1A_2024-05_SOL.pdf"
    },
    {
      "q_num": 17,
      "marks": 2,
      "topic": "regression_glm",
      "subtopics": [],
      "stem": "Based on the pairs A to F found in the printout, the regression equation for Y on X is –\n\n        A. Y = – 5 + 10X\n        B. Y = 5 + 10X\n        C. Y = 10 + 5X\n        D. Y = – 10 + 5X.                                                                                [2]",
      "parts": [],
      "solution": "Correct Answer is Option D.\n\n               Let us take any two pairs and solve the equations simultaneously to arrive at the Y on X\n               regression equation. Let us consider points A (10,40) and I (30, 140).\n\n               40 = α + 10β\n               140 = α + 30β\n\n               100 = 20 β\n               β=5\n\n               α = 40 – 10*5 = -10\n\n               So, fitted equation of Y on X is: Y = -10 + 5X.                                                      [2]",
      "has_math": false,
      "session": "2024-05",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2024-05_QP.pdf",
      "source_sol": "raw/CS1A_2024-05_SOL.pdf"
    },
    {
      "q_num": 18,
      "marks": 2,
      "topic": "regression_glm",
      "subtopics": [],
      "stem": "Using the plot depicted above, which one of the following options represents possible values\n       for ̅\n           X and ̅\n                 Y for the trial data set of 100 values?\n\n        A. ̅\n           X = 47.50 and ̅\n                         Y = 227.50\n           ̅ = 35.00 and Y\n        B. X             ̅ = 165.00\n        C. X = 30.00 and ̅\n           ̅             Y = 140.00\n        D. X = 45.00 and ̅\n           ̅             Y = 215.00\n\n        Hint: Regression line of Y on X passes through the point (𝑋̅, 𝑌̅).                               [2]",
      "parts": [],
      "solution": "Correct Answer is Option D.\n\n               Since the regression line of Y on X passes through the point (𝑋̅, 𝑌̅), among these 10 points\n               A to J, only that point which passes through the regression line can be a possible candidate\n               for the mean of X and Y. Only Point H passes through the regression line as seen in the plot\n               and hence it is possible that 𝑋̅ = 45.00 and 𝑌̅ = 215.00.                                            [2]",
      "has_math": true,
      "session": "2024-05",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2024-05_QP.pdf",
      "source_sol": "raw/CS1A_2024-05_SOL.pdf"
    },
    {
      "q_num": 19,
      "marks": 2,
      "topic": "regression_glm",
      "subtopics": [],
      "stem": "What will be the value of Adjusted R2 for this bivariate regression model?\n\n        A. 66.25%\n        B. 70.00%\n        C. 69.69%\n        D. 70.14%                                                                                        [2]",
      "parts": [],
      "solution": "Correct Answer is Option C.\n\n               Adjusted R2\n               = 1 – (n – 1) / (n – k – 1) * (1 – R2)\n               = 1 – (100 – 1) / (100 – 1 – 1) * (1 – 0.70)\n               = 0.696939\n               = 69.69%                                                                                             [2]",
      "has_math": false,
      "session": "2024-05",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2024-05_QP.pdf",
      "source_sol": "raw/CS1A_2024-05_SOL.pdf"
    },
    {
      "q_num": 20,
      "marks": 2,
      "topic": "regression_glm",
      "subtopics": [],
      "stem": "You decide to check whether the parameters are significant (i.e. non-zero) using one way\n       ANOVA (analysis of variance).\n\n        If the value of ANOVA F-statistic is calculated to be 229, what is your conclusion at 1% level\n        of significance?\n\n        A. β = 0\n        B. β ≠ 0\n        C. α ≠ 0\n        D. α = 0                                                                                         [2]",
      "parts": [],
      "solution": "Correct Answer is Option B.\n\n               We need to compare the observed value of F-statistic with F1,98 at 1% level of significance.\n               From tables, F1,60 = 7.077 and F1,120 = 6.851. So, value of F1,98 lies between 6.851 and 7.077.\n               Since, the observed value 229 is much higher than F1,98 we have sufficient evidence to reject\n               the null hypothesis (i.e. β = 0). Hence, we conclude that β ≠ 0.                                     [2]",
      "has_math": true,
      "session": "2024-05",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2024-05_QP.pdf",
      "source_sol": "raw/CS1A_2024-05_SOL.pdf"
    },
    {
      "q_num": 21,
      "marks": 10,
      "topic": "distributions",
      "subtopics": [],
      "stem": "A pair of two biased two-sided coins is tossed once. The random variables representing the\n       outcome on the first and the second coin are denoted by X and Y respectively. X and Y take\n       numerical values of 1 and 0 if the outcome is Heads (H) and Tails (T) respectively.\n\n        These two random variables X and Y, have the following joint probability function:\n\n                                                     X (Outcome on first coin)\n                                               Heads (H)                   Tails (T)\n                         Heads (H)     P (X = “H”, Y= “H”) = 2/9  P (X = “T”, Y= “H”) = 4/9\n          Y (Outcome\n           on second\n                          Tails (T)     P (X = “H”, Y= “T”) = 1/9       P (X = “T”, Y= “T”) = 2/9\n             coin)",
      "parts": [
        {
          "label": "i",
          "marks": 1,
          "text": "Determine conditional expectation: E(Y | X = “H”).                                          (1)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 2,
          "text": "Determine conditional variance: Var(Y | X = “H”).                                              (2)",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 2,
          "text": "State the probability functions of the marginal distributions of X and Y.                      (2)",
          "topic": null
        },
        {
          "label": "iv",
          "marks": 1,
          "text": "Determine E(Y) and Var(Y) using the marginal distribution of Y as determined in part\n             (iii).                                                                                          (1)",
          "topic": null
        },
        {
          "label": "v",
          "marks": 1,
          "text": "Based on the marginal distributions of X and Y as determined in part (iii), determine\n             which coin is biased in favour of Heads (H) and which coin is biased in favour of Tails\n             (T)?                                                                                            (1)",
          "topic": null
        },
        {
          "label": "vi",
          "marks": 3,
          "text": "Check whether the random variables X and Y are independent of each other and suitably\n             comment on your findings in parts (i), (ii) and (iv).                                           (3)",
          "topic": null
        }
      ],
      "solution": "i) E(Y | X = “H”)\n                  = 1 * (2/9)/(3/9) + 0 * (1/9)/(3/9)\n                  = 2/3                                                                                             (1)\n\n               ii) E2(Y | X = “H”)\n                   = 12 * (2/9)/(3/9) + 02 * (1/9)/(3/9)                                                            (1)\n\n                                                                                                     Page 5 of 12\n\fIAI                                                                                        CS1A-0524\n\n           = 2/3\n\n          Var(Y | X = “H”)\n          = E2(Y | X = “H”) – [E(Y | X = “H”)]2\n          = 2/3 – (2/3)2\n          =6/9 – 4/9\n          = 2/9                                                                                           (1)\n\n      iii) Marginal distribution of X:\n\n       P(X = 1) = 2/9 + 1/9 = 1/3\n       P(X = 0) = 4/9 + 2/9 = 2/3                                                                         (1)\n\n       Marginal distribution of Y:\n\n       P(Y = 1) = 2/9 + 4/9 = 2/3\n       P(Y = 0) = 1/9 + 2/9 = 1/3                                                                         (1)\n\n       iv)\n       E(Y) = 1 * 2/3 + 0 * 1/3 = 2/3                                                                (0.5)\n\n       E2(Y) = 12 * 2/3 + 02 * 1/3 = 2/3\n\n       Var(Y)\n       = E2(Y) – [E(Y)]2\n       = 2/3 – (2/3)2\n       = 2/9                                                                                         (0.5)\n\n      v)\n      As P(X = 0) > P(X = 1) (i.e. 2/3 > 1/3), first coin is biased in the favour of Tails (T).\n      As P(Y = 1) > P(Y = 0) (i.e. 2/3 > 1/3), second coin is biased in the favour of Head (H)            (1)\n\n      vi)\n       P(X = 1) * P(Y = 1) = 1/3 * 2/3 = 2/9 = P(X = 1,Y = 1) = P(X = “H”, Y = “H”)\n       P(X = 0) * P(Y = 1) = 2/3 * 2/3 = 4/9 = P(X = 0,Y = 1) = P(X = “T”, Y = “H”)\n       P(X = 1) * P(Y = 0) = 1/3 * 1/3 = 1/9 = P(X = 1, Y = 0) = P(X = “H”, Y = “T”)\n       P(X = 0) * P(Y = 0) = 1/3 * 2/3 = 2/9 = P(X = 0, X = 0) = P(X = “T”, Y = “T”)\n\n       Hence random variables X and Y can be considered independent of each other.                        (2)\n\n       As X and Y are independent, the conditional expectation of Y will be equal to the\n       unconditional expectation of Y and the conditional variance of Y will be equal to the\n       unconditional variance of Y.\n\n       E(Y | X = “H”) = E(Y)\n       Var(Y | X = “H”) = Var(Y)\n\n       From parts (i), (ii) and (iv), it can be validated that:\n       E(Y | X = “H”) = E(Y) = 2/3\n       Var(Y | X = “H”) = Var(Y) = 2/9                                                                    (1)\n\n                                                                                          Page 6 of 12\n\f  IAI                                                                                                    CS1A-0524",
      "has_math": false,
      "session": "2024-05",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2024-05_QP.pdf",
      "source_sol": "raw/CS1A_2024-05_SOL.pdf"
    },
    {
      "q_num": 22,
      "marks": 10,
      "topic": "distributions",
      "subtopics": [],
      "stem": "Total processing time for a unit of a catalytic converter for a car is normally distributed with a\n       mean of 20 minutes and a standard deviation of 5 minutes.\n\n         Let us consider a random sample of 5 units of the catalytic converter which comprises of X1,\n         X2, X3, X4 and X5 which are independent and identically distributed random variables from the\n         normal distribution as mentioned above.\n\n                                                                                    ̅) for these 5 units",
      "parts": [
        {
          "label": "i",
          "marks": 2,
          "text": "Determine the probability that the sample mean of the processing time (X\n             is less than 15 minutes.                                                                        (2)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 2,
          "text": "Determine the probability that the sample standard deviation of the processing time (S)\n              for these 5 units is greater than 6.65 minutes.                                                (2)",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 2,
          "text": "Determine the probability that both sample mean of processing time (X  ̅) is less than 15\n              minutes and sample standard deviation of processing time (S) is greater than 6.65 minutes\n              for this sample of 5 units of the catalytic converter.                                         (2)\n\n         Random variable [(n – 1) * S2 / σ2] is distributed as a chi-square variable with (n – 1) degrees\n         of freedom.\n\n         Mean of this random variable is (n – 1) and its variance is 2 * (n – 1).",
          "topic": null
        },
        {
          "label": "iv",
          "marks": 4,
          "text": "In light of the above information, calculate the value of E(S2) and Var(S2) for the random\n             sample of 5 units of the catalytic convertor. If the sample size is increased to 100 units\n             instead of 5 units, what will be the impact on the value of E(S2) and Var(S2)? Briefly\n             comment on the same.                                                                            (4)",
          "topic": null
        }
      ],
      "solution": "i)   X_bar ~ N(20, 5^2 / 5)\n\n                P(X_bar < 15)\n                = P(Z < (15 – 20) / (5 / sqrt(5))\n                = P(Z < -2.236)\n                = 1 – P(Z > 2.236)\n                = 0.0127                                                                                               (2)\n\n               ii) For a normal population,\n\n                (n-1) * S^2 / sigma^2 ~ chi_square with (n-1) degrees of freedom.\n\n                P(S > 6.65)\n                = P(S^2 > 6.65^2)\n                = P(4 * S^2 / 5^2 > (4 * 6.65^2) / 25)\n                = P(chi_square(4) > 7.0756)\n                = 0.1320                                                                                               (2)\n\n               iii) S^2 and X_bar are independent if we are sampling from a normal distribution.                       (1)\n\n                P(X_bar < 15 and S > 6.65)\n                = P(X_bar < 15 and S^2 > 6.65^2)\n                = 0.0127 * 0.1320\n                = 0.001676                                                                                             (1)\n\n               iv) E( (n – 1) *S^2 / sigma^2) ) = (n – 1)\n\n                E(S^2) = (n – 1) / (n – 1) * sigma^2 = sigma^2 = 25                                                    (1)\n\n                Var ((n – 1) *S^2 / sigma^2) ) = 2 (n – 1)\n\n                Var (S^2)\n                = 2 (n – 1) / (n – 1)^2 * sigma^4\n                = 2 * sigma^4 / (n – 1)\n                = 2 * 625 / 4\n                = 312.50                                                                                               (2)\n\n                If the value of n is increased to 100 instead of 5, E(S^2) will remain unchanged at 25, but\n                Var(S^2) will be reduced from 312.50 to 12.6263 (1250/99). This means that for a higher\n                sample size i.e. as n tends to infinity, the sample variance tends to be closer and closer to\n                the population variance (sigma^2) as the variance of the sample variance i.e. Var(S^2)\n                becomes smaller and smaller.                                                                           (1)",
      "has_math": true,
      "session": "2024-05",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2024-05_QP.pdf",
      "source_sol": "raw/CS1A_2024-05_SOL.pdf"
    },
    {
      "q_num": 23,
      "marks": 10,
      "topic": "regression_glm",
      "subtopics": [],
      "stem": "You are working in a firm of climate risk actuaries and your firm has developed an Actuarial\n       Climate Risk Index (ACRI) for your country which is determined based on five key elements\n       viz. temperature, high precipitation, drought, strong winds and sea level.\n\n         A high value of the index is associated with greater climate risk for the country.\n\n         Over the past few years, an increase in the value of the index has been consistently observed.\n         Every year the value of this index is increasing by one as compared to the previous year. You\n         have been given the responsibility to check whether there is any association between ACRI and\n         climate related deaths in the country.\n\n        The table below gives year-wise details of ACRI represented by x and the number of deaths\n        due to climate change represented by n. The population exposed to climate change related\n        events, denoted E, have also been given. The values of death rates y, where y = n / E, and the\n        log (death rates), denoted w, are also given.\n\n          Year                  ACRI (x)       Number of deaths     Exposure                        y=n/E            w = log(y)\n                                                    (n)               (E)\n          2016                       70             30                426                            0.07042          ̶ 2.6532\n          2017                       71             38                471                            0.08068           ̶ 2.5173\n          2018                       72             38                454                            0.08370            ̶ 2.4805\n          2019                       73             53                482                            0.10996             ̶ 2.2077\n          2020                       74             59                445                            0.13258              ̶ 2.0205\n          2021                       75             61                423                            0.14421               ̶ 1.9365\n          2022                       76             82                468                            0.17521                ̶ 1.7417\n          2023                       77             96                430                            0.22326                 ̶ 1.4994\n\n                           ∑ x = 588; ∑ x 2 = 43260; ∑ w = ̶ 17.0568; ∑ w 2 = 37.5173; ∑ xw = ̶ 1246.7879\n\n        A regression line is to be fit with death rates as the response variable and ACRI as the\n        explanatory variable. However, your teammates have suggested that logarithm of the death\n        rates i.e. w should be used instead of y.\n\n        Following scatter plots have been obtained for the data.\n\n                                     SCATTER PLOT                                                SCATTER PLOT\n                                        (Y ON X)                                                   (W ON X)\n                                                                                                              ACRI\n                        0.250                                                               -\n                                                                     LOG (DEATH RATE)\n           DEATH RATE\n\n                        0.200                                                           -0.500 68   70   72     74     76         78\n                        0.150                                                           -1.000\n                        0.100                                                           -1.500\n                        0.050                                                           -2.000\n                           -                                                            -2.500\n                                68        70   72    74   76   78\n                                                                                        -3.000\n                                                 ACRI",
      "parts": [
        {
          "label": "i",
          "marks": 2,
          "text": "Based on the scatter plots as depicted above, briefly explain why a logarithmic\n             transformation could be necessary in this case.                                                                                   (2)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 6,
          "text": "Calculate the least squares fit regression line in which w is modelled as response variable\n             and x is modelled as explanatory variable.                                                                                        (6)",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 2,
          "text": "Calculate the fitted value for the number of deaths in the year 2018.                                                            (2)",
          "topic": null
        }
      ],
      "solution": "i) From the scatter plots it can be observed that –\n                   • the relationship between Y and X is non-linear (shape of a curve)\n                   • the relationship between W=log(Y) and X seems to be linear\n\n                Hence, logarithmic transformation is justified in this case in order to fit a linear regression\n                model to the data.                                                                                     (2)\n               ii)\n                S(xw)                                                                                                  (1)\n                                                                                                       Page 7 of 12\n\f  IAI                                                                                   CS1A-0524\n\n               = Sum(xw) – (Sum(x) * Sum(w))/n\n               = -1246.7879 – (588 * - 17.0568)/8\n               = 6.8869\n\n               S(xx)\n               = Sum(x^2) – ((Sum(x))^2) / n\n               = 43260 – ((588)^2) / 8\n               = 42                                                                                   (1)\n\n               beta_hat\n               = S(xw) / S(xx)\n               = 6.8869 / 42\n               = 0.1640                                                                               (1)\n\n               x_bar = 588/8 = 73.50\n               w_bar = -17.0568/8 = -2.1321                                                           (1)\n\n               The regression line of W on X passes through the point (x_bar, w_bar)\n\n               alpha_hat\n               = w_bar – beta_hat * x_bar\n               = -2.1321 – 0.1640 * 73.50\n               = -14.1842                                                                             (1)\n\n               Least squares fit regression line of W on X is –\n               w_hat = -14.1842 + 0.1640 * x                                                          (1)\n\n               iii) For the year 2018,\n\n               w_hat\n               = -14.1842 + 0.1640 * 72\n               = -2.3762                                                                              (1)\n\n               w = log(y)\n               y_hat\n               = exp(w_hat)\n               = exp (-2.3762)\n\n               y_hat\n               = 0.0929                                                                           (0.5)\n\n               n_hat\n               = E * y_hat\n               = 454 * 0.0929\n               = 42.1766                                                                          (0.5)",
      "has_math": true,
      "session": "2024-05",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2024-05_QP.pdf",
      "source_sol": "raw/CS1A_2024-05_SOL.pdf"
    },
    {
      "q_num": 24,
      "marks": 10,
      "topic": "bayes_credibility",
      "subtopics": [
        "distributions"
      ],
      "stem": "Let X be the performance rating that an employee gets in Math Solutions Ltd. X is modelled\n       using normal distribution i.e. X ~ N (μ, σ2) where μ is unknown. The parameter μ follows a\n       uniform distribution over the interval (0, 10). Data for 150 employees who were part of the\n       current year’s performance appraisal exercise is collected and the average rating for these 150\n       employees is found out to be 5.\n\n      We want to find the posterior density of μ for which following workings have been done.\n\n      Step1: Prior density of μ\n      f(μ) = 1 / (A – B) = 1/10 = constant.\n\n      Step 2: Likelihood function of X ignoring any coefficient of proportionality:\n                                    −1\n                                   {        ∑(x− D)2 }\n      L(x, μ) = constant * e 2C ×                        .\n\n      Step 3: Arriving at the posterior density of μ:\n      f(μ | x)\n      = f(μ) * L(x, μ)\n                       −1\n                       {        ∑(x− D)2 }\n      = constant * e 2C ×\n                       −1          2\n                       {                    ∑ x+nμ2 )}\n      = constant * e 2C ×(∑ x −2μ\n                       −1\n                       {               ∑ x+E)}\n      = constant * e 2C ×(−2μ\n                            −1\n                       {       × (−2μx̅+μ2 )}\n                           2σ2\n      = constant * e        F\n\n                            −1\n                       {       × (μ2 −2μx̅+G)}\n                           2σ2\n      = constant * e        F\n\n                           −1\n                       {        (μ −I)2 }\n      = constant * e 2H ×\n\n      Hence, the posterior density of μ is as follows:\n\n      μ ~ N(5, σ2 / J).",
      "parts": [
        {
          "label": "i",
          "marks": 5,
          "text": "Determine the values of the missing terms A, B, C, D, E, F, G, H, I and J. You are NOT\n           required to provide any justification / supporting calculations.                                (5)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 2,
          "text": "Determine the Bayesian estimate for μ under –\n\n           a) Squared error loss;\n           b) All-or-nothing loss;\n           c) Absolute error loss.                                                                         (2)",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 2,
          "text": "Assuming σ2 to be 25, determine a 95% equal-tailed credible interval for μ.\n\n           Hint: An equal-tailed credible interval is a confidence interval determined using the\n           posterior distribution of μ.                                                                    (2)\n\n      An alternative to equal-tailed credible interval is highest posterior density interval for μ. This\n      interval is such that the minimum density of any point within this interval is equal to or higher\n      than the density outside this interval.",
          "topic": null
        },
        {
          "label": "iv",
          "marks": 1,
          "text": "Without performing any additional calculations, state the 95% highest posterior density\n          interval for μ. Justify your answer.                                                             (1)",
          "topic": null
        }
      ],
      "solution": "i)\n               Step1: Prior density of μ\n               f(μ) = 1 / (10 – 0) = 1/10 = constant.     So, A = 10 and B = 0\n\n                                                                                       Page 8 of 12\n\f  IAI                                                                                                    CS1A-0524\n\n               Step 2: Likelihood function of X ignoring any coefficient of proportionality:\n                                               −1                2\n                                           {        × ∑(x− 𝐦𝐮) }\n               L(x, μ) = constant * e 2𝐬𝐢𝐠𝐦𝐚^𝟐                       . So, C = sigma^2 and D = mu\n\n               Step 3: Arriving at the posterior density of μ:\n               f(μ | x)\n               = f(μ) * L(x, μ)\n                                      −1                 2\n                                {          × ∑(x− 𝐦𝐮) }\n               = constant * e 2𝐬𝐢𝐠𝐦𝐚^𝟐\n                                      −1\n                                {          ×(∑ x2 −2μ ∑ x+nμ2 )}\n               = constant * e 2𝐬𝐢𝐠𝐦𝐚^𝟐\n                                      −1\n                                {          ×(−2μ ∑ x+𝐧∗𝐦𝐮^𝟐)}\n               = constant * e 2𝐬𝐢𝐠𝐦𝐚^𝟐                               So, E = n * mu^2\n                                     −1\n                                {       × (−2μx̅+μ2 )}\n                                    2σ2\n               = constant * e        𝒏                               So, F = n\n                                     −1\n                                {       × (μ2 −2μx̅+𝐱_𝐛𝐚𝐫^𝟐)}\n                                    2σ2\n               = constant * e        𝒏                               So G = x_bar^2\n                                      −1\n                                {              × (μ −𝐱_𝐛𝐚𝐫)2 }\n               = constant * e   2(𝐬𝐢𝐠𝐦𝐚^𝟐 / 𝐧)                       So H = sigma^2 / n and I = x_bar\n\n               Hence, the posterior density of μ is as follows:\n\n               μ ~ N(5, σ2 / n).                                     So J = n = 150\n\n               ii)     For a normal distribution, mean = median = mode.\n                       For mu, mean = median = mode = 5                                                                (1)\n\n                     So, Bayesian Estimate under –\n                         (a) Squared error loss (mean) = 5\n                         (b) All-or-nothing loss (mode) = 5\n                         (c) Absolute error loss (median) = 5                                                          (1)\n\n               iii)     From tables we know that P(-1.96 < Z < 1.96) = 0.95\n\n                        95% equal tailed credible interval for mu\n                        = (5 – 1.96 * sqrt(25/150), 5 + 1.96 * sqrt(25/150))\n                        = (4.1998, 5.8002)                                                                             (2)\n\n                iv)     For symmetrical distributions like normal distribution, equal tailed credible\n                        interval is equal to highest posterior density interval as the minimum density of\n                        any point within this interval is higher than any density outside this interval.\n\n                        So, highest posterior density interval is (4.1998, 5.8002)                                  (1)",
      "has_math": true,
      "session": "2024-05",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2024-05_QP.pdf",
      "source_sol": "raw/CS1A_2024-05_SOL.pdf"
    },
    {
      "q_num": 25,
      "marks": 10,
      "topic": "inference",
      "subtopics": [],
      "stem": "Recently the results of December 2023 Diet Examinations conducted by the Institute of\n       Actuaries of Statistica (IAS) were declared. Following news was published in one of the local\n       newspaper:\n\n        “Candidates from Region X outperform candidates from Region Y in the December 2023\n        actuarial examinations with pass rates of 52% and 40% respectively.”\n\n        Your manager, Mrs. Numara is a reputed actuary in the industry and she belongs to Region Y.\n        She was not convinced with such a significant difference in pass rates for Region X and Region\n        Y. She has asked you to contact the officials of IAS and statistically test whether candidates\n        from Region X are smarter than candidates from Region Y when it comes to passing actuarial\n        examinations.\n\n        You reached out to officials of IAS and learned that 2050 candidates from Region X and 800\n        candidates from Region Y appeared in the examinations, with reported pass rates of 52% and\n        40% respectively, as reported in the newspaper. You have opted to use a non-parametric\n        method for conducting the statistical analysis.",
      "parts": [
        {
          "label": "i",
          "marks": 1,
          "text": "Name any two non-parametric approaches to hypothesis testing.                                    (1)\n\n        You decided to use contingency tables for performing this statistical test.",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 1,
          "text": "State the null hypothesis (H0) and alternate hypothesis (H1) for this test.                      (1)",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 1,
          "text": "Complete the observed frequencies table based on the data collected from the officials of\n              IAS. Write the values of XPO, XFO, YPO and YFO in the answer script. You are NOT\n              required to copy the table and NOT required to show any supporting calculations.\n                                              Observed Frequencies\n                                         Region X          Region Y                        Total\n               Pass                        XPO                YPO\n               Fail                        XFO                YFO\n               Total                                                                                          (1)\n\n        In the expected frequencies table corresponding to the above observed frequencies, let us define\n        four quantities viz. XPE, XFE, YPE and YFE representing candidates from Region X expected\n        to pass, candidates from Region X expected to fail, candidates from Region Y expected to pass\n        and candidates from Region Y expected to fail respectively.",
          "topic": null
        },
        {
          "label": "iv",
          "marks": 2,
          "text": "Write the values of XPE, XFE, YPE and YFE in the answer script (rounded off to the\n               nearest integer). You are NOT required to show any supporting calculations.                    (2)",
          "topic": null
        },
        {
          "label": "v",
          "marks": 3,
          "text": "Calculate the value of the χ2 test statistic. Based on the value of the test statistic, show\n               that the above test in fact strengthens the claim that candidates from Region X are smarter\n               than candidates from Region Y when it comes to passing actuarial examinations. Test at\n               0.5% level of significance.                                                                    (3)\n\n        Your manager is still not convinced with the results of the statistical test and makes a request\n        to the Examinations Controller of IAS under the Right to Information (RTI) Act asking for\n        detailed subject-wise and region-wise results data for December 2023 Diet Examinations.\n\n        Following data is obtained through RTI.\n\n                                       Region X                                Region Y\n               Subject    Appeared      Passed       Pass %        Appeared     Passed      Pass %\n                CS1          500         250          50%            100          60         60%\n                CM1          600         300          50%             50          40         80%\n                CB1          400         250          63%            100          90         90%\n                CB2          450         250          56%             50          35         70%\n                CM2           50          10          20%            200          50         25%\n                CS2           50           6          12%            300          45         15%\n                Total       2050         1066         52%            800         320         40%\n\n        Based on the above data, your manager has criticized the conclusion you presented in part (v)\n        of your analysis, claiming that your statistical test led to an incorrect conclusion. She argues\n        that across all subjects, the pass rate for Region Y consistently exceeds that of Region X based\n        on the data provided.",
          "topic": null
        },
        {
          "label": "vi",
          "marks": 2,
          "text": "Defend your test results obtained in part (v) clearly addressing the observation made by\n                Mrs. Numara. You are NOT required to perform any additional calculations.                    (2)",
          "topic": null
        }
      ],
      "solution": "i)     Examples of non-parametric approaches to hypothesis testing (any two names are\n                      sufficient):\n                     • Permutations approach\n                     • Chi-square goodness of fit tests\n                     • Contingency tables\n                     • Fisher’s Exact Test                                                                         (1)\n\n               ii)                                                                                                 (1)\n                                                                                                        Page 9 of 12\n\fIAI                                                                                           CS1A-0524\n\n              H0: Passing actuarial examinations is independent of region\n              H1: Passing actuarial examinations is not independent of region\n\n       iii)\n              XP(o) = Candidates from Region X who passed = 2050 * 52% = 1066\n              XF(o) = Candidates from Region X who failed = 2050 – 1066 = 984\n              YP(o) = Candidates from Region Y who passed = 800 * 40% = 320\n              YF(o) = Candidates from Region Y who failed = 800 – 320 = 480.                            (1)\n\n      iv)\n       Total passed candidates (P) = 1066+320 = 1386\n       Total failed candidates (F) = 984 + 480 = 1464\n       Total Candidates from Region X (B) = 2050\n       Total Candidates from Region Y (G) = 800\n       Total Candidates (T) = 2850\n\n       XP(e) = Candidates from Region X expected to pass = 2050 * 1386 / 2850 = 996.9474 =\n       997\n       XF(e) = Candidates from Region X expected to fail = 2050 * 1464 / 2850 = 1053.053 =\n       1053\n       YP(e) = Candidates from Region Y expected to pass = 800 * 1386 / 2850 = 389.0526 =\n       389\n       YF(e) = Candidates from Region Y expected to fail = 800 * 1464 / 2850 = 410.9474 =\n       411                                                                                              (2)\n\n      v)\n       Test statistic for chi-square test\n       = (1066 – 997)^2 / 997 + (984 – 1053)^2 / 1053 + (320 – 389)^2 / 389 + (480 – 411)^2 /\n       411\n       = 33.1197                                                                                        (1)\n\n       Number of degrees of freedom\n       = (2-1) * (2-1) = 1\n\n       We are carrying out a one-sided test. The upper 0.5% point of a chi_square distribution\n       with 1 degree of freedom is 7.879.                                                               (1)\n\n       As the observed value of the test statistic (33.1197) is in excess of 7.879, we have\n       sufficient evidence to reject the null hypothesis at 0.5% level of significance. Therefore,\n       it is reasonable to conclude that passing actuarial examinations is dependent on region to\n       which the candidate belongs (candidates from Region X seem to be better at passing these\n       examinations as compared to candidates from Region Y)                                            (1)\n\n       vi) The observation raised by Mrs. Numara is valid as it is clear from the table that in\n           every subject the pass percentage of Region Y is higher than that of Region X.\n\n       However, in simpler subjects like CS1, CM1, CB1, CB2 where the average pass\n       percentage is higher, the number of candidates from Region Y appearing for these papers\n       is relatively very low. Whereas for difficult subjects like CM2 and CS2 where the average\n       pass percentage is very low, the number of candidates from Region Y appearing for these\n       papers is relatively very high.\n\n       For Region X, it is exactly the opposite. Candidates from Region X seem to have attempted\n       simpler papers in large numbers and their participation in difficult papers is relatively low.   (2)\n\n                                                                                           Page 10 of 12\n\f IAI                                                                                                CS1A-0524\n\n               So, in case of Region X despite of an average performance across papers, due to higher\n               number of candidates attempting simpler papers, the overall pass percentage of Region X\n               tends to be higher.\n\n               On the contrary in case of Region Y, despite of performing above average across all\n               papers, due to lower number of candidates attempting simpler papers, the overall pass\n               percentage of Region Y tends to be lower.\n\n               However, there is no error in the test performed using contingency tables. With more\n               insights into the underlying data, a better picture has emerged.",
      "has_math": true,
      "session": "2024-05",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2024-05_QP.pdf",
      "source_sol": "raw/CS1A_2024-05_SOL.pdf"
    },
    {
      "q_num": 26,
      "marks": 10,
      "topic": "regression_glm",
      "subtopics": [],
      "stem": "Railway Ticketing Corporation of Actuaria (RTCA) is an online ticketing platform for booking\n       tickets for railway travels all throughout the country. It has recently started displaying\n       probability of confirmation (pc) in case of waiting list tickets.\n\n        Past data is available for the following explanatory variables:\n         Wknd : a categorical variable with value = 1 in case if it is a weekend and value = 0 in\n                case if it is a weekday\n         Kms : a numerical variable which captures the distance between the boarding station\n                and the destination in kms\n\n        This probability pc is being determined by using a binomial generalized linear model (GLM)\n        with the canonical link function. The linear predictor has the form:\n                                           g(pc) = α + βw i=0,1 + βk * kms\n\n                 where: βw i=0 is used in case of a weekday and βw i=1 is used in case of a weekend\n\n        The analysis of the past data gave the following estimates for the model:\n\n               Coefficient             Estimate           Standard Error\n                   α                    - 0.372                0.053\n                 βw i=0                   0.036                0.023\n                 βw i=1                 - 0.100                0.080\n                   βk                   - 0.003                0.001",
      "parts": [
        {
          "label": "i",
          "marks": 2,
          "text": "Determine whether the coefficients βw i=0, βw i=1 and βk are significant by comparing the\n             estimate of these coefficients with their respective standard errors. You are NOT required\n             to calculate p-values.                                                                          (2)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 2,
          "text": "Show that the probability pc that a waiting list ticket will be confirmed for a distance of\n           300 kms when the travel is on a weekend is 0.20. Use the canonical link function relevant\n           to a binomial model.                                                                           (3)\n\n      Mr. Traveller travels between two cities viz. “Thousand Palms” and “Forty Miles” every\n      weekend. Distance between these two cities is 300 kms. He books the tickets invariably late in\n      the evenings on Friday and always features in the waiting list. In case if his ticket is not\n      confirmed then he needs to travel by a private bus.\n\n      iii) Calculate the probability that Mr. Traveller travels at least 2 times by train during the\n           coming month. Use the estimate of pc from part (ii). Assume that a month has 4 weeks.          (3)\n\n      iv) State how the linear predictor as given above will change if an interaction term between\n          the covariates (wknd * kms) is also included in the model.                                      (2)",
          "topic": null
        }
      ],
      "solution": "i) For checking whether a parameter is significant (i.e. significantly different from zero),\n                  as a general rule we use –\n\n               mod(beta) > 2 * se(beta)                                                                       (0.5)\n\n               mod(beta_w_i=0) = 0.036\n               2 * se(beta_w_i=0) = 0.046 > 0.036\n               beta_w_i=0 is not significant                                                                  (0.5)\n\n               mod(beta_w_i=1) = 0.100\n               2 * se(beta_w_i=1) = 0.160 > 0.100\n               beta_w_i=1 is not significant                                                                  (0.5)\n\n               mod(beta_k) = 0.003\n               2 * se(beta_k) = 0.002 < 0.003\n               beta_k is significant                                                                          (0.5)\n               ii) In case of a weekend travel over a distance of 300 kms\n               g(pc)\n               = α + βw i=1 + βk * kms\n               = -0.372 – 0.100 – 0.003*300\n               = -1.372                                                                                        (1)\n\n               From Tables, canonical link function for a binomial model is given by:\n               g(pc) = log (pc / (1-pc))\n               -1.372 = log (pc / (1-pc))\n               exp(-1.372) = pc / (1-pc)\n               0.2536 = pc / (1-pc)\n               0.2536 = (1+0.2536) pc\n               pc = 0.20 as required                                                                           (2)\n\n               iii)\n                We are using a binomial distribution with –\n                n=4\n\n               success = travelling by train (due to confirmation of railway ticket)\n               p = pc = 0.20\n\n               failure = travelling by private bus (due to non-confirmation of railway ticket)\n               q = 1 – p = 0.80                                                                                (1)\n\n                                                                                                 Page 11 of 12\n\fIAI                                                                                    CS1A-0524\n\n      Required probability\n      = P(X>=2)\n      = 1 – P(X<= 1)\n      = 1 – P(X=1) – P(X=0)\n      = 1 – 0.4096 – 0.4096\n      = 0.1808                                                                                   (2)\n\n      iv)   If we add an interaction term (wknd * kms), the revised linear predictor would be\n            as follows:\n\n            g(pc) = alpha + beta_w_i=0,1 + beta_k * kms + beta_wk_i=0,1 * kms                    (1)\n\n      where:\n         1. beta_w_i=0 is used in case of a weekday and beta_w_i=1 is used in case of a\n             weekend\n         2. beta_wk_i=0 is used in case of a weekday and beta_wk_i=1 is used in case of a\n             weekend                                                                             (1)\n\n                                     *************\n\n                                                                                    Page 12 of 12",
      "has_math": true,
      "session": "2024-05",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2024-05_QP.pdf",
      "source_sol": "raw/CS1A_2024-05_SOL.pdf"
    },
    {
      "q_num": 1,
      "marks": 2,
      "topic": "regression_glm",
      "subtopics": [
        "distributions"
      ],
      "stem": "A multi-variate linear regression model has been fit which has the following equation:\n                                       Y = 8.928 + 0.993 X1 + 0.476 X2\n        Which of the following statements regarding the relationship between Y, X1 and X2 is\n        necessarily TRUE based on the above equation?\n\n        A. Y and X1 are negatively correlated.\n        B. Y and X2 are positively correlated.\n        C. Y and X2 are negatively correlated.\n        D. Y and X1 are un-correlated.                                                                        [2]",
      "parts": [],
      "solution": "Correct Answer is Option B.                                                                    [2]\n\n              We know that sample correlation coefficient = Sxy / sqrt (Sxx * Syy).\n\n              Beta coefficient in the regression line is given by Sxy / Sxx.\n\n              So, the sign (positive or negative) of both the correlation coefficient and beta coefficient\n              is dependent on the sign of Sxy. It follows that if beta coefficient in the regression line\n              is positive then correlation coefficient is also positive.\n\n              Looking at the each of the statements –\n\n                  A. β1 is positive, so Y and X1 cannot be negatively correlated.\n                  B. β2 is positive, so Y and X2 can be positively correlated.\n                  C. β2 is positive, so Y and X2 cannot be negatively correlated.\n                  D. β1 is positive (non-zero), so Y and X1 cannot be un-correlated.\n\n              Considering the above, only Option B is true.",
      "has_math": false,
      "session": "2024-11",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2024-11_QP.pdf",
      "source_sol": "raw/CS1A_2024-11_SOL.pdf"
    },
    {
      "q_num": 2,
      "marks": 2,
      "topic": "inference",
      "subtopics": [],
      "stem": "It is generally believed that as we age, we tend to put on weight. This is proposed to be\n        statistically tested using a students’ t-distribution based on the data collected.\n\n        Sample correlation coefficient between X1 and X2 based on the sample of 7 volunteers using\n        Pearson’s Method i.e. r PEARSON is determined as – 0.1599543.\n\n        Select the correct null hypothesis (H0) and the correct value of the test statistic under the null\n        hypothesis from the following options:\n\n           A. H0: “X1 and X2 are un-correlated”;\n              t-statistic = - 0.35318\n           B. H0: “X1 and X2 are correlated”;\n              t-statistic = 0.10162\n           C. H0: “X1 and X2 are un-correlated”;\n              t-statistic = - 0.36233\n           D. H0: “X1 and X2 are correlated”;\n              t-statistic = 0.32164.                                                                          [2]",
      "parts": [],
      "solution": "Correct Answer is Option C.                                                                    [2]\n\n              In order to test the population correlation coefficient, null hypothesis and alternative\n              hypothesis are defined as under –\n              H0: ρ = 0 (in other words, X1 and X2 are uncorrelated).\n              H1: ρ ≠ 0 (in other words, X1 and X2 are not uncorrelated).\n\n              The value of the test statistic under the null hypothesis\n              = r * sqrt (n – 2) / sqrt (1 – r2)\n              = – 0.1599543 * sqrt (7 – 2) / sqrt (1 – (– 0.1599543)2)\n              = – 0.36233.",
      "has_math": false,
      "session": "2024-11",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2024-11_QP.pdf",
      "source_sol": "raw/CS1A_2024-11_SOL.pdf"
    },
    {
      "q_num": 3,
      "marks": 2,
      "topic": "inference",
      "subtopics": [],
      "stem": "At 5% level of significance, determine the critical value from the students’ t-distribution and\n        the result of the t test using the null hypothesis and the test statistic calculated in question 2.\n        (Kindly note that it is a two-sided test).\n\n           A. Critical value = 2.365;\n              Result: With age, we tend to put on weight\n           B. Critical value = 2.571;\n              Result: With age, we tend to lose weight\n           C. Critical value = 2.365;\n              Result: We may put on or lose weight irrespective of age\n           D. Critical value = 2.571;\n              Result: We may put on or lose weight irrespective of age.                                       [2]",
      "parts": [],
      "solution": "Correct Answer is Option D.                                                                    [2]\n\n              Under H0, the test statistic has a students’ t-distribution with n – 2 degrees of freedom.\n\n              For a t5 distribution, the critical value at 5% level of significance (considering a two-\n              sided hypothesis) is 2.571.\n\n              As the value of test statistic – 0.36233 < 2.571, we do not have sufficient evidence to\n              reject the null hypothesis. Hence, we conclude that X1 and X2 are un-correlated. In other\n              words, we may put on weight or lose weight irrespective of age.",
      "has_math": false,
      "session": "2024-11",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2024-11_QP.pdf",
      "source_sol": "raw/CS1A_2024-11_SOL.pdf"
    },
    {
      "q_num": 4,
      "marks": 2,
      "topic": "data_analysis",
      "subtopics": [],
      "stem": "One of the research scholars Mr. Anshul has commented that if Spearman’s Rank method is\n         used, then the sample correlation coefficient between weight and age of the practitioner will\n         be exactly equal to 0.\n\n         What is the implicit value of the sum of squares of deviations in ranks when r SPEARMAN = 0?\n\n             A. 56\n             B. 336\n             C. 48\n             D. 120.                                                                                         [2]",
      "parts": [],
      "solution": "Correct Answer is Option A.                                                                    [2]\n\n              r SPEARMAN\n              = 1 – (6 ∑𝑛1 𝑑2 ) / (n * (n2 – 1))\n\n              As r SPEARMAN = 0,\n              0 = 1 – (6∑71 𝑑2 ) / (7 * (49 – 1))\n              1 = (6∑71 𝑑2 ) / 336\n              6∑71 𝑑2 = 336         ∑71 𝑑 2 = 336 / 6 = 56.",
      "has_math": false,
      "session": "2024-11",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2024-11_QP.pdf",
      "source_sol": "raw/CS1A_2024-11_SOL.pdf"
    },
    {
      "q_num": 5,
      "marks": 2,
      "topic": "data_analysis",
      "subtopics": [],
      "stem": "Another research scholar Mrs. Eesha has suggested that Kendall’s method should be used as\n         it has better statistical properties as compared to Spearman’s method. However, Mr. Anshul\n         is claiming that this is not required as the value of the sample correlation coefficient under\n         Kendall’s method would also be equal to 0 i.e. r KENDALL = 0.\n\n         Which of the following statements is correct in relation to the claim made by Mr. Anshul?\n\n         A. As total number of rank pairs is an even number, number of concordant pairs and\n            discordant pairs can be equal leading to r KENDALL = 0.\n         B. As total number of rank pairs is an odd number, number of concordant pairs and discordant\n            pairs can be equal leading to r KENDALL = 0.\n         C. As total number of rank pairs is an odd number, number of concordant pairs and discordant\n            pairs cannot be equal leading to r KENDALL ≠ 0.\n         D. As total number of rank pairs is an even number, number of concordant pairs and\n            discordant pairs cannot be equal leading to r KENDALL ≠ 0.                                       [2]\n\n         Use the following information for attempting questions 6 to 10:\n\n         Let X represent the height of one-year old infants in inches. It is assumed to be normally\n         distributed.\n\n         Y= eX is lognormally distributed with parameters μ = 25 and σ2 = 36.\n\n         A paediatrician believes that generally infant height lies in the interval of mean height plus /\n         minus 4 inches.",
      "parts": [],
      "solution": "Correct Answer is Option C.                                                                    [2]\n\n                                                                                                  Page 2 of 13\n\fIAI                                                                                               CS1A-1124\n\n              r KENDALL\n              = (nc – nd) / (nc + nd)\n\n              (nc + nd) represents the total number of pairs.\n\n              It can be also calculated as n * (n – 1) / 2.\n\n              As n = 7, the total number of pairs\n              = 7 * (7 – 1) / 2\n              =7*3\n              = 21.\n\n              As 21 is an odd number, the number of concordant pairs (nc) and number of discordant\n              pairs (nd) cannot be equal. Hence, the value of the numerator (nc – nd) will be a non-\n              zero number as nc ≠ nd. Consequently, the value of r KENDALL cannot be equal to 0.",
      "has_math": true,
      "session": "2024-11",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2024-11_QP.pdf",
      "source_sol": "raw/CS1A_2024-11_SOL.pdf"
    },
    {
      "q_num": 6,
      "marks": 2,
      "topic": "distributions",
      "subtopics": [],
      "stem": "The claim of the paediatrician is proposed to be quantified statistically. Which of the following\n         expression represents this correctly?\n\n         A. P ( | X – 4 | < 25 )\n         B. P ( | X – 25 | < 4 )\n         C. P ( | Y – 4 | < 25 )\n         D. P ( | Y – 25 | < 4 ).                                                                            [2]",
      "parts": [],
      "solution": "Correct Answer is Option B.                                                                   [2]\n\n              The paediatrician believes that generally infant height lies in the interval of mean height\n              plus / minus 4 inches.\n\n              This can be mathematically represented as\n\n              P(Mean(X) – 4 < X < Mean(X) + 4)\n\n              X follows a normal distribution and Y = eX ~ logN(μ = 25, σ2 = 36).\n\n              Hence, X ~ N(μ = 25, σ2 = 36).\n\n              So, P(Mean(X) – 4 < X < Mean(X) + 4)\n              = P( (25 – 4) < X < (25 + 4))\n              = P(21 < X < 29).\n\n              Statements C and D can be outrightly discarded as we are finding probability in terms\n              of X and not in terms of Y.\n\n              Let us consider the other two statements –\n\n              Statement A –\n              P ( | X – 4 | < 25 ) = P ( - 25 < (X – 4) < 25) = P (-21 < X < 29)\n              This does not match with the required probability and hence this statement is incorrect.\n\n              Statement B –\n              P ( | X – 25 | < 4 ) = P ( - 4 < (X – 25) < 4) = P (21 < X < 29)\n              This matches with the required probability and hence this statement is the correct\n              answer.",
      "has_math": false,
      "session": "2024-11",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2024-11_QP.pdf",
      "source_sol": "raw/CS1A_2024-11_SOL.pdf"
    },
    {
      "q_num": 7,
      "marks": 2,
      "topic": "inference",
      "subtopics": [],
      "stem": "How much probable is the claim made by the paediatrician?\n\n         A. 49.52%\n         B. 50.00%\n         C. 74.76%\n         D. 25.24%.                                                                                          [2]",
      "parts": [],
      "solution": "Correct Answer is Option A.                                                                   [2]\n\n              Required probability\n              = P(21 < X < 29)\n              = P(X < 29) – P(X < 21)\n              = P(Z < (29-25)/6) – P (Z < (21-25)/6)\n              = P(Z < 0.67) – P(Z < -0.67)\n              = P(Z < 0.67) – (1 – P(Z < 0.67))\n                                                                                                 Page 3 of 13\n\fIAI                                                                                          CS1A-1124\n\n               = 0.4952\n               = 49.52%.",
      "has_math": false,
      "session": "2024-11",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2024-11_QP.pdf",
      "source_sol": "raw/CS1A_2024-11_SOL.pdf"
    },
    {
      "q_num": 8,
      "marks": 2,
      "topic": "inference",
      "subtopics": [],
      "stem": "Which of the following is necessarily TRUE about E[Y]?\n\n         A. E[Y] = E[X]\n         B. E[Y] < E[X]\n         C. E[Y] = e ( E[X] )\n         D. E[Y] = E [ e X ]                                                                             [2]",
      "parts": [],
      "solution": "Correct Answer is Option D.                                                            [2]\n\n               E(Y) = exp(25 + 0.5 * 36) = exp(43) = 4.73 * 1018\n               E(X) = μ = 25\n\n               Let us consider each of the given statements.\n\n               A. E[Y] = E[X]. This is not true as E[Y] ≠ E[X] as seen above.\n               B. E[Y] < E[X]. This is not true as E[Y] > E[X] as seen above.\n               C. E[Y] = exp ( E[X] ). This is not true as exp (25) = 7.2 * 1010 ≠ E[Y]\n               D. E[Y] = E [ exp(X) ]. This statement seems to be true. As Y = exp(X), E [ exp(X) ]\n                  = E[Y].",
      "has_math": false,
      "session": "2024-11",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2024-11_QP.pdf",
      "source_sol": "raw/CS1A_2024-11_SOL.pdf"
    },
    {
      "q_num": 9,
      "marks": 2,
      "topic": "distributions",
      "subtopics": [],
      "stem": "Select the correct expression which represents the moment generating function of the random\n         variable X i.e. MX(t) –\n\n         A. MX(t) = exp (½ t2)\n         B. MX(t) = exp (25t + 36t2)\n         C. MX(t) = exp (25t + 18t2)\n         D. MX(t) does not exist in closed form.                                                         [2]",
      "parts": [],
      "solution": "Correct Answer is Option C.                                                            [2]\n\n               For a normal distribution,\n               Mx(t)\n               = exp (μt + ½ σ2 * t2)\n               = exp (25t + ½ * 36 * t2)\n               = exp (25t + 18t2).",
      "has_math": false,
      "session": "2024-11",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2024-11_QP.pdf",
      "source_sol": "raw/CS1A_2024-11_SOL.pdf"
    },
    {
      "q_num": 10,
      "marks": 2,
      "topic": "distributions",
      "subtopics": [
        "regression_glm"
      ],
      "stem": "Which of the following represents simulated value from the normal distribution of X\n         corresponding to the value 0.40 from U (0, 1)?\n\n         [Hint: For simulating from a normal distribution, following steps are to be followed:\n         1. Generate random number from U (0,1).\n         2. If u > 0.5, using standard normal tables, find z such that P (Z <= z) = u. Simulated value\n            is z in this case.\n         3. If u < 0.5, find z such that P (Z <= z) = 1 – u. Simulated value is – z in this case.\n         4. Depending on the simulated value generated in 2 or 3, generate the simulated x value as\n            x = μ + σz.]\n\n         A. 26.52\n         B. 6.52\n         C. 14.12\n         D. 23.48.                                                                                       [2]\n\n         Use the following information for attempting questions 11 to 15:\n\n         Green-City is known for having rainfall all throughout the year. During the year, a few days\n         are marked by torrential rainfall exceeding 5 centimetres.\n\n         A bi-variate linear regression model (Y = α + β * X + e) is being fit to understand the\n         dependency of –\n            • the number of trees falling in Green-City on a day with torrential rainfall (Y) on\n            • the rainfall in the city on that day in centimetres (X).\n\n         Data for 6 such days has been collected from the records of the municipal authorities.\n         Summary for that data is presented below:\n\n                ∑61 𝑥 = 68,     ∑61 𝑦 = 85,   ∑61 𝑥 2 = 842,   ∑61 𝑥𝑦 = 1062,   ∑61 𝑦 2 = 1463.",
      "parts": [],
      "solution": "Correct Answer is Option D.                                                            [2]\n\n               Let us follow the given steps:\n\n               Step 1: Random value u from U (0, 1) is 0.40\n\n               Step 2: Not applicable as u < 0.5\n\n               Step 3: We have to find z such that P(Z <= z) = 1 – u\n\n               P(Z <= z) = 1 – 0.40\n               P(Z <= z) = 0.60\n\n               This can be also expressed as\n               P(Z > z) = 0.40\n\n               From tables, z = 0.2533. Simulated z-value is – 0.2533.\n\n               Step 4: Simulated x value = 25 – 0.2533*6 = 23.4802.",
      "has_math": true,
      "session": "2024-11",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2024-11_QP.pdf",
      "source_sol": "raw/CS1A_2024-11_SOL.pdf"
    },
    {
      "q_num": 11,
      "marks": 2,
      "topic": "inference",
      "subtopics": [],
      "stem": "What is the value of the coefficient β based on the collected data?\n\n         A. 1.3832\n         B. – 1.5093\n         C. 0.7230\n         D. 0.1132.                                                                                      [2]",
      "parts": [],
      "solution": "Correct Answer is Option A.                                                            [2]\n\n               Sxy\n               = (∑61 𝑥𝑦 – n(x_bar)(y_bar))\n               = (1062 – 6 * 68/6 * 85 /6)\n               = 98.67\n\n               Sxx\n               = (∑61 𝑥 2 – n(x_bar)2)\n               = (842 – 6 * (68/6)2)\n               = 71.33\n\n                                                                                           Page 4 of 13\n\fIAI                                                                                                 CS1A-1124\n\n               β\n               = Sxy / Sxx\n               = 16.44 / 11.89\n               = 1.3832.",
      "has_math": true,
      "session": "2024-11",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2024-11_QP.pdf",
      "source_sol": "raw/CS1A_2024-11_SOL.pdf"
    },
    {
      "q_num": 12,
      "marks": 2,
      "topic": "regression_glm",
      "subtopics": [],
      "stem": "A cyclone “Pawan” is expected to pass from the coastline adjoining Green City and an average\n         daily rainfall of 15 centimetres is expected for the next three days in Green City.\n\n         The municipal authorities want to know the number of trees which are expected to fall during\n         the cyclone “Pawan” so that they can keep their disaster task force ready. How many trees are\n         expected to fall based on the fitted bi-variate regression model (round off your answer to the\n         next integer)?\n\n         A. 10\n         B. 20\n         C. 25\n         D. 58.                                                                                              [2]",
      "parts": [],
      "solution": "Correct Answer is Option D.                                                                 [2]\n\n               α\n               = y_bar – x_bar * β\n               = (85/6) – (68/6) * 1.3832\n               = - 1.5096\n\n               y_hat (i.e. number of trees expected to fall in one day of the cyclone Pawan)\n               = -1.5096 + 1.3832 * (15)\n               = 19.2384\n\n               So, total number of trees expected to fall in the next three days of cyclone Pawan\n               = 19.2384 * 3\n               = 57.7152\n               = 58 (rounded off to the next integer as instructed)",
      "has_math": false,
      "session": "2024-11",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2024-11_QP.pdf",
      "source_sol": "raw/CS1A_2024-11_SOL.pdf"
    },
    {
      "q_num": 13,
      "marks": 2,
      "topic": "regression_glm",
      "subtopics": [],
      "stem": "What is the value of residual sum of squares for the bi-variate linear regression model?\n\n         A. 258\n         B. 136\n         C. 122\n         D. 71.                                                                                              [2]",
      "parts": [],
      "solution": "Correct Answer is Option C.                                                                 [2]\n\n               Syy\n               = (∑61 𝑦 2 – n(y_bar)2)\n               = (1463 – 6 * (85/6)2)\n               = 258.83\n\n               Residual sum of squares\n               = Syy – S2xy / Sxx\n               = 258.33 – 98.672 / 71.33\n               = 122.36\n               = 122 (approximately).",
      "has_math": false,
      "session": "2024-11",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2024-11_QP.pdf",
      "source_sol": "raw/CS1A_2024-11_SOL.pdf"
    },
    {
      "q_num": 14,
      "marks": 2,
      "topic": "regression_glm",
      "subtopics": [],
      "stem": "For predicting the results more accurately, an alternative model is being fit to the data which\n         has the following regression equation:\n                                           Y = σ ∗ exp(λ ∗ X + e)\n\n         Following additional calculations are made based on the collected data:\n             • ∑61 ln yi = 15.039379,\n             • ∑61 xi ln yi = 178.519074.\n\n         What is the value of the coefficient λ for the alternative model?\n\n         A. 7.9256\n         B. 0.1132\n         C. 1.1345\n         D. 3.4007.                                                                                          [2]",
      "parts": [],
      "solution": "Correct Answer is Option B.                                                                 [2]\n\n               Y = σ ∗ exp(λ ∗ X + e)\n\n               ln Y = ln σ + λ ∗ X + e\n\n               Sxlny\n               = (∑61 xi ln yi – n(x_bar)(ln y_bar)) ……. where ln y_bar = (∑ ln yi) / 6\n               = (178.519074 – 6 * 68/6 * 15.039379 /6)\n               = 8.0728\n\n               λ\n               = Sxlny / Sxx\n               = 8.0728 / 71.33\n               = 0.113170.",
      "has_math": true,
      "session": "2024-11",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2024-11_QP.pdf",
      "source_sol": "raw/CS1A_2024-11_SOL.pdf"
    },
    {
      "q_num": 15,
      "marks": 2,
      "topic": "distributions",
      "subtopics": [],
      "stem": "What is the value of the coefficient σ for the alternative model fitted in question 14?\n\n         A. 1.0697\n         B. 3.4007\n         C. 0.7191\n         D. 2.0526.                                                                                          [2]\n\n         Use the following information for attempting questions 16 to 20:\n\n         A start-up insurance company offering bite-size insurance products has recently launched a\n         policy which offers “performer and entertainer liability insurance” for stand-up comedians.\n\n         Performer liability has two aspects which need to be statistically modelled –\n             • Number of people from the audience filing a legal suit against the comedian during\n                the year (N) and\n             • Total amount payable to the comedian under the policy during the year (Y).\n\n         Y = ∑N1 X = X1 + X2 + …………….. + XN ….. where Xi denotes the amount (INR in lakhs)\n         payable in respect of the ith claim for i = 1, 2, …………….., N.\n\n         N follows a Binomial distribution with parameter p = 0.005 and Xi follows a Gamma\n         Distribution with parameters α = ½ and β = 1/10.\n\n         Mr. Prakhar, a renowned stand-up comedian in the country who is one of the policyholders of\n         the company has participated in 10 stand-up events during the year with average audience size\n         of 100 persons per event.",
      "parts": [],
      "solution": "Correct Answer is Option B.                                                                 [2]\n\n               ln σ\n               = ln y_bar - λ * x_bar\n\n                                                                                               Page 5 of 13\n\fIAI                                                                          CS1A-1124\n\n               = 15.039379/6 – 0.113170 * 68/6\n               = 1.22397\n\n               σ\n               = exp(1.22397)\n               =3.400661.",
      "has_math": true,
      "session": "2024-11",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2024-11_QP.pdf",
      "source_sol": "raw/CS1A_2024-11_SOL.pdf"
    },
    {
      "q_num": 16,
      "marks": 2,
      "topic": "inference",
      "subtopics": [],
      "stem": "What is the value of E(Xi) and Var(Xi)?\n\n         A. E(Xi) = 5 and Var(Xi) = 50\n         B. E(Xi) = 10 and Var(Xi) = 100\n         C. E(Xi) = 20 and Var(Xi) = 100\n         D. E(Xi) = 10 and Var(Xi) = 50.                                                                 [2]",
      "parts": [],
      "solution": "Correct Answer is Option A.                                            [2]\n\n               E(Xi)\n               =α/β\n               = ½ / (1/10)\n               =5\n\n               Var(Xi)\n               = α / β2\n               = ½ / (1/10)2\n               = 50.",
      "has_math": false,
      "session": "2024-11",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2024-11_QP.pdf",
      "source_sol": "raw/CS1A_2024-11_SOL.pdf"
    },
    {
      "q_num": 17,
      "marks": 2,
      "topic": "inference",
      "subtopics": [],
      "stem": "What is the expected number of claims and the standard deviation in the number of claims\n         that Mr. Prakhar will file with the insurance company during the year?\n\n         A. E(N) = 5 and SD(N) = 4.975\n         B. E(N) = 5 and SD(N) = 2.230\n         C. E(N) = 10 and SD(N) = 9.950\n         D. E(N) = 10 and SD(N) = 3.154.                                                                 [2]",
      "parts": [],
      "solution": "Correct Answer is Option B.                                            [2]\n\n               E(N)\n               =n*p\n               = (10 * 100) * 0.005\n               =5\n\n               Var(N)\n               =n*p*q\n               = (10 * 100) * 0.005 * 0.995\n               = 4.975\n\n               SD(N)\n               = sqrt(4.975)\n               = 2.230.",
      "has_math": false,
      "session": "2024-11",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2024-11_QP.pdf",
      "source_sol": "raw/CS1A_2024-11_SOL.pdf"
    },
    {
      "q_num": 18,
      "marks": 2,
      "topic": "distributions",
      "subtopics": [],
      "stem": "What is the expected value of total claim amount from Mr. Prakhar during the year under the\n         policy (INR in lakhs)?\n\n         A. 5\n         B. 50\n         C. 25\n         D. 100.                                                                                         [2]",
      "parts": [],
      "solution": "Correct Answer is Option C.                                            [2]\n\n               E(Y)\n               = E ( E(Y|N) )\n               = E ( E(X1 + X2 +………….. + XN) )\n               = E ( N * E(Xi) )\n               = E ( N * 5)\n               = 5 * E(N)\n               =5*5\n               = 25 lakhs.",
      "has_math": false,
      "session": "2024-11",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2024-11_QP.pdf",
      "source_sol": "raw/CS1A_2024-11_SOL.pdf"
    },
    {
      "q_num": 19,
      "marks": 2,
      "topic": "distributions",
      "subtopics": [],
      "stem": "Which of the following represents the correct expression for the unconditional variance of Y\n         i.e. Var(Y)?\n\n         A. Var(Y) = E[Var(X|N)] + Var[E(X|N)]\n         B. Var(Y) = E[Var(N|Y)] + Var[E(N|Y)]\n         C. Var(Y) = E[Var(N|X)] + Var[E(N|X)]\n         D. Var(Y) = E[Var(Y|N)] + Var[E(Y|N)].                                                          [2]",
      "parts": [],
      "solution": "Correct Answer is Option D.                                            [2]\n\n               The correct expression for unconditional variance of Y is:\n\n               Var(Y) = E[Var(Y|N)] + Var[E(Y|N)].",
      "has_math": false,
      "session": "2024-11",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2024-11_QP.pdf",
      "source_sol": "raw/CS1A_2024-11_SOL.pdf"
    },
    {
      "q_num": 20,
      "marks": 2,
      "topic": "regression_glm",
      "subtopics": [],
      "stem": "What is the standard deviation of the total claim amount from Mr. Prakhar during the year\n         under the policy (INR in lakhs)?\n\n         A. 15.77\n         B. 19.35\n         C. 11.15\n         D. 26.92.                                                                                       [2]\n\n         Use the following information for attempting questions 21 to 25:\n\n         VIDEA Ltd is a leading telecommunication company in the country of Actuaria. Strategy\n         team of the company has developed a metric “Customer Value” which measures the worth of\n         the product or service to the customer.\n\n         Customer value (Y) is derived based on multiple explanatory variables:\n            • Usage (in seconds) (X1),\n\n               •    Subscription length (in months) (X2),\n               •    Age (in years) (X3),\n               •    Number of call failures (X4).\n\n         The analysts from the strategy team have developed a multi-variate linear regression model\n         based on data of 25 customers. Following information is presented to you:\n\n             Particulars Coefficient        Standard           99% confidence interval for the\n                                              Error                     coefficient\n                   β0        172.3588        9.6541\n                   β1          0.0885        0.0036                    (0.0783, 0.0987)\n                   β2          0.7825        0.1586                    (0.3313, 1.2337)\n                   β3         -8.0144        0.6283                   (-9.8019, -6.2269)\n                   β4         -0.4372        0.2537                   (-1.1590, 0.2846)\n\n         The analyst has also constructed the following ANOVA table:\n\n             Source of variation    Degrees of freedom      Sum of squares    Mean sum of squares\n                 Regression                  4                    ?              1602.748925\n                  Residual                   ?                2136.9986               ?\n                   Total                     ?                    ?                   ?\n         .",
      "parts": [],
      "solution": "Correct Answer is Option B.                                            [2]\n\n               Var(Y)\n               = E[Var(X1 + X2 +………… + XN)] + Var[E(X1 + X2 +……….. + XN)]\n                                                                            Page 6 of 13\n\fIAI                                                                                              CS1A-1124\n\n               = E[N * Var(Xi)] + Var [ N * E(Xi)] …………. As Xis are independent\n               = E[N * 50] + Var [N * 5]\n               = 50 * E[N] + 52 * Var[N]\n               = 50 * 5 + 25 * 4.975\n               = 374.375\n\n               SD(Y)\n               = sqrt(374.375)\n               = 19.35.",
      "has_math": true,
      "session": "2024-11",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2024-11_QP.pdf",
      "source_sol": "raw/CS1A_2024-11_SOL.pdf"
    },
    {
      "q_num": 21,
      "marks": 2,
      "topic": "inference",
      "subtopics": [],
      "stem": "Based on the 99% confidence intervals calculated by the analysts, which explanatory variables\n         can be considered to be significant?\n\n         A. X1 and X2\n         B. X1 , X2 and X3\n         C. X1 , X2, X3 and X4\n         D. None of the variables are significant.                                                       [2]",
      "parts": [],
      "solution": "Correct Answer is Option B.                                                                [2]\n\n               The 99% two-sided confidence intervals for coefficients β1, β2 and β3 do not contain the\n               value 0 and hence the explanatory variables corresponding to these coefficients viz. X1,\n               X2 and X3 can be considered to be significant.\n\n               The 99% two-sided confidence interval for coefficient β4 contains the value 0 and hence\n               the explanatory variable X4 is considered to be insignificant.",
      "has_math": false,
      "session": "2024-11",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2024-11_QP.pdf",
      "source_sol": "raw/CS1A_2024-11_SOL.pdf"
    },
    {
      "q_num": 22,
      "marks": 2,
      "topic": "regression_glm",
      "subtopics": [],
      "stem": "What is the value of the Coefficient of Determination (R2)? Answer using the ANOVA table\n         constructed by the analyst.\n\n         A. 75%\n         B. 43%\n         C. 80%\n         D. 25%.                                                                                         [2]",
      "parts": [],
      "solution": "Correct Answer is Option A.                                                                [2]\n\n               SSREG\n               = 1602.748925 * 4\n               = 6410.9957.\n\n               SSTOT\n               = 6410.9957 + 2136.9986\n               = 8547.9943\n\n               R2\n               = SSREG / SSTOT\n               = 6410.9957 / 8547.9943\n               = 75%.",
      "has_math": false,
      "session": "2024-11",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2024-11_QP.pdf",
      "source_sol": "raw/CS1A_2024-11_SOL.pdf"
    },
    {
      "q_num": 23,
      "marks": 2,
      "topic": "regression_glm",
      "subtopics": [],
      "stem": "What is the value of Adjusted R2?\n\n         A. 10%\n         B. 76%\n         C. 32%\n         D. 70%.                                                                                         [2]",
      "parts": [],
      "solution": "Correct Answer is Option D.                                                                [2]\n\n               Adjusted R2\n               = 1 – (n – 1) / (n – k – 1) * (1 – R2)\n               = 1 – (25 – 1) / (25 – 4 – 1) * (1 – 75%)\n               = 1 – 24/20 * 25%\n               = 70%.",
      "has_math": false,
      "session": "2024-11",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2024-11_QP.pdf",
      "source_sol": "raw/CS1A_2024-11_SOL.pdf"
    },
    {
      "q_num": 24,
      "marks": 2,
      "topic": "regression_glm",
      "subtopics": [
        "inference"
      ],
      "stem": "The Chief Operating Officer of the company is of the view that none of the explanatory\n         variables have any significant impact on the customer value. Analysts from the strategy team\n         are statistically testing this claim using one-sided analysis of variance (ANOVA).\n\n         What is the outcome of the ANOVA test at 5% level of significance?\n\n         A. View of the Chief Operating Officer is justified. None of the four explanatory variables\n            have any impact on the customer value.                                                     [2]\n\n         B. View of the Chief Operating Officer is not justified. At least one of the four explanatory\n            variables has a significant impact on the customer value.\n         C. View of the Chief Operating Officer is not justified. All the explanatory variables have a\n            significant impact on the customer value.\n         D. ANOVA test cannot be applied to the dataset.",
      "parts": [],
      "solution": "Correct Answer is Option B.                                                                [2]\n\n               Value of ANOVA Test Statistic\n               = regression mean square / residual mean square\n               = 1602.748925 / (2136.9986 / (25 – 4 – 1))\n               = 15\n\n               Critical value of Fk,(n – k – 1) = F4,20 at 5% level of significance is 2.866.\n\n               The value of the statistic is much greater than the critical value. Hence we have\n               sufficient evidence to reject the null hypothesis.\n\n                                                                                                Page 7 of 13\n\fIAI                                                                                               CS1A-1124\n\n               H0 states all βi = 0 (for i = 1,2,3,4). H1 states that at least one βi ≠ 0. Hence Option B\n               represents the correct answer. At least one of the explanatory variables has significant\n               impact on the customer value.",
      "has_math": false,
      "session": "2024-11",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2024-11_QP.pdf",
      "source_sol": "raw/CS1A_2024-11_SOL.pdf"
    },
    {
      "q_num": 25,
      "marks": 2,
      "topic": "regression_glm",
      "subtopics": [],
      "stem": "The Chief Strategy Officer has suggested that the model should be optimized and only those\n         explanatory variables which have a significant impact on the customer value must be retained.\n         Analysts have decided to use the process of backward selection for selecting the optimal\n         number of explanatory variables.\n\n         Some output of this process is shown in the below table:\n\n          Model                                         R2                         Adjusted R2\n          X1 + X2 + X3 + X4                             ?                               ?\n          X1 + X2 + X3                                 76%                            73%\n          X1 + X2                                      73%                            70%\n          X1                                           69%                            68%\n\n         Select the option which represents the optimal set of explanatory variables for estimating\n         customer value –\n\n         A. X1 only\n         B. X1 and X2\n         C. X1, X2 and X3\n         D. X1, X2, X3 and X4.                                                                              [2]\n\n         Use the following information for attempting questions 26 to 30:\n\n         A random variable X is a continuous random variable representing the age in years of a\n         historical artefact in a country’s premier museum. X is believed to have the following\n         probability density function: f(x) = 3λ3 (λ + x)-4 for x > 0.\n\n         In order to test the null hypothesis H0 : λ = 50 against the alternative hypothesis H1 : λ = 60,\n         a single value is observed. If this value is greater than 93.50, H0 is rejected.",
      "parts": [],
      "solution": "Correct Answer is Option C.                                                                  [2]\n\n               Selection of the optimal set of explanatory variables (using either forward selection or\n               backward selection) should be based on Adjusted R2.\n\n                Model                  Adjusted R2                      Decision\n                X1 + X2 + X3 + X4         70%\n                X1 + X2 + X3              73%           Adjusted R2 is increasing so X4 can be\n                                                        dropped.\n                X1 + X2                    70%          Adjusted R2 is reducing so X3 need not be\n                                                        dropped.\n                X1                         68%          Adjusted R2 is reducing so X2 need not be\n                                                        dropped.\n\n               So, the optimal set of explanatory variables will be X1, X2 and X3. This is consistent\n               with the solution obtained in question 21.",
      "has_math": true,
      "session": "2024-11",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2024-11_QP.pdf",
      "source_sol": "raw/CS1A_2024-11_SOL.pdf"
    },
    {
      "q_num": 26,
      "marks": 2,
      "topic": "inference",
      "subtopics": [],
      "stem": "In the context of the above hypothesis, what is definition of Type I Error and Type II Error?\n\n                Type I Error                                 Type II Error\n           A.   Fail to reject H0 when it is true            Reject H0 when it is false\n           B.   Fail to reject H0 when it is true            Fail to reject H0 when it is false\n           C.   Reject H0 when it is true                    Reject H0 when it is false\n           D.   Reject H0 when it is true                    Fail to reject H0 when it is false             [2]",
      "parts": [],
      "solution": "Correct Answer is Option D.                                                                  [2]\n\n               Type I Error: Reject H0 when it is true.\n               Type II Error: Fail to reject H0 when it is false.",
      "has_math": false,
      "session": "2024-11",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2024-11_QP.pdf",
      "source_sol": "raw/CS1A_2024-11_SOL.pdf"
    },
    {
      "q_num": 27,
      "marks": 2,
      "topic": "inference",
      "subtopics": [],
      "stem": "What is the probability of Type I Error?\n\n         A. 4.23%\n         B. 5.97%\n         C. 0.09%\n         D. 12.69%.                                                                                         [2]",
      "parts": [],
      "solution": "Correct Answer is Option A.                                                                  [2]\n\n               Probability of Type I Error\n               = P(X > 93.50) under H0: (λ = 50)\n               = integral from 93.50 to ∞ (3λ3 (λ + x)-4) dx\n               = -3λ3 / 3 * (from 93.50 to ∞ (λ + x)-3 )\n\n               = -503 * (from 93.50 to ∞ (50 + x)-3)\n               = 0.0423\n               = 4.23%",
      "has_math": false,
      "session": "2024-11",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2024-11_QP.pdf",
      "source_sol": "raw/CS1A_2024-11_SOL.pdf"
    },
    {
      "q_num": 28,
      "marks": 2,
      "topic": "inference",
      "subtopics": [],
      "stem": "What is the probability of Type II Error?\n\n         A. 4.23%                                                                                           [2]\n\n         B. 94.03%\n         C. 5.97%\n         D. 17.92%.",
      "parts": [],
      "solution": "Correct Answer is Option B.                                                                  [2]\n\n               Probability of Type II Error\n               = P(X <= 93.50) under H1: (λ = 60)\n               = integral from 0 to 93.50 (3λ3 (λ + x)-4) dx\n               = -3λ3 / 3 * (from 0 to 93.50 (λ + x)-3 )\n               = -603 * (from 0 to 93.50 (60 + x)-3)\n               = 0.9403\n               = 94.03%",
      "has_math": false,
      "session": "2024-11",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2024-11_QP.pdf",
      "source_sol": "raw/CS1A_2024-11_SOL.pdf"
    },
    {
      "q_num": 29,
      "marks": 2,
      "topic": "inference",
      "subtopics": [],
      "stem": "What is the definition of power of the test and size of the test?\n\n                 Power of Test                                Size of Test\n            A.   1 – Probability (Type I Error)               Probability (Type II Error)\n            B.   Probability (Type II Error)                  1 – Probability (Type I Error)\n            C.   1 – Probability (Type II Error)              Probability (Type I Error)\n            D.   1 – Probability (Type II Error)              1 – Probability (Type I Error)               [2]",
      "parts": [],
      "solution": "Correct Answer is Option C.                                                                  [2]\n\n               Power of Test = 1 – Probability of committing Type II Error\n               Size of Test = Probability of committing Type I Error",
      "has_math": false,
      "session": "2024-11",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2024-11_QP.pdf",
      "source_sol": "raw/CS1A_2024-11_SOL.pdf"
    },
    {
      "q_num": 30,
      "marks": 2,
      "topic": "inference",
      "subtopics": [],
      "stem": "What is the definition of sensitivity of the test and specificity of the test?\n\n                 Sensitivity of Test                          Specificity of Test\n            A.   1 – Probability (Type I Error)               1 – Probability (Type II Error)\n            B.   Probability (Type II Error)                  Probability (Type I Error)\n            C.   1 – Probability (Type II Error)              1 – Probability (Type I Error)\n            D.   Probability (Type I Error)                   Probability (Type II Error)                  [2]\n\n         Use the following information for attempting questions 31 to 35:\n\n         The civic administration of a metropolis is trying to implement preventive measures to reduce\n         road accidents due to drink and drive.\n\n         Following is a sample of results of Breath Alcohol Content (BAC) test of civilians who have\n         breached the legal limit of 8 basis points per 100 ml of blood. The sample (in basis points) is\n         assumed to be taken from a normal distribution with mean μ and variance of 20.\n\n         56, 32, 49, 57, 44.\n\n         The administration has made it compulsory for every lounge in the city to administer a drug\n         “Neurontin” as a part of the drink so that the alcohol absorption in the body will be delayed\n         and it will continue to promote better alertness and attentiveness while driving. This is\n         believed to reduce the BAC below the legally permissible level and thereby reduce the\n         accidents due to drink and drive.\n\n         For the same participants among the sample, the BAC test is taken again after the\n         administration of the drug and the results were found to be as follows:\n\n         8, 4, 7, 6, 5.",
      "parts": [],
      "solution": "Correct Answer is Option C.                                                                  [2]\n\n               Sensitivity of Test = Power of Test = = 1 – Probability of committing Type II Error\n\n                                                                                                 Page 8 of 13\n\fIAI                                                                                              CS1A-1124\n\n               Specificity of Test = 1 – Probability of committing Type I Error.",
      "has_math": true,
      "session": "2024-11",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2024-11_QP.pdf",
      "source_sol": "raw/CS1A_2024-11_SOL.pdf"
    },
    {
      "q_num": 31,
      "marks": 2,
      "topic": "inference",
      "subtopics": [],
      "stem": "Which of the following represents a symmetrical 95% confidence interval for μ based on the\n         original sample (in basis points)?\n\n         A. (43.68, 51.52)\n         B. (8.40, 8.68)\n         C. (42.95, 52.25)\n         D. (1.10, 94.10).                                                                                 [2]",
      "parts": [],
      "solution": "Correct Answer is Option A.                                                                 [2]\n\n               The 95% confidence interval for the original sample\n               = x_bar +/- 1.96 * σ / sqrt(n)\n               = (56 + 32 + 49 + 57 + 44) / 5 +/- 1.96 * sqrt(20/5)\n               = 47.6 +/- 3.92\n               = (43.68, 51.52).",
      "has_math": true,
      "session": "2024-11",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2024-11_QP.pdf",
      "source_sol": "raw/CS1A_2024-11_SOL.pdf"
    },
    {
      "q_num": 32,
      "marks": 2,
      "topic": "inference",
      "subtopics": [],
      "stem": "Which of the following represents a symmetrical 95% prediction interval for a single value\n         from this distribution based on the original sample (in basis points)?\n\n         A. (34.98, 60.22)\n         B. (36.20, 59.00)                                                                                 [2]\n\n         C. (38.00, 57.20)\n         D. (39.54, 55.66).",
      "parts": [],
      "solution": "Correct Answer is Option C.                                                                 [2]\n\n               The 95% prediction interval for the original sample\n               = x_bar +/- 1.96 * σ * sqrt(1 + 1/n)\n               = 47.6 +/- 1.96 * sqrt(20) * sqrt(1 + 1/5)\n               = 47.6 +/- 9.602\n               = (38.00, 57.20)",
      "has_math": false,
      "session": "2024-11",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2024-11_QP.pdf",
      "source_sol": "raw/CS1A_2024-11_SOL.pdf"
    },
    {
      "q_num": 33,
      "marks": 2,
      "topic": "inference",
      "subtopics": [],
      "stem": "We want to reduce the width of each of these intervals. What should be done to reduce their\n         width?\n\n         Choose the correct option from those given below which will result in reduction in the width\n         of both these intervals:\n\n                Reduce the Width of Symmetrical             Reduce the Width of Symmetrical 95%\n                95% Confidence Interval                     Prediction Interval\n           A.   Reduce the sample size n                    Reduce the sample size n\n           B.   Increase the sample size n                  Reduce the sample size n\n           C.   Reduce the sample size n                    Increase the sample size n\n           D.   Increase the sample size n                  Increase the sample size n                   [2]",
      "parts": [],
      "solution": "Correct Answer is Option D.                                                                 [2]\n\n               The width of the 95% confidence interval\n               = 2 * 1.96 * σ / sqrt(n)\n\n               Higher the value of n, lower will be the width of the 95% confidence interval\n\n               The width of the 95% prediction interval\n               = 2 * 1.96 * σ * sqrt(1 + 1/n)\n\n               Higher the value of n, lower will be the width of the 95% prediction interval.",
      "has_math": false,
      "session": "2024-11",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2024-11_QP.pdf",
      "source_sol": "raw/CS1A_2024-11_SOL.pdf"
    },
    {
      "q_num": 34,
      "marks": 2,
      "topic": "distributions",
      "subtopics": [],
      "stem": "Based on the two samples – before and after the administration of the drug “Neurontin”, the\n         mean value of reduction in BAC levels and the standard deviation of reduction in BAC levels\n         is –\n\n         A. Mean = 60 basis points, Standard Deviation = 1.41 basis points\n         B. Mean = 41.6 basis points, Standard Deviation = 8.01 basis points\n         C. Mean = 60 basis points, Standard Deviation = 1.58 basis points\n         D. Mean = 41.6 basis points, Standard Deviation = 8.96 basis points.                            [2]",
      "parts": [],
      "solution": "Correct Answer is Option D.                                                                 [2]\n\n               Average reduction in BAC levels\n               = (48 + 28 + 42 + 51 + 39) / 5\n               = 41.60 basis points\n\n               Variance of reduction in BAC levels\n               (Σ (x – x_bar)^2) / (n-1)\n               = (40.96 + 184.96 + 0.16 + 88.36 + 6.76) / 4\n               = 321.2 / 4\n               = 80.30\n\n               Standard deviation of reduction in BAC levels\n               = sqrt (80.30)\n               = 8.96 basis points.",
      "has_math": false,
      "session": "2024-11",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2024-11_QP.pdf",
      "source_sol": "raw/CS1A_2024-11_SOL.pdf"
    },
    {
      "q_num": 35,
      "marks": 2,
      "topic": "inference",
      "subtopics": [
        "distributions"
      ],
      "stem": "Which of the following represents a 90% symmetrical confidence interval for the reduction\n         in BAC levels (in basis points)?\n\n         A. (4.73, 72.70)\n         B. (34.38, 48.82)\n         C. (35.01, 48.19)\n         D. (33.06, 50.14).                                                                              [2]\n\n         Use the following information for attempting questions 36 to 40:\n\n         In the country Economa, general elections are being conducted for electing new government\n         for a period of next five years.\n\n         Ruling Alliance and Opposition Alliance are equally strong contenders and have an equal\n         chance of winning the poll.\n\n         The increase in the country’s stock market index in the next one year is assumed to follow a\n         Pareto distribution (two parameter version with α = 1) with the following probability density\n         function:\n\n                                             λ\n                                 f(x|λ) = (λ+x)2        0 < x < ∞,      λ > 0.\n\n         If the Ruling Alliance is re-elected to power, λ is expected to be equal to 100 and if the\n         Opposite Alliance is elected, λ is expected to be equal to 300.",
      "parts": [],
      "solution": "Correct Answer is Option D.                                                                 [2]\n\n               90% confidence interval for the reduction in BAC levels\n               = 41.60 +/- t4 (at 5%) * 8.96 / sqrt(5)\n               = 41.60 +/- 2.132 * 8.96 / sqrt(5)\n               = (33.06, 50.14).",
      "has_math": true,
      "session": "2024-11",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2024-11_QP.pdf",
      "source_sol": "raw/CS1A_2024-11_SOL.pdf"
    },
    {
      "q_num": 36,
      "marks": 2,
      "topic": "bayes_credibility",
      "subtopics": [],
      "stem": "Which of the following represents the correct prior distribution of λ?\n\n                      𝜆\n         A. f(λ) = (x+λ)2\n         B. P(λ = 100) = 0.5, P(λ = 300) = 0.5                                                           [2]\n\n                       𝑥\n         C.   f(λ) = (x+λ)2\n         D. P(λ= 100) = 0.75, P(λ = 300) = 0.25.",
      "parts": [],
      "solution": "Correct Answer is Option B.                                                                 [2]\n                                                                                                Page 9 of 13\n\fIAI                                                                                               CS1A-1124\n\n               Ruling Alliance and Opposition Alliance are equally strong contenders and have an\n               equal chance of winning the poll.\n\n               If the Ruling Alliance is re-elected to power, λ is expected to be equal to 100 and if the\n               Opposite Alliance is elected, λ is expected to be equal to 300.\n\n               Hence the prior distribution is: P(λ = 100) = 0.5, P(λ = 300) = 0.5",
      "has_math": true,
      "session": "2024-11",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2024-11_QP.pdf",
      "source_sol": "raw/CS1A_2024-11_SOL.pdf"
    },
    {
      "q_num": 37,
      "marks": 2,
      "topic": "distributions",
      "subtopics": [],
      "stem": "Which of the following correctly represents the distribution function of the two parameter\n         Pareto distribution? Answer based on Actuarial Formulae and Tables.\n\n                          λ\n         A. F(x|λ) = (λ+x)2\n                              λ\n         B. F(x|λ) = 1 − (λ+x)2\n                              λ\n         C. F(x|λ) = 1 − (λ+x)\n                         λ\n         D. F(x|λ) = (λ+x).                                                                                  [2]",
      "parts": [],
      "solution": "Correct Answer is Option C.                                                                  [2]\n\n               From Actuarial Tables, it can be seen that the distribution function of the two parameter\n               Pareto distribution is –\n                                                                       𝛼\n                                                                 λ\n                                               F(x) = 1 − (           )\n                                                              (λ + x)\n\n               We know that in the extant case, α = 1. Hence,\n\n                                                                      λ\n                                                    F(x|λ) = 1 −\n                                                                   (λ + x)",
      "has_math": true,
      "session": "2024-11",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2024-11_QP.pdf",
      "source_sol": "raw/CS1A_2024-11_SOL.pdf"
    },
    {
      "q_num": 38,
      "marks": 2,
      "topic": "bayes_credibility",
      "subtopics": [],
      "stem": "Technical analysts believe that the stock market index of the country will grow by at least 500\n         points by the end of the year. Which alliance has a better chance of winning the general\n         elections based on this information?\n\n         Hint: [Answer based on the posterior distribution of λ i.e. f (λ | x > 500)]\n\n         A. Ruling Alliance\n         B. Opposition Alliance\n         C. Both Ruling Alliance and Opposition Alliance are equally likely to win\n         D. None of them will win, elections will have to be re-conducted.                                   [2]",
      "parts": [],
      "solution": "Correct Answer is Option B.                                                                  [2]\n\n               Posterior probability that the Ruling Alliance wins the election\n               = f (λ=100 | x > 500)\n                              (1−F(500|λ=100))∗P(λ=100)\n               = (1−F(500|λ=100))∗P(λ=100)+(1−F(500|λ=300))∗P(λ=300)\n\n                    0.1667\n               = 0.1667+0.375\n\n               = 0.307692.\n\n               Posterior probability that the Opposition Alliance wins the election\n               = f (λ=300 | x > 500)\n                               (1−𝐹(500|λ=300))∗𝑃(λ=300)\n               = (1−𝐹(500|λ=100))∗𝑃(λ=100)+(1−𝐹(500|λ=300))∗𝑃(λ=300)\n                    0.375\n               = 0.1667+0.375\n\n               = 0.692308.\n\n               If the stock market index is expected to grow by at least 500 points by the end of the\n               year, Opposition Alliance has a better chance of winning the elections as compared to\n               the Ruling Alliance.",
      "has_math": true,
      "session": "2024-11",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2024-11_QP.pdf",
      "source_sol": "raw/CS1A_2024-11_SOL.pdf"
    },
    {
      "q_num": 39,
      "marks": 2,
      "topic": "inference",
      "subtopics": [],
      "stem": "What should be the minimum level of expected growth in the stock market index during the\n         year, which will make change in the government virtually certain?\n\n         A. 200 points\n         B. 300 points\n         C. 500 points\n         D. None of the above.                                                                               [2]",
      "parts": [],
      "solution": "Correct Answer is Option D.                                                                  [2]\n\n               We want to find the minimum level of growth in the stock market index which will\n               make the change in government virtually certain.\n\n               So, we want to find “a” such that:\n\n               f (λ=300 | x > a) = 1\n\n               Let us check each of the options.\n                                                                                               Page 10 of 13\n\fIAI                                                                                             CS1A-1124\n\n               A. At a = 200, f (λ=300 | x > 200) = 0.64287 < 1\n               B. At a = 300, f (λ=300 | x > 300) = 0.66667 < 1\n               C. At a = 500, we know that f (λ=300 | x > 500) = 0.692308 < 1\n               D. None of the above – this is the right answer.",
      "has_math": false,
      "session": "2024-11",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2024-11_QP.pdf",
      "source_sol": "raw/CS1A_2024-11_SOL.pdf"
    },
    {
      "q_num": 40,
      "marks": 2,
      "topic": "bayes_credibility",
      "subtopics": [
        "distributions"
      ],
      "stem": "In the context of Bayesian statistics, loss function L(g(x), θ) is a measure of the loss incurred\n         when g(x) is used as an estimator of θ. Bayesian estimator g(x) is that which minimises the\n         expected loss with respect to the posterior distribution.\n\n         Which of the following is the correct expression for “All-or-nothing” loss function?\n\n         A. L(g(x), θ) = 0 if g(x) = θ, L(g(x), θ) = 1 if g(x) ≠ θ\n         B. L(g(x), θ) = [g(x) – θ]2\n         C. L(g(x), θ) = | g(x) – θ |\n         D. L(g(x), θ) = g(x) – θ.                                                                           [2]\n\n         Use the following information for attempting questions 41 to 45:\n\n         Let θ denote the proportion of wrong answers given by an AI Chatbot. Prior beliefs about θ\n         are described by a beta distribution with parameters α and β. Let μ and σ2 denote the prior\n         mean and prior variance of θ.\n\n         A random sample of ‘n’ questions is taken and it is observed that wrong answers have been\n         given in respect of ‘w’ questions.",
      "parts": [],
      "solution": "Correct Answer is Option A.                                                               [2]\n\n               In case of All or Nothing Loss of 0/1 loss, the loss function is given by the following\n               expression:\n\n                                   L(g(x), θ) = 0 if g(x) = θ, L(g(x), θ) = 1 if g(x) ≠ θ",
      "has_math": true,
      "session": "2024-11",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2024-11_QP.pdf",
      "source_sol": "raw/CS1A_2024-11_SOL.pdf"
    },
    {
      "q_num": 41,
      "marks": 2,
      "topic": "distributions",
      "subtopics": [],
      "stem": "Which of the following is TRUE about μ and σ2?\n\n         A. μ = α / (α + β) and σ = μβ / (α + β)\n                                    2            2\n\n         B. μ = α / (α + β) and σ2 = μβ / [(α + β)2 + (α + β)]\n         C. μ = α / (α + β) and σ2 = μβ / (α + β)\n         D. μ = α / (α + β) and σ2 = μβ / (α – β) (α + β).",
      "parts": [],
      "solution": "Correct Answer is Option B.                                                               [2]\n\n               From Actuarial Tables, we know that for a beta distribution,\n\n               μ = α / (α + β)\n\n               σ2\n               = αβ / ((α + β)2 * (α + β+1))\n               = μβ / ((α + β) * (α + β+1))\n               = μβ / ((α + β)2 + (α + β)).",
      "has_math": true,
      "session": "2024-11",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2024-11_QP.pdf",
      "source_sol": "raw/CS1A_2024-11_SOL.pdf"
    },
    {
      "q_num": 42,
      "marks": 2,
      "topic": "inference",
      "subtopics": [],
      "stem": "Experts in the AI company have estimated μ and σ2 to be 10% and (9 / 1100) %% respectively.\n\n         The method of moments estimates for α and β are –\n\n            ̂ MOM = 3 and β̂ MOM = 3\n         A. α\n            ̂ MOM = 9 and β̂ MOM = 1\n         B. α\n            ̂ MOM = 1 and β̂ MOM = 9\n         C. α\n            ̂ MOM = 2 and β̂ MOM = 4.\n         D. α                                                                                            [2]",
      "parts": [],
      "solution": "Correct Answer is Option C.                                                               [2]\n\n               Let’s check each option and work the value of µ first.\n\n               A. µ = 3 / (3+3) = 50% - this is incorrect\n               B. µ = 9 / (9+1) = 90% - this is incorrect\n               C. µ = 1 / (1+9) = 10% - this is correct\n               D. µ = 2 / (2+4) = 66.67% - this is incorrect.\n\n               Now, for Option C, let us find out the value of σ2\n\n               σ2\n               = 9 / ((10)2 * (11))\n               = 9 / 1100 – this is correct\n\n               So, the right answer is Option C.",
      "has_math": true,
      "session": "2024-11",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2024-11_QP.pdf",
      "source_sol": "raw/CS1A_2024-11_SOL.pdf"
    },
    {
      "q_num": 43,
      "marks": 2,
      "topic": "bayes_credibility",
      "subtopics": [
        "distributions"
      ],
      "stem": "What is the posterior distribution of θ based on the random sample of n questions?\n\n         A. θ ~ Beta (α + n, β + n – w)\n         B. θ ~ Beta (α + n – w, β + w)\n         C. θ ~ Beta (α + n – w, β + n + w)\n         D. θ ~ Beta (α + w, β + n – w).                                                                 [2]",
      "parts": [],
      "solution": "Correct Answer is Option D.                                                               [2]\n\n               f prior (θ)\n               ∝ θα-1 * (1-θ)β-1      0<θ<1\n\n               Let X be the number of wrong answers given by the AI Chatbot.\n\n               X | θ ~ Binomial (n, θ)\n\n               We have observed w wrong answers. So the likelihood function is:\n\n               L(θ)\n               = P(x = w | θ)\n               = nCw * θw * (1 – θ)n-w\n               ∝ θw * (1 – θ)n-w\n\n                                                                                             Page 11 of 13\n\fIAI                                                                                               CS1A-1124\n\n               Combining the prior distribution and the sample data, we see that:\n\n               f posterior (θ)\n               ∝ θα-1 * (1-θ)β-1 * θw * (1 – θ)n-w\n               ∝ θα+w-1 * (1-θ)β+n-w-1\n\n               So, posterior distribution of θ is Beta (α+w, β+n-w).",
      "has_math": true,
      "session": "2024-11",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2024-11_QP.pdf",
      "source_sol": "raw/CS1A_2024-11_SOL.pdf"
    },
    {
      "q_num": 44,
      "marks": 2,
      "topic": "inference",
      "subtopics": [],
      "stem": "What is the maximum likelihood estimate of θ based on the random sample of n questions?\n\n         A. θ̂ MLE = w / n\n         B. θ̂ MLE = (n – w) / n\n         C. θ̂ MLE = n / w\n         D. θ̂ MLE = w / (n + w).                                                                        [2]",
      "parts": [],
      "solution": "Correct Answer is Option A.                                                                 [2]\n\n               L(θ)\n               = P(x = w | θ)\n               = nCw * θw * (1 – θ)n-w\n\n               LogL(θ)\n               = w * logθ + (n – w) * log(1-θ) + constant\n\n               d/dθ (LogL(θ)) = w/θ – (n-w)/(1-θ)\n\n               Equating this to 0.\n\n               w * (1 – θ) – (n – w) * θ = 0\n               w – wθ – nθ + wθ = 0\n               w – nθ = 0\n               θ̂ MLE = w / n.\n\n               We can also check that d2/dθ2 (LogL(θ)) < 0 indicating that this is maxima.",
      "has_math": true,
      "session": "2024-11",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2024-11_QP.pdf",
      "source_sol": "raw/CS1A_2024-11_SOL.pdf"
    },
    {
      "q_num": 45,
      "marks": 2,
      "topic": "bayes_credibility",
      "subtopics": [
        "inference"
      ],
      "stem": "Credibility factor Z in this case is derived as: n / (α + β + n). Which of the following is a\n         correct expression for the credibility estimate?\n\n         A. Z * μ + (1 – Z) * θ̂ MLE\n         B. Z * θ̂ MLE + (1 – Z) * μ\n         C. Z * μ + (1 – Z) * (1 / θ̂ MLE)\n         D. Z * θ̂ MLE + (1 – Z) * (1 / μ).                                                              [2]\n\n         Use the following information for attempting questions 46 to 50:\n\n         A matrimonial website calculates a compatibility score between profiles to assist individuals\n         in better shortlisting.\n\n         Compatibility score (Y) is calculated based on following two important factors:\n\n         •   Education:\n\n                                  Criteria                            Education Score (X1)\n              Both of them have professional qualifications                    1\n              Both do not have professional qualifications                     1\n              Other cases                                                      0\n\n         •   Mode of Earning:\n\n                                   Criteria                       Mode of Earning Score (X2)\n              Both of them are into public / private employment               1\n\n              Both of them are self-employed                                        0\n              One is employed and the other is self-employed                        1\n              Other cases                                                           0\n\n         An actuary has been employed by the website owner to predict the compatibility score. The\n         actuary has fit a generalised linear model to predict the compatibility score based on the values\n         of education score and mode of earning score.\n\n         Following information is available about the model:\n             • Poisson distribution has been used.\n             • Linear predictor is given by: g(μ) = β0 + β1 + β2\n             • “1” is the base value for Education Score (X1).\n             • “0” is the base value for Mode of Earning Score (X2).\n\n         Following is a snapshot of the model output achieved by the actuary:\n\n                              Particulars                                Coefficient\n                                   β0                                      0.7732\n                              β1 SCORE = 0                                - 0.5725\n                              β1 SCORE = 1                                    0\n                              β2 SCORE = 0                                    0\n                              β2 SCORE = 1                                 0.6931\n\n         In respect of profiles who have already met, both the individuals are asked to submit a\n         feedback form and rate their compatibility with the other individual. The rating should be\n         between 0% to 100%.",
      "parts": [],
      "solution": "Correct Answer is Option B.                                                                 [2]\n\n               Mean of the posterior distribution can be expressed in the form of a credibility estimate\n               with the credibility factor Z = n / (α + β + n).\n\n               Mean of the posterior distribution\n               = (α+w) / (α+w + β+n-w)\n               = (α+w) / (α+β+n)\n               = w / (α + β + n) + α / (α + β + n)\n               = n / (α + β + n) * (w / n) + (α + β) / (α + β + n) * α / (α + β)\n               = Z * θ̂ MLE + (1 – Z) * µ.",
      "has_math": true,
      "session": "2024-11",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2024-11_QP.pdf",
      "source_sol": "raw/CS1A_2024-11_SOL.pdf"
    },
    {
      "q_num": 46,
      "marks": 2,
      "topic": "regression_glm",
      "subtopics": [],
      "stem": "What is the canonical link function in case of a Poisson model?\n\n         A. g(μ) = log μ\n         B. g(μ) = μ\n         C. g(μ) = log (μ / (1 – μ))\n         D. g(μ) = 1 / μ.                                                                                    [2]",
      "parts": [],
      "solution": "Correct Answer is Option A.                                                                 [2]\n\n               For a Poisson distribution, the canonical link function is:\n\n               g(μ) = log μ",
      "has_math": true,
      "session": "2024-11",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2024-11_QP.pdf",
      "source_sol": "raw/CS1A_2024-11_SOL.pdf"
    },
    {
      "q_num": 47,
      "marks": 2,
      "topic": "regression_glm",
      "subtopics": [],
      "stem": "Why are the values of the coefficients of β1 SCORE = 1 and β2 SCORE = 0 equal to 0?\n\n         A. Based on the current profiles of individuals registered on the website none of the pairs\n            have got an education score of 1 or a mode of earning score of 0. Hence, these coefficients\n            are equal to 0.\n         B. As the Poisson model is used for fitting the GLM, so these coefficients are equal to 0. If a\n            normal model was used instead of Poisson, these coefficients wouldn’t have been equal\n            to 0.\n         C. Education score of 1 and mode of earning score of 0 are considered to be base values and\n            hence the true value of their coefficients is implicitly covered in the intercept coefficient\n            β0 .\n         D. There is an error in the model fit by the actuary and the same needs to be rectified.            [2]",
      "parts": [],
      "solution": "Correct Answer is Option C.                                                                 [2]\n\n               As mentioned in the question, education score of 1 and mode of earning score of 0 are\n               considered to be base values. Hence, the true value of their coefficients is implicitly\n               covered in the intercept coefficient β0.",
      "has_math": true,
      "session": "2024-11",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2024-11_QP.pdf",
      "source_sol": "raw/CS1A_2024-11_SOL.pdf"
    },
    {
      "q_num": 48,
      "marks": 2,
      "topic": "regression_glm",
      "subtopics": [],
      "stem": "For Mr. Fast and Ms. Steady, following information is presented based on their profile details\n         as per the website –\n             • Mr. Fast is a qualified Chartered Accountant whereas Ms. Steady is an Architect.\n             • Mr. Fast has his own CA practice whereas Ms. Steady works as a freelancer.\n\n         Select the correct value of compatibility score for Mr. Fast and Ms. Steady based on the fitted\n         model. Use the link function g(μ) = μ.\n\n         A. 89.38%\n         B. 20.07%\n         C. 77.32%\n         D. 46.63%.",
      "parts": [],
      "solution": "Correct Answer is Option C.                                                                 [2]\n\n               It is given that:\n\n                                                                                               Page 12 of 13\n\fIAI                                                                                                     CS1A-1124\n\n                   •   Mr. Fast is a qualified Chartered Accountant whereas Ms. Steady is an Architect.\n                       So, both are professionals.\n                   •   Mr. Fast has his own CA practice whereas Ms. Steady works as a freelancer. So,\n                       both are self-employed.\n\n               Education Score = 1\n               Mode of Earning Score = 0\n\n               g(μ)\n               = β0 + β1 + β2\n               = 0.7732 + 0 + 0\n               = 0.7732\n\n               Using the given link function, g(μ) = μ,\n\n               μ\n               = 0.7732\n               = 77.32%",
      "has_math": true,
      "session": "2024-11",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2024-11_QP.pdf",
      "source_sol": "raw/CS1A_2024-11_SOL.pdf"
    },
    {
      "q_num": 49,
      "marks": 2,
      "topic": "inference",
      "subtopics": [],
      "stem": "The log likelihood of the Poisson model is given by which of the following expressions?\n\n         A. log L = ∑n1 yi log yi + ∑n1 yi + ∑n1 log yi !\n         B. log L = ∑n1 yi log yi + ∑n1 yi − ∑n1 log yi !\n         C. log L = ∑n1 yi log yi − ∑n1 yi + ∑n1 log yi !\n         D. log L = ∑n1 yi log yi − ∑n1 yi − ∑n1 log yi !.                                                  [2]",
      "parts": [],
      "solution": "Correct Answer is Option D.                                                                     [2]\n\n               The log-likelihood of the Poisson model can be derived as follows:",
      "has_math": true,
      "session": "2024-11",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2024-11_QP.pdf",
      "source_sol": "raw/CS1A_2024-11_SOL.pdf"
    },
    {
      "q_num": 50,
      "marks": 2,
      "topic": "regression_glm",
      "subtopics": [],
      "stem": "Based on the fitted values of the compatibility score (𝑦̂) and the actual value of compatibility\n         score (y), following information has been calculated:\n\n                          ∑n1 yi log( yi / ŷi ) = 28.6132,   ∑n1( yi − ŷi ) = 16.2375.\n\n         Scaled deviance for the fitted model is –\n\n         A. 20.7672\n         B. 24.7514\n         C. 24.8475\n         D. 26.4799.                                                                                        [2]",
      "parts": [],
      "solution": "Correct Answer is Option B.                                                                     [2]\n\n               Scaled Deviance\n               =D\n               = 2(𝑙𝑜𝑔 𝐿𝑆 − 𝑙𝑜𝑔 𝐿𝑀 )\n               = 2( ∑𝑛1 𝑦𝑖 𝑙𝑜𝑔 𝑦𝑖 − ∑𝑛1 𝑦𝑖 − ∑𝑛1 log 𝑦𝑖 ! − ∑𝑛1 𝑦𝑖 𝑙𝑜𝑔 𝑦̂𝑖 + ∑𝑛1 𝑦̂𝑖 + ∑𝑛1 log 𝑦𝑖 ! )     =\n               2( ∑𝑛1 𝑦𝑖 𝑙𝑜𝑔( 𝑦𝑖 / 𝑦̂𝑖 ) − ∑𝑛1( 𝑦𝑖 − 𝑦̂𝑖 ) )\n               = 2 * (28.6132 – 16.2375)\n               = 24.7514.\n\n                                                *************\n\n                                                                                                   Page 13 of 13",
      "has_math": true,
      "session": "2024-11",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2024-11_QP.pdf",
      "source_sol": "raw/CS1A_2024-11_SOL.pdf"
    },
    {
      "q_num": 1,
      "marks": 2,
      "topic": "distributions",
      "subtopics": [],
      "stem": "Let X be a random variable denoting the number of hours you spent on scrolling random\n     social media reels during the day. It has the following cumulant generating function:\n\n                                          C X (t) = 2t + 3t2\n      Which one of the following is correct?\n      A. E (X) = 2 and Var (X) = 3\n      B. E (X) = 2 and Var (X) = 6\n      C. E (X) = 3 and Var (X) = 3\n      D. E (X) = 2 and Var (X) = 9\n      E. E (X) = 3 and Var (X) = 3",
      "parts": [],
      "solution": null,
      "has_math": false,
      "session": "2026-05",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2026-05_QP.pdf",
      "source_sol": null
    },
    {
      "q_num": 2,
      "marks": 2,
      "topic": "inference",
      "subtopics": [],
      "stem": "The game of musical chairs is played in the following way:\n\n      •   The first round of the game starts with (n – 1) chairs and n players\n      •   Music is played and players have to go in circle around the chairs\n      •   Music stops and players are supposed to occupy a chair nearest to them\n      •   One player out of the n players is eliminated in that round\n      •   Next round of the game continues with (n – 2) chairs and (n – 1) players\n      •   Game continues till we get the final winner\n\n      What is the probability of an individual winning the game of musical chairs?\n      A. 1 / (n – 1)\n      B. (n – 2) / (n – 1)\n\n      C. n / (n – 1)\n      D. (n – 1) / n\n      E. 1 / n",
      "parts": [],
      "solution": null,
      "has_math": false,
      "session": "2026-05",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2026-05_QP.pdf",
      "source_sol": null
    },
    {
      "q_num": 3,
      "marks": 2,
      "topic": "regression_glm",
      "subtopics": [],
      "stem": "A civil services aspirant has developed a multi-variate linear regression model which\n     predicts the UPSC mains score (Y) based on two variables:\n      •   X1: hours of study per week\n      •   X2: number of mock tests taken\n\n      The multiple linear regression model is given by:\n      Y = α + β1 X1 + β2 X2\n\n      If α = 150, β1 = 3 and β2 = 2, which one of the following statements is factually correct?\n      A. UPSC mains score would be exactly twice the number of mock tests taken by the\n         aspirant.\n      B. For every extra hour of study per week, UPSC mains score increases by 2 marks, holding\n         number of mock tests taken constant.\n      C. For every extra mock test taken, UPSC mains score increases by 2 marks, holding weekly\n         hours of study constant.\n      D. For every extra mock test taken, UPSC mains score increases by 2 marks, even when the\n         aspirant alters his weekly hours of study.\n      E. An aspirant who has not appeared for any mock test, will score zero marks in UPSC\n         mains examination.",
      "parts": [],
      "solution": null,
      "has_math": true,
      "session": "2026-05",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2026-05_QP.pdf",
      "source_sol": null
    },
    {
      "q_num": 4,
      "marks": 2,
      "topic": "inference",
      "subtopics": [],
      "stem": "In an online retail platform:\n      •   A = number of orders placed in an hour\n      •   B = number of orders with delivery complaints in that hour\n\n      The joint probability distribution of A and B is given by:\n      P(A = a, B = b) = (a+3b)/75,\n      where\n      a = 0, 1, 2, 3, 4 and b=0, 1, 2.\n      Given that at least one delivery complaint occurred (B≥1), what is the expected number of\n      complaints?\n      A. 1.38\n      B. 1.46\n      C. 1.54\n      D. 1.62\n      E. 1.70",
      "parts": [],
      "solution": null,
      "has_math": true,
      "session": "2026-05",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2026-05_QP.pdf",
      "source_sol": null
    },
    {
      "q_num": 5,
      "marks": 2,
      "topic": "inference",
      "subtopics": [
        "distributions"
      ],
      "stem": "A bank samples 36 customers and finds mean withdrawal = 2,800 with sample variance =\n     400,000.\n      What is the standard error of the sample mean?\n      A. 632.54\n      B. 105.41\n      C. 111.17\n      D. 95.02\n      E. 120.09",
      "parts": [],
      "solution": null,
      "has_math": false,
      "session": "2026-05",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2026-05_QP.pdf",
      "source_sol": null
    },
    {
      "q_num": 6,
      "marks": 2,
      "topic": "distributions",
      "subtopics": [
        "regression_glm"
      ],
      "stem": "Let the rental rate per square feet in a particular locality be denoted by random variable Y.\n     Y follows a lognormal distribution with a known parameter σ2 = 0.16.\n      Y ~ Lognormal (μ, 0.16) i.e. lnY ~ N (μ, 0.16)\n      Let the floor level in the building be denoted by random variable X.\n      We model the mean of lnY in a generalised linear model with a single predictor X using the\n      identity link function:\n\n                                        E[lnY] = μ = β0 + β1 * X\n       Akshar has finalised a flat on the second floor of a building and is comfortable with the\n      rental rate. However, in the same building there is a flat on the fourth floor which he likes\n      but may not be able to afford.\n\n      What is the percentage increase in expected rental rate per square feet i.e. E[Y], if Akshar\n      decides to go with the flat on the fourth floor instead of the flat on the second floor?\n      Assume β1 = 0.5.\n      A. 50%\n      B. 100%\n      C. 135%\n      D. 172%\n      E. 200%",
      "parts": [],
      "solution": null,
      "has_math": true,
      "session": "2026-05",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2026-05_QP.pdf",
      "source_sol": null
    },
    {
      "q_num": 7,
      "marks": 2,
      "topic": "distributions",
      "subtopics": [],
      "stem": "Let X denote the annual budgetary allowance for a particular department in the organization\n     and Y denote the actual spendings by the department during the year.\n      Let U = X – Y such that U denotes the amount of savings during the year.\n      Which of the following expressions correctly represents the moment generating function for\n      the random variable U i.e. M U (t)?\n      Hint: [By definition, joint MGF of (X, Y) is M X, Y (s1, s2) = E [exp (s1X + s2Y)].]\n      A. M X,Y (t,−t)\n      B. M X,Y (t, t)\n\n       C. M X (t) – M Y (t)\n       D. M X (t) * M Y (−t)\n       E. log e (M X,Y (t,−t))",
      "parts": [],
      "solution": null,
      "has_math": true,
      "session": "2026-05",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2026-05_QP.pdf",
      "source_sol": null
    },
    {
      "q_num": 8,
      "marks": 2,
      "topic": "inference",
      "subtopics": [],
      "stem": "A retinologist knows that 2% of patients undergoing a particular eye screening test have a\n     rare eye condition.\n       In a given year, she screens 2,000 patients. Using the Central Limit Theorem (CLT), what\n       is the probability that more than 50 patients are found to have this condition?\n       A. 0.15\n       B. 0.08\n       C. 0.05\n       D. 0.12\n       E. 0.20",
      "parts": [],
      "solution": null,
      "has_math": false,
      "session": "2026-05",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2026-05_QP.pdf",
      "source_sol": null
    },
    {
      "q_num": 9,
      "marks": 2,
      "topic": "inference",
      "subtopics": [],
      "stem": "In a game of ludo, a player needs to roll ‘6’ on a fair six-sided die to move the token out of\n     the starting square. Each roll is assumed to be independent.\n       Let X be the number of rolls needed to get the first ‘6’.\n       What is the probability that the player rolls the die exactly 3 times to roll the first ‘6’?\n       A. 1/36\n       B. 125/216\n       C. 5/18\n       D. 1/6\n       E. 25/216",
      "parts": [],
      "solution": null,
      "has_math": false,
      "session": "2026-05",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2026-05_QP.pdf",
      "source_sol": null
    },
    {
      "q_num": 10,
      "marks": 2,
      "topic": "regression_glm",
      "subtopics": [
        "distributions"
      ],
      "stem": "Consider the following pairs relating to canonical link functions of some of the common\n      distributions belonging to the exponential family of distributions:\n\n      1.   Poisson distribution – log link function\n      2.   Gamma distribution – inverse link function\n      3.   Binomial distribution – identity link function\n      4.   Normal distribution – logit link function\n      Which of the above pairs are correct?\n      A. 1 and 3 only\n      B. 1 and 2 only\n      C. 1, 2 and 3 only\n      D. 1, 3 and 4 only\n\n      E. 1, 2, 3 and 4",
      "parts": [],
      "solution": null,
      "has_math": false,
      "session": "2026-05",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2026-05_QP.pdf",
      "source_sol": null
    },
    {
      "q_num": 11,
      "marks": 2,
      "topic": "distributions",
      "subtopics": [],
      "stem": "An insurer provides cyber risk cover to small technology firms. For a particular firm –\n      •    The number of cyber incidents in a year follows a Poisson distribution with mean 2.70.\n      •    Each incident results in a loss amount independent of other incidents, with mean of 7,350\n           and standard deviation of 5,120.\n      If the insurer wants to set a premium equal to expected loss plus 30% of the standard\n      deviation, what premium should the insurer charge in respect of the above firm?\n\n      A. 22,154\n      B. 24,261\n      C. 23,502\n      D. 25,103\n      E. 26,000",
      "parts": [],
      "solution": null,
      "has_math": false,
      "session": "2026-05",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2026-05_QP.pdf",
      "source_sol": null
    },
    {
      "q_num": 12,
      "marks": 2,
      "topic": "distributions",
      "subtopics": [],
      "stem": "Let X denote the height of fellow actuaries attending the global conference of actuaries in\n      the year 2026.\n           A random sample of nine actuaries (X1, X2, …………, X9) is drawn from a N (0, σ2)\n          distribution. Let\n\n           What is the value of\n\n          A. 0.75%\n          B. 0.81%\n          C. 0.89%\n          D. 0.93%\n          E. 0.97%",
      "parts": [],
      "solution": null,
      "has_math": true,
      "session": "2026-05",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2026-05_QP.pdf",
      "source_sol": null
    },
    {
      "q_num": 13,
      "marks": 2,
      "topic": "distributions",
      "subtopics": [],
      "stem": "Let XS, XB and XJ represent the time taken by ratha yatras of Subhadra-ji, Balaram-ji and\n     Jagannath-ji during the annual ratha yatra festival in Puri. XS, XB and XJ are assumed to\n     follow normal distribution with μS = μB = μJ = 5 hours. However, they do not have equal\n     σ2 .\n          The probability density curves for each of the variables are given below:\n\n      As seen in the above graph, XS has the tallest and narrowest curve whereas XJ has the\n      shortest and widest curve and XB lies in between.\n      A new devotee wants to attend the procession this year and wants to know the probability\n      that the ratha yatra time exceeds 6 hours.\n      Which of the following is TRUE in this context?\n      A. P (XJ > 6) < P (XB > 6) < P (XS > 6)\n      B. P (XS > 6) > P (XJ > 6) > P (XB > 6)\n      C. P (XJ > 6) < P (XS > 6) < P (XB > 6)\n      D. P (XJ > 6) > P (XB > 6) > P (XS > 6)\n      E. P (XJ > 6) = P (XB > 6) = P (XS > 6)",
      "parts": [],
      "solution": null,
      "has_math": true,
      "session": "2026-05",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2026-05_QP.pdf",
      "source_sol": null
    },
    {
      "q_num": 14,
      "marks": 2,
      "topic": "bayes_credibility",
      "subtopics": [],
      "stem": "Charlie uses a dating app and each month he has a probability θ of getting a match,\n      independently from month to month.\n         His prior beliefs about θ follow a Beta (α, β) distribution. Over n months, he observes a\n         total of x matches.\n         It is known that the posterior distribution of θ is θ | x ~ Beta (α + x, β + n – x).\n         If α = 2, β = 4 and n = 12, the credibility factor for Charlie’s estimate of θ is\n         A. 0.6667\n         B. 0.8571\n         C. 1.0000\n         D. 0.2868\n         E. 0.8183",
      "parts": [],
      "solution": null,
      "has_math": true,
      "session": "2026-05",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2026-05_QP.pdf",
      "source_sol": null
    },
    {
      "q_num": 15,
      "marks": 2,
      "topic": "distributions",
      "subtopics": [],
      "stem": "In a cosmetics store, the joint probability density function of number of units of skincare\n     cream sold per day (X) and price per unit (P) is given by:\n\n                            f X,P (x,p)= (x+p) / 300,   0 ≤ x ≤ 10, 0 ≤ p ≤20\n\n        What is the marginal probability density function of X (no of units sold)?\n         A. (x + 20) / 30\n         B. (x + 20) / 300\n         C. (x + 200) / 300\n         D. (x + 10) / 300\n         E. (x + 10) / 15",
      "parts": [],
      "solution": null,
      "has_math": true,
      "session": "2026-05",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2026-05_QP.pdf",
      "source_sol": null
    },
    {
      "q_num": 16,
      "marks": 20,
      "topic": "bayes_credibility",
      "subtopics": [
        "distributions"
      ],
      "stem": "During an international conflict, a country's border security agency is monitoring the\n      arrival of unmanned reconnaissance drones crossing a sensitive air corridor.\n        Intelligence analysts are interested in modelling the time between consecutive drone\n        detections at a radar checkpoint.\n\n        Let T denote the time (in minutes) between successive drone detections. Based on\n        historical radar data, analysts model this time using an exponential distribution\n                                                T ∼ Exp (λ)\n          where λ is the unknown rate parameter representing the average detection rate of drones\n\n         To estimate λ analysts adopt a Bayesian approach. Before observing the new radar data,\n         they assume a Gamma prior distribution for the rate parameter.\n\n                                            λ ∼ Gamma (α, β)\n          where α and β are known parameters.\n\n      Suppose that over a monitoring period the radar system records the arrival times between n\n      consecutive drone detections as t1,t2,…...,tn.",
      "parts": [
        {
          "label": "i",
          "marks": 2,
          "text": "Write down the likelihood function for the sample t1, t2, …….., tn given λ.              [2]",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 4,
          "text": "Using the likelihood as determined in part (i) and the Gamma prior, show that the\n          posterior distribution of λ is also a Gamma distribution and verify that its parameters are:\n          λ | t ~ Gamma (α + n, β + Σ tᵢ)",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 2,
          "text": "Assume that 20 drone detections have been recorded during the observation period.\n\n         and gamma prior has the parameters α = 2 and β = 1.\n\n         Determine the Bayesian estimate for λ under quadratic loss.                              [2]",
          "topic": null
        },
        {
          "label": "iv",
          "marks": 10,
          "text": "A plot of the posterior distribution has been given below and two points A (xA, yA) and B\n          (xB, yB) have been marked on the graph:\n\nAt Point A, the value of the gamma density function is the highest, whereas the perpendicular on\nX axis from Point B divides the area under the gamma density curve into exactly two equal\nhalves.\n\nState, with reasons, whether the following statements are correct or incorrect:\n      •   xA represents the Bayesian estimate under all or nothing loss;\n      •   xB represents the Bayesian estimate under absolute error loss.",
          "topic": null
        }
      ],
      "solution": null,
      "has_math": true,
      "session": "2026-05",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2026-05_QP.pdf",
      "source_sol": null
    },
    {
      "q_num": 17,
      "marks": 20,
      "topic": "inference",
      "subtopics": [],
      "stem": "An international airport records the number of passengers passing through a particular\nsecurity checkpoint each hour. Over a monitoring period of 500 hours, the following data were\nobserved:\n            No of passengers       0   1   2 3 4 ≥ 5 Total\n            No of hours           280 150 50 15 4 1   500",
      "parts": [
        {
          "label": "i",
          "marks": 2,
          "text": "Calculate the sample mean and sample variance for the data.                               [2]",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 2,
          "text": "Using the method of moments, estimate the parameter of Poisson distribution used to fit\n          to the above data.\n          Hence, calculate the expected number of hours with 0, 1, 2, 3, 4, and ≥ 5 passengers\n          assuming a Poisson model.                                                        [2]",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 2,
          "text": "Based on your answers in parts (i) and (ii), briefly discuss why a Poisson model may not\n           be an appropriate fit for the data.                                                  [2]",
          "topic": null
        },
        {
          "label": "iv",
          "marks": 2,
          "text": "Instead of fitting a Poisson distribution to the data, it is decided to use bootstrapping for\n          the purpose of summarising the data.\n\n          Following five bootstrap samples viz. B1 to B5 have been generated for this purpose.\n\n           No of passengers          0        1         2         3        4       ≥5       Total\n                  Sample B1         278      152       49        14        5        2        500\n                  Sample B2         285      145       52        13        4        1        500\n         No of\n                  Sample B3         276      154       51        15        3        1        500\n         hours\n                  Sample B4         282      148       50        16        3        1        500\n                  Sample B5         279      151       48        15        6        1        500\n\n         Calculate the expected number of hours with 0, 1, 2, 3, 4, and ≥ 5 passengers using the\n         bootstrap samples generated above.                                                  [2]",
          "topic": null
        },
        {
          "label": "v",
          "marks": 10,
          "text": "Determine whether the method used for generating these samples is parametric (using\n         Poisson parameter determined in part (ii)) or non-parametric. Justify your answer.   [2]",
          "topic": null
        }
      ],
      "solution": null,
      "has_math": true,
      "session": "2026-05",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2026-05_QP.pdf",
      "source_sol": null
    },
    {
      "q_num": 18,
      "marks": 20,
      "topic": "inference",
      "subtopics": [],
      "stem": "A wildlife research unit studying orangutan rehabilitation in Arithmetica examines how\n     annual structured training hours (x) affect the number of minor survival incidents (y)\n     observed post-release. Data from 10 orangutans is given below:\n\n∑ x = 770,      ∑ y = 35,   ∑ x2 = 61,600,    ∑ y2 = 145,   ∑ xy = 2,470",
      "parts": [
        {
          "label": "i",
          "marks": 2,
          "text": "Determine the equation of the simple regression line.                                     [2]",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 4,
          "text": "Perform an ANOVA test to determine whether the slope of the regression line is\n          significantly different from zero. Perform the test at 1% level of significance. [4]",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 2,
          "text": "An actuarial student who has recently joined the research unit explores several\n           transformations of the standard linear regression model to improve model fit. Four such\n           transformations are considered, and the coefficient of determination R2is computed for\n           each.\n\n  Calculate R2 for the standard linear regression model and comment on which model out of the\n  five models (one standard model and four transformations of the standard model) is the best\n  fit to the orangutan rehabilitation data.                                               [2]",
          "topic": null
        },
        {
          "label": "iv",
          "marks": 10,
          "text": "Using the linear regression model obtained in part (i), determine the value of the residual\n          term for orangutan E.                                                                   [2]",
          "topic": null
        }
      ],
      "solution": null,
      "has_math": true,
      "session": "2026-05",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2026-05_QP.pdf",
      "source_sol": null
    },
    {
      "q_num": 19,
      "marks": 20,
      "topic": "inference",
      "subtopics": [],
      "stem": "At an amusement park in Mathematica, engineers are testing the reliability of a\n     giantwheel’s maximum height in metres. Due to variations in wind resistance, passenger\n     distribution and mechanical performance, the height fluctuates slightly between the rides.\n\nThe radius of the giant wheel is 20 metres and the perpendicular distance between the centre and\nground level is 22 metres. Based on a sample collected over the past 300 rides on the giant\nwheel, the mean distance between the centre and the maximum height reached by the giant\nwheel is 19 metres and it has a standard deviation of 3 metres.",
      "parts": [
        {
          "label": "i",
          "marks": 2,
          "text": "Calculate a 95% two-sided confidence interval for the maximum height achieved by the giant\n   wheel. (Assume critical value of t-distribution to be 1.96).                           [2]",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 3,
          "text": "Calculate a 95% two-sided prediction interval for the maximum height to be achieved in the\n     next ride of the giant wheel. Comment on why the 95% prediction interval is wider than the\n     95% confidence interval. (Assume critical value of t-distribution to be 1.96)             [3]",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 2,
          "text": "The Government of Mathematica has come up with a new guideline relating to places of\n     public entertainment which makes it mandatory for giant wheels to have a distance of at least\n     4 metres from the ground level and caps the maximum height that the giant wheel can\n     achieve to 36 metres.\n\n      The giant wheel is restructured in order to comply with the regulations and based on 300\n      rides on the restructured giant wheel, the mean maximum height reached is 35 metres with a\n      standard deviation of 2 metres.\n\n      Assuming normality, calculate the probability that a ride on the restructured wheel exceeds\n      36 metres, and comment on whether the restructured wheel complies with the new guideline\n      at the 5% level of significance.                                                        [2]",
          "topic": null
        },
        {
          "label": "iv",
          "marks": 10,
          "text": "Calculate a 95% two-sided confidence interval for the difference in the mean maximum\n    heights of the original and the restructured giant wheel. (Assume critical value of t-\n    distribution to be 1.96)                                                           [3]",
          "topic": null
        }
      ],
      "solution": null,
      "has_math": false,
      "session": "2026-05",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2026-05_QP.pdf",
      "source_sol": null
    },
    {
      "q_num": 20,
      "marks": 20,
      "topic": "inference",
      "subtopics": [],
      "stem": "A health actuary wants to model mental health score (MHS) which is measured on a\n     continuous scale of 0 to 100. He has chosen to use generalised linear models for this\n     purpose and has identified the following five predictors to be used for calculating MHS:\n\n      1. REM: Extent of REM sleep (hours / night)\n      2. DPM: Dopamine levels in blood (ng/mL)\n      3. MDT: Number of hours spent in mediation / mindfulness per day\n      4. MRT: Marital status (1 = single, 2 = married, 3 = divorced, 4 = widowed)\n      5. EMP: Employment status (1= student, 2 = employed, 3 = unemployed, 4 = retired, 5 =\n         self-employed)",
      "parts": [
        {
          "label": "i",
          "marks": 1,
          "text": "Which of the above predictors are factor variables? Justify your answer.               [1]",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 7,
          "text": "The actuary is trying to use five different models for estimating the mental health score\n          (MHS):\n\n      •   Model 1: REM + DPM\n      •   Model 2: REM + DPM + MDT + MRT\n      •   Model 3: REM + DPM + MDT + MRT + MDT:MRT\n      •   Model 4: REM + DPM + MDT + MRT + EMP + MDT:MRT + MRT:EMP + MDT:EMP\n      •   Model 5: MRT*EMP*REM\n\nwhere –\n      ⎯ predictor1: predictor2 means the interaction term between the two predictors\n      ⎯ predictor1*predictor2 includes the main effect as well as the interaction terms between\n        the two predictors\n      ⎯ base scenario reflects marital status = 1 and employment status = 1\n      ⎯ intercept parameter captures the main effects relating to base scenario\n      Determine the number of parameters in respect of each of the above models.                  [7]",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 10,
          "text": "Based on data collected for a cohort of policyholders, following values are generated for\n           each of the above models including a saturated model (denoted as Model 0) as presented\n           in the below table:\n\n                      Scaled          Difference in Scaled        Critical Point at 5% level of\n          Model\n                     Deviance               Deviance                       significance\n            0           99\n            1           91                       8                             5.99\n            2           80                      11                             9.49\n            3           71                       9                             7.82\n\n              4           51                      20                             31.41\n              5           34                      17                             18.31\n\n        Identify the model which is the best fit to the data at the 5% level of significance.     [2]",
          "topic": null
        }
      ],
      "solution": null,
      "has_math": false,
      "session": "2026-05",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2026-05_QP.pdf",
      "source_sol": null
    },
    {
      "q_num": 21,
      "marks": 20,
      "topic": "inference",
      "subtopics": [],
      "stem": "A non-life insurer introduces a fraud‑detection system to identify suspicious motor\n      insurance claims. The system flags a claim as “fraudulent” if it exceeds certain risk\n      thresholds. In reality, only 8% of claims are fraudulent.\n\n         During the observation period, a total of 10,000 claims (fraudulent and genuine\n         combined) are reported.\n\n         Following matrix has been developed in respect of the same:\n\n                               Predicted Fraud       Predicted Genuine           Total\n             Actual Fraud        True Positives        False Negatives      Actual Positives\n         Actual Genuine         False Positives        True Negatives       Actual Negatives\n                  Total        Predicted Positives Predicted Negatives           Total\n\n         The insurer has identified that –\n\n         •     Sensitivity = 85% (correctly flags fraud when present) and\n         •     Specificity = 90% (correctly clears genuine claims)",
      "parts": [
        {
          "label": "i",
          "marks": 1,
          "text": "Explain Type I error and Type II error in the context of this fraud‑detection system. What\n   do these errors mean for the insurer’s operations?                                     [1]",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 1,
          "text": "Calculate –\n\n      a) Expected number of fraudulent claims correctly flagged (true positives).\n      b) Expected number of genuine claims wrongly flagged (false positives)",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 4,
          "text": "The insurer wants to formally test whether the fraud‑detection system’s sensitivity is\n     significantly greater than 80% and is working on the following hypotheses:\n\n                                        H0: p = 0.80 v. H1: p > 0.80\n      where p is the true sensitivity\n\n      In the last year, out of a total of 800 actually fraudulent claims, if 680 were correctly\n      flagged, perform the hypothesis test at 1% level of significance.\n      Hint: [Use chi-square test.]                                                                [4]",
          "topic": null
        },
        {
          "label": "iv",
          "marks": 10,
          "text": "Construct a 95% two-sided confidence interval for the fraud detection system’s sensitivity.\n    Also comment briefly on whether this confidence interval supports the claim that sensitivity\n    exceeds 80%.                                                                             [4]",
          "topic": null
        }
      ],
      "solution": null,
      "has_math": false,
      "session": "2026-05",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2026-05_QP.pdf",
      "source_sol": null
    },
    {
      "q_num": 22,
      "marks": 20,
      "topic": "inference",
      "subtopics": [],
      "stem": "A space agency has collected data from 10 spacecraft missions to Jupiter to study the\n     factors affecting mission performance. The dataset includes the following variables:",
      "parts": [
        {
          "label": "i",
          "marks": 2,
          "text": "Identify the right option which correctly describes the nature of the dataset. Justify your\n   answer:\n\nA. Longitudinal data;\nB. Censored data;\nC. Cross-sectional data;\nD. Truncated data.",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 2,
          "text": "Calculate the Karl Pearson’s correlation coefficient between launch mass (x) and travel time\n    (y). Also, interpret the results in terms of mission planning.\n\n      ∑ x = 41.02,   ∑ y = 267.70,   ∑ x2 = 184.0878,   ∑ y2 = 7493.87,    ∑ xy = 1106.348",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 3,
          "text": "It is decided to calculate the Kendall’s correlation coefficient between number of instruments\n     and mission cost. In this context, it is observed that –\n\n      •   the number of instances “where higher number of instruments lead to higher mission\n          costs”\n                                                   is lower than\n      •   the number of instances “where higher number of instruments leading to lower mission\n          costs”\nWithout performing any additional calculations, answer the following questions:\n\n      a) What do you mean by concordant and discordant pairs in this context?\n\n      b) Will Kendal’s correlation coefficient between number of instruments and mission cost be\n         positive or negative in this case? Justify your answer.\n\n      c) In the next five years, another ten space missions are expected to be launched to Jupiter.\n         However, the space missions are expected to be designed with higher number of state-of-\n         art high-cost instruments. Will this improve the Kendal’s correlation coefficient?",
          "topic": null
        },
        {
          "label": "iv",
          "marks": 10,
          "text": "Without performing any calculations, explain how the presence of an outlier might affect the\n    slope of a simple linear regression line fitted between launch mass and travel time.     [3]",
          "topic": null
        }
      ],
      "solution": null,
      "has_math": true,
      "session": "2026-05",
      "subject": "CS1A",
      "source_qp": "raw/CS1A_2026-05_QP.pdf",
      "source_sol": null
    },
    {
      "q_num": 1,
      "marks": 30,
      "topic": "inference",
      "subtopics": [],
      "stem": "With reference to the dataset “AutoClaims.csv”, answer the following questions.",
      "parts": [
        {
          "label": "i",
          "marks": 10,
          "text": "Prepare a table with mean, standard deviation and coefficient of variance (CV) of the\n               claims paid\n               a)     for each STATE and identify the one with the least and the highest CV\n               b)     for each CLASS and sort the table in the ascending order of the CV\n\n               Hint: Coefficient of Variance (CV) can be computed as the ratio of standard deviation\n               to the mean.                                                                               (10)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 5,
          "text": "By using a box plot, identify the STATE(s) and CLASS(es) which have no outlier\n               values in terms of the claims paid.                                                         (5)",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 10,
          "text": "Prepare a contingency table identifying the proportion of insured across different\n               CLASS(es) for each gender. Test the null hypothesis “There is no relationship between\n               the GENDER and the CLASS of the insured” against an alternate hypothesis of\n               existence of significant relation at 95% confidence level. Perform an appropriate test\n               and comment.                                                                               (10)",
          "topic": null
        },
        {
          "label": "iv",
          "marks": 5,
          "text": "Use an appropriate test to test if the mean claim amount paid and variance of claim\n               amounts paid is uniform between males and females.                                          (5)",
          "topic": null
        }
      ],
      "solution": "i)\nstate_mean<-aggregate(PAID~STATE,data = AutoClaims,FUN = mean)\nnames(state_mean)<-c(\"State\",\"Mean\")\nstate_sd<-aggregate(PAID~STATE,data = AutoClaims,FUN = sd)\nnames(state_sd)<-c(\"State\",\"SD\")\nstate_summary<-merge(state_mean,state_sd)\nstate_summary$CV<-state_summary$SD/state_summary$Mean\n#Mean, Standard Deviation and Coefficient of Variance for each state\nstate_summary                                                                             [4]\n##    State Mean      SD     CV\n## 1 STATE 01 10235.800 10932.877 1.0681018\n## 2 STATE 02 7055.078 6327.473 0.8968678\n## 3 STATE 03 8714.932 6494.346 0.7451976\n## 4 STATE 04 8152.759 6985.210 0.8567910\n## 5 STATE 06 8786.739 10749.517 1.2233796\n## 6 STATE 07 4960.479 3065.092 0.6179023\n## 7 STATE 10 12340.643 14599.291 1.1830251\n## 8 STATE 12 6893.705 8634.955 1.2525856\n## 9 STATE 14 10399.313 8406.388 0.8083599\n## 10 STATE 15 3321.449 3364.269 1.0128920\n## 11 STATE 17 7886.282 7831.913 0.9931059\nNote: One mark is awarded for each correct column\n# a. The State with minimum coefficient of variance is\nstate_summary$State[state_summary$CV==min(state_summary$CV)]\n## [1] STATE 07                                                                         [0.5]\n## 11 Levels: STATE 01 STATE 02 STATE 03 STATE 04 STATE 06 ... STATE 17\n# b. The State with maximum coefficient of variance is\nstate_summary$State[state_summary$CV==max(state_summary$CV)]\n## [1] STATE 12                                                                         [0.5]\n## 11 Levels: STATE 01 STATE 02 STATE 03 STATE 04 STATE 06 ... STATE 17\nNote: Even if the code is not present for finding the state with max/min CV and the student\nhas written the answer by observation, mark can be awarded\nClass_mean<-aggregate(PAID~CLASS,data = AutoClaims,FUN = mean)\nnames(Class_mean)<-c(\"Class\",\"Mean\")\nClass_sd<-aggregate(PAID~CLASS,data = AutoClaims,FUN = sd)\nnames(Class_sd)<-c(\"Class\",\"SD\")\nClass_summary<-merge(Class_mean,Class_sd)\nClass_summary$CV<-Class_summary$SD/Class_summary$Mean\n#Mean, Standard Deviation and Coefficient of variance for each rating class\nClass_summary                                                                             [3]\n## Class Mean      SD    CV\n## 1 C1 17464.484 9613.323 0.5504499\n## 2 C11 5887.049 3434.371 0.5833774\n\n                                                                                Page 1 of 13\n\fIAI                                                                                CS1B-0619\n\n## 3 C6 2367.215 1533.807 0.6479373\n## 4 F6 17434.533 9188.423 0.5270243\n# Arranging the rating class in ascending order of CV\nClass_summary[order(Class_summary$CV),]                                                    [2]\n## Class Mean      SD    CV\n## 4 F6 17434.533 9188.423 0.5270243\n## 1 C1 17464.484 9613.323 0.5504499\n## 2 C11 5887.049 3434.371 0.5833774\n## 3 C6 2367.215 1533.807 0.6479373\nNote: Full marks (5) are awarded if the student directly shows the table in ascending of the\nCVs\nii)\nboxplot(PAID~STATE,data = AutoClaims)                                                    [1.5]\n\nStates with no outliers are State 14.                                                      [1]\nboxplot(PAID~CLASS,data = AutoClaims)                                                    [1.5]\n\n Class with no outlier is C11.                                                             [1]\n\n                                                                                 Page 2 of 13\n\fIAI                                                                                 CS1B-0619\n\niii)\n#Contingency Table for the distribution of Gender across rating classes\nState_Class_Freq<-table(AutoClaims$GENDER,AutoClaims$CLASS)                                 [2]\n#Contingency Table in terms of proportion of Rating classes for each gender\nState_Class_RowProp<-prop.table(State_Class_Freq,margin = 1)\nState_Class_RowProp                                                                         [2]\n##\n##      C1    C11     C6     F6\n## F 0.07312925 0.30102041 0.55272109 0.07312925\n## M 0.07840772 0.34620024 0.46200241 0.11338963\n#Chi-Square test for independence of two variables\nchisq.test(AutoClaims$GENDER,AutoClaims$CLASS)                                              [3]\n##\n## Pearson's Chi-squared test\n##\n## data: AutoClaims$GENDER and AutoClaims$CLASS\n## X-squared = 13.704, df = 3, p-value = 0.003338\nInterpretation: Chi-Squared values is 13.704 with 3 degrees of freedom. P-Value for the chi\nsquare test is 0.0033 < 0.05. Hence the null hypothesis of independence between the\ngender and rating class is rejected at 95% confidence level.                             [2]\nLooking at the proportions in the table, it is evident that more than 55% of the females are\nin rating class C6 whereas as only 46% of the males are in that rating class. Similarly 34.6%\nof the males are in C11 as against only 30% of females in that class. Same is the case with\nF6. Hence the difference in proportions                                                      [1]\niv)\n#Testing for mean amount paid between Male and Female\nt.test(AutoClaims$PAID[AutoClaims$GENDER==\"M\"],AutoClaims$PAID[AutoClaims$GENDE\nR==\"F\"])                                                                     [1.5]\n##\n## Welch Two Sample t-test\n##\n## data: AutoClaims$PAID[AutoClaims$GENDER == \"M\"] and\nAutoClaims$PAID[AutoClaims$GENDER == \"F\"]\n## t = -1.0808, df = 1165.9, p-value = 0.28\n## alternative hypothesis: true difference in means is not equal to 0\n## 95 percent confidence interval:\n## -1176.7935 340.7988\n## sample estimates:\n## mean of x mean of y\n## 5953.770 6371.767\nP-value = 0.28 > 0.05  Null hypothesis of “Mean claim paid is same between males and\nfemales” cannot be rejected at 95% confidence level                                   [1]\n#Testing for variance of amount paid between Male and Female\nvar.test(AutoClaims$PAID[AutoClaims$GENDER==\"M\"],AutoClaims$PAID[AutoClaims$GEND\nER==\"F\"])                                                                    [1.5]\n                                                                                   Page 3 of 13\n\fIAI                                                                                 CS1B-0619\n\n##\n## F test to compare two variances\n##\n## data: AutoClaims$PAID[AutoClaims$GENDER == \"M\"] and\nAutoClaims$PAID[AutoClaims$GENDER == \"F\"]\n## F = 0.78416, num df = 828, denom df = 587, p-value = 0.001337\n## alternative hypothesis: true ratio of variances is not equal to 1\n## 95 percent confidence interval:\n## 0.6744929 0.9098948\n## sample estimates:\n## ratio of variances\n##      0.7841591\nP-value = 0.0013 < 0.05  Null hypothesis of “Variance of claim paid is same between males\nand females” is rejected at 95% confidence level. Claims paid among females is more\nvolatile compared to that of males                                                      [1]\nPart (i) was well answered by the students who were successful. Students struggled with\ncomputing the coefficient of variance and interpreting based on that. In part (ii), the students\nwere able to use the R Code to generate the box plot. Only a few of them were able to identify\noutliers successfully. In Part (iii), many students were able to perform Chi square test and do\nthe interpretation successfully but majority of them failed in preparing contingency tables\nbased on proportions. Also majority of them failed in making detailed interpretation based on\nthe results of the chi-square test. Part (iv) was hardly attempted. Very few students were able\nto perform the correct tests and provide right interpretation. A number of students had\ndifficulty in identifying the correct test itself.",
      "has_math": false,
      "is_r_task": true,
      "session": "2019-06",
      "subject": "CS1B",
      "source_qp": "raw/CS1B_2019-06_QP.pdf",
      "source_sol": "raw/CS1B_2019-06_SOL.pdf"
    },
    {
      "q_num": 2,
      "marks": 25,
      "topic": "distributions",
      "subtopics": [],
      "stem": "Refer to the dataset “AutoClaims.csv” and answer the following questions.\n\n        You were asked to fit an appropriate distribution to the “PAID” data. You decided to fit\n        Normal distribution, Lognormal distribution, Exponential distribution and Gamma\n        Distribution based on the method of moments.",
      "parts": [
        {
          "label": "i",
          "marks": 8,
          "text": "Estimate the parameters of each of these distributions.                                     (8)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 8,
          "text": "Plot a histogram of “PAID” data with 30 equal class intervals. Superimpose the\n               histogram with the Probability Density Functions of the above four distributions using\n               their estimated parameters obtained in part (i). Mark each plot distinctly using an\n               appropriate legend.                                                                           (8)",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 5,
          "text": "Compute the 5th percentile, 1st quartile, median, 3rd quartile and 95th percentile of both\n               the actual claims paid as well as from the fitted distributions.                              (5)",
          "topic": null
        },
        {
          "label": "iv",
          "marks": 4,
          "text": "Using the results from (ii) and (iii), comment on goodness-of-fit of the models to the\n               data.                                                                                         (4)",
          "topic": null
        }
      ],
      "solution": "i)\nsampleMean<-mean(AutoClaims$PAID)\nsampleVariance<-var(AutoClaims$PAID)\n\n#Method of Moments Estimates - Normal distribution\nNormalmu <- sampleMean                                                                       [1]\nNormalsigma <- sqrt(sampleVariance)                                                          [1]\nNormalmu\n## [1] 6127.222\n\nNormalsigma\n\n## [1] 7027.434\n\n#Method of Moments Estimates - Log Normal distribution\nLNsigma<- sqrt(log(1+sampleVariance/sampleMean^2))                                          [1]\nLNmu<-log(sampleMean)-LNsigma^2/2                                                           [1]\nLNsigma\n\n## [1] 0.9162933\n\nLNmu\n\n## [1] 8.3007\n                                                                                   Page 4 of 13\n\fIAI                                                                                  CS1B-0619\n\n#Method of Moments Estimates - Exponential distribution\nExprate <- 1/sampleMean                                                                      [2]\nExprate\n\n## [1] 0.0001632061\n\n#Method of Moments Estimates - Gamma distribution\nGammaBeta<-sampleMean/sampleVariance                                                         [1]\nGammaAlpha<-GammaBeta*sampleMean                                                             [1]\nGammaBeta\n\n## [1] 0.0001240709\n\nGammaAlpha\n\n## [1] 0.7602103\nii)\n#Histogram of Paid Claims Data\nhist(AutoClaims$PAID,breaks = 30,freq = FALSE)                                       [2]\n#Superimposing a Normal distribution over the histogram\ncurve(dnorm(x,mean = Normalmu,sd = Normalsigma),from = min(AutoClaims$PAID), to = m\nax(AutoClaims$PAID), add = TRUE, col= \"red\")                                      [1.5]\n#Superimposing a Log Normal distribution over the histogram\ncurve(dlnorm(x,meanlog = LNmu,sdlog = LNsigma),from = min(AutoClaims$PAID), to = max(\nAutoClaims$PAID), add = TRUE, col= \"green\")                                        [1.5]\n#Superimposing a Exponential distribution over the histogram\ncurve(dexp(x,rate = Exprate),from = min(AutoClaims$PAID), to = max(AutoClaims$PAID), ad\nd = TRUE, col= \"blue\")                                                             [1.5]\n#Superimposing a Gamma distribution over the histogram\ncurve(dgamma(x,shape = GammaAlpha,rate = GammaBeta),from = min(AutoClaims$PAID), t\no = max(AutoClaims$PAID), add = TRUE, col= \"brown\")                                [1.5]\n\nlegend(\"topright\",legend = c(\"Normal\", \"Lognormal\", \"Exponential\", \"Gamma\"),lty = 1, col =\nc(\"red\",\"green\",\"blue\",\"brown\"))\n\niii)\n#Computed quantiles of the actual data as well as that of different distributions\nquantile(AutoClaims$PAID,c(0.05,0.25,0.5,0.75,0.95))                                         [1]\n\n                                                                                    Page 5 of 13\n\fIAI                                                                                       CS1B-0619\n\n##    5%    25%    50%    75%    95%\n## 1116.379 1610.660 3395.368 7774.382 20405.050\n\nqnorm(c(0.05,0.25,0.5,0.75,0.95),mean = Normalmu,sd = Normalsigma)                                 [1]\n\n## [1] -5431.878 1387.290 6127.222 10867.155 17686.323\n\nqlnorm(c(0.05,0.25,0.5,0.75,0.95),meanlog = LNmu,sdlog = LNsigma)                                  [1]\n\n## [1] 892.0584 2170.4062 4026.6904 7470.5996 18176.2032\n\nqexp(c(0.05,0.25,0.5,0.75,0.95),rate = Exprate)                                                    [1]\n\n## [1] 314.2854 1762.6920 4247.0670 8494.1339 18355.5180\n\nqgamma(c(0.05,0.25,0.5,0.75,0.95),shape = GammaAlpha,rate = GammaBeta)                             [1]\n\n## [1] 142.0733 1276.6130 3737.9364 8453.3371 20244.3595\niv)\n#Comment based on (ii) and (iii)\n\nFrom the histogram and the superimposed plots, it is clear that normal distribution does not\nfit the data well.                                                                        [1]\n\nThe other three curves are getting superimposed more or less similarly to the data. Even fro\nm the quantiles we observe that lower values are aptly modeled using lognormal distributio\nn (5th percentile of lognormal being close to actual values) whereas gamma distribution is m\nodeling the higher values more appropriately (95th percentile).                           [2]\n\nHence, just by looking at (ii) and (iii), best fitting distribution among Lognormal, Gamma and\nExponential distributions cannot be concluded. It requires additional analysis in the form of o\nther statistical tests to confirm the best fit\n\nPart (i) was well attempted though many students had difficulty in arriving at the parameter\ns corresponding to lognormal distribution and Gamma distribution. All those who attempted\nPart (i) successfully were able to attempt Part (ii) and Part (iii) as well, but the failure in fittin\ng a few distributions in part (i) resulted in only part answers for part (ii) and (iii). Many stude\nnts were able to plot the histogram and the distributions decently well but the labelling thro\nugh appropriate legends was not done well. Part (iv) was not attempted by many students a\nnd among them who attempted, complete interpretation was lagging. Very few wrote a det\nailed interpretation of the result.",
      "has_math": false,
      "is_r_task": true,
      "session": "2019-06",
      "subject": "CS1B",
      "source_qp": "raw/CS1B_2019-06_QP.pdf",
      "source_sol": "raw/CS1B_2019-06_SOL.pdf"
    },
    {
      "q_num": 3,
      "marks": 45,
      "topic": "regression_glm",
      "subtopics": [],
      "stem": "Refer to the dataset “AutoClaims.csv” and answer the following questions.",
      "parts": [
        {
          "label": "i",
          "marks": 12,
          "text": "Fit a linear regression model to predict the “PAID” claim amount based on other\n               variables (Consider the AGE as a numerical variable and all others as categorical).\n               Provide your interpretation of the model by explaining R-Squared, Adjusted R-\n               Squared, p-value of the model and p-value of each of the coefficients. Identify the\n               significant variables in the prediction of “PAID” claims.                                    (12)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 10,
          "text": "Comment on the applicability of the linear regression model by plotting “Residuals vs.\n               Fitted Values” and “QQ Plot of the residuals”.                                               (10)",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 7,
          "text": "Your actuarial friend has suggested you to use natural logarithm of “PAID” claims\n               instead of the actual “PAID” Claim amount because the loge(PAID) is more closer to\n               normal distribution than “PAID” Claims. Verify the statement made by your friend by\n               comparing the Skewness and Excess Kurtosis of both the PAID claims as well as\n               loge(PAID). Write appropriate custom functions to compute both of them.                       (7)",
          "topic": null
        },
        {
          "label": "iv",
          "marks": 6,
          "text": "Repeat the model in (i) above by considering the suggestion in (iii). Identify and\n               comment on the key differences between both the models.                                       (6)",
          "topic": null
        },
        {
          "label": "v",
          "marks": 45,
          "text": "Your Manager has suggested that the model can be improved by adding interaction\n               effects between STATE and CLASS, STATE and GENDER, CLASS and GENDER\n               as additional variables to the set of independent variables taken in (i). Evaluate the\n               worthiness of this suggestion.                                                               [10]\n\n         Link for Data Set:- https://actuariesindia.org/sites/default/files/2022-09/CS1B_AutoClaims.csv",
          "topic": null
        }
      ],
      "solution": "i)\n#Fitting a Linear Regression Model\nmodel1<-lm(PAID~.,data = AutoClaims)\nsummary(model1)                                                                                    [5]\n##\n## Call:\n## lm(formula = PAID ~ ., data = AutoClaims)\n\n                                                                                         Page 6 of 13\n\fIAI                                                                                 CS1B-0619\n\n##\n## Residuals:\n## Min 1Q Median 3Q Max\n## -10462 -2276 119 1611 36377\n##\n## Coefficients:\n##           Estimate Std. Error t value Pr(>|t|)\n## (Intercept) 19818.12 1391.58 14.242 < 2e-16 ***\n## STATESTATE 02 -2306.41 658.69 -3.502 0.000477 ***\n## STATESTATE 03 -580.09 761.36 -0.762 0.446242\n## STATESTATE 04 -689.08 702.04 -0.982 0.326495\n## STATESTATE 06 440.79 752.27 0.586 0.558010\n## STATESTATE 07 -1254.29 837.22 -1.498 0.134318\n## STATESTATE 10 2275.25 885.44 2.570 0.010284 *\n## STATESTATE 12 -752.99 850.10 -0.886 0.375897\n## STATESTATE 14 -404.69 842.90 -0.480 0.631216\n## STATESTATE 15 -4791.86 623.56 -7.685 2.87e-14 ***\n## STATESTATE 17 -883.67 704.58 -1.254 0.209982\n## CLASSC11 -11743.95 430.60 -27.274 < 2e-16 ***\n## CLASSC6       -14833.37 410.84 -36.105 < 2e-16 ***\n## CLASSF6        -225.16 517.68 -0.435 0.663670\n## GENDERM          -1193.01 215.50 -5.536 3.69e-08 ***\n## AGE           15.76 24.10 0.654 0.513418\n## ---\n## Signif. codes: 0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1\n##\n## Residual standard error: 3943 on 1401 degrees of freedom\n## Multiple R-squared: 0.6886, Adjusted R-squared: 0.6852\n## F-statistic: 206.5 on 15 and 1401 DF, p-value: < 2.2e-16\nNote: 1 Mark can be deducted if the output is not pasted\nanova(model1)\nAnalysis of Variance Table\n\nResponse: PAID\n        Df Sum Sq Mean Sq F value Pr(>F)\nSTATE      10 9.8027e+09 9.8027e+08 63.0629 < 2.2e-16 ***\nCLASS       3 3.7869e+10 1.2623e+10 812.0763 < 2.2e-16 ***\nGENDER        1 4.7271e+08 4.7271e+08 30.4106 4.158e-08 ***\nAGE        1 6.6423e+06 6.6423e+06 0.4273 0.5134\nResiduals 1401 2.1778e+10 1.5544e+07\n---\nSignif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1\nInterpretation\nR-Squared: 68.86% of the variation in the claims paid is explained by state, rating class,\ngender and age                                                                               [1]\nAdjusted R-Squared: 68.52% is used to compare with other models, adjusts for the number\nof terms in the model. We Use adjusted R-squared to compare the goodness-of-fit for\nregression models that contain differing numbers of independent variables.            [1]\n\n                                                                                   Page 7 of 13\n\fIAI                                                                                   CS1B-0619\n\np-value of the model is <2.2*E-16 which is less than 0.05 and hence the null hypothesis of\n“There is no significant linear relationship between the given independent variables X and a\ndependent variable Y” is rejected at 5% level of significance. Using this model to predict the\nDV is better than simply using the expected value of the DV as a predictor for the DV       [2]\np-value of the coefficients: While the model is overall significant, some of the variables may\nbe insignificant. As state 1, Rating class C1 and Gender female are taken as based states and\ntheir coefficients are clubbed in the intercept itself, we observe that coefficients of State 2\nand state 15 (Negative) and State 10 (Positive) are significantly different from state 1 (At\n95% Confidence level). Similarly rating classes C11 and C6 have significantly negative\ncoefficients compared to C1 indicating that the claim paid for those two rating classes is\nsignificantly lesser compared to that of C1. Males have significantly lesser claim paid\ncompared to females at 95% confidence level                                                    [2]\nFrom the ANOVA table, we can infer that except Age, all other variables are significant in\nprediction of claims paid                                                                  [1]\nii)\n#Plot of residuals vs. Fitted Values\nplot(model1$fitted.values,model1$residuals)                                                   [3]\n\nThe plot is used to detect non-linearity, unequal error variances, and outliers.\nThe residuals \"do not bounce randomly\" around the 0 line. This suggests that the\nassumption that the relationship is linear is not reasonable.\nThe residuals do not form a \"horizontal band\" around the 0 line. This suggests that the\nvariances of the error terms are not equal and exhibit heteroscedasticity\nA few residuals \"stands out\" from the basic random pattern of residuals. This suggests that\nthere are outliers.                                                                       [2]\n# QQ Plot of the residuals\nqqnorm(model1$residuals)                                                                      [3]\n\n                                                                                     Page 8 of 13\n\fIAI                                                                                 CS1B-0619\n\nA Q-Q plot is a scatterplot created by plotting two sets of quantiles against one another. If\nboth sets of quantiles came from the same distribution, we should see the points forming a\nline that’s roughly straight. Here it is not, indicating deviance of the residuals from\nnormality. Thus linear regression may not be a better fit to the data                        [2]\niii)\n#Reason for better model\n#Checking for the normality of Auto Claims Paid vs. Logairthm of Auto Claims Paid\n#Writing Functions for Skewness and Kurtosis\nskew<-function(x)mean((x-mean(x))^3)/sd(x)^3                                                 [2]\nkurt<-function(x)(mean((x-mean(x))^4)/sd(x)^4)-3                                             [2]\nskew(AutoClaims$PAID)                                                                      [0.5]\n## [1] 2.619422\nkurt(AutoClaims$PAID)                                                                      [0.5]\n## [1] 9.20876\nskew(log(AutoClaims$PAID))                                                                 [0.5]\n## [1] 0.4528057\nkurt(log(AutoClaims$PAID))                                                                 [0.5]\n## [1] -0.787689\nSkewness and Kurtosis of Log (Claims) are more close to Zero compared to those of actual\nclaims paid, thus indicating the possibility of using linear regression with this dependent\nvariable                                                                                    [1]\niv)\n# Using Natural Logarithm of the claims paid\nmodel2<-lm(log(PAID)~.,data = AutoClaims)\nsummary(model2)                                                                              [3]\n\n                                                                                   Page 9 of 13\n\fIAI                                                                              CS1B-0619\n\nanova(model2)\nAnalysis of Variance Table\n\nResponse: log(PAID)\n        Df Sum Sq Mean Sq F value Pr(>F)\nSTATE      10 246.26 24.626 107.6969 < 2.2e-16 ***\nCLASS       3 690.09 230.031 1006.0090 < 2.2e-16 ***\nGENDER        1 10.88 10.882 47.5908 7.927e-12 ***\nAGE        1 0.12 0.123 0.5384 0.4632\nResiduals 1401 320.35 0.229\n---\nSignif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1\n##\n## Call:\n## lm(formula = log(PAID) ~ ., data = AutoClaims)\n##\n## Residuals:\n## Min          1Q Median       3Q Max\n## -0.96098 -0.34264 -0.05047 0.36828 1.08237\n##\n## Coefficients:\n##           Estimate Std. Error t value Pr(>|t|)\n## (Intercept) 9.896308 0.168777 58.635 < 2e-16 ***\n## STATESTATE 02 -0.154804 0.079889 -1.938 0.0529 .\n## STATESTATE 03 0.110585 0.092342 1.198 0.2313\n## STATESTATE 04 0.049554 0.085147 0.582 0.5607\n## STATESTATE 06 0.116190 0.091239 1.273 0.2031\n## STATESTATE 07 0.142721 0.101543 1.406 0.1601\n## STATESTATE 10 0.098014 0.107391 0.913 0.3616\n## STATESTATE 12 0.027982 0.103105 0.271 0.7861\n## STATESTATE 14 0.090316 0.102231 0.883 0.3771\n## STATESTATE 15 -0.645918 0.075628 -8.541 < 2e-16 ***\n## STATESTATE 17 0.004611 0.085455 0.054 0.9570\n## CLASSC11 -1.203098 0.052225 -23.037 < 2e-16 ***\n## CLASSC6        -1.988743 0.049829 -39.911 < 2e-16 ***\n## CLASSF6       -0.034909 0.062787 -0.556 0.5783\n## GENDERM          -0.180923 0.026136 -6.922 6.75e-12 ***\n## AGE          0.002145 0.002923 0.734 0.4632\n## ---\n## Signif. codes: 0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1\n##\n## Residual standard error: 0.4782 on 1401 degrees of freedom\n## Multiple R-squared: 0.7473, Adjusted R-squared: 0.7446\n## F-statistic: 276.2 on 15 and 1401 DF, p-value: < 2.2e-16\n\nKey Differences\n      1. R-Squared and Adjusted R-Squared improved         and hence the model is a better\n         fit compared to the initial model                                             [1.5]\n\n                                                                              Page 10 of 13\n\fIAI                                                                                 CS1B-0619\n\n      2. While a the significance level of a few factor coefficients when compared with the\n         base categories changed, the overall significant variables did not change which can\n         be inferred from the ANOVA table                                                 [1.5]\nv)\n# Using Interaction effects in the model\nmodel3<-lm(PAID~.+STATE:CLASS+STATE:GENDER+CLASS:GENDER,data = AutoClaims)\nsummary(model3)                                                                             [5]\n##\n## Call:\n## lm(formula = PAID ~ . + STATE:CLASS + STATE:GENDER + CLASS:GENDER,\n## data = AutoClaims)\n##\n## Residuals:\n## Min        1Q Median         3Q Max\n## -13110.3 -1475.9 -377.5 1250.5 20442.8\n##\n## Coefficients: (3 not defined because of singularities)\n##               Estimate Std. Error t value Pr(>|t|)\n## (Intercept)        24373.09 3236.06 7.532 9.08e-14 ***\n## STATESTATE 02          -7868.94 3269.51 -2.407 0.016227 *\n## STATESTATE 03          -3695.87 3422.76 -1.080 0.280427\n## STATESTATE 04           6883.00 3545.51 1.941 0.052425 .\n## STATESTATE 06           5428.58 3262.57 1.664 0.096363 .\n## STATESTATE 07           -979.80 1184.59 -0.827 0.408314\n## STATESTATE 10           7340.08 3546.35 2.070 0.038664 *\n## STATESTATE 12           1048.26 3382.21 0.310 0.756659\n## STATESTATE 14          -2796.37 3538.40 -0.790 0.429494\n## STATESTATE 15         -14038.72 3164.04 -4.437 9.86e-06 ***\n## STATESTATE 17          -4266.43 3640.98 -1.172 0.241490\n## CLASSC11           -15321.38 3534.31 -4.335 1.56e-05 ***\n## CLASSC6           -20741.82 3262.53 -6.358 2.79e-10 ***\n## CLASSF6            4075.85 3386.90 1.203 0.229026\n## GENDERM              -5508.80 1305.95 -4.218 2.62e-05 ***\n## AGE               23.06 19.35 1.192 0.233581\n## STATESTATE 02:CLASSC11 4114.13 3666.38 1.122 0.262008\n## STATESTATE 03:CLASSC11 2022.51 3802.06 0.532 0.594847\n## STATESTATE 04:CLASSC11 -7719.04 3904.17 -1.977 0.048229 *\n## STATESTATE 06:CLASSC11 -6209.18 3743.46 -1.659 0.097412 .\n## STATESTATE 07:CLASSC11 -1049.15 1827.01 -0.574 0.565899\n## STATESTATE 10:CLASSC11 -6795.03 3956.71 -1.717 0.086144 .\n## STATESTATE 12:CLASSC11 -3756.14 3939.41 -0.953 0.340517\n## STATESTATE 14:CLASSC11 1668.66 3927.62 0.425 0.671012\n## STATESTATE 15:CLASSC11 7776.13 3581.01 2.171 0.030066 *\n## STATESTATE 17:CLASSC11 2227.20 4011.07 0.555 0.578806\n## STATESTATE 02:CLASSC6 5677.35 3416.32 1.662 0.096777 .\n## STATESTATE 03:CLASSC6 2122.55 3605.64 0.589 0.556177\n## STATESTATE 04:CLASSC6 -8231.17 3692.27 -2.229 0.025957 *\n## STATESTATE 06:CLASSC6 -6151.60 3417.36 -1.800 0.072066 .\n## STATESTATE 07:CLASSC6          NA       NA NA        NA\n## STATESTATE 10:CLASSC6 -8117.07 3687.43 -2.201 0.027884 *\n\n                                                                                 Page 11 of 13\n\fIAI                                                                            CS1B-0619\n\n## STATESTATE 12:CLASSC6 -2090.95 3561.06 -0.587 0.557187\n## STATESTATE 14:CLASSC6 1711.83 3617.52 0.473 0.636143\n## STATESTATE 15:CLASSC6 10891.63 3315.76 3.285 0.001046 **\n## STATESTATE 17:CLASSC6 2376.31 3776.77 0.629 0.529330\n## STATESTATE 02:CLASSF6 -1213.53 3542.37 -0.343 0.731971\n## STATESTATE 03:CLASSF6 -2272.17 3835.79 -0.592 0.553708\n## STATESTATE 04:CLASSF6 -11888.76 3838.95 -3.097 0.001996 **\n## STATESTATE 06:CLASSF6 -17556.04 4687.92 -3.745 0.000188 ***\n## STATESTATE 07:CLASSF6          NA       NA NA            NA\n## STATESTATE 10:CLASSF6 1619.22 3953.37 0.410 0.682178\n## STATESTATE 12:CLASSF6          NA       NA NA            NA\n## STATESTATE 14:CLASSF6 -2570.84 3769.07 -0.682 0.495299\n## STATESTATE 15:CLASSF6 -3790.79 3441.11 -1.102 0.270822\n## STATESTATE 17:CLASSF6 362.96 3914.01 0.093 0.926130\n## STATESTATE 02:GENDERM 3404.19 1199.24 2.839 0.004598 **\n## STATESTATE 03:GENDERM 3061.44 1347.32 2.272 0.023227 *\n## STATESTATE 04:GENDERM 1514.18 1279.13 1.184 0.236716\n## STATESTATE 06:GENDERM 2052.42 1337.41 1.535 0.125108\n## STATESTATE 07:GENDERM 3094.31 1450.30 2.134 0.033057 *\n## STATESTATE 10:GENDERM -73.99 1540.53 -0.048 0.961699\n## STATESTATE 12:GENDERM 2476.44 1474.81 1.679 0.093351 .\n## STATESTATE 14:GENDERM 2684.73 1624.58 1.653 0.098650 .\n## STATESTATE 15:GENDERM 3280.15 1148.55 2.856 0.004357 **\n## STATESTATE 17:GENDERM 3033.26 1261.58 2.404 0.016335 *\n## CLASSC11:GENDERM            1171.02 730.27 1.604 0.109047\n## CLASSC6 :GENDERM           2372.86 692.55 3.426 0.000630 ***\n## CLASSF6 :GENDERM          -1786.40 905.11 -1.974 0.048620 *\n## ---\n## Signif. codes: 0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1\n##\n## Residual standard error: 3111 on 1361 degrees of freedom\n## Multiple R-squared: 0.8116, Adjusted R-squared: 0.804\n## F-statistic: 106.6 on 55 and 1361 DF, p-value: < 2.2e-16\nInterpretation                                                                        [5]\nanova(model3)\nAnalysis of Variance Table\n\nResponse: PAID\n          Df Sum Sq Mean Sq F value Pr(>F)\nSTATE       10 9.8027e+09 9.8027e+08 101.2664 < 2.2e-16 ***\nCLASS        3 3.7869e+10 1.2623e+10 1304.0318 < 2.2e-16 ***\nGENDER         1 4.7271e+08 4.7271e+08 48.8334 4.350e-12 ***\nAGE         1 6.6423e+06 6.6423e+06 0.6862 0.407612\nSTATE:CLASS 27 7.8659e+09 2.9133e+08 30.0960 < 2.2e-16 ***\nSTATE:GENDER 10 2.7570e+08 2.7570e+07 2.8482 0.001634 **\nCLASS:GENDER 3 4.6128e+08 1.5376e+08 15.8841 3.733e-10 ***\nResiduals 1361 1.3175e+10 9.6801e+06\n---\nSignif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1\n      1. R-Squared and Adjusted R-Squared increased to above 80% and hence the model is a\n         better fit compared to the earlier models                                    [2]\n                                                                            Page 12 of 13\n\fIAI                                                                                      CS1B-0619\n\n      2. Interaction effect between a few classes and states emerged out to be very\n         significant (State 6 and Class F6 came out to be significantly negative). Though State\n         15 came out to be significantly negative when main effects alone were considered,\n         the interaction effects compensated that negative significantly when interacted with\n         class C6 and Class 11 whereas the interaction coefficient is not significant between\n         State 15 and Class F6 indicating that the claims paid is significantly lesser when the\n         state is 15 and class is F6 compared to other rating classes. Digging deeper into the\n         relationships is possible with the interaction effect. Similarly main effect of Gender is\n         significantly negative compared to females but that is offset to some extent for some\n         states (2,3,7,15) and for some rating classes (C6) whereas it is further negative in\n         case of F6. So the differences can be magnified by considering the interaction\n         effects, improving the predictability of the model                                    [1]\n      3.     ANOVA table for the model suggests that except the age all the main effects and\n            their interaction effects are significant at 5% significance level indicating their\n            contribution to the predictability of the model                                     [2]\n\n      Part (i) and Part (ii) were attempted by maximum number of students. Many of them\n      were successful in writing the code to fit a linear regression model but only successful\n      students were able to provide a good interpretation of the results. Likewise many\n      students were able to write the code for plotting but very few of them ended up in\n      writing the interpretations based on the graphs. Part (iii) was not attempted by many\n      students and those who attempted also ended up in writing wrong functions. Overall,\n      this question was very poorly done. Part (iv) was attempted well, code was written well\n      but a few ended up in writing correct interpretations. Part (v) was attempted by a few\n      students only and many of them failed to provide appropriate interpretation.\n\n          Performance of students in CS1B varied drastically. Only a few questions were\n           answered successfully by majority of the participants\n          R code was used well by majority of the participants but they failed to make\n           interpretations based on the results of the code\n          Topics wise, students showed a good understanding of the regression models, data\n           analysis and visualizations but the topics on distributions, Hypothesis testing were not\n           addressed well.\n          The level of interpretation and the comments provided next to the R Code varied\n           significantly among the students.\n          A good number of students failed in submitting/pasting the output for a few R\n           Functions which resulted in them losing a few marks.\n                                        *******************\n\n                                                                                      Page 13 of 13",
      "has_math": false,
      "is_r_task": true,
      "session": "2019-06",
      "subject": "CS1B",
      "source_qp": "raw/CS1B_2019-06_QP.pdf",
      "source_sol": "raw/CS1B_2019-06_SOL.pdf"
    },
    {
      "q_num": 1,
      "marks": 12,
      "topic": "data_analysis",
      "subtopics": [],
      "stem": "The probability that India will win a cricket match against South Africa is 0.7",
      "parts": [
        {
          "label": "i",
          "marks": 4,
          "text": "Prepare a probability distribution table of number of wins if Indians are going to play 10\n           matches against the South Africans.                                                             (4)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 4,
          "text": "Plot a bar chart of the probabilities of number of wins from 0 to 10.                          (4)",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 4,
          "text": "Find the mean and median number of wins for India against South Africa.                       (4)",
          "topic": null
        }
      ],
      "solution": "i)\n\n       p<-0.7\n\n       probability distribution table of number of number of wins\n\n       prob<-dbinom(0:10,size = 10,prob = 0.7)                                       [3]\n       probability_Densities<-data.frame(No_Successes = 0:10,prob)\n       probability_Densities                                                         [1]\n\n       No_Successes           prob\n       0                   0.0000059049\n       1                   0.0001377810\n       2                   0.0014467005\n       3                   0.0090016920\n       4                   0.0367569090\n       5                   0.1029193452\n       6                   0.2001209490\n       7                   0.2668279320\n       8                   0.2334744405\n       9                   0.1210608210\n       10                  0.0282475249\n\nii)\n       bar chart of the probabilities of number of wins\n\n       barplot(prob,main = \"Bar Chart of Probability of Successes\", xlab = \"Numbe\n       r of Successes\", names.arg = 0:10)                                     [4]\n\niii)\n             mean and median number of wins for India against South Africa\n\n             mean<-10*0.7 #or\n                                                                             Page 2 of 13\n\f      IAI                                                                 CS1B-1119\n            mean<-sum(probability_Densities$No_Successes*probability_Densities$prob\n            )\n            mean                                                                [1]\n            ## [1] 7                                                            [1]\n            median<-qbinom(0.5,size = 10,prob = 0.7)\n            median                                                              [1]\n            ## [1] 7                                                            [1]",
      "has_math": false,
      "is_r_task": true,
      "session": "2019-11",
      "subject": "CS1B",
      "source_qp": "raw/CS1B_2019-11_QP.pdf",
      "source_sol": "raw/CS1B_2019-11_SOL.pdf"
    },
    {
      "q_num": 2,
      "marks": 18,
      "topic": "distributions",
      "subtopics": [],
      "stem": "Comment on the appropriateness of the central limit theorem by implementing the following\n        steps:",
      "parts": [
        {
          "label": "i",
          "marks": 2,
          "text": "Generate a sample of 10000 random observations following Lognormal distribution with\n           parameters µ = 2 and σ2 = 0.25 (µ and σ are the parameters of the distribution as defined\n           in the Actuarial Tables provided) and display the first few simulated observations using\n           the head (...) function. [Use a seed value of 100 to generate random numbers]                   (2)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 8,
          "text": "Compute the sample mean, median and variance from the generated sample and compare\n            the values with those of a population following a lognormal distribution with the given\n            parameters.                                                                                    (8)",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 5,
          "text": "Treat the data generated in (i) as the population. Generate 500 different random samples\n             of size 200 from the above population and compute the sample mean for each sample.\n             [Use a seed value of 100 to generate random numbers]                                          (5)",
          "topic": null
        },
        {
          "label": "iv",
          "marks": 3,
          "text": "Plot the histogram of sample means generated from part (iii) and interpret the distribution\n            of sample means.                                                                               (3)",
          "topic": null
        }
      ],
      "solution": "i)\n            Generate a random sample from a Lognormal distribution\n            set.seed(100)\n            data1<-rlnorm(10000,meanlog = 2,sdlog = 0.5)                         [1]\n            # First 6 observations are shown below\n            head(data1)                                                          [1]\n            ## [1] 5.748298 7.891337 7.103172 11.512028 7.834097     8.665200\n\nii)\n\n            Compute the mean, median and variance of the sample\n            mean(data1)                                                          [1]\n            ## [1] 8.375649\n            median(data1)                                                        [1]\n            ## [1] 7.398463\n            var(data1)                                                           [1]\n            ## [1] 19.51361\n\n            # Formula based mean values\n            mean<-exp(2+0.25/2)\n            median<-exp(2)\n            var<-(exp(0.25)-1)*exp(2*2+0.25)\n\n            mean                                                                 [1]\n            ## [1] 8.372897\n\n            Median                                                               [1]\n\n            ## [1] 7.389056\n\n            var                                                                  [1]\n\n            ## [1] 19.91172\n\n            Interpretation: Mean, Median and Variance of the generated sample and t\n            hose computed based on the parameters are almost equal because the samp\n            le size is 10,000 which is pretty large. Generating a much larger sampl\n            e will bridge those smaller differences existing between them as well\n\n                                                                         Page 3 of 13\n\f       IAI                                                                               CS1B-1119\n\niii)\n             Generating 500 different samples of size 200 and computing their sample\n             means\n\n             means<-c()\n             set.seed(100)\n             for (i in 1:500){\n               selected_rows<-sample(1:10000,200,FALSE)\n               selected_data<-data1[selected_rows]\n               sample_mean<-mean(selected_data)\n               means<-c(means,sample_mean)\n             }                                                                                  [5]\n\niv)\n             Histogram of Sample Means\n             hist(means,breaks = 30)                                                            [1]\n\n       Interpretation: The sample means tend to follow a normal distribution though the actual data\n       comes from lognormal distribution. Central limit theorem can be verified through this\n       exercise that sample means tend to follow a normal distribution as the sample size increases.\n       Increase in Sample size from 200 to much higher can ensure better normality of the sample\n       means                                                                                    [2]\n\n                                                                                       Page 4 of 13\n\f       IAI                                                                   CS1B-1119",
      "has_math": true,
      "is_r_task": true,
      "session": "2019-11",
      "subject": "CS1B",
      "source_qp": "raw/CS1B_2019-11_QP.pdf",
      "source_sol": "raw/CS1B_2019-11_SOL.pdf"
    },
    {
      "q_num": 3,
      "marks": 20,
      "topic": "distributions",
      "subtopics": [],
      "stem": "Refer to the data file “Indices_Returns.csv” and answer the following questions:\n\n        Indices_Returns.csv file is provided in the system.",
      "parts": [
        {
          "label": "i",
          "marks": 3,
          "text": "Compute the pairwise Pearson correlation coefficient between the returns of 10 sectors\n           (BM, CD, EN, FM, FI, HC, IN, IT, TE and UT) rounded to three digits after the decimal\n           point. Display the correlation matrix in the output.                                            (3)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 2,
          "text": "Identify the pair with the highest correlation coefficient and the pair with the least\n            correlation coefficient.                                                                       (2)",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 4,
          "text": "Perform Principal component analysis on the returns values of the 10 sectors.                 (4)",
          "topic": null
        },
        {
          "label": "iv",
          "marks": 1,
          "text": "How many principal components have an Eigen value of more than 1?                              (1)",
          "topic": null
        },
        {
          "label": "v",
          "marks": 1,
          "text": "What is the approximate proportion of total variation explained by the first two principal\n           components?                                                                                     (1)",
          "topic": null
        },
        {
          "label": "vi",
          "marks": 5,
          "text": "Compute the pair wise correlations among the 10 principal components (Round them to\n            3 digits after the decimal point) and display the results. What do you infer about the\n            resulting correlations?                                                                        (5)",
          "topic": null
        },
        {
          "label": "vii",
          "marks": 4,
          "text": "Using a scree plot comment on the number of significant components in the model.              (4)",
          "topic": null
        }
      ],
      "solution": "# Load the data file\n       indices<-read.csv(\"D:\\\\Indices_Returns.csv\")\n\n i)\n       Compute pearson correlation coefficient and finding the most\n       correlated and least correlated pair\n       correlation<-cor(indices[,3:12], method = \"pearson\")                         [1]\n       correlation<-round(correlation,3)                                            [1]\n\n       correlation                                                                  [1]\n       ##       BM    CD    EN    FM    FI    HC    IN    IT    TE    UT\n       ## BM 1.000 0.882 0.823 0.534 0.819 0.605 0.898 0.447 0.646 0.861\n       ## CD 0.882 1.000 0.783 0.581 0.878 0.637 0.915 0.434 0.696 0.847\n       ## EN 0.823 0.783 1.000 0.446 0.745 0.506 0.799 0.359 0.622 0.793\n       ## FM 0.534 0.581 0.446 1.000 0.485 0.510 0.511 0.303 0.410 0.520\n       ## FI 0.819 0.878 0.745 0.485 1.000 0.502 0.902 0.349 0.623 0.838\n       ## HC 0.605 0.637 0.506 0.510 0.502 1.000 0.588 0.525 0.489 0.530\n       ## IN 0.898 0.915 0.799 0.511 0.902 0.588 1.000 0.370 0.676 0.882\n       ## IT 0.447 0.434 0.359 0.303 0.349 0.525 0.370 1.000 0.291 0.317\n       ## TE 0.646 0.696 0.622 0.410 0.623 0.489 0.676 0.291 1.000 0.669\n       ## UT 0.861 0.847 0.793 0.520 0.838 0.530 0.882 0.317 0.669 1.000\n\nii)\n             Pair with Minimum Correlation and Pair with maximum correlation\n             min_cor_location<-which(correlation == min(correlation))[1]\n             min_cor_pair<-paste(rownames(correlation)[ceiling(min_cor_location/10)]\n             ,colnames(correlation)[ceiling(min_cor_location%%10)])\n             min_cor_pair\n\n             ## [1] \"IT TE\"                                                         [1]\n\n             max_cor_location<-which(correlation == max(correlation[correlation!=1])\n             )[1]\n             max_cor_pair<-paste(rownames(correlation)[ceiling(max_cor_location/10)]\n             ,colnames(correlation)[ceiling(max_cor_location%%10)])\n             max_cor_pair\n\n             ## [1] \"CD IN\"                                                         [1]\n\niii)\n             Perform a Principal component analysis of the sectoral return values\n             PCA_corr<-princomp(indices[,3:12])\n             summary(PCA_corr)                                                    [4]\n             ## Importance of components:\n             ##                           Comp.1     Comp.2     Comp.3     Comp.4\n             ## Standard deviation     0.2106142 0.06763728 0.05825406 0.04631607\n\n                                                                           Page 5 of 13\n\f      IAI                                                                    CS1B-1119\n            ## Proportion of Variance 0.7294045 0.07522554 0.05580143 0.03527414\n            ## Cumulative Proportion 0.7294045 0.80463008 0.86043151 0.89570565\n            ##                            Comp.5     Comp.6     Comp.7     Comp.8\n            ## Standard deviation     0.04255566 0.03735608 0.03326497 0.03093065\n            ## Proportion of Variance 0.02977884 0.02294646 0.01819564 0.01573154\n            ## Cumulative Proportion 0.92548448 0.94843094 0.96662658 0.98235812\n            ##                             Comp.9     Comp.10\n            ## Standard deviation     0.023803102 0.022500987\n            ## Proportion of Variance 0.009316657 0.008325228\n            ## Cumulative Proportion 0.991674772 1.000000000\n\n            Alternatively instead of using princomp, the student can use prcomp as\n            well.\n\n            PCA_corr_1<-prcomp(indices[,3:12])\n            summary(PCA_corr_1)\n\n            ## Importance of components:\n            ##                           PC1     PC2     PC3     PC4     PC5     P6\n            ## Standard deviation     0.2113 0.06784 0.05843 0.04646 0.04269 0.0377\n            ## Proportion of Variance 0.7294 0.07523 0.05580 0.03527 0.02978 0.0225\n            ## Cumulative Proportion 0.7294 0.80463 0.86043 0.89571 0.92548 0.9483\n            ##                            PC7     PC8     PC9    PC10\n            ## Standard deviation     0.03337 0.03103 0.02388 0.02257\n            ## Proportion of Variance 0.01820 0.01573 0.00932 0.00833\n            ## Cumulative Proportion 0.96663 0.98236 0.99167 1.00000\n\niv)\n            Number of PCA components with Eigen value more than 1\n            sum(PCA_corr$sdev^2/sum(PCA_corr$sdev^2)>(1/10))\n\n            ## [1] 1                                                                [1]\n            OR Alternatives\n\n            sum(PCA_corr_1$sdev^2/sum(PCA_corr_1$sdev^2)>(1/10))\n\n            ## [1] 1\n\nv)\n            proportion of total variation explained by the first two principal comp\n            onents\n            sum(PCA_corr$sdev[1:2]^2)/sum(PCA_corr$sdev^2)\n\n            ## [1] 0.8046301                                                        [1]\n            OR Alternatively\n\n            sum(PCA_corr_1$sdev[1:2]^2)/sum(PCA_corr_1$sdev^2)\n\n            ## [1] 0.8046301\n\nvi)\n            Paiwise correlations of the transformed components\n            round(cor(PCA_corr$scores),3)                                           [3]\n\n      ##               Comp.1 Comp.2 Comp.3 Comp.4 Comp.5 Comp.6 Comp.7 Comp.8 Comp.9\n      ## Comp.1             1      0      0      0      0      0      0      0      0\n      ## Comp.2             0      1      0      0      0      0      0      0      0\n\n                                                                            Page 6 of 13\n\f       IAI                                                              CS1B-1119\n       ## Comp.3       0         0     1      0    0      0      0     0       0\n       ## Comp.4       0         0     0      1    0      0      0     0       0\n       ## Comp.5       0         0     0      0    1      0      0     0       0\n       ## Comp.6       0         0     0      0    0      1      0     0       0\n       ## Comp.7       0         0     0      0    0      0      1     0       0\n       ## Comp.8       0         0     0      0    0      0      0     1       0\n       ## Comp.9       0         0     0      0    0      0      0     0       1\n       ## Comp.10      0         0     0      0    0      0      0     0       0\n       ##         Comp.10\n\n       ## Comp.1           0\n       ## Comp.2           0\n       ## Comp.3           0\n       ## Comp.4           0\n       ## Comp.5           0\n       ## Comp.6           0\n       ## Comp.7           0\n       ## Comp.8           0\n       ## Comp.9           0\n       ## Comp.10          1\n\n       OR Alternatively\n\n       round(cor(PCA_corr_1$x),3)\n\n       ##      PC1 PC2 PC3 PC4 PC5 PC6 PC7 PC8 PC9 PC10\n       ## PC1    1   0   0   0   0   0   0   0   0    0\n       ## PC2    0   1   0   0   0   0   0   0   0    0\n       ## PC3    0   0   1   0   0   0   0   0   0    0\n       ## PC4    0   0   0   1   0   0   0   0   0    0\n       ## PC5    0   0   0   0   1   0   0   0   0    0\n       ## PC6    0   0   0   0   0   1   0   0   0    0\n       ## PC7    0   0   0   0   0   0   1   0   0    0\n       ## PC8    0   0   0   0   0   0   0   1   0    0\n       ## PC9    0   0   0   0   0   0   0   0   1    0\n       ## PC10   0   0   0   0   0   0   0   0   0    1\n\n       Interpretation\n\n       The pairwise correlation between the components after the PCA is performed\n       should be zero as PCA is a way to deal with highly correlated variables. I\n       f N variables are highly correlated than they will all load out on the SA\n       ME Principal Component (Eigenvector) and they will be uncorrelated with ot\n       her components (All these components are orthogonal). Hence the correlatio\n       ns will be zero between the components                                 [2]\n\nvii)\n             Scree Plot\n             screeplot(PCA_corr,type = \"l\")                                    [3]\n\n                                                                       Page 7 of 13\n\f     IAI                                                                  CS1B-1119\n\n           OR Alternatively\n           screeplot(PCA_corr_1,type = \"l\")\n\n           Interpretation: Number of significant components is 1 as the scree plot\n           almost flattened out after the second component                     [1]",
      "has_math": false,
      "is_r_task": true,
      "session": "2019-11",
      "subject": "CS1B",
      "source_qp": "raw/CS1B_2019-11_QP.pdf",
      "source_sol": "raw/CS1B_2019-11_SOL.pdf"
    },
    {
      "q_num": 4,
      "marks": 20,
      "topic": "inference",
      "subtopics": [],
      "stem": "Refer to the data file “Indices_Returns.csv” and answer the following questions:\n\n        Indices_Returns.csv file is provided in the system.",
      "parts": [
        {
          "label": "i",
          "marks": 2,
          "text": "Express the number of months with negative Sensex returns as a proportion of total\n           number of months.                                                                               (2)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 4,
          "text": "Test whether the proportion of months with negative Sensex returns is less than 50% at\n            95% confidence level as well as at 99% confidence level.                                       (4)",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 5,
          "text": "Classify the monthly returns of FI and IT sectors as follows and prepare a contingency\n             table of counts.                                                                              (5)\n\n         Monthly Return           Classification\n         ≤ First Quartile value   Low\n         > Third Quartile value   High\n         All others               Medium",
          "topic": null
        },
        {
          "label": "iv",
          "marks": 5,
          "text": "Use the contingency table provided in 4(iii) above and test for the independence of\n            monthly returns between FI and IT sector returns using an appropriate test and provide an\n            interpretation of the test results.                                                            (5)",
          "topic": null
        },
        {
          "label": "v",
          "marks": 4,
          "text": "Test whether the returns of FI sector are significantly higher compared to that of IT Sector\n           at 95% confidence level using appropriate test.                                                 (4)",
          "topic": null
        }
      ],
      "solution": "# Load the data file\n     indices<-read.csv(\"D:\\\\Indices_Returns.csv\")\n\ni)\n           number of months with negative Sensex returns as a proportion of total\n           number of months\n           negative_Sensex<-sum(indices$Sensex<0)\n           negative_Sensex                                                      [1]\n\n           ## [1] 69\n\n                                                                         Page 8 of 13\n\f       IAI                                                                    CS1B-1119\n             proportion_neg<-negative_Sensex/nrow(indices)\n             proportion_neg                                                        [1]\n\n             ## [1] 0.4207317\n\nii)\n             Test whether the proportion of months with negative Sensex returns is l\n             ess than 50% at 95% confidence level as well as at 99% confidence level\n\n             binom.test(negative_Sensex,nrow(indices),p=0.5,alternative   =    \"less\")\n\n             ##\n             ## Exact binomial test\n             ##\n             ## data: negative_Sensex and nrow(indices)\n             ## number of successes = 69, number of trials = 164, p-value =\n             ## 0.02529\n             ## alternative hypothesis: true probability of success is less than 0.5\n             ## 95 percent confidence interval:\n             ## 0.0000000 0.4878846\n             ## sample estimates:\n             ## probability of success\n             ##              0.4207317\n\n             Interpretation: p-value corresponding to the test is 0.02529. So the nu\n             ll hypothesis of Proportion of Negatives is at least 50% is rejected at\n             95% Confidence level but is failed to be rejected at 99% Confidence lev\n             el                                                                  [1]\n\niii)\n             Classify the monthly returns of FI and IT\n             FI<-ifelse(indices$FI<=quantile(indices$FI,0.25),\"Low\",ifelse(indices$F\n             I>quantile(indices$FI,0.75),\"High\",\"Medium\"))                      [1.5]\n             IT<-ifelse(indices$IT<=quantile(indices$IT,0.25),\"Low\",ifelse(indices$I\n             T>quantile(indices$IT,0.75),\"High\",\"Medium\"))                      [1.5]\n             table(FI,IT)                                                         [2]\n\n             ##           IT\n             ## FI         High Low Medium\n             ##    High      14   8     19\n             ##    Low        3 18      20\n             ##    Medium    24 15      43\n\niv)\n\n       Test if the returns of FI and IT sectors are independent of each other\n       chisq.test(FI,IT)                                                      [3]\n\n       ##\n       ## Pearson's Chi-squared test\n       ##\n       ## data: FI and IT\n       ## X-squared = 15.146, df = 4, p-value = 0.004407\n\n       Interpretation                                                              [2]\n\n                                                                           Page 9 of 13\n\f      IAI                                                                  CS1B-1119\n\n           p-value < 0.05 indicating the rejection of null hypothesis of independe\n            nce of returns of both the sectors\n\n           There is a lot of interdependence in the sectoral returns\n\n           There were very few instances where one sector's returns were below the\n            Q1 and the other sector's returns were above Q3\n\n  The numbers in the diagonals are much higher to the ones in the off\n  diagonals indicating the strength of the relationship\nv)\n      Test whether the returns of FI sector are significantly higher compared to\n      that of IT Sector\n\n      t.test(indices$FI,indices$IT, alternative = “greater”)                      [3]\n\n      ##\n      ## Welch Two Sample t-test\n      ##\n      ## data: indices$FI and indices$IT\n      ## t = 0.22935, df = 314.81, p-value = 0.4094\n      ## alternative hypothesis: true difference in means is greater than 0\n      ## 95 percent confidence interval:\n      ## -0.012103 Inf\n      ## sample estimates:\n      ##   mean of x   mean of y\n      ## 0.010674390 0.008720122\n\n      Interpretation                                                              [1]\n\n           p-value =0.4094 > 0.05 indicates failure to reject the null hypothesis\n            of the return of FI not greater than that of IT sector at 95% confidenc\n            e level\n\n           There is no sufficient evidence to infer that the returns of FI sector\n            are significantly higher than that of IT Sector",
      "has_math": true,
      "is_r_task": true,
      "session": "2019-11",
      "subject": "CS1B",
      "source_qp": "raw/CS1B_2019-11_QP.pdf",
      "source_sol": "raw/CS1B_2019-11_SOL.pdf"
    },
    {
      "q_num": 5,
      "marks": 30,
      "topic": "regression_glm",
      "subtopics": [],
      "stem": "Refer to the data file “Indices_Returns.csv” and answer the following questions:\n\n        Indices_Returns.csv file is provided in the system.",
      "parts": [
        {
          "label": "i",
          "marks": 3,
          "text": "Load the csv file into R and create a new column called “Sensex_Direction”. The value\n           of this column will be “Positive” when the Sensex returns are positive and “Negative”\n           when they are negative and convert the variable as a factor variable.                           (3)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 6,
          "text": "Fit an appropriate generalized linear model (GLM) to with a ‘logit’ link function to relate\n            the “Sensex_Direction” with the returns of 10 sectors as a multivariate model and display\n            the summary of the model.                                                                      (6)",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 4,
          "text": "Identify which sectors have significantly impacted the direction of Sensex returns at 95%\n             and 99% confidence level.                                                                     (4)",
          "topic": null
        },
        {
          "label": "iv",
          "marks": 3,
          "text": "Verify the relationship between residual deviance of the model and the Akaike\n          Information Criteria (AIC).                                                                    (3)",
          "topic": null
        },
        {
          "label": "v",
          "marks": 4,
          "text": "Plot the residuals of the fitted model and identify which month is the most significant\n         outlier in the residuals.                                                                       (4)",
          "topic": null
        },
        {
          "label": "vi",
          "marks": 2,
          "text": "Comment on the appropriateness of the model fitted.                                            (2)\n\n      Your actuarial friend has suggested that the current model can be improved by removing the\n      variables which do not impact the direction of Sensex returns at 95% confidence level and\n      refitting the GLM with ‘logit’ link function.",
          "topic": null
        },
        {
          "label": "vii",
          "marks": 4,
          "text": "Update the model fitted in (ii) above, as suggested by your friend and display the summary\n           of the model.                                                                                 (4)",
          "topic": null
        },
        {
          "label": "viii",
          "marks": 4,
          "text": "Compare the models in (ii) and (vii) using an appropriate test and comment on the\n            difference in the residual deviances between the two models.                                 (4)",
          "topic": null
        }
      ],
      "solution": "i)\n            Creation of new column called Sensex_Direction\n            indices<-read.csv(\"D:\\\\Indices_Returns.csv\")\n            indices$Sensex_Direction<-ifelse(indices$Sensex>0,\"Positive\",\"Negative\"\n            )                                                                   [2]\n            indices$Sensex_Direction<-as.factor(indices$Sensex_Direction)        [1]\n\nii)\n            Fit the model and display the summary\n\n            model1<-glm(Sensex_Direction~BM+CD+EN+FI+FM+HC+IN+IT+TE+UT,data = indic\n            es, family = binomial(link = \"logit\"))                               [4]\n\n            ## Warning: glm.fit: fitted probabilities numerically 0 or 1 occurred\n\n            summary(model1)                                                       [2]\n                                                                         Page 10 of 13\n\f       IAI                                                                  CS1B-1119\n             ##\n             ## Call:\n             ## glm(formula = Sensex_Direction ~ BM + CD + EN + FI + FM + HC +\n             ##     IN + IT + TE + UT, family = binomial(link = \"logit\"), data = ind\n             ices)\n             ##\n             ## Deviance Residuals:\n             ##       Min        1Q    Median        3Q      Max\n             ## -2.27544 -0.00117     0.00000   0.01354  1.75651\n             ##\n             ## Coefficients:\n             ##             Estimate Std. Error z value Pr(>|z|)\n             ## (Intercept) -1.0086       0.7315 -1.379 0.16796\n             ## BM             7.7977    16.5255   0.472 0.63703\n             ## CD          -87.5335     42.6785 -2.051 0.04027 *\n             ## EN            93.9675    38.3193   2.452 0.01420 *\n             ## FI          172.8807     60.8192   2.843 0.00448 **\n             ## FM            41.1745    20.1436   2.044 0.04095 *\n             ## HC            -6.4294    13.9394 -0.461 0.64463\n             ## IN             4.1735    18.2152   0.229 0.81877\n             ## IT            78.3494    30.9307   2.533 0.01131 *\n             ## TE            29.9111    13.4184   2.229 0.02581 *\n             ## UT          -14.4767     23.0602 -0.628 0.53015\n             ## ---\n             ## Signif. codes: 0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1\n             ##\n             ## (Dispersion parameter for binomial family taken to be 1)\n             ##\n             ##     Null deviance: 223.213 on 163 degrees of freedom\n             ## Residual deviance: 32.905 on 153 degrees of freedom\n             ## AIC: 54.905\n             ##\n             ## Number of Fisher Scoring iterations: 11\n\niii)\n       Sectors significantly impacted\n\n            Sectors which have significantly impacted the direction of Sensex retur\n             ns are CD, EN, FI, FM, IT and TE at 95% Confidence level            [2]\n\n            But only FI has impacted the Sensex direction at 99% Confidence level\niv)\n       Relationship between the residual deviance and AIC\n\n            all.equal(AIC(model1), model1$deviance+2*11)                          [3]\n\n            (11 number of model parameters (the number of variables in the model pl\n             us the intercept))\n\n       ## [1] TRUE\n\n v)\n             Plot the residuals of Model\n             plot(model1$residuals)                                                [2]\n\n                                                                          Page 11 of 13\n\f       IAI                                                                   CS1B-1119\n\n             which(model1$residuals==min(model1$residuals))\n\n             ## 41\n             ## 41\n\n             indices$Month[which(model1$residuals==min(model1$residuals))]        [2]\n\n             ## [1] Jun-09\n\n             ## 164 Levels: Apr-06 Apr-07 Apr-08 Apr-09 Apr-10 Apr-11 Apr-12 ... Sep\n             -19\n\nvi)\n\n       Interpretation                                                             [2]\n\n            As the residual deviance came down significantly from Null Deviance of\n             223.21 to 32.90, the variables are able to classify the direction appro\n             priately\n\n            One huge outlier (Jun-09) can impact the accuracy of the result (Removi\n             ng this may reduce the residual deviance further)\n\n            The independent variables are not independent and they are interdepende\n             nt (Correlations are very high among the sectors). Hence the standard e\n             rrors may not be appropriate\n\n       As this data is a time series data serial correlation between the observat\n       ions need to be considered and the model may have to be fitted after remov\n       ing serial correlation.\n\nvii)\n             Remove the variables that do not impact the model\n             model2<-update(model1,~.-BM-HC-IN-UT)                                [3]\n\n             ## Warning: glm.fit: fitted probabilities numerically 0 or 1 occurred\n\n             summary(model2)                                                      [1]\n\n                                                                         Page 12 of 13\n\f        IAI                                                                 CS1B-1119\n              ##\n              ## Call:\n              ## glm(formula = Sensex_Direction ~ CD + EN + FI + FM + IT + TE,\n              ##     family = binomial(link = \"logit\"), data = indices)\n              ##\n              ## Deviance Residuals:\n              ##       Min        1Q    Median        3Q      Max\n              ## -2.08262 -0.00132     0.00000   0.01757  1.84630\n              ##\n              ## Coefficients:\n              ##             Estimate Std. Error z value Pr(>|z|)\n              ## (Intercept) -0.9623       0.6803 -1.414 0.15722\n              ## CD          -98.1771     37.5952 -2.611 0.00902 **\n              ## EN            94.2632    33.9930   2.773 0.00555 **\n              ## FI          174.4732     53.3659   3.269 0.00108 **\n              ## FM            37.7105    15.1808   2.484 0.01299 *\n              ## IT            78.7460    26.1203   3.015 0.00257 **\n              ## TE            30.7688    11.8875   2.588 0.00964 **\n              ## ---\n              ## Signif. codes: 0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1\n              ##\n              ## (Dispersion parameter for binomial family taken to be 1)\n              ##\n              ##     Null deviance: 223.213 on 163 degrees of freedom\n              ## Residual deviance: 33.646 on 157 degrees of freedom\n              ## AIC: 47.646\n              ##\n              ## Number of Fisher Scoring iterations: 10\n\nviii)\n        anova(model1,model2,test = \"Chisq\")                                        [3]\n\n        ## Analysis of Deviance Table\n        ##\n        ## Model 1: Sensex_Direction ~ BM + CD + EN + FI + FM + HC + IN + IT + TE+\n        ##     UT\n        ## Model 2: Sensex_Direction ~ CD + EN + FI + FM + IT + TE\n        ##   Resid. Df Resid. Dev Df Deviance Pr(>Chi)\n        ## 1       153     32.905\n        ## 2       157     33.646 -4 -0.74102   0.9462\n\n        Interpretation                                                             [1]\n\n             p-value of the comparison is 0.94 > 0.05 thus not rejecting the null hy\n              pothesis of no significant difference between the two models. So the mo\n              del did not improve significantly based on the friend's suggestion\n\n                                      *******************\n\n                                                                          Page 13 of 13",
      "has_math": false,
      "is_r_task": true,
      "session": "2019-11",
      "subject": "CS1B",
      "source_qp": "raw/CS1B_2019-11_QP.pdf",
      "source_sol": "raw/CS1B_2019-11_SOL.pdf"
    },
    {
      "q_num": 1,
      "marks": 24,
      "topic": "distributions",
      "subtopics": [],
      "stem": "Use the ‘MotorClaims’        dataset   provided   to   answer    the   following   questions:\n        (MotorClaims.csv)",
      "parts": [
        {
          "label": "i",
          "marks": 8,
          "text": "Fit Gamma distribution on the dataset provided by determining its scale and shape\n               parameters. State clearly the distribution with the parameters.                            (8)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 5,
          "text": "Simulate 1000 values from the distribution obtained in question (i) and print the first\n               six values. (Set seed to 100)                                                              (5)",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 2,
          "text": "Calculate the mean and variance of simulated values.                                       (2)",
          "topic": null
        },
        {
          "label": "iv",
          "marks": 4,
          "text": "Obtain a QQ plot for the simulations of 1000 values and a normal distribution.             (4)",
          "topic": null
        },
        {
          "label": "v",
          "marks": 2,
          "text": "Add a line to the above plot to show the true position of normal distribution.             (2)",
          "topic": null
        },
        {
          "label": "vi",
          "marks": 3,
          "text": "Comment on the shape of the distribution and how close it is to a normal distribution.     (3)",
          "topic": null
        }
      ],
      "solution": "i)\n> claims <- read.csv(\"MotorClaims.csv\")\n> mean = mean(claims$Claims)\n> mean\n[1] 18672.76                                                                 [1]\n> stddev = sd(claims$Claims)\n> variance = stddev ^ 2\n> variance\n[1] 161323921                                                                [1]\n> lambda <- mean/variance\n> lambda\n[1] 0.000115747                                                              [2]\n> alpha <- mean * lambda\n> alpha\n[1] 2.161316                                                                 [2]\n\nX ~ Gamma (2.16, 0.0001)                                                     [2]\n\nii)\n> set.seed(100)\n> samples <- rgamma(1000,alpha,lambda)                                       [2]\n> head(samples,6)                                                            [1]\n[1] 9305.461 2125.292 25926.442 15685.099 18120.436   8605.442               [2]\n\niii)\n> mean(samples)\n[1] 18423.47\n> variance <- sd(samples) ^ 2\n> variance\n[1] 153958637                                                                [2]\n\niv)\n> qqnorm(samples)\n\n                                                                     Page 2 of 8\n\fIAI                                                                                          CS1B-1120\nv)\n> qqline(samples,col=\"red\")\n\nvi)\nClose to normal…(1 mark) in the middle values…(1 mark).\n‘Banana-shaped’ indicates positively skewed… (1 mark).                                                  [3]",
      "has_math": false,
      "is_r_task": true,
      "session": "2020-11",
      "subject": "CS1B",
      "source_qp": "raw/CS1B_2020-11_QP.pdf",
      "source_sol": "raw/CS1B_2020-11_SOL.pdf"
    },
    {
      "q_num": 2,
      "marks": 19,
      "topic": "distributions",
      "subtopics": [
        "data_analysis"
      ],
      "stem": "The dataset ‘mtcars’ (built into R) consists of data on various models of car, taken from an\n        American motoring magazine (1974 Motor Trend magazine). For each car, there are certain\n        features expressed in varying units.",
      "parts": [
        {
          "label": "i",
          "marks": 4,
          "text": "Load the ‘mtcars’ dataset which is built into R. How many observations and variables\n               are there in this dataset? Your answer should include the R output.                        (4)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 5,
          "text": "Identify the categorical variables from the dataset ‘mtcars’ and create a dataset\n               excluding the categorical variables.                                                       (5)",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 2,
          "text": "How many observations and variables are there in the new dataset? Your answer\n               should include the R output.                                                               (2)",
          "topic": null
        },
        {
          "label": "iv",
          "marks": 5,
          "text": "Carry out a principal component analysis on the new dataset of mtcars by passing two\n               arguments, ‘center’ and ‘scale’ to be TRUE. Your answer should include a summary\n               of the principal component analysis.                                                       (5)",
          "topic": null
        },
        {
          "label": "v",
          "marks": 3,
          "text": "How many components of the reduced data should be retained using the output\n               derived in question (iv)? Also state the reason for the same.                              (3)",
          "topic": null
        }
      ],
      "solution": "i)\ndata(\"mtcars\")\n> str(mtcars)\n'data.frame': 32 obs. of 11 variables:\n $ mpg : num 21 21 22.8 21.4 18.7 18.1 14.3 24.4 22.8 19.2 ...\n $ cyl : num 6 6 4 6 8 6 8 4 4 6 ...\n $ disp: num 160 160 108 258 360 ...\n $ hp : num 110 110 93 110 175 105 245 62 95 123 ...\n $ drat: num 3.9 3.9 3.85 3.08 3.15 2.76 3.21 3.69 3.92 3.92 ...\n $ wt : num 2.62 2.88 2.32 3.21 3.44 ...\n $ qsec: num 16.5 17 18.6 19.4 17 ...\n $ vs : num 0 0 1 1 0 1 0 1 1 1 ...\n $ am : num 1 1 1 0 0 0 0 0 0 0 ...\n $ gear: num 4 4 4 3 3 3 3 4 4 4 ...\n $ carb: num 4 4 1 1 2 1 4 2 2 4 ...\n\nThere are 32 observations (car models) and 11 variables (car features) in the dataset.                   [4]\n\nii)\nsummary(mtcars)\n      mpg                   cyl                  disp                   hp              drat\n Min.   :10.40         Min.   :4.000        Min.   : 71.1         Min.   : 52.0    Min.   :2.760\n 1st Qu.:15.43         1st Qu.:4.000        1st Qu.:120.8         1st Qu.: 96.5    1st Qu.:3.080\n Median :19.20         Median :6.000        Median :196.3         Median :123.0    Median :3.695\n Mean   :20.09         Mean   :6.188        Mean   :230.7         Mean   :146.7    Mean   :3.597\n 3rd Qu.:22.80         3rd Qu.:8.000        3rd Qu.:326.0         3rd Qu.:180.0    3rd Qu.:3.920\n Max.   :33.90         Max.   :8.000        Max.   :472.0         Max.   :335.0    Max.   :4.930\n       wt                   qsec                  vs                     am               gear\n Min.   :1.513         Min.   :14.50        Min.   :0.0000         Min.    :0.0000   Min.    :3.000\n 1st Qu.:2.581         1st Qu.:16.89        1st Qu.:0.0000         1st Qu.:0.0000    1st Qu.:3.000\n Median :3.325         Median :17.71        Median :0.0000         Median :0.0000    Median :4.000\n Mean   :3.217         Mean   :17.85        Mean   :0.4375         Mean    :0.4062   Mean    :3.688\n 3rd Qu.:3.610         3rd Qu.:18.90        3rd Qu.:1.0000         3rd Qu.:1.0000    3rd Qu.:4.000\n\n                                                                                                 Page 3 of 8\n\fIAI                                                                                                    CS1B-1120\n  Max.   :5.424        Max.       :22.90        Max.    :1.0000       Max.      :1.0000     Max.      :5.000\n       carb\n  Min.   :1.000\n  1st Qu.:2.000\n  Median :2.000\n  Mean   :2.812\n  3rd Qu.:4.000\n  Max.   :8.000\n\nThe two variables ‘vs’ and ‘am’ are categorical variables. (This can be identified using str or summary function)\nmtcars1 <- mtcars[,c(1:7,10,11)]\n\niii)\n> str(mtcars1)\n'data.frame': 32 obs. of 9 variables:\n $ mpg : num 21 21 22.8 21.4 18.7 18.1 14.3 24.4 22.8 19.2 ...\n $ cyl : num 6 6 4 6 8 6 8 4 4 6 ...\n $ disp: num 160 160 108 258 360 ...\n $ hp : num 110 110 93 110 175 105 245 62 95 123 ...\n $ drat: num 3.9 3.9 3.85 3.08 3.15 2.76 3.21 3.69 3.92 3.92 ...\n $ wt : num 2.62 2.88 2.32 3.21 3.44 ...\n $ qsec: num 16.5 17 18.6 19.4 17 ...\n $ gear: num 4 4 4 3 3 3 3 4 4 4 ...\n $ carb: num 4 4 1 1 2 1 4 2 2 4 ...\n\nThere are 32 observations (car models) and 9 variables (car features) in the dataset.                                [2]\n\niv)\nmtcars1.pca <- prcomp(mtcars1,center = TRUE,scale=TRUE)                                                              [2]\n> summary(mtcars1.pca)                                                                                               [1]\nImportance of components:\n\n               PC1 PC2      PC3    PC4    PC5    PC6    PC7 PC8      PC9\n\nStandard deviation    2.3782 1.4429 0.71008 0.51481 0.42797 0.35184 0.32413 0.2419 0.14896\n\nProportion of Variance 0.6284 0.2313 0.05602 0.02945 0.02035 0.01375 0.01167 0.0065 0.00247\n\nCumulative Proportion 0.6284 0.8598 0.91581 0.94525 0.96560 0.97936 0.99103 0.9975 1.00000                           [2]\n\nv)\n\nThe R analysis shows that the proportion of variance explained by first three principal components is 91.5% and by\nfirst four variables is 94.5%.\n\nThus, it will be appropriate to retain the first three (or four) principal components.                             [3]",
      "has_math": false,
      "is_r_task": true,
      "session": "2020-11",
      "subject": "CS1B",
      "source_qp": "raw/CS1B_2020-11_QP.pdf",
      "source_sol": "raw/CS1B_2020-11_SOL.pdf"
    },
    {
      "q_num": 3,
      "marks": 24,
      "topic": "inference",
      "subtopics": [],
      "stem": "Data is provided for ‘BMIClaims’ of 150 policyholders and corresponding claim count.\n           (BMIClaims.csv)",
      "parts": [
        {
          "label": "i",
          "marks": 6,
          "text": "Obtain 95% confidence interval for the standard deviation of BMI (using qchisq).           (6)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 6,
          "text": "Further test the standard deviation of BMI to be equal to 4 by obtaining p value. State\n               your conclusion of the test.                                                               (6)",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 6,
          "text": "If obese is defined to be BMI above 30, use binom.test to calculate 99% confidence\n                interval for proportion of obese people and comment on the likelihood if more than\n                20 pc are obese.                                                                              (6)",
          "topic": null
        },
        {
          "label": "iv",
          "marks": 6,
          "text": "Claim frequency can be calculated as claim count divided by number of policyholders.\n                Test whether claim frequency is different between obese and others.                           (6)",
          "topic": null
        }
      ],
      "solution": "i)\n> BMI <- read.csv(\"BMIClaims.csv\")\n> n <- length(BMI$BMI)\n> alpha <- 0.05                                                   …                  [2]\n> sqrt(c((n-1)*var(BMI$BMI)/qchisq(1-alpha/2,df=n-1),(n-1)*var(BMI$BMI)/qchisq(alpha/2\n,df=n-1)))                                                                           [2]\n[1] 5.920028 7.434763                                                                [2]\n\n                                                                                                            Page 4 of 8\n\fIAI                                                                        CS1B-1120\nii)\n> sigma <- 4\n> statistic <- (n-1)*var(BMI$BMI)/sigma^2                                                [1]\n> statistic\n[1] 404.5421\n> qchisq(alpha/2,n-1)\n[1] 117.098\n> qchisq(alpha/2,n-1,lower=FALSE)\n[1] 184.687\n> 2*(pchisq((n-1)*var(BMI$BMI)/sigma^2,df=n-1,lower.tail=FALSE))                         [2]\n[1] 3.564503e-25                                                                         [1]\nSince p-value is less than 5%, there is sufficient evidence to reject the hypothesis,\ni.e. the standard deviation of BMI is not equal to 4.                                [2]\n\niii)\n> x <- nrow(BMI[BMI$BMI>30,])                                                            [1]\n> binom.test(x,n,conf.level = 0.99)                                                      [2]\n            Exact binomial test\n\ndata: x and n\nnumber of successes = 10, number of trials = 150, p-value < 2.2e-16\nalternative hypothesis: true probability of success is not equal to 0.5\n99 percent confidence interval:\n 0.02522882 0.13728337                                                                   [1]\nsample estimates:\nprobability of success\n            0.06666667\n\nSince 99% CI for p doesn’t contain p=0.2                                                 [1]\nit is unlikely that the proportion of obese policyholders is more than 20%.. ..          [1]\n\niv)\n> table(BMI$BMI>30,BMI$ClaimCount)\n\n               0   1\n       FALSE 133   7\n       TRUE    7   3\n\n> y <- c(3,7)\n> m <- c(10,140)                                                                         [2]\n> poisson.test(y,m)                                                                      [1]\n            Comparison of Poisson rates\n\ndata: y time base: m\ncount1 = 3, expected count1 = 0.66667, p-value = 0.02493\nalternative hypothesis: true rate ratio is not equal to 1\n95 percent confidence interval:\n  1.001171 26.282304\nsample estimates:\nrate ratio\n         6\n\nSince p-value is less than 5% i.e. 2.5%, there is sufficient evidence to reject the hy\npothesis, i.e. Claim frequency is different between obese and others.\n(Alternatively, can use prop.test)                                                       [6]\n\n                                                                                  Page 5 of 8\n\fIAI                                                                                                     CS1B-1120",
      "has_math": false,
      "is_r_task": true,
      "session": "2020-11",
      "subject": "CS1B",
      "source_qp": "raw/CS1B_2020-11_QP.pdf",
      "source_sol": "raw/CS1B_2020-11_SOL.pdf"
    },
    {
      "q_num": 4,
      "marks": 33,
      "topic": "regression_glm",
      "subtopics": [
        "distributions"
      ],
      "stem": "The CDC and EISS detect influenza activity through clinical data including Influenza-like\n        Illness (ILI) physician visits on weekly basis. The objective of this question is to estimate\n        influenza-like illness (ILI) activity using Google web search logs.\n\n        The csv file FluTrain.csv aggregates this data from January 1, 2004 until December 31, 2011\n        as follows:\n\n        \"Week\" - The range of dates represented by this observation, in year/month/day format.\n\n        \"ILI\" - This column lists the percentage of ILI-related physician visits for the corresponding\n        week.\n\n        \"Queries\" - This column lists the fraction of queries that are ILI-related for the corresponding\n        week, adjusted to be between 0 and 1 (higher values correspond to more ILI-related search\n        queries).",
      "parts": [
        {
          "label": "i",
          "marks": 3,
          "text": "Plot a histogram of the dependent variable, ILI. Comment on the shape of the\n                distribution.                                                                                 (3)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 4,
          "text": "Plot the natural logarithm of ILI versus Queries. What does the plot suggest?                 (4)",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 6,
          "text": "Fit a linear regression model for dependent variable log(ILI). Summarize it.                  (6)",
          "topic": null
        },
        {
          "label": "iv",
          "marks": 3,
          "text": "State the formula of the model fitted in part (iii), explaining all the terms used.           (3)",
          "topic": null
        },
        {
          "label": "v",
          "marks": 5,
          "text": "Calculate R-squared and the correlation between the independent and dependent\n                variable. What is the relationship between the two values?                                    (5)",
          "topic": null
        },
        {
          "label": "vi",
          "marks": 4,
          "text": "Looking at the time period 2004-2011, which week corresponds to the highest\n                percentage of ILI-related physician visits?                                                   (4)",
          "topic": null
        },
        {
          "label": "vii",
          "marks": 4,
          "text": "Based on the linear regression model fitted in question (iii), what is the estimate for\n                the percentage of ILI-related physician visits for the week computed in question (vi)?        (4)",
          "topic": null
        },
        {
          "label": "viii",
          "marks": 4,
          "text": "What is the relative error between the estimate (prediction calculated in question (vii))\n                and the actual observed value for the week computed in question (vi)?                         (4)\n\n      CS1B_BMIClaims\n\n      https://actuariesindia.org/sites/default/files/2022-10/CS1B_BMIClaims.csv\n\n      CS1B_FluTrain\n\n      https://actuariesindia.org/sites/default/files/2022-10/CS1B_FluTrain.csv\n\n      CS1B_MotorClaims\n\n      https://actuariesindia.org/sites/default/files/2022-10/CS1B_MotorClaims.csv",
          "topic": null
        }
      ],
      "solution": "i)\nsetwd(\"C:/Users/shrey/Downloads\")\nFluTrain <- read.csv(\"FluTrain.csv\")\n> str(FluTrain)\n'data.frame': 417 obs. of 3 variables:\n  $ Week   : Factor w/ 417 levels \"2004-01-04 - 2004-01-10\",..: 1 2 3 4 5 6 7 8 9 10 ..\n.\n  $ ILI    : num 2.42 1.81 1.71 1.54 1.44 ...\n  $ Queries: num 0.238 0.22 0.226 0.238 0.224 ...\nhist(FluTrain$ILI)\n\nThe data is positively skewed. Most of the ILI values are small, with a relatively small number of much larger values.\n\nii)\nplot(FluTrain$Queries,log(FluTrain$ILI))\n\nThere is a positive linear relationship between log(ILI) and Queries.\n\ni.e. more the number of the Google search queries, higher the number of ILI-related physician visits.\n\n                                                                                                            Page 6 of 8\n\fIAI                                                                                                 CS1B-1120\niii)\n\nFluTrend1 = lm(log(ILI) ~ Queries, data = FluTrain)                                                              [3]\n> summary (FluTrend1)                                                                                            [1]\nCall:\nlm(formula = log(ILI) ~ Queries, data = FluTrain)\n\nResiduals:\n     Min       1Q   Median                  3Q         Max\n-0.76003 -0.19696 -0.01657             0.18685     1.06450\n\nCoefficients:\n            Estimate Std. Error t value Pr(>|t|)\n(Intercept) -0.49934    0.03041 -16.42    <2e-16 ***\nQueries       2.96129   0.09312   31.80   <2e-16 ***\n---\nSignif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1\n\nResidual standard error: 0.2995 on 415 degrees of freedom\nMultiple R-squared: 0.709,    Adjusted R-squared: 0.7083\nF-statistic: 1011 on 1 and 415 DF, p-value: < 2.2e-16                                                           [2]\niv)\n\nln y = -0.49934 +2.96129x                                                                                       [2]\n\nwhere x is the google search queries and y is the percentage of ILI related physician visits.                   [1]\n\nv)\n\nFrom the R output, R-squared value is 0.709.                                                                    [1]\n\ncorrelation <- cor(FluTrain$Queries,log(FluTrain$ILI))                                                          [1]\n> correlation\n[1] 0.8420333                                                                                                   [1]\n> correlation ^ 2\n[1] 0.7090201\nHence, R-squared = Correlation ^ 2                                                                              [2]\n\nvi)\nwhich.max(FluTrain$ILI)\n[1] 303\n> FluTrain$Week[303]\n[1] 2009-10-18 - 2009-10-24\n417 Levels: 2004-01-04 - 2004-01-10 2004-01-11 - 2004-01-17 ... 2011-12-25 - 2011-12-3\n1\nWeek of 18th October 2009 to 24th October 2009 corresponds to the highest percentage of ILI-related physician visits.\n\nvii)\nPredTest1 = exp(predict(FluTrend1,newdata = FluTrain))                                                           [2]\n> PredTest1[303]\n     303\n11.72765                                                                                                         [2]\n\n                                                                                                         Page 7 of 8\n\fIAI                                                      CS1B-1120\nviii)\nFluTrain$ILI[303]\n[1] 7.618892                                                         [2]\n(7.618892-11.72765)/7.618892\n[1] -0.5392855                                                     [2]\n\n                               ***********************\n\n                                                             Page 8 of 8",
      "has_math": false,
      "is_r_task": true,
      "session": "2020-11",
      "subject": "CS1B",
      "source_qp": "raw/CS1B_2020-11_QP.pdf",
      "source_sol": "raw/CS1B_2020-11_SOL.pdf"
    },
    {
      "q_num": 1,
      "marks": 22,
      "topic": "inference",
      "subtopics": [],
      "stem": "The following amounts are the sizes of claims (in INR) on house insurance policies for a\n          certain type of repair.\n\n          1990, 2400, 2150, 2090, 2300, 2100, 2180, 2150, 2030, 2100, 2180, 2010, 2060, 2160, 2120",
      "parts": [
        {
          "label": "i",
          "marks": 1,
          "text": "Enter data in R.                                                                               (1)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 4,
          "text": "Calculate Q1, Q2, Q3 and Inter-quartile range.                                                (4)",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 2,
          "text": "Determine the sample mean and variance of the data.                                          (2)",
          "topic": null
        },
        {
          "label": "iv",
          "marks": 5,
          "text": "Test the hypothesis whether the mean claim amount is equal to INR 2000 and comment\n              on the results.                                                                               (5)",
          "topic": null
        },
        {
          "label": "v",
          "marks": 4,
          "text": "Assuming the data to be normally distributed, calculate the probability of a claim amount\n             exceeding INR 2300.                                                                            (4)",
          "topic": null
        },
        {
          "label": "vi",
          "marks": 6,
          "text": "Calculate the revised mean and median after removing the largest two values from the\n              dataset. Comment on the result.                                                               (6)",
          "topic": null
        }
      ],
      "solution": "i)\n> data = c(1990,2400,2150,2090,2300,2100,2180,2150,2030,2100,2180,2010,2060,2\n160,2120)…                                                                  [1]\n\nii)\n> quantile(data)\n  0% 25% 50% 75% 100%\n1990 2075 2120 2170 2400\n> IQR(data)\n[1] 95\nQ1 – INR 2,075; Q2 – INR 2,120; Q3 – INR 2,170 and IQR – INR 95\n\n                                                                 [1 MARK EACH FOR QUARTILE AND IQR]\niii)\n> mean(data)\n[1] 2134.667\n> var(data)\n[1] 11469.52\n                                                                  [1 MARK EACH FOR MEAN & VARIANCE]\niv)\n\nHo: The mean claim amount is INR 2000\nH1: Mean claim amount is not equal to INR 2000\n> t.test(data,mean=2000)\n\n          One Sample t-test\n\ndata: data\nt = 77.197, df = 14, p-value < 2.2e-16\nalternative hypothesis: true mean is not equal to 0\n95 percent confidence interval:\n 2075.359 2193.974\nsample estimates:\nmean of x\n 2134.667\n\nSince the p-value is 2.2*10^-16 is less than 5%, there is sufficient evidence to reject Ho of mean equal to\nINR 2000.\n\n[1 MARK FOR HYPOTHESIS, 1 MARK FOR T-TEST, 1 MARK FOR RESULTS, 1 MARK FOR MENTIONING P-\nVALUE & 1 MARK FOR CONCLUSION]\n\nv)\n> pnorm(2300,mean,sqrt(var),lower.tail = FALSE)\n[1] 0.06131982\n                                                              [3 MARKS FOR CODE, 1 MARK FOR RESULT]\n\nvi)\n\n                                                                                               Page 2 of 9\n\fIAI                                                                                         CS1B-0321\n\n> data1 = data[data < max(data)]\n> data1\n [1] 1990 2150 2090 2300 2100 2180 2150 2030 2100 2180 2010 2060 2160 2120\n> data2 = data1[data1 < max(data1)]\n> data2\n [1] 1990 2150 2090 2100 2180 2150 2030 2100 2180 2010 2060 2160 2120\n> mean(data2)\n[1] 2101.538\n> median(data2)\n[1] 2100\n\nImpact of removing outliers from the data has led to the mean and median being almost equal. Earlier the\nmean was higher than median, which shows that the mean is more likely to be affected by outliers.\n\n[2 MARKS FOR CALCULATING REVISED DATA, 1 MARK FOR MEAN, 1 MARK FOR MEDIAN, 2 MARKS FOR\nCOMMENTS]",
      "has_math": false,
      "is_r_task": true,
      "session": "2021-03",
      "subject": "CS1B",
      "source_qp": "raw/CS1B_2021-03_QP.pdf",
      "source_sol": "raw/CS1B_2021-03_SOL.pdf"
    },
    {
      "q_num": 2,
      "marks": 20,
      "topic": "bayes_credibility",
      "subtopics": [],
      "stem": "The prior and posterior distribution for values of systolic blood pressure follows Normal\n          distribution. Prior distribution of systolic blood pressure (x) has a mean of 120 and standard\n          deviation of 10.",
      "parts": [
        {
          "label": "i",
          "marks": 3,
          "text": "Generate range of values of x in the interval [80,160] and using len = 100.                    (3)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 6,
          "text": "Plot the posterior probability density function of x using answer from (i).                   (6)",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 2,
          "text": "Perform a simulation of 1000 posterior samples for the parameter x.                          (2)",
          "topic": null
        },
        {
          "label": "iv",
          "marks": 3,
          "text": "Plot a histogram of the posterior distribution of x.                                          (3)",
          "topic": null
        },
        {
          "label": "v",
          "marks": 2,
          "text": "What is the mean and standard deviation of posterior distribution of x?                        (2)",
          "topic": null
        },
        {
          "label": "vi",
          "marks": 4,
          "text": "Calculate the 95% confidence interval for systolic blood pressure using the posterior\n              distribution.                                                                                 (4)",
          "topic": null
        }
      ],
      "solution": "i)\n> priormean <- 120\n> priorsd <- 10\n> x <- seq(80,160,len = 100)\n> x\n  [1] 80.00000 80.80808 81.61616 82.42424 83.23232 84.04040 84.84848\n85.65657 86.46465\n [10] 87.27273 88.08081 88.88889 89.69697 90.50505 91.31313 92.12121\n92.92929 93.73737\n [19] 94.54545 95.35354 96.16162 96.96970 97.77778 98.58586 99.39394 1\n00.20202 101.01010\n [28] 101.81818 102.62626 103.43434 104.24242 105.05051 105.85859 106.66667 1\n07.47475 108.28283\n [37] 109.09091 109.89899 110.70707 111.51515 112.32323 113.13131 113.93939 1\n14.74747 115.55556\n [46] 116.36364 117.17172 117.97980 118.78788 119.59596 120.40404 121.21212 1\n22.02020 122.82828\n [55] 123.63636 124.44444 125.25253 126.06061 126.86869 127.67677 128.48485 1\n29.29293 130.10101\n [64] 130.90909 131.71717 132.52525 133.33333 134.14141 134.94949 135.75758 1\n36.56566 137.37374\n [73] 138.18182 138.98990 139.79798 140.60606 141.41414 142.22222 143.03030 1\n43.83838 144.64646\n [82] 145.45455 146.26263 147.07071 147.87879 148.68687 149.49495 150.30303 1\n51.11111 151.91919\n [91] 152.72727 153.53535 154.34343 155.15152 155.95960 156.76768 157.57576 1\n58.38384 159.19192\n[100] 160.00000\n\n                                                          [2 MARKS FOR CODE, 1 MARK FOR RESULTS]\nii)\n\n> y<-dnorm(x,mean = priormean,sd=priorsd)….                                                          [4]\n\n                                                                                             Page 3 of 9\n\fIAI                                                                     CS1B-0321\n\n> plot(x,y,type = \"l\")….                                                      [1]\n\niii)\n> z <- rnorm(1000,priormean,priorsd)…                                          [2]\n\niv)\n> hist(z,main = \"Posterior distribution of x\")\n\n                                             [2 MARKS FOR CODE, 1 MARK FOR GRAPH]\n\nv)\n> mean(z)\n[1] 119.8336\n> sd(z)\n[1] 10.06614\n                                                 [1 MARK FOR MEAN, 1 MARK FOR SD]\n\n                                                                        Page 4 of 9\n\fIAI                                                                                              CS1B-0321\n\nvi)\n> sbp <- mean(z) + qnorm(c(0.025,0.975)) * sd(z)\n> sbp\n[1] 100.1044 139.5629\n\n                                             [2 MARKS FOR CODE, 2 MARKS FOR CONFIDENCE INTERVAL]",
      "has_math": false,
      "is_r_task": true,
      "session": "2021-03",
      "subject": "CS1B",
      "source_qp": "raw/CS1B_2021-03_QP.pdf",
      "source_sol": "raw/CS1B_2021-03_SOL.pdf"
    },
    {
      "q_num": 3,
      "marks": 19,
      "topic": "inference",
      "subtopics": [],
      "stem": "An agency has collected data on the number of COVID19 cases of two cities in order to\n          analyse the similarities & differences between them. Below is the data for two cities on\n          monthly basis.\n           Month City A City B\n              1       9150       8919\n              2       9418       9095\n              3       9218       9046\n              4       9539       9321\n              5       9179       9719\n              6       8907       9704\n              7       9472       9107\n              8       8921       9275",
      "parts": [
        {
          "label": "i",
          "marks": 1,
          "text": "Enter data in R.                                                                                      (1)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 8,
          "text": "Test at 5% level with clearly mentioning the hypothesis, if there is a difference in the\n              mean of the two sample data assuming equal & unknown variance.                                       (8)",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 5,
          "text": "Test whether the variances are equal at 5% level and comment on the results.                        (5)",
          "topic": null
        },
        {
          "label": "iv",
          "marks": 3,
          "text": "Calculate the 95% confidence interval for the difference in means.                                   (3)",
          "topic": null
        },
        {
          "label": "v",
          "marks": 2,
          "text": "Comment on your findings in part (ii) and part (iv).                                                  (2)",
          "topic": null
        }
      ],
      "solution": "i)\n> city1 = c(9150,9418,9218,9539,9179,8907,9472,8921)\n> city2 = c(8919,9095,9046,9321,9719,9704,9107,9275)…                                                        [1]\n\nii)\n\nHo: There is no difference in the average number of monthly COVID19 cases between two cities.\n\nH1: There is a difference in the average number of monthly COVID19 cases between two cities.\n> t.test(x=city1,y=city2,var.equal = TRUE,conf.level = 0.95)\n\n          Two Sample t-test\n\ndata: city1 and city2\nt = -0.35359, df = 14, p-value = 0.7289\nalternative hypothesis: true difference in means is not equal to 0\n95 percent confidence interval:\n -337.3886 241.8886\nsample estimates:\nmean of x mean of y\n  9225.50   9273.25\n\nSince the p-value is 0.7289 is significantly greater than 5%, there is insufficient evidence to reject Ho.\nThus, we have no evidence to suggest that the means are different between the two samples.\n\n[2 MARKS FOR HYPOTHESIS, 2 MARKS FOR T-TEST, 1 MARK FOR RESULTS, 1 MARK FOR MENTIONING\nP-VALUE, 2 MARKS FOR CONCLUSION]\niii)\n> var.test(x=city1,y=city2,conf.level = 0.95)\n\n          F test to compare two variances\n\ndata: city1 and city2\nF = 0.63907, num df = 7, denom df = 7, p-value = 0.5691\nalternative hypothesis: true ratio of variances is not equal to 1\n95 percent confidence interval:\n 0.1279433 3.1920724\nsample estimates:\nratio of variances\n         0.6390651\n\nThe p-value is 0.5691 > 5%, so there is insufficient evidence to reject the assumption of equal variance.\n\n[2 MARKS FOR VAR TEST, 1 MARK FOR RESULTS, 1 MARK FOR P-VALUE, 1 MARK FOR CONCLUSION]\n\n                                                                                                  Page 5 of 9\n\fIAI                                                                                        CS1B-0321\n\niv)\n\nConfidence interval can be read from Part (b) or can be derived as below:\n> t.test(x=city1,y=city2,var.equal = TRUE,conf.level = 0.95)$conf.int\n[1] -337.3886 241.8886\nattr(,\"conf.level\")\n[1] 0.95\ni.e. 95% CI is (-337.39, 241.89)\n\n                                          [1 MARKS FOR CODE, 2 MARKS FOR CONFIDENCE INTERVAL]\n\nv)\n\nThe confidence interval (-337,241) contains 0, therefore the assumption of equal means holds.\n\n                                  [1 MARK FOR MENTIONING CONTAINS 0, 1 MARK FOR CONCLUSION]",
      "has_math": false,
      "is_r_task": true,
      "session": "2021-03",
      "subject": "CS1B",
      "source_qp": "raw/CS1B_2021-03_QP.pdf",
      "source_sol": "raw/CS1B_2021-03_SOL.pdf"
    },
    {
      "q_num": 4,
      "marks": 39,
      "topic": "inference",
      "subtopics": [],
      "stem": "Five years of marketing spend and company sales by month",
      "parts": [
        {
          "label": "i",
          "marks": 4,
          "text": "Construct a scatterplot of the data. Comment on the relationship between the Sales &\n             Spend based on the plot.                                                                              (4)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 2,
          "text": "Calculate Pearson’s correlation coefficient between Sales and Spend of the company.                  (2)",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 5,
          "text": "Perform a hypothesis test for the null hypothesis that Pearson’s population correlation\n               coefficient is equal to zero, against the alternative that it is positive. You should report the\n               p-value of the test and a clear conclusion.                                                         (5)",
          "topic": null
        },
        {
          "label": "iv",
          "marks": 6,
          "text": "Perform a simple linear regression analysis on the data. Your answer should report the\n              estimate of parameter sigma.                                                                         (6)",
          "topic": null
        },
        {
          "label": "v",
          "marks": 2,
          "text": "Plot the fitted line on the data scatterplot.                                                         (2)",
          "topic": null
        },
        {
          "label": "vi",
          "marks": 1,
          "text": "State the proportion of the total variability of the responses explained by the model based\n              on your output in (iv).                                                                              (1)",
          "topic": null
        },
        {
          "label": "vii",
          "marks": 2,
          "text": "Plot a graph of the residuals of the model fitted in (iv) against the explanatory variable.         (2)",
          "topic": null
        },
        {
          "label": "viii",
          "marks": 4,
          "text": "Obtain a 99% confidence interval for parameter sigma.                                              (4)",
          "topic": null
        },
        {
          "label": "ix",
          "marks": 2,
          "text": "Comment on the validity of the model based on results in part (vii) and part (viii).                 (2)",
          "topic": null
        },
        {
          "label": "x",
          "marks": 7,
          "text": "Calculate the p-value of a hypothesis test for this suggestion (slope equal to 10), by\n             creating a suitable test statistic.                                                                   (7)",
          "topic": null
        },
        {
          "label": "xi",
          "marks": 2,
          "text": "Comment on the suggestion in point (x).                                                              (2)",
          "topic": null
        },
        {
          "label": "xii",
          "marks": 2,
          "text": "Calculate the predicted amount of sales when the marketing spend is INR 4500.                       (2)",
          "topic": null
        }
      ],
      "solution": "i)\n> budget = read.csv(\"marketingbudget.csv\")\n> plot(budget$Spend,budget$Sales)\n\nThe above scatter plot shows a positive linear relationship between marketing Spend and Sales data.\n\n[1 MARK FOR CODE, 1 MARK FOR SCATTER PLOT, 1 MARK FOR MENTIONING POSITIVE, 1 MARK FOR\nLINEAR]\n\nii)\n> cor = cor(budget$Sales,budget$Spend)\n\n> cor\n[1] 0.9701669\n                                                             [1 MARK FOR CODE, 1 MARK FOR RESULT]\n\n                                                                                            Page 6 of 9\n\fIAI                                                                                            CS1B-0321\n\niii)\n> cor.test(budget$Spend,budget$Sales,method=\"pearson\",alternative = \"greater\"\n)\n\n          Pearson's product-moment correlation\n\ndata: budget$Spend and budget$Sales\nt = 30.476, df = 58, p-value < 2.2e-16\nalternative hypothesis: true correlation is greater than 0\n95 percent confidence interval:\n 0.9542479 1.0000000\nsample estimates:\n      cor\n0.9701669\nThe p-value is 2.2 X 10^-16, showing very strong evidence against the null hypothesis. Thus, we reject that\nthe Pearson’s correlation coefficient is equal to 0 and conclude that it is positive.\n\n[2 MARKS FOR COR TEST, 1 MARK FOR RESULTS, 1 MARK FOR P-VALUE, 1 MARK FOR CONCLUSION]\n\niv)\n> reg = lm(Sales ~ Spend, data = budget)\n> summary(reg)\n\nCall:\nlm(formula = Sales ~ Spend, data = budget)\n\nResiduals:\n     Min      1Q           Median          3Q         Max\n-25331.9 -6783.1           -844.5      7965.9     25320.1\n\nCoefficients:\n             Estimate Std. Error t value Pr(>|t|)\n(Intercept) 3431.5592 3245.9169    1.057    0.295\nSpend         10.5310     0.3455 30.476    <2e-16 ***\n---\nSignif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1\n\nResidual standard error: 10650 on 58 degrees of freedom\nMultiple R-squared: 0.9412, Adjusted R-squared: 0.9402\nF-statistic: 928.8 on 1 and 58 DF, p-value: < 2.2e-16\n\nFrom the output, the estimate of parameter sigma is 10,650.\n\n                 [3 MARKS FOR REG CODE, 2 MARKS FOR RESULTS, 1 MARK FOR ESTIMATE OF SIGMA]\n\nv)\n> abline(reg)\n\n                                                                                               Page 7 of 9\n\fIAI                                                                                          CS1B-0321\n\n                                                              [1 MARK FOR CODE, 1 MARK FOR GRAPH]\nvi)\n\nFrom the R output, the proportion of total variability of the responses explained by the model is 94.12%.\n\nvii)\n> plot(budget$Spend,residuals(reg))\n\n                                                              [1 MARK FOR CODE, 1 MARK FOR GRAPH]\nviii)\nes = resid(reg)\n> t.test(es,conf.level = 0.99)$conf.int\n[1] -3630.146 3630.146\nattr(,\"conf.level\")\n[1] 0.99\n\n                                                                                              Page 8 of 9\n\fIAI                                                                                              CS1B-0321\n\nFrom the above, the confidence interval for parameter sigma is (-3630.15, 3630.15)\n\n[1 MARK FOR EVALUATING RESIDUALS, 1 MARK FOR T-TEST, 2 MARKS FOR CONFIDENCE INTERVAL]\n\nix)\n\nBased on the results in both part (vii) and part (viii), the errors seem to be close to zero and the\nconfidence interval of residuals also contains 0. Hence the model seems to be a good fit.\n\nx)\nLet Ho: Beta = 10 and H1: Beta not equal to 10\n\n> b1 = (coef(reg))[['Spend']]…                                                                            [1]\n> n = 60\n> s = sqrt(sum(es^2)/(n-2))\n> SE = s/sqrt(sum((budget$Spend-mean(budget$Spend))^2))..                                                 [2]\n> t = (b1-10)/SE …                                                                                        [1]\n> pt(t,58,lower.tail = FALSE)…                                                                            [1]\n[1] 0.06489565\n> pvalue = 2*pt(t,58,lower.tail = FALSE)..                                                                [1]\n> pvalue\n[1] 0.1297913..                                                                                           [1]\n\nxi)\n\nThere is insufficient evidence to reject the null hypothesis at 5% level of significance. The slope is equal\nto 10 for this data.\n\nxii)\n> y = 3431.5592 + (b1*4500)\n> y\n[1] 50821.17\nWith a marketing spend of INR 4,500, the Sales would be INR 50,821.\n\n[1 MARK FOR CODE, 1 MARK FOR RESULTS, DEDUCT ½ MARK FOR NOT MENTIONING INR]\n\n                                            *****************\n\n                                                                                                  Page 9 of 9",
      "has_math": false,
      "is_r_task": true,
      "session": "2021-03",
      "subject": "CS1B",
      "source_qp": "raw/CS1B_2021-03_QP.pdf",
      "source_sol": "raw/CS1B_2021-03_SOL.pdf"
    },
    {
      "q_num": 1,
      "marks": 5,
      "topic": "distributions",
      "subtopics": [],
      "stem": "Assume X is a random variable which follows Poi (2) where X = 1,2,3,4,5,6,7,8,9,10 i.e.\n          values from 1 to 10",
      "parts": [
        {
          "label": "i",
          "marks": 2,
          "text": "Calculate the probability for each of the values of X.                                   (2)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 2,
          "text": "Calculate the cumulative probability for each of the values of X.                       (2)",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 1,
          "text": "Plot a graph of the distribution function of X.                                        (1)",
          "topic": null
        }
      ],
      "solution": "X = c(1:10)\n         Lambda = 2                                                                             (1)\n\n   i)     dpois (x, Lambda)\n          >>Outcome:\n          2.706706e-01 2.706706e-01 1.804470e-01 9.022352e-02 3.608941e-02 1.202980e-02\n          3.437087e-03 8.592716e-04 1.909493e-04 3.818985e-05                             (1.5)\n   ii)    ppois (x, Lambda)\n          >> Outcome\n          0.4060058 0.6766764 0.8571235 0.9473470 0.9834364 0.9954662 0.9989033 0.9997626\n          0.9999535 0.9999917                                                             (1.5)\n   iii)   plot (ppois(x,Lambda)",
      "has_math": false,
      "is_r_task": true,
      "session": "2021-09",
      "subject": "CS1B",
      "source_qp": "raw/CS1B_2021-09_QP.pdf",
      "source_sol": "raw/CS1B_2021-09_SOL.pdf"
    },
    {
      "q_num": 2,
      "marks": 12,
      "topic": "data_analysis",
      "subtopics": [
        "distributions"
      ],
      "stem": "A student is performing study to understand the correlation of temperature among\n          different days of the week. Temperature on weekdays is recorded for 15 weeks in the file\n          Temperature_data.csv.",
      "parts": [
        {
          "label": "i",
          "marks": 3,
          "text": "Generate,\n\n                  a)      Pearson’s correlation matrix                                                 (3)\n\n                  b)      Kendall’s rank                                                               (3)\n\n                  c)      Spearman’s rank                                                              (3)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 3,
          "text": "Comment on the outcomes of the results in a, b and c.                                   (3)",
          "topic": null
        }
      ],
      "solution": "i)    a) Setwd()\n         Data = read.csv(“Temperature_data.csv”)\n\n          cor (Data, method = “pearson”)\n                         Monday Tuesday Wednesday Thursday Friday\n           Monday        1.0000 0.9566   0.5393    0.8236 0.9271\n           Tuesday       0.9566 1.0000   0.7612    0.9484 0.9776\n           Wednesday 0.5393 0.7612       1.0000    0.9080 0.7620\n           Thursday      0.8236 0.9484   0.9080    1.0000 0.9632\n           Friday        0.9271 0.9776   0.7620    0.9632 1.0000                                (3)\n\n          b) cor (Data, method = “kendall”)\n                         Monday Tuesday Wednesday Thursday Friday\n           Monday        1.0000 0.6952      0.2952 0.4667  0.6762\n           Tuesday       0.6952 1.0000      0.6000 0.7714  0.8286\n           Wednesday 0.2952 0.6000          1.0000 0.7905  0.5810\n           Thursday      0.4667 0.7714      0.7905 1.0000  0.7905\n           Friday        0.6762 0.8286      0.5810 0.7905  1.0000                               (3)\n\n          c) cor (Data, method = “spearman”)\n                                                                                 Page 2 of 15\n\f   IAI                                                                                             CS1B-0921\n                     Monday Tuesday Wednesday Thursday Friday\n           Monday    1.0000 0.8643   0.4036    0.6250  0.8250\n           Tuesday   0.8643 1.0000   0.7679    0.9036  0.9536\n           Wednesday 0.4036 0.7679   1.0000    0.9393  0.7821\n           Thursday  0.6250 0.9036   0.9393    1.0000  0.8964\n           Friday    0.8250 0.9536   0.7821    0.8964  1.0000                                                    (3)\n\n   ii)    The outcome of Pearson method is based on the values of the data whereas Kendall and\n          Spearman correlation matrix is based on the rank of the values in the dataset.\n          Diagonal in each matrix represents the correlation of temperature with respect to the individual\n          days, for example Monday with Monday.\n          Pearson shows strong positive correlation between temperatures on Monday, Tuesday,\n          Thursday and Friday.\n          Kendall and Spearman shows lesser positive correlation as compared to the Pearson’s result         (3)",
      "has_math": false,
      "is_r_task": true,
      "session": "2021-09",
      "subject": "CS1B",
      "source_qp": "raw/CS1B_2021-09_QP.pdf",
      "source_sol": "raw/CS1B_2021-09_SOL.pdf"
    },
    {
      "q_num": 3,
      "marks": 7,
      "topic": "distributions",
      "subtopics": [
        "data_analysis"
      ],
      "stem": "A child tosses n coins and the outcome of heads and tails are recorded in n samples as\n          X1, X2, …… Xn, where, Xi’s are independent Bernoulli variables with p = 0.5. The total\n          outcome of n variables is Y = X1 +……. + Xn",
      "parts": [
        {
          "label": "i",
          "marks": 2,
          "text": "Specify the distribution of Y                                                             (2)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 2,
          "text": "Simulate a sample of 10 values for Y                                                     (2)",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 2,
          "text": "Assess the value of Y for the sample created in (ii)                                    (2)",
          "topic": null
        },
        {
          "label": "iv",
          "marks": 1,
          "text": "what is the probability of Y = 4                                                         (1)",
          "topic": null
        }
      ],
      "solution": "i)    If X1, X2, …… Xn follows Bernoulli distribution with p = 0.5                                            (1)\n         Then Y follows binomial distribution ~ Bin(n,p) as sum of independent Bernoulli distribution\n         results in binomial distribution with n trials and p is the probability of success.                     (1)\n\n   ii)    Since ‘n’ is not specified in the question, sample simulations of Y can be generated assuming any\n          value of ‘n’. The answer below is calculated using n = 10, however marks are allotted to\n          simulations of Y, calculated using any value of ‘n’.\n\n          R Code\n          n = 10\n          p = 0.5\n          Y ~ rbinom(10, n, p)\n          Y\n\n          Values of y from R (since these are random values, answers for each individual may vary from\n          the numbers stated below, please evaluate accordingly)\n          5445265425                                                                                             (2)\n\n   iii)   The sample simulated in part (ii) are representative value of variable Y itself. Hence, no further\n          calculation required to assess the value of variable Y.                                                (2)\n\n   iv)    dbinom (4,n,p)\n          The solution is based on assumption that n = 10 (as assumed in part(ii)). Marks are allotted to\n          solutions calculated using any value for “n”.\n          0.2050781                                                                                         (1)",
      "has_math": false,
      "is_r_task": true,
      "session": "2021-09",
      "subject": "CS1B",
      "source_qp": "raw/CS1B_2021-09_QP.pdf",
      "source_sol": "raw/CS1B_2021-09_SOL.pdf"
    },
    {
      "q_num": 4,
      "marks": 6,
      "topic": "data_analysis",
      "subtopics": [],
      "stem": "For a sample data X = 0.5820, 0.04981, 0.1552, 0.1555, 0.9036, 0.8501, 0.9288, 0.4408,\n          0.9688, 0.6300",
      "parts": [
        {
          "label": "i",
          "marks": 2,
          "text": "Find the mean value of X                                                                  (2)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 1,
          "text": "Calculate the standard deviation of X                                                    (1)",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 3,
          "text": "Find the median of X                                                                    (3)",
          "topic": null
        }
      ],
      "solution": "i)    X = c(0.5820, 0.04981, 0.1552, 0.1555, 0.9036, 0.8501, 0.9288, 0.4408, 0.9688, 0.6300)\n         > mean(x)\n         0.566461                                                                                                (2)\n\n   ii)    Sd(x)\n                                                                                                  Page 3 of 15\n\f   IAI                                                                                            CS1B-0921\n          0.3515525                                                                                             (1)\n\n   iii)   median(x)\n          0.606                                                                                               (3)",
      "has_math": false,
      "is_r_task": true,
      "session": "2021-09",
      "subject": "CS1B",
      "source_qp": "raw/CS1B_2021-09_QP.pdf",
      "source_sol": "raw/CS1B_2021-09_SOL.pdf"
    },
    {
      "q_num": 5,
      "marks": 15,
      "topic": "bayes_credibility",
      "subtopics": [
        "data_analysis"
      ],
      "stem": "(Students can copy the R code as provided file ‘Reference_RCode’)\n\n      Claim payment data (Year 1 to Year 4) in INR Crores\n\n                 Year 1 Year 2 Year 3 Year 4\n       Insurer A   112    130    178    150\n       Insurer B    38     50     80     68\n       Insurer C    89    127    210    150\n       Insurer D    70     75     77     80",
      "parts": [
        {
          "label": "i",
          "marks": 2,
          "text": "Analyse the data using EBCT Model 1 and calculate the expected total claim\n         payment to be made by each insurer (A, B, C and D) in the Year 5 by replacing\n         question marks‘?’ with appropriate function in the R Code shared in file\n         “Reference_RCode”.                                                                           (2)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 3,
          "text": "You are now required to calculate the expected total claim payment to be made by\n          each insurer (A, B, C, D and E) in year 5, when additional data of insurer E is made\n          available to you (By using EBCT Model 1).\n\n      Additional claim payment data for Insurer E in INR Crores\n\n                 Year 1 Year 2 Year 3 Year 4\n       Insurer E    73     87    113    112\n\n      a) Provide the new R code so that ‘data’ includes insurer E and calculate the expected\n         claim payment to be paid by each insurer in year 5.                                          (2)\n\n      b) Comment on the values of Z and ‘s/v’ ratio in part (ii) compared to part (i).                (3)\n\n      c) Comment on the change in expected claim payment for insurers A, B C and D in\n         year 5 in part (ii) compared to part (i).                                                    (3)",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 3,
          "text": "Your manager is not happy with results based on EBCT 1 model. He has now\n           provided you additional information about volume measure for these 5 insurers over\n           Year 1 to Year 5.\n\n      Volume measure (Year 1 to Year 5)\n\n                 Year 1    Year 2    Year 3    Year 4    Year 5\n       Insurer A       165       186       198       200       210\n       Insurer B        51        63        78        83        91\n       Insurer C       119       159       219       188       192\n       Insurer D       108       112       122       133       144\n       Insurer E       103       116       126       151       181\n\n      Analyse the data using EBCT Model 2 to calculate the expected total claim payment to\n      be made by each insurer (A, B, C, D and E) in Year 5.\n\n      a) Complete the reference R Code (by replacing question marks‘?’ with appropriate\n         function for data and zi, where data is as used in part (ii) to arrive at the ‘claims5i’\n         which represents the estimated claim payment for each insurer in Year 5.                     (2)\n\n          b) Comment on change in expected claim payment for insurer A, B, C, D and E in year\n             5 in part (iii) compared to part (ii).                                              (3)",
          "topic": null
        }
      ],
      "solution": "i)    data=matrix(c(112,38,89,70,130,50,127,75,178,80,210,77,150,68,150,80),4,4)\n\n          n<-ncol(data)\n          m<-mean(rowMeans(data))\n          s<-mean(apply(data,1,var))\n          v<-var(rowMeans(data))-mean(apply(data,1,var))/n\n          Z<-n/(n+s/v)\n          Z*rowMeans(data)+(1-Z)*m\n\n          [1] 138.08805 64.47793 139.41039 79.02364\n\n          Expected total claim payment from\n           insurer A is INR 138.09 Crores\n           insurer B is INR 64.48 Crores\n           insurer C is INR 139.41 Crores\n           insurer D is INR 79.02 Crores\n          (1 Mark for correct Z and 1 mark for identifying the correct claim amount for each insurer)\n          As required for other parts\n              >n\n              [1] 4\n              >m\n              [1] 105.25\n              >s\n              [1] 933.8333\n              >v\n              [1] 1737.625\n              >Z\n              [1] 0.8815584                                                                                     (2)\n\n   ii)    a) R Code\n\n          data=matrix(c(112,38,89,70,73,130,50,127,75,87,178,80,210,77,113,150,68,150,80,112),5,4)\n          #no change in other code from part i\n          n<-ncol(data)\n          m<-mean(rowMeans(data))\n          s<-mean(apply(data,1,var))\n          v<-var(rowMeans(data))-mean(apply(data,1,var))/n\n          Z<-n/(n+s/v)\n          Z*rowMeans(data)+(1-Z)*m\n\n          [1] 137.11712 65.12725 138.41035 79.35279 97.24249\n          Expected total claim payment from\n              insurer A is INR 137.12 Crores (reduced compared to part i)\n              insurer B is INR 65.13 Crores (increased compared to part i)\n\n                                                                                                 Page 4 of 15\n\fIAI                                                                                                 CS1B-0921\n             insurer C is INR 138.41 Crores (reduced compared to part i)\n             insurer D is INR 79.35 Crores (increased compared to part i)\n             insurer E is INR 97.24 Crores (newly added insurer)\n              (1 Mark for correct data and 1 mark for identifying the correct claim amount for each\n              insurer)\n         required for other parts\n         >n\n         [1] 4\n         >m\n         [1] 103.45\n         >s\n         [1] 824.05\n         >v\n         [1] 1288.5\n         >Z\n         [1] 0.862154                                                                                             (2)\n\n       b) Comparison of calculated quantities\n\n        With the introduction of insurer E, credibility factor Z has reduced.\n       o E[s2(Theta)] (i.e. ‘s’ in the R code used above) has reduced marginally – this has an effect of\n         increasing Z.\n       o var[m(theta)] has reduced significantly (i.e. ‘v’ in the R code used above) - this has an effect\n         of reducing Z.\n        Overall the ratio - ‘s’ / ‘v’ is 0.54 in part i; ‘s’ / ‘v’ increases to 0.64 in part ii – as value of\n         denominator (in the calculation of Z = n/ (n + (s/v))) has increased in part ii compared to part\n         i – hence, credibility factor Z has reduced in part ii compared to part i.                               (3)\n\n       c) Change in expected claim payment for each insurer\n         Expected claim payment for insurer is based on credibility value and equal to\n         Z*individual average + (1-Z) * total average\n         (0.5 mark for comment on formula)\n\n             Total average ‘m’ has reduced in part ii compared to part i as average of insurer E is lower\n         than total average (m = 105.25) in part i\n         (0.5 mark for comment on ‘m’)\n\n             Expected claim payment for insurers A and C has reduced in part ii; as lower weight to\n         individual experience compared to part i (A and C have experienced higher claim payments in\n         the past and their average is higher than the total average of 5 insurers 103.45 - value of ‘m’\n         as calculated above) brings down the total expected value for these insurers towards 103.45\n         (1 mark for comment on A and C - where reduction observed)\n\n             Expected claim payment for insurers B and D has increased in part ii; as lower weight to\n         individual experience compared to part i (B and D’s individual average values are lower than\n         103.45) and higher weight to total average (103.45) pulls up the total expected value for these\n         insurers towards 103.45.\n         (1 mark for comment on B and D – where increase observed)\n\niii)   a) Use data as used in part ii\n\n                                                                                                   Page 5 of 15\n\fIAI                                                                                          CS1B-0921\n      data=matrix(c(112,38,89,70,73,130,50,127,75,87,178,80,210,77,113,150,68,150,80,112),5,4)\n      #all other code lines as provide in part iii reference R code\n      volume<-\n      matrix(c(165,51,119,108,103,186,63,159,112,116,198,78,219,122,126,200,83,188,133,151),5,4)\n      n<-ncol(data)\n      N<-nrow(data)\n      X<-data/volume #claim payment per unit of volume measure\n      Xibar<-rowSums(data)/rowSums(volume)\n      Pi<-rowSums(volume) #volume measure for each insurer\n      P<-sum(Pi)\n      Pstar<-sum(Pi*(1-Pi/P))/(N*n-1)\n      m<-sum(data)/P #average claim payment per unit volume measure across all insurers\n      s<-mean(rowSums(volume*(X-Xibar)^2)/(n-1))\n      v<-(sum(rowSums(volume*(X-m)^2))/(n*N-1)-s)/Pstar\n      zi<-Pi/(Pi+s/v) #credibility factor for each insurer\n      credibilityi<-zi*Xibar+(1-zi)*m #expected claim payment per unit of volume measure\n      volume5i<-matrix(c(210,91,192,144,181),5,1) #’volume5i’ is Year 5 volume measure\n      #okay to have 1,5 instead of 5,1 in volume5i (row , column interchange)\n      claims5i<-credibilityi*volume5i #claims5i is the expected claim payment in year 5\n\n      Additional output (marks not to be deducted in case if other parts are not provided by student)\n\n      >n\n      [1] 4\n      >N\n      [1] 5\n      >X\n             [,1] [,2] [,3] [,4]\n      [1,] 0.6787879 0.6989247 0.8989899 0.7500000\n      [2,] 0.7450980 0.7936508 1.0256410 0.8192771\n      [3,] 0.7478992 0.7987421 0.9589041 0.7978723\n      [4,] 0.6481481 0.6696429 0.6311475 0.6015038\n      [5,] 0.7087379 0.7500000 0.8968254 0.7417219\n      > Xibar\n      [1] 0.7610147 0.8581818 0.8408759 0.6357895 0.7762097\n      > Pi\n      [1] 749 275 685 475 496\n      >P\n      [1] 2680\n      > Pstar\n      [1] 110.0728\n      >m\n      [1] 0.7720149\n      >s\n      [1] 1.095221\n      >v\n      [1] 0.00469698\n      > zi\n      [1] 0.7625928 0.5411516 0.7460447 0.6707376 0.6802203\n      > credibility\n      [1] 0.7636262 0.8186443 0.8233883 0.6806434 0.7748683\n      > volume5\n                                                                                            Page 6 of 15\n\f   IAI                                                                                         CS1B-0921\n            [,1]\n         [1,] 210\n         [2,] 91\n         [3,] 192\n         [4,] 144\n         [5,] 181                                                                                            (2)\n\n         b)\n         > claims5\n               [,1]\n         [1,] 160.36151\n         [2,] 74.49663\n         [3,] 158.09055\n         [4,] 98.01265\n         [5,] 140.25116\n\n         Claims5 gives expected claim payment based on EBCT 2\n              INR 160.36 Crores for insurer A\n              INR 74.50 Crores for insurer B\n              INR 158.09 Crores for insurer C\n              INR 98.01 Crores for insurer D\n              INR 140.25 Crores for insurer E\n         (1 mark for identifying expected claim payment)\n\n         Expected claim payments for all insurers has increased compared to part ii as it makes use of\n         Year 5 volume which is higher for all insurers.\n         (1 mark for comment on increase and 1 mark for comment on Year 5 volume)                           (3)",
      "has_math": false,
      "is_r_task": true,
      "session": "2021-09",
      "subject": "CS1B",
      "source_qp": "raw/CS1B_2021-09_QP.pdf",
      "source_sol": "raw/CS1B_2021-09_SOL.pdf"
    },
    {
      "q_num": 6,
      "marks": 20,
      "topic": "inference",
      "subtopics": [],
      "stem": "You are investigating the level of premium charged by two companies for certain group.\n          Random samples of 10 policies from Company1 and Company2 are compared. Below R\n          code provides the premiums charged by Company1 and Company2 in current year for\n          10 sample policies.\n\n                 Company1<-c(1350,1790,1500,1150,2100,2350,1550,1800,1650,1450)\n                 Company2<-c(1500,1200,1300,1700,1800,2400,1450,1950,1850,2100)\n\n          Students are required to provide hypothesis, R code, Output and conclusion",
      "parts": [
        {
          "label": "i",
          "marks": 4,
          "text": "Assuming that the premiums are normally distributed, carry out a statistical test to\n              check equal variance assumption so that it is appropriate to apply a two-sample t\n              test to these data.\n\n              Use R code – var.test                                                                   (4)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 4,
          "text": "Test whether the level of premiums charged by Company1 and Company2 are\n               same.\n\n               Use R code – t.test (use var.equal =TRUE)                                              (4)",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 4,
          "text": "The average premium charged by Company2 in the previous year was INR 1500.\n                Test whether Company2 appears to have increased its premiums since the previous\n                year.\n\n              Use R code -t.test (#one sided)                                                         (4)",
          "topic": null
        },
        {
          "label": "iv",
          "marks": 4,
          "text": "It was found that out of sample of 200 policies of Company1 and 100 policies of\n               Company2 sold at the start of the year, 60 policies of Company1 and 50 policies of\n               Company2 resulted in claim. Carry out a hypothesis test for the difference in\n               proportions.\n\n              Use R code – prop.test (use correct=FALSE)                                              (4)",
          "topic": null
        },
        {
          "label": "v",
          "marks": 4,
          "text": "Company2 wants to study the claim frequency between group1 policies having\n              premium less than INR 1500 and remaining policyholders. There were 65 claims\n              out of 250 policies of group1 and 45 claims out of 110 other policies in a year.\n              Assuming number of claims having Poisson distribution, test at 2.5% level whether\n              the ratio of claim frequency between group1 and other policyholders is less than 1.\n\n              Use R code – poisson.test (#one sided)                                                  (4)",
          "topic": null
        }
      ],
      "solution": "i)    H0 – Company1 and Company2 having equal variance (i.e., ratio of variance =1)\n         H1 – There is a difference in variance between Company1 and Company2.\n\n         Company1<-c(1350,1790,1500,1150,2100,2350,1550,1800,1650,1450)\n         Company2<-c(1500,1200,1300,1700,1800,2400,1450,1950,1850,2100)\n\n         var.test(Company1,Company2)\n\n         F test to compare two variances\n\n         data: Company1 and Company2\n         F = 0.91388, num df = 9, denom df = 9, p-value = 0.8955\n         alternative hypothesis: true ratio of variances is not equal to 1\n         95 percent confidence interval:\n          0.2269944 3.6792680\n         sample estimates:\n         ratio of variances\n               0.9138781\n\n                                                                                              Page 7 of 15\n\f IAI                                                                                            CS1B-0921\n        Conclusion – As the confidence interval for the ratio between variance includes 1, we do not\n        have enough evidence to reject H0 at 5% significance level. Hence, we can conclude that ratio of\n        variance between Company1 and Company2 is 1 i.e., Company1 and Company2 having equal\n        variance at 5% significance level.                                                                    (4)\n\n ii)    H0 – No difference in mean premium charged by Company1 and Company2.\n        H1 – difference in mean premium charged by Company1 and Company2.\n\n        t.test(Company1, Company2, var.equal = TRUE, conf=0.95)\n\n           Two Sample t-test\n\n        data: Company1 and Company2\n        t = -0.3433, df = 18, p-value = 0.7353\n        alternative hypothesis: true difference in means is not equal to 0\n        95 percent confidence interval:\n         -398.703 286.703\n        sample estimates:\n        mean of x mean of y\n           1669 1725\n\n        Conclusion – As the confidence interval for the difference in mean premium includes 0, we do\n        not have enough evidence to reject H0 at 5% significance level. Hence, we can conclude that\n        there is no difference between mean premium charged by Company1 and Company2 at 5%\n        significance level.                                                                                   (4)\n\n iii)   H0 – mean premium charged by Company2 is INR 1500.\n        H1 – mean premium charged by Company2 is greater than 1500.\n\n        t.test(Company2, mu=1500, alternative= c(\"greater\"), conf=0.95)\n\n            One Sample t-test\n\n        data: Company2\n        t = 1.9082, df = 9, p-value = 0.04436\n        alternative hypothesis: true mean is greater than 1500\n        95 percent confidence interval:\n         1508.858 Inf\n        sample estimates:\n        mean of x\n           1725\n\n        Conclusion – As the confidence interval for the mean premium does not include 1500, at 5%\n        significance level we have enough evidence to reject null hypothesis.\n        (p-value less than 5% also confirms that we have enough evidence to reject null hypothesis at\n        5% significance level).\n        Hence, we can conclude that the mean premium charged by Company2 is greater than INR 1500\n        at 5% significance level.                                                                             (4)\n\niv)     H0 – There is no difference in proportion of (policies that result in) claim between Company1 and\n        Company2.\n\n                                                                                               Page 8 of 15\n\f   IAI                                                                                             CS1B-0921\n         H1 – There is difference in proportion of (policies that result in) claim between Company1 and\n         Company2.\n\n         prop.test(c(60,50),c(200,100),correct=FALSE)\n\n             2-sample test for equality of proportions without continuity correction\n\n         data: c(60, 50) out of c(200, 100)\n         X-squared = 11.483, df = 1, p-value = 0.0007023\n         alternative hypothesis: two.sided\n         95 percent confidence interval:\n          -0.31677833 -0.08322167\n         sample estimates:\n         prop 1 prop 2\n           0.3 0.5\n\n         Conclusion – As the confidence interval does not include 0 (or as p-value is less than 5%), we\n         have sufficient evidence to reject null hypothesis at 5% significance level. Hence, we can\n         conclude that there is difference in proportion of (policies that result in) claims between\n         Company1 and Company2 at 5% significance level.                                                         (4)\n\n  v)     H0 – There is no difference in ratio of claim frequency between group1 and other groups in\n         Company2.\n         H1 – Claim frequency of group1 is less than other groups in Company2.\n         poisson.test(c(65,45),c(250,110), conf=.975, alternative=\"less\")\n\n             Comparison of Poisson rates\n\n         data: c(65, 45) time base: c(250, 110)\n         count1 = 65, expected count1 = 76.389, p-value = 0.01358\n         alternative hypothesis: true rate ratio is less than 1\n         97.5 percent confidence interval:\n          0.0000000 0.9511753\n         sample estimates:\n         rate ratio\n          0.6355556\n\n         Conclusion – As the confidence interval for the ratio does not include 1, at 2.5% significance level\n         we have enough evidence to reject null hypothesis.\n         (p-value less than 2.5% also confirms that we have enough evidence to reject null hypothesis at\n         2.5% significance level).\n         Hence, we can conclude that the claim frequency of group1 (having premium less than INR 1500)\n         is less than other groups in Company2.                                                                 (4)",
      "has_math": false,
      "is_r_task": true,
      "session": "2021-09",
      "subject": "CS1B",
      "source_qp": "raw/CS1B_2021-09_QP.pdf",
      "source_sol": "raw/CS1B_2021-09_SOL.pdf"
    },
    {
      "q_num": 7,
      "marks": 35,
      "topic": "regression_glm",
      "subtopics": [],
      "stem": "You are working as an Actuarial analyst in a medium size health insurance company in\n          India. Your manager has asked you to analyse the claim amounts paid data of the past\n          six months.\n\n          You have received the data set “HealthClaims.csv” from the claims department of your\n          company with the following explanations of the data fields\n\n      GEOGRAPHY: The geographical region of residence of the Insured\n      PROFESSION: Profession of the Insured\n      GENDER: Gender of the Insured\n      AGE: Age of the Insured\n      CLAIM_AMOUNT: Amount of health claim paid by the Insurer\n\n      Refer to the data set “HealthClaims.csv”.",
      "parts": [
        {
          "label": "i",
          "marks": 10,
          "text": "Fit a linear regression model to the data with “CLAIM_AMOUNT” as the response\n          and other variables as explanatory variables (consider “Age” as numerical variable\n          and others as categorical variables).\n\n           Provide your interpretation of the model by explaining R-Squared, Adjusted R-\n           Squared & p-value of the model. Identify the significant variables in the prediction\n           of claim “CLAIM_AMOUNT”.                                                                 (10)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 3,
          "text": "Determine 95% confidence intervals for the parameters of the regression model.            (3)",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 4,
          "text": "Plot “QQ plot of the residuals” and comment on applicability of linear regression\n            model.                                                                                   (4)",
          "topic": null
        },
        {
          "label": "iv",
          "marks": 10,
          "text": "Your manager has also suggested you to include the interaction effects between\n           Geography and Profession, Profession and Gender, Gender and Geography as\n           additional explanatory variables to the set of independent variables taken in (i) for\n           the purpose of improvement of the model. Comment on the suitability of inclusion\n           of interaction effects for the purpose of improvement of the model.                      (10)",
          "topic": null
        },
        {
          "label": "v",
          "marks": 8,
          "text": "One of your friend working with Actuarial consulting firm told you that natural\n          logarithm{loge(CLAIM_AMOUNT)} is better fit to normal distribution than\n          “CLAIM_AMOUNT”. You wanted to validate the same by fitting a linear\n          regression model to the data with loge(CLAIM_AMOUNT) as the response and\n          other variables as explanatory variables. Identify and comment on the key\n          differences between the models in (i) and (iv).                                            (8)",
          "topic": null
        }
      ],
      "solution": "i)    # Fitting Linear Regression Model\n         R Code:\n         Data_claim = read.csv('HealthClaims.csv')\n         model = lm(CLAIM_AMOUNT~.,data = Data_claim)\n         summary(model)\n\n                                                                                                  Page 9 of 15\n\fIAI                                                                     CS1B-0921\n      R Output:\n\n      Call:\n      lm(formula = CLAIM_AMOUNT ~ ., data = Data_claim)\n\n      Residuals:\n        Min 1Q Median 3Q Max\n      -155636 -22909 1558 19047 546017\n\n      Coefficients:\n                    Estimate Std. Error t value Pr(>|t|)\n      (Intercept)        -32487.4 12119.7 -2.681 0.007435 **\n      GEOGRAPHYREGION 02 -34130.3 9237.3 -3.695 0.000228 ***\n      GEOGRAPHYREGION 03 -18468.4 10664.6 -1.732 0.083534 .\n      GEOGRAPHYREGION 04 -17925.8 9866.1 -1.817 0.069441 .\n      GEOGRAPHYREGION 05 -18160.8 9895.2 -1.835 0.066668 .\n      GEOGRAPHYREGION 06           1325.1 10558.3 0.125 0.900146\n      GEOGRAPHYREGION 07 -30443.9 11716.4 -2.598 0.009463 **\n      GEOGRAPHYREGION 08 -57724.8 8801.0 -6.559 7.57e-11 ***\n      GEOGRAPHYREGION 09 -10506.6 11595.9 -0.906 0.365056\n      GEOGRAPHYREGION 10 30219.2 12332.9 2.450 0.014394 *\n      GEOGRAPHYREGION 11 -14969.7 11932.6 -1.255 0.209858\n      PROFESSIONBusiness 158911.2 6700.4 23.717 < 2e-16 ***\n      PROFESSIONSelfemployed 15605.8 4094.8 3.811 0.000144 ***\n      PROFESSIONService 163694.3 7081.3 23.116 < 2e-16 ***\n      GENDERM              -13556.3 3024.6 -4.482 7.99e-06 ***\n      AGE               2827.4 197.8 14.292 < 2e-16 ***\n      ---\n      Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1\n\n      Residual standard error: 55350 on 1420 degrees of freedom\n      Multiple R-squared: 0.7275, Adjusted R-squared: 0.7246\n      F-statistic: 252.8 on 15 and 1420 DF, p-value: < 2.2e-16\n\n      R code:\n      anova(model)\n\n      R Output:\n\n      Analysis of Variance Table\n\n      Response: CLAIM_AMOUNT\n               Df Sum Sq Mean Sq F value Pr(>F)\n      GEOGRAPHY 10 2.2059e+12 2.2059e+11 71.997 < 2.2e-16 ***\n      PROFESSION 3 8.6687e+12 2.8896e+12 943.124 < 2.2e-16 ***\n      GENDER        1 1.1570e+11 1.1570e+11 37.763 1.036e-09 ***\n      AGE        1 6.2585e+11 6.2585e+11 204.271 < 2.2e-16 ***\n      Residuals 1420 4.3506e+12 3.0638e+09\n      ---\n      Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1\n\n                                                                      Page 10 of 15\n\f  IAI                                                                                               CS1B-0921\n        R-Squared : 72.75% of the variation of the claim amount is explained by geographical region,\n        profession, gender and age\n        Adjusted R- Squared: 72.46% is used to compare with the other models . Adjusted R-Squared is\n        used to compare the goodness of fit for regression models that contain different number of\n        independent variables. The model which maximises the ‘adjusted R2 statistics can be regarded\n        in some sense as the best model.\n\n        p-value of the model is < 2.2e-16 which is less than 0.05, hence the null hypothesis that “there\n        is no significant relationship between independent variable and dependent variable” is rejected\n        at 5% level of significance.\n\n        P-value of the coefficients: Some of the dependent variables appeared to be insignificant\n        although the overall model is significant. As Geographical region 1, Profession Agriculture and\n        Gender female are insignificant hence clubbed in the intercept itself. It is also observed that\n        Geographical region 10 has a positive coefficient and Geographical region 2 has negative\n        coefficient indicating that claim paid for Geographical region 2 is significantly lower than\n        Geographical region 10. Similarly claim paid for profession business is much higher compared\n        to profession self-employed.\n\n        From ANOVA table, it can be concluded that all the variables are significant in terms of claim\n        paid as the p values of all the parameters are less than 0.05 so they are all significantly different\n        from zero                                                                                               (10)\n\nii)     R Code:\n        confint(model,level=0.95)\n\n        R output:\n                    2.5 % 97.5 %\n        (Intercept)     -56261.888 -8712.970\n        GEOGRAPHYREGION 02 -52250.448 -16010.129\n        GEOGRAPHYREGION 03 -39388.438 2451.580\n        GEOGRAPHYREGION 04 -37279.438 1427.815\n        GEOGRAPHYREGION 05 -37571.441 1249.937\n        GEOGRAPHYREGION 06 -19386.522 22036.649\n        GEOGRAPHYREGION 07 -53427.302 -7460.556\n        GEOGRAPHYREGION 08 -74989.181 -40460.356\n        GEOGRAPHYREGION 09 -33253.405 12240.270\n        GEOGRAPHYREGION 10       6026.559 54411.771\n        GEOGRAPHYREGION 11 -38377.145 8437.686\n        PROFESSIONBusiness 145767.428 172054.915\n        PROFESSIONSelfemployed 7573.347 23638.179\n        PROFESSIONService 149803.318 177585.265\n        GENDERM           -19489.519 -7623.148\n        AGE           2439.328 3215.450                                                                          (3)\n\niii)    R code:\n        qqnorm(model$residuals)\n\n        Output:\n\n                                                                                                 Page 11 of 15\n\f IAI                                                                                             CS1B-0921\n\n       If the residuals are normally distributed, it is expected the Q-Q plot to be along the diagonal.\n       Here it is not , indicating residuals are not forming normal distribution. Hence linear regression\n       may not be a good fit to the data                                                                       (4)\n\niv)    R Code:\n       model2 = lm(CLAIM_AMOUNT~.+GEOGRAPHY:PROFESSION + PROFESSION:GENDER + GENDER:\n       GEOGRAPHY,data = Data_claim)\n       summary(model2)\n       Output:\n       Call:\n       lm(formula = CLAIM_AMOUNT ~ . + GEOGRAPHY:PROFESSION + PROFESSION:GENDER +\n         GENDER:GEOGRAPHY, data = Data_claim)\n\n       Residuals:\n         Min 1Q Median 3Q Max\n       -194405 -12955 -1034 10489 302012\n\n       Coefficients: (3 not defined because of singularities)\n                                Estimate Std. Error t value Pr(>|t|)\n       (Intercept)                   -64594.5 14131.6 -4.571 5.29e-06 ***\n       GEOGRAPHYREGION 02                     -31544.8 13652.9 -2.310 0.02101 *\n       GEOGRAPHYREGION 03                     -28731.6 16539.4 -1.737 0.08258 .\n       GEOGRAPHYREGION 04                     -20891.3 14966.3 -1.396 0.16297\n       GEOGRAPHYREGION 05                     -32366.2 14299.3 -2.263 0.02376 *\n       GEOGRAPHYREGION 06                     -18604.1 15100.7 -1.232 0.21816\n       GEOGRAPHYREGION 07                     -29002.4 15235.1 -1.904 0.05716 .\n       GEOGRAPHYREGION 08                     -33692.9 12994.4 -2.593 0.00962 **\n       GEOGRAPHYREGION 09                     -24352.8 18909.3 -1.288 0.19801\n       GEOGRAPHYREGION 10                      -8309.3 17192.1 -0.483 0.62894\n       GEOGRAPHYREGION 11                     -18947.6 16282.5 -1.164 0.24476\n       PROFESSIONBusiness                  308385.3 15720.6 19.617 < 2e-16 ***\n       PROFESSIONSelfemployed                  30126.2 20435.0 1.474 0.14064\n\n                                                                                               Page 12 of 15\n\fIAI                                                                                      CS1B-0921\n      PROFESSIONService                   249218.9 42323.2 5.888 4.89e-09 ***\n      GENDERM                         -43175.6 14519.3 -2.974 0.00299 **\n      AGE                          3281.3 150.7 21.770 < 2e-16 ***\n      GEOGRAPHYREGION 02:PROFESSIONBusiness -103364.5 16534.2 -6.252 5.41e-10 ***\n      GEOGRAPHYREGION 03:PROFESSIONBusiness -61296.8 23584.0 -2.599 0.00945 **\n      GEOGRAPHYREGION 04:PROFESSIONBusiness -63068.3 19388.6 -3.253 0.00117 **\n      GEOGRAPHYREGION 05:PROFESSIONBusiness -29786.4 18110.5 -1.645 0.10026\n      GEOGRAPHYREGION 06:PROFESSIONBusiness -172656.7 43119.4 -4.004 6.56e-05 ***\n      GEOGRAPHYREGION 07:PROFESSIONBusiness                     NA   NA NA       NA\n      GEOGRAPHYREGION 08:PROFESSIONBusiness -223496.4 15218.8 -14.686 < 2e-16 ***\n      GEOGRAPHYREGION 09:PROFESSIONBusiness -47277.9 20406.5 -2.317 0.02066 *\n      GEOGRAPHYREGION 10:PROFESSIONBusiness 135659.4 24962.1 5.435 6.48e-08 ***\n      GEOGRAPHYREGION 11:PROFESSIONBusiness                     NA   NA NA       NA\n      GEOGRAPHYREGION 02:PROFESSIONSelfemployed -14904.8 20843.9 -0.715 0.47469\n      GEOGRAPHYREGION 03:PROFESSIONSelfemployed -5376.4 22863.8 -0.235 0.81413\n      GEOGRAPHYREGION 04:PROFESSIONSelfemployed -1332.0 21623.4 -0.062 0.95089\n      GEOGRAPHYREGION 05:PROFESSIONSelfemployed -715.2 21573.5 -0.033 0.97356\n      GEOGRAPHYREGION 06:PROFESSIONSelfemployed 2571.7 23542.6 0.109 0.91303\n      GEOGRAPHYREGION 07:PROFESSIONSelfemployed 2053.5 23634.2 0.087 0.93077\n      GEOGRAPHYREGION 08:PROFESSIONSelfemployed -21048.6 20323.1 -1.036 0.30053\n      GEOGRAPHYREGION 09:PROFESSIONSelfemployed 3725.3 25207.1 0.148 0.88253\n      GEOGRAPHYREGION 10:PROFESSIONSelfemployed -1907.5 25341.3 -0.075 0.94001\n      GEOGRAPHYREGION 11:PROFESSIONSelfemployed -10631.8 26202.1 -0.406 0.68498\n      GEOGRAPHYREGION 02:PROFESSIONService                 -91701.8 44213.1 -2.074 0.03826 *\n      GEOGRAPHYREGION 03:PROFESSIONService                 -32757.7 46667.3 -0.702 0.48283\n      GEOGRAPHYREGION 04:PROFESSIONService                 108004.4 47768.4 2.261 0.02391 *\n      GEOGRAPHYREGION 05:PROFESSIONService                 -39108.5 48885.0 -0.800 0.42384\n      GEOGRAPHYREGION 06:PROFESSIONService                  93708.9 44227.6 2.119 0.03429 *\n      GEOGRAPHYREGION 07:PROFESSIONService                     NA   NA NA       NA\n      GEOGRAPHYREGION 08:PROFESSIONService -174071.2 42916.8 -4.056 5.27e-05 ***\n      GEOGRAPHYREGION 09:PROFESSIONService                 -24666.4 46726.3 -0.528 0.59766\n      GEOGRAPHYREGION 10:PROFESSIONService                 108119.5 47621.0 2.270 0.02334 *\n      GEOGRAPHYREGION 11:PROFESSIONService                  30088.2 46092.1 0.653 0.51400\n      PROFESSIONBusiness:GENDERM                   -69347.9 8414.3 -8.242 3.91e-16 ***\n      PROFESSIONSelfemployed:GENDERM                  -11465.1 5210.5 -2.200 0.02794 *\n      PROFESSIONService:GENDERM                   -35812.2 8955.2 -3.999 6.70e-05 ***\n      GEOGRAPHYREGION 02:GENDERM                      48893.3 15503.6 3.154 0.00165 **\n      GEOGRAPHYREGION 03:GENDERM                      47556.3 17387.1 2.735 0.00632 **\n      GEOGRAPHYREGION 04:GENDERM                      29240.4 16539.5 1.768 0.07730 .\n      GEOGRAPHYREGION 05:GENDERM                      47373.9 16304.1 2.906 0.00372 **\n      GEOGRAPHYREGION 06:GENDERM                      31561.7 17303.9 1.824 0.06837 .\n      GEOGRAPHYREGION 07:GENDERM                      41765.9 18709.7 2.232 0.02575 *\n      GEOGRAPHYREGION 08:GENDERM                      47600.7 14846.0 3.206 0.00138 **\n      GEOGRAPHYREGION 09:GENDERM                      36413.1 20344.7 1.790 0.07370 .\n      GEOGRAPHYREGION 10:GENDERM                       9577.3 19749.4 0.485 0.62779\n      GEOGRAPHYREGION 11:GENDERM                      31022.5 19076.4 1.626 0.10413\n      ---\n      Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1\n\n      Residual standard error: 40270 on 1380 degrees of freedom\n      Multiple R-squared: 0.8598, Adjusted R-squared: 0.8542\n      F-statistic: 153.9 on 55 and 1380 DF, p-value: < 2.2e-16\n                                                                                       Page 13 of 15\n\f IAI                                                                                      CS1B-0921\n       anova(model3)\n\n       Analysis of Variance Table\n\n       Response: CLAIM_AMOUNT\n                     Df Sum Sq Mean Sq F value Pr(>F)\n       GEOGRAPHY             10 2.2059e+12 2.2059e+11 136.0055 < 2.2e-16 ***\n       PROFESSION            3 8.6687e+12 2.8896e+12 1781.5995 < 2.2e-16 ***\n       GENDER              1 1.1570e+11 1.1570e+11 71.3357 < 2.2e-16 ***\n       AGE              1 6.2585e+11 6.2585e+11 385.8770 < 2.2e-16 ***\n       GEOGRAPHY:PROFESSION 27 1.9432e+12 7.1969e+10 44.3731 < 2.2e-16 ***\n       PROFESSION:GENDER          3 1.2973e+11 4.3244e+10 26.6624 < 2.2e-16 ***\n       GEOGRAPHY:GENDER           10 3.9538e+10 3.9538e+09 2.4378 0.007001 **\n       Residuals        1380 2.2382e+12 1.6219e+09\n       ---\n       Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1\n\n       Comments:\n           R-Squared and Adjusted R-Squared improved and hence model is a better fit compared\n            to the initial model\n           Interaction effect between geographical region and profession, profession and gender,\n            gender and geographical region along with their main effects emerged out to be\n            significant as demonstrated by ANOVA table\nv)     R code:\n       model3 = lm(log(CLAIM_AMOUNT)~.,data = Data_claim)\n       summary(model3)\n\n       Output:\n       Call:\n       lm(formula = log(CLAIM_AMOUNT) ~ ., data = Data_claim)\n\n       Residuals:\n          Min     1Q Median     3Q Max\n       -0.86313 -0.18493 -0.00393 0.17478 1.02901\n\n       Coefficients:\n                   Estimate Std. Error t value Pr(>|t|)\n       (Intercept)     8.4529628 0.0580014 145.737 < 2e-16 ***\n       GEOGRAPHYREGION 02 -0.1408025 0.0442069 -3.185 0.00148 **\n       GEOGRAPHYREGION 03 -0.0774399 0.0510375 -1.517 0.12941\n       GEOGRAPHYREGION 04 -0.0889472 0.0472161 -1.884 0.05979 .\n       GEOGRAPHYREGION 05 -0.0829154 0.0473553 -1.751 0.08018 .\n       GEOGRAPHYREGION 06 0.0231982 0.0505290 0.459 0.64623\n       GEOGRAPHYREGION 07 -0.0654231 0.0560714 -1.167 0.24349\n       GEOGRAPHYREGION 08 -0.3693971 0.0421191 -8.770 < 2e-16 ***\n       GEOGRAPHYREGION 09 -0.0568913 0.0554943 -1.025 0.30546\n       GEOGRAPHYREGION 10 0.0540411 0.0590215 0.916 0.36002\n       GEOGRAPHYREGION 11 -0.0345033 0.0571059 -0.604 0.54581\n       PROFESSIONBusiness 0.8217253 0.0320661 25.626 < 2e-16 ***\n       PROFESSIONSelfemployed 0.1956622 0.0195963 9.985 < 2e-16 ***\n\n                                                                                        Page 14 of 15\n\fIAI                                                                                        CS1B-0921\n      PROFESSIONService        0.8736364 0.0338891 25.779 < 2e-16 ***\n      GENDERM             -0.0889058 0.0144749 -6.142 1.06e-09 ***\n      AGE              0.0536827 0.0009467 56.703 < 2e-16 ***\n      ---\n      Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1\n\n      Residual standard error: 0.2649 on 1420 degrees of freedom\n      Multiple R-squared: 0.9227, Adjusted R-squared: 0.9219\n      F-statistic: 1130 on 15 and 1420 DF, p-value: < 2.2e-16\n      R code:\n      anova(model3)\n      Analysis of Variance Table\n      Response: log(CLAIM_AMOUNT)\n               Df Sum Sq Mean Sq F value Pr(>F)\n      GEOGRAPHY 10 246.08 24.608 350.69 < 2.2e-16 ***\n      PROFESSION 3 706.84 235.613 3357.70 < 2.2e-16 ***\n      GENDER         1 11.25 11.255 160.39 < 2.2e-16 ***\n      AGE         1 225.62 225.616 3215.24 < 2.2e-16 ***\n      Residuals 1420 99.64 0.070\n      ---\n      Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1\n\n      Key differences:\n          R-Squared and Adjusted R-Squared increased to above 90% and hence model is a better\n              fit compared to the earlier models\n          The significance level of few factor coefficients changed when compared with the model\n              in (i) however ANOVA table shows there is no change in number of significant variables\n\n                                         *********************\n\n                                                                                         Page 15 of 15",
      "has_math": false,
      "is_r_task": true,
      "session": "2021-09",
      "subject": "CS1B",
      "source_qp": "raw/CS1B_2021-09_QP.pdf",
      "source_sol": "raw/CS1B_2021-09_SOL.pdf"
    },
    {
      "q_num": 1,
      "marks": 30,
      "topic": "inference",
      "subtopics": [
        "regression_glm"
      ],
      "stem": "An Actuarial student fits following simple regression model to the data\n          yi = alpha + beta * xi + ei ; i =1 to 12\n          where ei are independent normal random variables with mean 0 and variance sigma2\n\n          Use following 12 data points for x and y\n          where y is response variable while x is explanatory variable\n\n          x = c(5,10,15,20,25,30,35,40,45,50,55,60)\n          y = c(15,12,25,23,35,36,33,38,43,45,50,53)\n\n          Note: Do not use standard ‘model fitting related’ R codes - lm, glm, fitted, residuals,\n          predict, anova - to answer parts of this question.",
      "parts": [
        {
          "label": "i",
          "marks": 7,
          "text": "Calculate Sxx, Sxy and Syy.                                                                    (7)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 3,
          "text": "Calculate alpha, beta and sigma2 using results in part (i).                                   (3)",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 1,
          "text": "Calculate fitted values of y using results in part (ii).                                     (1)",
          "topic": null
        },
        {
          "label": "iv",
          "marks": 3,
          "text": "Calculate residuals using results of earlier parts and Calculate mean and variance of\n              residuals and comment on the result.                                                          (3)",
          "topic": null
        },
        {
          "label": "v",
          "marks": 7,
          "text": "Calculate 95% confidence interval for beta and comment if we can conclude that beta\n             is not zero stating the Null and alternate hypothesis. Calculate 95% confidence interval\n             for sigma2.\n             Note: Candidates are expected to find tabulated value using R.                                 (7)",
          "topic": null
        },
        {
          "label": "vi",
          "marks": 5,
          "text": "State SSTOT and Calculate SSREG and SSRES. Hence calculate the proportion of\n              variability explained by the model using SSREG and SSRES and comment on the result.\n              Calculate adjusted R2 and compare it with R^2 to explain result.                              (5)",
          "topic": null
        },
        {
          "label": "vii",
          "marks": 4,
          "text": "Calculate mean predicted response when x=52 and 95% confidence interval for the\n               same. Note: Candidates are expected to find tabulated value using R.                         (4)",
          "topic": null
        }
      ],
      "solution": "i)\nx= c(5,10,15,20,25,30,35,40,45,50,55,60)\ny=c(15,12,25,23,35,36,33,38,43,45,50,53)\n\n> meanx = mean(x)\n> meany=mean(y)\n> meanx\n[1] 32.5\n> meany\n[1] 34\n\n> x_sq=x*x\n> x_sq\n[1] 25 100 225 400 625 900 1225 1600 2025 2500 3025 3600\n> y_sq=y*y\n> xy=x*y\n> xy\n[1] 75 120 375 460 875 1080 1155 1520 1935 2250 2750 3180\n\n> sumx_sq=sum(x_sq)\n> sumy_sq=sum(y_sq)\n> sumxy=sum(xy)\n\n> sumx_sq\n[1] 16250\n> sumy_sq\n[1] 15760\n> sumxy\n[1] 15775\n\n> Sxx=Sumx_sq-12*meanx^2\n> Sxx\n[1] 3575\n> Sxy=sumxy-12*meanx*meany\n> Sxy\n[1] 2515\n> Syy=sumy_sq-12*meany^2\n> Syy\n[1] 1888\n\nii)\n\n> beta=Sxy/Sxx\n> beta\n[1] 0.7034965\n\n> alpha=meany-beta*meanx\n> alpha\n[1] 11.13636\n\n                                                            Page 2 of 13\n\fIAI                                                                                         CS1B-0322\n\n> sigmasq=(1/(12-2))*(Syy-Sxy^2/Sxx)\n> sigmasq\n[1] 11.87063\n\niii)\n\n> expectedy=alpha+beta*x\n> expectedy\n[1] 14.65385 18.17133 21.68881 25.20629 28.72378 32.24126 35.75874 39.27622 42.79371\n46.31119 49.82867 53.34615\n\niv)\n> e=y-alpha-beta*x\n>e\n [1] 0.3461538 -6.1713287 3.3111888 -2.2062937 6.2762238 3.7587413\n [7] -2.7587413 -1.2762238 0.2062937 -1.3111888 0.1713287 -0.3461538\n\n> meane=mean(e)\n> meane\n[1] -1.702344e-15\n\n> var(e)\n[1] 10.79148\n\nMean value of residuals is close to zero as expected as e~N(0,sigma^2)\n\n(Otherwise, “e” could be calculated as e = y-expectedy)\n\nVar of e is slightly lower than sigma square as calculated in part ii – as denominator is not adjusted\nWhen denominator of 10 gets used instead of 11 we see that var of residuals = sigma^2\n> var(e)*11/10\n[1] 11.87063\n\nv) 95% confidence interval for beta\n\nHo: Beta is zero (i.e. no linear relationship between x and y)\nH1: Beta is not equal to zero\n\n(Beta_cap-0)/ sqrt(sigma^2_cap/ Sxx) ~ t10\n\nWe use t distribution with n-2 i.e. 10 degrees of freedom\n\n> qt(p=0.025, lower.tail = T, df=10)\n[1] -2.228139\nBeing symmetric distribution, 97.5% point would be 2.228139\n> sqrt(sigmasq/Sxx)\n[1] 0.0576234\n\nHence, endpoints of CI would be\n> end1=beta+sqrt(sigmasq/Sxx)*qt(p=0.025, lower.tail = T, df=10)\n> end1\n[1] 0.5751036\n\n                                                                                           Page 3 of 13\n\fIAI                                                                                        CS1B-0322\n\n> end2=beta-sqrt(sigmasq/Sxx)*qt(p=0.025, lower.tail = T, df=10)\n> end2\n[1] 0.8318894\n\nHence 95% Confidence interval for beta is (0.5751, 0.8319)\n\nAs confidence interval for beta does not include zero, we can reject null hypothesis (viz. beta=0) and\nHence, can conclude that beta is not equal to zero at 5% level.\n\n95% CI for sigma^2\n\n(n-2)sigmacap^2/ sigma^2 ~ Chi sq distribution with 10 degrees of freedom\n\nTabulated values of Chi square having 10 df can be obtained as\n> chitenend1=qchisq(df=10, p=0.025)\n> chitenend2=qchisq(df=10,p=0.975)\n> chitenend1\n[1] 3.246973\n> chitenend2\n[1] 20.48318\n\nEnd points of CI would be\n> sigmasqend1=(12-2)*sigmasq/chitenend1\n> sigmasqend2=(12-2)*sigmasq/chitenend2\n> sigmasqend1\n[1] 36.55907\n> sigmasqend2\n[1] 5.795307\n\nHence 95% Confidence interval for sigma^2 is (5.795,36.559)\n\nvi)\nSSTOT = Syy = 1888 (as calculated in part i)\n\nSSREG = Sxy^2/Sxx\n> ss_reg=Sxy^2/Sxx\n> ss_reg\n[1] 1769.294\n\nSSRES = SSTOT - SSREG\n\n> ss_res=Syy-ss_reg\n> ss_res\n[1] 118.7063\n\nR^2 denotes the % of variability explained by the model\nR^2 = SSREG / (SSREG + SSRES)\n\n> Rsq = ss_reg/(ss_reg+ss_res)\n> Rsq\n[1] 0.9371259\n\nModel is a good fit as 93.7% of the variability is explained by the model.\n\n                                                                                          Page 4 of 13\n\fIAI                                                                                        CS1B-0322\n\n> adj_Rsq = 1-((12-1)/(12-1-1))*(1-Rsq)\n> adj_Rsq\n[1] 0.9308385\n\nAdjusted R^2 (93.08%) is lower than R^2 (93.71%) as adjusted R square penalises for extra predictors\nand\nhence is better suited to assess the adequacy of the model (or for comparison between models)\ncompared to just using R^2 for model comparison as\nR^2 cannot decrease on addition of more explanatory variables which can be undesirable (as it may\npromote too many explanatory variables though not adding significant improvement in the predicted\nvalue)\n\nvii)\nusing results from earlier parts mean predicted response is calculated (using regression line\nequation)\n> Emean52=alpha+beta*52\n> Emean52\n[1] 47.71818\n\nExpected value of mean predicted response is 47.718 when x=52\n\nvarofmean52=((1/12)+(52-meanx)^2/Sxx)*sigmasq\n> varofmean52\n[1] 2.251822\n\n> mean52end1=Emean52+qt(p=0.025, lower.tail = T, df=10)*sqrt(varofmean52)\n> mean52end2=Emean52-qt(0.025,lower.tail = T, df=10)*sqrt(varofmean52)\n> mean52end1\n[1] 44.37462\n> mean52end2\n[1] 51.06174\n\nHence 95% confidence interval for the mean predicted response is (44.3746,51.0617)",
      "has_math": false,
      "is_r_task": true,
      "session": "2022-03",
      "subject": "CS1B",
      "source_qp": "raw/CS1B_2022-03_QP.pdf",
      "source_sol": "raw/CS1B_2022-03_SOL.pdf"
    },
    {
      "q_num": 2,
      "marks": 35,
      "topic": "inference",
      "subtopics": [],
      "stem": "Policy and claims information (PolicyData.csv) of 650 policies is provided to you. The\n          data contains following fields:\n\n          Policy : Policy Number\n          Claim : Number of claims corresponding to each policy\n          Cust_Exp: Policyholder’s experience (VS = Very satisfied, SA= Satisfied, DS=\n          Disappointed, VD= Very Disappointed) at the end of the policy tenure.\n          Amount: Claim amount per policy. Note that amount is set 0 if there is no claim.",
      "parts": [
        {
          "label": "i",
          "marks": 2,
          "text": "Create a frequency table of claim and share how many policies don’t have any claim.            (2)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 2,
          "text": "Plot a histogram of claim count and suggest 2 distributions that can be a good fit.           (2)",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 3,
          "text": "Given that claim count follows Poisson distribution with following two possible values\n               for Poisson parameter:\n\n              a) 0.35 and\n              b) 0.30\n\n          Compute confidence interval at 95% confidence level to assess which value is more\n          suitable for the given data.                                                        (3)",
          "topic": null
        },
        {
          "label": "iv",
          "marks": 4,
          "text": "Compute mean, variance and median of log of claim amount. (Name it: log amount).\n              Make sure to exclude policies with no claim.                                               (4)",
          "topic": null
        },
        {
          "label": "v",
          "marks": 3,
          "text": "Obtain histogram and Normal QQ Plot of log amount. Add a line to the QQ plot for\n             normal distribution                                                                         (3)",
          "topic": null
        },
        {
          "label": "vi",
          "marks": 3,
          "text": "Indicate which distribution the claim amount might be following using evidence from\n              (iv) and (v).                                                                              (3)",
          "topic": null
        },
        {
          "label": "vii",
          "marks": 4,
          "text": "Assuming the log amount (i.e. log of claim amount) follows a Normal distribution, test\n               if mean of log amount is greater than 10 at 90% level of confidence. State the\n               hypothesis and conclusion clearly.                                                        (4)",
          "topic": null
        },
        {
          "label": "viii",
          "marks": 3,
          "text": "Assess whether the policyholder experience (i.e. Cust_Exp) changes with more\n              number of claims. Create contingency table and perform test to check the above\n              assertion. State the hypothesis clearly.                                                   (3)",
          "topic": null
        },
        {
          "label": "ix",
          "marks": 1,
          "text": "Please explain why warning message appears while performing the above test.                (1)",
          "topic": null
        },
        {
          "label": "x",
          "marks": 5,
          "text": "Again, perform the above test by combining\n             a. Very Satisfied and Satisfied customers\n             b. Disappointed and very disappointed customers\n             and\n             c. 2 or more claims.\n\n             Please provide your conclusion on association of policyholder’s experience with\n             number of claims.                                                                           (5)",
          "topic": null
        },
        {
          "label": "xi",
          "marks": 5,
          "text": "Amount is defined to be large if the amount is greater than 100,000. Calculate 95%\n              confidence interval for proportion of large claim, and comment on the likelihood if\n              more than 25% of claims are large.                                                         (5)",
          "topic": null
        }
      ],
      "solution": "i)\nlibrary(dplyr)\n\n> str(policydata)\n'data.frame':          650 obs. of 4 variables:\n $ Policy : int 1 2 3 4 5 6 7 8 9 10 ...\n $ Claim : int 0 0 0 2 1 0 0 0 0 0 ...\n $ Cust_Exp: chr \"SA\" \"SA\" \"SA\" \"DS\" ...\n $ Amount : int 0 0 0 52601 56174 0 0 0 0 0 ...\n>\n> #a\n> table(policydata$Claim)\n\n 0 1 2 3\n458 149 36 7\n\n                                                                                          Page 5 of 13\n\fIAI                                                                                       CS1B-0322\n\n> #Alternative, if dplyr installed\n> #count(policydata,Claim)\n\n> print(\"458 Policies don't have any claim\")\n[1] \"458 Policies don't have any claim\"\n\nii)\n> hist(policydata$Claim)\n> #poisson and negative binomial distribution\n\n                               Histogram of policydata$Claim\n             400\n Frequency\n\n             200\n             0\n\n                   0.0   0.5         1.0         1.5          2.0   2.5        3.0\n\n                                           policydata$Claim\n\niii)\n> poisson.test(x=sum(policydata$Claim),T=length(policydata$Policy))\n\nExact Poisson test\n\ndata: sum(policydata$Claim) time base: length(policydata$Policy)\nnumber of events = 242, time base = 650, p-value < 2.2e-16\nalternative hypothesis: true event rate is not equal to 1\n95 percent confidence interval:\n 0.3268739 0.4222903\nsample estimates:\nevent rate\n 0.3723077\n> #0.35 is more suitable value of parameter since it lies between confidence interval.\n\niv)\n> lx=log(policydata$Amount[policydata$Amount>0])\n> #Alternative, if dplyr installed\n> #lx=log(filter(policydata,Amount >0)$Amount)\n\n> mean(lx)\n[1] 9.835205\n> median(lx)\n[1] 9.774659\n> sd(lx)^2\n[1] 3.425705\n\n                                                                                         Page 6 of 13\n\fIAI                                                                                         CS1B-0322\n\nv)\n> par(mfrow=c(2,1))\n> hist(lx)\n> qqnorm(lx)\n> qqline(lx)\n\n                                                 Histogram of lx\n                   40\n                   30\nFrequency\n\n                   20\n                   10\n                   0\n\n                               6            8             10             12       14\n\n                                                          lx\n\n                                                 Normal Q-Q Plot\n                   14\nSample Quantiles\n\n                   10\n                   8\n                   6\n\n                        -3         -2       -1            0              1    2        3\n\n                                                 Theoretical Quantiles\n\nvi)\n> # From Histogram and QQPlot it seems log amount closely follows normal distribution.\n> # To add, the mean and median are very close indicating symmtery. One of the characterstics of Z.\n> # Hence,Claim amount might be following log normal distribution.\n\nvii)\n> #Null Hypothesis : mu = 10 , alternate hypothesis mu >10\n> t.test(lx,mu=10,alternative=\"greater\", conf.level = .9)\n\n                        One Sample t-test\n\ndata: lx\nt = -1.2337, df = 191, p-value = 0.8906\nalternative hypothesis: true mean is greater than 10\n90 percent confidence interval:\n 9.663428 Inf\nsample estimates:\nmean of x\n\n                                                                                           Page 7 of 13\n\fIAI                                                                               CS1B-0322\n\n9.835205\n\n> #Given p-value greater than 10% null hypothesis can not be rejected.\n\nviii)\n> ct=table(policydata$Claim,policydata$Cust_Exp)\n> ct\n\n   DS SA VD VS\n 0 63 306 20 69\n 1 36 90 9 14\n 2 16 14 6 0\n 3 3 3 1 0\n> # Null Hypothesis: No association between Policyholder's experience and Claim\n> chisq.test(ct)\n\n           Pearson's Chi-squared test\n\ndata: ct\nX-squared = 47.749, df = 9, p-value = 2.846e-07\n\nWarning message:\nIn chisq.test(ct) : Chi-squared approximation may be incorrect\n\nix)\n> #There are cells where the number of observations are less than 5.                      (1)\n\nx)\n> policydata$Claim2=ifelse(policydata$Claim >2,2,policydata$Claim)\n> policydata$Cust_Exp2=ifelse(policydata$Cust_Exp %in% c(\"DS\",\"VD\"),\"DS\",\"SA\")\n> ct2=table(policydata$Claim2,policydata$Cust_Exp2)\n> ct2\n\n  DS SA\n 0 83 375\n 1 45 104\n 2 26 17\n\n> chisq.test(ct2)\n\n           Pearson's Chi-squared test\n\ndata: ct2\nX-squared = 43.514, df = 2, p-value = 3.557e-10\n\n> # There is a strong reason to reject null hypothesis.\n> # Hence, it can concluded that policyholder's experience gets worse as\nclaim count increases\n\nxi)\n> summary(policydata$Amount)\n\n                                                                                  Page 8 of 13\n\f    IAI                                                                       CS1B-0322\n\n      Min. 1st Qu. Median Mean 3rd Qu. Max.\n        0    0   0 29501 3232 1848069\n    > policydata$large= ifelse(policydata$Amount >100000,1,0)\n    > x = sum(policydata$large)\n\n    > n = length(policydata$Amount[policydata$Amount>0])\n    > #Alternative, if dplyr installed\n    > #n = length(filter(policydata,Amount >0)$Amount)\n\n    > binom.test(x,n)\n\n               Exact binomial test\n\n    data: x and n\n    number of successes = 35, number of trials = 192, p-value < 2.2e-16\n    alternative hypothesis: true probability of success is not equal to 0.5\n    95 percent confidence interval:\n     0.1303796 0.2442928\n    sample estimates:\n    probability of success\n           0.1822917\n\n    >\n    > # Since upper bound of c.i is less that .25, it is unlikely that more\n    that\n    > #25% claims are large",
      "has_math": false,
      "is_r_task": true,
      "session": "2022-03",
      "subject": "CS1B",
      "source_qp": "raw/CS1B_2022-03_QP.pdf",
      "source_sol": "raw/CS1B_2022-03_SOL.pdf"
    },
    {
      "q_num": 3,
      "marks": 35,
      "topic": "distributions",
      "subtopics": [],
      "stem": "A General Insurance company is trying to analyse the two-wheeler motor insurance claims\n          reported over last one quarter.\n\n          The data is provided herewith the file MotorClaim.csv which contains the following fields\n\n          POLICY: Policy Number\n          CLAIM : Amount of claim reported for a policy\n\n          Insurance Company is interested to find out an appropriate distribution to fit the “CLAIM”\n          data. You are being asked to find out the appropriateness of the following distributions\n          based on method of moments:\n\n             1.   Normal distribution\n             2.   Lognormal distribution\n             3.   Exponential distribution\n             4.   Gamma distribution.",
      "parts": [
        {
          "label": "i",
          "marks": 8,
          "text": "Estimate the parameters of each of the above distributions.                                     (8)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 8,
          "text": "Plot a histogram of “CLAIM” data with 35 equal class intervals. Superimpose the\n          histogram with the probability density function of the above four distributions using\n          their estimated parameters as obtained in part (i). Mark each plot distinctly using\n          appropriate legend.                                                                            (8)",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 5,
          "text": "Compute the 5th percentile, 1st quartile, median, 3rd quartile and 95th percentile of both\n           the actual claim paid as well as the fitted distributions.",
          "topic": null
        },
        {
          "label": "iv",
          "marks": 4,
          "text": "Using the results from (ii) and (iii) comment on goodness of fit of the models to the\n          data.                                                                                          (4)",
          "topic": null
        },
        {
          "label": "v",
          "marks": 2,
          "text": "Assuming Gamma distribution to be the right fit to the data, simulate 20,000 values of\n         claim amounts using the Gamma distribution based on the parameter estimates\n         obtained in part (i) and print first 10 values of claim amounts. (Set seed to 2022)             (2)",
          "topic": null
        },
        {
          "label": "vi",
          "marks": 5,
          "text": "Generate 700 different random samples of size 400 from the simulated data obtained\n          in part (v) and compute sample mean for each of the samples. (Set seed to 2022)                (5)",
          "topic": null
        },
        {
          "label": "vii",
          "marks": 3,
          "text": "Plot the histogram of sample means generated from part (vi) and comment on the\n           distribution of the sample means from the point of view of central limit theorem.             (3)",
          "topic": null
        }
      ],
      "solution": "# Sample mean and variance\n    Motorclaim = read.csv(\"Motorclaim.CSV\")\n    Mean_Claim<-mean(Motorclaim$CLAIM)\n    Var_Claim<-var(Motorclaim$CLAIM)\n\n    i)\n    # Method of moments estimate\n\n    # Normal Distribution\n\n    Normal_mu <- Mean_Claim\n    Normal_sigma <- sqrt(Var_Claim)\n\n    Normal_mu\n    [1] 6357.314\n\n    Normal_sigma\n    [1] 6986.523\n\n    # Log Normal Distribution\n\n    LogNormal_sigma<- sqrt(log(1+Var_Claim/Mean_Claim^2))\n    LogNormal_mu<-log(Mean_Claim)-LogNormal_sigma^2/2\n\n                                                                              Page 9 of 13\n\fIAI                                                                                     CS1B-0322\n\nLogNormal_sigma\n[1] 0.8899276\n\nLogNormal_mu\n[1] 8.361376\n\n# Exponential Distribution\n\nExp_lamda <- 1/Mean_Claim\n\nExp_lamda\n[1] 0.0001572991\n\n# Gamma Distribution\n\nGamma_lamda<-Mean_Claim/Var_Claim\nGamma_alpha<-Gamma_lamda*Mean_Claim\n\nGamma_lamda\n[1] 0.0001302421\n\nGamma_alpha\n[1] 0.82799\n\nii)\n# Histogram\n\nhist(Motorclaim$CLAIM,breaks = 35,freq = FALSE)\n\n#Superimpose Normal distribution\n\ncurve(dnorm(x,mean = Normal_mu,sd = Normal_sigma),from = min(Motorclaim$CLAIM), to =\nmax(Motorclaim$CLAIM), add = TRUE, col= \"blue\")\n\n#Superimpose Log Normal distribution\n\ncurve(dlnorm(x,meanlog = LogNormal_mu,sdlog = LogNormal_sigma),from =\nmin(Motorclaim$CLAIM), to = max(Motorclaim$CLAIM), add = TRUE, col= \"green\")\n\n#Superimpose Exponential distribution\n\ncurve(dexp(x,rate = Exp_lamda),from = min(Motorclaim$CLAIM), to = max(Motorclaim$CLAIM), add\n= TRUE, col= \"red\")\n\n#Superimpose Gamma distribution\n\ncurve(dgamma(x,shape = Gamma_alpha,rate = Gamma_lamda),from = min(Motorclaim$CLAIM), to\n= max(Motorclaim$CLAIM), add = TRUE, col= \"yellow\")\n\nlegend(\"topright\",legend = c(\"Normal\", \"Lognormal\", \"Exponential\", \"Gamma\"),lty = 1, col =\nc(\"blue\",\"green\",\"red\",\"yellow\"))\n\n                                                                                     Page 10 of 13\n\fIAI                                                                                 CS1B-0322\n\niii)\n\n# Quantiles\n\n# Actual Claim Data\n\nquantile(Motorclaim$CLAIM,c(0.05,0.25,0.5,0.75,0.95))\n\n   5%     25%    50%       75%    95%\n1324.561 1934.876 3631.070 7870.028 21246.913\n\n# Normal Distribution\n\nqnorm(c(0.05,0.25,0.5,0.75,0.95),mean = Normal_mu,sd = Normal_sigma)\n\n[1] -5134.494 1644.976 6357.314 11069.653 17849.123\n\n# Log Normal Distribution\n\nqlnorm(c(0.05,0.25,0.5,0.75,0.95),meanlog = LogNormal_mu,sdlog = LogNormal_sigma)\n\n[1] 989.8714 2347.5526 4278.5767 7798.0014 18493.5327\n\n# Exponential Distribution\n\nqexp(c(0.05,0.25,0.5,0.75,0.95),rate = Exp_lamda)\n\n[1] 326.0876 1828.8853 4406.5544 8813.1089 19044.8114\n\n                                                                                Page 11 of 13\n\fIAI                                                                                         CS1B-0322\n\n# Gamma Distribution\n\nqgamma(c(0.05,0.25,0.5,0.75,0.95),shape = Gamma_alpha,rate = Gamma_lamda)\n\n[1] 193.6261 1479.4200 4053.4299 8797.0450 20369.6614\n\niv) From the histogram and superimposed plots it is clear that normal distribution is not good fit to\nthe data.\n\nThe other three plots are getting superimposed more or less similar to the data. From the quantiles it\nis observed that lower value(5th percentile) of lognormal is closed to actual value and higher\nvalues(95th percentile) of gamma distribution is closed to actual value\n\nThe best fitting distribution among Lognormal, exponential & Gamma can not be decided basis of\nobservations from (ii) & (iii). Further statistical tests need to be carried out to confirm best fit\n\nv)\n# Simulation from Gamma distribution\n\nset.seed(2022)\nSim_samples <- rgamma(20000,Gamma_alpha,Gamma_lamda)\n\nhead(Sim_samples,10)\n[1] 9505.735311 1376.831631 458.302589 3189.065594 5.340363 5821.017458\n [7] 11122.004509 5372.490004 43002.362493 3557.086406\nvi)\n# Generating 700 random samples of size 400 and computing sample means\n\nmeans<-c()\nset.seed(2022)\nfor (i in 1:700){\nselected_data_point<-sample(1:20000,400,FALSE)\nrandom_sample<- Sim_samples[selected_data_point]\nsample_mean<-mean(random_sample)\nmeans<-c(means,sample_mean)\n}\n\nvii)\n# Histogram of the sample means\nhist(means,breaks = 40)\n\n                                                                                          Page 12 of 13\n\fIAI                                                                                        CS1B-0322\n\nComment:\nThe distribution of sample means tend to follow normal distribution however the actual data comes\nfrom gamma distribution. Central Limit Theorem states that the sample means tend to follow\nnormal distribution as the sample size increases. The distribution of sample means will be closer to\nnormal distribution by increasing the sample size from its current level of 400.\n\n                                  ***************************\n\n                                                                                         Page 13 of 13",
      "has_math": false,
      "is_r_task": true,
      "session": "2022-03",
      "subject": "CS1B",
      "source_qp": "raw/CS1B_2022-03_QP.pdf",
      "source_sol": "raw/CS1B_2022-03_SOL.pdf"
    },
    {
      "q_num": 1,
      "marks": 26,
      "topic": "distributions",
      "subtopics": [],
      "stem": "Identify the probability distribution that best fits the below questions. Calculate the\n          following using R functions:",
      "parts": [
        {
          "label": "i",
          "marks": 4,
          "text": "Assume golf balls from the driving range next door lands in your yard at an average\n             rate of 3 balls per hour during the day. What is the probability that 10 or fewer golf\n             balls will land in your yard during the afternoon, assuming the afternoon is 5 hours\n             long?                                                                                          (4)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 5,
          "text": "You are surveying people exiting from a polling booth and asking them if they voted\n              independently. The probability that a person voted independently is 20%. What is\n              the probability that 70 people must be asked before you can find 5 people who voted\n              independently?                                                                                (5)",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 4,
          "text": "Assume you flip a fair coin 100 times. What is the number N such that, 90% of the\n               time, the number of heads is less than or equal to N?                                        (4)",
          "topic": null
        },
        {
          "label": "iv",
          "marks": 4,
          "text": "A researcher is waiting outside of a library to ask people if they support a certain\n              law. The probability that a given person supports the law is p = 0.2. What is the\n              probability that the fourth person the researcher talks to is the first person to support\n              the law?                                                                                      (4)",
          "topic": null
        },
        {
          "label": "v",
          "marks": 5,
          "text": "Assume that a light bulb has a mean lifetime of 1000 hours. What is the probability\n             that the light bulb survives to 2000 hours?                                                    (5)",
          "topic": null
        },
        {
          "label": "vi",
          "marks": 4,
          "text": "Assume a random variable Z is distributed according to the normal distribution with\n              mean 6 and standard deviation 4. What is the probability that Z takes on a value\n              between -1 and 3?                                                                            (4)",
          "topic": null
        }
      ],
      "solution": "i)    X ~ Poisson (15)\n      ppois(10, 15)\n      [1] 0.1184644\nii) X ~ NB(5,0.2)\n    dnbinom(65,5,0.2)\n    [1] 0.00013892\niii) X ~ Binom (100,0.5)\n     qbinom(0.9, 100, 0.5)\n     [1] 56\niv) X ~ Geometric (0.2)\n    dgeom(x=3, prob=0.2)\n    [1] 0.1024\nv) X ~ Exp(1/1000)\n   1 - pexp(2000, 0.001)\n   [1] 0.1353353\nvi) X ~ N(6,16)\n    pnorm(3, 6, 4) - pnorm(-1, 6, 4)\n    [1] 0.1865682",
      "has_math": false,
      "is_r_task": true,
      "session": "2022-07",
      "subject": "CS1B",
      "source_qp": "raw/CS1B_2022-07_QP.pdf",
      "source_sol": "raw/CS1B_2022-07_SOL.pdf"
    },
    {
      "q_num": 2,
      "marks": 24,
      "topic": "inference",
      "subtopics": [],
      "stem": "In a study done at the National Institute of Science and Technology, asbestos fibers on\n          filters were counted as part of a project to develop measurement standards for asbestos\n          concentration.\n\n          An operator counted the number of fibers in each of 23 grid squares, yielding the\n          following counts:\n\n          31,29,19,18,31,28, 34,27,34,30,16,18, 26,27,27,18,24,22, 28,24,21,17,24\n\n          Assume that the Poisson distribution with unknown parameter lambda describes the\n          variability from each of the grid squares.",
      "parts": [
        {
          "label": "i",
          "marks": 4,
          "text": "Calculate Q1, Q3 and Inter-quartile range.                                                     (4)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 3,
          "text": "Plot histogram of sample data and label it appropriately.                                     (3)",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 3,
          "text": "Use the method of maximum likelihood to estimate the parameter lambda.                       (3)",
          "topic": null
        },
        {
          "label": "iv",
          "marks": 5,
          "text": "Test the hypothesis whether the mean fiber count is equal to 25. Comment on the\n              results.                                                                                      (5)",
          "topic": null
        },
        {
          "label": "v",
          "marks": 2,
          "text": "Calculate the standard error of parameter lambda.                                              (2)",
          "topic": null
        },
        {
          "label": "vi",
          "marks": 3,
          "text": "Calculate the 90% confidence interval for standard error.                                     (3)",
          "topic": null
        },
        {
          "label": "vii",
          "marks": 4,
          "text": "Calculate the probability of fiber count exceeding 30, with the help of Central Limit\n               Theorem.                                                                                 (4)",
          "topic": null
        }
      ],
      "solution": "i)    count <- c(31,29,19,18,31,28, 34,27,34,30,16,18, 26,27,27,18,24,22, 28,24,21,17,24)\n\n> quantile(count,0.25)\n\n25%\n\n20\n\n> quantile(count,0.75)\n\n75%\n\n28.5\n\n> IQR(count)\n[1] 8.5\nii) hist(count)\n\n                                                                                            Page 2 of 7\n\fIAI                                                                                           CS1B-0722\n\niii) lambda.hat=mean(x)\n\nprint(lambda.hat)\n[1] 24.91304\niv) Ho: The mean fiber count is 25\nH1: Mean fiber count is not equal to 25\n\n> t.test(count,mu=25)\n\n                One Sample t-test\n\ndata: count\nt = -0.076034, df = 22, p-value = 0.9401\nalternative hypothesis: true mean is not equal to 25\n95 percent confidence interval:\n 22.54124 27.28485\nsample estimates:\nmean of x\n 24.91304\n\nBased on the p-value the null hypothesis Ho that “the mean fiber count is 25” cannot be rejected. Also 25\nlies within the 95% confidence interval.\nv) lambda.hat.sterror=sqrt(lambda.hat/length(x))\n\nprint(lambda.hat.sterror)\n\n[1] 1.040757\n\n                                                                                             Page 3 of 7\n\fIAI                                                                                       CS1B-0722\n\nvi) lambda.CI.Limits=lambda.hat + c(-1,1)*qnorm(.95)*lambda.hat.sterror\n\nprint(lambda.CI.Limits)\n\n[1] 23.20115 26.62494\nvii) > pnorm(30,lambda.hat,sqrt(lambda.hat),lower.tail = FALSE)\n\n[1] 0.1540622",
      "has_math": false,
      "is_r_task": true,
      "session": "2022-07",
      "subject": "CS1B",
      "source_qp": "raw/CS1B_2022-07_QP.pdf",
      "source_sol": "raw/CS1B_2022-07_SOL.pdf"
    },
    {
      "q_num": 3,
      "marks": 17,
      "topic": "inference",
      "subtopics": [],
      "stem": "An analysis was carried out to investigate the annual average rainfall of two countries.\n          The data is as below:\n\n          Iran: 128,125,133,104,146,132,125,118,129,124\n\n          Belgium:160,128,169,105,151,164,162,177,185,150,182,158,156,123,141,176,162,172",
      "parts": [
        {
          "label": "i",
          "marks": 6,
          "text": "Perform a suitable test to determine whether the rainfall in both countries has equal\n             variance or not, at the 5% confidence level.                                                (6)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 6,
          "text": "Test whether the mean rainfall in both countries is equal or not, at 5% confidence\n              level.                                                                                     (6)",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 2,
          "text": "Calculate the 95% confidence interval for difference in means.                            (2)",
          "topic": null
        },
        {
          "label": "iv",
          "marks": 3,
          "text": "Comment on the results in part (ii) and (iii).                                            (3)",
          "topic": null
        }
      ],
      "solution": "i)\n\nH0: Annual average rainfall of Belgium and Iran are same\n\nH1: Annual average rainfall of Belgium and Iran are not same\n\nIran <- c(128,125,133,104,146,132,125,118,129,124)\n\nBelgium <- c(160,128,169,105,151,164,162,177,185,150,182,158,156,123,141,176,162,172)\n\nvar.test(Iran, Belgium)\n\nF test to compare two variances\n\ndata: Iran and Belgium\n\nF = 0.25802, num df = 9, denom df = 17, p-value = 0.04385\n\nalternative hypothesis: true ratio of variances is not equal to 1\n\n95 percent confidence interval:\n\n0.0864436 0.9602591\n\nsample estimates:\n\nratio of variances\n\n      0.258022\n\nSince p-value < 0.05 we fail the variance test thus we reject the null hypothesis that both have equal\nvariance\nii)\n\nt.test(Iran, Belgium, var.equal = FALSE)\n\nOUTPUT\n\ndata: Iran and Belgium\n\nt = -4.9984, df = 25.904, p-value = 3.407e-05\n\nalternative hypothesis: true difference in means is not equal to 0\n\n                                                                                           Page 4 of 7\n\fIAI                                                                                           CS1B-0722\n\n95 percent confidence interval:\n\n-42.79403 -17.85041\n\nsample estimates:\n\nmean of x mean of y\n\n126.4000 156.7222\n\nSince, P-value<0.05 we reject the null hypothesis and can conclude that both cities have different amount\nof rainfall with 95% confidence.\n\niii) Confidence interval can be read from part b\n\n95 percent confidence interval:\n-42.79403 -17.85041\n\niv) The confidence interval (-42.8,-17.8) does not contain 0, therefore the assumption of equal means is\n    not true. This result is in line with the conclusion in part (b).",
      "has_math": false,
      "is_r_task": true,
      "session": "2022-07",
      "subject": "CS1B",
      "source_qp": "raw/CS1B_2022-07_QP.pdf",
      "source_sol": "raw/CS1B_2022-07_SOL.pdf"
    },
    {
      "q_num": 4,
      "marks": 33,
      "topic": "regression_glm",
      "subtopics": [
        "inference"
      ],
      "stem": "The marketing dataset contains the impact of three advertising medias\n          (youtube, facebook and newspaper) on sales. The first three columns are the\n          advertising budget in thousands of dollars along with the fourth column as sales. The\n          advertising experiment has been repeated 200 times.\n\n          The marketing data is provided in a csv file “data.csv”.",
      "parts": [
        {
          "label": "i",
          "marks": 5,
          "text": "Plot the data. Analyze the trend of how sales varies with the advertising budget of all\n             3 advertising medias and comment on the same.                                               (5)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 6,
          "text": "Perform a simple linear regression analysis on the data. Your answer should include\n              summary of the data.                                                                       (6)",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 3,
          "text": "Comment on the significance of the parameters of the model and justify your\n               observations from point (i).                                                              (3)",
          "topic": null
        },
        {
          "label": "iv",
          "marks": 3,
          "text": "Calculate the correlation between independent and dependent variables.                     (3)",
          "topic": null
        },
        {
          "label": "v",
          "marks": 4,
          "text": "Fit an improved model for the model in part (ii), using your answer in part (iv).\n             State the linear regression formula clearly explaining all parameters.                      (4)",
          "topic": null
        },
        {
          "label": "vi",
          "marks": 3,
          "text": "Which model is better between part (ii) and (v) and why?                                   (3)",
          "topic": null
        },
        {
          "label": "vii",
          "marks": 2,
          "text": "What is the maximum sales generated?                                                      (2)",
          "topic": null
        },
        {
          "label": "viii",
          "marks": 4,
          "text": "Based on the linear regression model fitted in question (v), what is the predicted\n                value for maximum sales generated in question (vii) .                                    (4)",
          "topic": null
        },
        {
          "label": "ix",
          "marks": 3,
          "text": "What is the relative error between the estimated (prediction calculated in question\n              (viii)) and the actual sales computed in question (vii)?                                  (3)",
          "topic": null
        }
      ],
      "solution": "marketing = read.csv(\"data.csv\")\n\ndata_size = dim(marketing)\n\ni)    plot(marketing)\n\n                                                                                             Page 5 of 7\n\fIAI                                                                                            CS1B-0722\n\n         The last row of the plot indicates how various advertising channel budgets impact the sales. We\n         can clearly see that youtube and facebook sales increase linearly with increase in the advertising\n         budget. The newspaper (3rd plot) sales however shows no particular trend.\n\nii)      > Model <- lm(sales ~ youtube + facebook + newspaper, data = marketing)\n         > summary (Model)\n\n         Call:\n         lm(formula = sales ~ youtube + facebook + newspaper, data = marketing)\n\n         Residuals:\n            Min     1Q Median   3Q Max\n         -10.5932 -1.0690 0.2902 1.4272 3.3951\n\n         Coefficients:\n                  Estimate Std. Error t value Pr(>|t|)\n         (Intercept) 3.526667 0.374290 9.422 <2e-16 ***\n         youtube 0.045765 0.001395 32.809 <2e-16 ***\n         facebook 0.188530 0.008611 21.893 <2e-16 ***\n         newspaper -0.001037 0.005871 -0.177 0.86\n         ---\n         Signif. codes:\n         0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1\n\n         Residual standard error: 2.023 on 196 degrees of freedom\n         Multiple R-squared: 0.8972,      Adjusted R-squared: 0.8956\n         F-statistic: 570.3 on 3 and 196 DF, p-value: < 2.2e-16\n\n  iii)   It can be seen that from the estimates column and from p values that, changes in the youtube\n         and facebook advertising budgets are significantly associated to changes in sales while changes in\n         the newspaper budget is not.                                                                  [3]\n\n  iv)    > cor(marketing$youtube,marketing$sales)\n         [1] 0.7822244\n         > cor(marketing$facebook,marketing$sales)\n         [1] 0.5762226\n         > cor(marketing$newspaper,marketing$sales)\n         [1] 0.228299\n\n         The pairwise plot and the above correlation indicated the same conclusion on newspaper having\n         a very low / no particular trend with respect to sales.\n\n      v) > Model1 <- lm(sales ~ youtube + facebook , data = marketing)\n         > summary(Model1)\n\n                                                                                               Page 6 of 7\n\fIAI                                                                                             CS1B-0722\n\n          Call:\n          lm(formula = sales ~ youtube + facebook, data = marketing)\n\n          Residuals:\n             Min     1Q Median   3Q Max\n          -10.5572 -1.0502 0.2906 1.4049 3.3994\n\n          Coefficients:\n                  Estimate Std. Error t value Pr(>|t|)\n          (Intercept) 3.50532 0.35339 9.919 <2e-16 ***\n          youtube 0.04575 0.00139 32.909 <2e-16 ***\n          facebook 0.18799 0.00804 23.382 <2e-16 ***\n          ---\n          Signif. codes:\n          0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1\n\n          Residual standard error: 2.018 on 197 degrees of freedom\n          Multiple R-squared: 0.8972,      Adjusted R-squared: 0.8962\n          F-statistic: 859.6 on 2 and 197 DF, p-value: < 2.2e-16\n\n          Sales = 3.5 + 0.045*youtube + 0.187*facebook\n\n      vi) Adjusted R squared for Model in part (b) and that of part (e) is 0.89, hence there is no particular\n          improvement after removing newspaper parameter. However, a model with less parameters is\n          considered better, hence we can consider Model 1 calculated in part (f) to be a good fit.       [3]\n\n      vii) > marketing[which.max(marketing$sales),]\n\n  youtube facebook newspaper sales\n\n          176 332.28 58.68      50.16 32.4\n\n          Maximum sales generated is 32.4 thousand dollars.                                              [2]\n\n      viii) > PredTest = predict(Model1)\n            > PredTest[176]\n\n          29.74023\n\n      ix) (Observed ILI - Estimated ILI)/Observed ILI\n          > (32.4-29.74023)/32.4\n          [1] 0.08209167\n                                    *************************************\n\n                                                                                                 Page 7 of 7",
      "has_math": false,
      "is_r_task": true,
      "session": "2022-07",
      "subject": "CS1B",
      "source_qp": "raw/CS1B_2022-07_QP.pdf",
      "source_sol": "raw/CS1B_2022-07_SOL.pdf"
    },
    {
      "q_num": 1,
      "marks": 19,
      "topic": "data_analysis",
      "subtopics": [],
      "stem": "",
      "parts": [
        {
          "label": "i",
          "marks": 2,
          "text": "Use the following code to generate 150 random numbers from uniform [0,1] distribution.\n              Show that mean of the sample ~= 0.52 at 2 decimal places.\n              set.seed(052023)\n              u<-runif(150)                                                                             (2)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 2,
          "text": "Using the random numbers generated in part (i), simulate sample with sample size = 150\n            from chi-square distribution with 2 degrees of freedom.\n            No need to print the sample. Store the sample.\n            [Hint: useful functions : qchisq]                                                           (2)",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 3,
          "text": "Using the random numbers generated in part (i) simulate another sample with sample\n             size = 150 from Gamma (1, ½) distribution.\n              Show and Explain why samples generated in part (ii) and (iii) are same by establishing\n              link between the chisquare and gamma distributions.                                       (3)",
          "topic": null
        },
        {
          "label": "iv",
          "marks": 2,
          "text": "a) Plot histogram of sample generated in part (ii) and comment on the shape of the\n                 distribution.                                                                          (3)\n\n              b) Compute mean & median of sample and explain why mean is greater than median.           (2)",
          "topic": null
        },
        {
          "label": "v",
          "marks": 4,
          "text": "Simulate 1000 values of sum of samples of size 150 from chi-square distribution with 2\n           degrees of freedom.\n              Store the value of sum of samples. Also, make sure to set.seed(052023) again before\n              generating samples.\n              No need to print the data.                                                                (4)",
          "topic": null
        },
        {
          "label": "vi",
          "marks": 3,
          "text": "Plot histogram of 1000 samples sum generated in part (v) and comment on the shape of\n            the distribution in the context of central limit theorem.                                   (3)",
          "topic": null
        }
      ],
      "solution": "i) set.seed(052023)\n        i                                                                                                (1)\n       u<-runif(150)\n        .\n       round(mean(u),2)                                                                                  (1)\n       [1] 0.52\n\n         This shows sample mean ~0.52.                                                                  [2]\n\n ii)     chi<-qchisq(u,2)\n         i                                                                                               [2]\n\niii)     gam<-qgamma(u,1,1/2)\n         i                                                                                               (1)\n          i\n         sum(chi-gam)\n         i                                                                                             (0.5)\n         . 0\n         We know the property that if X ~ Gamma(α,ʎ) then 2ʎX has χ2 distribution with 2α degrees of     (1)\n         freedom.\n\n         We have X ~ χ2 with 2 degrees of freedom.\n         Above can be written as (2ʎX/2ʎ) ~ χ2 with 2α degrees of freedom where α=1,ʎ =1/2             (0.5)\n         Thus,\n         X ~ Gamma(α,ʎ) with α=1,ʎ =1/2\n         X ~ Gamma(1,1/2)                                                                              (0.5)\n\n         This is why both samples are same.\n                                                                                                   [Max 3]\n\n iv)      .\n      a) hist(chi,\n         a         main =\"Histogram of chi square distribution sample\")                                  (1)\n\n         #Histogram shows positive skewed distribution                                                  (1)\n\n      b) >bsummary(chi)                                                                                  (1)\n          . Min. 1st Qu. Median    Mean 3rd Qu. Max.\n          0.006184 0.603187 1.534968 2.217946 2.983161 12.350562\n\n         Alternate:\n         mean(chi)                                                                                     (0.5)\n         median(chi)                                                                                   (0.5)\n         #mean = 2.218\n         #median = 1.535\n\n         #Mean is greater than median since it is positively skewed distribution                        (1)\n                                                                                                  Page 2 of 9\n\fIAI                                                                          CS1B-0523\n\n v)      set.seed(052023)\n         v                                                                          (0.5)\n         y = rep(0,1000)                                                              (1)\n         for(i in 1:1000){                                                            (1)\n           y[i] = sum(rchisq(150,2))                                                  (2)\n         }\n                                                                                 [Max 4]\n\n vi)     hist(y,\n         .       main =\"Histogram of Samples Sum\")                                  (0.5)\n\n         #The distribution of sample sums is roughly symmetrical.                    (1)\n         #This displays Central Limit Theorem property. As the sample size           (1)\n         #gets large , distribution move towards normality.                          (1)\n                                                                                 [Max 3]",
      "has_math": false,
      "is_r_task": true,
      "session": "2023-05",
      "subject": "CS1B",
      "source_qp": "raw/CS1B_2023-05_QP.pdf",
      "source_sol": "raw/CS1B_2023-05_SOL.pdf"
    },
    {
      "q_num": 2,
      "marks": 61,
      "topic": "regression_glm",
      "subtopics": [
        "inference",
        "distributions"
      ],
      "stem": "An investment firm is planning to invest in a new start-up GPT. To assess the valuations,\n         investment firm asked you to study the sales of GPT.\n\n         SalesData.csv contains the sample sales data of 60 repeat customers having the following\n         information:\n\n         Order: Number of orders made by customers in last month\n         Value: Average order value per order of customers\n         Device: Device used to order\n         Age: Age of customers\n         City: Location of customers (grouped into 2 segments)",
      "parts": [
        {
          "label": "i",
          "marks": 2,
          "text": "Using read.csv load the Sales data and compute sample mean of value.                         (2)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 2,
          "text": "Compute\n\n              a) Kendall correlation between Value and City.                                            (1)\n\n              b) Covariance between Value and Age.                                                      (2)",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 3,
          "text": "a) Create a scatterplot between Value and Age. Without performing any test, state\n                whether the hypothesis that correlation between Value and City equals 0 can be\n                accepted or not. Provide reason for the same.                                            (3)\n\n             b) Compute the confidence interval of correlation coefficient between Value and Age\n                and test the assertion that correlation coefficient is 0.                                (3)",
          "topic": null
        },
        {
          "label": "iv",
          "marks": 3,
          "text": "Project manager asked you to fit a multiple linear regression in which Value is modelled\n          as response and customer details (Device, Age and City) as explanatory variables.\n          However, due to time constraints only 2 models can be tested.\n\n             a) Using part (iii), suggest, including reason, which explanatory variable (customer\n                detail) should not be included while fitting the regression model.                       (2)\n\n             b) Fit regression model with Value as response and customer details (Device, Age and\n                City) as explanatory variables.\n                Comment on the significance of the parameters including linkage with your\n                suggestion in above part iv(a).                                                          (4)\n\n             c) Fit simple regression model Value ~Device and compare with the above model using\n                ANOVA.\n                Explain the result of ANOVA including the hypothesis considered, degrees of\n                freedom shown in output and inference drawn.                                             (5)\n\n             d) Write down the regression equation for the above model, fitted in part iv(c).\n                Please use the parameter values and not the symbols.                                     (3)",
          "topic": null
        },
        {
          "label": "v",
          "marks": 5,
          "text": "Model Value ~Device, fitted in part iv(c), is selected for further use.\n\n             Calculate a 95% confidence interval for\n\n             a) β , the true underlying slope parameter.                                                 (2)\n\n             b) ϭ2 , the true underlying error variance.                                                 (5)\n\n             Hint: Some useful functions : qchisq ,qt, confint",
          "topic": null
        },
        {
          "label": "vi",
          "marks": 8,
          "text": "Investor wants to test if number of orders follows Poisson distribution with mean µ.\n\n             a) Estimate µ using the Sales data.                                                         (2)\n\n             b) Using Sales data, create a frequency table showing number of customers by orders.        (2)\n\n             c) Perform goodness of test to assess Order follows Poisson distribution.\n                You can combine orders equal or more than 5 to ensure actual frequency is at least\n                5.\n\n                 You can use as.numeric and apply on frequency table to store the frequency of\n                 customers in a vector.\n                 Ensure that total sum of expected probabilities equals 1.                               (8)\n\n             Hint : Useful function : chisq.test",
          "topic": null
        },
        {
          "label": "vii",
          "marks": 2,
          "text": "a) Assuming the number of orders follow Poisson distribution, fit a poisson\n                Generalised Linear Model (GLM) specified as Order~Device and print its summary.\n\n                 Make sure to specify the family as Poisson and appropriate link function.                 (3)\n\n              b) State the link function you have used in above model and provide reason why it is\n                 appropriate for poisson distribution.                                                     (2)",
          "topic": null
        },
        {
          "label": "viii",
          "marks": 2,
          "text": "a) Project manager has asked you to predict the average number of order for following\n               2 customers using this fitted poisson model.\n\n                   Device     Age     City\n                   Mobile     18       1\n                   Mobile     28       2\n\n        [Hint: type=”response” provides prediction on the scale of the response variable. The\n        default is on the scale of the linear predictors].                                                 (4)\n\n              b) After observing the predicted average number of orders for both customers, project\n                 manager doubts if there is any error in the model. Explain to project manager that\n                 model is correct and predicted values are as expected.                                    (2)",
          "topic": null
        },
        {
          "label": "ix",
          "marks": 6,
          "text": "Investment firm wants to close one of the channels – Website (Laptop) or App (Mobile).\n            Project manager asks you to provide total value of these channels.\n\n              Total value is defined as Average order value x Average number of order.\n\n        Predict total value for laptop & mobile and suggest which channel to close.\n\n        Use models Value ~Device, fitted in part iv(c) and Order ~ Device fitted in part vii(a) for\n        predicting total value.                                                                            (6)",
          "topic": null
        }
      ],
      "solution": "i) Sales<-read.csv(<>)\n        i                                                                             (1)\n         .\n          > mean(Sales$Value)                                                         (1)\n         [1] 1307.167\n\n ii) i\n    a) cor(Sales$Value,Sales$City,method\n       a                                 = \"kendall\")                                 [1]\n       . -0.2130327\n\n      b) r<-cor(Sales$Value,Sales$Age)\n          b                                                                           (1)\n         >. r*sd(Sales$Value)*sd(Sales$Age)                                           (1)\n         [1] -278.5169\n         Alternate:\n         cov(sales$Value,sales$Age)                                                 (1.5)\n         [1] -278.5169                                                              (0.5)\n         Credit is given if Kendall covariance is computed.\niii) i\n    a) plot(Sales$Value,Sales$Age)\n       a                                                                              (1)\n       .\n\n                                                                               Page 3 of 9\n\fIAI                                                                                                     CS1B-0523\n\n         # No trend (showing linear relationship) is visible from the scatter plot.                                (1)\n         # Most likely it indicates that correlation is zero.\n\n      b) >bcor.test(Sales$Value,Sales$Age)                                                                         (1)\n          .\n         Pearson's  product-moment correlation\n         data: Sales$Value and Sales$Age\n         t = -0.95277, df = 58, p-value = 0.3447\n         alternative hypothesis: true correlation is not equal to 0\n         95 percent confidence interval:\n          -0.3665088 0.1340120                                                                                    (1)\n         sample estimates:                                                                                       for CI\n              cor\n         -0.1241369\n         Confidence Interval is ( -0.367,0.134)\n         Since 0 lies in the confidence interval, we cannot reject the hypothesis that correlation coefficient     (1)\n         = 0.\n\n iv)\n      a) Since\n         a     correlation between Value and Age is (close to) 0, age can be excluded.                             (2)\n\n      b) >bmodel1<-lm(data = Sales,Value~Device+City+Age)                                                          (1)\n         >. summary(model1)                                                                                        (1)\n\n         Call:\n         lm(formula = Value ~ Device + City + Age, data = Sales)\n\n         Residuals:\n           Min 1Q Median 3Q Max\n         -768.21 -239.97 -19.05 236.74 959.32\n\n         Coefficients:\n                  Estimate Std. Error t value Pr(>|t|)\n         (Intercept) 2887.26 378.29 7.632 3.12e-10 ***\n         DeviceMobile -1022.97 110.18 -9.284 6.28e-13 ***\n         City       -135.98 100.03 -1.359 0.1795\n         Age          -26.25 13.41 -1.958 0.0552 .\n         ---\n         Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1\n\n                                                                                                            Page 4 of 9\n\fIAI                                                                                              CS1B-0523\n         Residual standard error: 377 on 56 degrees of freedom\n         Multiple R-squared: 0.6388,      Adjusted R-squared: 0.6195\n         F-statistic: 33.01 on 3 and 56 DF, p-value: 2.019e-12\n\n         Device = 1 for Mobile and 0 for Laptop. Parameter for this variable is significant.\n         Alternate : Device is significant                                                              (0.5)\n\n         Age and City are not significant …                                                              (1)\n         …….since p-value > 0.05.                                                                      (0.5)\n         Age is expected to insignificant per the earlier analysis.                                    (0.5)\n                                                                                                     [Max 4]\n\n      c) model2<-lm(data\n          c              = Sales,Value~Device)                                                            (1)\n         >.\n         > anova(model1,model2)                                                                           (1)\n         Analysis of Variance Table\n\n         Model 1: Value ~ Device + City + Age\n         Model 2: Value ~ Device\n          Res.Df RSS Df Sum of Sq F Pr(>F)\n         1 56 7958529\n         2 58 8716151 -2 -757622 2.6655 0.07838 .\n         ---\n         Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1\n\n         H0: β (City) = β (Age) =0 against H1: atleast one of β (City) or β (Age) is non-zero.           (1)\n         In model 2, there are 2 less parameters thus -2 degrees of freedom in Anova analysis            (1)\n         p-value >0.05 showing we can’t reject H0: β (City) = β (Age) =0.                                (1)\n         This indicates neither of the covariates have strong relationship with Order value.           (0.5)\n                                                                                                     [Max 5]\n\n      d) summary(model2)\n         d                                                                                                (1)\n         .\n         Call:\n         lm(formula = Value ~ Device, data = Sales)\n\n         Residuals:\n           Min 1Q Median 3Q Max\n         -810.9 -240.1 -8.7 271.6 889.1\n\n         Coefficients:\n                  Estimate Std. Error t value Pr(>|t|)\n         (Intercept) 2056.47 94.02 21.873 < 2e-16 ***\n         DeviceMobile -1045.54 111.06 -9.414 2.77e-13 ***\n         ---\n         Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1\n\n         Residual standard error: 387.7 on 58 degrees of freedom\n         Multiple R-squared: 0.6044,      Adjusted R-squared: 0.5976\n         F-statistic: 88.62 on 1 and 58 DF, p-value: 2.77e-13\n\n         > # Value = 2056.47 - 1045.54 X Device,                                                        (1.5)\n                  where Device= 1 for Mobile else 0                                                     (0.5)\n         Alternate:\n         > # Value = 2056.47 - 1045.54 X Device Mobile                                                    (2)\n\n                                                                                                     [Max 3]\n\n                                                                                                   Page 5 of 9\n\fIAI                                                                                               CS1B-0523\n\n v) .\n   a) confint(model2,level=.95)\n      a                                                                                                       (1)\n      .        2.5 % 97.5 %\n         (Intercept) 1868.268 2244.6737\n         DeviceMobile -1267.855 -823.2257\n         > # C.I. (-1267.8,-823.2)                                                                            (1)\n\n      b) >bresidual<-model2$residuals                                                                         (1)\n         >. n<-length(residual)                                                                             (0.5)\n         > varhat<-var(residual)                                                                            (1.5)\n         > (n-2)*varhat/qchisq(c(0.975,.025),58)                                                              (2)\n         [1] 105867.1 220588.2\n         > # C.I. (105867,220588)                                                                             (1)\n                                                                                                          [Max 5]\n\n vi)    .\n    a) m<-mean(Sales$Order)\n        a                                                                                                     (1)\n       >. m\n         [1] 2.383333                                                                                         (1)\n          Mu hat = 2.383\n\n      b) table(Sales$Order)\n         b                                                                                                    (1)\n         .\n          0 1 2 3 4 5 7\n          8 13 13 8 12 5 1                                                                                    (1)\n\n      c) a<-as.numeric(table(Sales$Order))\n         c                                                                                                    (1)\n         #using\n         .      above table, combine order 5 and 5+\n         a[6]=sum(a[6:7])                                                                                     (1)\n         a<-a[-7] #to remove 5+ as combine above                                                            (0.5)\n\n         e<-dpois(c(0:4),m)                                                                                   (1)\n         sum(e)                                                                                             (0.5)\n         [1] 0.9062099\n         e[6]<-1 - sum(e)                                                                                     (1)\n         sum(e)                                                                                               (1)\n         [1] 1\n\n         chisq.test(x=a,p=e)                                                                                  (2)\n                 Chi-squared test for given probabilities\n         data: a\n         X-squared = 6.0026, df = 5, p-value = 0.306\n\n         Since p-value is greater than 0.05 , we can reject the hypothesis that number of order follows       (1)\n         poisson distribution.\n                                                                                                          [Max 8]\n\nvii)    .\n    a) glm2<-glm(data=Sales,Order~Device,family=poisson(link=\"log\"))\n        a                                                                                                     (2)\n       >. summary(glm2)                                                                                       (1)\n\n         Call:\n         glm(formula = Order ~ Device, family = poisson(lin = \"log\"),\n                                                                                                     Page 6 of 9\n\fIAI                                                                                                CS1B-0523\n           data = Sales)\n\n         Deviance Residuals:\n           Min    1Q Median     3Q Max\n         -2.4495 -0.6149 0.0000 0.5491 1.9652\n\n         Coefficients:\n                  Estimate Std. Error z value Pr(>|z|)\n         (Intercept) -0.1942 0.2673 -0.726 0.468\n         DeviceMobile 1.2928 0.2814 4.594 4.34e-06 ***\n         ---\n         Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1\n\n         (Dispersion parameter for poisson family taken to be 1)\n\n           Null deviance: 81.185 on 59 degrees of freedom\n         Residual deviance: 51.570 on 58 degrees of freedom\n         AIC: 199.88\n\n         Number of Fisher Scoring iterations: 5                                                              [3]\n\n      b) Log\n         b link function is used in above model.                                                            (0.5)\n         .\n         Log link implied log mu = linear predictor. Inverting to mu leads to mu = exp(linear predictor).\n         Exponent will make sure mu always remain greater than 0, an essential feature for poisson\n         distribution with mean mu.                                                                         (1.5)\n\nviii)     .\n      a) Customers<-\n          a            data.frame(Device =c(\"Mobile\",\"Mobile\"),Age =c(18,28), City = c(1,2))                (1.5)\n         >. predict.glm(glm2,newdata = Customers,type= \"response\")                                            (2)\n         12\n         33                                                                                                 (0.5)\n\n      b) Only\n         b device is used in the model and for both customers, device is same and thus, the predicted\n         value\n         .     is same for both customers.                                                                    [2]\n\n ix)     Channel <-data.frame(Device =c(\"Mobile\",\"Laptop\"))                                                 (1.5)\n         > pred_order <- predict.glm(glm2,newdata = Channel,type= \"response\")                               (1.5)\n         > pred_value <- predict(model2,newdata = Channel)                                                    (1)\n         >\n         > totalvalue<-pred_order * pred_value                                                              (1.5)\n         > totalvalue\n             1     2\n         3032.791 1693.564                                                                                    (1)\n         > #total value for Mobile = 3032.8\n         > # and for Laptop = 1693.6\n                                                                                                         [Max 6]",
      "has_math": true,
      "is_r_task": true,
      "session": "2023-05",
      "subject": "CS1B",
      "source_qp": "raw/CS1B_2023-05_QP.pdf",
      "source_sol": "raw/CS1B_2023-05_SOL.pdf"
    },
    {
      "q_num": 3,
      "marks": 20,
      "topic": "distributions",
      "subtopics": [],
      "stem": "Total sales per month on a particular cloud kitchen follow a normal distribution with\n        unknown mean θ and variance 202. Sales x1, x2,…..,xn are observed over n months. Prior\n        beliefs that θ follows normal distribution with mean 60 and variance 52.\n\n        Sales for last 5 months are extracted for analysis. Total sales for last 5 months were 340.",
      "parts": [
        {
          "label": "i",
          "marks": 3,
          "text": "Mθ(t) denotes MGF of θ. State and compute M’θ (0) and M” θ (0).\n           Use m1t0 for M’θ (0) and m2t0 for M” θ (0).                                                     (3)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 3,
          "text": "a) State the posterior distribution of θ and compute parameters of the distribution using\n                 the sales of last 5 months.                                                               (6)\n\n              b) After 50 months same analysis is performed using the 50 months sales data. Total\n                 sales for last 50 months were 3400.\n\n                 Compute the parameters of posterior distribution of θ using 50 months sales data.         (3)",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 3,
          "text": "a) Use the below code to plot the prior distribution of θ\n\n                 x<-60+seq(-3,3,by=0.2)*5\n                 y<-dnorm(x,mean=60,sd=5)\n                 plot(x,y,ylim=c(0,.02))                                                                       (1)\n\n             b) Add the line to overlay the posterior distribution of θ derived in part iii(a).                (2)\n\n             c) Add another line to overlay the posterior distribution of θ derived in part iii(b).            (2)\n\n             d) Place the final graph and comment on it.                                                       (3)\n\n       Hint: useful function : lines",
          "topic": null
        }
      ],
      "solution": "i) #m1x0\n        i   = E[X] and                                                                                      (0.5)\n       #m2x0=E[x^2]\n        .           and var = E[x^2] - E[x]^2 >> m2x0 = var + E[x]^2                                        (0.5)\n\n         prior_mean=60\n         prior_sd=5\n                                                                                                       Page 7 of 9\n\fIAI                                                                                             CS1B-0523\n\n         m1x0= prior_mean                                                                              (0.5)\n         m2x0 = prior_sd^2 + prior_mean^2                                                                (1)\n         > m1x0                                                                                        (0.5)\n         [1] 60\n         > m2x0\n         [1] 3625                                                                                     (0.5)\n                                                                                                    [Max 3]\n\n ii) .\n    a) #a theta follows N(prior mean,prior variance)\n       #. Random Variable X follows N(theta,variance)\n         # posterior distribution of theta follows Normal with                                         (0.5)\n         # post mean = (n*sample mean/variance +prior mean/prior variance)/                            (1.5)\n         #              (n/variance + 1/prior variance)\n         #post variance = 1/(n/variance + 1/prior variance)                                              (1)\n         Max 2\n         n<-5\n         sample_mean<- 340/5\n         sdev<-20\n\n         post_mean = (n*sample_mean/sdev^2 + prior_mean/prior_sd^2)/(n/sdev^2 + 1/prior_sd^2)            (2)\n         post_var= 1/(n/sdev^2 + 1/prior_sd^2)                                                           (1)\n\n         > post_mean                                                                                   (0.5)\n         [1] 61.90476\n         > post_var                                                                                    (0.5)\n         [1] 19.04762\n         > sqrt(post_var)                                                                              (0.5)\n         [1] 4.364358\n                                                                                                    [Max 6]\n\n      b) sample2_mean<-3400/50\n         b                                                                                             (0.5)\n         n2<-50\n         .\n         post2_mean = (n2*sample2_mean/sdev^2 + prior_mean/prior_sd^2)/(n2/sdev^2 + 1/prior_sd^2)      (1.5)\n\n         post2_var= 1/(n2/sdev^2 + 1/prior_sd^2)                                                       (0.5)\n         > post2_mean                                                                                  (0.5)\n         [1] 66.06061\n         > post2_var                                                                                   (0.5)\n         [1] 6.060606\n         > sqrt(post2_var)\n         [1] 2.46183\n                                                                                                    [Max 3]\niii) .\n    a) x<-60+seq(-3,3,by=0.2)*5\n       a\n       y<-dnorm(x,mean=60,sd=5)\n       .\n         plot(x,y,ylim=c(0,.2))                                                                          [1]\n\n      b) px1<-post_mean+seq(-3,3,by=0.2)*sqrt(post_var)\n         b                                                                                             (0.5)\n         py1<-dnorm(px1,mean=post_mean,sd=sqrt(post_var))\n         .                                                                                             (0.5)\n         lines(px1,py1)                                                                                  (1)\n\n      c) px2<-post2_mean+seq(-3,3,by=0.2)*sqrt(post2_var)\n         c                                                                                             (0.5)\n         py2<-dnorm(px2,mean=post2_mean,sd=sqrt(post2_var))\n         .                                                                                             (0.5)\n         lines(px2,py2)                                                                                  (1)\n                                                                                                  Page 8 of 9\n\fIAI                                                                                                      CS1B-0523\n\n      d) d\n         .\n\n        The posterior distribution with sample size =5 is close to prior distribution. There is slight shift to     (1)\n        mean towards sample mean and similar dispersion.\n        When the sample size increased, the posterior distribution moves towards sample mean and                    (1)\n        dispersion.\n        More weight is given to sample where sample is big. Further, the variation reduced with larger              (1)\n        sample size.\n                                                                                                              [Max 3]\n\n                                               *****************\n\n                                                                                                             Page 9 of 9",
      "has_math": true,
      "is_r_task": true,
      "session": "2023-05",
      "subject": "CS1B",
      "source_qp": "raw/CS1B_2023-05_QP.pdf",
      "source_sol": "raw/CS1B_2023-05_SOL.pdf"
    },
    {
      "q_num": 1,
      "marks": 26,
      "topic": "inference",
      "subtopics": [],
      "stem": "",
      "parts": [
        {
          "label": "i",
          "marks": 5,
          "text": "Hotel room prices in Mumbai are normally distributed with a mean of Rs. 5400 and\n            standard deviation of Rs. 900, whereas prices in Kolkata are normally distributed with\n            a mean of Rs. 3600 and standard deviation of Rs. 1500. Compute the probability that a\n            hotel room price in Mumbai is atleast twice the price in Kolkata. Note that the hotel\n            prices in Mumbai and Kolkata are independent of each other.\n\n              Store mean and sd values as x.mean, x.sd, y.mean,y.sd,etc.\n              Hint: p**** is a useful function to ascertain CDF.                                          (5)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 1234,
          "text": "a) Show that the population mean and standard deviation of “differences in hotel prices\n                 between Mumbai and Kolkata” (i.e., Mumbai hotel prices – Kolkata hotel Prices)\n                 are 1800 and 1749.86.                                                                    (3)\n\n              b) Generate a sample of size 50 for the ” differences in hotel prices between Mumbai\n                 and Kolkata”.                                                                            (2)\n\n              c) Draw qqplot and qqline for the sample generated and comment on the results.              (4)\n\n        Store population mean and standard deviation of differences as dif.mean and dif.sd\n        Make sure to set the seed value to 1234 i.e use set.seed(1234) before generating the sample.\n        Store the sample as dif.sample\n        Don’t paste the sample",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 3,
          "text": "Test using 5% significance level whether the mean of “difference in hotel prices” is less\n             than 1375 based on the sample generated in part (ii).\n\n              a) Perform z-test and comment.                                                              (5)\n\n              b) Perform t-test (assuming population is not known) and comment.                           (3)\n\n        Please ensure to include test conclusion.",
          "topic": null
        },
        {
          "label": "iv",
          "marks": 3,
          "text": "a) Generate another sample but with size=1000 and store as dif.sample2.                     (1)\n\n              b) Draw qqplot and qqline for the sample generated in (iv.a) and comment on the\n                 results.                                                                                 (3)",
          "topic": null
        }
      ],
      "solution": "i)         > #X~N(5400,900^2)\n              > x.mean=5400\n              > x.sd=900\n              > #Y~N(3600,1500^2)\n              > y.mean= 3600\n              > y.sd=1500\n              > #P(X>2Y) = P(X-2Y>0)\n              > # X-2Y ~ N(x.mean - 2y.mean, x.sd^2+4y.sd^2)\n              >\n              > f.mean<-x.mean - 2*y.mean\n              > f.var <- x.sd^2+4*y.sd^2\n              >\n              > 1-pnorm(0,mean=f.mean,sd=sqrt(f.var))\n              [1] 0.2827485\n                                                                                                                  [Max 5]\n\n  ii)\n  a)          > > dif.mean<-x.mean - y.mean\n              > dif.mean\n              [1] 1800\n              >\n              > dif.sd <-sqrt(x.sd^2+y.sd^2)\n              > dif.sd\n              [1] 1749.286                                                                                            [3]\n\n  b)          > set.seed(1234)\n              > dif.sample<-rnorm(50,dif.mean,dif.sd)                                                                 [2]\n  c)          qnorm(dif.sample)\n              qqline(dif.sample)\n\n              Above results indicates that despite sample generated from normal distribution it doesn’t seem to       [4]\n\n                                                                                                        Page 2 of 8\n\fIAI                                                                                                 CS1B-1123\n\n       align with normality probably due to low sample size and high variability.\n\niii)\n a)    > sample.mean<-mean(dif.sample)\n       > sample.mean\n       [1] 1007.481\n       > z<- (sample.mean - 1375)/(dif.sd/sqrt(50))\n       > pnorm(z)\n       [1] 0.06869143\n\n       Since p-value is >0.05 we cannot reject the null hypothesis and thus, don’t have sufficient evidence to\n       say that mean is less than 1375.                                                                          [Max 5]\n\n b)    t.test(dif.sample, mu=1375, alternative = \"less\")\n                One Sample t-test\n\n       data: dif.sample\n       t = -1.6786, df = 49, p-value = 0.0498\n       alternative hypothesis: true mean is less than 1375\n       95 percent confidence interval:\n           -Inf 1374.558\n       sample estimates:\n       mean of x\n        1007.481\n\n       Since p-value is < 0.05 we can reject the null hypothesis and can imply that mean is less than 1375.          [3]\niv)\n a)    dif.sample2<-rnorm(1000,dif.mean,dif.sd)                                                                      [1]\n\n b)    qqnorm(dif.sample2)\n       qqline(dif.sample2)\n\n       With larger sample size, it indicates normality.                                                              [3]\n\n                                                                                                       Page 3 of 8\n\f IAI                                                                                                   CS1B-1123",
      "has_math": false,
      "is_r_task": true,
      "session": "2023-11",
      "subject": "CS1B",
      "source_qp": "raw/CS1B_2023-11_QP.pdf",
      "source_sol": "raw/CS1B_2023-11_SOL.pdf"
    },
    {
      "q_num": 2,
      "marks": 34,
      "topic": "inference",
      "subtopics": [
        "regression_glm"
      ],
      "stem": "You are a Research Analyst and working on a short term assignment. Your professor\n        provided dance data (dance.csv) containing scores of 20 contestants of a famous dance show\n        and asked you to perform the following tasks.",
      "parts": [
        {
          "label": "i",
          "marks": 2,
          "text": "Using read.csv load the data and use head command to view the first few rows of the           (2)\n            data.\n            No need to paste the output of head command.\n\n        dance.csv contains following fields:\n              -   Judges: It indicates the score provided by judges\n              -   Audience: It indicates the sum of scores provided by audience\n\n          -     Final: It indicates the final score of the contestant",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 4,
          "text": "Plot a scatterplot for each pair of data.                                                      (4)\n          Make sure to paste the scatterplot in your answer scripts.",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 3,
          "text": "Using scatterplot of part (ii), comment on the relationship between the pairs of data.        (3)",
          "topic": null
        },
        {
          "label": "iv",
          "marks": 3,
          "text": "Fit a multiple linear regression model with\n          -     Final score as response variable and\n          -     Judges and Audience score as explanatory variables.\n\n          a) Store the model as m1.                                                                      (2)\n\n          b) Show summary of the output and also write the equation of the fitted model.                 (4)\n\n          c) Compute confidence interval of all parameters.                                              (3)\n\n          d) Comment on the significance of the each explanatory variable either referring to\n             above confidence interval or other statistics.                                              (3)\n\n      You published the report showing the above results and reference data on your research\n      website. After few months, a journalist accused the dance show of ignoring audience scores\n      and cited your report.\n\n      The sponsors of dance show stated that incomplete data is used for analysis as audience\n      score depends upon the number of audiences provided scores. They shared supplement data\n      providing the number of audiences for each participant.",
          "topic": null
        },
        {
          "label": "v",
          "marks": 2,
          "text": "Using the following code to load the number of audiences for 20 contestants\n          audience.count<-c(110,100,90,120,100,100,100,100,110,110,100,\n                    100,110,90,100,110,120,120,100,100)\n          Also, verify total count is equal to 2090.                                                     (2)",
          "topic": null
        },
        {
          "label": "vi",
          "marks": 2,
          "text": "Compute the new audience score by dividing the sum of score provided by audience              (2)\n              (“Audience”) with audience count (“audience.count”).\n           Store as Audience2 and make sure to attach to the dance data.\n           Please don’t paste Audience2 values.",
          "topic": null
        },
        {
          "label": "vii",
          "marks": 3,
          "text": "Perform correlation test to check whether any correlation exits between final score and\n           new audience score.                                                                           (3)",
          "topic": null
        },
        {
          "label": "viii",
          "marks": 3,
          "text": "Fit a new multiple linear regression model with\n          -     Final score as response variable and\n          -     Judges and New Audience score as explanatory variables.\n\n      Store this model as m2 and show the summary of the output.                                         (3)",
          "topic": null
        },
        {
          "label": "ix",
          "marks": 3,
          "text": "Using a suitable statistic from part (iv.b) and (viii). outputs, compare models m1 and    (3)\n            m2 and suggest which model is better.\n            Please write the figures of the statistic while answering.",
          "topic": null
        }
      ],
      "solution": "i)         dance<-read.csv(\"dance.csv\")\n              head(dance)                                                                                                [2]\n\n  ii)         plot(dance)\n\n  iii)        Judges score and Final score seems to have a linear relationship\n              Audience score is quite scattered and doesn't show any strong linear relationship with either Judges\n              or Final score                                                                                             [3]\n\n  iv)\n   a)         m1<-lm(Final~Judges+Audience,data=dance)                                                                   [2]\n\n  b)          > summary(m1)\n\n              Call:\n              lm(formula = Final ~ Judges + Audience, data = dance)\n\n              Residuals:\n                 Min     1Q Median    3Q Max\n              -5.2783 -0.7971 0.1841 1.6334 3.7680\n\n              Coefficients:\n                      Estimate Std. Error t value Pr(>|t|)\n              (Intercept) 10.617827 15.129865 0.702 0.492\n              Judges      0.720273 0.128870 5.589 3.26e-05 ***\n              Audience 0.001449 0.001294 1.120 0.278\n\n                                                                                                           Page 4 of 8\n\fIAI                                                                                      CS1B-1123\n\n        ---\n        Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1\n\n        Residual standard error: 2.604 on 17 degrees of freedom\n        Multiple R-squared: 0.6601, Adjusted R-squared: 0.6201\n        F-statistic: 16.51 on 2 and 17 DF, p-value: 0.0001038\n\n        Final = 10.617827 + 0.720273 * Judges + 0.001449 *Audience                                       [4]\n\n  c)    > confint(m1)\n                    2.5 % 97.5 %\n        (Intercept) -21.3033977 42.5390512\n        Judges       0.4483811 0.9921649\n        Audience -0.0012812 0.0041793\n        >\n  d)    > #Judges score seems significant since confidence interval doesn't contain 0.\n        Sum of audience score doesn’t seem significant sine it contains 0.\n        Alternate:\n        > #P-value is <.01 for judges score showing significance.                                        [3]\n\n  v)    audience.count<-c(110,100,90,120,100,100,100,100,110,110,100,\n                   100,110,90,100,110,120,120,100,100)\n        sum(audience.count)\n        [1] 2090                                                                                         [2]\n\n vi)    dance$Audience2<-dance$Audience/audience.count                                                   [2]\n\nvii)    > cor.test(dance$Final,dance$Audience2)\n\n               Pearson's product-moment correlation\n\n        data: dance$Final and dance$Audience2\n        t = 4.9045, df = 18, p-value = 0.0001142\n        alternative hypothesis: true correlation is not equal to 0\n        95 percent confidence interval:\n         0.4716109 0.8982071\n        sample estimates:\n            cor\n        0.7562948\n        > #correlation between audience score and final score is quite high\n\nviii)   m2<-lm(Final~Judges+Audience2,data=dance)\n        > summary(m2)\n\n        Call:\n        lm(formula = Final ~ Judges + Audience2, data = dance)\n\n        Residuals:\n           Min     1Q Median    3Q Max\n        -1.9408 -1.0269 0.1129 1.0466 1.6075\n\n        Coefficients:\n                                                                                           Page 5 of 8\n\f    IAI                                                                                            CS1B-1123\n\n                      Estimate Std. Error t value Pr(>|t|)\n              (Intercept) 1.20323 5.84491 0.206 0.839\n              Judges      0.56604 0.06545 8.648 1.24e-07 ***\n              Audience2 0.42002 0.05366 7.827 4.91e-07 ***\n              ---\n              Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1\n\n              Residual standard error: 1.258 on 17 degrees of freedom\n              Multiple R-squared: 0.9207, Adjusted R-squared: 0.9114\n              F-statistic: 98.73 on 2 and 17 DF, p-value: 4.39e-10                                                    [3]\n\n     ix)      #Adjusted R-square =0.9114 for model2 vs 0.6201 for model1.\n              #This indicates model 2 is better\n\n              Alternate:\n              # R square can be used                                                                                  [3]",
      "has_math": false,
      "is_r_task": true,
      "session": "2023-11",
      "subject": "CS1B",
      "source_qp": "raw/CS1B_2023-11_QP.pdf",
      "source_sol": "raw/CS1B_2023-11_SOL.pdf"
    },
    {
      "q_num": 3,
      "marks": 20,
      "topic": "regression_glm",
      "subtopics": [],
      "stem": "An analyst fitted various models on the data containing claims information of 20 policies\n        and shared the output to conduct analysis.\n\n        Output 1:\n\n        Call:\n        glm(formula = Claim ~ 1, family = poisson(lin = \"log\"), data = q3)\n\n        Deviance Residuals:\n           Min    1Q Median     3Q Max\n        -1.8439 -0.8975 -0.1791 0.3925 2.5561\n\n        Coefficients:\n                Estimate Std. Error z value Pr(>|z|)\n        (Intercept) 0.5306 0.1715 3.094 0.00197 **\n        ---\n        Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1\n\n        (Dispersion parameter for poisson family taken to be 1)\n\n          Null deviance: 30.147 on 19 degrees of freedom\n        Residual deviance: 30.147 on 19 degrees of freedom\n        AIC: 71.114\n\n        Number of Fisher Scoring iterations: 5",
      "parts": [
        {
          "label": "i",
          "marks": 2,
          "text": "Write the equation of the fitted model using Output 1 as above and specify which\n               distribution is used to model the response variable.                                   (2)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 2,
          "text": "Using output 1, show that the sample mean of the response variable is 1.7 (when\n               rounded to one decimal place).                                                         (2)\n\n        Output 2:\n        Call:\n        glm(formula = Claim ~ Gender + Health - 1, family = poisson(lin = \"log\"),\n          data = q3)\n\n        Deviance Residuals:\n           Min     1Q Median       3Q    Max\n        -1.22474 -0.81754 -0.07119 0.27453 1.44149\n\n        Coefficients:\n                   Estimate Std. Error z value Pr(>|z|)\n        GenderF          1.1394 0.3490 3.265 0.001096 **\n        GenderM           1.1394 0.2216 5.143 2.71e-07 ***\n        HealthNonDiabetic -1.4271 0.3939 -3.623 0.000291 ***\n        ---\n        Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1\n\n        (Dispersion parameter for poisson family taken to be 1)\n\n          Null deviance: 38.229 on 20 degrees of freedom\n        Residual deviance: 14.436 on 17 degrees of freedom\n        AIC: 59.403\n\n        Number of Fisher Scoring iterations: 5",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 4,
          "text": "Write the equation of the fitted model using Output 2.                                    (4)\n                   Make sure to define the explanatory (categorical) variables.",
          "topic": null
        },
        {
          "label": "iv",
          "marks": 2,
          "text": "Output 2 doesn’t have an intercept as a coefficient unlike output1. Please provide the\n                   reason for the same by comparing glm R formulas.                                          (2)",
          "topic": null
        },
        {
          "label": "v",
          "marks": 2,
          "text": "Compare output 1 and output 2, and suggest which model is better using suitable\n                   statistics.                                                                               (2)",
          "topic": null
        },
        {
          "label": "vi",
          "marks": 3,
          "text": "Compute log likelihood of the model given in output 2.                                    (3)",
          "topic": null
        },
        {
          "label": "vii",
          "marks": 5,
          "text": "Claim follows Poisson distribution. Test for the parameter to be equal to 1.5 at 1%\n                   level of significance.                                                                    (5)\n                   Hint: Use outputs and various sub-parts to determine x for poisson test.",
          "topic": null
        }
      ],
      "solution": "i)      log(y) = 0.5306\n              poisson distribution is used to model response variable.                                                [2]\n\n      ii)     > Claim.mean<-round(exp(.5306),1)\n              > Claim.mean\n              [1] 1.7                                                                                                 [2]\n\n     iii)     log(y) = 1.1394 * GenderF + 1.1394 * GenderM -1.4271 *HealthNonDiabetic\n              where GenderF =1 if Gender = F else 0\n              GenderM =1 if Gender = M else 0\n              HealthNonDiabetic =1 if Health= NonDiabetic else 0\n                                                                                                                  [Max 4]\n\n     iv)      For model 2, “-1” is used in Glm R formula to not take the intercepts while fitting the model.\n              Thus, no intercept exits.                                                                               [2]\n\n      v)      AIC (=59.4) of model 2 which is lower than AIC (71.1) of model 1. This indicates that model\n              2 is better.                                                                                            [2]\n\n     vi)      #AIC = - 2LogL(Model) + 2*Parameters\n              > #LogL(Model) = Parameters - AIC/2\n              > aic<-59.403\n              > par<-3\n              >\n              > L<- par - aic/2\n              >L\n              [1] -26.7015\n                                                                                                                  [Max 3]\n\n    vii)      #Total claims = Mean claims * Total policies\n              > x<-Claim.mean*20\n              > poisson.test(x=X,T=20,r=1.5,conf.level = 0.99)\n\n                     Exact Poisson test\n\n                                                                                                      Page 6 of 8\n\f    IAI                                                                                  CS1B-1123\n\n              data: X time base: 20\n              number of events = 34, time base = 20, p-value = 0.4639\n              alternative hypothesis: true event rate is not equal to 1.5\n              99 percent confidence interval:\n               1.042837 2.605372\n              sample estimates:\n              event rate\n                   1.7\n\n              >\n              > # We cannot reject the null hypothesis that parameter is equal to 1.5.               [Max 5]",
      "has_math": false,
      "is_r_task": true,
      "session": "2023-11",
      "subject": "CS1B",
      "source_qp": "raw/CS1B_2023-11_QP.pdf",
      "source_sol": "raw/CS1B_2023-11_SOL.pdf"
    },
    {
      "q_num": 4,
      "marks": 20,
      "topic": "inference",
      "subtopics": [],
      "stem": "The table below shows the total claim number (cancellations) per year, Xij, for 4 travel\n        companies over last 4 years.\n\n                                                   Years,j\n\n                                     1         2             3     4\n           Companies, i\n\n                          Make     455        458        587      531\n             Travel\n\n                          Ease     251        322        292      340\n                           Go      309        246        217      120\n                          Clear    400        426        470      547",
      "parts": [
        {
          "label": "i",
          "marks": 2,
          "text": "Using Empirical Bayes Credibility Theory (EBCT) Model 1, compute the following\n\n               a) Copy the below code to load the data:                                                      (1)\n                   q4<-matrix(c(455,251,309,400,\n                    458,322,246,426,\n                    587,292,217,470,\n                    531,340,120,547),\n                  ncol=4,nrow=4)\n\n               b) E[m(θ)]                                                                                    (2)\n\n               c) E[s2(θ)]                                                                                   (2)\n\n               d) Var[m(θ)]                                                                                  (3)\n\n               e) Z                                                                                          (2)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 5,
          "text": "Using part (i), calculate the expected claim number for Go and Clear.                   (3)\n\n      iii)   What additional information is required to use EBCT Model 2.                            (2)\n\n      iv)    Travel company “Ease” launched a membership program last year providing full\n             refund on cancellations. Number of cancellations believed to follow binomial\n             distribution with parameters n=3 and 0.20.\n             Number of cancellations in last year on 150 memberships are as follows:\n\n                Cancellations          0             1                2                3\n                 Members              61            71               15                3\n\n      Carry out goodness of fit test for the binomial model specified for number of cancellations\n      on each membership.                                                                            (5)",
          "topic": null
        }
      ],
      "solution": "i)     q4<-matrix(c(455,251,309,400,\n                    458,322,246,426,\n                    587,292,217,470,\n                    531,340,120,547),\n                   ncol=4,nrow=4)\n\n              n<-ncol(q4)\n              m<-mean(rowMeans(q4))\n              s<-mean(apply(q4,1,var))\n              v<-var(rowMeans(q4)) - mean(apply(q4,1,var))/n\n              Z<- n/(n+s/v)\n\n              n\n              [1] 4\n              >m\n              [1] 373.1875\n              >s\n              [1] 3967.854\n              >v\n              [1] 16843.22\n              >Z\n              [1] 0.9443816\n                                                                                               [1+2+2+3+2]\n\n      ii)     Z*rowMeans(q4[3:4,])+(1-Z)*m\n              [1] 231.3532 455.8799                                                                      [3]\n\n     iii)     Risk Volumes are required to apply EBCT2                                                   [2]\n\n     iv)      obs<-c(61,71,15,3)\n              >\n              > #Combine 2 and 3 since value of 3 is less than 5\n              > obs.comb<-c(61,71,15 + 3)\n              >\n              > p<-0.2\n              > exp<-dbinom(c(0:1),3,p)\n              > exp[3]<-1-pbinom(1,3,p)\n              > sum(exp)\n              [1] 1\n              >\n\n                                                                                           Page 7 of 8\n\fIAI                                                                                               CS1B-1123\n\n      > chisq.test(x=obs.comb,p=exp)\n\n              Chi-squared test for given probabilities\n\n      data: obs.comb\n      X-squared = 6.7371, df = 2, p-value = 0.03444\n      Since p-value <0.5, we have sufficient evidence to reject the null hypothesis that cancellation follows\n      bin(3,0.2)\n                                                                                                               [Max 5]\n\n                                         ***********************\n\n                                                                                                     Page 8 of 8",
      "has_math": true,
      "is_r_task": true,
      "session": "2023-11",
      "subject": "CS1B",
      "source_qp": "raw/CS1B_2023-11_QP.pdf",
      "source_sol": "raw/CS1B_2023-11_SOL.pdf"
    },
    {
      "q_num": 1,
      "marks": 20,
      "topic": "distributions",
      "subtopics": [],
      "stem": "Consider a random variable X which represents the project completion time for a process\n        automation project for a FMCG company (in years). Three scenarios have been contemplated\n        for completion of the project which are as follows:\n\n         Scenario                                          Probability Distribution\n         Scenario 1                                   :    X has a beta distribution with parameters α =5, β=1;\n         Scenario 2                                   :    X has a beta distribution with parameters α =1, β=5;\n         Scenario 3                                   :    X has a beta distribution with parameters α =3, β=3.",
      "parts": [
        {
          "label": "i",
          "marks": 2,
          "text": "Calculate the following probabilities for X under Scenario 1:\n\n                            a) Use pbeta function to calculate P(0.2 < X < 0.8);                                                                                                                                              (2)\n\n                            b) Use qbeta to find x such that P(X > x) = 0.65.                                                                                                                                                 (2)\n\n         ii) Use the following code to calculate the coefficient of skewness of X under all the three\n             scenarios:\n\n                            Skew = 2*((β-α)/(α+β+2))* sqrt((α+β+1)/(α*β))                                                                                                                                                     (3)\n\n        iii) Under all the three scenarios, write a code to simulate a sample of 12,000 values from a\n             Beta distribution. Use the command set.seed(421967) to initialize the random number\n             generator, before you start the simulation. Simulate this sample by defining vectors “x1”,\n             “x2” and “x3” for Scenario1, Scenario2 and Scenario3 and write code to draw histogram\n             for each of the scenarios.\n\n                            You are NOT required to execute the code or print the result of these vectors or reproduce\n                            the histograms.                                                                                                                                                                                   (3)\n\n        Histograms for these simulated samples have been given below:\n\n                            Scenario 1                                                             Scenario 2                                                              Scenario 3\n                       X1 ~ Beta (α =5, β=1)                                                  X2 ~ Beta (α =1, β=5)                                                   X3 ~ Beta (α =3, β=3)\n                       Histogram of 12000 values from Beta (5,1) Distribution                 Histogram of 12000 values from Beta (1,5) Distribution                 Histogram of 12000 values from Beta (3,3) Distribution\n                     2500\n\n                                                                                            2500\n\n                                                                                                                                                                   1000\n         Frequency\n\n                                                                                                                                                       Frequency\n                                                                                Frequency\n                     1500\n\n                                                                                            1500\n\n                                                                                                                                                                   600\n                     500\n\n                                                                                            500\n\n                                                                                                                                                                   200\n                     0\n\n                                                                                                                                                                   0\n                                                                                            0\n\n                             0.2       0.4        0.6       0.8        1.0                         0.0     0.2        0.4        0.6       0.8                            0.0    0.2      0.4        0.6    0.8      1.0\n\n                                                 x1                                                                     x2                                                                      x3\n\n        iv) Comment on the shape of the histograms plotted above by taking into consideration\n            coefficient of skewness calculated in part (ii).                                                                                                                                                                  (3)\n\n         v) Write a code to perform 1,200 repetitions of the simulations in part (iii) for all the three\n            scenarios. You should compute and store the value of the mean of the sample (x1_bar),\n            (x2_bar) and (x3_bar) for each repetition. Use the same command set.seed(421967) to\n            initialize the random number generator, before you start the simulations.\n\n                            You are NOT required to execute the code or print the output of these simulations or\n                            histogram of these simulations.                                                                                                                                                                   (5)\n\n        Histograms of sample means for the three scenarios have been given below:\n\n                           Scenario 1: x1_bar                                                        Scenario 2: x2_bar                                                 Scenario 3: x3_bar\n                             Histogram of 1200 simulations of x1_bar                                     Histogram of 1200 simulations of x2_bar                        Histogram of 1200 simulations of x3_bar\n\n                                                                                                                                                                 250\n                                                                                           150\n                     150\n\n                                                                                                                                                                 200\n                                                                               Frequency\n         Frequency\n\n                                                                                                                                                     Frequency\n\n                                                                                                                                                                 150\n                                                                                           100\n                     100\n\n                                                                                                                                                                 100\n                                                                                           50\n                     50\n\n                                                                                                                                                                 50\n                                                                                           0\n                     0\n\n                                                                                                                                                                 0\n                                                                                                 0.162        0.164      0.166     0.168     0.170\n                           0.830     0.832     0.834     0.836         0.838                                                                                            0.494   0.496   0.498   0.500   0.502    0.504\n                                                                                                                          x2_bar\n                                             x1_bar                                                                                                                                        x3_bar\n\n        vi) Comment on the shape of the histograms, by referring to the central limit theorem. Also,\n            compare and contrast with your observations in part (iv).                                                                                                                                                     (2)",
          "topic": null
        }
      ],
      "solution": "(i)\na) > pbeta(0.8,5,1) - pbeta(0.2,5,1)                                         (1)\n         [1] 0.32736                                                         (1)\n\nb) > qbeta(0.65,5,1,lower= FALSE) or Alternate: qbeta(0.35,5,1\n)\n          [1] 0.8106131                                                      (1)\n\n(ii)\n> a=c(5,1,3)\n> b=c(1,5,3)\n> skew=2*((b-a)/(a+b+2))*sqrt((a+b+1)/(a*b))                                (2)\n> skew\n[1] -1.183216   1.183216   0.000000                                         (1)\n\nOr Alternatively\n\n> skew1=2*((1-5)/(5+1+2))*sqrt((5+1+1)/5)\n> skew1\n[1] -1.183216\n\n> skew2=2*((5-1)/(1+5+2))*sqrt((5+1+1)/5)\n> skew2\n[1] 1.183216\n\n> skew3=2*((3-3)/(3+3+2))*sqrt((3+3+1)/9)\n> skew3\n[1] 0\n\n(iii)\n> set.seed(421967)                                                         (0.5)\n> x1=rbeta(12000,5,1)                                                       (1)\n> hist(x1)                                                                (0.5)\n> x2=rbeta(12000,1,5)\n> hist(x2)                                                                (0.5)\n> x3=rbeta(12000,3,3)\n> hist(x3)                                                                (0.5)\n\n(iv)\n         As Alpha is greater than 1 and Beta is equal to 1, the histogram is he\navily negatively skewed and strictly increasing as evident from the result obt\nained in (ii) above and from the shape of the graph.\n         As Alpha is equal to 1 and Beta is greater than 1, the histogram is he\navily positively skewed and strictly decreasing as evident from the result obt\nained in (ii) above and from the shape of the graph.\n\n                                                                     Page 2 of 11\n\fIAI                                                                  CS1B-0524\n\n          As both the parameters alpha and beta are equal, the graph is roughly\nsymmetrical as evident from the graph and the value of the skewness obtained\nin (ii) above.\n(v)\n> set.seed(421967)\n> x1_bar <-replicate (1200, mean(rbeta (12000,5,1)))\n\nOr alternatively\n> set.seed(421967)\n> x1_bar=rep(0,1200)\n> for(i in 1:1200){x1<-rbeta(12000,5,1);x1_bar[i]<-mean(x1)}\n\n> set.seed(421967)\n> x2_bar <-replicate (1200, mean (rbeta (12000,1,5)))\n\nOr alternatively\n> set.seed(421967)\n> x2_bar=rep(0,1200)\n> for(i in 1:1200){x2<-rbeta(12000,1,5);x2_bar[i]<-mean(x2)}\n> set.seed(421967)\n> x3_bar <-replicate (1200, mean(rbeta (12000,3,3)))\n\nOr alternatively\n> set.seed(421967)\n> x3_bar=rep(0,1200)\n> for(i in 1:1200){x3<-rbeta(12000,1,5);x3_bar[i]<-mean(x3)}\n\n(vi)The distribution of sample mean is roughly symmetrical in all the three ca\nses irrespective of the values of the shape parameters (alpha and beta).These\nshape parameters(alpha and beta) do not significantly affect     the sample mean\nof large sample size, which is in line with the central limit theorem. Irrespe\nctive of the population distribution of the random variable from which the sam\nple is selected, for a large sample size the distribution of the sample means\nis approximately normal.",
      "has_math": true,
      "is_r_task": true,
      "session": "2024-05",
      "subject": "CS1B",
      "source_qp": "raw/CS1B_2024-05_QP.pdf",
      "source_sol": "raw/CS1B_2024-05_SOL.pdf"
    },
    {
      "q_num": 2,
      "marks": 30,
      "topic": "distributions",
      "subtopics": [
        "inference",
        "data_analysis"
      ],
      "stem": "As per a recent research, the maximum systolic blood pressure for a person is related to age and\n        can be expressed in terms of the following equation:\n\n        maximum systolic blood pressure = 100 + age (in years)",
      "parts": [
        {
          "label": "i",
          "marks": 2,
          "text": "If we decide to fit a regression line with maximum systolic blood pressure as the response\n              variable (Y) and age as the explanatory variable (X), what should be the values of the\n              regression coefficients α and β in light of the above equation?                                                                                                                                             (2)\n\n        Suppose this is to be empirically proven and 20 people of varying ages are tested for their\n        maximum systolic blood pressure. The following data has been collected:\n\n                             Age (X)                             Maximum Systolic                                                    Age (X)                           Maximum Systolic\n                                                                 Blood Pressure (Y)                                                                                    Blood Pressure (Y)\n                                   28                                   132                                                                60                                 149\n                                   37                                   140                                                                55                                 154\n                                   41                                   155                                                                29                                 117\n                                   52                                   160                                                                43                                 146\n                                   57                                   167                                                                36                                 142\n                                   49                                   148                                                                50                                 168\n                                   38                                   128                                                                34                                 144\n                                   25                                   131                                                                40                                 156\n                                   23                                   118                                                                26                                 114\n                                   48                                   139                                                                33                                 133\n\n        The data of age in years (X) and corresponding maximum systolic blood pressure (Y) can be\n        entered in R using the below code:\n\n        x=c(28, 37, 41, 52, 57, 49, 38, 25, 23, 48, 60, 55, 29, 43, 36, 50, 34, 40, 26, 33)\n\n        y=c(132, 140, 155, 160, 167, 148, 128, 131, 118, 139, 149, 154, 117, 146, 142, 168, 144, 156,\n        114, 133)\n\n        Enter the data in R by copying the above code.",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 3,
          "text": "Prepare a scatter plot of the data and briefly comment on the same.                                                                                                                                          (3)",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 4,
          "text": "Plot the fitted line for regression of Y on X.                                                      (4)",
          "topic": null
        },
        {
          "label": "iv",
          "marks": 3,
          "text": "Using anova function check whether the slope parameter is significant.                               (3)",
          "topic": null
        },
        {
          "label": "v",
          "marks": 3,
          "text": "Obtain a summary for the linear regression model fitted in part (iv) and clearly state the\n             estimates for the values of the coefficients α and β.                                                (3)",
          "topic": null
        },
        {
          "label": "vi",
          "marks": 2,
          "text": "Test the veracity of the above equation as promulgated in the recent research, by referring\n             to your answers in part (i) and part (v).                                                            (2)",
          "topic": null
        },
        {
          "label": "vii",
          "marks": 3,
          "text": "Calculate the sample Pearson correlation coefficient, sample Kendall correlation\n             coefficient and sample Spearman correlation coefficient.                                             (3)",
          "topic": null
        },
        {
          "label": "viii",
          "marks": 4,
          "text": "Comment on each method of calculating correlation coefficient and also comment on the\n              sample correlation coefficients obtained in part (vii).                                             (4)\n\n        One of your colleagues says that there exists perfect positive correlation between age and\n        maximum systolic blood pressure.",
          "topic": null
        },
        {
          "label": "ix",
          "marks": 6,
          "text": "Perform a hypothesis test to test the above statement at 5% level of significance. You\n             should state the null and alternate hypotheses, report the p-value of the test and arrive at a\n             clear conclusion.(Hint: Use Pearson Method)                                                          (6)",
          "topic": null
        }
      ],
      "solution": "(i)\nThe given equation is\n        Maximum Systolic Blood Pressure = 100 + Age (in years)\n         It can be written as\n                   y = 100 + x\n                   y = 100 + 1*(x)\n                   y = α + βx\n                   𝛼 = 100 𝑎𝑛𝑑 𝛽 = 1\n\n                                                                      Page 3 of 11\n\fIAI                                                                CS1B-0524\n\n(ii)\n\n> x=c(28, 37, 41, 52, 57, 49, 38, 25, 23, 48, 60, 55, 29, 43, 36, 50, 34, 40,\n26, 33)\n> y=c(132, 140, 155, 160, 167, 148, 128, 131, 118, 139, 149, 154, 117, 146, 1\n42, 168, 144, 156, 114, 133)\n\n> plot(x,y)\n\nThe age in years(x) and the systolic blood pressure(y) are positively correlat\ned.\n\n(iii)\n> lm.result=lm(y~x)                                                        (1)\n\n> abline(lm(y~x))                                                          (1)\n\n(iv)\n>anova(lm.result)\nAnalysis of Variance Table\nResponse: y\n          Df Sum Sq Mean Sq F value    Pr(>F)\nx           1 3082.9 3082.94 33.591 1.717e-05 ***\nResiduals 18 1652.0    91.78\n---\nSignif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1\n\nFrom the above, it is clear that the slope parameter is significant.\n\n                                                                   Page 4 of 11\n\fIAI                                                                  CS1B-0524\n\n(v)\n>summary(lm.result)                                                          (1)\n\nCall:\nlm(formula = y ~ x)\n\nResiduals:\n    Min     1Q     Median      3Q       Max\n-15.485 -6.504      1.177   5.979    14.846\n\nCoefficients:\n            Estimate Std. Error t value Pr(>|t|)\n(Intercept) 96.4994      8.1460 11.846 6.21e-10 ***\nx             1.1331     0.1955   5.796 1.72e-05 ***\n---\nSignif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1\n\nResidual standard error: 9.58 on 18 degrees of freedom\nMultiple R-squared: 0.6511, Adjusted R-squared: 0.6317\nF-statistic: 33.59 on 1 and 18 DF, p-value: 1.717e-05\nThe value of the estimates of the co-efficient are\n                Alpha(α) = 96.4994    Beta(β) = 1.1331                        (1)\n\n(vi) The values of alpha and beta are expected to be 100 and 1 respectively a\ns per (i).    But empirical test results fetch the values as 96.4994 and 1.1331\nrespectively, which is close to the expected values of 100 and 1 respectively\n. Hence when empirically tested, we find that the claim made by the research\nabout the maximum systolic blood pressure is valid.\n\n(vii)\n> cor(x,y,method=\"pearson\")\n[1] 0.8069094\n\n> cor(x,y,method=\"spearman\")\n[1] 0.8180451\n\n> cor(x,y,method=\"kendall\")\n[1] 0.6105263\n\n(viii)\n        Pearson’s correlation co-efficient measures the strength of the linear r\nelationship between the two variables, whereas Spearman Correlation method mea\nsures the strength of monotonic but not necessarily linearity between two vari\nables.\n        Since Spearman considers the rank than the actual values, the value of t\nhe coefficient is less affected by extreme values/outliers in the data than Pe\narson’s Correlation Coefficient. Hence it is more robust.\n      Kendall’s correlation coefficient is considered to have better statistical\nproperties when      the data set is small and have more tied ranks, though it c\nonsiders the relative values between the data set and not actual values.\n\n                                                                      Page 5 of 11\n\fIAI                                                                  CS1B-0524\n\nGenerally, the value of Kendall’s coefficient is lower than the Spearman’s\nrank coefficient.\n      Based on the sample correlation coefficients calculated in part (vii), we\nconclude Spearman Rank Coefficient > Pearson’s Coefficient > Kendall’s Coeffi\ncient\n(ix)\n𝐻0 : ρ = 1\n𝐻1 : ρ ≠ 1                                                                  (1)\n\n> cor.test(x,y,method=\"pearson\")                                             (1)\n\n          Pearson's product-moment correlation\ndata: x and y\nt = 5.7958, df = 18, p-value = 1.717e-05\nalternative hypothesis: true correlation is not equal to 0\n95 percent confidence interval:\n 0.5667661 0.9206793\nsample estimates:\n      cor\n0.8069094                                                                    (1)\n\np-value is 1.717e-05                                                         (1)\n\n          Since 95% confidence interval (0.5667661, 0.9206793) does not include\nthe value 1, there is sufficient evidence to reject the hypothesis that there\nis perfect correlation between Age and Systolic Blood Pressure though there is\nstrong positive correlation between the age and systolic Blood Pressure.",
      "has_math": true,
      "is_r_task": true,
      "session": "2024-05",
      "subject": "CS1B",
      "source_qp": "raw/CS1B_2024-05_QP.pdf",
      "source_sol": "raw/CS1B_2024-05_SOL.pdf"
    },
    {
      "q_num": 3,
      "marks": 30,
      "topic": "regression_glm",
      "subtopics": [
        "inference"
      ],
      "stem": "Consider a portfolio of fire insurance policies. The data relating to 192 policies which have\n        claimed at least once till now, is given in the file Firepolicies.csv. Confirm through output that the\n        file contains data regarding four variables:\n\n           •   Occupancy:          The claim has arisen in five different Occupancies:\n                                   TM – for Textile Mills\n                                   DW – for Dwellings\n                                   TG – for Transporters’ Godowns\n                                   CS – for Cold Storage Premises\n                                   HG – for Hazardous Goods Storage\n\n           •   Location:           It relates to the loss state:\n                                   M – for Maharashtra\n                                   G – for Gujarat\n                                   T – for Telangana\n                                   K – for Karnataka\n                                   A – for Andhra Pradesh\n\n           •   Claim.size:         It refers to the amount of the claim in INR lakhs.\n\n           •   Claimed:            This is an indicator variable which captures whether there is an\n                                   incidence of claim in the past one year.\n                                   0 – refers to NO claim in the past one year\n                                   1 – refers to ONE claim in the past one year",
      "parts": [
        {
          "label": "i",
          "marks": 1,
          "text": "View the data Firepolicies.csv for 192 entries using read.csv function. Columns\n             Claim.Size, and Claimed have numeric data type, all other columns i.e. Location and\n             Occupancy have character data type.                                                                  (1)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 1,
          "text": "Using the Firepolicies.csv, calculate the proportion of:\n\n            a) Policyholders with no claims in the past one year from the state of Maharahstra                (3)\n\n            b) Policyholders with no claims in the past one year from the state of Gujarat                    (1)",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 5,
          "text": "Test the hypothesis that the proportion of policyholders with no claims in the past one year\n            is equal in both states (Maharashtra and Gujarat) at 5% level of significance. You should\n            state the null and alternate hypotheses, report the p-value of the test and arrive at a clear\n            conclusion.                                                                                       (5)",
          "topic": null
        },
        {
          "label": "iv",
          "marks": 6,
          "text": "Test the hypothesis that there is no significant difference in the average claim size for\n           Textile Mills and Transporters’ Godowns at 5% level of significance. You should state the\n           null and alternate hypotheses, report the p-value of the test and arrive at a clear conclusion.\n           Assume that population variances are equal.                                                        (6)",
          "topic": null
        },
        {
          "label": "v",
          "marks": 2,
          "text": "The number of policies with at least one claim (X) for the fire insurance portfolio is\n           modelled as a random variable with a Binomial distribution X ~ Binomial (n, p).\n\n            An Actuary wishes to fit different Generalized Linear Models (GLMs) to the data,\n            assuming that the number of policies with one submitted claim has a Binomial distribution\n            and the link function of the GLM is the logit function.\n\n            a) Fit a GLM to the data such that p depends on the Location and Claim Size i.e., Loss\n               State and the Claim size and report the summary.                                               (4)\n\n            b) Fit a GLM to the data such that p depends on the Occupancy and Claim Size and report\n               the summary.                                                                                   (2)",
          "topic": null
        },
        {
          "label": "vi",
          "marks": 3,
          "text": "Compare the fit of the models in parts (v)(a) and (v)(b) using Akaike’s Information\n           Criterion (AIC) and comment on which model is preferable.                                          (3)",
          "topic": null
        },
        {
          "label": "vii",
          "marks": 5,
          "text": "Among the four variables as given in the Firepolicies.csv data, which of them are numerical\n            variables and which of them are factor variables? What is the difference between the two\n            in the context of generalised linear models?                                                      (5)",
          "topic": null
        }
      ],
      "solution": "(i) > Firepolicies<-read_csv(\"Firepolicies.csv\")\n\n(ii)\na)\n\n>     Maha<-Firepolicies[Firepolicies$Location==\"M\",]\n> Maha_0_Claims<-Maha[Maha$Claimed == 0,]\n> ProporMaha_0_Claims<-nrow(Maha_0_Claims)/nrow(Maha)\n> ProporMaha_0_Claims\n[1] 0.7346939\nb)\n>     Gujarat<-Firepolicies[Firepolicies$Location==\"G\",]\n>     Gujarat_0_Claims<-Gujarat[Gujarat$Claimed == 0,]\n>     ProporGuj_0_claims <-nrow(Gujarat_0_Claims)/nrow(Gujarat)\n> ProporGuj_0_claims\n[1] 0.5909091\n(iii)\n𝐻0 (Null Hypothesis):\n                                                                     Page 6 of 11\n\fIAI                                                                      CS1B-0524\n\nProportion of No claims in the past one year in Maharashtra is equal       to Propo\nrtion of NO Claims in the past one year in Gujarat\n\n𝐻1 (Alternative Hypothesis):\nProportion of No claims in the past one year in Maharashtra is NOT equal to Pr\noportion of NO Claims in the past one year in Gujarat\n\n> prop.test (c(nrow( Maha_0_Claims),nrow(Gujarat_0_Claims)),c(nrow(Maha),nrow\n(Gujarat)),correct = FALSE)\n\n2-sample test for equality of proportions without continuity correction\n\ndata: c(nrow(Maha_0_Claims), nrow(Gujarat_0_Claims)) out of c(nrow(Maha), nr\now(Gujarat))\nX-squared = 2.1568, df = 1, p-value = 0.1419\nalternative hypothesis: two. sided\n95 percent confidence interval:\n -0.04696637 0.33453595\nsample estimates:\n   prop 1    prop 2\n0.7346939 0.5909091\n\n       p-value is 0.1419                                                         (1)\n\nAt 95% confidence interval (-0.04696637, 0.33453595), which contains “0”, we h\nave insufficient evidence to reject null hypothesis and can conclude that ther\ne is no significant difference between Maharashtra and Gujarat in respect of t\nhe proportion of No claims in the previous year.                           (1)\n(iv)\n𝐻0 (Null Hypothesis):\nPopulation mean of Textile Mills Claims is equal to Population mean of Transp\norters’ Godowns Claims\n\n𝐻1 (Alternative Hypothesis):\nPopulation mean of Textile Mills Claims is NOT equal to Population mean of Tra\nnsporters’ Godowns Claims\n\n> Textile<-Firepolicies[Firepolicies$Occupancy==\"TM\",]\n\n> Transporter<-Firepolicies[Firepolicies$Occupancy==\"TG\",]\n\n>     t.test(Textile$Claim.Size,Transporter$Claim.Size,var.equal=TRUE)\n\n          Two Sample t-test\ndata:     Textile$Claim.Size and Transporter$Claim.Size\nt = 7.877, df = 103, p-value = 3.586e-12\nalternative hypothesis: true difference in means is not equal to 0\n95 percent confidence interval:\n\n                                                                         Page 7 of 11\n\fIAI                                                                 CS1B-0524\n\n 24.27889 40.61871\nsample estimates:\nmean of x mean of y\n 98.07843   65.62963\n       p-value is 3.586e-12                                                  (1)\n\n       At 95% Confidence interval as the confidence interval (24.27889 40.618\n71) does not contain the value 0, we have sufficient evidence to reject Null H\nypothesis and conclude that there is significant difference between the averag\ne claim size of Textile Mills and Transporters’ Godowns.\n\n(V)\n(a)\n> model1=glm(Firepolicies$Claimed~Firepolicies$Claim.Size+Firepolicies$Locati\non,family = binomial())                                                      (2)\n\n> summary(model1)                                                            (1)\n\nCall:\nglm(formula = Firepolicies$Claimed ~ Firepolicies$Claim.Size +\n    Firepolicies$Location, family = binomial())\n\nCoefficients:\n                         Estimate Std. Error z value Pr(>|z|)\n(Intercept)             -0.840330   0.453954 -1.851    0.0642 .\nFirepolicies$Claim.Size 0.003711    0.004935   0.752   0.4521\nFirepolicies$LocationG   0.198437   0.494674   0.401   0.6883\nFirepolicies$LocationK   0.205169   0.506205   0.405   0.6853\nFirepolicies$LocationM -0.418141    0.498094 -0.839    0.4012\nFirepolicies$LocationT   0.448865   0.519037   0.865   0.3871\n---\nSignif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1\n\n(Dispersion parameter for binomial family taken to be 1)\n\n    Null deviance: 251.91    on 191 degrees of freedom\nResidual deviance: 247.65   on 186 degrees of freedom\nAIC: 259.65\n\nNumber of Fisher Scoring iterations: 4                                       (1)\n\nOr Alternatively\n\n> model1=glm(Firepolicies$Claimed~Firepolicies$Claim.Size+Firepolicies$Locati\non,family = binomial(link=logit))\n> model1\n\nCall: glm(formula = Firepolicies$Claimed ~ Firepolicies$Claim.Size +\n    Firepolicies$Location, family = binomial (link = logit))\n\nCoefficients:\n            (Intercept)   Firepolicies$Claim.Size   Firepolicies$LocationG\n              -0.840330                  0.003711                 0.198437\n Firepolicies$LocationK    Firepolicies$LocationM   Firepolicies$LocationT\n               0.205169                 -0.418141                 0.448865\n\nDegrees of Freedom: 191 Total (i.e. Null);   186 Residual\n\n                                                                    Page 8 of 11\n\fIAI                                                                     CS1B-0524\n\nNull Deviance:     251.9\nResidual Deviance: 247.6         AIC: 259.6\n\nb)\n> model2=glm(Firepolicies$Claimed~Firepolicies$Claim.Size+Firepolicies$Occupa\nncy,family = binomial())\n> model2\n\nCall: glm(formula = Firepolicies$Claimed ~ Firepolicies$Claim.Size +\n    Firepolicies$Occupancy, family = binomial())\n\nCoefficients:\n             (Intercept)      Firepolicies$Claim.Size    Firepolicies$OccupancyDW\n               -0.492623                    -0.009519                   -0.266777\nFirepolicies$OccupancyHG     Firepolicies$OccupancyTG    FirepoliciesOccupancyTG F\nirepolicies$OccupancyTM\n                0.100853                      0.817776                   1.063207\n\nDegrees of Freedom: 191 Total (i.e. Null);     186 Residual\nNull Deviance:     251.9\nResidual Deviance: 247.3      AIC: 259.28                                       (1)\n\nOr Alternatively\n\n> model2=glm(Firepolicies$Claimed~Firepolicies$Claim.Size+Firepolicies$Occup\nancy,family = binomial(logit))\n> model2\n\nCall: glm(formula = Firepolicies$Claimed ~ Firepolicies$Claim.Size +\n    Firepolicies$Occupancy, family = binomial(logit))\n\nCoefficients:\n             (Intercept)      Firepolicies$Claim.Size    Firepolicies$OccupancyDW\n               -0.492623                    -0.009519                   -0.266777\nFirepolicies$OccupancyHG     Firepolicies$OccupancyTG    Firepolicies$OccupancyTM\n                0.100853                     0.817776                    1.063207\n\nDegrees of Freedom: 191 Total (i.e. Null);     186 Residual\nNull Deviance:     251.9\nResidual Deviance: 247.3      AIC: 259.3\n\n(vi)\nAIC for Model 1 = 259.65 and AIC for Model 2 = 259.28                          (1)\nThe AIC is smaller for the model2 as compared to model1                        (1)\nSo Claim Size and Occupancy model2 seems to be a better predictor than\n      Claim Size and Location, and we would choose the model2.                 (1)\n\nAlternatively,\nSince both the models have a very minor difference in AICs, one can conclude\nthat both models 1 and 2 are equally good.\n\n(vii)\nClaimed is a numerical variable\nClaim.Size is a numerical variable\nLocation is a factor variable\nOccupancy is a factor variable\n\n                                                                        Page 9 of 11\n\fIAI                                                                CS1B-0524\n\nNumerical variables are continuous variables which can take numerical values.\nClaim size and Claimed (0 for no claims in the last year or 1 for claim in     t\nhe last year) are examples in the context of this GLM.\nFactor / categorical variables are variables which only take categories.\nLocation is a factor variable which takes values of 5 states – M, G, T, K and\nA. Occupancy is also a factor variable which takes 5 values viz.TM, TG, DW,\nCS and HG.",
      "has_math": false,
      "is_r_task": true,
      "session": "2024-05",
      "subject": "CS1B",
      "source_qp": "raw/CS1B_2024-05_QP.pdf",
      "source_sol": "raw/CS1B_2024-05_SOL.pdf"
    },
    {
      "q_num": 4,
      "marks": 20,
      "topic": "bayes_credibility",
      "subtopics": [
        "inference"
      ],
      "stem": "The following data represents the average claim size (motor third party - accident injury) under\n       motor insurance policies settled in six different states across a country.\n\n                                                        Year j\n           State i\n                         2018-2019      2019-2020     2020-2021      2021-2022      2022-2023\n           Assam          219458         240371        289307         264439         279704\n           Bihar          216594         231311        261915         286211         291934\n         Chattisgarh      213871         231461        264519         261279         270058\n            Delhi         389197         400926        393130         391921         373129\n           Odisha         224879         243361        276718         292564         300135\n           Kerala         194814         230113        258101         276876         287973\n\n       Enter the data in R in the form of a matrix using the following code:\n\n       Claims<-matrix(c(219458, 216594, 213871, 389197,224879,194814,240371,231311,231461,4\n       00926,243361,230113, 289307, 261915, 264519, 393130, 276718, 258101, 264439,286211, 2\n       61279, 391921, 292564, 276876, 279704,291934, 270058, 373129, 300135, 287973), nrow=6\n       , ncol=5)\n\n      Or\n\n      (You can copy the R code from provided file ‘Q4 Reference_RCode.docx)",
      "parts": [
        {
          "label": "i",
          "marks": 4,
          "text": "Calculate, using Empirical Bayes Credibility Theory (EBCT) Model 1, the following:\n\n           a) E[m(ɵ)]                                                                                      (2)\n\n           b) E[𝑠 2 (ɵ)]                                                                                   (2)\n\n           c) Var [m(ɵ)]                                                                                   (4)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 5,
          "text": "Calculate the credibility factors Zi and the credibility premiums for Delhi and Kerala.         (5)",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 2,
          "text": "Comment on the relationship between n and Zi in case of EBCT Model 1.                           (2)\n\n      The number of motor third party claims per policy is modelled as a random variable X with a\n      Poisson distribution with unknown parameter ʎ. The log likelihood function for estimating ʎ is\n      given by:\n                                      L(ʎ) = log(ʎ) × ∑ x − ʎn\n\n                                    where n is the number of observations in the sample data.\n\n      There is data for a total of 150 observations and the total number of motor third party claims is\n      280.\n\n      The log likelihood function for the values of ʎ = 0, 0.1, 0.2, ………….., 1.8, 1.9, 3.3 has been\n      plotted below:\n\n                                                       Log-likelihood of lambda\n                                    -200\n                   Log_likelihood\n\n                                    -400\n                                    -600\n\n                                           0.0   0.5    1.0    1.5       2.0   2.5   3.0\n\n                                                                lambda",
          "topic": null
        },
        {
          "label": "iv",
          "marks": 2,
          "text": "Determine an approximate maximum likelihood estimate for ʎ using this plot.                      (2)",
          "topic": null
        },
        {
          "label": "v",
          "marks": 3,
          "text": "Determine the exact maximum likelihood estimate for ʎ and compare your answer with\n          the approximate estimate obtained in part (iv).                                                  (3)",
          "topic": null
        }
      ],
      "solution": "(i)\na) > m <- mean(rowMeans(Claims))                                             (1)\n   > m\n   [1] 278542.3                                                              (1)\n\nb) > s<-mean(apply(Claims,1,var))                                            (1)\n      > s\n      [1] 846425572                                                          (1)\n\nc) > n <-ncol(Claims)\n   > n\n   [1] 5\n\n  > v<-var(rowMeans(Claims))-mean(apply(Claims,1,var))/n                     (2)\n  > v\n  [1] 2842778626                                                             (1)\n\n(ii)\n\n> Z <- n/(n+s/v)                                                             (1)\n> Z\n[1] 0.9437976                                                                (1)\n\n> PurePremium <-Z*rowMeans(Claims)+(1-Z)*m                                   (1)\n> PurePremium\n[1] 259773.5 258770.4 249940.8 383415.5 268150.2 251203.4                    (1)\nCredibility premium for Delhi is 383415.5 and for Kerala is 251203.40        (1)\n\n(iii)\nZ is an increasing function of n. In the formula for credibility factor Z = n\n/ (n + s/v), with an increasing value of n, Z will tend to increase. Intuitive\nly also, it is true because as the number of observations for the particular r\nisk under consideration are more, more reliable is the specific data from that\nparticular risk and hence credibility factor Z would be higher indicative of m\n\n                                                                  Page 10 of 11\n\fIAI                                                                CS1B-0524\n\nore weightage to the specific data(mean for the specific risk) rather than the\ncollateral data (overall mean for all risks)\n\n(iv)\nBased on the graph, the approximate maximum likelihood estimate i.e. the value\nat which the log likelihood is maximum is around 1.8 to 1.9.\n(v)\nExact Maximum Likelihood Estimate λ is\n\n                                  λ = Σ𝑥𝑖 /𝑛\n                                   = 280/150\n                                   = 1.87                                 (2)\nSo, the actual maximum likelihood estimate calculated using first principles\nis close to the approximate maximum likelihood determined based on the graph.\n\n                               **************\n\n                                                                  Page 11 of 11",
      "has_math": true,
      "is_r_task": true,
      "session": "2024-05",
      "subject": "CS1B",
      "source_qp": "raw/CS1B_2024-05_QP.pdf",
      "source_sol": "raw/CS1B_2024-05_SOL.pdf"
    },
    {
      "q_num": 1,
      "marks": 30,
      "topic": "inference",
      "subtopics": [
        "distributions"
      ],
      "stem": "Consider a random sample X1, X2, …….., Xn from a Chi-Square distribution with 2\n        degrees of freedom and define Y = ∑ni=1 Xi",
      "parts": [
        {
          "label": "i",
          "marks": 3,
          "text": "State the distribution of Y, giving all the parameters of the distribution.\n\n        [Hint: If W ~ Gamma (α,ʎ) then 2ʎW has χ2 distribution with 2α degrees of freedom.]             (3)\n\n        Following 15 random numbers have been generated from a U(0,1) distribution using R.\n        0.07991847,    0.82064314,     0.33683219,    0.93005953,  0.31919393,\n        0.92695533,0.76263949, 0.51740370, 0.49224880, 0.46354694, 0.89832157,\n        0.21920729, 0.20471780, 0.19055074, 0.69137537",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 3,
          "text": "Using the 15 random numbers provided above, simulate a sample x1, x2,……., x15 from a\n        chi-square distribution with 2 degrees of freedom. Based on this sample, calculate the\n        value of Y.\n\n        Note: You can directly copy these random numbers from Codes.docx file in your R Script.         (3)",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 6,
          "text": "Using the values of x1, ……., xn generated in part (ii), test whether the standard deviation\n        of X is equal to 2.5 from scratch using qchisq() to determine the critical values and\n        pchisq() function to determine the p-value. Clearly state the null and alternate\n        hypothesis, the p-value and the conclusion of the test. Perform the test at 5% level of\n        significance.                                                                                   (6)",
          "topic": null
        },
        {
          "label": "iv",
          "marks": 4,
          "text": "Write and execute R script to generate 1,000 samples of x of size 15. Then, calculate the\n        sum of 1,000 corresponding values of Y based on these samples. Use set.seed(47) and\n        print the sum of 1,000 y values.                                                                (4)",
          "topic": null
        },
        {
          "label": "v",
          "marks": 3,
          "text": "Histograms have been plotted showing the relative frequencies of y1 to y1000 with\n        underlying sample sizes of 15 and 10,000 respectively. They are given below:\n\n               Underlying Sample Size of 15                 Underlying Sample Size of 10,000\n\n        Comment on the difference between the two histograms of Y for sample sizes 15 and\n        10,000, particularly in relation to the central limit theorem and how the sample size\n        affects the shape of the distribution.                                                          (3)",
          "topic": null
        },
        {
          "label": "vi",
          "marks": 5,
          "text": "By appropriately modifying the code written in part (iv) or otherwise, write code in R to\n        generate sample means ̅X and sample variances S2 for 1,000 random samples of Y each\n        having 10,000 values of x. Use set.seed(47). You are NOT required to execute the\n        code.                                                                                           (5)",
          "topic": null
        },
        {
          "label": "vii",
          "marks": 4,
          "text": "Histograms have been plotted for sample means X    ̅ and sample variannces S2 using the\n        code written in part (vi). They have been given overleaf:\n\n                       Sample means                                 Sample variances\n\n         By doing visual inspection of these histograms or otherwise, check whether –\n\n                ̅ is an unbiased estimator of the population mean µ\n            (a) X\n            (b) S2 is an unbiased estimator of the population variance δ2                                (4)",
          "topic": null
        },
        {
          "label": "viii",
          "marks": 2,
          "text": "Also comment on the normality of the distributions of ̅\n                                                               X and S2 using the above\n         histograms.                                                                                     (2)",
          "topic": null
        }
      ],
      "solution": "i)    We know the property that if W ~ Gamma(α,ʎ) then 2ʎW has χ2 distribution with 2α degrees\n        of freedom.\n\n        We have Xi ~ χ2 with 2 degrees of freedom.\n\n        Xi / 2ʎ ~ Gamma(1,ʎ)\n\n        We further know that for a chi-square distribution, ʎ = ½\n\n        Xi ~ Gamma(1,½)\n\n        Hence, Xi ~ Exp(½)\n                                              ……… Gamma distribution with parameters α=1 and ʎ\n                                              is an exponential distribution with parameter ʎ\n\n         Y = ∑𝑛𝑖=1 𝑋𝑖\n\n        Hence, Y ~ Gamma (n,½                 ……….. Sum of ‘n’ exponential variables is a gamma\n        variable                                                                                   (3)\n\n  ii)   First copy the numbers in R and define a vector u_ran.\n        Then using qchisq function convert these uniform random numbers into random numbers\n        from a chi-square distribution with 2 degrees of freedom.\n\n        R Code and Output:\n\n        > u_ran <-c(0.07991847, 0.82064314, 0.33683219, 0.93005953, 0.31919393, 0.92695533,0.7\n        6263949, 0.51740370, 0.49224880, 0.46354694, 0.89832157, 0.21920729, 0.20471780, 0.190\n        55074, 0.69137537)\n        > print(u_ran)\n         [1] 0.07991847 0.82064314 0.33683219 0.93005953 0.31919393 0.92695533\n         [7] 0.76263949 0.51740370 0.49224880 0.46354694 0.89832157 0.21920729\n        [13] 0.20471780 0.19055074 0.69137537\n\n        > x <- qchisq(u_ran,2)\n         [1] 0.1665860 3.4367557 0.8214544 5.3202217 0.7689556 5.2333682 2.8763503\n         [8] 1.4571496 1.3555274 1.2455524 4.5718802 0.4948912 0.4581165 0.4228024\n        [15] 2.3512591\n\n        > y=sum(x)\n        > print(y)\n        [1] 30.98087                                                                               (3)\n\n iii)   H0: Standard deviation of X is equal to 2.5.\n        H1: Standard deviation of X is not equal to 2.5.\n\n        > n=length(x)\n        > sigma=2.5\n        > alpha=0.05\n        >\n        > statistic <- (n-1)*var(x)/sigma^2\n\n                                                                                        Page 2 of 11\n\fIAI                                                                                              CS1B-1124\n\n       > statistic\n       [1] 7.305598\n       >\n       > #critical value\n       > qchisq(alpha/2,n-1)\n       [1] 5.628726\n       > qchisq(alpha/2,n-1,lower=FALSE)\n       [1] 26.11895\n       >\n       > #p-value\n       > 2*(pchisq((n-1)*var(x)/sigma^2,df=n-1))\n       [1] 0.1554235\n\n       As the p-value is > 5%, we do not have sufficient evidence to reject the null hypothesis and\n       hence we can conclude that the standard deviation of X is equal to 2.5.                            (6)\n\n iv)   We need to write a code to obtain sample for y with 1,000 simulations. Following code as\n       given the question can be used for that:\n\n       R Code and Output:\n\n       > set.seed(47)\n       > y = 0*(1:1000)\n       > for(i in 1:1000){\n       + y[i] = sum(rchisq(15,2))\n       +}\n\n       > sum(y)\n       [1] 30508.66                                                                                       (4)\n\n v)    The distribution of Y when sample size = 15, is not perfectly symmetrical like a normal\n       distribution. This is given because Y seems to accept only positive values as gamma\n       distribution is defined only when y>0.\n\n       However, when sample size = 10,000 this histogram is symmetrical and is closer to a normal\n       distribution.\n\n       As n tends to infinity, using central limit theorem, the distribution of the sample approaches a\n       normal distribution.\n\n       For a larger sample size of n (changed from 15 to 10,000) the central limit theorem ensures\n       that the distribution of Y becomes approximately normal.                                           (3)\n\n vi)   By modifying the code in part (iv), we generate 10,000 values of x and calculate the sample\n       mean and sample variance for each sample.\n\n       R Code and Output:\n\n       > set.seed(47)\n       > x_bar = 0*(1:1000)\n       > s_squared = 0*(1:1000)\n       > for(i in 1:1000){\n       + x_bar[i] = mean(rchisq(10000,2))\n       + s_squared[i] = var(rchisq(10000,2))\n                                                                                               Page 3 of 11\n\fIAI                                                                                                CS1B-1124\n\n         +}                                                                                                  (5)\n\n vii)    >\n         > print(mean(x_bar))\n         [1] 1.998798\n\n         Also, by visual inspection we can see that the mean of the sample means is close to 2.\n\n         Population mean is 2 (number of degrees of freedom of the variable Xi which has chi-square\n         distribution). Since E(X_bar) = mu, we can conclude that X_bar is an unbiased estimator of\n         population mean mu.\n\n         > print(mean(s_squared))\n         [1] 3.995956\n\n         Also, by visual inspection we can see that the mean of the sample variances is close to 4.\n\n         Population variance is 4 (2 times the number of degrees of freedom of the variable Xi which\n         has chi-square distribution). Since E(S2) = sigma2, we can conclude that S2 is an unbiased\n         estimator of population variance sigma2.                                                            (4)\n\n viii)   Comment:\n         The plot of X_bar is indicative of normality. This is true as for large sample size, X_bar ~\n         N(mu, sigma2 / n).\n\n         However, plot of Sigma_squared is relatively less normal as compared to plot of X_bar.              (2)",
      "has_math": true,
      "is_r_task": true,
      "session": "2024-11",
      "subject": "CS1B",
      "source_qp": "raw/CS1B_2024-11_QP.pdf",
      "source_sol": "raw/CS1B_2024-11_SOL.pdf"
    },
    {
      "q_num": 2,
      "marks": 15,
      "topic": "bayes_credibility",
      "subtopics": [
        "distributions"
      ],
      "stem": "There are three life insurance companies having below summary of claims for last 4\n         years:\n\n         Aggregate claims in millions:\n             Company                                         Year\n                                   1                2                   3                 4\n                A                14.2              15.8               22.7               19\n                 B               58.6              63.1                81               64.2\n                C                 123              132                161               133\n\n         No of claims:\n             Company                                         Year\n                                    1                2                  3                  4\n                  A                163              189                252                199\n                  B               4435             4761               5576               4581\n                  C              16184            17443              20102              18000\n\n         We want to calculate credibility premiums and estimate risk premiums under the\n         assumptions of the Empirical Bayes Credibility Theory (EBCT) Model 2 using the\n         following code in R:\n\n         claims<-data.frame(\n           Year1 = c(14.2,58.6,123),\n            Year2 = c(15.8,63.1,132),\n            Year3 = c(22.7,81.0,161),\n            Year4 = c(19,64.2,133)\n          )\n         nopols<-data.frame(\n           Year1 = c(163,4435,16184),\n           Year2 = c(189,4761,17443),\n           Year3 = c(252,5576,20102),\n           Year4 = c(199,4581,18000)\n         )\n\n         n <- ncol(claims)\n         N <-nrow(claims)\n\n        X <- claims/nopols\n        Xibar <-rowSums(claims) / rowSums(nopols)\n        Pibar <- rowSums(nopols) #………………… A\n        Pbar <-sum(Pibar)\n        Pstar <-sum(Pibar * (1-Pibar/Pbar))/(N*n-1) #………………… B\n        m <-sum(claims) / Pbar #…………………… C\n        s <- mean(rowSums(nopols *(X-Xibar)^2)/(n-1))               #……………… D\n        v <- (sum(rowSums(nopols*(X-m)^2))/(n*N-1)-s)/Pstar #………………… E",
      "parts": [
        {
          "label": "i",
          "marks": 5,
          "text": "Briefly explain the meaning of the quantities calculated in code lines A to E in simple\n        terms.                                                                                          (5)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 3,
          "text": "Use the code provided to calculate the estimates of E[m(θ)], E[s2(θ)] and Var[m(θ)]\n        under EBCT Model 2.\n\n        Note: You can directly copy these code lines from Codes.docx file in your R Script.             (3)",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 2,
          "text": "Calculate the credibility factors Zi.                                                           (2)",
          "topic": null
        },
        {
          "label": "iv",
          "marks": 3,
          "text": "If the number of claims for the next year are 5000, 4800 and 4200 respectively for\n        insurers A, B and C, then estimate the risk premiums for the next year using the\n        credibility factors determined in part (iii).                                                   (3)",
          "topic": null
        },
        {
          "label": "v",
          "marks": 2,
          "text": "Based on your calculations in part (iii), comment on whether more emphasis is given to\n        direct data or collateral data when calculating risk premiums.                                  (2)",
          "topic": null
        }
      ],
      "solution": "i)     A. Pibar represents the total number of claims for every insurer (total of no of claims across\n            all the years)\n         B. Pstar is a quantity calculated by taking sum over all insurers of (total number of claims per\n            insurer multiplied by proportion of total no of claims of that insurer to total) and divided\n            by total number of cells minus 1. This quantity is used in the denominator while\n            determining V(m(θ)).\n         C. m or E(m(θ)) is the expected value of average claims per policy based on the collateral\n            data.\n         D. s or E(s2(θ)) is the expected value of the variance of claims per policy based on the\n            collateral data.\n         E. v or V(m(θ)) is the variance of average claims per policy based on the collateral data           (5)\n\n  ii)    > claims<-data.frame(\n         + Year1 = c(14.2,58.6,123),\n         + Year2 = c(15.8,63.1,132),\n         + Year3 = c(22.7,81.0,161),\n         + Year4 = c(19,64.2,133)\n         +)\n         >\n         > nopols<-data.frame(\n         + Year1 = c(163,4435,16184),\n         + Year2 = c(189,4761,17443),\n                                                                                                  Page 4 of 11\n\fIAI                                                                                            CS1B-1124\n\n        + Year3 = c(252,5576,20102),\n        + Year4 = c(199,4581,18000)\n        +)\n        >\n        > n <- ncol(claims) ## This stands for n\n        > N <-nrow(claims) ## This stands for N\n        > X <- claims/nopols ## This stands for Xij\n        > Xibar <-rowSums(claims) / rowSums(nopols) # Xibar\n        > Pibar <- rowSums(nopols) # Pibar\n        > Pbar <-sum(Pibar) # Pbar\n        > Pstar <-sum(Pibar * (1-Pibar/Pbar))/(N*n-1) # Pstar\n        > m <-sum(claims) / Pbar # E[m(θ)]\n        > print(m)\n        [1] 0.009659901\n        >\n        > s <- mean(rowSums(nopols *(X-Xibar)^2)/(n-1)) # E[s2(θ)]\n        > print(s)\n        [1] 0.002749923\n\n        > v <- (sum(rowSums(nopols*(X-m)^2))/(n*N-1)-s)/Pstar\n        > print(v)\n        [1] 0.0001793695                                                                                 (3)\n\n iii)   > Zi <- Pibar / (Pibar + s/v)\n        > print(Zi)\n        [1] 0.9812655 0.9992084 0.9997863                                                                (2)\n\n iv)    > Premi <- Zi * Xibar + (1-Zi) * m\n        > print(Premi)\n        [1] 0.087798326 0.013787873 0.007654237\n\n        > nopols_y5 <- c(5000,4800,4200)\n        > claims_y5 <- Premi * nopols_y5\n\n        > print(claims_y5)\n        [1] 438.99163 66.18179 32.14779                                                                  (3)\n\n  v)    Since the credibility premiums are high and close to 1 in case of all insurers, we are giving\n        more importance to the direct data (related to that particular insurer) and ignoring the\n        collateral data (data related to other insurers)\n\n        This is reasonable as there is wide variation in the claims amount for various insurers and\n        hence more emphasis should be given on direct data.                                              (2)",
      "has_math": true,
      "is_r_task": true,
      "session": "2024-11",
      "subject": "CS1B",
      "source_qp": "raw/CS1B_2024-11_QP.pdf",
      "source_sol": "raw/CS1B_2024-11_SOL.pdf"
    },
    {
      "q_num": 3,
      "marks": 40,
      "topic": "regression_glm",
      "subtopics": [
        "distributions"
      ],
      "stem": "A National Sports Meet was organised and 27 States participated in the meet. There are\n        three categories of sports:\n\n        1. Running : 100m, 400m and 110m hurdle\n        2. Jumping : High, Long and Pole Vault\n        3. Throwing: Shot put, Javeline and Discuss throw\n\n        Data, named Sports.csv, was collected of the event and shared with National Academy\n        to analyse and prepare for International events.",
      "parts": [
        {
          "label": "i",
          "marks": 2,
          "text": "(a) Fit a multiple regression model between winning points (as the response variable) and\n        the scores in the nine sports (as explanatory variables). Display the summary of the fitted\n        model and identify which explanatory variables are significant at 5% level of\n        significance.                                                                                   (6)\n\n        (b) Prepare a plot of the residuals of the multiple regression model.                           (2)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 4,
          "text": "Fit a Generalised Linear Model (GLM) to the data using winning points as the response\n        variable and scores obtained in the nine sports as the explanatory variables, assuming a\n        Poisson distribution for the response variable. Your answer should include the estimated\n        coefficients and the Akaike’s Information Criteria (AIC) of the fitted model.                   (4)",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 2,
          "text": "Explain why scaled deviance cannot be used to compare the fit of the models in parts (i)        (2)\n\n         and (ii).",
          "topic": null
        },
        {
          "label": "iv",
          "marks": 4,
          "text": "Fit, by choosing a suitable argument for family in the glm command, a GLM to the data\n         that would be equivalent to the model fitted in part (i). Your answer should include the\n         estimated coefficients and the AIC of the fitted model.                                         (4)",
          "topic": null
        },
        {
          "label": "v",
          "marks": 2,
          "text": "Compare the fit of the models fitted in parts (i) and (ii) using AIC for comparison.            (2)",
          "topic": null
        },
        {
          "label": "vi",
          "marks": 7,
          "text": "Perform Principal Components Analysis (PCA) separately for the above three categories\n         of sports viz. Running Sports, Jumping Sports and Throwing Sports and share how much\n         variation is captured by first principal component (PC1) for each category.\n\n         [Hint: Use prcomp and do scaling by using scale.=TRUE parameter.]                               (7)",
          "topic": null
        },
        {
          "label": "vii",
          "marks": 3,
          "text": "Under which sports category, principal components will be most useful in reducing\n         dimensionality of the dataset while capturing 90% of variance?                                  (3)",
          "topic": null
        },
        {
          "label": "viii",
          "marks": 5,
          "text": "The organisers informed that pole vault sports winning points are not properly captured.\n         Test whether there is any correlation between pole vault points and overall winning points\n         by calculating a 95% confidence interval for the Slope Coefficient Beta. Clearly state the\n         null and alternative hypotheses and the conclusion of the test.                                 (5)",
          "topic": null
        },
        {
          "label": "ix",
          "marks": 5,
          "text": "Players for racing sports have now been shortlisted for the international event and trials\n         have started for the upcoming international event. However due to unforeseen\n         circumstances, trials for 110 metres hurdle race got cancelled for last 5 runners. An\n         analyst suggests that their scores can be predicted using their 100m trial scores.\n\n         Below are the 100 metres scores of last 5 runners.\n         1. 10.68\n         2. 10.42\n         3. 11.68\n         4. 11.62\n         5. 10.54\n\n         Fit a linear regression model to predict the scores in 110 metres hurdle race using scores\n         in 100 metres race as the explanatory variable using the national sports meet data. Clearly\n         state the equation of this linear regression model. Using this equation, predict the\n         expected 110m hurdle race scores of the last five runners.                                      (5)",
          "topic": null
        }
      ],
      "solution": "i)    #a > Sports <- read.csv(Sports.csv\")\n\n        > Model1<-lm(Sports$Points~Sports$X100m+Sports$X400m+Sports$X110m.hurdle+Sports\n        $High.jump+Sports$Long.jump+Sports$Pole.vault+Sports$Shot.put+Sports$Javeline+Sports\n        $Discus)\n        > summary(Model1)\n\n                                                                                             Page 5 of 11\n\fIAI                                                                                            CS1B-1124\n\n       Call:\n       lm(formula = Sports$Points ~ Sports$X100m + Sports$X400m + Sports$X110m.hurdle +\n         Sports$High.jump + Sports$Long.jump + Sports$Pole.vault +\n         Sports$Shot.put + Sports$Javeline + Sports$Discus)\n\n       Residuals:\n         Min      1Q Median    3Q Max\n       -97.106 -25.043 -7.748 33.856 119.528\n\n       Coefficients:\n                   Estimate Std. Error t value Pr(>|t|)\n       (Intercept)     8632.632 1368.083 6.310 7.83e-06 ***\n       Sports$X100m        -163.414 76.208 -2.144 0.046756 *\n       Sports$X400m         -78.141 16.269 -4.803 0.000166 ***\n       Sports$X110m.hurdle -120.901 41.818 -2.891 0.010152 *\n       Sports$High.jump 810.196 211.532 3.830 0.001340 **\n       Sports$Long.jump 209.187 63.472 3.296 0.004269 **\n       Sports$Pole.vault 183.345 66.876 2.742 0.013912 *\n       Sports$Shot.put     90.349 28.200 3.204 0.005204 **\n       Sports$Javeline     17.757     2.975 5.969 1.52e-05 ***\n       Sports$Discus       10.986     5.781 1.900 0.074491 .\n       ---\n       Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1\n\n       Residual standard error: 60.74 on 17 degrees of freedom\n       Multiple R-squared: 0.9794,           Adjusted R-squared: 0.9685\n       F-statistic: 89.91 on 9 and 17 DF, p-value: 1.491e-12\n\n       Using a general rule if the p-value is less than 0.05, then the concerned explanatory variable\n       is considered to be significant. This is shown by *, **, *** signs in the R summary output .\n\n       So, based on the R-output, only points for Discus Throw are not significant. All other\n       explanatory variables are considered to be significant.\n       #b Plot of residuals for the model\n                             residuals(Model1)\n\n                                                 0 50\n                                                 -100\n\n                                                        0   5   10      15   20   25\n\n                                                                     Index\n\n ii)   > Model2_poisson<-glm(Sports$Points~Sports$X100m+Sports$X400m+Sports$X110m.hurd\n       le+Sports$High.jump+Sports$Long.jump+Sports$Pole.vault+Sports$Shot.put+Sports$Javelin\n       e+Sports$Discus,family=\"poisson\")\n\n       > summary(Model2_poisson)\n\n       Call:\n       glm(formula = Sports$Points ~ Sports$X100m + Sports$X400m + Sports$X110m.hurdle +\n         Sports$High.jump + Sports$Long.jump + Sports$Pole.vault +\n         Sports$Shot.put + Sports$Javeline + Sports$Discus, family = \"poisson\")\n\n       Deviance Residuals:\n                                                                                             Page 6 of 11\n\fIAI                                                                                           CS1B-1124\n\n           Min     1Q Median      3Q    Max\n        -1.02319 -0.25491 0.03473 0.36449 1.31328\n\n        Coefficients:\n                      Estimate Std. Error z value Pr(>|z|)\n        (Intercept)      9.0988649 0.2501151 36.379 < 2e-16 ***\n        Sports$X100m         -0.0211126 0.0139153 -1.517 0.12921\n        Sports$X400m         -0.0093734 0.0029665 -3.160 0.00158 **\n        Sports$X110m.hurdle -0.0161147 0.0076612 -2.103 0.03543 *\n        Sports$High.jump 0.1009034 0.0387525 2.604 0.00922 **\n        Sports$Long.jump 0.0239972 0.0116951 2.052 0.04018 *\n        Sports$Pole.vault 0.0225249 0.0122296 1.842 0.06550 .\n        Sports$Shot.put     0.0111537 0.0051471 2.167 0.03024 *\n        Sports$Javeline 0.0021487 0.0005379 3.995 6.48e-05 ***\n        Sports$Discus       0.0012340 0.0010555 1.169 0.24236\n        ---\n        Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1\n\n        (Dispersion parameter for poisson family taken to be 1)\n\n          Null deviance: 374.4158 on 26 degrees of freedom\n        Residual deviance: 8.3483 on 17 degrees of freedom\n        AIC: 321\n\n        Number of Fisher Scoring iterations: 3                                                          (4)\n\n iii)   Scaled deviance can be used to compare only nested models.\n\n        Since, Model 1 and Model2_Poisson are models with different distributional assumptions.\n\n        Model 1 assumes normal distribution and Model 2 assumes Poisson distribution.\n        They are not nested models.\n\n        Hence, scaled deviance cannot be used to compare the models in parts (i) and (ii).              (2)\n\n iv)    An equivalent model to a linear regression model will be a GLM with normal distribution.\n\n        > Model2_normal<-glm(Sports$Points~Sports$X100m+Sports$X400m+Sports$X110m.hurdl\n        e+Sports$High.jump+Sports$Long.jump+Sports$Pole.vault+Sports$Shot.put+Sports$Javelin\n        e+Sports$Discus,family=gaussian())\n        > summary(Model2_normal)\n\n        Call:\n        glm(formula = Sports$Points ~ Sports$X100m + Sports$X400m + Sports$X110m.hurdle +\n          Sports$High.jump + Sports$Long.jump + Sports$Pole.vault +\n          Sports$Shot.put + Sports$Javeline + Sports$Discus, family = gaussian())\n\n        Deviance Residuals:\n          Min     1Q Median     3Q Max\n        -97.106 -25.043 -7.748 33.856 119.528\n\n        Coefficients:\n                    Estimate Std. Error t value Pr(>|t|)\n        (Intercept)     8632.632 1368.083 6.310 7.83e-06 ***\n                                                                                             Page 7 of 11\n\fIAI                                                                                              CS1B-1124\n\n       Sports$X100m        -163.414 76.208 -2.144 0.046756 *\n       Sports$X400m         -78.141 16.269 -4.803 0.000166 ***\n       Sports$X110m.hurdle -120.901 41.818 -2.891 0.010152 *\n       Sports$High.jump 810.196 211.532 3.830 0.001340 **\n       Sports$Long.jump 209.187 63.472 3.296 0.004269 **\n       Sports$Pole.vault 183.345 66.876 2.742 0.013912 *\n       Sports$Shot.put     90.349 28.200 3.204 0.005204 **\n       Sports$Javeline     17.757     2.975 5.969 1.52e-05 ***\n       Sports$Discus       10.986     5.781 1.900 0.074491 .\n       ---\n       Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1\n\n       (Dispersion parameter for gaussian family taken to be 3689.4)\n\n         Null deviance: 3048179 on 26 degrees of freedom\n       Residual deviance: 62720 on 17 degrees of freedom\n       AIC: 307.89\n\n       Number of Fisher Scoring iterations: 2                                                             (4)\n\n v)    Model 1 in part (i) and Model 2 (Normal) in part (iv) are exact equivalents of each other. The\n       same can be seen from the estimates of coefficients in the R summary output.\n\n       So, for comparing the fit of Model 1 and Model 2 (Poisson), Model 2 (Normal) can be used as\n       a proxy for Model 1. And then, the AIC of Model 2 (Poisson) can be compared with the AIC\n       of Model 2 (Normal).\n\n       AIC (Model 2 Normal) = 307.89\n       AIC (Model 2 Poisson) = 321\n\n       Smaller the AIC, better is the fit. So, Model 2 Normal is a better fit as compared to Model 2\n       Poisson.\n\n       Consequentially, the linear multiple regression model in part (i) is a better fit to the data as\n       compared to the GLM fitted in part (ii).                                                           (2)\n\n vi)   > run<-data.frame(Sports$X100m,Sports$X400m,Sports$X110m.hurdle)\n       > pr_run<-prcomp(run,scale. = TRUE)\n       >\n       > summary(pr_run)\n       Importance of components:\n                      PC1 PC2 PC3\n       Standard deviation 1.492 0.6680 0.5726\n       Proportion of Variance 0.742 0.1487 0.1093\n       Cumulative Proportion 0.742 0.8907 1.0000\n\n       >\n       > jump<-data.frame(Sports$High.jump,Sports$Long.jump,Sports$Pole.vault)\n       > pr_jump<-prcomp(jump,scale. = TRUE)\n       >\n       > summary(pr_jump)\n       Importance of components:\n                      PC1 PC2 PC3\n                                                                                               Page 8 of 11\n\fIAI                                                                                               CS1B-1124\n\n        Standard deviation 1.2588 1.0307 0.5941\n        Proportion of Variance 0.5282 0.3541 0.1176\n        Cumulative Proportion 0.5282 0.8824 1.0000\n\n        >\n        > throw<-data.frame(Sports$Shot.put,Sports$Javeline,Sports$Discus)\n        > pr_throw<-prcomp(throw,scale. = TRUE)\n        >\n        > summary(pr_throw)\n        Importance of components:\n                        PC1 PC2 PC3\n        Standard deviation 1.4141 0.8727 0.48853\n        Proportion of Variance 0.6665 0.2539 0.07955\n        Cumulative Proportion 0.6665 0.9204 1.00000\n\n        For Run Sports Category, 74.2% variance is captured by PC1.\n        For Jumping Sports Category, 52.82% variance is captured by PC1.\n        For Throw Sports Category, 66.65% variance is captured by PC1.                                     (7)\n\nvii)    Considering the summary R output generated in part (vi),\n\n        In case of Run Sports Category, all three PCs cumulatively capture at least 90% of the total\n        variance of the data.\n\n        In case of Jump Sports Category, all three PCs cumulatively capture at least 90% of the total\n        variance of the data.\n\n        In case of Throw Sports Category, first two PCs cumulatively capture 92.04% (at least 90%)\n        of the total variance of the data.\n\n        So, if PCs capturing at least 90% of the variance is a criterion for reducing the dimensionality\n        of the data set, then Throw Sports Category satisfies it. For Throw Sports, PC1 and PC2 can\n        be retained and PC3 can be dropped thus reducing the dimensionality of the data set. In case\n        of Run Sports and Jumping Sports, all three PCs need to be retained and thus the\n        dimensionality of the data set will not be reduced.                                                (3)\n\nviii)   > Model3<-lm(Sports$Points~Sports$Pole.vault)\n\n        H0: Beta coefficient is equal to 0\n        H1: Beta coefficient is not equal to 0.\n\n        > confint(Model3,level=0.95)\n                     2.5 % 97.5 %\n        (Intercept)   5654.3838 10895.1124\n        Sports$Pole.vault -573.3559 508.9227\n\n        As the 95% confidence interval for beta (-573.3559, 508.9277) contains the value 0, we do\n        not have sufficient evidence to reject the null hypothesis at 5% level of significance. Hence,\n        based on this test, one can conclude that there is no correlation between pole vault score and\n        winning points.                                                                                    (5)\n\n ix)    > Model4<-lm(Sports$X110m.hurdle~Sports$X100m)\n        > summary(Model4)\n\n                                                                                                Page 9 of 11\n\fIAI                                                                                               CS1B-1124\n\n        Call:\n        lm(formula = Sports$X110m.hurdle ~ Sports$X100m)\n\n        Residuals:\n           Min     1Q Median      3Q Max\n        -0.48075 -0.28936 -0.05353 0.21428 0.76203\n\n        Coefficients:\n                 Estimate Std. Error t value Pr(>|t|)\n        (Intercept) 2.2036 2.7252 0.809 0.426379\n        Sports$X100m 1.1183 0.2478 4.512 0.000132 ***\n        ---\n        Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1\n\n        Residual standard error: 0.356 on 25 degrees of freedom\n        Multiple R-squared: 0.4489,           Adjusted R-squared: 0.4268\n        F-statistic: 20.36 on 1 and 25 DF, p-value: 0.0001319\n\n        > score<-c(10.68,10.42,11.68,11.62,10.54)\n        > hurdle_score<-2.2036+1.1183*score\n        > hurdle_score\n        [1] 14.14704 13.85629 15.26534 15.19825 13.99048",
      "has_math": false,
      "is_r_task": true,
      "session": "2024-11",
      "subject": "CS1B",
      "source_qp": "raw/CS1B_2024-11_QP.pdf",
      "source_sol": "raw/CS1B_2024-11_SOL.pdf"
    },
    {
      "q_num": 4,
      "marks": 15,
      "topic": "inference",
      "subtopics": [],
      "stem": "A researcher has collected the following data on a group of students, regarding whether\n         they passed or failed an exam and whether or not they attended tutorials:\n\n            Number of students                  Exam passed                    Exam failed\n             Attended tutorials                    132                            27\n           Did not attend tutorials                120                            51\n\n         The data can be entered into R in matrix form using the following code:\n\n         exam.success = matrix(c(132,120,27,51),ncol=2,nrow=2)\n\n         Note: You can directly copy the above code from Codes.docx file in your R Script.\n\n         The researcher wants to establish whether tutorial attendance is independent of exam\n\n        success, using a chi-square test.\n\n        Load the data in R and check for errors by displaying it.",
      "parts": [
        {
          "label": "i",
          "marks": 1,
          "text": "State the hypothesis of this test.                                                              (1)",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 2,
          "text": "Calculate the expected frequencies for the data under the null hypotheses in part (i) using\n        an appropriate function in R.                                                                   (2)",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 3,
          "text": "Perform the test with continuity correction\n\n        Clearly state your conclusions at 1% level of significance.                                     (3)",
          "topic": null
        },
        {
          "label": "iv",
          "marks": 3,
          "text": "What would be the conclusion of the test at 1% level of significance if Fisher’s test is\n        used instead of chi-square test? The null and alternative hypotheses need NOT be re-\n        stated.\n\n        [Hint: Use fisher.test function in R.]                                                          (3)",
          "topic": null
        },
        {
          "label": "v",
          "marks": 2,
          "text": "In the context of the two tests performed in parts (iii) and (iv) above –\n\n            (a) Which one is an exact test and which one is an approximation?\n\n            (b) Which test is suitable for only 2×2 datasets and which test is suitable for any N×N\n                dataset?                                                                                (2)",
          "topic": null
        },
        {
          "label": "vi",
          "marks": 4,
          "text": "Using binomial test, test whether the proportion of students passing the examination is\n        equal to 60%. You are required to clearly state the null and alternative hypotheses, p-\n        value of the test and your conclusion at 5% level of significance.                              (4)",
          "topic": null
        }
      ],
      "solution": "i)    > exam.success = matrix(c(132,120,27,51),ncol=2,nrow=2)\n        > exam.success\n            [,1] [,2]\n        [1,] 132 27\n        [2,] 120 51\n\n        Data is properly getting displayed.\n\n        H0: tutorial attendance and exam success are independent, against\n        H1: tutorial attendance and exam success are not independent                                         (1)\n\n  ii)   > chisq.test(exam.success)$expected\n              [,1] [,2]\n        [1,] 121.4182 37.58182\n        [2,] 130.5818 40.41818                                                                               (2)\n\n iii)   > chisq.test(exam.success)\n\n                 Pearson's Chi-squared test with Yates' continuity correction\n\n        data: exam.success\n        X-squared = 6.8349, df = 1, p-value = 0.008939\n\n        The p-value is significant (e.g. at the 1% level), since 0.008939 < 0.01 – therefore there is evi\n        dence to reject the null hypothesis and we conclude that tutorial attendance and exam success\n        are not independent.                                                                                 (3)\n\n                                                                                               Page 10 of 11\n\fIAI                                                                                              CS1B-1124\n\n iv)   > fisher.test(exam.success)\n\n                 Fisher's Exact Test for Count Data\n\n       data: exam.success\n       p-value = 0.006544\n       alternative hypothesis: true odds ratio is not equal to 1\n       95 percent confidence interval:\n        1.190372 3.671876\n       sample estimates:\n       odds ratio\n         2.073216\n\n       The p-value is significant (e.g. at the 1% level), 0.006544 < 0.01 – therefore there is evidence\n       to reject the null hypothesis and we conclude that tutorial attendance and exam success are not\n       independent. Conclusion under Fisher’s exact test is similar to conclusion under contingency\n       table test.                                                                                         (3)\n\n v)    (a) Fisher’s test is an exact test whereas chi-square test is an approximation\n\n       (b) Fisher’s test is suitable for 2×2 datasets whereas chi-square test can be used for N×N datas\n       ets.                                                                                                (2)\n\n vi)   H0: the proportion of students passing the exam is 60% (p = 0.60)\n       H1: the proportion of students passing the exam is not equal to 60% (p <> 0.60)\n\n       > x=132+120\n       > n=x+27+51\n       > binom.test(x,n,conf.level = 0.95)\n\n                 Exact binomial test\n\n       data: x and n\n       number of successes = 252, number of trials = 330, p-value <\n       2.2e-16\n       alternative hypothesis: true probability of success is not equal to 0.5\n       95 percent confidence interval:\n        0.7140288 0.8084419\n       sample estimates:\n       probability of success\n                0.7636364\n\n       The p-value is < 2.2e-16 which is definitely less than 5% and hence we have sufficient eviden\n       ce to reject the null hypothesis and hence we can conclude that the proportion of students pass\n       ing the examination is not equal to 60%.\n\n                                               *************\n\n                                                                                              Page 11 of 11",
      "has_math": true,
      "is_r_task": true,
      "session": "2024-11",
      "subject": "CS1B",
      "source_qp": "raw/CS1B_2024-11_QP.pdf",
      "source_sol": "raw/CS1B_2024-11_SOL.pdf"
    },
    {
      "q_num": 1,
      "marks": 40,
      "topic": "bayes_credibility",
      "subtopics": [],
      "stem": "Following table presents the data on the number of matches played and total runs scored in\n     respect of three batsmen viz. Rohit, Virat and Shivam over four seasons of a premier cricketing\n     league:\n\n       Number of matches Rohit Virat Shivam\n             Season 1              14    16        18\n             Season 2              15    17        19\n             Season 3              16    18        20\n             Season 4              14    17        19\n\n       Runs Scored Rohit Virat Shivam\n        Season 1    420   640    720\n        Season 2    450   680    760\n        Season 3    500   720    800\n        Season 4    430   660    750\n\n      Sponsoring partners of the premier cricketing league want to estimate the runs to be\n      scored in respect of these players in the upcoming season 5 of the league. It is decided to\n      use Empirical Bayes Credibility Theory Model 2 (EBCT Model 2) for this purpose.\n      Following code has been written in R for the same:\n      matches <- data.frame(\n        Season1 = c(14, 16, 18),\n        Season2 = c(15, 17, 19),\n        Season3 = c(16, 18, 20),\n        Season4 = c(14, 17, 19)\n      )\n\n      runs <- data.frame(\n        Season1 = c(420, 640, 720),\n        Season2 = c(450, 680, 760),\n        Season3 = c(500, 720, 800),\n        Season4 = c(430, 660, 750)\n      )\n\n      n <- ncol(runs)\n      N <- nrow(runs)\n\n      X <- runs / matches\n      Xibar <- rowSums(runs) / rowSums(matches)\n\n      Pibar <- rowSums(matches)\n      Pbar <- sum(Pibar)\n\n      Pstar <- sum(Pibar * (1 - Pibar / Pbar)) / (N * n - 1)\n\n      m <- sum(runs) / Pbar\n\n      s <- mean(rowSums(matches * (X - Xibar)^2) / (n - 1))\n\n       v <- (sum(rowSums(matches * (X - m)^2)) / (n * N - 1) - s) / Pstar",
      "parts": [
        {
          "label": "i",
          "marks": 1,
          "text": "Copy the above code lines and display the batting averages (i.e. X) for each player across\n              all the four seasons of the premier cricketing league.\n               Note: You can directly copy these code lines from Codes.docx file in your R Script.",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 3,
          "text": "Calculate the estimates of E[m(θ)], E[s2 (θ)] and Var[m(θ)] under EBCT Model 2.        [3]",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 2,
          "text": "Calculate the credibility factors Zi.                                                 [2]",
          "topic": null
        },
        {
          "label": "iv",
          "marks": 3,
          "text": "If all the three players are estimated to play 15 matches each in the upcoming season of\n               the premier cricketing league, then estimate the runs to be scored during this season using\n               the assumptions of EBCT Model 2.                                                        [3]",
          "topic": null
        },
        {
          "label": "v",
          "marks": 1,
          "text": "Based on your calculations in part (iii), comment on whether we are giving more\n              emphasis on the direct data or on the collateral data while calculating risk premiums. [1]",
          "topic": null
        },
        {
          "label": "vi",
          "marks": 4,
          "text": "Instead of estimating the runs scored using the assumptions of EBCT Model 2, it is\n               decided to estimate the runs to be scored using the assumptions of Empirical Bayes\n               Credibility Theory Model 1 (EBCT Model 1). By making appropriate changes in the\n               code given earlier or otherwise, determine the credibility factor under the assumptions of\n               EBCT Model 1.                                                                          [4]",
          "topic": null
        },
        {
          "label": "vii",
          "marks": 3,
          "text": "Hence estimate the number of runs to be scored for Rohit, Virat and Shivam during the\n                fifth season under the assumptions of EBCT Model 1.                               [3]",
          "topic": null
        },
        {
          "label": "viii",
          "marks": 20,
          "text": "Briefly comment on the differences in the answers obtained in parts (iv) and (vii). Also,\n                comment on which model is the most appropriate in this case.                          [3]",
          "topic": null
        }
      ],
      "solution": null,
      "has_math": true,
      "is_r_task": true,
      "session": "2026-05",
      "subject": "CS1B",
      "source_qp": "raw/CS1B_2026-05_QP.pdf",
      "source_sol": null
    },
    {
      "q_num": 2,
      "marks": 50,
      "topic": "data_analysis",
      "subtopics": [],
      "stem": "Researchers are studying ectoplasmic density in the brain – a substance believed to influence\n          cognitive resonance.\n\n          Following data is available relating to Brain Resonance Index (BRI) and Ectoplasmic Load\n          (EL) measured in grams for 10 individuals:\n                           Brain Resonance Index Ectoplasmic Load\n             Individual\n                                    (X)                 (Y)\n                  A                169.6               71.2\n                  B                166.8               58.2\n                  C                157.1                 56\n                  D                181.1               64.5\n                  E                158.4                 53\n                  F                165.6               52.4\n                  G                166.7               56.8\n                  H                156.5               49.2\n                  I                168.1               55.6\n                  J                165.3               77.8\n\n           Use the following R code to load the data in R.\n\n       BRI <- c(169.6,166.8,157.1,181.1,158.4,165.6,166.7,156.5,168.1,165.3)\n       EL <- c(71.2,58.2,56.0,64.5,53.0,52.4,56.8,49.2,55.6,77.8)\n\n       Note: You can directly copy these code lines from Codes.docx file in your R Script.",
      "parts": [
        {
          "label": "i",
          "marks": 3,
          "text": "Draw a labelled scatterplot of ectoplasmic load on the vertical axis versus brain\n              resonance index on the horizontal axis. Comment on the scatterplot.           [3]",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 2,
          "text": "Determine the sample correlation coefficient using the Karl Pearson’s Method.                [2]",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 4,
          "text": "Fit a regression of Y on X. Print the summary of the regression model in R. Also state the\n                estimated standard errors for slope and intercept parameters.                          [4]",
          "topic": null
        },
        {
          "label": "iv",
          "marks": 2,
          "text": "Draw the fitted line on your scatterplot.                                                    [2]",
          "topic": null
        },
        {
          "label": "v",
          "marks": 2,
          "text": "Predict the ectoplasmic load of an individual whose brain resonance index is 160.4.           [2]",
          "topic": null
        },
        {
          "label": "vi",
          "marks": 2,
          "text": "Estimate the covariance between the slope and intercept parameter of the regression\n               model.\n               Hint: [Use vcov function in R.]                                                 [2]",
          "topic": null
        },
        {
          "label": "vii",
          "marks": 4,
          "text": "Some researchers still believe that ectoplasmic load does not depend on brain resonance\n                index. By calculating a two-sided 95% confidence interval for the slope parameter of the\n                regression line, test this claim of the researchers. You are required to clearly state the null\n                and alternative hypothesis, the confidence interval and the conclusion of the test.         [4]",
          "topic": null
        },
        {
          "label": "viii",
          "marks": 4,
          "text": "Some others believe that instead of testing for the slope coefficient, one should test\n                whether there exists any non-zero correlation between the two variables. Perform the test\n                at 5% level of significance. Clearly state the null and alternative hypotheses, p-value and\n                the conclusion of the test.                                                             [4]",
          "topic": null
        },
        {
          "label": "ix",
          "marks": 25,
          "text": "“If a particular dataset passes the test in part (vii), then passing the test in part (viii) is a\n               foregone conclusion.” Briefly comment on the statement giving suitable reasons to justify\n               your answer.                                                                                  [2]",
          "topic": null
        }
      ],
      "solution": null,
      "has_math": false,
      "is_r_task": true,
      "session": "2026-05",
      "subject": "CS1B",
      "source_qp": "raw/CS1B_2026-05_QP.pdf",
      "source_sol": null
    },
    {
      "q_num": 3,
      "marks": 50,
      "topic": "regression_glm",
      "subtopics": [],
      "stem": "An insurance company wants to analyse its motor insurance portfolio over a period of 7\n          years. For each policyholder, the number of insurance claims during the observation period is\n          recorded and is available in the file MotorClaims.csv.\n          The drivers in the data are classified according to the following risk factors:\n           •   Sex - of the driver (1= Male, 2=Female)\n           •   Region – of residence (1 = Urban, 2 = Semi-Urban, 3 = Rural))\n           •   Car_type – type of car insured (1 = Hatchback,2 = Sedan,3 = SUV)\n\n           •   Job – exposure to driving based on job (1= Low, 2 = Medium, 3= High)",
      "parts": [
        {
          "label": "i",
          "marks": 2,
          "text": "Load the data and using lapply() or otherwise convert the risk factors into factors in R.   [2]",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 4,
          "text": "Fit a Generalised Linear Model (GLM) ‘model_1’ to the data assuming a Poisson distribution\n           with a log link function, using claims as the response variable and sex, region, job, and\n           car_type as explanatory variables, without interaction terms. Print the summary of this model\n           in R.                                                                                     [4]",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 2,
          "text": "Comment on the significance of each risk factor.                                          [2]",
          "topic": null
        },
        {
          "label": "iv",
          "marks": 3,
          "text": "Estimate the expected number of claims for a male policy holder driving a Sedan in Urban\n           area with high exposure to driving.                                                  [3]",
          "topic": null
        },
        {
          "label": "v",
          "marks": 3,
          "text": "Due to regulatory reasons, sex cannot be used as a risk factor. Refit a new model ‘model_2’\n          removing sex as risk factor and print the summary of this model.                       [3]",
          "topic": null
        },
        {
          "label": "vi",
          "marks": 1,
          "text": "State with reason which risk factor(s) can be further removed from the model_2.            [1]",
          "topic": null
        },
        {
          "label": "vii",
          "marks": 3,
          "text": "Fit a new model ‘model_3’ removing the risk factor(s) identified in part (vi) and print the\n            summary of the new model.                                                               [3]",
          "topic": null
        },
        {
          "label": "viii",
          "marks": 3,
          "text": "The analyst now wishes to investigate whether there is any interaction between the remaining\n            significant risk factors.\n          Refit model_3 by including interaction between the significant risk factors. Label this as\n          model_4. Print the summary of this model.                                              [3]",
          "topic": null
        },
        {
          "label": "ix",
          "marks": 2,
          "text": "Following is a QQ-plot of the deviance residuals for model_4.\n\n       Briefly comment on the above plot in the context of the goodness of fit of model_4.            [2]",
          "topic": null
        },
        {
          "label": "x",
          "marks": 25,
          "text": "Using Akaike Information Criterion (AIC), comment on which model is better fit among the\n          four models.                                                                         [2]",
          "topic": null
        }
      ],
      "solution": null,
      "has_math": false,
      "is_r_task": true,
      "session": "2026-05",
      "subject": "CS1B",
      "source_qp": "raw/CS1B_2026-05_QP.pdf",
      "source_sol": null
    },
    {
      "q_num": 4,
      "marks": 30,
      "topic": "distributions",
      "subtopics": [],
      "stem": "The number of car accidents at a fixed point on a road were recorded for 20 consecutive\n      months. The results are as shown below, with 𝑥𝑖 being the number of accidents at the fixed\n      point in the 𝑖th month\n       Note: You can directly copy these code lines from Codes.docx file in your R Script.\n\n          Month        1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19                              20\n           No of\n                       2 2 1 1 0 4 2 1 2                  1   1   1   3   1   2   2   3   2    3    4\n         accidents\n\n        x <- c(2,2,1,1,0,4,2,1,2,1,1,1,3,1,2,2,3,2,3,4)\n\n       X is assumed to follow a Poisson distribution with parameter 𝜆 with probability function given\n       as –\n       P (X = x) = λx * exp(-λ) / x!",
      "parts": [
        {
          "label": "i",
          "marks": 2,
          "text": "Calculate the sample mean and sample variance for the number of car accidents.                 [2]",
          "topic": null
        },
        {
          "label": "ii",
          "marks": 1,
          "text": "Create and store a sequence of 1000 values for the population parameter 𝜆 defining it as\n           vector lambda assuming plausible values between 1 and 3.\n           You are NOT required to print these values.                                                   [1]",
          "topic": null
        },
        {
          "label": "iii",
          "marks": 2,
          "text": "Likelihood function for the Poisson distribution can be coded in R using the following code:\n\n       lik <- (lambda^(sum(x))*exp(-length(x)*lambda))/prod(factorial(x))\n\n       Note: You can directly copy these code lines from Codes.docx file in your R Script.\n\n       Calculate the maximum value of likelihood function based on the values of 𝜆 generated in part\n       (ii) and the value of 𝜆 corresponding to it.                                                      [2]",
          "topic": null
        },
        {
          "label": "iv",
          "marks": 3,
          "text": "Calculate the maximum value of the log likelihood function based on the values of 𝜆\n           generated in part (ii) and the value of 𝜆 corresponding to it.                  [3]",
          "topic": null
        },
        {
          "label": "v",
          "marks": 2,
          "text": "By first principles, state what will be the maximum likelihood estimate (MLE) of the Poisson\n          parameter 𝜆. Assume that the second order derivative is less than zero thus leading to\n          maxima.                                                                                  [2]",
          "topic": null
        },
        {
          "label": "vi",
          "marks": 2,
          "text": "Plot suitable graph for the likelihood function against plausible range of 𝜆. Also, add a line\n           for MLE in the above graph.                                                                [2]",
          "topic": null
        },
        {
          "label": "vii",
          "marks": 2,
          "text": "Based on the plot in part (vi) and based on the maximum value calculated in part (iii), show\n            that the maximum likelihood estimate of 𝜆 is the point where the likelihood function is\n            maximized.                                                                                [2]",
          "topic": null
        },
        {
          "label": "viii",
          "marks": 2,
          "text": "Cramér–Rao Lower Bound (CRLB) for the parameter 𝜆 of this distribution is given as 𝜆MLE /\n            n.\n           Calculate an approximate 95% confidence interval for 𝜆 using the normal distribution with\n           mean 𝜆MLE and variance given by the CRLB.                                             [2]",
          "topic": null
        },
        {
          "label": "ix",
          "marks": 3,
          "text": "Calculate an exact two-sided 95% confidence interval for the Poisson parameter 𝜆 and\n          compare with the confidence interval calculated in part (viii).                  [3]",
          "topic": null
        },
        {
          "label": "x",
          "marks": 3,
          "text": "Using the estimated value of 𝜆 from part (v), calculate the probability that:\n         a) No accident occurs in a month\n         b) More than 5 accidents happen in a month\n         c) Exactly 7 accidents happen in a month",
          "topic": null
        },
        {
          "label": "xi",
          "marks": 2026,
          "text": "Using the estimated value of the parameter 𝜆 from part (v), generate 1,000 random samples\n          containing 100 values each from a Poisson distribution.\n         In order to initialise the random number generator, use set.seed(2026).\n         You are NOT required to display the output.                                                   [4]",
          "topic": null
        },
        {
          "label": "xii",
          "marks": 30,
          "text": "Calculate the values of E(X_bar) and E (S2) based on the random samples generated in part\n           (xi).\n         Hence, show that X_bar and S2 are unbiased estimators of the population mean and\n         population variance respectively.",
          "topic": null
        }
      ],
      "solution": null,
      "has_math": true,
      "is_r_task": true,
      "session": "2026-05",
      "subject": "CS1B",
      "source_qp": "raw/CS1B_2026-05_QP.pdf",
      "source_sol": null
    }
  ]
}