Marks-versus-Percentile Data: A New Method of Analysis to Determine the Preparation Level of Students
Sudipto Roy
Department of Physics, St. Xavier’s College, Kolkata, West Bengal, India
Email: roy.sudipto@sxccal.edu
ORCID URL: https://orcid.org/0000-0002-8811-2511
ABSTRACT
The aim of this study is to formulate a new method for an analysis of examination-related data (marks vs. percentile) to determine the average preparation level of the examinees. For this purpose, we have used a small set of marks vs. percentile data (obtained from the internet) of the Joint Entrance Examination – Main [JEE (Main)] conducted in April 2024. An empirical expression has been chosen here to represent the percentile score as a function of marks. Based on this function, we have derived an expression for the probability of scoring any marks and shown its behaviour graphically. The most probable value, the expectation value and the standard deviation of marks have been calculated for the examination. We have shown how to calculate the probability of scoring marks equal to or greater than any marks, and depicted its behaviour graphically. The method of data analysis of this study can be used for any examination, not only to evaluate the average preparation level of the examinees but also to be used as a tool by the examinees of the next session of the examination to predict the percentile score based on marks. The accuracy of such predictions depends upon the size of the data sample used to determine the constant parameters associated with the empirical expression representing the percentile score as a function of marks.
Keywords: Percentile Score, JEE (Main), NTA, Most Probable Marks, Marks-versus-Percentile
INTRODUCTION
Percentile score is a measure of one’s performance relative to the rest of the candidates who have taken a certain examination. Percentile scores are also calculated to evaluate journals, books, research papers and researchers, in various ways, to obtain relative measures of performance [1-4]. The merit of a publication can be judged by finding the percentage of publications (in the same field) which have received equal and less number of citations. Apart from the marks scored in an examination, one often needs to know the percentile score to find his/her chance to get admission to different institutions. There should be a mathematical way to determine the percentile score when one already knows the marks obtained by an examinee. In a study carried out in 2023, a function was mathematically derived (based on probabilistic considerations), using which the percentile score of an examinee can be calculated from marks obtained in an examination [5].
In the present study, we have attempted to find a much simpler way to analyze marks vs. percentile data with the help of an empirical function. Apart from calculating the percentile score, it might often be necessary for a candidate to estimate the chance of scoring marks beyond a certain percentage of the full marks for an examination. It can be done only if there is a function available to them, using which, one can calculate the probability of scoring any marks. To formulate a function of that kind we have used a set of marks vs. percentile data. Based on the behavior of how the percentile score increases with marks, we have proposed here an empirical formula which is meant to serve as an expression for the percentile score as a function of marks. For this formula to work properly in certain year, the constant parameters of the formula have to be determined using a reasonably large dataset of marks vs. percentile of the previous year. We have used a small dataset regarding the results of an examination named JEE (Main) conducted by the National Testing Agency (NTA) [6-8]. We have determined the parameter values required for a best fit of our empirical formula to the dataset. Based on this formula, a function f(m) has been derived, which is the probability of obtaining m marks in the examination. The values of f(m) as a function of m has given us the most probable marks (m) and the corresponding probability. Using f(m), one can estimate the probability of getting marks equal to or greater than a certain value of m. An important finding of the present study is that the sum of f(m), over all possible values of m, has come out to be very close to unity, confirming its validity as representing probability, although we have used a small dataset to formulate this function. We have shown how the expectation value and standard deviation can be calculated from marks vs. percentile data. The method of data analysis is very simple and it can be used to make predictions with sufficient accuracy if one uses a large volume of data corresponding to previous years’ results of the examination.
METHODOLOGY
Let us assume that there exists a function P(m) representing the percentile score of a candidate who has scored m marks in an examination. The definition of percentile score is as follows [5].
█(P(m)=(number of candidates with marks≤m)/(total number of candidates)×100 #(1) )
Let f(m) be the fraction of examinees who have scored m marks each. This fraction can be expressed as,
█(f(m)=(number of candidates with marks=m)/(total number of candidates) #(2) )
Based on the definition of P(m), represented by equation (1), f(m) can be expressed as,
█(f(m)=(P(m)-P(m-β))/100 #(3) )
In writing equation (3), we have assumed that the variable m changes in steps of β. Thus, (m-β) is the possible score just below m. Equation (3) cannot be the definition of f(m) when m is the lowest score obtainable in the examination. In that case, we have, f(m)=P(m)/100, which is valid according to equations (1) and (2). If the number of candidates taking the examination is sufficiently large, the fraction of candidates, each having m marks, can approximately be regarded as the probability of scoring that marks in the examination. In the present study, we have used the function f(m) as the probability that an examinee scores m marks.
The probability of obtaining marks between any two values of m, such as q and r (i.e., q≤m≤r), can be expressed as [9-12],
█(P(q,r)=∑_(m=q)^r▒f(m) #(4) )
Let g(m) be the probability that a candidate scores marks ≥m in the examination. It can be expressed as [9-12],
█(g(m)=∑_(x=m)^(m_max)▒f(x) #(5) )
In equation (5), m_max is the highest marks obtainable in the examination.
The expectation value of marks, denoted by m_ex here, is expressed as [9-12],
█(m_ex=∑_(m=m_min)^(m_max)▒〖m f(m) 〗 #(6) )
In equation (6), m_min stands for the lowest marks obtainable in the examination.
The standard deviation of marks, denoted here by m_sd, is given by [9-12],
█(m_sd=[∑_(m=m_min)^(m_max)▒〖(m-m_ex )^2 f(m) 〗]^(1/2) #(7) )
A simpler formula for m_sd is given below [9-12].
█(m_sd=[(∑_(m=m_min)^(m_max)▒〖m^2 f(m) 〗)-(m_ex )^2 ]^(1/2) #(8) )
DATA ANALYSIS
For the present study, we have used a set of marks vs. percentile data of the examination named JEE (Main), held in April 2024, obtained from a certain website [8]. Table-1 of this article contains that dataset.
Table 1. Marks vs. Percentile Data of April 2024 session of JEE (Main)
Marks Percentile Marks Percentile
82 80 158 96.5
96 85 163 97
113 90 170 97.5
119 91 177 98
124 92 188 98.5
130 93 199 99
137 94 220 99.5
145 95 226 99.6
149 95.5 234 99.7
153 96 271 99.8
According to this table, the percentile score increases by 13 for an increase in marks from 82 to 130. For almost the same change in marks, i.e., from 130 to 177, the percentile score increases by 5. Based on this nature of dependence of the percentile score upon marks, we have assumed the following empirical expression to represent the percentile score P(m) as a function of marks (m).
█(P(m)=A[1-Exp{-b(m+c)}]^n#(9) )
The JEE (Main) question paper has 75 questions carrying 4 marks each, the full marks being 300. Each wrong answer causes a deduction of 1 mark. Therefore, the value of the parameter β (of eqn. 3) is 1.
The value of P(m) should be positive and it should be increasing with marks (m), as per equation (1). Due to negative marking, the lowest value of P(m) (i.e., zero) should correspond to a negative value of marks (m). For these conditions to be satisfied, we must have A,b,c,n> 0, in equation (9).
Since the maximum marks obtainable in the JEE (Mains) is 300 (the full marks), equation (9) must yield P(m)=100 (the highest percentile score) for m=300. This fact leads to the following value for the constant parameter A which belongs to the expression for P(m) (eqn. 9).
█(A=100/[1-Exp{-b(300+c)}]^n #(10) )
As per equation (9), P(m)=0 for m=-c. It means that, there is no examinee for whom m≤-c. Thus, the lowest marks scored is, m_min=-c+1, since m changes in steps of 1. For any analysis using equation (9), one should consider the values of m lying in the range: -c+1≤m≤300.
The constant parameter b in equation (9) determines how rapidly P(m) approaches its maximum value (i.e., 100, the highest percentile score) as m increases. It has been observed that, as the difficulty level rises, one gets a greater percentile score for the same marks secured in the examination. This is due to a fall in the number of candidates scoring higher marks. The value of P(m), for a certain m, decreases as the parameter n increases. Thus, both b & n need to be changed, as the difficulty level changes.
Using equation (10) in equation (9), one gets,
█(P(m)=100 [1-Exp{-b(m+c)}]^n/[1-Exp{-b(300+c)}]^n #(11) )
Using equation (11) in equation (3), we can write,
█(f(m)=([1-Exp{-b(m+c)}]^n-[1-Exp{-b(m-1+c)}]^n)/[1-Exp{-b(300+c)}]^n #(12) )
The number of examinees for each session of JEE (Main) is close to 1 million. Due this large sample size, the function f(m) can be regarded as the probability that an examinee obtains m marks in the examination. Here, m can be treated as a discrete random variable, with values in the range: -c+1≤m≤300.
Using equation (12) in equation (4), we get the following expression for P(q,r),
█(P(q,r)=∑_(m=q)^r▒([1-Exp{-b(m+c)}]^n-[1-Exp{-b(m-1+c)}]^n)/[1-Exp{-b(300+c)}]^n #(13) )
Using equation (12) in equation (5), we get the following expression for g(m),
█(g(m)=∑_(x=m)^300▒([1-Exp{-b(x+c)}]^n-[1-Exp{-b(x-1+c)}]^n)/[1-Exp{-b(300+c)}]^n #(14) )
Using equation (12) in equation (6), we get the following expression for m_ex,
█(m_ex=∑_(m=-c+1)^300▒〖m [([1-Exp{-b(m+c)}]^n-[1-Exp{-b(m-1+c)}]^n)/[1-Exp{-b(300+c)}]^n ] 〗 #(15) )
Using equations (12) and (15) in equation (8), one gets the following expression for m_sd,
█(m_sd=[█((∑_(m=-c+1)^300▒〖m^2 ([1-Exp{-b(m+c)}]^n-[1-Exp{-b(m-1+c)}]^n)/[1-Exp{-b(300+c)}]^n 〗)-@(∑_(m=-c+1)^300▒〖m [([1-Exp{-b(m+c)}]^n-[1-Exp{-b(m-1+c)}]^n)/[1-Exp{-b(300+c)}]^n ] 〗)^2 )]^(1/2) #(16) )
RESULTS AND DISCUSSION
Fitting equation (11) to the marks vs. percentile dataset of Table-1, we have obtained, b=0.02344, c=46 and n=4.44651. For this curve fitting, the coefficient of determination is, R^2=0.99894. Using the values of b, c and n, one obtains A=100.134 from equation (10). For m=-c≡-46, we have P(m)=0 (as per eqn. 11), which implies that no examinee has obtained marks ≤-46. Thus, m_min=-c+1=-45. For no analysis, one should consider the values of m for which m≤m_min-1 or, m≤-46.
Figure 1 shows the best fit of the function P(m) (eqn. 11) to the dataset of Table-1. Figure 2 shows the same P(m)-versus-m curve for the entire range of values of m.
Figure 3 shows the variation of the probability f(m) as a function of marks (m), based on equation (12). It shows that a candidate has the highest probability of scoring 19 marks and the corresponding probability is close to 0.00981.
Figure 4 shows the variation of the cumulative probability g(m) as a function of marks (m), based on equation (14). It seems to be falling at a faster rate beyond m=250. For the clarity of data, we have chosen logarithmic scales for showing the values of both f(m) and g(m), in Figures 3 and 4 respectively.
Using equation (13), the probability of scoring marks between 100 & 200 (including these two values) comes out to be P(100,200)=0.126371, which means that 12.6371% of the examinees are likely to score marks in this range.
Figure 1. Percentile score versus marks. The blue curve is the best fit to the Table-1 data (circles), based on the function P(m) of equation (11). Figure 2. Percentile score versus marks, using equation (11), over the entire range of m, for the parameter values required for the best fit in Fig. 1.
Using equations (15) and (16), respectively, we have obtained m_ex=47.26 and m_SD=49.74.
If the function f(m) (of eqn. 12) has to serve as the probability of scoring m marks in the examination, its sum over all values of m should be unity. Our calculation shows, ∑_(m=-45)^300▒〖f(m)〗=1.000002. The deviation of the sum from unity may be due to the smallness of the dataset (Table-1) used in this study. If we had used a sufficiently large dataset (i.e., a set of marks vs. percentile data having almost all possible values of m), we could probably have obtained a sum closer to unity.
The expectation value, the standard deviation and the most probable value of m, constitute a set of numbers which may collectively be regarded as an index for the average preparation level of the examinees who took the JEE (Main) in the session of April 2024.
One can use equations (11) to (16) for making predictions for any session of the JEE (Main) with an accuracy which depends upon the values of the constant parameters (b, c & n) associated with P(m). To determine their values properly, one must use a reasonably large data of marks vs. percentile.
The coefficient of variation (or CV, which is, standard deviation / mean) can be regarded as an index of the overall performance of the examinees. The smaller the value of CV, the better would be the collective performance of the candidates who have taken the examination. For the JEE (Main) of April 2023, the mean and standard deviation of marks are respectively 47.26 and 49.74. Therefore, CV=49.74/47.26=1.052 or 105.2%, which is quite large, indicating a poor collective performance by the examinees.
Figure 3. Probability versus marks. The highest probability corresponds to m=19. The scale along the vertical axis is logarithmic. Figure 4. Cumulative probability versus marks. g(m) is the probability of scoring marks ≥m. The scale along the vertical axis is logarithmic.
CONCLUDING REMARKS
In the present study, we have formulated a new method to determine the average preparation level of the examinees, based on the marks vs. percentile data of an examination. The accuracy of this determination depends upon the volume of data used for this purpose. Unfortunately, the data available to us through the internet regarding the JEE (Main) were inadequate. But there might be examinations whose entire marks vs. percentile data are available online. The same analysis can be carried out based on them, presumably with a greater accuracy. A candidate, who has somehow calculated the marks to be obtained, can estimate the percentile score using the expression for P(m), if the values of the parameters b, c and n have already been determined based on a sufficiently large volume of data obtained from previous years’ results. The empirical expression for P(m), chosen for the present study (eqn. 9), can be replaced by a new function. One can do the entire analysis using that function and compare the findings with those of the present study.
Conflicts of Interest: The author declares no conflict of interest.
REFERENCES
Waltman, L., & Schreiber, M. (2013). On the calculation of percentile-based bibliometric indicators. Journal of the Association for Information Science and Technology, 64(2), 372–379. https://doi.org/10.1002/asi.22775
Bornmann, L., & Williams, R. (2020). An evaluation of percentile measures of citation impact, and a proposal for making them better. Scientometrics, 124(2), 1457–1478. https://doi.org/10.1007/s11192-020-03512-7
Bornmann, L., Leydesdorff, L., & Mutz, R. (2013). The use of percentiles and percentile rank classes in the analysis of bibliometric data: Opportunities and limits. Journal of Informetrics, 7(1), 158–165. https://doi.org/10.1016/j.joi.2012.10.001
Yamamoto, K., & Yasunaga, T. (2022). A percentile rank score of group productivity: An evaluation of publication productivity for researchers from various fields. Scientometrics, 127(4), 1737–1754. https://doi.org/10.1007/s11192-022-04278-w
Roy, S. (2023). Estimation of percentile score based on marks obtained in an examination: A simple mathematical model. International Journal of Physics and Mathematics, 5(1), 25–32. https://doi.org/10.33545/26648636.2023.v5.i1a.48
National Testing Agency. (n.d.). National Testing Agency. https://nta.ac.in/
National Testing Agency. (n.d.). Joint Entrance Examination (Main). https://jeemain.nta.ac.in/
CollegeDunia. (n.d.). JEE Main rank vs marks 2024. https://collegedunia.com/exams/jee-main/rank-vs-marks
Boas, M. L. (2006). Mathematical methods in the physical sciences (3rd ed.). John Wiley & Sons.
Arfken, G. B., Weber, H. J., & Harris, F. E. (2011). Mathematical methods for physicists: A comprehensive guide (7th ed.). Academic Press.
Riley, K. F., Hobson, M. P., & Bence, S. J. (2006). Mathematical methods for physics and engineering: A comprehensive guide (3rd ed.). Cambridge University Press.
Ross, S. M. (1976). A first course in probability. Macmillan.
Get EduPub Publication Services Now