githubtemp5 2 years ago
parent cbe0dc3d7e
commit 08b59dab18

File diff suppressed because it is too large Load Diff

File diff suppressed because one or more lines are too long

@ -0,0 +1,620 @@
{
"cells": [
{
"cell_type": "markdown",
"metadata": {},
"source": [
"# Exploring data with Python - real world data\n",
"\n",
"In the last notebook, we looked at grades for our student data and investigated the data visually with histograms and box plots. Now we'll look into more complex cases, describe the data more fully, and discuss how to make basic comparisons between data.\n",
"\n",
"### Real world data distributions\n",
"\n",
"Previously, we looked at grades for our student data and estimated from this sample what the full population of grades might look like. Let's refresh our memory and take a look at this data again.\n",
"\n",
"Run the following code to print out the data and make a histogram plus box plot that shows the grades for our sample of students."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"import pandas as pd\n",
"from matplotlib import pyplot as plt\n",
"\n",
"# Load data from a text file\n",
"!wget https://raw.githubusercontent.com/MicrosoftDocs/mslearn-introduction-to-machine-learning/main/Data/ml-basics/grades.csv\n",
"df_students = pd.read_csv('grades.csv',delimiter=',',header='infer')\n",
"\n",
"# Remove any rows with missing data\n",
"df_students = df_students.dropna(axis=0, how='any')\n",
"\n",
"# Calculate who passed, assuming '60' is the grade needed to pass\n",
"passes = pd.Series(df_students['Grade'] >= 60)\n",
"\n",
"# Save who passed to the Pandas dataframe\n",
"df_students = pd.concat([df_students, passes.rename(\"Pass\")], axis=1)\n",
"\n",
"\n",
"# Print the result out into this notebook\n",
"print(df_students)\n",
"\n",
"\n",
"# Create a function that we can re-use\n",
"def show_distribution(var_data):\n",
" '''\n",
" This function will make a distribution (graph) and display it\n",
" '''\n",
"\n",
" # Get statistics\n",
" min_val = var_data.min()\n",
" max_val = var_data.max()\n",
" mean_val = var_data.mean()\n",
" med_val = var_data.median()\n",
" mod_val = var_data.mode()[0]\n",
"\n",
" print('Minimum:{:.2f}\\nMean:{:.2f}\\nMedian:{:.2f}\\nMode:{:.2f}\\nMaximum:{:.2f}\\n'.format(min_val,\n",
" mean_val,\n",
" med_val,\n",
" mod_val,\n",
" max_val))\n",
"\n",
" # Create a figure for 2 subplots (2 rows, 1 column)\n",
" fig, ax = plt.subplots(2, 1, figsize = (10,4))\n",
"\n",
" # Plot the histogram \n",
" ax[0].hist(var_data)\n",
" ax[0].set_ylabel('Frequency')\n",
"\n",
" # Add lines for the mean, median, and mode\n",
" ax[0].axvline(x=min_val, color = 'gray', linestyle='dashed', linewidth = 2)\n",
" ax[0].axvline(x=mean_val, color = 'cyan', linestyle='dashed', linewidth = 2)\n",
" ax[0].axvline(x=med_val, color = 'red', linestyle='dashed', linewidth = 2)\n",
" ax[0].axvline(x=mod_val, color = 'yellow', linestyle='dashed', linewidth = 2)\n",
" ax[0].axvline(x=max_val, color = 'gray', linestyle='dashed', linewidth = 2)\n",
"\n",
" # Plot the boxplot \n",
" ax[1].boxplot(var_data, vert=False)\n",
" ax[1].set_xlabel('Value')\n",
"\n",
" # Add a title to the Figure\n",
" fig.suptitle('Data Distribution')\n",
"\n",
" # Show the figure\n",
" fig.show()\n",
"\n",
"\n",
"show_distribution(df_students['Grade'])"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"As you might recall, our data had the mean and mode at the center, with data spread symmetrically from there.\n",
"\n",
"Now let's take a look at the distribution of the study hours data."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"tags": []
},
"outputs": [],
"source": [
"# Get the variable to examine\n",
"col = df_students['StudyHours']\n",
"# Call the function\n",
"show_distribution(col)"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"The distribution of the study time data is significantly different from that of the grades.\n",
"\n",
"Note that the whiskers of the box plot only begin at around 6.0, indicating that the vast majority of the first quarter of the data is above this value. The minimum is marked with an **o**, indicating that it is statistically an *outlier*: a value that lies significantly outside the range of the rest of the distribution.\n",
"\n",
"Outliers can occur for many reasons. Maybe a student meant to record \"10\" hours of study time, but entered \"1\" and missed the \"0\". Or maybe the student was abnormally lazy when it comes to studying! Either way, it's a statistical anomaly that doesn't represent a typical student. Let's see what the distribution looks like without it."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"tags": []
},
"outputs": [],
"source": [
"# Get the variable to examine\n",
"# We will only get students who have studied more than one hour\n",
"col = df_students[df_students.StudyHours>1]['StudyHours']\n",
"\n",
"# Call the function\n",
"show_distribution(col)"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"For learning purposes, we've just treated the value **1** as a true outlier here and excluded it. In the real world, it would be unusual to exclude data at the extremes without more justification when our sample size is so small. This is because the smaller our sample size, the more likely it is that our sampling is a bad representation of the whole population. (Here, the population means grades for all students, not just our 22.) For example, if we sampled study time for another 1,000 students, we might find that it's actually quite common to not study much!\n",
"\n",
"When we have more data available, our sample becomes more reliable. This makes it easier to consider outliers as being values that fall below or above percentiles within which most of the data lie. For example, the following code uses the Pandas **quantile** function to exclude observations below the 0.01th percentile (the value above which 99% of the data reside)."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"# calculate the 0.01th percentile\n",
"q01 = df_students.StudyHours.quantile(0.01)\n",
"# Get the variable to examine\n",
"col = df_students[df_students.StudyHours>q01]['StudyHours']\n",
"# Call the function\n",
"show_distribution(col)"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"> **Tip**: You can also eliminate outliers at the upper end of the distribution by defining a threshold at a high percentile value. For example, you could use the **quantile** function to find the 0.99 percentile, below which 99% of the data reside.\n",
"\n",
"With the outliers removed, the box plot shows all data within the four quartiles. Note that the distribution is not symmetric like it is for the grade data. There are some students with very high study times of around 16 hours, but the bulk of the data is between 7 and 13 hours. The few extremely high values pull the mean towards the higher end of the scale.\n",
"\n",
"Let's look at the density for this distribution."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"def show_density(var_data):\n",
" fig = plt.figure(figsize=(10,4))\n",
"\n",
" # Plot density\n",
" var_data.plot.density()\n",
"\n",
" # Add titles and labels\n",
" plt.title('Data Density')\n",
"\n",
" # Show the mean, median, and mode\n",
" plt.axvline(x=var_data.mean(), color = 'cyan', linestyle='dashed', linewidth = 2)\n",
" plt.axvline(x=var_data.median(), color = 'red', linestyle='dashed', linewidth = 2)\n",
" plt.axvline(x=var_data.mode()[0], color = 'yellow', linestyle='dashed', linewidth = 2)\n",
"\n",
" # Show the figure\n",
" plt.show()\n",
"\n",
"# Get the density of StudyHours\n",
"show_density(col)"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"This kind of distribution is called *right skewed*. The mass of the data is on the left side of the distribution, creating a long tail to the right because of the values at the extreme high end, which pull the mean to the right.\n",
"\n",
"#### Measures of variance\n",
"\n",
"So now we have a good idea where the middle of the grade and study hours data distributions are. However, there's another aspect of the distributions we should examine: how much variability is there in the data?\n",
"\n",
"Typical statistics that measure variability in the data include:\n",
"\n",
"- **Range**: The difference between the maximum and minimum. There's no built-in function for this, but it's easy to calculate using the **min** and **max** functions.\n",
"- **Variance**: The average of the squared difference from the mean. You can use the built-in **var** function to find this.\n",
"- **Standard Deviation**: The square root of the variance. You can use the built-in **std** function to find this."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"tags": []
},
"outputs": [],
"source": [
"for col_name in ['Grade','StudyHours']:\n",
" col = df_students[col_name]\n",
" rng = col.max() - col.min()\n",
" var = col.var()\n",
" std = col.std()\n",
" print('\\n{}:\\n - Range: {:.2f}\\n - Variance: {:.2f}\\n - Std.Dev: {:.2f}'.format(col_name, rng, var, std))"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"Of these statistics, the standard deviation is generally the most useful. It provides a measure of variance in the data on the same scale as the data itself (so grade points for the Grade distribution and hours for the StudyHours distribution). The higher the standard deviation, the more variance there is when comparing values in the distribution to the distribution mean; in other words, the data is more spread out.\n",
"\n",
"When working with a *normal* distribution, the standard deviation works with the particular characteristics of a normal distribution to provide even greater insight. Run the following cell to see the relationship between standard deviations and the data in the normal distribution."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"import scipy.stats as stats\n",
"\n",
"# Get the Grade column\n",
"col = df_students['Grade']\n",
"\n",
"# get the density\n",
"density = stats.gaussian_kde(col)\n",
"\n",
"# Plot the density\n",
"col.plot.density()\n",
"\n",
"# Get the mean and standard deviation\n",
"s = col.std()\n",
"m = col.mean()\n",
"\n",
"# Annotate 1 stdev\n",
"x1 = [m-s, m+s]\n",
"y1 = density(x1)\n",
"plt.plot(x1,y1, color='magenta')\n",
"plt.annotate('1 std (68.26%)', (x1[1],y1[1]))\n",
"\n",
"# Annotate 2 stdevs\n",
"x2 = [m-(s*2), m+(s*2)]\n",
"y2 = density(x2)\n",
"plt.plot(x2,y2, color='green')\n",
"plt.annotate('2 std (95.45%)', (x2[1],y2[1]))\n",
"\n",
"# Annotate 3 stdevs\n",
"x3 = [m-(s*3), m+(s*3)]\n",
"y3 = density(x3)\n",
"plt.plot(x3,y3, color='orange')\n",
"plt.annotate('3 std (99.73%)', (x3[1],y3[1]))\n",
"\n",
"# Show the location of the mean\n",
"plt.axvline(col.mean(), color='cyan', linestyle='dashed', linewidth=1)\n",
"\n",
"plt.axis('off')\n",
"\n",
"plt.show()"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"The horizontal lines show the percentage of data within one, two, and three standard deviations of the mean (plus or minus).\n",
"\n",
"In any normal distribution:\n",
"- Approximately 68.26% of values fall within one standard deviation from the mean.\n",
"- Approximately 95.45% of values fall within two standard deviations from the mean.\n",
"- Approximately 99.73% of values fall within three standard deviations from the mean.\n",
"\n",
"So, because we know that the mean grade is 49.18, the standard deviation is 21.74, and distribution of grades is approximately normal, we can calculate that 68.26% of students should achieve a grade between 27.44 and 70.92.\n",
"\n",
"The descriptive statistics we've used to understand the distribution of the student data variables are the basis of statistical analysis. Because they're such an important part of exploring your data, there's a built-in `describe` method of the DataFrame object that returns the main descriptive statistics for all numeric columns."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"df_students.describe()"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Comparing data\n",
"\n",
"Now that you know something about the statistical distribution of the data in your dataset, you're ready to examine your data to identify any apparent relationships between variables.\n",
"\n",
"First of all, let's get rid of any rows that contain outliers so that we have a sample that is representative of a typical class of students. We identified that the StudyHours column contains some outliers with extremely low values, so we'll remove those rows."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"df_sample = df_students[df_students['StudyHours']>1]\n",
"df_sample"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"### Comparing numeric and categorical variables\n",
"\n",
"The data includes two *numeric* variables (**StudyHours** and **Grade**) and two *categorical* variables (**Name** and **Pass**). Let's start by comparing the numeric **StudyHours** column to the categorical **Pass** column to see if there's an apparent relationship between the number of hours studied and a passing grade.\n",
"\n",
"To make this comparison, let's create box plots showing the distribution of StudyHours for each possible Pass value (true and false)."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"df_sample.boxplot(column='StudyHours', by='Pass', figsize=(8,5))"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"Comparing the StudyHours distributions, it's immediately apparent (if not particularly surprising) that students who passed the course tended to study for more hours than students who didn't. So if you wanted to predict whether or not a student is likely to pass the course, the amount of time they spend studying may be a good predictive indicator.\n",
"\n",
"### Comparing numeric variables\n",
"\n",
"Now let's compare two numeric variables. We'll start by creating a bar chart that shows both grade and study hours."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"# Create a bar plot of name vs grade and study hours\n",
"df_sample.plot(x='Name', y=['Grade','StudyHours'], kind='bar', figsize=(8,5))"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"The chart shows bars for both grade and study hours for each student, but it's not easy to compare because the values are on different scales. A grade is measured in grade points (and ranges from 3 to 97), and study time is measured in hours (and ranges from 1 to 16).\n",
"\n",
"A common technique when dealing with numeric data in different scales is to *normalize* the data so that the values retain their proportional distribution but are measured on the same scale. To accomplish this, we'll use a technique called *MinMax* scaling that distributes the values proportionally on a scale of 0 to 1. You could write the code to apply this transformation, but the **Scikit-Learn** library provides a scaler to do it for you."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"from sklearn.preprocessing import MinMaxScaler\n",
"\n",
"# Get a scaler object\n",
"scaler = MinMaxScaler()\n",
"\n",
"# Create a new dataframe for the scaled values\n",
"df_normalized = df_sample[['Name', 'Grade', 'StudyHours']].copy()\n",
"\n",
"# Normalize the numeric columns\n",
"df_normalized[['Grade','StudyHours']] = scaler.fit_transform(df_normalized[['Grade','StudyHours']])\n",
"\n",
"# Plot the normalized values\n",
"df_normalized.plot(x='Name', y=['Grade','StudyHours'], kind='bar', figsize=(8,5))"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"With the data normalized, it's easier to see an apparent relationship between grade and study time. It's not an exact match, but it definitely seems like students with higher grades tend to have studied more.\n",
"\n",
"So there seems to be a correlation between study time and grade. In fact, there's a statistical *correlation* measurement we can use to quantify the relationship between these columns."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"df_normalized.Grade.corr(df_normalized.StudyHours)"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"The correlation statistic is a value between -1 and 1 that indicates the strength of a relationship. Values above 0 indicate a *positive* correlation (high values of one variable tend to coincide with high values of the other), while values below 0 indicate a *negative* correlation (high values of one variable tend to coincide with low values of the other). In this case, the correlation value is close to 1, showing a strongly positive correlation between study time and grade.\n",
"\n",
"> **Note**: Data scientists often quote the maxim \"*correlation* is not *causation*.\" In other words, as tempting as it might be, you shouldn't interpret the statistical correlation as explaining *why* one of the values is high. In the case of the student data, the statistics demonstrate that students with high grades tend to also have high amounts of study time, but this is not the same as proving that they achieved high grades *because* they studied a lot. You could equally use the statistic as evidence to support the nonsensical conclusion that the students studied a lot *because* their grades were going to be high.\n",
"\n",
"Another way to visualize the apparent correlation between two numeric columns is to use a *scatter* plot."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"# Create a scatter plot\n",
"df_sample.plot.scatter(title='Study Time vs Grade', x='StudyHours', y='Grade')"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"Again, it looks like there's a discernible pattern in which the students who studied the most hours are also the students who got the highest grades.\n",
"\n",
"We can see this more clearly by adding a *regression* line (or a *line of best fit*) to the plot that shows the general trend in the data. To do this, we'll use a statistical technique called *least squares regression*.\n",
"\n",
"Remember when you were learning how to solve linear equations in school, and recall that the *slope-intercept* form of a linear equation looks like this: \n",
"\n",
"$ y = mx + b $\n",
"\n",
"In this equation, *y* and *x* are the coordinate variables, *m* is the slope of the line, and *b* is the y-intercept (where the line goes through the Y-axis).\n",
"\n",
"In the case of our scatter plot for our student data, we already have our values for *x* (*StudyHours*) and *y* (*Grade*), so we just need to calculate the intercept and slope of the straight line that lies closest to those points. Then, we can form a linear equation that calculates a new *y* value on that line for each of our *x* (*StudyHours*) values. To avoid confusion, we'll call this new *y* value *f(x)* (because it's the output from a linear equation ***f***unction based on *x*). The difference between the original *y* (*Grade*) value and the *f(x)* value is the *error* between our regression line and the actual *Grade* achieved by the student. Our goal is to calculate the slope and intercept for a line with the lowest overall error.\n",
"\n",
"Specifically, we define the overall error by taking the error for each point, squaring it, and adding all the squared errors together. The line of best fit is the line that gives us the lowest value for the sum of the squared errors, hence the name *least squares regression*.\n",
"\n",
"Fortunately, you don't need to code the regression calculation yourself. The **SciPy** package includes a **stats** class that provides a **linregress** method to do the hard work for you. This returns (among other things) the coefficients you need for the slope equation: slope (*m*) and intercept (*b*) based on a given pair of variable samples you want to compare."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"tags": []
},
"outputs": [],
"source": [
"from scipy import stats\n",
"\n",
"#\n",
"df_regression = df_sample[['Grade', 'StudyHours']].copy()\n",
"\n",
"# Get the regression slope and intercept\n",
"m, b, r, p, se = stats.linregress(df_regression['StudyHours'], df_regression['Grade'])\n",
"print('slope: {:.4f}\\ny-intercept: {:.4f}'.format(m,b))\n",
"print('so...\\n f(x) = {:.4f}x + {:.4f}'.format(m,b))\n",
"\n",
"# Use the function (mx + b) to calculate f(x) for each x (StudyHours) value\n",
"df_regression['fx'] = (m * df_regression['StudyHours']) + b\n",
"\n",
"# Calculate the error between f(x) and the actual y (Grade) value\n",
"df_regression['error'] = df_regression['fx'] - df_regression['Grade']\n",
"\n",
"# Create a scatter plot of Grade vs StudyHours\n",
"df_regression.plot.scatter(x='StudyHours', y='Grade')\n",
"\n",
"# Plot the regression line\n",
"plt.plot(df_regression['StudyHours'],df_regression['fx'], color='cyan')\n",
"\n",
"# Display the plot\n",
"plt.show()"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"Note that this time, the code plotted two distinct things: the scatter plot of the sample study hours and grades is plotted as before, and then a line of best fit based on the least squares regression coefficients is plotted.\n",
"\n",
"The slope and intercept coefficients calculated for the regression line are shown above the plot.\n",
"\n",
"The line is based on the ***f*(x)** values calculated for each **StudyHours** value. Run the following cell to see a table that includes the following values:\n",
"\n",
"- The **StudyHours** for each student\n",
"- The **Grade** achieved by each student\n",
"- The ***f(x)*** value calculated using the regression line coefficients\n",
"- The *error* between the calculated ***f(x)*** value and the actual **Grade** value\n",
"\n",
"Some of the errors, particularly at the extreme ends, are quite large (up to more than 17.5 grade points). But, in general, the line is pretty close to the actual grades."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"# Show the original x,y values, the f(x) value, and the error\n",
"df_regression[['StudyHours', 'Grade', 'fx', 'error']]"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"### Using the regression coefficients for prediction\n",
"\n",
"Now that you have the regression coefficients for the study time and grade relationship, you can use them in a function to estimate the expected grade for a given amount of study."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {
"tags": []
},
"outputs": [],
"source": [
"# Define a function based on our regression coefficients\n",
"def f(x):\n",
" m = 6.3134\n",
" b = -17.9164\n",
" return m*x + b\n",
"\n",
"study_time = 14\n",
"\n",
"# Get f(x) for study time\n",
"prediction = f(study_time)\n",
"\n",
"# Grade can't be less than 0 or more than 100\n",
"expected_grade = max(0,min(100,prediction))\n",
"\n",
"#Print the estimated grade\n",
"print ('Studying for {} hours per week may result in a grade of {:.0f}'.format(study_time, expected_grade))"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"So by applying statistics to sample data, you've determined a relationship between study time and grade and encapsulated that relationship in a general function that can be used to predict a grade for a given amount of study time.\n",
"\n",
"This technique is, in fact, the basic premise of machine learning. You can take a set of sample data that includes one or more *features* (in this case, the number of hours studied) and a known *label* value (in this case, the grade achieved) and use the sample data to derive a function that calculates predicted label values for any given set of features.\n",
"\n",
"## Summary\n",
"\n",
"Here we've looked at:\n",
"\n",
"1. What an outlier is and how to remove them\n",
"2. How data can be skewed\n",
"3. How to look at the spread of data\n",
"4. Basic ways to compare variables, such as grades and study time \n",
"\n",
"## Further Reading\n",
"\n",
"To learn more about the Python packages you explored in this notebook, see the following documentation:\n",
"\n",
"- [NumPy](https://numpy.org/doc/stable/)\n",
"- [Pandas](https://pandas.pydata.org/pandas-docs/stable/)\n",
"- [Matplotlib](https://matplotlib.org/contents.html)\n"
]
}
],
"metadata": {
"kernel_info": {
"name": "conda-env-azureml_py38-py"
},
"kernelspec": {
"display_name": "azureml_py38",
"language": "python",
"name": "conda-env-azureml_py38-py"
},
"language_info": {
"codemirror_mode": {
"name": "ipython",
"version": 3
},
"file_extension": ".py",
"mimetype": "text/x-python",
"name": "python",
"nbconvert_exporter": "python",
"pygments_lexer": "ipython3",
"version": "3.7.9"
},
"nteract": {
"version": "nteract-front-end@1.0.0"
}
},
"nbformat": 4,
"nbformat_minor": 2
}

@ -0,0 +1,25 @@
Name,StudyHours,Grade
Dan,10,50
Joann,11.5,50
Pedro,9,47
Rosie,16,97
Ethan,9.25,49
Vicky,1,3
Frederic,11.5,53
Jimmie,9,42
Rhonda,8.5,26
Giovanni,14.5,74
Francesca,15.5,82
Rajab,13.75,62
Naiyana,9,37
Kian,8,15
Jenny,15.5,70
Jakeem,8,27
Helena,9,36
Ismat,6,35
Anila,10,48
Skye,12,52
Daniel,12.5,63
Aisha,12,64
Bill,8,
Ted,,
1 Name StudyHours Grade
2 Dan 10 50
3 Joann 11.5 50
4 Pedro 9 47
5 Rosie 16 97
6 Ethan 9.25 49
7 Vicky 1 3
8 Frederic 11.5 53
9 Jimmie 9 42
10 Rhonda 8.5 26
11 Giovanni 14.5 74
12 Francesca 15.5 82
13 Rajab 13.75 62
14 Naiyana 9 37
15 Kian 8 15
16 Jenny 15.5 70
17 Jakeem 8 27
18 Helena 9 36
19 Ismat 6 35
20 Anila 10 48
21 Skye 12 52
22 Daniel 12.5 63
23 Aisha 12 64
24 Bill 8
25 Ted

File diff suppressed because one or more lines are too long
Loading…
Cancel
Save