Showing posts with label Study Statistics. Show all posts
Showing posts with label Study Statistics. Show all posts

Sunday, May 25, 2014

Why Study Statistics

Decision Making in an Uncertain Environment
Everyday decisions are made based on incomplete information. For example, in any business, decisions are made regularly in an environment where decision makers cannot be certain of the future behavior of those factors that will eventually affect the outcome resulting from various options under consideration.

In submitting a bid for a contract, a manufacturer will not be completely certain of total future costs involved nor have knowledge about bids to be submitted by competitors. In spite of this uncertainty, a bid must be made. An investor, deciding how to balance a future market movements are unknown. This investor does not know with certainty whether the market will be buoyant, steady, or depressed. 

Journey to Making Decisions
To think statistically will involve a journey from problem definition to the transformation of data to information, and then the transformation of information to knowledge. Finally, knowledge should lead to better decision making.  

DATA, INFORMATION, KNOWLEDGE

1.    Data: specific observations of measured numbers.
2.    Information: processed and summarized data yielding facts and ideas.
3.    Knowledge: selected and organized information that provides understanding, recommendations, and the basis for decision.

After a problem is identified and defined, data produced by various processes is collected according to a design. These data are analyzed by using one or more statistical procedures. From this analysis, information is obtained and converted to knowledge using understanding based on specific experience, theory, literature, and additional statistical procedures. Knowledge then leads to decision making.


  
DESCRIPTIVE AND INFERENTIAL STATISTICS
Descriptive statistics include graphical and numerical procedures that summarize and process data and are used to transform data to information. Inferential statistics provide the bases for predictions, forecasts, and estimates that are used to transform information to knowledge.


POPUALTION AND SAMPLE

A population is the complete set of all items in which an investigator is interested. A population is the set of all of the outcomes from a system or process that is to be studied. N represents the population size.

Other examples of populations might be:
•    Names of all the registered voters in Bangladesh.
•    Income of all families living in Dhaka city.
•    Costs of all claims for medical insurance coverage received by a company in a given year.
•    Grade point average (GPA) of all students in NSU.

Sample:

A sample is an observed subset of population values with sample size given by n.

Random Sampling
Simple random sampling is a procedure used to select a sample of n objects from a population in such a way that each member of the population is chosen strictly by chance, each member of the population is equally likely to be chosen, and every possible sample of a given size, n, has same chance of selection.

Parameter and Statistic

A parameter is a specific characteristic of a population. A statistic is a specific characteristic of a sample.

Data:

The raw material of statistics is data. We may define data as numbers. Data are collections of any number of related observations.

CLASSIFICATION OF VARIABLES

Variable:
A characteristic that varies from one person or thing to another is called a variable. Examples of variables for humans are height, weight, number of siblings, sex, marital status, and eye color.

Numerical or Quantitative Variables:

A numerically valued variable is called quantitative variable. Height, weight and number of siblings yield numerical information and are examples of quantitative variable.

Numerical variables include both discrete and continuous variables. A discrete numerical variable produces a response that comes from a counting process.
•    A professor is asked the number of students that are enrolled in a class;
•    A student is asked how many times a week an assignment is completed in the in the in the computer lab;
•    An investor is asked how stocks of Microsoft are in investor’s portfolio;
•    An insurance company wants to know how many claims were filled following a hurricane.

Responses to each these questions are 0, 1, 2, and so forth.

A continuous numerical variable produces a response that is the outcome of a measurement process. For example, you might tell someone, that you are 6 feet (or 72 inches) tall, but your height could actually be 72.1 inches or 71.8 inches, or some other similar numbers depending on the accuracy of the instrument used to measure your height. You might state your age “20 years,” but you are constantly aging. You are actually 20 years, 3 months, 2 days, 5 hours, 4 minutes, 12 (now it’s 13) second old. Sounds absurd. So you just say 20 years. The actual weight of packages could also deviate within a certain amount depending on the precision of the instrument used to weight each package.  
 

Categorical or Qualitative Variables: 

A non-numerically valued variable is called qualitative variable. Sex, marital status, and eye color yield non-numerical information and are examples of qualitative variables.

Categorical variables produce responses that belong to groups (sometimes called “classes”) or categories. For example, responses to yes/no type questions belong in this category. “Do you own a cellular phone” or “Do you have any homework tonight” or “Have you ever been in London” are limited to yes or no answers. A health care insurance company may ask if a claim was incorrectly processed. The operations manager at a water bottling company asks if bottles are properly filled. Sometimes, categorical variables include a range of choices such as “strongly disagree” to “strongly agree”.

Saturday, May 24, 2014

Measurement Levels (Scales)

Data could be described as qualitative (includes nominal and ordinal levels of measurement) or quantitative (includes interval and ratio levels of measurement).
   
Nominal and Ordinal Levels of Measurement refer  to data obtained from categorical questions. A nominal scale indicates assignments to groups or classes such as gender (male-female), geographic region (Dhaka, Rajshahi, Chittagong, etc), the model of car you own, or simply the yes or no responses to questions such as the ownership of a cellular phone. Nominal data is considered the lowest or weakest type of data, since numerical identification is chosen strictly for convenience.   

The values of nominal variables are words that describe the categories or classes of responses. The values of the gender variable are male and female; the values of “Did you ever visit Oslo, Norway?” are “yes” and “no.” We arbitrarily assign a code or number to each response. However, this number has no meaning other than for categorizing. For example, we could gender responses or yes/no responses as

                      1 = Male                                   1 = Yes
                      2 = Female                                2 = No

Ordinal data indicates rank ordering of items. Whenever observations are not only different from category to category but can be ranked according to some criteria. Examples include product quality ratings (1: poor; 2: average; 3: good) or satisfaction ratings with university food service, academic advising, or your new car (1: very unsatisfied; 2: somewhat unsatisfied; 3: neutral; 4: somewhat satisfied; 5: very satisfied). At the end of the semester you complete questionnaire about your course and your instructor. Often students are asked to respond to a statement such as “The instructor in this course was an effective teacher” with number from 1 to 5 (with 1, strongly disagree; 2, slightly disagree; 3, neutral; 4, slightly agree; and 5, strongly agree). In these examples, the responses are ordinal, or put into a rank order but there is no measurable meaning to the “difference” between responses. That is, the difference between your first and second choices may not be the same as the difference between your second and third choices.  

Interval and Ratio Levels of measurement refer to data on an ordered scale where meaning is given to the difference between measurements. An interval scale indicates rank and distance from an arbitrary zero measured in unit intervals. Temperature is a classic example of this level of measurement, with arbitrarily determined benchmarks generally based on either Fahrenheit or Celsius degrees. Suppose that it is 80 degrees Fahrenheit in Otlando, Florida, only 20 degrees Fahrenheit in St. Paul, Minesota. We can conclude that the difference in temperature is 60 degrees, but we can not that it is four times as warm in Orlando as it is in St. Paul.  

Ratio scale data does indicate both rank and the distance from a natural zero, with ratios of two measures having meaning.  A person who weighs 200 pounds is twice the weight of a person who weighs only 100 pounds; a person who is 40 years old is twice as old as someone who is 20 years of age.

After you have defined the problem of interest, you will need to carefully design an instrument, such as a survey, collect needed information. Or perhaps you will gather information from available data sources. In either case, before you can proceed to summarize or design data, you will first need to classify responses as categorical or numerical, or by measurement scale. Certain graphs and descriptive measures are used for numerical variables. Analysts use different graphs and descriptive measures for categorical data.