Showing posts with label median. Show all posts
Showing posts with label median. Show all posts

Friday, October 10, 2014

5 Years of Continuous Blogging, Part 3

This article is the third one in my series about having blogged continuously for five years. It describes how I collected the basic data for my five years of continuous blogging, used some elementary Excel calculations, and graphed some results.

The pixstrip shows a truncated Excel tally sheet and a partial Word-paste window. The Excel portion has a red rectangle around the recipe article that I did the word count on. The partial Word paste shows truncation of the recipe text. The lower left of the Word window displays the word count without needing to open menu options or use shortcut keys. How convenient!

Creating a spreadsheet with initial data
I tiled the Word file of my catalog next to a browser view of my blog site to create the following columns in Excel:
  1. Article titles ("Article Title")
  2. Word count ("# Words")—more on that farther down the article
  3. Image, y/n ("Image, 0 or 1")
  4. Recipe, y/n ("recipe, 0 or 1")
I differentiated the time periods (September to August over five years) by changing font colors. If you choose to try assessing your own blog over years or other time periods, you can try differentiating by text color, bolding, italics, border looks, etc. (For recipes, I flood-filling rows for quick visual cues.)

In retrospect, the catalog was helpful for quickly confirming recipe status, but not crucial. The most important compilation tool was the index in the blog. It grouped by years and months, and linked to all the articles.

Although I created a column for word count early on, I collected the word counts after creating the rest of the spreadsheet. It seemed more sensible to do word count as a separate process than going through each cell of each row. The word count process follows.

EZ Word Counting
For multiple article word tallying, it was helpful using a wide-screen monitor for tiling my blogsite on one side, the Excel tally sheet on opposite side, and MS Word somewhere else. I performed the following steps:
  1. Opened the blog article webpage.
  2. Selected the content—excluding title, superfluous info, images, keywords, and reader comments, and pasted it into the clipboard (Ctrl+C).
  3. Opened Word, pasted the clipboard contents onto the page (Ctrl+V).
  4. Noted the word count at the bottom left of Word window, and typed it into the appropriate Excel cell for word count.
I repeated the process of opening a blog article, copying, pasting into Word, and entering the word count into Excel. However, when bringing Word back to the foreground, I selected the existing text and pasted the new text in its place. I continued with this copy/paste process till finished.

Using Some Very Basic Excel Commands
Using formulas is not necessary in the info gathering, but interesting to calculate totals, averages, highest and lowest word counts (endpoints), and medians (midpoints). Creating graphs require two main layouts of info in Excel—1) calculated values in cells in grid form and 2) and "physical" cell distribution.

The only formula that is necessary for creating the graphs for contrasting numbers of articles over n years is summing. In the Excel window, select the cells, and select the summing icon—
Formula tab > Autosum icon.

Expanding the icon displays options for Sum, Average, Count Numbers, Max, Min, and More Functions. Besides using the summation formula for obtaining the data for creating the graphs in Part 1 of my article series, I also used the percentage formula for creating the table.

Note: I copied some of the data and placed them in another clear area so I could select adjacent cells and use graphing features Insert tab > Chart > 3-D Clustered Column and Chart > Line, respectively. Right-clicking in various areas provided additional options, such as displaying data labels, and modifying titles and other labels.

In case you are a blogger and have not yet tried collecting info on your own articles but might want to try—well, the tools are in the three parts I've written. Click "1" and "2" for the first two articles. For "5 Years Continuous Blogging, Part 4", I will share article links and stats that resulted from my trip down memory lane, my memory lane that I created for myself.

October 31, 2014: Links to the series
  1. 5 Years of Continuous Blogging, Part 1
    Focus on single- and five-year views for total articles, articles with images, and recipe articles.
  2. 5 Years of Continuous Blogging, Part 2
    Focus on numbers of words in articles and graphical representations.
  3. 5 Years of Continuous Blogging, Part 3
    More details on collecting the data.
  4. 5 Years of Continuous Blogging, Part 4
    Emphasis on data sorting and distribution of word count groupings.

Tuesday, September 30, 2014

5 Years of Continuous Blogging, Part 2


In my article about blogging 5 years (part 1), I emphasize quantities of articles over the last five years. This time, I emphasize numbers of words for articles. The images indicate distribution in scatter plots, a line chart, and a box-and-whiskers chart.

The scatter plot for the third year indicates a narrower stream of word counts than the other years. However, it's also the year with the fewest articles. The spread between high and low for word count is as follows:

Sep 2009 to Aug 2010
1323
Sep 2010 to Aug 2011
895
Sep 2011 to Aug 2012
613
Sep 2012 to Aug 2013
1277
Sep 2013 to Aug 2014
852

The line chart shows data points for high, average (mean), median, and low word counts. The box-and-whiskers chart displays the bunching of data. The "whiskers" in the box-and-whiskers graph display the end points to the "box". For each unit:
  • The box represents the middle 50% of the word-count range.
  • The whiskers each depict the other 50%—25% at each end.
  • The line inside the box depicts the median of the word count (spreadsheet-sort of word counts by article and establishing the midpoint).
For each year's period, the line chart numbers for most words, fewest words, and medians coincide with end points and medians in the box-and-whiskers chart. Note that the fifth year "box" (smallest of the five year periods) shows that half the articles fall between approximately 400 and 600 words, with the median around 450.

Some handy resources on graphing, with the first two being the most helpful for me:
Excerpt from Box plot that summarizes "whiskers":
lines extending vertically from the boxes (whiskers) indicating variability outside the upper and lower quartiles
For "5 Years Continuous Blogging, Part 3", I will go into more detail about the road I traveled in obtaining and processing my data.

October 31, 2014: Links to the series
  1. 5 Years of Continuous Blogging, Part 1
    Focus on single- and five-year views for total articles, articles with images, and recipe articles.
  2. 5 Years of Continuous Blogging, Part 2
    Focus on numbers of words in articles and graphical representations.
  3. 5 Years of Continuous Blogging, Part 3
    More details on collecting the data.
  4. 5 Years of Continuous Blogging, Part 4
    Emphasis on data sorting and distribution of word count groupings.

Thursday, June 10, 2010

SurveyMonkeying-A-Round

This year, I used SurveyMonkey for the first time. Some people in my professional writing organization (STC, Society for Technical Communication) had ushered in its usage a few years ago for the annual salary survey for the area. The version we subscribe to is Unlimited. (The free version allows only 10 questions and up to 100 responses.) The pricing page shows a page of features for Free, Pro, and Unlimited. Except for the billing commitment, there appears to be scant difference between Pro and Unlimited.

In February, I finally dived into using the tool when my request (plea?) for a volunteer to conduct the survey yielded a few polite declines. I logged into our SurveyMonkey's account to see what was there. Fortunately, I did receive a guideline document; otherwise I would have been totally lost. In any case, SurveyMonkey's online help was extremely helpful. I felt the tool itself was extremely well-designed and user friendly—a lot of intuitiveness built in for the intermediate tool user. There were ways to clone surveys, copy questions, move questions, select answer options, …. There are expansive explanations of all the features.

At SurveyMonkey's Design Survey section, I navigated a prototype survey and explored response types. I checked out behaviors for multiple choice (check-all-that-apply), single-selection of multiple options ("radio button", dropdown listbox), and open-end answers ("other", fill-in). There were options for presenting possible responses horizontally, vertically, and grid configuration. The advantage for an x-by-y matrix is saving vertical space.

Another response setup that saved vertical space was using the dropdown listbox for single-selection of multiple options, if there were at least four response possibilities. If fewer than four, there was virtually no vertical space savings. For few-answer-option questions, I felt it was more user friendly that all possible answers were visible at the same time.

Why did I concern myself with vertical space? I wanted to present a survey that appeared to be shorter than if all answer possibilities made it visually long. I sensed an overly long survey might fatigue the participants and maybe decrease the chances they would take or complete the survey. Putting some response possibilities in grids and some others in dropdown listboxes required less vertical space for the entire survey than listing line-by-line selections.

After setting up the survey in March, I sent a dry-run version to the board using the Collect Responses feature, then did a cursory analysis using Analyze Results. After minor tweaking, I set up the survey for the public and launched it in early April, sending out several emails over about three weeks requesting participation from technical communicators.

After I finished collecting the data, having set a shut-off date in SurveyMonkey, I moved to the data analysis stage. At the Analyze Results section, I re-acquainted myself with graph types and appropriate uses. The graphs I used were pie, column, and line. Later, when I presented the results, I concluded that in some instances, bar graphs would have better conveyed information than column graphs. But there needs to be judicious use and consistency for bars rather than columns—the biggest reason being if the column charts showed vertically rotated text.

SurveyMonkey's graphing options were actually fun. At the Create Chart option, the following choices were available:

  • Chart shape/type
  • Number of answer choices to show
  • Sort by answer quantity
  • Show or hide a chart title, the default text being the question itself
  • Labels for response number, percent, both, or none
  • Location of the labels, inside or outside the graphic

Clicking Download Chart created the chart. Right-clicking the selection to create the chart in a new window was effective. If I wanted to vary my selections, it was easy to return to the Create Chart option and try something else. On the other hand, I found that SurveyMonkey's graphing capability lacking with regard to answers to open-ended questions. To make suitable graphs for those responses, I used Excel formulas.

In May, I presented the salary results to the STC Austin chapter meeting. My presentation showed a hybrid of a few updated parts from the previous year's presentation, but mostly graphs of each question's responses for this year. I handed out hardcopies of the report that reflected mainly a subset of responses—those pertaining to salaried, full-time technical communicators.

During the week after my presentation, I created a supplement document that went into detail about the questions that required "other" responses and fill-ins. I also wrote about survey design changes from the previous year and specific areas for possible future handling of some questions. All three survey documents—report, presentation, supplement—are available at http://www.stcaustin.org/employment-mainmenu-30/13-salaries/2-salary-survey-results-available.

Conducting this year's salary survey was eye-opening for methodology, learning SurveyMonkey, writeups, and coming up ways for improving the 2011 version. I have listed some survey resources below:

Excel formulas I used are as follows:


=MIN([cellposition1]:[cellposition2]) <- lowest value

=Quartile([cellposition1]:[cellposition2], 1) <- 1st quartile, aka 25th percentile

=AVERAGE([cellposition1]:[cellposition2]) <- mean

=MEDIAN([cellposition1]:[cellposition2]) <- median, midpoint, aka 50th percentile, also calculable using 2nd quartile

=Quartile([cellposition1]:[cellposition2], 3) <- 3rd quartile, aka 75th percentile

=MAX([cellposition1]:[cellposition2]) <- highest value

=COUNT([cellposition1]:[cellposition2]) <- quantity of responses