Thursday, February 12, 2009

Analysis of MusicWeb compression test

Following my earlier comment about analysing the collective data, Kirk and I spoke and he has provided me with the raw data. Before undertaking the analysis I provided him with a protocol for how it would be done and interpreted, based on the approach suggested in that comment. I am happy to share that with anyone who is interested in the details.

In essence, the analysis is based on comparing the actual ranking from 1 (= lossless) down to 5 (=96 kbps) with the rankings provided by each individual for each sample. There were 6 samples so the total value for each level which would be expected if there was no ability to discriminate would be 6 x 3 (the middle rank) = 18. It is important to bear in mind that the lowest value possible is 6 and the highest is 30.

I used the information to look at 3 separate possibilities:


  • That there was some collective ability to discriminate the lossless file (level 1) - this was defined as an average score for this level of 12 or less, 12-15 being a marginal effect and over 15 no effect

  • That there was some collective ability to discriminate the lowest grade file (level 5) - this was defined as an average score for this level of 24 or more, 21-24 being a marginal effect and under 21 being no effect

  • That there might be some ability to discriminate continuously across the range - this was defined as there being at least marginal effects at the top and bottom and exactly the correct order in the middle ranks.


Seven responses were submitted to Kirk but 3 did not contain complete data, broadly because these people felt that the test was impossible. It would have been messy to include partial data and therefore they were excluded. I will come back to the implication of this later.

So 4 complete responses were left which is a very small sample-size and whatever the results, it was unlikely to be completely conclusive. Nevertheless Kirk and I agreed it was still worthwhile doing the analysis, mainly to help in planning any further exercises that might be done in the future.

The average scores for each level turned out to be as follows:

1 (lossless) - 16.6

2 (320 kbps) - 17.6

3 (160 kbps) - 17.3

4 (128 kbps) - 16.6

5 (96 kbps) - 21.9


By the pre-defined criteria, this meets one of the three pre-specified criteria i.e. it suggests a marginal ability to discriminate at the bottom of the range. For all four individuals the scores for the lowest quality sample (level 5) were the highest of their five levels and the actual numbers were pretty consistent i.e. 22, 21.5, 23 and 21. This adds a little credence to the idea that this finding might be real. It is also, in my view, the most plausible of the three possibilities tested - we know that if you drop the bit rate low enough (say to 20 kbps) a difference can be heard but we don't really know at what point this starts to "kick in".

Had there been more data I would have looked systematically at the individual samples to see whether it looked as though the level of discrimination at the lower level might be related to particular types of music. For interest I did have a look at the harpsichord sample (no. 4) which was supposed to be the "easiest" but there was nothing to suggest that.

Coming back to the effect of the excluded participants, the study group divided into two i.e. those who threw up their hands at some point and those who soldiered on regardless. Whether or not these groups might be different in discriminatory power is hard to say for sure. It is certainly possible that if those who opted out had struggled on they would have diluted away the marginal effect which was observed but to say any more on the data we have would be pure speculation.

The broad conclusion from the data collected is similar to our subjective observations – there is little evidence of ability to discriminate between the files. However, there is a suggestion that at least some people may be able to discriminate, albeit far from perfectly, once the bit rate is lowered to 96 kbps. It is worth noting that this is below the bit rate of anything that one is now likely to download commercially.

13 comments:

  1. I am surprised there were so few entrants. I have said to Kirk that I am willing to offer the test to the MusicWeb weekly reviews mailing list (just under 1000 people) and/or to advertise it on the site itself. This should produce a more statistically valid number of responses. We could restrict it to Kirk's first download which included the first three samplings as I would have thought that was a sufficient number. Excessive downloading would begine to cost money!

    The results are so interesting that they must be published, in which case there needs to be some certainty of them being regarded as accurate. We would not be taken seriously if we published based on only four sets of results.

    ReplyDelete
  2. I agree with the idea to try to get a bigger statistical sample. However, I still think we need to do the following:

    (1) Use extracts of no more than about 1 minute duration for the samples (reduces download size and hence discourages fewer people from taking part, reduces strain on concentration and memory, but is still long enough to get "feel" of music)

    (2) Extend the bit-rate range down to a value where artefacts are blatantly obvious to everybody (will give us a measure of both the "threshold" below which people CAN detect the effects of compression and the limit above which the differently-compressed samples sound the same, will include widest possible range of hearing and discriminatory abilities - and hence maximise number of usable returns). Perhaps compensate by widening the bit-rate steps?

    (3) preface each set with the uncompressed sample (so we all have a benchmark, i.e. know what is the "best" sound"), BUT

    (4) also include the uncompressed sample somewhere within the set (so the test itself remains blind, because you still have to identify this version)

    ReplyDelete
  3. Let me assure Len that there was no intention that the above posting would be published more widely. When I offered to do the analysis I had expected there to quite a few more completions but, nevertheless, I do believe that it has been useful to look systematically at what was generated. The problem with the data is not that they are not "valid" or "accurate", just that there is not enough of it (lack of statistical "power" is the technical term).

    I certainly agree with Paul's first point - Kirk and I have already discussed reducing the duration of extracts. However, I would not agree with limiting the number sound of samples for two reasons: (1) It will reduce the amount of data by half and we may not get so many completions, even when offered to 1,000 people (2) It will reduce the possibility of seeing whether the type of music makes a difference - I would regard this as still being possible.

    Whether or not to look at lower bit rates depends on what question we want to answer.

    ReplyDelete
  4. My thought is that if we do a larger tests we should indeed use shorter samples. But I don't think we should include _more_ samples to have lower bit rates tested. This is not a test to show what low bit rates sound like; anyone can do that. This is a test to determine whether "experienced" listeners can tell the difference between uncompressed music and compressed music at bit rates they are likely to encounter. Even now, it's getting increasingly rare to find music below about 192 kpbs (eMusic). iTunes and Amazon use 256 kbps (though in different formats). Mots classical sites that sell music for download use 256 or higher.

    I _don't_ agree with the "benchmark" idea. That would be telling people "here's the good one; now find the bad one". I'm interested in people telling whether what they hear sounds good or not. Again, if people want to compare, they can do so after the results are published.

    My thought for another test is three 30-second samples of 5 or 6 bits of music. That would probably come to about a 100 MB download, which would be much more practical for all involved.

    ReplyDelete
  5. In general, I agree with Kirk. The only thing I would say is that the external scepticism which Len was rightly sensitive to could work two ways. What I mean is that people could not only be sceptical of whether our existing data do show an effect at the lowest level but also whether they exclude an effect at the other end. Indeed, the real sceptics out there who haven’t tried this are more likely to be worried about the latter. This is where we came in originally I seem to remember – the “anyone who can’t tell a lossless file from a compressed one whatever the bit rate can’t be listening properly” brigade. For that reason I think it important to keep a top of the range compressed file and (barring shortening the extracts) would see merit in changing as little as possible (to save having to work out another analysis plan if nothing else!). If push comes to shove the 160 kbps files could go but surely downloading about 200 megabytes is not a big deal for most people?

    I also think we should encourage people to complete the whole test, tell them that “ties” are allowed but ask them to use them as little as possible. We should also ask them to do it individually rather than in a group - listening as a group would be OK if there is no conferring and everyone just records their rankings and sends them in. I say all this just to maximise the amount of data we get. Calculating the ideal sample-size is very tricky here and therefore we will have to go for the “as much as possible” approach. Based on instinct more than reason, I would like to see at least 20 completed tests.

    ReplyDelete
  6. Considering the numbers of these "experienced" listeners who struggled to differentiate between the compressed and uncompressed, I'd have thought that there'd be little or nothing to be gained by offering the same to a more general set of listeners. You'd just end up with a load of people who've been unable to find any difference that they can put their fingers on, or even "given up in despair". We cannot do a statistical analysis of rankings if hardly anybody has been able to provide any. The situation would be similar to asking people to do an unaided sight test using only the bottom two lines!

    That's why, in this sort of test, it is very important that EVERYONE should be able to perceive some difference - then everybody will have at least one ranking per sample to report back. It may add up to a bit more than the original intention, but I think we would gain a lot more knowledge for very little extra effort, simply (say) including versions at 48 and 24 kbps.

    Sorry, Kirk, but we'll have to differ over the "benchmark" issue. Nevertheless, it's fair to point out that the results have already shown that, without a defined point of reference, people have no option but to work from, not the uncompressed version (the actual "best"), but from whatever comes closest to their PREFERRED sound (their personal "best").

    It's of paramount importance to note that the statistical analysis WILL NOT BE VALID if there is no proper, common reference, that is if people are free to choose whatever version they want as "the" benchmark (and I wouldn't advise going public with a statistical analysis containing such a blooper - we'd be laughed out of court).

    By both setting the benchmark AND secreting a copy of it in the set, we would be saying, "This is what it should sound like; now arrange these into order of increasing departure from this standard" (not "better than", not "worse than", but just "further away from"). They still have the challenge of finding and (if they are very dicriminating) ranking "first" the concealed copy of the uncompressed version. In effect, it's a test of "fidelity", as opposed to "quality".

    ReplyDelete
  7. Paul, cats can be skinned in many different ways. I don’t have any problem with you arguing in favour of benchmarking but I cannot accept this:

    “It's of paramount importance to note that the statistical analysis WILL NOT BE VALID if there is no proper, common reference, that is if people are free to choose whatever version they want as "the" benchmark (and I wouldn't advise going public with a statistical analysis containing such a blooper - we'd be laughed out of court).”

    The files have an intrinsic order and comparing that with a perceived order using an ordinal scale is a standard research technique used the world over in many fields for decades. So, apart from rubbishing what has been done so far, you seem to be rubbishing huge swathes of published scientific research.

    Everything comes back to the questions that we want answered. Our views are all coloured by participation in the test. If you want to convince people who haven’t taken it, then there will have to be some numerical analysis.

    If you refer back to the bullets in my original analysis, the possibilities outlined all relate to specific practical questions facing people who might now download recorded music i.e. (1) If a lossless file is available, it is worth the extra money, bandwidth and storage space? Presumably, the answer can only be “Yes” if it is possible to tell the difference (2) Is there a threshold of compression at lower levels at which I might notice the difference? (3) Between these extremes might it be worth paying a bit more for a higher bit rate? The experiment Kirk designed and implemented, and the simple numerical analysis (there were no statistical tests) I performed addresses these questions.

    There were two problems – not enough people completed the test (logical solution – get some more) and some people didn’t complete the test because it was so difficult. The latter can be dealt with in different ways but I don’t think we should get too hung up about it. One approach would be that if people don’t provide data because they can’t tell any differences they should have all their data points assigned as rank three. The other extreme is to exclude them (as I did above), in which case the data only apply to people who, on listening, are prepared to accept that there might be differences. It isn’t unreasonable to look at the data both ways to see whether it matters (a “sensitivity analysis”). Of seven participants, four provided complete data, two provided half the data and one provided none. I think that, with due encouragement to soldier on, there is no reason why most people should not provide complete data.

    Above all, I think the word “validity” (or its derivatives) should be used with extreme caution. The most important validity issue in such an exercise relates to “blinding”, and, in this respect, it is clearly a strong experiment. It is very easy to damn something as “invalid” for any number of reasons but if you want to do that we are all clearly wasting our time.

    ReplyDelete
  8. This comment has been removed by the author.

    ReplyDelete
  9. Patrick - I did not intend to "rubbish" anything. I pointed out a problem, and presented what I considered to be (dare I say this?) perfectly "valid" reasoning. Even a standard technique must be properly applied, and the proper application clearly requires participants to have a common reference point. Often, this will be implicit (in that there's only the one), but in the present case it isn't. The difficulty lies with the criterion for differentiation: that woed "best".

    Let me illustrate with a simple example. Suppose a sound sample whose versions have been created by varying the treble response. You ask people to put them in the order "best" to "worst". Which is "best"? Is it the one with the most "top"? Or is it one of the others? You have no way of knowing. People will choose the "best" according to their individual taste, and the results will not be comparable.

    However, if you aske them to rank the versions by "brightness" or "brilliance", you remove the inherent ambiguity, because there is a common understanding of the criterion, and everyone (where they can discriminate) will put the ones with more treble higher up the list.

    Our "problem" is compounded by the fact that compression is more complicated than a treble control - it is a complex of different processes, producing a mixture of different perceptual artefacts, which are present in degrees that may vary with the amount of compression, and to each of which different individuals may be variously sensitive. This is compounded by the fact that most folk will have no idea what audible effects these various artefacts produce! Assuming that folk can hear differences, their individual CHOICEs of "best" are going to be much more widely dispersed than in the simple "treble" example.

    Sure, you CAN make a valid analysis of such results, but only if you take each person's choice of "best" as correct (which, by definition, it IS) - and I wouldn't fancy trying to figure out how you'd handle the maths! It all hinges on providing a well-defined criterion. For something as complicated as compression, this is going to be almost impossible, so one must adopt the alternative: define the criterion by example - i.e. give them a "benchmark" sample.

    Eeeh - I hope this clears up the misunderstanding!

    ReplyDelete
  10. (There seems to be no way of editing comments so I have deleted and resubmitted. You can see above where this comment was originally located)

    With regard to the above I would agree with Patrick that it did not work well having three of us do the test together as we had to reach some sort of concensus so where one of us thought we heard a difference it was over-ridden by the other two. So individual testing is the only way to go.

    I agree with much of what Paul says and I would suggest we offer:
    ® three music samples.
    ® a named "best" which is also included in the test anonymously.
    ® a 48kbps so everyone can feel satisfied they have at least found something.
    It was very frustrating not being able to hear any clear differences.

    I do not think the track length matters. We never listened to the whole track but heard about 30 secs and then jumped to the next sample. But there is no harm (other than download size) in actually including a full track each time.

    ReplyDelete
  11. If we only offer three samples, I can't see the use of one of them being 48 kbps...

    ReplyDelete
  12. Three different music samples each at a range of compressions; not three compressions. Just as in your first download file but with modification

    ReplyDelete
  13. This comment has been removed by a blog administrator.

    ReplyDelete