Piano technology

Two sounds compared

38 replies 7,503 views

Started by

Pianoaccordatore

Opening post by Pianoaccordatore

Here is a very pleasant comparison for all of you between 2 Japanese upright models; Yamaha and Kawai

This is the section of the forum where piano construction technology is discussed.

Which sound do you prefer?

http://www.youtube.com/watch?v=rMbq4U61CbE

Member

Reply 2 by Andry

Posts
86
City
Roma

The Kawai is too bright; I listened without seeing which piano came first, but it's worth noting that trying out pianos in a shop seems absurd to me, the first one was better, they all seem the same except for clear differences in quality; besides, small spaces and many soundboards close together will inevitably cause some effect..

Administrator

Reply 3 by Thesimon

Posts
4,504
City
Ariccia, RM

It seems to me that the sound, just by ear, has been recorded differently. For the Kawai, it seems to me that the microphones are placed much closer to the hammers. It is impossible to make a comparison of the sound this way.

Member

Reply 4 by Boesendorfer

Posts
103
City
Tirol

I agree, TheSimon.

A comparison that is anything but pleasant: mediocre sound, bad music.

I prefer German pianos, not just out of patriotism.

Member

Reply 5 by Andry

Posts
86
City
Roma

Uhm yes, many of us think so.. but if asked, I would have liked to answer whether I prefer Kawai or Yamaha sound; unfortunately, they all sound so damnly alike when you're in a shop that it makes you want to go crazy! and I've already gone through the process of choosing a piano. Anyway, assuming new pianos are homogeneous in quality, I would like to make the choice again with recordings in halls with more impartial acoustics

Moderator

Reply 6 by pianoexpert

Posts
3,876
City
rocca di papa

It is difficult to make a comparison without taking the hammer voicing into account! Both Houses feature hard, saturated hammers that must be voiced over time, otherwise the sound becomes too "hard".

Member

Reply 7 by lory

Posts
106

I don't think one can reason about hammer voicing when the audio is recorded with a microphone found in an Easter egg and the recording level adapts to the microphone's input level. I agree with TheSimon when he says that it is impossible to make a comparison this way; I believe he will agree with me if I say that this recording is flat without dynamics, and the amplitude can only be recognized by the change in timbre and not as it should be from the change in timbre given by higher volume. Am I right?

Administrator

Reply 9 by Thesimon

Posts
4,504
City
Ariccia, RM

Wow! 10 out of 10 for Lory... Here I have to do a proper analysis, complete with FFT and phase comparison.

General information about the audio track:

We remind ourselves that the bitrate is calculated using the formula: b = 8 * f / t where f is the file size expressed in bytes (1MByte = 1024KByte and 1KByte = 1024Byte), therefore 1.87MB = 196083712 Bytes, and t is the time in seconds. In this case, there is a reading error regarding the track information for the time, which is nonetheless 3:59, so 239 seconds.

A bitrate this low already gives us an idea of the poor audio quality, even if the channel mode is Joint Stereo (which allows us to recover something in terms of bits per channel at such low bitrates).

What we must focus on, which is quite important, is the Sample Rate, i.e., the number of samples per second measured by the audio sampler.

22050 samples per second means that, according to the sampling theorem, the maximum physical frequency measurable and samplable by the sampler is approximately 10000 Hz. This is a very important piece of data. Given that the nominal human hearing range is between 20Hz and 20 KHz—even though in reality it is already good to consider 15 KHz as a good audible threshold—the higher sounds, even if they can no longer be recognized on their own, still produce psychoacoustic effects when combined with other sounds. Having a sample that records up to a threshold of 10000 Hz means losing a lot of information. Not only that... In a piano context, this means that when playing the highest C(8) on the keyboard, the recording is only able to capture the second harmonic at about 8500 Hz. The third harmonic is already out of reach.

A spectral analysis confirms these expectations...

For this spectral analysis, I used a native plugin within Logic. What you see are the maximum peaks after playing the entire piece. On the x-axis are the frequencies, and on the y-axis are the amplitudes.

We observe that the audible frequencies above the -35 threshold start from 50 Hz to 8 KHz; however, let's consider the real sonic sludge starting from -20dB onwards. Our field narrows to frequencies in the order of 150 Hz up to 4KHz. In short, the recording has already eaten up a large part of the audible frequencies.

Let's analyze the phase correlation to see how Joint Stereo encoding has affected spatiality.

As we can see, the file has no spatiality whatsoever. It is mono stereo!

In other words, it makes us think that at the time of capture, the audio was captured by a camera microphone, split from mono to stereo. At the moment of conversion, there were no encoding bits at low frequencies that could be differentiated; therefore, Joint Stereo did not produce any useful result for its characteristic. We can thus consider the file MONO, or technically false stereo with a 32KBit rate per channel.

From a dynamics standpoint, Lory hit the nail on the head. All "entry-level" cameras have an autogain system that focuses on the frequency spectrum with greater amplitude by decreasing others (a sort of adaptive band-pass filter); this technique serves to mask unwanted noises in the recording (e.g., wind, rough sea, background noise) in order to enhance the volume of the person speaking, who is most likely in front of the camera microphone. The problem is that this continuous rising and falling of dynamics vs. frequency also destroys frequencies that could be interesting in music because they characterize the timbre of the instrument (let's not forget that timbre is determined by the Fourier summation of the fundamental with all its partials, but if the partials are cut, the timbre is inevitably modified).

The demonstration of this adaptation comes from an non-existent dynamic range (low sound pressure sounds recorded at the same level as high sound pressure sounds); as can be seen in the following slide, the average dynamics is very close to the peak.

Tall these words just to say that, in my opinion, this video has no reason to exist and provides no qualitative indication regarding the quality of the sounds being compared.

To pose the same question on the level of images, it would be like saying...

Which photo among these two is the most beautiful of the Colosseum in Rome?

It goes without saying that the low quality of the photos does not allow us to express an opinion.

mp3info.pngfft.pngcorrelation.pngwaveform.pngcolosseo.jpg

Member

Reply 10 by lory

Posts
106

Goodness! I didn't follow you very well on some things, like when you talk about the second harmonic being out of reach and also when you talk about the sampling theorem, what do you mean?

Administrator

Reply 11 by Thesimon

Posts
4,504
City
Ariccia, RM

Forgive me, you pointed out an unforgivable error that I had to correct, otherwise my considerations wouldn't have made sense. It is not the second harmonic but the third harmonic. The first harmonic is the fundamental itself, the second harmonic has double the frequency of the fundamental, the third harmonic has triple the frequency of the fundamental, and so on.

The sampling theorem, also known as the Nyquist-Shannon theorem, defines analog-to-digital sampling by utilizing the Fourier transform and tells us that an analog signal whose bandwidth is limited by a certain frequency we call Fmax can be described by the sum of different frequencies (Fourier). It is a very complex theorem; if you have solid mathematical foundations, you can take a look at this link, otherwise, in the following lines, I will greatly summarize the concept to make it accessible to everyone. Before being converted to digital, it must be sampled with a certain sampling frequency that we will call Fc. Fc must be sufficiently > 2 Fmax, otherwise one encounters an effect called Aliasing, and when reconverting the signal to analog, there will be data loss. In other words, if the Nyquist theorem is satisfied, the original signal will be equal to the sampled signal reconverted to analog. It cannot be strictly greater than 2 Fmax but reasonably greater than Fmax because low-pass filters are not ideal; they have a certain signal attenuation slope as a function of frequency, therefore it is preferable to sample at frequencies higher than 2 Fmax. Since the nominal maximum audible threshold for a human being is in the order of 20,000 Hertz, according to Nyquist we should sample at a frequency greater than 2*20,000 = 40,000 Hz, but as we said, this frequency would be valid if we were working with ideal components; since we have to work with real components, this frequency is raised to 44,100 to be safe from encountering sampling aliasing phenomena. This is where the need to work at 44,100 Hz for audio arises. If we had to produce a CD for bats, this sampling frequency would be decidedly unsatisfactory. Bats have a maximum audible threshold in the order of 80,000 Hz and we would have to sample at more than 160,000 Hz; current 192KHz samplers would work for us. Unfortunately, however, there are no microphones capable of perceiving such high frequencies, also because for our use they would be too expensive and perfectly useless, just as it is useless to produce a disc for bats. It was obviously just an example to see how to use the theorem.

I hope the explanation is clear to you; if there is anything you didn't understand, feel free to ask...

questo link

Administrator

Reply 14 by Thesimon

Posts
4,504
City
Ariccia, RM

Exactly so. You understood perfectly. If they didn't have this sampling frequency, the reconstruction of the analog output wave would not be comparable (at those frequencies we want to reproduce) to the acoustic input wave.

Administrator

Reply 17 by Thesimon

Posts
4,504
City
Ariccia, RM
Fra wrote:

And what about information loss?

Great question. Information loss occurs for 2000 different reasons. We will only analyze a few of them so as not to get too technical and heavy; then, if you think it's necessary, we can delve deeper into the matter as much as you want, or as far as I am able to go because it falls within my expertise...

The first loss occurs before sampling, as the signal is passed through a low-pass filter. To describe nature, we use simpler models because it would be impossible to consider all factors together. We therefore model the low-pass filter in the simple case with a resistor and a capacitor in series (the capacitor is in parallel with respect to the signal). The filter is not at all an ideal element that sharply cuts off all frequencies; therefore, looking at the frequency response graph of the filter, we will see a cutoff curve that descends with a certain slope and certainly not, as in the ideal case, at a right angle.

Another reason is the background noise introduced by the connections and the circuit components which, although they may be excellent, are again not ideal; thus, regarding the cables, we find a resistive component, an inductive component, and a capacitive component. The same applies to the components. In capacitors, we have a capacitive quantity that is presumed to be so prevalent as to consider its resistive and inductive components negligible. The same reasoning applies to resistors and inductors.

Other losses stem from the fact that sampling at 24 bits means dividing the ordinate axis (the amplitude axis) into 2^24 divisions, or 16,777,216 divisions. With an input signal of 10 Volts, I am measuring differences on the order of 596 nV (This is the magnitude of the "Conversion Quantum") and we remember that 600nV is the Johnson noise threshold in a 2.2K Ohm resistor at a temperature of 25°C (Noise introduced by electrons moving due to thermal effect inside the resistor) measured at a frequency of 10KHz. With this resolution, we expect to hear a range of 144dB (found by calculating 20 times the log(1/2^n) with n = number of sampling bits); in reality, it reaches 120/122dB because there are unusable bits due to non-ideal components.

Other problems are related to the imperfect clock scanning; so much so that to improve this condition, people usually use so-called World Clocks, which are external machines that synchronize the converters of different devices. As you will have understood, AD/DA converters are extremely complex machines. There are only 3 manufacturing companies in the world that have the funds to invest in designing such complex machinery. This is why AD/DA converters that minimize these errors cost in the order of 8000 euros (see the Prism). The sound quality is thus profoundly influenced by the quality of these machines.

Member

Reply 18 by Fra

Posts
266

Come come, I was joking! I just wanted to provoke you, as an electronic engineer that I unfortunately am

Compliments on the passion!

PS. I haven't read the reply, I will read it

Member

Reply 19 by Fra

Posts
266

I held my breath and read .... All correct considerations!!

In all this process, what would you say about the most metallic component of the sound (since we are talking about pianos)... where does it end up?

And compared to recordings from the past (those of the great masters), what changes?

Is it better to capture the sound near the hammers or far away?

Administrator

Reply 20 by Thesimon

Posts
4,504
City
Ariccia, RM

Ahahah, what a delinquent!

The metallic component of the sound comes from partials quite distant from the fundamental; it is indeed these that provide brightness, if that is what you are referring to, which is given by the summation in n of n times A with n all over the summation in n of A with n. (No stress dictating formulas like this)... n would be the number of the partial and A with n is the amplitude of the nth partial we are considering.

There are also other parameters that influence brightness and therefore the Fourier sum. For example, the waveform. The square wave, for instance, is made up of only odd components. If we wanted to simulate an analogous behavior on a piano, we would have to strike the string exactly in the center with the hammer so as to make the formation of all even harmonics impossible.

Capturing the sound near the hammers or far away is just a matter of taste. It would be like asking us when a noise becomes sound and vice versa when a sound becomes noise. Even banging elbows on a table can be a sound to me if it becomes contextualized in a certain setting. Certainly, recording the sound near the hammers provides much more detail regardless of the type of piano; the resulting sound is more

Log in to participate