So, I haven't gone out anymore; as usual, sleep is not assisting me, so I started writing this explanation. I will try to make it as simple as possible.
What is a sound: it is the result of a body set into vibration that excites the state of rest of the air in its immediate vicinity. The air particles collide with one another and thus the sound propagates through space. For the moment, we are not interested in knowing what the directions are because that would complicate things too much. This mechanical energy, which is transferred from particle to particle, reaches our eardrum, which vibrates coherently with the excited air particles, and through a complex system—which I won't describe to avoid turning this into an encyclopedia—reaches our brain in the form of an electrical impulse, and finally, we perceive this sound.
One can therefore think of the concept of a wave. The air particles that move thanks to the vibrating body break their state of rest; in particular, they will move back and forth relative to the position of rest. This determines zones of high concentration and zones of low concentration (the concept of a wave). However, sound cannot be described simply with the concept of a wave (a pure sinusoid), otherwise, in everyday life, we would only hear pure sounds; rather, it consists of a set of pure sounds that determine the sonic timbre. We can therefore think mathematically of a wave as a function in time, meaning particles that move through space in a certain amount of time, determining, as we said before, zones of high and low concentration, but we can also think of it in an infinitely small amount of time as a set of pure tones. If we think of sound as a set of pure tones, we are looking at it in the frequency domain; that is, I am saying that at that precise temporal instant, the timbre of the sound is given by the sum of certain frequencies reaching our ear.
Below are two images to clarify the concept.
With these plugins, we have photographed the frequencies at play at a certain instant in time in a sound or set of sounds.
On the x-axis, we have the frequencies; on the y-axis, the "weight," improperly called volume, of each frequency. (The correct term would be Sound Pressure Level, or SPL, which is a logarithmic quantity consistent with our auditory system, which is also logarithmic).
I believe Autoflip is based exactly on this principle, or at least I cannot imagine any other way to make it work that way.
Using the self-learning function on Autoflip, we record the audio file we are listening to, and the app stores the point where we turned the page while following the score. Once the self-learning is finished, the recorded audio file is taken, and snapshots are taken—just like the one above—at regular intervals; even 0.5 seconds might suffice.
Why did I initially make the premise of the concept of a wave seen in time and frequency? Because at this moment we are seeing the sound in frequency, therefore the temporal variable no longer appears. Frequency recognition becomes independent of time; that is why even when playing very slowly, Autoflip will still turn the page at the right point.
Having analyzed the entire score, let's assume that Autoflip recognizes that between the 56th and 57th frequency photo, the page must be turned. How does it achieve this? With markers.
Let's place our iPad with Autoflip open on the piano and begin playing the score for which we previously performed the self-learning.
At this point, Autofflip begins to take snapshots of the frequencies (Fast Fourier Transform) in real time—snapshots of what we are playing—and if the frequency played on the piano corresponds to the first snapshot taken during the self-learning phase, it places a marker (marking it as recognized) and waits for the frequencies of the second snapshot to be played. When the frequencies of the second snapshot are played, it places the recognition marker on the second snapshot and waits for the third, and so on. When we reach the 56th snapshot, it places the marker and turns the page (since we remembered that we had posited that the page should be turned between snapshots 56 and 57 of the self-learning analysis). At this point, the same procedure continues, comparing the frequencies we play in real time with those obtained from the self-learning and placing markers, until the end of the score.
It is obvious that for this app to work well, some calibrations will have been necessary. Probably, a frequency triggering was performed, meaning a selection of the"}
most important frequencies, disregarding involuntary ones (a vacuum cleaner in the other room, a passing plane, a train), so this is necessary to avoid considering noise that can differ from place to place but also the timbral difference of different pianos. In short, identical frequencies will never be found even in the same instrument, so to guarantee the positioning of a marker, some tolerance thresholds must be given.
Obviously, I cannot know how autoflip was developed and therefore these are just my personal considerations. That is to say, if I were to implement it, I would develop it this way, also because I don't see any other possibilities... But if someone finds some documentation on the internet refuting this thesis, all welcome.
I have kept it very down-to-earth, so experts forgive some inaccuracies in terminology, but if we started talking like engineers, perhaps the message would truly reach only a few; I hope nonetheless to have cleared up your ideas.