A method for detecting finger tapping sound to realize switch

Through the neural network model, the problem of switching operation relies on external devices in the existing technology is solved, and the switching operation is achieved without external devices directly contacting the machine, which improves the recognition rate and reduces false triggering.

CN116007140BActive Publication Date: 2025-05-16QUANZHOU MUSIC OPERATOR TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211636106.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-20
Publication Date
2025-05-16
Estimated Expiration
2042-12-20

AI Technical Summary

Technical Problem

In the prior art, switch operation depends on external devices, such as remote controls, touch screens, buttons, etc., which have problems such as battery exhaustion, inconvenient operation, and unhygienicity. Voice control depends on the network and has a low recognition rate.

Method used

Through the neural network model, the sound-touching sound is analyzed, the sound-sensing area is divided, and the data is collected and marked. The convolutional neural network is used for training to realize the recognition of the finger-touching sound, and the sound signal is received by the microphone for pre-judgment classification, and whether it is an effective knocking is started to start the switch operation.

Benefits of technology

It can directly contact the machine for switching operations without borrowing external devices, improve the recognition rate of finger tapping sounds, reduce the possibility of false triggering, and is suitable for various complex sound field environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116007140B_ABST
    Figure CN116007140B_ABST
Patent Text Reader

Abstract

A method for detecting finger tapping sound to realize switch, belongs to the field of sound processing, including: S1, sound sensing area division; S2, data collection; S3, data labeling; S5, neural network model design and training; S6, loading and compiling finger tapping detection neural network model; S7, receiving sound signal, the model pre-judges and classifies the sound signal; S8, the model decodes the recognition result of each frame of the classified sound signal, and caches the decoding sequence; S9, parses the sequence to determine whether it is a valid tap; if so, the switch operation of the controlled target object is started; if not, the original state of the controlled target object is maintained. The present invention realizes switch control by tapping the sound sensing area of ​​the controlled target object with a finger, without borrowing external equipment; the finger tapping detection neural network model predicts which label the received sound signal is, and combines effective decoding and parsing strategies, which not only has a high recognition rate, but also avoids false triggering.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of sound processing, and in particular relates to a method for detecting finger tapping sound to realize switching. Background Art

[0002] Switch operation is one of the most frequently used functions for human-computer interaction. The most common interaction methods are through buttons, touch screens or touch panels, remote controls or mobile phone APP universal remote controls. Now, there are also direct human-computer dialogues through voice, and switch control can also be achieved through various signals such as radio waves, and sensors receiving ultrasonic, infrared or laser signals.

[0003] Since remote controls, touch screens or touch panels, buttons, and mobile phone apps are all external devices, they may face various problems. For example, remote controls require batteries or may not be found for a while; touch screens or touch panels have specific locations and may fail when wet hands, or the contact surface is too small to operate at night; buttons may be unhygienic (especially during the epidemic) and are more likely to fail; mobile phone apps may rely on the Internet and need to be aligned with the device; voice control requires the Internet online, and the recognition rate of voice offline may be poor. Therefore, it is necessary to develop a way to directly contact the machine to achieve switch operations without borrowing external devices. Summary of the invention

[0004] In view of the deficiencies in the prior art, the purpose of the present invention is to provide a method for detecting finger tapping sounds to realize switching, which uses a neural network model to analyze whether it is a finger tapping signal sound, thereby simplifying and realizing the "switch function" in human-computer interaction.

[0005] To achieve the above object, the present invention adopts the following technical solution: a method for detecting finger tapping sound to realize switch, comprising:

[0006] S1. Sound sensing area division: the outer surface of the controlled target object is divided into several sound sensing areas for finger tapping;

[0007] S2. Data collection: Install a microphone inside the controlled target object, and let multiple tappers tap the sound-sensitive area with their fingers multiple times. The microphone collects the silence, the starting sound of the tap, the tail sound of the tap, and the noise;

[0008] S3, data labeling: label silence, the starting sound of knocking, and the tailing sound of knocking as positive sample data, with silence labeled as the first label, the starting sound of knocking as the second label, and the tailing sound of knocking as the third label; label noise as negative sample data, with noise labeled as the fourth label; the labeled positive sample data and negative sample data constitute a data set;

[0009] S4, data set noise addition, feature extraction, shuffling, and segmentation: add noise to the data set, extract features, and obtain data features with labeled data; after shuffling the data features, split it into a training set, a verification set, and a test set;

[0010] S5. Neural network model design and training: A convolutional neural network model is used, and its loss function adopts a multi-classification cross entropy loss function; the training set, the verification set, and the test set are put into the convolutional neural network model for training to obtain a finger tapping detection neural network model;

[0011] S6, loading and compiling the neural network model for finger tapping detection;

[0012] S7, receiving a sound signal through a microphone, and using a finger tapping detection neural network model to predict and classify the sound signal into silence, a tapping start sound, a tapping tail sound, or noise;

[0013] S8, decoding each frame recognition result of the classified sound signal by the finger tapping detection neural network model, and caching the decoded sequence;

[0014] S9, the post-filter of the decoder analyzes the sequence to determine whether it is a valid tap; if so, the switch operation of the controlled target object is started; if not, the original state of the controlled target object is maintained.

[0015] Preferably, the microphone cover in step S2 is provided with a silicone cover.

[0016] Preferably, the finger tapping in step S2 includes finger tip tapping, finger web tapping and finger joint tapping.

[0017] Preferably, the strength of the finger tapping in step S2 is that each tapper taps at will in a manner that he or she considers comfortable.

[0018] Preferably, the finger tapping method and number of times in step S2 are:

[0019] (1) Finger tapping: Each tapper taps M times in each sound-sensing area, including N consecutive taps of two times and (MN) consecutive taps of three times;

[0020] (2) Finger tapping: Each tapper taps M times in each sound-sensing area, including N consecutive taps of two times and (MN) consecutive taps of three times;

[0021] (3) Finger joint tapping: Each tapper taps M times in each sound-sensitive area, including N consecutive taps of two times and (MN) consecutive taps of three times.

[0022] Preferably, the continuous tapping means that the time between two adjacent tappings does not exceed 200ms.

[0023] Preferably, the feature extraction in step S4 adopts a balanced Fbank feature extraction method.

[0024] Preferably, the specific process of the balanced Fbank feature extraction method is:

[0025] (1) Assume that the original input data is x, the pi is FPI, the number of Fourier transformation points is N, the highest frequency is Max_freq, the lowest frequency is Min_freq, and the number of generated containers is M;

[0026] (2) Generate the jitter factor as follows:

[0027] Use the cosine function to limit the jitter factor to a range, that is, d_feat = sqrt(-2 * log(d_) *cos(2*FPI* d_), where d_ is a Gaussian distribution with a mean of 0 and a variance of 1;

[0028] The dither factor is dither = (d_feat * 2.0f) - 1.0f), where f represents a floating point number;

[0029] (3) Original data plus dither factor: x_D = x + dither;

[0030] (4) After processing, take the average of every 400 points (i.e. 25ms), and then perform regularization, as follows:

[0031] Take the mean of x_D: x_bar = mean(x_D);

[0032] Subtract the DC component and make regularization to prevent large changes: x_f = x_D - x_bar;

[0033] (5) Pre-emphasis:

[0034] At the first point, the DC component is x_pre(i) = x_f(i) * 0.05, where i=0;

[0035] From the second point to the following points, make the difference: x_pre(i) = x_f(i) - x_f(i-1) *0.95, where i=1,2,3,……,N;

[0036] (6) Add a Hamming window, as follows:

[0037] The numerical value of the Hamming window is: ham = 0.54 - 0.46*cos((2FPI / (N-1)) * i), where i=0,2,3,…,N;

[0038] Data windowing: x_H = x_pre * ham;

[0039] (7) Short-time Fourier transform of N points:

[0040] Transform from time domain to frequency domain, that is, X_H = fft(x_H, N);

[0041] (8) Obtain energy spectrum:

[0042] X_E = real(X_H) ^ 2 + imag(X_H) ^ 2;

[0043] (9) Balanced generation of corresponding containers:

[0044] Spectrum bandwidth: SPAN_freq = (Max_freq- Min_freq);

[0045] Average container size: B_s = SPAN_freq / (M + 1);

[0046] Multiply the corresponding points: X_c =X_E* SPAN_freq[j * B_s:(j + 1)* B_s ], where j=0,2,3,…,M;

[0047] (10) Add the points in each container and take the natural logarithm to get the balanced Fbank:

[0048] Unibank[j] =Log(sum( X_c[j* B_s:(j + 1)* B_s ])), where j=j=0,2,3,……,M.

[0049] Preferably, the post-filtering algorithm process of the post-filter in step S9 is as follows:

[0050] (1) Put the sequence decoded by the neural network and the corresponding probability into buffers A and A_PRO;

[0051] (2) Analyze buffers A and A_PROB, merge the same and adjacent categories, put them into buffer B, record the length of consecutive intervals of the same category, put them into buffer B_SPAN, and calculate the mean probability within the interval length and put it into B_PROB;

[0052] (3) According to buffers B, B_SPAN and B_PROB: if it is the second label of 2-5 consecutive frames, put the second label into buffer C; if it is the third label of 4-10 consecutive frames, put the third label into C; otherwise, put the first label into buffer C and put the probabilities corresponding to the respective categories into C_PROB;

[0053] (4) Perform secondary decoding on buffers C, C_PROB, and B_SPAN. If the adjacent second and third tags are found in buffer C and their probabilities exceed the set thresholds, and the knocking duration does not exceed 27 frames and contains two or three knocking sounds, it is judged as a valid knocking. Otherwise, it is an invalid knocking.

[0054] Preferably, the effective tapping in step S9 is 2 or 3 consecutive tappings.

[0055] Compared with the prior art, the present invention has the following beneficial effects: switch control is achieved by tapping the sound sensing area of ​​the controlled target object with fingers, and the switch operation can be achieved by directly contacting the controlled target object without borrowing external equipment; through special data collection and labeling methods, feature extraction methods, neural network model learning and prediction of which label the received sound signal is (silent, starting sound, tailing sound, noise), combined with effective decoding and analysis strategies, not only the recognition rate of finger tapping sound is very high, but also unnecessary false triggering is avoided.

[0056] If we use traditional signal processing algorithms, such as detecting the sound of finger tapping by energy, the recognition rate is poor and the rejection of recognition is also poor (frequent false detection). For example, the sound of knocking, collision, decoration, wall breaking machine, wall knocking, and hammer hitting outside may trigger the switch function inexplicably. Because the finger taps the controlled target, the sound is transmitted in solid and air. The solid contains various materials in the cavity. Different materials have different propagation speeds for sound. At this time, the sound field environment becomes extremely complex. It is far from enough to rely solely on a single signal attribute, such as energy size. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] Figure 1 is a flow chart of an embodiment of the present invention.

[0058] Figure 2 It is a structural schematic diagram of a vertical air conditioner for a living room in an embodiment of the present invention.

[0059] Figure 3 The figure is a schematic diagram of data annotation in an embodiment of the present invention.

[0060] Figure 4 4 is a comparison diagram between the balanced Fbank feature extraction and the conventional Fbank feature extraction in an embodiment of the present invention.

[0061] Figure 5 This is a specific process diagram of equalization filtering in an embodiment of the present invention.

[0062] Figure 6 4 is a flow chart of a post-filtering algorithm in an embodiment of the present invention.

[0063] Markings in the figure: 11, sound sensing area; 12, microphone; 13, air conditioning touch panel; 14, air conditioning outlet. DETAILED DESCRIPTION

[0064] In order to make the above features and advantages of the present invention more obvious and easy to understand, embodiments are given below with reference to the accompanying drawings for detailed description as follows.

[0065] like Figure 1 As shown, a method for detecting finger tapping sound to realize switching includes:

[0066] S1. Sound sensing area division: the outer surface of the controlled target object is divided into several sound sensing areas for finger tapping;

[0067] S2. Data collection: Install microphones inside the controlled object (in the form of a single microphone or multiple microphones in an array), and let multiple tappers tap the sound-sensing area with their fingers multiple times. The microphones collect the silence, the starting sound of the tap, the tail sound of the tap, and the noise.

[0068] S3, data labeling: label silence, the starting sound of knocking, and the tailing sound of knocking as positive sample data, with silence labeled as the first label, the starting sound of knocking as the second label, and the tailing sound of knocking as the third label; label noise as negative sample data, with noise labeled as the fourth label; the labeled positive sample data and negative sample data constitute a data set;

[0069] S4, data set noise addition, feature extraction, shuffling, and segmentation: add noise to the data set, extract features, and obtain data features with labeled data; after shuffling the data features, split them into training set, verification set, and test set (for example, in a ratio of 7:2:1);

[0070] S5. Neural network model design and training: A convolutional neural network model is used, and its loss function adopts a multi-classification cross entropy loss function; the training set, the verification set, and the test set are put into the convolutional neural network model for training to obtain a finger tapping detection neural network model;

[0071] S6, loading and compiling the neural network model for finger tapping detection;

[0072] S7, receiving a sound signal through a microphone, and using a finger tapping detection neural network model to predict and classify the sound signal into silence, a tapping start sound, a tapping tail sound, or noise;

[0073] S8, decoding each frame recognition result of the classified sound signal by the finger tapping detection neural network model, and caching the decoded sequence;

[0074] S9, the post-filter of the decoder analyzes the sequence to determine whether it is a valid tap; if so, the switch operation of the controlled target object is started; if not, the original state of the controlled target object is maintained.

[0075] The microphone in step S2 may be covered with a silicone sleeve to ensure the stability of the sound field, and the sound of finger tapping is transmitted into the microphone through the shell and inner cavity air of the controlled target object.

[0076] The finger tapping in step S2 includes finger tip tapping, finger web tapping and finger joint tapping.

[0077] The strength of the finger tapping in step S2 is determined by each tapper in a way that he or she considers comfortable.

[0078] The finger tapping method and number of times in step S2 are preferably but not limited to:

[0079] (1) Finger tapping: Each tapper taps M times in each sound-sensing area, including N consecutive taps of two times and (MN) consecutive taps of three times;

[0080] (2) Finger tapping: Each tapper taps M times in each sound-sensing area, including N consecutive taps of two times and (MN) consecutive taps of three times;

[0081] (3) Finger joint tapping: Each tapper taps M times in each sound-sensitive area, including N consecutive taps of two times and (MN) consecutive taps of three times.

[0082] The time between two consecutive taps is preferably, but not limited to, no more than 200ms.

[0083] The noise addition in step S4 is preferably but not limited to being performed on the positive sample data, while adding noise to the negative sample data or not adding noise to the negative sample data has little effect on the anti-noise performance of the algorithm.

[0084] Among them, the feature extraction in step S4 is preferably but not limited to using a balanced Fbank feature extraction method, such as Figure 4 As shown, the conventional Fbank feature extraction method uses nonlinear triangular filtering, while the balanced Fbank feature extraction method of the present invention uses linear triangular filtering, such as Figure 5 As shown, the specific process is:

[0085] (1) Assume that the original input data is x, the pi is FPI, the number of Fourier transformation points is N, the highest frequency is Max_freq, the lowest frequency is Min_freq, and the number of generated containers is M;

[0086] (2) Generate the jitter factor as follows:

[0087] Use the cosine function to limit the jitter factor to a range, that is, d_feat = sqrt(-2 * log(d_) *cos(2*FPI* d_), where d_ is a Gaussian distribution with a mean of 0 and a variance of 1;

[0088] The dither factor is dither = (d_feat * 2.0f) - 1.0f), where f represents a floating point number;

[0089] (3) Add a dither factor to the original data: x_D = x + dither; adding a dither factor is similar to adding micro noise, which can enhance robustness;

[0090] (4) After processing, take the average of every 400 points (i.e. 25ms), and then perform regularization, as follows:

[0091] Take the mean of x_D: x_bar = mean(x_D);

[0092] Subtract the DC component and make regularization to prevent large changes: x_f = x_D - x_bar;

[0093] (5) Pre-emphasis:

[0094] At the first point, the DC component is x_pre(i) = x_f(i) * 0.05, where i=0;

[0095] From the second point to the following points, make the difference: x_pre(i) = x_f(i) - x_f(i-1) *0.95, where i=1,2,3,……,N;

[0096] (6) Add a Hamming window, as follows:

[0097] The numerical value of the Hamming window is: ham = 0.54 - 0.46*cos((2FPI / (N-1)) * i), where i=0,2,3,…,N;

[0098] Data windowing: x_H = x_pre * ham, which can prevent spectrum leakage;

[0099] (7) Short-time Fourier transform of N points:

[0100] Transform from time domain to frequency domain, that is, X_H = fft(x_H, N);

[0101] (8) Obtain energy spectrum:

[0102] X_E = real(X_H) ^ 2 + imag(X_H) ^ 2, where real is the real part of the frequency domain point and imag is the imaginary part of the frequency domain point;

[0103] (9) Balanced generation of corresponding containers:

[0104] Spectrum bandwidth: SPAN_freq = (Max_freq- Min_freq);

[0105] Average container size: B_s = SPAN_freq / (M + 1);

[0106] Multiply the corresponding points: X_c =X_E* SPAN_freq[j * B_s:(j + 1)* B_s ], where j=0,2,3,…,M;

[0107] (10) Add the points in each container and then take the natural logarithm. Taking the logarithm can prevent the feature from jumping too much and can also observe the regularity of the feature, that is, obtain a balanced Fbank:

[0108] Unibank[j] =Log(sum( X_c[j* B_s:(j + 1)* B_s ])), where j=j=0,2,3,……,M.

[0109] Although there are many ways to extract features, such as MFCC, it has been tested that the balanced Fbank algorithm of the present invention is more universal for knocking sounds. The energy of human speech is mainly concentrated in the low-frequency band, while the energy of the high frequency is less. Taking into full consideration the human auditory characteristics and speech characteristics, Fbank was born. However, the purpose of the present invention is to improve the recognition rate of knocking detection. The low-frequency, medium-frequency and high-frequency energy distribution of knocking sounds is relatively uniform. As long as the machine can recognize it, there is no need to consider the human hearing sense, so a balanced Fbank is designed. There are two advantages to extracting features in this way: (1) The recognition rate is improved because it better fits the characteristics of the knocking sound; (2) When extracting features, the complexity of the balanced triangular filter is lower than that of the conventional triangular filter.

[0110] The post-filtering algorithm flow of the post-filter in step S9 is as follows:

[0111] (1) Put the sequence decoded by the neural network and the corresponding probability into buffers A and A_PRO;

[0112] (2) Analyze buffers A and A_PROB, merge the same and adjacent categories, put them into buffer B, record the length of consecutive intervals of the same category, put them into buffer B_SPAN, and calculate the mean probability within the interval length and put it into B_PROB;

[0113] (3) According to buffers B, B_SPAN and B_PROB: if it is the second label of 2-5 consecutive frames, put the second label into buffer C; if it is the third label of 4-10 consecutive frames, put the third label into C; otherwise, put the first label into buffer C and put the probabilities corresponding to the respective categories into C_PROB;

[0114] (4) Perform secondary decoding on buffers C, C_PROB, and B_SPAN. If the adjacent second and third tags are found in buffer C and their probabilities exceed the set thresholds, and the knocking duration does not exceed 27 frames and contains two or three knocking sounds, it is judged as a valid knocking. Otherwise, it is an invalid knocking.

[0115] Specific embodiment: Taking a living room vertical air conditioner as an example, the embodiment of the present invention is described.

[0116] like Figure 1~2 As shown, a method for detecting finger tapping sound to realize the switch of the living room vertical air conditioner includes:

[0117] S1. Sound sensing area division: the outer surface of the vertical air conditioner in the living room is divided into 10 sound sensing areas 11 for finger tapping, which are marked as A1, A2, A3, A4, A5, A6, A7, A8, A9, and A10 respectively;

[0118] S2. Data collection: A single microphone 12 is installed inside the vertical air conditioner in the living room. The microphone 12 is provided with a silicone cover. 26 tappers (half male and half female) are asked to tap the sound sensing area 11 with their fingers 10 times respectively. The microphone 12 collects the silence, the starting sound of the tapping, the tail sound of the tapping and the noise.

[0119] Finger tapping includes finger tip tapping, finger web tapping and finger joint tapping;

[0120] The intensity of finger tapping is determined by each tapper in a way that he or she finds comfortable.

[0121] The finger tapping method and number of times are preferably but not limited to:

[0122] (1) Fingertip tapping: Each tapper taps 10 times on each sound-sensing area 11, including 5 consecutive taps twice and 5 consecutive taps three times;

[0123] (2) Finger tapping: Each tapper taps 10 times on each sound-sensing area 11, including 5 consecutive taps twice and 5 consecutive taps three times;

[0124] (3) Finger joint tapping: Each tapper taps 10 times on each sound-sensing area 11, including 5 consecutive taps twice and 5 consecutive taps three times;

[0125] The continuous tapping means that the time between two adjacent tappings does not exceed 200ms;

[0126] S3, data annotation: mark the silence, the starting sound of the knock, and the tailing sound of the knock as positive sample data, such as Figure 3 As shown in the figure, silence is marked as 0, the start sound of a tap is marked as 1, and the tail sound of a tap is marked as 2; the noise is marked as negative sample data through a conventional speech recognition model, and the noise is marked as 3; the positive sample data and negative sample data after marking form a data set; among them, the speech recognition model mainly determines the time length of the noise audio for marking, and can also calculate the specific size of the noise audio to obtain the length of time, and then calculate the total number of frames (in phonetics, the professional unit of audio duration is generally expressed in frames), and each frame will be marked with the noise label 3;

[0127] S4, data set noise addition, feature extraction, shuffling, and segmentation: the data set is noised and feature extracted to obtain data features with labeled data; after the data features are shuffled, they are segmented into a training set, a verification set, and a test set in a ratio of 7:2:1; wherein the feature extraction adopts the above-mentioned balanced Fbank feature extraction method of 40 dimensions;

[0128] S5. Neural network model design and training: A convolutional neural network model is used, and its loss function adopts a multi-classification cross entropy loss function; the training set, verification set, and test set are put into the convolutional neural network model for training, and iterate 30 rounds to obtain a finger tapping detection neural network model;

[0129] S6, loading and compiling the neural network model for finger tapping detection;

[0130] S7, receiving a sound signal through a microphone, and using a finger tapping detection neural network model to predict and classify the sound signal into silence, a tapping start sound, a tapping tail sound, or noise;

[0131] S8, the finger tapping detection neural network model decodes each frame recognition result of the classified sound signal, and caches the decoded sequence, such as the sequence of 000011112222111222000, 0000111222000111222, 0001112220000000033333, etc., which can be set according to actual conditions such as tapping and marking;

[0132] S9. The post-filter of the universal back-end decoder parses the sequence to determine whether it is a valid knock. A valid knock is 2 or 3 consecutive knocks, and the rest are invalid knocks, such as one or more knocks and non-knock sounds. If so (such as the sequence is 000011112222111222000, 0000111222000111222, etc., which are valid knocks), the switch operation of the living room floor-standing air conditioner is started; if not (such as the sequence is 0001112220000000033333, etc., which are invalid knocks), the original state of the living room floor-standing air conditioner is maintained.

[0133] Since the data is labeled using a special labeling method, the knocking sound is composed of two labels (label 1 and label 2), so it is impossible to make a correct judgment by relying solely on the neural network classification results and conventional post-filtering, and special post-filtering is required. Figure 6 As shown, the post-filtering algorithm flow of the post-filter is as follows:

[0134] (1) Put the sequence decoded by the neural network and the corresponding probability into buffers A and A_PRO;

[0135] (2) Analyze buffers A and A_PROB, merge the same and adjacent categories, put them into buffer B, record the length of consecutive intervals of the same category, put them into buffer B_SPAN, and calculate the mean probability within the interval length and put it into B_PROB;

[0136] (3) According to buffers B, B_SPAN, and B_PROB: if the label is 1 for 2-5 consecutive frames, put label 1 into buffer C; if the label is 2 for 4-10 consecutive frames, put label 2 into C; otherwise, put label 0 into buffer C, and put the probabilities corresponding to the respective categories into C_PROB;

[0137] (4) Perform secondary decoding on buffers C, C_PROB, and B_SPAN. Find adjacent tags 1 and 2 in buffer C, and their probabilities exceed the set thresholds. If the knocking duration does not exceed 27 frames and contains two or three knocking sounds, it is judged as a valid knocking. Otherwise, it is an invalid knocking.

[0138] In step S1, the definition of the sound sensing area 11 refers to the area that allows effective tapping when designing the product. The number of divided sound sensing areas 11 can vary depending on the appearance of the target machine.

[0139] In step S2, the number of people collecting samples, the ratio of men to women, the number of taps, etc. can all be changed according to the specific situation. The above description is the effective data volume in the actual experimental process.

[0140] In step S3, generally, the starting sound is 3-5 frames, the tailing sound is 10-20 frames, and each frame is 10ms, which is directly related to the machine itself (such as material) and the installation position of the microphone 12. The noise of the negative sample data includes speech, music, TV news and various other noises.

[0141] In the actual test, the code was compiled into DSP and burned into the voice chip board. The microphone of the voice chip was installed on the upper inner wall of the vertical air conditioner in the living room. The test results are shown in the following table.

[0142] Table 1: Test results

[0143]

[0144] It can be seen from Table 1 above that in extreme environments, when the air conditioner is playing TV in the living room and the signal-to-noise ratio is very low, the recognition rate is still very high. When the external noise is very strong, false triggering is very rare, which meets the usage standard.

[0145] From the experimental results, the present invention divides the knocking sound (a target class) into the starting sound of the knocking and the tailing sound of the knocking (two subclasses). Such annotation can effectively improve the recognition rate of the decoding algorithm and reduce false awakening. From the algorithm level analysis, because the knocking sound (target class) is divided into two subclasses, the neural network learns more carefully, and the tailing sound of the knocking (subclass 2) plays a buffering transition role, which can effectively distinguish the target class (subclass 1) from the noise class (class 3, i.e. non-target sound or similar target sound). If the existing conventional annotation is used, there will be a high probability of false awakening when encountering a sound similar to a shock wave but not a knocking sound.

[0146] The above description is only a preferred embodiment of the present invention and does not limit the present invention in any form. Any simple modification, equalization change and modification made to the above embodiment by any technician familiar with the field without departing from the content of the technical solution of the present invention according to the technical essence of the present invention shall fall within the scope of the present invention.

Claims

1. A method for detecting finger tapping sound to realize switching, characterized in that: include: S1. Sound sensing area division: the outer surface of the controlled target object is divided into several sound sensing areas for finger tapping; S2. Data collection: Install a microphone inside the controlled target object, and let multiple tappers tap the sound-sensitive area with their fingers multiple times. The microphone collects the silence, the starting sound of the tap, the tail sound of the tap, and the noise; S3, data labeling: label silence, the starting sound of knocking, and the tailing sound of knocking as positive sample data, with silence labeled as the first label, the starting sound of knocking as the second label, and the tailing sound of knocking as the third label; label noise as negative sample data, with noise labeled as the fourth label; the labeled positive sample data and negative sample data constitute a data set; S4, data set noise addition, feature extraction, shuffling, and segmentation: add noise to the data set, extract features, and obtain data features with labeled data; after shuffling the data features, split it into a training set, a verification set, and a test set; S5. Neural network model design and training: A convolutional neural network model is used, and its loss function adopts a multi-classification cross entropy loss function; the training set, the verification set, and the test set are put into the convolutional neural network model for training to obtain a finger tapping detection neural network model; S6, loading and compiling the neural network model for finger tapping detection; S7, receiving a sound signal through a microphone, and using a finger tapping detection neural network model to predict and classify the sound signal into silence, a tapping start sound, a tapping tail sound, or noise; S8, decoding each frame recognition result of the classified sound signal by the finger tapping detection neural network model, and caching the decoded sequence; S9, the post-filter of the decoder analyzes the sequence to determine whether it is a valid tap; if so, the switch operation of the controlled target object is started; If not, the original state of the controlled object is maintained.

2. A method for detecting finger tapping sound to realize switching according to claim 1, characterized in that: The microphone cover in step S2 is provided with a silicone cover.

3. The method for detecting finger tapping sound to realize switching according to claim 1, characterized in that: The finger tapping in step S2 includes finger tip tapping, finger web tapping and finger joint tapping.

4. The method for detecting finger tapping sound to realize switching according to claim 1, characterized in that: The strength of the finger tapping in step S2 is determined by each tapper in a manner that he or she considers comfortable.

5. The method for detecting finger tapping sound to realize switching according to claim 1, characterized in that: The finger tapping method and number of times in step S2 are: 1) Finger tapping: Each tapper taps M times in each sound-sensing area, including N consecutive taps of two times and (MN) consecutive taps of three times; 2) Finger tapping: Each tapper taps M times in each sound-sensing area, including N consecutive taps twice and (MN) consecutive taps three times; 3) Finger joint tapping: Each tapper taps M times in each sound-sensitive area, including N consecutive taps twice and (MN) consecutive taps three times.

6. A method for detecting finger tapping sound to realize switching according to claim 5, characterized in that: The continuous tapping means that the time between two adjacent tappings does not exceed 200ms.

7. The method for detecting finger tapping sound to realize switching according to claim 1, characterized in that: The feature extraction in step S4 adopts the balanced Fbank feature extraction method.

8. The method for detecting finger tapping sound to realize switching according to claim 7, characterized in that: The specific process of the balanced Fbank feature extraction method is as follows: 1) Assume that the original input data is x, the pi is FPI, the number of Fourier transformation points is N, the highest frequency is Max_freq, the lowest frequency is Min_freq, and the number of generated containers is M; 2) Generate the jitter factor as follows: Use the cosine function to limit the jitter factor to a range, that is, d_feat = sqrt(-2 * log(d_) * cos(2*FPI* d_), where d_ is a Gaussian distribution with a mean of 0 and a variance of 1; The dither factor is dither = (d_feat * 2.0f) - 1.0f), where f represents a floating point number; 3) Original data plus dither factor: x_D = x + dither; 4) After processing, take the average of every 400 points, i.e. 25ms, and then do regularization, as follows: Take the mean of x_D: x_bar = mean(x_D); Subtract the DC component and make regularization to prevent large changes: x_f = x_D - x_bar; 5) Pre-emphasis: At the first point, the DC component is x_pre(i) = x_f(i) * 0.05, where i=0; From the second point to the following points, make the difference: x_pre(i) = x_f(i) - x_f(i-1) *0.95, where i=1,2,3,……,N; 6) Add a Hamming window as follows: The numerical value of the Hamming window is: ham = 0.54 - 0.46*cos((2FPI / (N-1)) * i), where i=0,2,3,…,N; Data windowing: x_H = x_pre * ham; 7) Short-time Fourier transform of N points: Transform from time domain to frequency domain, that is, X_H = fft(x_H, N); 8) Get the energy spectrum: X_E = real(X_H) ^ 2 + imag(X_H) ^ 2; 9) Balanced generation of corresponding containers: Spectrum bandwidth: SPAN_freq = (Max_freq- Min_freq); Average container size: B_s = SPAN_freq / (M + 1); Multiply the corresponding points: X_c =X_E* SPAN_freq[j * B_s:(j + 1)* B_s ], where j=0,2,3,…,M; 10) Add the points in each container and take the natural logarithm to get the balanced Fbank: Unibank[j] =Log(sum( X_c[j* B_s:(j + 1)* B_s ])), where j=j=0,2,3,……,M.

9. The method for detecting finger tapping sound to realize switching according to claim 1, characterized in that: The post-filtering algorithm flow of the post-filter in step S9 is as follows: 1) Put the sequence decoded by the neural network and the corresponding probability into buffers A and A_PRO; 2) Analyze buffers A and A_PROB, merge the same and adjacent categories, put them into buffer B, record the length of consecutive intervals of the same category, put them into buffer B_SPAN, and calculate the mean probability within the interval length and put it into B_PROB; 3) According to buffers B, B_SPAN and B_PROB: if it is the second label of 2-5 consecutive frames, put the second label into buffer C; if it is the third label of 4-10 consecutive frames, put the third label into C; otherwise, put the first label into buffer C and put the probabilities corresponding to the respective categories into C_PROB; 4) Perform secondary decoding on buffer C, C_PROB and B_SPAN, and find the adjacent second and third tags in buffer C, and their probabilities exceed the set thresholds respectively, and the knocking duration does not exceed 27 frames and contains two or three knocking sounds, it is judged as a valid knocking, otherwise it is an invalid knocking.

10. The method for detecting finger tapping sound to realize switching according to claim 1, characterized in that: The effective tapping in step S9 is 2 or 3 consecutive tappings.

Citation Information

Patent Citations

  • Control method and device of sound control tapping switch, household appliance and storage medium

    CN112272019A

  • Intelligent knocking detection method and system for structural defects

    CN114324580A