An interactive analysis instrument control system and control method integrating speech recognition
Through spectrum difference analysis and frequency domain fusion recognition technology of multi-microphone system, combined with operation feedback correction, the accuracy and consistency of speech recognition of interactive analytical instruments in noisy environments is solved, and the stability and user experience of the system are improved.
Patent Information
- Application Number
- CN202411502165.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-25
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2044-10-25
AI Technical Summary
Existing interactive analytical instruments have low speech recognition accuracy in multi-noise environments, and lack of reference comparison for single subject microphone collection, resulting in misidentification and operation command deviation.
The multi-microphone system is used to analyze the spectrum difference, select the main body and auxiliary microphone, divide the voice segments through voice activity detection, and perform frequency domain fusion recognition. The final operation instructions are determined through comparison and secondary correction is performed with operation feedback.
It improves the accuracy and robustness of speech recognition, reduces misidentification, improves signal quality and consistency of operation instructions, and provides better user experience and system stability.
Smart Images

Figure CN119541481B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of voice interaction control, and particularly relates to an interactive analysis instrument control system and control method integrating voice recognition. Background Art
[0002] An interactive analysis instrument refers to a scientific instrument that has the ability to communicate bidirectionally with users and can automatically adjust its operating behavior according to user input or instructions. Such instruments are widely used in laboratory tests in multiple fields such as chemistry, biology, and physics (such as analytical instruments), and industrial production fields (such as production line equipment). By directly interacting with the device through voice commands without manual operation or keyboard input, the operation efficiency is significantly improved, especially in cases where repetitive tasks need to be performed frequently.
[0003] In industrial and laboratory environments, interactive analysis instruments usually face interference from multiple noise sources, such as the operating sounds of fans, pumps, and other mechanical equipment. These noise sources will cause significant interference to the acquisition of voice signals and reduce the accuracy of voice recognition. In addition, due to the large volume of such devices, the distance between the microphone and the operator may be relatively far, resulting in serious attenuation of voice signals during transmission. To overcome these problems, it is currently common to arrange multiple microphones on the device to locate the sound source direction to reduce the interference of voice acquisition.
[0004] However, existing multi-microphone systems are usually only used to determine the main microphone and operate according to the voice commands collected by the main microphone. The selection of the main microphone is usually based on the principle of the closest distance to the sound source, but this method has some inherent limitations. On the one hand, even the microphone closest to the sound source may have poor signal quality at its location, thus affecting the effect of voice recognition. On the other hand, since the voice signals collected by a single main microphone lack reference and comparison, it is difficult to determine whether the collected voice is missing, and there is a hidden danger of not fully reflecting all the characteristics of the voice, which may lead to misrecognition of the voice recognition system in some cases, resulting in a deviation between the generated operation instruction text and the actual situation, and further affecting the interaction effect of the system. Summary of the Invention
[0005] The purpose of the present invention is to improve the deficiencies existing in the prior art, and provide an interactive analysis instrument control system and control method integrating voice recognition. By optimizing the way of collecting voice signals by the multi-microphone system, the defects existing in voice control based only on the main microphone can be minimized.
[0006] The object of the present invention can be achieved by the following technical solutions: In the first aspect of the present invention, an interactive analysis instrument control system integrating speech recognition is provided, including the following modules: a voice signal induction and acquisition module, which is used to perform human body induction by using an infrared induction terminal arranged on the interactive analysis instrument, and when a human body is sensed to be approaching, start multiple microphones arranged on the interactive analysis instrument to collect voice signals emitted by the operator.
[0007] A microphone selection module, which is used to perform spectral difference analysis on the voice signals collected by each microphone, thereby locating the position of the sound source, and selecting a main microphone and an auxiliary microphone with the sound source position as the center.
[0008] A voice segment division module, which is used to divide the voice signals collected by the main microphone and the auxiliary microphone into voice segments and non-voice segments by using a voice activity detection algorithm, and retain the voice segments.
[0009] A single-voice recognition module, which is used to input the voice segments retained by the main microphone and the auxiliary microphone into a voice recognition system to generate operation instruction texts generated by the main microphone and the auxiliary microphone.
[0010] A fused speech recognition module, which is used to perform frequency-domain fusion on the voice segments retained by the main microphone and the auxiliary microphone, and input the fused voice into a voice recognition system to generate an operation instruction text.
[0011] An operation instruction determination module, which is used to compare the operation instruction text generated by voice fusion with the operation instruction texts generated by the main microphone and the auxiliary microphone, thereby determining the final operation instruction.
[0012] An operation execution feedback module, which is used to obtain the feedback information of the operator during the operation of the interactive analysis instrument according to the final operation instruction, and when the feedback indicates that the operation fails, select an effective operation instruction text from the operation instruction texts previously generated by the main microphone and the auxiliary microphone for secondary operation.
[0013] In the second aspect of the present invention, a control method for an interactive analysis instrument integrating speech recognition is proposed, including the following steps: S1. Perform human body induction by using an infrared induction terminal arranged on the interactive analysis instrument, and when a human body is sensed to be approaching, start multiple microphones arranged on the interactive analysis instrument to collect voice signals emitted by the operator.
[0014] S2. Perform spectral difference analysis on the voice signals collected by each microphone, thereby locating the position of the sound source, and selecting a main microphone and an auxiliary microphone with the sound source position as the center.
[0015] S3. Divide the voice signals collected by the main microphone and the auxiliary microphone into voice segments and non-voice segments by using a voice activity detection algorithm, and retain the voice segments.
[0016] S4. Input the voice segments retained by the main microphone and the auxiliary microphone into the speech recognition system to generate the operation instruction text generated by the main microphone and the auxiliary microphone.
[0017] S5. Perform frequency-domain fusion on the voice segments retained by the main microphone and the auxiliary microphone, and input the fused voice into the speech recognition system to generate the operation instruction text.
[0018] S6. Compare the operation instruction text generated by voice fusion with the operation instruction text generated by the main microphone and the auxiliary microphone to determine the final operation instruction.
[0019] S7. Obtain the operator's feedback information during the operation of the interactive analysis instrument according to the final operation instruction.
[0020] S8. When the feedback indicates that the operation fails, select the valid operation instruction text from the operation instruction text previously generated by the main microphone and the auxiliary microphone for a secondary operation.
[0021] Combining all the above technical solutions, the positive effects of the present invention are as follows: 1. By using the voice signals collected by multiple microphones on the interactive analysis instrument to perform spectral difference analysis, the main microphone and the auxiliary microphone are selected, and then the voice signals collected by the main microphone and the auxiliary microphone are subjected to single recognition and fusion recognition, so as to determine the operation instruction by combining the recognition results of the two, which can significantly improve the accuracy, robustness and reliability of speech recognition, reduce misrecognition, improve the signal quality and the consistency of operation instructions, and thus provide a better user experience and interaction effect.
[0022] 2. The present invention obtains the operator's feedback information during the operation of the interactive analysis instrument according to the determined operation instruction. When the feedback indicates that the operation fails, select the valid operation instruction text from the operation instruction text previously generated by the main microphone and the auxiliary microphone for a secondary operation. In this way, the situation of operation failure can be quickly discovered by obtaining the operator's immediate feedback, and measures can be taken to correct it, avoiding the chain errors of subsequent operations caused by a failed operation instruction. At the same time, a secondary operation can be quickly performed after the first operation fails, reducing the operation interruption time and improving the continuity and stability of the system. Description of the Drawings
[0023] The present invention is further described with reference to the accompanying drawings, but the embodiments in the drawings do not constitute any limitation to the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the following drawings without creative efforts.
[0024] Figure 1 It is a schematic diagram of the system module connection of the present invention.
[0025] Figure 2 This is a schematic diagram of voice signal induction and acquisition in the present invention.
[0026] Figure 3 This is a flowchart of the method implementation steps of the present invention. Specific implementation manners
[0027] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0028] Embodiment 1
[0029] The present invention provides an interactive analysis instrument control system integrating voice recognition, including a voice signal induction and acquisition module, a microphone selection module, a voice segment division module, a single voice recognition module, a fused voice recognition module, an operation instruction determination module, and an operation execution feedback module. The voice signal induction and acquisition module is connected to the microphone selection module, the microphone selection module is connected to the voice segment division module, the voice segment division module is respectively connected to the single voice recognition module and the fused voice recognition module, both the single voice recognition module and the fused voice recognition module are connected to the operation instruction determination module, and both the operation instruction determination module and the single voice recognition module are connected to the operation execution feedback module.
[0030] The voice signal induction and acquisition module is used to perform human body induction by using an infrared induction terminal provided on the interactive analysis instrument, and when a human body is sensed to be approaching, it activates a multi-microphone arranged on the interactive analysis instrument to collect voice signals emitted by the operator.
[0031] For the activation and triggering of the multi-microphone by the above infrared induction terminal, refer to Figure 2 as shown.
[0032] In the operation of the above solution, the infrared induction terminal can be an infrared sensor, exemplarily a pyroelectric infrared sensor, which mainly senses the presence of a human body by detecting infrared rays radiated by the human body.
[0033] It should be understood that the purpose of setting the infrared induction terminal on the interactive analysis instrument is to detect the presence of a human body. When a human body approaches the instrument, the sensor will trigger the activation of the microphone, which can ensure that the microphone is ready to receive voice commands when the operator is ready to interact with the instrument. In addition, by using the infrared induction terminal to activate the microphone system only when someone is approaching, power resources can be saved and the service life of the device can be extended.
[0034] The microphone selection module is used to perform spectral difference analysis on the voice signals collected by each microphone, thereby locating the position of the sound source, and selecting the main microphone and the auxiliary microphone with the sound source position as the center.
[0035] As a preferred implementation of the above solution, the spectral difference analysis of the voice signals collected by each microphone is as follows: convert the voice signals collected by each microphone into the frequency domain to obtain a spectrogram. The spectrogram of the voice signal is constructed with the frequency as the horizontal axis and the amplitude (or energy) as the vertical axis. The spectrogram shows the energy distribution of the voice signal at different frequencies, and marks the position of the energy peak from the spectrogram, and then extracts the frequency component corresponding to the energy peak as the main frequency component of the voice signal collected by each microphone.
[0036] It should be noted that the peak position in the spectrogram indicates that there is a strong component of the voice signal at a certain frequency, that is, the main frequency component.
[0037] Overlap the main frequency components of the voice signals collected by each microphone, and extract the overlapping main frequency components.
[0038] Extract the phase values of each frequency point from the spectrograms of the voice signals collected by each microphone within the overlapping main frequency components.
[0039] In the optimized implementation of the above solution, before extracting the phase values of each frequency point, it is necessary to divide the frequency points of the overlapping main frequency components. The specific division can determine the frequency division interval according to the sampling frequency of the spectrum and the window size. The sampling frequency refers to the number of signal samples collected per second, and the frequency resolution refers to the minimum frequency interval that the spectrum can distinguish. The frequency resolution depends on the sampling frequency of the spectrum and the length of the window. Specifically, where FR represents the frequency resolution, f s represents the sampling frequency, N represents the number of samples in the spectrum analysis window, and the frequency division interval is usually the frequency resolution. Exemplarily, when using 1024 sample points for spectrum analysis and the sampling frequency is 8000 Hz, the frequency division interval is That is, the interval between each frequency point is about 7.81 Hz.
[0040] Number the microphones arranged on the interactive analysis instrument, and form the microphones into pairs in the order of the numbers to obtain a number of microphone groups.
[0041] In the example of the above solution, it is assumed that there are 4 microphones arranged on the interactive analysis instrument. The microphones can be numbered with a certain point on the interactive analysis instrument as the reference point, such as the center position of the instrument. According to the order of the distances from the arranged positions of each microphone to the reference point from near to far, the microphones are numbered. According to the numbering order, the microphones are grouped in pairs to obtain several microphone groups. The specific grouping method is as follows: Group 1: Microphone 1 and Microphone 2; Group 2: Microphone 1 and Microphone 3; Group 3: Microphone 1 and Microphone 4; Group 4: Microphone 2 and Microphone 3; Group 5: Microphone 2 and Microphone 4; Group 6: Microphone 3 and Microphone 4.
[0042] The phase difference values of the two microphones in each microphone group at each frequency point are subtracted to form a phase difference set of each microphone group, and the average value of the existing phase differences in the set is taken to obtain the average phase difference corresponding to each microphone group.
[0043] It should be explained that for different microphones at the same frequency point, there will be different phase values. By calculating the phase difference, the phase difference between two microphones at the same frequency point can be reflected.
[0044] Furthermore, it should be explained that due to the influence of factors such as noise, the phase difference at a single frequency point may not be accurate enough. By calculating the average value of the phase differences at multiple frequency points, the influence of errors can be reduced to a certain extent.
[0045] As another preferred implementation of the above solution, the location of the sound source is implemented as follows: A three-dimensional coordinate system is constructed on the interactive analysis instrument, and thus the coordinates of the arranged positions of each microphone are obtained.
[0046] Specifically, the process of constructing the three-dimensional coordinate system is as follows: First, a fixed reference point is selected as the origin o of the three-dimensional coordinate system. This reference point can be the center point of the instrument. Next, the three coordinate axes x, y, and z of the three-dimensional coordinate system are defined. Among them, the x-axis is defined as the horizontal direction and can be along the length direction of the instrument. The y-axis is defined as the width direction and is perpendicular to the x-axis. The z-axis is defined as the height direction and is perpendicular to the x-axis and the y-axis.
[0047] More specifically, for each microphone, the horizontal distance of the microphone along the x-axis direction is measured, that is, the distance from the reference point o to the horizontal projection point of the microphone. The horizontal distance of the microphone along the y-axis direction is measured, that is, the distance from the reference point o to the horizontal projection point of the microphone. The height of the microphone is measured, that is, the height distance from the reference point o to the microphone. Thus, the measured x, y, and z coordinates are recorded to form the position coordinates of the microphone.
[0048] The present invention determines the position of the microphone by defining a three-dimensional coordinate system, which can significantly improve the accuracy and robustness of the system, improve the user experience, and provide a solid foundation for subsequent sound source positioning. This coordinate system definition method has wide applicability and superiority in practical applications due to its advantages such as clear direction definition, convenient measurement and recording, and improved positioning accuracy.
[0049] The average phase difference of each microphone group is converted into the time difference of the sound wave arriving at the two microphones.
[0050] It should be noted that the present invention uses a positioning technology based on time difference to locate the sound source. This method uses the time difference of sound waves propagating between different receiving points (such as microphones) to determine the position of the sound source.
[0051] In the above specific embodiment of the time difference, assuming that the position coordinates of a microphone group are (x1, y1, z1) and (x2, y2, z2), the average phase difference is Then the time difference of the sound wave reaching the microphone group is Where f is the frequency of the sound wave signal and π represents pi.
[0052] The time difference of the sound waves reaching the two microphones in each microphone group is combined with the propagation speed of the sound waves to calculate the sound wave propagation distance difference corresponding to each microphone group.
[0053] In the implementation of the above scheme, since the propagation speed c of the sound wave is known (approximately 342m / s under standard atmospheric pressure), the time difference between the sound waves reaching the two microphones is combined with the sound wave propagation speed c to calculate the sound wave propagation distance difference, and the specific calculation formula is Δd=c*Δt.
[0054] A hypersphere equation is established based on the coordinates of the locations of the two microphones in each microphone group and the difference in sound wave propagation distances.
[0055] It should be noted that when the difference in sound wave propagation distance of a microphone group is known, the sound source must be located on a hypersphere with the line connecting the two microphones as the diameter. For example, the equation of the hypersphere is (x-x1) 2 +(y-y1) 2 +(z-z1) 2 =(x-x2) 2 +(y-y2) 2 +(z-z2) 2 +(Δd) 2 If more microphone groups are added, multiple such hyperspheres can be obtained, and the intersection of these hyperspheres is the location of the sound source.
[0056] The hyper-sphere equations established by each microphone group form a non-linear equation system, and then the position coordinates of the sound source are obtained by numerically solving this equation system.
[0057] In the above, the least squares method can be used to solve the equation system.
[0058] As a further preferred implementation of the above solution, the main microphone is selected centered on the position of the sound source as follows: The position coordinates of the sound source are compared with the position coordinates of each microphone to obtain the separation distances between the sound source and each microphone, and the microphone with the shortest separation distance is extracted as the main microphone.
[0059] As a further preferred implementation of the above solution, the auxiliary microphones are selected as follows: A sphere is made with the position coordinates of the sound source as the center and a defined distance as the radius, and the area inside the sphere is the selection range. Then, the microphones other than the main microphone that fall within the selection range are used as alternative microphones.
[0060] In particular, when determining the defined distance, multiple factors need to be considered, including the arrangement spacing of the microphones, the expected position of the sound source, etc. It is usually recommended to be 3 meters (laboratory environment) or 5 meters (industrial environment).
[0061] Based on the spectrogram of the voice signals collected by the alternative microphones, the signal-to-noise ratio of the alternative microphones is calculated. The specific operation is as follows: The collected voice signals are divided into voice segments and non-voice segments using the voice activity detection algorithm. Among them, the non-voice segments are the noise segments.
[0062] The main frequency components of the voice segments and their corresponding energy distribution characteristics are extracted from the spectrogram, and thus the signal energy of the voice segments is calculated. The calculation of the signal energy of the voice segments can use the short-time energy calculation method, that is, the sum of the squared amplitudes of the signals within the voice segments is calculated.
[0063] The main frequency distribution of the non-voice segments and their corresponding energy distribution characteristics are extracted from the spectrogram, and similarly, the signal energy of the non-voice segments is calculated by referring to the signal energy calculation of the voice segments.
[0064] The mean values of the signal energies of all voice segments and the signal energies of all non-voice segments in the voice signals collected by the alternative microphones are calculated, and the signal-to-noise ratio of the voice signals collected by the alternative microphones is calculated using the calculation results. Specifically, the signal-to-noise ratio calculation formula is In the formula respectively represent the average signal energy of the voice segments and the average signal energy of the non-voice segments.
[0065] The signal-to-noise ratio of the alternative microphones is compared with the preset critical signal-to-noise ratio. The purpose of setting the critical signal-to-noise ratio is to provide assistance for screening auxiliary microphones with good signal quality, and the alternative microphones that reach the critical signal-to-noise ratio are extracted as auxiliary microphones.
[0066] In the selection range formed by taking the position coordinates of the sound source as the center and the defined distance as the radius to form a sphere, the present invention screens out the auxiliary microphones with good signal-to-noise ratio according to the signal-to-noise ratio of the collected voice signals, rather than directly selecting the microphones falling within the selection range, which can significantly improve the reliability of voice signal acquisition. This method can better cope with the challenges of speech recognition in complex environments, improve the quality of voice signals, and reduce the risk of misrecognition.
[0067] The voice segment division module is used to divide the voice signals collected by the main microphone and the auxiliary microphones into voice segments and non-voice segments by using the voice activity detection algorithm, and retain the voice segments.
[0068] The single-voice recognition module is used to input the voice segments retained by the main microphone and the auxiliary microphones into the speech recognition system to generate the operation instruction texts generated by the main microphone and the auxiliary microphones.
[0069] The fused speech recognition module is used to perform frequency-domain fusion on the voice segments retained by the main microphone and the auxiliary microphones, and input the fused voice into the speech recognition system to generate the operation instruction text.
[0070] In an implementable manner, the frequency-domain fusion of the voice segments retained by the main microphone and the auxiliary microphones is as follows: Locate the frequency region where the voice segment is located in the spectrogram corresponding to the main microphone as the main frequency region, and divide the frequency region into frequency points, and then extract the frequency values of each frequency point as the main spectrum values.
[0071] Locate the main frequency region in the spectrogram corresponding to the auxiliary microphone, and similarly extract the spectrum values of each frequency point as the auxiliary frequency values.
[0072] Fuse the main spectrum values and the auxiliary spectrum values of each frequency point to obtain the fused frequency-domain signal.
[0073] Convert the fused frequency-domain signal back to the time-domain signal through the inverse fast Fourier transform to obtain the fused voice signal.
[0074] It should be understood that performing spectrum fusion on the voice signals collected by the main microphone and the auxiliary microphones can play a role in signal enhancement. The fused voice signal has a higher signal-to-noise ratio, better voice quality, and lower misrecognition risk.
[0075] The operation instruction determination module is used to compare the operation instruction text generated by voice fusion with the operation instruction texts generated by the main microphone and the auxiliary microphone, so as to determine the final operation instruction. The specific operations are as follows: The operation instruction text generated by voice fusion and the operation instruction texts generated by the main microphone and the auxiliary microphone are respectively segmented to obtain several segments of each operation instruction text, and the segmented segments are numbered according to their positions in the operation instruction text.
[0076] In the example of the above solution, the operation instruction text generated by voice fusion is "Turn on the device and increase the temperature", the operation instruction text generated by the main microphone is "Turn on the device and increase the humidity", and the operation instruction text generated by the auxiliary microphone is "Turn on the device and decrease the pressure". The segmented results of the operation instruction text generated by voice fusion in this example are ["Turn on", "device", "and", "increase", "temperature"], the segmented results of the operation instruction text generated by the main microphone are ["Turn on", "device", "and", "increase", "humidity"], and the segmented results of the operation instruction text generated by the auxiliary microphone are ["Turn on", "device", "and", "decrease", "pressure"]. The segmented segments are numbered 1, 2, 3, 4, 5 in sequence according to the order.
[0077] Extract the segments with the same number in the operation instruction text generated by voice fusion, the operation instruction text of the main microphone, and the operation instruction text of the auxiliary microphone in sequence according to the number order for comparison. If the segments in multiple texts are the same, select the same segments as the final content. If they are different, select the segment with the most occurrences as the final content. If the occurrence times are evenly distributed, select the segment in the operation instruction text generated by voice fusion as the final content.
[0078] In a further example, the 1st segment "Turn on" in the segmented segments of the operation instruction text generated by voice fusion, the operation instruction text of the main microphone, and the operation instruction text of the auxiliary microphone is the same, and the final content is "Turn on"; the 2nd segment "device" is the same, and the final content is "device"; the 3rd segment "and" is the same, and the final content is "and"; the 4th segments "increase" and "decrease" are different, but "increase" appears the most times, so the final content is "increase"; the 5th segments are "temperature", "humidity", and "pressure", which are different and the occurrence times are evenly distributed, so the final content is "temperature".
[0079] Combine the final content of each segment into the final operation instruction. In the above example, the final operation instruction is "Turn on the device and increase the temperature".
[0080] It should be explained that when there are inconsistent word segmentations in different operation instruction texts, the word segmentation content in the voice fusion operation instruction text is preferentially adopted because the risk of misrecognition of the fused voice signal is more advantageous than that of the main microphone operation instruction text and the auxiliary microphone operation instruction text.
[0081] The operation execution feedback module is used to obtain the operator's feedback information during the operation of the interactive analysis instrument according to the final operation instruction. When the feedback indicates that the operation fails, a valid operation instruction text is selected from the operation instruction texts previously generated by the main microphone and the auxiliary microphone for secondary operation.
[0082] In the innovative implementation of the above solution, the selected valid operation instruction text is: select the operation instruction text generated by the main microphone from the operation instruction texts previously generated by the main microphone and the auxiliary microphone as the valid operation instruction text.
[0083] It should be understood that the reason for using the operation instruction text generated by the main microphone as the valid operation instruction text is that the main microphone is usually the closest to the sound source, so its signal quality is usually better and the reliability is higher. When the operation fails, directly using the operation instruction text generated by the main microphone for secondary operation can simplify the processing logic and reduce the algorithm complexity.
[0084] Embodiment 2
[0085] Refer to Figure 3 As shown, the present invention proposes a control method for an interactive analysis instrument integrating voice recognition, including the following steps: S1. Use the infrared induction terminal set on the interactive analysis instrument for human body induction, and start the multi-microphones arranged on the interactive analysis instrument to collect the voice signals emitted by the operator when a human body is sensed to be approaching.
[0086] S2. Perform spectral difference analysis on the voice signals collected by each microphone, thereby locating the position of the sound source, and selecting the main microphone and the auxiliary microphone with the sound source position as the center.
[0087] S3. Use the voice activity detection algorithm to divide the voice signals collected by the main microphone and the auxiliary microphone into voice segments and non-voice segments, and retain the voice segments.
[0088] S4. Input the voice segments retained by the main microphone and the auxiliary microphone into the voice recognition system to generate the operation instruction texts generated by the main microphone and the auxiliary microphone.
[0089] S5. Perform frequency domain fusion on the voice segments retained by the main microphone and the auxiliary microphone, and input the fused voice into the voice recognition system to generate the operation instruction text.
[0090] S6. Compare the operation instruction text generated by voice fusion with the operation instruction text generated by the main microphone and the auxiliary microphone, thereby determining the final operation instruction.
[0091] S7. During the operation of the interactive analysis instrument according to the final operation instruction, obtain the feedback information of the operator.
[0092] S8. If the feedback indicates that the operation fails, select the valid operation instruction text from the operation instruction text previously generated by the main microphone and the auxiliary microphone for secondary operation.
[0093] The above content is only an example and illustration of the structure of the present invention. Those skilled in the art of the present technology can make various modifications or supplements to the described specific embodiments or use similar methods for substitution. As long as they do not deviate from the structure of the invention or exceed the scope defined by the present invention, they should fall within the protection scope of the present invention.
Claims
1. An interactive analysis instrument control system integrating speech recognition, characterized in that, It includes the following modules: A voice signal sensing and acquisition module, which is used to sense a human body by using an infrared sensing terminal set on an interactive analysis instrument, and when a human body is sensed to approach, start multiple microphones arranged on the interactive analysis instrument to collect voice signals emitted by an operator; A microphone selection module, which is used to perform spectral difference analysis on the voice signals collected by each microphone, thereby locating the position of the sound source, and selecting a main microphone and an auxiliary microphone with the sound source position as the center; A voice segment division module, which is used to divide the voice signals collected by the main microphone and the auxiliary microphone into voice segments and non-voice segments by using a voice activity detection algorithm, and retain the voice segments; A single voice recognition module, which is used to input the voice segments retained by the main microphone and the auxiliary microphone into a voice recognition system to generate an operation instruction text generated by the main microphone and the auxiliary microphone; A fused voice recognition module, which is used to perform frequency domain fusion on the voice segments retained by the main microphone and the auxiliary microphone, and input the fused voice into a voice recognition system to generate an operation instruction text; An operation instruction determination module, which is used to compare the operation instruction text generated by voice fusion with the operation instruction texts generated by the main microphone and the auxiliary microphone, thereby determining the final operation instruction; An operation execution feedback module, which is used to obtain the feedback information of an operator during the process that the interactive analysis instrument executes an operation according to the final operation instruction, and when the feedback indicates that the operation fails, select a valid operation instruction text from the operation instruction texts previously generated by the main microphone and the auxiliary microphone for a secondary operation; The process of determining the final operation instruction is as follows: Perform word segmentation on the operation instruction text generated by voice fusion, the operation instruction text generated by the main microphone, and the operation instruction text generated by the auxiliary microphone respectively to obtain several word segments divided from each operation instruction text, and number the divided word segments according to their positions in the operation instruction text; Extract the word segments with the same number in the voice fusion operation instruction text, the main microphone operation instruction text, and the auxiliary microphone operation instruction text in sequence according to the number order for comparison. If the word segments in multiple texts are the same, select the identical word segments as the final content. If they are inconsistent, select the word segment with the most occurrences as the final content. If the occurrence distribution is uniform, select the word segment in the voice fusion operation instruction text as the final content; Combine the final content of each word segment into the final operation instruction.
2. The interactive analysis instrument control system integrating speech recognition according to claim 1, characterized in that: The process of performing spectral difference analysis on the voice signals collected by each microphone is as follows: Convert the voice signals collected by each microphone to the frequency domain to obtain a spectrogram, and mark the position of the energy peak from the spectrogram, and then extract the frequency components corresponding to the energy peak as the main frequency components of the voice signals collected by each microphone; Overlap the main frequency components of the voice signals collected by each microphone, and extract the overlapping main frequency components; Extract the phase values of each frequency point from the spectrograms of the voice signals collected by each microphone within the overlapping main frequency components; Number the microphones arranged on the interactive analysis instrument, and form several microphone groups by pairing the microphones two by two in the order of the numbers; The phase values of two microphones in each microphone group at each frequency point are subtracted to form a phase difference set of each microphone group, and the phase differences in the set are averaged to obtain an average phase difference corresponding to each microphone group.
3. The interactive analysis instrument control system integrating speech recognition according to claim 2, characterized in that: The location of the localization sound source is implemented as follows: A three-dimensional coordinate system is constructed on the interactive analysis instrument, thereby obtaining the coordinates of the layout positions of each microphone; Convert the average phase difference of each microphone group into the time difference of the sound wave arriving at the two microphones; The time difference between the sound waves in each microphone group reaching the two microphones is combined with the propagation speed of the sound waves to calculate the sound wave propagation distance difference corresponding to each microphone group; A hypersphere equation is established based on the coordinates of the locations of the two microphones in each microphone group and the difference in sound wave propagation distances; The hypersphere equations established by each microphone group are formed into a nonlinear equation group, and then the position coordinates of the sound source are obtained by solving the equation group through numerical calculation.
4. The interactive analysis instrument control system integrating speech recognition according to claim 3, characterized in that: The process of selecting the main microphone based on the sound source position is as follows: The position coordinates of the sound source are compared with the position coordinates of each microphone to obtain the distance between the sound source and each microphone, and the microphone with the shortest distance is extracted as the main microphone.
5. The interactive analysis instrument control system integrating speech recognition according to claim 3, characterized in that: The auxiliary microphone is selected as follows: A sphere is constructed with the position coordinates of the sound source as the center and the limited distance as the radius. The area inside the sphere is the selection range, and the microphones other than the main microphone within the selection range are selected as candidate microphones. Calculate the signal-to-noise ratio of the candidate microphone based on the spectrum of the speech signal collected by the candidate microphone; The signal-to-noise ratio of the candidate microphones is compared with a preset critical signal-to-noise ratio, and the candidate microphones reaching the critical signal-to-noise ratio are extracted as auxiliary microphones.
6. The interactive analysis instrument control system integrating speech recognition according to claim 5, characterized in that: The calculation of the signal-to-noise ratio of the candidate microphone based on the spectrum of the speech signal collected by the candidate microphone is implemented as follows: Based on the divided speech segments and non-speech segments, the speech segments and non-speech segments are marked in the spectrum graph of the candidate microphone, and then the energy distribution characteristics of the speech segments and non-speech segments are respectively extracted; The signal energy of the speech segment is calculated based on the energy distribution characteristics of the speech segment, and the signal energy of the non-speech segment is calculated similarly; The signal energies of all speech segments and the signal energies of all non-speech segments corresponding to each candidate microphone are averaged to thereby calculate the signal-to-noise ratio of the candidate microphone.
7. The interactive analysis instrument control system integrating speech recognition according to claim 1, characterized in that: The frequency domain fusion of the voice segments retained by the main microphone and the auxiliary microphone is performed as follows: Locate the frequency region where the speech segment is located in the spectrum graph corresponding to the subject microphone as the subject frequency region, divide the frequency region into frequency points, and then extract the frequency value of each frequency point as the subject spectrum value; Locate the main frequency area in the spectrum diagram corresponding to the auxiliary microphone, and similarly extract the auxiliary spectrum value of each frequency point; The main spectrum value and the auxiliary spectrum value of each frequency point are fused to obtain a fused frequency domain signal; The fused frequency domain signal is converted back to the time domain signal through inverse fast Fourier transform to obtain the fused speech signal.
8. The interactive analysis instrument control system integrating speech recognition according to claim 1, wherein: The process of selecting the valid operation instruction text is as follows: The operation instruction text generated by the main microphone is selected as the valid operation instruction text from the operation instruction texts previously generated by the main microphone and the auxiliary microphone.
9. An interactive analysis instrument control method integrating speech recognition, characterized in that, The following steps are involved: S1. Use the infrared induction terminal set on the interactive analysis instrument to perform human body induction. When a human body is sensed approaching, start the multi-microphone arranged on the interactive analysis instrument to collect the voice signal emitted by the operator. S2. Perform spectrum difference analysis on the voice signals collected by each microphone, thereby locating the position of the sound source, and select the main microphone and auxiliary microphone with the sound source position as the center. S3. Use the voice activity detection algorithm to divide the voice signals collected by the main microphone and the auxiliary microphone into voice segments and non-voice segments, and retain the voice segments. S4. Input the voice segments retained by the main microphone and the auxiliary microphone into the voice recognition system to generate the operation instruction text generated by the main microphone and the auxiliary microphone. S5. Perform frequency domain fusion on the voice segments retained by the main microphone and the auxiliary microphone, and input the fused voice into the voice recognition system to generate the operation instruction text. S6. Compare the operation instruction text generated by voice fusion with the operation instruction texts generated by the main microphone and the auxiliary microphone, thereby determining the final operation instruction. The process of determining the final operation instruction is as follows: Perform word segmentation on the operation instruction text generated by voice fusion, the operation instruction texts generated by the main microphone and the auxiliary microphone respectively, to obtain several word segments divided by each operation instruction text, and number the divided word segments according to their positions in the operation instruction text. Extract the word segments with the same number in the voice fusion operation instruction text, the main microphone operation instruction text, and the auxiliary microphone operation instruction text in sequence according to the number order for comparison. If the word segments in multiple texts are the same, select the same word segments as the final content. If they are inconsistent, select the word segment with the most occurrences as the final content. If the occurrence distribution is uniform, select the word segment in the voice fusion operation instruction text as the final content. Combine the final content of each word segment into the final operation instruction. S7. Obtain the feedback information of the operator during the operation of the interactive analysis instrument according to the final operation instruction. S8. When the feedback indicates that the operation fails, select the valid operation instruction text from the operation instruction texts previously generated by the main microphone and the auxiliary microphone for secondary operation.
Citation Information
Patent Citations
Microphone reception method, mobile terminal and computer readable memory medium
CN108737615A
Voice signal processing method, assembly, device and medium
CN109509465A