Intelligent voice interaction system and method

By using a multi-channel acoustic sensor array and intelligent semantic repair technology, the problems of noise interference and semantic understanding in traditional intelligent voice interaction in complex environments have been solved, achieving efficient and accurate voice interaction in complex environments and improving user experience and device control convenience.

CN121583247BActive Publication Date: 2026-05-12FUJIAN REIDA PRECISION
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
FUJIAN REIDA PRECISION
Filing Date
2026-01-29
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Traditional intelligent voice interaction suffers from insufficient speech recognition capabilities in complex environments, severe noise interference, and inadequate semantic understanding, resulting in choppy and inefficient interaction.

Method used

A multi-channel acoustic sensor array is used to collect mixed audio signals. Geometric calibration is performed by setting a quadrilateral structure, noise type is identified and noise reduction strategy is dynamically adjusted. Semantic repair and fusion processing are performed by combining contextual scene knowledge base to generate complete voice commands.

Benefits of technology

Improve speech recognition accuracy and interaction fluency in complex and noisy environments, reduce misunderstandings, provide a personalized and user-friendly intelligent voice interaction experience, and support rich functional modules such as information query, smart home control, and voice navigation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121583247B_ABST
    Figure CN121583247B_ABST
Patent Text Reader

Abstract

The application provides a kind of intelligent voice interaction system and method, it is related to electronic information technical field, including: acquisition calibration module, for by multichannel acoustic sensor array real-time acquisition mixed audio signal in environment, and set four reference positioning points located at corner point on sensor array, form quadrilateral structure;Based on the area characteristics of quadrilateral generation geometric correction value, and synchronous extraction noise frequency band feature and speech short-time energy distribution characteristics, and utilize geometric correction value to extract the feature and calibrate;Noise identification module is used for based on noise frequency band feature and speech short-time energy distribution characteristics, identify noise type, including wide-band high-intensity noise, low-frequency persistent noise, multi-source human voice interference and burst transient noise, generate label vector, and dynamically adjust noise reduction strategy according to label vector.The application is collected and calibrated audio signal by multichannel acoustic sensor array, improves the accuracy and fluency of voice interaction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of electronic information technology, and in particular to an intelligent voice interaction system and method. Background Technology

[0002] Traditional smart voice interaction systems need improvement in their voice recognition capabilities in complex environments. In a home environment, there are often various background noises, such as television playback and the operation of kitchen appliances. These noises can interfere with the accurate recognition of user voice commands. For example, when a user is cooking in the kitchen and setting an alarm clock via smart voice, the noise from the range hood and the boiling kettle on the stove can blend together, making it difficult to accurately capture the specific alarm time spoken by the user, leading to incorrect settings and inconvenience.

[0003] Furthermore, traditional technologies lack a deep understanding of the complex semantics and contextual relationships in natural language. In a home environment, conversations between users and intelligent voice assistants are often casual and conversational, and often involve contextual relationships. For example, if a user asks, "How's the weather today?" and the intelligent voice assistant answers, and then the user replies, "And tomorrow?", the assistant may not accurately understand that "And tomorrow?" is a continuation of the previous question about the weather, but rather misinterprets it as having another meaning, or may be unable to provide a direct and relevant answer, requiring the user to rephrase the question, thus reducing the fluency and efficiency of the interaction. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to provide an intelligent voice interaction system and method, which solves the problems of voice recognition being easily interfered with and inaccurate semantic understanding, and achieves smooth and natural voice interaction.

[0005] To solve the above-mentioned technical problems, the technical solution of the present invention is as follows:

[0006] Firstly, an intelligent voice interaction system includes:

[0007] The acquisition and calibration module is used to acquire mixed audio signals in the environment in real time through a multi-channel acoustic sensor array, and set four reference positioning points located at the corners of the sensor array to form a quadrilateral structure; it generates geometric correction values ​​based on the area characteristics of the quadrilateral, and simultaneously extracts noise frequency band features and speech short-time energy distribution features, and uses the geometric correction values ​​to calibrate the extracted features;

[0008] The noise identification module is used to identify noise types based on noise frequency band characteristics and short-time energy distribution characteristics of speech, including wideband high-intensity noise, low-frequency continuous noise, multi-source human voice interference and sudden transient noise, generate label vectors, and dynamically adjust the noise reduction strategy according to the label vectors;

[0009] The semantic repair module is used to perform semantic repair and fusion processing on speech signals based on a hierarchical noise reduction strategy. It repairs missing syllable segments through a syllable coherence detection mechanism, performs probabilistic reasoning to complete keywords by combining a contextual scene knowledge base, and dynamically adjusts the temporal relationship between semantic segments to generate complete speech commands.

[0010] The instruction execution module is used to adjust the threshold range of speech endpoint detection in real time based on the noise type information in the label vector, and extend the speech buffer time window to obtain speech segments covered by noise, so as to obtain the final recognition instruction and execute the corresponding operation.

[0011] Furthermore, a multi-channel acoustic sensor array is used to acquire mixed audio signals from the environment in real time, and four reference positioning points located at the corners are set on the sensor array to form a quadrilateral structure. Geometric correction values ​​are generated based on the area characteristics of the quadrilateral, and noise frequency band features and short-time speech energy distribution features are extracted simultaneously. The extracted features are then calibrated using the geometric correction values, including:

[0012] Reference positioning points are set at the four corners of the effective sound pickup area of ​​the sensor array, and the spatial distance between adjacent positioning points is tracked in real time to construct a dynamic quadrilateral structure.

[0013] Based on the real-time area change of the dynamic quadrilateral structure, the deviation ratio between the current area and the preset standard area is calculated, and a geometric correction value reflecting the degree of deformation of the sensor array is generated.

[0014] The mixed audio signal is input into a multi-band filter bank. During the dynamic quadrilateral construction process, the energy ratio features of the 0-300Hz low frequency band and the 4-8kHz mid-high frequency band are extracted simultaneously to generate the initial noise frequency band feature vector.

[0015] During the geometric correction value generation process, the mixed audio signal is synchronously segmented into a short-time frame sequence, the temporal energy value of each frame is calculated, and the energy fluctuation variance between consecutive frames is statistically analyzed to form the initial speech short-time energy distribution characteristics.

[0016] The amplitude of the initial noise frequency band feature vector is adjusted using geometric correction values, and the energy fluctuation variance in the initial speech short-time energy distribution feature is distorted and compensated for, so as to obtain the calibrated noise frequency band feature and speech short-time energy distribution feature.

[0017] Furthermore, based on noise frequency band characteristics and short-time energy distribution characteristics of speech, noise types are identified, including broadband high-intensity noise, low-frequency continuous noise, multi-source human voice interference, and sudden transient noise. Label vectors are generated, and noise reduction strategies are dynamically adjusted based on these label vectors, including:

[0018] Based on noise frequency band features and speech short-time energy distribution features, the noise frequency band features and speech energy distribution in each audio frame are time-synchronized and associated to construct a time-aligned feature association table.

[0019] Based on the feature association table, continuous frame sequences with energy > first threshold are detected in the low-frequency and mid-to-high-frequency bands. If the duration of the continuous frame sequence is > 50ms, it is determined that there is broadband high-intensity noise, and a first noise intensity parameter is generated.

[0020] In the remaining frame sequence, the variance of low-frequency energy fluctuation is statistically analyzed. If the proportion of consecutive frames with variance values ​​less than the second threshold reaches 80%, it is determined that there is persistent low-frequency noise, and a second noise stability parameter is generated.

[0021] Based on the feature association table and the first noise intensity parameter and the second noise stability parameter, the energy peak distribution in the 4-8kHz frequency band is analyzed in the low-frequency noise frame. If there are multiple energy peaks in a single frame and the peak spacing is less than the critical bandwidth, and the variance mutation amount is greater than the third threshold, then it is determined that there is multi-source human voice interference, and the third interference source quantity parameter is generated.

[0022] In the unlabeled multi-source human voice interference noise type, calculate the energy difference between adjacent frames, count the frequency of abrupt events where the difference is greater than the fourth threshold within a 100ms window, and if the abrupt event frequency reaches 3 times / 100ms, it is determined that there is sudden transient noise and the fourth transient density parameter is generated.

[0023] The first noise intensity parameter, the second noise stability parameter, the third interference source quantity parameter, and the fourth transient density parameter are integrated to generate the dimension parameter value of the label vector, and the noise reduction strategy is dynamically adjusted according to the dimension parameter value.

[0024] Furthermore, the noise reduction strategy includes frequency domain masking and syllable energy compensation for broadband high-intensity noise; adaptive notch filtering for low-frequency noise; source separation and directional beam focusing for multi-source human voice interference; and a delayed decision mechanism for triggering sudden transient noise.

[0025] Furthermore, based on a hierarchical noise reduction strategy, semantic repair and fusion processing are performed on the speech signal. Missing syllable segments are repaired through a syllable coherence detection mechanism, and keywords are probabilistically inferred and completed using a contextual knowledge base. Finally, dynamic adjustments are made based on the temporal relationships between semantic segments to generate complete speech commands, including:

[0026] Syllable coherence detection is performed on the speech signal processed by the layered noise reduction strategy. This includes segmenting the speech signal into discrete syllable units, detecting the energy continuity and time interval between adjacent syllable units, and inserting synthetic syllable segments at the interval positions when the time interval between adjacent syllable units is greater than a preset coherence threshold to generate a repaired syllable sequence.

[0027] Based on the repaired syllable sequence, keyword extraction and completion are performed in conjunction with a pre-built contextual scene knowledge base. When a missing keyword in the current scene is detected, probabilistic reasoning is performed based on the semantic template and keyword probability distribution to generate a set of semantic fragments with keyword completion.

[0028] The temporal relationship of the semantic fragment set after keyword completion is dynamically adjusted, including parsing the timestamp information of each semantic fragment and rearranging the temporal position of the semantic fragments according to the requirements of instruction logical coherence, so as to generate a complete voice instruction whose temporal relationship conforms to the preset logical rules.

[0029] Furthermore, based on the noise type information in the label vector, the threshold range for speech endpoint detection is adjusted in real time, and the speech buffer time window is extended to acquire speech segments covered by noise, so as to obtain the final recognition instruction and execute corresponding operations, including:

[0030] Based on the noise type information in the label vector, a noise type label is generated, and the speech endpoint detection is adjusted in real time according to the noise type label to obtain the adjusted speech endpoint detection threshold range.

[0031] When the noise type is marked as broadband high-intensity noise and sudden transient noise, the speech buffer time window is extended to the first preset value to obtain the speech segment covered by noise; when it is marked as low-frequency continuous noise and multi-source human voice interference, the speech buffer time window is extended to the second preset value to obtain the integrity of the speech segment.

[0032] Based on the adjusted voice endpoint detection threshold range and the extended buffer time window, complete voice commands are extracted and corresponding voice command operations are executed.

[0033] Furthermore, the threshold range for voice endpoint detection includes reducing the threshold range when the noise type is labeled as broadband high-intensity noise and sudden transient noise, and increasing the threshold range when it is labeled as low-frequency continuous noise and multi-source human voice interference.

[0034] Secondly, an intelligent voice interaction method includes:

[0035] Mixed audio signals in the environment are acquired in real time by a multi-channel acoustic sensor array, and four reference positioning points located at the corners are set on the sensor array to form a quadrilateral structure. Geometric correction values ​​are generated based on the area characteristics of the quadrilateral, and noise frequency band features and speech short-time energy distribution features are extracted simultaneously. The extracted features are then calibrated using the geometric correction values.

[0036] Based on noise frequency band characteristics and speech short-time energy distribution characteristics, noise types are identified, including wideband high-intensity noise, low-frequency continuous noise, multi-source human voice interference, and sudden transient noise. A label vector is generated, and the noise reduction strategy is dynamically adjusted according to the label vector.

[0037] Based on a hierarchical noise reduction strategy, semantic repair and fusion processing are performed on the speech signal. Missing syllable segments are repaired through a syllable coherence detection mechanism, keywords are probabilistically inferred and completed by combining a contextual scene knowledge base, and the temporal relationship between semantic segments is dynamically adjusted to generate complete speech commands.

[0038] Based on the noise type information in the label vector, the threshold range of speech endpoint detection is adjusted in real time, and the speech buffer time window is extended to obtain speech segments covered by noise, so as to obtain the final recognition instruction and execute the corresponding operation.

[0039] Thirdly, a computing device, comprising:

[0040] One or more processors;

[0041] A storage device for storing one or more programs that, when executed by one or more processors, enable the one or more processors to implement the system.

[0042] Fourthly, a computer-readable storage medium storing a program that, when executed by a processor, implements the system.

[0043] The above-described solution of the present invention has at least the following beneficial effects:

[0044] Intelligent voice interaction systems break the limitations of traditional manual input, allowing users to complete operations simply through voice commands, improving the efficiency of information acquisition and task execution. In scenarios where hands are occupied, such as driving or doing housework, users can perform functions like checking the weather, playing music, and setting reminders via voice without manually operating the device, enhancing the convenience and safety of interaction. Employing advanced speech recognition algorithms and natural language processing technology, it can quickly and accurately recognize speech information with different accents and speaking speeds, and deeply understand semantics. Whether it's everyday conversation, professional terminology, or vague expressions, it can accurately interpret user intent, reducing interaction errors and providing users with a smoother and more natural interactive experience. It integrates a wealth of functional modules, covering multiple areas such as information query, smart home control, voice navigation, and voice translation. Users can control smart devices in their homes via voice commands to adjust lighting and temperature; they can also use the voice navigation function when traveling to obtain real-time traffic conditions and final routes. It possesses strong learning capabilities, can proactively adapt to user needs based on user habits and preferences, and provide personalized interactive services. It can intelligently recommend news, music, and film content that match user interests, and can also provide corresponding responses and suggestions based on the user's emotional state, making the interaction more human and warm. Attached Figure Description

[0045] Figure 1 This is a schematic diagram of an intelligent voice interaction system provided by an embodiment of the present invention.

[0046] Figure 2 This is a flowchart illustrating an intelligent voice interaction method provided by an embodiment of the present invention. Detailed Implementation

[0047] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.

[0048] like Figure 1 As shown, an embodiment of the present invention proposes an intelligent voice interaction system, comprising:

[0049] The acquisition and calibration module is used to acquire mixed audio signals in the environment in real time through a multi-channel acoustic sensor array, and set four reference positioning points located at the corners of the sensor array to form a quadrilateral structure; it generates geometric correction values ​​based on the area characteristics of the quadrilateral, and simultaneously extracts noise frequency band features and speech short-time energy distribution features, and uses the geometric correction values ​​to calibrate the extracted features;

[0050] The noise identification module is used to identify noise types based on noise frequency band characteristics and short-time energy distribution characteristics of speech, including wideband high-intensity noise, low-frequency continuous noise, multi-source human voice interference and sudden transient noise, generate label vectors, and dynamically adjust the noise reduction strategy according to the label vectors;

[0051] The semantic repair module is used to perform semantic repair and fusion processing on speech signals based on a hierarchical noise reduction strategy. It repairs missing syllable segments through a syllable coherence detection mechanism, performs probabilistic reasoning to complete keywords by combining a contextual scene knowledge base, and dynamically adjusts the temporal relationship between semantic segments to generate complete speech commands.

[0052] The instruction execution module is used to adjust the threshold range of speech endpoint detection in real time based on the noise type information in the label vector, and extend the speech buffer time window to obtain speech segments covered by noise, so as to obtain the final recognition instruction and execute the corresponding operation.

[0053] In this embodiment of the invention, a multi-channel acoustic sensor array can be used to collect mixed audio signals in the environment in real time and comprehensively. Compared with single-sensor acquisition, it has a wider coverage and richer information acquisition. By setting reference positioning points at the corners of the sensor array to form a quadrilateral structure, and generating geometric correction values ​​based on its area characteristics, audio acquisition errors caused by sensor position, angle deviation, and other factors can be effectively corrected, ensuring the accuracy of the original audio data. At the same time, noise frequency band features and short-time energy distribution features of speech are extracted synchronously and calibrated using geometric correction values, making the feature information on which processing is based more accurate and reliable, and reducing interaction errors caused by acquisition errors. Based on the provided noise frequency band features and short-time energy distribution features of speech, various noise types can be accurately identified, including broadband high-intensity noise, low-frequency continuous noise, multi-source human voice interference, and sudden transient noise. By generating a label vector, the noise type information is quantified and identified, providing a clear basis for dynamically adjusting the noise reduction strategy. For different types of noise, appropriate noise reduction algorithms can be flexibly selected to effectively suppress noise interference, preserve the integrity of the speech signal, improve the clarity and recognizability of speech in complex noise environments, and ensure that speech interaction is not affected by noise.

[0054] The speech signal processed by the hierarchical noise reduction strategy, through a syllable coherence detection mechanism, can accurately capture and repair missing syllable segments caused by noise interference or transmission problems, making speech expression smoother. Combining a contextual knowledge base, keyword probabilistic reasoning is used for completion, fully utilizing semantic environment information. Even if some keywords are masked by noise, reasonable content can be inferred based on contextual logic, avoiding semantic comprehension bias. The temporal relationship between semantic segments is dynamically adjusted to ensure logical coherence and clear meaning of speech commands, ultimately generating complete and accurate speech commands. This process improves the completeness and accuracy of speech semantics, enabling better understanding of the user's true intent, reducing interaction misunderstandings, and enhancing user experience. Based on the noise type information in the generated label vector, the threshold range of speech endpoint detection is adjusted in real time. For noise of different intensities and characteristics, the start and end positions of speech are accurately determined, avoiding misjudgments caused by noise interference. Extending the speech buffer time window allows for the acquisition of speech segments covered by noise, maximizing the collection of effective speech information and preventing the omission of key command content. These operations yield the final accurate recognition command and execute the corresponding operation, ensuring that the execution of user commands closely matches the user's actual needs, thereby improving the practicality and reliability of the intelligent voice interaction system, and enabling users to smoothly complete tasks through voice control devices in various complex and noisy environments.

[0055] In a preferred embodiment of the present invention, a multi-channel acoustic sensor array is used to collect mixed audio signals in the environment in real time, and four reference positioning points located at the corners are set on the sensor array to form a quadrilateral structure; a geometric correction value is generated based on the area characteristics of the quadrilateral, and noise frequency band features and short-time energy distribution features of speech are extracted simultaneously, and the extracted features are calibrated using the geometric correction value, which may include:

[0056] Reference positioning points are set at the four corners of the effective sound pickup area of ​​the sensor array, and the spatial distance between adjacent positioning points is tracked in real time to construct a dynamic quadrilateral structure.

[0057] Based on the real-time area change of the dynamic quadrilateral structure, the deviation ratio between the current area and the preset standard area is calculated, and a geometric correction value reflecting the degree of deformation of the sensor array is generated.

[0058] The mixed audio signal is input into a multi-band filter bank. During the dynamic quadrilateral construction process, the energy ratio features of the 0-300Hz low frequency band and the 4-8kHz mid-high frequency band are extracted simultaneously to generate the initial noise frequency band feature vector.

[0059] During the geometric correction value generation process, the mixed audio signal is synchronously segmented into a short-time frame sequence, the temporal energy value of each frame is calculated, and the energy fluctuation variance between consecutive frames is statistically analyzed to form the initial speech short-time energy distribution characteristics.

[0060] The amplitude of the initial noise frequency band feature vector is adjusted using geometric correction values, and the energy fluctuation variance in the initial speech short-time energy distribution feature is distorted and compensated for, so as to obtain the calibrated noise frequency band feature and speech short-time energy distribution feature.

[0061] In this embodiment of the invention, reference positioning points are set at the four corners of the area where the multi-channel acoustic sensor array can effectively collect sound, i.e., the effective sound pickup area. During operation, the spatial distance changes between every two adjacent positioning points are continuously monitored and recorded. These distances may change over time and due to environmental changes. Based on these real-time distance data, the four reference positioning points are connected sequentially to construct a dynamic quadrilateral structure whose shape and size can change in real time. The shape change of this quadrilateral can intuitively reflect the spatial state change of the sensor array. A preset standard area is determined, which is based on the area of ​​the quadrilateral formed by the four reference positioning points under ideal conditions of the sensor array. After the dynamic quadrilateral structure is formed, the actual area of ​​the current dynamic quadrilateral is calculated in real time. Then, by comparing the current area with the preset standard area, the deviation ratio between the two is calculated. For example, if the preset standard area is 100 and the current area is 80, then the deviation ratio is (100-80)÷100=20%. This deviation ratio is used to generate a geometric correction value, which clearly reflects the degree of spatial deformation of the sensor array, i.e., how much the sensor array has changed compared to its ideal state. The acquired mixed audio signal contains various sound information and is input into a multi-band filter bank, which acts like a "sound classifier," dividing the mixed audio signal according to different frequency ranges. While constructing the dynamic quadrilateral structure, the focus is on the low-frequency band (0-300Hz) and the mid-high frequency band (4-8kHz), calculating the energy proportion of these two frequency bands in the entire audio signal. For example, in the audio signal acquired at a certain moment, the energy of the 0-300Hz low-frequency band accounts for 30% of the total energy, and the energy of the 4-8kHz mid-high frequency band accounts for 40% of the total energy. These energy proportion data are recorded and combined into a vector, which is the initial noise frequency band feature vector, reflecting the distribution characteristics of noise energy in different frequency bands in the current audio signal.

[0062] During the calculation of geometric correction values, the acquired mixed audio signal undergoes fine processing. First, the mixed audio signal is divided into multiple short, equal-length time segments in chronological order; these segments constitute a short-time frame sequence. The duration of each short-time frame is pre-set, typically between a few milliseconds and tens of milliseconds, ensuring sufficient speech information within each frame while meeting real-time processing requirements. For each short-time frame, its energy value in the time domain is calculated. This time-domain energy value is obtained by summing the squares of the signal amplitudes at all sampling points within the short-time frame and then taking the average. Specifically, assuming a short-time frame contains... There are 1 sampling point, and the signal amplitude of each sampling point is 1. ( =1, 2, ..., Then the energy value of that short frame. The calculation method is to first calculate the amplitude of each sampling point. Squaring, we get Then put this Adding the squares of each value together, that is Finally, divide this sum by the number of sampling points. ,get ,this The value represents the strength of the sound signal within that short frame. After calculating the energy value for each short frame, the fluctuation of the energy values ​​between consecutive frames is statistically analyzed, and variance is used to measure the degree of this fluctuation. Taking three consecutive short frames as an example, assuming their energy values ​​are... , , First, calculate the average of these three energy values. Then, calculate the square of the difference between each energy value and the average value, i.e. , , Then add these three squared differences together to get + + Finally, dividing this sum by the frame number 3 yields the variance of the energy fluctuations in these three consecutive short frames. For the entire short frame sequence, the variance of every three consecutive short frames is calculated sequentially, thus comprehensively reflecting the fluctuations in the speech signal energy at different points in time. Combining the energy values ​​of all short frames and the variance data of their energy fluctuations forms the initial short-time energy distribution characteristics of the speech signal.

[0063] After obtaining the geometric correction value, the initial noise frequency band feature vector, and the initial short-time speech energy distribution characteristics, the amplitude of the initial noise frequency band feature vector is adjusted using the geometric correction value. If the geometric correction value indicates that the sensor array has undergone a certain degree of deformation, the amplitude of the energy proportion of each frequency band in the noise frequency band feature vector is adjusted accordingly to better reflect the actual situation. Simultaneously, the energy fluctuation variance in the initial short-time speech energy distribution characteristics is also compensated for distortion using the geometric correction value. For example, if sensor array deformation causes deviations in the energy fluctuation data, these deviations are corrected using the geometric correction value, ultimately yielding the calibrated noise frequency band features and short-time speech energy distribution characteristics.

[0064] By setting reference positioning points at the corners of the effective sound pickup area of ​​the sensor array and constructing a dynamic quadrilateral structure, the spatial distance changes between adjacent positioning points can be tracked in real time, enabling timely perception of the sensor array's spatial deformation. This monitoring method based on actual spatial conditions avoids audio acquisition errors caused by sensor position offsets and angle changes, ensuring that the acquired mixed audio signals accurately reflect the environmental sound conditions. Geometric correction values ​​are calculated based on the area changes of the dynamic quadrilateral structure, quantifying the degree of sensor array deformation. This mechanism is adaptable to different environmental conditions. Whether it's physical deformation caused by external force or temperature changes, or minor changes in installation position, the acquired data can be calibrated using geometric correction values, ensuring accurate acquisition and processing of audio signals even in complex and changing environments. During the construction of the dynamic quadrilateral and the calculation of geometric correction values, noise frequency band features and short-time speech energy distribution features are extracted simultaneously, and these features are calibrated using geometric correction values, achieving precise capture and optimization of noise and speech features. This parallel processing and calibration mechanism ensures that the extracted features accurately reflect the actual noise and speech in the audio signal, avoiding feature deviations caused by sensor deformation.

[0065] In a preferred embodiment of the present invention, based on noise frequency band characteristics and short-time energy distribution characteristics of speech, noise types are identified, including broadband high-intensity noise, low-frequency continuous noise, multi-source human voice interference, and sudden transient noise. A label vector is generated, and the noise reduction strategy is dynamically adjusted based on the label vector, which may include:

[0066] Based on noise frequency band features and speech short-time energy distribution features, the noise frequency band features and speech energy distribution in each audio frame are time-synchronized and associated to construct a time-aligned feature association table.

[0067] Based on the feature association table, continuous frame sequences with energy > first threshold are detected in the low-frequency and mid-to-high-frequency bands. If the duration of the continuous frame sequence is > 50ms, it is determined that there is broadband high-intensity noise, and a first noise intensity parameter is generated.

[0068] In the remaining frame sequence, the variance of low-frequency energy fluctuation is statistically analyzed. If the proportion of consecutive frames with variance values ​​less than the second threshold reaches 80%, it is determined that there is persistent low-frequency noise, and a second noise stability parameter is generated.

[0069] Based on the feature association table and the first noise intensity parameter and the second noise stability parameter, the energy peak distribution in the 4-8kHz frequency band is analyzed in the low-frequency noise frame. If there are multiple energy peaks in a single frame and the peak spacing is less than the critical bandwidth, and the variance mutation amount is greater than the third threshold, then it is determined that there is multi-source human voice interference, and the third interference source quantity parameter is generated.

[0070] In the unlabeled multi-source human voice interference noise type, calculate the energy difference between adjacent frames, count the frequency of abrupt events where the difference is greater than the fourth threshold within a 100ms window, and if the abrupt event frequency reaches 3 times / 100ms, it is determined that there is sudden transient noise and the fourth transient density parameter is generated.

[0071] The first noise intensity parameter, the second noise stability parameter, the third interference source quantity parameter, and the fourth transient density parameter are integrated to generate the dimension parameter value of the marker vector. Based on the dimension parameter value, the noise reduction strategy is dynamically adjusted, specifically including: frequency domain masking and syllable energy compensation for broadband high-intensity noise; adaptive notch filtering for low-frequency noise; source separation and directional beam focusing for multi-source human voice interference; and a delayed decision mechanism for triggering sudden transient noise.

[0072] In this embodiment of the invention, after obtaining the noise frequency band features and the short-term energy distribution features of the speech, for each audio frame, the noise frequency band feature data and the speech energy distribution data are precisely matched and associated in chronological order. For example, the noise frequency band feature information, which shows the energy proportion of the 0-300Hz low-frequency band and the 4-8kHz mid-high-frequency band in an audio frame at a certain moment, is correlated with the short-term energy distribution feature information, which shows the strength and fluctuation of the speech energy in that frame. In this way, the two types of feature data within each audio frame are integrated together, ultimately constructing a complete, time-aligned feature association table. This table clearly presents the specific state and interrelationship of the noise frequency band and speech energy distribution at different time points. Based on the constructed feature association table, the energy status of the low-frequency band (0-300Hz) and the mid-high-frequency band (4-8kHz) is detected, with a preset first threshold used to measure whether the energy has reached a high level. The system checks the energy values ​​of low-frequency and mid-to-high-frequency bands frame by frame. When a continuous audio frame sequence is found, and the energy of both the low-frequency and mid-to-high-frequency bands in these frames exceeds a first threshold, the duration of this continuous frame sequence is recorded. Once the duration of this continuous frame sequence exceeds 50 milliseconds, it is determined that broadband high-intensity noise exists in the current environment. Simultaneously, based on the degree to which the energy in these continuous frames exceeds the threshold, a first noise intensity parameter is generated to represent the noise intensity. This parameter reflects the strength of the broadband high-intensity noise.

[0073] After identifying broadband high-intensity noise, the focus will shift to the remaining frame sequences not identified as broadband high-intensity noise, where the energy fluctuations in the low-frequency band (0-300Hz) will be analyzed in depth. For each audio frame, the low-frequency energy values ​​will be extracted and correlated with the low-frequency energy values ​​of adjacent frames. Specifically, when calculating the variance of low-frequency energy fluctuations, it is assumed that a certain frame is the... The low-frequency energy value of a frame is denoted as... Its previous frame (the first) -1 frame) Low frequency band energy value -1, the next frame (the first) +1 frame) Low-frequency energy value +1. First, calculate the average of the low-frequency energy values ​​for these three frames. Then, calculate the square of the difference between each energy value and the average value, i.e. , , Then add these three squared differences together to get + Finally, dividing this sum by the frame number 3 gives the result of the first frame. Variance of low-frequency energy fluctuations between a frame and its two adjacent frames Following this method, the variance of low-frequency energy fluctuations for each frame in the remaining frame sequence is calculated sequentially, forming a complete variance sequence to comprehensively measure the stability of low-frequency energy fluctuations across consecutive frames. A second threshold is preset as a standard for judging whether low-frequency energy fluctuations are stable. After calculating the variance sequence, the number of consecutive frames in the remaining frame sequence with low-frequency energy fluctuation variance values ​​less than the second threshold is counted. When the number of these qualified consecutive frames accounts for 80% of the total number of frames in the remaining frame sequence, it is determined that persistent low-frequency noise exists in the current environment. After confirming the existence of persistent low-frequency noise, a second noise stability parameter is generated based on the relevant data of these low-frequency energy stable frames.

[0074] Based on the feature association table again, combined with the previously generated first noise intensity parameter and second noise stability parameter, a focused analysis is conducted on the audio data of frames identified as low-frequency noise. Attention is paid to the energy peak distribution in the 4-8kHz frequency band, with a preset critical bandwidth used to determine if the distance between energy peaks is sufficiently close. When multiple energy peaks are detected within an audio frame, and the distance between these peaks is less than the critical bandwidth, the abrupt change in the low-frequency energy fluctuation variance of that frame is also checked. If the abrupt change is greater than a third threshold, multi-source human voice interference is identified. Then, based on the number of detected energy peaks, a third interference source quantity parameter is generated, which roughly represents the number of interference sources in multi-source human voice interference. For audio frames not yet identified as multi-source human voice interference, further analysis is performed to identify the presence of sudden transient noise. First, adjacent audio frames are processed one by one, calculating the energy difference between them. Specifically, for the first... The energy value of a frame audio is denoted as... The adjacent first The audio energy value of +1 frame is denoted as +1, then the energy difference between these two frames It is equal to | +1- By taking the absolute value, the difference is ensured to be non-negative, thus accurately reflecting the magnitude of energy change between adjacent frames.

[0075] Following this method, the energy difference between all adjacent audio frames is calculated sequentially, forming a complete energy difference sequence to capture the instantaneous changes in audio signal energy over time. Next, a 100-millisecond time window is set, and the energy difference sequence is divided and analyzed within this window during audio frame processing. Within each 100-millisecond time window, a fourth threshold is preset to determine whether the energy difference represents a sudden change event. The energy differences within the window are checked one by one, and when a certain energy difference is detected... When the frequency exceeds the fourth threshold, it is considered a mutation event and counted. The number of mutation events occurring within each 100-millisecond time window is continuously recorded to obtain the frequency of mutation events within that time window. When the frequency of mutation events reaches 3 times within a certain 100-millisecond time window, it is determined that there is sudden transient noise in the current environment. After confirming the existence of sudden transient noise, a fourth transient density parameter is generated based on the relevant data of mutation events within that 100-millisecond time window, including the specific occurrence time of the mutation events and the magnitude of the energy difference at each mutation. This parameter, by comprehensively considering the number of mutation events and the length of the time window, can accurately reflect the frequency of sudden transient noise occurring per unit time.

[0076] After identifying broadband high-intensity noise, low-frequency continuous noise, multi-source human voice interference, and sudden transient noise, and generating a first noise intensity parameter, a second noise stability parameter, a third interference source quantity parameter, and a fourth transient density parameter, these parameters are integrated. A predefined label vector with multiple dimensions, each corresponding to a characteristic parameter of a different type of noise, is used. The first noise intensity parameter, the second noise stability parameter, the third interference source quantity parameter, and the fourth transient density parameter are sequentially assigned to the corresponding dimensions of the label vector to determine the parameter values ​​for each dimension, generating a complete label vector. This label vector comprehensively and accurately carries detailed information about the noise type and related characteristics in the current environment. Based on the generated marker vector, the specific type and characteristics of the current noise can be quickly identified. If the marker vector indicates the presence of high-intensity noise in a wide frequency band, a filtering algorithm with suppression capability will be determined from the preset noise reduction algorithm library to process the audio signal and reduce the interference of high-intensity noise. If multi-source human voice interference is detected, a separation algorithm will be invoked. By utilizing its ability to analyze the frequency and energy characteristics of different human voices, the target speech and the interfering human voices will be distinguished and separated, preserving useful speech information to the greatest extent, reducing the impact of noise on voice interaction, and realizing the dynamic and precise adjustment of the noise reduction strategy.

[0077] By synchronously correlating the frequency band characteristics of noise with the short-term energy distribution characteristics of speech, and based on different judgment conditions, it can accurately identify various complex noise types, including broadband high-intensity noise, low-frequency continuous noise, multi-source human voice interference, and sudden transient noise. This meticulous analysis and judgment method avoids misjudgment and omission of noise types, enabling a precise grasp of the noise situation in the environment. During the determination of each noise type, corresponding noise intensity, stability, number of interference sources, and transient density parameters are generated, quantifying various characteristics of the noise. These quantified parameters can more accurately describe the characteristics of the noise, providing richer and more detailed noise information compared to simply identifying the noise type. Based on the generated label vector, the noise reduction strategy can be adjusted dynamically in real time. Since different types and levels of noise require different processing methods, this dynamic adjustment mechanism can employ the noise reduction algorithm most suitable for the current noise environment, improving the noise reduction effect. It can suppress noise interference while preserving the integrity of the speech signal, improving the clarity and accuracy of voice interaction, and providing users with a better user experience. The noise identification and processing mechanism can adapt to various complex and changing noise environments. Whether it is continuous and stable low-frequency noise, sudden and variable high-intensity noise, or complex multi-source human voice interference, it can quickly and accurately identify and respond to adjust the noise reduction strategy.

[0078] In a preferred embodiment of the present invention, based on a hierarchical noise reduction strategy, semantic repair and fusion processing is performed on the speech signal. Missing syllable segments are repaired through a syllable coherence detection mechanism, keywords are probabilistically inferred and completed using a contextual scene knowledge base, and the temporal relationship between semantic segments is dynamically adjusted to generate complete speech instructions. This may include:

[0079] Syllable coherence detection is performed on the speech signal processed by the layered noise reduction strategy. This includes segmenting the speech signal into discrete syllable units, detecting the energy continuity and time interval between adjacent syllable units, and inserting synthetic syllable segments at the interval positions when the time interval between adjacent syllable units is greater than a preset coherence threshold to generate a repaired syllable sequence.

[0080] Based on the repaired syllable sequence, keyword extraction and completion are performed in conjunction with a pre-built contextual scene knowledge base. When a missing keyword in the current scene is detected, probabilistic reasoning is performed based on the semantic template and keyword probability distribution to generate a set of semantic fragments with keyword completion.

[0081] The temporal relationship of the semantic fragment set after keyword completion is dynamically adjusted, including parsing the timestamp information of each semantic fragment and rearranging the temporal position of the semantic fragments according to the requirements of instruction logical coherence, so as to generate a complete voice instruction whose temporal relationship conforms to the preset logical rules.

[0082] In this embodiment of the invention, the speech signal after layered noise reduction is analyzed in detail, firstly segmented into individual syllable units. During the segmentation process, multiple factors are considered, including phoneme features, energy changes, and spectral characteristics of the speech signal, to accurately divide different syllables. After segmentation, the energy continuity and time interval between adjacent syllable units are detected. A coherence threshold is preset to determine whether the time interval between adjacent syllables is reasonable. When the time interval between two adjacent syllable units is detected to be greater than this coherence threshold, it is considered that there may be a missing syllable segment between these two syllables. At this time, a synthesized syllable segment is inserted at this interval based on the pronunciation features of the preceding and following syllables, the prosodic information of the speech, and the pronunciation rules of the language. The pronunciation and prosody of this synthesized syllable segment will match the preceding and following syllables, thereby generating a repaired syllable sequence, making the speech more coherent and natural to the ear.

[0083] Based on the repaired syllable sequence, keyword extraction and completion are performed using a pre-built contextual knowledge base. This knowledge base stores a large number of commonly used keywords, semantic templates, and probability distribution relationships between keywords in different scenarios. First, semantic analysis is performed on the repaired syllable sequence to attempt to extract keywords relevant to the current scenario. When a keyword that should appear in the current scenario is not extracted (i.e., a keyword is missing), probabilistic reasoning is performed based on the semantic templates and keyword probability distributions in the knowledge base. The current semantic environment is analyzed to infer the most likely missing keyword, which is then added to the semantic fragment, generating a set of semantic fragments with keyword completion. Each semantic fragment in this set contains complete keyword information, enabling a more accurate expression of the user's intent.

[0084] The temporal order of the semantic fragment set after keyword completion is dynamically adjusted. First, the timestamp information carried by each semantic fragment is parsed; this timestamp records the time position of the semantic fragment in the original speech signal. Then, according to preset command logical coherence requirements, the temporal positions of these semantic fragments are rearranged. The logical relationships between the semantic fragments are analyzed to determine the correct order in which they should be arranged to form a complete speech command that conforms to logical rules. For example, if one semantic fragment represents "turn on" and another represents "light," they will be arranged in the order of "turn on the light," which conforms to normal language expression habits. Through this temporal rearrangement, a complete speech command with a temporal relationship conforming to preset logical rules is finally generated, ensuring the logicality and comprehensibility of the command.

[0085] Through a syllable continuity detection mechanism, missing syllable segments in the speech signal can be accurately identified and repaired in a timely manner. This process effectively solves the problem of speech discontinuity caused by noise interference and transmission distortion, making the repaired syllable sequence sound smoother and more natural. Regardless of noisy environments or poor speech signal quality, it ensures the integrity of the speech, making the speech content heard by the user more complete and clear. Combining contextual scene knowledge base for probabilistic keyword inference and completion improves the understanding of speech semantics. When keywords are missing in the speech, the system can accurately infer and complete the missing keywords based on the current semantic environment and information in the knowledge base. This reduces semantic misunderstandings caused by missing keywords, enabling a more accurate grasp of the user's true intent and improving the accuracy and reliability of voice interaction. Dynamic adjustment of the temporal relationship of semantic segments ensures that the generated speech commands conform to normal language logic and expression habits. By parsing timestamp information and rearranging the order of semantic segments, it avoids unclear command logic caused by temporal disorder.

[0086] In a preferred embodiment of the present invention, the threshold range for speech endpoint detection is adjusted in real time based on the noise type information in the marker vector, and the speech buffer time window is extended to acquire speech segments covered by noise, so as to obtain the final recognition instruction and execute the corresponding operation, which may include:

[0087] Based on the noise type information in the label vector, a noise type label is generated, and the speech endpoint detection is adjusted in real time according to the noise type label to obtain the adjusted speech endpoint detection threshold range. Specifically, the threshold range of the speech endpoint detection includes: when the noise type label is wideband high-intensity noise and sudden transient noise, the threshold range of the speech endpoint detection is reduced; when the label is low-frequency continuous noise and multi-source human voice interference, the threshold range of the speech endpoint detection is increased.

[0088] When the noise type is marked as broadband high-intensity noise and sudden transient noise, the speech buffer time window is extended to the first preset value to obtain the speech segment covered by noise; when it is marked as low-frequency continuous noise and multi-source human voice interference, the speech buffer time window is extended to the second preset value to obtain the integrity of the speech segment.

[0089] Based on the adjusted voice endpoint detection threshold range and the extended buffer time window, complete voice commands are extracted and corresponding voice command operations are executed.

[0090] In this embodiment of the invention, noise type information is first extracted from the label vector and converted into corresponding noise type labels, such as "wideband high-intensity noise," "low-frequency continuous noise," "multi-source human voice interference," and "sudden transient noise." Then, based on different noise type labels, preset endpoint detection threshold adjustment rules are invoked. These rules are pre-defined based on the influence characteristics of different noises on the endpoint features of speech signals. For example, when the noise type is labeled as wideband high-intensity noise, since this type of noise may obscure the start and end boundaries of the speech signal, the starting threshold for speech endpoint detection is automatically lowered, while the ending threshold is raised, thereby expanding the threshold range and preventing speech endpoints from being misjudged as non-speech regions due to noise interference. Through this real-time adjustment, a speech endpoint detection threshold range adapted to the current noise environment is obtained. Different speech buffer time window extension strategies are executed according to different noise type labels. When the noise type is labeled as wideband high-intensity noise or sudden transient noise, this type of noise is characterized by high intensity and suddenness, easily covering the beginning and end of speech segments. At this point, the voice buffer time window is extended to a first preset value (e.g., 500 milliseconds). This means that after detecting a voice endpoint, an additional 500 milliseconds of audio data are retained to ensure that the beginning or end segments of voice briefly covered by noise can be captured completely. When the noise type is labeled as low-frequency persistent noise or multi-source human voice interference, although this type of noise is persistent or contains multiple voices, its impact on the integrity of the voice segment is relatively uniform. Therefore, the voice buffer time window is extended to a second preset value (e.g., 300 milliseconds) to ensure the acquisition of complete voice segments while avoiding the introduction of irrelevant noise due to excessive buffer extension.

[0091] After adjusting the speech endpoint detection threshold range and extending the buffer time window, the audio signal is processed based on the adjusted parameters. First, the start and end endpoints of the speech are re-detected using the adjusted threshold range to accurately locate the boundaries of speech segments. Then, based on the extended buffer time window, segments containing complete speech commands are extracted from the audio stream, ensuring that even if there are noise-covered parts before and after the speech segment, the complete content can still be obtained through the extended buffer. Finally, the extracted speech segments are input into the speech recognition module for decoding, generating the final recognition command, and driving relevant devices to perform corresponding operations, such as controlling smart home devices or returning information query results.

[0092] By adjusting the speech endpoint detection threshold range in real time based on noise type markings, the system can adapt to the impact of different noise environments on speech signal boundaries. For high-intensity noise or complex interference scenarios, the dynamically adjusted threshold range effectively avoids missed or false detections of speech endpoints, ensuring accurate capture of the start and end positions of speech. Differentiated buffer time window extension strategies for different noise types can accurately address the issue of various noises covering speech segments. For sudden and high-intensity noise, a longer buffer time window can effectively salvage speech content momentarily covered by noise; for multi-source interference noise, a moderate buffer time window ensures speech integrity while reducing the introduction of irrelevant noise, improving the purity of the speech signal. Through the synergistic effect of adjusting the endpoint detection threshold and extending the buffer time window, speech commands containing complete semantics can be extracted, avoiding missing command segments or boundary errors due to noise interference. This enables the speech recognition module to accurately decode user intent, thereby driving the device to perform correct operations and improving the reliability and user experience of the intelligent voice interaction system in complex noise environments. This mechanism can dynamically adjust the processing strategy according to real-time noise type, maintaining stable performance in different noise scenarios without manual intervention. Whether in noisy public places, low-frequency industrial environments, or social scenarios with multiple people talking, the system can achieve accurate processing of voice signals through adaptive adjustments, thus expanding the application scenarios of intelligent voice interaction systems.

[0093] like Figure 2 As shown, embodiments of the present invention also provide an intelligent voice interaction method, the method comprising:

[0094] Mixed audio signals in the environment are acquired in real time by a multi-channel acoustic sensor array, and four reference positioning points located at the corners are set on the sensor array to form a quadrilateral structure. Geometric correction values ​​are generated based on the area characteristics of the quadrilateral, and noise frequency band features and speech short-time energy distribution features are extracted simultaneously. The extracted features are then calibrated using the geometric correction values.

[0095] Based on noise frequency band characteristics and speech short-time energy distribution characteristics, noise types are identified, including wideband high-intensity noise, low-frequency continuous noise, multi-source human voice interference, and sudden transient noise. A label vector is generated, and the noise reduction strategy is dynamically adjusted according to the label vector.

[0096] Based on a hierarchical noise reduction strategy, semantic repair and fusion processing are performed on the speech signal. Missing syllable segments are repaired through a syllable coherence detection mechanism, keywords are probabilistically inferred and completed by combining a contextual scene knowledge base, and the temporal relationship between semantic segments is dynamically adjusted to generate complete speech commands.

[0097] Based on the noise type information in the label vector, the threshold range of speech endpoint detection is adjusted in real time, and the speech buffer time window is extended to obtain speech segments covered by noise, so as to obtain the final recognition instruction and execute the corresponding operation.

[0098] It should be noted that this method is the same as the method described above for the system. All implementation methods in the above system embodiments are applicable to this embodiment and can achieve the same technical effect.

[0099] Embodiments of the present invention also provide a computing device, including: a processor and a memory storing a computer program, wherein the computer program, when executed by the processor, performs the method described above. All implementations in the above method embodiments are applicable to this embodiment and can achieve the same technical effects.

[0100] Embodiments of the present invention also provide a computer-readable storage medium storing instructions that, when executed on a computer, cause the computer to perform the system as described above. All implementations in the above system embodiments are applicable to this embodiment and can achieve the same technical effects.

[0101] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. An intelligent voice interaction system, characterized in that, include: The acquisition and calibration module is used to set reference positioning points at the four corners of the effective sound pickup area of ​​the sensor array and track the changes in spatial distance between adjacent positioning points in real time to construct a dynamic quadrilateral structure. Based on the real-time area change of the dynamic quadrilateral structure, the deviation ratio between the current area and the preset standard area is calculated to generate a geometric correction value reflecting the degree of sensor array deformation. The mixed audio signal is input into a multi-band filter bank, and the energy proportion features of the 0-300Hz low-frequency band and the 4-8kHz mid-high frequency band are extracted simultaneously during the construction of the dynamic quadrilateral to generate an initial noise frequency band feature vector. During the generation of the geometric correction value, the mixed audio signal is simultaneously segmented into a short-time frame sequence, the time-domain energy value of each frame is calculated, and the energy fluctuation variance between consecutive frames is statistically analyzed to form the initial speech short-time energy distribution features. The amplitude of the initial noise frequency band feature vector is adjusted using the geometric correction value, and the distortion compensation is performed on the energy fluctuation variance in the initial speech short-time energy distribution features to finally obtain the calibrated noise frequency band features and speech short-time energy distribution features. The noise identification module is used to identify noise types based on noise frequency band characteristics and short-time energy distribution characteristics of speech, including wideband high-intensity noise, low-frequency continuous noise, multi-source human voice interference and sudden transient noise, generate label vectors, and dynamically adjust the noise reduction strategy according to the label vectors; The semantic repair module is used to perform semantic repair and fusion processing on speech signals based on a hierarchical noise reduction strategy. It repairs missing syllable segments through a syllable coherence detection mechanism, performs probabilistic reasoning to complete keywords by combining a contextual scene knowledge base, and dynamically adjusts the temporal relationship between semantic segments to generate complete speech commands. The instruction execution module is used to adjust the threshold range of speech endpoint detection in real time based on the noise type information in the label vector, and extend the speech buffer time window to obtain speech segments covered by noise, so as to obtain the final recognition instruction and execute the corresponding operation.

2. The intelligent voice interaction system according to claim 1, characterized in that, Based on noise frequency band characteristics and short-time energy distribution characteristics of speech, noise types are identified, including broadband high-intensity noise, low-frequency continuous noise, multi-source human voice interference, and sudden transient noise. Label vectors are generated, and noise reduction strategies are dynamically adjusted based on these vectors, including: Based on noise frequency band features and speech short-time energy distribution features, the noise frequency band features and speech energy distribution in each audio frame are time-synchronized and associated to construct a time-aligned feature association table. Based on the feature association table, continuous frame sequences with energy > first threshold are detected in the low-frequency and mid-to-high-frequency bands. If the duration of the continuous frame sequence is > 50ms, it is determined that there is broadband high-intensity noise, and a first noise intensity parameter is generated. In the remaining frame sequence, the variance of low-frequency energy fluctuation is statistically analyzed. If the proportion of consecutive frames with variance values ​​less than the second threshold reaches 80%, it is determined that there is persistent low-frequency noise, and a second noise stability parameter is generated. Based on the feature association table and the first noise intensity parameter and the second noise stability parameter, the energy peak distribution in the 4-8kHz frequency band is analyzed in the low-frequency noise frame. If there are multiple energy peaks in a single frame and the peak spacing is less than the critical bandwidth, and the variance mutation amount is greater than the third threshold, then it is determined that there is multi-source human voice interference, and the third interference source quantity parameter is generated. In the unlabeled multi-source human voice interference noise type, calculate the energy difference between adjacent frames, count the frequency of abrupt events where the difference is greater than the fourth threshold within a 100ms window, and if the abrupt event frequency reaches 3 times / 100ms, it is determined that there is sudden transient noise and the fourth transient density parameter is generated. The first noise intensity parameter, the second noise stability parameter, the third interference source quantity parameter, and the fourth transient density parameter are integrated to generate the dimension parameter value of the label vector, and the noise reduction strategy is dynamically adjusted according to the dimension parameter value.

3. The intelligent voice interaction system according to claim 2, characterized in that, The noise reduction strategy includes frequency domain masking and syllable energy compensation for broadband high-intensity noise; adaptive notch filtering for low-frequency noise; source separation and directional beam focusing for multi-source human voice interference; and a delayed decision mechanism for triggering sudden transient noise.

4. The intelligent voice interaction system according to claim 3, characterized in that, Based on a hierarchical noise reduction strategy, semantic repair and fusion processing are performed on the speech signal. Missing syllable segments are repaired through a syllable coherence detection mechanism, and keywords are probabilistically inferred and completed using a contextual knowledge base. Furthermore, the temporal relationships between semantic segments are dynamically adjusted to generate complete speech commands, including: Syllable coherence detection is performed on the speech signal processed by the layered noise reduction strategy. This includes segmenting the speech signal into discrete syllable units, detecting the energy continuity and time interval between adjacent syllable units, and inserting synthetic syllable segments at the interval positions when the time interval between adjacent syllable units is greater than a preset coherence threshold to generate a repaired syllable sequence. Based on the repaired syllable sequence, keyword extraction and completion are performed in conjunction with a pre-built contextual scene knowledge base. When a missing keyword in the current scene is detected, probabilistic reasoning is performed based on the semantic template and keyword probability distribution to generate a set of semantic fragments with keyword completion. The temporal relationship of the semantic fragment set after keyword completion is dynamically adjusted, including parsing the timestamp information of each semantic fragment and rearranging the temporal position of the semantic fragments according to the requirements of instruction logical coherence, so as to generate a complete voice instruction whose temporal relationship conforms to the preset logical rules.

5. The intelligent voice interaction system according to claim 4, characterized in that, Based on the noise type information in the label vector, the threshold range for speech endpoint detection is adjusted in real time, and the speech buffer time window is extended to acquire speech segments covered by noise, so as to obtain the final recognition instruction and execute corresponding operations, including: Based on the noise type information in the label vector, a noise type label is generated, and the speech endpoint detection is adjusted in real time according to the noise type label to obtain the adjusted speech endpoint detection threshold range. When the noise type is marked as broadband high-intensity noise and sudden transient noise, the speech buffer time window is extended to the first preset value to obtain the speech segment covered by noise; when it is marked as low-frequency continuous noise and multi-source human voice interference, the speech buffer time window is extended to the second preset value to obtain the integrity of the speech segment. Based on the adjusted voice endpoint detection threshold range and the extended buffer time window, complete voice commands are extracted and corresponding voice command operations are executed.

6. The intelligent voice interaction system according to claim 5, characterized in that, The threshold range for voice endpoint detection includes reducing the threshold range when the noise type is labeled as broadband high-intensity noise and sudden transient noise, and increasing the threshold range when it is labeled as low-frequency continuous noise and multi-source human voice interference.

7. An intelligent voice interaction method, wherein the method implements the system as described in any one of claims 1 to 6, characterized in that, include: Reference positioning points are set at the four corners of the effective sound pickup area of ​​the sensor array, and the spatial distance between adjacent positioning points is tracked in real time to construct a dynamic quadrilateral structure. Based on the real-time area change of the dynamic quadrilateral structure, the deviation ratio between the current area and the preset standard area is calculated to generate a geometric correction value reflecting the degree of sensor array deformation. The mixed audio signal is input into a multi-band filter bank, and the energy proportion features of the 0-300Hz low-frequency band and the 4-8kHz mid-high frequency band are extracted simultaneously during the construction of the dynamic quadrilateral to generate an initial noise frequency band feature vector. During the generation of the geometric correction value, the mixed audio signal is simultaneously segmented into a short-time frame sequence, the time-domain energy value of each frame is calculated, and the energy fluctuation variance between consecutive frames is statistically analyzed to form the initial speech short-time energy distribution features. The amplitude of the initial noise frequency band feature vector is adjusted using the geometric correction value, and the distortion compensation is performed on the energy fluctuation variance in the initial speech short-time energy distribution features to finally obtain the calibrated noise frequency band features and speech short-time energy distribution features. Based on noise frequency band characteristics and speech short-time energy distribution characteristics, noise types are identified, including wideband high-intensity noise, low-frequency continuous noise, multi-source human voice interference, and sudden transient noise. A label vector is generated, and the noise reduction strategy is dynamically adjusted according to the label vector. Based on a hierarchical noise reduction strategy, semantic repair and fusion processing are performed on the speech signal. Missing syllable segments are repaired through a syllable coherence detection mechanism, keywords are probabilistically inferred and completed by combining a contextual scene knowledge base, and the temporal relationship between semantic segments is dynamically adjusted to generate complete speech commands. Based on the noise type information in the label vector, the threshold range of speech endpoint detection is adjusted in real time, and the speech buffer time window is extended to obtain speech segments covered by noise, so as to obtain the final recognition instruction and execute the corresponding operation.

8. A computing device, characterized in that, include: One or more processors; A storage device for storing one or more programs, which, when executed by one or more processors, cause the one or more processors to implement the system as described in any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program that, when executed by a processor, implements the system as described in any one of claims 1 to 6.