Intelligent voice recognition and analysis system for offline recording

Through preset voiceprint classification and feature hierarchical analysis, combined with semantic matching and device scheduling, the problems of noise interference and device differences in offline recording environments are solved, high-precision speech recognition and intelligent response are achieved, and user interaction experience is improved.

CN120472892AActive Publication Date: 2025-08-12BEIJING SHUZHUO INFORMATION TECHNOLOGY CO LTD +1

Patent Information

Application Number
CN202510618980.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-14
Publication Date
2025-08-12
Estimated Expiration
2045-05-14

AI Technical Summary

Technical Problem

The existing voice recognition and analysis system faces problems such as severe noise interference, inconsistent data format and quality, and inadequate inadequate exploration of the information behind the voice and inflexible intention decisions, resulting in poor recognition accuracy and interactive experience.

Method used

The original audio data stream is divided into basic voiceprint segments and extended voice data segments using preset voiceprint classification rules. The voiceprint pattern and emotional fluctuation characteristics are extracted through the feature hierarchical analysis module, combined with the semantic matching module for joint encoding, the central processing unit performs logic verification, generates final control instructions, and optimizes device calls through the voice resource scheduling module.

Benefits of technology

It improves the accuracy of speech recognition in a noisy environment, deeply explores the information behind the voice, adapts to different recording devices, realizes intelligent intention decision-making and response, and improves user interaction experience and system universality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120472892A_ABST
    Figure CN120472892A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of intelligent voice recognition and analysis, and discloses an intelligent voice recognition and analysis system for offline recording. The system comprises a recording data acquisition module, a feature hierarchical analysis module, a semantic matching module, an intention decision module, a central processing unit and the like. The recording data acquisition module divides an original audio data stream according to a preset voiceprint classification rule; the feature hierarchical analysis module extracts different semantic feature sets; the semantic matching module maps the feature set into an intention classification label; the intention decision module calls a target execution scheme; and the central processing unit checks the parameters to generate a final control instruction. In addition, the system is further provided with a voice resource scheduling module, an instruction conversion module, a voice archiving module and the like which are used for equipment scheduling, instruction conversion and data storage. The system can effectively process offline complex recording data, accurately recognize voice intentions, improve the intelligent level of voice interaction, and is suitable for various offline voice application scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent speech recognition and analysis, and in particular to an intelligent speech recognition and analysis system for offline recording. Background Art

[0002] With the rapid development of artificial intelligence (AI), speech recognition and analysis technologies have been widely applied in numerous fields. From simple voice command recognition in the early days to today's complex voice interaction and speech content analysis, this technology has continuously innovated, bringing significant convenience to people's lives and work. However, in the current field of speech recognition and analysis, the processing of offline recordings still faces many challenges.

[0003] In offline environments, the complexity of recorded data far exceeds that of online recordings. First, offline scenarios are subject to diverse noise interference. In public places such as shopping malls and train stations, the clamor of crowds and the roar of traffic can seriously affect the accuracy of speech recognition. Even in relatively quiet indoor environments, there may be the sound of operating electrical equipment and ambient noise from outdoors. This noise not only degrades the quality of the speech signal but can also cause the speech recognition system to mistakenly identify noise as speech content, resulting in erroneous recognition results. For example, when recording speech in a restaurant, the sounds of clinking cutlery and conversations intertwine, making it difficult to clearly distinguish the target speech.

[0004] Secondly, existing speech recognition systems have a relatively simple approach to classifying and analyzing voice data when processing offline recordings. Most systems simply identify the speech content and fail to deeply explore the rich information behind the speech, such as the speaker's identity, the recording context, and their emotional state. This significantly limits the application scenarios of speech recognition and makes it difficult to meet the growing and diverse needs of users. For example, traditional systems for meeting minutes can only record the content of speeches, but cannot automatically identify different speakers or perceive changes in speakers' emotions. This significantly limits their ability to grasp the meeting's key points and discussion atmosphere.

[0005] Furthermore, offline recording involves different types of recording equipment, each with significantly varying performance and parameters, leading to varying formats and quality of the collected raw audio data. Different devices also vary in sampling rate, bit rate, and number of channels, which creates significant inconvenience for subsequent speech processing. Furthermore, existing speech recognition and analysis systems often lack effective adaptation mechanisms to these device differences, resulting in significant fluctuations in recognition accuracy and analysis results when processing data collected from different devices.

[0006] At the same time, current speech recognition and analysis systems also have shortcomings in intent determination and response execution. They typically make decisions based on simple, preset rules and are unable to make flexible and intelligent decisions based on the specific semantics of speech and complex scenarios. As a result, in practical applications, the system's response may not accurately meet user needs, resulting in a poor interactive experience. For example, in smart home control scenarios, when users issue control commands using different tones and expressions, the system may not accurately understand the user's intentions, resulting in control errors or unresponsiveness. Summary of the Invention

[0007] The purpose of the present invention is to provide an intelligent speech recognition and analysis system for offline recording to solve the problems raised in the above background technology.

[0008] To achieve the above objectives, the present invention provides the following technical solution: an intelligent speech recognition and analysis system for offline recording, the system comprising:

[0009] The recording data acquisition module is used to obtain the original audio data stream of the offline scene and divide the original audio data stream into basic voice data segments and extended voice data segments based on preset voiceprint classification rules;

[0010] a feature hierarchical analysis module, configured to extract voiceprint features from the basic voice data segment to generate a first semantic feature set, and extract emotional fluctuation features from the extended voice data segment to generate a second semantic feature set;

[0011] A semantic matching module, configured to jointly encode the first semantic feature set and the second semantic feature set according to a preset semantic mapping table and map them to corresponding intent classification labels, and use the intent classification labels as parsing parameters for the current speech;

[0012] An intention decision module is used to call a target execution scheme in a preset strategy library based on the intention classification label, and use the target execution scheme as a response parameter of the current voice;

[0013] a central processing unit, configured to send the original audio data stream to the feature hierarchical parsing module, send the first semantic feature set and the second semantic feature set to the semantic matching module, and perform logical verification on the parsing parameters and the response parameters to generate a final control instruction;

[0014] Preferably, the feature hierarchical analysis module extracts emotional fluctuation features from the extended voice data segment, including:

[0015] Dividing the continuous speech frames in the extended speech data segment into valid speech segments and invalid speech segments, and performing short-time energy spectrum feature extraction on the valid speech segments based on a preset time-frequency transformation window to generate a time-frequency feature matrix;

[0016] Performing fundamental frequency tracking processing on the background noise segment in the extended voice data segment, extracting waveform distortion features of each fundamental frequency period and constructing a fundamental frequency feature vector;

[0017] Perform multi-dimensional feature fusion on the time-frequency feature matrix and the fundamental frequency feature vector to generate the second semantic feature set.

[0018] Preferably, the preset voiceprint classification rules include a basic acoustic type set and an extended acoustic type set; the basic acoustic type set includes a voice content identifier, an environmental noise identifier, and a device fingerprint identifier; the extended acoustic type set includes a user identity identifier and a recording scene identifier, and each acoustic identifier corresponds to an independent data parsing protocol.

[0019] Preferably, the system further comprises a recording interface module, which is used to realize the physical connection between the recording data acquisition module, the feature hierarchical analysis module, the semantic matching module and the intention decision module and the local recording device respectively;

[0020] The recording data acquisition module divides the original audio data stream based on the preset voiceprint classification rule, including:

[0021] Acquiring raw waveform data from a local recording device in real time through the recording interface module, and matching head features of the raw waveform data according to identifiers in the basic acoustic type set to separate basic audio segments;

[0022] Scanning tail features of the original waveform data according to identifiers in the extended acoustic type set to extract an extended audio segment;

[0023] The basic audio segment and the extended audio segment are synchronized according to the sampling point sequence number and then written into the basic voiceprint buffer area and the extended voiceprint buffer area respectively.

[0024] Preferably, when the preset semantic mapping table adopts a linear matching model, the intention classification label is a linear mapping result of the normalized fusion value of the first semantic feature set and the second semantic feature set;

[0025] When the preset semantic mapping table adopts a nonlinear matching model, the intention classification label is a discrete semantic label that discriminates a nonlinear combination result of the first semantic feature set and the second semantic feature set through a classification function.

[0026] Preferably, it further comprises a voice resource scheduling module connected to the central processing unit, wherein the voice resource scheduling module is connected to the local voice device database via the recording interface module;

[0027] The voice resource scheduling module is used to match the available device list from the local voice device database according to the device call request in the final control instruction, and generate a device deployment topology diagram to optimize the voiceprint matching path.

[0028] Preferably, the voice resource scheduling module generates a device deployment topology diagram including:

[0029] Obtaining a plane model of the recording space, and marking the real-time voiceprint feature points of each device in the available device list in the plane model;

[0030] Calculating the optimal matching path from the feature point of each device to the target voice area based on a path search algorithm, and prioritizing the list of available devices according to the path matching degree;

[0031] The optimal matching path and the priority ranking are superimposed on the plane model to generate a device deployment topology diagram updated in real time.

[0032] Preferably, when the central processing unit performs logical verification on the parsing parameters and the response parameters, a dual verification mechanism of integrity verification rules and logical consistency detection rules is adopted, wherein the integrity verification rules are used to confirm the validity of the parameters, and the logical consistency detection rules are used to eliminate semantic conflicts between parameters.

[0033] Preferably, it also includes an instruction conversion module connected to the central processing unit, which is used to convert the final control instruction into a voice drive signal, and send the voice drive signal to the target voice device through the recording interface module to trigger the execution action.

[0034] Preferably, it also includes a voice archiving module connected to the central processing unit, and the voice archiving module is used to store the original audio data stream, the first semantic feature set, the second semantic feature set, the intention classification label and the final control instruction, and generate a complete voice processing record chain according to the recording time sequence.

[0035] Compared with the prior art, the present invention has the following beneficial effects:

[0036] In terms of processing complex offline environments, the system divides the original audio data stream into basic voice data segments and extended voice data segments through preset voiceprint classification rules. This classification method can process different types of data in a targeted manner, greatly improving the processing capabilities of complex offline recordings. In offline scenarios with a lot of background noise, the basic voice data segment can focus on the voice content itself, while the extended voice data segment can perform fundamental frequency tracking processing on the background noise segment, extract waveform distortion features, and combine the time-frequency feature matrix extracted by time-frequency transformation to achieve multi-dimensional feature fusion to generate a second semantic feature set. In this way, even in the case of severe noise interference, the system can capture voice features more accurately, improve the accuracy of voice recognition, and avoid recognition errors caused by noise.

[0037] When it comes to mining deep information from voice data, the system's hierarchical feature parsing module not only extracts voiceprint features from basic voice data segments to generate a first semantic feature set, but also extracts emotional fluctuation features from extended voice data segments to generate a second semantic feature set. This enables the system to deeply mine the rich information behind the speech, such as the speaker's identity, the recording scene, and their emotional state. In a meeting scenario, voiceprint features can accurately identify different speakers, while emotional fluctuation features can reflect the speaker's emotional changes, helping to fully grasp the meeting's focus and discussion atmosphere, providing richer and more accurate data support for subsequent meeting summaries and decision-making.

[0038] To address the differences between different recording devices, the system's recording interface module achieves a physical connection to the local recording device. During data acquisition, it can match head features and scan tail features based on the raw waveform data collected by different devices, according to the identifiers in the basic acoustic type set and the extended acoustic type set, separate and extract the basic audio segment and the extended audio segment, and then synchronously write them to the corresponding buffer area. This approach effectively adapts to the diverse data collected by different recording devices, ensuring that regardless of the device used for recording, the system can process data stably and efficiently, improving the system's versatility and compatibility.

[0039] In terms of the intelligence of intent decision-making and response execution, the semantic matching module jointly encodes the two semantic feature sets according to the preset semantic mapping table and maps them to the intent classification label. Based on this, the intent decision module calls the target execution plan in the preset strategy library. This process realizes the intelligent transformation from speech semantics to actual execution plan. In addition, the central processing unit adopts a dual verification mechanism of integrity verification rules and logical consistency detection rules to perform logical verification on parsing parameters and response parameters to generate accurate final control instructions. In the smart home control scenario, when the user issues control instructions in different tones and expressions, the system can accurately understand the user's intentions and quickly make correct responses, thereby improving the user's interactive experience.

[0040] Furthermore, the system's voice resource scheduling module matches available devices from the local voice device database based on the device call request in the final control command, and generates a device deployment topology to optimize the voiceprint matching path. This significantly improves the system's efficiency and accuracy when calling devices, ensuring the effectiveness of voice interaction. Furthermore, the voice archiving module stores the raw audio data stream, various semantic feature sets, intent classification labels, and the final control command in chronological order, generating a complete voice processing record chain. This facilitates subsequent data query, analysis, and backtracking, providing strong support for system optimization and improvement. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 This is a working principle diagram of the intelligent speech recognition and analysis system for offline recording according to the present invention;

[0042] Figure 2 This is a flowchart of the feature hierarchical parsing module extracting emotional fluctuation features of the extended speech data segment;

[0043] Figure 3 The preset voiceprint classification rules and data analysis protocol association flow chart;

[0044] Figure 4 A flowchart for dividing the audio data stream according to preset rules for the recording data acquisition module;

[0045] Figure 5 Flowchart for generating intent classification labels based on different semantic mapping models. DETAILED DESCRIPTION

[0046] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0047] See also Figure 1-Figure 5 The present invention provides an intelligent speech recognition and analysis system for offline recording, and its specific implementation method is described in detail below.

[0048] The system as a whole includes a recording data acquisition module, a feature hierarchical analysis module, a semantic matching module, an intention decision module and a central processing unit.

[0049] The function of the recording data acquisition module is to obtain the original audio data stream of the offline scene and divide it into basic voice data segments and extended voice data segments based on the preset voiceprint classification rules. This module collects the original audio data stream from various offline scenes, such as conference rooms, classrooms, homes, etc., with the help of professional recording equipment or terminal devices with integrated recording functions. The collected audio data will be divided according to the preset voiceprint classification rules. For example, by analyzing the acoustic characteristics of the audio, based on the identifiers in the basic acoustic type set and the extended acoustic type set, the part containing basic information such as voice content, environmental noise, and device fingerprints is divided into the basic voice data segment, while the part containing extended information such as user identity and recording scene is divided into the extended voice data segment.

[0050] The feature layered parsing module receives the basic and extended speech data segments from the audio recording data acquisition module. For the basic speech data segments, this module extracts voiceprint features to generate a first semantic feature set. This process analyzes the acoustic characteristics of the speech, such as frequency and amplitude, to identify unique voiceprint features related to the speech content. For the extended speech data segments, the module extracts emotional fluctuation features to generate a second semantic feature set. For example, by analyzing speech speed and intonation, the module determines the speaker's emotional state and extracts these emotion-related features to form the second semantic feature set.

[0051] The semantic matching module jointly encodes the first and second semantic feature sets according to a preset semantic mapping table. The resulting joint encoding is mapped to corresponding intent classification labels, which serve as parsing parameters for the current speech. The preset semantic mapping table can be a pre-trained model, trained on a large amount of speech data, to determine the intent classification labels corresponding to different semantic feature combinations.

[0052] Based on the intent classification label, the intent decision module calls the target execution plan from the preset strategy library. The preset strategy library stores multiple execution plans for different intent classification labels. The module selects the most appropriate target execution plan based on the current intent classification label and uses it as the response parameter for the current speech.

[0053] The central processing unit (CPU) plays a central control role in the entire system. It sends the raw audio data stream to the feature layer parsing module and simultaneously sends the first and second semantic feature sets to the semantic matching module. Furthermore, the CPU performs logical verification on parsed and response parameters, confirming parameter validity using integrity check rules and eliminating semantic conflicts between parameters using logical consistency check rules, ultimately generating final control instructions.

[0054] The specific implementation of the present invention is further described in detail below through five examples.

[0055] Example 1:

[0056] This embodiment relates to a specific method in which a feature hierarchical analysis module in the system extracts emotional fluctuation features from an extended speech data segment.

[0057] When processing extended speech data segments, the system divides the continuous speech frames within the segment into valid and invalid segments. This division is based on characteristics such as the energy and frequency of the speech signal. For example, speech frames with high energy and frequencies within the normal speech range are considered valid segments, while speech frames with low energy or abnormal frequencies are considered invalid segments.

[0058] Short-time energy spectrum features are extracted from valid speech segments based on a preset time-frequency transform window. This window is set based on the actual application scenario and requirements. Within this window, valid speech segments are analyzed segment by segment, and the short-time energy spectrum of each segment is calculated to generate a time-frequency feature matrix. This matrix contains information about the energy distribution of the valid speech segments at different times and frequencies.

[0059] For the background noise segments within the extended speech data, fundamental frequency tracking is performed. This process extracts waveform distortion features for each fundamental frequency period. Because background noise interferes with the speech signal, causing distortion in the speech waveform, these distortion features are extracted and constructed into a fundamental frequency feature vector.

[0060] Perform multi-dimensional feature fusion on the time-frequency feature matrix and the fundamental frequency feature vector. This fusion can be performed using methods such as concatenation and weighted summation, organically combining the two different types of features to generate a second semantic feature set. This second semantic feature set integrates the time-frequency characteristics of speech and the impact of background noise on the speech waveform, more comprehensively reflecting the emotional fluctuations of the extended speech data segment.

[0061] In an actual offline meeting scenario, professional recording equipment is used at the meeting site to record the entire meeting process. The recording data acquisition module continuously obtains the original audio data stream and passes it to the feature hierarchical analysis module.

[0062] When the feature hierarchical parsing module processes the expanded speech data segments, the first step is to classify them into valid and invalid speech segments. During a meeting, when participants speak, the speech signal energy is relatively high and the frequency is within the normal human speech range. These speech frames are considered valid speech segments. However, when there are brief silences, the sound of shuffling papers, or the slight hum of electrical current, the corresponding speech frames have low energy or abnormal frequencies and are considered invalid speech segments. For example, when participant A is explaining his or her point of view, the speech portion of his or her speech is considered valid, while the occasional brief pauses are considered invalid speech segments.

[0063] After the division is completed, the short-time energy spectrum features of the effective speech segment are extracted based on the preset time-frequency transformation window to generate a time-frequency feature matrix. Assuming that the duration of the preset time-frequency transformation window is 0.1 seconds, the effective speech segment spoken by participant A is analyzed segment by segment at intervals of 0.1 seconds. For each segment, its energy distribution at different frequencies is calculated. For example, in a certain 0.1-second speech segment, the energy in the frequency range of 200Hz-500Hz is higher, while the energy in other frequency ranges is lower. By analyzing multiple such 0.1-second speech segments, a time-frequency feature matrix reflecting the energy distribution of the effective speech segment at different times and frequencies can be generated.

[0064] Next, fundamental frequency tracking is performed on the background noise segments within the extended speech data segment. During a meeting, background noise includes the sound of the air conditioner running in the conference room and traffic noise from the distant street. Taking the air conditioner sound as an example, fundamental frequency tracking technology analyzes the waveform distortion characteristics of each fundamental frequency cycle. Since the frequency and amplitude of the air conditioner sound fluctuate, this causes distortion in the speech waveform. These distortion features, such as changes in the peaks and valleys of the waveform, are extracted and constructed into a fundamental frequency feature vector.

[0065] Finally, a multi-dimensional feature fusion is performed on the time-frequency feature matrix and the fundamental frequency feature vector to generate a second semantic feature set. This fusion is performed using a splicing approach, where the time-frequency feature matrix and the fundamental frequency feature vector are connected in a certain order to form a new feature set, which is the second semantic feature set. This fusion integrates the time-frequency characteristics of speech and the impact of background noise on the speech waveform, more comprehensively reflecting the emotional fluctuation characteristics of the extended speech data segment. For example, when participant A is emotionally excited while speaking, their speech speed increases and their tone rises. At the same time, changes in background noise also affect the speech. This information is included in the generated second semantic feature set, providing rich data support for subsequent semantic analysis and intent judgment.

[0066] Example 2:

[0067] This embodiment mainly describes the preset voiceprint classification rules and the specific process of the recording data acquisition module dividing the original audio data stream based on the rules.

[0068] The preset voiceprint classification rules include a basic acoustic type set and an extended acoustic type set. The basic acoustic type set includes voice content identification, environmental noise identification, and device fingerprint identification. Voice content identification is used to identify specific speech information in the voice, environmental noise identification is used to distinguish noise characteristics in different environments, and device fingerprint identification can identify the characteristics of the device used to collect audio. The extended acoustic type set includes user identity identification and recording scene identification. Each acoustic identification corresponds to an independent data parsing protocol.

[0069] When dividing the raw audio data stream, the recording data acquisition module uses the recording interface module to obtain raw waveform data from the local recording device in real time. The recording interface module establishes a physical connection between the system and the local recording device, ensuring stable data transmission. After obtaining the raw waveform data, the header features of the raw waveform data are matched according to the identifiers in the basic acoustic type set. For example, by identifying the specific frequency, amplitude, and other characteristic patterns in the waveform data header, basic audio segments related to speech content, ambient noise, and device fingerprints can be separated.

[0070] Then, the tail features of the original waveform data are scanned based on the identifiers in the extended acoustic type set to extract the extended audio segment. For example, the user identity identifier and recording scene identifier are identified based on specific acoustic patterns in the tail features to extract the corresponding extended audio segment.

[0071] Finally, the separated basic audio segment and extended audio segment are synchronized according to the sampling point sequence to ensure the time sequence consistency of the data. After synchronization, they are written into the basic voiceprint buffer and the extended voiceprint buffer respectively, providing an orderly data foundation for subsequent feature extraction and analysis.

[0072] Suppose a teaching activity is taking place in a classroom environment. A recording device is installed in the classroom to record the sounds of the teacher's lectures and the interactions between teachers and students. The original audio data stream formed by these sounds is acquired by the recording data acquisition module.

[0073] Within the set of basic acoustic types in the preset voiceprint classification rules, speech content identification is used to identify the specific content of speech between teachers and students. For example, when a teacher says, "Today we will learn Newton's first law," the specific pronunciation, vocabulary, and other features of this speech are manifestations of speech content identification. Ambient noise identification can identify ambient sounds in the classroom, such as the hum of a fan or the sounds of desks and chairs moving. These sounds have unique frequency and amplitude characteristics. Device fingerprint identification can distinguish the specific recording device used in the classroom. Each device leaves unique signal signatures during audio collection, such as weak signal interference generated by specific electronic components.

[0074] Expanding the user identifiers in the acoustic type set, in this classroom scenario, the teacher or student can be identified by analyzing voice characteristics such as timbre and pronunciation habits. For example, Teacher A's pronunciation habits are moderate speed and certain words have a unique intonation, which constitutes Teacher A's user identifier. The recording scene identifier identifies the classroom teaching environment. The acoustic reflection characteristics of the classroom space and the sound reverberation caused by the density of people are all part of the recording scene identifier. Each acoustic identifier corresponds to an independent data parsing protocol to ensure accurate extraction and analysis of relevant information.

[0075] The recording data acquisition module begins dividing the raw audio data stream. It uses the recording interface module to obtain raw waveform data in real time from the classroom's recording equipment. The recording interface module establishes a stable physical connection between the system and the recording equipment, ensuring reliable data transmission. After obtaining the raw waveform data, the head features of the raw waveform data are matched according to the identifiers in the basic acoustic type set. For example, by identifying the starting frequency and amplitude change pattern of the waveform data header, it can determine which parts belong to the teacher or student's speech content and which parts are environmental noise or signals generated by the device itself, thereby separating the basic audio segments. In this process, if the waveform header exhibits clear speech frequency characteristics and the amplitude is within the normal speech range, it can be determined to be a basic audio segment related to the speech content.

[0076] The extended audio segment is extracted by scanning the tail features of the original waveform data based on the identifiers in the extended acoustic type set. For example, by analyzing the specific timbre characteristics at the end of the waveform data, it can be determined whether the speaker is teacher A or a student, thereby extracting the extended audio segment associated with the user's identity. Furthermore, based on the sound reverberation characteristics and changes in the ambient noise at the end, it can be determined that the recording was made in a classroom setting, thereby extracting the extended audio segment associated with the recording scene identifier.

[0077] The separated basic audio segment and extended audio segment are synchronized according to the sampling point sequence. This ensures the temporal consistency of the two audio segments and the accuracy of subsequent analysis. After synchronization, the basic audio segment is written to the basic voiceprint cache, and the extended audio segment is written to the extended voiceprint cache. This provides an organized and clearly classified data foundation for subsequent feature extraction and analysis, facilitating the system's further in-depth processing of the voice data, such as analyzing the teacher's teaching effectiveness and student engagement.

[0078] Example 3:

[0079] This embodiment focuses on different types of preset semantic mapping tables and methods for generating intent classification labels.

[0080] When the preset semantic mapping table uses a linear matching model, the first and second semantic feature sets are first normalized and fused. Normalization eliminates differences in numerical range and dimension between feature sets, ensuring that they can be fused at the same scale. Fusion can be performed using simple addition or weighted addition to obtain a normalized fusion value.

[0081] This normalized fusion value is then input into a linear mapping model. This model, trained using a large amount of speech data, establishes a linear relationship between the normalized fusion value and the intent classification label. Through this linear mapping, the corresponding intent classification label is obtained, which serves as one of the parsing parameters for the current speech.

[0082] When the preset semantic mapping table adopts a nonlinear matching model, the first semantic feature set and the second semantic feature set are nonlinearly combined. The nonlinear combination can be performed in various ways, such as the neuron connection method in a neural network, to combine the two feature sets in a complex nonlinear relationship.

[0083] Next, the nonlinear combination results are identified using a classification function. This trained classification function can identify the discrete semantic label associated with the nonlinear combination results based on their characteristic patterns. This discrete semantic label is the intent classification label. This nonlinear matching model can more flexibly handle complex semantic relationships and improve the accuracy of intent classification.

[0084] Assume that in a smart home environment, a user interacts with a smart speaker through voice. When the preset semantic mapping table uses a linear matching model:

[0085] The user says "turn on the living room light" to the smart speaker. After the recording data acquisition module collects this voice, the feature hierarchical analysis module extracts the voiceprint pattern features of the basic voice data segment from it to generate the first semantic feature set, recorded as F1, which contains features corresponding to key information such as "living room light" and "turn on" in the voice; at the same time, the emotional fluctuation features of the extended voice data segment are extracted to generate the second semantic feature set, recorded as F2, such as the user's urgency when speaking, the tone of voice, and other features.

[0086] First, F1 and F2 are normalized and fused. The purpose of normalization is to make different feature sets have the same scale to facilitate subsequent calculations. Here, the maximum-minimum normalization method is used, and the formula is: Where X represents the original eigenvalue, X min is the minimum value of the feature set, X max is the maximum value of the feature set, X norm is the normalized eigenvalue. Each eigenvalue in F1 and F2 is normalized to obtain the normalized feature set F 1n and F 2n , and then perform fusion, here simply add and fuse to get the normalized fusion value S, that is, S = F 1n +F 2n .

[0087] The preset linear mapping model is trained using a large amount of smart home voice command data. Assume the linear mapping relationship is Y = aS + b, where Y is the intent classification label, and a and b are coefficients determined through training. In this example, after calculation, Y corresponds to the intent classification label "Turn on the living room lights." This becomes the parsing parameter for the current voice command, and the system then takes appropriate actions based on this parsing parameter.

[0088] When the preset semantic mapping table adopts a nonlinear matching model:

[0089] If the user says "turn on the living room light," the extraction process for F1 and F2 remains unchanged. F1 and F2 are then nonlinearly combined, for example, using a simple neural network structure. This neural network has several neurons, and the eigenvalues of F1 and F2 serve as input. The neurons then perform complex nonlinear calculations using specific weights and activation functions.

[0090] After the nonlinear combination, the results are then evaluated using a classification function. Assuming the classification function is based on a decision tree algorithm, the decision tree performs layer-by-layer judgments based on various feature patterns within the nonlinear combination results. For example, it first determines whether features related to "light" are present, then determines whether the action is "on" or "off," and ultimately outputs a discrete semantic label, "Turn on the living room lights," as the intent classification label. This determines the user's intent and allows the smart speaker to perform the corresponding light-on operation.

[0091] Example 4:

[0092] This embodiment focuses on the functions of the voice resource scheduling module and the specific process of generating a device deployment topology diagram.

[0093] The voice resource scheduling module is connected to the central processing unit and is connected to the local voice device database through the recording interface module. When receiving the device call request in the final control instruction of the central processing unit, the voice resource scheduling module starts working.

[0094] It matches the available device list with the local voice device database. The local voice device database stores information about various voice devices, including device type, function, and current status. The voice resource scheduling module selects eligible available devices based on the requirements in the device call request, such as device functionality and usage scenario, to form a list of available devices.

[0095] When generating a device deployment topology, first obtain a recording space plan model. This model can be obtained through on-site measurements, map data, or a preset space template. It reflects the layout and dimensions of the recording space.

[0096] Mark the real-time voiceprint feature points of each device in the available device list on the plane model. These voiceprint feature points are determined based on the voiceprint characteristics of the device's current location and environment, and are obtained by analyzing the test signals sent by the device or the actual voiceprint data collected.

[0097] Next, a path search algorithm is used to calculate the optimal matching path from each device's feature point to the target voice area. This path search algorithm, which can employ common algorithms such as Dijkstra or A*, calculates the optimal signal propagation path from the device to the target voice area based on spatial layout and signal propagation characteristics. The list of available devices is prioritized by path matching, with devices with higher path matching ranking first.

[0098] Finally, the optimal matching path and priority ranking are superimposed on the plane model to generate a real-time updated device deployment topology map. This device deployment topology map can intuitively display the location of each available device in the recording space, the optimal path to the target voice area, and the device priority, providing a visual basis for optimizing the voiceprint matching path.

[0099] Imagine a scenario where a voice monitoring and equipment scheduling system based on the present invention is deployed in a large shopping mall's monitoring center. Multiple recording devices are installed in the center to collect real-time sound information from various areas within the mall. When an unusual sound, such as an argument or alarm, is detected in a particular area, the central processing unit (CPU) sends a device call request to the voice resource scheduling module based on the results of the speech recognition and analysis system.

[0100] After receiving the request, the voice resource scheduling module first matches the available device list from the local voice device database. The local voice device database stores detailed information about all relevant voice devices in the mall, including the location, model, and current operating status of each broadcast speaker and surveillance camera (some of which have voice functions). Assuming that the abnormal sound occurs in the clothing area on the third floor of the mall, the voice resource scheduling module will use the location information ("clothing area on the third floor of the mall") in the device call request to filter out broadcast speakers and surveillance cameras that are close to this area and in normal working condition, forming a list of available devices. For example, broadcast speakers numbered A1 and A2 and surveillance cameras numbered C3 and C4 are included in the list.

[0101] Next, we generate a device deployment topology diagram. The first step is to obtain a floor plan model of the recording space. In this shopping mall scenario, we can use the mall's architectural drawings to construct a floor plan model. This model accurately displays the layout of each floor, aisle locations, store distribution, and the installation locations of each voice device.

[0102] Mark the real-time voiceprint feature points of each device in the list of available devices in the plane model. Due to the complex environment of the shopping mall and the different sound propagation characteristics in different locations, the voiceprint features received by the devices will also vary. By sending a specific test signal to each device and analyzing the reflected signal received by the device and the actual collected environmental voiceprint data, the real-time voiceprint feature points of each device are determined. For example, the A1 broadcast speaker is located near the entrance to the clothing area, and its voiceprint feature points reflect the unique characteristics of sound propagation at that location, including signal strength, frequency attenuation and other information; the C3 surveillance camera is installed in the corner of the clothing area, and its voiceprint feature points include the impact of factors such as acoustic reflection and occlusion in that corner on the sound.

[0103] Next, a path search algorithm is used to calculate the optimal matching path from each device's feature point to the target voice area (i.e., the clothing section on the third floor of the mall, where the abnormal sound is occurring). Algorithm A is used for path calculation. Algorithm A comprehensively considers factors such as the distance from the device to the target area and the sound propagation loss along the path. For the A1 speaker, the algorithm calculates multiple possible paths from its location to the target point in the clothing section based on the mall's layout. It then evaluates the combined cost of each path and ultimately determines the optimal path, which may follow the shortest aisle and avoid large obstacles. Similarly, for the C3 surveillance camera, the algorithm calculates the optimal path from its installation location to the target area, which may require navigating around obstructions such as shelves and walls. The list of available devices is prioritized based on path matching. A high path matching score indicates better signal transmission from the device to the target area with less interference. For example, the A1 speaker has the highest path matching score due to its shorter path distance and relatively low sound propagation loss, placing it first. The C3 surveillance camera has the second highest path matching score, placing it second, and so on.

[0104] Finally, the optimal matching path and priority ranking are overlaid onto the plane model to generate a real-time device deployment topology map. This topology map allows monitoring center staff to visually visualize the location of each available device within the mall, the optimal path to the target area, and the device priority order. This allows staff to quickly identify which devices are most effective in responding to abnormal situations, such as prioritizing the use of A1 speakers to play warning messages while simultaneously utilizing C3 surveillance cameras to capture on-site footage. This allows for timely action to address abnormal situations within the mall, achieving efficient monitoring and management of the mall environment.

[0105] Example 5:

[0106] This embodiment mainly describes the specific working processes of the instruction conversion module and the voice archiving module.

[0107] The command conversion module is connected to the central processing unit (CPU). Its function is to convert the final control command generated by the CPU into a voice-driven signal. The final control command is a control signal generated after logical verification and contains the operating requirements for the target voice device. The command conversion module converts the final control command into the corresponding voice-driven signal based on the type and interface specifications of the target voice device. For example, if the target voice device is a smart speaker, the command conversion module will convert the control command into a voice-driven signal that complies with the smart speaker's communication protocol.

[0108] The converted voice-driven signal is sent to the target voice device through the recording interface module, triggering it to perform the corresponding action. The recording interface module ensures that the voice-driven signal is accurately and stably transmitted to the target voice device, enabling the target voice device to play the voice, perform the operation, and other actions as instructed.

[0109] The voice archiving module, also connected to the central processing unit, stores the raw audio data stream, the first semantic feature set, the second semantic feature set, the intent classification label, and the final control instruction. This data is stored sequentially according to the recording time sequence, forming a complete voice processing record chain. This allows for convenient access to relevant data from the voice archiving module when querying or analyzing the voice processing process. For example, when optimizing a voice recognition algorithm or troubleshooting a problem, reviewing the historical voice processing record chain allows understanding the data at different stages, providing a basis for improvement and debugging.

[0110] Imagine a smart office scenario where employee Xiao Zhang uses the company's intelligent voice assistant to schedule and collaborate on work tasks. When Xiao Zhang says, "Please schedule a meeting with client Mr. Li at 3 p.m. tomorrow and prepare the relevant project materials," the various system modules begin to work together.

[0111] After completing voice data recognition, analysis, and logical verification, the central processing unit generates a final control instruction. This instruction contains specific operational requirements for scheduling a meeting and preparing project materials, such as the meeting time, attendees, and specific information about the project materials to be prepared.

[0112] The command conversion module receives the final control command from the central processing unit. Because the communication protocols used by the company's meeting reservation system and document management system differ from the command format generated by the central processing unit, the command conversion module needs to convert the final control command into voice-activated signals that the corresponding systems can recognize. For example, the meeting reservation system uses a specific API to receive reservation information. The command conversion module re-encodes and encapsulates information such as the meeting time (tomorrow at 3 p.m.) and the attendees (client Mr. Li) in the format required by the API, converting it into a voice-activated signal that complies with the meeting reservation system's communication protocol. For the document management system, the command conversion module similarly converts instructions for preparing project materials into a signal format recognizable by the document management system. After conversion, the command conversion module transmits these voice-activated signals to the corresponding target voice devices—the relevant terminal devices of the meeting reservation system and document management system—through the recording interface module, triggering them to perform the corresponding actions. Upon receiving the signal, the meeting reservation system automatically creates a meeting reservation record for tomorrow at 3 p.m. with client Mr. Li. The document management system, in response to the command, begins collecting and organizing the project materials specified by Xiao Zhang.

[0113] At the same time, the voice archiving module begins operating. Connected to the central processing unit, the voice archiving module stores key data generated throughout the entire voice processing process. First, the raw audio data stream—the voice data captured by the recording device as Xiao Zhang spoke the command—is stored intact for possible subsequent review and analysis. Next, the first and second semantic feature sets are generated. These include the voiceprint and emotional fluctuation features extracted from Xiao Zhang's voice, respectively. These features provide a basis for understanding Xiao Zhang's intentions and emotional state. The intent classification label is also stored. In this example, the intent classification label might be "meeting appointment and document preparation," which clearly defines the core intent of Xiao Zhang's voice command. Finally, the final control command generated by the central processing unit is also stored in the voice archiving module. The voice archiving module organizes and stores this data in chronological order, forming a complete voice processing record chain. For example, in subsequent work, if there is a meeting appointment error or incomplete document preparation, managers can query the record chain in the voice archiving module to trace the entire voice processing process and check which link has the problem, whether it is a voice recognition error or an error in the command conversion or execution process, so as to quickly locate the problem and correct it. At the same time, it can also provide actual data support for optimizing the intelligent voice assistant system.

[0114] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus.

[0115] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.

Claims

1. An intelligent speech recognition and analysis system for offline recording, characterized in that: include: The recording data acquisition module is used to obtain the original audio data stream of the offline scene and divide the original audio data stream into basic voice data segments and extended voice data segments based on preset voiceprint classification rules; a feature hierarchical analysis module, configured to extract voiceprint features from the basic voice data segment to generate a first semantic feature set, and extract emotional fluctuation features from the extended voice data segment to generate a second semantic feature set; A semantic matching module, configured to jointly encode the first semantic feature set and the second semantic feature set according to a preset semantic mapping table and map them to corresponding intent classification labels, and use the intent classification labels as parsing parameters for the current speech; An intention decision module is used to call a target execution scheme in a preset strategy library based on the intention classification label, and use the target execution scheme as a response parameter of the current voice; A central processing unit is used to send the original audio data stream to the feature hierarchical parsing module, send the first semantic feature set and the second semantic feature set to the semantic matching module, and perform logical verification on the parsing parameters and the response parameters to generate a final control instruction.

2. The speech recognition analysis system according to claim 1, characterized in that The feature hierarchical analysis module extracts emotional fluctuation features from the extended voice data segment, including: Dividing the continuous speech frames in the extended speech data segment into valid speech segments and invalid speech segments, and performing short-time energy spectrum feature extraction on the valid speech segments based on a preset time-frequency transformation window to generate a time-frequency feature matrix; Performing fundamental frequency tracking processing on the background noise segment in the extended voice data segment, extracting waveform distortion features of each fundamental frequency period and constructing a fundamental frequency feature vector; Perform multi-dimensional feature fusion on the time-frequency feature matrix and the fundamental frequency feature vector to generate the second semantic feature set.

3. The speech recognition analysis system according to claim 1, wherein: The preset voiceprint classification rules include a basic acoustic type set and an extended acoustic type set; the basic acoustic type set includes a voice content identifier, an environmental noise identifier, and a device fingerprint identifier; the extended acoustic type set includes a user identity identifier and a recording scene identifier, and each acoustic identifier corresponds to an independent data parsing protocol.

4. The speech recognition analysis system according to claim 3, characterized in that: The system further includes a recording interface module, which is used to realize the physical connection between the recording data acquisition module, the feature layer analysis module, the semantic matching module and the intention decision module and the local recording device respectively; The recording data acquisition module divides the original audio data stream based on the preset voiceprint classification rule, including: Acquiring raw waveform data from a local recording device in real time through the recording interface module, and matching head features of the raw waveform data according to identifiers in the basic acoustic type set to separate basic audio segments; Scanning tail features of the original waveform data according to identifiers in the extended acoustic type set to extract an extended audio segment; The basic audio segment and the extended audio segment are synchronized according to the sampling point sequence number and then written into the basic voiceprint buffer area and the extended voiceprint buffer area respectively.

5. The speech recognition analysis system according to claim 1, characterized in that: When the preset semantic mapping table adopts a linear matching model, the intention classification label is a linear mapping result of the normalized fusion value of the first semantic feature set and the second semantic feature set; When the preset semantic mapping table adopts a nonlinear matching model, the intention classification label is a discrete semantic label that discriminates a nonlinear combination result of the first semantic feature set and the second semantic feature set through a classification function.

6. The speech recognition analysis system according to claim 1, characterized in that It also includes a voice resource scheduling module connected to the central processing unit, and the voice resource scheduling module is connected to the local voice device database through the recording interface module; The voice resource scheduling module is used to match the available device list from the local voice device database according to the device call request in the final control instruction, and generate a device deployment topology diagram to optimize the voiceprint matching path.

7. The speech recognition analysis system according to claim 6, characterized in that: The voice resource scheduling module generates a device deployment topology diagram including: Obtaining a plane model of the recording space, and marking the real-time voiceprint feature points of each device in the available device list in the plane model; Calculating the optimal matching path from the feature point of each device to the target voice area based on a path search algorithm, and prioritizing the list of available devices according to the path matching degree; The optimal matching path and the priority ranking are superimposed on the plane model to generate a device deployment topology diagram updated in real time.

8. The speech recognition analysis system according to claim 1, characterized in that: When the central processing unit performs logical verification on the parsing parameters and the response parameters, a dual verification mechanism of integrity verification rules and logical consistency detection rules is adopted. The integrity verification rules are used to confirm the validity of the parameters, and the logical consistency detection rules are used to eliminate semantic conflicts between parameters.

9. The speech recognition analysis system according to claim 1, characterized in that: It also includes an instruction conversion module connected to the central processing unit, which is used to convert the final control instruction into a voice drive signal and send the voice drive signal to the target voice device through the recording interface module to trigger the execution action.

10. The speech recognition analysis system according to claim 1, characterized in that: It also includes a voice archiving module connected to the central processing unit, which is used to store the original audio data stream, the first semantic feature set, the second semantic feature set, the intention classification label and the final control instruction, and generate a complete voice processing record chain according to the recording time sequence.

Citation Information

Patent Citations

  • Voice emotion interaction method, computer equipment and computer readable storage medium

    CN110085221A

  • Speech emotion recognition method

    CN113409824A

  • Abnormal sound recognition model construction method and abnormal sound detection method and system

    CN114724584A

  • Voice control method and device of equipment, storage medium and electronic device

    CN116504225A

  • System and method for multi-modal focus detection, referential ambiguity resolution and mood classification using multi-modal input

    US20020135618A1

Cited By

  • Remote acquisition system based on four diagnosis methods of traditional Chinese medicine

    CN121171535A