Robot voice interaction signal recognition and processing method and system

By training the basic acoustic model and combining it with lightweight processing and parallelization technology, the problems of low speech recognition rate and insufficient hardware resources in low-resource language scenarios are solved, and efficient and low-latency speech interaction recognition is achieved.

CN120673748AActive Publication Date: 2025-09-19BEIJING QINGFEI TECH CO LTD

Patent Information

Application Number
CN202510772944.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-11
Publication Date
2025-09-19
Estimated Expiration
2045-06-11

AI Technical Summary

Technical Problem

Existing speech recognition models have low recognition rates in low-resource language scenarios and have difficulty handling complex conversations or multi-round interactions. In addition, deep learning models have high requirements for embedded robot hardware resources, resulting in response delays and increased energy consumption.

Method used

The basic acoustic model is trained based on high-resource speech data, the underlying parameters are frozen and the top-level language-specific parameters are fine-tuned. A lightweight end-to-end speech recognition model is used for parallel processing, and the model is adaptively optimized based on historical human-computer interaction data.

Benefits of technology

It achieves accurate recognition of mixed Chinese and English input, dialects and low-resource languages, reduces latency and improves real-time performance, and adapts to embedded robot hardware resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120673748A_ABST
    Figure CN120673748A_ABST
Patent Text Reader

Abstract

The invention discloses a robot voice interaction signal identification and processing method and system, and relates to the technical field of electric digital processing, and the method comprises the steps: training a basic acoustic model based on high-resource voice data; freezing bottom layer parameters of the basic acoustic model, and finely adjusting top layer language specificity parameters by using low-resource voice data to obtain a multilingual voice signal recognition model; carrying out acceleration processing on the multilingual speech signal recognition model; splitting processing tasks of the voice signals, and parallelizing the processing tasks by using hardware resources of the robot; and adaptively optimizing the multilingual speech signal recognition model according to the man-machine interaction historical data. Chinese and English mixed input and accurate recognition of dialects and low-resource languages are supported, a lightweight end-to-end speech recognition model and a parallel processing architecture are adopted, and low delay and high real-time performance are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of electronic digital processing technology, and in particular to a method and system for recognizing and processing robot voice interaction signals. Background Art

[0002] With the rapid development of artificial intelligence and robotics, voice interaction has become a core method of human-computer interaction. Compared to traditional touch or command input, voice interaction offers advantages such as naturalness, efficiency, and the absence of physical contact. It has been widely used in fields such as service robotics, smart homes, and industrial automation. However, traditional speech recognition models (such as hidden Markov models combined with Gaussian mixture models) have limited generalization capabilities for pronunciation habits, dialects, and accents, and recognition rates drop significantly in low-resource language scenarios. Furthermore, most systems rely on pre-set grammatical templates or finite-state semantic parsing frameworks, lacking support for dynamic reasoning about contextual relevance and user intent, making them difficult to handle complex conversations or multi-turn interactions. Furthermore, existing speech processing pipelines face a significant conflict between real-time performance and computational efficiency. While deep learning models (such as end-to-end speech recognition networks) have improved recognition accuracy, their large number of parameters places high demands on the computing power and memory resources of embedded robot hardware platforms, resulting in response delays and increased energy consumption. Therefore, designing a highly robust, low-latency, and resource-efficient speech signal recognition and processing method is crucial for advancing the intelligentization of robotic interactions. Summary of the Invention

[0003] The present invention provides a method for recognizing and processing robot voice interaction signals, comprising: Step 1: Train a basic acoustic model based on high-resource speech data; Step 2: Freeze the underlying parameters of the basic acoustic model and fine-tune the top-level language-specific parameters using low-resource speech data to obtain a multilingual speech signal recognition model. Step 3: Accelerate the multilingual speech signal recognition model; Step 4: Split the voice signal processing tasks and parallelize them using the robot's hardware resources. Step 5: Adaptively optimize the multilingual speech signal recognition model based on historical human-computer interaction data.

[0004] The above-mentioned method for recognizing and processing robot voice interaction signals, in which a basic acoustic model is trained based on high-resource voice data, is specifically divided into the following sub-steps: The high-resource speech data is converted into a numerical sequence as the model input, and the corresponding phoneme sequence is used as the model output to construct a first training data set; Use the first training dataset to pre-train the basic acoustic model to enhance the model's ability to extract common cross-language features; The introduction of a joint training strategy enables the model to gradually transition to the phoneme classification task.

[0005] The method for recognizing and processing robot voice interaction signals described above, wherein the underlying parameters of the basic acoustic model are frozen and the top-level language-specific parameters are fine-tuned using low-resource speech data, is specifically divided into the following sub-steps: Adaptively select the number of frozen layers based on the amount of low-resource data; Processing the low-resource speech data into a second training dataset; The unfrozen top model parameters are fine-tuned using the second training dataset.

[0006] The above-mentioned method for recognizing and processing robot voice interaction signals, wherein a multilingual voice signal recognition model is adaptively optimized based on historical human-computer interaction data, is specifically divided into the following sub-steps: Filter samples with confidence levels below a threshold in historical human-computer interaction data to generate incremental training datasets; Use incremental training datasets to incrementally train multilingual speech signal recognition models.

[0007] The present invention also provides a robot voice interaction signal recognition and processing system, comprising: an acoustic model preparation module, an acoustic model fine-tuning module, an acoustic model acceleration module, a signal processing task acceleration module, and an acoustic model enhancement module; Acoustic model preparation module, used to train a basic acoustic model based on high-resource speech data; The acoustic model fine-tuning module freezes the underlying parameters of the basic acoustic model and uses low-resource speech data to fine-tune the top-level language-specific parameters to obtain a multilingual speech signal recognition model. Acoustic model acceleration module, used to accelerate the processing of multilingual speech signal recognition models; The signal processing task acceleration module is used to split the voice signal processing tasks and parallelize the processing tasks by utilizing the robot's hardware resources; The acoustic model enhancement module is used to adaptively optimize the multilingual speech signal recognition model based on historical human-computer interaction data.

[0008] The beneficial effects achieved by the present invention are as follows: it supports mixed Chinese and English input, accurate recognition of dialects and low-resource languages, and adopts a lightweight end-to-end speech recognition model and parallel processing architecture to achieve low latency and high real-time performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments described in the present invention. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.

[0010] Figure 1 This is a flow chart of a method for identifying and processing robot voice interaction signals provided in Example 1 of the present application; Figure 2 This is a schematic diagram of a robot voice interaction signal recognition and processing system provided in Example 2 of the present application. DETAILED DESCRIPTION

[0011] The following is a clear and complete description of the technical solutions in the embodiments of the present invention, in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.

[0012] Example 1 like Figure 1 As shown, the first embodiment of the present application provides a method for recognizing and processing robot voice interaction signals, including: Step S10: training a basic acoustic model based on high-resource speech data; Prepare a multilingual speech database that includes both high-resource speech data such as Chinese and English, as well as low-resource speech data such as various local dialects and Tibetan and Mongolian. The text annotations for the speech data in the database must include phoneme-level alignment (e.g., generated using a forced alignment tool). Train a basic acoustic model based on the high-resource speech data, which is divided into the following sub-steps: Step S11: converting high-resource speech data into a numerical sequence as model input and the corresponding phoneme sequence as model output to construct a first training data set; Bring the voice frames of high-resource voice data into the conversion function one by one The converted numerical sequence A' is used as the model input, and the corresponding phoneme sequence is used as the model output. After sorting, the first training data set is obtained. The conversion function Expressed as: , in is the i-th frame speech signal in a high-resource speech data A, is the time-varying steepness parameter, is the phase enhancement coefficient, and They are the i-th frame and the i-th frame in A respectively. The fundamental frequency of the frame speech signal, is the floor function, and n is the total number of speech frames in A.

[0013] Step S12: pre-training the basic acoustic model using the first training data set to enhance the model's ability to extract cross-language common features; The basic acoustic model adopts the Transformer encoder structure and first optimizes the contrast loss , enhance the model's ability to extract cross-language common features, input the data samples in the first training dataset into the basic acoustic model in batches, and fine-tune the model based on the contrast loss calculation results; Contrastive loss The calculation formula is: , in is the cosine similarity function, is the phoneme sequence output by the basic acoustic model for the u-th input sample, Indicates the phoneme sequence actually corresponding to the u-th input sample, Indicates the phoneme sequence actually corresponding to the k-th input sample in the current training batch, is an adjustable parameter, k takes values ​​from 1 to , N is the number of samples in the current training batch, is the number of samples in the current training batch that actually correspond to the phoneme sequence of the u-th input sample.

[0014] Step S13: Introducing a joint training strategy to gradually transition the model to the phoneme classification task; When the contrast loss curve no longer has a clear downward trend, a joint training strategy is introduced to gradually transition the model to the phoneme classification task. At this time, the model needs to be further trained based on the total loss calculation results. The total loss at the current stage is The calculation formula is: , in is an adjustable weight, is the contrast loss, is an input sample calculated by dynamic programming To model output The sum of all possible alignment path probabilities, The number of phonemes replaced by the model output phoneme sequence for the actual phoneme sequence, The number of phonemes deleted from the model output phoneme sequence for the actual phoneme sequence, Output the number of newly inserted phonemes for the model's phoneme sequence to the actual phoneme. is the total number of factors in the actual phoneme sequence.

[0015] Step S20: Freeze the bottom-level parameters of the basic acoustic model and fine-tune the top-level language-specific parameters using low-resource speech data to obtain a multilingual speech signal recognition model; Freeze the underlying parameters of the basic acoustic model and fine-tune only the top-level parameters. This preserves the general acoustic features learned at the bottom level while adapting the top level to the pronunciation characteristics of low-resource languages. This is done in the following sub-steps: Step S21: adaptively selecting the number of frozen layers according to the amount of low resource data; Use the formula: Adaptively select the number of layers that should be frozen for the current model ,in is the total number of layers of the basic acoustic model, For low-resource voice data, The data volume of high-resource voice data.

[0016] Step S22: Processing the low-resource speech data into a second training data set; Using the aforementioned conversion function The low-resource speech data is converted into a numerical sequence as the model input, and the corresponding phoneme sequence is used as the output. After sorting, the second training dataset is obtained.

[0017] Step S23: fine-tuning the unfrozen top-level model parameters using the second training data set; Total loss during the top-level fine-tuning phase The calculation formula is expressed as: , in Is the current model for the input sample The cross entropy loss of the t-th data in , t ranges from 1 to T, and T is the input sample The total number of data items contained in , 、 is an adjustable weight, is the KL divergence calculation function, Represents the basic acoustic model for the input sample The output probability distribution of Indicates the current model for the input sample The output probability distribution of , is the j-th model parameter of the current model, yes The diagonal values ​​of the Fisher information matrix, is the jth model parameter of the basic acoustic model, j ranges from 1 to m, and m is the total number of model parameters.

[0018] During the first 1000 steps of fine-tuning training, are all set to 0, and the model's generalization ability for low-resource speech data is improved based only on the cross entropy loss. During subsequent fine-tuning training, the model is adjusted according to the actual situation. Assign values ​​to avoid overfitting and catastrophic forgetting. Fine-tuning is stopped when the total loss value stops decreasing or shows no clear downward trend, resulting in a multilingual speech signal recognition model.

[0019] Step S30: performing acceleration processing on the multilingual speech signal recognition model; To improve the real-time performance of the multilingual speech signal recognition model, the model needs to be accelerated. This is divided into the following sub-steps: Step S31: compressing the multilingual speech signal recognition model; Through AQW quantization technology, the model can be compressed several times with almost no impact on performance.

[0020] Step S32: Divide the input data into small segments and input them into the model recognition segment by segment; Input is performed segment by segment to achieve simultaneous input and processing of voice interaction signals, saving time waiting for complete sentences and improving computing efficiency.

[0021] Step S40: splitting the voice signal processing tasks and parallelizing the processing tasks using the robot's hardware resources; The multilingual speech signal recognition model, after accelerated processing, is now adaptable to the hardware resources of the embedded robot. Now, the complete speech signal processing task needs to be parallelized to further improve the real-time performance of speech recognition and interaction. First, the speech signal processing task is split into multiple subtasks. Each of these subtasks is then deployed to different computing chips for parallel processing. The specific parallel pipeline design is as follows: Stage 1 (DSP): signal acquisition → framing → noise suppression; Stage 2 (CPU+FPGA): Signal conversion (CPU) + baseband extraction (FPGA); Stage 3 (GPU): Context splicing → Phoneme recognition → Streaming decoding.

[0022] Double buffering is used to further optimize the parallel speed: Each stage maintains two buffers (A / B). When Stage 1 writes A, Stage 2 processes B, and Stage 3 reads the results of the previous cycle.

[0023] Step S50: Adaptively optimizing the multilingual speech signal recognition model based on historical human-computer interaction data; Historical human-computer interaction data is key to identifying model deficiencies. Integrating this data into the training dataset for incremental model training enables adaptive optimization of the model, allowing the model to grow during use. Specifically: Step S51: Filtering samples with confidence levels lower than a threshold in the human-computer interaction history data to generate an incremental training data set; The confidence level is the model's predicted probability for a certain interactive speech recognition result. The threshold is the system preset value. The selected samples are labeled with the correct factor sequence and organized into a data set to obtain an incremental training data set.

[0024] Step S52: performing incremental training on the multilingual speech signal recognition model using the incremental training dataset; The incremental training process is the same as that of fine-tuning the basic acoustic model, and the bottom layer needs to be frozen first. The model parameters of the layer are then fine-tuned using the incremental training dataset to fine-tune the unfrozen model parameters of the top layer. The loss function used is also .

[0025] Example 2 like Figure 2 As shown, the second embodiment of the present application provides a robot voice interaction signal recognition and processing system, including: an acoustic model preparation module 21, an acoustic model fine-tuning module 22, an acoustic model acceleration module 23, a signal processing task acceleration module 24, and an acoustic model enhancement module 25; The acoustic model preparation module 21 is used to train a basic acoustic model based on high-resource speech data; specifically, it includes: a first training data set construction submodule, a model pre-training submodule; 1. A first training dataset construction submodule, which is used to convert high-resource speech data into a numerical sequence as model input and the corresponding phoneme sequence as model output to construct the first training dataset; Bring the voice frames of high-resource voice data into the conversion function one by one The converted numerical sequence A' is used as the model input, and the corresponding phoneme sequence is used as the model output. After sorting, the first training data set is obtained. The conversion function Expressed as: ,in is the i-th frame speech signal in a high-resource speech data A, is the time-varying steepness parameter, is the phase enhancement coefficient, and They are the i-th frame and the i-th frame in A respectively. The fundamental frequency of the frame speech signal, is the floor function, and n is the total number of speech frames in A.

[0026] 2. Model pre-training submodule, which is used to pre-train the basic acoustic model using the first training dataset, enhance the model's ability to extract common cross-language features, and introduce a joint training strategy during training to gradually transition the model to phoneme classification tasks; The basic acoustic model adopts the Transformer encoder structure and first optimizes the contrast loss , enhance the model's ability to extract cross-language common features, input the data samples in the first training dataset into the basic acoustic model in batches, and fine-tune the model based on the contrast loss calculation results; Contrastive loss The calculation formula is: , in is the cosine similarity function, is the phoneme sequence output by the basic acoustic model for the u-th input sample, Indicates the phoneme sequence actually corresponding to the u-th input sample, Indicates the phoneme sequence actually corresponding to the k-th input sample in the current training batch, is an adjustable parameter, k takes values ​​from 1 to , N is the number of samples in the current training batch, is the number of samples in the current training batch that actually correspond to the phoneme sequence of the u-th input sample.

[0027] Step S13: Introducing a joint training strategy to gradually transition the model to the phoneme classification task; When the contrast loss curve no longer has a clear downward trend, a joint training strategy is introduced to gradually transition the model to the phoneme classification task. At this time, the model needs to be further trained based on the total loss calculation results. The total loss at the current stage is The calculation formula is: , in is an adjustable weight, is the contrast loss, is an input sample calculated by dynamic programming To model output The sum of all possible alignment path probabilities, The number of phonemes replaced by the model output phoneme sequence for the actual phoneme sequence, The number of phonemes deleted from the model output phoneme sequence for the actual phoneme sequence, Output the number of newly inserted phonemes for the model's phoneme sequence to the actual phoneme. is the total number of factors in the actual phoneme sequence.

[0028] The acoustic model fine-tuning module 22 is used to freeze the bottom-level parameters of the basic acoustic model and fine-tune the top-level language-specific parameters using low-resource speech data to obtain a multilingual speech signal recognition model. It specifically includes: a freezing layer number selection submodule, a second training dataset construction submodule, and a top-level parameter adjustment submodule; 1. The freezing layer number selection submodule is used to adaptively select the freezing layer number according to the low resource data volume; Use the formula: Adaptively select the number of layers that should be frozen for the current model ,in is the total number of layers of the basic acoustic model, For low-resource voice data, The data volume of high-resource voice data.

[0029] Step S22: Processing the low-resource speech data into a second training data set; Using the aforementioned conversion function The low-resource speech data is converted into a numerical sequence as the model input, and the corresponding phoneme sequence is used as the output. After sorting, the second training dataset is obtained.

[0030] Step S23: fine-tuning the unfrozen top-level model parameters using the second training data set; Total loss during the top-level fine-tuning phase The calculation formula is expressed as: , in Is the current model for the input sample The cross entropy loss of the t-th data in , t ranges from 1 to T, and T is the input sample The total number of data items contained in , 、 is an adjustable weight, is the KL divergence calculation function, Represents the basic acoustic model for the input sample The output probability distribution of Indicates the current model for the input sample The output probability distribution of , is the j-th model parameter of the current model, yes The diagonal values ​​of the Fisher information matrix, is the jth model parameter of the basic acoustic model, j ranges from 1 to m, and m is the total number of model parameters.

[0031] During the first 1000 steps of fine-tuning training, are all set to 0, and the model's generalization ability for low-resource speech data is improved based only on the cross entropy loss. During subsequent fine-tuning training, the model is adjusted according to the actual situation. Assign values ​​to avoid overfitting and catastrophic forgetting. Fine-tuning is stopped when the total loss value stops decreasing or shows no clear downward trend, resulting in a multilingual speech signal recognition model.

[0032] The acoustic model acceleration module 23 is used to accelerate the processing of the multilingual speech signal recognition model; specifically, it includes: a model compression submodule and an input data segmentation submodule; 1. Model compression submodule, used to compress the multilingual speech signal recognition model; Through AQW quantization technology, the model can be compressed several times with almost no impact on performance.

[0033] 2. Input data segmentation submodule, which is used to segment the input data into small segments and input them into the model recognition segment by segment; Input is performed segment by segment to achieve simultaneous input and processing of voice interaction signals, saving time waiting for complete sentences and improving computing efficiency.

[0034] The signal processing task acceleration module 24 is used to split the voice signal processing task and parallelize the processing task by using the robot's hardware resources.

[0035] The acoustic model enhancement module 25 is used to adaptively optimize the multilingual speech signal recognition model based on historical human-computer interaction data; specifically, it includes: an incremental training data set generation submodule and an incremental training submodule; 1. Incremental training dataset generation submodule, used to filter samples with confidence levels below a threshold in historical human-computer interaction data to generate incremental training datasets; The confidence level is the model's predicted probability for a certain interactive speech recognition result. The threshold is the system preset value. The selected samples are labeled with the correct factor sequence and organized into a data set to obtain an incremental training data set.

[0036] 2. Incremental training submodule, used to perform incremental training on the multilingual speech signal recognition model using incremental training datasets; The incremental training process is the same as that of fine-tuning the basic acoustic model, and the bottom layer needs to be frozen first. The model parameters of the layer are then fine-tuned using the incremental training dataset to fine-tune the unfrozen model parameters of the top layer. The loss function used is also .

[0037] Corresponding to the above embodiment, an embodiment of the present invention provides a computer storage medium, comprising: at least one memory and at least one processor; The memory is used to store one or more program instructions; The processor is used to run one or more program instructions to execute a method for recognizing and processing robot voice interaction signals.

[0038] Corresponding to the above embodiment, an embodiment of the present invention provides a computer-readable storage medium, which contains one or more program instructions, and the one or more program instructions are used by a processor to execute a method for recognizing and processing robot voice interaction signals.

[0039] The embodiments disclosed in the present invention provide a computer-readable storage medium, in which computer program instructions are stored. When the computer program instructions are executed on a computer, the computer executes the above-mentioned method for recognizing and processing robot voice interaction signals.

[0040] In the embodiments of the present invention, the processor may be an integrated circuit chip having signal processing capabilities. The processor may be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0041] The methods, steps, and logic diagrams disclosed in the embodiments of the present invention can be implemented or executed. A general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of the present invention can be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules within the decoding processor. The software modules can be located in a storage medium well-established in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. The processor reads the information from the storage medium and, in conjunction with its hardware, completes the steps of the aforementioned methods.

[0042] The storage medium may be a memory and may be, for example, a volatile memory or a nonvolatile memory, or may include both volatile and nonvolatile memory.

[0043] Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory.

[0044] Volatile memory may be random access memory (RAM), which is used as an external cache memory. By way of example and not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct RAM bus random access memory (DRRAM).

[0045] The storage media described in the embodiments of the present invention are intended to include, but are not limited to, these and any other suitable types of memory.

[0046] Those skilled in the art will appreciate that in one or more of the above examples, the functions described herein can be implemented using a combination of hardware and software. When software is used, the corresponding functions can be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium. Computer-readable media include computer storage media and communication media, wherein communication media includes any medium that facilitates the transmission of computer programs from one place to another. The storage medium can be any available medium that can be accessed by a general-purpose or special-purpose computer.

[0047] The specific implementation methods described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only a specific implementation method of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solution of the present invention should be included in the scope of protection of the present invention.

Claims

1. A method for recognizing and processing robot voice interaction signals, characterized in that: include: Step 1: Train a basic acoustic model based on high-resource speech data; Step 2: Freeze the underlying parameters of the basic acoustic model and fine-tune the top-level language-specific parameters using low-resource speech data to obtain a multilingual speech signal recognition model. Step 3: Accelerate the multilingual speech signal recognition model; Step 4: Split the voice signal processing tasks and parallelize them using the robot's hardware resources. Step 5: Adaptively optimize the multilingual speech signal recognition model based on historical human-computer interaction data.

2. The method for recognizing and processing robot voice interaction signals according to claim 1, characterized in that: Training a basic acoustic model based on high-resource speech data is divided into the following sub-steps: The high-resource speech data is converted into a numerical sequence as the model input, and the corresponding phoneme sequence is used as the model output to construct a first training data set; Use the first training dataset to pre-train the basic acoustic model to enhance the model's ability to extract common cross-language features; The introduction of a joint training strategy enables the model to gradually transition to the phoneme classification task.

3. The method for recognizing and processing robot voice interaction signals according to claim 1, characterized in that: Freeze the underlying parameters of the base acoustic model and fine-tune the top-level language-specific parameters using low-resource speech data. This is done in the following sub-steps: Adaptively select the number of frozen layers based on the amount of low-resource data; Processing the low-resource speech data into a second training dataset; The unfrozen top model parameters are fine-tuned using the second training dataset.

4. The method for recognizing and processing robot voice interaction signals according to claim 1, characterized in that: Accelerate the multilingual speech signal recognition model, which is divided into the following sub-steps: Compress multilingual speech signal recognition models; Divide the input data into small segments and input them into the model recognition segment by segment.

5. The method for recognizing and processing robot voice interaction signals according to claim 1, characterized in that: Adaptively optimize the multilingual speech signal recognition model based on historical human-computer interaction data. This is divided into the following sub-steps: Filter samples with confidence levels below a threshold in historical human-computer interaction data to generate incremental training datasets; Use incremental training datasets to incrementally train multilingual speech signal recognition models.

6. A robot voice interaction signal recognition and processing system, characterized in that: include: Acoustic model preparation module, acoustic model fine-tuning module, acoustic model acceleration module, signal processing task acceleration module, acoustic model enhancement module; Acoustic model preparation module, used to train a basic acoustic model based on high-resource speech data; The acoustic model fine-tuning module freezes the underlying parameters of the basic acoustic model and uses low-resource speech data to fine-tune the top-level language-specific parameters to obtain a multilingual speech signal recognition model. Acoustic model acceleration module, used to accelerate the processing of multilingual speech signal recognition models; The signal processing task acceleration module is used to split the voice signal processing tasks and parallelize the processing tasks by utilizing the robot's hardware resources; The acoustic model enhancement module is used to adaptively optimize the multilingual speech signal recognition model based on historical human-computer interaction data.

7. A robot voice interaction signal recognition and processing system according to claim 6, characterized in that: The acoustic model preparation module specifically includes: a first training data set construction submodule and a model pre-training submodule; A first training data set construction submodule is used to convert high-resource speech data into a numerical sequence as a model input and the corresponding phoneme sequence as a model output to construct a first training data set; The model pre-training submodule is used to pre-train the basic acoustic model using the first training dataset, enhance the model's ability to extract common cross-language features, and introduce a joint training strategy during training to gradually transition the model to the phoneme classification task.

8. The robot voice interaction signal recognition and processing system according to claim 6, characterized in that: The acoustic model fine-tuning module specifically includes: a frozen layer number selection submodule, a second training dataset construction submodule, and a top-level parameter adjustment submodule; A freezing layer number selection submodule is used to adaptively select the number of freezing layers according to the amount of low-resource data; A second training data set construction submodule is used to process the low-resource speech data into a second training data set; The top-level parameter adjustment submodule is used to fine-tune the unfrozen top-level model parameters using the second training data set.

9. The robot voice interaction signal recognition and processing system according to claim 6, characterized in that: The acoustic model acceleration module specifically includes: a model compression submodule and an input data segmentation submodule; Model compression submodule, used to compress the multilingual speech signal recognition model; The input data segmentation submodule is used to divide the input data into small segments and input them into the model recognition segment by segment.

10. The robot voice interaction signal recognition and processing system according to claim 6, characterized in that: The acoustic model enhancement module specifically includes: an incremental training dataset generation submodule and an incremental training submodule; The incremental training data set generation submodule is used to filter samples with confidence levels below a threshold in the historical human-computer interaction data to generate an incremental training data set; The incremental training submodule is used to perform incremental training on the multilingual speech signal recognition model using the incremental training dataset.

Citation Information

Patent Citations

  • Voice acquisition and identification method for online incrementation

    CN104464721A

  • Airborne infrared small target detection method and device based on model migration

    CN116597325A

  • Image processing method and device, computer equipment and storage medium

    CN116842479A

  • Image reconstruction method and system based on differential output

    CN117455774A

  • Information analysis model training method and information analysis method

    CN117574981A

Cited By

  • Man-machine interaction control method and system for multi-language and dialect recognition

    CN122224146A