Speech recognition method, system, device, medium and product

By adopting intelligent dynamic time domain downsampling and asynchronous decoding mechanisms in the speech recognition system, the problem that traditional systems are difficult to ensure recognition accuracy, real-time and low power consumption on resource-constrained devices is solved, and more efficient speech recognition performance is achieved.

CN119943039AInactive Publication Date: 2025-05-06IFLYTEK CO LTD

Patent Information

Application Number
CN202510423624.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-05-06
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional voice recognition systems are difficult to ensure recognition accuracy, real-time and low power consumption at the same time on end-side devices with limited resources.

Method used

By performing intelligent dynamic time domain downsampling based on the time-frequency characteristics of voice segments, combined with the asynchronous decoding mechanism, multi-threaded asynchronous concurrent processing is realized to reduce unnecessary high-frequency sampling and computing resources, and improve recognition accuracy and real-timeness.

Benefits of technology

Under limited resource conditions, the accuracy, real-time and energy efficiency of speech recognition are improved, meeting the needs of end-side devices for high-performance speech recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119943039A_ABST
    Figure CN119943039A_ABST
Patent Text Reader

Abstract

The invention provides a voice recognition method, system and device, a medium and a product, and relates to the technical field of voice processing, and the method comprises the steps: carrying out the down-sampling of each voice segment according to the time-frequency characteristics of each voice segment in a current voice data flow, and obtaining a to-be-recognized voice sequence; encoding each data unit in the speech sequence to be recognized, and caching encoding features corresponding to the encoded data units to a target cache interval; and asynchronously loading the plurality of target coding features from the target cache interval through a decoding thread, and decoding the plurality of target coding features to obtain a real-time voice recognition result of the current voice data stream. According to the invention, voice recognition is carried out through a dynamic down-sampling and multi-thread asynchronous concurrent processing mechanism, and the recognition precision, the real-time performance and the energy efficiency can be effectively improved under the condition of effectively guaranteeing limited resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of speech processing technology, and in particular to a speech recognition method, system, device, medium and product. Background Art

[0002] Speech recognition technology provides a more convenient and natural way of interaction for human-computer interaction. It is one of the key technologies for realizing human-computer interaction and is widely used in various electronic devices.

[0003] At present, traditional speech recognition systems mostly rely on high-frequency sampling and full-stream synchronous decoding to perform speech recognition. However, on resource-constrained end-side devices, due to their limited computing power, if this method is used for speech recognition, it is difficult to ensure recognition accuracy while meeting real-time and low power consumption requirements.

[0004] Therefore, there is an urgent need to provide a speech recognition method, system, device, medium and product to solve the above technical problems. Summary of the invention

[0005] The present invention provides a speech recognition method, system, device, medium and product to solve the defect in the prior art that the traditional speech recognition system is difficult to ensure recognition accuracy while meeting the real-time and low power consumption requirements under resource-constrained scenarios, thereby improving the recognition accuracy, real-time and energy efficiency under the condition of limited resources.

[0006] The present invention provides a speech recognition method, comprising: According to the time-frequency characteristics of each speech segment in the current speech data stream, down-sampling each speech segment to obtain a speech sequence to be recognized; Encode each data unit in the speech sequence to be recognized, and cache the encoding features corresponding to the encoded data units in a target cache interval; A plurality of target coding features are asynchronously loaded from the target cache interval through a decoding thread, and the plurality of target coding features are decoded to obtain a real-time speech recognition result of the current speech data stream; the target coding features are coding features cached in the target cache interval.

[0007] According to a speech recognition method provided by the present invention, downsampling each speech segment according to the time-frequency characteristics of each speech segment in the current speech data stream to obtain a speech sequence to be recognized includes: According to the time-frequency features, obtaining the importance level of each of the speech segments; According to the importance level, obtaining a target sampling frequency for each of the speech segments; Each of the speech segments is downsampled according to the target sampling frequency to obtain the speech sequence to be recognized.

[0008] According to a speech recognition method provided by the present invention, obtaining a target sampling frequency of each speech segment according to the importance level includes: Acquire a first sampling frequency of each of the voice segments according to a voice recognition time consumption corresponding to at least one historical voice data stream; According to the importance level, obtaining a second sampling frequency of each of the voice segments; According to the first sampling frequency and the second sampling frequency, a target sampling frequency of each of the speech segments is obtained.

[0009] According to a speech recognition method provided by the present invention, the step of asynchronously loading a plurality of target coding features from the target buffer interval through a decoding thread and decoding the plurality of target coding features to obtain a real-time speech recognition result of the current speech data stream includes: Asynchronously loading a plurality of the target coding features from the target buffer interval through the decoding thread, determining coding features corresponding to each speech segment to be processed in the target speech data stream according to the plurality of the target coding features, and decoding the coding features corresponding to each speech segment to be processed according to the priority level corresponding to each speech segment to be processed, to obtain the real-time speech recognition result; The target voice data stream is formed by combining multiple data units to which the target coding features belong.

[0010] According to a speech recognition method provided by the present invention, each of the speech segments to be processed is obtained by segmenting the target speech data stream according to the semantic features and / or speech features of the target speech data stream.

[0011] According to a speech recognition method provided by the present invention, the priority level corresponding to each of the to-be-processed speech segments is determined according to the instruction type and / or importance level corresponding to each of the to-be-processed speech segments.

[0012] According to a speech recognition method provided by the present invention, the step of caching the coding features corresponding to the coded data units into a target buffer interval includes: According to the importance level of the speech sequence to be recognized and the current load information of the decoding thread, obtaining the cache capacity corresponding to the speech sequence to be recognized; The encoding features corresponding to the encoded data units are asynchronously cached to the target cache interval according to the cache capacity through the cache thread.

[0013] According to a speech recognition method provided by the present invention, encoding each data unit in the speech sequence to be recognized includes: Detecting the frequency range and energy level of background noise in each of the data units according to the time-frequency characteristics of each of the data units; Performing noise reduction processing on each of the data units according to the frequency range and the energy level; The speech features and target sampling frequency of each data unit after noise reduction processing are encoded to obtain encoding features corresponding to each data unit.

[0014] The present invention also provides a speech recognition system, comprising: A sampling unit, used for down-sampling each speech segment according to the time-frequency characteristics of each speech segment in the current speech data stream to obtain a speech sequence to be recognized; A coding processing unit, used for coding each data unit in the speech sequence to be recognized, and caching the coding features corresponding to the coded data units into a target buffer interval; A decoding processing unit is used to asynchronously load multiple target coding features from the target cache interval through a decoding thread, and decode the multiple target coding features to obtain a real-time speech recognition result of the current voice data stream; the target coding features are coding features cached in the target cache interval.

[0015] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, any of the above-mentioned speech recognition methods is implemented.

[0016] The present invention also provides a non-transitory computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the speech recognition method described in any one of the above is implemented.

[0017] The present invention also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the speech recognition method described above is implemented.

[0018] The speech recognition method, system, device, medium and product provided by the present invention can adaptively adjust the sampling frequency according to the time-frequency characteristics of each speech segment to perform intelligent dynamic time-domain down-sampling on each speech segment, which can not only reduce high-frequency sampling of unnecessary data and reduce power consumption, but also ensure that the speech sequence to be recognized obtained thereby can still accurately reflect the characteristics of the original signal, and ensure the integrity of the information of the key speech segment, avoid the loss of important information, and effectively guarantee the accuracy of speech recognition; and, by introducing an asynchronous decoding mechanism, the encoding and caching of each data unit in the speech sequence to be recognized are decoupled from the decoding, and multi-threaded asynchronous concurrent processing is realized to reduce the delay caused by long waiting, improve the real-time and response speed of speech recognition, and thus ensure that under the condition of limited resources, the recognition accuracy, real-time and energy efficiency can be effectively improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0020] Figure 1 This is one of the flow charts of the speech recognition method provided by the present invention.

[0021] Figure 2 One of the flowchart diagrams of a specific example of the speech recognition method provided by the present invention.

[0022] Figure 3 It is a schematic diagram of the process of down-sampling a speech segment provided by the present invention.

[0023] Figure 4 The second flowchart of the specific example of the speech recognition method provided by the present invention.

[0024] Figure 5 This is the second flow chart of the speech recognition method provided by the present invention.

[0025] Figure 6 The third flowchart of the specific example of the speech recognition method provided by the present invention.

[0026] Figure 7 It is a structural schematic diagram of the speech recognition system provided by the present invention.

[0027] Figure 8 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION

[0028] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0029] Speech recognition technology provides a more convenient and natural way of interaction for human-computer interaction. It is one of the key technologies for realizing human-computer interaction. It is widely used in electronic devices such as smart devices, mobile terminals, and vehicle-mounted systems.

[0030] At present, traditional speech recognition systems mostly rely on high-frequency sampling and full-stream synchronous decoding to perform speech recognition. Specifically, during the sampling process, after the traditional speech recognition system collects the speech through the microphone, it usually performs digital processing of the speech at a fixed sampling rate (such as 16kHz or 44.1kHz). Regardless of the complexity of the content or volume of the signal in the speech, all audio segments of the entire speech are processed at the same high-frequency sampling rate. As a result, even in non-critical speech segments (such as silence, environmental noise segments), the system still consumes a lot of computing resources for high-frequency sampling and processing, and cannot adaptively adjust the sampling rate according to the complexity or importance of the audio signal, resulting in resource waste and excessive system energy consumption, especially in electronic devices with limited resources (such as smartphones, smart watches or some embedded systems). During the decoding process, the traditional speech recognition system transmits the sampled audio segment to the processing unit so that the processing unit uses the speech recognition model on the terminal side to perform full-stream synchronous decoding of the sampled audio segment, that is, the system will wait until the complete audio segment is collected and encoded before unified decoding processing, and immediately return the decoded text or command to the application layer for processing. However, this synchronous decoding requires a lot of waiting time and large real-time computing resources, resulting in significant decoding delays and reduced recognition accuracy in long-duration speech or complex environments. Especially in real-time interactive scenarios, the response speed and accuracy of the speech recognition system are crucial, and synchronous decoding will lead to poor user experience and even recognition jams in high-load environments.

[0031] In summary, on resource-constrained electronic devices (such as smartphones, smart watches or some embedded systems, etc.), due to the limitations of their computing power and battery life, they rely on high-frequency sampling and full-stream synchronous decoding to perform speech recognition. It is difficult to ensure recognition accuracy while meeting real-time and low power requirements. Especially in long-duration speech, complex background noise and multi-speaker environments, traditional speech recognition systems are more prone to delays, missed recognition or excessive power consumption.

[0032] In this regard, in order to improve the performance of speech recognition under the condition of limited resources, the present application provides a speech recognition method, which performs speech recognition by combining a collaborative mode of intelligent dynamic time domain downsampling, an audio caching mechanism, and an asynchronous decoding strategy, significantly optimizing the energy efficiency, real-time performance, and accuracy of the recognition system, and can be effectively applied to speech recognition in long-term voice interactions and complex audio environments.

[0033] It should be noted that the speech recognition method provided in this application can be applied to electronic devices deployed with speech recognition systems, which can be mobile phones, tablet computers, laptops, PDAs, vehicle-mounted electronic devices, wearable devices, embedded systems or servers, etc. The speech recognition system includes multiple threads, such as a main thread, a cache thread and a decoding thread, etc., so as to realize the concurrent execution of the decoding process, audio acquisition encoding and caching through multi-threaded collaboration, and ensure that even when some audio clips have not been completely acquired and encoded, the system can decode the encoded data in advance, improve instant responsiveness and accuracy, and ensure dynamic time domain downsampling of audio data, so as to reduce unnecessary high-frequency sampling and computing resources, effectively extend the battery life of mobile devices, and reduce system power consumption. Among them, the execution subject of this method is the main thread.

[0034] Figure 1 FIG. 1 is one of the flow charts of the speech recognition method provided by the present invention. Figure 1 As shown, the method includes step 110 , step 120 and step 130 .

[0035] Step 110: downsample each speech segment according to the time-frequency characteristics of each speech segment in the current speech data stream to obtain a speech sequence to be recognized.

[0036] The current voice data stream here is the voice data stream currently required for voice recognition, which includes multiple voice segments. The voice segments can be obtained by dividing according to time windows, or by symbol positioning according to the semantics of the current voice data stream and dividing with symbols as boundaries, etc. This embodiment does not make specific limitations on this.

[0037] Optionally, when the speech recognition trigger conditions are met (such as the user clicks a speech recognition button in a graphical user interface, selects a speech recognition menu item, or enters a speech recognition command), the main thread may perform time domain feature extraction on each speech segment in the current speech data stream, where the time domain features may include energy features, etc. These time domain features can be used to detect whether the speech segment has important speech information; in addition, the main thread may also perform frequency domain feature extraction on each speech segment in the current voice data stream, where the frequency domain features may include spectrum features, etc. These frequency domain features can be used to judge the complexity of the speech signal in the speech segment.

[0038] By extracting and integrating time domain features and frequency domain features, we can construct a time-frequency feature that comprehensively distinguishes the importance of speech information and the complexity of speech signals in each speech segment from multiple levels, so as to dynamically downsample the speech segments by analyzing the time-frequency features, and integrate the sampling signals (that is, data units or tokens) in the multiple speech segments obtained by downsampling to form a speech sequence to be recognized, so as to ensure that there are audio segments with important speech information and complex speech signals, and collect data at a higher sampling rate, thereby ensuring that the downsampled speech segments can still accurately reflect the characteristics of the original signal, and ensure the integrity of the information of the key speech segments, avoid the loss of important information, and thus ensure the recognition accuracy.

[0039] For example, for complex and / or important speech segments, a higher sampling frequency is selected to retain more detailed information, and for low-energy and stable segments, a lower sampling frequency is selected to filter silent segments and reduce the loss of computing resources. For segments with simple background noise, a lower sampling frequency is selected to filter noise and reduce the loss of computing resources. Compared with the sampling method with a fixed sampling rate, the dynamic sampling method provided in this embodiment can ensure that the downsampled speech segments can still accurately reflect the characteristics of the original signal, and ensure the integrity of the information of the key speech segments, avoid the loss of important information, and thereby ensure recognition accuracy while reducing unnecessary high-frequency sampling, reducing the loss of computing resources, and effectively extending the battery life of electronic devices.

[0040] It should be noted that, in the process of dynamic downsampling of speech segments, each speech segment can be dynamically downsampled by matching the sampling frequency that matches the time-frequency features of each speech segment based on the direct correlation mapping relationship between the time-frequency features and the sampling frequency; or intermediate features, such as the importance level of the speech segment, can be first obtained based on a comprehensive evaluation of the time-frequency features, and then the sampling frequency that matches the intermediate features of each speech segment can be obtained based on the intermediate features to dynamically downsample each speech segment, etc. This embodiment does not make any specific limitations on this.

[0041] Step 120, encode each data unit in the speech sequence to be recognized, and cache the encoding features corresponding to the encoded data units in the target cache interval.

[0042] Optionally, after obtaining the speech sequence to be recognized (also called speech recognition token sequence), an encoder (such as an encoder constructed by a multi-head self-attention mechanism and a feedforward neural network, an encoder constructed based on a recurrent network, etc.) can be used to perform streaming encoding on each data unit in the speech sequence to be recognized, that is, only a single data unit is encoded in real time each time, so as to obtain the encoding features corresponding to the data unit in real time, and according to the cache capacity corresponding to the speech sequence to be recognized, the encoding features corresponding to each encoded data unit are cached in real time to the target cache interval in a target cache interval, so that subsequent decoding operations can quickly access these data.

[0043] The target cache interval here refers to the interval used to cache voice data; the cache capacity refers to the maximum amount of data that can be cached in the target cache interval for the voice sequence to be recognized, which can be dynamically adjusted according to the importance level of the voice sequence to be recognized and / or the processing efficiency of the decoding thread, so as to optimize resource utilization and improve system adaptability.

[0044] It should be noted that, during the encoding and caching process, the next encoding may be performed after the local end (main thread end) completes one encoding, that is, the local end synchronously caches the encoding features corresponding to the data unit currently encoded to the target cache interval; or, after the local end completes one encoding, the local end continues to perform the next encoding, and at the same time, asynchronously caches the encoding features corresponding to the data unit currently encoded to the target cache interval through the cache thread, etc. This embodiment does not specifically limit this.

[0045] The coding features here can be features used to characterize the essential features and semantic content of speech, including but not limited to semantic features, context-dependent features, speech features (such as phonemes, syllables, etc.) and other features that assist speech recognition, such as sampling frequency, etc. This embodiment does not specifically limit this.

[0046] Step 130, asynchronously load multiple target coding features from the target cache interval through a decoding thread, and decode the multiple target coding features to obtain a real-time speech recognition result of the current voice data stream; the target coding features are coding features cached in the target cache interval.

[0047] The decoding thread here is used to decode the encoded features output by the encoder using a decoder (such as a decoder built by a multi-head self-attention mechanism and a feedforward neural network, a decoder built based on a recurrent network, etc.) when the speech recognition task is triggered to obtain the corresponding speech recognition results.

[0048] Optionally, while caching the coding features corresponding to the encoded data units into the target cache interval, the decoding thread can be called to asynchronously stream load multiple cached coding features from the target cache interval to obtain multiple target coding features, and decode the multiple target coding features currently loaded to instantly output the real-time speech recognition results corresponding to the multiple target coding features currently loaded. That is, the encoding and decoding processes can be carried out simultaneously, and there is no need to wait until all coding features are cached before starting decoding. Compared with the synchronous decoding mode of the prior art, the processing delay is greatly reduced, the decoding rate, the real-time performance and response speed of speech recognition are improved, and it can be effectively applied to scenarios that require instant response, such as intelligent voice assistants, in-vehicle voice control, real-time command recognition scenarios, etc.

[0049] It should be noted that, during the decoding process, the multiple target coding features currently loaded may be decoded as a whole to output the speech recognition results corresponding to the multiple target coding features as a whole as the real-time speech recognition results of the current voice data stream; or, the multiple target coding features currently loaded may be decoded in segments to stream-output the speech recognition results corresponding to each segment coding feature as the real-time speech recognition results of the current voice data stream according to the priority of the segment coding, etc. This embodiment does not specifically limit this.

[0050] The method provided in this embodiment can not only reduce high-frequency sampling of unnecessary data and power consumption, but also ensure that the speech sequence to be recognized obtained thereby can still accurately reflect the characteristics of the original signal, and ensure the integrity of the information of the key speech segments, avoid the loss of important information, and effectively guarantee the accuracy of speech recognition; and, by introducing an asynchronous decoding mechanism, the encoding and caching of each data unit in the speech sequence to be recognized are decoupled from the decoding, and multi-threaded asynchronous concurrent processing is realized, so as to reduce the delay caused by long waiting time, improve the real-time and response speed of speech recognition, and thereby ensure that the recognition accuracy, real-time and energy efficiency can be effectively improved under the condition of limited resources.

[0051] Based on the speech recognition method provided in the above embodiment, a specific embodiment of the speech recognition method is given below. Figure 2 One of the flow charts of a specific example of the speech recognition method provided by the present invention; Figure 2 As shown, this example specifically includes step 210, step 220 and step 230.

[0052] Step 210: down-sample each speech segment according to the time-frequency characteristics of each speech segment in the current speech data stream to obtain a speech sequence to be recognized.

[0053] In a possible implementation, the steps of downsampling each speech segment according to the time-frequency characteristics of each speech segment in the current speech data stream to obtain the speech sequence to be recognized specifically include: obtaining the importance level of each speech segment according to the time-frequency characteristics; obtaining the target sampling frequency of each speech segment according to the importance level; and downsampling each speech segment according to the target sampling frequency to obtain the speech sequence to be recognized.

[0054] Figure 3 It is a schematic diagram of the process of down-sampling a speech segment provided by the present invention.

[0055] like Figure 3 As shown, optionally, when downsampling is performed, a speech energy analysis may be performed on each speech segment in the current speech data stream to obtain the energy characteristics of each speech segment, and a speech frequency domain analysis may be performed on each speech segment in the current speech data stream to obtain the spectrum characteristics of each speech segment, and the energy characteristics and the spectrum characteristics may be integrated to obtain the time-frequency characteristics. The spectrum characteristics may be obtained by fast Fourier transform, or may be obtained by further calculating Mel-frequency cepstrum coefficients based on fast Fourier transform, etc., which is not specifically limited in this embodiment.

[0056] Subsequently, the importance level of each speech segment is divided according to the energy features and spectrum features in the time-frequency features of each speech segment to obtain the importance level of each speech segment. For example, the energy features of each speech segment can be matched with the energy intervals of multiple energy levels to determine the energy level to which each speech segment belongs, and the spectrum features of each speech segment can be matched with the spectrum features of multiple complexity levels to determine the complexity level to which each speech segment belongs, and the importance level of each speech segment is obtained by mapping according to the correlation between the combined information between the energy level and the complexity level and the importance level. For example, the importance level of the speech segment with the highest energy level and the highest complexity level (including the important speech information part) is classified as the highest, the importance level of the speech segment with the medium energy level and the lowest complexity level (such as the transition sound or the light voice part) is classified as the second highest, and the importance level of the speech segment with the lowest energy level and the lowest complexity level (such as the background noise or the silent segment) is classified as the lowest.

[0057] Subsequently, the sampling frequency is optimized according to the importance level of each voice segment to obtain the target sampling frequency of each voice segment, thereby avoiding unnecessary high-frequency sampling of non-critical audio segments, so as to significantly reduce power consumption while ensuring efficient recognition and recognition quality, so as to be effectively suitable for mobile devices with limited batteries or smart devices that need to be used for a long time. For example, the target sampling frequency of each voice segment can be directly obtained based on the correlation mapping relationship between the importance level and the sampling frequency, or, on the basis of the importance level, combined with other performance parameters, such as the speech recognition time corresponding to the historical voice data stream, to jointly obtain the target sampling frequency of each voice segment.

[0058] In a possible implementation, the step of obtaining the target sampling frequency of each of the voice segments according to the importance level specifically includes: obtaining the first sampling frequency of each of the voice segments according to the speech recognition time corresponding to at least one historical voice data stream; obtaining the second sampling frequency of each of the voice segments according to the importance level; and obtaining the target sampling frequency of each of the voice segments according to the first sampling frequency and the second sampling frequency.

[0059] Optionally, the speech recognition time corresponding to at least one historical speech data stream can be calculated and fused to obtain the fused speech recognition time, and the first sampling frequency of each speech segment can be mapped and obtained based on the correlation between the fused speech recognition time and the sampling frequency; and the second sampling frequency of each speech segment can be mapped and obtained based on the correlation between the importance level and the sampling frequency; then, the first sampling frequency and the second sampling frequency can be fused to obtain the target sampling frequency of each speech segment. The fusion can be weighted fusion or one selected from multiple ones, etc., which is not specifically limited in this embodiment.

[0060] After obtaining the target sampling frequency, each voice segment can be dynamically downsampled in the time domain according to the target sampling frequency to ensure that a high sampling rate is maintained for recognition of key voice segments, thereby improving the accuracy of voice recognition and reducing unnecessary high-frequency sampling of non-key segments. This reduces the amount of calculation and time consumption on resource-constrained devices, effectively extending the battery life of electronic devices, and effectively solving the problem of power waste caused by traditional fixed sampling rates. While ensuring the recognition effect, it significantly improves the real-time, efficiency and energy efficiency of the system, thereby meeting the high performance requirements of terminal devices for voice recognition.

[0061] Step 220, encode each data unit in the speech sequence to be recognized, so as to cache the encoding features corresponding to the encoded data units in the target cache interval in real time.

[0062] After obtaining the speech sequence to be recognized, an encoder can be used to perform streaming encoding on each data unit in the speech sequence to be recognized, that is, only a single data unit is encoded in real time each time, so as to obtain the encoding features corresponding to the data unit in real time, and cache the encoding features corresponding to each encoded data unit in real time to the target cache interval, so that subsequent decoding operations can quickly access these data.

[0063] The target cache interval here refers to the interval used to cache voice data; the cache capacity refers to the maximum amount of data that can be cached in the target cache interval for the voice sequence to be recognized, which can be dynamically adjusted according to the importance level of the voice sequence to be recognized and / or the processing efficiency of the decoding thread, so as to optimize resource utilization and improve system adaptability.

[0064] During the encoding and caching process, the next encoding may be performed after the encoding is completed once at the local end (main thread end), that is, the encoding features corresponding to the data unit currently encoded are synchronously cached to the target cache interval at the local end; or, after the encoding is completed once at the local end, the local end continues to perform the next encoding, and at the same time, the encoding features corresponding to the data unit currently encoded are asynchronously cached to the target cache interval through the cache thread. During the encoding and caching process, this embodiment gives priority to completing the encoding once at the local end, and then the local end continues to perform the next encoding, and at the same time, the encoding features corresponding to the data unit currently encoded are asynchronously cached to the target cache interval through the cache thread. This embodiment preferably performs encoding feature caching through the cache thread asynchronous caching mechanism.

[0065] Step 230, calling the decoding thread to asynchronously load multiple target encoding features from the target buffer interval for decoding to obtain real-time speech recognition results.

[0066] Optionally, while caching the coding features corresponding to the encoded data units into the target cache interval, the decoding thread can be called to asynchronously stream load multiple cached target coding features from the target cache interval, and decode the multiple target coding features currently loaded to instantly output the real-time speech recognition results corresponding to the multiple target coding features currently loaded. That is, the decoding and encoding processes are decoupled so that the encoding and decoding processes can be carried out simultaneously without waiting for all coding features to be cached before starting decoding, which effectively reduces processing delays and improves the decoding rate and real-time performance of speech recognition.

[0067] In this embodiment, importance levels are divided according to the time-frequency characteristics of each voice segment, and the sampling frequency of each voice segment is dynamically adjusted in combination with the importance level and the voice recognition time consumption corresponding to the historical voice data stream, so as to perform intelligent dynamic time-domain down-sampling on each voice segment to ensure that a high sampling rate is maintained for recognition of key voice segments, and the sampling frequency of non-key segments is reduced, so as to improve the accuracy of voice recognition while reducing unnecessary high-frequency sampling, so as to reduce the amount of calculation and time consumption on resource-constrained devices; and, by introducing an asynchronous decoding mechanism, the encoding and caching of each data unit in the voice sequence to be recognized are decoupled from the decoding, and multi-threaded asynchronous concurrent processing is realized to reduce the delay caused by long waiting, improve the real-time and response speed of voice recognition, and further ensure that under the condition of limited resources, the recognition accuracy, real-time and energy efficiency can be effectively improved, thereby meeting the high performance requirements of the terminal device for voice recognition.

[0068] Figure 4 This is a flow chart of a specific example of the speech recognition method provided by the present invention; in order to improve the speech recognition effect, in another specific embodiment, multiple target coding features can be decoded in sections. Figure 4 As shown, this embodiment specifically includes step 410, step 420 and step 430.

[0069] Step 410: downsample each speech segment in the current speech data stream according to its time-frequency features to obtain a speech sequence to be recognized that includes downsampled signals of multiple speech segments.

[0070] Optionally, when the speech recognition trigger condition is met, the main thread may perform time domain feature extraction on each speech segment in the current speech data stream, where the time domain features may include energy features, etc., and these time domain features can be used to detect whether the speech segment contains important speech information; in addition, the main thread may also perform frequency domain feature extraction on each speech segment in the current voice data stream, where the frequency domain features may include spectrum features, etc., and these frequency domain features can be used to judge the complexity of the speech signal in the speech segment.

[0071] By extracting and integrating time domain features and frequency domain features, we can construct a time-frequency feature that comprehensively distinguishes the importance of speech information and the complexity of speech signals in each speech segment from multiple levels, so as to dynamically downsample the speech segments by analyzing the time-frequency features, and integrate the sampling signals (that is, data units or tokens) in the multiple speech segments obtained by downsampling to form a speech sequence to be recognized, so as to ensure that there are audio segments with important speech information and complex speech signals, and collect data at a higher sampling rate, thereby ensuring that the downsampled speech segments can still accurately reflect the characteristics of the original signal, and ensure the integrity of the information of the key speech segments, avoid the loss of important information, and thus ensure the recognition accuracy.

[0072] It should be noted that, in the process of dynamically downsampling the speech segments, each speech segment may be dynamically downsampled by matching and obtaining the sampling frequency that matches the time-frequency features of each speech segment based on the direct correlation mapping relationship between the time-frequency features and the sampling frequency; or by first obtaining intermediate features, such as the importance level of the speech segment, based on the comprehensive evaluation of the time-frequency features, and then matching and obtaining the sampling frequency that matches the intermediate features of each speech segment based on the correlation mapping relationship between the intermediate features and the sampling frequency, to dynamically downsample each speech segment, etc. This embodiment does not specifically limit this. This embodiment preferably uses the comprehensive evaluation of the time-frequency features to obtain the importance level, and then obtains the sampling frequency based on the importance level association to dynamically downsample each speech segment.

[0073] Step 420, encode each data unit in the speech sequence to be recognized, and cache the encoding features corresponding to the encoded data units in the target cache interval.

[0074] Optionally, after obtaining the speech sequence to be recognized, an encoder can be used to perform streaming encoding on each data unit in the speech sequence to be recognized, that is, only a single data unit is encoded in real time each time, so as to obtain the encoding features corresponding to the data unit in real time, and cache the encoding features corresponding to each encoded data unit in real time to the target cache interval, so that subsequent decoding operations can quickly access these data.

[0075] The target cache interval here refers to the interval used to cache voice data; the cache capacity refers to the maximum amount of data that can be cached in the target cache interval for the voice sequence to be recognized, which can be dynamically adjusted according to the importance level of the voice sequence to be recognized and / or the processing efficiency of the decoding thread, so as to optimize resource utilization and improve system adaptability.

[0076] During the encoding and caching process, the next encoding can be performed after the encoding is completed once at the local end (main thread end), that is, the encoding features corresponding to the data unit currently encoded are synchronously cached to the target cache interval at the local end; or, after the encoding is completed once at the local end, the local end continues to perform the next encoding, and at the same time, the encoding features corresponding to the data unit currently encoded are asynchronously cached to the target cache interval through the cache thread. During the encoding and caching process, this embodiment preferably performs the encoding and caching steps using the strategy of asynchronous execution of encoding and caching.

[0077] Step 430 , a plurality of target coding features are asynchronously loaded from a target buffer interval by a decoding thread for decoding, thereby outputting a real-time speech recognition result of the current speech data stream in real time.

[0078] In a possible implementation, a decoding thread asynchronously loads multiple target coding features from a target cache interval for decoding, thereby outputting a real-time speech recognition result of the current voice data stream in real time. Specifically, the step includes: asynchronously loading multiple target coding features from the target cache interval by the decoding thread, determining the coding features corresponding to each voice segment to be processed in the target voice data stream according to the multiple target coding features, and decoding the coding features corresponding to each voice segment to be processed according to the priority level corresponding to each voice segment to be processed to obtain the real-time speech recognition result; wherein, the target voice data stream is formed by a combination of data units to which multiple target coding features belong.

[0079] Figure 5 FIG. 2 is a flow chart of the speech recognition method provided by the present invention; Figure 5 As shown, during the decoding process, the decoding thread can be used to asynchronously execute the following steps to achieve asynchronous decoding: Load multiple target coding features that have been cached from the target cache interval, and segment the target voice data stream formed by the combination of data units to which the target coding features belong, so as to obtain multiple voice segments to be processed. The segmentation here can be implemented by segmenting according to the time boundary predicted by the information entropy window, or by segmenting according to the symbol boundary corresponding to the semantic feature of the target voice data stream and / or the silence boundary corresponding to the voice feature, etc., which is not specifically limited in this embodiment.

[0080] In a possible implementation, each of the to-be-processed speech segments is obtained by segmenting the target speech data stream according to semantic features and / or speech features of the target speech data stream.

[0081] Optionally, during the segmentation process, the symbol boundaries can be determined by locating symbols (such as periods, question marks, etc.) based on the semantic features of the target voice data stream, thereby dividing the target voice data stream into multiple different paragraphs according to the symbol boundaries, or the silence boundaries can be determined by locating silence segments (such as voice features below an energy threshold or a zero-crossing rate threshold) based on the voice features of the target voice data stream, thereby dividing the target voice data stream into multiple different paragraphs according to the silence boundaries; or, the segmentation time boundaries can be determined by combining symbol boundaries and silence boundaries to divide the target voice data stream into multiple different paragraphs, thereby accurately dividing the target voice data stream into multiple complete voice segments through the semantic features and / or voice features of the target voice data stream, so as to ensure that the voice segments are processed on demand, significantly improve recognition efficiency and accuracy, reduce system resource waste, and are effectively applicable to scenarios that require long-term continuous recognition (such as meeting records, conversation transcription, etc.), enhancing the system's adaptability to different voice scenarios.

[0082] Then, according to the correspondence between each target coding feature and each voice segment to be processed, the coding feature corresponding to each voice segment to be processed in the target voice data stream is determined from multiple target coding features, and the decoding order is dynamically adjusted according to the priority level corresponding to each voice segment to be processed, and according to the adjusted decoding order, the decoder is used to perform streaming decoding on the coding features corresponding to the voice segment to be processed, so that the speech recognition result corresponding to the decoded voice segment to be processed is output immediately as a real-time speech recognition result, thereby ensuring that the high priority level is responded to and processed quickly first, while optimizing system resource allocation, improving the real-time and efficiency of speech recognition, and providing users with a smoother interactive experience.

[0083] In a possible implementation, the priority level corresponding to each of the voice segments to be processed is determined according to the instruction type and / or importance level corresponding to each of the voice segments to be processed.

[0084] Optionally, in the process of determining the priority level corresponding to the voice segment to be processed, the priority level corresponding to each voice segment to be processed can be mapped based on the association between the importance level and the priority level, such as a high importance level corresponds to a high priority level, and a low importance level corresponds to a low priority level, so as to ensure that the voice segment to be processed with a high importance level is recognized first; or, based on the association between the instruction type and the priority level, the priority level corresponding to each voice segment to be processed can be mapped, such as a voice command and a user request correspond to a high priority level, while background noise and low importance segments correspond to a low priority level, so as to ensure that important instructions (such as voice commands, user requests) are decoded first, while background noise or low importance segments will delay decoding; or, the association between the importance level and the priority level, as well as the association between the instruction type and the priority level, are combined to map and obtain the priority level corresponding to each voice segment to be processed, thereby optimizing the allocation of decoding resources, ensuring that important instructions are responded to and processed quickly, and improving the efficiency of the speech recognition system and user experience.

[0085] In this example, by adaptively adjusting the sampling frequency according to the time-frequency characteristics of each voice segment to perform intelligent dynamic time-domain down-sampling on each voice segment, not only can high-frequency sampling of unnecessary data be reduced and power consumption be reduced, but also it can be ensured that the voice sequence to be recognized obtained thereby can still accurately reflect the characteristics of the original signal, and the integrity of the information of the key voice segment is ensured to avoid the loss of important information, thereby effectively ensuring the accuracy of voice recognition; and, by introducing an asynchronous decoding mechanism, the encoding and caching of each data unit in the voice sequence to be recognized are decoupled from the decoding, and multi-threaded asynchronous concurrent processing is realized to reduce the delay caused by long waiting, improve the real-time and response speed of voice recognition, and perform segmented decoding according to priority levels to optimize the allocation of decoding resources, ensure that important instructions are quickly responded to and processed, and improve the efficiency and user experience of the voice recognition system, thereby ensuring that under the condition of limited resources, the recognition accuracy, real-time and energy efficiency and user experience can be effectively improved.

[0086] Figure 6 The third flowchart of the specific example of the speech recognition method provided by the present invention. In order to improve the speech recognition effect, in another specific embodiment, speech recognition can be performed by combining dynamic caching and anti-noise and dynamic feature fusion mechanisms on the basis of dynamic downsampling and asynchronous coding. Figure 6 As shown, this embodiment specifically includes step 610, step 620 and step 630.

[0087] Step 610: down-sample the speech segments according to the time-frequency characteristics of each speech segment in the current speech data stream to obtain a speech sequence to be recognized.

[0088] Optionally, when the speech recognition trigger condition is met, the main thread may perform time domain feature extraction on each speech segment in the current speech data stream, where the time domain features may include energy features, etc., and these time domain features can be used to detect whether the speech segment contains important speech information; in addition, the main thread may also perform frequency domain feature extraction on each speech segment in the current voice data stream, where the frequency domain features may include spectrum features, etc., and these frequency domain features can be used to judge the complexity of the speech signal in the speech segment.

[0089] By extracting and integrating time domain features and frequency domain features, we can construct a time-frequency feature that comprehensively distinguishes the importance of speech information and the complexity of speech signals in each speech segment from multiple levels, so as to dynamically downsample the speech segments by analyzing the time-frequency features, and integrate the sampling signals (that is, data units or tokens) in the multiple speech segments obtained by downsampling to form a speech sequence to be recognized, so as to ensure that there are audio segments with important speech information and complex speech signals, and collect data at a higher sampling rate, thereby ensuring that the downsampled speech segments can still accurately reflect the characteristics of the original signal, and ensure the integrity of the information of the key speech segments, avoid the loss of important information, and thus ensure the recognition accuracy.

[0090] It should be noted that, in the process of dynamically downsampling the speech segments, the sampling frequency matching the time-frequency features of each speech segment can be matched and obtained based on the direct correlation mapping relationship between the time-frequency features and the sampling frequency to dynamically downsample each speech segment; or intermediate features such as the importance level of the speech segment are first obtained based on the comprehensive evaluation of the time-frequency features, and then each speech segment is dynamically downsampled based on the intermediate features, etc. This embodiment does not specifically limit this. This example preferably uses the comprehensive evaluation of the time-frequency features to obtain the importance level to dynamically downsample each speech segment.

[0091] Step 620, encode each data unit in the speech sequence to be recognized, and cache the encoding features corresponding to the encoded data units in the target cache interval.

[0092] In a possible implementation, the step of encoding each data unit in the speech sequence to be recognized specifically includes: detecting the frequency range and energy level of the background noise in each data unit according to the time-frequency characteristics of each data unit; performing noise reduction processing on each data unit according to the frequency range and the energy level; encoding the speech characteristics and target sampling frequency of each data unit after the noise reduction processing to obtain the coding characteristics corresponding to each data unit.

[0093] Optionally, after obtaining the speech sequence to be recognized, an encoder may be used to stream encode each data unit in the speech sequence to be recognized. Since noise interference is the main challenge to accurate speech recognition in a complex environment, in each encoding process, the time-frequency characteristics of the currently encoded data unit may be analyzed to detect the frequency range and energy level of the background noise in the data unit, and the background noise in the data unit may be denoised based on the frequency range and energy level of the background noise to effectively filter and suppress the noise. Then, on the basis of noise reduction processing, speech features are extracted from the data unit after noise reduction processing, and the target sampling frequency corresponding to the data unit is determined, and the extracted speech features and the target sampling frequency are fused and encoded. Compared with the existing system, the recognition accuracy in noisy environment is reduced. In this embodiment, a dynamic feature fusion mechanism of anti-noise, speech features and target sampling frequency is used to effectively filter and suppress background noise, enhance noise resistance, enhance speech clarity and recognizability, and more comprehensively retain key information of speech signals, reduce information loss, thereby maintaining recognition accuracy in a variable noise environment, that is, enhancing recognition accuracy in a low signal-to-noise ratio environment, further improving recognition accuracy in complex backgrounds, and being able to effectively adapt to speech recognition in complex scenarios such as vehicle-mounted systems and speech recognition in public places, thereby improving speech recognition adaptability.

[0094] In one possible implementation, the step of caching the coding features corresponding to the encoded data units into the target cache interval specifically includes: obtaining the cache capacity corresponding to the speech sequence to be recognized according to the importance level of the speech sequence to be recognized and the current load information of the decoding thread; and asynchronously caching the coding features corresponding to the encoded data units into the target cache interval according to the cache capacity through the cache thread.

[0095] Optionally, considering that in the prior art, a cache mechanism is introduced when processing long-duration speech, the audio stream is temporarily stored to ensure processing stability. However, these cache mechanisms are often fixed-length buffers, and the buffer size is dynamically adjusted without considering the importance of the audio segment or the current system's computing resource status, resulting in poor performance in complex scenarios. In particular, when long-duration speech or continuous conversations need to be processed, the buffer size and processing strategy cannot be flexibly adjusted, which is prone to delays and freezes, resulting in reduced system performance.

[0096] To solve the above problems, in this embodiment, the buffer size is dynamically adjusted according to the importance level and the current load information of the decoding thread, and asynchronous caching is combined for voice processing to ensure processing stability while optimizing system resource allocation and improving adaptability and overall performance in complex scenarios.

[0097] Specifically, in the process of encoding feature caching, the cache capacity of the speech sequence to be recognized can be dynamically adjusted according to the importance level of the speech sequence to be recognized and the current load information of the decoding thread. For example, the cache capacity of the speech sequence to be recognized can be obtained by mapping the associated mapping relationship between the combined encoding information between the importance level and the current load information and the cache capacity. For example, for speech sequences with high importance levels, a larger cache capacity is allocated to ensure that these key speech data can be processed and recognized in a timely manner. For speech sequences with low importance levels, the cache capacity is dynamically adjusted according to the load of the decoding thread to balance the use of system resources.

[0098] Then, after obtaining the cache capacity corresponding to the speech sequence to be recognized, the cache thread can be called to asynchronously cache the encoding features corresponding to the encoded data unit to the target cache interval according to the cache capacity. That is, while the main thread is encoding, the cache thread can asynchronously cache the encoding features corresponding to the data unit encoded by the main thread to the target cache interval, thereby improving the concurrent processing capability and response speed of the system. Therefore, by dynamically adjusting the cache capacity according to the importance level of the speech sequence and the system load, and using the asynchronous cache thread to store the encoding features in the target interval, frame loss due to system busyness is avoided, which effectively improves the efficiency, real-time and system stability of speech recognition, while optimizing resource utilization, and can be effectively applied to scenarios that require long-term continuous recognition (such as meeting records, conversation transcription, etc.), thereby enhancing the adaptability and overall performance of the system in complex scenarios.

[0099] Step 630 , asynchronously load multiple target coding features from the target buffer interval through a decoding thread and decode them to obtain a real-time speech recognition result.

[0100] Optionally, while caching the coding features corresponding to the encoded data units into the target cache interval, the decoding thread can be called to asynchronously stream load multiple cached target coding features from the target cache interval, and decode the multiple target coding features currently loaded to instantly output the real-time speech recognition results corresponding to the multiple target coding features currently loaded. That is, the encoding and decoding processes can be carried out simultaneously, and there is no need to wait until all coding features are cached before starting decoding, which can improve the decoding rate and the real-time performance of speech recognition.

[0101] In this embodiment, each voice segment is intelligently and dynamically down-sampled in the time domain according to the time-frequency characteristics of each voice segment to ensure that a high sampling rate is maintained for recognition in key voice segments, thereby improving the accuracy of voice recognition and reducing unnecessary high-frequency sampling, so as to reduce the amount of calculation and time consumption on resource-constrained devices; and, by introducing dynamic cache capacity adjustment, asynchronous caching and asynchronous decoding mechanisms, the encoding, caching and decoding of each data unit in the voice sequence to be recognized are decoupled to achieve multi-threaded asynchronous concurrent processing, so as to reduce the delay caused by long waiting time, improve the real-time and response speed of voice recognition, and thus ensure that under the condition of limited resources, the recognition accuracy, real-time and energy efficiency can be effectively improved, thereby meeting the high performance requirements of the terminal device for voice recognition.

[0102] The speech recognition system provided by the present invention is described below. The speech recognition system described below and the speech recognition method described above can be referenced to each other.

[0103] Figure 7 Schematic diagram of the structure of the speech recognition system provided by the present invention; Figure 7 As shown, the system includes: The sampling unit 710 is used to downsample each speech segment according to the time-frequency characteristics of each speech segment in the current speech data stream to obtain a speech sequence to be recognized; The encoding processing unit 720 is used to encode each data unit in the speech sequence to be recognized, and cache the encoding features corresponding to the encoded data unit in the target cache interval; The decoding processing unit 730 is used to asynchronously load multiple target coding features from the target cache interval through a decoding thread, and decode the multiple target coding features to obtain real-time speech recognition results of the current voice data stream; the target coding features are coding features cached in the target cache interval.

[0104] The system provided by the present invention can perform intelligent dynamic time-domain down-sampling on each voice segment by adaptively adjusting the sampling frequency according to the time-frequency characteristics of each voice segment, thereby reducing high-frequency sampling of unnecessary data and reducing power consumption, and can also ensure that the voice sequence to be recognized obtained thereby can still accurately reflect the characteristics of the original signal, and ensure the integrity of the information of the key voice segment, avoid the loss of important information, and effectively guarantee the accuracy of voice recognition; and, by introducing an asynchronous decoding mechanism, the encoding and caching of each data unit in the voice sequence to be recognized are decoupled from the decoding, and multi-threaded asynchronous concurrent processing is realized, so as to reduce the delay caused by long waiting, improve the real-time and response speed of voice recognition, and further ensure that under the condition of limited resources, the recognition accuracy, real-time and energy efficiency can be effectively improved.

[0105] The system provided by the present invention is used to execute the above-mentioned method embodiments. Please refer to the above-mentioned embodiments for the specific processes and detailed contents, which will not be repeated here.

[0106] Figure 8 An example of a physical structure diagram of an electronic device is shown in FIG. Figure 8 As shown, the electronic device may include: a processor 810, a communication interface 820, a memory 830 and a communication bus 840, wherein the processor 810, the communication interface 820 and the memory 830 communicate with each other through the communication bus 840. The processor 810 may call the logic instructions in the memory 830 to execute the speech recognition method, which includes: according to the time-frequency characteristics of each speech segment in the current speech data stream, down-sampling each speech segment to obtain a speech sequence to be recognized; encoding each data unit in the speech sequence to be recognized, and caching the encoding features corresponding to the encoded data unit into the target cache interval; asynchronously loading multiple target encoding features from the target cache interval through a decoding thread, and decoding the multiple target encoding features to obtain the real-time speech recognition result of the current speech data stream; the target encoding feature is the encoding feature cached in the target cache interval.

[0107] In addition, the logic instructions in the above-mentioned memory 830 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art or the part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program codes.

[0108] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the speech recognition method provided by the above-mentioned methods, which includes: downsampling each speech segment according to the time-frequency characteristics of each speech segment in the current speech data stream to obtain a speech sequence to be recognized; encoding each data unit in the speech sequence to be recognized, and caching the coding features corresponding to the encoded data units into a target cache interval; asynchronously loading multiple target coding features from the target cache interval through a decoding thread, and decoding the multiple target coding features to obtain a real-time speech recognition result of the current speech data stream; the target coding features are coding features cached in the target cache interval.

[0109] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the speech recognition method provided by the above-mentioned methods, the method comprising: downsampling each speech segment in the current speech data stream according to the time-frequency characteristics of each speech segment, to obtain a speech sequence to be recognized; encoding each data unit in the speech sequence to be recognized, and caching the coding features corresponding to the encoded data units into a target cache interval; asynchronously loading multiple target coding features from the target cache interval through a decoding thread, and decoding the multiple target coding features to obtain a real-time speech recognition result of the current speech data stream; the target coding features are coding features cached in the target cache interval.

[0110] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.

[0111] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0112] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A speech recognition method, characterized in that: include: According to the time-frequency characteristics of each speech segment in the current speech data stream, down-sampling each speech segment to obtain a speech sequence to be recognized; Encode each data unit in the speech sequence to be recognized, and cache the encoding features corresponding to the encoded data units in a target cache interval; A plurality of target coding features are asynchronously loaded from the target cache interval through a decoding thread, and the plurality of target coding features are decoded to obtain a real-time speech recognition result of the current speech data stream; the target coding features are coding features cached in the target cache interval.

2. The speech recognition method according to claim 1, characterized in that: The method of downsampling each voice segment according to the time-frequency characteristics of each voice segment in the current voice data stream to obtain a voice sequence to be recognized includes: According to the time-frequency features, obtaining the importance level of each of the speech segments; According to the importance level, obtaining a target sampling frequency for each of the speech segments; Each of the speech segments is downsampled according to the target sampling frequency to obtain the speech sequence to be recognized.

3. The speech recognition method according to claim 2, characterized in that: The step of obtaining a target sampling frequency for each of the voice segments according to the importance level includes: Acquire a first sampling frequency of each of the voice segments according to a voice recognition time consumption corresponding to at least one historical voice data stream; According to the importance level, obtaining a second sampling frequency of each of the voice segments; According to the first sampling frequency and the second sampling frequency, a target sampling frequency of each of the speech segments is obtained.

4. The speech recognition method according to any one of claims 1 to 3, characterized in that: The step of asynchronously loading a plurality of target coding features from the target buffer interval through a decoding thread and decoding the plurality of target coding features to obtain a real-time speech recognition result of the current speech data stream includes: Asynchronously loading a plurality of the target coding features from the target buffer interval through the decoding thread, determining coding features corresponding to each speech segment to be processed in the target speech data stream according to the plurality of the target coding features, and decoding the coding features corresponding to each speech segment to be processed according to the priority level corresponding to each speech segment to be processed, to obtain the real-time speech recognition result; The target voice data stream is formed by combining multiple data units to which the target coding features belong.

5. The speech recognition method according to claim 4, characterized in that: Each of the to-be-processed speech segments is obtained by segmenting the target speech data stream according to the semantic features and / or speech features of the target speech data stream.

6. The speech recognition method according to claim 4, characterized in that: The priority level corresponding to each of the voice segments to be processed is determined according to the instruction type and / or importance level corresponding to each of the voice segments to be processed.

7. The speech recognition method according to any one of claims 1 to 3, characterized in that: The step of caching the encoding features corresponding to the encoded data units into a target cache interval includes: According to the importance level of the speech sequence to be recognized and the current load information of the decoding thread, obtaining the cache capacity corresponding to the speech sequence to be recognized; The encoding features corresponding to the encoded data units are asynchronously cached to the target cache interval according to the cache capacity through the cache thread.

8. The speech recognition method according to any one of claims 1 to 3, characterized in that: The step of encoding each data unit in the to-be-recognized speech sequence comprises: Detecting the frequency range and energy level of background noise in each of the data units according to the time-frequency characteristics of each of the data units; Performing noise reduction processing on each of the data units according to the frequency range and the energy level; The speech features and target sampling frequency of each data unit after noise reduction processing are encoded to obtain encoding features corresponding to each data unit.

9. A speech recognition system, characterized in that: include: A sampling unit, used for down-sampling each speech segment according to the time-frequency characteristics of each speech segment in the current speech data stream to obtain a speech sequence to be recognized; A coding processing unit, used for coding each data unit in the speech sequence to be recognized, and caching the coding features corresponding to the coded data units into a target buffer interval; A decoding processing unit is used to asynchronously load multiple target coding features from the target cache interval through a decoding thread, and decode the multiple target coding features to obtain a real-time speech recognition result of the current voice data stream; the target coding features are coding features cached in the target cache interval.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the speech recognition method according to any one of claims 1 to 8 is implemented.

11. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the speech recognition method according to any one of claims 1 to 8 is implemented.

12. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the speech recognition method according to any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Adaptive 3D video coding and decoding method based on compressed sensing

    CN107509074A

  • Voice processing method and device, electronic equipment and storage medium

    CN111402908A

  • Method and apparatus for processing command audio for virtual personal assistant

    CN117636844A

  • Video coding method and device, computer equipment and storage medium

    CN117979085A

  • Frame preemption method and device, frame verification method and device, medium and electronic equipment

    CN118301104A

Cited By

  • Real-time structured extraction method and device for streaming voice

    CN121789687A