Voice Data Processing Method, Apparatus, Electronic Device, and Storage Medium

By framed and time stamping the voice data stream, the user experience problem caused by untimely data processing in voice wake-up technology is solved, and the effect of fast processing of voice data lag is achieved, and the system robustness is enhanced.

CN114822586BActive Publication Date: 2025-06-27MIDEA GRP (SHANGHAI) CO LTD +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210449017.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-26
Publication Date
2025-06-27
Estimated Expiration
2042-04-26

AI Technical Summary

Technical Problem

In voice wake-up technology, due to cost limitations, local equipment lacks computing power and limited memory, resulting in untimely processing of voice data and data accumulation. The wake-up word judgment module cannot process real-time data in a timely manner, which makes it take several seconds for the user to get a response after saying the wake-up word, affecting the user experience.

Method used

By performing frame-based processing of the voice data stream and timestamping each voice data frame, extracting the time stamp of the target voice data frame in the voice processing module, and calculating the time difference between the timestamp and the current system time. If the time difference is greater than the preset threshold, the data processing process is judged and the resource scheduling strategy is determined according to the processing steps of the voice processing module, and the voice data stream is processed first.

Benefits of technology

It can detect abnormalities in voice data processing in a timely manner, quickly handle voice data lag problems, enhance system robustness, and improve user voice wake-up interactive experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114822586B_ABST
    Figure CN114822586B_ABST
Patent Text Reader

Abstract

The present application relates to the field of data processing, and provides a method, apparatus, electronic device, and storage medium for processing voice data. The method includes: performing frame division processing on a voice data stream through a first voice processing module, and marking a timestamp for each voice data frame; extracting the timestamp of a target voice data frame in the voice processing module; if the time difference between the timestamp and the current system time is greater than a preset threshold, determining that an abnormality has occurred in the data processing process of the voice data stream. The method of the present application can timely detect an abnormality in voice data processing by extracting the timestamp corresponding to the voice data frame in the voice processing module, calculating the time difference between the timestamp and the current system time, and comparing the time difference with a set threshold to determine whether there is a blockage of the voice data stream in the current system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data processing, and in particular, to a method, apparatus, electronic device, and storage medium for processing voice data. Background Art

[0002] Voice wake-up technology pre-sets a wake-up word in a device or software. When the user issues this voice command, the device is awakened from the sleep state and makes a specified response, greatly improving the efficiency of human-computer interaction. To protect user privacy, voice data cannot be uploaded before the device is awakened. Therefore, voice wake-up often needs to be implemented on a local device.

[0003] Before performing voice wake-up, multiple voice signal processing processes are required, including voice acquisition, voice preprocessing, voice endpoint detection, wake-up word judgment, and finally obtaining a wake-up result. In this pipeline, each processing module processes voice data frame by frame, and voice data is transmitted between different modules through a buffer.

[0004] Limited by cost, local devices often have insufficient computing power and limited memory. In some special cases, data cannot be processed in time, resulting in data accumulation. The wake-up word judgment module cannot process real-time data in time, causing the user to wait several seconds to get a response after saying the wake-up word, resulting in a poor actual use experience.

[0005] Therefore, in the voice wake-up process, how to timely detect the data flow blockage situation in the voice wake-up process has become the key to affecting system stability. Summary of the Invention

[0006] This application aims to at least solve one of the technical problems existing in the related art. For this purpose, this application proposes a method for processing voice data, which can timely detect the data flow blockage situation in the voice wake-up process.

[0007] This application also proposes a voice data processing apparatus.

[0008] This application also proposes an electronic device.

[0009] This application also proposes a storage medium.

[0010] This application also proposes a computer program product.

[0011] According to the voice data processing method of the first aspect embodiment of this application, it includes:

[0012] Performing frame division processing on a voice data stream through a first voice processing module, and performing timestamp marking on each voice data frame obtained after the frame division processing according to the current system time;

[0013] Extract the timestamp of the target voice data frame in the voice processing module;

[0014] If the time difference between the timestamp and the current system time is greater than a preset threshold, it is determined that an abnormality has occurred in the data processing process of the voice data stream;

[0015] Wherein, the target voice data frame is one of the voice data frames obtained after the voice data stream is frame-processed.

[0016] According to the voice data processing method of the embodiments of the present application, by extracting the timestamp corresponding to the voice data frame in the voice processing module and calculating the time difference between the timestamp and the current system time, and comparing the time difference with the set threshold to determine whether the voice data stream is blocked in the current system, it is possible to detect abnormalities in voice data processing in a timely manner.

[0017] According to an embodiment of the present application, the data processing process of the voice data stream sequentially includes a plurality of processing steps, there are a plurality of voice processing modules, and the plurality of processing steps and the plurality of voice processing modules have a one-to-one correspondence;

[0018] Determining that an abnormality has occurred in the data processing process of the voice data stream when the time difference between the timestamp and the current system time is greater than a preset threshold includes:

[0019] Determine a first threshold corresponding to the voice processing module according to the order of the processing steps corresponding to the voice processing module in the data processing process;

[0020] If the time difference between the timestamp and the current system time is greater than the first threshold, it is determined that an abnormality has occurred in the data processing process of the voice data stream.

[0021] According to the voice data processing method of the embodiments of the present application, the voice data frames of the voice data stream are processed by a plurality of sequentially arranged voice processing modules respectively. Since the voice data frames have different time delays when flowing to different voice processing modules, the corresponding judgment thresholds are determined according to the links where the voice processing modules are located in the entire data processing process, and then it is judged whether the time delay of the voice data frames in each voice processing module is too large, which can improve the accuracy of monitoring the transmission or processing time delay of the voice data frames.

[0022] According to an embodiment of the present application, after determining that an abnormality has occurred in the data processing process of the voice data stream when the time difference between the timestamp and the current system time is greater than a preset threshold, it further includes:

[0023] Determine a resource scheduling strategy according to the order of the processing steps corresponding to the voice processing module in the data processing process;

[0024] Execute the resource scheduling policy so that multiple said voice processing modules preferentially process voice data frames of the voice data stream.

[0025] According to the voice data processing method of the embodiments of the present application, by specifically determining the resource scheduling policy based on the order of the steps corresponding to the voice processing modules in the entire processing process, it is possible to quickly perform emergency processing on the currently processed voice data stream, thereby being able to promptly discover and handle the problem of voice data lag, enhancing the system robustness.

[0026] According to an embodiment of the present application, the first voice processing module is the voice processing module corresponding to the starting step of the data processing process.

[0027] According to the voice data processing method of the embodiments of the present application, by using the voice processing module corresponding to the starting step of the data processing process to perform frame splitting on the voice data stream and attach a timestamp identifier to each voice data frame, and determining the timestamp of the data frame by the source closest to the voice data acquisition, it is possible to further improve the accuracy of monitoring the transmission or processing delay of voice data frames.

[0028] According to an embodiment of the present application, extracting the timestamp of the target voice data frame in the voice processing module includes:

[0029] Determine that the voice processing module has not performed data processing on the target voice data frame, and extract the timestamp of the target voice data frame in the voice processing module.

[0030] According to the voice data processing method of the embodiments of the present application, for each voice processing module, before it processes the voice data frame, that is, extract the timestamp of the voice data frame for delay judgment, so as to be able to more promptly discover the data lag problem in the voice data processing process.

[0031] According to an embodiment of the present application, the multiple processing steps sequentially include a voice acquisition step, a voice preprocessing step, a voice endpoint detection step, and a voice data judgment step.

[0032] According to the voice data processing method of the embodiments of the present application, the processing of voice data including voice wake-up data is respectively realized through the voice acquisition step, the voice preprocessing step, the voice endpoint detection step, and the voice data judgment step, further improving the efficiency and accuracy of voice data processing.

[0033] According to an embodiment of the present application, the voice data processing method is applied to an edge computing end corresponding to a voice processing system.

[0034] According to the voice data processing method of the embodiments of the present application, this method is implemented through an edge computing terminal, avoiding additional computing overhead on the local devices of the voice processing system, so as not to affect the operation of the original voice processing system while monitoring the abnormal processing of voice data, and improving the robustness of the voice processing system.

[0035] The voice data processing device according to the second aspect embodiment of the present application includes:

[0036] A marking module, configured to perform frame segmentation processing on a voice data stream through a first voice processing module, and perform timestamp marking on each voice data frame obtained after the frame segmentation processing according to the current system time;

[0037] An extraction module, configured to extract the timestamp of a target voice data frame in the voice processing module;

[0038] A determination module, configured to determine that an abnormality occurs in the data processing process of the voice data stream when the time difference between the timestamp and the current system time is greater than a preset threshold;

[0039] Wherein, the target voice data frame is one of the voice data frames obtained after the voice data stream is subjected to frame segmentation processing.

[0040] The electronic device according to the third aspect embodiment of the present application includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the voice data processing method as described in any one of the above.

[0041] The non-transitory computer-readable storage medium according to the fourth aspect embodiment of the present application stores a computer program thereon. When the computer program is executed by a processor, it implements the voice data processing method as described in any one of the above.

[0042] The computer program product according to the fifth aspect embodiment of the present application includes a computer program. When the computer program is executed by a processor, it implements the voice data processing method as described in any one of the above.

[0043] One or more of the above technical solutions in the embodiments of the present application have at least one of the following technical effects:

[0044] By extracting the timestamp corresponding to the voice data frame in the voice processing module and calculating the time difference between the timestamp and the current system time, and comparing the time difference with a set threshold to determine whether the voice data stream is blocked in the current system, the abnormality of voice data processing can be detected in a timely manner.

[0045] The additional aspects and advantages of the present application will be partially given in the following description, partially become obvious from the following description, or be understood through the practice of the present application. Description of the Drawings

[0046] To more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the accompanying drawings required for the description of the embodiments or related technologies. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.

[0047] Figure 1 is a schematic flowchart of the voice data processing method provided by the embodiments of the present application;

[0048] Figure 2 is a schematic flowchart of the voice wake-up result acquisition process provided by the embodiments of the present application;

[0049] Figure 3 is a schematic flowchart of the timestamp processing process provided by the embodiments of the present application;

[0050] Figure 4 is a schematic structural diagram of the voice data processing device provided by the embodiments of the present application;

[0051] Figure 5 is a schematic structural diagram of the electronic device provided by the embodiments of the present application. Detailed implementation manners

[0052] The following will further describe in detail the implementation manners of the present application in conjunction with the accompanying drawings and embodiments. The following embodiments are used to illustrate the present application, but cannot be used to limit the scope of the present application.

[0053] Figure 1 is one of the schematic flowcharts of the voice data processing method provided by the embodiments of the present application. Referring to Figure 1 , the embodiments of the present application provide a voice data processing method, which may include the steps:

[0054] S1. Perform frame segmentation processing on the voice data stream through the first voice processing module, and perform timestamp marking on each voice data frame obtained after the frame segmentation processing according to the current system time.

[0055] In the embodiments of the present application, the processing process of the voice data stream can be completed sequentially by multiple voice processing modules. Step S1 is to perform timestamping on each frame of voice data frame by the first voice processing module according to the current system time; among them, the first voice processing module can be one of the voice processing modules closer to the source of the voice data.

[0056] S2. Extract the timestamp of the target voice data frame in the voice processing module. Among them, the target voice data frame is one of the voice data frames obtained after the voice data stream is frame-segmented.

[0057] It should be noted that voice signals have strong varying characteristics. During the process of a person speaking, different mouth shapes, breaths, and pronunciation positions will generate different voice signals, and the duration of each voice signal is very short. Therefore, voice signals have very strong time-varying characteristics. From the overall perspective of the voice waveform diagram, it can be seen that voice signals change rapidly over time. As mentioned above, the changes in voice signals rely on different mouth shapes, breaths, and pronunciation positions, and the changes in these parts of a person are relatively slow compared to the vibration speed of sound. Through the voice waveform, it can be seen that the voice changes within 20 ms are very small compared to the overall voice changes. Therefore, the analysis and processing of short-time voice signals can be carried out in the way of processing stationary process data. The time during which a voice signal can remain relatively stable is from 10 ms to 30 ms. Therefore, the processing and analysis of voice signals usually adopt the short-time analysis method, that is, the so-called frame division operation (frame division processing).

[0058] In the embodiment of the present application, first, a frame division operation is performed on the voice data stream to be processed, and data processing and transmission are carried out in units of voice data frames. Overall, the voice processing system can be composed of one or more voice processing modules. For each voice processing module, after obtaining a complete frame of voice data, it is processed and then output after the processing is completed. For a certain voice data frame in the voice processing module, the timestamp identifier carried by it can be extracted, where the timestamp can be marked according to the system time by this voice processing module or the processing module before this voice processing module.

[0059] S3. If the time difference between the timestamp and the current system time is greater than a preset threshold, it is determined that an abnormality has occurred in the data processing process of the voice data stream;

[0060] It should be noted that due to cost limitations, local devices often have insufficient computing power and limited memory, and in some special cases, they cannot process data in a timely manner, resulting in data accumulation. Therefore, exploring how to timely identify data stream blockages, detect system abnormalities, and avoid deterioration of the actual user experience has become the key to affecting system stability.

[0061] In the embodiment of the present application, after the timestamp of the target voice data frame is extracted, the time difference between the timestamp and the current system time can be calculated, and then the time difference can be compared with the preset threshold to determine whether an abnormality has occurred in the processing process of the entire voice data stream. For example, when it is determined that the time difference is greater than the preset threshold, it means that the time delay of the voice data frame from the moment when the data is initially obtained (the moment when the timestamp is marked) to the current moment is relatively long, and it can be considered that an abnormality has occurred in the processing process of the voice data stream corresponding to the voice data frame (indicating data accumulation in the system).

[0062] According to the voice data processing method of the embodiment of the present application, by extracting the timestamp corresponding to the voice data frame processed by the voice processing module and calculating the time difference between the timestamp and the current system time, it is determined whether the voice data flow is blocked in the current system by comparing the time difference with the set threshold, thereby enabling timely detection of abnormalities in voice data processing.

[0063] In one embodiment, the data processing process of the voice data stream includes a plurality of processing steps in sequence, there are a plurality of voice processing modules, and the plurality of processing steps and the plurality of voice processing modules are in a one-to-one correspondence;

[0064] Step S3 may include the steps of:

[0065] S31, determining a first threshold corresponding to the speech processing module according to the order of the processing steps corresponding to the speech processing module in the data processing process;

[0066] S32: If the time difference between the timestamp and the current system time is greater than the first threshold, it is determined that an abnormality occurs in the data processing process of the voice data stream.

[0067] In an embodiment of the present application, the voice data stream is processed in sequence by multiple voice processing modules. For example, the processing of the voice data stream requires five steps, and a voice processing module is used to process each step. Taking voice wake-up data as an example, before responding to the voice wake-up of the device, it is necessary to go through the steps of voice acquisition, voice preprocessing, voice endpoint detection, wake-up word judgment, and finally obtain the wake-up result. In this pipeline, each processing module will process the voice data frame by frame, and different modules will transmit voice data through cache.

[0068] It is understandable that for the same voice data frame, its timestamp is fixed after being marked, and the voice data frame has different delays when reaching different voice processing modules. Therefore, the link in the entire voice data stream processing process can be determined according to the processing steps corresponding to the voice processing module, and the corresponding first threshold can be determined. Then, according to the first threshold, it can be judged whether the voice data frame has too much delay when passing through the voice processing module. The data transmission or processing delay can be judged in a targeted manner according to different voice processing modules. Specifically, if the time difference between the timestamp and the current system time is greater than the threshold, it is considered that the data processing process is abnormal, and if the time difference between the timestamp and the current system time is not greater than the threshold, it is considered that the data processing process is normal.

[0069] According to the voice data processing method of the embodiment of the present application, the voice data frames of the voice data stream are processed respectively by multiple voice processing modules arranged in sequence. Since the voice data frames have different delays when flowing to different voice processing modules, the corresponding judgment threshold is determined by the link where the voice processing module is located in the entire data processing process, and then it is determined whether the voice data frame has a too large delay in each voice processing module, which can further improve the accuracy of voice data frame transmission or processing delay monitoring.

[0070] In one embodiment, after step S3, the method may further include the following steps:

[0071] S4, determining a resource scheduling strategy according to the order of the processing steps corresponding to the voice processing module in the data processing process;

[0072] S5. Execute the resource scheduling strategy to enable the plurality of voice processing modules to preferentially process the voice data frames of the voice data stream.

[0073] It should be noted that in the voice data processing pipeline, each voice processing module will process the voice data frame by frame, and different modules will transmit voice data through cache. Due to cost constraints, local devices often have insufficient computing power and limited memory. In some special cases, data cannot be processed in time, resulting in data accumulation. For example, the wake-up word judgment module cannot obtain and process voice data in time, resulting in users having to wait several seconds to get a response after saying the wake-up word, which deteriorates the actual user experience.

[0074] In an embodiment of the present application, when it is determined that an abnormality has occurred in the data processing process of a voice data stream, based on the voice processing module that has determined that the time difference is greater than a threshold, a resource scheduling strategy can be formulated in a targeted manner according to the processing steps corresponding to the voice processing module in the entire voice processing process. For example, system computing resources are allocated preferentially to the voice processing module and subsequent voice processing modules, or other tasks with higher priorities are suspended, and the tasks corresponding to the current voice data stream are adjusted to a higher priority, so as to maximize the processing efficiency of voice data frames in the voice processing module and subsequent modules, and overcome the problem of abnormal data processing of the voice data stream.

[0075] According to the voice data processing method of the embodiment of the present application, the resource scheduling strategy is determined in a targeted manner through the order of the steps corresponding to the voice processing module in the entire processing process, and the currently processed voice data stream can be quickly processed urgently, so that the problem of voice data lag can be discovered and processed in time, thereby enhancing the robustness of the system.

[0076] According to an embodiment of the present application, the first voice processing module is a voice processing module corresponding to the starting step of the data processing process.

[0077] In the embodiment of the present application, for the first voice processing module corresponding to the starting step of the data processing process, since this first voice processing module is closest to the source of the data, the first voice processing module is used to perform voice framing processing, and a timestamp identifier is added to each voice data frame according to the system time at this time.

[0078] According to the voice data processing method of the embodiment of the present application, by using the voice processing module corresponding to the starting step of the data processing process to perform framing processing on the voice data stream and adding a timestamp identifier to each voice data frame, and determining the timestamp of the data frame by the source closest to the acquisition of the voice data, the accuracy of monitoring the transmission or processing delay of the voice data frame can be further improved.

[0079] In one embodiment, step S2 may include the steps of:

[0080] S21. Determine that the voice processing module has not processed the target voice data frame, and extract the timestamp of the target voice data frame in the voice processing module.

[0081] It should be noted that the operation of extracting the timestamp of the target voice data frame in the voice processing module can be performed at any time when the target voice data frame is in the voice processing module. In the embodiment of the present application, further, before the voice processing module starts to process the target voice data frame, that is, the timestamp of the target voice data frame is first extracted, and subsequent data delay judgment is performed, so that it can be more timely to find whether the delay of the target voice data frame is too large during the process of flowing from the previous voice processing module to the current voice processing module, and thus it can be more timely to find the data lag problem in the voice data processing process.

[0082] According to the voice data processing method of the embodiment of the present application, for each voice processing module, before it processes the voice data frame, that is, the timestamp of the voice data frame is extracted for delay judgment, so that it can be more timely to find the data lag problem in the voice data processing process, and further improve the timeliness and accuracy of data delay judgment.

[0083] In one embodiment, the multiple processing steps sequentially include a voice acquisition step, a voice preprocessing step, a voice endpoint detection step, and a voice data judgment step.

[0084] It should be noted that for the processing of voice data such as voice wake-up data, steps such as voice acquisition, voice preprocessing, voice endpoint detection, and wake-up word judgment are required. Specifically, the voice processing module corresponding to the voice acquisition step can be used to acquire voice data, perform frame division operations, and mark timestamps; the voice processing module corresponding to the voice preprocessing step can be used to perform filtering, noise reduction, and other processing on the voice data; the voice processing module corresponding to the voice endpoint detection step can be used to perform endpoint detection on the voice data; the voice processing module corresponding to the voice data judgment step can be used to perform operations such as voice recognition and data judgment on the voice data.

[0085] According to the voice data processing method of the embodiments of the present application, the processing of voice data including voice wake-up data is realized through the voice acquisition step, the voice preprocessing step, the voice endpoint detection step, and the voice data judgment step respectively, realizing the division of labor and collaborative processing of voice wake-up data and the like, and further improving the efficiency and accuracy of voice data processing.

[0086] In one embodiment, the voice data processing method is applied to an edge computing end corresponding to a voice processing system; the voice processing system is used to execute the data processing process of the voice data stream.

[0087] It should be noted that due to cost limitations, local devices often have insufficient computing power and limited memory, and in some special cases, they cannot process data in a timely manner, resulting in data accumulation. The wake-up word judgment module cannot process real-time data in a timely manner, resulting in a situation where users have to wait several seconds to get a response after speaking the wake-up word, causing the actual usage experience to deteriorate.

[0088] In the embodiments of the present application, the data processing process of the voice data stream includes steps such as voice acquisition, voice preprocessing, voice endpoint detection, and wake-up word judgment, and this process is realized through a voice processing system; in addition, the processing anomaly detection process of the voice data stream such as extracting timestamps and time difference judgment is realized through an edge computing end corresponding to the voice processing system.

[0089] While realizing the monitoring of voice data processing anomalies in the embodiments of the present application, by using edge computing technology to realize the monitoring processing process, it is not necessary for the local device of the voice processing system to bear the operation of the monitoring process, thereby avoiding increasing the computing power burden on the local device with limited computing power, and effectively improving the robustness of the voice processing system.

[0090] Please refer to Figure 2-3 , based on the above scheme, for better understanding of the voice data processing method provided by the embodiments of the present application, the following takes the processing of voice wake-up data as an example for detailed description:

[0091] It should be noted that before voice wake-up, the voice signal needs to be processed first, including voice acquisition, voice preprocessing, voice endpoint detection, wake-word judgment, and finally obtaining the wake-up result. The process is as Figure 2 shown. In this pipeline, each voice processing module processes the voice data frame by frame, and the voice data is transmitted between different modules through a buffer. Limited by cost, local devices often have insufficient computing power and limited memory. In some special cases, data cannot be processed in time, resulting in data accumulation, and the wake-word judgment result cannot be obtained in time. As a result, it takes several seconds for the user to get a response after saying the wake-word, which deteriorates the actual usage experience. Therefore, in the voice wake-up process, exploring how to timely identify data stream blockages, detect system anomalies, and avoid the deterioration of the actual user experience has become the key to affecting system stability.

[0092] To solve the problem that voice data cannot be processed in time due to system anomalies or high-priority tasks preempting system resources such as the CPU, resulting in a slow wake-up response in the case of data accumulation. Through the solution of the embodiments of the present application, anomalies in the voice data transmission process can be timely identified, and then emergency measures can be taken to process the accumulated data in time, so that the system can quickly return to the normal state and improve the user interaction experience.

[0093] In the embodiments of the present application, by adding timestamps to the voice data frames, different modules and different threads can compare the voice frame timestamps with the current system time when processing the voice data frames, timely discover the problem of voice data lag caused by untimely processing, enhance the system robustness, and improve the user voice wake-up interaction experience.

[0094] As Figure 2 shown, for the data stream of voice wake-up, first, the voice data is acquired. After preprocessing and endpoint detection of the original voice signal to obtain valid voice data, it is sent to the wake-word judgment module for discrimination to obtain the wake-word discrimination result.

[0095] Perform a frame splitting operation on the voice data stream, process and transmit the data in units of voice data frames, and by adding timestamps to the voice data frames, the monitoring of the transfer process of the voice data frames in each module can be realized. The specific process is as follows:

[0096] The first voice processing module (voice acquisition module) acquires the original voice data. In the entire data processing system, this module is closest to the data source, so the system time at this time is used to timestamp the voice data frame.

[0097] Therefore, when acquiring the 0th frame of voice, the 0th frame of voice will be timestamped according to the system time at the current moment; when acquiring the 1st frame of voice, the 1st frame of voice will be timestamped according to the system time at the current moment, and so on.

[0098] When performing subsequent steps, this timestamp is applied to analyze and determine whether the speech frames are delayed.

[0099] For example, during the speech preprocessing process, after obtaining a complete speech data frame, the timestamp identifier of the speech data frame is extracted, and the current system time is combined with the timestamp of the speech data frame for comparison and judgment. Different coping strategies are adopted for different results, such as Figure 3 as shown:

[0100] By comparing the current system time with the timestamp of the speech data frame, if the time difference is too large, it means that there is data accumulation in the entire system, and the data cannot be processed in time, resulting in serious delay phenomena. At this time, an emergency strategy should be adopted for processing to enable the system to return to the normal state in time.

[0101] Similarly, in the subsequent endpoint detection and wake word judgment modules, the input time of the speech data frame can be obtained through the timestamp of the speech frame. By comparing with the system time, a conclusion can be drawn on whether there is an abnormality in the transmission process of the speech data, and then corresponding coping strategies can be adopted in a timely manner.

[0102] It should be noted that through the embodiments of the present application, the problem of speech data lag caused by untimely processing can be discovered in time. The time difference between the system time and the timestamp of the speech data frame is used to judge whether the data is delayed. After the delay is too large, an emergency strategy is adopted in time, avoiding the phenomenon that the user does not get a response within a short time after waking up and then suddenly gets a response after a few seconds, thereby enhancing the system robustness and improving the interactive experience of user speech wake-up.

[0103] Refer to Figure 4 , Figure 4 is a schematic diagram of the modules of the speech data processing device provided by the embodiments of the present application. The speech data processing device provided by the embodiments of the present application includes:

[0104] Marking module 1, configured to perform frame splitting processing on the speech data stream through the first speech processing module, and perform timestamp marking on each speech data frame obtained after the frame splitting processing according to the current system time;

[0105] Extraction module 2, configured to extract the timestamp of the target speech data frame in the speech processing module;

[0106] Determination module 3, configured to determine that an abnormality has occurred in the data processing process of the speech data stream when the time difference between the timestamp and the current system time is greater than a preset threshold;

[0107] Wherein, the target speech data frame is one of the speech data frames obtained after the speech data stream is subjected to frame splitting processing.

[0108] In one embodiment, the data processing process of the voice data stream sequentially includes a plurality of processing steps. There are a plurality of voice processing modules, and the plurality of processing steps and the plurality of voice processing modules have a one-to-one correspondence relationship;

[0109] The determining module 3 is specifically configured to:

[0110] Determine a first threshold corresponding to the voice processing module according to the order of the processing steps corresponding to the voice processing module in the data processing process;

[0111] If the time difference between the timestamp and the current system time is greater than the first threshold, it is determined that an abnormality occurs in the data processing process of the voice data stream.

[0112] In one embodiment, the voice data processing device further includes a scheduling module, which is used for:

[0113] Determine a resource scheduling policy according to the order of the processing steps corresponding to the voice processing module in the data processing process;

[0114] Execute the resource scheduling policy so that the plurality of voice processing modules preferentially process the voice data frames of the voice data stream.

[0115] In one embodiment, the first voice processing module is the voice processing module corresponding to the starting step of the data processing process.

[0116] In one embodiment, the extraction module 2 is specifically configured to:

[0117] Determine that the voice processing module has not performed data processing on the target voice data frame, and extract the timestamp of the target voice data frame in the voice processing module.

[0118] In one embodiment, the plurality of processing steps sequentially include a voice acquisition step, a voice preprocessing step, a voice endpoint detection step, and a voice data judgment step.

[0119] In one embodiment, the voice data processing method is applied to an edge computing end corresponding to a voice processing system.

[0120] Figure 5 Illustrates a schematic diagram of the physical structure of an electronic device, such as Figure 5As shown in the figure, the electronic device may include: a processor 510, a communications interface 520, a memory 530, and a communication bus 540. Among them, the processor 510, the communications interface 520, and the memory 530 communicate with each other through the communication bus 540. The processor 310 may call the logical instructions in the memory 530 to execute the following method:

[0121] S1. Perform frame splitting processing on the voice data stream through the first voice processing module, and perform timestamp marking on each voice data frame obtained after the frame splitting processing according to the current system time;

[0122] S2. Extract the timestamp of the target voice data frame in the voice processing module;

[0123] S3. If the time difference between the timestamp and the current system time is greater than a preset threshold, it is determined that an abnormality has occurred in the data processing process of the voice data stream;

[0124] Among them, the target voice data frame is one of the voice data frames obtained after the voice data stream is subjected to frame splitting processing.

[0125] In addition, when the logical instructions in the above-mentioned memory 530 are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the related technology, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0126] On the other hand, the embodiments of the present application disclose a computer program product. The computer program product includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the methods provided in the above-mentioned method embodiments, for example, including:

[0127] S1. Perform frame splitting processing on the voice data stream through the first voice processing module, and perform timestamp marking on each voice data frame obtained after the frame splitting processing according to the current system time;

[0128] S2. Extract the timestamp of the target voice data frame in the voice processing module;

[0129] S3. If the time difference between the timestamp and the current system time is greater than a preset threshold, determine that an abnormality has occurred in the data processing process of the voice data stream;

[0130] Wherein, the target voice data frame is one of the voice data frames obtained after the voice data stream is frame-processed.

[0131] In another aspect, an embodiment of the present application further provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is configured to execute the transmission methods provided in the above embodiments, for example, including:

[0132] S1. Perform frame processing on the voice data stream through the first voice processing module, and mark the timestamp of each voice data frame obtained after the frame processing according to the current system time;

[0133] S2. Extract the timestamp of the target voice data frame in the voice processing module;

[0134] S3. If the time difference between the timestamp and the current system time is greater than a preset threshold, determine that an abnormality has occurred in the data processing process of the voice data stream;

[0135] Wherein, the target voice data frame is one of the voice data frames obtained after the voice data stream is frame-processed.

[0136] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0137] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the above technical solutions, in essence, or the part that contributes to the related technologies can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0138] Finally, it should be noted that the above embodiments are only used to illustrate the present application, rather than limiting the present application. Although the present application has been described in detail with reference to the embodiments, those of ordinary skill in the art should understand that various combinations, modifications, or equivalent replacements of the technical solutions of the present application do not depart from the spirit and scope of the technical solutions of the present application, and should all be covered within the scope of the claims of the present application.

Claims

1. A method for processing voice data, characterized in that, Including: Performing frame segmentation processing on the voice data stream through the first voice processing module, and performing timestamp marking on each voice data frame obtained after the frame segmentation processing according to the current system time; Extracting the timestamp of the target voice data frame in the voice processing module; When the time difference between the timestamp and the current system time is greater than a preset threshold, determining that an abnormality has occurred in the data processing process of the voice data stream; Wherein, the target voice data frame is one of the voice data frames obtained after the voice data stream is frame-segmented; The data processing process of the voice data stream sequentially includes multiple processing steps, there are multiple voice processing modules, and the multiple processing steps and the multiple voice processing modules are in one-to-one correspondence; When the time difference between the timestamp and the current system time is greater than a preset threshold, determining that an abnormality has occurred in the data processing process of the voice data stream, including: Determining a first threshold corresponding to the voice processing module according to the order of the processing steps corresponding to the voice processing module in the data processing process; When the time difference between the timestamp and the current system time is greater than the first threshold, determining that an abnormality has occurred in the data processing process of the voice data stream.

2. The voice data processing method according to claim 1, characterized in that After determining that an abnormality has occurred in the data processing process of the voice data stream when the time difference between the timestamp and the current system time is greater than a preset threshold, further including: Determining a resource scheduling strategy according to the order of the processing steps corresponding to the voice processing module in the data processing process; Executing the resource scheduling strategy so that the multiple voice processing modules preferentially process the voice data frames of the voice data stream.

3. The voice data processing method according to claim 1, wherein The first voice processing module is the voice processing module corresponding to the starting step of the data processing process.

4. The voice data processing method according to claim 1, wherein The extracting the timestamp of the target voice data frame in the voice processing module includes: Determining that the voice processing module has not performed data processing on the target voice data frame, and extracting the timestamp of the target voice data frame in the voice processing module.

5. The voice data processing method according to claim 1, wherein , The multiple processing steps sequentially include a voice acquisition step, a voice preprocessing step, a voice endpoint detection step, and a voice data judgment step.

6. The voice data processing method according to any one of claims 1 to 5, characterized in that , The voice data processing method is applied to an edge computing end corresponding to a voice processing system.

7. A voice data processing device, characterized in that, Including: A marking module, configured to perform frame segmentation processing on the voice data stream through the first voice processing module, and perform timestamp marking on each voice data frame obtained after the frame segmentation processing according to the current system time; An extraction module, configured to extract the timestamp of the target voice data frame in the voice processing module; A determination module, configured to determine that an abnormality has occurred in the data processing process of the voice data stream when the time difference between the timestamp and the current system time is greater than a preset threshold; Wherein, the target voice data frame is one of the voice data frames obtained after the voice data stream is frame-segmented; The data processing process of the voice data stream sequentially includes multiple processing steps, there are multiple voice processing modules, and the multiple processing steps and the multiple voice processing modules are in one-to-one correspondence; The determination module is specifically configured to: Determine a first threshold corresponding to the voice processing module according to the order of the processing steps corresponding to the voice processing module in the data processing process; If the time difference between the timestamp and the current system time is greater than the first threshold, it is determined that an abnormality has occurred in the data processing process of the voice data stream.

8. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the voice data processing method according to any one of claims 1 to 6.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the voice data processing method according to any one of claims 1 to 6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the voice data processing method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Reducing speech recognition latency

    US9514747B1