A real-time speech recognition and network anomaly detection system, method and apparatus

CN122821990APending Publication Date: 2026-09-25李虎 +4
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611014394.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-08
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0006]本申请提供了一种实时语音识别与网络异常检测系统、方法、设备及介质,用于解决现有复杂多变的网络环境对语音数据的上传和处理产生不利影响,进而破坏字幕与语音之间的同步性,影响实时字幕的显示效果,降低用户体验的问题

Benefits of technology

[0022]1、通过语音计时模块动态监测语音数据包的时间流逝时长,结合预设时长阈值判定当前网络状态,如网络正常、网络轻度异常和网络重度异常,实现网络波动与语音处理流程的精准适配。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122821990A_ABST
    Figure CN122821990A_ABST
Patent Text Reader

Abstract

The application discloses a real-time speech recognition and network anomaly detection and speech recognition system, method and equipment. A speech timing module dynamically monitors the time duration, combines a preset threshold to determine the current network state, the current network state includes network normal, network mild anomaly and network severe anomaly, and realizes accurate adaptation of network fluctuation and speech processing. The speech continuity judgment module distinguishes packet loss and delay when the network is severely abnormal, and provides differentiated processing basis for the control module. The control module directly calls the speech recognition module when the network is normal; when the network is abnormal, the recognition is kept and prompt information is superimposed to ensure uninterrupted service and meet the scene demand of high continuity requirement such as conference and customer service. The prompt information corresponding to different abnormal reasons enables users to intuitively perceive the network quality, and the forced processing of abnormal speech data packets ensures the completeness of the speech stream, solves the problems of recognition delay and misrecognition, and improves the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of artificial intelligence technology and speech recognition, and in particular to a real-time speech recognition and network anomaly detection system, method and device. Background Technology

[0002] With the rapid development of information technology, speech recognition technology has been widely applied in many fields, greatly improving the efficiency and convenience of information interaction. In remote audio-visual systems, such as remote conferencing systems and live streaming systems, as well as voice input systems, speech recognition technology plays a crucial role. It can accurately convert speech signals into text information, meeting the diverse needs for information recording, display, and interaction in different scenarios.

[0003] Taking remote audio and video systems as an example, these systems play an irreplaceable role in many fields such as business communication, online education, and live entertainment. Figure 1 This diagram illustrates a typical remote audio-visual system architecture provided in this application. This architecture typically consists of multiple components, including a data acquisition computer, a camera, a microphone, a server, and a client. The data acquisition computer connects to the camera and microphone, and is responsible for acquiring video and audio data. This data is then uploaded to the server via a network. Upon receiving the data, the server uses a speech recognition service to convert the speech content into text, generating subtitles. Finally, real-time audio-visual content with subtitles is displayed on the client, providing users with richer and more intuitive information.

[0004] In practical applications, audio acquisition follows certain rules, and once a certain amount of voice data is collected, it is uploaded to the server for processing. Figure 2 This diagram illustrates the subtitle functionality of a remote audio / video system under normal network conditions. In an ideal network environment, the audio upload process takes only a very short time, and the server's audio recognition processing time is also minimal. Because the continuous acquisition, uploading, and recognition of audio are seamlessly integrated, the time difference between subtitles and audio is extremely short, only a few milliseconds, enabling real-time subtitles and significantly improving the user experience and information transmission efficiency of the remote audio / video system.

[0005] However, real-world network environments are complex and ever-changing, and are not always ideal. Network fluctuations and latency issues occur frequently, which adversely affect the uploading and processing of voice data, thereby disrupting the synchronization between subtitles and audio, impacting the display quality of real-time subtitles, and reducing user experience. Therefore, ensuring real-time synchronization between subtitles and audio in a speech recognition system under complex network environments has become a critical issue that urgently needs to be addressed. Summary of the Invention

[0006] This application provides a real-time speech recognition and network anomaly detection system, method, device, and medium to address the problem that the complex and ever-changing network environment adversely affects the uploading and processing of speech data, thereby disrupting the synchronization between subtitles and speech, affecting the display effect of real-time subtitles, and reducing user experience.

[0007] In a first aspect, this application provides a real-time speech recognition and network anomaly detection system, the system comprising:

[0008] The speech recognition module is configured to provide streaming speech recognition services;

[0009] The voice timing module is configured to determine the reception timestamp of any voice data packet received in real time; determine the elapsed time based on the reception timestamp and the cumulative total duration of the voice data packet and all previous voice data packets in the same voice stream; determine the current network status based on the comparison result of the elapsed time and a preset duration threshold; wherein the current network status includes normal network, slightly abnormal network, and severely abnormal network; if the current network status is determined to be severely abnormal, the voice continuity judgment module is invoked.

[0010] The voice continuity determination module is configured to provide a voice continuity determination detection service in response to the voice timing module, so as to determine the cause of the anomaly in the current network state; wherein the cause of the anomaly includes network latency and network packet loss;

[0011] The control module is configured to, if it is determined that the current network status is normal, call the speech recognition module to process the speech data packet; otherwise, output the prompt information corresponding to the abnormal reason of the current network status and call the speech recognition module to process the speech data packet; wherein, the abnormal reason for a minor network abnormality is network latency.

[0012] Secondly, this application also provides a real-time speech recognition and network anomaly detection method, the method comprising:

[0013] For any voice data packet received in real time, determine the timestamp of the voice data packet's reception;

[0014] The elapsed time is determined based on the received timestamp and the total cumulative duration of the voice data packet and all previous voice data packets in the same voice stream.

[0015] The current network status is determined based on the comparison between the elapsed time and the preset time threshold; wherein the current network status includes normal network, slightly abnormal network, and severely abnormal network.

[0016] If the current network status is determined to be normal, then the voice data packet is directly processed for speech recognition.

[0017] If the current network status is determined to be a slight network anomaly, then output the prompt message corresponding to the network delay and perform speech recognition processing on the voice data packet;

[0018] If the current network state is determined to be severely abnormal, then the voice data packet is subjected to voice continuity judgment and detection to determine the cause of the abnormality of the current network state; wherein, the cause of the abnormality includes network latency and network packet loss; the prompt information corresponding to the cause of the abnormality of the current network state is output and the voice data packet is subjected to voice recognition processing.

[0019] Thirdly, this application provides a computer device including a processor, which executes a computer program stored in a memory to implement the steps of the real-time speech recognition and network anomaly detection method described above.

[0020] Fourthly, this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the real-time speech recognition and network anomaly detection method described above.

[0021] The beneficial effects of this application are as follows:

[0022] 1. The voice timing module dynamically monitors the elapsed time of voice data packets and determines the current network status by combining preset duration thresholds, such as normal network, slightly abnormal network, and severely abnormal network, so as to achieve accurate adaptation between network fluctuations and voice processing flow.

[0023] 2. By introducing a voice continuity judgment module, packet loss and delay can be further distinguished when the network is severely abnormal, providing a basis for differentiated processing of the control module.

[0024] 3. When the network is normal, the control module directly calls the speech recognition module to process data packets. Even in cases of network issues, the recognition process continues, with added prompts such as "Network delay, results may be inaccurate," ensuring uninterrupted service and effectively mitigating the impact of network problems on the real-time performance of speech recognition. Especially in scenarios with minor network anomalies, users can obtain recognition results and network status in real time without waiting for network recovery, meeting the stringent continuity requirements of scenarios such as conferences and customer service.

[0025] 4. The system achieves functional decoupling through modular design. Each module runs independently and triggers subsequent processes only when necessary, significantly reducing CPU and memory usage. It is suitable for deployment in resource-constrained embedded devices or mobile terminals.

[0026] 5. By providing prompts indicating the reasons for current network anomalies, users can intuitively perceive changes in network quality. At the same time, the control module's forced processing of abnormal data packets, rather than discarding them, ensures the integrity of the voice stream. It effectively detects and resolves recognition delays and misrecognition issues caused by factors such as network fluctuations and bandwidth limitations in voice recognition scenarios, thereby improving the user experience in different voice recognition application scenarios. Attached Figure Description

[0027] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0028] Figure 1 This application provides a schematic diagram of the architecture of a typical remote audio and video system.

[0029] Figure 2 This is a diagram illustrating the subtitle display of a remote audio-visual system under normal network conditions.

[0030] Figure 3 A schematic diagram of the structure of a real-time speech recognition and network anomaly detection system provided in an embodiment of this application;

[0031] Figure 4 This is a diagram illustrating the subtitle display in a remote conferencing scenario where network latency exists.

[0032] Figure 5 This is a diagram illustrating the subtitle display in a remote conferencing scenario where network packet loss occurs.

[0033] Figure 6 A schematic diagram of the spectrum of two adjacent voice data packets provided in an embodiment of this application;

[0034] Figure 7 This is a schematic diagram illustrating the determination of continuity between two adjacent voice data packets, provided as an embodiment of this application.

[0035] Figure 8 A schematic diagram illustrating the workflow of a specific real-time speech recognition and network anomaly detection system provided in this application embodiment;

[0036] Figure 9 This application provides a schematic diagram of a real-time speech recognition and network anomaly detection process.

[0037] Figure 10 This is a schematic diagram of the structure of a computer device provided in an optional embodiment of this application. Detailed Implementation

[0038] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0039] To effectively ensure real-time synchronization of subtitles and speech in a speech recognition system under complex network environments and improve user experience, this application provides a real-time speech recognition and network anomaly detection system, method, device, and medium.

[0040] Example 1:

[0041] This application provides a real-time speech recognition and network anomaly detection system. Figure 3 A schematic diagram of a real-time speech recognition and network anomaly detection system provided in this application embodiment is shown. The system includes:

[0042] The speech recognition module 31 is configured to provide streaming speech recognition services;

[0043] The voice timing module 32 is configured to determine the reception timestamp of any voice data packet received in real time; determine the elapsed time based on the reception timestamp and the cumulative total duration of the voice data packet and all previous voice data packets in the same voice stream; determine the current network status based on the comparison result of the elapsed time and a preset duration threshold; wherein the current network status includes normal network, slightly abnormal network, and severely abnormal network; if the current network status is determined to be severely abnormal, the voice continuity judgment module 33 is invoked;

[0044] The voice continuity judgment module 33 is configured to respond to the voice timing module 32 and provide a voice continuity judgment detection service to determine the cause of the anomaly in the current network state; wherein the cause of the anomaly includes network latency and network packet loss;

[0045] The control module 34 is configured to, if it is determined that the current network status is normal, call the voice recognition module 31 to process the voice data packet; otherwise, output the prompt information corresponding to the abnormal reason of the current network status and call the voice recognition module 31 to process the voice data packet; wherein, the abnormal reason for a minor network abnormality is network latency.

[0046] In existing remote audio and video systems such as remote conferencing and live streaming systems, as well as voice input scenarios, while speech recognition technology has achieved real-time captioning, it has significant drawbacks in complex network environments: when network fluctuations occur, such as latency or packet loss, the time difference between uploading and processing voice data packets increases dramatically, disrupting the synchronization between captions and speech and significantly degrading the user experience. Specifically, in network latency scenarios... Figure 4 This diagram illustrates the subtitle display in a remote conferencing scenario with network latency. Because the audio data packets take too long to transmit, the subtitles lag behind the audio by several seconds after the server receives the packets, completes recognition, and generates the subtitles. This causes a mismatch between the audio and subtitles seen by the user, especially in fast-paced remote meetings where this delay severely impacts information comprehension. Furthermore, in scenarios with network packet loss… Figure 5 This diagram illustrates the subtitle situation in a remote conferencing scenario where network packet loss occurs. Some voice data packets are lost during transmission, preventing the server from receiving complete voice data and resulting in incomplete subtitles. If the lost data packets are subsequently retransmitted, the subtitle order may be disordered, leading to disjointed content and making it difficult for users to accurately grasp the complete information. To address these issues, this application provides a real-time speech recognition and network anomaly detection system (hereinafter referred to as the system). This system can accurately detect network anomalies and determine their causes while performing real-time speech recognition, thus providing users with a more stable and reliable speech recognition service. It is particularly suitable for scenarios with high real-time requirements, such as remote conferencing and live streaming. The system includes a speech recognition module 31, a speech timing module 32, a speech continuity judgment module 33, and a control module 34. Through the cooperation of these modules, real-time speech recognition and network anomaly detection and processing functions are achieved. The functions of each module in the system are described in detail below:

[0047] A. Speech recognition module 31

[0048] The speech recognition module 31, as one of the core functional modules of the system, is configured to provide streaming speech recognition services. This module can receive real-time transmitted voice data packets and, based on deep learning-based streaming speech recognition technology, perform continuous and efficient speech-to-text processing, outputting the corresponding recognition results to achieve low-latency speech-to-text functionality. For example, in a remote conferencing scenario, this module can convert participants' speech into text information in real time, providing basic support for generating real-time subtitles. The streaming speech recognition technology used in this module 31 can segment the continuously incoming voice data stream, outputting recognition results gradually without waiting for the complete transmission of the voice data, ensuring the real-time performance of speech recognition.

[0049] In one possible implementation, the speech recognition module 31 can load a pre-trained speech recognition model, such as a TDNN-F model trained on the Kaldi framework, a Transformer model trained on the ESPnet framework, and recurrent neural networks (RNNs) and their variants (such as Long Short-Term Memory networks (LSTM) and gated recurrent units (GRUs). This speech recognition model, trained on a large amount of speech data, has the ability to accurately recognize speech content in at least one language, such as Chinese or English. In practical use, upon receiving a speech data packet, the speech recognition module 31 immediately extracts features from the data packet and inputs the feature sequence into the speech recognition model for real-time decoding. A frame synchronization strategy can be used during decoding to gradually output the recognition results, ensuring the synchronization of subtitles and speech. For example, in a remote conferencing scenario, when participants' speech data packets are continuously transmitted to the server, the speech recognition module 31 can output the corresponding text subtitles within 20ms of receiving the speech data packets and display them in real-time through the client.

[0050] B. Voice timing module 32

[0051] The voice timing module 32 is primarily responsible for recording the reception timestamps of each received voice data packet and determining the current network status based on this. Specifically, for any voice data packet received in real time, the voice timing module 32 immediately records and determines the reception timestamp of that voice data packet. Based on the determined reception timestamp and the cumulative total duration of the voice data packet and all previous voice data packets in the same voice stream, the elapsed time is calculated. The total duration of all previous voice data packets should theoretically be the sum of the actual duration of the voice content contained in all successfully received voice data packets before the current data packet in the same voice stream. For example, if two 200ms voice data packets from the same voice stream were received before the current reception timestamp, then the total duration of all previous voice data packets is 400ms.

[0052] In one example, if the timestamp of the current voice data packet reception is determined based on actual time, such as milliseconds based on the system clock, then the timestamp recorded at the start of timing is used as the starting time. The difference between the current voice data packet reception timestamp and the starting time, minus the accumulated total duration, yields the elapsed time. This elapsed time reflects, to some extent, the accumulated additional time consumed by network transmission. For example, if the starting time is T0, such as 10:00:00.000, the current voice data packet reception timestamp is T3, such as 10:00:00.600, and the accumulated total duration is N, such as 400ms, then the elapsed time is T3 - T0 - N = (600ms - 0ms) - 400ms = 200ms.

[0053] In another example, if the start timer timestamp is set to a base time of 0ms, then the reception timestamps of the current voice data packet and subsequent voice data packets are all cumulative durations relative to this base time. For example, in the same voice stream, the first voice data packet is received 210ms after the voice timing module 32 starts timing, with a timestamp of 210ms; the second voice data packet is received 10ms later, with a timestamp of 220ms. In this case, the calculation logic for the elapsed time is: directly subtract the cumulative total duration from the reception timestamp of the current voice data packet. For example, if the timestamp of the current voice data packet is 500ms, and the cumulative total duration is 400ms, then the elapsed time is 500ms - 400ms = 100ms.

[0054] In one example, the total cumulative duration can be determined as follows:

[0055] Method 1: For voice data packets within the same voice stream, the voice timing module 32 can maintain an internal data structure, such as an array or linked list, to store the duration information of successfully received voice data packets. Each time a new voice data packet is received, the voice timing module 32 adds that duration to the previously stored total duration. For example, initially, the total duration is 0. If the first voice data packet is received with a duration of 200ms, the total duration is updated to 200ms. Then, if the second voice data packet is received with a duration of 250ms, the total duration is updated to 200ms + 250ms = 450ms, and so on, thus obtaining the total duration of the current voice data packet and all previous voice data packets within the same voice stream.

[0056] Method 2: Determine the total cumulative duration by using the sequence number of the voice data packets and the preset duration of the voice data packets. From a voice coding perspective, many common voice coding technologies (such as PCM, ADPCM, etc.) typically use a fixed duration to divide voice data packets during design and implementation. For example, in a real-time voice communication system based on PCM coding, to ensure the continuity and real-time performance of the voice while also considering network transmission efficiency, the duration of each voice data packet is generally set to a fixed length to ensure the stability and consistency of the voice data during encoding and decoding. Furthermore, in scenarios with extremely high real-time requirements, such as game voice chat and remote conferencing, the preset duration of the voice data packets allows the server to quickly calculate the total duration, promptly assess the network status, and take corresponding optimization measures to ensure smooth voice transmission. For example, the acquisition computer collects voice data at a 16kHz sampling rate, generating a voice data packet every 200ms. Therefore, in this application, the acquisition end can pre-agree with the server on parameters such as encoding method, sampling rate, and preset voice data packet duration. The acquisition end can upload each voice data packet acquired according to these parameters to the system's voice timing module 32. After receiving any voice data packet, the system's voice timing module 32 can determine the sequence number of the voice data packet based on the maximum sequence number of the previously received voice data packet in the same voice stream. The cumulative total duration is determined by multiplying the sequence number by the preset voice data packet duration. For example, assuming the preset voice data packet duration is 200ms and the sequence number of the currently received voice data packet in the voice stream is 5, then the cumulative total duration is 5 * 200 = 1000ms.

[0057] By determining the total cumulative duration using the aforementioned sequence number and preset duration, the system can quickly obtain the total cumulative duration through simple arithmetic operations without parsing the specific content of each voice data packet to obtain duration information. This significantly reduces the complexity of data processing and further improves the system's processing efficiency. This method is suitable for scenarios where voice data packet durations are fixed and the transmission order is stable, such as standardized remote conferencing systems and fixed-format live voice streams. It can meet the system's high real-time requirements while ensuring the accuracy of duration calculation.

[0058] After obtaining the elapsed duration of the current voice data packet based on the above embodiments, the voice timing module 32 can compare the elapsed duration with a preset duration threshold and determine the current network status based on the comparison result.

[0059] The preset duration threshold can be flexibly set according to the actual network environment and application requirements. Specific setting methods include, but are not limited to, the following:

[0060] 1) Setting based on normal network transmission characteristics: This can be achieved by statistically analyzing the time required for normal voice data packet transmission over a period of time, and then calculating the average of these times (denoted as the average transmission duration). The average transmission duration of voice data packets under normal network conditions can then be used as a benchmark to set a preset duration threshold. For example, a threshold 1.5 times the average transmission duration can be set. For instance, if the average transmission duration under normal conditions is 50ms, the preset duration threshold could be set to 125ms.

[0061] 2) Setting based on human experience: Thresholds can also be set directly based on historical operating data and maintenance experience for specific application scenarios. For example, in remote conferencing scenarios with extremely high real-time requirements, the preset duration threshold can be set to 50ms; while in voice message transcription scenarios with relatively relaxed real-time requirements, the threshold can be appropriately increased, such as setting the preset duration threshold to 100ms.

[0062] During the setup process, if you want the system to be more sensitive to network anomalies, i.e., to identify anomalies earlier, you can set the preset duration threshold to be smaller; if you want to reduce false positives, i.e., allow for a certain degree of network fluctuation, you can set the preset duration threshold to be larger. The specific value needs to be determined based on the actual application scenario, and is not limited here.

[0063] In one example, when the preset duration threshold is a single value, the voice timing module 32 can compare the elapsed time with the preset duration threshold. If it is less than the preset duration threshold, the current network status is determined to be normal; if it is greater than or equal to the preset duration threshold, the current network status is determined to be abnormal. If the current network status is determined to be abnormal, the degree to which the elapsed time exceeds the preset duration threshold can be used to determine whether the current network status is slightly abnormal or severely abnormal. For example, if the elapsed time is greater than or equal to the preset duration threshold but less than a preset multiple of the preset duration threshold, it indicates that the current network status may have slight latency, and is therefore determined to be slightly abnormal. If the elapsed time is greater than or equal to a preset multiple of the preset duration threshold, it indicates that the current network status is highly likely to have packet loss, and is therefore determined to be severely abnormal.

[0064] In another example, the voice timing module 32 is specifically configured as follows:

[0065] If the elapsed time is less than a first preset time threshold, then the current network status is determined to be normal.

[0066] If the elapsed time is greater than or equal to the first preset time threshold and less than the second preset time threshold, then the current network state is determined to be a slight network anomaly.

[0067] If the elapsed time is greater than or equal to the second preset time threshold, then the current network state is determined to be a severely abnormal network.

[0068] The preset duration threshold can also include a first preset duration threshold and a second preset duration threshold. The first preset duration threshold can be understood as a threshold for minor anomalies, and the second preset duration threshold can be understood as a threshold for severe anomalies. If the elapsed time is less than the first preset duration threshold, the network is considered normal. If the elapsed time is greater than or equal to the first preset duration threshold but less than the second preset duration threshold, it indicates that the current network status may have minor latency, and is therefore considered a minor network anomaly. If the elapsed time is greater than or equal to the second preset duration threshold, it indicates that the current network status is highly likely to have packet loss or other issues, and is therefore considered a severe network anomaly.

[0069] When the current network status is determined to be severely abnormal, the voice timing module 32 will actively call the voice continuity judgment module 33 to further analyze the cause of the abnormality.

[0070] C. Voice continuity judgment module 33

[0071] The speech continuity judgment module 33 responds to the call of the speech timing module 32 and provides speech continuity judgment and detection services. Its core function is to determine the cause of the current network state (severe network anomaly), specifically including network latency and network packet loss. This speech continuity judgment module 33 can determine the continuity of speech by performing acoustic feature analysis on the speech content of two adjacent speech data packets, and then determine the cause of the current network state anomaly based on the analysis results. The acoustic features include, but are not limited to, speech energy (such as root mean square energy), spectral features (such as Mel-frequency cepstral coefficients and spectral centroid), fundamental frequency information (such as F0), and zero-crossing rate.

[0072] For example, the speech continuity determination module 33 can employ a preset continuity determination algorithm to analyze the acoustic features of the current speech data packet and the previous speech data packet. The preset continuity determination algorithm includes, but is not limited to, determination based on curve smoothness (derivative), determination based on speech segment energy difference, zero-crossing rate analysis statistics, and dynamic time warping (DTW) algorithm. Then, based on the feature analysis results, the continuity of the speech is determined, thereby determining the cause of the current network state anomaly. Specifically, it can be determined whether the feature analysis results meet preset continuity detection conditions. If the feature analysis results meet the preset continuity detection conditions, the cause of the current network state anomaly is determined to be network latency; if the feature analysis results do not meet the preset continuity detection conditions, the cause of the current network state anomaly is determined to be network packet loss.

[0073] Taking a curve smoothness-based judgment algorithm as an example, Figure 6This is a schematic diagram of the spectrum of two adjacent voice data packets provided in an embodiment of this application. Figure 7 This diagram illustrates a method for determining the continuity of two adjacent voice data packets, as provided in this embodiment. Since continuous voice segments exhibit continuity at the boundary between two data packets, a sliding window is established. The window size can be set according to the voice sampling rate and data packet duration, such as 16kHz sampling rate or 100 sampling points corresponding to a 200ms data packet. The window is moved from the end of the current voice data packet to the beginning of the next voice data packet, and the voice data within the window is differentiated to calculate the first derivative. A preset continuity detection condition is that the difference in the first derivative values ​​of adjacent windows is within a preset difference threshold range. If the first derivatives at the boundary of two data packets are continuous, meaning the difference in the first derivative values ​​of adjacent windows is within a preset minimum threshold range (e.g., a threshold of 0.01), it indicates that the voice data is continuous at the boundary, satisfying the preset continuity detection condition, and no voice data packets are lost. In this case, the cause of the severe network anomaly is determined to be network latency. If the first derivatives are discontinuous, meaning the difference in the first derivative values ​​of adjacent windows exceeds the preset threshold, it indicates that the voice data is interrupted at the boundary, failing to satisfy the preset continuity detection condition. In this case, the cause of the severe network anomaly is determined to be network packet loss leading to partial data loss, thus preventing the two data packets from being connected coherently.

[0074] D. Control Module 34

[0075] The control module 34, acting as the system's scheduling center, controls the invocation and information output of each module based on the determined current network status. For example, if the voice timing module 32 determines that the current network status is normal, the control module 34 can directly invoke the voice recognition module 31 to process the currently received voice data packets, ensuring the smooth progress of the voice recognition process. If the voice timing module 32 determines that the current network status is slightly abnormal or severely abnormal, the control module 34 will first output a prompt message corresponding to the cause of the abnormality. For example, for a slightly abnormal network, it outputs "There is a delay in the current network, and the subtitles may be slightly delayed"; for a severely abnormal network with network delay, it outputs "The current network delay is severe, and the subtitles may be delayed for a long time"; for a severely abnormal network with packet loss, it outputs "There is packet loss in the current network, and the subtitles may be missing content." Subsequently, it still invokes the voice recognition module 31 to process the voice data packet, ensuring that voice recognition continues as much as possible while allowing the user to be aware of the network abnormality in a timely manner. Among them, the cause of network minor anomalies is always network latency. This is because in the case of minor anomalies, the main problem of the network is the slowdown in data transmission speed, which has not yet reached the level of severity that leads to packet loss.

[0076] In one possible implementation, the control module 34 is specifically configured as follows:

[0077] If the abnormality of the current network status is due to network packet loss, then available voice data is extracted from the cache based on the voice data packet and its adjacent voice data packets.

[0078] The speech recognition module 31 is invoked to process the available speech data to generate alternative recognition results;

[0079] The alternative identification result is output by displaying the preset highlighted mark.

[0080] In this application, when the network experiences severe anomalies caused by packet loss, the control module 34 can extract usable voice data from the system cache based on the currently received voice data packets and their adjacent voice data packets. The system cache stores complete data for each received voice data packet, including voice content, sequence number, timestamp, etc. The control module 34 can locate adjacent voice data packets based on their sequence number or timestamp and extract the portion of voice data unaffected by packet loss. The voice recognition module 31 is then invoked to process the extracted usable voice data. Due to the loss of some data packets, the usable voice data may be incomplete. The voice recognition module 31 will perform maximum recognition based on the existing data to generate alternative recognition results, filling in the content gaps caused by packet loss as much as possible. The alternative recognition results are then output using a preset highlighted display method. The preset highlighted marks may include, but are not limited to: using a different color than the normal recognition result, such as red; adding special symbols, such as "【】" or "<>"; using italics or bold fonts, etc. For example, displaying the alternative identification result as "[Content may be missing here due to network packet loss: ...]" allows users to intuitively distinguish between normal identification results and alternative identification results caused by packet loss, and clearly understand the integrity status of the content.

[0081] The beneficial effects of this application are as follows:

[0082] 1. The voice timing module 32 dynamically monitors the elapsed time of voice data packets and determines the current network status by combining the preset duration threshold, such as normal network, slightly abnormal network, and severely abnormal network, so as to achieve accurate adaptation between network fluctuations and voice processing flow.

[0083] 2. By introducing the voice continuity judgment module 33, the two types of abnormal causes, packet loss and delay, are further distinguished when the network is severely abnormal, providing a basis for the control module 34 to handle the abnormalities differently.

[0084] 3. When the network is normal, the control module 34 directly calls the speech recognition module 31 to process data packets. Even in cases of network issues, the recognition process continues to run with added prompts such as "Network delay, results may be inaccurate," ensuring uninterrupted service and effectively mitigating the impact of network problems on the real-time performance of speech recognition. Especially in scenarios with minor network anomalies, users can obtain recognition results and network status in real time without waiting for network recovery, meeting the stringent continuity requirements of scenarios such as meetings and customer service.

[0085] 4. The system achieves functional decoupling through modular design. Each module runs independently and triggers subsequent processes only when necessary, significantly reducing CPU and memory usage. It is suitable for deployment in resource-constrained embedded devices or mobile terminals.

[0086] 5. By providing prompts indicating the reasons for the current network status anomalies, users can intuitively perceive changes in network quality. At the same time, the control module 34's forced processing of abnormal data packets, rather than discarding them, ensures the integrity of the voice stream, effectively detects and resolves recognition delays and misrecognition issues caused by factors such as network fluctuations and bandwidth limitations in voice recognition scenarios, and improves the user experience in different voice recognition application scenarios.

[0087] Example 2:

[0088] The workflow of the real-time speech recognition and network anomaly detection system provided in this application will be described below through specific embodiments. Figure 8 This application provides a schematic diagram of the workflow of a specific real-time speech recognition and network anomaly detection system, which includes the following steps:

[0089] When the system starts running, the voice timing module 32 starts timing, the voice acquisition device collects voice data, and when the voice data reaches the preset time length k, it packages the voice data of the preset time length k and sends it to the voice timing module 32 on the server through the network.

[0090] After receiving a voice data packet, the voice timing module 32 determines and stores the sequence number n of the voice data packet. This sequence number represents the number of voice data packets from the same voice stream that the voice timing module 32 has currently received. Based on the timing of receiving the voice data packet, it determines and stores the reception timestamp Tn of the voice data packet. Based on the reception timestamp Tn, the sequence number n of the voice data packet, and the preset time length k, the elapsed time length (Tn-n*k) ms is determined. Ideally, the elapsed time length (Tn-n*k) ms should be between a few milliseconds and tens of milliseconds. Therefore, based on this empirical value, a first preset duration threshold of 50 ms and a second preset duration threshold of 250 ms can be set. If (Tn-n*k) ms < 50 ms, it indicates that no network packet loss or delay has occurred. The voice timing module 32 then determines that the current network status is normal and calls the voice recognition module 31 through the control module 34 for subsequent processing. If 50ms ≤ (Tn - n*k)ms < 250ms, it indicates a data delay, a relatively large network delay but no packet loss, and the data delay is within one packet's range. In this case, the delay is within the range imperceptible to human vision and hearing and is still acceptable. The voice timing module 32 determines the current network state as slightly abnormal, and the cause of the abnormality is network delay. The control module 34 outputs a prompt message corresponding to the network delay, indicating that there is currently a network delay and attention is needed. Simultaneously, the control module 34 calls the voice recognition module 31 for subsequent processing. If (Tn - n*k)ms ≥ 250ms, it indicates packet loss or data delay. The voice timing module 32 determines the current network state as severely abnormal, retrieves the previous voice data packet and sends it along with the current voice data packet to the voice continuity judgment module 33 for continuity judgment.

[0091] The voice continuity judgment module 33 uses a preset continuity judgment algorithm to perform feature analysis on the current voice data packet and the previous voice data packet. If the feature analysis results determine that the current voice data packet is continuous with the previous voice data packet, indicating no packet loss but a significant delay, the voice continuity judgment module 33 determines that the abnormality in the current network state is due to network latency. The control module 34 outputs a corresponding network latency warning, indicating that network latency exists and requires attention. Simultaneously, the control module 34 calls the voice recognition module 31 for further processing. If the feature analysis results determine that the current voice data packet is discontinuous with the previous voice data packet, indicating a high probability of packet loss, the voice continuity judgment module 33 determines that the abnormality in the current network state is due to network packet loss. The control module 34 outputs a corresponding network packet loss warning, indicating that network packet loss exists and subtitle content may be missing, requiring attention. Simultaneously, the control module 34 calls the voice recognition module 31 for further processing. When the control module 34 calls the speech recognition module 31 to process speech data packets whose abnormality is due to network packet loss, it can extract available speech data from the system cache based on the currently received speech data packets and their adjacent speech data packets. The system cache stores complete data of each received speech data packet, including speech content, sequence number, timestamp, etc. The control module 34 can locate adjacent speech data packets based on the packet sequence number or timestamp and extract the speech data portion that is not affected by packet loss. The speech recognition module 31 is then called to process the extracted available speech data. Since some data packets are lost, the available speech data may be incomplete. The speech recognition module 31 will perform maximum recognition based on the existing data to generate alternative recognition results to fill in the content gaps caused by packet loss as much as possible. Then, the alternative recognition results are output through a preset highlighting display method. The preset highlighting may include, but is not limited to: using a different color than the normal recognition result, such as red, adding special symbols, such as "【】" "<>", using italics or bold fonts, etc. For example, displaying the alternative identification result as "[Content may be missing here due to network packet loss: ...]" allows users to intuitively distinguish between normal identification results and alternative identification results caused by packet loss, and clearly understand the integrity status of the content.

[0092] It should be noted that the system provided in this application can be used not only for speech recognition in audio and video systems, but also for real-time detection in simple speech recognition systems.

[0093] In a remote conferencing scenario, the acquisition computer collects voice data at a sampling rate of 16kbps, generating a voice data packet every 200ms. The first preset duration threshold is 50ms, and the second preset duration threshold is 250ms.

[0094] After the meeting begins, the voice timing module 32 is activated and starts timing.

[0095] The acquisition device collects voice data. After the first 200ms voice data packet is generated, it is sent to the voice timing module 32 on the server via the network.

[0096] The voice timing module 32 receives the voice data packet, records the data number n=1, and adds a timestamp T1. At this time, if the elapsed time is (T1-1*200)ms=10ms, and is less than the first preset duration threshold of 50ms, the voice timing module 32 determines that the current network status is normal, and calls the voice recognition module 31 through the control module 34 for subsequent processing.

[0097] After the acquisition device sends the first 200ms voice data packet, it continues to acquire the next 200ms voice data.

[0098] As the meeting progressed, a network fluctuation occurred at a certain point. When the fifth 200ms voice data packet arrived at the server's voice timing module 32, the voice timing module 32 recorded a timestamp T5 and determined the elapsed time to be (T5-5*200)ms=100ms. Since 50ms≤(T5-5*200)ms<250ms, the voice timing module 32 determined the current network status to be a minor network anomaly, caused by network latency. The control module 34 then output a corresponding network latency warning, indicating that network latency exists and requires attention. Simultaneously, the control module 34 invoked the voice recognition module 31 for further processing.

[0099] If at some later time, the 8th 200ms voice data packet arrives at the server's voice timing module 32, the voice timing module 32 records the timestamp T8, determines the elapsed time (T8-8*200)ms=300ms, and (T8-8*200)ms≥250ms, then the voice timing module 32 determines that the current network status is a severe network anomaly, retrieves the 7th and 8th voice data packets, and sends them to the voice continuity judgment module 33 to call the voice continuity judgment module 33 for continuity judgment.

[0100] The speech continuity judgment module 33 uses a preset continuity judgment algorithm to perform feature analysis on the 7th and 8th speech data packets. If the feature analysis results indicate that the 7th and 8th speech data packets are continuous, it means there is only a severe delay, and the control module 34 sends the 8th speech data packet to the speech recognition module 31 for further processing. If the feature analysis results indicate that the 7th and 8th speech data packets are not continuous, it means there is packet loss, and the speech continuity judgment module 33 determines that the abnormality of the current network status is due to network packet loss. The control module 34 outputs a prompt message corresponding to network packet loss, indicating that there is currently a network packet loss problem and that the subtitle content may be missing, requiring attention. At the same time, the control module 34 calls the speech recognition module 31 for further processing. When the control module 34 calls the speech recognition module 31 to process speech data packets whose abnormality is due to network packet loss, it can search for available speech data in the cache, and call the speech recognition module 31 to process the extracted available speech data, generate alternative recognition results, and present them in a preset highlighted display mode.

[0101] It should be noted that the various modules in this system can be deployed in a distributed manner on different servers or clusters, or they can be deployed on the same server or server cluster; no specific restrictions are imposed here.

[0102] Example 3:

[0103] This application also provides a method for real-time speech recognition and network anomaly detection. Figure 9 This application provides a schematic diagram of a real-time speech recognition and network anomaly detection process, which includes:

[0104] S801: For any voice data packet received in real time, determine the timestamp of the voice data packet reception.

[0105] S802: Determine the elapsed time based on the received timestamp and the total duration of the voice data packet and all previous voice data packets in the same voice stream.

[0106] S803: Determine the current network status based on the comparison result between the elapsed time and the preset time threshold; wherein, the current network status includes normal network, slightly abnormal network, and severely abnormal network.

[0107] S804: If it is determined that the current network status is normal, then the voice data packet is directly processed for voice recognition.

[0108] S805: If it is determined that the current network status is a slight network anomaly, then output the prompt information corresponding to the network delay and perform speech recognition processing on the voice data packet.

[0109] S806: If the current network state is determined to be severely abnormal, then the voice data packet is subjected to voice continuity judgment and detection to determine the cause of the abnormality of the current network state; wherein, the cause of the abnormality includes network latency and network packet loss; output the prompt information corresponding to the cause of the abnormality of the current network state and perform voice recognition processing on the voice data packet.

[0110] The real-time speech recognition and network anomaly detection method provided in this application is applied to computer equipment, which can be a smart device, such as a mobile device or a computer, or a server, such as a business server or an application server.

[0111] It should be noted that the principle of this real-time speech recognition and network anomaly detection method in solving technical problems is similar to that in the above system embodiments. For details, please refer to the description in the above system embodiments, which will not be elaborated here.

[0112] Example 4:

[0113] Please see Figure 10 , Figure 10 This is a schematic diagram of the structure of a computer device provided in an optional embodiment of this application, such as... Figure 10 As shown, the computer device includes one or more processors 10, memory 20, and interfaces for connecting the components, including high-speed interfaces and low-speed interfaces. The six components are interconnected via different buses and can be mounted on a common motherboard or otherwise installed as needed. The processors can process instructions executed within the computer device, including instructions stored in or on memory to display graphical information of a GUI on external input / output devices (such as display devices coupled to the interfaces). In some alternative implementations, multiple processors and / or multiple buses can be used with multiple memories and multiple memory modules, if desired. Similarly, multiple computer devices can be connected, each providing some of the necessary operations (e.g., as a server array, a group of blade servers, or a multiprocessor system). Figure 10 Take a processor 10 as an example.

[0114] Processor 10 may be a central processing unit, a network processor, or a combination thereof. Processor 10 may further include a hardware chip. The hardware chip may be an integrated circuit, a programmable logic device, or a combination thereof. The programmable logic device may be a complex programmable logic device, a field-programmable gate array (FPGA), a general-purpose array logic (GPA), or any combination thereof.

[0115] The memory 20 stores instructions executable by at least one processor 10 to cause at least one processor 10 to perform the method shown in the above embodiments.

[0116] The memory 20 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the computer device as shown by a landing page for an app. Furthermore, the memory 20 may include high-speed random access memory and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some alternative embodiments, the memory 20 may optionally include memory remotely located relative to the processor 10, which can be connected to the computer device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0117] The memory 20 may include volatile memory, such as random access memory; the memory may also include non-volatile memory, such as flash memory, hard disk or solid-state drive; the memory 20 may also include a combination of the above types of memory.

[0118] The computer device also includes an input device 30 and an output device 40. The processor 10, memory 20, input device 30, and output device 40 can be connected via a bus or other means. Figure 10 Taking the example of a connection between China and Israel via a bus.

[0119] Input device 30 can receive input numerical or character information, and generate key signal inputs related to user settings and function control of the computer device, such as a touchscreen, keypad, mouse, trackpad, touchpad, joystick, one or more mouse buttons, trackball, joystick, etc. Output device 40 may include display devices, auxiliary lighting devices (e.g., LEDs), and haptic feedback devices (e.g., vibration motors). The aforementioned display devices include, but are not limited to, liquid crystal displays, light-emitting diodes, displays, and plasma displays. In some alternative embodiments, the display device may be a touchscreen.

[0120] Example 5:

[0121] Based on the above embodiments, this application also provides a computer-readable storage medium storing a computer program executable by a processor. When the program runs on the processor, it causes the processor to perform the following steps:

[0122] For any voice data packet received in real time, determine the timestamp of the voice data packet's reception;

[0123] The elapsed time is determined based on the received timestamp and the total cumulative duration of the voice data packet and all previous voice data packets in the same voice stream.

[0124] The current network status is determined based on the comparison between the elapsed time and the preset time threshold; wherein the current network status includes normal network, slightly abnormal network, and severely abnormal network.

[0125] If the current network status is determined to be normal, then the voice data packet is directly processed for speech recognition.

[0126] If the current network status is determined to be a slight network anomaly, then output the prompt message corresponding to the network delay and perform speech recognition processing on the voice data packet;

[0127] If the current network state is determined to be severely abnormal, then the voice data packet is subjected to voice continuity judgment and detection to determine the cause of the abnormality of the current network state; wherein, the cause of the abnormality includes network latency and network packet loss; the prompt information corresponding to the cause of the abnormality of the current network state is output and the voice data packet is subjected to voice recognition processing.

[0128] Since the principle of the computer-readable storage medium in solving the problem is similar to that of a real-time speech recognition and network anomaly detection system, the implementation of the computer-readable storage medium can be found in Embodiment 3 of the method, and the repetition will not be repeated.

[0129] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.

Claims

1. A real-time speech recognition and network anomaly detection system, characterized in that, The system includes: The speech recognition module is configured to provide streaming speech recognition services; The voice timing module is configured to determine the reception timestamp of any voice data packet received in real time; determine the elapsed time based on the reception timestamp and the cumulative total duration of the voice data packet and all previous voice data packets in the same voice stream; determine the current network status based on the comparison result of the elapsed time and a preset duration threshold; wherein the current network status includes normal network, slightly abnormal network, and severely abnormal network; if the current network status is determined to be severely abnormal, the voice continuity judgment module is invoked. The voice continuity determination module is configured to provide a voice continuity determination detection service in response to the voice timing module, so as to determine the cause of the anomaly in the current network state; wherein the cause of the anomaly includes network latency and network packet loss; The control module is configured to, if it is determined that the current network status is normal, call the speech recognition module to process the speech data packet; otherwise, output the prompt information corresponding to the abnormal reason of the current network status and call the speech recognition module to process the speech data packet; wherein, the abnormal reason for a minor network abnormality is network latency.

2. The system according to claim 1, characterized in that, The voice timing module is specifically configured as follows: If the elapsed time is less than a first preset time threshold, then the current network status is determined to be normal. If the elapsed time is greater than or equal to the first preset time threshold and less than the second preset time threshold, then the current network state is determined to be a slight network anomaly. If the elapsed time is greater than or equal to the second preset time threshold, then the current network state is determined to be a severely abnormal network.

3. The system according to claim 1, characterized in that, The voice continuity determination module is specifically configured as follows: A preset continuity judgment algorithm is used to perform feature analysis on the current voice data packet and the previous voice data packet; Based on the feature analysis results, the cause of the anomaly in the current network state is determined.

4. The system according to claim 1, characterized in that, The control module is specifically configured as follows: If the abnormality of the current network status is due to network packet loss, then based on the voice data packet and the previous voice data packet, retrieve the available voice data from the cache; The speech recognition module is invoked to process the available speech data to generate alternative recognition results; The alternative identification result is output by displaying the preset highlighted mark.

5. The system according to claim 1, characterized in that, The voice timing module is specifically configured as follows: Determine the sequence number of the voice data packet in the voice stream; determine the total cumulative duration based on the sequence number and the preset duration of the voice data packet.

6. A real-time speech recognition and network anomaly detection method, characterized in that, The method includes: For any voice data packet received in real time, determine the timestamp of the voice data packet's reception; The elapsed time is determined based on the received timestamp and the total cumulative duration of the voice data packet and all previous voice data packets in the same voice stream. The current network status is determined based on the comparison between the elapsed time and the preset time threshold; wherein the current network status includes normal network, slightly abnormal network, and severely abnormal network. If the current network status is determined to be normal, then the voice data packet is directly processed for speech recognition. If the current network status is determined to be a slight network anomaly, then output the prompt message corresponding to the network delay and perform speech recognition processing on the voice data packet; If the current network state is determined to be severely abnormal, then the voice data packet is subjected to voice continuity judgment and detection to determine the cause of the abnormality of the current network state; wherein, the cause of the abnormality includes network latency and network packet loss; the prompt information corresponding to the cause of the abnormality of the current network state is output and the voice data packet is subjected to voice recognition processing.

7. The method according to claim 6, characterized in that, The step of determining the elapsed time based on the received timestamp and the cumulative total duration of the voice data packet and all previous voice data packets in the same voice stream includes: Determine the sequence number of the voice data packet in the voice stream; The total cumulative duration is determined based on the sequence number and the preset voice data packet duration.

8. The method according to claim 6, characterized in that, The step of performing voice continuity detection on the voice data packet to determine the cause of the current network state anomaly includes: A preset continuity judgment algorithm is used to perform feature analysis on the current voice data packet and the previous voice data packet; Based on the feature analysis results, the cause of the anomaly in the current network state is determined.

9. The method according to claim 6, characterized in that, If the abnormality of the current network status is due to network packet loss, the speech recognition processing of the voice data packet includes: Based on the current voice data packet and the previous voice data packet, extract available voice data from the cache; The speech recognition module is invoked to process the available speech data to generate alternative recognition results; The alternative identification result is output by displaying the preset highlighted mark.

10. A computer device, characterized in that, The computer device includes a processor that executes a computer program stored in a memory to implement the steps of the real-time speech recognition and network anomaly detection method as described in any one of claims 6-9.