An audio detection method and device, electronic equipment and storage medium

By detecting the transmission rate of audio data in the vehicle system and identifying abnormal transmission nodes, the problem of difficulty in locating audio stuttering in the vehicle system is solved, and the smoothness and stability of audio playback are improved.

CN119676513BActive Publication Date: 2025-11-18CHONGQING SELIS PHOENIX INTELLIGENT INNOVATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411610501.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-12
Publication Date
2025-11-18
Estimated Expiration
2044-11-12

AI Technical Summary

Technical Problem

In vehicle systems, due to the long audio transmission and processing links, audio stuttering issues are difficult to pinpoint, resulting in poor playback quality.

Method used

By acquiring the audio data packets sent by the client, the target node of the audio data from the client to the playback end is determined, the transmission time between each node is detected, the transmission rate is calculated, the transmission rate is compared with a preset threshold, and an anomaly detection log is generated to lock the abnormal transmission node.

Benefits of technology

It enables rapid and accurate location of audio stuttering issues, improves the accuracy and efficiency of troubleshooting, reduces stuttering and latency in audio playback, and ensures the smoothness and stability of audio playback.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119676513B_ABST
    Figure CN119676513B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide an audio detection method and device, electronic equipment and storage medium, comprising: obtaining an audio data packet sent by a client, determining a target node of audio data transmission from the client to a playback end, detecting a time of audio data transmission to the target node, obtaining a first transmission time length of audio data from the client to the server, a second transmission time length of audio data from the server to a decoder, and a third transmission time length of audio data from the decoder to the playback end, calculating a data length and the first transmission time length, the second transmission time length and the third transmission time length to obtain a transmission rate of the audio data between the target nodes, comparing the transmission rate between the target nodes with a preset rate threshold, determining an abnormal transmission node of the audio data, and generating an anomaly detection log. The present application detects the transmission rate of each transmission stage of the audio data, realizes fast and accurate locking of the abnormal node of the audio data transmission, and improves the troubleshooting accuracy of the audio lag.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of audio processing technology, specifically to an audio detection method, apparatus, electronic device, and storage medium. Background Technology

[0002] With the rapid development of internet technology, real-time audio streaming applications such as live streaming, voice calls, telephone calls, and screen mirroring are becoming increasingly popular. The core of these applications lies in the real-time transmission and playback of audio data to ensure that users can obtain a smooth audio experience.

[0003] Currently, most in-vehicle systems on the market use Android, which buffers audio data of a fixed duration before playing it. If the audio is insufficient, it is supplemented from the buffer. However, when the audio data in an Android in-vehicle system is blocked, delayed, or the data in the buffer is exhausted during transmission, audio stuttering occurs. Because the audio transmission and processing links are long, the causes of audio stuttering vary, and the causes of stuttering also vary in different scenarios. It is impossible to accurately detect at which stage the audio stuttering occurs, making it difficult to locate the audio stuttering problem and further affecting the audio playback effect. Summary of the Invention

[0004] In view of this, the present invention aims to provide an audio detection method, device, electronic device and storage medium to solve the problem that it is difficult to locate audio stuttering during real-time audio streaming in vehicle systems due to the long audio transmission and processing links.

[0005] According to a first aspect of the present invention, an audio detection method is provided, applied to a server, the server being connected to both a client and a playback end, the method comprising:

[0006] Obtain the audio data packet sent by the client; wherein the audio data packet includes pre-encoded audio data and data length;

[0007] The target node for transmitting the audio data from the client to the playback terminal is determined; the target node includes the client, server, decoder, and playback terminal.

[0008] The time it takes for the audio data to be transmitted to the target node is detected to obtain the first transmission time from the client to the server, the second transmission time from the server to the decoder, and the third transmission time from the decoder to the playback end.

[0009] The transmission rate of the audio data between the target nodes is calculated by combining the data length with the first transmission duration, the second transmission duration, and the third transmission duration.

[0010] The transmission rate between target nodes is compared with a preset rate threshold to identify abnormal audio data transmission nodes and generate an anomaly detection log.

[0011] Optionally, the step of detecting the time it takes for the audio data to be transmitted to the target node, and obtaining the first transmission time of the audio data from the client to the server, the second transmission time from the server to the decoder, and the third transmission time from the decoder to the playback end, includes:

[0012] The audio data packet is unpacked to obtain pre-encoded audio data and data packet information, wherein the data packet information includes data length and encapsulation time;

[0013] Determine the reception time of the audio data packet received by the server;

[0014] The first transmission duration of the audio data from the client to the server is calculated based on the encapsulation time and the reception time.

[0015] Optionally, the step of detecting the time it takes for the audio data to be transmitted to the target node, and obtaining the first transmission time of the audio data from the client to the server, the second transmission time from the server to the decoder, and the third transmission time from the decoder to the playback end, includes:

[0016] A decoder is used to decode the pre-encoded audio data, and the start time of decoding the audio data is detected.

[0017] Once the audio data decoding is complete, the decoding time of the audio data is detected.

[0018] The second transmission duration of the audio data from the server to the decoder is calculated from the start decoding time and the completion decoding time.

[0019] Optionally, the step of detecting the time it takes for the audio data to be transmitted to the target node, and obtaining the first transmission time of the audio data from the client to the server, the second transmission time from the server to the decoder, and the third transmission time from the decoder to the playback end, includes:

[0020] The decoded audio data is transmitted to the playback terminal for playback, and the start time of the audio data is detected.

[0021] The third transmission duration of the audio data from the decoder to the playback end is calculated based on the completion decoding time and the start playback time.

[0022] Optionally, the step of comparing the transmission rate between target nodes with a preset rate threshold to determine abnormal audio data transmission nodes and generating an anomaly detection log includes:

[0023] Obtain the pre-set transmission rate threshold between adjacent target nodes;

[0024] The transmission rate between target nodes is compared with the transmission rate threshold to determine whether the transmission rate is less than the transmission rate threshold.

[0025] If so, then at least one abnormal transmission node of the audio data is identified;

[0026] An anomaly detection log for the audio data is generated, and the anomaly detection log is pushed to the client.

[0027] Optionally, before comparing the transmission rate between target nodes with a preset rate threshold to determine abnormal audio data transmission nodes and generating an anomaly detection log, the method further includes:

[0028] The total transmission time of the audio data is obtained by summing the first transmission time, the second transmission time, and the third transmission time.

[0029] The total transmission time is compared with the predetermined standard transmission time of the audio data to determine whether the audio data transmission is abnormal.

[0030] Optionally, comparing the total transmission time with the predetermined standard transmission time of the audio data to determine whether the audio data transmission is abnormal includes:

[0031] The unit transmission data volume of the audio data is obtained in advance, and the standard transmission duration of the audio data is calculated using the data length of the audio data and the unit transmission data volume.

[0032] The total transmission time of the audio data is compared with the standard transmission time;

[0033] If the total transmission duration exceeds the standard transmission duration, then the audio data transmission is determined to be abnormal.

[0034] According to a second aspect of the present invention, an audio detection device is provided, applied to a server, the server being connected to a client and a playback end respectively, the device comprising:

[0035] An audio acquisition module is used to acquire audio data packets sent by the client; wherein the audio data packets include pre-encoded audio data and data length;

[0036] The node determination module is used to determine the target node from which the audio data is transmitted from the client to the playback terminal; the target node includes the client, server, decoder, and playback terminal;

[0037] The transmission detection module is used to detect the time it takes for the audio data to be transmitted to the target node, and to obtain the first transmission time of the audio data from the client to the server, the second transmission time from the server to the decoder, and the third transmission time from the decoder to the playback end.

[0038] A rate determination module is used to calculate the transmission rate of the audio data between target nodes by combining the data length with the first transmission duration, the second transmission duration, and the third transmission duration.

[0039] The anomaly detection module is used to compare the transmission rate between target nodes with a preset rate threshold, identify abnormal transmission nodes of audio data, and generate anomaly detection logs.

[0040] According to another aspect of the present invention, an electronic device is also provided, comprising:

[0041] processor;

[0042] Memory used to store the processor's executable instructions;

[0043] The processor is configured to execute the instructions to implement the audio detection method described above.

[0044] According to another aspect of the present invention, a readable storage medium is also provided, on which a computer program is stored, which, when executed by a processor, implements the steps of the audio detection method as described above.

[0045] The audio detection method provided in this invention obtains audio data packets sent by the client, determines the target node from the client to the playback end, detects the time it takes for the audio data to travel to the target node, and obtains the first transmission time from the client to the server, the second transmission time from the server to the decoder, and the third transmission time from the decoder to the playback end. The transmission rate of the audio data between the target nodes is calculated by comparing the data length with the first, second, and third transmission times. The transmission rate between the target nodes is compared with a preset rate threshold to identify abnormal transmission nodes and generate an anomaly detection log. This invention, by detecting the transmission rate of audio data at each stage from the client to the playback end, can accurately detect abnormal transmission nodes, quickly locate which transmission node is blocked or malfunctioning, thereby narrowing down the scope of audio stuttering problems. By monitoring the audio data transmission process in real time, it can promptly and accurately pinpoint abnormal audio data transmission nodes, achieving automatic testing and anomaly detection of audio playback time, further improving the accuracy and efficiency of troubleshooting audio stuttering problems, effectively reducing stuttering and latency in audio playback, and ensuring the smoothness and stability of audio playback.

[0046] The above description is merely an overview of the technical solution of the present invention. In order to better understand the technical means of the present invention and to implement it in accordance with the contents of the specification, and in order to make the above and other objects, features and advantages of the present invention more apparent and understandable, specific embodiments of the present invention are described below. Attached Figure Description

[0047] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings:

[0048] Figure 1 This is a flowchart of the steps of an audio detection method provided in an embodiment of the present invention;

[0049] Figure 2 yes Figure 1 The flow of step 103 in the audio detection method provided in this embodiment of the invention Figure 1 ;

[0050] Figure 3 yes Figure 1 The flow of step 103 in the audio detection method provided in this embodiment of the invention Figure 2 ;

[0051] Figure 4 yes Figure 1 The flow of step 103 in the audio detection method provided in this embodiment of the invention Figure 3 ;

[0052] Figure 5 yes Figure 1 A flowchart of step 105 in the audio detection method provided in this embodiment of the invention;

[0053] Figure 6 This is a flowchart of another audio detection method provided in an embodiment of the present invention;

[0054] Figure 7 This is a schematic diagram of a scenario for the audio detection method provided in an embodiment of the present invention;

[0055] Figure 8 This is a schematic diagram of the structure of an audio detection device provided in an embodiment of the present invention;

[0056] Figure 9 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0057] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the various embodiments of the present invention will be described in detail below with reference to the accompanying drawings. However, those skilled in the art will understand that many technical details are presented in the various embodiments of the present invention to facilitate a better understanding of this application. However, the technical solutions claimed in this application can be implemented even without these technical details and various changes and modifications based on the following embodiments. The division of the various embodiments below is for ease of description and should not constitute any limitation on the specific implementation of the present invention. The various embodiments can be combined with and referenced by each other without contradiction.

[0058] Reference Figure 1 This diagram illustrates a flowchart of the audio detection method provided in an embodiment of the present invention. The method is applied to a server, which is connected to both a client and a playback terminal. The method may include:

[0059] Step 101: Obtain the audio data packet sent by the client; wherein the audio data packet includes pre-encoded audio data and data length.

[0060] In this embodiment of the invention, to address the difficulty in locating audio stuttering issues during real-time audio streaming in in-vehicle systems due to the long audio transmission and processing links, the server detects the audio transmission rate from the start of audio data transmission from the client to the end of playback on the playback end. This allows for timely and accurate identification of abnormal nodes in audio data transmission, enabling automatic testing and anomaly detection of audio playback time. This further improves the accuracy and efficiency of troubleshooting audio stuttering issues, effectively reduces stuttering and latency in audio playback, and ensures the smoothness and stability of audio playback.

[0061] Reference Figure 7This illustration shows a scenario diagram of the audio detection method provided by an embodiment of the present invention. Specifically, the encoder generates encoded audio data, the client is responsible for collecting pre-encoded audio data, encapsulating the audio data into audio data packets, and sending the audio data packets to the server. The server receives the audio data packets sent by the client and forwards the data packets to the decoder. The decoder receives the audio data packets forwarded by the server, performs decoding processing, and sends the decoded data to the playback end. The playback end receives the audio data sent by the decoder and plays it. Therefore, playing audio data from the client to the playback end involves a first transmission stage (from the encoder to the client), a second transmission stage (from the client to the server), a third transmission stage (from the server to the decoder), and a fourth transmission stage (from the decoder to the playback end). Since the encoder outputs a fixed amount of audio data per second according to the audio parameters, the transmission time from the encoder to the client is negligible. This embodiment calculates the transmission rates from the client to the server, from the server to the decoder, and from the decoder to the playback end and determines the abnormal transmission stages to accurately locate the audio stuttering problem nodes.

[0062] Specifically, the server obtains the audio data packets sent by the client. The audio data packets include pre-encoded audio data and data length. The audio data packets are obtained by the client by encapsulating the audio data encoded by the encoder with data packet information, which will not be elaborated here.

[0063] Step 102: Determine the target node from which the audio data is transmitted from the client to the playback end; the target node includes the client, server, decoder, and playback end.

[0064] Step 103: Detect the time it takes for the audio data to be transmitted to the target node, and obtain the first transmission time of the audio data from the client to the server, the second transmission time from the server to the decoder, and the third transmission time from the decoder to the playback end.

[0065] In this embodiment of the invention, by detecting the time when the audio data is transmitted to the target node, the transmission duration between adjacent target nodes is determined based on the time difference of the audio data arriving at the target node. Specifically, the first transmission duration of the audio data from the client to the server, the second transmission duration from the server to the decoder, and the third transmission duration from the decoder to the playback end are obtained.

[0066] Step 104: Calculate the transmission rate of audio data between target nodes by combining the data length with the first transmission duration, the second transmission duration, and the third transmission duration.

[0067] In this embodiment of the invention, the ratios of data length to a first transmission duration, data length to a second transmission duration, and data length to a third transmission duration are calculated respectively to determine the transmission rate of audio data between target nodes, including a first transmission rate from the client to the server, a second transmission rate from the server to the decoder, and a third transmission rate from the decoder to the playback end.

[0068] Step 105: Compare the transmission rate between target nodes with the preset rate threshold to identify abnormal transmission nodes of audio data and generate an anomaly detection log.

[0069] In this embodiment of the invention, a transmission rate threshold between adjacent target nodes is preset. The transmission rate between target nodes is compared with the transmission rate threshold to determine whether the transmission rate is less than the threshold. If so, at least one abnormal transmission node of audio data is identified, thereby generating an audio data anomaly detection log. The anomaly detection log is pushed to the client for developers to view, which can promptly and accurately locate abnormal nodes in audio data transmission. This realizes automatic testing and anomaly detection of audio playback time, improving the accuracy and efficiency of troubleshooting audio stuttering problems.

[0070] The audio detection method provided in this invention obtains audio data packets sent by the client, determines the target node from the client to the playback end, detects the time it takes for the audio data to travel to the target node, and obtains the first transmission time from the client to the server, the second transmission time from the server to the decoder, and the third transmission time from the decoder to the playback end. The transmission rate of the audio data between the target nodes is calculated by comparing the data length with the first, second, and third transmission times. The transmission rate between the target nodes is compared with a preset rate threshold to identify abnormal transmission nodes and generate an anomaly detection log. This invention, by detecting the transmission rate of audio data at each stage from the client to the playback end, can accurately detect abnormal transmission nodes, quickly locate which transmission node is blocked or malfunctioning, thereby narrowing down the scope of audio stuttering problems. By monitoring the audio data transmission process in real time, it can promptly and accurately pinpoint abnormal audio data transmission nodes, achieving automatic testing and anomaly detection of audio playback time, further improving the accuracy and efficiency of troubleshooting audio stuttering problems, effectively reducing stuttering and latency in audio playback, and ensuring the smoothness and stability of audio playback.

[0071] Furthermore, refer to Figure 2 , showed Figure 1 The process of step 103 in the provided audio detection method Figure 1 This method is basically the same as the audio detection method provided in the first embodiment of the present invention, and step 103 may include:

[0072] Step 201: Unpack the audio data packets to obtain pre-encoded audio data and data packet information, including data length and encapsulation time.

[0073] Step 202: Determine the reception time of the audio data packets received by the server.

[0074] Step 203: Calculate the first transmission time of audio data from the client to the server based on the encapsulation time and the reception time.

[0075] In this embodiment of the invention, the client is responsible for collecting audio data, encapsulating the audio data into audio data packets, and sending the audio data packets to the server. The server receives the audio data packets sent by the client and processes them. The client encapsulates the audio data. Each audio data packet contains pre-encoded audio data and data packet information. The data packet information includes the data length and encapsulation time. The encapsulated audio data packet adds data packet information to the pre-encoded audio data. The data packet information includes a data header, data length, and encapsulation time. The data header is a marker header for the audio data, used to distinguish different audio data types. For example, the header of call audio data can be marked as 0x1, the header of media audio data can be marked as 0x2, and so on. Assuming that the audio type to be transmitted is media audio, the data header is 0x2. The data length is the length of the audio data. In this embodiment, the data length is 10240 bytes as an example. The encapsulation time is the current Greenwich Mean Time (referring to the standard time of the Royal Greenwich Observatory in the suburbs of London, England). The data header, data length, and encapsulation time each occupy 4 bytes. The size of the audio data depends on the size of the encoded output and is not specifically limited here.

[0076] It should be noted that audio data is audio in a certain audio format generated by an encoder. For example, the original audio generated by a phone recording in Android is generally in PCM format. The encoder will encode the PCM audio into audio formats such as WAV and MP3. The phone outputs audio evenly according to a fixed time axis, and outputs a certain amount of audio data per second according to the audio parameters. Therefore, the transmission time from the encoder to the client is negligible.

[0077] Specifically, after receiving the audio data packet, the server unpacks it to obtain pre-encoded audio data and data packet information, namely, the pre-encoded audio data, data length, and encapsulation time. When the server receives the audio data packet, it records the current time as the reception time and calculates the time difference between the encapsulation time and the reception time to obtain the first transmission time of the audio data from the client to the server. Data transmission from the client to the server can be performed using either TCP / IP or USB. This embodiment proposes to use TCP / IP. Since the client and server reside in two different processes, it is necessary to determine the transmission time from the client to the server.

[0078] This invention obtains the encapsulation time through unpacking operations and records the reception time, which can accurately calculate the transmission time of audio data from the client to the server. By recording and calculating the transmission time in real time, the transmission process of audio data can be monitored in real time. By recording the transmission time of each stage, it helps to accurately evaluate the transmission efficiency of audio data from the client to the server, so as to quickly locate and troubleshoot problems in the transmission process.

[0079] Furthermore, refer to Figure 3 , showed Figure 1 The process of step 103 in the provided audio detection method Figure 2 This method is basically the same as the audio detection method provided in the first embodiment of the present invention, and step 103 may include:

[0080] Step 301: Use a decoder to decode the pre-encoded audio data and detect the start time of the audio data decoding.

[0081] Step 302: The audio data decoding is detected as complete, and the audio data decoding time is detected.

[0082] Step 303: Calculate the second transmission time of audio data from the server to the decoder based on the start decoding time and the completion decoding time.

[0083] It should be noted that in this embodiment of the invention, the server receives the audio data packet sent by the client and sends the pre-encoded audio data to the decoder. The decoder receives the pre-encoded audio data sent by the server and performs decoding processing. Specifically, after receiving the pre-encoded audio data, the decoder starts the decoding operation and records the current time as the start decoding time. After the decoder completes the decoding of the audio data, it records the current time as the completion decoding time. The time difference between the start decoding time and the completion decoding time is calculated to obtain the second transmission time of the audio data from the server to the decoder.

[0084] By recording the start and end times of decoding, this invention can accurately calculate the transmission time of audio data from the server to the decoder. Real-time recording and calculation of decoding time helps to accurately assess the transmission efficiency of audio data from the server to the decoder, and facilitates quick location and troubleshooting of problems in the decoding process.

[0085] Furthermore, refer to Figure 4 , showed Figure 1 The process of step 103 in the provided audio detection method Figure 3 This method is basically the same as the audio detection method provided in the first embodiment of the present invention, and step 103 may include:

[0086] Step 401: Transmit the decoded audio data to the playback terminal for playback and detect the start time of the audio data playback.

[0087] Step 402: Calculate the third transmission time of audio data from the decoder to the playback end based on the completion decoding time and the start playback time.

[0088] It should be noted that in this embodiment of the invention, the decoder receives pre-encoded audio data sent by the server and performs decoding processing. The server sends the decoded audio data to the playback end, which receives the audio data sent by the decoder and plays it. Specifically, the decoded audio data is transmitted to the playback end, which starts playing the audio data and records the current time as the start time. The time difference between the completion time of decoding and the start time of playback is calculated to obtain the third transmission duration of the audio data from the decoder to the playback end. The playback end can be AudioTrack, a class provided by the Android platform for audio data playback. AudioTrack allows developers to directly write audio data to audio hardware for playback, and is suitable for in-vehicle infotainment system audio playback scenarios requiring low latency and high efficiency. AudioTrack is typically used to play PCM format audio data.

[0089] In this embodiment, the audio data in PCM format is played from the decoder to the AudioTrack player. The decoder needs to decode the audio data from WAV format to PCM format first, and then transmit the PCM format audio data to the player for playback. Therefore, it is necessary to determine the transmission time from the decoder to the player.

[0090] This invention, by recording the start time of playback and the completion time of decoding, accurately calculates the transmission time of audio data from the decoder to the playback end. This helps to accurately assess the transmission efficiency of audio data from the decoder to the playback end, and facilitates quick location and troubleshooting of problems during playback.

[0091] Furthermore, refer to Figure 5 , showed Figure 1 A flowchart of step 105 in an audio detection method is provided. This method is basically the same as the audio detection method provided in the first embodiment of the present invention. Step 105 may include:

[0092] Step 501: Obtain the pre-set transmission rate threshold between adjacent target nodes.

[0093] Step 502: Compare the transmission rate between target nodes with the transmission rate threshold to determine whether the transmission rate is less than the transmission rate threshold.

[0094] Step 503: If yes, then at least one abnormal transmission node of the audio data is identified.

[0095] Step 504: Generate an anomaly detection log for the audio data and push the anomaly detection log to the client.

[0096] It should be noted that in this embodiment of the invention, the client is responsible for collecting audio data and sending the audio data packets to the server. At the same time, the client receives and displays the anomaly detection log. The server receives the audio data packets sent by the client and forwards the data packets to the decoder. The decoder receives the audio data packets forwarded by the server, performs decoding processing, and sends the decoded data to the playback end. The playback end receives the audio data sent by the decoder and plays it.

[0097] Specifically, the system obtains a pre-set transmission rate threshold between adjacent target nodes. This threshold is the minimum rate at which audio data is transmitted between nodes, and is typically set based on system performance and network environment. It is used to determine whether the transmission rate is normal. For example, the transmission rate threshold from the client to the server is 1024 bytes / ms, the transmission rate threshold from the server to the decoder is 2048 bytes / ms, and the transmission rate threshold from the decoder to the playback end is 1024 bytes / ms. If the transmission rate in this stage exceeds the transmission rate threshold between target nodes, at least one abnormal transmission node of the audio data is identified, and an abnormal detection log of the audio data is generated.

[0098] Specifically, the calculated transmission rate between target nodes is compared with a preset transmission rate threshold to determine if the transmission rate is less than the threshold. If the transmission rate is less than the threshold, at least one abnormal transmission node of the audio data is identified. The abnormal transmission node can be one or more of the client, server, decoder, or playback end. Based on the abnormal transmission node, the abnormal transmission stage is determined, and an abnormal detection log of the audio data is generated, recording information such as the abnormal transmission node, abnormal time, and abnormal transmission rate. The generated abnormal detection log is pushed to the client, which receives and displays the abnormal detection log for developers to view and analyze.

[0099] This invention compares the transmission rate between target nodes with a preset transmission rate threshold to accurately detect abnormal audio data transmission. Once the transmission rate is found to be below the threshold, the abnormal transmission node is immediately locked and an anomaly detection log is generated, enabling rapid location and troubleshooting of problems in the audio transmission process. This further improves the accuracy and efficiency of audio stuttering anomaly detection and helps reduce stuttering and latency in audio playback.

[0100] Reference Figure 6 The diagram illustrates a flowchart of another audio detection method provided by an embodiment of the present invention. This method is basically the same as the audio detection method provided by the first embodiment of the present invention, except that the method may further include:

[0101] Step 101: Obtain the audio data packet sent by the client; wherein the audio data packet includes pre-encoded audio data and data length.

[0102] Step 102: Determine the target node from which the audio data is transmitted from the client to the playback end; the target node includes the client, server, decoder, and playback end.

[0103] Step 103: Detect the time it takes for the audio data to be transmitted to the target node, and obtain the first transmission time of the audio data from the client to the server, the second transmission time from the server to the decoder, and the third transmission time from the decoder to the playback end.

[0104] Step 104: Calculate the transmission rate of audio data between target nodes by combining the data length with the first transmission duration, the second transmission duration, and the third transmission duration.

[0105] Step 106: Sum the first transmission duration, the second transmission duration, and the third transmission duration to obtain the total transmission duration of the audio data.

[0106] In this embodiment of the invention, the client is responsible for collecting audio data and sending the audio data packets to the server. The server receives the audio data packets sent by the client and forwards them to the decoder. The decoder receives the audio data packets forwarded by the server, performs decoding processing, and sends the decoded data to the playback end. The playback end receives the audio data sent by the decoder and plays it. This embodiment divides the target nodes according to the above transmission process and calculates the transmission time between the target nodes. The first transmission time (time from the client to the server), the second transmission time (time from the server to the decoder's decoding completion), and the third transmission time (time from the decoder to the playback end) are summed to generate the total transmission time of the audio data.

[0107] Step 107: Compare the total transmission time with the predetermined standard transmission time of the audio data to determine whether there is a transmission abnormality in the audio data.

[0108] In this embodiment of the invention, the total transmission time is compared with a predetermined standard transmission time for audio data to determine whether the audio data transmission is abnormal. The standard transmission time is calculated based on the size of the audio data and the standard amount of data transmitted per second. By comparing the total transmission time with the predetermined standard transmission time for audio data, it is possible to directly determine whether the audio data transmission is abnormal. In the case of transmission abnormality, the transmission rate between target nodes is compared with a preset rate threshold to identify the abnormal transmission node of the audio data and generate an abnormality detection log.

[0109] Step 105: Compare the transmission rate between target nodes with the preset rate threshold to identify abnormal transmission nodes of audio data and generate an anomaly detection log.

[0110] Steps 101 to 105 described above are the same as those previously mentioned and will not be repeated here.

[0111] Compared with the prior art, the embodiments of the present invention, while achieving the beneficial effects of the first embodiment, can monitor audio data transmission anomalies in real time by comparing the total transmission time with the preset standard transmission time. Once the total transmission time is found to exceed the standard transmission time, it can be immediately determined that the audio data transmission is abnormal, thereby reducing manual intervention and improving the accuracy and efficiency of anomaly detection.

[0112] Specifically, step 107 compares the total transmission time with the predetermined standard transmission time of the audio data to determine whether there is a transmission abnormality in the audio data. This may include:

[0113] First, the unit transmission data volume of the audio data is obtained in advance, and the standard transmission duration of the audio data is calculated using the data length of the audio data and the unit transmission data volume.

[0114] Secondly, the total transmission time of the audio data is compared with the standard transmission time.

[0115] Secondly, if the total transmission time exceeds the standard transmission time, it is determined that there is a transmission abnormality in the audio data.

[0116] It should be noted that in the above steps of this embodiment, the unit transmission data volume of the audio data is obtained in advance, and the standard transmission duration of the audio data is calculated by using the data length of the audio data and the unit transmission data volume. The size of the audio data is the "audio data" field in the data packet, which is 10240 bytes. The amount of data transmitted per unit is the product of the sampling rate, the number of channels, and the sampling bit depth. The sampling rate is the number of audio samples per second, typically 16,000 bytes, 36,000 bytes, or 44,100 bytes. Here, we'll use a sampling rate of 44,100 bytes as an example. The number of channels is the number of audio channels, typically mono or stereo. Here, we'll use stereo as an example. The sampling bit depth is the size of each sample, typically 8 bits, 16 bits, or 32 bits. Here, we'll assume it's 16 bits. The standard audio playback time is the ratio of the audio data length to the amount of data transmitted per unit, approximately 58 milliseconds. That is, 10,240 bytes of audio data requires 58 milliseconds to play. Comparing the total transmission time of the audio data with the standard transmission time, if the total transmission time is longer than the standard transmission time, an audio data transmission anomaly is determined. In other words, if the next set of audio data does not arrive within 58 milliseconds, it indicates that the audio playback is stuttering.

[0117] This invention enables real-time monitoring and detection of audio data transmission anomalies, allowing for timely discovery and handling of audio stuttering issues. It achieves real-time monitoring and judgment of audio data transmission anomalies, facilitating automated anomaly detection and problem troubleshooting.

[0118] Reference Figure 8 The diagram illustrates a structural schematic of an audio detection device according to an embodiment of the present invention, applied to a server. The server is connected to both a client and a playback terminal. The device includes:

[0119] An audio acquisition module 601 is used to acquire audio data packets sent by the client; wherein the audio data packets include pre-encoded audio data and data length;

[0120] The node determination module 602 is used to determine the target node from which the audio data is transmitted from the client to the playback terminal; the target node includes the client, server, decoder, and playback terminal;

[0121] The transmission detection module 603 is used to detect the time when the audio data is transmitted to the target node, and to obtain the first transmission time of the audio data from the client to the server, the second transmission time from the server to the decoder, and the third transmission time from the decoder to the playback end.

[0122] The rate determination module 604 is used to calculate the transmission rate of the audio data between target nodes by combining the data length with the first transmission duration, the second transmission duration, and the third transmission duration.

[0123] The anomaly detection module 605 is used to compare the transmission rate between target nodes with a preset rate threshold, identify abnormal transmission nodes of audio data, and generate anomaly detection logs.

[0124] Furthermore, the transmission detection module 603 includes:

[0125] The processing submodule is used to unpack the audio data packets to obtain pre-encoded audio data and data packet information, wherein the data packet information includes data length and encapsulation time;

[0126] The first determining submodule is used to determine the reception time of the audio data packet received by the server;

[0127] The first calculation submodule is used to calculate the first transmission duration of the audio data from the client to the server based on the encapsulation time and the reception time.

[0128] Furthermore, the transmission detection module 603 includes:

[0129] The first detection submodule is used to decode the pre-encoded audio data using a decoder and detect the start time of the audio data decoding.

[0130] The second detection submodule is used to detect when the audio data decoding is completed and to detect the completion time of the audio data decoding.

[0131] The second calculation submodule is used to calculate the second transmission duration of the audio data from the server to the decoder based on the start decoding time and the completion decoding time.

[0132] Furthermore, the transmission detection module 603 includes:

[0133] The third detection submodule is used to transmit the decoded audio data to the playback terminal for playback and to detect the start time of the audio data.

[0134] The third calculation submodule is used to calculate the third transmission duration of the audio data from the decoder to the playback end based on the completion decoding time and the start playback time.

[0135] Optionally, the anomaly detection module 605 includes:

[0136] The acquisition submodule is used to acquire a pre-set transmission rate threshold between adjacent target nodes;

[0137] The comparison submodule is used to compare the transmission rate between target nodes with the transmission rate threshold to determine whether the transmission rate is less than the transmission rate threshold.

[0138] The second determining submodule is used to determine at least one abnormal transmission node of the audio data if the condition is met.

[0139] A generation submodule is used to generate anomaly detection logs for the audio data and push the anomaly detection logs to the client.

[0140] Furthermore, the device also includes:

[0141] The total duration determination module is used to sum the first transmission duration, the second transmission duration, and the third transmission duration to obtain the total transmission duration of the audio data;

[0142] The anomaly detection module is used to compare the total transmission time with the predetermined standard transmission time of the audio data to determine whether the audio data has a transmission anomaly.

[0143] Furthermore, the anomaly detection module includes:

[0144] The fourth calculation submodule is used to pre-obtain the unit transmission data volume of the audio data, and calculate the standard transmission duration of the audio data using the data length of the audio data and the unit transmission data volume;

[0145] The comparison submodule is used to compare the total transmission time of the audio data with the standard transmission time;

[0146] An anomaly detection submodule is used to determine that the audio data has a transmission anomaly if the total transmission duration exceeds the standard transmission duration.

[0147] The audio detection device provided in this invention obtains audio data packets sent by the client, determines the target node from the client to the playback end, detects the time it takes for the audio data to travel to the target node, and obtains the first transmission time from the client to the server, the second transmission time from the server to the decoder, and the third transmission time from the decoder to the playback end. The data length is then used to calculate the transmission rate of the audio data between the target nodes using the first, second, and third transmission times. This transmission rate between the target nodes is compared with a preset rate threshold to identify abnormal transmission nodes and generate an anomaly detection log. By detecting the transmission rate of audio data at each stage from the client to the playback end, this invention can accurately detect abnormal transmission nodes, quickly pinpointing which transmission node is blocked or malfunctioning, thus narrowing down the scope of audio stuttering problems. Through real-time monitoring of the audio data transmission process, it can promptly and accurately locate abnormal audio data transmission nodes, achieving automatic testing and anomaly detection of audio playback time, further improving the accuracy and efficiency of troubleshooting audio stuttering problems, effectively reducing stuttering and latency in audio playback, and ensuring the smoothness and stability of audio playback.

[0148] Reference Figure 9 The present invention also provides an electronic device, such as... Figure 9 As shown, it includes a processor 701, a communication interface 702, a memory 703, and a communication bus 704, wherein the processor 701, the communication interface 702, and the memory 703 communicate with each other through the communication bus 704.

[0149] Processor 701, memory 703 for storing processor-executable instructions;

[0150] The processor 701 is configured to execute the instructions to implement the audio detection method as described below:

[0151] Obtain the audio data packet sent by the client; wherein the audio data packet includes pre-encoded audio data and data length;

[0152] The target node for transmitting the audio data from the client to the playback terminal is determined; the target node includes the client, server, decoder, and playback terminal.

[0153] The time it takes for the audio data to be transmitted to the target node is detected to obtain the first transmission time from the client to the server, the second transmission time from the server to the decoder, and the third transmission time from the decoder to the playback end.

[0154] The transmission rate of the audio data between the target nodes is calculated by combining the data length with the first transmission duration, the second transmission duration, and the third transmission duration.

[0155] The transmission rate between target nodes is compared with a preset rate threshold to identify abnormal audio data transmission nodes and generate an anomaly detection log.

[0156] The communication bus mentioned above can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used to represent it in the diagram, but this does not mean that there is only one bus or one type of bus.

[0157] The communication interface is used for communication between the aforementioned terminal and other devices.

[0158] The memory may include random access memory (RAM) or non-volatile memory, such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.

[0159] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0160] In another embodiment of the present invention, a computer-readable storage medium is also provided, on which a computer program is stored, which, when executed by a processor, implements any of the audio detection methods described in the above embodiments.

[0161] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid state disk (SSD)).

[0162] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0163] The various embodiments in this specification are described in a related manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0164] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention are included within the scope of protection of the present invention.

Claims

1. An audio detection method, characterized in that, Applied to the server side, wherein the server is connected to both the client and the playback terminal, the method includes: Obtain the audio data packet sent by the client; wherein the audio data packet includes pre-encoded audio data and data length; The target node for transmitting the audio data from the client to the playback terminal is determined; the target node includes the client, server, decoder, and playback terminal. The time it takes for the audio data to be transmitted to the target node is detected to obtain the first transmission time from the client to the server, the second transmission time from the server to the decoder, and the third transmission time from the decoder to the playback end. The transmission rate of the audio data between the target nodes is calculated by combining the data length with the first transmission duration, the second transmission duration, and the third transmission duration. The transmission rate between target nodes is compared with a preset rate threshold to identify abnormal audio data transmission nodes and generate an anomaly detection log.

2. The method according to claim 1, characterized in that, The process of detecting the time it takes for the audio data to be transmitted to the target node, and obtaining the first transmission time of the audio data from the client to the server, the second transmission time from the server to the decoder, and the third transmission time from the decoder to the playback end, includes: The audio data packet is unpacked to obtain pre-encoded audio data and data packet information, wherein the data packet information includes data length and encapsulation time; Determine the reception time of the audio data packet received by the server; The first transmission duration of the audio data from the client to the server is calculated based on the encapsulation time and the reception time.

3. The method according to claim 1, characterized in that, The process of detecting the time it takes for the audio data to be transmitted to the target node, and obtaining the first transmission time of the audio data from the client to the server, the second transmission time from the server to the decoder, and the third transmission time from the decoder to the playback end, includes: A decoder is used to decode the pre-encoded audio data, and the start time of decoding the audio data is detected. Once the audio data decoding is complete, the decoding time of the audio data is detected. The second transmission duration of the audio data from the server to the decoder is calculated based on the start decoding time and the completion decoding time.

4. The method according to claim 1, characterized in that, The process of detecting the time it takes for the audio data to be transmitted to the target node, and obtaining the first transmission time of the audio data from the client to the server, the second transmission time from the server to the decoder, and the third transmission time from the decoder to the playback end, includes: The decoded audio data is transmitted to the playback terminal for playback, and the start time of the audio data is detected. The third transmission duration of the audio data from the decoder to the playback end is calculated based on the completion decoding time and the start playback time.

5. The method according to claim 1, characterized in that, The step of comparing the transmission rate between target nodes with a preset rate threshold to determine abnormal audio data transmission nodes and generating an anomaly detection log includes: Obtain the pre-set transmission rate threshold between adjacent target nodes; The transmission rate between target nodes is compared with the transmission rate threshold to determine whether the transmission rate is less than the transmission rate threshold. If so, then at least one abnormal transmission node of the audio data is identified; An anomaly detection log for the audio data is generated, and the anomaly detection log is pushed to the client.

6. The method according to claim 1, characterized in that, Before comparing the transmission rate between target nodes with a preset rate threshold to determine abnormal audio data transmission nodes and generating an anomaly detection log, the process also includes: The total transmission time of the audio data is obtained by summing the first transmission time, the second transmission time, and the third transmission time. The total transmission time is compared with the predetermined standard transmission time of the audio data to determine whether the audio data transmission is abnormal.

7. The method according to claim 6, characterized in that, The step of comparing the total transmission time with the predetermined standard transmission time of the audio data to determine whether the audio data transmission is abnormal includes: The unit transmission data volume of the audio data is obtained in advance, and the standard transmission duration of the audio data is calculated using the data length of the audio data and the unit transmission data volume. The total transmission time of the audio data is compared with the standard transmission time; If the total transmission duration exceeds the standard transmission duration, then the audio data transmission is determined to be abnormal.

8. An audio detection device, characterized in that, The device is applied to a server, which is connected to both the client and the playback terminal. The device includes: An audio acquisition module is used to acquire audio data packets sent by the client; wherein the audio data packets include pre-encoded audio data and data length; The node determination module is used to determine the target node from which the audio data is transmitted from the client to the playback terminal; the target node includes the client, server, decoder, and playback terminal; The transmission detection module is used to detect the time it takes for the audio data to be transmitted to the target node, and to obtain the first transmission time of the audio data from the client to the server, the second transmission time from the server to the decoder, and the third transmission time from the decoder to the playback end. A rate determination module is used to calculate the transmission rate of the audio data between target nodes by combining the data length with the first transmission duration, the second transmission duration, and the third transmission duration. The anomaly detection module is used to compare the transmission rate between target nodes with a preset rate threshold, identify abnormal transmission nodes of audio data, and generate anomaly detection logs.

9. An electronic device, characterized in that, It includes a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; The memory is used to store computer programs; When the processor executes the program stored in the memory, it implements the audio detection method as described in any one of claims 1-7.

10. A computer-readable storage medium having instructions stored thereon that, when executed by one or more processors, cause the processors to perform the audio detection method as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Method and device for processing play abnormality of audio stream, computer device and computer readable storage medium

    CN107566890A

  • Audio transmission delay test method and device, storage medium and electronic equipment

    CN110430104A