A voice data transmission method and system based on cloud platform
By calculating semantic importance and loss rate before voice data transmission and optimizing the data transmission path, the semantic loss problem during voice data transmission is solved, ensuring accurate data transmission and user experience in multi-party interactive scenarios.
Patent Information
- Application Number
- CN202510697069.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2045-05-28
AI Technical Summary
The prior art has the problem of semantic loss in the transmission of voice data. The speech speed detection results cannot reflect the timing distribution of semantic information in the speech data, resulting in semantic loss in the speech data during transmission.
By converting the speech data into text data and inputting the semantic extraction model, erasing the audio information of any time frame in the speech data, calculating semantic importance, and optimizing data transmission on the communication link based on the semantic loss rate and retransmission priority, obtaining the final received data to avoid semantic loss.
It realizes accurate transmission of voice data in multi-party interactive scenarios, avoids semantic loss of voice data during transmission, ensures that the receiving node can receive data containing semantic information at the same time, and improves the user experience.
Smart Images

Figure CN120238247B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of voice transmission technology, and in particular to a method and system for voice transmission based on a cloud platform. Background Art
[0002] With the continuous development of artificial intelligence technology, online voice interaction technology is becoming more and more mature. Online meetings, online education, or online sales all require online voice interaction, and the interaction process involves the transmission of voice data.
[0003] At present, the patent application document with application publication number CN116996622A discloses a method, apparatus, device, medium and program product for transmitting voice data, wherein the method includes: real-time collection of first voice data, where the first voice data is voice audio data to be transmitted to a second terminal in real time; performing speech rate detection on the first voice data to obtain a speech rate detection result corresponding to the first voice data, where the speech rate detection result is used to characterize the density distribution of voice expressions in the first voice data in time series; determining a target transmission link from multiple candidate transmission link configurations based on the speech rate detection result; and transmitting the first voice data from the first terminal to the second terminal through the target transmission link; wherein the speech rate detection result is used to characterize the density distribution of voice expressions in the first voice data in time series.
[0004] The above method performs speech rate detection on the first voice data and determines the target transmission link based on the speech rate detection result, thereby reducing the probability of packet loss of voice audio data during transmission. However, the purpose of voice data transmission is to enable the receiver to understand the semantic information in the voice data, and the speech rate detection result cannot reflect the temporal distribution of the semantic information in the voice data, resulting in semantic loss of the voice data during transmission. Summary of the Invention
[0005] In order to solve the technical problem of semantic loss of voice data during transmission, the present application provides a voice data transmission method and system based on a cloud platform, which can avoid semantic loss of voice data during transmission.
[0006] In a first aspect, the present application provides a voice data transmission method based on a cloud platform, the transmission method comprising: converting the voice data of a speaking node into text data and then inputting the text into a semantic extraction model, erasing the audio information of any time frame in the voice data, and calculating the semantic importance of each time frame based on the change in the output result of the semantic extraction model; transmitting the voice data to each receiving node via a communication link between the speaking node and the receiving node, and calculating the semantic loss rate of the first received data of each receiving node based on the semantic importance; in response to the semantic loss rate of any receiving node being no greater than a loss rate threshold, taking the first received data as the final received data, and conversely, calculating the retransmission priority of other nodes of the receiving node, receiving the first received data of the other nodes in descending order of the retransmission priority, obtaining the second received data of the receiving node, obtaining fusion data of the first received data and the second received data, until the final received data is obtained, thereby completing the data transmission of the receiving node.
[0007] In an online multi-party interaction scenario, the voice data of a speaking node needs to be simultaneously transmitted to multiple receiving nodes. Before data transmission, the semantic importance of each time frame in the voice data of the speaking node is accurately measured using the change in the output result of a semantic extraction model. After the data is transmitted to each receiving node via a communication link between the speaking node and the receiving node, the semantic loss rate of each receiving node is calculated based on the semantic importance of each time frame. If the semantic loss rate of the receiving node is not greater than a loss rate threshold, it indicates that the first received data of the receiving node contains accurate semantic information, and the first received data is directly used as the final received data. If the semantic loss rate of the receiving node is greater than the loss rate threshold, it indicates that accurate semantic information cannot be obtained based on the first received data of the receiving node. The retransmission priorities of other nodes of the receiving node are calculated, and the first received data of the other nodes are received in descending order of retransmission priority to obtain the second received data of the receiving node, and then a fusion of the first received data and the second received data is obtained. The second received data is continuously obtained to update the fusion data until the fusion data contains accurate semantic information, and the final received data is obtained, thereby avoiding semantic loss of the voice data during transmission.
[0008] Preferably, erasing the audio information of any time frame in the voice data includes: after deleting the audio information of the any time frame, supplementing the voice data by using an interpolation method to complete the erasing operation of the time frame.
[0009] Preferably, calculating the semantic importance of each time frame includes: inputting the text data of the speech data into the semantic extraction model and recording the output result as the benchmark semantic ; will erase the time frame The text data of the post-speech data is input into the semantic extraction model to obtain the time frame The second semantics ; Timeframe The semantic importance of for: ; for and The Euclidean distance of It is the sum of the Euclidean distances between the second semantics and the baseline semantics in each time frame.
[0010] Accurately quantify the semantic importance of each time frame in the speech data. The greater the semantic importance of a time frame, the more likely the loss of the time frame will cause a significant deviation in the semantic information during the transmission of the semantic data.
[0011] Preferably, the receiving node Semantic loss rate of the first received data for:
[0012] , is the total number of time frames in the speech data; For time frame The semantic importance of is an indicator function, if the time frame in the first received data The audio data is lost, If the time frame of the first received data The audio data is not lost. .
[0013] During the transmission of voice data along the communication link, there is a certain packet loss rate, which causes the audio information of some time frames in the first received data to be lost. The semantic loss rate is quantified by the semantic importance of each time frame. The larger the semantic loss rate, the less accurate the first received data of the receiving node can be in reflecting the accurate semantic information of the speaking node.
[0014] Preferably, the receiving node Other nodes Retransmission priority for:
[0015] ; For receiving nodes Other nodes The semantic loss rate, For receiving nodes Receive the transmission duration of the first received data, For other nodes To the receiving node The predicted transmission time of the communication link, is the average transmission time for each receiving node with a semantic loss rate not greater than the loss rate threshold to receive the first received data, For other nodes To the receiving node The predicted packet loss rate of the communication link is obtained, wherein the predicted transmission duration is the average of multiple adjacent historical transmission durations, and the predicted packet loss rate is the average of multiple adjacent historical packet loss rates.
[0016] From the receiving node Other nodes The first received data contains semantic information, other nodes To the receiving node The retransmission priority is calculated comprehensively based on the predicted packet loss rate of the communication link and whether the time length for each receiving node to receive the final received data containing semantic information is consistent, so as to achieve accurate quantification of the retransmission priority of other nodes of the receiving node.
[0017] Preferably, obtaining the fusion data of the first received data and the second received data includes: locating multiple lost time frames where audio data is lost in the first received data; if any lost time frame has audio data in the second received data, the lost time frame in the first received data is supplemented based on the audio data of the lost time frame in the second received data; conversely, the lost time frame in the first received data is supplemented using interpolation to obtain fused data.
[0018] Preferably, after completing the data transmission of the receiving node, the transmission method further comprises: converting the final received data of the receiving node into a broadcast voice for broadcasting.
[0019] Preferably, completing the data transmission of the receiving node also includes: obtaining the number of second received data of any receiving node, and issuing a fault warning in response to the number of second received data being greater than a number threshold and the semantic loss rate of the fused data being greater than a loss rate threshold.
[0020] It can prevent the speech data of the speaking node from being transmitted to the receiving node for too long due to the communication link failure of the receiving node, thereby realizing the early warning of the communication link failure.
[0021] Preferably, the semantic extraction model is BERT or LSTM.
[0022] In a second aspect of the present application, a cloud platform-based voice data transmission system is also provided, comprising a processor and a memory, wherein the memory stores computer program instructions, and when the computer program instructions are executed by the processor, a cloud platform-based voice data transmission method according to the first aspect of the present application is implemented.
[0023] The technical solution of this application has the following beneficial technical effects:
[0024] In an online multi-party interaction scenario, the voice data of a speaking node needs to be simultaneously transmitted to multiple receiving nodes. Before data transmission, the semantic importance of each time frame in the voice data of the speaking node is accurately measured using the change in the output result of a semantic extraction model. After the data is transmitted to each receiving node via a communication link between the speaking node and the receiving node, the semantic loss rate of each receiving node is calculated based on the semantic importance of each time frame. If the semantic loss rate of the receiving node is not greater than a loss rate threshold, it indicates that the first received data of the receiving node contains accurate semantic information, and the first received data is directly used as the final received data. If the semantic loss rate of the receiving node is greater than the loss rate threshold, it indicates that accurate semantic information cannot be obtained based on the first received data of the receiving node. The retransmission priorities of other nodes of the receiving node are calculated, and the first received data of the other nodes are received in descending order of retransmission priority to obtain the second received data of the receiving node, and then a fusion of the first received data and the second received data is obtained. The second received data is continuously obtained to update the fusion data until the fusion data contains accurate semantic information, and the final received data is obtained, thereby avoiding semantic loss of the voice data during transmission. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 This is a flowchart of a cloud platform-based voice data transmission method according to an embodiment of the present application.
[0026] Figure 2 This is a structural block diagram of a cloud platform-based voice data transmission system according to an embodiment of the present application. DETAILED DESCRIPTION
[0027] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without making creative efforts are within the scope of protection of this application.
[0028] According to the first aspect of the present application, the present application provides a cloud platform-based voice data transmission method for realizing the transmission of voice data in an online multi-party interaction scenario. The online multi-party interaction scenario can be a scenario with multiple user nodes, such as an online classroom or an online meeting. In such a scenario, the voice data of any user node needs to be transmitted to all user nodes other than that user. The user node is the user's terminal device.
[0029] Figure 1 This is a flow chart of a voice data transmission method based on a cloud platform according to an embodiment of the present application. Figure 1As shown, the cloud platform-based voice data transmission method includes steps S101 to S103, which are described in detail below.
[0030] S101, converting the speech data of the speaking node into text data and inputting it into a semantic extraction model, erasing the audio information of any time frame in the speech data, and calculating the semantic importance of each time frame based on the change in the output result of the semantic extraction model.
[0031] In one embodiment, there are multiple user nodes in an online multi-party interaction scenario. When any user node speaks, the speaking user node is defined as a speaking node, and the user nodes other than the speaking node are defined as receiving nodes. At this time, the voice data of the speaking node needs to be transmitted to each receiving node.
[0032] It can be understood that the purpose of transmitting the voice data of the speaking node to the receiving node is to enable the user of the receiving node to correctly understand the semantic information of the speaking node user. Therefore, before transmission, the voice data of the speaking node is analyzed to determine the semantic importance of each time frame in the voice data of the speaking node.
[0033] Specifically, the speech data of the speaking node is converted into text data and then input into a semantic extraction model. The semantic extraction model is a recurrent neural network such as Bert or LSTM, which is used to extract semantic features from the text data. The output result of the semantic extraction model is the semantic information of the speech data.
[0034] The conversion of voice data into text data is a well-known technique to those skilled in the art and will not be elaborated here.
[0035] In one embodiment, erasing audio information of any time frame in the voice data simulates packet loss during voice data transmission. For example, erasing audio information of time frame 2 in the voice data indicates that the audio information of time frame 2 was lost during the voice data transmission process. The semantic importance of time frame 2 can be determined based on the change in semantic information of the voice data after the loss of the audio information of time frame 2. Specifically, erasing audio information of any time frame in the voice data includes: deleting the audio information of the any time frame and then supplementing the voice data using interpolation to complete the erasing operation of the time frame.
[0036] In one embodiment, calculating the semantic importance of each time frame includes: inputting the text data of the speech data into the semantic extraction model and recording the output result as the benchmark semantic ; will erase the time frame The text data of the post-speech data is input into the semantic extraction model to obtain the time frame The second semantics ; Timeframe The semantic importance of for: ; for and The Euclidean distance of It is the sum of the Euclidean distances between the second semantics and the baseline semantics in each time frame.
[0037] in, It can be regarded as erasing the time frame in the speech data The change in semantic information caused by the audio information of the time frame is larger. The loss of semantic information has a greater impact on the time frame The greater the semantic importance.
[0038] In this way, accurate quantification of the semantic importance of each time frame in the speech data is achieved. The greater the semantic importance of a time frame, the more likely the loss of the time frame will cause a significant deviation in the semantic information during the transmission of the semantic data.
[0039] S102 : transmitting the speech data to each receiving node via the communication link between the speaking node and the receiving node, and calculating the semantic loss rate of the first received data of each receiving node according to the semantic importance.
[0040] In one embodiment, a communication link exists between the speaking node and each receiving node, as well as between any two receiving nodes, enabling data transmission between the nodes. After voice data from the speaking node is collected, the voice data is preferentially transmitted to each receiving node via the communication link between the speaking node and the receiving nodes, thereby obtaining first received data from each receiving node.
[0041] It should be noted that during the transmission of voice data along the communication link, there is a certain packet loss rate, which causes the audio information of some time frames in the first received data to be lost. The loss of time frame audio information will cause deviations in semantic information. Therefore, the semantic loss rate is calculated based on the time frames in which audio information can be received in the first received data.
[0042] Specifically, the receiving node Semantic loss rate of the first received data for:
[0043] , is the total number of time frames in the speech data; For time frame The semantic importance of is an indicator function, if the time frame in the first received data The audio data is lost, If the time frame of the first received data The audio data is not lost. .
[0044] Thus, the greater the semantic loss rate is, the less likely the first received data of the receiving node can accurately reflect the accurate semantic information of the speaking node.
[0045] S103, in response to the semantic loss rate of any receiving node being no greater than the loss rate threshold, the first received data is taken as the final received data, otherwise, the retransmission priorities of other nodes of the receiving node are calculated, and the first received data of other nodes are received in descending order according to the retransmission priorities to obtain the second received data of the receiving node, and the fusion data of the first received data and the second received data is obtained until the final received data is obtained, thereby completing the data transmission of the receiving node.
[0046] In one embodiment, in response to the semantic loss rate of any receiving node being no greater than a loss rate threshold, indicating that the first received data of the receiving node can truly reflect the semantic information of the speech data of the speaking node, the first received data is used as the final received data.
[0047] In response to the semantic loss rate of any receiving node being greater than the loss rate threshold, it indicates that the first received data of the receiving node cannot accurately reflect the semantic information of the voice data of the speaking node. At this time, it is necessary to calculate the retransmission priority of other nodes of the receiving node, and the receiving node receives the voice data again from other nodes to ensure that the receiving node can obtain the final received data containing semantic information; wherein, the other nodes include the speaking node and receiving nodes other than the receiving node.
[0048] The loss rate threshold is set to 0.3.
[0049] On the one hand, one other node corresponds to a semantic loss rate. The smaller the semantic loss rate, the more the first received data of the other node can reflect the semantic information of the speaking node. After the first received data of the other node is transmitted to the receiving node, it can effectively supplement the semantic information in the receiving node. The greater the retransmission priority of the other node, that is, the retransmission priority is negatively correlated with the semantic loss rate. It can be understood that if the other node is a speaking node, the first received data is the voice data of the speaking node, and the semantic loss rate is 0.
[0050] On the other hand, there is a communication link between each other node and the receiving node, and the transmission speed and packet loss rate in each communication link are different. In the online multi-party interaction scenario, the purpose of voice data transmission is to allow all receiving nodes to simultaneously receive the final received data containing semantic information. Therefore, for any other node, the total transmission time of the receiving node receiving the first received data and the transmission time of the second received data from the other node are obtained. If the total time and the average transmission time of each receiving node receiving the first received data with a semantic loss rate not greater than the loss rate threshold are not much different, it means that the time when each receiving node receives the final received data containing semantic information is basically the same. Therefore, in order to ensure that each receiving node in the online multi-party interaction scenario can simultaneously receive the final received data containing semantic information, the retransmission priority of other nodes is calculated based on the difference between the total time and the average transmission time.
[0051] Specifically, the receiving node Other nodes Retransmission priority for:
[0052] ; For receiving nodes Other nodes The semantic loss rate, For receiving nodes Receive the transmission duration of the first received data, For other nodes To the receiving node The predicted transmission time of the communication link, is the average transmission time for each receiving node with a semantic loss rate not greater than the loss rate threshold to receive the first received data, For other nodes To the receiving node The predicted packet loss rate of the communication link is obtained, wherein the predicted transmission duration is the average of multiple adjacent historical transmission durations, and the predicted packet loss rate is the average of multiple adjacent historical packet loss rates.
[0053] Understandably, the semantic loss rate The smaller the receiving node Other nodes The more semantic information the first received data contains, the more likely the receiving node is to Other nodes Retransmission priority The bigger; For receiving nodes The length of time it takes to receive the final received data containing semantic information, The average time length of each receiving node to directly use the first received data as the final received data is: The smaller the value of is, the more consistent the time length for each receiving node to receive the final receiving data containing semantic information is. In other words, all receiving nodes can receive the final receiving data containing semantic information at the same time, so the receiving nodes Other nodes Retransmission priority The other nodes will also be larger. To the receiving node The predicted packet loss rate of the communication link The larger the The first received data is transmitted to the receiving node In the process, it will further lead to the loss of semantic information, so the receiving node Other nodes Retransmission priority In this way, the accurate quantification of the retransmission priority of other nodes of the receiving node is achieved.
[0054] Among them, the other nodes To the receiving node The predicted transmission duration of the communication link is the average of multiple adjacent historical transmission durations, and the predicted packet loss rate is the average of multiple adjacent historical packet loss rates.
[0055] In one embodiment, after obtaining the retransmission priority of other nodes of the receiving node, the other nodes corresponding to the maximum retransmission priority are first used as target nodes, and the first received data of the target node is transmitted to the receiving node via the communication link between the target node and the receiving node to obtain the second received data of the receiving node.
[0056] In one embodiment, obtaining fused data of first received data and second received data includes: locating multiple lost time frames where audio data is lost in the first received data; if any lost time frame has audio data in the second received data, the lost time frame in the first received data is supplemented based on the audio data of the lost time frame in the second received data; otherwise, the lost time frame in the first received data is supplemented using interpolation to obtain fused data.
[0057] If the semantic loss rate of the fused data is greater than the loss rate threshold, it means that the fused data still cannot obtain accurate semantic information. The first received data of other nodes are received in order of retransmission priority from large to small, and the second received data of the receiving node is continuously obtained, and then the fused data is continuously updated until the semantic loss rate of the fused data is no greater than the loss rate threshold, indicating that the fused data can obtain accurate semantic information, obtain the final received data, and complete the data transmission of the receiving node.
[0058] In one embodiment, after the data transmission of the receiving node is completed, the transmission method further includes: converting the final received data of the receiving node into a broadcast voice for broadcasting.
[0059] In one embodiment, in order to prevent the voice data of the speaking node from being transmitted to the receiving node for too long due to a communication link failure of the receiving node, the transmission method also includes: obtaining the number of second received data of any receiving node, and issuing a fault warning in response to the number of second received data being greater than a number threshold and the semantic loss rate of the fused data being greater than a loss rate threshold.
[0060] The value of the quantity threshold is 5. The greater the quantity of the second received data, the more times other nodes send the first received data to the receiving node, which will increase the time it takes to transmit the voice data of the speaking node to the receiving node.
[0061] In this way, it is ensured that in the online multi-party interaction scenario, all receiving nodes can simultaneously receive voice data containing semantic information, ensuring the smooth progress of online multi-party interaction and improving user experience.
[0062] According to the second aspect of the present application, the present application also provides a voice data transmission system based on a cloud platform. Figure 2 This is a structural block diagram of a voice data transmission system based on a cloud platform according to an embodiment of the present application. Figure 2 As shown, the system 50 includes a processor and a memory. The memory stores computer program instructions. When the computer program instructions are executed by the processor, the cloud-based voice data transmission method according to the first aspect of the present application is implemented. The system also includes other components well known to those skilled in the art, such as a communication bus and a communication interface. The configuration and functions of these components are well known in the art and are therefore not described in detail here.
[0063] It should be pointed out that, for ordinary technicians in this field, several variations and improvements can be made without departing from the concept of the present application, and these all fall within the scope of protection of the present application.
Claims
1. A voice data transmission method based on a cloud platform, characterized in that: The transmission method includes: converting the speech data of the speaking node into text data and then inputting it into the semantic extraction model, erasing the audio information of any time frame in the speech data, including: after deleting the audio information of any time frame, and using the interpolation method to complete the speech data, completing the erasing operation of the time frame, calculating the semantic importance of each time frame according to the change in the output result of the semantic extraction model, including: recording the output result of the text data of the speech data after inputting the semantic extraction model as the baseline semantic ; will erase the time frame The text data of the post-speech data is input into the semantic extraction model to obtain the time frame The second semantics ; Timeframe The semantic importance of for: ; for and The Euclidean distance of is the sum of the Euclidean distances between the second semantics and the baseline semantics in each time frame; The speech data is transmitted to each receiving node via a communication link between a speaking node and a receiving node, and a semantic loss rate of the first received data of each receiving node is calculated based on the semantic importance; In response to a semantic loss rate of any receiving node being no greater than a loss rate threshold, taking the first received data as the final received data; conversely, calculating the retransmission priorities of other nodes of the receiving node, sequentially receiving the first received data of the other nodes in descending order of the retransmission priorities, obtaining the second received data of the receiving node, obtaining fused data of the first received data and the second received data, until the final received data is obtained, thereby completing the data transmission of the receiving node; Receiving Node Other nodes Retransmission priority for: ; For receiving nodes Other nodes The semantic loss rate, For receiving nodes Receive the transmission duration of the first received data, For other nodes To the receiving node The predicted transmission time of the communication link, is the average transmission time for each receiving node with a semantic loss rate not greater than the loss rate threshold to receive the first received data, For other nodes To the receiving node The predicted packet loss rate of the communication link is obtained, wherein the predicted transmission duration is the average of multiple adjacent historical transmission durations, and the predicted packet loss rate is the average of multiple adjacent historical packet loss rates.
2. The method for transmitting voice data based on a cloud platform according to claim 1, characterized in that: Receiving Node Semantic loss rate of the first received data for: , is the total number of time frames in the speech data; For time frame The semantic importance of is an indicator function, if the time frame in the first received data The audio data is lost, If the time frame of the first received data The audio data is not lost. .
3. The method for transmitting voice data based on a cloud platform according to claim 1, wherein: Acquiring fusion data of the first received data and the second received data includes: locating a plurality of lost time frames in which audio data is lost in the first received data; If there is audio data for any lost time frame in the second received data, the lost time frame in the first received data is supplemented based on the audio data of the lost time frame in the second received data; otherwise, the lost time frame in the first received data is supplemented using interpolation to obtain fused data.
4. The method for transmitting voice data based on a cloud platform according to claim 1, wherein: After completing the data transmission of the receiving node, the transmission method further includes: converting the final received data of the receiving node into a broadcast voice for broadcasting.
5. The method for transmitting voice data based on a cloud platform according to claim 1, characterized in that: Completing the data transmission of the receiving node also includes: obtaining the number of second received data of any receiving node, and issuing a fault warning in response to the number of second received data being greater than a number threshold and the semantic loss rate of the fused data being greater than a loss rate threshold.
6. The method for transmitting voice data based on a cloud platform according to claim 1, characterized in that: The semantic extraction model is Bert or LSTM.
7. A voice data transmission system based on a cloud platform, characterized in that: The method comprises a processor and a memory, wherein the memory stores computer program instructions, and when the computer program instructions are executed by the processor, a cloud platform-based voice data transmission method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Data processing method and device, medium and electronic equipment
CN111464262A
Voice data transmission method and device, equipment, medium and program product
CN116996622A