Voice data transmission method and system based on cloud platform

By calculating semantic importance and loss rate before voice data transmission, using semantic extraction model and interpolation method to complete the lost information, the problem of semantic loss in voice data transmission is solved, and accurate voice data transmission in multi-party interactive scenarios is realized.

CN120238247AActive Publication Date: 2025-07-01GUANGZHOU JIUSI INTELLIGENT TECH CO LTD

Patent Information

Application Number
CN202510697069.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-07-01
Estimated Expiration
2045-05-28

AI Technical Summary

Technical Problem

The prior art cannot effectively avoid semantic loss during voice data transmission, and the speech speed detection results cannot reflect the timing distribution of semantic information in voice data, resulting in semantic loss during voice data transmission.

Method used

After converting the speech data into text data, the semantic importance of each time frame is calculated using the semantic extraction model, the data transmission sequence is determined based on the semantic loss rate and retransmission priority, and the lost audio information is completed by interpolation method to obtain the final received data to avoid semantic loss.

Benefits of technology

In multi-party interaction scenarios, it is ensured that the receiving node can accurately receive semantic information of voice data, avoid semantic loss of voice data during transmission, and improve user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120238247A_ABST
    Figure CN120238247A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of voice transmission, in particular to a voice data transmission method and system based on a cloud platform, and the method comprises the steps: calculating the semantic importance of each time frame in voice data of a speaking node; after transmitting the voice data to each receiving node, calculating a semantic loss rate of first receiving data of each receiving node; and in response to the condition that the semantic loss rate of any receiving node is not greater than a loss rate threshold value, taking the first receiving data as final receiving data, otherwise, calculating retransmission priorities of other nodes, sequentially receiving the first receiving data of the other nodes according to a descending order of the retransmission priorities, and obtaining fusion data until the final receiving data is obtained. And completing data transmission. According to the technical scheme, semantic loss of the voice data in the transmission process can be avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of voice transmission technology, and particularly to a voice data transmission method and system based on a cloud platform. Background Art

[0002] With the continuous development of artificial intelligence technology, online voice interaction technology has become increasingly mature. Online meetings, online education, or online sales all require online voice interaction, and the transmission process of voice data is involved in the interaction process.

[0003] Currently, the patent application document with the publication number CN116996622A discloses a method, device, equipment, medium, and program product for transmitting voice data. The method includes: real-time collecting first voice data, where the first voice data is voice audio data to be transmitted to a second terminal in real time; performing a speech rate detection on the first voice data to obtain a speech rate detection result corresponding to the first voice data, where the speech rate detection result is used to characterize the density distribution of the speech expression in the first voice data in terms of time sequence; determining a target transmission link from multiple candidate transmission link configurations based on the speech rate detection result; and transmitting the first voice data from a first terminal to the second terminal through the target transmission link; where the speech rate detection result is used to characterize the density distribution of the speech expression in the first voice data in terms of time sequence.

[0004] The above method reduces the packet loss probability of voice audio data during the transmission process by performing a speech rate detection on the first voice data and determining the target transmission link according to the speech rate detection result. However, the purpose of voice data transmission is to enable the receiving party to understand the semantic information in the voice data, and the speech rate detection result cannot reflect the distribution of the semantic information in the voice data in terms of time sequence, resulting in semantic loss of the voice data during the transmission process. Summary of the Invention

[0005] In order to solve the technical problem of semantic loss of voice data during the transmission process, this application provides a voice data transmission method and system based on a cloud platform, which can avoid semantic loss of voice data during the transmission process.

[0006] In the first aspect of the present application, a voice data transmission method based on a cloud platform is provided. The transmission method includes: converting the voice data of a speaking node into text data, inputting it into a semantic extraction model, erasing the audio information of any time frame in the voice data, and calculating the semantic importance of each time frame according to the change amount of the output result of the semantic extraction model; transmitting the voice data to each receiving node through the communication link between the speaking node and the receiving node, and calculating the semantic loss rate of the first received data of each receiving node according to the semantic importance; in response to the semantic loss rate of any receiving node not being greater than the loss rate threshold, taking the first received data as the final received data. Otherwise, calculating the retransmission priorities of other nodes of the receiving node, receiving the first received data of other nodes in descending order of the retransmission priorities, obtaining the second received data of the receiving node, and obtaining the fusion data of the first received data and the second received data until the final received data is obtained, completing the data transmission of the receiving node.

[0007] In the scenario of online multi-party interaction, it is necessary to transmit the voice data of the speaking node to multiple receiving nodes simultaneously; before data transmission, the change amount of the output result of the semantic extraction model is used to accurately measure the semantic importance of each time frame in the voice data of the speaking node. After transmitting it to each receiving node through the communication link between the speaking node and the receiving node, the semantic loss rate of each receiving node is calculated according to the semantic importance of each time frame. If the semantic loss rate of the receiving node is not greater than the loss rate threshold, it means that the first received data of the receiving node contains accurate semantic information, and the first received data is directly used as the final received data; if the semantic loss rate of the receiving node is greater than the loss rate threshold, it means that accurate semantic information cannot be obtained based on the first received data of the receiving node. Then, calculate the retransmission priorities of other nodes of the receiving node, receive the first received data of other nodes in descending order of the retransmission priorities, obtain the second received data of the receiving node, and obtain the fusion data of the first received data and the second received data; continuously obtain the second received data to update the fusion data until the fusion data contains accurate semantic information, and then obtain the final received data, avoiding semantic loss during the transmission of voice data.

[0008] Preferably, erasing the audio information of any time frame in the voice data includes: deleting the audio information of any time frame, and using the interpolation method to complete the erasure operation of the time frame by complementing the voice data.

[0009] Preferably, calculating the semantic importance of each time frame includes: recording the output result after inputting the text data of the voice data into the semantic extraction model as the reference semantics ; inputting the text data of the voice data after erasing the time frame into the semantic extraction model to obtain the second semantics of the time frame ; the time frame ; the time frame Semantic importance is as follows: ; is and the Euclidean distance, is the sum of the Euclidean distances between the second semantics and the reference semantics of each time frame.

[0010] Precisely quantify the semantic importance of each time frame in the voice data. If the semantic importance of a time frame is greater, it means that during the transmission of the semantic data, the loss of this time frame will cause a large deviation in the semantic information.

[0011] Preferably, the receiving node The semantic loss rate of the first received data is as follows: , is the total number of time frames in the voice data; is the time frame semantic importance; is an indicator function. If the audio data of the time frame in the first received data is lost, , if the audio data of the time frame in the first received data is not lost, .

[0012] During the transmission of the voice data along the communication link, there is a certain packet loss rate, resulting in the loss of audio information of some time frames in the first received data. Quantify the semantic loss rate through the semantic importance of each time frame. The greater the semantic loss rate, the less accurately the first received data of the receiving node can reflect the accurate semantic information of the speaking node.

[0013] Preferably, the receiving node of other nodes retransmission priority is as follows: ; is the receiving node other nodes semantic loss rate, is the receiving node transmission duration of receiving the first received data, is other nodes to the receiving node predicted transmission duration of the communication link, is the average transmission duration of receiving the first received data by each receiving node with a semantic loss rate not greater than the loss rate threshold, is other nodes to the receiving node The predicted packet loss rate of the communication link, where the predicted transmission duration is the average of multiple adjacent historical transmission durations, and the predicted packet loss rate is the average of multiple adjacent historical packet loss rates.

[0014] From the receiving node Other nodes The semantic information included in the first received data of other nodes To the receiving node The predicted packet loss rate of the communication link, and whether the time lengths of the final received data containing semantic information received by each receiving node are consistent. The retransmission priority is comprehensively calculated from these three aspects to achieve accurate quantification of the retransmission priority of other nodes of the receiving node.

[0015] Preferably, obtaining the fusion data of the first received data and the second received data includes: locating multiple lost time frames where audio data is lost in the first received data; if any lost time frame has audio data in the second received data, complement the lost time frame in the first received data according to the audio data of the lost time frame in the second received data. Otherwise, use the interpolation method to complement the lost time frame in the first received data to obtain the fusion data.

[0016] Preferably, after completing the data transmission of the receiving node, the transmission method further includes: converting the final received data of the receiving node into a broadcast voice for broadcasting.

[0017] Preferably, completing the data transmission of the receiving node further includes: obtaining the quantity of the second received data of any receiving node, and issuing a fault warning when the quantity of the second received data is greater than the quantity threshold and the semantic loss rate of the fusion data is greater than the loss rate threshold.

[0018] It can prevent the voice data of the speaking node from being transmitted to the receiving node for too long due to the communication link failure of the receiving node, and realize the fault warning of the communication link.

[0019] Preferably, the semantic extraction model is Bert or LSTM.

[0020] In the second aspect of the present application, a voice data transmission system based on a cloud platform is further provided, including a processor and a memory. The memory stores computer program instructions, and when the computer program instructions are executed by the processor, a voice data transmission method based on the cloud platform according to the first aspect of the present application is implemented.

[0021] The technical solution of the present application has the following beneficial technical effects: In the scenario of online multi-party interaction, it is necessary to transmit the voice data of the speaking node to multiple receiving nodes simultaneously. Before data transmission, the change amount of the output result of the semantic extraction model is used to accurately measure the semantic importance of each time frame in the voice data of the speaking node. After being transmitted to each receiving node through the communication link between the speaking node and the receiving nodes, the semantic loss rate of each receiving node is calculated according to the semantic importance of each time frame. If the semantic loss rate of the receiving node is not greater than the loss rate threshold, it means that the first received data of the receiving node contains accurate semantic information, and the first received data is directly used as the final received data. If the semantic loss rate of the receiving node is greater than the loss rate threshold, it means that accurate semantic information cannot be obtained based on the first received data of the receiving node. Then, the retransmission priority of other nodes of the receiving node is calculated, and the first received data of other nodes is received in descending order of the retransmission priority to obtain the second received data of the receiving node, and the fusion data of the first received data and the second received data is obtained. The second received data is continuously obtained to update the fusion data until the fusion data contains accurate semantic information, and the final received data is obtained, avoiding semantic loss during the transmission of voice data. Description of the Drawings

[0022] Figure 1 is a flowchart of a method for transmitting voice data based on a cloud platform according to an embodiment of the present application.

[0023] Figure 2 is a structural block diagram of a system for transmitting voice data based on a cloud platform according to an embodiment of the present application. Detailed Embodiments

[0024] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application. All other embodiments obtained by those skilled in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0025] According to the first aspect of the present application, the present application provides a method for transmitting voice data based on a cloud platform, which is used to implement the transmission of voice data in the scenario of online multi-party interaction. The scenario of online multi-party interaction can be scenarios with multiple user nodes such as online classrooms and online meetings. In such scenarios, the voice data of any one user node needs to be transmitted to all user nodes other than the user. Among them, the user node is the terminal device of the user.

[0026] Figure 1 is a flowchart of a method for transmitting voice data based on a cloud platform according to an embodiment of the present application. As Figure 1As shown, the voice data transmission method based on a cloud platform includes steps S101 to S103, which are described in detail below.

[0027] S101, convert the voice data of the speaking node into text data, then input it into the semantic extraction model. After erasing the audio information of any time frame in the voice data, calculate the semantic importance of each time frame according to the change amount of the output result of the semantic extraction model.

[0028] In one embodiment, in a multi-party online interaction scenario, there are multiple user nodes. When any one user node speaks, the speaking user node is defined as the speaking node, and the user nodes other than the speaking node are defined as receiving nodes. At this time, it is necessary to transmit the voice data of the speaking node to each receiving node.

[0029] It can be understood that the purpose of transmitting the voice data of the speaking node to the receiving node is to enable the users of the receiving nodes to correctly understand the semantic information of the speaking node user. Therefore, before transmission, the voice data of the speaking node is analyzed to determine the semantic importance of each time frame in the voice data of the speaking node.

[0030] Specifically, after converting the voice data of the speaking node into text data, it is input into the semantic extraction model. The semantic extraction model is a recurrent neural network such as Bert or LSTM, which is used to extract semantic features from the text data. The output result of the semantic extraction model is the semantic information of the voice data.

[0031] Among them, the conversion of voice data into text data is a well-known technology to those skilled in the art and will not be elaborated here.

[0032] In one embodiment, erasing the audio information of any time frame in the voice data is to simulate the process of packet loss during the voice data transmission. For example, erasing the audio information data of time frame 2 in the voice data means that the audio information of time frame 2 is lost during the voice data transmission. According to the change amount of the semantic information of the voice data after the audio information of time frame 2 is lost, the semantic importance of time frame 2 can be judged. Specifically, erasing the audio information of any time frame in the voice data includes: after deleting the audio information of any time frame, using the interpolation method to complete the erasure operation of the time frame.

[0033] In one embodiment, calculating the semantic importance of each time frame includes: recording the output result after inputting the text data of the voice data into the semantic extraction model as the reference semantics ; inputting the text data of the voice data after erasing the time frame into the semantic extraction model to obtain the second semantics of the time frame ; time frame ; time frame Semantic importance is: ; is and the Euclidean distance of, is the sum of the Euclidean distances between the second semantics and the reference semantics of each time frame.

[0034] Among them, can be regarded as the change amount of semantic information caused by erasing the audio information of the time frame in the voice data. The larger this change amount is, the greater the impact of the loss of the time frame on the semantic information, and the greater the semantic importance of the time frame .

[0035] In this way, the accurate quantification of the semantic importance of each time frame in the voice data is realized. If the semantic importance of a time frame is greater, it means that during the transmission of semantic data, the loss of this time frame will cause a large deviation in semantic information.

[0036] S102. Transmit the voice data to each receiving node through the communication link between the speaking node and the receiving node, and calculate the semantic loss rate of the first received data of each receiving node according to the semantic importance.

[0037] In one embodiment, there is a communication link between the speaking node and each receiving node, and between any two receiving nodes, and data can be transmitted between all nodes. After collecting the voice data of the speaking node, the voice data is preferentially transmitted to each receiving node through the communication link between the speaking node and the receiving node to obtain the first received data of each receiving node.

[0038] It should be noted that during the transmission of voice data along the communication link, there is a certain packet loss rate, which causes the audio information of some time frames in the first received data to be lost, and the loss of the audio information of the time frame will cause a deviation in semantic information. Therefore, the semantic loss rate is calculated according to the time frames in the first received data that can receive audio information.

[0039] Specifically, the receiving node The semantic loss rate of the first received data is: , is the total number of time frames in the voice data; is the time frame Semantic importance of; is an indicator function. If the audio data of the time frame in the first received data is lost, , if the time frame The audio data is not lost, .

[0040] Thus, the greater the semantic loss rate, the less accurately the first received data at the receiving node can reflect the accurate semantic information of the speaking node.

[0041] S103, in response to the semantic loss rate of any receiving node being not greater than the loss rate threshold, use the first received data as the final received data; otherwise, calculate the retransmission priorities of other nodes of the receiving node, and receive the first received data of other nodes in descending order of the retransmission priorities to obtain the second received data of the receiving node, and obtain the fusion data of the first received data and the second received data until the final received data is obtained, completing the data transmission of the receiving node.

[0042] In one embodiment, in response to the semantic loss rate of any receiving node being not greater than the loss rate threshold, it means that the first received data of this receiving node can truly reflect the semantic information of the voice data of the speaking node, and the first received data is used as the final received data.

[0043] In response to the semantic loss rate of any receiving node being greater than the loss rate threshold, it means that the first received data of this receiving node cannot accurately reflect the semantic information of the voice data of the speaking node. At this time, it is necessary to calculate the retransmission priorities of other nodes of this receiving node, and this receiving node receives the voice data from other nodes again to ensure that the receiving node can obtain the final received data containing semantic information; wherein, the other nodes include the speaking node and the receiving nodes other than this receiving node.

[0044] Among them, the value of the loss rate threshold is 0.3.

[0045] On the one hand, one other node corresponds to one semantic loss rate. The smaller the semantic loss rate, the more accurately the first received data of this other node can reflect the semantic information of the speaking node. After the first received data of this other node is transmitted to the receiving node, it can effectively supplement the semantic information in the receiving node, and the retransmission priority of this other node is greater. That is to say, the retransmission priority is negatively correlated with the semantic loss rate; it can be understood that when the other node is the speaking node, the first received data is the voice data of the speaking node, and the semantic loss rate is 0.

[0046] On the other hand, there are communication links between each of the other nodes and the receiving node, and there are differences in the transmission speed and packet loss rate in each communication link. In the online multi-party interaction scenario, the purpose of voice data transmission is to enable all receiving nodes to receive the final received data containing semantic information simultaneously. Therefore, for any other node, the total duration of obtaining the transmission duration of the first received data received by the receiving node and the transmission duration of the second received data transmitted from this other node, if the difference between this total duration and the average transmission duration of the first received data received by the receiving nodes with a semantic loss rate not greater than the loss rate threshold is not significant, it indicates that the time for each receiving node to receive the final received data containing semantic information is basically the same. Therefore, to ensure that each receiving node in the online multi-party interaction scenario can receive the final received data containing semantic information simultaneously, the retransmission priority with other nodes is calculated based on the difference between the total duration and the average transmission duration.

[0047] Specifically, the retransmission priority of other nodes of the receiving node is: ; is the semantic loss rate of the receiving node for other nodes , is the transmission duration for the receiving node to receive the first received data, is the predicted transmission duration of the communication link from other nodes to the receiving node , is the average transmission duration of the first received data received by each receiving node with a semantic loss rate not greater than the loss rate threshold, is the predicted packet loss rate of the communication link from other nodes to the receiving node , and the predicted transmission duration is the average value of multiple adjacent historical transmission durations, and the predicted packet loss rate is the average value of multiple adjacent historical packet loss rates.

[0048] It can be understood that the smaller the semantic loss rate , the more semantic information is contained in the first received data of other nodes of the receiving node , and then the retransmission priority of other nodes of the receiving node is greater; is the time length for the receiving node to receive the final received data containing semantic information, is the average time length of each receiving node that directly uses the first received data as the final received data, The smaller the value, the more consistent the time lengths for each receiving node to receive the final received data containing semantic information. That is to say, all receiving nodes can receive the final received data containing semantic information simultaneously. Therefore, for other nodes of the receiving node of other nodes retransmission priority will also be larger; for other nodes to the receiving node the predicted packet loss rate of the communication link is larger. When transmitting the first received data of other nodes to the receiving node it will further cause the loss of semantic information. Therefore, for other nodes of the receiving node of other nodes retransmission priority will also be smaller. In this way, the accurate quantification of the retransmission priority of other nodes of the receiving node is realized.

[0049] Among them, the predicted transmission duration of the communication link from the other node to the receiving node is the average value of multiple adjacent historical transmission durations, and the predicted packet loss rate is the average value of multiple adjacent historical packet loss rates.

[0050] In one embodiment, after obtaining the retransmission priority of other nodes of the receiving node, first use the other node corresponding to the maximum retransmission priority as the target node, and transmit the first received data of the target node to the receiving node through the communication link between the target node and the receiving node to obtain the second received data of the receiving node.

[0051] In one embodiment, obtaining the fusion data of the first received data and the second received data includes: locating multiple lost time frames where audio data is lost in the first received data; if any lost time frame has audio data in the second received data, complete the lost time frame in the first received data according to the audio data of the lost time frame in the second received data. Otherwise, use the interpolation method to complete the lost time frame in the first received data to obtain the fusion data.

[0052] If the semantic loss rate of the fusion data is greater than the loss rate threshold, it means that the fusion data still cannot obtain accurate semantic information. Receive the first received data of other nodes in descending order of retransmission priority, continuously obtain the second received data of the receiving node, and then continuously update the fusion data until the semantic loss rate of the fusion data is not greater than the loss rate threshold, indicating that the fusion data can obtain accurate semantic information, and obtain the final received data to complete the data transmission of the receiving node.

[0053] In one embodiment, after the data transmission of the receiving node is completed, the transmission method further includes: converting the final received data of the receiving node into a broadcast voice for broadcast.

[0054] In one embodiment, to prevent the voice data of the speaking node from being transmitted to the receiving node for too long due to a communication link failure of the receiving node, the transmission method further includes: obtaining the quantity of the second received data of any receiving node, and issuing a fault warning when the quantity of the second received data is greater than a quantity threshold and the semantic loss rate of the fused data is greater than a loss rate threshold.

[0055] Wherein, the value of the quantity threshold is 5. The more the quantity of the second received data, the more times other nodes send the first received data to the receiving node, which will in turn increase the time for the voice data of the speaking node to be transmitted to the receiving node.

[0056] In this way, it is ensured that in the online multi-party interaction scenario, all receiving nodes can receive the voice data containing semantic information simultaneously, ensuring the smooth progress of online multi-party interaction and improving the user experience.

[0057] According to the second aspect of the present application, the present application further provides a voice data transmission system based on a cloud platform. Figure 2 It is a structural block diagram of a voice data transmission system based on a cloud platform according to an embodiment of the present application. As Figure 2 shown, the system 50 includes a processor and a memory. The memory stores computer program instructions, and when the computer program instructions are executed by the processor, a voice data transmission method based on the first aspect of the present application is implemented. The system further includes a communication bus, a communication interface and other components well known to those skilled in the art. Their settings and functions are known in the art, so they will not be described in detail here.

[0058] It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several deformations and improvements can be made, and these all belong to the protection scope of the present application.

Claims

1. A voice data transmission method based on a cloud platform, characterized in that, The described transmission method includes: converting the speech data of the speaking node into text data and then inputting it into the semantic extraction model. After erasing the audio information of any time frame in the speech data, calculate the semantic importance of each time frame according to the change amount of the output result of the semantic extraction model; Transmit the speech data to each receiving node through the communication link between the speaking node and the receiving node, and calculate the semantic loss rate of the first received data of each receiving node according to the semantic importance; In response to the semantic loss rate of any receiving node not being greater than the loss rate threshold, use the first received data as the final received data. Otherwise, calculate the retransmission priorities of other nodes of the receiving node, and receive the first received data of other nodes in descending order of the retransmission priorities to obtain the second received data of the receiving node. Obtain the fusion data of the first received data and the second received data until the final received data is obtained, and complete the data transmission of the receiving node.

2. The method for transmitting voice data based on a cloud platform according to claim 1, wherein Erasing the audio information of any time frame in the speech data includes: After deleting the audio information of any time frame, use the interpolation method to complete the speech data to complete the erasing operation of the time frame.

3. A voice data transmission method based on a cloud platform according to claim 1, characterized in that, Calculating the semantic importance of each time frame includes: The output result after inputting the text data of the speech data into the semantic extraction model is denoted as the reference semantics ; Input the text data of the speech data after erasing the time frame into the semantic extraction model to obtain the second semantics of the time frame ; ; Time frame Semantic importance is as follows: ; is and the Euclidean distance, which is the sum of the Euclidean distances between the second semantics and the reference semantics of each time frame.

4. A voice data transmission method based on a cloud platform according to claim 1, characterized in that, Receiving node Semantic loss rate of the first received data is as follows: , is the total number of time frames in the voice data; is the time frame 's semantic importance; is an indicator function. If the audio data of the time frame in the first received data is lost, , if the audio data of the time frame in the first received data is not lost, .

5. A method for voice data transmission based on a cloud platform according to claim 1, characterized in that Receiving node Other nodes Re - transmission priority Is: ; is the receiving node other nodes 's semantic loss rate, is the receiving node transmission duration for receiving the first received data, is other nodes to the receiving node predicted transmission duration of the communication link, is the average transmission duration for each receiving node with a semantic loss rate not greater than the loss rate threshold to receive the first received data, is other nodes to the receiving node predicted packet loss rate of the communication link, where the predicted transmission duration is the mean of multiple adjacent historical transmission durations, and the predicted packet loss rate is the mean of multiple adjacent historical packet loss rates.

6. A method for voice data transmission based on a cloud platform according to claim 1, wherein Obtaining the fusion data of the first received data and the second received data includes: Locate multiple lost time frames where the audio data is lost in the first received data; If there is audio data for the lost time frame in the second received data, complete the lost time frame in the first received data according to the audio data of the lost time frame in the second received data. Otherwise, use the interpolation method to complete the lost time frame in the first received data to obtain the fusion data.

7. A method for voice data transmission based on a cloud platform according to claim 1, characterized in that After completing the data transmission of the receiving node, the transmission method further includes: converting the final received data of the receiving node into a broadcast speech for broadcasting.

8. A method for voice data transmission based on a cloud platform according to claim 1, characterized in that, Completing the data transmission of the receiving node further includes: obtaining the quantity of the second received data of any receiving node. In response to the quantity of the second received data being greater than the quantity threshold and the semantic loss rate of the fusion data being greater than the loss rate threshold, issue a fault warning.

9. A method for voice data transmission based on a cloud platform according to claim 1, wherein, The semantic extraction model is Bert or LSTM.

10. A voice data transmission system based on a cloud platform, characterized in that, It includes a processor and a memory. The memory stores computer program instructions. When the computer program instructions are executed by the processor, it implements a voice data transmission method based on a cloud platform according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Data transmission method and device, terminal and storage medium

    CN110890945A

  • Data processing method and device, medium and electronic equipment

    CN111464262A

  • Audio packet loss compensation processing method and device and electronic equipment

    CN113035205A

  • Voice data transmission method and device, equipment, medium and program product

    CN116996622A

  • Audio-to-text conversion method and device, electronic equipment and storage medium

    CN119380719A

Cited By

  • Voice transmission method and device, storage medium and program product

    CN122290580A