Packet loss concealment method and apparatus, computing device, and storage medium

By establishing a packet loss concealment model based on multiple objects and a correction model for the target object, the voice parameters of the lost frames are optimized, solving the problem of poor packet loss recovery in existing technologies and improving the quality of voice calls.

CN116074298BActive Publication Date: 2026-05-05TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
TENCENT TECHNOLOGY (SHENZHEN) CO LTD
Filing Date
2021-11-04
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Existing packet loss concealment solutions are ineffective in recovering from packet loss due to significant differences in each person's pronunciation model and habits, thus failing to effectively improve the quality of voice calls.

Method used

By establishing a packet loss hiding model and a correction model, the packet loss hiding model is used to predict based on historical speech data of multiple objects, and the correction model is combined with the historical speech data of the target object to correct the speech parameters of the lost frames, so as to better match the speech characteristics of the target object.

Benefits of technology

The effect and quality of packet loss concealment have been improved, making the corrected voice parameters more consistent with the speaker's characteristics and improving the quality of voice calls.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116074298B_ABST
    Figure CN116074298B_ABST
Patent Text Reader

Abstract

A packet loss concealment method is described, comprising: determining the current frame in a speech data stream of a target object, the speech data stream including multiple frames of speech data with target speech parameters; responding to packet loss in the current frame, performing a first data processing operation on the current frame, the first data processing operation including the following steps: identifying the current frame as a packet-loss frame; determining predicted speech parameters of the packet-loss frame based on one or more preceding frames using a packet loss concealment model, wherein the packet loss concealment model is established based on historical speech data of multiple objects; correcting the predicted speech parameters of the packet-loss frame using a correction model to obtain corrected speech parameters of the packet-loss frame, wherein the correction model is established based on historical speech data of the target object; and determining the speech data of the packet-loss frame based on the corrected speech parameters of the packet-loss frame. Embodiments of this invention can be applied to various scenarios such as packet loss concealment, data transmission, and voice calls.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of network communication technology, and in particular to methods and apparatus for packet loss concealment, computing devices, and storage media. Background Technology

[0002] Packet loss is an unavoidable problem during network transmission, and it is one of the main reasons affecting voice call quality. Packet loss concealment (PLC) is a method that reconstructs the signal at the location of packet loss based on the audio signal information before and after the packet loss location, thereby reducing the impact of packet loss on voice call quality during network transmission.

[0003] In relevant packet loss concealment schemes, when packet loss is detected, the speech parameters of the lost frame are predicted based on the speech parameters of one or more normal speech frames preceding the packet loss, thereby recovering the lost frame signal. However, because everyone's pronunciation models and habits differ significantly, and the pronunciation methods of different languages ​​vary even more markedly, such packet loss concealment schemes result in poor speech recovery effectiveness (i.e., packet loss concealment effect). Summary of the Invention

[0004] In view of this, the present disclosure provides a method and apparatus for packet loss concealment, which is intended to overcome some or all of the defects mentioned above, as well as other possible defects.

[0005] According to a first aspect of this disclosure, a method for packet loss concealment is disclosed, comprising: determining a current frame in a speech data stream of a target object, the speech data stream including multiple frames of speech data having target speech parameters; in response to packet loss in the current frame, performing a first data processing operation on the current frame, the first data processing operation including the following steps: determining the current frame as a packet-loss frame; using a packet loss concealment model to determine predicted speech parameters of the packet-loss frame based on one or more previous frames of the packet-loss frame, wherein the packet loss concealment model is established based on historical speech data of multiple objects; using a correction model to correct the predicted speech parameters of the packet-loss frame to obtain corrected speech parameters of the packet-loss frame, wherein the correction model is established based on historical speech data of the target object; and determining the speech data of the packet-loss frame based on the corrected speech parameters of the packet-loss frame.

[0006] In some embodiments, the correction model is obtained by training a deep learning model based on the historical speech data of the target object through the following training steps: determining the target frame and the target speech parameters of the target frame in the historical speech data of the target object; using a packet loss hiding model to determine the predicted speech parameters of the target frame based on one or more previous frames of the target frame; using a deep learning model to correct the predicted speech parameters of the target frame to obtain the corrected speech parameters of the target frame; adjusting the parameters of the deep learning model to minimize the error between the corrected speech parameters of the target frame and the target speech parameters of the target frame, thereby obtaining the correction model.

[0007] In some embodiments, the method further includes: in response to no packet loss in the current frame, performing a second data processing operation on the current frame, the second data processing operation including the following steps: using a packet loss concealment model to determine the predicted speech parameters of the current frame based on one or more previous frames; using a correction model to correct the predicted speech parameters of the current frame to obtain the corrected speech parameters of the current frame; updating the parameters of the correction model to minimize the error between the corrected speech parameters of the current frame and the target speech parameters of the current frame.

[0008] In some embodiments, the method further includes: determining the type of the current frame, wherein the type of the current frame represents the trend of change of target speech parameters in one or more previous frames; wherein, correcting the predicted speech parameters of the lost frame using a correction model to obtain corrected speech parameters of the lost frame includes: correcting the predicted speech parameters of the lost frame using a correction model corresponding to the type of the current frame to obtain corrected speech parameters of the lost frame.

[0009] In some embodiments, the method further includes: in response to no packet loss in the current frame, performing a third data processing operation on the current frame, the third data processing operation including the following steps: using a packet loss concealment model to determine the predicted speech parameters of the current frame based on one or more previous frames; using a correction model corresponding to the type of the current frame to correct the predicted speech parameters of the current frame to obtain the corrected speech parameters of the current frame; updating the parameters of the correction model corresponding to the type of the current frame to minimize the error between the corrected speech parameters of the current frame and the target speech parameters of the current frame.

[0010] In some embodiments, for a correction model corresponding to the type of the current frame, there are multiple historical prediction errors. Each historical prediction error is the error between the predicted speech parameters of a historical frame obtained by a packet loss hiding model and the target speech parameters of the historical frame, wherein the historical frame is of the same type as the current frame. Updating the parameters of the correction model corresponding to the type of the current frame to minimize the error between the corrected speech parameters of the current frame and the target speech parameters of the current frame includes: using the multiple historical prediction errors to update the parameters of the correction model corresponding to the type of the current frame to minimize the error between the corrected speech parameters of the current frame and the target speech parameters of the current frame.

[0011] In some embodiments, updating the parameters of the correction model corresponding to the type of the current frame using multiple historical prediction errors to minimize the error between the corrected speech parameters of the current frame and the target speech parameters of the current frame includes: updating the parameters of the correction model corresponding to the type of the current frame to minimize the difference between the error between the corrected speech parameters of the current frame and the predicted speech parameters of the current frame, and the center error of the multiple historical prediction errors, so as to minimize the error between the corrected speech parameters of the current frame and the target speech parameters of the current frame.

[0012] In some embodiments, the central error of multiple historical prediction errors is determined by the following error determination steps, the error determination steps including: establishing a point group based on the multiple historical prediction errors, where each point in the point group represents a historical prediction error; clustering the point group to obtain a target point set; determining the central value of the historical prediction error corresponding to the point in the target point set; and determining the central value as the central error of the multiple historical prediction errors.

[0013] In some embodiments, clustering a group of points to obtain a target point set includes: removing outliers from the group of points, where outliers are points in the group whose average distance from other points in the group is greater than a preset threshold; and determining a target point set such that the target point set includes points other than the outliers in the group of points.

[0014] In some embodiments, the target speech parameters include one or more of the fundamental frequency, line spectrum pairs, and gain of the speech data.

[0015] According to a second aspect of this disclosure, an apparatus for packet loss concealment is provided, comprising: a determining module configured to determine a current frame in a speech data stream of a target object, the speech data stream including multiple frames of speech data having target speech parameters; and a processing module configured to perform a first data processing operation on the current frame in response to packet loss, the first data processing operation including the following steps: determining the current frame as a packet-loss frame; determining predicted speech parameters of the packet-loss frame based on one or more preceding frames of the packet-loss frame using a packet loss concealment model, wherein the packet loss concealment model is established based on historical speech data of multiple objects; correcting the predicted speech parameters of the packet-loss frame using a correction model to obtain corrected speech parameters of the packet-loss frame, wherein the correction model is established based on historical speech data of the target object; and determining the speech data of the packet-loss frame based on the corrected speech parameters of the packet-loss frame.

[0016] In some embodiments, the processing module is further configured to: in response to the fact that no packet loss has occurred in the current frame, perform a second data processing operation on the current frame, the second data processing operation including the following steps: using a packet loss concealment model to determine the predicted speech parameters of the current frame based on one or more previous frames; using a correction model to correct the predicted speech parameters of the current frame to obtain the corrected speech parameters of the current frame; and updating the parameters of the correction model to minimize the error between the corrected speech parameters of the current frame and the speech parameters of the current frame.

[0017] According to a third aspect of this disclosure, a computing device is provided, including a processor; and a memory configured to store computer-executable instructions thereon, which, when executed by the processor, perform any of the methods described above.

[0018] According to a fourth aspect of this disclosure, a computer-readable storage medium is provided that stores computer-executable instructions that, when executed, perform any of the methods described above.

[0019] According to a fifth aspect of this disclosure, a computer program product is provided, including computer-executable instructions, wherein the computer-executable instructions, when executed by a processor, perform any of the methods described above.

[0020] In the packet loss concealment method and apparatus claimed in this disclosure, the current frame in the speech data stream of the target object is first determined. The speech data stream includes multiple frames of speech data with target speech parameters. Then, in response to packet loss in the current frame, a first data processing operation is performed on the current frame. In the first data processing operation, a packet loss concealment model is first used to determine the predicted speech parameters of the lost frame based on one or more previous frames. Since the packet loss concealment model is based on historical speech data of multiple objects, the predicted speech parameters of the lost frame have wider applicability but are also more moderate. A correction model is then used to correct the predicted speech parameters of the lost frame to obtain the corrected speech parameters of the lost frame. Because the correction model is based on the historical speech data of the target object, i.e., the correction model is targeted at the target object, the corrected speech parameters better match the target speech data of the target object. Thus, when determining the speech data of the lost packet frame based on the corrected speech parameters, the corrected speech parameters better match the speech characteristics of the target object. Therefore, the speech data of the lost packet frame determined based on the corrected speech parameters better matches the speech data before the packet loss occurred. In this way, in the packet loss hiding method of this application, by establishing and using a targeted correction model for the target object, the speech parameters of the packet loss hiding model, which are biased towards the general public and lack specific predictions, are corrected, making the corrected speech parameters more consistent with the speaker's characteristics, thereby improving the effect and quality of packet loss hiding.

[0021] These and other advantages of this disclosure will become clear from the embodiments described below, and will be illustrated with reference to the embodiments described below. Attached Figure Description

[0022] Embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings, in which:

[0023] Figure 1 The illustrations depict exemplary application scenarios in which the technical solutions according to embodiments of this disclosure can be implemented;

[0024] Figure 2 The illustration shows a schematic flowchart of a packet loss concealment method according to an embodiment of the present disclosure;

[0025] Figure 3 The illustration shows a schematic flowchart of a method for determining a modified model according to an embodiment of the present disclosure;

[0026] Figure 4 A schematic flowchart illustrating a method for updating and correcting a model according to an embodiment of the present disclosure is shown.

[0027] Figure 5The illustration shows a schematic flowchart of a method for updating a correction model corresponding to the type of the current frame according to an embodiment of the present disclosure;

[0028] Figure 6 The illustration shows a schematic diagram of a packet loss hiding method according to an embodiment of the present disclosure;

[0029] Figure 7 The illustration shows the spectrogram of a speech signal according to an embodiment of the present disclosure;

[0030] Figure 8 An exemplary structural block diagram of a packet loss hiding device according to an embodiment of the present disclosure is shown;

[0031] Figure 9 An example system is illustrated, which includes an example computing device representing one or more systems and / or devices that can implement the various technologies described herein. Detailed Implementation

[0032] The following description provides specific details of various embodiments of this disclosure to enable those skilled in the art to fully understand and implement the various embodiments of this disclosure. It should be understood that the technical solutions of this disclosure can be implemented without some of these details. In some cases, this disclosure does not show or describe in detail some well-known structures or functions to avoid such unnecessary descriptions that would obscure the description of the embodiments of this disclosure. The terminology used in this disclosure should be understood in its broadest and most reasonable manner, even when used in conjunction with specific embodiments of this disclosure.

[0033] First, some of the terms used in the embodiments of this application will be explained to facilitate understanding by those skilled in the art.

[0034] Artificial Intelligence (AI): AI is the theory, methods, technology, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to possess the functions of perception, reasoning, and decision-making.

[0035] Network communication: Network protocols act as bridges for communication and exchange between networks. Only computers using the same network protocol can communicate and exchange information. This is similar to the different languages ​​used by people to communicate; only by using the same language can communication be normal and smooth. From a professional perspective, a network protocol is the agreement that computers must follow when communicating on a network; it is a communication protocol. It mainly specifies and standardizes information transmission rates, transmission codes, code structure, transmission control procedures, error control, and other aspects.

[0036] Data transmission: Data transmission is the communication process of transferring data from one place to another. A data transmission system typically consists of a transmission channel and data circuit terminating equipment at both ends of the channel; in some cases, it may also include multiplexing equipment at both ends of the channel. The transmission channel can be a dedicated communication channel or provided by a data switching network, telephone switching network, or other types of switching networks. The input and output devices of a data transmission system are terminals or computers, collectively referred to as data terminal equipment. The data information it sends is generally a combination of letters, numbers, and symbols. To transmit this information, each letter, number, or symbol must be represented using binary code.

[0037] Data packet: A packet is a unit of data in TCP / IP protocol communication, and is also generally called a "data packet".

[0038] Voice calls: Voice calls are a form of communication that uses voice and a transmission medium. Common examples include landline calls, mobile phone calls, walkie-talkie calls, and online voice chat. They can be categorized into two types: those that consume data and those that consume phone bills.

[0039] Packet loss: Packet loss refers to the failure of one or more data packets to reach their destination over the network. Packet loss, along with bit errors and spurious packets caused by noise, are the three main causes of digital communication errors.

[0040] Neural Networks: Artificial Neural Networks (ANNs), also known as neural networks or connection models, are mathematical models that mimic the behavioral characteristics of animal neural networks to perform distributed parallel information processing. These networks rely on the complexity of the system to adjust the connections between a large number of internal nodes, thereby achieving the purpose of information processing.

[0041] Deep Learning (DL) is a new research direction in the field of Machine Learning (ML). It was introduced into machine learning to bring it closer to its original goal—Artificial Intelligence (AI). Deep learning learns the inherent patterns and hierarchical representations of sample data. The information gained during this learning process greatly helps in interpreting data such as text, images, and sound. Its ultimate goal is to enable machines to have analytical and learning capabilities like humans, and to recognize data such as text, images, and sound.

[0042] Machine Learning (ML) is a multidisciplinary field involving probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence; its applications span all areas of artificial intelligence. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and instructional learning. With the research and advancement of artificial intelligence technology, it is being researched and applied in multiple fields, such as smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, autonomous driving, drones, robots, smart healthcare, smart customer service, connected vehicles, and intelligent transportation. It is believed that with technological development, artificial intelligence will be applied in even more fields and play an increasingly important role.

[0043] In related technologies, it is possible to recover packet loss during voice transmission by hiding the packet loss. For example, when packet loss is detected, the speech parameters of the lost frame are predicted based on the speech parameters of one or more normal speech frames preceding the packet loss, thereby recovering the lost frame signal. However, since everyone's pronunciation models and habits differ significantly, and the pronunciation methods of different languages ​​vary even more markedly, such packet loss hiding schemes result in poor speech recovery effects (i.e., packet loss hiding effectiveness).

[0044] Figure 1 The illustration shows an exemplary application scenario 100 in which the technical solutions according to embodiments of this disclosure can be implemented. For example... Figure 1As shown, the application scenario 100 includes a server 110, terminals 120 and 130, and a network 140. Terminals 120 and 130 are communicatively coupled to the server 110 via the network 140. As an example, a voice data stream of a target object, which can be an individual user, a robot, or other speech object, can be transmitted between applications or clients of the server 110, terminals 120, and 130 via the network 140. The voice data stream includes multiple frames of voice data with target voice parameters. Only two terminals are shown here; in reality, three or more terminals can exist.

[0045] As an example, during the reception of a voice data stream, terminal 120 may respond to packet loss in the current frame of the voice data stream. For example, if packet loss is detected in the current frame when receiving a voice data stream from a target object of terminal 130, a first data processing operation is performed on the current frame. The first data processing operation includes the following steps: identifying the current frame as a packet-loss frame; using a packet loss concealment model to determine the predicted voice parameters of the packet-loss frame based on one or more previous frames, for example, the packet loss concealment model determines the predicted voice parameters of the packet-loss frame based on the previous three frames, and the packet loss concealment model is established based on historical voice data of multiple objects; using a correction model to correct the predicted voice parameters of the packet-loss frame to obtain the corrected voice parameters of the packet-loss frame, wherein the correction model is established based on the historical voice data of the target object, for example, if the target object is an individual, the correction model is established based on the historical voice data of that individual, and is targeted to that individual; and determining the voice data of the packet-loss frame based on the corrected voice parameters of the packet-loss frame.

[0046] As an example, the server 110, terminal 120, 130, or other terminal devices may also perform a first data processing operation on the current frame in response to packet loss. Furthermore, the packet loss concealment model and correction model may be stored locally on the server 110, terminal 120, or 130, or stored on a cloud server, and transmitted and invoked via network 140.

[0047] Optionally, server 110 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The aforementioned terminals 120 and 130 can include, but are not limited to, at least one of the following: mobile phones, tablets, laptops, desktop PCs, digital televisions, and other terminals capable of displaying content. The network 140 can be, for example, a wide area network (WAN), a local area network (LAN), a wireless network, a public telephone network, an intranet, and any other type of network well known to those skilled in the art. It should also be noted that the scenario described above is merely one example in which embodiments of this disclosure can be implemented and is not restrictive.

[0048] It should be noted that the scenario described above is merely one example in which the embodiments of this disclosure can be implemented, and is not restrictive. For example, in some exemplary scenarios, packet loss concealment may also be implemented on specific terminals.

[0049] Figure 2 A schematic flowchart of a packet loss concealment method 200 according to an embodiment of the present disclosure is illustrated. The method 200 can be implemented, for example, in applications such as... Figure 1 On server 110 or terminal 120 as the receiver, this is not a limitation. Figure 2 As shown, the method 200 includes the following steps.

[0050] In step 201, the current frame in the speech data stream of the target object is determined, the speech data stream comprising multiple frames of speech data with target speech parameters. As an example, when the target object can be a personal object, the speech data stream of the target object can be a data stream obtained by encoding a segment of the target object's speech. In some embodiments, the speech parameters include one or more of the fundamental frequency, LSP (Line Spectral Pair), and gain of the speech data. As an example, the genetic frequency of the speech data can be selected as the speech parameter.

[0051] In step 202, it is determined whether packet loss has occurred in the current frame. As an example, packet loss in the current frame can be analyzed based on the communication protocol (or transport protocol) of the transmitted data stream. In response to packet loss in the current frame, a first data processing operation is performed on the current frame, which may include steps 203, 204, 205, and 206. Optionally, in response to no packet loss in the current frame, step 207 is executed. In step 207, a first other operation is performed on the current frame. In some embodiments, the first other operation may include the second data processing operation described in method 400 below. In some embodiments, the first other operation may include decoding the current frame, during which speech model parameters may be parsed to obtain the speech parameters of the current frame.

[0052] In step 203, the current frame is determined to be a lost frame.

[0053] In step 204, the predicted speech parameters of the lost frame are determined using a packet loss concealment model based on one or more preceding frames. This model is built upon historical speech data from multiple objects. The historical speech data used to build the model may or may not include the historical speech data of the target object. Therefore, since the packet loss concealment model is based on the historical speech data of multiple objects, the predicted speech parameters of the lost frame have broader applicability but are not specific to any particular object. For this reason, step 205 is then executed.

[0054] In step 205, the predicted speech parameters of the lost frame are corrected using a correction model to obtain the corrected speech parameters of the lost frame. The correction model is based on the historical speech data of the target object. For example, when the target object is an individual, the historical speech data can be a segment of their historical voice conversation or a recorded historical voice conversation, etc., encoded as such. It can be seen that since the correction model is based on the historical speech data of the target object, the corrected speech parameters better match the target object's speech data, making the correction model more targeted to the target object.

[0055] In step 206, the speech data of the lost frame is determined based on the corrected speech parameters of the lost frame. Since the corrected speech parameters better match the speech characteristics of the target object, the speech data of the lost frame determined based on the corrected speech parameters of the lost frame better matches the speech data before the packet loss occurred.

[0056] In method 200, a first data processing operation is performed on the current frame in response to packet loss. In this first data processing operation, the predicted speech parameters of the lost frame are corrected using a correction model to obtain corrected speech parameters for the lost frame. Since the correction model is based on the historical speech data of the target object, meaning it is targeted at the target object, the corrected speech parameters better match the target object's speech data. Therefore, the speech data of the lost frame determined based on the corrected speech parameters of the lost frame better matches the speech data before the packet loss occurred. Thus, packet loss hiding method 200 improves the effectiveness and quality of packet loss hiding by establishing and using a target-target-specific correction model to correct the biased and untargeted predicted speech parameters of the packet loss hiding model, making the corrected speech parameters more consistent with the speaker's characteristics.

[0057] Figure 3 A schematic flowchart illustrating a method 300 for determining a modified model according to an embodiment of the present disclosure is shown. The modified model is, for example, the modified model in the embodiment described with reference to 2. The method 300 can be implemented, for example, in applications such as... Figure 1 This can be done on server 110 or terminal 120 as the receiving end, though this is not a limitation. Figure 3 As shown, the method 300 includes the following steps.

[0058] In step 301, the target frame and the target speech parameters of the target frame are determined from the historical speech data of the target object. As an example, when the target object is an individual, the historical speech data of the target object may include speech data encoded from the individual's conversations, recordings, interviews, and other historical speech.

[0059] In step 302, the predicted speech parameters of the target frame are determined using a packet loss concealment model based on one or more preceding frames. The packet loss concealment model is, for example, the packet loss concealment model described in the embodiment referred to in 2. The packet loss concealment model is built upon historical speech data from multiple objects; therefore, the predicted speech parameters of the lost frame have broad applicability but are not specific to any particular object.

[0060] In step 303, the predicted speech parameters of the target frame are corrected using a deep learning model to obtain the corrected speech parameters of the target frame. As an example, the deep learning model can be a neural network model, a convolutional network model, etc., and is not limited here.

[0061] In step 304, the parameters of the deep learning model are adjusted to minimize the error between the corrected speech parameters of the target frame and the target speech parameters of the target frame, thereby obtaining the corrected model. The corrected speech parameters of the target frame output by the deep learning model change with the parameters of the deep learning model. By adjusting the parameters of the deep learning model to minimize the error between the corrected speech parameters of the target frame and the target speech parameters of the target frame, a deep learning model for the target object is obtained. As an example, when a neural network model is selected as the deep learning model, the output of the neural network model can be adjusted by adjusting the weights of the neurons in each layer of the neural network model to minimize the error between the corrected speech parameters of the target frame and the target speech parameters of the target frame, thereby obtaining a neural network model for the target object.

[0062] It should be noted that the steps for determining the correction model described above can be performed for each target frame in the historical speech data of the target object to obtain the correction model. Optionally, the steps for determining the correction model described above can also be performed for a subset of target frames in the historical speech data of the target object to obtain the correction model.

[0063] The method 300 utilizes historical speech data of the target object to train a correction model, making the correction model specific to the target object. It can correct the predicted speech parameters output by the packet loss hiding model, overcoming the shortcoming that the output of the packet loss hiding model cannot effectively match the speech features of the target object. The method 300 can be performed on server 110 or client 120, 130, and after generating the correction model, it can be invoked by server 110 or client 120, 130 via network 140.

[0064] Figure 4 A schematic flowchart illustrating a method 400 for updating and correcting a model according to an embodiment of the present disclosure is shown. The method 400 can be implemented, for example, in applications such as... Figure 1 This can be done on server 110 or terminal 120 as the receiving end, though this is not a limitation. Figure 4 As shown, the method 400 includes the following steps.

[0065] In step 401, the current frame in the speech data stream of the target object is determined, the speech data stream comprising multiple frames of speech data having target speech parameters. In some embodiments, step 401 may be the same as step 201 of method 200.

[0066] In step 402, it is determined whether packet loss has occurred in the current frame. As an example, packet loss in the current frame can be analyzed based on the communication protocol (or transport protocol) of the transmitted data stream. In response to the absence of packet loss in the current frame, a second data processing operation is performed on the current frame, including steps 403, 404, and 405.

[0067] Optionally, in response to packet loss in the current frame, step 406 is executed. In step 406, a second additional operation is performed on the current frame. In some embodiments, the second additional operation may include the first data processing operation in method 200 above.

[0068] In step 403, the packet loss hiding model is used to determine the predicted speech parameters of the current frame based on one or more previous frames. For example, the packet loss hiding model can output the predicted speech parameters of the current frame based on the three previous frames. In some embodiments, the packet loss hiding model is built based on historical speech data from multiple objects, thus the predicted speech parameters of the lost frame have broader applicability but are generally not specific to any particular object. Therefore, steps 404 and 405 are performed.

[0069] In step 404, the predicted speech parameters of the current frame are corrected using a correction model to obtain the corrected speech parameters for the current frame. Since the correction model is based on the historical speech data of the target object, it is specifically designed for that object. Therefore, the corrected speech parameters better match the target object's speech data compared to the predicted speech parameters. Furthermore, because the historical speech data of the target object can be continuously enriched—for example, when the target object is an individual—the historical speech data will increase as more frames without packet loss are received during a voice call. Therefore, the parameters of the correction model can be updated in step 405.

[0070] In step 405, the parameters of the correction model are updated to minimize the error between the corrected speech parameters of the current frame and the speech parameters of the current frame. The corrected speech parameters of the target frame output by the correction model change as the parameters of the correction model change. By adjusting the parameters of the correction model to minimize the error between the corrected speech parameters of the target frame and the target speech parameters of the target frame, a correction model for the target object is obtained. As an example, when a neural network model is selected as the correction model, the output of the neural network model can be adjusted by adjusting the weights of the neurons in each layer of the neural network model to minimize the error between the corrected speech parameters of the target frame and the target speech parameters of the target frame, thus obtaining a neural network model for the target object. For example, when the target object is a person, the correction training can be continuously updated during the person's normal conversation, so that the corrected speech parameters output by the correction model become closer and closer to the target speech parameters of the person, that is, the correction model becomes more and more suitable for the person.

[0071] In some embodiments, the type of the current frame can be determined based on the changing trend of the target speech parameters of the previous one or more frames in the speech data stream. Then, a correction model corresponding to the type of the current frame is used to correct the predicted speech parameters of the lost frame to obtain the corrected speech parameters of the lost frame. For example, the type of the current frame can be determined based on the changing trend of the target speech parameters of the previous 8 frames in the speech data stream. For instance, if the i-th frame in the speech data stream is larger than the (i-1)-th frame, 0 represents the changing trend of the i-th frame in the speech data stream; if the i-th frame in the speech data stream is unchanged compared to the (i-1)-th frame, 1 represents the changing trend of the i-th frame in the speech data stream; if the i-th frame in the speech data stream is smaller than the (i-1)-th frame, 2 represents the changing trend of the i-th frame in the speech data stream. For example, 00000000 indicates the type of the current frame where the speech parameters of the previous 8 frames increase sequentially. 00001111 indicates the type of the current frame where the speech parameters of the previous 8 frames first increase sequentially and then remain unchanged. As an example, after determining that the type of the current frame is 00001111, the predicted speech parameters of the lost frame are corrected using the correction model corresponding to the type 00001111 of the current frame to obtain the corrected speech parameters of the lost frame. Generally, the variation trend of the target speech parameters of the previous N frames in the speech data stream can be 3 to the power of N, that is, the type corresponding to the current frame can be 3 to the power of N.

[0072] In some embodiments, the correction model corresponding to the type of the current frame can be updated according to the type of the current frame. Figure 5 A schematic flowchart of a method 500 for updating a correction model corresponding to the type of the current frame according to an embodiment of the present disclosure is shown. The method 500 includes the following steps.

[0073] In step 501, the current frame in the speech data stream of the target object is determined, the speech data stream comprising multiple frames of speech data having target speech parameters. In some embodiments, step 501 may be the same as step 201 of method 200.

[0074] In step 502, the type of the current frame is determined, whereby the type of the current frame represents the trend of change in the target speech parameters of the previous one or more frames. For example, based on the trend of change in the target speech parameters of the previous four frames in the speech data stream, which shows an initial increase followed by a plateau, the type of the current frame is determined to be 0011.

[0075] In step 503, it is determined whether packet loss has occurred in the current frame. As an example, packet loss in the current frame can be analyzed based on the communication protocol (or transport protocol) of the transmitted data stream. In response to the current frame not experiencing packet loss, a third data processing operation is performed on the current frame, including steps 504, 505, and 506. Optionally, in response to the current frame experiencing packet loss, step 507 is executed. In step 507, a third additional operation is performed on the current frame. In some embodiments, the third additional operation includes correcting the predicted speech parameters of the current frame output by the packet loss concealment model using a correction model corresponding to the type of the current frame, to obtain corrected speech parameters for the current frame. In step 504, the predicted speech parameters of the current frame are determined using the packet loss concealment model based on one or more previous frames. In some embodiments, step 504 can be the same as step 403 of method 400.

[0076] In step 505, the predicted speech parameters of the current frame are corrected using a correction model corresponding to the type of the current frame to obtain the corrected speech parameters of the current frame. For example, when the type of the current frame is determined to be 0011, the predicted speech parameters of the current frame are corrected using a correction model corresponding to the type 0011 of the current frame to obtain the corrected speech parameters of the current frame.

[0077] In step 506, the parameters of the correction model corresponding to the type of the current frame are updated to minimize the error between the corrected speech parameters of the current frame and the speech parameters that the current frame already has. For example, when the type of the current frame is determined to be 0011, the parameters of the correction model corresponding to the type 0011 of the current frame are updated to minimize the error between the corrected speech parameters of the current frame and the speech parameters that the current frame already has.

[0078] In some embodiments, for the correction model corresponding to the type of the current frame, there are multiple historical prediction errors. Each historical prediction error is the error between the predicted speech parameters of the historical frame obtained by the packet loss hiding model and the target speech parameters of the historical frame, wherein the historical frame has the same type as the current frame. For example, when the type of the current frame is determined to be 0011, the type of the historical frame should also be 0011. In some embodiments, updating the parameters of the correction model corresponding to the type of the current frame to minimize the error between the corrected speech parameters of the current frame and the target speech parameters of the current frame includes: using the multiple historical prediction errors to update the parameters of the correction model corresponding to the type of the current frame to minimize the error between the corrected speech parameters of the current frame and the target speech parameters of the current frame. For example, if the trend of the previous four frames remains unchanged, i.e., the type of the current frame is determined to be 1111, the parameters of the correction model corresponding to the type 1111 of the current frame are updated using the multiple historical prediction errors to minimize the error between the corrected speech parameters of the current frame and the target speech parameters of the current frame.

[0079] In some embodiments, updating the parameters of the correction model corresponding to the type of the current frame using multiple historical prediction errors to minimize the error between the corrected speech parameters of the current frame and the target speech parameters of the current frame includes: updating the parameters of the correction model corresponding to the type of the current frame such that the difference between the error between the corrected speech parameters of the current frame and the predicted speech parameters of the current frame, and the center error of the multiple historical prediction errors, is minimized, thereby minimizing the error between the corrected speech parameters of the current frame and the target speech parameters of the current frame. For example, if the type of the current frame is 1111, and the historical prediction errors include five historical prediction errors of the correction model corresponding to the type 1111 of the current frame, and the center error of the five historical prediction errors is 3.6, then updating the parameters of the correction model corresponding to the type 1111 of the current frame minimizes the difference between the error between the corrected speech parameters of the current frame and the predicted speech parameters of the current frame, and the center error of the five historical prediction errors (3.6), thereby minimizing the error between the corrected speech parameters of the current frame and the target speech parameters of the current frame.

[0080] In some embodiments, the central error of multiple historical prediction errors can be determined through the following error determination steps, which include: establishing a point cluster based on the multiple historical prediction errors, where each point in the point cluster represents a historical prediction error; clustering the point cluster to obtain a target point set, for example, the clustering can employ algorithms such as k-means clustering, hierarchical clustering, SOM clustering, FCM clustering, etc., without limitation; determining the central value of the historical prediction errors corresponding to the points in the target point set; and determining the central value as the central error of the multiple historical prediction errors. In some embodiments, clustering the point cluster to obtain the target point set includes: removing outliers from the point cluster, where outliers are points in the point cluster whose average distance from other points in the point cluster is greater than a preset threshold; and determining the target point set such that the target point set includes points other than the outliers in the point cluster. Since the determined target point set does not contain outliers, the central error of the multiple historical prediction errors determined based on the target point set is more representative.

[0081] Figure 6 The illustration shows a schematic diagram of a packet loss concealment method according to an embodiment of the present disclosure. Figure 6 As shown, the original recording (e.g., a telephone conversation with an individual) is encoded into a voice data stream and then sent to the receiving end via a channel (which could be...). Figure 1 The receiving end (as shown in server 110 or terminal 120, 130, etc.) confirms whether packet loss has occurred in the current frame of the voice data stream. In response to packet loss in the current frame, a packet loss concealment model is used to determine the predicted speech parameters of the current frame based on one or more previous frames (e.g., the frame preceding the current frame). Then, a correction model is used to correct the predicted speech parameters of the current frame, and the corrected speech parameters of the current frame are output. Finally, the corrected speech parameters are decoded and output, thus recovering the speech parameters of the lost current frame.

[0082] If no packet loss occurs in the current frame, the speech model parameters of the current frame are parsed (and optionally decoded) to obtain the target speech parameters (i.e., the true speech parameters of the current frame). The packet loss hiding model is then used to determine the predicted speech parameters of the current frame based on one or more previous frames (e.g., the frame before the current frame). The predicted speech parameters of the current frame are then corrected using a correction model to obtain the corrected speech parameters. The parameters of the correction model are then updated to minimize the error between the corrected speech parameters and the target speech parameters of the current frame. Ultimately, this update of the correction model ensures that its output better matches the speech characteristics of the target object.

[0083] As an example, the speech parameters (also known as speech model parameters) of the current frame in the speech data stream can be the fundamental frequency, LSP, gain, etc. Figure 7 The illustration shows a spectrogram when gene frequencies were selected as the speech parameter. For example... Figure 7 As shown, the horizontal axis represents time, and the vertical axis represents frequency. The darker the stripe color, the higher the energy value at that point. The dark lines represent the fundamental tone and corresponding harmonic positions of voiced sounds. The fundamental tone frequency is the lowest value (the first dark curve in the figure, such as in...). Figure 7 The frequency values ​​shown in the figure indicate that the fundamental frequency is continuous within a morpheme, and therefore its trend is predictable. Thus, in some embodiments, the speech parameters of the current frame are often selected from the gene frequency of the current frame.

[0084] Figure 8 An exemplary structural block diagram of a packet loss concealment device 800 according to an embodiment of the present disclosure is shown. Figure 8 As shown, the packet loss hiding device 800 includes a determination module 810 and a processing module 820.

[0085] The determining module 810 is configured to determine the current frame in the speech data stream of the target object, the speech data stream comprising multiple frames of speech data with target speech parameters. As an example, when the target object can be a personal object, the speech data stream of the target object can be obtained by encoding a segment of the target object's speech. In some embodiments, the speech parameters include one or more of the fundamental frequency, LSP, and gain of the speech data; as an example, the fundamental frequency of the speech data can be selected as the speech parameter.

[0086] The processing module 820 is configured to perform a first data processing operation on the current frame in response to packet loss. The first data processing operation includes the following steps: identifying the current frame as a packet-loss frame; determining the predicted speech parameters of the packet-loss frame based on one or more previous frames using a packet loss hiding model, wherein the packet loss hiding model is established based on historical speech data of multiple objects. In some embodiments, the historical speech data of the multiple objects used to establish the packet loss hiding model may include the historical speech data of the target object, or may not include the historical speech data of the target object; and correcting the predicted speech parameters of the packet-loss frame using a correction model to obtain the corrected speech parameters of the packet-loss frame, wherein the correction model is established based on the historical speech data of the target object. As an example, when the target object is an individual, the historical speech data may be a segment of the individual's historical voice calls or encoded speech data such as historical recordings. The correction model is built based on the historical speech data of the target object. Therefore, the corrected speech parameters output by the correction model can better match the target speech data of the target object, making the correction model targeted to the target object. The speech data of the lost frame is determined according to the corrected speech parameters of the lost frame. Since the corrected speech parameters are more in line with the speech characteristics of the target object, the speech data of the lost frame determined according to the corrected speech parameters of the lost frame is more in line with the speech data before the loss of the lost frame.

[0087] The processing module 820 is further configured to perform a second data processing operation on the current frame in response to the absence of packet loss in the current frame. The second data processing operation includes the following steps: A packet loss concealment model is used to determine the predicted speech parameters of the current frame based on one or more previous frames. For example, the model can output the predicted speech parameters based on the first three frames. Since this model is built upon historical speech data from multiple objects, the predicted speech parameters for a lost frame may not perfectly match the target speech data of the target object. A correction model is then used to correct the predicted speech parameters of the current frame, resulting in corrected speech parameters. Because this correction model is based on the historical speech data of the target object, it is targeted to that object, and therefore the corrected speech parameters better match the target speech data compared to the predicted parameters. Furthermore, the historical speech data of the target object can be continuously enriched; for example, when the target object is an individual, the historical speech data increases as more frames without packet loss are received during a voice call. The parameters of the correction model are updated to minimize the error between the corrected speech parameters and the existing speech parameters of the current frame. The corrected speech parameters of the target frame output by the correction model change with the parameters of the correction model. By adjusting the parameters of the correction model to minimize the error between the corrected speech parameters of the target frame and the target speech parameters of the target frame, a correction model for the target object is obtained. As an example, when a neural network model is selected as the correction model, the output of the neural network model can be adjusted by changing the weights of neurons in each layer, minimizing the error between the corrected speech parameters of the output target frame and the target speech parameters of the target frame, thus obtaining a neural network model for the target object. For example, when the target object is a person, the correction training can be continuously updated during the person's normal conversation, making the corrected speech parameters output by the correction model increasingly closer to the target speech parameters of the person, i.e., the correction model becomes increasingly suitable for the person.

[0088] Figure 9 The illustration depicts an example system 900, which includes an example computing device 910 representing one or more systems and / or devices that can implement the various technologies described herein. The computing device 910 may be, for example, a server of a service provider, a device associated with a server, a system-on-a-chip, and / or any other suitable computing device or computing system. (Refer to above) Figure 8 The packet loss hiding device 800 described can take the form of a computing device 910. Alternatively, the packet loss hiding device 800 can be implemented as a computer program as an application 916.

[0089] The example computing device 910 shown includes a processing system 911, one or more computer-readable media 912, and one or more I / O interfaces 913, all communicatively coupled to each other. Although not shown, the computing device 910 may also include a system bus or other data and command transfer system that couples the various components to each other. The system bus may include any or a combination of different bus architectures, such as a memory bus or memory controller, a peripheral bus, a universal serial bus, and / or a processor or local bus utilizing any of these various bus architectures. Various other examples, such as control and data lines, are also conceived.

[0090] Processing system 911 represents the functionality of performing one or more operations using hardware. Therefore, processing system 911 is illustrated as including hardware elements 914 that can be configured as processors, function blocks, etc. This may include other logic devices implemented in hardware as application-specific integrated circuits (ASICs) or formed using one or more semiconductors. Hardware element 914 is not limited by the materials in which it is formed or the processing mechanism employed therein. For example, a processor may consist of semiconductors and / or transistors (e.g., integrated circuits (ICs)). In such a context, processor-executable instructions may be electronically executable instructions.

[0091] Computer-readable medium 912 is illustrated as including memory / storage device 915. Memory / storage device 915 represents a memory / storage capacity associated with one or more computer-readable media. Memory / storage device 915 may include volatile media (such as random access memory (RAM)) and / or non-volatile media (such as read-only memory (ROM), flash memory, optical disk, magnetic disk, etc.). Memory / storage device 915 may include fixed media (e.g., RAM, ROM, fixed hard disk drive, etc.) and removable media (e.g., flash memory, removable hard disk drive, optical disk, etc.). Computer-readable medium 912 may be configured in various other ways as further described below.

[0092] One or more I / O interfaces 913 represent the functionality to allow users to input commands and information to the computing device 910 using various input devices and optionally also to present information to the user and / or other components or devices using various output devices. Examples of input devices include keyboards, cursor control devices (e.g., mice), microphones (e.g., for voice input), scanners, touch functionality (e.g., capacitive or other sensors configured to detect physical touch), cameras (e.g., capable of detecting non-touch-related movements as gestures using visible or invisible wavelengths (such as infrared frequencies), etc. Examples of output devices include display devices (e.g., monitors or projectors), speakers, printers, network interface cards, haptic-responsive devices, etc. Therefore, the computing device 910 can be configured to support user interaction in various ways as further described below.

[0093] The computing device 910 also includes an application 916. The application 916 may be, for example, a software instance of the packet loss hiding device 800, and may be used in combination with other elements in the computing device 910 to implement the techniques described herein.

[0094] This document describes various technologies within the general context of software and hardware components or program modules. Generally, these modules include routines, programs, objects, elements, components, data structures, etc., that perform specific tasks or implement specific abstract data types. As used herein, the terms "module," "function," and "component" generally refer to software, firmware, hardware, or a combination thereof. The technologies described herein are characterized as platform-independent, meaning that these technologies can be implemented on a variety of computing platforms with various processors.

[0095] Implementations of the described modules and technologies may be stored on or transmitted across some form of computer-readable medium. The computer-readable medium may include a variety of media accessible by the computing device 910. By way of example and not limitation, the computer-readable medium may include "computer-readable storage media" and "computer-readable signal media".

[0096] In contrast to simple signal transmission, carrier waves, or signals themselves, a "computer-readable storage medium" refers to a medium and / or device capable of persistently storing information, and / or a tangible storage device. Therefore, a computer-readable storage medium refers to a non-signal-bearing medium. Computer-readable storage media include hardware such as volatile and non-volatile, removable and non-removable media and / or storage devices implemented using methods or techniques suitable for storing information (such as computer-readable instructions, data structures, program modules, logic elements / circuits, or other data). Examples of computer-readable storage media may include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disc (DVD) or other optical storage devices, hard disks, cassette tapes, magnetic tapes, disk storage devices or other magnetic storage devices, or other storage devices, tangible media, or articles of art suitable for storing desired information and accessible by a computer.

[0097] "Computer-readable signal medium" refers to a signal-bearing medium configured to transmit instructions, such as via a network, to computing device 910. A signal medium typically embodies computer-readable instructions, data structures, program modules, or other data in a modulated data signal, such as a carrier wave, data signal, or other transmission mechanism. Signal media also include any information transmission medium. The term "modulated data signal" refers to a signal in which one or more of its characteristics are set or altered to encode information. By way of example and not limitation, communication media include wired media such as wired networks or direct connections, and wireless media such as acoustic, RF, infrared, and other wireless media.

[0098] As previously described, hardware element 914 and computer-readable medium 912 represent instructions, modules, programmable device logic, and / or fixed device logic implemented in hardware, which in some embodiments can be used to implement at least some aspects of the techniques described herein. Hardware elements may include components of integrated circuits or systems-on-a-chip, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), complex programmable logic devices (CPLDs), and other implementations or other hardware devices in silicon. In this context, hardware elements can serve as processing devices for executing program tasks defined by instructions, modules, and / or logic embodied by the hardware element, and as hardware devices for storing instructions for execution, such as the previously described computer-readable storage medium.

[0099] The foregoing combinations can also be used to implement the various techniques and modules described herein. Therefore, software, hardware, or program modules and other program modules can be implemented as one or more instructions and / or logic embodied on some form of computer-readable storage medium and / or by one or more hardware elements 914. The computing device 910 can be configured to implement specific instructions and / or functions corresponding to the software and / or hardware modules. Thus, for example, by using the computer-readable storage medium and / or hardware elements 914 of a processing system, modules can be implemented at least partially in hardware as modules executable as software by the computing device 910. Instructions and / or functions can be executable / operable by one or more articles of art (e.g., one or more computing devices 910 and / or processing systems 911) to implement the techniques, modules, and examples described herein.

[0100] In various embodiments, the computing device 910 can be configured in various ways. For example, the computing device 910 can be implemented as a computer-type device, including personal computers, desktop computers, multi-screen computers, laptop computers, netbooks, etc. The computing device 910 can also be implemented as a mobile device, including mobile devices such as mobile phones, portable music players, portable gaming devices, tablet computers, multi-screen computers, etc. The computing device 910 can also be implemented as a television-type device, including devices with or connected to a generally large screen in a leisure viewing environment. These devices include televisions, set-top boxes, game consoles, etc.

[0101] The techniques described herein can be supported by these various configurations of computing device 910, and are not limited to specific examples of the techniques described herein. Functionality can also be implemented, wholly or partially, on the “cloud” 920 using distributed systems, such as through platform 922 as described below.

[0102] Cloud 920 includes and / or represents platform 922 for resource 924. Platform 922 abstracts the underlying functionality of the hardware (e.g., server) and software resources of cloud 920. Resource 924 may include applications and / or data that can be used when performing computer processing on a server located remotely from computing device 910. Resource 924 may also include services provided via the Internet and / or via subscriber networks such as cellular or Wi-Fi networks.

[0103] Platform 922 can abstract resources and functions to connect computing device 910 to other computing devices. Platform 922 can also be used to abstract resource hierarchy to provide a corresponding level of hierarchy for any encountered needs for resource 924 implemented via platform 922. Therefore, in interconnect device embodiments, the implementation of the functions described herein can be distributed throughout system 900. For example, functions can be implemented partly on computing device 910 and partly through platform 922, which abstracts the functions of cloud 920.

[0104] This disclosure provides a computer-readable storage medium having computer-readable instructions stored thereon, which, when executed, implement the methods described above for determining associated applications or for determining recommended content.

[0105] This disclosure provides a computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computing device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computing device to perform the methods provided in the various optional implementations above for determining associated applications or for determining recommended content.

[0106] It should be understood that, for clarity, embodiments of this disclosure have been described with reference to different functional units. However, it will be apparent that, without departing from this disclosure, the functionality of each functional unit may be implemented in a single unit, in multiple units, or as part of other functional units. For example, functionality described as being performed by a single unit may be performed by multiple different units. Therefore, references to a particular functional unit are considered merely as references to the appropriate unit used to provide the described functionality, and not as indicating a strict logical or physical structure or organization. Thus, this disclosure may be implemented in a single unit, or may be physically and functionally distributed among different units and circuits.

[0107] It will be understood that although the terms first, second, third, etc., may be used herein to describe various devices, elements, components, or parts, these devices, elements, components, or parts should not be limited by these terms. These terms are used only to distinguish one device, element, component, or part from another device, element, component, or part.

[0108] Although this disclosure has been described in conjunction with some embodiments, it is not intended to be limited to the specific forms set forth herein. Rather, the scope of this disclosure is limited only by the appended claims. Additionally, although individual features may be included in different claims, these may be advantageously combined, and inclusion in different claims does not imply that such a combination of features is not feasible and / or advantageous. The order of features in the claims does not imply that the features must be in any particular order of their operation. Furthermore, in the claims, the word "comprising" does not exclude other elements, and the terms "a" or "an" do not exclude a plurality. Reference numerals in the claims are provided only by way of explicit example and should not be construed as limiting the scope of the claims in any way.

Claims

1. A method for packet loss hiding, comprising: Determine the current frame in the speech data stream of the target object, wherein the speech data stream includes multiple frames of speech data with target speech parameters; In response to packet loss in the current frame, a first data processing operation is performed on the current frame, the first data processing operation including the following steps: The current frame is identified as a lost frame; The packet loss hiding model is used to determine the predicted speech parameters of the lost frame based on one or more frames preceding the lost frame. The packet loss hiding model is built based on historical speech data of multiple objects. The predicted speech parameters of the lost frame are corrected using a correction model to obtain the corrected speech parameters of the lost frame, wherein the correction model is established based on the historical speech data of the target object. The speech data of the lost frame is determined based on the corrected speech parameters of the lost frame; The correction model is obtained by training a deep learning model based on the historical speech data of the target object through the following training steps: Determine the target frame and the target speech parameters of the target frame from the historical speech data of the target object; The predicted speech parameters of the target frame are determined using a packet loss hiding model based on one or more preceding frames. The predicted speech parameters of the target frame are corrected using a deep learning model to obtain the corrected speech parameters of the target frame. The parameters of the deep learning model are adjusted to minimize the error between the corrected speech parameters of the target frame and the target speech parameters of the target frame, thereby obtaining the corrected model.

2. The method according to claim 1, further comprising: In response to the absence of packet loss in the current frame, a second data processing operation is performed on the current frame, the second data processing operation including the following steps: The packet loss hiding model is used to determine the predicted speech parameters of the current frame based on one or more previous frames. The predicted speech parameters of the current frame are corrected using a correction model to obtain the corrected speech parameters of the current frame. Update the parameters of the correction model to minimize the error between the corrected speech parameters of the current frame and the target speech parameters of the current frame.

3. The method according to claim 1, further comprising: Determine the type of the current frame, where the type of the current frame represents the changing trend of the target speech parameters in the previous one or more frames; Specifically, the predicted speech parameters of the lost packet frame are corrected using a correction model to obtain the corrected speech parameters of the lost packet frame, including: The predicted speech parameters of the lost frame are corrected using a correction model corresponding to the type of the current frame, so as to obtain the corrected speech parameters of the lost frame.

4. The method according to claim 3, further comprising: In response to the absence of packet loss in the current frame, a third data processing operation is performed on the current frame, which includes the following steps: The packet loss hiding model is used to determine the predicted speech parameters of the current frame based on one or more previous frames. The predicted speech parameters of the current frame are corrected using a correction model corresponding to the type of the current frame, so as to obtain the corrected speech parameters of the current frame. Update the parameters of the corrected model corresponding to the type of the current frame, so as to minimize the error between the corrected speech parameters of the current frame and the target speech parameters of the current frame.

5. The method according to claim 4, wherein, For the correction model corresponding to the type of the current frame, there are multiple historical prediction errors. Each historical prediction error is the error between the predicted speech parameters of the historical frame obtained by the packet loss hiding model and the target speech parameters of the historical frame. The historical frame is of the same type as the current frame. The updating of the parameters of the correction model corresponding to the type of the current frame, so as to minimize the error between the corrected speech parameters of the current frame and the target speech parameters of the current frame, includes: The parameters of the corrected model corresponding to the type of the current frame are updated using the multiple historical prediction errors, so as to minimize the error between the corrected speech parameters of the current frame and the target speech parameters of the current frame.

6. The method according to claim 5, wherein, Updating the parameters of the corrected model corresponding to the type of the current frame using the multiple historical prediction errors minimizes the error between the corrected speech parameters of the current frame and the target speech parameters of the current frame, including: Update the parameters of the corrected model corresponding to the type of the current frame, so that the difference between the error between the corrected speech parameters of the current frame and the predicted speech parameters of the current frame, and the difference between the central error of the plurality of historical prediction errors, is minimized, so that the error between the corrected speech parameters of the current frame and the target speech parameters of the current frame is minimized.

7. The method according to claim 6, wherein, The central error of the plurality of historical prediction errors is determined by the following error determination steps, which include: A point cluster is established based on the multiple historical prediction errors, where each point in the point cluster represents a historical prediction error. Cluster the point group to obtain the target point set; Determine the center value of the historical prediction error corresponding to the points in the target point set; The central value is determined as the central error of the plurality of historical prediction errors.

8. The method according to claim 7, wherein, Clustering the point group to obtain the target point set includes: Remove outliers from the point group, where an outlier is a point in the point group whose average distance from other points in the point group is greater than a preset threshold. Determine a target point set such that the target point set includes points other than the outliers in the point group.

9. The method according to claim 1, wherein the target speech parameters include one or more of the fundamental frequency, line spectrum pairs, and gain of the speech data.

10. A device for hiding packet loss, comprising: A determination module is configured to determine the current frame in the speech data stream of the target object, the speech data stream comprising multiple frames of speech data having target speech parameters. The processing module is configured to perform a first data processing operation on the current frame in response to packet loss, the first data processing operation including the following steps: The current frame is identified as a lost frame; The packet loss hiding model is used to determine the predicted speech parameters of the lost frame based on one or more frames preceding the lost frame. The packet loss hiding model is built based on historical speech data of multiple objects. The predicted speech parameters of the lost frame are corrected using a correction model to obtain the corrected speech parameters of the lost frame, wherein the correction model is established based on the historical speech data of the target object. The speech data of the lost frame is determined based on the corrected speech parameters of the lost frame; The processing module is further configured to: Determine the target frame and the target speech parameters of the target frame from the historical speech data of the target object; The predicted speech parameters of the target frame are determined using a packet loss hiding model based on one or more preceding frames. The predicted speech parameters of the target frame are corrected using a deep learning model to obtain the corrected speech parameters of the target frame. The parameters of the deep learning model are adjusted to minimize the error between the corrected speech parameters of the target frame and the target speech parameters of the target frame, thereby obtaining the corrected model.

11. The apparatus of claim 10, wherein the processing module is further configured to: In response to the absence of packet loss in the current frame, a second data processing operation is performed on the current frame, the second data processing operation including the following steps: The packet loss hiding model is used to determine the predicted speech parameters of the current frame based on one or more previous frames. The predicted speech parameters of the current frame are corrected using a correction model to obtain the corrected speech parameters of the current frame. Update the parameters of the correction model to minimize the error between the corrected speech parameters of the current frame and the speech parameters of the current frame.

12. A computing device, comprising: Memory, which is configured to store computer-executable instructions; A processor configured to perform the method according to any one of claims 1-9 when the computer-executable instructions are executed by the processor.

13. A computer-readable storage medium storing computer-executable instructions that, when executed, perform the method according to any one of claims 1-9.

14. A computer program product comprising computer-executable instructions, wherein the computer-executable instructions, when executed by a processor, perform the method according to any one of claims 1-9.

Citation Information

Patent Citations

  • Speech recognition method, device, electronic device and storage medium

    CN109243468A

  • Voice processing method, device and equipment and storage medium

    CN111554322A