Voice processing method and device, equipment, storage medium and program product

The receiving device receives and processes voices whose voice delay is greater than the specified delay, and uses the sound source probability to accelerate playback processing, solving the problem of degraded voice playback quality, realizing the reduction of voice delay and improved intelligibility.

CN120302081APending Publication Date: 2025-07-11TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410041540.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-10
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

In a scenario where voice is transmitted over the network and voice playback is performed, when the voice delay is greater than the specified time delay, periodic discarding some voices will affect the intelligibility of the voice, resulting in a degradation of the playback quality.

Method used

The receiving device receives the pending voice sequence and its sound source probability, and when the current voice delay is greater than the specified time delay, the playback acceleration process is performed on the pending voice whose sound source probability is less than the specified probability, including discarding or accelerating the playback speed, and adjusting the playback strategy in combination with the voice energy.

Benefits of technology

Effectively reduce voice delay, reduce the impact on voice intelligibility, improve voice playback quality, especially in multi-source playback scenarios to reduce the impact on non-main sound sources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120302081A_ABST
    Figure CN120302081A_ABST
Patent Text Reader

Abstract

The invention provides a voice processing method and device, equipment, a storage medium and a program product, which are applied to various voice playing scenes such as cloud technology, artificial intelligence, smart traffic, maps, vehicle-mounted and games. The voice processing method comprises: receiving a to-be-processed voice sequence sent by a sending end device for a first sound source and a sound source probability corresponding to each to-be-processed voice in the to-be-processed voice sequence, the sound source probability being a probability that the to-be-processed voice is a sound emitted by the first sound source; when the current voice time delay is greater than the specified voice time delay, performing playing acceleration processing on the to-be-processed voice of which the sound source probability is less than the specified probability in the to-be-processed voice sequence to obtain a to-be-played voice sequence; and for the first sound source, performing first voice playing based on the to-be-played voice sequence. According to the invention, the voice playing quality can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to voice processing technology in the field of computer applications, and in particular, to a voice processing method, apparatus, device, storage medium, and program product. Background Art

[0002] In a scenario where voice is transmitted over a network and played, when the current voice delay is greater than the specified voice delay, in order to reduce the voice delay, some voices are usually discarded periodically; however, periodically discarding some voices loses the original voice and affects the intelligibility of the voice. Therefore, the voice playback quality is affected. Summary of the Invention

[0003] Embodiments of this application provide a voice processing method, apparatus, device, storage medium, and program product, which can improve the voice playback quality.

[0004] The technical solution of the embodiments of this application is implemented as follows:

[0005] Embodiments of this application provide a voice processing method, and the method includes:

[0006] Receiving a to-be-processed voice sequence sent by a sending-end device for a first sound source and a sound source probability corresponding to each to-be-processed voice in the to-be-processed voice sequence, where the sound source probability refers to the probability that the to-be-processed voice is a sound emitted by the first sound source;

[0007] When the current voice delay is greater than the specified voice delay, in the to-be-processed voice sequence, performing playback acceleration processing on the to-be-processed voice whose sound source probability is less than the specified probability to obtain a to-be-played voice sequence;

[0008] Performing first voice playback based on the to-be-played voice sequence for the first sound source.

[0009] Embodiments of this application also provide a voice processing method, and the method includes:

[0010] In response to a voice collection instruction, collecting voice for a first sound source to obtain a to-be-processed voice sequence;

[0011] Determining a sound source probability based on a to-be-processed feature corresponding to each to-be-processed voice in the to-be-processed voice sequence, where the sound source probability refers to the probability that the to-be-processed voice is a sound emitted by the first sound source;

[0012] Send the to-be-processed voice sequence and the sound source probability of each to-be-processed voice to a receiving-end device, where the receiving-end device is configured to perform a playback acceleration process on the to-be-processed voice with a sound source probability less than a specified probability to implement first voice playback when a current voice delay is greater than a specified voice delay.

[0013] An embodiment of the present application provides a first voice processing device, where the first voice processing device includes:

[0014] A voice receiving module, configured to receive a to-be-processed voice sequence sent by a sending-end device for a first sound source and a sound source probability corresponding to each to-be-processed voice in the to-be-processed voice sequence, where the sound source probability refers to the probability that the to-be-processed voice is a sound emitted by the first sound source;

[0015] A voice processing module, configured to perform a playback acceleration process on the to-be-processed voice with a sound source probability less than a specified probability in the to-be-processed voice sequence to obtain a to-be-played voice sequence when a current voice delay is greater than a specified voice delay;

[0016] A voice playback module, configured to perform first voice playback for the first sound source based on the to-be-played voice sequence.

[0017] In an embodiment of the present application, the playback acceleration process is discarding or increasing the playback speed, where discarding means discarding the to-be-processed voice with a sound source probability less than a specified probability, and increasing the playback speed means increasing the original playback speed of the to-be-processed voice with a sound source probability less than a specified probability at a target acceleration.

[0018] In an embodiment of the present application, the voice receiving module is further configured to receive the to-be-processed voice sequence sent by the sending-end device for the first sound source and the sound source probability corresponding to each to-be-processed voice in the to-be-processed voice sequence during a process of performing second voice playback for a second sound source.

[0019] In an embodiment of the present application, the voice processing module is further configured to obtain a current voice energy of the second voice when the current voice delay is greater than the specified voice delay; and perform a playback acceleration process on the to-be-processed voice with a sound source probability less than the specified probability in the to-be-processed voice sequence based on the current voice energy to obtain the to-be-played voice sequence.

[0020] In an embodiment of the present application, the voice processing module is further configured to, when the current voice energy is less than the first specified energy, in the to-be-processed voice sequence, accelerate the playing speed of the to-be-processed voice with a sound source probability less than the specified probability, to obtain the to-be-played voice sequence, where the playing acceleration process is the acceleration of the playing speed.

[0021] In an embodiment of the present application, the voice processing module is further configured to, when the current voice energy is greater than the second specified energy, discard, from the to-be-processed voice sequence, the to-be-processed voice with a sound source probability less than the specified probability, to obtain the to-be-played voice sequence, where the second specified energy is less than or equal to the first specified energy, and the playing acceleration process is the discarding.

[0022] In an embodiment of the present application, the voice processing module is further configured to obtain the previous discard time, where the previous discard time is the time when the voice discarding process was last executed; when the duration between the previous discard time and the current time is greater than or equal to the specified cycle duration, discard, from the to-be-processed voice sequence, the to-be-processed voice with a sound source probability less than the specified probability, to obtain the to-be-played voice sequence.

[0023] In an embodiment of the present application, the voice playing module is further configured to perform voice decoding on the to-be-played voice sequence to obtain a first voice sequence; for the first sound source, perform first voice playing based on the first voice sequence.

[0024] In an embodiment of the present application, the voice processing module is further configured to, when the current voice delay is less than or equal to the specified voice delay, perform voice decoding on the to-be-processed voice sequence to obtain a second voice sequence; for the first sound source, perform first voice playing based on the second voice sequence.

[0025] An embodiment of the present application provides a second voice processing device, where the second voice processing device includes:

[0026] A voice acquisition module, configured to, in response to a voice acquisition instruction, acquire voice of a first sound source to obtain a to-be-processed voice sequence;

[0027] A probability determination module, configured to determine a sound source probability based on the to-be-processed features corresponding to each to-be-processed voice in the to-be-processed voice sequence, where the sound source probability refers to the probability that the to-be-processed voice is the voice emitted by the first sound source;

[0028] A voice sending module, configured to send the to-be-processed voice sequence and the sound source probability of each to-be-processed voice to a receiving-end device, where the receiving-end device is configured to, when a current voice delay is greater than a specified voice delay, perform a playback acceleration process on the to-be-processed voice with a sound source probability less than a specified probability to implement a first voice playback.

[0029] In an embodiment of the present application, the voice acquisition module is configured to, in response to the voice acquisition instruction, perform voice acquisition on the first sound source to obtain an initial voice sequence; perform noise reduction on the initial voice sequence to obtain a to-be-encoded voice sequence; and perform voice encoding on the to-be-encoded voice sequence to obtain the to-be-processed voice sequence.

[0030] In an embodiment of the present application, the acquisition of the sound source probability of each to-be-processed voice is implemented through a sound source classification model. The second voice processing device further includes a model training module, configured to obtain initial training data of a to-be-trained model, where the to-be-trained model is a neural network model to be trained for determining the probability that a voice is emitted by the first sound source; perform scene classification on the initial training data to obtain a plurality of scene training data; perform sound source classification on the plurality of scene training data by using the to-be-trained model to obtain a plurality of sound source classification results; and adjust model parameters of the to-be-trained model based on the plurality of sound source classification results to obtain the sound source classification model.

[0031] An embodiment of the present application provides a receiving-end device for voice processing. The receiving-end device includes:

[0032] A first memory, configured to store computer-executable instructions or a computer program;

[0033] A first processor, configured to, when executing the computer-executable instructions or the computer program stored in the first memory, implement the voice processing method applied to the receiving-end device provided in the embodiment of the present application.

[0034] An embodiment of the present application provides a sending-end device for voice processing. The sending-end device includes:

[0035] A second memory, configured to store computer-executable instructions or a computer program;

[0036] A second processor, configured to, when executing the computer-executable instructions or the computer program stored in the second memory, implement the voice processing method applied to the sending-end device provided in the embodiment of the present application.

[0037] An embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions or a computer program. When the computer-executable instructions or the computer program are executed by a first processor, a voice processing method applied to a receiving-end device provided by the embodiment of the present application is implemented; or, when the computer-executable instructions or the computer program are executed by a second processor, a voice processing method applied to a sending-end device provided by the embodiment of the present application is implemented.

[0038] An embodiment of the present application provides a computer program product including computer-executable instructions or a computer program. When the computer-executable instructions or the computer program are executed by a first processor, a voice processing method applied to a receiving-end device provided by the embodiment of the present application is implemented; or, when the computer-executable instructions or the computer program are executed by a second processor, a voice processing method applied to a sending-end device provided by the embodiment of the present application is implemented.

[0039] The embodiment of the present application has at least the following beneficial effects: Since the data to be played of the first sound source received includes not only the voice sequence to be processed sent by the sending-end device, but also the sound source probability of each voice to be processed sent by the sending-end device; therefore, when the current voice delay is large (greater than the specified voice delay), the voice to be processed with a sound source probability less than the specified sound source probability can be processed for playback acceleration, thereby being able to improve the voice consumption speed of the first sound source, reduce the voice delay, and the object of the playback acceleration processing is the voice to be processed with a sound source probability less than the specified sound source probability, so the influence on the voice of the first sound source can be reduced, and thus the influence on the intelligibility of the first sound source can be reduced; thereby, the voice playback quality can be improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 is a schematic structural diagram of a voice processing system provided by an embodiment of the present application;

[0041] Figure 2 is provided by an embodiment of the present application Figure 1 a schematic structural diagram of one of the terminals in;

[0042] Figure 3 is provided by an embodiment of the present application Figure 1 a schematic structural diagram of another terminal in;

[0043] Figure 4 is a schematic flowchart of a voice processing method provided by an embodiment of the present application Figure 1 ;

[0044] Figure 5 is a schematic flowchart of a voice processing method provided by an embodiment of the present application Figure 2 ;

[0045] Figure 6 It is a flowchart showing the voice processing method provided by an embodiment of the present application Figure 3 ;

[0046] Figure 7 It is a schematic diagram of an exemplary game voice interaction scenario provided by an embodiment of the present application;

[0047] Figure 8 It is an architectural diagram of an exemplary game voice interaction scenario provided by an embodiment of the present application;

[0048] Figure 9 It is a schematic diagram of an exemplary sending process provided by an embodiment of the present application;

[0049] Figure 10 It is a schematic diagram of an exemplary model training provided by an embodiment of the present application;

[0050] Figure 11 It is a schematic diagram of an exemplary model application provided by an embodiment of the present application;

[0051] Figure 12 It is a schematic diagram of an exemplary packaging result provided by an embodiment of the present application;

[0052] Figure 13 It is a processing flowchart of an exemplary server provided by an embodiment of the present application;

[0053] Figure 14 It is a processing flowchart of an exemplary quality server provided by an embodiment of the present application;

[0054] Figure 15 It is a schematic diagram of an exemplary sending process provided by an embodiment of the present application;

[0055] Figure 16 It is a schematic diagram of an exemplary voice processing flowchart provided by an embodiment of the present application. Detailed implementation manners

[0056] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be construed as limitations on the present application. All other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present application.

[0057] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.

[0058] In the following description, the terms "first / second" involved are used to distinguish similar objects and do not represent a specific order for the objects. Understandably, "first / second" can be interchanged with a specific order or sequence when permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0059] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the function of the module or unit.

[0060] Unless otherwise defined, all technical and scientific terms used in the embodiments of the present application have the same meaning as commonly understood by those skilled in the art to which the present application belongs. The terms used in the embodiments of the present application are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.

[0061] Before further elaborating on the embodiments of the present application, the nouns and terms involved in the embodiments of the present application are described, and the nouns and terms involved in the embodiments of the present application are subject to the following explanations.

[0062] 1) Artificial Intelligence (AI) is the theory, method, technology, and application system that uses a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. That is to say, artificial intelligence is a comprehensive technology of computer science, used to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence enables machines to have the functions of perception, reasoning, and decision-making by studying the design principles and implementation methods of various intelligent machines.

[0063] It should be noted that artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include sensors, special artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operation / interaction systems, mechatronics, etc. Among them, the pre-trained model is also known as the large model and the basic model; after fine-tuning, the pre-trained model can be widely applied to downstream tasks in various directions of artificial intelligence. Artificial intelligence software technologies include computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning and other major directions. In the embodiments of the present application, the determination of the sound source probability can be achieved through artificial intelligence technology.

[0064] 2) Machine Learning (ML) is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It is used to study the computer simulation or realization of human learning behaviors to acquire new knowledge or skills, and reorganize the existing knowledge structure to continuously improve its own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Machine learning applications cover all fields of artificial intelligence. Machine learning / deep learning usually includes technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, and inductive learning. The large model is the latest development result of machine learning / deep learning, which integrates the above technologies. In the embodiments of the present application, the determination of the sound source probability can be achieved in combination with machine learning / deep learning.

[0065] 3) An artificial neural network is a mathematical model that imitates the structure and function of a biological neural network. The exemplary structures of the artificial neural network in the embodiments of the present application include a graph convolutional network (Graph Convolutional Network, GCN, a neural network for processing graph-structured data), a deep neural network (Deep Neural Networks, DNN), a convolutional neural network (Convolutional Neural Network, CNN), a recurrent neural network (Recurrent Neural Network, RNN), a neural state machine (Neural State Machine, NSM), and a phase-functioned neural network (Phase-Functioned Neural Network, PFNN), etc. In the embodiments of the present application, the determination of the sound source probability can be achieved through an artificial neural network model (abbreviated as a neural network model).

[0066] It should be noted that with the research and progress of artificial intelligence technology, artificial intelligence technology has been studied and applied in multiple fields. For example, smart home, smart wearable devices, virtual assistants, smart speakers, smart marketing, driverless, autonomous driving, drones, digital twins, virtual humans, robots, artificial intelligence-generated content (AIGC), conversational interaction, smart healthcare, smart customer service, and virtual scenarios, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role. The voice processing method provided by the embodiments of this application involves technologies such as machine learning / deep learning in artificial intelligence, which will be specifically described through the following embodiments.

[0067] 4) The probability of human voice refers to the probability that a piece of voice is emitted by a human; the probability of human voice is an example of the probability of the target sound source.

[0068] 5) Voice delay is the time interval between the sound emission time of the voice data collected by the sending device and the time when the receiving device plays the sound based on the voice data.

[0069] 6) Voice encoding, abbreviated as encoding, refers to the process of compressing and encoding the original voice data according to a compression encoding algorithm; correspondingly, the reverse process of voice encoding is voice decoding, abbreviated as decoding, which refers to the process of restoring the encoded voice to the original voice data according to the compression encoding algorithm.

[0070] 7) Preprocessing refers to the processing performed before voice encoding to improve the clarity and intelligibility of the voice; among them, intelligibility is used to measure the understandability of the voice, and it can generally be quantified by calculating the number of correctly recognized words or phonemes.

[0071] 8) The multi-source playback scenario refers to an application scenario where an application plays the voices of multiple sound sources. For example, a scenario where local music is played and voice communication is carried out.

[0072] 9) In response to is used to represent the conditions or states on which the executed operations depend. When the dependent conditions or states are met, one or more operations to be executed can be real-time or can have a set delay; without special instructions, there is no restriction on the execution order of the multiple operations to be executed.

[0073] 10) A control is a triggerable piece of information presented in forms such as areas, buttons, icons, links, text, selection boxes, input boxes, and tabs, etc.; among them, the triggering method can be contact triggering, non-contact triggering, or instruction-receiving triggering, etc.; in addition, various controls in the embodiments of this application can be a single control or the general term for multiple controls.

[0074] 11) An operation is a way to trigger a device to perform processing. For example, a click operation, a double - click operation, a long - press operation, a swipe operation, a gesture operation, a received trigger instruction, etc. Additionally, in the embodiments of this application, various operations can be a single operation or a collective term for multiple operations; and various operations in the embodiments of this application can be touch operations or non - touch operations.

[0075] 12) A client is an application program running on a device to provide various services. For example, a game client, an instant messaging client, and so on.

[0076] It should be noted that in a scenario where voice is transmitted over a network and played, when the current voice delay is greater than the specified voice delay, in order to reduce the current voice delay, usually some voice frames are periodically discarded or the received voice is played at an accelerated speed; however, periodically discarding some voice frames loses the original voice. For example, the voice of the target sound source may be discarded, affecting the intelligibility of the voice, and playing the received voice at an accelerated speed also affects the intelligibility of the voice. Therefore, both periodically discarding some voice frames and playing the received voice at an accelerated speed affect the voice - playing effect. Also, in a multi - source playback scenario, when the voice energy of other sound sources outside the current sound source is greater than the specified energy, the impact on voice intelligibility is increased.

[0077] In addition, in order to reduce the voice delay, the receiving - end device can also perform voice decoding on the received voice data and discard the voice with a sound - source probability lower than the specified probability from the decoding result; thus, it increases the voice - decoding consumption of the discarded voice, affects the degree of reduction of the voice delay, and cannot effectively reduce the voice delay.

[0078] Based on this, the embodiments of this application provide a voice - processing method, device, equipment, computer - readable storage medium, and computer program product, which can effectively reduce the voice delay, improve the voice intelligibility and voice - playing quality. The following describes the exemplary applications of the receiving - end device and the sending - end device for voice processing provided by the embodiments of this application. The receiving - end device and the sending - end device provided by the embodiments of this application can be implemented as various types of terminals such as smartphones, smart watches, laptop computers, tablet computers, desktop computers, smart home appliances, set - top boxes, intelligent in - vehicle devices, portable music players, personal digital assistants, dedicated messaging devices, intelligent voice - interaction devices, portable game devices, and smart speakers, or can be implemented as servers, or can be a combination of both. Next, the exemplary application when both the receiving - end device and the sending - end device are implemented as terminals will be described.

[0079] See Figure 1 , Figure 1 is the schematic diagram of the architecture of the voice - processing system provided by the embodiments of this application; as Figure 1As shown, to support a voice processing application, in a voice processing system 100, a terminal 400 and a terminal 200 are connected to a server 600 through a network 300. The network 300 can be a wide area network, a local area network, or a combination of both. In addition, the voice processing system 100 further includes a database 500 for providing data support to the server 600; and, Figure 1 The figure shows a case where the database 500 is independent of the server 600. In addition, the database 500 can also be integrated in the server 600, which is not limited in the embodiments of the present application.

[0080] The terminal 400 is configured to receive, through the network 300, a to-be-processed voice sequence sent by the terminal 200 for a first sound source and a sound source probability corresponding to each to-be-processed voice in the to-be-processed voice sequence, where the sound source probability refers to the probability that the to-be-processed voice is a sound emitted by the first sound source; when the current voice delay is greater than a specified voice delay, in the to-be-processed voice sequence, perform a playback acceleration process on the to-be-processed voice whose sound source probability is less than the specified probability to obtain a to-be-played voice sequence; and perform a first voice playback based on the to-be-played voice sequence for the first sound source (graphical interface 410-1 is exemplarily shown).

[0081] The terminal 200 is configured to, in response to a voice collection instruction, collect voice for the first sound source (graphical interface 210-1 is exemplarily shown) to obtain a to-be-processed voice sequence; determine a sound source probability based on a to-be-processed feature corresponding to each to-be-processed voice in the to-be-processed voice sequence; and send the to-be-processed voice sequence and the sound source probability of each to-be-processed voice to the terminal 400 through the network 300.

[0082] In some embodiments, the server 600 can be an independent physical server, or a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms. The terminal and the server can be directly or indirectly connected through wired or wireless communication methods, which are not limited in the embodiments of the present application.

[0083] See Figure 2 , Figure 2 which is a schematic structural diagram of a terminal provided by the embodiments of the present application; as Figure 1 in Figure 2As shown, the terminal 400 includes: at least one first processor 410, a first memory 450, at least one first network interface 420, and a first user interface 430. Each component in the terminal 400 is coupled together via a first bus system 440. It can be understood that the first bus system 440 is used to enable connection and communication between these components. In addition to a data bus, the first bus system 440 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clear illustration, in Figure 2 all kinds of buses are labeled as the first bus system 440.

[0084] The first processor 410 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or any conventional processor, etc.

[0085] The first user interface 430 includes one or more first output devices 431 that enable the presentation of media content, including one or more speakers and / or one or more visual display screens. The first user interface 430 also includes one or more first input devices 432, including user interface components that assist user input, such as a keyboard, a mouse, a microphone, a touch screen display, a camera, other input buttons, and controls.

[0086] The first memory 450 can be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard disk drives, optical disc drives, etc. The first memory 450 optionally includes one or more storage devices that are physically located away from the first processor 410.

[0087] The first memory 450 includes volatile memory or non-volatile memory, and can also include both volatile and non-volatile memory. The non-volatile memory can be a read-only memory (ROM), and the volatile memory can be a random access memory (RAM). The first memory 450 described in the embodiments of the present application is intended to include any suitable type of memory.

[0088] In some embodiments, the first memory 450 is capable of storing data to support various operations. Examples of such data include programs, modules, and data structures, or subsets or supersets thereof, which are illustrated below.

[0089] The first operating system 451 includes system programs for processing various basic system services and performing hardware-related tasks, such as the framework layer, the core library layer, the driver layer, etc., for implementing various basic services and processing hardware-based tasks;

[0090] The first network communication module 452 is used to reach other electronic devices via one or more (wired or wireless) first network interfaces 420. Exemplary first network interfaces 420 include: Bluetooth, Wi-Fi (Wireless Fidelity), Universal Serial Bus (USB), etc.;

[0091] The first presentation module 453 is used to enable the presentation of information (such as a user interface for operating peripheral devices and displaying content and information) via one or more first output devices 431 associated with the first user interface 430 (such as a display screen, a speaker, etc.);

[0092] The first input processing module 454 is used to detect one or more user inputs or interactions from one of the one or more first input devices 432 and translate the detected inputs or interactions.

[0093] In some embodiments, the first voice processing device provided by the embodiments of the present application can be implemented in software. Figure 2 Shown is the first voice processing device 455 stored in the first memory 450, which can be software in the form of programs and plugins, etc., including the following software modules: a voice receiving module 4551, a voice processing module 4552, and a voice playing module 4553. These modules are logical, so they can be combined arbitrarily or further split according to the functions to be implemented. The functions of each module will be described below.

[0094] See Figure 3 , Figure 3 is the structure schematic diagram of another terminal provided by the embodiments of the present application; as Figure 1 shown in Figure 3 , the terminal 200 includes: at least one second processor 210, a second memory 250, at least one second network interface 220, and a second user interface 230. Each component in the terminal 200 is coupled together through a second bus system 240. It can be understood that the second bus system 240 is used to realize the connection and communication between these components. The second bus system 240 includes, in addition to a data bus, a power bus, a control bus, and a status signal bus. However, for the sake of clear illustration, in Figure 3 all kinds of buses are labeled as the second bus system 240.

[0095] The second processor 210 may be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or any conventional processor, etc.

[0096] The second user interface 230 includes one or more second output devices 231 that enable the presentation of media content, including one or more speakers and / or one or more visual display screens. The second user interface 230 further includes one or more second input devices 232, including user interface components that facilitate user input, such as a keyboard, a mouse, a microphone, a touch screen display, a camera, other input buttons, and controls.

[0097] The second memory 250 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard disk drives, optical disc drives, etc. The second memory 250 optionally includes one or more storage devices that are physically located away from the second processor 210.

[0098] The second memory 250 includes volatile memory or non-volatile memory, and may also include both volatile and non-volatile memory. The non-volatile memory may be read-only memory, and the volatile memory may be random access memory. The second memory 250 described in the embodiments of the present application is intended to include any suitable type of memory.

[0099] In some embodiments, the second memory 250 is capable of storing data to support various operations. Examples of such data include programs, modules, and data structures, or subsets or supersets thereof, which are illustratively described below.

[0100] The second operating system 251, including system programs for processing various basic system services and performing hardware-related tasks, such as a framework layer, a core library layer, a driver layer, etc., for implementing various basic services and processing hardware-based tasks;

[0101] The second network communication module 252 is used to reach other electronic devices via one or more (wired or wireless) second network interfaces 220. Exemplary second network interfaces 220 include: Bluetooth, Wi-Fi, and Universal Serial Bus, etc.;

[0102] The second presentation module 253 is used to enable the presentation of information (such as a user interface for operating peripheral devices and displaying content and information) via one or more second output devices 231 associated with the second user interface 230 (e.g., a display screen, a speaker, etc.);

[0103] A second input processing module 254 is configured to detect one or more user inputs or interactions from one of one or more second input devices 232 and translate the detected inputs or interactions.

[0104] In some embodiments, the second voice processing device provided in the embodiments of the present application may be implemented in software. Figure 3 Shown is a second voice processing device 255 stored in a second memory 250, which may be software in the form of programs and plugins, etc., including the following software modules: a voice acquisition module 2551, a probability determination module 2552, a voice sending module 2553, and a model training module 2554. These modules are logical, and thus can be combined arbitrarily or further split according to the functions implemented. The functions of each module will be described below.

[0105] In some embodiments, the first voice processing device and the second voice processing device provided in the embodiments of the present application may be implemented in hardware. As an example, the first voice processing device and the second voice processing device provided in the embodiments of the present application may be processors in the form of hardware decoding processors, which are programmed to execute the voice processing method provided in the embodiments of the present application. For example, a processor in the form of a hardware decoding processor may employ one or more application specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components.

[0106] In some embodiments, a terminal or a server may implement the voice processing method provided in the embodiments of the present application by running various computer-executable instructions or computer programs. For example, the computer-executable instructions may be commands at the microprogram level, machine instructions, or software instructions. The computer program may be a native program or software module in an operating system; it may be a native application (APPlication, APP), that is, a program that needs to be installed in the operating system to run, such as a game APP or an instant messaging APP; it may also be a small program that can be embedded in any APP, that is, a program that only needs to be downloaded into a browser environment to run. In short, the above computer-executable instructions may be instructions in any form, and the above computer programs may be application programs, modules, or plugins in any form.

[0107] Next, in combination with the exemplary applications and implementations of the voice processing device provided in the embodiments of the present application, the voice processing method provided in the embodiments of the present application will be described. In addition, the voice processing method provided in the embodiments of the present application is applied to various voice playback scenarios such as cloud technology, artificial intelligence, intelligent transportation, maps, in-vehicle, and games.

[0108] See Figure 4 , Figure 4 which is a flowchart of the voice processing method provided in the embodiments of the present application Figure 1 ; Next, it will be described in combination with Figure 4 the steps shown.

[0109] Step 101: The sending device responds to the voice collection instruction, collects the voice of the first sound source, and obtains a voice sequence to be processed.

[0110] In the embodiments of the present application, when an operation for collecting the voice of the first sound source is triggered, the sending device receives the voice collection instruction; since the voice collection instruction is used to indicate the collection of the voice of the first sound source, at this time, the sending device responds to the voice collection instruction, executes the processing indicated by the voice collection instruction, collects the voice of the first sound source, collects the voice signal of the first sound source, and combines the collected voice signals of the first sound source into a voice sequence to be processed. Here, the sending device can encode one voice frame and multiple consecutive voice frames in the collected voice signal into a voice to be processed, and combine at least one voice to be processed obtained based on the voice frame order into a voice sequence to be processed.

[0111] It should be noted that the voice collection instruction can be generated when an operation for collecting the voice of the first sound source is received, for example, a voice communication link establishment operation, a video communication link establishment operation, a microphone turn-on operation. The first sound source is the sound source of the voice signal to be collected on the sending device side, for example, the call object in a voice call. The voice sequence to be processed can be the voice signal of the first sound source collected within one collection period, which refers to a kind of real-time information; of course, the voice sequence to be processed can also be obtained based on the historical voice signal of the first sound source; the embodiments of the present application do not limit this.

[0112] In step 101 of the embodiments of the present application, the sending device responds to the voice collection instruction, collects the voice of the first sound source, and obtains a voice sequence to be processed, including: the sending device first responds to the voice collection instruction, collects the voice of the first sound source, and obtains an initial voice sequence; then denoises the initial voice sequence to obtain a voice sequence to be encoded; finally, performs voice encoding on the voice sequence to be encoded to obtain a voice sequence to be processed.

[0113] It should be noted that the initial voice sequence is the result of voice acquisition of the first sound source. The sending device can directly encode the initial voice sequence into the voice sequence to be processed, or can perform preprocessing such as denoising on the initial voice sequence, and encode the obtained result into the voice sequence to be encoded. Here, the sending device takes a voice frame and multiple consecutive voice frames in the collected voice signal as an initial voice, and combines at least one initial voice obtained based on the voice frame order into an initial voice sequence.

[0114] Step 102: The sending device determines the sound source probability based on the to-be-processed features corresponding to each to-be-processed voice in the to-be-processed voice sequence.

[0115] In the embodiment of the present application, the sending device first extracts features from each to-be-processed voice in the to-be-processed voice sequence, and uses the extracted features as the to-be-processed features; then determines the probability that the to-be-processed voice belongs to the first sound source based on the to-be-processed features, and thus obtains the sound source probability.

[0116] It should be noted that the sending device can extract features from the to-be-processed voice based on a specified coding rule (such as one-hot code, etc.), and can also use a neural network model to extract features from the to-be-processed voice, etc. The embodiment of the present application does not limit this. In addition, the sound source probability refers to the probability that the to-be-processed voice is the sound emitted by the first sound source, that is to say, the sound source probability refers to the probability that the to-be-processed voice belongs to the first sound source.

[0117] It should also be noted that the to-be-processed features refer to the voice features of the to-be-processed voice, which are used to represent the sound features of the first sound source included in the to-be-processed voice, and can be at least one of the following: semantic features, energy features, spectral features, power spectral features, cepstral features, spectral envelope features, and timbre features, etc. Here, when the sending device determines the sound source probability based on the to-be-processed features, it can be implemented based on a trained neural network model, can also be implemented based on a specified mapping function between features and sound source probability, and can also be implemented through a large prediction model, etc. The embodiment of the present application does not limit this.

[0118] Step 103: The receiving device receives the to-be-processed voice sequence sent by the sending device for the first sound source and the sound source probability corresponding to each to-be-processed voice in the to-be-processed voice sequence.

[0119] In the embodiment of the present application, when the sending device has determined the sound source probability for each to-be-processed voice in the to-be-processed voice sequence, the sending device sends the to-be-processed voice sequence and the sound source probability of each to-be-processed voice to the receiving device; at this time, the receiving device has also received the to-be-processed voice sequence sent by the sending device for the first sound source and the sound source probability corresponding to each to-be-processed voice in the to-be-processed voice sequence.

[0120] It should be noted that an interactive connection including voice communication is established at least between the receiving device and the transmitting device, and the receiving device and the transmitting device can interchange device roles.

[0121] See Figure 5 , Figure 5 is a flowchart of the voice processing method provided by an embodiment of the present application Figure 2 ; as Figure 5 shown, the voice processing method provided by an embodiment of the present application can be applied to a single-source playback scenario and can also be applied to a multi-source playback scenario. The embodiments of the present application do not limit this; however, when applied to a multi-source playback scenario, the receiving device in step 103 receives the to-be-processed voice sequence sent by the transmitting device for the first sound source and the sound source probability corresponding to each to-be-processed voice in the to-be-processed voice sequence, including step 1031. The following explains each step.

[0122] Step 1031: During the process of playing the second voice for the second sound source, the receiving device receives the to-be-processed voice sequence sent by the transmitting device for the first sound source and the sound source probability corresponding to each to-be-processed voice in the to-be-processed voice sequence.

[0123] It should be noted that before receiving the to-be-processed voice sequence, the receiving device plays the second voice for the second sound source. Among them, the second sound source is different from the first sound source. For example, in a game scenario, the first sound source can be a player, and the second sound source can be a sound source corresponding to background music or sound effects. And the first sound source and the second sound source are multiple sound sources in a multi-source playback scenario.

[0124] Step 104: When the current voice delay is greater than the specified voice delay, the receiving device performs a playback acceleration process on the to-be-processed voices in the to-be-processed voice sequence with a sound source probability less than the specified probability to obtain a to-be-played voice sequence.

[0125] In the embodiment of the present application, after the receiving device receives the to-be-processed voice sequence and the sound source probability corresponding to each to-be-processed voice, it calculates the current voice delay and compares the current voice delay with the specified voice delay; when it is determined through the comparison that the current voice delay is greater than the specified voice delay, it indicates that the current voice delay is large; thus, to reduce the current voice delay, the receiving device performs a playback acceleration process on the to-be-processed voice sequence based on the sound source probability; that is, the receiving device performs a playback acceleration process on the to-be-processed voices in the to-be-processed voice sequence with a sound source probability less than the specified probability, and the processing result is called the to-be-played voice sequence. Here, the receiving device can obtain the current voice delay by calculating the voice duration in the current to-be-played voice queue.

[0126] It should be noted that the playback acceleration process is either discarding or increasing the playback speed. Here, discarding means discarding the to-be-processed speech with a sound source probability less than the specified probability, and increasing the playback speed means increasing the original playback speed of the to-be-processed speech with a sound source probability less than the specified probability at the target acceleration.

[0127] Continue to refer to Figure 5 , in the embodiment of the present application, step 104 can be implemented through step 1041 and step 1042; that is, when the current voice delay is greater than the specified voice delay, the receiving-end device performs a playback acceleration process on the to-be-processed speech with a sound source probability less than the specified probability in the to-be-processed speech sequence to obtain the to-be-played speech sequence, including step 1041 and step 1042. Each step will be described separately below.

[0128] Step 1041: When the current voice delay is greater than the specified voice delay, the receiving-end device obtains the current voice energy of the second voice.

[0129] In the embodiment of the present application, when the current voice delay is greater than the specified voice delay, if it is a multi-source playback scenario, the receiving-end device also adjusts the current voice delay in combination with the current voice energy of the second voice of the second sound source. Here, the receiving-end device can obtain the current voice energy by obtaining the sum of the frame voice energies of a continuous multiple frames (for example, 5 frames, 4 frames, 6 frames, etc.) to be played at the current moment; among them, the frame voice energy refers to the energy corresponding to one frame of the second voice.

[0130] Step 1042: In the to-be-processed speech sequence, the receiving-end device performs a playback acceleration process on the to-be-processed speech with a sound source probability less than the specified probability based on the current voice energy to obtain the to-be-played speech sequence.

[0131] In the embodiment of the present application, since the to-be-processed speech sequence is received during the playback of the second voice, thus, the receiving-end device can perform a playback acceleration process in combination with the current voice energy of the second voice to improve the voice playback quality in the multi-source playback scenario.

[0132] In the embodiment of the present application, the receiving-end device performs a playback acceleration process on the to-be-processed speech with a sound source probability less than the specified probability based on the current voice energy in the to-be-processed speech sequence to obtain the to-be-played speech sequence, including: when the current voice energy is less than the first specified energy, the receiving-end device increases the playback speed of the to-be-processed speech with a sound source probability less than the specified probability in the to-be-processed speech sequence to obtain the to-be-played speech sequence. At this time, the playback acceleration process is increasing the playback speed. And when the current voice energy is greater than the second specified energy, the receiving-end device discards the to-be-processed speech with a sound source probability less than the specified probability from the to-be-processed speech sequence to obtain the to-be-played speech sequence. At this time, the playback acceleration process is discarding. Among them, the second specified energy is less than or equal to the first specified energy.

[0133] It should be noted that in a multi-source playback scenario, to improve the voice playback quality, when the current voice energy is less than the first specified energy, it indicates that the voice playback sound of the second sound source is relatively small and has a relatively small impact on the voice playback of the first sound source. Therefore, to ensure the voice playback effect of the first sound source, the playback speed of the voice to be processed with a sound source probability less than the specified probability is increased, so that the voice of the first sound source can be effectively played. When the current voice energy is greater than the second specified energy, it indicates that the voice playback sound of the second sound source is relatively large and has a relatively large impact on the voice playback of the first sound source. Therefore, even if the voice to be processed with a sound source probability less than the specified probability is discarded, the voice playback effect of the first sound source can still be ensured. Additionally, when the second specified energy is less than the first specified energy, if the current voice energy is between the second specified energy and the first specified energy, the receiving device can perform any playback acceleration processing on the voice to be processed with a sound source probability less than the specified probability, and the embodiments of the present application do not limit this.

[0134] In the embodiments of the present application, when the playback acceleration processing is discarding, the receiving device discards the voice to be processed with a sound source probability less than the specified probability from the voice sequence to be processed to obtain the voice sequence to be played, including: the receiving device first obtains the previous discard time; then when the duration between the previous discard time and the current time is greater than or equal to the specified cycle duration, it discards the voice to be processed with a sound source probability less than the specified probability from the voice sequence to be processed to obtain the voice sequence to be played; when the duration between the previous discard time and the current time is less than the specified cycle duration, it can perform the playback acceleration processing of increasing the playback speed.

[0135] It should be noted that a specified cycle duration is pre-set in the receiving device, and this specified cycle duration is used to determine the discard cycle of the voice to be processed; the previous discard time is the time when the voice discard processing was last performed.

[0136] It can be understood that when it is determined to discard the voice to be processed with a sound source probability less than the specified probability, periodically discarding the voice to be processed with a sound source probability less than the specified probability can reduce the impact on the voice playback effect of the first sound source, and thus can improve the voice playback quality.

[0137] In the embodiments of the present application, when the receiving device performs play acceleration processing on the to-be-played voice with a sound source probability less than a specified probability in the to-be-processed voice sequence based on the current voice energy, it may be in the to-be-processed voice sequence to increase the play speed of the to-be-played voice with a sound source probability less than the specified probability and increase the play speed of the to-be-played voice with a sound source probability greater than or equal to the specified probability; or, in the to-be-processed voice sequence, increase the play speed of the to-be-played voice with a sound source probability less than the specified probability and maintain the play speed of the to-be-played voice with a sound source probability greater than or equal to the specified probability; the embodiments of the present application do not limit this.

[0138] Step 105: The receiving device performs first voice playback for the first sound source based on the to-be-played voice sequence.

[0139] It should be noted that when the current voice delay is relatively large, the receiving device realizes the playback of the first voice corresponding to the first sound source by playing the to-be-played voice sequence. Among them, each to-be-played voice in the to-be-played voice sequence may be a to-be-processed voice with a sound source probability greater than or equal to the specified probability, or a to-be-processed voice whose original play speed has been accelerated; when the to-be-played voice is a to-be-processed voice whose original play speed has been accelerated, the receiving device plays the to-be-played voice at the accelerated play speed when playing the to-be-played voice.

[0140] It can be understood that the receiving device is used to realize the first voice playback by performing play acceleration processing on the to-be-processed voice with a sound source probability less than the specified probability when the current voice delay is greater than the specified voice delay.

[0141] Continue to refer to Figure 5 , in the embodiments of the present application, step 105 can be implemented through step 1051 and step 1052; that is, the receiving device performs first voice playback for the first sound source based on the to-be-played voice sequence, including step 1051 and step 1052, and the following will explain each step separately.

[0142] Step 1051: The receiving device performs voice decoding on the to-be-played voice sequence to obtain a first voice sequence.

[0143] In the embodiments of the present application, the to-be-processed voice sequence is the data directly received from the sending device, which is the voice coding result, that is, the data before voice decoding. That is to say, the receiving device performs play acceleration processing on the to-be-processed voice with a sound source probability less than the specified probability before voice decoding. Among them, the first voice sequence is the voice decoding result of the to-be-played voice sequence.

[0144] Step 1052: The receiving device performs first voice playback for the first sound source based on the first voice sequence.

[0145] It should be noted that the first voice sequence is the voice decoding result, and the first voice playback for the first sound source can be realized by playing the first voice sequence.

[0146] It can be understood that by adjusting the current voice delay in combination with the sound source probability before voice decoding, when the adjustment method is to discard, the amount of voice decoding data can be reduced, and thus the voice decoding efficiency can be improved.

[0147] In the embodiments of the present application, the acquisition of the sound source probability of each voice to be processed can be realized through a sound source classification model; wherein, the sound source classification model is used to determine the probability that the voice is the sound emitted by the first sound source, and can be obtained through the following steps of training: the sending device first obtains the initial training data of the model to be trained; then performs scene classification on the initial training data to obtain multiple scene training data; then, uses the model to be trained to perform sound source classification on the multiple scene training data to obtain multiple sound source classification results; finally, adjusts the model parameters of the model to be trained based on the multiple sound source classification results to obtain the sound source classification model.

[0148] It should be noted that the model to be trained is a neural network model to be trained for determining the probability that the voice is the sound emitted by the first sound source. Here, the sending device performs scene classification on the initial training data. For example, the initial training data is divided into training data for indoor scenes, training data for outdoor scenes, etc. Among them, the sound source classification result refers to the classification result of whether each scene training data is the first sound source. The sending device uses the classification distribution corresponding to the sound source classification result to adjust the model parameters of the model to be trained so that the adjusted model parameters are adapted to the classification distribution. The adjustment of the model parameters is carried out iteratively. When the adjustment amplitude of the model parameters is less than the specified amplitude, it is determined that the adjustment of the model parameters ends, and then the adjusted model parameters are determined as the final model parameters.

[0149] In the embodiments of the present application, the receiving device can reduce the current voice delay through playback acceleration processing when the current voice delay is greater than the specified voice delay, and stop the adjustment processing until the voice delay is reduced to the lowest adjustment threshold, and the lowest adjustment threshold is less than the specified delay threshold.

[0150] See Figure 6 , Figure 6 is the flow schematic of the voice processing method provided by the embodiments of the present application Figure 3 ; as Figure 6 shown, in the embodiments of the present application, after the receiving device receives the voice sequence to be processed sent by the sending device for the first sound source and the sound source probability corresponding to each voice to be processed in the voice sequence to be processed, the voice processing method further includes step 106 and step 107, and the following is an explanation of each step.

[0151] Step 106: When the current voice delay is less than or equal to the specified voice delay, the receiving device decodes the voice sequence to be processed to obtain a second voice sequence.

[0152] In the embodiment of the present application, when the current voice delay is less than or equal to the specified voice delay, it indicates that the current voice delay is small. Thus, the receiving device can directly decode the voice sequence to be processed for the first voice playback. Here, the second voice sequence is the voice decoding result of the voice sequence to be processed.

[0153] Step 107: The receiving device performs the first voice playback based on the second voice sequence for the first sound source.

[0154] It should be noted that when the current voice delay is small, the receiving device plays the first voice corresponding to the first sound source by playing the second voice sequence.

[0155] Next, the exemplary application of the embodiment of the present application in an actual application scenario will be described. This exemplary application describes the process of performing voice playback in a game voice interaction scenario. It is easy to know that the voice processing method provided by the embodiment of the present application is applicable to voice playback in any voice transmission scenario. Here, the voice playback in the game voice interaction scenario is taken as an example for description.

[0156] It should be noted that when the user turns on the microphone for voice interaction in the game, a microphone turn-on request is received; at this time, the sending end (for example, Figure 1 the terminal 200 in it, called the sending device) responds to the microphone turn-on request, collects voice through the microphone; then preprocesses the collected voice signal (called the initial voice sequence), calculates the voice probability (called the sound source probability) for each frame of the preprocessing result (called the voice sequence to be encoded), and performs voice encoding on the preprocessing result; finally, packs and transmits the voice encoding result (called the voice sequence to be processed) and the voice probability of each frame to the server (for example, Figure 1 the server 600 in it), and the server forwards the packed result to the receiving end (for example, Figure 1 the terminal 400 in it, called the receiving device). When the receiving end plays the voice based on the received packed result, since the game background music (or sound effects) is also played locally at the receiving end, thus, this is a multi-source playback scenario at this time; when the current voice delay is high (greater than the specified voice delay), and the energy of the game background music (or sound effects) is large (greater than the second specified energy), the receiving end discards the voice frames with low voice probability (called the voice to be processed) to reduce the delay. And when the current voice delay is high, and the energy of the game background music (or sound effects) is small (less than the first specified energy), the receiving end accelerates the playback of each received voice frame to accelerate the consumption of voice data and reduce the voice delay.

[0157] Exemplarily, refer to Figure 7 , Figure 7 which is a schematic diagram of an exemplary game voice interaction scenario provided by an embodiment of the present application; as Figure 7 shown, in the interface 7-1 corresponding to the game scenario, when the microphone 7-11 is triggered, a microphone open request is received; the receiving end performs voice playback through the playback control 7-12.

[0158] Refer to Figure 8 , Figure 8 which is an architecture diagram of an exemplary game voice interaction scenario provided by an embodiment of the present application; as Figure 8 shown, the sending end 8-1 (exemplarily showing the terminal 8-11 and the terminal 8-12) sends a packing result to the receiving end 8-3 (exemplarily showing the terminal 8-31 and the terminal 8-32) through the server 8-2, and the receiving end 8-3 is used to receive the packing result and perform voice playback based on the packing result. Among them, both the sending end 8-1 and the receiving end 8-3 are terminals. The sending end 8-1 is used for voice acquisition, preprocessing, encoding processing and sending, and the receiving end 8-3 is used for receiving, decoding and playing; the server 8-2 includes a forwarding server and a quality server. The forwarding server is used to receive data sent by the terminal and forward it to the indicated forwarding server or terminal. The quality server is used to confirm the forward error correction algorithm with the sending end 8-1 according to the received network quality report. Among them, the forward error correction algorithm is a kind of error control method, so that the receiving end can determine the error code generated during the transmission process and correct the error code.

[0159] The processing process of the sending end will be described first below.

[0160] Refer to Figure 9 , Figure 9 which is a schematic diagram of an exemplary sending process provided by an embodiment of the present application; as Figure 9 shown, this exemplary sending process schematic diagram can be executed by the terminal 200 in Figure 1 , including steps 901 to 907, and each step will be described separately below.

[0161] Step 901: Initialize game resources to run the game.

[0162] Step 902: Collect voice signals during the running of the game.

[0163] It should be noted that the data to be sent at least includes voice signals, and may also include other data to be sent to the receiving end through the network, such as video signals, game data, etc.

[0164] Step 903: Preprocess the voice signals.

[0165] It should be noted that pre - processing refers to the processing carried out before voice transmission, such as noise reduction processing for noise elimination, etc.

[0166] Step 904: Calculate the vocal probability of each frame for the pre - processing result.

[0167] It should be noted that the calculation of the vocal probability can be achieved through a trained Gaussian Mixture Model (GMM).

[0168] It can be understood that since the Gaussian Mixture Model uses the Gaussian probability density function (normal distribution curve) to precisely quantify things, a thing can be decomposed into a model formed based on the Gaussian probability density function (normal distribution curve); thus, based on the Gaussian mixture model, the human voice can be represented as a normal distribution. When calculating the vocal probability using the Gaussian mixture model, the corresponding vocal probability can be determined based on the normal distribution of the voice signal, improving the accuracy of vocal probability calculation.

[0169] Regarding the training process of the GMM model, by way of example, refer to Figure 10 , Figure 10 which is an exemplary model training schematic diagram provided by an embodiment of the present application; as Figure 10 shown, perform clustering analysis 10 - 21 on training data 10 - 11 (referred to as initial training data) to obtain a clustering analysis result 10 - 12 (referred to as multiple scenario training data); use the Expectation - Maximum (EM) algorithm 10 - 22 to estimate the model parameters of the Gaussian mixture model (referred to as the model to be trained) 10 - 3 for the clustering analysis result 10 - 12; extract features from the clustering analysis result 10 - 12 using the model parameters to obtain features 10 - 13; calculate the vocal probability 10 - 14 using the likelihood function 10 - 4. Then, adjust the model parameters of the Gaussian mixture model 10 - 3 based on the vocal probability 10 - 14, and then use the adjusted Gaussian mixture model 10 - 3 to sequentially perform feature extraction and vocal probability calculation on the clustering analysis result 10 - 12 until the model training is completed to obtain a trained Gaussian mixture model (referred to as a sound source classification model).

[0170] It should be noted that the training data 10-11 can be a voice signal including human voices, or a voice signal including other sounds except human voices; in addition, the training data 10-11 can be voice signals in various voice scenarios (such as outdoor scenarios like playgrounds and squares, and indoor scenarios like offices). Cluster analysis is a method of simplifying data through data modeling; cluster analysis methods include hierarchical clustering method, decomposition method, addition method, dynamic clustering method, ordered sample clustering, overlapping clustering, and fuzzy clustering, etc. The embodiment of the present application can adopt the dynamic clustering method; thus, by performing cluster analysis on the training data 10-11, the voice scenarios of the training data can be classified, and sub-training data corresponding to different voice scenarios can be obtained. Therefore, the cluster analysis result 10-12 includes each sub-training data corresponding to various voice scenarios.

[0171] It can be understood that since the energy of voice signals is different in different voice scenarios; by performing cluster analysis on the training data, the training data is divided into sub-training data in different voice scenarios, and then subsequent processing is performed on the sub-training data, which can improve the accuracy of voice signal processing.

[0172] Here, when using the expectation maximization algorithm 10-22 to estimate the model parameters of the Gaussian mixture model 10-3 for the cluster analysis result 10-12, for each sub-training data, first initialize the normal distribution parameters of human voices, and then calculate whether each voice signal in the sub-training data belongs to human voices or non-human voices, so as to re-determine the normal distribution parameters of human voices based on the full amount of voice signals belonging to human voices in the sub-training data; and so on until the difference between the re-determined normal distribution parameters of human voices and the previously determined normal distribution parameters of human voices is less than the specified difference, the re-determined normal distribution parameters of human voices this time are determined as the model parameters of the Gaussian mixture model. Therefore, the obtained model parameters can be the parameters of the Gaussian mixture model in different voice scenarios.

[0173] It should be noted that the model parameters are used to extract the features of human voice or non-human voice. The likelihood function is a function used to statistically model parameters. When the feature x is given, the likelihood function L(θ|x) of the model parameter θ (numerically) is equal to the probability P(X=x|θ) of the variable X after the model parameter θ is given, as shown in Equation (1).

[0174] L(θ|x) = P(X=x|θ) (1);

[0175] When obtaining the probability of human voices through the trained Gaussian mixture model, by way of example, refer to Figure 11 , Figure 11 which is an exemplary model application schematic diagram provided by the embodiment of the present application; as Figure 11As shown, the Gaussian mixture model 11-1 is used to extract features from the speech to be detected 11-2 (referred to as the speech to be processed), obtaining features 11-3 (referred to as the features to be processed); the likelihood function 11-4 is used to calculate the features 11-3, obtaining the human voice probability 11-5.

[0176] Step 905: Perform speech encoding on the preprocessing result.

[0177] It should be noted that speech encoding is performed on the preprocessing result based on a compression algorithm to reduce the amount of data transmitted over the network.

[0178] Step 906: Package the speech encoding result and the human voice probability of each frame.

[0179] It should be noted that the speech encoding result and the human voice probability of each frame are packaged based on service rules, and the packaging is performed in combination with the forward error correction algorithm returned by the quality server. The packaged data includes the speech encoding result and the human voice probability of each frame in this segment of speech.

[0180] Exemplarily, refer to Figure 12 , Figure 12 which is a schematic diagram of an exemplary packaging result provided by an embodiment of the present application; as Figure 12 shown, the packaging result 12-1 includes the speech encoding result 12-11 and the human voice probability 12-12 of each frame corresponding to the speech encoding result 12-11.

[0181] Step 907: Send the packaging result to the receiving end. Then execute step 901.

[0182] It should be noted that when sending the packaging result to the receiving end, the packaging result can be sent to the receiving end through a forwarding server.

[0183] Next, the processing flow of the server will be described.

[0184] Refer to Figure 13 , Figure 13 which is a schematic flowchart of the processing of an exemplary server provided by an embodiment of the present application; as Figure 13 shown, for the server, the forwarding server 13-1, the forwarding server 13-2, and the quality server 13-3 are exemplarily shown; among them, the forwarding server is used to receive data from the sending end 13-4, and is also used to receive data from other forwarding servers, and is also used to send the received data to the receiving end 13-5. The quality server 13-3 is used to receive the network quality reports (including at least one parameter such as network packet loss rate, network jitter, network type, etc.) of the sending end 13-4 and the receiving end 13-5, and adjust the data sending related algorithms (such as the forward error correction algorithm) of the sending end according to the network quality reports.

[0185] See Figure 14 , Figure 14 is a processing flowchart of an exemplary quality server provided by an embodiment of the present application; as Figure 14 shown, the processing flow of the exemplary quality server includes steps 1401 to 1405, and each step will be described separately below.

[0186] Step 1401: Determine whether a network quality report sent by a sending end is received. If so, execute Step 1402; if not, execute Step 1403.

[0187] Step 1402: Select a forward error correction algorithm for the sending end based on the network quality report sent by the sending end. Then execute Step 1405.

[0188] It should be noted that the quality server receives and parses the network sending report sent by the sending end (such as quality data such as the number of sent packets, the number of arrived packets, network type, calculated packet loss rate, network jitter, etc.), selects a forward error correction algorithm for the sending end, and returns the selected forward error correction algorithm to the sending end.

[0189] Step 1403: Receive the network quality report sent by the receiving end.

[0190] Step 1404: Select a forward error correction algorithm for the sending end based on the network quality report sent by the receiving end. Then execute Step 1405.

[0191] It should be noted that the quality server receives and parses the network sending report sent by the receiving end (such as quality data such as the number of sent packets, the number of arrived packets, network type, calculated packet loss rate, network jitter, etc.), selects a forward error correction algorithm for the sending end, and returns the selected forward error correction algorithm to the sending end.

[0192] Step 1405: Return the forward error correction algorithm to the sending end. Then execute Step 1401.

[0193] Next, the processing flow of the receiving end will be described.

[0194] See Figure 15 , Figure 15 is a schematic diagram of an exemplary sending process provided by an embodiment of the present application; as Figure 15 shown, the schematic diagram of the exemplary sending process is executed by the terminal 400 in Figure 1 and includes steps 1501 to 1507, and each step will be described separately below.

[0195] Step 1501: Receive the packaging result sent by the forwarding server.

[0196] It should be noted that when receiving the data sent by the forwarding server on the network, the packaging result is also received.

[0197] Step 1502: Analyze the received packaging result.

[0198] It should be noted that by analyzing the received packaging result, the voice encoding result and the voice probability of each frame can be parsed out.

[0199] Step 1503: Perform data recovery on the analysis result.

[0200] It should be noted that during network transmission, when data is lost or arrives late within a certain lifecycle, the inverse operation of the forward error correction algorithm can be used to perform data recovery on the lost or late data.

[0201] Step 1504: Perform voice buffering on the data recovery result.

[0202] It should be noted that by performing voice buffering on the analysis result, the received voice frames are sorted, and then the sorting result is stored in the voice buffer queue. In this way, the out-of-order data after network transmission can be stored in an orderly manner, and the impact of network jitter on data transmission can be eliminated.

[0203] Step 1505: Perform voice delay adjustment based on the voice probability.

[0204] Step 1506: Perform voice decoding.

[0205] Step 1507: Play the voice. Then execute Step 1501.

[0206] It should be noted that the processing procedures corresponding to Step 1505 to Step 1507 can be referred to Figure 16 , Figure 16 which is an exemplary voice processing flowchart provided by an embodiment of the present application; as Figure 16 shown, this exemplary voice processing flow includes Step 1601 to Step 1610, and each step will be described separately below.

[0207] Step 1601: Determine whether the current voice delay is greater than the specified voice delay (for example, 800 milliseconds (ms)). If so, execute Step 1602; if not, execute Step 1607.

[0208] It should be noted that when calculating the current voice delay, the sum of the voice duration corresponding to the buffered voice queue and the voice duration corresponding to the voice playback queue can be used as the current voice delay; among them, the voice in the voice playback queue comes from the buffered voice queue.

[0209] Step 1602: Calculate the speech energy (referred to as the current speech energy) of other sound sources (referred to as the second sound sources).

[0210] It should be noted that when calculating the speech energy of other sound sources, based on the energy of each frame (for example, 20 milliseconds) of speech, the sum of the energies of multiple consecutive frames (for example, 5 frames) is calculated to obtain the speech energy. In the game voice interaction scenario, other sound sources refer to the sound sources used to play game background music (or sound effects).

[0211] Step 1603: Determine whether the speech energy is less than the first specified energy (for example, 16000). If so, execute Step 1607, and after executing Step 1607, execute Step 1609; if not, execute Step 1604.

[0212] Step 1604: Determine whether the speech energy is greater than the second specified energy (for example, 6400). If so, execute Step 1605; if not, execute Step 1610.

[0213] Step 1605: Determine whether the probability of human voice is less than the specified probability (for example, 0.5). If so, execute Step 1606; if not, execute Step 1607.

[0214] Step 1606: Periodically discard speech frames.

[0215] Step 1607: Speech decoding.

[0216] Step 1608: Based on the speech decoding result, play the speech at the original playback speed.

[0217] Step 1609: Based on the speech decoding result, play the speech at an accelerated speed relative to the original playback speed.

[0218] Step 1610: Determine whether the current speech delay is less than the normal delay threshold (for example, 500 ms). If so, execute Step 1607; if not, execute Step 1602.

[0219] It should be noted that the processing for reducing speech delay provided in the embodiments of this application is an end-to-end processing, that is, from the sending end to the receiving end, and is the speech processing performed before speech decoding.

[0220] It can be understood that in a multi-source playback scenario, at the sending end, the voice probability of the voice frame is synchronously packed and sent to the receiving end together with the data after voice coding, so that when the voice delay at the receiving end is relatively high (greater than the specified voice delay), the voice playback delay can be adjusted based on the voice probability to reduce the voice delay. Moreover, when the playback energy of other sound sources is relatively low (lower than the first specified energy) and the voice delay is relatively high, the received voice is played back at an accelerated speed, thereby accelerating the voice consumption speed and reducing the voice delay; when the playback energy of other sound sources is relatively high (greater than the second specified energy) and the voice delay is relatively high, the voice probability of the voice in the voice frame is read, and some voice frames with relatively low voice probability (lower than the specified probability) are periodically discarded, thereby reducing the voice delay and reducing the overall voice delay with the least loss, and improving the voice playback quality. In addition, the voice frames discarded in the embodiments of the present application are voice frames before decoding, which reduces the decoding consumption, and thus can improve the voice playback efficiency and reduce the voice delay.

[0221] The following continues to describe the exemplary structure of the first voice processing device 455 provided in the embodiments of the present application as a software module. In some embodiments, as Figure 2 shown, the software module stored in the first voice processing device 455 in the first memory 450 may include:

[0222] A voice receiving module 4551, configured to receive a to-be-processed voice sequence sent by a sending-end device for a first sound source and the sound source probability corresponding to each to-be-processed voice in the to-be-processed voice sequence, where the sound source probability refers to the probability that the to-be-processed voice is the sound emitted by the first sound source;

[0223] A voice processing module 4552, configured to perform playback acceleration processing on the to-be-processed voice with a sound source probability less than a specified probability in the to-be-processed voice sequence when the current voice delay is greater than the specified voice delay, to obtain a to-be-played voice sequence;

[0224] A voice playback module 4553, configured to perform first voice playback for the first sound source based on the to-be-played voice sequence.

[0225] In the embodiments of the present application, the playback acceleration processing is discarding or accelerating the playback speed, where discarding means discarding the to-be-processed voice with a sound source probability less than a specified probability, and accelerating the playback speed means accelerating the original playback speed of the to-be-processed voice with a sound source probability less than a specified probability at a target acceleration.

[0226] In an embodiment of the present application, the voice receiving module 4551 is further configured to receive the to-be-processed voice sequence sent by the sending device for the first sound source and the sound source probability corresponding to each to-be-processed voice in the to-be-processed voice sequence during the process of playing the second voice for the second sound source.

[0227] In an embodiment of the present application, the voice processing module 4552 is further configured to obtain the current voice energy of the second voice when the current voice delay is greater than the specified voice delay; in the to-be-processed voice sequence, perform play acceleration processing on the to-be-processed voice whose sound source probability is less than the specified probability based on the current voice energy to obtain the to-be-played voice sequence.

[0228] In an embodiment of the present application, the voice processing module 4552 is further configured to, when the current voice energy is less than the first specified energy, in the to-be-processed voice sequence, speed up the playing speed of the to-be-processed voice whose sound source probability is less than the specified probability to obtain the to-be-played voice sequence, where the play acceleration processing is the speed up of the playing speed.

[0229] In an embodiment of the present application, the voice processing module 4552 is further configured to, when the current voice energy is greater than the second specified energy, discard the to-be-processed voice whose sound source probability is less than the specified probability from the to-be-processed voice sequence to obtain the to-be-played voice sequence, where the second specified energy is less than or equal to the first specified energy, and the play acceleration processing is the discard.

[0230] In an embodiment of the present application, the voice processing module 4552 is further configured to obtain the previous discard time, where the previous discard time is the time when the voice discard processing was last performed; when the duration between the previous discard time and the current time is greater than or equal to the specified cycle duration, discard the to-be-processed voice whose sound source probability is less than the specified probability from the to-be-processed voice sequence to obtain the to-be-played voice sequence.

[0231] In an embodiment of the present application, the voice playing module 4553 is further configured to perform voice decoding on the to-be-played voice sequence to obtain a first voice sequence; for the first sound source, perform first voice playing based on the first voice sequence.

[0232] In an embodiment of the present application, the voice processing module 4552 is further configured to, when the current voice delay is less than or equal to the specified voice delay, perform voice decoding on the to-be-processed voice sequence to obtain a second voice sequence; for the first sound source, perform first voice playing based on the second voice sequence.

[0233] Next, the exemplary structure of the second voice processing device 255 provided in the embodiments of the present application implemented as a software module will be further described. In some embodiments, as Figure 3 shown, the software module stored in the second voice processing device 255 in the second memory 250 may include:

[0234] A voice acquisition module 2551, configured to perform voice acquisition on a first sound source in response to a voice acquisition instruction, so as to obtain a voice sequence to be processed;

[0235] A probability determination module 2552, configured to determine a sound source probability based on the to-be-processed features corresponding to each to-be-processed voice in the to-be-processed voice sequence, where the sound source probability refers to the probability that the to-be-processed voice is the sound emitted by the first sound source;

[0236] A voice sending module 2553, configured to send the to-be-processed voice sequence and the sound source probability of each to-be-processed voice to a receiving end device, where the receiving end device is configured to perform a playback acceleration process on the to-be-processed voice with a sound source probability less than a specified probability when the current voice delay is greater than a specified voice delay to implement first voice playback.

[0237] In the embodiments of the present application, the voice acquisition module 2551 is configured to perform voice acquisition on the first sound source in response to the voice acquisition instruction to obtain an initial voice sequence; perform denoising on the initial voice sequence to obtain a voice sequence to be encoded; and perform voice encoding on the voice sequence to be encoded to obtain the to-be-processed voice sequence.

[0238] In the embodiments of the present application, the acquisition of the sound source probability of each to-be-processed voice is implemented through a sound source classification model. The second voice processing device 255 further includes a model training module 2554, configured to obtain initial training data of a to-be-trained model, where the to-be-trained model is a neural network model to be trained for determining the probability that a voice is the sound emitted by the first sound source; perform scene classification on the initial training data to obtain a plurality of scene training data; perform sound source classification on the plurality of scene training data by using the to-be-trained model to obtain a plurality of sound source classification results; and adjust model parameters of the to-be-trained model based on the plurality of sound source classification results to obtain the sound source classification model.

[0239] An embodiment of the present application provides a computer program product, which includes computer-executable instructions or a computer program. The computer-executable instructions or the computer program are stored in a computer-readable storage medium. A first processor of a receiving-end device reads the computer-executable instructions or the computer program from the computer-readable storage medium, and the first processor executes the computer-executable instructions or the computer program, so that the receiving-end device executes the voice processing method applied to the receiving-end device in the embodiment of the present application. A second processor of a sending-end device reads the computer-executable instructions or the computer program from the computer-readable storage medium, and the second processor executes the computer-executable instructions or the computer program, so that the sending-end device executes the voice processing method applied to the sending-end device in the embodiment of the present application.

[0240] An embodiment of the present application provides a computer-readable storage medium, in which computer-executable instructions or a computer program are stored. When the computer-executable instructions or the computer program are executed by a first processor, it will cause the first processor to execute the voice processing method applied to the receiving-end device provided in the embodiment of the present application; or, when the computer-executable instructions or the computer program are executed by a second processor, it will cause the second processor to execute the voice processing method applied to the sending-end device provided in the embodiment of the present application; for example, as Figure 4 the shown voice processing method.

[0241] In some embodiments, the computer-readable storage medium may be a memory such as FRAM, ROM, flash memory, magnetic surface memory, optical disc, or CD-ROM; or may be various devices including one or any combination of the above memories.

[0242] In some embodiments, the computer-executable instructions may be in the form of a program, software, software module, script, or code, and may be written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including being deployed as an independent program or being deployed as a module, component, subroutine, or other unit suitable for use in a computing environment.

[0243] As an example, the computer-executable instructions may or may not correspond to a file in a file system, and may be stored as part of a file that stores other programs or data. For example, they may be stored in one or more scripts in a Hyper Text Markup Language (HTML) document, stored in a single file dedicated to the program being discussed, or stored in multiple cooperating files (for example, files that store one or more modules, subroutines, or code portions).

[0244] As an example, the computer-executable instructions may be deployed to be executed on multiple electronic devices located at one location (in this case, the multiple electronic devices located at one location are the sending device and the receiving device), or on multiple electronic devices distributed at multiple locations and interconnected through a communication network (in this case, the multiple electronic devices distributed at multiple locations and interconnected through a communication network are the sending device and the receiving device).

[0245] It can be understood that in the embodiments of the present application, for data related to voice and the like, when the embodiments of the present application are applied to specific products or technologies, user permission or consent needs to be obtained, and the collection, use, and processing of the relevant data need to comply with the relevant laws, regulations, and standards of the relevant countries and regions. When the relevant data collection and processing in this application are applied in practical examples, the informed consent or separate consent of the personal information subject should be obtained strictly in accordance with the requirements of the relevant national laws and regulations, and subsequent data use and processing behaviors should be carried out within the scope authorized by the laws, regulations, and the personal information subject. In addition, in the present application, for the implementation of the data scraping technical solution related to the training data, when the above embodiments of the present application are applied to specific products or technologies, the process of collecting, using, and processing the relevant data should comply with the requirements of national laws and regulations, conform to the principles of legality, legitimacy, and necessity, not involve obtaining data types prohibited or restricted by laws and regulations, and will not interfere with the normal operation of the target website.

[0246] In summary, since the data to be played of the first sound source received in the embodiments of the present application not only includes the voice sequence to be processed sent by the sending device, but also includes the sound source probability of each voice to be processed sent by the sending device; therefore, when the current voice delay is large, it is possible to perform playback acceleration processing on the voice data to be processed whose sound source probability is less than the specified sound source probability, thereby being able to improve the voice consumption speed of the first sound source, reduce the voice delay, and be able to reduce the impact on the voice of the first sound source, that is, reduce the impact on the intelligibility of the first sound source; thus, the voice playback quality can be improved. In addition, in a multi-source playback scenario, by determining the playback acceleration processing in combination with the voice energy of the second sound source, it is possible to improve the playback quality of the first sound source while reducing the playback impact on the second sound source.

[0247] The above is only the embodiments of the present application and is not intended to limit the protection scope of the present application. Any modifications, equivalent replacements, and improvements made within the spirit and scope of the present application are included in the protection scope of the present application.

Claims

1. A voice processing method, characterized in that, The method includes: Receiving a to-be-processed voice sequence sent by a sending-end device for a first sound source and a sound-source probability corresponding to each to-be-processed voice in the to-be-processed voice sequence, where the sound-source probability refers to the probability that the to-be-processed voice is a sound emitted by the first sound source; When a current voice delay is greater than a specified voice delay, performing a playback acceleration process on the to-be-processed voices in the to-be-processed voice sequence whose sound-source probabilities are less than a specified probability to obtain a to-be-played voice sequence; Performing first voice playback for the first sound source based on the to-be-played voice sequence.

2. The method according to claim 1, characterized in that, The playback acceleration process is discarding or increasing the playback speed. Here, discarding means discarding the to-be-processed voices in the to-be-processed voice sequence whose sound-source probabilities are less than the specified probability, and increasing the playback speed means increasing the original playback speed of the to-be-processed voices in the to-be-processed voice sequence whose sound-source probabilities are less than the specified probability at a target acceleration.

3. The method according to claim 1 or 2, characterized in that, The receiving the to-be-processed voice sequence sent by the sending-end device for the first sound source and the sound-source probability corresponding to each to-be-processed voice in the to-be-processed voice sequence includes: During the process of performing second voice playback for a second sound source, receiving the to-be-processed voice sequence sent by the sending-end device for the first sound source and the sound-source probability corresponding to each to-be-processed voice in the to-be-processed voice sequence; The performing, when the current voice delay is greater than the specified voice delay, the playback acceleration process on the to-be-processed voices in the to-be-processed voice sequence whose sound-source probabilities are less than the specified probability to obtain the to-be-played voice sequence includes: When the current voice delay is greater than the specified voice delay, obtaining the current voice energy of the second voice; In the to-be-processed voice sequence, performing the playback acceleration process on the to-be-processed voices in the to-be-processed voice sequence whose sound-source probabilities are less than the specified probability based on the current voice energy to obtain the to-be-played voice sequence.

4. The method according to claim 3, wherein The performing, in the to-be-processed voice sequence, the playback acceleration process on the to-be-processed voices in the to-be-processed voice sequence whose sound-source probabilities are less than the specified probability based on the current voice energy to obtain the to-be-played voice sequence includes: When the current voice energy is less than a first specified energy, in the to-be-processed voice sequence, increasing the playback speed of the to-be-processed voices in the to-be-processed voice sequence whose sound-source probabilities are less than the specified probability to obtain the to-be-played voice sequence, where the playback acceleration process is increasing the playback speed.

5. The method according to claim 3, characterized in that, The performing, in the to-be-processed voice sequence, the playback acceleration process on the to-be-processed voices in the to-be-processed voice sequence whose sound-source probabilities are less than the specified probability based on the current voice energy to obtain the to-be-played voice sequence includes: When the current voice energy is greater than a second specified energy, discarding, from the to-be-processed voice sequence, the to-be-processed voices in the to-be-processed voice sequence whose sound-source probabilities are less than the specified probability to obtain the to-be-played voice sequence, where the second specified energy is less than or equal to the first specified energy, and the playback acceleration process is discarding.

6. The method according to claim 5, wherein The discarding, from the to-be-processed voice sequence, the to-be-processed voices in the to-be-processed voice sequence whose sound-source probabilities are less than the specified probability to obtain the to-be-played voice sequence includes: Obtain the previous discard time, where the previous discard time is the time when the voice discard process was last executed; When the duration between the previous discard time and the current time is greater than or equal to the specified cycle duration, discard the voice to be processed with a sound source probability less than the specified probability from the voice sequence to be processed, to obtain the voice sequence to be played.

7. The method according to claim 1 or 2, characterized in that, The first voice playback based on the voice sequence to be played for the first sound source includes: Perform voice decoding on the voice sequence to be played to obtain a first voice sequence; For the first sound source, perform first voice playback based on the first voice sequence.

8. The method according to claim 1 or 2, characterized in that, After receiving the voice sequence to be processed sent by the sending device for the first sound source and the sound source probability corresponding to each voice to be processed in the voice sequence to be processed, the method further includes: When the current voice delay is less than or equal to the specified voice delay, perform voice decoding on the voice sequence to be processed to obtain a second voice sequence; For the first sound source, perform first voice playback based on the second voice sequence.

9. A voice processing method, characterized in that The method includes: In response to a voice collection instruction, perform voice collection on the first sound source to obtain a voice sequence to be processed; Based on the processing features corresponding to each voice to be processed in the voice sequence to be processed, determine the sound source probability, where the sound source probability refers to the probability that the voice to be processed is the sound emitted by the first sound source; Send the voice sequence to be processed and the sound source probability of each voice to be processed to a receiving device, where the receiving device is used to perform first voice playback by performing playback acceleration processing on the voice to be processed with a sound source probability less than the specified probability when the current voice delay is greater than the specified voice delay.

10. The method according to claim 9, wherein The performing voice collection on the first sound source in response to a voice collection instruction to obtain a voice sequence to be processed includes: In response to the voice collection instruction, perform voice collection on the first sound source to obtain an initial voice sequence; Denoise the initial voice sequence to obtain a voice sequence to be encoded; Perform voice encoding on the voice sequence to be encoded to obtain the voice sequence to be processed.

11. The method according to claim 9 or 10, characterized in that The obtaining of the sound source probability of each voice to be processed is implemented through a sound source classification model, where the sound source classification model is obtained through training by the following steps: Obtain the initial training data of the model to be trained, where the model to be trained is a neural network model to be trained for determining the probability that a voice is the sound emitted by the first sound source; Perform scene classification on the initial training data to obtain multiple scene training data; Use the model to be trained to perform sound source classification on multiple scene training data to obtain multiple sound source classification results; Adjust the model parameters of the model to be trained based on multiple sound source classification results to obtain the sound source classification model.

12. A first voice processing device, characterized in that, The first voice processing device includes: A voice receiving module, configured to receive a to-be-processed voice sequence sent by a sending-end device for a first sound source and a sound source probability corresponding to each to-be-processed voice in the to-be-processed voice sequence, where the sound source probability refers to the probability that the to-be-processed voice is a sound emitted by the first sound source; A voice processing module, configured to, when a current voice delay is greater than a specified voice delay, perform a playback acceleration process on the to-be-processed voice whose sound source probability is less than a specified probability in the to-be-processed voice sequence to obtain a to-be-played voice sequence; A voice playback module, configured to perform a first voice playback for the first sound source based on the to-be-played voice sequence.

13. A second voice processing device, characterized in that, The second voice processing device includes: A voice acquisition module, configured to, in response to a voice acquisition instruction, acquire voice of a first sound source to obtain a to-be-processed voice sequence; A probability determination module, configured to determine a sound source probability based on to-be-processed features corresponding to each to-be-processed voice in the to-be-processed voice sequence, where the sound source probability refers to the probability that the to-be-processed voice is a sound emitted by the first sound source; A voice sending module, configured to send the to-be-processed voice sequence and the sound source probability of each to-be-processed voice to a receiving-end device, where the receiving-end device is configured to, when a current voice delay is greater than a specified voice delay, implement a first voice playback by performing a playback acceleration process on the to-be-processed voice whose sound source probability is less than a specified probability.

14. A receiving-end device for speech processing, characterized in that, The receiving-end device includes: A first memory, configured to store computer-executable instructions or a computer program; A first processor, configured to, when executing the computer-executable instructions or the computer program stored in the first memory, implement the voice processing method according to any one of claims 1 to 8.

15. A transmitting end device for speech processing, characterized in that The sending-end device includes: A second memory, configured to store computer-executable instructions or a computer program; A second processor, configured to, when executing the computer-executable instructions or the computer program stored in the second memory, implement the voice processing method according to any one of claims 9 to 11.

16. A computer-readable storage medium storing computer-executable instructions or a computer program, characterized in that, When the computer-executable instructions or the computer program are executed by the first processor, the voice processing method according to any one of claims 1 to 8 is implemented; or when the computer-executable instructions or the computer program are executed by the second processor, the voice processing method according to any one of claims 9 to 11 is implemented.

17. A computer program product, comprising computer-executable instructions or a computer program, characterized in that, When the computer-executable instructions or the computer program are executed by the first processor, the voice processing method according to any one of claims 1 to 8 is implemented; when the computer-executable instructions or the computer program are executed by the second processor, the voice processing method according to any one of claims 9 to 11 is implemented.