Speech processing method and related device

By using bandwidth usage data and peak prediction models to predict bandwidth peaks in voice communication scenarios, and reducing voice quality when approaching thresholds, network congestion caused by bandwidth peaks in voice communication is solved, and the voice communication experience is improved.

CN120223549APending Publication Date: 2025-06-27TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311802303.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-25
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The prior art cannot effectively deal with bandwidth peak problems in voice communication scenarios, resulting in network congestion and affecting the voice communication experience.

Method used

Through the bandwidth usage data and peak prediction model for a period of time, the predicted bandwidth peak after that period of time is quickly, effectively and accurately predicted. When the predicted bandwidth peak is close to the preset bandwidth threshold, the speech quality of voice information is reduced to reduce bandwidth peaks and avoid network congestion.

Benefits of technology

Effectively reduce bandwidth peaks, avoid network congestion, improve voice communication effect, and improve voice communication experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120223549A_ABST
    Figure CN120223549A_ABST
Patent Text Reader

Abstract

The invention discloses a voice processing method and a related device. The method comprises the following steps: after obtaining first voice information of a first target time period, obtaining bandwidth use data of the first target time period; and inputting the bandwidth usage data of the first target time period into a peak prediction model, and predicting a bandwidth peak of a second target time period after the first target time period to obtain a predicted bandwidth peak of the second target time period. When it is judged that a first bandwidth difference between the predicted bandwidth peak value of the second target time period and a preset bandwidth threshold value is smaller than a first preset difference, and a second bandwidth difference between the bandwidth usage amount in the bandwidth usage data of the first target time period and the predicted bandwidth peak value of the second target time period is smaller than a second preset difference; and adjusting the voice quality of the first voice information to obtain second voice information of which the voice quality is lower than that of the first voice information. The method can avoid network congestion and improve the voice communication effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and particularly to a voice processing method and related devices. Background Art

[0002] In a voice communication scenario, a target user can achieve voice communication with other users through voice communication; especially in a game voice communication scenario, a target player can achieve voice communication with other players through voice communication.

[0003] In related technologies, in order to improve the voice communication effect, usually the change situation of the user traffic in the voice communication scenario is detected. When the user traffic is large, the voice quality of the voice information uploaded by some users is reduced, so as to improve the voice communication effect.

[0004] However, the above method cannot effectively handle the bandwidth peak problem in the voice communication scenario. When the bandwidth peak in the voice communication scenario is large, it is easy to cause network congestion, resulting in poor voice communication effect and seriously affecting the voice communication experience. Summary of the Invention

[0005] To solve the above technical problems, this application provides a voice processing method and related devices. Through the bandwidth usage data of a period of time and a peak prediction model, the predicted bandwidth peak after this period of time is quickly, effectively and accurately predicted. When the predicted bandwidth peak is close to the preset bandwidth threshold and the bandwidth usage of this period of time is close to the predicted bandwidth peak, the voice quality of the voice information of this period of time is reduced, quickly, effectively and accurately reducing the bandwidth peak to avoid network congestion, improving the voice communication effect, and thus improving the voice communication experience.

[0006] Embodiments of this application disclose the following technical solutions:

[0007] On the one hand, an embodiment of this application provides a voice processing method, and the method includes:

[0008] After obtaining the first voice information of the first target time period, obtain the bandwidth usage data of the first target time period;

[0009] Predict the bandwidth peak of the second target time period according to the bandwidth usage data of the first target time period through a peak prediction model, and obtain the predicted bandwidth peak of the second target time period; the second target time period is after the first target time period;

[0010] If a first bandwidth difference between a predicted bandwidth peak value of the second target time period and a preset bandwidth threshold is less than a first preset difference, and a second bandwidth difference between a bandwidth usage amount in the bandwidth usage data of the first target time period and the predicted bandwidth peak value of the second target time period is less than a second preset difference, adjust a voice quality of the first voice message to obtain a second voice message; the voice quality of the second voice message is lower than the voice quality of the first voice message.

[0011] On the other hand, an embodiment of the present application provides a voice processing device, which includes: an acquisition unit, a prediction unit, and an adjustment unit;

[0012] The acquisition unit is configured to, after acquiring a first voice message of a first target time period, acquire bandwidth usage data of the first target time period;

[0013] The prediction unit is configured to predict a bandwidth peak value of a second target time period based on the bandwidth usage data of the first target time period through a peak prediction model to obtain a predicted bandwidth peak value of the second target time period; the second target time period is after the first target time period;

[0014] The adjustment unit is configured to, if a first bandwidth difference between a predicted bandwidth peak value of the second target time period and a preset bandwidth threshold is less than a first preset difference, and a second bandwidth difference between a bandwidth usage amount in the bandwidth usage data of the first target time period and the predicted bandwidth peak value of the second target time period is less than a second preset difference, adjust a voice quality of the first voice message to obtain a second voice message; the voice quality of the second voice message is lower than the voice quality of the first voice message.

[0015] On the other hand, an embodiment of the present application provides a computer device, which includes a processor and a memory:

[0016] The memory is configured to store a computer program and transmit the computer program to the processor;

[0017] The processor is configured to execute the method described in any of the foregoing aspects according to instructions in the computer program.

[0018] On the other hand, an embodiment of the present application provides a computer-readable storage medium, which is configured to store a computer program, and when the computer program runs on a computer device, enable the computer device to execute the method described in any of the foregoing aspects.

[0019] On the other hand, an embodiment of the present application provides a computer program product, including a computer program, which, when running on a computer device, causes the computer device to execute the method described in any of the foregoing aspects.

[0020] As can be seen from the above technical solution, first, after obtaining the first voice information of the first target time period, the bandwidth usage data of the first target time period is obtained. This step obtains the bandwidth usage data of the target time period for the case of obtaining the voice information of the target time period, so as to pay attention to the bandwidth usage situation of the target time period.

[0021] Then, the bandwidth usage data of the first target time period is input into the peak prediction model to predict the bandwidth peak of the second target time period after the first target time period, and the predicted bandwidth peak of the second target time period is obtained. This step can quickly, effectively and accurately predict the predicted bandwidth peak of the time period after the target time period based on the bandwidth usage data of the target time period through the peak prediction model.

[0022] Finally, when it is determined that the first bandwidth difference between the predicted bandwidth peak of the second target time period and the preset bandwidth threshold is less than the first preset difference, and the second bandwidth difference between the bandwidth usage amount in the bandwidth usage data of the first target time period and the predicted bandwidth peak of the second target time period is less than the second preset difference, the voice quality of the first voice information is adjusted to obtain the second voice information, so that the voice quality of the second voice information is lower than that of the first voice information. This step determines that the predicted bandwidth peak of the time period after the target time period is close to the preset bandwidth threshold, and the bandwidth usage amount of the target time period is close to the predicted bandwidth peak of the time period after the target time period. By reducing the voice quality of the voice information of the target time period, the bandwidth peak of the time period after the target time period can be quickly, effectively and accurately reduced to avoid network congestion.

[0023] Based on this, the method can quickly, effectively and accurately predict the predicted bandwidth peak after a period of time through the bandwidth usage data and the peak prediction model of this period of time. When the predicted bandwidth peak is close to the preset bandwidth threshold and the bandwidth usage amount of this period of time is close to the predicted bandwidth peak, the voice quality of the voice information of this period of time is reduced, and the bandwidth peak is quickly, effectively and accurately reduced to avoid network congestion and improve the voice communication effect, thereby improving the voice communication experience. Description of the Drawings

[0024] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0025] Figure 1 Schematic diagram of a system for a voice processing method provided by an embodiment of the present application;

[0026] Figure 2 Flowchart of a voice processing method provided by an embodiment of the present application;

[0027] Figure 3 Schematic diagram of the system architecture of a voice processing method provided by an embodiment of the present application;

[0028] Figure 4 Schematic diagram of voice processing interaction based on the system architecture of a voice processing method provided by an embodiment of the present application;

[0029] Figure 5 Flowchart of another voice processing method provided by an embodiment of the present application;

[0030] Figure 6 Schematic diagram of the voice communication scenario of a voice processing method provided by an embodiment of the present application;

[0031] Figure 7 Schematic diagram of the voice communication scenario of another voice processing method provided by an embodiment of the present application;

[0032] Figure 8 Structural diagram of a voice processing device provided by an embodiment of the present application;

[0033] Figure 9 Structural diagram of a server provided by an embodiment of the present application;

[0034] Figure 10 Structural diagram of a terminal provided by an embodiment of the present application. Detailed implementation manners

[0035] The following will describe the embodiments of the present application in conjunction with the accompanying drawings.

[0036] At present, in the voice communication scenario, the target user can achieve voice communication with other users through voice communication; for example, in the game voice communication scenario, the target player can achieve voice communication with other players through voice communication. To improve the voice communication effect, usually the change of user traffic in the voice communication scenario is detected. When the user traffic is large, the voice quality of the voice information uploaded by some users is reduced; for example, when it is detected that the player traffic in the game voice communication scenario is large, the voice quality of the voice information uploaded by some players is reduced.

[0037] However, through research, it is found that the above method cannot effectively handle the bandwidth peak problem in the voice communication scenario. When the bandwidth peak in the voice communication scenario is large, it is easy to cause network congestion, resulting in problems such as voice information loss, voice communication delay, and voice communication interruption, thereby resulting in poor voice communication effect and seriously affecting the voice communication experience. For example, the above method cannot effectively handle the bandwidth peak problem in the game voice communication scenario. When the bandwidth peak in the game voice communication scenario is large, it is easy to cause network congestion, resulting in problems such as voice information loss, voice communication delay, and voice communication interruption, thereby resulting in poor voice communication effect and seriously affecting the voice communication experience of players.

[0038] The embodiment of the present application provides a voice processing method. Through the bandwidth usage data of a period of time and a peak prediction model, the predicted bandwidth peak after this period of time is quickly, effectively, and accurately predicted. When the predicted bandwidth peak is close to the preset bandwidth threshold and the bandwidth usage of this period of time is close to the predicted bandwidth peak, the voice quality of the voice information of this period of time is reduced, and the bandwidth peak is quickly, effectively, and accurately reduced to avoid network congestion, improve the voice communication effect, and thus improve the voice communication experience.

[0039] Next, the system architecture of the voice processing method will be introduced. Refer to Figure 1 , Figure 1 FIG. is a schematic diagram of a system for a voice processing method provided by an embodiment of the present application. The system includes a computer device 100, and the computer device 100 is used to execute the voice processing method.

[0040] After the computer device 100 obtains the first voice information of the first target time period, it obtains the bandwidth usage data of the first target time period.

[0041] As an example, the first target time period is the X time period, and the first voice information is game voice information; then after the computer device 100 obtains the game voice information of the X time period, it obtains the bandwidth usage data of the X time period.

[0042] The computer device 100 predicts the bandwidth peak of the second target time period based on the bandwidth usage data of the first target time period through a peak prediction model, and obtains the predicted bandwidth peak of the second target time period; the second target time period is after the first target time period.

[0043] As an example, based on the above example, the second target time period is the Y time period, the Y time period is after the X time period, and the computer device 100 inputs the bandwidth usage data of the X time period into the peak prediction model to predict the bandwidth peak of the Y time period, and obtains that the predicted bandwidth peak of the Y time period is PB of the Y time period. pre 。

[0044] If the first bandwidth difference between the predicted bandwidth peak of the second target time period and the preset bandwidth threshold is less than the first preset difference, and the second bandwidth difference between the bandwidth usage amount in the bandwidth usage data of the first target time period and the predicted bandwidth peak of the second target time period is less than the second preset difference, the computer device 100 adjusts the voice quality of the first voice information to obtain the second voice information; the voice quality of the second voice information is lower than that of the first voice information.

[0045] As an example, the bandwidth usage amount is B, and the preset bandwidth threshold is B. thr Based on the above example, when it is judged that the first bandwidth difference between PB of the Y time period pre and B thr is less than the first preset difference, and the second bandwidth difference between B in the bandwidth usage data of the X time period and PB of the Y time period pre is less than the second preset difference, the computer device 100 adjusts the voice quality of the game voice information of the X time period to obtain the second voice information as the adjusted game voice information, and the voice quality of the adjusted game voice information is lower than that of the game voice information of the X time period.

[0046] That is to say, for the case of obtaining voice information in a target time period, obtain the bandwidth usage data in the target time period to focus on the bandwidth usage situation in the target time period; through the peak prediction model, it is possible to quickly, effectively, and accurately predict the predicted bandwidth peak in the time period after the target time period based on the bandwidth usage data in the target time period; determine that the predicted bandwidth peak in the time period after the target time period is close to the preset bandwidth threshold, and the bandwidth usage amount in the target time period is close to the predicted bandwidth peak in the time period after the target time period. By reducing the voice quality of the voice information in the target time period, it is possible to quickly, effectively, and accurately reduce the bandwidth peak in the time period after the target time period to avoid network congestion. That is, through the bandwidth usage data and the peak prediction model for a period of time, quickly, effectively, and accurately predict the predicted bandwidth peak after this period of time. When the predicted bandwidth peak is close to the preset bandwidth threshold and the bandwidth usage amount in this period of time is close to the predicted bandwidth peak, reduce the voice quality of the voice information in this period of time, quickly, effectively, and accurately reduce the bandwidth peak to avoid network congestion and improve the voice communication effect, thereby enhancing the voice communication experience.

[0047] It should be noted that the voice processing method in the embodiments of the present application involves artificial intelligence. Artificial intelligence is to use a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, and is a theory, method, technology, and application system that perceives the environment, acquires knowledge, and uses knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence is also to study the design principles and implementation methods of various intelligent machines to enable the machines to have the functions of perception, reasoning, and decision-making.

[0048] Artificial intelligence technology is a comprehensive discipline that involves a wide range of fields, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. In the embodiments of the present application, artificial intelligence technology mainly involves natural language processing technology and machine learning / deep learning and other technologies.

[0049] Natural language processing is an important direction in the field of computer science and artificial intelligence. It studies various theories and methods that can realize effective communication between humans and computers in natural language. Natural language processing is a science that integrates linguistics, computer science, and mathematics. Therefore, the research in this field will involve natural language, that is, the language that people use in daily life, so it has a close connection with the research of linguistics. In the embodiments of the present application, natural language processing technology mainly involves technologies such as text processing, semantic understanding, and robot question answering.

[0050] Machine learning / deep learning is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning / deep learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence, and its applications cover all fields of artificial intelligence. Machine learning / deep learning technologies usually include artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, teaching learning and other technologies.

[0051] It should be noted that in the embodiments of the present application, the computer device can be a server or a terminal. The method provided in the embodiments of the present application can be executed independently by the terminal or the server, or can be executed in cooperation by the terminal and the server. Among them, when the method provided in the embodiments of the present application is executed independently by the terminal or the server, its execution method is similar to Figure 1 the corresponding embodiment, mainly replacing the computer device with the terminal or the server. In addition, when the method provided in the embodiments of the present application is executed in cooperation by the terminal and the server, the steps that need to be reflected on the front-end interface can be executed by the terminal, while some steps that require background calculation and do not need to be reflected on the front-end interface can be executed by the server.

[0052] Among them, the terminal can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart voice interaction device, a vehicle-mounted terminal, an extended reality device or an aircraft, etc., but is not limited thereto. The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services, but is not limited thereto. The terminal and the server can be directly or indirectly connected through wired or wireless communication methods, and the present application does not make any restrictions here. For example, the terminal and the server can be connected through a network, and the network can be a wired or wireless network.

[0053] In addition, the embodiments of the present application can be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, intelligent transportation, assisted driving, autonomous driving, digital humans, virtual humans, virtual reality, augmented reality, mixed reality, audio and video, etc.

[0054] Next, taking the computer device executing the method provided in the embodiments of the present application as an example, the voice processing method provided in the embodiments of the present application will be introduced in detail in combination with the accompanying drawings. See Figure 2 , Figure 2 which is a flowchart of a voice processing method provided in the embodiments of the present application. The method includes:

[0055] S201: After obtaining the first voice information of the first target time period, obtain the bandwidth usage data of the first target time period.

[0056] In the related art, generally, the change of user traffic in the voice communication scenario is detected. When the user traffic is large, the voice quality of the voice information uploaded by some users is reduced. However, this method cannot effectively handle the bandwidth peak problem in the voice communication scenario. When the bandwidth peak in the voice communication scenario is large, it is easy to cause network congestion, resulting in problems such as voice information loss, voice communication delay, and voice communication interruption, thus resulting in poor voice communication effect and seriously affecting the voice communication experience.

[0057] Therefore, in the embodiments of the present application, to solve the above problems, considering that the bandwidth peak problem in the voice communication scenario is related to the bandwidth usage data of the time period, in the voice communication scenario for a period of time, after obtaining the voice information of this period of time, it is necessary to obtain the bandwidth usage data of this period of time, so as to effectively handle the bandwidth peak problem in the voice communication scenario based on the bandwidth usage data of this period of time in the subsequent stage, avoid network congestion and improve the voice communication effect.

[0058] That is, after obtaining the first voice information of the first target time period, first obtain the bandwidth usage data of the first target time period. Among them, the first voice information of the first target time period refers to the voice information sent by the object through the device in the first target time period in the voice communication scenario; the bandwidth usage data of the first target time period refers to the bandwidth-related data detected corresponding to the first target time period in the voice communication scenario.

[0059] This S201 obtains the bandwidth usage data of a period of time for the situation of obtaining the voice information of a period of time, so as to pay attention to the bandwidth usage situation of this period of time; it provides basic data for predicting the predicted bandwidth peak after this period of time in the subsequent stage to reduce the bandwidth peak and avoid network congestion.

[0060] As an example of S201, the first target time period is the real-time time period, and the first voice information is the team voice information; after the computer device obtains the team voice information of the real-time time period, it obtains the bandwidth usage data of the real-time time period, and the bandwidth usage data of the real-time time period is the network delay data, bandwidth usage amount, bandwidth growth rate, etc. of the real-time time period.

[0061] S202: Use the peak prediction model to predict the bandwidth peak of the second target time period according to the bandwidth usage data of the first target time period, and obtain the predicted bandwidth peak of the second target time period; the second target time period is after the first target time period.

[0062] In the embodiments of the present application, to solve the above problems, after obtaining the bandwidth usage data for a period of time, since the bandwidth usage data for this period of time is correlated with the bandwidth peak after this period of time, it is necessary to predict the predicted bandwidth peak after this period of time based on the bandwidth usage data for this period of time through a peak prediction model, so as to effectively handle the bandwidth peak problem in the voice communication scenario based on the bandwidth usage data for this period, avoid network congestion, and improve the voice communication effect.

[0063] That is, after executing the above S201 to obtain the bandwidth usage data for the first target time period, the bandwidth usage data for the first target time period is input into the peak prediction model to predict the bandwidth peak for the second target time period after the first target time period, and the predicted bandwidth peak for the second target time period is obtained. Among them, the predicted bandwidth peak for the second target time period refers to the predicted value of the maximum bandwidth value for the second target time period.

[0064] This S202 can quickly, effectively, and accurately predict the predicted bandwidth peak after a period of time based on the bandwidth usage data for a period of time through the peak prediction model; it provides a processing basis for reducing the bandwidth peak and avoiding network congestion in the future.

[0065] As an example of S202, on the basis of the above S201 example, the second target time period is a future time period, and the future time period is after the real-time time period. The computer device inputs the bandwidth usage data for the real-time time period into the peak prediction model to predict the bandwidth peak for the future time period, and the predicted bandwidth peak for the future time period is obtained as the future bandwidth peak PB. pre .

[0066] S203: If the first bandwidth difference between the predicted bandwidth peak for the second target time period and the preset bandwidth threshold is less than the first preset difference, and the second bandwidth difference between the bandwidth usage amount in the bandwidth usage data for the first target time period and the predicted bandwidth peak for the second target time period is less than the second preset difference, adjust the voice quality of the first voice message to obtain a second voice message; the voice quality of the second voice message is lower than that of the first voice message.

[0067] In the embodiments of the present application, to solve the above problems, based on the predicted bandwidth peak value after a certain period of time, considering that effectively handling the bandwidth peak value problem in the voice communication scenario requires controlling the bandwidth peak value to be less than a certain bandwidth threshold, that is, a preset bandwidth threshold; it is necessary to determine whether the predicted bandwidth peak value after the certain period of time is close to the preset bandwidth threshold, and whether the bandwidth usage amount in the bandwidth usage data of the certain period of time is close to the predicted bandwidth peak value after the certain period of time. If so, it means that the maximum bandwidth value in the voice communication scenario is about to reach the preset bandwidth threshold, and it is necessary to reduce the voice quality of the voice information in the certain period of time to effectively handle the bandwidth peak value problem in the voice communication scenario and avoid network congestion to improve the voice communication effect.

[0068] That is, after predicting the predicted bandwidth peak value of the second target time period after the first target time period in the above S202, it is determined whether the first bandwidth difference between the predicted bandwidth peak value of the second target time period and the preset bandwidth threshold is less than the first preset difference, and whether the second bandwidth difference between the bandwidth usage amount in the bandwidth usage data of the first target time period and the predicted bandwidth peak value of the second target time period is less than the second preset difference; if so, it means that the maximum bandwidth value in the voice communication scenario is about to reach the preset bandwidth threshold, and the voice quality of the first voice information is adjusted to obtain the second voice information, so that the voice quality of the second voice information is lower than that of the first voice information.

[0069] Among them, the preset bandwidth threshold refers to the upper limit threshold representing the maximum bandwidth value in the voice communication scenario; the bandwidth usage amount of the first target time period refers to the true value of the bandwidth value of the first target time period; the first preset difference refers to the first upper limit value of the bandwidth difference; the second preset difference refers to the second upper limit value of the bandwidth difference, and the second preset difference may be equal to or not equal to the first preset difference; the second voice information refers to the first voice information with reduced voice quality.

[0070] When specifically implemented, when the first bandwidth difference between the predicted bandwidth peak value of the second target time period and the preset bandwidth threshold is specifically the first bandwidth difference value between the preset bandwidth threshold and the predicted bandwidth peak value of the second target time period, the first bandwidth difference being less than the first preset difference is specifically the first bandwidth difference value being less than the first preset difference value; when the first bandwidth difference between the predicted bandwidth peak value of the second target time period and the preset bandwidth threshold is specifically the first bandwidth ratio of the bandwidth difference value between the preset bandwidth threshold and the predicted bandwidth peak value of the second target time period to the preset bandwidth threshold, the first bandwidth difference being less than the first preset difference is specifically the first bandwidth ratio being less than the first preset ratio.

[0071] The first bandwidth difference between the bandwidth usage in the first target time period and the predicted bandwidth peak in the second target time period, specifically when the second bandwidth difference is the second bandwidth difference between the predicted bandwidth peak in the second target time period and the bandwidth usage in the first target time period, the second bandwidth difference is less than the second preset difference, specifically the second bandwidth difference is less than the second preset difference; the first bandwidth difference between the bandwidth usage in the first target time period and the predicted bandwidth peak in the second target time period, specifically when the second bandwidth difference is the bandwidth difference between the predicted bandwidth peak in the second target time period and the bandwidth usage in the first target time period, and the second bandwidth ratio of the predicted bandwidth peak in the second target time period, the second bandwidth difference is less than the second preset difference, specifically the second bandwidth ratio is less than the second preset ratio.

[0072] It is determined in S203 that the predicted bandwidth peak after a period of time is close to the preset bandwidth threshold, and the bandwidth usage in this period of time is close to the predicted bandwidth peak after this period of time. By reducing the voice quality of the voice information in this period of time, the bandwidth peak after this period of time can be quickly, effectively and accurately reduced; to avoid network congestion, improve the voice communication effect, and thus improve the voice communication experience.

[0073] As an example of S203, the bandwidth usage is B, and the preset bandwidth threshold is B thr , based on the above example of S202, when it is determined that the first bandwidth difference between the future bandwidth peak PB pre and B thr is less than the first preset difference, and the second bandwidth difference between B and PB pre in the bandwidth usage data of the real-time time period is less than the second preset difference, the computer device adjusts the voice quality of the team voice information in the real-time time period to obtain the second voice information as the adjusted team voice information, and the voice quality of the adjusted team voice information is lower than the voice quality of the team voice information in the real-time time period.

[0074] In addition, it is determined whether the first bandwidth difference between the predicted bandwidth peak in the second target time period and the preset bandwidth threshold is less than the first preset difference, and whether the second bandwidth difference between the bandwidth usage and the predicted bandwidth peak in the second target time period in the bandwidth usage data of the first target time period is less than the second preset difference; if not, it means that there is still a certain gap between the maximum bandwidth in the voice communication scenario and the preset bandwidth threshold, and there is no need to adjust the voice quality of the first voice information, and the first target time period is updated to continue to execute the above voice processing method.

[0075] As can be seen from the above technical solution, first, after obtaining the first voice information of the first target time period, the bandwidth usage data of the first target time period is obtained. This step obtains the bandwidth usage data of the target time period for the case of obtaining the voice information of the target time period, so as to pay attention to the bandwidth usage situation of the target time period.

[0076] Then, the bandwidth usage data of the first target time period is input into the peak prediction model to predict the bandwidth peak of the second target time period after the first target time period, and the predicted bandwidth peak of the second target time period is obtained. This step can quickly, effectively and accurately predict the predicted bandwidth peak of the time period after the target time period based on the bandwidth usage data of the target time period through the peak prediction model.

[0077] Finally, when it is determined that the first bandwidth difference between the predicted bandwidth peak of the second target time period and the preset bandwidth threshold is less than the first preset difference, and the second bandwidth difference between the bandwidth usage amount in the bandwidth usage data of the first target time period and the predicted bandwidth peak of the second target time period is less than the second preset difference, the voice quality of the first voice information is adjusted to obtain the second voice information, so that the voice quality of the second voice information is lower than that of the first voice information. This step determines that the predicted bandwidth peak of the time period after the target time period is close to the preset bandwidth threshold, and the bandwidth usage amount of the target time period is close to the predicted bandwidth peak of the time period after the target time period. By reducing the voice quality of the voice information of the target time period, the bandwidth peak of the time period after the target time period can be quickly, effectively and accurately reduced to avoid network congestion.

[0078] Based on this, the method can quickly, effectively and accurately predict the predicted bandwidth peak after a period of time through the bandwidth usage data and the peak prediction model of this period of time. When the predicted bandwidth peak is close to the preset bandwidth threshold and the bandwidth usage amount of this period of time is close to the predicted bandwidth peak, the voice quality of the voice information of this period of time is reduced, and the bandwidth peak is quickly, effectively and accurately reduced to avoid network congestion and improve the voice communication effect, thereby improving the voice communication experience.

[0079] In addition, in the case where most cloud service providers charge according to the bandwidth peak, quickly, effectively and accurately reducing the bandwidth peak can effectively reduce the bandwidth peak billing cost.

[0080] In the embodiment of the present application, for the peak prediction model in the above S202, the initial prediction model is trained in advance through the bandwidth usage data of the first historical time period and the corresponding historical bandwidth peak of the second historical time period after the first historical time period to obtain the peak prediction model, and the peak prediction model is used to predict the bandwidth peak of the time period after a period of time based on the bandwidth usage data of this period of time.

[0081] By using the bandwidth usage data in the first historical period and the corresponding historical bandwidth peak in the second historical period after the first historical period to train the initial prediction model to obtain the peak prediction model, it actually means: inputting the bandwidth usage data in the first historical period into the initial prediction model to predict the bandwidth peak in the second historical period, obtaining the predicted bandwidth peak in the second historical period, and the predicted bandwidth peak in the second historical period is the predicted value of the maximum bandwidth value in the second historical period predicted based on the bandwidth usage data in the first historical period; combining the true value of the maximum bandwidth value in the second historical period, that is, the historical bandwidth peak in the second historical period, to adjust the model parameters of the initial prediction model to achieve training, so that the predicted bandwidth peak in the second historical period is close to the historical bandwidth peak in the second historical period to complete the model training, and taking the trained initial prediction model as the peak prediction model. Therefore, the present application provides a possible implementation manner, and the training steps of the peak prediction model in the above S202 include the following S1-S2 (not shown in the figure):

[0082] S1: Use the initial prediction model to predict the bandwidth peak in the second historical period according to the bandwidth usage data in the first historical period, and obtain the predicted bandwidth peak in the second historical period.

[0083] Among them, the bandwidth usage data in the first historical period refers to the bandwidth-related data corresponding to voice communication detected in the first historical period; the historical bandwidth peak in the second historical period refers to the true value of the maximum bandwidth value in the second historical period; the initial prediction model refers to a prediction model constructed based on the time series prediction algorithm; the peak prediction model refers to the initial prediction model trained based on the bandwidth usage data in the first historical period and the corresponding historical bandwidth peak in the second historical period.

[0084] S2: According to the predicted bandwidth peak in the second historical period and the historical bandwidth peak in the second historical period, adjust the model parameters of the initial prediction model to obtain the peak prediction model.

[0085] In the specific implementation of S2, according to the loss function of the initial prediction model, minimize the predicted bandwidth peak in the second historical period and the historical bandwidth peak in the second historical period, and adjust the model parameters of the initial prediction model to obtain the peak prediction model.

[0086] It should be noted that in the process of training the initial prediction model to obtain the peak prediction model in S1-S2, it is necessary to pay attention to the loss function, optimization algorithm, learning rate, batch size, number of iterations, etc. of the initial prediction model to prevent the obtained peak prediction model from overfitting, so as to improve the prediction performance and generalization ability of the peak prediction model.

[0087] The S1-S2 uses an initial prediction model to predict the bandwidth peak in the second historical period based on the bandwidth usage data in the first historical period, obtaining the predicted bandwidth peak in the second historical period. According to the training direction that makes the predicted bandwidth peak in the second historical period close to the historical bandwidth peak in the second historical period, the model parameters of the initial prediction model are adjusted to quickly, effectively, and accurately learn the correlation between the bandwidth usage data in the first historical period and the historical bandwidth peak in the second historical period, completing the model training to obtain the peak prediction model; providing a prediction model for quickly, effectively, and accurately predicting the bandwidth peak after a period of time based on the bandwidth usage data for that period in the future.

[0088] As an example of S1-S2, the first historical period is historical period 1, the second historical period is historical period 2, historical period 2 is after historical period 1, and the historical bandwidth peak is PB. his , and the initial prediction model is a Long Short-Term Memory Network (LSTM); based on the example of S202 above, the computer device uses the bandwidth usage data in historical period 1 and the corresponding PB in historical period 2 his to train the LSTM to obtain the peak prediction model.

[0089] Specifically, the computer device inputs the bandwidth usage data in historical period 1 into the LSTM to predict the bandwidth peak in historical period 2, obtaining the predicted bandwidth peak in historical period 2 as the PB in historical period 2 pre ; combined with the PB in historical period 2 his , the model parameters of the LSTM are adjusted to achieve training, making the PB in historical period 2 pre close to the PB in historical period 2 his Completing the model training, the trained LSTM is used as the peak prediction model.

[0090] Among them, considering the required format of the input data of the LSTM, the bandwidth usage data in historical period 1 is preprocessed time series data, for example, time series data after normalization processing. The LSTM is a multi-layer LSTM network quickly constructed using the deep learning framework PyTorch. Each layer of the LSTM network includes a certain number of LSTM cells. The number of layers of the LSTM network and the number of cells of the LSTM cells can be determined through experiments to enable the LSTM to achieve the best prediction effect.

[0091] LSTM is a variant of the recurrent neural network, which can capture the non - linear relationships and long - term dependencies in time - series data and can process and learn multivariate time - series data more effectively. In addition, the initial prediction model can also be an Autoregressive Integrated Moving Average Model (ARIMA) or a Prophet model. Among them, both the ARIMA model and the Prophet model are prediction models constructed based on time - series prediction algorithms.

[0092] Based on the above description, refer to Figure 3 , Figure 3 which is a schematic diagram of the system architecture of a voice processing method provided by an embodiment of the present application; the system architecture includes a voice control center platform, a bandwidth detection unit, a bandwidth database, a peak prediction unit, and a quality control unit. Among them, the voice control center platform is connected to the bandwidth detection unit, the bandwidth detection unit is connected to the bandwidth database and the peak prediction unit, the bandwidth database is connected to the peak prediction unit, the peak prediction unit is connected to the quality control unit, and the quality control unit is connected to the voice control center platform.

[0093] On the basis of Figure 3 , refer to Figure 4 , Figure 4 which is a schematic diagram of voice processing interaction based on the system architecture of a voice processing method provided by an embodiment of the present application; the voice control center platform receives the real - time voice information sent by the object through the device and requests bandwidth usage data from the bandwidth detection unit; on the one hand, the bandwidth detection unit requests historical bandwidth usage data from the bandwidth database so that the bandwidth database sends the historical bandwidth usage data to the peak prediction unit, and on the other hand, the bandwidth detection unit sends the real - time bandwidth usage data to the peak prediction unit; the peak prediction unit, based on training the initial prediction model with historical bandwidth usage data to obtain the peak prediction model, predicts the future bandwidth peak based on the real - time bandwidth usage data through the peak prediction model and sends the future bandwidth peak to the quality control unit; the quality control unit determines whether the future bandwidth peak is close to the preset bandwidth threshold to determine whether to adjust the voice quality of the real - time voice information. If so, it sends a voice quality adjustment instruction to the voice control center platform; the voice control center platform adjusts the voice quality of the real - time voice information so that the voice quality of the adjusted real - time voice information is lower than that of the real - time voice information and sends the adjusted real - time voice information to the device used by the object.

[0094] In the embodiments of the present application, considering that when the packet loss rate associated with the voice quality is relatively stable, adjusting the coding complexity or coding redundancy of the voice information can not only adjust the voice quality of the voice information, but also does not affect the voice information to achieve voice communication. Among them, the coding complexity of the voice information refers to the amount of calculation required to complete the coding of the voice information, and the coding redundancy of the voice information refers to the amount of redundancy added for error correction during the coding process of the voice information. Therefore, the step of adjusting the voice quality of the first voice information to obtain the second voice information in S203 above may include the following various implementation manners:

[0095] One implementation manner of the above S203 is that: the higher the coding complexity of the voice information, the higher the voice quality of the voice information, and the lower the coding complexity of the voice information, the lower the voice quality of the voice information. Then, adjusting the voice quality of the first voice information to obtain the second voice information in S203 above may be: reducing the coding complexity of the first voice information to obtain the second voice information. That is, first determine the first coding complexity of the first voice information, and then adjust the voice quality of the first voice information with a second coding complexity lower than the first coding complexity to obtain the second voice information. Therefore, the present application provides a possible implementation manner. The step of adjusting the voice quality of the first voice information to obtain the second voice information in S203 above includes the following S3-S4 (not shown in the figure):

[0096] S3: Determine the first coding complexity of the first voice information.

[0097] S4: Adjust the voice quality of the first voice information according to the second coding complexity to obtain the second voice information; the second coding complexity is lower than the first coding complexity.

[0098] Among them, the first coding complexity of the first voice information refers to the first amount of calculation required to complete the coding of the first voice information; the second coding complexity of the second voice information refers to the second amount of calculation required to complete the coding of the second voice information.

[0099] By reducing the coding complexity of the voice information for a period of time, S3-S4 can quickly, effectively, and accurately reduce the amount of calculation required for voice information coding, so as to quickly, effectively, and accurately reduce the voice quality of the voice information for that period of time, thereby reducing the bandwidth peak after that period of time. In addition, reducing the coding complexity of the voice information for a period of time can reduce the load of the computer device and improve the system performance of the computer device.

[0100] As an example of S3-S4, based on the example of S203 above, the computer device first determines that the first coding complexity of the team voice information is coding complexity 1, then determines that the second coding complexity lower than coding complexity 1 is coding complexity 2, and adjusts the voice quality of the team voice information in the real-time time period through coding complexity 2 to obtain the adjusted team voice information.

[0101] Among them, considering that the coding complexity is associated with the coding configuration and coding bit rate of the encoder, adjusting the coding configuration or coding bit rate of the encoder can adjust the coding complexity of the voice information; among them, the coding configuration of the encoder refers to the coding parameters pre-configured by the encoder, and the coding bit rate refers to the amount of data coded per unit time. Therefore, the above S4 can include the following multiple implementation manners:

[0102] One implementation manner of the above S4 is that: the higher the coding configuration of the encoder, the higher the coding complexity of the voice information, and the lower the coding configuration of the encoder, the lower the coding complexity of the voice information; then S4 can be: reducing the coding configuration of the encoder corresponding to the first coding complexity of the first voice information to obtain the second voice information; that is, first determining the first encoder corresponding to the first coding complexity, and then adjusting the voice quality of the first voice information through a second encoder whose coding configuration is lower than that of the first encoder to obtain the second voice information. Therefore, the present application provides a possible implementation manner, and the above S4 includes the following S41-S42 (not shown in the figure):

[0103] S41: Determine the first encoder corresponding to the first coding complexity.

[0104] S42: Adjust the voice quality of the first voice information according to the second encoder to obtain the second voice information; the coding configuration of the second encoder is lower than that of the first encoder.

[0105] Among them, the coding configuration of the first encoder refers to the first coding parameters pre-configured by the first encoder; the coding configuration of the second encoder refers to the second coding parameters pre-configured by the second encoder.

[0106] By reducing the coding configuration of the encoder of the voice information for a period of time and using an encoder with a lower coding configuration, S41-S42 can quickly, effectively, and accurately reduce the amount of calculation required for coding the voice information, so as to quickly, effectively, and accurately reduce the voice quality of the voice information for this period of time, thereby reducing the bandwidth peak after this period of time.

[0107] As an example of S41-S42, based on the above S4 example, the computer device first determines that the first encoder corresponding to the encoding complexity 1 of the team voice information is HE-AAC, and then determines that the second encoder with an encoding configuration lower than that of HE-AAC is LC-AAC. The voice quality of the team voice information in the real-time time period is adjusted by LC-AAC to obtain the adjusted team voice information.

[0108] Another implementation of the above S4 means that the higher the encoding bit rate, the higher the encoding complexity of the voice information, and the lower the encoding bit rate, the lower the encoding complexity of the voice information. Then, the above S4 can be: reducing the encoding bit rate corresponding to the first encoding complexity of the first voice information to obtain the second voice information. That is, first determine the first encoding bit rate corresponding to the first encoding complexity, and then adjust the voice quality of the first voice information through the second encoding bit rate lower than the first encoding bit rate to obtain the second voice information. Therefore, the present application provides a possible implementation manner, and the above S4 includes the following S43-S44 (not shown in the figure):

[0109] S43: Determine the first encoding bit rate corresponding to the first encoding complexity.

[0110] S44: Adjust the voice quality of the first voice information according to the second encoding bit rate to obtain the second voice information; the second encoding bit rate is lower than the first encoding bit rate.

[0111] Among them, the first encoding bit rate refers to the first data volume encoded per unit time; the second encoding bit rate refers to the second data volume encoded per unit time.

[0112] This S43-S44 reduces the encoding bit rate of the voice information in this period of time. Using a lower encoding bit rate can quickly, effectively, and accurately reduce the computational amount required for voice information encoding, so as to quickly, effectively, and accurately reduce the voice quality of the voice information in this period of time, thereby reducing the bandwidth peak after this period of time.

[0113] As an example of S43-S44, based on the above S4 example, the computer device first determines that the first encoding bit rate corresponding to the encoding complexity 1 of the team voice information is the encoding bit rate 1, and then determines that the second encoding bit rate lower than the encoding bit rate 1 is the encoding bit rate 2. The voice quality of the team voice information in the real-time time period is adjusted by the encoding bit rate 2 to obtain the adjusted team voice information.

[0114] Among them, obtaining the second voice message by reducing the coding bit rate corresponding to the first coding complexity of the first voice message may involve flexibly reducing the coding bit rate of some voice sub-messages in the first voice message according to different voice complexities of the voice sub-messages in the first voice message; or flexibly reducing the coding bit rate of some voice sub-messages in the first voice message according to different voice activity degrees of the voice sub-messages in the first voice message; or reducing the coding bit rate of each voice sub-message in the first voice message to obtain the second voice message. Here, the voice complexity of the voice message refers to the complexity degree of the voice signal corresponding to the voice message; the voice activity degree of the voice message refers to the effective degree of the voice signal corresponding to the voice message. Therefore, the above S44 may include the following multiple implementation manners:

[0115] One implementation manner of the above S44 is that, in order to reduce the loss of the voice quality of the first voice message on the basis of reducing the voice quality of the first voice message, instead of reducing the coding bit rate of each voice sub-message in the first voice message, the coding bit rate of the voice sub-message corresponding to the lower voice complexity in the first voice message is reduced to obtain the second voice message. That is, the voice sub-message corresponding to the first voice complexity and the voice sub-message corresponding to the second voice complexity lower than the first voice complexity in the first voice message are determined; the voice quality of the voice sub-message corresponding to the second voice complexity in the first voice message is adjusted by the second coding bit rate lower than the first coding bit rate to obtain the second voice message. Therefore, the present application provides a possible implementation manner, and the above S44 includes the following S44a - S44b (not shown in the figure):

[0116] S44a: Determine the voice sub-message corresponding to the first voice complexity and the voice sub-message corresponding to the second voice complexity in the first voice message; the second voice complexity is lower than the first voice complexity.

[0117] S44b: Adjust the voice quality of the voice sub-message corresponding to the second voice complexity in the first voice message according to the second coding bit rate to obtain the second voice message.

[0118] Among them, the first voice complexity refers to the first complexity degree of the voice signal corresponding to a part of the voice sub-messages in the first voice message; the second voice complexity refers to the second complexity degree of the voice signal corresponding to another part of the voice sub-messages in the first voice message.

[0119] The S44a-S44b reduces the encoding bit rate of the speech sub-information corresponding to the lower speech complexity in the speech information for a period of time. Not only can it quickly, effectively, and accurately reduce the computational amount required for encoding the speech information with a lower encoding bit rate, quickly, effectively, and accurately reduce the speech quality of the speech information for this period of time, but also mitigate the loss of the speech quality of the speech information for this period of time. Thus, on the basis of reducing the bandwidth peak after this period of time, it avoids excessively reducing the speech quality of the speech information for this period of time.

[0120] As an example of S44a-S44b, on the basis of the above S44 example, the computer device determines that the speech sub-information corresponding to the first speech complexity in the team speech information is team speech sub-information 1, and the speech sub-information corresponding to the second speech complexity lower than the first speech complexity is team speech sub-information 2; adjusts the speech quality of team speech sub-information 2 in the team speech information of the real-time time period with an encoding bit rate 2 lower than encoding bit rate 1 to obtain the adjusted team speech information.

[0121] Another implementation manner of the above S44 is: in order to mitigate the loss of the speech quality of the first speech information on the basis of reducing the speech quality of the first speech information; instead of reducing the encoding bit rate of each speech sub-information in the first speech information, the encoding bit rate of the speech sub-information corresponding to the lower speech activity in the first speech information is reduced to obtain the second speech information. That is, determine the speech sub-information corresponding to the first speech activity and the speech sub-information corresponding to the second speech activity lower than the first speech activity in the first speech information; adjust the speech quality of the speech sub-information corresponding to the second speech activity in the first speech information with a second encoding bit rate lower than the first encoding bit rate to obtain the second speech information. Therefore, the present application provides a possible implementation manner, and the above S44 includes the following S44c-S44d (not shown in the figure):

[0122] S44c: Determine the speech sub-information corresponding to the first speech activity and the speech sub-information corresponding to the second speech activity in the first speech information; the second speech activity is lower than the first speech activity.

[0123] S44d: Adjust the speech quality of the speech sub-information corresponding to the second speech activity in the first speech information according to the second encoding bit rate to obtain the second speech information.

[0124] Wherein, the first speech activity refers to the first effective degree of the speech signals corresponding to a part of the speech sub-information in the first speech information; the second speech activity refers to the second effective degree of the speech signals corresponding to another part of the speech sub-information in the first speech information.

[0125] The S44c-S44d not only can quickly, effectively, and accurately reduce the computational amount required for encoding voice information by using a lower encoding bit rate by reducing the encoding bit rate of the voice sub-information corresponding to the lower voice activity in the voice information for a period of time, quickly, effectively, and accurately reduce the voice quality of the voice information for a period of time, but also mitigate the loss of the voice quality of the voice information for a period of time, so as to avoid excessively reducing the voice quality of the voice information for a period of time on the basis of reducing the bandwidth peak after a period of time.

[0126] As an example of S44c-S44d, on the basis of the above S44 example, the computer device determines that the voice sub-information corresponding to the first voice activity in the team voice information is team voice sub-information 3, and the voice sub-information corresponding to the second voice activity lower than the first voice activity is team voice sub-information 4; the voice quality of team voice sub-information 4 in the team voice information in the real-time time period is adjusted by the encoding bit rate 2 lower than the encoding bit rate 1 to obtain the adjusted team voice information.

[0127] Another implementation manner of the above S44 means that in order to more efficiently and comprehensively reduce the voice quality of the first voice information, it is necessary to reduce the encoding bit rate of each voice sub-information in the first voice information; that is, the voice quality of each voice sub-information in the first voice information is adjusted by the second encoding bit rate lower than the first encoding bit rate to obtain the second voice information. Therefore, the present application provides a possible implementation manner, and the above S44 includes the following S44e (not shown in the figure): the voice quality of each voice sub-information in the first voice information is adjusted according to the second encoding bit rate to obtain the second voice information.

[0128] The S44e can quickly, effectively, accurately, and comprehensively reduce the computational amount required for encoding voice information by reducing the encoding bit rate of each voice sub-information in the voice information for a period of time, and use a lower encoding bit rate to quickly, effectively, accurately, and comprehensively reduce the voice quality of the voice information for a period of time, thereby reducing the bandwidth peak after a period of time.

[0129] As an example of S44e, on the basis of the above S44 example, each voice sub-information in the team voice information is each team voice sub-information, and the computer device adjusts the voice quality of each team voice sub-information in the team voice information in the real-time time period by the encoding bit rate 2 lower than the encoding bit rate 1 to obtain the adjusted team voice information.

[0130] Among them, considering that the encoding bit rate is determined by the sampling rate and the bit depth, to reduce the first encoding bit rate corresponding to the first encoding complexity of the first voice information, it is necessary to reduce the sampling rate and the bit depth used to determine the first encoding bit rate, so as to determine a second encoding bit rate lower than the first encoding bit rate; wherein, the sampling rate refers to the number of times the voice signal corresponding to the voice information is sampled per unit time, and the bit depth refers to the number of bits recorded by the sampling per unit time. That is, first determine the first sampling rate and the first bit depth corresponding to the first encoding bit rate, adjust the first sampling rate and the first bit depth to obtain a second sampling rate and a second bit depth, so as to meet one or more of the following conditions: the second sampling rate is lower than the first sampling rate, the second bit depth is lower than the first bit depth, the second sampling rate is lower than the first sampling rate and the second bit depth is lower than the first bit depth; then determine the second encoding bit rate through the second sampling rate and the second bit depth. Therefore, the present application provides a possible implementation manner. The steps for determining the second encoding bit rate in S44, S44a-S44b, S44c-S44d, or S44e above include the following S44f-S44h:

[0131] S44f: Determine the first sampling rate and the first bit depth corresponding to the first encoding bit rate.

[0132] S44g: Adjust the first sampling rate and the first bit depth to obtain a second sampling rate and a second bit depth; the second sampling rate and the second bit depth meet one or more of the following conditions: the second sampling rate is lower than the first sampling rate, the second bit depth is lower than the first bit depth, the second sampling rate is lower than the first sampling rate and the second bit depth is lower than the first bit depth.

[0133] S44h: Determine the second encoding bit rate according to the second sampling rate and the second bit depth.

[0134] Among them, the first sampling rate refers to the first number of times the voice signal corresponding to the voice information is sampled per unit time, and the first bit depth refers to the first number of bits recorded by the sampling per unit time; the second sampling rate refers to the second number of times the voice signal corresponding to the voice information is sampled per unit time, and the second bit depth refers to the second number of bits recorded by the sampling per unit time.

[0135] These S44f-S44h determine a lower encoding bit rate by reducing one or more of the sampling rate and the bit depth used to determine the encoding bit rate of the voice information for a period of time, laying a foundation for quickly, effectively, and accurately reducing the computational amount required for encoding the voice information by using the lower encoding bit rate, so as to quickly, effectively, and accurately reduce the voice quality of the voice information for this period of time, and provide the encoding bit rate.

[0136] As an example of S44f - S44h, based on the above examples of S44, S44a - S44b, S44c - S44d, or S44e, the computer device first determines that the first sampling rate corresponding to the encoding bit rate 1 is 44.1 Hz and the first bit depth is 16 bits, adjusts 44.1 Hz and 16 bits, and obtains the second sampling rate of 22.05 Hz and the second bit depth of 8 bits; then determines the encoding bit rate 2 through 22.05 Hz and 8 bits.

[0137] Another implementation manner of the above S203 means that the higher the encoding redundancy of the voice information, the higher the voice quality of the voice information, and the lower the encoding redundancy of the voice information, the lower the voice quality of the voice information; then adjusting the voice quality of the first voice information to obtain the second voice information in S203 can be: reducing the encoding redundancy of the first voice information to obtain the second voice information, that is, first determining the first encoding redundancy of the first voice information, and then adjusting the voice quality of the first voice information through the second encoding redundancy lower than the first encoding redundancy to obtain the second voice information. Therefore, the present application provides a possible implementation manner, and the prediction steps of multiple preset output probabilities in the above S203 include the following S5 - S6:

[0138] S5: Determine the first encoding redundancy of the first voice information.

[0139] S6: Adjust the voice quality of the first voice information according to the second encoding redundancy to obtain the second voice information; the second encoding redundancy is lower than the first encoding redundancy.

[0140] Among them, the first encoding redundancy of the first voice information refers to the first redundancy amount added for error correction during the encoding process of the first voice information; the first encoding redundancy of the second voice information refers to the second redundancy amount added for error correction during the encoding process of the second voice information.

[0141] By reducing the encoding redundancy of the voice information for this period of time, S5 - S6 can quickly, effectively, and accurately reduce the redundancy amount added for error correction during the encoding process of the voice information, so as to quickly, effectively, and accurately reduce the voice quality of the voice information for this period of time, thereby reducing the bandwidth peak value after this period of time.

[0142] However, reducing the encoding redundancy of the voice information for a period of time may cause the lower encoding redundancy of the voice information for this period of time to make it more difficult to recover the lost data when data loss occurs; therefore, when reducing the encoding redundancy of the voice information for this period of time, it is necessary to balance the relationship between the bandwidth usage amount and the voice quality of the voice information.

[0143] As an example of S5 - S6, based on the above S203 example, the computer device first determines that the first coding redundancy of the team voice information is coding redundancy 1, then determines that the second coding redundancy lower than coding redundancy 1 is coding redundancy 2, and adjusts the voice quality of the team voice information in the real - time time period through coding redundancy 2 to obtain the adjusted team voice information.

[0144] Among them, considering that the coding redundancy is associated with the error - correcting coding level and the redundant data ratio, adjusting the error - correcting coding level or the redundant data ratio can adjust the coding complexity of the voice information; among them, the error - correcting coding level of the voice information refers to the error - correcting ability of the redundant amount added for error correction during the coding process of the voice information, and the redundant data ratio of the voice information refers to the ratio of the redundant amount added for error correction during the coding process of the voice information to the voice information. Therefore, the above S6 can include the following multiple implementation manners:

[0145] One implementation manner of the above S6 is that: the higher the error - correcting coding level of the voice information, the higher the coding redundancy of the voice information, and the lower the error - correcting coding level of the voice information, the lower the coding redundancy of the voice information; then the above S6 can be: reducing the error - correcting coding level corresponding to the first coding redundancy of the first voice information to obtain the second voice information; that is, first determining the first error - correcting coding level corresponding to the first coding redundancy, and then adjusting the voice quality of the first voice information through the second error - correcting coding level lower than the first error - correcting coding level to obtain the second voice information. Therefore, the present application provides a possible implementation manner, and the above S6 includes the following S61 - S62 (not shown in the figure):

[0146] S61: Determine the first error - correcting coding level corresponding to the first coding redundancy.

[0147] S62: Adjust the voice quality of the first voice information according to the second error - correcting coding level to obtain the second voice information; the second error - correcting coding level is lower than the first error - correcting coding level.

[0148] Among them, the first error - correcting coding level refers to the first error - correcting ability of the redundant amount added for error correction during the coding process of the first voice information; the second error - correcting coding level refers to the second error - correcting ability of the redundant amount added for error correction during the coding process of the second voice information.

[0149] The S61 - S62 can quickly, effectively, and accurately reduce the redundant amount added for error correction during the coding process of the voice information by reducing the error - correcting coding level of the voice information for a period of time, so as to quickly, effectively, and accurately reduce the voice quality of the voice information for this period of time, thereby reducing the bandwidth peak after this period of time.

[0150] As an example of S61-S62, based on the above S6 example, the computer device first determines that the first error correction coding level corresponding to the coding redundancy 1 of the team voice information is Reed-Solomon, and then determines that the second error correction coding level lower than Reed-Solomon is parity check. The voice quality of the team voice information in the real-time time period is adjusted through parity check to obtain the adjusted team voice information. Among them, Reed-Solomon refers to an error correction coding level based on blocks, and parity check is an error correction coding level based on the number.

[0151] Another implementation manner of the above S6 is: the higher the redundant data ratio of the voice information, the higher the coding redundancy of the voice information, and the lower the redundant data ratio of the voice information, the lower the coding redundancy of the voice information; then the above S6 can be: reducing the redundant data ratio corresponding to the first coding redundancy of the first voice information to obtain the second voice information; that is, first determining the first redundant data ratio corresponding to the first coding redundancy, and then adjusting the voice quality of the first voice information through the second redundant data ratio lower than the first redundant data ratio to obtain the second voice information. Therefore, the present application provides a possible implementation manner, and the above S6 includes the following S63-S64 (not shown in the figure):

[0152] S61: Determine the first redundant data ratio corresponding to the first coding redundancy.

[0153] S62: Adjust the voice quality of the first voice information according to the second redundant data ratio to obtain the second voice information; the second redundant data ratio is lower than the first redundant data ratio.

[0154] The first redundant data ratio refers to the first ratio of the redundant amount added for error correction in the coding process of the first voice information to the voice information; the second redundant data ratio refers to the second ratio of the redundant amount added for error correction in the coding process of the second voice information to the voice information.

[0155] The S63-S64 can quickly, effectively, and accurately reduce the redundant amount added for error correction in the coding process of the voice information by reducing the redundant data ratio of the voice information for a period of time, so as to quickly, effectively, and accurately reduce the voice quality of the voice information for this period of time, thereby reducing the bandwidth peak after this period of time.

[0156] As an example of S63-S64, based on the above S6 example, the computer device first determines that the first redundant data ratio corresponding to the coding redundancy 1 of the team voice information is 20%, and then determines that the second redundant data ratio lower than 20% is 10%. The voice quality of the team voice information in the real-time time period is adjusted by 10% to obtain the adjusted team voice information.

[0157] In addition, in the embodiments of the present application, considering that when the predicted bandwidth peak after another period of time obtained by subsequent prediction is far from the above preset bandwidth threshold, it is necessary to improve the voice quality of the voice information in that another period of time, and more intelligently and dynamically adjust the voice information in different periods of time based on the bandwidth peak predicted by the bandwidth usage data in different periods of time, so as to further improve the voice communication effect and voice communication experience. Based on this, after the above S201-S203, for obtaining the third voice information in the third target period after the first target period, first, obtain the bandwidth usage data of the third target period; then, input the bandwidth usage data of the third target period into the peak prediction model to predict the bandwidth peak in the fourth target period after the third target period, and obtain the predicted bandwidth peak of the fourth target period; finally, determine whether the third bandwidth difference between the predicted bandwidth peak of the fourth target period and the preset bandwidth threshold is greater than or equal to the first preset difference; if so, it means that the maximum bandwidth in the voice communication scenario is far from the preset bandwidth threshold, and adjust the voice quality of the third voice information to obtain the fourth voice information, so that the voice quality of the fourth voice information is higher than that of the third voice information. Therefore, the present application provides a possible implementation manner, and the method further includes the following S7-S9 (not shown in the figure):

[0158] S7: After obtaining the third voice information of the third target period, obtain the bandwidth usage data of the third target period; the third target period is after the first target period.

[0159] The third voice information of the third target period refers to the voice information sent by the object through the device in the third target period in the voice communication scenario; the bandwidth usage data of the third target period refers to the bandwidth-related data detected in the third target period corresponding to the voice communication in the voice communication scenario.

[0160] S8: Use the peak prediction model to predict the bandwidth peak of the fourth target period according to the bandwidth usage data of the third target period, and obtain the predicted bandwidth peak of the fourth target period; the fourth target period is after the third target period.

[0161] The predicted bandwidth peak of the fourth target period refers to the predicted value of the maximum bandwidth value of the fourth target period.

[0162] S9: If the third bandwidth difference between the predicted bandwidth peak of the fourth target period and the preset bandwidth threshold is greater than or equal to the first preset difference, adjust the voice quality of the third voice information to obtain the fourth voice information; the voice quality of the fourth voice information is higher than that of the third voice information.

[0163] Wherein, when the third bandwidth difference between the predicted bandwidth peak of the fourth target time period and the preset bandwidth threshold is the third bandwidth difference value between the preset bandwidth threshold and the predicted bandwidth peak of the fourth target time period, the third bandwidth difference being greater than or equal to the first preset difference means the third bandwidth difference value being greater than or equal to the first preset difference value; when the third bandwidth difference between the predicted bandwidth peak of the fourth target time period and the preset bandwidth threshold is the third bandwidth ratio of the bandwidth difference between the preset bandwidth threshold and the predicted bandwidth peak of the fourth target time period to the preset bandwidth threshold, the third bandwidth difference being greater than or equal to the first preset difference means the third bandwidth ratio being greater than or equal to the first preset ratio.

[0164] The S7 - S9 quickly, effectively, and accurately predicts the predicted bandwidth peak after another period of time through the bandwidth usage data and peak prediction model of another period of time. In the case where the predicted bandwidth peak is far from the preset bandwidth threshold, there is no need to reduce the bandwidth peak, improving the voice quality of the voice information in that another period of time and further enhancing the voice communication effect, thereby further enhancing the voice communication experience.

[0165] As an example of S7 - S9, based on the above S201 - S203 example, after the computer device continues to obtain the team voice information of the real - time time period, it continues to obtain the bandwidth usage data of the real - time time period; it continues to input the bandwidth usage data of the real - time time period into the peak prediction model to predict the bandwidth peak of the future time period, and obtains the predicted bandwidth peak PB of the future time period. pre ; when it is determined that the first bandwidth difference between PB pre and B thr is greater than or equal to the first preset difference, it continues to adjust the voice quality of the team voice information in the real - time time period to obtain the adjusted team voice information, so that the voice quality of the adjusted team voice information is higher than the voice quality of the team voice information.

[0166] In summary, referring to Figure 5 Figure 5 is a flowchart of another voice processing method provided by an embodiment of the present application. The method includes:

[0167] S501: After obtaining the real - time voice information of the real - time time period, obtain the real - time bandwidth usage data of the real - time time period.

[0168] S502: Use the peak prediction model to predict the bandwidth peak of the future time period according to the real - time bandwidth usage data to obtain the future bandwidth peak.

[0169] S503: Determine whether the first bandwidth difference between the future bandwidth peak and the preset bandwidth threshold is less than the first preset difference, and whether the second bandwidth difference between the real-time bandwidth usage and the future bandwidth peak is less than the second preset difference. If so, execute S504; if not, return to execute S501.

[0170] S504: Adjust the voice quality of the real-time voice information to obtain the adjusted real-time voice information; the voice quality of the adjusted voice information is lower than that of the real-time voice information.

[0171] S505: After continuing to obtain the real-time voice information in the real-time time period, continue to obtain the real-time bandwidth usage data in the real-time time period.

[0172] S506: Continue to predict the bandwidth peak in the future time period based on the real-time bandwidth usage data through the peak prediction model to obtain the future bandwidth peak.

[0173] S507: Continue to determine whether the third bandwidth difference between the future bandwidth peak and the preset bandwidth threshold is greater than the first preset difference. If so, execute S508; if not, return to execute S501.

[0174] S508: Adjust the voice quality of the real-time voice information to obtain the adjusted real-time voice information; the voice quality of the adjusted voice information is higher than that of the real-time voice information.

[0175] It can be seen from the above technical solutions that through the real-time bandwidth usage data and the peak prediction model, the future bandwidth peak can be predicted quickly, effectively, and accurately. When the future bandwidth peak is close to the preset bandwidth threshold and the real-time bandwidth usage is close to the future bandwidth peak, the voice quality of the real-time voice information is reduced, and the bandwidth peak is reduced quickly, effectively, and accurately to avoid network congestion and improve the voice communication effect, thereby enhancing the voice communication experience. In addition, after reducing the voice quality of the real-time voice information, when it is continuously detected that the future bandwidth peak is far from the preset bandwidth threshold, the voice quality of the real-time voice information is improved to further enhance the voice communication effect, thereby further enhancing the voice communication experience.

[0176] It should be noted that on the basis of the implementation manners provided in the above aspects of the present application, further combinations can be made to provide more implementation manners.

[0177] See Figure 6 , Figure 6 which is a schematic diagram of the voice communication scenario of a voice processing method provided by an embodiment of the present application, and see Figure 7 , Figure 7 which is a schematic diagram of the voice communication scenario of another voice processing method provided by an embodiment of the present application. The Figure 6 andFigure 7 The voice communication scenario is a multi-player voice communication scenario within different games; after obtaining real-time game voice information in the multi-player voice communication scenario, the future bandwidth peak is quickly, effectively, and accurately predicted through real-time bandwidth usage data and a peak prediction model. When the future bandwidth peak is close to the preset bandwidth threshold and the real-time bandwidth usage is close to the future bandwidth peak, the voice quality of the real-time game voice information of some players is reduced, and the bandwidth peak is quickly, effectively, and accurately reduced to avoid network congestion, improve the multi-player voice communication effect, and thus enhance the multi-player voice communication experience.

[0178] Based on Figure 2 For the voice processing method provided in the corresponding embodiment, an embodiment of the present application further provides a voice processing device. Refer to Figure 8 , Figure 8 It is a structural diagram of a voice processing device provided in an embodiment of the present application. The voice processing device 800 includes: an acquisition unit 801, a prediction unit 802, and an adjustment unit 803;

[0179] The acquisition unit 801 is configured to obtain the bandwidth usage data of the first target time period after obtaining the first voice information of the first target time period;

[0180] The prediction unit 802 is configured to predict the bandwidth peak of the second target time period through the peak prediction model according to the bandwidth usage data of the first target time period, and obtain the predicted bandwidth peak of the second target time period; the second target time period is after the first target time period;

[0181] The adjustment unit 803 is configured to adjust the voice quality of the first voice information to obtain the second voice information if the first bandwidth difference between the predicted bandwidth peak of the second target time period and the preset bandwidth threshold is less than the first preset difference, and the second bandwidth difference between the bandwidth usage amount in the bandwidth usage data of the first target time period and the predicted bandwidth peak of the second target time period is less than the second preset difference; the voice quality of the second voice information is lower than that of the first voice information.

[0182] In a possible implementation manner, the peak prediction model is obtained by training an initial prediction model according to the bandwidth usage data of the first historical time period and the corresponding historical bandwidth peak of the second historical time period; the second historical time period is after the first historical time period; the device further includes: a training unit;

[0183] The training unit is specifically configured to:

[0184] Predict the bandwidth peak of the second historical time period through the initial prediction model according to the bandwidth usage data of the first historical time period, and obtain the predicted bandwidth peak of the second historical time period;

[0185] Adjust the model parameters of the initial prediction model according to the predicted bandwidth peak in the second historical time period and the historical bandwidth peak in the second historical time period to obtain a peak prediction model.

[0186] In a possible implementation, the adjustment unit 803 is specifically configured to:

[0187] Determine the first coding complexity of the first voice information;

[0188] Adjust the voice quality of the first voice information according to the second coding complexity to obtain second voice information; the second coding complexity is lower than the first coding complexity.

[0189] In a possible implementation, the adjustment unit 803 is specifically configured to:

[0190] Determine the first encoder corresponding to the first coding complexity;

[0191] Adjust the voice quality of the first voice information according to the second encoder to obtain second voice information; the coding configuration of the second encoder is lower than the coding configuration of the first encoder.

[0192] In a possible implementation, the adjustment unit 803 is specifically configured to:

[0193] Determine the first coding bit rate corresponding to the first coding complexity;

[0194] Adjust the voice quality of the first voice information according to the second coding bit rate to obtain second voice information; the second coding bit rate is lower than the first coding bit rate.

[0195] In a possible implementation, the adjustment unit 803 is specifically configured to:

[0196] Determine the voice sub-information corresponding to the first voice complexity and the voice sub-information corresponding to the second voice complexity in the first voice information; the second voice complexity is lower than the first voice complexity;

[0197] Adjust the voice quality of the voice sub-information corresponding to the second voice complexity in the first voice information according to the second coding bit rate to obtain second voice information.

[0198] In a possible implementation, the adjustment unit 803 is specifically configured to:

[0199] Determine the voice sub-information corresponding to the first voice activity and the voice sub-information corresponding to the second voice activity in the first voice information; the second voice activity is lower than the first voice activity;

[0200] Adjust the voice quality of the voice sub - information corresponding to the second voice activity degree in the first voice information according to the second coding bit rate to obtain the second voice information.

[0201] In a possible implementation manner, the adjustment unit 803 is specifically configured to:

[0202] Adjust the voice quality of each voice sub - information in the first voice information according to the second coding bit rate to obtain the second voice information.

[0203] In a possible implementation manner, the device further includes: a determination unit;

[0204] The determination unit is specifically configured to:

[0205] Determine the first sampling rate and the first bit depth corresponding to the first coding bit rate;

[0206] Adjust the first sampling rate and the first bit depth to obtain a second sampling rate and a second bit depth; the second sampling rate and the second bit depth meet one or more of the conditions that the second sampling rate is lower than the first sampling rate, the second bit depth is lower than the first bit depth, and the second sampling rate is lower than the first sampling rate and the second bit depth is lower than the first bit depth;

[0207] Determine the second coding bit rate according to the second sampling rate and the second bit depth.

[0208] In a possible implementation manner, the adjustment unit 803 is specifically configured to:

[0209] Determine the first coding redundancy of the first voice information;

[0210] Adjust the voice quality of the first voice information according to the second coding redundancy to obtain the second voice information; the second coding redundancy is lower than the first coding redundancy.

[0211] In a possible implementation manner, the adjustment unit 803 is specifically configured to:

[0212] Determine the first error - correction coding level corresponding to the first coding redundancy;

[0213] Adjust the voice quality of the first voice information according to the second error - correction coding level to obtain the second voice information; the second error - correction coding level is lower than the first error - correction coding level.

[0214] In a possible implementation manner, the adjustment unit 803 is specifically configured to:

[0215] Determine the first redundancy data ratio corresponding to the first coding redundancy;

[0216] Adjust the voice quality of the first voice information according to the second redundancy data ratio to obtain the second voice information; the second redundancy data ratio is lower than the first redundancy data ratio.

[0217] In a possible implementation manner, the obtaining unit 801 is further configured to:

[0218] After obtaining the third voice information of the third target time period, obtain the bandwidth usage data of the third target time period; the third target time period is after the first target time period;

[0219] The prediction unit 802 is further configured to:

[0220] Predict the bandwidth peak value of the fourth target time period according to the bandwidth usage data of the third target time period through the peak prediction model to obtain the predicted bandwidth peak value of the fourth target time period; the fourth target time period is after the third target time period;

[0221] The adjustment unit 803 is further configured to:

[0222] If the third bandwidth difference between the predicted bandwidth peak value of the fourth target time period and the preset bandwidth threshold is greater than or equal to the first preset difference, adjust the voice quality of the third voice information to obtain the fourth voice information; the voice quality of the fourth voice information is higher than that of the third voice information.

[0223] It can be seen from the above technical solutions that the voice processing device includes an obtaining unit, a prediction unit, and an adjustment unit. Among them, the obtaining unit obtains the bandwidth usage data of the first target time period after obtaining the first voice information of the first target time period. That is, in the case of obtaining the voice information of the target time period, the bandwidth usage data of the target time period is obtained to pay attention to the bandwidth usage situation of the target time period.

[0224] The prediction unit inputs the bandwidth usage data of the first target time period into the peak prediction model to predict the bandwidth peak value of the second target time period after the first target time period, and obtains the predicted bandwidth peak value of the second target time period. That is, through the peak prediction model, the predicted bandwidth peak value of the time period after the target time period can be quickly, effectively, and accurately predicted based on the bandwidth usage data of the target time period.

[0225] When the adjustment unit determines that the first bandwidth difference between the predicted bandwidth peak value in the second target time period and the preset bandwidth threshold is less than the first preset difference, and the second bandwidth difference between the bandwidth usage amount in the bandwidth usage data of the first target time period and the predicted bandwidth peak value in the second target time period is less than the second preset difference, the voice quality of the first voice information is adjusted to obtain the second voice information, so that the voice quality of the second voice information is lower than that of the first voice information. That is, it is determined that the predicted bandwidth peak value in the time period after the target time period is close to the preset bandwidth threshold, and the bandwidth usage amount in the target time period is close to the predicted bandwidth peak value in the time period after the target time period. By reducing the voice quality of the voice information in the target time period, the bandwidth peak value in the time period after the target time period can be quickly, effectively, and accurately reduced to avoid network congestion.

[0226] Based on this, the device can quickly, effectively, and accurately predict the predicted bandwidth peak value after a period of time through the bandwidth usage data and the peak prediction model of this period of time. When the predicted bandwidth peak value is close to the preset bandwidth threshold and the bandwidth usage amount of this period of time is close to the predicted bandwidth peak value, the voice quality of the voice information of this period of time is reduced, and the bandwidth peak value is quickly, effectively, and accurately reduced to avoid network congestion and improve the voice communication effect, thereby enhancing the voice communication experience.

[0227] The embodiment of the present application also provides a computer device, which can be a server. Refer to Figure 9 , Figure 9 which is a structural diagram of a server provided by the embodiment of the present application. The server 900 may vary greatly due to configuration or performance differences, and may include one or more processors, such as the CPU 922, and a memory 932, and one or more storage media 930 (such as one or more mass storage devices) for storing application programs 942 or data 944. Among them, the memory 932 and the storage media 930 can be transient storage or persistent storage. The program stored in the storage media 930 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server. Further, the central processing unit 922 can be set to communicate with the storage media 930 and execute a series of instruction operations in the storage media 930 on the server 900.

[0228] The server 900 may further include one or more power supplies 926, one or more wired or wireless network interfaces 950, one or more input / output interfaces 958, and / or one or more operating systems 941, such as Windows Server TM , Mac OS X TM , Unix TM , LinuxTM , FreeBSD TM and so on.

[0229] In this embodiment, the central processing unit 922 in the server 900 can execute the methods provided in various alternative implementations of the foregoing embodiments.

[0230] The computer device provided in the embodiments of the present application may also be a terminal. Refer to Figure 10 , Figure 10 , which is a structural diagram of a terminal provided in the embodiments of the present application. Taking the terminal as a smart phone as an example, the smart phone includes: a Radio Frequency (RF) circuit 1010, a memory 1020, an input unit 1030, a display unit 1040, a sensor 1050, an audio circuit 1060, a Wireless Fidelity (WiFi) module 1070, a processor 1080, and a power supply 1090 and other components. The input unit 1030 may include a touch panel 1031 and other input devices 1032. The display unit 1040 may include a display panel 1041. The audio circuit 1060 may include a speaker 1061 and a microphone 1062. Those skilled in the art can understand that Figure 10 the structure of the smart phone shown in

[0231] does not limit the smart phone, and may include more or fewer components than shown in the figure, or combine some components, or have different component arrangements.

[0232] The processor 1080 is the control center of the smart phone, connecting various parts of the entire smart phone using various interfaces and circuits. By running or executing software programs and / or modules stored in the memory 1020, and by invoking the data stored in the memory 1020, it performs various functions of the smart phone and processes data. Optionally, the processor 1080 may include one or more processing units; preferably, the processor 1080 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor may not be integrated into the processor 1080 either.

[0233] In this embodiment, the processor 1080 in the smart phone can execute the methods provided in various optional implementation manners of the above embodiment.

[0234] According to one aspect of the present application, there is provided a computer-readable storage medium for storing a computer program. When the computer program runs on a computer device, the computer device is caused to execute the methods provided in various optional implementation manners of the above embodiment.

[0235] According to one aspect of the present application, there is provided a computer program product. The computer program product includes a computer program, and the computer program is stored in a computer-readable storage medium. The processor of the computer device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, causing the computer device to execute the methods provided in various optional implementation manners of the above embodiment.

[0236] The descriptions of the processes or structures corresponding to the above respective drawings each have their own focuses. For parts not detailed in a certain process or structure, reference may be made to the relevant descriptions of other processes or structures.

[0237] Terms such as "first" and "second" in the specification of the present application and the above drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0238] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections between each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.

[0239] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0240] In addition, in each embodiment of the present application, the functional units can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0241] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device to execute all or part of the steps of the methods described in each embodiment of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), RAM, magnetic disks, or optical discs that can store computer programs.

[0242] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of each embodiment of the present application.

Claims

1. A voice processing method, characterized in that, The method includes: After obtaining the first voice information of the first target time period, obtaining the bandwidth usage data of the first target time period; Predicting the bandwidth peak of the second target time period based on the bandwidth usage data of the first target time period through a peak prediction model to obtain the predicted bandwidth peak of the second target time period; the second target time period is after the first target time period; If the first bandwidth difference between the predicted bandwidth peak of the second target time period and the preset bandwidth threshold is less than the first preset difference, and the second bandwidth difference between the bandwidth usage amount in the bandwidth usage data of the first target time period and the predicted bandwidth peak of the second target time period is less than the second preset difference, adjusting the voice quality of the first voice information to obtain the second voice information; the voice quality of the second voice information is lower than that of the first voice information.

2. The method according to claim 1, characterized in that, The peak prediction model is obtained by training an initial prediction model based on the bandwidth usage data of the first historical time period and the corresponding historical bandwidth peak of the second historical time period, and the second historical time period is after the first historical time period; the training steps of the peak prediction model include: Predicting the bandwidth peak of the second historical time period based on the bandwidth usage data of the first historical time period through the initial prediction model to obtain the predicted bandwidth peak of the second historical time period; Adjusting the model parameters of the initial prediction model according to the predicted bandwidth peak of the second historical time period and the historical bandwidth peak of the second historical time period to obtain the peak prediction model.

3. The method according to claim 1, characterized in that, The adjusting the voice quality of the first voice information to obtain the second voice information includes: Determining the first coding complexity of the first voice information; Adjusting the voice quality of the first voice information according to the second coding complexity to obtain the second voice information; the second coding complexity is lower than the first coding complexity.

4. The method according to claim 3, characterized in that The adjusting the voice quality of the first voice information according to the second coding complexity to obtain the second voice information includes: Determining the first encoder corresponding to the first coding complexity; Adjusting the voice quality of the first voice information according to the second encoder to obtain the second voice information; the coding configuration of the second encoder is lower than that of the first encoder.

5. The method according to claim 3, characterized in that, The adjusting the voice quality of the first voice information according to the second coding complexity to obtain the second voice information includes: Determining the first coding bit rate corresponding to the first coding complexity; Adjusting the voice quality of the first voice information according to the second coding bit rate to obtain the second voice information; the second coding bit rate is lower than the first coding bit rate.

6. The method according to claim 5, characterized in that The adjusting the voice quality of the first voice information according to the second coding bit rate to obtain the second voice information includes: Determining the voice sub-information corresponding to the first voice complexity and the voice sub-information corresponding to the second voice complexity in the first voice information; the second voice complexity is lower than the first voice complexity; Adjust the voice quality of the voice sub - information corresponding to the second voice complexity in the first voice information according to the second coding bit rate to obtain the second voice information.

7. The method according to claim 5, characterized in that The step of adjusting the voice quality of the first voice information according to the second coding bit rate to obtain the second voice information includes: Determine the voice sub - information corresponding to the first voice activity and the voice sub - information corresponding to the second voice activity in the first voice information; the second voice activity is lower than the first voice activity. Adjust the voice quality of the voice sub - information corresponding to the second voice activity in the first voice information according to the second coding bit rate to obtain the second voice information.

8. The method according to claim 5, wherein The step of adjusting the voice quality of the first voice information according to the second coding bit rate to obtain the second voice information includes: Adjust the voice quality of each voice sub - information in the first voice information according to the second coding bit rate to obtain the second voice information.

9. The method according to any one of claims 5 - 8, characterized in that The step of determining the second coding bit rate includes: Determine the first sampling rate and the first bit depth corresponding to the first coding bit rate. Adjust the first sampling rate and the first bit depth to obtain a second sampling rate and a second bit depth; the second sampling rate and the second bit depth meet one or more of the following conditions: the second sampling rate is lower than the first sampling rate, the second bit depth is lower than the first bit depth, and the second sampling rate is lower than the first sampling rate and the second bit depth is lower than the first bit depth. Determine the second coding bit rate according to the second sampling rate and the second bit depth.

10. The method according to claim 1, wherein The step of adjusting the voice quality of the first voice information to obtain the second voice information includes: Determine the first coding redundancy of the first voice information. Adjust the voice quality of the first voice information according to the second coding redundancy to obtain the second voice information; the second coding redundancy is lower than the first coding redundancy.

11. The method according to claim 10, wherein The step of adjusting the voice quality of the first voice information according to the second coding redundancy to obtain the second voice information includes: Determine the first error - correction coding level corresponding to the first coding redundancy. Adjust the voice quality of the first voice information according to the second error - correction coding level to obtain the second voice information; the second error - correction coding level is lower than the first error - correction coding level.

12. The method according to claim 10, wherein The step of adjusting the voice quality of the first voice information according to the second coding redundancy to obtain the second voice information includes: Determine the first redundancy data ratio corresponding to the first coding redundancy. Adjust the voice quality of the first voice information according to the second redundancy data ratio to obtain the second voice information; the second redundancy data ratio is lower than the first redundancy data ratio.

13. The method according to claim 1, characterized in that, The method further includes: After obtaining the third voice information in the third target time period, obtain the bandwidth usage data in the third target time period; the third target time period is after the first target time period. Predict the bandwidth peak of a fourth target time period based on the bandwidth usage data of the third target time period through the peak prediction model, and obtain the predicted bandwidth peak of the fourth target time period; the fourth target time period is after the third target time period; If the third bandwidth difference between the predicted bandwidth peak of the fourth target time period and the preset bandwidth threshold is greater than or equal to the first preset difference, adjust the voice quality of the third voice information to obtain fourth voice information; the voice quality of the fourth voice information is higher than that of the third voice information.

14. A voice processing device, characterized in that, The device includes: an acquisition unit, a prediction unit, and an adjustment unit; The acquisition unit is configured to acquire the bandwidth usage data of the first target time period after acquiring the first voice information of the first target time period; The prediction unit is configured to predict the bandwidth peak of a second target time period based on the bandwidth usage data of the first target time period through a peak prediction model, and obtain the predicted bandwidth peak of the second target time period; the second target time period is after the first target time period; The adjustment unit is configured to, if the first bandwidth difference between the predicted bandwidth peak of the second target time period and the preset bandwidth threshold is less than the first preset difference, and the second bandwidth difference between the bandwidth usage amount in the bandwidth usage data of the first target time period and the predicted bandwidth peak of the second target time period is less than the second preset difference, adjust the voice quality of the first voice information to obtain second voice information; the voice quality of the second voice information is lower than that of the first voice information.

15. A computer device, characterized in that, The computer device includes a processor and a memory: The memory is used to store a computer program and transmit the computer program to the processor; The processor is configured to execute the method according to any one of claims 1-13 based on the instructions in the computer program.

16. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store a computer program, and when the computer program runs on a computer device, the computer device is caused to execute the method according to any one of claims 1-13.

17. A computer program product, comprising a computer program, characterized in that, When the computer program runs on a computer device, the computer device is caused to execute the method according to any one of claims 1-13.