Audio signal processing method and device, electronic equipment and storage medium

By calculating and adjusting the audio signal energy difference between the target remote end and other remote ends in multi-terminal voice calls, the problem of cumbersome mute operation for users in multi-terminal voice calls is solved, and the convenience and accuracy of hearing specific users' speech are improved.

CN120977324APending Publication Date: 2025-11-18GUANGZHOU TENCENT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410617259.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-05-17
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

In multi-terminal voice call scenarios, users need to frequently mute in order to hear the speech of a specific user, which is cumbersome and prone to errors.

Method used

By calculating the difference in audio signal energy between the target call remote end and other call remote ends, the audio signal energy of the target call remote end and other call remote ends is adjusted so that the audio signal energy value of the target call remote end is greater than that of other call remote ends. After mixing processing, it is played at the near end of the call to reduce the interference of audio signals from other call remote ends.

Benefits of technology

It reduces the need for users to mute other remote devices during multi-device voice calls, improves the convenience of clearly hearing specific users' speech, and reduces the probability of accidental operation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120977324A_ABST
    Figure CN120977324A_ABST
Patent Text Reader

Abstract

The invention discloses an audio signal processing method and device, electronic equipment and a storage medium. The embodiment of the invention can be applied to various scenes including but not limited to cloud technology and the like. The method comprises the following steps: acquiring respective audio signals of a plurality of call remote ends; subtracting the energy value of the audio signal of the target call far end from the energy value of the audio signal of each other call far end to obtain an energy difference value; according to the energy difference value, performing energy adjustment on at least one of the audio signals of each call far-end to obtain a first audio signal of the target call far-end and second audio signals of each other call far-end; performing stream mixing processing on the first audio signal of the target call far end and the second audio signals of the other call far ends to obtain a first target audio signal; and playing the first target audio signal at the call near end. According to the method provided by the invention, the user operation is less, and the probability of misoperation is also lower.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of electronic information, and more particularly, to an audio signal processing method and device, an electronic device, and a storage medium. BACKGROUND

[0002] With the development of audio technology, voice calls can be made between different terminals (or clients running on the terminals). In a multi-terminal voice call scenario, there are often multiple users of the call far ends speaking, and in order to hear the speech of a specific user clearly, the users of the other call far ends except the specific user need to be muted. If the number of the call far ends participating in the voice call is large, the user needs to trigger the mute operation multiple times, which is complicated and prone to misoperation. SUMMARY

[0003] Therefore, the embodiments of the present application provide an audio signal processing method and device, an electronic device, and a storage medium.

[0004] In a first aspect, the embodiments of the present application provide an audio signal processing method, which includes: obtaining audio signals of multiple call far ends in a target voice call in which a call near end participates; the multiple call far ends include a target call far end selected for the call near end; subtracting an energy value of the audio signal of the target call far end from energy values of the audio signals of the other call far ends to obtain energy difference values between the target call far end and the other call far ends, the other call far end being a call far end other than the target call far end in the multiple call far ends; performing energy adjustment on at least one of the audio signal of the target call far end and the audio signals of the other call far ends according to the energy difference values between the target call far end and the other call far ends to obtain a first audio signal of the target call far end and second audio signals of the other call far ends; the energy value of the first audio signal is greater than the energy values of the second audio signals; performing mixed stream processing on the first audio signal of the target call far end and the second audio signals of the other call far ends to obtain a first target audio signal; and playing the first target audio signal at the call near end.

[0005] In a second aspect, an audio signal processing apparatus is provided. The apparatus comprises: an obtaining module configured to obtain audio signals of a plurality of far-end terminals in a target voice call in which a near-end terminal is involved; the plurality of far-end terminals comprises a target far-end terminal selected for the near-end terminal; a calculating module configured to subtract an energy value of the audio signal of the target far-end terminal from energy values of the audio signals of each of the other far-end terminals to obtain energy difference values between the target far-end terminal and each of the other far-end terminals, the other far-end terminals being the far-end terminals in the plurality of far-end terminals except the target far-end terminal; an adjusting module configured to perform energy adjustment on at least one of the audio signal of the target far-end terminal and the audio signals of the other far-end terminals according to the energy difference values between the target far-end terminal and each of the other far-end terminals to obtain a first audio signal of the target far-end terminal and second audio signals of each of the other far-end terminals; the energy value of the first audio signal is greater than the energy values of the second audio signals; a mixing module configured to perform mixing processing on the first audio signal of the target far-end terminal and the second audio signals of each of the other far-end terminals to obtain a first target audio signal; and a playing module configured to play the first target audio signal at the near-end terminal.

[0006] Optionally, the adjusting module is further configured to, if it is determined according to the energy difference values between the target far-end terminal and each of the other far-end terminals that there is a first far-end terminal in the other far-end terminals corresponding to an energy difference value less than an energy threshold value, obtain a first expected energy value of the target far-end terminal and a second expected energy value of the first far-end terminal; the first expected energy value is greater than the second expected energy value; adjust the energy value of the audio signal of the target far-end terminal to the first expected energy value to obtain the first audio signal of the target far-end terminal; and adjust the energy value of the audio signal of the first far-end terminal to the second expected energy value to obtain a second audio signal of the first far-end terminal.

[0007] Optionally, the adjusting module is further configured to obtain the energy value of the audio signal of the target far-end terminal as the first expected energy value of the target far-end terminal; calculate a difference between the energy value of the audio signal of the target far-end terminal and the energy threshold value to obtain a first energy value; and determine the second expected energy value of the first far-end terminal based on the first energy value, the second expected energy value being not greater than the first energy value.

[0008] Optionally, the adjusting module is further configured to obtain the energy value of the audio signal of the first far-end terminal as the second expected energy value of the first far-end terminal; calculate a sum of the energy threshold value and the energy value of the audio signal of the second far-end terminal to obtain a second energy value; the second far-end terminal is the first far-end terminal with the highest energy value of the audio signal; and determine the first expected energy value of the target far-end terminal based on the second energy value, the first expected energy value being not less than the second energy value.

[0009] Optionally, the adjusting module is further configured to calculate a sum of the energy value of the audio signal of the target far-end and a third energy value to obtain a first expected energy value of the target far-end, and calculate a difference between the energy value of the audio signal of the second far-end and a fourth energy value to obtain a second expected energy value of the first far-end, wherein the second far-end is the first far-end with the highest energy value of the audio signal.

[0010] Optionally, the adjusting module is further configured to adjust the energy value of the target audio segment containing the speech signal in the audio signal of the target far-end to the first expected energy value, and keep the energy values of other audio segments in the audio signal of the target far-end unchanged except the target audio segment to obtain the first audio signal of the target far-end.

[0011] Optionally, the adjusting module is further configured to, if it is determined according to the energy difference values between the target far-end and each of the other far-ends that there is a third far-end in the other far-ends with a corresponding energy difference value not less than the energy threshold, obtain the audio signal of the third far-end as the second audio signal of the third far-end.

[0012] Optionally, the playing module is further configured to, if the energy difference values between the energy value of the audio signal of the target far-end and the energy values of the audio signals of each of the other far-ends are all not less than the energy threshold, play the audio signal of the target far-end as the first audio signal of the target far-end, and play the audio signals of each of the other far-ends as the corresponding second audio signals.

[0013] Optionally, the calculating module is further configured to perform speech detection on the audio signal of the target far-end to obtain a speech detection result, and perform the step of subtracting the energy value of the audio signal of the target far-end from the energy values of the audio signals of each of the other far-ends if the speech detection result indicates that the audio signal of the target far-end includes a speech signal.

[0014] Optionally, the playing module is further configured to, if the speech detection result indicates that the audio signal of the target far-end does not include a speech signal, perform stream mixing processing on the audio signal of the target far-end and the audio signals of the other far-ends to obtain a second target audio signal, and play the second target audio signal at the near-end.

[0015] Optionally, the apparatus further comprises a selecting module configured to, in response to a selection operation triggered by an object at the near-end side on the multiple far-ends, select a far-end selected by the selection operation as the target far-end.

[0016] In a third aspect, an electronic device is provided, including a processor and a memory. The memory stores computer readable instructions. When the computer readable instructions are executed by the processor, the method described above is implemented.

[0017] In a fourth aspect, a computer readable storage medium is provided, which stores computer readable instructions. When the computer readable instructions are executed by a processor, the method described above is implemented.

[0018] In a fifth aspect, a computer program product is provided, which includes computer readable instructions. When the computer readable instructions are executed by a processor, the method described above is implemented.

[0019] The audio signal processing method and device, electronic device and storage medium provided in the embodiments of the present application. In the present application, the energy difference between the energy value of the audio signal of the target call far end selected by the call near end from the plurality of call far ends of the target voice call and the energy value of the audio signal of each other call far end is used to adjust the energy of at least one of the audio signal of the target call far end and the audio signal of the other call far end, to obtain the first audio signal of the target call far end and the second audio signal of each other call far end. Since the energy value of the first audio signal of the target call far end is greater than the energy value of the second audio signal of the other call far end, in the process of playing the first target audio signal obtained by mixing the first audio signal and the second audio signal at the call near end, the audio signal of the other call far end interferes with the user at the call near end to hear the audio signal of the target call far end, ensuring that the first audio signal of the target call far end is more easily heard and distinguished by the user at the call near end. Moreover, even if there are more call far ends relative to the call near end, since the user at the call near end specifies the call far end as the target call far end, there is no need to mute each of the other call far ends except the call far end of interest, the user operation is less, and the probability of misoperation is also lower. BRIEF DESCRIPTION OF DRAWINGS

[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creating any creative labor.

[0021] Figure 1 A schematic diagram suitable for the application scenario to which the embodiments of the present application are applied is shown;

[0022] Figure 2 A flowchart of an audio signal processing method according to an embodiment of the present application is shown;

[0023] Figure 3 Fig. 1 shows a flowchart of steps in an embodiment of the method according to the present application; Figure 2 Fig. 2 shows a flowchart of steps in an embodiment of the method according to the present application;

[0024] Figure 4 Fig. 3 shows a schematic diagram of a process of detecting a voice signal in an embodiment of the present application;

[0025] Figure 5 Fig. 4 shows a flowchart of steps in an embodiment of the method according to the present application; Figure 2 Fig. 5 shows a flowchart of steps in an embodiment of the method according to the present application;

[0026] Figure 6 Fig. 6 shows a schematic diagram of a process of processing an audio signal in an embodiment of the present application;

[0027] Figure 7 Fig. 7 shows a schematic diagram of a process of voice communication in an embodiment of the present application;

[0028] Figure 8 Fig. 8 shows a schematic diagram of a process of energy adjustment in an embodiment of the present application;

[0029] Figure 9 Fig. 9 shows a block diagram of an audio signal processing apparatus in an embodiment of the present application;

[0030] Figure 10 Fig. 10 shows a structural block diagram of an electronic device for performing the audio signal processing method according to an embodiment of the present application. DETAILED DESCRIPTION

[0031] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, but not all of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of the present application.

[0032] In the following description, the terms "first\second" are only to distinguish similar objects, and do not represent a specific order of the objects. It can be understood that "first\second" can be interchanged in a specific order or sequence as allowed, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0033] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used in the description herein is for describing particular embodiments only and is not intended to be limiting of the application. It is to be understood that the use of "a", "an", or "the" herein includes singular and plural referents unless the context clearly dictates otherwise. The use of "multiple", "plurality", or "a plurality" herein refers to two or more. The use of "and / or" in describing a relationship between two or more objects refers to the relationship that three conditions exist, for example, A and / or B can mean A exists alone, A and B exist together, and B exists alone. The character " / " generally represents an "or" relationship between the objects before and after the " / ".

[0034] The application discloses an audio signal processing method and device, electronic equipment and storage medium, and relates to cloud technology.

[0035] Cloud technology refers to a hosting technology that unifies a series of resources such as hardware, software, network, etc. in a wide area network or a local area network to realize data calculation, storage, processing and sharing.

[0036] Cloud technology is a general term for network technology, information technology, integration technology, management platform technology, application technology, etc. based on the cloud computing business model application, which can form a resource pool, and be used on demand, flexibly and conveniently. Cloud computing technology will become an important support. The background service of a technical network system needs a large amount of calculation and storage resources, such as a video website, a picture website and more portals. With the high development and application of the Internet industry, every item may have its own identification mark in the future, and needs to be transmitted to the background system for logical processing. Different levels of data will be processed separately, and various industry data need strong system support, which can only be realized through cloud computing.

[0037] A cloud call center is a call center system based on cloud computing technology. An enterprise does not need to purchase any software and hardware system, but only needs to have basic conditions such as personnel and site to quickly have its own call center. The software and hardware platform, communication resources, daily maintenance and service are provided by the server provider. The cloud call center has many characteristics such as short construction period, low investment, low risk, flexible deployment, strong system capacity scalability and low operation and maintenance cost. Whether it is a telephone marketing center or a customer service center, an enterprise only needs to rent services as needed to establish a call center system that is comprehensive, stable, reliable, has distributed seats nationwide and has nationwide call access.

[0038] Cloud conference is a kind of efficient, convenient and low-cost conference form based on cloud computing technology. Users only need to operate through an Internet interface, and can quickly and efficiently share voice, data files and video with teams and customers around the world. The complex technology of data transmission and processing in the conference is operated by the cloud conference service provider.

[0039] At present, the domestic cloud conference mainly focuses on the service content based on the SaaS (Software as a Service) mode, including telephone, network, video and other service forms. The video conference based on cloud computing is called cloud conference.

[0040] In the cloud conference era, data transmission, processing and storage are all handled by the computer resources of the video conference manufacturer. Users no longer need to purchase expensive hardware and install complicated software. They only need to open a browser and log in to the corresponding interface to conduct efficient remote conferences.

[0041] The cloud conference system supports multi-server dynamic cluster deployment and provides multiple high-performance servers, greatly improving the stability, security and availability of the conference. In recent years, video conferences have been widely used in government, military, transportation, finance, operators, education and enterprises due to their ability to greatly improve communication efficiency, continuously reduce communication costs and upgrade internal management levels. There is no doubt that video conferences using cloud computing will have stronger appeal in convenience, speed and ease of use, and will stimulate a new high tide of video conference applications.

[0042] The scheme of the present application can be applied to voice calls in a cloud conference scenario.

[0043] Reference Figure 1 , Figure 1 A schematic diagram suitable for the application scenario to which the embodiments of the present application are applicable is shown. The terminal device 400 is connected to the server 200 through the network 300, wherein the network 300 can be a wide area network or a local area network, or a combination of the two.

[0044] In some embodiments, at least two terminal devices 400 participate in a target voice call (one of the terminal devices 400 can initiate a voice call request, and at least one other terminal device 400 can accept the voice call request initiated by the terminal device 400 to create a target voice call). Each terminal device 400 can act as a near-end of the call. When a terminal device 400 acts as a near-end of the call, the other terminal devices 400 participating in the target voice call act as the corresponding far-ends of the call for the terminal device 400. That is, the far-end and the near-end of the call are relative, and the same terminal device 400 can act as a far-end and a near-end of the call.

[0045] In an embodiment, for any one terminal device 400 participating in the target voice call, the audio signal can be collected and sent to the server 200, and the server 200 obtains the mixed first target audio signal according to the audio signal of the target call far end selected by each terminal device 400 and the audio signal of the other call far end corresponding to the terminal device 400, and sends the mixed first target audio signal to the terminal device 400 as the call near end, so that the terminal device 400 plays the mixed first target audio signal.

[0046] In another embodiment, for any one terminal device 400 participating in the target voice call, the audio signal can be collected and sent to the server 200, and the server 200 sends the audio signal of the call far end corresponding to each terminal device 400 to the terminal device 400 as the call near end, and the terminal device 400 as the call near end obtains the mixed first target audio signal according to the audio signal of the target call far end selected and the audio signal of the corresponding other call far end, and plays the mixed first target audio signal.

[0047] In some embodiments, the server 200 can be a standalone physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content distribution networks (CDN), and big data and artificial intelligence platforms, etc. Basic cloud computing services. The terminal device 400 can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart voice interaction device, a smart home, a vehicle-mounted terminal, an aircraft, etc., but is not limited thereto. The terminal device and the server can be connected directly or indirectly through wired or wireless communication, and the present application is not limited in this embodiment.

[0048] For convenience of description, the following embodiments are described by taking the audio signal processing method executed by the electronic device as an example.

[0049] Please refer to Figure 2 , Figure 2 The flowchart of an audio signal processing method according to an embodiment of the present application is shown, which can be used in an electronic device, which can be a terminal device 400 in Figure 1 , or a terminal 400 and a server 200 interacting to implement the method of the present application, which can include:

[0050] S110, acquire audio signals of multiple call far ends in a target voice call in which the call near end participates.

[0051] The multiple call far ends include a target call far end selected for the call near end.

[0052] In the present application, a voice call in which a device currently serving as a call near end participates is referred to as a target voice call. The target voice call can be a pure voice call (for example, a voice call in an instant messaging application, a voice call in a live broadcast application, a voice call in a game application, a voice call in a conference application, etc.), or a voice call in a video call, etc. Any terminal device participating in the target voice call can serve as a call near end, and other terminal devices participating in the target voice call except the call near end serve as call far ends corresponding to the call near end.

[0053] It can be understood that in the same voice call, the call near end is different, and correspondingly, the call far end relative to the call near end is also different. For example, if users participating in the target voice call include user A, user B, and user C, on the side of user A, the terminal device (or client, hereinafter described by taking a terminal as an example) in which user A is located can serve as a call near end, and correspondingly, the terminal device in which user B is located and the terminal device in which user C is located serve as call far ends relative to the terminal device in which user A is located; on the side of user B, the terminal device in which user B is located can serve as a call near end, and correspondingly, the terminal device in which user A is located and the terminal device in which user C is located serve as call far ends relative to the terminal device in which user B is located; on the side of user C, the terminal device in which user C is located can serve as a call near end, and correspondingly, the terminal device in which user A is located and the terminal device in which user B is located serve as call far ends relative to the terminal device in which user C is located.

[0054] Therefore, in the target voice call, the client in which each user participating in the target voice call is located can serve as a call near end in the present application, and the audio signal of the call far end received by the call near end is processed according to the method of the present application.

[0055] The multiple terminal devices participating in the target voice call can be installed with an application (for example, the instant messaging application, the live broadcast application, the conference application, the game application, etc. listed above) having a voice call function, and each terminal device running the application can log in with a registered account to indicate the object (i.e., the user) participating in the target voice call through the logged-in account. Any terminal device running the application having a voice call function can initiate a voice call request, and other terminal devices running the application having a voice call function can accept the voice call request after receiving the voice call request to establish the target voice call between the multiple terminal devices. The application having a voice call function can be a desktop application, a mini-program application, a webpage application, etc.

[0056] For example, the application having a voice call function is a video conference application, the terminal device d1 installed with the video conference application sends a voice call request to the terminal devices d2 and d3 installed with the video conference application, the d2 and d3 accept the voice call request, and a video conference is created, and the voice interaction in the video conference is the target voice call. When the d1 is the call near end, the d2 and d3 are the call far ends relative to the d1.

[0057] The target call far end can be one selected by the user on the call near end side from the multiple call far ends participating in the target voice call, and the target call far end selected by the user on the call near end side refers to the call far end focused on by the user on the call near end side in the target voice call, that is, the user on the call near end side expects to hear the voice of the user on the target call far end side clearly during the target voice call. The selection of the target call far end can include: in response to a selection operation triggered by the object on the call near end side on the multiple call far ends, selecting the call far end selected by the selection operation as the target call far end. The object refers to the user on the call near end side.

[0058] It is worth mentioning that the object on the call near end side can select a call far end as the target call far end when initiating a voice call request or accepting a voice call request, or during a voice call, that is, the object can select the corresponding call far end as the target call far end in the voice call initiation interface of the application, or the voice call acceptance interface of the application, or the voice call interface.

[0059] After the target call far end is selected by the call near end, the audio signal processing can be performed on the call near end in the manner of the present application.

[0060] It is worth mentioning that, in the target voice call process, the call near-end can also change the target call far-end in response to a far-end changing operation for the target voice call, and perform the audio signal processing method of the present application according to the changed target call far-end.

[0061] The audio signal of the call far-end refers to the sound signal collected by the call far-end in the target voice call process. The audio signal of the call far-end can or can not include speech. For example, the object on the far-end side is speaking, and the audio signal of the call far-end includes speech. The object on the far-end side is not speaking, and the audio signal of the call far-end does not include speech.

[0062] In S120, the energy value of the audio signal of the target call far-end is subtracted from the energy value of the audio signal of each other call far-end to obtain an energy difference value between the target call far-end and each other call far-end.

[0063] In the present application, the other call far-end refers to the call far-end other than the target call far-end in the plurality of call far-ends.

[0064] The other call far-end can be one or more. If the other call far-end is one, the energy value of the audio signal of the target call far-end is subtracted from the energy value of the audio signal of the other call far-end to obtain one energy difference value. If the other call far-end is more than one, the energy value of the audio signal of the target call far-end is subtracted from the energy value of the audio signal of each other call far-end to obtain a plurality of energy difference values, one other call far-end corresponding to one energy difference value.

[0065] The energy value of the audio signal refers to a value representing the energy of the audio signal. The energy value of the audio signal collected by the terminal device is generally an integer of 2 bytes, the maximum value of the energy value is 32767, and the minimum value is 0. The greater the energy value of the audio signal, the greater the energy of the audio signal, and the easier the audio signal is to be heard by the user. When playing audio signals with different energies at the same volume, the greater the energy of the audio signal, the easier the audio signal is to be heard and distinguished by the user. Conversely, the smaller the energy value of the audio signal, the less likely the audio signal is to be heard and distinguished by the user.

[0066] In some embodiments, the square of the signal amplitude of the audio signal can be integrated to obtain an energy value of the audio signal. In other embodiments, the audio signal can be converted to a frequency domain signal, and the square of the amplitude of the frequency domain signal can be integrated to obtain an energy value of the audio signal. In yet other embodiments, the audio signal can be converted to a 16-bit value, and the maximum value among 16 values involved in the 16-bit value can be obtained as the energy value of the audio signal, i.e., the maximum value among 16 values involved in the 16-bit value after the audio signal is converted to the 16-bit value is the energy value.

[0067] The energy values of the audio signals of the target far-end and the other far-ends are determined according to the foregoing energy value determination manner, and the energy difference values between the target far-end and the other far-ends are calculated by calculating the difference between the energy value of the audio signal of the target far-end and the energy value of the audio signal of each of the other far-ends. It can be understood that if the energy difference value between the target far-end and one of the other far-ends is positive, it indicates that the energy value of the audio signal of the other far-end is less than the energy value of the audio signal of the target far-end; if the energy difference value is zero, it indicates that the energy value of the audio signal of the other far-end is equal to the energy value of the audio signal of the target far-end; and if the energy difference value is negative, it indicates that the energy value of the audio signal of the other far-end is greater than the energy value of the audio signal of the target far-end.

[0068] The energy difference values between the target far-end and the other far-ends are determined, the positive or negative of the energy difference values is used to indicate the size between the energy value of the audio signal of the target far-end and the energy value of the audio signal of each of the other far-ends, and the absolute value of the energy difference value is used to indicate the size of the gap between the energy value of the audio signal of the target far-end and the energy value of the audio signal of each of the other far-ends, so as to subsequently determine the energy adjustment direction of the audio signal of the target far-end and the energy adjustment direction of the audio signal of the other far-ends according to the energy difference value.

[0069] In S130, at least one of the audio signal of the target far-end and the audio signal of the other far-ends is energy-adjusted according to the energy difference value between the target far-end and the other far-ends, to obtain a first audio signal of the target far-end and a second audio signal of each of the other far-ends.

[0070] In the first audio signal, the energy value is greater than the energy value of the second audio signal.

[0071] In this application, for the sake of distinction, the audio signal of the target far-end obtained after energy adjustment is referred to as the first audio signal, and the audio signal of each of the other far-ends obtained after energy adjustment is referred to as the second audio signal.

[0072] After obtaining the energy difference between the target far-end and each other far-end, at least one of the audio signal of the target far-end and the audio signal of the other far-end is adjusted in energy according to the energy difference between the target far-end and each other far-end, so that the energy value of the first audio signal obtained after adjustment is higher than the energy value of the second audio signal, so that after mixing, the first audio signal is more easily heard by the user on the near-end side, and the user on the near-end side is more easily able to hear the voice of the user on the target far-end in the target voice call.

[0073] If the energy difference between the target far-end and one other far-end is greater than zero and the absolute value of the energy difference is large, it indicates that the energy value of the audio signal of the target far-end is relatively large with respect to the energy value of the audio signal of the other far-end, and if no energy adjustment is performed, the audio signal of the other far-end will not affect the user on the near-end side to clearly hear the audio signal of the target far-end. Therefore, in this case, no energy adjustment is needed for the audio signal of the target far-end and the audio signal of the other far-end.

[0074] If the energy difference between the target far-end and one other far-end is not less than zero but the absolute value of the energy difference is small, although the energy of the audio signal of the target far-end is higher than the energy of the audio signal of the other far-end, the difference between the energy of the audio signal of the target far-end and the energy of the audio signal of the other far-end is small, and if no energy adjustment is performed, the audio signal of the other far-end can cause the user on the near-end side to not clearly hear the audio signal of the target far-end.

[0075] If the energy difference between the target far-end and one other far-end is less than zero, it indicates that the energy of the audio signal of the target far-end is lower than the energy of the audio signal of the other far-end, and if no energy adjustment is performed, the audio signal of the other far-end will cause the user on the near-end side to not clearly hear the audio signal of the target far-end.

[0076] Based on this, an energy threshold value can be set, which is a positive number, and the energy threshold value ensures that the user can accurately distinguish different audio signals, i.e., the energy threshold value is at least greater than the minimum energy difference of the audio signals that the user can distinguish.

[0077] If the energy difference between the target talkspurt and another talkspurt is not less than the energy threshold, it can be determined that energy adjustment is not needed for the audio signal of the target talkspurt and the audio signal of the other talkspurt; if the energy difference between the target talkspurt and another talkspurt is not less than the energy threshold, it can be determined that energy adjustment is needed for at least one of the audio signal of the target talkspurt and the audio signal of the other talkspurt (e.g., the audio signal of the target talkspurt is enhanced, and / or the audio signal of the other talkspurt is suppressed), so that the first audio signal of the target talkspurt is higher than the second audio signal of the other talkspurt after adjustment; and the energy difference between the energy value of the first audio signal and the energy value of the second audio signal can be further ensured to be not less than the energy threshold.

[0078] In some embodiments, the other talkspurt corresponding to the energy difference less than the energy threshold can be determined as the first talkspurt, and then energy adjustment is performed on at least one of the audio signal of the target talkspurt and the audio signal of the first talkspurt, so that the energy value of the first audio signal of the target talkspurt is greater than the energy value of the second audio signal of the first talkspurt after adjustment. In some embodiments, the energy value of the audio signal of the target talkspurt can be increased to obtain the first audio signal of the target talkspurt, and the energy value of the audio signal of the first talkspurt can be kept unchanged, and the audio signal of the first talkspurt is taken as the second audio signal of the first talkspurt. In this case, if there is another talkspurt corresponding to the energy difference not less than the energy threshold (referred to as the third talkspurt) in the other talkspurts, energy adjustment can not be needed for the audio signal of the third talkspurt.

[0079] For example, the first target talkspurt with the maximum energy value of the audio signal in the first talkspurt is determined, the first adjustment energy value greater than the energy value of the audio signal of the first target talkspurt is determined based on the energy value of the audio signal of the first target talkspurt. Then, the energy value of the audio signal of the target talkspurt is adjusted to the first adjustment energy value, the adjusted audio signal is taken as the first audio signal of the target talkspurt, and the audio signal of each first talkspurt is taken as the second audio signal of each first talkspurt.

[0080] In some embodiments, the energy value of the audio signal of the first far-end call can be reduced to obtain a reduced energy value audio signal as the second audio signal of the first far-end call, and the audio signal of the target far-end call can be obtained as the first audio signal of the target far-end call. The reduction of the energy value of the audio signal of the first far-end call is not limited in the present application, as long as the energy value of the second audio signal of the first far-end call is lower than that of the first audio signal. Further, the difference between the energy value of the first audio signal of the target far-end call and the energy value of the second audio signal of the first far-end call can be constrained to be not less than an energy threshold.

[0081] In some other embodiments, the audio signal of the target far-end call can be energy enhanced (i.e. the energy value is increased) to obtain the first audio signal of the target far-end call, and the energy value of the audio signal of the first far-end call can be reduced to obtain the second audio signal of the first far-end call, so that the energy value of the adjusted first audio signal of the target far-end call is greater than that of the second audio signal of the first far-end call. Similarly, the difference between the energy value of the first audio signal of the target far-end call and the energy value of the second audio signal of the first far-end call can be constrained to be not less than an energy threshold. If there are multiple first far-end calls, the energy values of the second audio signals of the multiple first far-end calls can be the same or different, which is not limited herein.

[0082] It can be understood that if the energy value of the audio signal of the target far-end call is high, energy enhancement of the audio signal of the target far-end call will result in that the first audio signal heard by the user at the near-end call side is too strong, which exceeds the acceptable audio signal strength of the user and affects the auditory experience of the user. If the energy value of the audio signal of the target far-end call is low, energy suppression of the audio signal of the first far-end call while keeping the energy value of the audio signal of the target far-end call unchanged will result in that the first audio signal heard by the user at the near-end call side is more obvious, but the user may not be able to hear the content of the audio signal of the target far-end call clearly due to the low energy value of the audio signal of the target far-end call.

[0083] Therefore, based on the above considerations, in some embodiments, the manner of adjustment can also be determined with reference to the energy value of the audio signal of the target far-end talker and the energy value of the audio signal of the first far-end talker. The candidate manners of adjustment include the three manners mentioned above, i.e., manner 1 - increasing the energy value of the audio signal of the target far-end talker (i.e., performing energy enhancement on the audio signal of the target far-end talker), keeping the energy value of the audio signal of the first far-end talker unchanged; manner 2 - decreasing the energy value of the audio signal of the first far-end talker (i.e., performing energy suppression on the audio signal of the first far-end talker), keeping the energy value of the audio signal of the target far-end talker unchanged; and manner 3 - increasing the energy value of the audio signal of the target far-end talker and decreasing the energy value of the audio signal of the first far-end talker.

[0084] For example, after determining the first far-end talker according to the energy difference value, if the energy value of the audio signal of the target far-end talker is greater than the first energy threshold and does not exceed the second energy threshold, and the energy value of the audio signal of the first far-end talker is greater than the first energy threshold, in order not to reduce the listening experience of the user, the audio signal of the first far-end talker is suppressed in energy according to the aforementioned manner 2, and the energy value of the audio signal of the target far-end talker is kept unchanged, so that the energy difference value between the energy value of the audio signal of the target far-end talker and the energy value of the adjusted audio signal of the first far-end talker is not less than the energy threshold.

[0085] The first energy threshold is less than the second energy threshold. The first energy threshold can be a lower limit value of a comfortable energy value range in auditory perception, and the second energy threshold can be an upper limit value of the comfortable energy value range. If the energy value of an audio signal is within the comfortable energy value range, the user is comfortable in auditory perception.

[0086] After determining the first far-end talker according to the energy difference value, if the energy value of the audio signal of the target far-end talker is not greater than the first energy threshold, and the energy value of the audio signal of the first far-end talker is not greater than the first energy threshold, at this time, since the energy value of the audio signal of the first far-end talker is not greater than the first energy threshold, i.e., the energy value of the audio signal of the first far-end talker is not high itself, and the energy value of the audio signal of the target far-end talker is low, in order to enhance the listening experience of the user to the audio signal of the target far-end talker, the aforementioned manner 1 can be used for energy adjustment, so that the energy value of the first audio signal of the target far-end talker is higher than the energy value of the first audio signal of the first far-end talker, and the energy difference value between them is not less than the energy threshold. In some embodiments, in the case of adjustment according to manner 1, the energy value of the first audio signal of the target far-end talker after adjustment can also be constrained to be within the comfortable energy value range.

[0087] After the first call far end is determined according to the energy difference, if the energy value of the audio signal of the target call far end is not greater than the first energy threshold value, and the energy value of the audio signal of the first call far end is greater than the first energy threshold value, at this time, in order to increase the listening experience of the user to the audio signal of the target call far end, and reduce the influence of the audio signal of the first call far end, the energy value of the audio signal can be adjusted in the aforementioned manner 3, the energy of the audio signal of the target call far end is enhanced, and the energy of the audio signal of the first call far end is suppressed, so that the energy value of the first audio signal of the target call far end is greater than the energy value of the second audio signal of the first call far end, further, it can also be ensured that the energy difference between the energy value of the first audio signal and the energy value of the second audio signal of the first call far end is not less than the energy threshold value. In some embodiments, in the case of adjusting according to manner 3, the energy value of the first audio signal of the target call far end after adjustment can also be constrained to be within the comfortable energy value range.

[0088] S140, mix the first audio signal of the target call far end and the second audio signals of the other call far ends to obtain a first target audio signal; and playing the first target audio signal at the call near end.

[0089] After obtaining the first audio signal of the target call far end and the second audio signals of the other call far ends, the first audio signal of the target call far end and the second audio signals of the other call far ends can be mixed to combine the first audio signal of the target call far end and the second audio signals of the other call far ends into one audio signal, to obtain a combined audio signal as a first target audio signal. After obtaining the first target audio signal, the first target audio signal is played at the call near end, so that the object at the call near end side listens to the first target audio signal.

[0090] In the embodiment, the energy difference between the energy value of the audio signal of the target call far end selected by the call near end from the plurality of call far ends of the target voice call and the energy value of the audio signal of each other call far end is used to adjust the energy of at least one of the audio signal of the target call far end and the audio signal of the other call far end, to obtain the first audio signal of the target call far end and the second audio signal of each other call far end. Since the energy value of the first audio signal of the target call far end is greater than the energy value of the second audio signal of the other call far end, in the process of playing the first target audio signal obtained by mixing the first audio signal and the second audio signal at the call near end, the audio signal of the other call far end interferes with the user at the call near end to listen to the audio signal of the target call far end, so that the first audio signal of the target call far end is more easily heard and distinguished by the user at the call near end. Moreover, even if the number of call far ends relative to the call near end is large, since the user at the call near end specifies the call far end as the target call far end, the user does not need to perform a mute operation on each of the other call far ends except for the call far end of interest, the user operation is less, and the probability of misoperation is also low.

[0091] Meanwhile, compared with the processing scheme in the related art for distinguishing the audio signal of the user of interest and performing voice enhancement on the audio signal of the user of interest by using the voiceprint information in the voice signal, the embodiment only needs to extract the energy value of the audio signal, and can adjust the audio signal according to the energy value of the audio signal, without the need to extract the voiceprint by using a large amount of computing resources, thereby greatly reducing the data processing amount and resource consumption, and reducing the cost of voice call.

[0092] In an embodiment, as shown in FIG. 1, Figure 3 Before S120, the method further includes:

[0093] S210, performing voice detection on the audio signal of the target call far end to obtain a voice detection result.

[0094] In the embodiment, the voice detection result is used to indicate whether the audio signal of the target call far end includes a voice signal.

[0095] Voice detection (Voice Activity Detection, VAD for short) refers to voice detection on an audio signal. If the audio signal includes human voice, it is determined that the audio signal includes a voice signal. If the audio signal does not include human voice, it is determined that the audio signal does not include a voice signal.

[0096] In some embodiments, the audio signal of the target call far end can be input into a voice detection model to obtain a voice detection result output by the voice detection model. The voice detection model can include sample audio signals labeled with sample labels and sample voice detection results corresponding to the sample audio signals, and is trained based on a neural network model initialized with parameters. The sample labels are used to indicate whether the sample audio signals include voice signals or not. For example, a sample label of 1 indicates that the sample audio signal includes a voice signal, and a sample label of 0 indicates that the sample audio signal does not include a voice signal.

[0097] In some other embodiments, the audio signal of the target call far end can be converted into a frequency domain signal. If the ratio of the energy value of the frequency domain signal below the target frequency to the total energy value of the frequency domain signal converted from the audio signal of the target call far end is greater than an energy ratio threshold, it is determined that the voice detection result of the audio signal of the target call far end includes a voice signal. The energy ratio threshold can be set based on requirements. For example, the energy ratio threshold is 0.5. The frequency of human voice is generally concentrated in 50-2000 Hz, and the sound with a frequency lower than 2000 Hz can be considered as human voice. Therefore, the target frequency can be set to 2 kHz.

[0098] In some other embodiments, the target audio signal of the target call far end can be obtained as an intermediate audio signal within a target time period. The target time period refers to the time period in which the audio signal of the target call far end is received. The target audio signal of the target call far end refers to the audio signal sent by the target call far end during the target voice call. If the proportion of the intermediate audio signal including a voice signal is greater than a proportion threshold, the voice detection result of the audio signal of the target call far end is obtained as including a voice signal. If the proportion of the intermediate audio signal including a voice signal is not greater than the proportion threshold, the voice detection result of the audio signal of the target call far end is obtained as not including a voice signal. The length of the target time period and the proportion threshold can be set based on requirements, which are not limited in the present application.

[0099] For example, the receiving time of the audio signal of the target call far end can be set as the cutoff time of the target time period, and a target time period with a target length, for example, 1 s or 2 s, is determined.

[0100] It can be understood that the identification method of whether the intermediate audio signal includes a voice signal can refer to the aforementioned two identification methods of whether a voice signal includes a voice signal, which will not be described again.

[0101] In a case where it is determined that the proportion of the intermediate audio signals including the speech signal is greater than the proportion threshold, it is indicated that the number of the intermediate audio signals including the speech signal is large, most of the intermediate audio signals include the speech signal, and the possibility of the object on the far-end side of the target call speaking in the target time period is high. Therefore, it is determined that the audio signal of the target call far-end includes the speech signal. In a case where it is determined that the proportion of the intermediate audio signals including the speech signal is not greater than the proportion threshold, it is indicated that the number of the intermediate audio signals including the speech signal is small, most of the intermediate audio signals do not include the speech signal, and the possibility of the object on the far-end side of the target call speaking in the target time period is low. Therefore, it is determined that the audio signal of the target call far-end does not include the speech signal.

[0102] In this embodiment, the detection process of the speech signal is as shown in Figure 4 The intermediate audio signals are determined, and then speech detection is performed on the intermediate audio signals to obtain the speech detection result of the intermediate audio signals. Then, based on the speech detection result of each intermediate audio signal, the speech detection result of the audio signal of the target call far-end is determined. If the proportion of the intermediate audio signals including the speech signal is greater than the proportion threshold, the speech detection result of the audio signal of the target call far-end including the speech signal is obtained. If the proportion of the intermediate audio signals including the speech signal is not greater than the proportion threshold, the speech detection result of the audio signal of the target call far-end not including the speech signal is obtained.

[0103] If the speech detection result indicates that the audio signal of the target call far-end includes the speech signal, S120 is performed. If the speech detection result indicates that the audio signal of the target call far-end does not include the speech signal, S220 is performed.

[0104] S220, performing stream mixing processing on the audio signal of the target call far-end and the audio signal of the other call far-end to obtain a second target audio signal; and playing the second target audio signal at the call near-end.

[0105] Since the call near-end wants to clearly hear and recognize the speech signal in the audio signal of the target call far-end, if it is determined that the audio signal of the target call far-end does not include the speech signal, whether the audio signal of the target call far-end is easily heard and recognized is not important to the call near-end. At this time, it is not necessary to perform audio energy adjustment on the audio signal of the target call far-end and the audio signal of the other call far-end, perform stream mixing processing on the audio signal of the target call far-end and the audio signal of the other call far-end to obtain a second target audio signal, and then play the second target audio signal at the call near-end.

[0106] In the embodiment, the voice detection is performed on the audio signal of the target call far end. If the voice detection result indicates that the audio signal of the target call far end does not include a voice signal, the energy adjustment is not performed, and the audio signal of the target call far end and the audio signals of the other call far ends are directly mixed to obtain the second target audio signal. Therefore, in the case that the audio signal of the target call far end does not include a voice signal, the energy adjustment is not required, the computing resource is saved, and the waste of the computing resource is avoided.

[0107] In an embodiment, the step S130 further includes: if the energy difference between the energy value of the audio signal of the target call far end and the energy value of the audio signal of each of the other call far ends is not less than the energy threshold, taking the audio signal of the target call far end as the first audio signal of the target call far end, and taking the audio signal of each of the other call far ends as the corresponding second audio signal.

[0108] If the energy difference between the energy value of the audio signal of the target call far end and the energy value of the audio signal of each of the other call far ends is not less than the energy threshold, it indicates that the influence of the audio signal of each of the other call far ends on the audio signal of the target call far end is small. Therefore, the audio signal of the target call far end is taken as the first audio signal of the target call far end, and the audio signal of each of the other call far ends is taken as the corresponding second audio signal. Then, the first audio signal of the target call far end and the second audio signals of the other call far ends are mixed to obtain the first target audio signal, and the first target audio signal is played at the call near end.

[0109] In the embodiment, the energy difference between the energy value of the audio signal of the target call far end and the energy value of the audio signal of each of the other call far ends is not less than the energy threshold, and the voice signal is not adjusted, but directly mixed. Therefore, the resource consumption for adjusting the voice signal is saved, and the waste of the resource is avoided. Moreover, the time for adjusting the voice signal is saved, the efficiency of obtaining the first target audio signal is improved, and the real-time performance of the voice call is improved.

[0110] In an embodiment, as shown in FIG. 3, the step S130 can include the following steps S310-S320. Figure 5

[0111] S310, if it is determined according to the energy difference between the target call far end and each of the other call far ends that there is a first call far end with a corresponding energy difference less than the energy threshold in the other call far ends, obtaining a first expected energy value of the target call far end and a second expected energy value of the first call far end.

[0112] ​S320, adjust the energy value of the audio signal of the target far-end talker to a first expected energy value to obtain a first audio signal of the target far-end talker; and adjust the energy value of the audio signal of the first far-end talker to a second expected energy value to obtain a second audio signal of the first far-end talker.

[0113] The first expected energy value is greater than the second expected energy value, and the energy threshold is a value greater than 0.

[0114] For ease of description, the other far-end talker corresponding to the energy difference value less than the energy threshold is referred to as the first far-end talker.

[0115] Since the energy difference value corresponding to the first far-end talker is less than the energy threshold, it indicates that the energy of the audio signal of the first far-end talker is higher than the energy of the audio signal of the target far-end talker. If no energy adjustment is performed, the audio signal of the first far-end talker is relatively high in energy, which can cause the user on the near-end side to be unable to clearly hear the audio signal on the target far-end side. Therefore, in this case, it is determined that the energy value of the audio signal of the target far-end talker and / or the energy value of the audio signal of the first far-end talker needs to be adjusted, so that the energy value of the first audio signal of the target far-end talker is greater than the energy value of the second audio signal of the first far-end talker after adjustment, and the difference between the energy value of the first audio signal of the target far-end talker and the energy value of the second audio signal of the first far-end talker is sufficient to ensure that the user on the near-end side can clearly hear the audio signal of the target far-end talker.

[0116] In the above embodiment, the first expected energy value of the target far-end talker and the second expected energy value of the first far-end talker are used to adjust the corresponding audio signal, and the difference between the first expected energy value and the second expected energy value can be greater than or equal to the energy threshold. This can ensure that the energy value of the first audio signal is greater than the energy value of the second audio signal after adjustment.

[0117] The second expected energy values corresponding to different first far-end talkers can be the same or different, which is not specifically limited herein. The first expected energy value of the target far-end talker and the second expected energy value of the first far-end talker can be pre-set or determined in a targeted manner with reference to the energy value of the audio signal of the target far-end talker and the energy value of the audio signal of the first far-end talker.

[0118] After the first expected energy value and the second expected energy value are obtained, the first gain of the target far-end speech can be calculated according to the energy value of the audio signal of the target far-end speech and the first expected energy value, and then the energy value of the audio signal of the target far-end speech is adjusted by the first gain of the target far-end speech to obtain the first audio signal of the target far-end speech. Meanwhile, the second gain of the other far-end speech can be calculated according to the energy value of the audio signal of the other far-end speech and the second expected energy value, and then the energy value of the audio signal of the other far-end speech is adjusted by the second gain of the other far-end speech to obtain the second audio signal of the other far-end speech, so as to realize the amplification of the energy value of the audio signal of the target far-end speech (relative to the energy value of the audio signal of the other far-end speech) and the reduction of the energy value of the audio signal of the other far-end speech (relative to the energy value of the audio signal of the target far-end speech).

[0119] As an implementation manner, the first expected energy value of the target far-end speech and the second expected energy value of the first far-end speech are obtained, including: obtaining the energy value of the audio signal of the target far-end speech as the first expected energy value of the target far-end speech; calculating the difference between the energy value of the audio signal of the target far-end speech and the energy threshold to obtain a first energy value; determining the second expected energy value of the first far-end speech based on the first energy value, and the second expected energy value is not greater than the first energy value.

[0120] That is, the energy value of the audio signal of the target far-end speech can be kept unchanged, and the energy value of the audio signal of the first far-end speech is reduced, so that the difference between the energy value of the audio signal of the target far-end speech and the energy value of the audio signal of the first far-end speech is not less than the energy threshold, and the purpose of increasing the difference between the energy value of the audio signal of the target far-end speech and the energy value of the audio signal of the first far-end speech is achieved.

[0121] As another implementation manner, the first expected energy value of the target far-end speech and the second expected energy value of the first far-end speech are obtained, including: obtaining the energy value of the audio signal of the first far-end speech as the second expected energy value of the first far-end speech; calculating the sum of the energy threshold and the energy value of the audio signal of the second far-end speech to obtain a second energy value; the second far-end speech is the first far-end speech with the highest energy value of the audio signal; and determining the first expected energy value of the target far-end speech based on the second energy value, and the first expected energy value is not less than the second energy value.

[0122] First, the first call remote with the highest energy value of the audio signal is determined as the second call remote, and then the sum of the energy value of the audio signal of the second call remote and the energy threshold is calculated to obtain a second energy value, at this time, the difference between the second energy value and the energy value of the audio signal of each second call remote is not less than the energy threshold, at this time, a first expected energy value not less than the second energy value is determined, and the difference between the first expected energy value and the energy value of the audio signal of each second call remote is also not less than the energy threshold.

[0123] That is, the energy value of the audio signal of the first call remote can be kept unchanged, and the energy value of the audio signal of the target call remote is increased, so that the difference between the energy value of the audio signal of the target call remote and the energy value of the audio signal of the first call remote is not less than the energy threshold, and the purpose of increasing the difference between the energy value of the audio signal of the target call remote and the energy value of the audio signal of the first call remote is achieved.

[0124] In still another embodiment, the first expected energy value of the target call remote and the second expected energy value of the first call remote are obtained, including: calculating the sum of the energy value of the audio signal of the target call remote and a third energy value to obtain the first expected energy value of the target call remote; calculating the difference between the energy value of the audio signal of the second call remote and a fourth energy value to obtain the second expected energy value of the first call remote; the second call remote is the first call remote with the highest energy value of the audio signal; the sum of the third energy value and the fourth energy value is not less than the energy threshold.

[0125] That is, the energy value of the audio signal of the first call remote can be reduced (the reduction amplitude is the fourth energy value), and the energy value of the audio signal of the target call remote is increased (the increase is the third energy value), so that the difference between the energy value of the audio signal of the target call remote and the energy value of the audio signal of the first call remote is not less than the energy threshold, and the purpose of increasing the difference between the energy value of the audio signal of the target call remote and the energy value of the audio signal of the first call remote is achieved. In this application, the specific values of the third energy value and the fourth energy value are not limited, as long as the sum of the third energy value and the fourth energy value is not less than the energy threshold. For example, the third energy value and the fourth energy value are each half of the energy threshold.

[0126] After obtaining the first expected energy value and the second expected energy value, the energy value of the audio signal of the target call remote can be adjusted to the first expected energy value to obtain the first audio signal of the target call remote, and the energy value of the audio signal of the first call remote is adjusted to the second expected energy value to obtain the second audio signal of the first call remote, at this time, the energy difference between the first audio signal and the second audio signal of the first call remote is not less than the energy threshold, and the difference between the first audio signal and the second audio signal of the first call remote is large.

[0127] In some embodiments, adjusting the energy value of the audio signal of the target far-end of the call to the first desired energy value to obtain the first audio signal of the target far-end of the call comprises: adjusting the energy value of a target audio segment containing the speech signal in the audio signal of the target far-end of the call to the first desired energy value, and keeping the energy values of other audio segments in the audio signal of the target far-end of the call except the target audio segment unchanged to obtain the first audio signal of the target far-end of the call.

[0128] The audio signal of the target far-end of the call can include a plurality of audio segments. The audio segments can be subjected to speech detection in the aforementioned manner to determine whether the audio segments include speech signals. If an audio segment includes a speech signal, the audio segment is determined to be a target audio segment. The energy value of the target audio segment containing the speech signal in the audio signal of the target far-end of the call is adjusted to the first desired energy value, and the energy values of other audio segments in the audio signal of the target far-end of the call except the target audio segment are kept unchanged to obtain the first audio signal of the target far-end of the call.

[0129] Since the near-end of the call wants to clearly hear and recognize the speech signal in the audio signal of the target far-end of the call, the energy value of an audio segment not containing a speech signal does not affect the user's listening to the speech, and thus the energy value of only the target audio segment containing the speech signal in the audio signal of the target far-end of the call can be adjusted, and the energy values of other audio segments in the audio signal of the target far-end of the call except the target audio segment do not need to be adjusted, thereby saving resource consumption for adjusting the energy values of other audio segments in the audio signal of the target far-end of the call except the target audio segment and reducing resource waste.

[0130] In the present application, the audio signal of the target far-end of the call can be subjected to speech detection by a speech detection model to obtain the position of the target audio segment containing the speech signal in the audio signal of the target far-end of the call, so as to determine the target audio segment from the audio signal of the target far-end of the call. Correspondingly, the speech detection model can be trained by a neural network model initialized with parameters based on sample audio signals labeled with sample labels, sample speech detection results corresponding to the sample audio signals, and sample position information of audio segments containing speech signals in the sample audio signals.

[0131] In an embodiment, adjusting the energy value of the audio signal of the first far-end of the call to the first desired energy value to obtain the first audio signal of the target far-end of the call comprises: adjusting the energy value of a relevant audio segment in the audio signal of the first far-end of the call to the second desired energy value, and keeping the energy values of other audio segments in the audio signal of the first far-end of the call except the relevant audio segment unchanged to obtain the second audio signal of the first far-end of the call.

[0132] The relevant audio segment can refer to an audio segment in the audio signal of the first call far end that is time-aligned with the target audio segment. Time-aligned can refer to that the start time and the end time of the two audio segments are the same.

[0133] Since the call near end wants to clearly hear and recognize the voice signal in the audio signal of the target call far end, the part of the audio signal of the first call far end that can affect the user's listening is only the relevant audio segment in the audio signal of the first call far end. For other audio segments in the audio signal of the first call far end except the relevant audio segment, the high or low energy value does not affect the user's listening voice. Therefore, only the energy value of the target audio segment containing the voice signal in the audio signal of the first call far end needs to be adjusted, and the energy value of other audio segments in the audio signal of the first call far end except the relevant audio segment does not need to be adjusted, thereby saving the resource consumption for adjusting the energy value of other audio segments in the audio signal of the first call far end except the relevant audio segment, and reducing resource waste.

[0134] In some embodiments, please continue to refer to Figure 5 Step S130 can further include the following step S330:

[0135] S330, if it is determined according to the energy difference between the target call far end and each other call far end that there is a third call far end in the other call far ends whose corresponding energy difference is not less than the energy threshold, obtaining the audio signal of the third call far end as the second audio signal of the third call far end.

[0136] For the third call far end whose energy difference is less than or not less than the energy threshold, the energy value of the audio signal of the third call far end does not affect the user on the call near end side to clearly hear the audio signal on the target call far end side, and the audio signal of the third call far end can be directly obtained as the second audio signal of the third call far end without adjustment, thereby avoiding resource waste.

[0137] For example, the audio signal processing process is as shown in Figure 6 , including:

[0138] S501: Voice detection. First, voice detection is performed on the audio signal of the target call far end to obtain a voice detection result of the audio signal of the target call far end.

[0139] S502: Whether the audio signal includes a voice signal. Whether the audio signal of the target call far end includes a voice signal is determined according to the voice detection result of the audio signal of the target call far end. If yes, S503 is performed, and if no, S507 is performed.

[0140] S503: Energy value detection of the audio signal. The energy values of the audio signals of each call far end (including the target call far end participating in the target voice call and other call far ends) are determined.

[0141] S504: Whether the energy difference values are all not less than the energy threshold. The difference values between the energy value of the audio signal of the target voice call and the energy values of the audio signals of each other call far end are calculated as the energy difference values corresponding to the audio signals of each other call far end, and it is determined whether each energy difference value is not less than the energy threshold. If yes, S507 is executed, and if no (indicating that at least one energy difference value is less than the energy threshold), S505 is executed.

[0142] S505: Determination of the expected energy value. A first expected energy value is determined for the target call far end, and a second expected energy value is determined for the first call far end (the other call far end with the corresponding energy difference value less than the energy threshold).

[0143] S506: Adjustment of the audio signal by the expected energy value. The energy value of the audio signal of the target call far end is adjusted to the first expected energy value to obtain a first audio signal of the target call far end, and the energy value of the audio signal of the first call far end is adjusted to the second expected energy value to obtain a second audio signal of the first call far end; meanwhile, the audio signal of the third call far end (the other call far end with the corresponding energy difference value not less than the energy threshold) can also be obtained as a second audio signal of the third call far end (i.e., the energy value of the audio signal of the third call far end is kept unchanged).

[0144] S507: Mixed flow playing. After the adjustment steps of S505 and S506, the first target audio signal is obtained by mixed flow processing of the first voice signal and the second voice signal, and the first target audio signal is played at the call near end; in the case where the adjustment steps of S505 and S506 are not performed, the second target audio signal is obtained by directly mixed flow processing of the audio signal of the target call far end and the audio signals of each other call far end, and the second target audio signal is played at the call near end.

[0145] In this embodiment, the energy value of the audio signal of the first call far end with the corresponding energy difference value less than the energy threshold is adjusted according to the second expected energy value, and the energy value of the audio signal of the target call far end is adjusted according to the first expected energy value, so that the difference between the audio signal of the target call far end and the audio signal of the first call far end is larger, and the interference of the audio signal of the first call far end on the audio signal of the target call far end is smaller, thereby making the voice of the target call far end in the first target audio signal after mixed flow more easily heard and distinguished by the user at the call near end. Meanwhile, the energy value of the audio signal of the third call far end with the corresponding energy difference value not less than the energy threshold is not adjusted, avoiding waste of processing resources.

[0146] To more clearly explain the solution of this application, the signal processing method of this application will be explained below with reference to an exemplary scenario.

[0147] like Figure 7 As shown, in a cloud conferencing scenario, the near end of the call ( Figure 7 (Not shown in the image) A conference request is sent to remote ends a, b, and c. Remote ends a, b, and c accept the conference request from the near end and establish a cloud conference, which can then serve as the target voice call. At this point, the remote ends participating in the target voice call include remote ends a, b, and c, as well as the near end (not shown in the image). Figure 7 (not shown in the image), where the remote end b is the target remote end specified for the near end of the call.

[0148] The audio signals acquired by remote end 'a' are audio signal a0, audio signals acquired by remote end 'b' are audio signal b0, and audio signals acquired by remote end 'c' are audio signal c0. Remote ends 'a', 'b', and 'c' each send their acquired audio signals to the server. The server performs PCM encoding on audio signals a0, b0, and c0 to obtain audio signals a1 (encoded from a0), b1 (encoded from b0), and c1 (encoded from c0). In this example, the encoding can refer to PCM (Pulse Code Modulation), which is digitally converted audio information and is the most primitive audio format; essentially, it's a series of digitally converted audio signals.

[0149] Then, the server sends the encoded audio signals a1, b1, and c1 to the near end of the call. The near end decodes the audio signals a1, b1, and c1 to obtain the decoded audio signal a2 from audio signal a1, the decoded audio signal b2 from audio signal b1, and the decoded audio signal c2 from audio signal c1. In this example, the encoding could refer to PCM.

[0150] Pulse Code Modulation (PCM) is audio information converted from digital signals. It is also the most primitive audio format, essentially a series of audio signals converted from digital values.

[0151] Afterwards, the server sends the voice signal a2, the voice signal b2 and the voice signal c2 to the talk near-end, the talk near-end performs energy adjustment to obtain the first voice signal of the target talk far-end and the second voice signal of the other talk far-ends, and then the talk near-end performs mixed stream playing: the first voice signal of the target talk far-end and the second voice signal of the other talk far-ends are mixed to obtain the first target voice signal, and then the first target voice signal is played at the talk near-end.

[0152] The foregoing energy adjustment process is as shown in Figure 8 The user at the talk near-end side can first specify a target talk far-end, and after the target talk far-end is specified, voice detection is performed on the audio signal of the target talk far-end to determine whether the audio signal b2 of the target talk far-end includes a voice signal, and after it is determined that the audio signal of the target talk far-end includes a voice signal, energy value detection is performed on the audio signal to determine the energy values of the audio signal a2, the audio signal b2 and the audio signal c2. For example, if it is determined that the energy value of the audio signal a2 is higher than the energy value of the audio signal b2, the energy value of the audio signal b2 is higher than the energy value of the audio signal c2, and the energy difference between the energy value of the audio signal b2 and the energy value of the audio signal c2 is greater than an energy threshold value. Afterwards, energy value adjustment is performed: the energy value of the audio signal b2 is increased to obtain the audio signal b3, and the energy value of the audio signal a2 is decreased to obtain the audio signal a3, and the energy difference between the energy value of the audio signal b3 and the energy value of the audio signal a3 is greater than the energy threshold value. At this time, the audio signal a3 is determined as the second audio signal of the talk far-end a, the audio signal b3 is determined as the second audio signal of the talk far-end b, and the audio signal c2 is determined as the second audio signal of the talk far-end c. After the respective second audio signals of the three talk far-ends are obtained, the three second audio signals are mixed to obtain the first target voice signal.

[0153] Please refer to Figure 9 , Figure 9 A block diagram of an audio signal processing device according to an embodiment of the present application is shown in

[0154] The acquisition module 1110 is configured to acquire audio signals of multiple talk far-ends in a target voice call in which the talk near-end participates; the multiple talk far-ends include a target talk far-end selected for the talk near-end;

[0155] The calculation module 1120 is configured to subtract the energy value of the audio signal of the target talk far-end from the energy values of the audio signals of the other talk far-ends to obtain energy differences between the target talk far-end and the other talk far-ends, the other talk far-ends being talk far-ends other than the target talk far-end in the multiple talk far-ends;

[0156] The adjusting module 1130 is configured to perform energy adjustment on at least one of the audio signal of the target talk far end and the audio signal of each of the other talk far ends according to the energy difference between the target talk far end and each of the other talk far ends, to obtain a first audio signal of the target talk far end and a second audio signal of each of the other talk far ends; the energy value of the first audio signal is greater than the energy value of the second audio signal.

[0157] The mixing module 1140 is configured to perform mixing processing on the first audio signal of the target talk far end and the second audio signal of each of the other talk far ends, to obtain a first target audio signal.

[0158] The playing module 1150 is configured to play the first target audio signal at the talk near end.

[0159] Optionally, the adjusting module is further configured to, if it is determined, according to the energy difference between the target talk far end and each of the other talk far ends, that there is a first talk far end corresponding to an energy difference less than the energy threshold among the other talk far ends, obtain a first expected energy value of the target talk far end and a second expected energy value of the first talk far end; the first expected energy value is greater than the second expected energy value; adjust the energy value of the audio signal of the target talk far end to the first expected energy value, to obtain the first audio signal of the target talk far end; and adjust the energy value of the audio signal of the first talk far end to the second expected energy value, to obtain a second audio signal of the first talk far end.

[0160] Optionally, the adjusting module is further configured to obtain the energy value of the audio signal of the target talk far end as the first expected energy value of the target talk far end; calculate a difference between the energy value of the audio signal of the target talk far end and the energy threshold, to obtain a first energy value; and determine the second expected energy value of the first talk far end based on the first energy value, the second expected energy value being not greater than the first energy value.

[0161] Optionally, the adjusting module is further configured to obtain the energy value of the audio signal of the first talk far end as the second expected energy value of the first talk far end; calculate a sum of the energy threshold and the energy value of the audio signal of a second talk far end, to obtain a second energy value; the second talk far end is the first talk far end with the highest energy value of the audio signal; and determine the first expected energy value of the target talk far end based on the second energy value, the first expected energy value being not less than the second energy value.

[0162] Optionally, the adjusting module is further configured to calculate a sum of the energy value of the audio signal of the target talk far end and a third energy value, to obtain the first expected energy value of the target talk far end; calculate a difference between the energy value of the audio signal of the second talk far end and a fourth energy value, to obtain the second expected energy value of the first talk far end; the second talk far end is the first talk far end with the highest energy value of the audio signal; and a sum of the third energy value and the fourth energy value is not less than the energy threshold.

[0163] Optionally, the adjusting module is further configured to adjust an energy value of a target audio segment containing the speech signal in the audio signal of the target far-end to a first expected energy value, and keep energy values of other audio segments in the audio signal of the target far-end unchanged except the target audio segment, to obtain a first audio signal of the target far-end.

[0164] Optionally, the adjusting module is further configured to, if it is determined according to the energy difference between the target far-end and each of the other far-ends that there is a third far-end in the other far-ends whose corresponding energy difference is not less than the energy threshold, obtain an audio signal of the third far-end as a second audio signal of the third far-end.

[0165] Optionally, the playing module is further configured to, if the energy difference between the energy value of the audio signal of the target far-end and the energy value of the audio signal of each of the other far-ends is not less than the energy threshold, play the audio signal of the target far-end as a first audio signal of the target far-end, and play the audio signal of each of the other far-ends as a corresponding second audio signal.

[0166] Optionally, the computing module is further configured to perform speech detection on the audio signal of the target far-end to obtain a speech detection result, and perform the step of subtracting the energy value of the audio signal of the target far-end from the energy value of the audio signal of each of the other far-ends if the speech detection result indicates that the audio signal of the target far-end includes a speech signal.

[0167] Optionally, the playing module is further configured to, if the speech detection result indicates that the audio signal of the target far-end does not include a speech signal, perform a stream mixing process on the audio signal of the target far-end and the audio signal of the other far-ends to obtain a second target audio signal, and play the second target audio signal at the near-end.

[0168] Optionally, the apparatus further includes a selecting module configured to, in response to a selection operation triggered by an object at the near-end on the plurality of far-ends, select a far-end selected by the selection operation as the target far-end.

[0169] Figure 9 A structural block diagram of an electronic device for performing an audio signal processing method according to an embodiment of the present application is shown. The electronic device can be a server 200 or a terminal device 400, etc., and it should be noted that Figure 1 The computer system 1200 of the electronic device shown is only an example and should not impose any limitation on the functions and use range of the embodiments of the present application. Figure 9 The computer system 1200 of the electronic device shown is only an example and should not impose any limitation on the functions and use range of the embodiments of the present application.

[0170] As shown in FIG. 12, the computer system 1200 can include a processor 1201, a memory 1202, a storage 1203, an input device 1204, an output device 1205, and a communication device 1206. Figure 9As shown, the computer system 1200 includes a central processing unit (CPU) 1201 which can execute various appropriate actions and processes in accordance with programs stored in a read-only memory (ROM) 1202 or loaded from a storage section 1208 into a random access memory (RAM) 1203, such as executing the methods in the above-described embodiments. Various programs and data required for system operation are also stored in the RAM 1203. The CPU 1201, the ROM 1202, and the RAM 1203 are connected to each other through a bus 1204. An input / output (I / O) interface 1205 is also connected to the bus 1204.

[0171] Connected to the I / O interface 1205 are an input section 1206 including a keyboard, a mouse, etc.; an output section 1207 including a display such as a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 1208 including a hard disk, etc.; and a communication section 1209 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 1209 performs communication processing via a network such as the Internet. A drive 1210 is also connected to the I / O interface 1205 as necessary. A removable recording medium 1211 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc. is attached to the drive 1210 as necessary, so that a computer program read therefrom is installed into the storage section 1208 as necessary.

[0172] In particular, according to embodiments of the present application, the processes described above with reference to the flowcharts can be implemented as a computer software program. For example, embodiments of the present application include a computer program product comprising a computer program carried on a computer readable medium, the computer program containing program code for executing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via the communication section 1209, and / or installed from the removable recording medium 1211. When the computer program is executed by the central processing unit (CPU) 1201, various functions defined in the systems of the present application are executed.

[0173] It should be noted that the computer-readable medium in the embodiments of the present application can be a computer-readable signal medium or a computer-readable storage medium or any combination thereof. The computer-readable storage medium may, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or apparatus, or any combination thereof. More specific examples of the computer-readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disk read-only memory (Compact Disc Read-Only Memory, CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this application, the computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device or apparatus. In this application, the computer-readable signal medium can include a data signal carrying a computer-readable program code in a baseband or as a part of a carrier wave. Such a propagated data signal can take on various forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium that can transmit, propagate or transport a program for use by or in connection with an instruction execution system, device or apparatus. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to wireless, wired, or the like, or any suitable combination thereof.

[0174] The flowcharts and block diagrams in the drawings illustrate the possible implementation architectures, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In the flowcharts or block diagrams, each block can represent a module, a program segment or a part of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur in different orders than that shown in the drawings. For example, two blocks that are shown in succession can actually be executed substantially in parallel, and sometimes in reverse order, depending on the involved functions. It should also be noted that each block in the block diagrams or flowcharts, and the combination of blocks in the block diagrams or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer-readable instructions.

[0175] The units described in the embodiments of the present application can be implemented by software, or by hardware, or by a combination of software and hardware. The units described can also be located in a single processor. In some cases, the names of the units do not limit the units themselves.

[0176] As another aspect, the present application also provides a computer readable storage medium, which can be included in the electronic device described in the above embodiments, or can exist separately without being assembled into the electronic device. The computer readable storage medium stores computer readable instructions, which, when executed by a processor, implement the method in any of the above embodiments.

[0177] According to an aspect of the embodiments of the present application, a computer program product is provided, which includes computer readable instructions stored in a computer readable storage medium. A processor of an electronic device reads the computer readable instructions from the computer readable storage medium, and the processor executes the computer readable instructions to cause the electronic device to perform the method in any of the above embodiments.

[0178] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program with a predetermined function, and works together with other related parts to achieve a predetermined target, and can be implemented entirely or partially by using software, hardware (such as a processing circuit or a memory), or a combination thereof, and similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a whole module or unit of the function of the module or unit, or a part of the module or unit.

[0179] It should be noted that although several modules or units of the device for action execution are mentioned in the above detailed description, such a division is not mandatory. In fact, according to the embodiments of the present application, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided into multiple modules or units.

[0180] Those skilled in the art can clearly understand the example embodiments described herein through the above description of the embodiments, which can be implemented by software or by software in combination with necessary hardware. Therefore, the technical solution according to the embodiments of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a U disk, a mobile hard disk, etc.) or on a network, and includes a number of instructions to enable an electronic device (which can be a personal computer, a server, a touch terminal, or a network device, etc.) to perform the method according to the embodiments of the present application.

[0181] Other embodiments of the present application will be apparent to those skilled in the art from consideration of the specification and practice of the embodiments disclosed herein. The present application is intended to cover any variations, uses, or adaptations of the application following the general principles thereof and including such departures from the present disclosure as come within known use or custom in the art. It is to be understood that the application is not limited to the exact details of construction, and the arrangement of the exact components as set forth herein and as such changes as are well known and obvious to one skilled in the art are intended to be encompassed by the present application. The scope of the application is limited only by the claims that follow.

[0182] Finally, it should be noted that: the above examples are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing examples, those skilled in the art will understand that they can still modify the technical solutions recorded in the foregoing examples, or make equivalent replacement for part of the technical features; and these modifications or replacements do not drive the essence of the corresponding technical solutions out of the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. An audio signal processing method, characterized in that, The method includes: Acquire audio signals from multiple far ends in a target voice call in which the near end of the call participates; the multiple far ends include a target far end selected for the near end of the call. The energy value of the audio signal of the target call terminal is subtracted from the energy value of the audio signals of each of the other call terminals to obtain the energy difference between the target call terminal and each of the other call terminals. The other call terminals refer to the call terminals other than the target call terminal among the plurality of call terminals. Based on the energy difference between the target remote caller and each of the other remote callers, at least one of the audio signals of the target remote caller and the audio signals of the other remote callers is energy adjusted to obtain a first audio signal of the target remote caller and a second audio signal of each of the other remote callers; the energy value of the first audio signal is greater than the energy value of the second audio signal. The first audio signal from the target call remote end and the second audio signals from each of the other call remote ends are mixed to obtain the first target audio signal; The first target audio signal is played at the near end of the call.

2. The method according to claim 1, characterized in that, The step of adjusting the energy of at least one of the audio signals of the target remote call terminal and the audio signals of the other remote call terminals based on the energy difference between the target remote call terminal and each of the other remote call terminals to obtain a first audio signal of the target remote call terminal and a second audio signal of each of the other remote call terminals includes: If, based on the energy difference between the target remote caller and each of the other remote callers, it is determined that there is a first remote caller among the other remote callers with a corresponding energy difference less than an energy threshold, then the first expected energy value of the target remote caller and the second expected energy value of the first remote caller are obtained; the first expected energy value is greater than the second expected energy value. The energy value of the audio signal at the target call remote end is adjusted to the first desired energy value to obtain the first audio signal at the target call remote end; The energy value of the audio signal at the far end of the first call is adjusted to the second desired energy value to obtain the second audio signal at the far end of the first call.

3. The method according to claim 2, characterized in that, The step of obtaining the first expected energy value of the target call remote end and the second expected energy value of the first call remote end includes: The energy value of the audio signal from the target remote end of the call is obtained as the first expected energy value of the target remote end of the call. The difference between the energy value of the audio signal at the target call's far end and the energy threshold is calculated to obtain a first energy value; A second expected energy value is determined for the remote end of the first call based on the first energy value, wherein the second expected energy value is not greater than the first energy value.

4. The method according to claim 2, characterized in that, The step of obtaining the first expected energy value of the target call remote end and the second expected energy value of the first call remote end includes: Obtain the energy value of the audio signal from the far end of the first call, and use it as the second expected energy value of the far end of the first call; The energy threshold is calculated by summing it with the energy value of the audio signal at the second call terminal to obtain the second energy value; the second call terminal is the first call terminal with the highest audio signal energy value. Based on the second energy value, a first expected energy value is determined for the target call remote end, and the first expected energy value is not less than the second energy value.

5. The method according to claim 2, characterized in that, The step of obtaining the first expected energy value of the target call remote end and the second expected energy value of the first call remote end includes: The sum of the energy value of the audio signal at the target call remote end and the third energy value is calculated to obtain the first expected energy value of the target call remote end; The difference between the energy value of the audio signal at the second call terminal and the fourth energy value is calculated to obtain the second expected energy value of the first call terminal; the second call terminal is the first call terminal with the highest audio signal energy value; the sum of the third energy value and the fourth energy value is not less than the energy threshold.

6. The method according to claim 2, characterized in that, The step of adjusting the energy value of the audio signal at the target call's far end to the first desired energy value to obtain the first audio signal at the target call's far end includes: The energy value of the target audio segment containing the voice signal in the audio signal of the target call remote end is adjusted to the first desired energy value, while keeping the energy values ​​of other audio segments in the audio signal of the target call remote end, excluding the target audio segment, unchanged, to obtain the first audio signal of the target call remote end.

7. The method according to claim 2, characterized in that, The step of adjusting the energy of at least one of the audio signals of the target remote call terminal and the audio signals of the other remote call terminals based on the energy difference between the target remote call terminal and each of the other remote call terminals to obtain a first audio signal of the target remote call terminal and a second audio signal of each of the other remote call terminals further includes: If, based on the energy difference between the target remote caller and each of the other remote callers, it is determined that there is a third remote caller among the other remote callers whose corresponding energy difference is not less than an energy threshold, the audio signal of the third remote caller is obtained as the second audio signal of the third remote caller.

8. The method according to claim 1, characterized in that, The step of adjusting the energy of at least one of the audio signals of the target remote call terminal and the audio signals of the other remote call terminals based on the energy difference between the target remote call terminal and each of the other remote call terminals to obtain a first audio signal of the target remote call terminal and a second audio signal of each of the other remote call terminals includes: If the energy difference between the energy value of the audio signal of the target call terminal and the energy value of the audio signals of each of the other call terminals is not less than the energy threshold, the audio signal of the target call terminal is taken as the first audio signal of the target call terminal; and the audio signals of each of the other call terminals are taken as the corresponding second audio signals.

9. The method according to any one of claims 1 to 8, characterized in that, The method further includes: Speech detection is performed on the audio signal from the far end of the target call to obtain the speech detection result; If the voice detection result indicates that the audio signal of the target call remote end includes a voice signal, the step of subtracting the energy value of the audio signal of the target call remote end from the energy value of the audio signals of each of the other call remote ends is performed to obtain the energy difference between the target call remote end and each of the other call remote ends.

10. The method according to claim 9, characterized in that, After performing voice detection on the audio signal of the target call's far end and obtaining the voice detection result, the method further includes: If the voice detection result indicates that the audio signal of the target call remote end does not include a voice signal, the audio signal of the target call remote end and the audio signals of other call remote ends are mixed to obtain a second target audio signal. The second target audio signal is played at the near end of the call.

11. The method according to claim 1, characterized in that, Before acquiring the audio signals of each of the multiple far ends in the target voice call in which the near end of the call participates, the method further includes: In response to a selection operation triggered by an object on the near side of the call to the plurality of far ends of the call, the far end of the call selected by the selection operation is taken as the target far end of the call.

12. An audio signal processing device, characterized in that, The device includes: The acquisition module is used to acquire the audio signals of multiple far ends in a target voice call in which the near end of the call participates; the multiple far ends include the target far end selected for the near end of the call; The calculation module is used to subtract the energy value of the audio signal of the target call remote end from the energy value of the audio signal of each other call remote end to obtain the energy difference between the target call remote end and each of the other call remote ends, wherein the other call remote ends refer to the call remote ends other than the target call remote end among the plurality of call remote ends. An adjustment module is configured to adjust the energy of at least one of the audio signals of the target call remote end and the audio signals of the other call remote ends based on the energy difference between the target call remote end and each of the other call remote ends, to obtain a first audio signal of the target call remote end and a second audio signal of each of the other call remote ends; the energy value of the first audio signal is greater than the energy value of the second audio signal; A mixing module is used to mix the first audio signal of the target call remote end and the second audio signals of each of the other call remote ends to obtain the first target audio signal. A playback module is used to play the first target audio signal at the near end of the call.

13. An electronic device, characterized in that, include: processor; A memory storing computer-readable instructions that, when executed by the processor, implement the method as described in any one of claims 1-11.

14. A computer-readable storage medium, characterized in that, It stores computer-readable instructions that, when executed by a processor, implement the method as described in any one of claims 1-11.

15. A computer program product, characterized in that, Includes computer-readable instructions that, when executed by a processor, implement the method of any one of claims 1-11.