Video Call Voice Processing Method, Communication Terminal, and Readable Storage Medium

By identifying the target person in the video image and adjusting the microphone gain according to their answering parameters, the problem in the prior art that the microphone gain cannot be adjusted according to the other party's situation is solved, and the quality of voice calls is improved.

CN114078466BActive Publication Date: 2025-06-10GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010848931.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-08-21
Publication Date
2025-06-10
Estimated Expiration
2040-08-21

AI Technical Summary

Technical Problem

The existing voice call scenarios cannot adjust the input gain of the microphone according to the other party's situation, making it difficult for the responder to hear the caller's voice clearly.

Method used

By identifying the target person in the video image and adjusting the gain of the sound acquisition device according to its position, action and answer parameters in the video image, the input gain of the microphone is adjusted in real time.

Benefits of technology

The microphone gain is adjusted according to the other party's situation, and the quality of voice calls is improved, so that the responder can hear the caller's voice more clearly.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114078466B_ABST
    Figure CN114078466B_ABST
Patent Text Reader

Abstract

The present application discloses a video call voice processing method, a communication terminal, and a readable storage medium. The video call voice processing method includes: identifying a target person in a current video image; obtaining an answering parameter of the target person, where the answering parameter is used to identify the sound intensity heard by the target person and includes at least one of the position of the target person in the current video image, the action of the target person, and the voice of the target person; adjusting the gain of a sound collection device according to the answering parameter, and transmitting a voice signal collected by the adjusted sound collection device to a target terminal. The present application can adjust the input gain of its own microphone according to the situation of the other party in a call scenario, thereby facilitating the provision of high-quality voice call services for users.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of communication and electronic devices, and particularly to a video call voice processing method, a communication terminal, and a readable storage medium. Background Art

[0002] In order to improve the quality of voice calls, most communication terminals with communication functions such as mobile phones are equipped with a call voice processing mechanism to ensure that both parties in a call can clearly hear each other's voices when the communication terminal makes a call in a noisy environment such as a bustling downtown area, thereby improving the quality of voice calls.

[0003] Currently, the most commonly used call voice processing method for communication terminals mainly focuses on noise reduction. For example, the microphone noise cancellation method is adopted, that is, the voice signal collected by the microphone is subjected to noise reduction processing. However, it does not consider the situation of the answering party. For example, when the answering party is far away from the terminal they hold, it is still very difficult for the answering party to clearly hear the call voice sent by the calling party. Summary of the Invention

[0004] In view of this, this application provides a video call voice processing method, a communication terminal, and a readable storage medium to solve the problem that the existing voice call scenario cannot adjust the input gain of its own microphone according to the situation of the other party.

[0005] A video call voice processing method provided by this application includes:

[0006] Identifying a target person in the current video image;

[0007] Obtaining an answering parameter of the target person, where the answering parameter is used to identify the sound intensity heard by the target person and includes at least one of the position of the target person in the current video image, the action of the target person, and the voice of the target person;

[0008] Adjusting the gain of the sound collection device according to the answering parameter, and transmitting the voice signal collected by the sound collection device after adjustment to the target terminal.

[0009] Optionally, the obtaining of the answering parameter of the target person includes:

[0010] Setting priorities for various answering parameters; and

[0011] When the sound intensities indicated by multiple answering parameters conflict, selecting the answering parameter with the highest priority and discarding the answering parameters with lower priorities;

[0012] When the sound intensities indicated by multiple answering parameters do not conflict, performing the step of adjusting the gain of the sound collection device according to the answering parameter.

[0013] Optionally, obtaining the answering parameter of the target person includes:

[0014] Obtaining the corresponding relationship between the shooting focal length and the position of the target terminal;

[0015] Obtaining the shooting focal length of the target terminal when imaging the current video image, and obtaining the position of the target person in the current video image according to the corresponding relationship.

[0016] Optionally, the call voice processing method further includes:

[0017] Detecting whether the target person continues to be displayed in the current video image;

[0018] When it is detected that the target person disappears in the current video image, collecting the current voice of the target person and determining the position of the target person in the current video image accordingly.

[0019] Optionally, the position of the target person in the video image includes: the face of the target person is located in the central area, the left half, or the right half of the video image;

[0020] The answering parameter is the position of the target person in the video image. Adjusting the gain of the sound collection device according to the answering parameter includes:

[0021] When the face of the target person is located in the left half of the video image, increasing the gain of the left channel of the sound collection device; when the face of the target person is located in the right half of the video image, increasing the gain of the right channel of the sound collection device; when the face of the target person is located in the central area of the video image, keeping the gains of the left channel and the right channel of the sound collection device unchanged.

[0022] Optionally, the action of the target person includes: the ear of the target person is facing the target terminal;

[0023] The answering parameter is the action of the target person. Adjusting the gain of the sound collection device according to the answering parameter includes: increasing the gain of the sound collection device.

[0024] Optionally, the voice of the target person includes: a language segment indicating the sound size;

[0025] The answering parameter is the voice of the target person. Adjusting the gain of the sound collection device according to the answering parameter includes:

[0026] When a language segment indicating a small sound is obtained, increasing the gain of the sound collection device; when a language segment indicating a large sound is obtained, decreasing the gain of the sound collection device.

[0027] A communication terminal provided by the present application, the communication terminal includes an application processor, a digital signal processor, and a sound collection device.

[0028] The application processor is used to obtain the current video image.

[0029] The digital signal processor is used to identify the target person in the current video image and obtain the answering parameter of the target person. The answering parameter is used to identify the sound intensity heard by the target person and includes at least one of the position of the target person in the current video image, the action of the target person, and the voice of the target person; and,

[0030] The application processor is further used to adjust the gain of the sound collection device according to the answering parameter and transmit the voice signal collected by the adjusted sound collection device to the target terminal.

[0031] Optionally, the application processor is further used to set priorities for various answering parameters.

[0032] When the sound strengths indicated by multiple answering parameters conflict, select the answering parameter with the highest priority and discard the answering parameters with lower priorities.

[0033] When the sound strengths indicated by multiple answering parameters do not conflict, adjust the gain of the sound collection device according to the answering parameter.

[0034] A readable storage medium provided by the present application stores a program, and this program is used to be run by a processor to execute one or more steps in any of the above video call voice processing methods.

[0035] The present application adjusts the gain of the sound collection device according to the answering parameter related to the sound of the target person in the video image, including at least one of the position of the target person in the current video image, the action of the target person, and the voice of the target person. This answering parameter identifies the real-time feedback of the other party on the call volume in the voice call scenario, so that the input gain of its own microphone can be adjusted according to the situation of the other party, which is beneficial to providing high-quality voice call services. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those skilled in the art, without creative efforts, other drawings can be obtained according to these drawings.

[0037] Figure 1 is a schematic flowchart of a video call voice processing method according to an embodiment of the present application.

[0038] Figure 2 It is a schematic diagram of the position of the target person in the current video image;

[0039] Figure 3 It is a schematic flow chart of the video call voice processing method according to another embodiment of the present application;

[0040] Figure 4 It is a schematic structural diagram of a communication terminal according to an embodiment of the present application. Detailed implementation manners

[0041] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments, rather than all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts fall within the protection scope of the present application. Without conflict, the following various embodiments and their technical features can be combined with each other.

[0042] It should be noted that in the description herein, step codes such as S11 and S12 are used. The purpose is to more clearly and briefly express the corresponding content and does not constitute a substantial limitation in order. Those skilled in the art may execute S12 first and then S11 in specific implementation, etc., but these should all be within the protection scope of the present application.

[0043] Figure 1 It is a schematic flow chart of the video call voice processing method according to an embodiment of the present application. The video call voice processing method can be applied to mobile Internet devices (MID) such as smart phones (Android phones, iOS phones, etc.), tablet computers, PDAs (Personal Digital Assistants), learning machines, etc., or wearable devices with audio and video call functions that can be worn on human limbs, prosthetics or embedded in clothing, jewelry, accessories, etc. The embodiments of the present application are not limited thereto.

[0044] In the audio and video call scenario, the execution subject of each step in this method can be any one of the two devices in the call. When the execution subject is the calling device, the called device is the target terminal; when the execution subject is the called device, the calling device is the target terminal. For the convenience of distinction and description, the execution subject of each step is referred to as the subject terminal in this article.

[0045] Please refer to Figure 1 , the video call voice processing method may include steps S11 to S13.

[0046] S11: Identify the target person in the current video image.

[0047] The current video image is the image displayed in real time on the main terminal. The target person is the person displayed in real time in the current video image, i.e., the called party. The target terminal captures the target person and their current environment through its own (front or rear) camera and forms an image, and transmits it to the main terminal. The main terminal can identify the target person in the image through face recognition technology, human body recognition technology, human body detection technology, and human body posture / behavior / motion recognition technology.

[0048] During a call, when multiple people appear in the current video image, the main terminal obtains the mouth features of each person through face recognition technology, determines the person currently speaking based on the mouth features, and uses their face as the target face. For example, the face with the mouth opening and closing is used as the target face. Accordingly, each person is allowed to become the target person.

[0049] Alternatively, the main terminal selects one of these faces as the target face through face recognition technology and discards the other faces. Accordingly, only one person is the target person. For example, when the caller is on a call with person A 0 and person B 0 enters the camera view range of the target terminal, and at this time person B who is speaking 0 may just be passing by and not on a call with the caller. In this case, the main terminal can only use person A 0 as the target face.

[0050] Among them, in addition to facial features, the reference basis for the main terminal to select the target person can also be combined with parameters that can identify the unique identity of the person, such as voiceprint features.

[0051] S12: Obtain the answering parameter of the target person. The answering parameter is used to identify the sound intensity heard by the target person and includes at least one of the position of the target person in the current video image, the action of the target person, and the voice of the target person.

[0052] S13: Adjust the gain of the sound collection device according to the answering parameter, and transmit the voice signal collected by the adjusted sound collection device to the target terminal.

[0053] The so-called answering parameter refers to the parameter that can identify the sound intensity (or the volume of the speaking voice), which is for the sound of the caller heard by the target person. If the answering parameter indicates that the currently heard sound is small, it means that the sound of the caller collected by the main terminal is small, then the main terminal can increase the gain of its own sound collection device (such as a microphone). If the answering parameter indicates that the currently heard sound is large, it means that the sound of the caller collected by the main terminal is large, then the main terminal can reduce the gain of its own sound collection device.

[0054] The answering parameter identifies the real-time feedback of the other party on the call volume in a voice call scenario, so that the embodiments of the present application can adjust the input gain of its own microphone according to the situation of the other party, which is beneficial to providing high-quality voice call services.

[0055] Among them, the specific value of the gain to be adjusted can be determined according to the value of the answering parameter, and the specific algorithm adopted is not limited by the embodiments of the present application.

[0056] For the three types of answering parameters exemplified in the above step S12, the following describes how to adjust the gain of the sound collection device according to each type of answering parameter.

[0057] Please refer to Figure 2 , the position of the target person in the video image can be divided into three types: the face of the target person (hereinafter referred to as the target face) is located in the central area, the left half, and the right half of the video image.

[0058] When the target face is located in the left half of the video image, for example Figure 2 at the position A shown in

[0059] Figure 2 at the position B shown in

[0060] When the target face is located in the right half of the video image, for example Figure 2 at the position C shown in

[0061] When the target face is located in the central area of the video image, for example the main body terminal can keep the gains of the left and right channels of the sound collection device unchanged.

[0062] The action of the target person can be a body action that can reflect the target person's perception of the listening sound size, and can be that the ear (left ear or right ear) of the target person faces the target terminal. Further, there may also be accompanying body actions such as covering the ear with a hand or shaking the head.

[0063] If in the current video image, the left ear or right ear of the target person faces the target terminal, the main body terminal increases the gain of the sound collection device.

[0064] When a language segment indicating a low voice is obtained, such as "The voice is too low, I can't hear clearly" or "Louder", the host terminal increases the gain of the sound collection device. When a language segment indicating a high voice is obtained, the host terminal decreases the gain of the sound collection device.

[0065] It should be understood that in the dimension of voice, the voice intensity of the target person can also be used as an answering parameter. For example, the current voice intensity of the target person can be compared with the previous voice intensity. Specifically, if the voice intensity becomes smaller, the gain of the sound collection device is increased; if the voice intensity becomes larger, the gain of the sound collection device is decreased.

[0066] For other types of answering parameters, the host terminal can obtain them in a suitable manner. For example, for the position of the target face in the current video image, the host terminal can obtain the corresponding relationship between the shooting focal length and the position of the target terminal from the target terminal, then obtain the shooting focal length of the target terminal when imaging the current video image, and obtain the position of the target face in the current video image according to the corresponding relationship; or, obtain its position in the video image according to the voice intensity of the target person. Specifically, the host terminal obtains the corresponding relationship between the voice intensity and the distance in advance, and then obtains the corresponding position according to the currently obtained voice intensity.

[0067] In the foregoing embodiment, the host terminal adjusts the gain of the sound collection device according to the face image in the video image. For the situation where the target person leaves the camera view range during a call, that is, when it is detected that the target person disappears in the current video image, the host terminal cannot obtain the position of the target person in the video image and the actions of the target person. At this time, the gain of the sound collection device can be adjusted through the voice of the target person. For example, compare the current voice intensity of the target person with the previous voice intensity, and directly adjust the gain through the change in voice intensity, or obtain the position of the target face in the current video image through the change in voice intensity, and then adjust the gain according to the position change.

[0068] In the embodiments of the present application, the host terminal can obtain multiple types of answering parameters and comprehensively adjust the gain according to these answering parameters. However, there may be conflicts in the volume sizes indicated by these answering parameters. For example, when the target person emits a language segment such as "The voice is too loud, don't shout at me", the target person turns his head sideways towards the target terminal because of itchy ears and has a limb movement of covering his ears with his hand. At this time, the host terminal needs to determine which parameter to adjust the gain according to.

[0069] In response to this, the embodiments of the present application can provide the following Figure 3 method as shown. As Figure 3 shown, the video call voice processing method includes steps S21 to S25.

[0070] S21: Identify the target person in the current video image.

[0071] S22: Obtain the answering parameters of the target person, where the answering parameters are used to identify the sound intensity heard by the target person and include at least two of the position of the target person in the current video image, the actions of the target person, and the speech of the target person.

[0072] S23: Set priorities for various answering parameters.

[0073] S24: When the sound strengths identified by multiple answering parameters conflict, select the answering parameter with the highest priority and discard the answering parameters with lower priorities; when the sound strengths identified by multiple answering parameters do not conflict, select all the collected answering parameters.

[0074] S25: Adjust the gain of the sound collection device according to the answering parameters, and transmit the voice signal collected by the adjusted sound collection device to the target terminal.

[0075] Based on the description of the foregoing embodiments, this embodiment can avoid the situation of mis-adjusting the gain caused by misjudgment of the main terminal, and more accurately feedback the answering situation of the target person.

[0076] Figure 4 is a schematic structural diagram of a communication terminal according to an embodiment of the present application. Please refer to Figure 4 As shown, the communication terminal 40 may be a device of one of the two parties in a video call, such as the foregoing main terminal. The communication terminal 40 includes an application processor 41, a digital signal processor 42, a sound collection device 43, a camera 44, a left speaker 451, a right speaker 452, and an antenna 46. The application processor 41 and the digital signal processor 42 can be regarded as the core of the communication terminal 40, and are connected to the respective structural elements to implement corresponding functions during the video call process.

[0077] The camera 44 is used to capture the imaging of the call participants and their surrounding environment.

[0078] The antenna 46 is, for example, a Wi-Fi antenna, etc., and is used to receive and transmit electromagnetic waves, realize the mutual conversion between electromagnetic waves and electrical signals, so as to communicate with the device of the other party in the video call.

[0079] The sound collection device 43, such as a microphone, is used to collect the speech of the call participants.

[0080] The left speaker 451 and the right speaker 452 are used to play the speech of the other party in the video call, and respectively play the left-channel voice signal and the right-channel voice signal.

[0081] The application processor 41 is used to obtain the current video image and transmit it to the digital signal processor 42.

[0082] The digital signal processor 42 is used to activate the corresponding algorithm to identify the target person in the current video image and obtain the answering parameters of the target person. The answering parameters are used to identify the sound intensity heard by the target person and include at least one of the position of the target person in the current video image, the action of the target person, and the voice of the target person.

[0083] Specifically, the application processor 41 sends the current video image to the digital signal processor 42, and the digital signal processor 42 activates the corresponding algorithm to identify the position of the target person in the current video image and the action of the target person. The application processor 41 transmits the voice of the target person to the digital signal processor 42, and the digital signal processor 42 activates the speech recognition algorithm to obtain the corresponding parameters, such as identifying whether there are special language segments in the speech signal, such as "your voice is too low", etc., and returns the detection result to the application processor 41.

[0084] The application processor 41 is used to generate a gain scheme according to the answering parameters, adjust the gain of the sound collection device 43 accordingly, and transmit the voice signal collected by the adjusted sound collection device to the target terminal through the antenna 46.

[0085] For the specific working modes of each structural element, reference can be made to the steps of the foregoing method, which will not be elaborated here one by one. For example, the application processor 41 is also used to set priorities for various answering parameters. When the sound strengths indicated by multiple answering parameters conflict, select the answering parameter with the highest priority and discard the answering parameters with lower priorities; when the sound strengths indicated by multiple answering parameters do not conflict, adjust the gain of the sound collection device 43 according to the answering parameters.

[0086] Herein, the communication terminal 40 has the beneficial effects achievable by the foregoing method.

[0087] It should be understood that in actual implementation in practical application scenarios, according to the device type to which the communication terminal 40 belongs, the execution entities of the above steps may not be the foregoing structural elements, but are respectively implemented by other modules and units.

[0088] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructions, or by controlling relevant hardware through instructions. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor. To this end, an embodiment of the present application provides a readable storage medium. Multiple instructions are stored in the readable storage medium, and these instructions can be loaded by a processor to execute the steps in any one of the video call voice processing methods provided by the embodiments of the present application.

[0089] Among them, the storage medium may include a read-only memory (ROM, Read Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, an optical disk, etc.

[0090] Since the instructions stored in the storage medium can execute the steps in any one of the video call voice processing methods provided by the embodiments of the present application, the beneficial effects achievable by any one of the video call voice processing methods can be realized. For details, see the previous embodiments.

[0091] An embodiment of the present application also provides a computer program product. The computer program product includes computer program code. When the computer program code runs on a computer, it causes the computer to execute the methods described in the above various possible embodiments.

[0092] An embodiment of the present application also provides a chip, including a memory and a processor. The memory is used to store a computer program, and the processor is used to call and run the computer program from the memory, so that a device installed with the chip executes the methods in the above various possible embodiments.

[0093] It should be noted that the term "including" or "comprising" or any other variant thereof in this article is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including that element. In addition, components, features, and elements with the same name in different embodiments may have the same meaning or different meanings, and their specific meanings need to be determined by their explanations in the specific embodiments or further in combination with the context of the specific embodiments.

[0094] In addition, although the terms "first", "second", "third", etc. are used herein to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of this document, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information, depending on the context. The term "if" can be interpreted as "when" or "while" or "in response to a determination". Furthermore, as used herein, the singular forms "a", "an", and "the" are intended to also include the plural forms, unless the context indicates otherwise. The terms "or" and "and / or" are interpreted inclusively, meaning either one or any combination. Thus, "A, B, or C" or "A, B, and / or C" means "any of the following: A; B; C; A and B; A and C; B and C; A, B, and C". An exception to this definition occurs only when the combination of elements, functions, steps, or operations is inherently mutually exclusive in some way.

[0095] Furthermore, although the steps in the flowcharts herein are shown sequentially as indicated by the arrows, they are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this document, these steps are not executed strictly in sequence and may be executed in other orders. Moreover, at least some of the steps in the figures may include multiple sub-steps or multiple phases, and these sub-steps or phases are not necessarily executed at the same time but can be executed at different times, and their execution order is not necessarily sequential but can be executed alternately or in rotation with at least a part of other steps or sub-steps or phases of other steps.

[0096] Although the present application has been shown and described with respect to one or more implementations, equivalent variations and modifications will occur to those skilled in the art based on a reading and understanding of this specification and the drawings. The present application includes all such modifications and variations and is supported by the technical solutions of the foregoing embodiments. That is, the above are only some embodiments of the present application, and thus do not limit the patent scope of the present application. Any equivalent structural transformation made using the content of this specification and the drawings, such as the combination of technical features between embodiments, or direct or indirect application in other related technical fields, is similarly included within the patent protection scope of the present application.

Claims

1. A method for processing video call voice, characterized in that, it includes: Identifying the target person in the current video image; Obtaining the answering parameter of the target person, where the answering parameter is used to identify the sound intensity heard by the target person, and includes at least one of the position of the target person in the current video image, the action of the target person, and the voice of the target person; Adjusting the gain of the sound collection device according to the answering parameter, and transmitting the voice signal collected by the adjusted sound collection device to the target terminal, where obtaining the answering parameter of the target person includes: Setting priorities for various answering parameters; and When the sound strengths identified by multiple answering parameters conflict, selecting the answering parameter with the highest priority and discarding the answering parameters with lower priorities; When the sound strengths identified by multiple answering parameters do not conflict, performing the step of adjusting the gain of the sound collection device according to the answering parameter.

2. The video call voice processing method according to claim 1, characterized in that, obtaining the answering parameter of the target person includes: Obtaining the correspondence between the shooting focal length and position of the target terminal; Obtaining the shooting focal length of the target terminal when imaging the current video image, and obtaining the position of the target person in the current video image according to the correspondence.

3. The video call voice processing method according to claim 1, characterized in that, the video call voice processing method further includes: Detecting whether the target person continues to be displayed in the current video image; When it is detected that the target person disappears in the current video image, collecting the current voice of the target person, and determining the voice intensity change according to the current voice and the historical voice collected when it is detected that the target person disappears in the current video image, so as to determine the position of the target person in the current video image based on this voice intensity change.

4. The video call voice processing method according to claim 1, characterized in that, the position of the target person in the video image includes: the face of the target person is located in the central area, the left half, or the right half of the video image; when the answering parameter is the position of the target person in the video image, adjusting the gain of the sound collection device according to the answering parameter includes: When the face of the target person is located in the left half of the video image, increasing the gain of the left channel of the sound collection device; when the face of the target person is located in the right half of the video image, increasing the gain of the right channel of the sound collection device; when the face of the target person is located in the central area of the video image, keeping the gains of the left channel and the right channel of the sound collection device unchanged.

5. The video call voice processing method according to claim 1, characterized in that, the action of the target person includes: the ear of the target person is facing the target terminal; when the answering parameter is the action of the target person, adjusting the gain of the sound collection device according to the answering parameter includes: increasing the gain of the sound collection device.

6. The video call voice processing method according to claim 1, characterized in that, the voice of the target person includes: a language segment indicating the sound size; The answering parameter is the voice of the target person, and adjusting the gain of the sound acquisition device according to the answering parameter includes: When a language segment indicating a low sound is obtained, increasing the gain of the sound acquisition device; when a language segment indicating a high sound is obtained, decreasing the gain of the sound acquisition device.

7. A communication terminal, characterized in that, the communication terminal includes an application processor, a digital signal processor, and a sound acquisition device, the application processor is configured to obtain a current video image; the digital signal processor is configured to identify a target person in the current video image and obtain an answering parameter of the target person, where the answering parameter is used to indicate the sound intensity heard by the target person and includes at least one of the position of the target person in the current video image, the action of the target person, and the voice of the target person; and, the application processor is further configured to adjust the gain of the sound acquisition device according to the answering parameter and transmit the voice signal collected by the sound acquisition device after adjustment to the target terminal, where the digital signal processor is further configured to: set priorities for various answering parameters; and when the sound strengths indicated by multiple answering parameters conflict, select the answering parameter with the highest priority and discard the answering parameters with lower priorities; when the sound strengths indicated by multiple answering parameters do not conflict, perform the step of adjusting the gain of the sound acquisition device according to the answering parameter.

8. A readable storage medium, characterized in that, a program is stored in the readable storage medium, and the program is used to be run by a processor to execute one or more steps in the video call voice processing method according to any one of the above claims 1 to 6.

Citation Information

Patent Citations

  • Voice signal processing method, device, and mobile terminal

    CN105825854A

  • Audio device control method and device

    CN109144466A