Sound processing method and related apparatus

By responding to speech enhancement operations in electronic devices, the pickup area is accurately located and non-target area sounds are suppressed, solving the problem of speech interference by noise in complex acoustic environments and improving the accuracy of speech separation and recognition.

WO2026158054A1PCT designated stage Publication Date: 2026-07-30HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
HUAWEI TECH CO LTD
Filing Date
2026-01-09
Publication Date
2026-07-30

AI Technical Summary

Technical Problem

In complex acoustic environments, the speaker's voice is easily affected by reverberation noise and interfering human voices in the surrounding environment, leading to a decrease in the listening experience of remote participants and the accuracy of speech-to-text conversion.

Method used

By responding to speech enhancement operations through electronic devices, the pickup area containing the location of the target object is determined, and the sound from outside the pickup area in the sound collected by the audio acquisition module is suppressed. By combining a microphone array and an image acquisition module, the pickup area is accurately located, thereby enhancing the speech separation effect of the target object.

Benefits of technology

It effectively reduces the impact of ambient noise, improves speech separation and speech content recognition accuracy, and enhances the listening experience and speech-to-text accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2026071680_30072026_PF_FP_ABST
    Figure CN2026071680_30072026_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of computers. Disclosed are a sound processing method and a related apparatus. The method comprises: in response to a received speech enhancement operation, an electronic device determining a sound pickup area, wherein the sound pickup area includes the position of a target object; and the electronic device performing suppression on sounds outside the sound pickup area among sounds collected by an audio collection module, so as to obtain speech of the target object. A sound pickup area determined in the present application includes the position of a target object, such that a situation where the voice of the target object is suppressed can be avoided, and sounds outside a sound pickup area can be suppressed, thereby reducing the impact of reverberation noise and interfering human voices in a surrounding environment. Relatively speaking, speech of the target object, who is located within the sound pickup area, can be enhanced, thereby improving the speech separation effect and facilitating an improvement in the accuracy of speech content recognition.
Need to check novelty before this filing date? Find Prior Art

Description

A sound processing method and related apparatus

[0001] Cross-references to related applications

[0002] This application claims priority to Chinese Patent Application No. 202510112375.5, filed on January 22, 2025, entitled "A Sound Processing Method and Related Apparatus", the entire contents of which are incorporated herein by reference.

[0003] This application claims priority to Chinese Patent Application No. 202610019319.1, filed on January 7, 2026, entitled "A Sound Processing Method and Related Apparatus", the entire contents of which are incorporated herein by reference. Technical Field

[0004] This application relates to the field of computer technology, and in particular to a sound processing method and related apparatus. Background Technology

[0005] In conference scenarios centered around a large conference screen or in open-space discussion settings, audio capture modules, such as microphones, can pick up the speaker's voice for communication or interaction. For example, in a remote conference scenario centered around a large conference screen, the speaker's voice, after being captured by the audio capture module, can be transmitted over the network to the remote participants' terminals for communication; or, after being captured by the audio capture module, the speaker's voice can be transcribed into text and saved, preserving the content of the speaker's speech during the meeting.

[0006] When the speaker's voice is captured by the audio acquisition module, the speaker's voice is often affected by reverberation noise and interference from the surrounding environment. In complex acoustic environments, the listening experience of remote participants will be affected, and the accuracy of speech-to-text conversion will also decrease. Summary of the Invention

[0007] This application provides a sound processing method and related apparatus, which can reduce the impact of noise on the speech of the target object, improve the speech separation effect, and help improve the accuracy of speech content recognition.

[0008] In a first aspect, this application provides a sound processing method, which can be executed by an electronic device, or by a chip, chip system, or circuit in the electronic device, wherein the electronic device includes an audio acquisition module or is connected to an audio acquisition module. The sound processing method may include: the electronic device, in response to a received speech enhancement operation, determining a pickup area, the pickup area including the location of a target object; and the electronic device suppressing sounds outside the pickup area from the sound collected by the audio acquisition module to obtain the speech of the target object.

[0009] The sound processing method provided in this application, after receiving a speech enhancement operation, determines a pickup area containing the location of the target object, and suppresses sounds outside the pickup area to obtain the speech of the target object. The pickup area determined by this application includes the location of the target object, which avoids suppressing the target object's sound. Suppressing sounds outside the pickup area reduces the influence of reverberation noise and interference from surrounding voices. Conversely, it enhances the speech of the target object located within the pickup area, improves speech separation, and enhances the accuracy of speech content recognition.

[0010] In one alternative implementation, the voice enhancement operation is a press operation on the voice enhancement control; or, the voice enhancement operation is a press operation on the display screen.

[0011] The above implementation method allows users to quickly activate the voice enhancement function by pressing the voice enhancement control or by pressing the display screen with their hands. This solves the problem of inconvenience and slowness when users frequently turn the voice enhancement function on and off, allowing users to turn the voice enhancement function on and off through simple interactive operations.

[0012] In one alternative implementation, the voice enhancement operation is a press operation on the display screen. When the electronic device receives the press operation on the display screen, and the duration of the press operation reaches a set duration threshold, the sound pickup area is determined.

[0013] The above implementation determines the sound pickup area only when the operation of pressing the display screen is received and the duration of the operation of pressing the display screen reaches the set duration threshold, which can avoid the activation of the voice enhancement function due to user misoperation.

[0014] In one alternative implementation, the electronic device, in response to a voice enhancement operation, can determine a pickup area based on a first set value and a second set value. The first set value characterizes the arm length of the target object, and the second set value characterizes the width of the display screen.

[0015] In the above implementation method, the range of the sound pickup area is determined based on the length of the target object's arm and the width of the display screen, which can ensure that the target object is within the sound pickup area and avoid the situation of suppressing the target object's voice.

[0016] In one alternative implementation, the audio acquisition module is a microphone array; the pickup area is a column, which includes a cross-section of a target shape, the target shape being circular or semi-circular, and the center of the target shape being the center of the microphone array; the radius of the target shape is determined based on a first set value and a second set value.

[0017] In one alternative implementation, the electronic device may respond to a speech enhancement operation, determine the interaction location of the speech enhancement operation, and determine a pickup area based on the interaction location, which is used to characterize the location of the target object.

[0018] In the above implementation method, the range of the sound pickup area is accurately determined based on the interaction position of the voice enhancement operation, which can ensure that the target object is within the sound pickup area and avoid suppressing the sound of the target object while ensuring that the sound pickup area is not too large.

[0019] In one optional implementation, the audio acquisition module is a microphone array; the pickup area is a cylinder, the cylinder includes a fan-shaped cross section, the vertex of the fan-shaped cross section is the center of the microphone array, the fan-shaped cross section is the outer sector of the target shape; the target shape is a circle or a semi-circle, the center of the target shape is the interactive position of the speech enhancement operation, and the target shape is used to characterize the activity range of the target object.

[0020] In the above implementation, the cross section of the pickup area is the circumscribed sector of the target shape. The target shape is used to characterize the activity range of the target object. Based on the activity range of the target object, the range of the pickup area is adaptively reduced. By suppressing the sound outside the pickup area, interference noise can be further removed. For scenarios where the interference sound is near the speaker, better speech separation effect can be achieved.

[0021] In one alternative implementation, the pickup area is a cylinder, which includes a cross-section of a target shape. The target shape is circular or semi-circular, and the center of the target shape is the interaction position. The target shape is used to characterize the activity range of the target object.

[0022] In the above implementation, the cross-section of the pickup area is a target shape used to characterize the activity range of the target object. The pickup area can be further reduced to achieve better speech separation effect.

[0023] In one alternative implementation, the microphone array includes omnidirectional microphones and the target shape is circular.

[0024] In another alternative implementation, the microphone array includes directional microphones with a semi-circular target shape.

[0025] The above implementation uses a directional microphone. By utilizing the directional microphone's ability to distinguish between front and back, the fan-shaped angle of the pickup area can be reduced by half, which can further improve the effect of speech separation.

[0026] In one alternative implementation, when determining the sound pickup area, the electronic device can perform facial recognition on the image acquired by the image acquisition module to determine the position of the target face, and determine the sound pickup area based on the interaction position and the position of the target face.

[0027] In the above implementation, by performing facial recognition on the image acquired by the image acquisition module, the position of the target face can be determined. The position of the target face represents the position of the target object. Based on the positional relationship between the interaction position and the position of the target object, the sound pickup area can be determined, and the sound pickup area can be further narrowed.

[0028] In one alternative implementation, the microphone array includes an omnidirectional microphone, the pickup area is a cylinder, the cylinder has a semi-circular cross-section, the center of the semi-circle is the interaction position for the voice enhancement operation, the axis of symmetry of the semi-circle is parallel to the arrangement direction of the microphone array, and the position of the target face is located within the pickup area.

[0029] In one alternative implementation, the microphone array includes a directional microphone, the pickup area is a cylinder, the cylinder has a fan-shaped cross-section, the center of the fan is the interaction position of the voice enhancement operation, and the target face is located within the pickup area.

[0030] In one alternative implementation, when the electronic device suppresses sounds from outside the pickup area in the sound collected by the audio acquisition module, it can perform speech separation processing on the sound collected by the audio acquisition module based on the range parameters of the pickup area and the sound characteristics of speech and interference noise to obtain the speech of the target object.

[0031] Secondly, this application provides a sound processing method, which can be executed by an electronic device, or by a chip, chip system, or circuit in the electronic device. The electronic device includes an audio acquisition module or is connected to an audio acquisition module. The sound processing method may include: when the electronic device detects a target object touching the display screen, determining a pickup area based on the interaction position of the touch operation; and the electronic device suppressing sounds outside the pickup area from the sound collected by the audio acquisition module to obtain the speech of the target object.

[0032] In one alternative implementation, the pickup area is a cylinder, and the cross-section of the pickup area is the outer shape of the target object's activity range, or the target object's activity range; the target object's activity range is determined based on the interaction position of the contact operation.

[0033] In one optional implementation, the activity range of the target object is a first shape, which is a circle, a semicircle, or a sector; the center of the first shape is the interaction position of the contact operation, and the radius of the first shape is a first set value. The first set value is greater than or equal to the arm length of the target object.

[0034] In one alternative implementation, the cross-section of the pickup area is the outer shape of the target object's active range; the outer shape can be a circle, a semi-circle, a fan shape, or a rectangle.

[0035] In one alternative implementation, the circumscribed shape is a rectangle, and the center of the circumscribed shape is the interaction position of the contact operation.

[0036] In one alternative implementation, the audio acquisition module is a microphone array; the external shape is circular, semi-circular, or fan-shaped, and the center of the external shape is the center of the microphone array.

[0037] In one alternative implementation, the interaction point for the touch operation is located at the edge of the display screen, and the radius of the outer shape is determined based on a first set value and a second set value. The first set value is greater than or equal to the arm length of the target object; the second set value is greater than or equal to half the width of the display screen.

[0038] In one alternative implementation, the electronic device can perform facial recognition on the image acquired by the image acquisition module to determine the location of the target face; and determine the sound pickup area based on the interaction location of the touch operation and the location of the target face.

[0039] In one alternative implementation, the target face is located within the pickup area.

[0040] In one alternative implementation, the electronic device determines that a contact operation has been detected when the duration of contact between the target object and the display screen reaches a set duration threshold.

[0041] In one alternative implementation, the electronic device can perform speech separation processing on the sound collected by the audio acquisition module based on the range parameters of the pickup area and the sound characteristics of speech and interference noise, to obtain the speech of the target object.

[0042] Thirdly, this application provides a sound processing method, which can be executed by an electronic device, or by a chip, chip system, or circuit in the electronic device. The electronic device includes an audio acquisition module or is connected to an audio acquisition module. The sound processing method may include: when the electronic device detects a target object touching the display screen, determining a pickup area based on set parameters; and the electronic device suppressing sounds outside the pickup area collected by the audio acquisition module to obtain the speech of the target object. The set parameters include a first set value and a second set value.

[0043] In one alternative implementation, the electronic device controls the pickup area to be fully open when it detects that the target object has stopped touching the display screen.

[0044] In one alternative implementation, the electronic device determines the sound pickup area based on a second preset value when no target object is detected touching the display screen; and determines the sound pickup area based on the second preset value when the target object stops touching the display screen.

[0045] In one alternative implementation, the pickup area is a cylinder, and the cross-section of the pickup area can be of any shape.

[0046] In one alternative implementation, the cross-section of the pickup area has a second shape, which can be circular, semi-circular, fan-shaped, or rectangular.

[0047] In one alternative implementation, the audio acquisition module is a microphone array; the second shape is circular, semi-circular, or fan-shaped, with the center of the second shape being the center of the microphone array.

[0048] In one alternative implementation, the electronic device can perform speech separation processing on the sound collected by the audio acquisition module based on the range parameters of the pickup area and the sound characteristics of speech and interference noise, to obtain the speech of the target object.

[0049] Fourthly, this application provides a sound processing apparatus that can be applied to electronic devices, the sound processing apparatus including:

[0050] The pickup area determination unit is used to determine the pickup area in response to the received speech enhancement operation; the pickup area includes the location of the target object;

[0051] The sound processing unit is used to suppress sounds from outside the pickup area in the sound collected by the audio acquisition module to obtain the speech of the target object.

[0052] In one alternative implementation, the voice enhancement operation is a press operation on the voice enhancement control; or, the voice enhancement operation is a press operation on the display screen.

[0053] In one alternative implementation, the voice enhancement operation is a pressing operation on the display screen. The pickup area determination unit is specifically used to: determine the pickup area when a pressing operation on the display screen is received, and the duration of the pressing operation reaches a set duration threshold.

[0054] In one alternative implementation, the pickup area determination unit is specifically used to: determine the pickup area based on a first set value and a second set value in response to a voice enhancement operation. The first set value is used to characterize the arm length of the target object, and the second set value is used to characterize the width of the display screen.

[0055] In one alternative implementation, the audio acquisition module is a microphone array; the pickup area is a column, which includes a cross-section of a target shape, the target shape being circular or semi-circular, and the center of the target shape being the center of the microphone array; the radius of the target shape is determined based on a first set value and a second set value.

[0056] In one alternative implementation, the pickup area determination unit is specifically used for:

[0057] In response to a speech enhancement operation, determine the interaction location of the speech enhancement operation; the interaction location is used to characterize the location of the target object.

[0058] The pickup area is determined based on the interaction location.

[0059] In one optional implementation, the audio acquisition module is a microphone array; the pickup area is a column, the column includes a fan-shaped cross section, the vertex of the fan-shaped cross section is the center of the microphone array, the fan-shaped cross section is the circumscribed fan of the target shape; the target shape is a circle or a semi-circle, the center of the target shape is the interaction position, and the target shape is used to characterize the activity range of the target object.

[0060] In one alternative implementation, the pickup area is a cylinder, which includes a cross-section of a target shape. The target shape is circular or semi-circular, and the center of the target shape is the interaction position. The target shape is used to characterize the activity range of the target object.

[0061] In one alternative implementation, the microphone array includes omnidirectional microphones with a circular target shape; or, the microphone array includes directional microphones with a semi-circular target shape.

[0062] In one alternative implementation, the pickup area determination unit is specifically used for:

[0063] The image acquisition module performs facial recognition on the images it acquires to determine the location of the target face. Based on the interaction location and the location of the target face, the sound pickup area is determined.

[0064] In one alternative implementation, the microphone array includes an omnidirectional microphone, the pickup area is a cylinder, the cylinder has a semi-circular cross-section, the center of the semi-circle is the interaction position for the voice enhancement operation, the axis of symmetry of the semi-circle is parallel to the arrangement direction of the microphone array, and the position of the target face is located within the pickup area.

[0065] In one alternative implementation, the microphone array includes a directional microphone, the pickup area is a cylinder, the cylinder has a fan-shaped cross-section, the center of the fan is the interaction position of the voice enhancement operation, and the target face is located within the pickup area.

[0066] In one alternative implementation, the sound processing unit is specifically used for:

[0067] Based on the range parameters of the pickup area and the sound characteristics of speech and interference noise, the sound collected by the audio acquisition module is processed by speech separation to obtain the speech of the target object.

[0068] Fifthly, this application provides a sound processing apparatus that can be applied to electronic devices, the sound processing apparatus including:

[0069] The sound pickup area determination unit is used to detect the contact operation of the target object on the display screen and determine the sound pickup area based on the interaction position of the contact operation;

[0070] The sound processing unit is used to suppress sounds from outside the pickup area in the sound collected by the audio acquisition module to obtain the speech of the target object.

[0071] In one alternative implementation, the pickup area is a cylinder, and the cross-section of the pickup area is the outer shape of the target object's activity range, or the target object's activity range; the target object's activity range is determined based on the interaction position of the contact operation.

[0072] In one alternative implementation, the activity range of the target object is a first shape, which is a circle, a semicircle, or a fan shape; the center of the first shape is the interaction position of the contact operation, and the radius of the first shape is a first set value.

[0073] Sixthly, this application provides a sound processing apparatus that can be applied to electronic devices, the sound processing apparatus including:

[0074] The sound pickup area determination unit is used to detect the contact operation of the target object with the display screen and determine the sound pickup area based on the set parameters;

[0075] The sound processing unit is used to suppress sounds from outside the pickup area in the sound collected by the audio acquisition module to obtain the speech of the target object.

[0076] In one alternative implementation, the setting parameters include a first setting value and a second setting value.

[0077] In one alternative implementation, the pickup area determination unit is further configured to: control the pickup area to be fully open when the target object stops touching the display screen.

[0078] In one optional implementation, the set parameters include a first preset value; the pickup area determination unit is further configured to: determine the pickup area based on a second preset value when no target object is detected touching the display screen; determine the pickup area based on the first preset value when a contact operation is detected; and determine the pickup area based on the second preset value when the target object stops touching the display screen.

[0079] In a seventh aspect, this application provides an electronic device, including a processor and a memory; the memory stores a computer program or instructions; the processor executes the computer program or instructions stored in the memory to cause the electronic device to perform any of the sound processing methods provided in the first, second, or third aspects above.

[0080] Eighthly, this application provides a system including a processor and a power supply circuit; the power supply circuit is used to supply power to the processor, and the processor is used to execute a computer program to implement any of the sound processing methods provided in the first, second, or third aspects above.

[0081] Ninthly, this application provides a sound pickup system, which includes the electronic device provided in the third aspect and an audio acquisition module connected to the electronic device.

[0082] In one alternative implementation, the sound pickup system also includes a display screen connected to an electronic device.

[0083] In a tenth aspect, embodiments of this application provide a computer-readable storage medium storing computer-executable instructions, which are used to cause a computer to perform any of the sound processing methods provided in the first, second, or third aspects.

[0084] Eleventhly, embodiments of this application provide a computer program product comprising computer-executable instructions, the computer-executable instructions being used to cause a computer to perform any of the sound processing methods provided in the first, second, or third aspects.

[0085] The technical effects that can be achieved by any of the second to eleventh aspects mentioned above can be referred to the description of the beneficial effects in the first aspect mentioned above, and will not be repeated here. Attached Figure Description

[0086] Figure 1 is a schematic diagram of a sound pickup system provided in an embodiment of this application;

[0087] Figure 2 is a schematic diagram of another sound pickup system provided in an embodiment of this application;

[0088] Figure 3 is a schematic diagram of a sound pickup area in a related technology;

[0089] Figure 4 is a flowchart of a sound processing method provided in an embodiment of this application;

[0090] Figure 5 is a schematic diagram of a display screen provided in an embodiment of this application;

[0091] Figure 6 is a schematic diagram of another display screen provided in an embodiment of this application;

[0092] Figure 7 is a flowchart of another sound processing method provided in an embodiment of this application;

[0093] Figure 8 is a schematic diagram of the judgment process of a voice enhancement operation provided in an embodiment of this application;

[0094] Figure 9 is a schematic diagram of a sound pickup area provided in an embodiment of this application;

[0095] Figure 10 is a schematic diagram of a speech processing procedure provided in an embodiment of this application;

[0096] Figure 11 is a schematic diagram of a face detection process provided in an embodiment of this application;

[0097] Figure 12 is a schematic diagram of another pickup area provided in an embodiment of this application;

[0098] Figure 13 is a schematic diagram of another pickup area provided in an embodiment of this application;

[0099] Figure 14 is a flowchart of another sound processing method provided in an embodiment of this application;

[0100] Figure 15 is a schematic diagram of another pickup area provided in an embodiment of this application;

[0101] Figure 16 is a schematic diagram of another pickup area provided in an embodiment of this application;

[0102] Figure 17 is a schematic diagram of another pickup area provided in an embodiment of this application;

[0103] Figure 18 is a schematic diagram of a sound processing effect provided in an embodiment of this application;

[0104] Figure 19 is a flowchart of another sound processing method provided in an embodiment of this application;

[0105] Figure 20 is a schematic diagram of another pickup area provided in an embodiment of this application;

[0106] Figure 21 is a schematic diagram of another pickup area provided in an embodiment of this application;

[0107] Figure 22 is a flowchart of another sound processing method provided in an embodiment of this application;

[0108] Figure 23 is a schematic diagram of another pickup area provided in an embodiment of this application;

[0109] Figure 24 is a flowchart of another sound processing method provided in an embodiment of this application;

[0110] Figure 25 is a schematic diagram of another pickup area provided in an embodiment of this application;

[0111] Figure 26 is a schematic diagram of another pickup area provided in an embodiment of this application;

[0112] Figure 27 is a schematic diagram of another pickup area provided in an embodiment of this application;

[0113] Figure 28 is a schematic diagram of a data synchronization device provided in an embodiment of this application;

[0114] Figure 29 is a schematic diagram of the structure of a chip provided in an embodiment of this application. Detailed Implementation

[0115] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the embodiments of this application will be described in detail below with reference to the accompanying drawings. The terminology used in the implementation section of this application is only for explaining specific embodiments of this application and is not intended to limit this application.

[0116] Before introducing the specific solutions provided in the embodiments of this application, some terms used in this application will be explained to facilitate understanding by those skilled in the art, but the terms used in this application are not limited.

[0117] (1) Microphone array: A microphone is an acoustic sensor used to collect sound and convert it into electrical signals. It is also called a microphone or transducer. A microphone array is made up of a certain number of microphones arranged in a set pattern. It is used to sample and process the spatial characteristics of the sound field.

[0118] (2) Omnidirectional microphone: It can capture sound from the surrounding environment in 360°, and is suitable for occasions such as meetings and interviews that require all-round sound recording.

[0119] (3) Directional microphone: Used to collect sound from a specified direction, mainly capturing sound from the front of the microphone. It is suitable for occasions such as lectures and singing where the speaker's or singer's voice needs to be highlighted. A directional microphone has a pickup diaphragm with an opening at each end. It can pick up sound based on the pressure difference between the two ends and is more sensitive to sound from the front than sound from the back.

[0120] In this application embodiment, "multiple" refers to two or more. Therefore, in this application embodiment, "multiple" can also be understood as "at least two". "At least one" can be understood as one or more, such as one, two, or more. For example, "including at least one" means including one, two, or more, and it does not limit which ones are included. For example, including at least one of A, B, and C, then it could include A, B, C, A and B, A and C, B and C, or A and B and C. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / ", unless otherwise specified, generally indicates that the preceding and following related objects have an "or" relationship.

[0121] Unless otherwise stated, the ordinal numbers such as "first" and "second" mentioned in the embodiments of this application are used to distinguish multiple objects, and are not used to limit the order, sequence, priority or importance of multiple objects.

[0122] The sound processing method provided in this application embodiment can be applied to the sound pickup system shown in Figure 1. This sound pickup system can be set up in a conference scenario, an open space discussion scenario, or a remote conference scenario. As shown in Figure 1, the sound pickup system 100 may include an electronic device 110, an audio acquisition module 120 connected to the electronic device 110, and a display screen 130. The audio acquisition module 120 may be a microphone array, which can be an omnidirectional microphone array or a directional microphone array. The audio acquisition module 120 may be installed at the top center of the display screen 130, or at the bottom center of the display screen 130, for collecting sound from the meeting room, converting the sound into electrical signals, and transmitting them to the electronic device 110.

[0123] The display screen 130 may include a display panel. The display panel may be a liquid crystal display (LCD), a light-emitting diode (LED), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a quantum dot light-emitting diode (QLED), etc. The display panel can display images or videos to the audience in the venue.

[0124] In some embodiments, the display screen 130 may be a touch screen, including a touch sensor and the aforementioned display panel. The touch sensor may also be referred to as a "touch device." The touch sensor can be used to detect touch operations applied to the touch sensor and determine the location of the touch operation.

[0125] Electronic device 110 may include processor 111 and memory 112. Processor 111 may include one or more processing units, such as: application processor (AP), modem processor, graphics processing unit (GPU), image signal processor (ISP), controller, video codec, digital signal processor (DSP), baseband processor, and / or neural network processing unit (NPU), etc. The controller can generate operation control signals based on instruction opcodes and timing signals to control instruction fetching and execution. Different processing units can be independent devices or integrated into one or more processors. Processor 111 can receive sound transmitted from audio acquisition module 120 and process the sound.

[0126] The memory 112, also known as internal memory, can be used to store computer executable program code, including instructions. The internal memory may include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function, etc. The data storage area may store data (such as sound data) created during the use of the processor 111. Furthermore, the internal memory may include high-speed random access memory and non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc. The processor 111 executes various functional applications and data processing of the electronic device 110, such as the sound processing method provided in the embodiments of this application, by running instructions stored in the internal memory and / or instructions stored in memory disposed within the processor.

[0127] In some embodiments, the electronic device 110 may further include an external memory interface, which can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 110. The external memory card communicates with the processor 111 through the external memory interface to perform data storage functions. For example, after the speaker's voice is picked up by the audio acquisition module, it can be recorded by converting speech to text and saved to the external memory card.

[0128] In some embodiments, the electronic device 110 may also include a speaker or an external speaker, also known as a "loudspeaker". The speaker is connected to the processor 111 and can amplify and play the sound processed by the processor 111 so that the audience in the venue can hear the speaker's voice.

[0129] In some optional embodiments, multiple meeting rooms may exist, and the electronic devices in the multiple meeting rooms can communicate with each other via a network. As shown in Figure 2, electronic device 110 may include a communication module 113. For example, electronic device 110 is located in a first meeting room, and electronic device 210 is located in a second meeting room. After electronic device 110 and electronic device 210 establish a communication connection, electronic device 110 can send processed sound to electronic device 210 through communication module 113. Electronic device 210 can then play the sound through a speaker in the second meeting room, allowing the audience in the second meeting room to remotely participate in the meeting and hear the speaker's voice in the first meeting room. In one embodiment, electronic device 210 may not be located in a meeting room, but rather as a terminal of a remote participant. Electronic device 210 plays sound from electronic device 110, allowing that participant to remotely participate in the meeting and hear the speaker's voice in the meeting room.

[0130] In some embodiments, the sound pickup system 100 may further include an image acquisition module 140 connected to the electronic device 110. The image acquisition module 140 may be a camera, used to acquire images of the venue and transmit them to the electronic device 110 for processing by the processor 111 within the electronic device 110. Exemplarily, the image acquisition module 140 may include a lens and a photosensitive element. An object is projected onto the photosensitive element through the lens to generate an optical image. The photosensitive element may be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, which is then transmitted to the processor 111 of the electronic device 110 for processing.

[0131] It is understood that the application scenarios shown in Figures 1 and 2 are provided as examples and do not constitute a specific limitation on the application scenarios. The sound processing method provided in this application embodiment can also be applied to other application scenarios that require sound noise reduction processing.

[0132] Currently, when capturing a speaker's voice using an audio acquisition module, the speaker's voice is often affected by reverberation noise and interfering human voices in the surrounding environment. In complex acoustic environments, the listening experience of the audience in the venue and remote participants is affected, and the accuracy of voice input also decreases. To improve the speaker's voice quality, related technologies can capture only the sound within a pickup area, but the range of the pickup area is usually fixed. For example, as shown in Figure 3, the angle of the pickup area is fixed, and the microphone array can pick up sound within a set angle range. When the speaker is standing to one side of the display screen, the pickup area cannot cover the speaker's location, and therefore, the speaker's voice cannot be collected.

[0133] Based on this, this application provides a sound processing method. This method can be executed by an electronic device in a sound pickup system. The method may include the electronic device, in response to a received speech enhancement operation, determining a pickup area containing the location of a target object, and suppressing sounds from outside the pickup area collected by the audio acquisition module to obtain the speech of the target object. In this application embodiment, the pickup area determined includes the location of the target object, thus avoiding the suppression of the target object's voice. Suppressing sounds from outside the pickup area reduces the influence of reverberation noise and interfering human voices in the surrounding environment, enhances the speech of the target object located within the pickup area, improves the listening experience of participants, and increases the accuracy of speech-to-text conversion.

[0134] Figure 4 illustrates a flowchart of a sound processing method provided in an embodiment of this application. This sound processing method can be executed by the electronic device 110 shown in Figure 1 or Figure 2, or by a module with data processing capabilities in a sound pickup system. As shown in Figure 4, the method may include the following steps:

[0135] S401, in response to the received speech enhancement operation, determines the pickup area.

[0136] In a meeting setting, the speaker can choose to enable the voice enhancement function, which can suppress reverberation noise and interference from the surrounding environment. When participants are discussing together, the speaker can choose to disable the voice enhancement function.

[0137] A display screen is installed on the podium at the front of the venue. In some embodiments, as shown in Figure 5, the display screen is equipped with a voice enhancement control, which can be a physical button or a toggle switch; alternatively, the display screen may include a display panel, on which the voice enhancement control can be displayed at a designated location. The voice enhancement operation can be a press operation of the voice enhancement control by a target object. The target object can be the speaker. When the speaker is speaking, they can choose to enable the voice enhancement function. When the speaker needs to enable the voice enhancement function, they can press the voice enhancement control. The electronic device receives the press operation on the voice enhancement control and confirms that a voice enhancement operation has been received. In response to the received voice enhancement operation, the electronic device determines the sound pickup area, which includes the location of the target object.

[0138] In other embodiments, the display screen can be a touch screen, as shown in Figure 6. The voice enhancement operation can be an operation where the target object presses the display screen with its hand. The target object can be a speaker. When the speaker is speaking, they can choose to enable the voice enhancement function. When the voice enhancement function needs to be enabled, the speaker can press any position on the display screen with their hand. The electronic device receives the operation of pressing the display screen and determines that it has received the voice enhancement operation. In response to the received voice enhancement operation, the electronic device determines the pickup area. In some embodiments, to avoid accidental operation, the electronic device only determines the pickup area when it receives the operation of pressing the display screen and the duration of the operation of pressing the display screen reaches a set duration threshold. The pickup area includes the location of the target object. The set duration threshold can be a preset value. For example, the set duration threshold can be 2 seconds, 3 seconds, or 5 seconds.

[0139] In some embodiments, the electronic device may determine the pickup area based on a first set value and a second set value. The first set value characterizes the arm length of the target object, and the second set value characterizes the width of the display screen. When the target object inputs a voice enhancement operation, it is located within the pickup area determined based on the first and second set values.

[0140] In other embodiments, the electronic device can determine the interaction location of the voice enhancement operation and determine a pickup area based on the interaction location, which can also be referred to as an audio gaiter. For example, the pickup area can be determined based on the interaction location and the aforementioned first set value. The interaction location is used to characterize the position of the target object. When the voice enhancement operation is a press operation of the target object on a voice enhancement control, the interaction location can be the position of the voice enhancement control; when the voice enhancement operation is an operation of the target object's hand pressing the display screen, the interaction location can be the position where the target object's hand presses the display screen.

[0141] S402, suppresses sounds from outside the pickup area in the sound collected by the audio acquisition module to obtain the speech of the target object.

[0142] After determining the pickup area, the electronic device can suppress sounds from outside the pickup area in the sound collected by the audio acquisition module to obtain the speech of the target object. For example, the electronic device can perform speech separation processing on the sound collected by the audio acquisition module based on the range parameters of the pickup area and the sound characteristics of the speech and interfering noise to obtain the speech of the target object.

[0143] For ease of understanding, the sound processing method provided in this application will be described in detail below through several specific embodiments. Each of the following embodiments is illustrated using a touch screen as an example.

[0144] In some optional embodiments, the electronic device may determine the pickup area based on a first set value and a second set value. As shown in FIG7, the sound processing method provided in this application embodiment may include the following steps:

[0145] S701 detects the contact area covered by the touch points on the display screen.

[0146] In a conference setting, when the speaker is speaking, the voice enhancement function can be activated, allowing the speaker to press any position on the display screen with their hand. When the speaker's hand presses the screen, multiple contact points exist at the point of contact between the hand and the screen. The electronic device can detect the positions of these contact points. Assuming the electronic device detects n contact points, the coordinates of these n contact points are obtained based on the display screen coordinate system. These n contact points form the basic shape of the hand. For example, as shown in Figure 8, the display screen coordinate system can be a coordinate system with the lower left corner of the display screen as the origin. The coordinates of the n contact points can be represented as {(x1(t),y1(t)),(x2(t),y2(t)),…,(x i (t),y i (t)),…,(x n (t),y n (t))}, where x i(t) and y i (t) represents the x-coordinate and y-coordinate of the i-th contact point at time t.

[0147] Based on the contact point coordinates of n contact points, the centroid coordinates of the shape formed by the n contact points can be determined ((x... c (t),y c (t)), the centroid coordinates can be expressed as:

[0148] Obtain the centroid coordinates ((x c (t),y c After (t)), the distance from each of the n contact points to the centroid can be determined, resulting in n distance values. The distance from the i-th contact point to the centroid can be expressed as:

[0149] As shown in Figure 8, with the centroid as the center, and the maximum value among n distance values ​​max(d) i Given a radius R1, generate an equivalent circle for the contact points, where each of the n contact points is enclosed within the equivalent circle. This circle is based on the maximum value among the n distance values, max(d). i (t)), using the formula πR1 2 Determine the area S(t) of the equivalent circle of the contact point, and use the area S(t) of the equivalent circle of the contact point as the contact area covered by the contact point.

[0150] S702, based on the contact area covered by the touch point, determine whether it is a voice enhancement operation; if yes, proceed to step S703; if no, return to step S701.

[0151] If the contact area S(t) covered by the contact point is within the preset area range, that is, S min <S(t)<S max It can be assumed that there is a hand pressing operation on the display screen, and the electronic device can determine that it has received a voice enhancement operation, where S min and S max It can be a preset value, S min S is set based on the estimated minimum palm size. max It is set based on the estimated maximum palm size. In some optional embodiments, when S min <S(t)<S max When t = t0 ~ t1, if t1 - t0 > t min In other words, the duration of hand-pressing the display screen reaches the set duration threshold t. min At that time, the electronic device can determine that it has received a voice enhancement operation.

[0152] S703, in response to the received voice enhancement operation, determines the pickup area based on a first setting value and a second setting value.

[0153] Upon receiving a voice enhancement operation, the electronic device can activate the voice enhancement function in response. The electronic device can determine the pickup area based on a first set value and a second set value. Both the first and second set values ​​are preset values. The first set value represents the speaker's arm length, and the second set value represents the width of the display screen. For example, the estimated approximate length of a human arm (1m) can be used as the first set value; the first set value could also be 1.5m. The second set value can be set based on the width of the display screen shown in Figure 8. Both the first and second set values ​​are adjustable, allowing the user to adjust the pickup area according to actual needs.

[0154] In this embodiment of the application, the audio acquisition module may be a linear microphone array, or simply a microphone array, which may be installed at the top of the display screen.

[0155] Figure 9 is a top view of the display screen. In some embodiments, as shown in Figure 9(a), the microphone array may include omnidirectional microphones, that is, the microphone array is an array composed of omnidirectional microphones. The pickup area may be a cylinder, and the height of the cylinder is not particularly limited, but may be limited by the height of the space in which the microphone array is located. The cylinder may include a circular cross-section, the center of which is the center of the microphone array. The radius R2 of the circular cross-section is determined based on a first set value and a second set value, where the first set value L... arm Used to represent the speaker's arm length, assuming the screen width is X. D The second setting value can be X. D / 2. The radius R2 of the circular cross-section can be expressed as:

[0156] In other embodiments, as shown in Figure 9(b), the microphone array may include directional microphones, i.e., the microphone array is an array of directional microphones used to collect sound from the area in front of the display screen. The pickup area may be a column, the height of which is not particularly limited and can be limited by the height of the space where the microphone array is located. The column may include a semi-circular cross-section, the center of which is the center of the microphone array. The diameter of the semi-circular cross-section is parallel to the plane where the display screen is located, and the radius of the semi-circular cross-section may also be R2, the same as the radius of the circular cross-section shown in Figure 9(a).

[0157] S704 suppresses sounds from outside the pickup area in the sound collected by the audio acquisition module to obtain the speech of the target object.

[0158] After determining the pickup area, the electronic device can perform speech separation processing on the sound collected by the audio acquisition module based on the range parameters of the pickup area and the sound characteristics of speech and interference noise, to obtain the speech of the target object. The range parameters of the pickup area may include the shape of the pickup area, the center or center of the cross-section, and the radius of the cross-section.

[0159] For example, as shown in Figure 10, the audio acquisition module is a microphone array, and the sound collected by the audio acquisition module can be a multi-channel audio signal acquired by the microphone array. The electronic device acquires the multi-channel audio signal acquired by the microphone array and can use an acoustic echo cancellation (AEC) algorithm to eliminate the linear echo component in the multi-channel audio signal, obtaining the multi-channel audio to be processed. To fully utilize the multi-channel information of the microphone array, the electronic device can perform multi-channel feature extraction on the multi-channel audio to be processed, obtaining a feature expression sequence. Then, an artificial intelligence (AI) model processes the feature expression sequence to separate the speech of the target object, i.e., the speaker's speech. The AI ​​model can adopt a two-level architecture, or in other words, the AI ​​model can include two parts: a first-level sub-model and a second-level sub-model. The first-level sub-model can be used to classify the input feature expression sequence to obtain the sound features of the speech and the sound features of interfering noise. The speech features and interference noise features are input into the second-level sub-model to obtain a speech separator. The range parameters of the pickup area and the multi-channel audio signals collected by the microphone array are then input into the speech separator. The speech separator suppresses sounds outside the pickup area and interference noise to obtain the speech of the target object. The second-level sub-model is trained on a deep learning model based on sample audio data, various pickup areas, and the sound features of different sounds. In some embodiments, the audio signal output by the second-level sub-model of the AI ​​model may still contain a small amount of far-end echo remnants. In this case, a residual echo suppression model can be used to further eliminate the echoes in the audio signal, ultimately obtaining clean speech of the target object. The residual echo suppression model can be a trained neural network model.

[0160] When the speaker finishes speaking and the discussion begins, the speaker can turn off the voice enhancement function. For example, the speaker can stop the voice enhancement operation and turn off the voice enhancement function by stopping pressing the display screen or releasing the voice enhancement control. The electronic device, upon detecting the cessation of voice enhancement operation, can turn off the voice enhancement function and pause voice enhancement processing of the sound collected by the audio acquisition module.

[0161] The above embodiments can solve the problem that the sound pickup area cannot cover the speaker's location when the speaker is standing on one side of the display screen, ensuring the stability of the speaker's voice enhancement effect using speech separation technology. In related technologies, when enabling the speech enhancement function, the speaker needs to first open the function menu, find the speech enhancement function control in the function menu, and manually enable the speech enhancement function by operating the control. This requires navigating through multiple pages to find the speech enhancement function control. When the speaker and participants are discussing, there will be repeated operations of turning the speech enhancement function on and off, resulting in poor usability and convenience. The embodiments of this application allow the speaker to quickly enable the speech enhancement function by pressing the speech enhancement control or pressing the display screen with their hand. This solves the problem of inconvenience and slowness when the speaker frequently turns the speech enhancement function on and off, allowing the speaker to turn the speech enhancement function on and off through simple interactive operations.

[0162] In some embodiments, the electronic device can also be connected to an image acquisition module, which can be a camera mounted above the display screen to capture images of the venue. When determining the sound pickup area, the electronic device, in response to a received voice enhancement operation, can perform face recognition on the image acquired by the image acquisition module to determine the location of the target face, and determine the sound pickup area based on the location of the target face. For example, as shown in Figure 11, the electronic device acquires the image acquired by the image acquisition module, preprocesses the image, including color space conversion, noise reduction, and normalization of image brightness and contrast, to obtain a processed image. The processed image is then input into a face detection module, which detects the image using a face detection algorithm. For example, the face detection module can be a deep learning model, such as any one of a multi-task convolutional neural network (MTCNN), a faster region-convolutional neural network (Faster R-CNN), or a YOLO model. If multiple face information is detected, the area of ​​each face in the image can be determined using the key point information of the face, and the face with the largest area can be selected as the target face to determine its location. The location of the target face is its coordinates within the image. Through coordinate transformation, these coordinates can be converted to corresponding display screen coordinates, determining the x-coordinate value of the target face's display screen coordinates. If the x-coordinate value of the target face's display screen coordinates is greater than a second preset value, i.e., X... D / 2, the electronic device can determine that the target object is located on the right half of the display screen. If the horizontal coordinate value of the display screen corresponding to the target face is less than the second set value, the electronic device can determine that the target object is located on the left half of the display screen.

[0163] In other embodiments, the electronic device can also determine whether the target object is located on the left or right half of the display screen based on the interaction location of the voice-enhanced operation, such as the centroid coordinates mentioned above. For example, this can be achieved when the width of the display screen meets a certain condition, such as when the width of the display screen is greater than 2 × L. arm When the x-coordinate in the centroid coordinate system is greater than X... D / 2+L arm This determines that the target object is located on the right half of the display screen if the x-coordinate in the centroid coordinate system is less than X. D / 2-L arm If the target object is located on the left half of the display screen, then it can be determined that the target object is located on the left or right half of the display screen. If the x-coordinate in the centroid coordinate system cannot determine whether the target object is located on the left or right half of the display screen, the pickup area can be determined by referring to the method in step S701.

[0164] In one alternative embodiment, the microphone array may include omnidirectional microphones, i.e., the microphone array is an array composed of omnidirectional microphones. The pickup area may be a cylinder, the height of which is not particularly limited and may be limited by the height of the space in which the microphone array is located. The cylinder may include a semi-circular cross-section, the center of which is the center of the microphone array. The diameter of the semi-circular cross-section is perpendicular to the plane of the display screen, and the radius of the circular cross-section may also be R2, the same as the radius of the circular cross-section shown in Figure 9(a). As shown in Figure 12(a), if the target object is determined to be located on the right half of the display screen, the pickup area is located to the right of the central axis of the display screen; as shown in Figure 12(b), if the target object is determined to be located on the left half of the display screen, the pickup area is located to the left of the central axis of the display screen.

[0165] In another alternative embodiment, the microphone array may include directional microphones, i.e., the microphone array is an array of directional microphones used to collect sound from the area in front of the display screen. The pickup area may be a column, the height of which is not particularly limited and may be limited to the height of the space in which the microphone array is located. The column may include a sector-shaped cross-section, the sector being a quarter circle, the center of which is the center of the microphone array, one side of which is perpendicular to the plane of the display screen, and the other side being parallel to the plane of the display screen. The radius of the sector-shaped cross-section may also be R2, the same as the radius of the circular cross-section shown in Figure 9(a). As shown in Figure 13(a), if the target object is determined to be located on the right half of the display screen, the pickup area is located to the right of the central axis of the display screen; as shown in Figure 13(b), if the target object is determined to be located on the left half of the display screen, the pickup area is located to the left of the central axis of the display screen.

[0166] If no face is detected after face detection in the image, the pickup area can be determined by referring to the method in step S701, which will not be elaborated here.

[0167] In some alternative embodiments, the electronic device may determine the pickup area based on the interaction location of the voice enhancement operation. As shown in FIG14, the sound processing method provided in this application embodiment may include the following steps:

[0168] S1401, in response to the received voice enhancement operation, determine the interaction location of the voice enhancement operation.

[0169] Taking the voice enhancement operation as an example of a hand pressing the display screen, the process of the electronic device detecting whether it has received a voice enhancement operation can be performed with reference to the embodiment shown in Figure 7, and will not be repeated here. After receiving the voice enhancement operation, the electronic device can obtain the centroid coordinates (x, y, x) of the shape formed by the multiple touch points of the hand pressing the display screen. c (t),y c (t)), using the centroid coordinates as the interaction position for the speech enhancement operation.

[0170] S1402, determine the pickup area based on the interaction position of the voice enhancement operation.

[0171] In some embodiments, the microphone array may include omnidirectional microphones, i.e., the microphone array is an array composed of omnidirectional microphones. As shown in Figure 15(a), the electronic device can estimate the speaker's location area based on the interactive location of the voice enhancement operation, i.e., using (x c (t),y c A circular region centered at (t) with radius R3, where radius R3 can be based on a first set value L. arm It is determined that, in some embodiments, the first set value L can be... arm As the radius R3, that is, R3 = L arm Users can adjust the first setting value L according to the actual situation. arm Size.

[0172] The pickup area can be a cylinder, with no particular height limit, but can be limited by the height of the space where the microphone array is located. The cylinder can include a sector-shaped cross-section, with the apex of the sector being the center of the microphone array. The sector-shaped cross-section is the circumscribing sector of a circle with a radius of R3. The radius R of the sector-shaped cross-section... sec It can be represented as:

[0173] The direction of the sector-shaped cross-section is the end-fire direction of the microphone array, and the included angle α of the sector is... sec1 It can be represented as:

[0174] From the formula above, we can see that when the x-coordinate of the centroid is... c (t) equals the radius R3 and X D When the sum is 2 / 2, the outer sector of the circle with radius R3 is a semicircle, and the dot of the semicircle is the center of the microphone array.

[0175] For x c <R3+X D In the case of / 2, the cross-section of the cylinder in the pickup area is a circle centered at the center of the microphone array, and the radius of the circle can be x. c (t)-X D / 2+R3.

[0176] This embodiment uses the interactive position of the voice enhancement operation and the speaker's arm length to estimate the speaker's range of motion. Based on the speaker's range of motion, it adaptively adjusts the range of the pickup area, further reducing the pickup area to a fan-shaped area with distinguishable distance. By suppressing sounds outside the pickup area, interference noise can be further removed. In scenarios where interference sounds are near the speaker, better voice separation can be achieved.

[0177] In other embodiments, the microphone array may include directional microphones, i.e., the microphone array is an array of directional microphones used to collect sound from the area in front of the display screen. As shown in Figure 15(b), the electronic device can estimate the speaker's location area based on the interaction location of the voice enhancement operation, i.e., using (x... c (t),y c (t) is a semicircular region with center R3 and radius R3, where radius R3 can be a first set value L. arm Users can adjust the first setting value L according to the actual situation. arm Size.

[0178] The pickup area can be a cylinder, with no particular height limit, but can be limited by the height of the space where the microphone array is located. The cylinder can include a sector-shaped cross-section, with the apex of the sector being the center of the microphone array. The sector-shaped cross-section is the circumcircle of a semicircle with a radius of R3. The radius R of the sector-shaped cross-section... sec It has the same radius as the sector section shown in (a) of Figure 15.

[0179] The included angle α of the sector sec2 It can be represented as:

[0180] From the formula above, we can see that when the x-coordinate of the centroid is... c (t) equals the radius R3 and X DWhen the sum is 2 / 2, the circumscribed sector of the circle with radius R3 is a quarter circle, and the vertex of the quarter circle is the center of the microphone array.

[0181] The angle α between the direction of the fan-shaped beam and the end-fire direction of the microphone array sec3 It can be represented as:

[0182] For x c <R3+X D In the case of / 2, the cross-section of the cylinder in the pickup area is a semicircle centered at the center of the microphone array, and the radius of the circle can be x. c (t)-X D / 2+R3.

[0183] In this embodiment, by using a directional microphone and leveraging its ability to distinguish between front and back, the fan-shaped angle of the pickup area is reduced by half, which can further improve the effect of speech separation.

[0184] In other embodiments, the microphone array may include omnidirectional microphones, i.e., the microphone array is an array composed of omnidirectional microphones. As shown in Figure 16(a), the electronic device can estimate the speaker's location area based on the interaction location of the voice enhancement operation, i.e., using (x... c (t),y c A circular region centered at (t) with radius R3, where radius R3 can be based on a first set value L. arm It is determined that, in some embodiments, R3 = L arm Users can adjust the first setting value L according to the actual situation. arm Size.

[0185] The pickup area can be a cylinder, and the height of the cylinder is not particularly limited, but can be limited by the height of the space where the microphone array is located. The cylinder can include a circular cross-section, the center of which is the interactive position of the speech enhancement operation. The circular cross-section is a circular area with a radius of R3 as shown in Figure 16(a).

[0186] This embodiment sets the center of the pickup area at the interactive position of the speech enhancement operation, allowing the closed pickup area to follow the speaker and dynamically adjust the position of the pickup area. Furthermore, regardless of whether the speaker is standing near or far from the center of the microphone array, the size of the pickup area is the same, achieving uniformity of the pickup range at each position. Moreover, the smaller pickup range is beneficial for further improving the speech separation effect.

[0187] In other embodiments, the microphone array may include directional microphones, i.e., the microphone array is an array of directional microphones used to collect sound from the area in front of the display screen. As shown in Figure 16(b), the electronic device can estimate the speaker's location area based on the interaction location of the voice enhancement operation, i.e., using (x c (t),y c A semi-circular region centered at (t) with radius R3, where radius R3 can be based on a first set value L. arm It is determined that, in some embodiments, R3 = L arm Users can adjust the first setting value L according to the actual situation. arm Size.

[0188] The pickup area can be a cylinder, and the height of the cylinder is not particularly limited, but can be limited by the height of the space where the microphone array is located. The cylinder can include a semi-circular cross-section, the center of which is the interactive position for voice enhancement operation. The diameter of the semi-circular cross-section is parallel to or coincides with the plane of the display screen. The semi-circular cross-section is a semi-circular area with a radius of R3 as shown in Figure 16(b). This embodiment, by introducing a directional microphone, reduces the pickup area from a circular cross-section to a semi-circular cross-section, further reducing the range of the pickup area, which is beneficial to improving the effect of voice separation.

[0189] In other embodiments, the cross-section of the pickup area can be a rectangle circumscribed by a circle with a radius of R3. Exemplarily, the microphone array can include omnidirectional microphones, i.e., the microphone array is an array composed of omnidirectional microphones. As shown in Figure 17(a), the electronic device can estimate the speaker's location area based on the interactive location of the voice enhancement operation, i.e., using (x... c (t),y c A circular region centered at (t) with radius R3, where radius R3 can be based on a first set value L. arm It is determined that, in some embodiments, R3 = L arm Users can adjust the first setting value L according to the actual situation. arm Size.

[0190] The pickup area can be a cylinder, and there is no special restriction on the height of the cylinder, which can be limited by the height of the space where the microphone array is located. The cylinder can include a rectangular cross-section, with the center of the rectangular cross-section being the interactive position for the speech enhancement operation. The length and width of the rectangular cross-section can be the same, both being 2*R3.

[0191] In other embodiments, the microphone array may include directional microphones, i.e., the microphone array is an array of directional microphones used to collect sound from the area in front of the display screen. As shown in Figure 17(b), the electronic device can estimate the speaker's location area based on the interaction location of the voice enhancement operation, i.e., using (x... c (t),y c A semi-circular region centered at (t) with radius R3, where radius R3 can be based on a first set value L. arm It is determined that, in some embodiments, R3 = L arm Users can adjust the first setting value L according to the actual situation. arm Size.

[0192] The pickup area can be a cylinder, and the height of the cylinder is not particularly limited, but can be limited by the height of the space where the microphone array is located. The cylinder can include a rectangular cross-section, and the midpoint of one of the long sides of the rectangular cross-section is the interactive position for the voice enhancement operation. The length of the rectangular cross-section can be 2*R3, and the width of the rectangular cross-section can be R3.

[0193] The shape of the cross-section of the pickup area in the above embodiments is only an illustrative example. In other embodiments, the cross-section of the pickup area may also be other shapes. This application does not limit the shape of the pickup area.

[0194] S1403, suppress the sound from outside the pickup area in the sound collected by the audio acquisition module to obtain the speech of the target object.

[0195] After determining the pickup area, the electronic device can perform speech separation processing on the sound collected by the audio acquisition module based on the range parameters of the pickup area and the sound characteristics of speech and interference noise, to obtain the speech of the target object. For a pickup area with a fan-shaped cross-section, the range parameters of the pickup area can include the shape of the pickup area, the vertex and radius of the fan-shaped cross-section, as well as the included angle and direction of the fan. For a pickup area with a circular or semi-circular cross-section, the range parameters of the pickup area can include the shape of the pickup area, the center and radius of the cross-section, etc.

[0196] The process of the electronic device processing the sound collected by the audio acquisition module can be referred to Figure 10, and will not be described in detail here.

[0197] After the speaker finishes speaking and the free discussion begins, the speaker can stop the voice enhancement operation and turn off the voice enhancement function by stopping pressing the display screen or releasing the voice enhancement control. If the electronic device detects that the voice enhancement operation has stopped, it can turn off the voice enhancement function and pause the voice enhancement processing of the sound collected by the audio acquisition module.

[0198] In the above embodiments, the range of the pickup area is determined based on the interaction position of the speech enhancement operation, further narrowing the pickup area. By suppressing sounds outside the pickup area, better speech separation can be achieved in scenarios where interfering sounds are near the speaker. For example, as shown in Figure 18, the cross-section of the pickup area has a radius of R. sec When the display is in a fan-shaped pattern, if someone is speaking in the area near the screen, their voice can be suppressed, and only the speaker's voice will be picked up.

[0199] In some alternative embodiments, the voice enhancement operation received by the electronic device can perform face recognition on the image acquired by the image acquisition module to determine the location of the target face, and determine the pickup area based on the location of the target face and the interaction location of the voice enhancement operation. As shown in Figure 19, the sound processing method provided in this application embodiment may include the following steps:

[0200] S1901, in response to the received speech enhancement operation, determine the interaction position of the speech enhancement operation.

[0201] Taking the voice enhancement operation as an example of a hand pressing the display screen, the process of the electronic device detecting whether it has received a voice enhancement operation can be performed with reference to the embodiment shown in Figure 7, and will not be repeated here. After receiving the voice enhancement operation, the electronic device can obtain the centroid coordinates (x, y, x) of the shape formed by the multiple touch points of the hand pressing the display screen. c (t),y c (t)), using the centroid coordinates as the interaction position for the speech enhancement operation.

[0202] S1902, perform face recognition on the image acquired by the image acquisition module to determine the location of the target face.

[0203] The electronic device can also be connected to an image acquisition module, which can be a camera mounted above the display screen to capture images of the venue. When determining the sound pickup area, the electronic device, in response to a received voice enhancement operation, can perform facial recognition on the image captured by the image acquisition module to determine the location of the target face. The process by which the electronic device determines the location of the target face can be executed according to the process shown in Figure 11, and will not be described in detail here.

[0204] S1903, determine the pickup area based on the interaction location of the voice enhancement operation and the location of the target face.

[0205] After the electronic device determines the location of the target face, the location of the target face is defined as its coordinates in the image. Through coordinate transformation, these coordinates can be converted into corresponding display screen coordinates. The x-coordinate value of the display screen coordinates corresponding to the target face is determined. This x-coordinate value is then compared with the x-coordinate value of the interaction position for the voice enhancement operation. c (t) is compared. If the x-coordinate of the display screen corresponding to the target face is greater than the x-coordinate of the interaction position... c (t), the electronic device can determine that the target object is located to the right of the interaction position if the x-coordinate of the display screen corresponding to the target face is less than the x-coordinate of the interaction position. c (t), the electronic device can determine that the target object is located to the left of the interaction position.

[0206] In some embodiments, the microphone array may include omnidirectional microphones, i.e., the microphone array is an array composed of omnidirectional microphones. The pickup area may be a column, and the height of the column is not particularly limited, but may be limited by the height of the space in which the microphone array is located. The column may include a semi-circular cross-section, the center of which is the interactive position for the voice enhancement operation, and the radius R3 of the semi-circular cross-section may be based on a first set value L. arm It is determined that, in some embodiments, R3 = L arm Users can adjust the first setting value L according to the actual situation. arm The size of the microphone array is as follows: The axis of symmetry of the semi-circular cross-section is parallel to the orientation of the microphone array. As shown in Figure 20(a), if the target object is located to the right of the interaction position, the pickup area is located to the right of the interaction position; as shown in Figure 20(b), if the target object is located to the left of the interaction position, the pickup area is located to the left of the interaction position, so as to ensure that the target's face is located within the pickup area.

[0207] This embodiment determines the position of the target face by performing face detection. By combining the left and right information of the target face position relative to the interaction position, the sound pickup area can be reduced to a semi-circular cross-section, further improving the effect of speech separation.

[0208] After performing face detection on the image, if no face is detected, the pickup area shown in Figure 16(a) can be determined by referring to the method in step S1402, which will not be elaborated here.

[0209] In other embodiments, the microphone array may include directional microphones, i.e., an array of directional microphones used to collect sound from the area in front of the display screen. The pickup area may be a column, the height of which is not particularly limited and may be limited by the height of the space in which the microphone array is located. The column may include a sector-shaped cross-section, which is a quarter-circle, with the apex of the quarter-circle being the interaction position for the voice enhancement operation, and the radius R3 of the quarter-circle may be based on a first set value L. arm It is determined that, in some embodiments, R3 = L arm Users can adjust the first setting value L according to the actual situation. arm The size of the quarter circle is such that one side is parallel to the direction of the microphone array, and the other side is perpendicular to the direction of the microphone array. As shown in Figure 21(a), if the target object is located to the right of the interaction position, the pickup area is located to the right of the interaction position; as shown in Figure 21(b), if the target object is located to the left of the interaction position, the pickup area is located to the left of the interaction position, so as to ensure that the position of the target's face is within the pickup area.

[0210] This embodiment combines the forward and backward judgment of the directional microphone with the left and right information of the target face position relative to the interaction position, which can reduce the sound pickup area to a quarter circle cross-section, further improving the effect of speech separation.

[0211] After performing face detection on the image, if no face is detected, the pickup area shown in Figure 16(b) can be determined by referring to the method in step S1402, which will not be elaborated here.

[0212] S1904, suppresses sounds from outside the pickup area in the sound collected by the audio acquisition module to obtain the speech of the target object.

[0213] After determining the pickup area, the electronic device can perform speech separation processing on the sound collected by the audio acquisition module based on the range parameters of the pickup area and the sound characteristics of speech and interference noise, to obtain the speech of the target object. For a pickup area with a fan-shaped cross-section, the range parameters of the pickup area can include the shape of the pickup area, the center and radius of the cross-section, and the directional relationship between the pickup area and the interaction position.

[0214] The process of the electronic device processing the sound collected by the audio acquisition module can be referred to Figure 10, and will not be described in detail here.

[0215] After the speaker finishes speaking and the free discussion begins, the speaker can stop the voice enhancement operation and turn off the voice enhancement function by stopping pressing the display screen or releasing the voice enhancement control. If the electronic device detects that the voice enhancement operation has stopped, it can turn off the voice enhancement function and pause the voice enhancement processing of the sound collected by the audio acquisition module.

[0216] In other embodiments, the electronic device detects a target object's contact with the display screen and determines the sound pickup area based on set parameters. Figure 22 exemplarily illustrates a flowchart of a sound processing method provided in an embodiment of this application. This sound processing method can be executed by the electronic device 110 shown in Figure 1 or Figure 2, or by a module with data processing capabilities in the sound pickup system. As shown in Figure 22, the method may include the following steps:

[0217] S2201, The target object's contact operation with the display screen is detected, and the sound pickup area is determined based on the set parameters.

[0218] Here, "display screen" refers to a large screen, such as a conference screen. The target object's contact operation with the display screen can include, but is not limited to, the target object touching the display screen, the target object pressing a certain position on the display screen, the target object sliding their hand on the display screen, or the target object clicking or double-clicking a certain position on the display screen. Taking the contact operation of pressing the display screen as an example, the electronic device detects the target object's contact operation with the display screen and can determine the sound pickup area based on set parameters. The set parameters are used to characterize the predicted range of motion of the target object. For example, in one embodiment, the set parameters may include a first set value and a second set value. The first set value is related to the length of the target object's arm and can be greater than or equal to the maximum length of a human arm; the second set value is related to the width of the display screen and can be greater than or equal to the width of the display screen. The target object's range of motion lies within the sound pickup area determined based on the first and second set values. As shown in Figure 9(a), the sound pickup area can be a cylinder, the cross-section of the cylinder can be circular, the center of the circle can be the center of the microphone array, and the radius of the cross-section of the cylinder can be determined based on a first set value and a second set value; as shown in Figure 9(b), the sound pickup area can be a cylinder, the cross-section of the cylinder can be semi-circular, the center of the circle can be the center of the microphone array, and the radius of the cross-section of the cylinder can be determined based on a first set value and a second set value.

[0219] In another embodiment, the setting parameters may include preset values, and the electronic device may determine the pickup area based on these preset values. The target object's activity range lies within the pickup area determined based on the preset values. For example, as shown in Figure 23(a), the pickup area may be a cylinder, the cross-section of which may be circular, the center of which may be the center of the microphone array, and the radius R2 of the cylinder's cross-section may be a preset value, which may be a value determined based on the target object's activity range. Exemplarily, the preset value may be 2.5m or 3m. As shown in Figure 23(b), the pickup area may be a cylinder, the cross-section of which may be semi-circular, the center of which may be the center of the microphone array, and the radius R2 of the cylinder's cross-section may be based on a preset value, which may be a value determined based on the target object's activity range.

[0220] In another embodiment, the electronic device can adjust the range of the sound pickup area. For example, when no target object is detected touching the display screen, the sound pickup area of ​​the electronic device can be in a fully open state. When a target object is detected touching the display screen, the electronic device can determine the sound pickup area based on the aforementioned settings. When the target object stops touching the display screen, the electronic device can control the sound pickup area to return to the fully open state.

[0221] In another embodiment, the setting parameters may include preset values. The number of preset values ​​can be multiple, which can be understood as the electronic device having multiple preset levels to adjust the range of the sound pickup area. For example, taking two preset values ​​as an example, the first preset value could be 2m, and the second preset value could be 3m. When a target object is detected touching the display screen, the electronic device predicts that the target object is near the display screen and can determine the sound pickup area based on the first preset value to define a relatively small sound pickup area, improving the voice separation effect. When the target object stops touching the display screen, the electronic device predicts that the target object may be slightly farther away from the display screen and can determine the sound pickup area based on the second preset value to define a relatively large sound pickup area, ensuring that the sound pickup area includes the target object's range of motion. In the above embodiment, the first preset value is less than the second preset value. In other embodiments, the first preset value may also be greater than the second preset value. The above-mentioned number of preset values ​​of 2 is only an illustrative example. In other embodiments, the number of preset values ​​may be greater than 2, for example, the number of preset values ​​may be 3, 4, or 5, etc.

[0222] S2202, suppresses sounds from outside the pickup area in the sound collected by the audio acquisition module to obtain the speech of the target object.

[0223] After determining the pickup area, the electronic device can suppress sounds from outside the pickup area in the sound collected by the audio acquisition module to obtain the speech of the target object. For example, the electronic device can perform speech separation processing on the sound collected by the audio acquisition module based on the range parameters of the pickup area and the sound characteristics of the speech and interfering noise to obtain the speech of the target object.

[0224] In other embodiments, the electronic device detects a touch operation of a target object on the display screen and determines the sound pickup area based on the interaction position of the touch operation. Figure 24 exemplarily illustrates a flowchart of a sound processing method provided in an embodiment of this application. This sound processing method can be executed by the electronic device 110 shown in Figure 1 or Figure 2, or by a module with data processing capabilities in the sound pickup system. As shown in Figure 24, the method may include the following steps:

[0225] S2401, The target object's touch operation on the display screen is detected, and the sound pickup area is determined based on the interaction position of the touch operation.

[0226] The contact operation of the target object with the display screen may include, but is not limited to, the target object touching the display screen, the target object pressing a certain position on the display screen, the target object sliding its hand on the display screen, or the target object clicking or double-clicking a certain position on the display screen. When the electronic device detects the target object's contact operation with the display screen, it can determine the sound pickup area based on the interaction position of the contact operation. For example, the sound pickup area can be a cylinder, and the cross-section of the sound pickup area can be the outer shape of the target object's range of motion; alternatively, the cross-section of the sound pickup area can be the target object's range of motion; the target object's range of motion is determined based on the interaction position of the contact operation.

[0227] In some embodiments, the electronic device can determine the interaction position between the target object and the display screen, and determine the sound pickup area based on the interaction position and the aforementioned first preset value. The interaction position characterizes the location of the target object, and the first preset value is greater than or equal to the arm length of the target object. For example, the range of motion of the target object is a first shape, which is a circle, a semicircle, or a fan shape; the center of the first shape is the interaction position of the contact operation, and the radius of the first shape is the first preset value.

[0228] For example, the cross-section of the sound pickup area can be the outer shape of the target object's range of motion. In an alternative embodiment, the interaction position of the touch operation can be located at the edge of the display screen by default, and the radius of the outer shape is determined based on a first set value and a second set value.

[0229] In one optional embodiment, the microphone array may include omnidirectional microphones, i.e., the microphone array is an array composed of omnidirectional microphones. The pickup area may be a cylinder, which may include a semi-circular cross-section. The center of the semi-circular cross-section is the center of the microphone array. The diameter of the semi-circular cross-section is perpendicular to the plane where the display screen is located. The radius of the circular cross-section may be a preset value, and the number of preset values ​​may be one or more. Based on the interaction position between the target object and the display screen, the electronic device can determine whether the target object is located on the right or left half of the display screen. Similar to the embodiment shown in Figure 12, if the target object is determined to be located on the right half of the display screen, the pickup area is located to the right of the central axis of the display screen; if the target object is determined to be located on the left half of the display screen, the pickup area is located to the left of the central axis of the display screen.

[0230] In another alternative embodiment, the microphone array may include directional microphones, i.e., the microphone array is an array of directional microphones used to collect sound from the area in front of the display screen. The pickup area may be a cylinder, which may include a sector-shaped cross-section, the sector being a quarter circle, with the center of the sector-shaped cross-section being the center of the microphone array. One side of the sector-shaped cross-section is perpendicular to the plane where the display screen is located, and the other side is parallel to the plane where the display screen is located. The radius of the sector-shaped cross-section may also be a preset value, and the number of preset values ​​may be one or more. Based on the interaction position between the target object and the display screen, the electronic device can determine whether the target object is located on the right or left half of the display screen. Similar to the embodiment shown in Figure 13, if the target object is determined to be located on the right half of the display screen, the pickup area is located to the right of the central axis of the display screen; if the target object is determined to be located on the left half of the display screen, the pickup area is located to the left of the central axis of the display screen.

[0231] In another alternative embodiment, the microphone array may include omnidirectional microphones, i.e., the microphone array is an array composed of omnidirectional microphones. As shown in Figure 25(a), the electronic device can estimate the speaker's location area based on the interaction position between the target object and the display screen, i.e., using (x... c (t),y c A circular region centered at (t) with radius R3, where radius R3 can be based on a first set value L. arm To determine, for example, the first set value L can be... arm As the radius R3, that is, R3 = L arm Users can adjust the first setting value L according to the actual situation. arm The size of the pickup area. The pickup area can be a cylinder, which can include a circular cross-section. The center of the circular cross-section is the center of the microphone array, and the radius of the circular cross-section is the distance x from the interaction position to the center of the microphone array. c (t) and the first set value L arm sum.

[0232] In another alternative embodiment, the microphone array may include directional microphones, i.e., the microphone array is an array of directional microphones used to collect sound from the area in front of the display screen. As shown in Figure 25(b), the electronic device can estimate the speaker's location area based on the interaction position between the target object and the display screen, i.e., using (x... c (t),y c A semi-circular region centered at (t) with radius R3, where radius R3 can be based on a first set value L. arm To determine, for example, the first set value L can be... arm As the radius R3, that is, R3 = L arm Users can adjust the first setting value L according to the actual situation. arm The size of the pickup area. The pickup area can be a cylinder, which can include a semi-circular cross-section. The center of the semi-circular cross-section is the center of the microphone array, and the radius of the semi-circular cross-section is the distance x from the interaction position to the center of the microphone array. c (t) and the first set value L arm sum.

[0233] Taking the scenario shown in Figure 25 as an example, in one embodiment, when a target object is detected touching the display screen, the electronic device predicts that the target object is near the display screen and can calculate the distance x from the interaction location to the center of the microphone array. c (t) and the first set value L arm The sum of these values ​​is used as the radius of the circle or semicircle to determine the pickup area, thus defining a relatively small pickup area and improving speech separation. When the target object stops touching the display screen, the electronic device predicts that the target object may be slightly farther away from the display screen, and can use the distance x from the interaction position to the center of the microphone array. c (t), first set value L arm The system defines a pickup area by setting a constant b as the radius of the circle or semicircle, where b is greater than 0 to determine a relatively large pickup area that ensures the pickup area encompasses the target object's range of motion. In one embodiment, when the electronic device detects movement in the interaction position between the target object and the display screen, such as when the target object moves while pressing on the display screen, the interaction position between the target object and the display screen moves accordingly. The electronic device can adjust the radius of the circle or semicircle in real time as the interaction position moves to adjust the size of the pickup area.

[0234] In another alternative embodiment, the microphone array may include omnidirectional microphones, i.e., the microphone array is an array composed of omnidirectional microphones. The electronic device can estimate the speaker's location area based on the interaction position of the target object and the display screen, i.e., using (x... c (t),yc A circular region centered at (t) with radius R3, where radius R3 can be based on a first set value L. arm To determine, for example, the first set value L can be... arm As the radius R3, that is, R3 = L arm Users can adjust the first setting value L according to the actual situation. arm The size of the target object. Based on the interaction position between the target object and the display screen, the electronic device can determine whether the target object is located on the right or left half of the display screen. As shown in Figure 26(a), if the target object is determined to be located on the right half of the display screen, the pickup area is located to the right of the central axis of the display screen; the pickup area can be a cylinder, which can include a semi-circular cross-section, the center of which is the center of the microphone array, and the radius of which is the distance x from the interaction position to the center of the microphone array. c (t) and the first set value L arm The sum. As shown in Figure 26(b), if the target object is determined to be located on the left half of the display screen, then the pickup area is located to the left of the central axis of the display screen; the pickup area can be a cylinder, which can include a semi-circular cross-section, the center of which is the center of the microphone array, and the radius of which is the distance x from the interaction position to the center of the microphone array. c (t) and the first set value L arm sum.

[0235] Taking the scenario shown in Figure 26 as an example, in one embodiment, when a target object is detected touching the display screen, the electronic device predicts that the target object is near the display screen and can calculate the distance x from the interaction location to the center of the microphone array. c (t) and the first set value L arm The sum of these values ​​is used as the radius of the semicircle to determine the pickup area, thus defining a relatively small pickup area and improving speech separation. When the target object stops touching the display screen, the electronic device predicts that the target object may be slightly farther away from the display screen, and can use the distance x from the interaction position to the center of the microphone array. c (t), first set value L arm The system sets a constant b as the radius of the semicircle to determine the sound pickup area, where b is greater than 0, to define a relatively large sound pickup area that ensures the area includes the target object's range of motion. In one embodiment, when the electronic device detects movement in the interaction position between the target object and the display screen, such as when the target object moves while pressing on the display screen, the interaction position between the target object and the display screen moves accordingly. The electronic device can adjust the radius of the semicircle in real time as the interaction position moves.

[0236] In another alternative embodiment, the microphone array may include directional microphones, i.e., the microphone array is an array of directional microphones used to collect sound from the area in front of the display screen. The electronic device can estimate the speaker's location area based on the interaction position of the target object with the display screen, i.e., using (x... c (t),y c A semi-circular region centered at (t) with radius R3, where radius R3 can be based on a first set value L. arm To determine, for example, the first set value L can be... arm As the radius R3, that is, R3 = L arm Users can adjust the first setting value L according to the actual situation. arm The size of the target object. Based on the interaction position between the target object and the display screen, the electronic device can determine whether the target object is located on the right or left half of the display screen. As shown in Figure 27(a), if the target object is determined to be located on the right half of the display screen, the pickup area is located to the right of the central axis of the display screen; the pickup area can be a cylinder, which can include a sector-shaped cross-section, the vertex of which is the center of the microphone array, and the radius of which is the distance x from the interaction position to the center of the microphone array. c (t) and the first set value L arm The sum. As shown in Figure 27(b), if the target object is determined to be located on the left half of the display screen, then the pickup area is located to the left of the central axis of the display screen; the pickup area can be a cylinder, which can include a sector-shaped cross-section, the center of which is the center of the microphone array, and the radius of which is the distance x from the interaction position to the center of the microphone array. c (t) and the first set value L arm sum.

[0237] Taking the scenario shown in Figure 27 as an example, in one embodiment, when a target object is detected touching the display screen, the electronic device predicts that the target object is near the display screen and can calculate the distance x from the interaction location to the center of the microphone array. c (t) and the first set value L arm The sum of these values ​​is used as the radius of the sector to determine the pickup area, thus establishing a relatively small pickup area and improving speech separation. When the target object stops touching the display screen, the electronic device predicts that the target object may be slightly farther away from the display screen, and can use the distance x from the interaction position to the center of the microphone array. c (t), first set value L armThe system sets a constant b as the radius of the sector to determine the pickup area, where b is greater than 0, to define a relatively large pickup area that ensures the pickup area includes the target object's range of motion. In one embodiment, when the electronic device detects movement in the interaction position between the target object and the display screen, such as when the target object moves while pressing on the display screen, the interaction position between the target object and the display screen moves accordingly. The electronic device can adjust the radius of the sector in real time as the interaction position moves to adjust the size of the pickup area.

[0238] S2402, suppresses sounds from outside the pickup area in the sound collected by the audio acquisition module to obtain the speech of the target object.

[0239] After determining the pickup area, the electronic device can suppress sounds from outside the pickup area in the sound collected by the audio acquisition module to obtain the speech of the target object. For example, the electronic device can perform speech separation processing on the sound collected by the audio acquisition module based on the range parameters of the pickup area and the sound characteristics of the speech and interfering noise to obtain the speech of the target object.

[0240] In conjunction with the above method embodiments, this application also provides a sound processing device that can be applied to the electronic device shown in FIG1 or FIG2. This sound processing device can be used to implement the functions of the above method embodiments, and therefore can achieve the beneficial effects of the above method embodiments. As shown in FIG28, the sound processing device 2800 may include a pickup area determination unit 2801 and a sound processing unit 2802.

[0241] In some embodiments, the pickup area determination unit 2801 can be used to determine the pickup area in response to a received voice enhancement operation; the pickup area includes the location of the target object.

[0242] The sound processing unit 2802 can be used to suppress sounds from outside the pickup area in the sound collected by the audio acquisition module to obtain the speech of the target object.

[0243] In other embodiments, the pickup area determination unit 2801 can be used to determine the pickup area based on the interaction position of the touch operation when a touch operation of a target object on the display screen is detected.

[0244] The sound processing unit 2802 can be used to suppress sounds from outside the pickup area in the sound collected by the audio acquisition module to obtain the speech of the target object.

[0245] In other embodiments, the pickup area determination unit 2801 can be used to determine the pickup area based on set parameters when a touch operation of a target object on the display screen is detected.

[0246] The sound processing unit 2802 can be used to suppress sounds from outside the pickup area in the sound collected by the audio acquisition module to obtain the speech of the target object.

[0247] It should be noted that, in some embodiments, the pickup area determination unit 2801 can be used to execute any step in the sound processing method, and the sound processing unit 2802 can be used to execute any step in the sound processing method. The steps implemented by the pickup area determination unit 2801 and the sound processing unit 2802 can be specified as needed. The pickup area determination unit 2801 and the sound processing unit 2802 respectively implement different steps in the sound processing method to achieve all the functions of the sound processing device. The sound processing device 2800 can also employ more or fewer functional modules to achieve the functions of the sound processing device 2800.

[0248] In the embodiments of this application, the functional modules can be integrated into a single processor, or each module can exist physically separately, or two or more modules can be integrated into a single module. The integrated modules described above can be implemented in hardware or as software functional units.

[0249] Based on the same technical concept as the above-described method embodiments, this application also provides an electronic device. This electronic device can be the one shown in FIG1 or FIG2. This electronic device can be used to implement the functions of the method embodiments shown in FIG4, FIG7, or FIG14, and therefore can achieve the beneficial effects of the above-described method embodiments. The structure of this electronic device can be as shown in FIG1 or FIG2, and will not be described again here.

[0250] Based on the same technical concept as the above-described method embodiments, this application also provides a sound pickup system. This sound pickup system can be used to implement the functions of the method embodiments shown in Figures 4, 7, or 14, and thus can achieve the beneficial effects of the above-described method embodiments. This sound pickup system can be the sound pickup system shown in Figure 1 or Figure 2, which will not be described further here.

[0251] This application also provides a chip that can be applied to any electronic device. This chip can be used to implement the functions of the above method embodiments, and therefore can achieve the beneficial effects of the above method embodiments.

[0252] In some embodiments, the structure of the chip 2900 can be as shown in Figure 29, including a processor 2901 and a power supply circuit 2902 connected to the processor 2901. The processor 2901 and the power supply circuit 2902 can be interconnected via a bus. The processor 2901 can be a digital signal processor (DSP), ASIC, field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, or other specific integrated circuits. The bus can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus. The bus can be divided into an address bus, a data bus, a control bus, etc. The power supply circuit 2902 is used to supply power to the processor 2901 via the bus.

[0253] The processor 2901 can be connected to a memory located outside the chip or to a memory located inside the chip, and run software programs and modules stored in the memory to perform various functional applications and data processing of the chip 2900, such as the sound processing method provided in the embodiments of this application.

[0254] In some embodiments, the processor 2901 may include one or more processing units, which may be independent devices or integrated into one or more processors. The processor 2901 may also include a controller, which can generate operation control signals based on instruction opcodes and timing signals to control instruction fetching and execution.

[0255] This application also provides a computer program product comprising computer-executable instructions. In one embodiment, the computer-executable instructions are used to cause a computer to perform the functions described in the method embodiments above.

[0256] Computer-executable instructions can be stored in a computer-readable storage medium. This application also provides a computer-readable storage medium storing executable instructions. In one embodiment, the computer-executable instructions are used to cause a computer to perform the functions described in the method embodiments above.

[0257] The computer-readable storage medium provided in the embodiments of this application may be random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), register, hard disk, portable hard disk, CD-ROM, or any other form of computer-readable storage medium known in the art.

[0258] Computer-executable instructions can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another. For example, the computer program or instructions can be transferred from one website, computer, server, or data center to another website, computer, server, or data center via wired or wireless means. The computer-readable storage medium can be any available medium that a computer can access, or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium, such as a floppy disk, hard disk, or magnetic tape; it can also be an optical medium, such as a digital video disc (DVD); or it can be a semiconductor medium, such as a solid-state drive (SSD).

[0259] In the various embodiments of this application, unless otherwise specified or logically conflicting, the terminology and / or descriptions between different embodiments are consistent and can be referenced mutually. Technical features in different embodiments can be combined to form new embodiments based on their inherent logical relationships. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, such as including a series of steps or units. A method, system, product, or device is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to these processes, methods, products, or devices.

[0260] Although this application has been described in conjunction with specific features and embodiments, it is obvious that various modifications and combinations can be made therein without departing from the spirit and scope of this application. Accordingly, this specification and drawings are merely illustrative examples of the solutions defined by the appended claims and are to be considered as covering any and all modifications, variations, combinations, or equivalents within the scope of this application.

[0261] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the scope of this application. Therefore, if these modifications and variations of the embodiments of this application fall within the scope of the claims of this application and their equivalents, this application also intends to include these modifications and variations.

Claims

1. A sound processing method, characterized in that, The method includes: The system detects a touch operation of the target object on the display screen and determines the sound pickup area based on the interaction position of the touch operation. The audio acquisition module suppresses sounds from outside the pickup area to obtain the speech of the target object.

2. The method according to claim 1, characterized in that, The pickup area is a cylinder, and the cross-section of the pickup area is the outer shape of the target object's range of motion, or the target object's range of motion; the target object's range of motion is determined based on the interaction position of the contact operation.

3. The method according to claim 2, characterized in that, The activity range of the target object is a first shape, which is a circle, a semicircle, or a fan shape; the center of the first shape is the interaction position of the contact operation, and the radius of the first shape is a first set value.

4. The method according to claim 3, characterized in that, The first set value is greater than or equal to the arm length of the target object.

5. The method according to any one of claims 1 to 3, characterized in that, The cross-section of the pickup area is the outer shape of the target object's range of motion; the outer shape is a circle, a semi-circle, a fan shape, or a rectangle.

6. The method according to claim 5, characterized in that, The outer shape is rectangular, and the center of the outer shape is the interaction position of the contact operation.

7. The method according to claim 5, characterized in that, The audio acquisition module is a microphone array; the external shape is circular, semi-circular, or fan-shaped, and the center of the external shape is the center of the microphone array.

8. The method according to claim 7, characterized in that, The interaction position of the touch operation is located at the edge of the display screen, and the radius of the outer shape is determined based on a first set value and a second set value.

9. The method according to claim 8, characterized in that, The first setting value is greater than or equal to the arm length of the target object; the second setting value is greater than or equal to half the width of the display screen.

10. The method according to any one of claims 1 to 9, characterized in that, Its features are, Determining the pickup area based on the interaction position of the contact operation includes: The image acquisition module performs facial recognition on the images it acquires to determine the location of the target face. The sound pickup area is determined based on the interaction position of the contact operation and the position of the target face.

11. The method according to claim 10, characterized in that, Its features are, The target face is located within the sound pickup area.

12. The method according to any one of claims 1 to 11, characterized in that, The method further includes: When the duration of contact between the target object and the display screen reaches a set duration threshold, it is determined that the contact operation has been detected.

13. The method according to any one of claims 1 to 12, characterized in that, The process of suppressing sounds from outside the pickup area in the audio acquisition module to obtain the speech of the target object includes: Based on the range parameters of the pickup area and the sound characteristics of speech and interference noise, the sound collected by the audio acquisition module is processed by speech separation to obtain the speech of the target object.

14. A sound processing method, characterized in that, The method includes: The system detects the target object's contact with the display screen and determines the sound pickup area based on set parameters. The audio acquisition module suppresses sounds from outside the pickup area to obtain the speech of the target object.

15. The method according to claim 14, characterized in that, The method further includes: When the target object stops contacting the display screen, the sound pickup area is controlled to be fully open.

16. The method according to claim 14, characterized in that, The setting parameters include a first preset value; the method further includes: When no contact between the target object and the display screen is detected, the sound pickup area is determined based on the second preset value; The step of determining the pickup area based on set parameters includes: when the contact operation is detected, determining the pickup area based on the first preset value; The method further includes: When the target object stops contacting the display screen, the sound pickup area is determined based on the second preset value.

17. The method according to claim 14, characterized in that, The setting parameters include a first setting value and a second setting value.

18. The method according to any one of claims 14 to 17, characterized in that, The pickup area is a cylinder, and the cross-section of the pickup area can be of any shape.

19. The method according to claim 18, characterized in that, The cross-section of the pickup area has a second shape, which can be circular, semi-circular, fan-shaped, or rectangular.

20. The method according to claim 19, characterized in that, The audio acquisition module is a microphone array; the second shape is circular, semi-circular, or fan-shaped, and the center of the second shape is the center of the microphone array.

21. The method according to any one of claims 14 to 20, characterized in that, The process of suppressing sounds from outside the pickup area in the audio acquisition module to obtain the speech of the target object includes: Based on the range parameters of the pickup area and the sound characteristics of speech and interference noise, the sound collected by the audio acquisition module is processed by speech separation to obtain the speech of the target object.

22. A sound processing device, characterized in that, The device includes: The sound pickup area determination unit is used to detect the contact operation of the target object on the display screen and determine the sound pickup area based on the interaction position of the contact operation; The sound processing unit is used to suppress sounds from outside the pickup area in the sound collected by the audio acquisition module to obtain the speech of the target object.

23. A sound processing device, characterized in that, The device includes: The sound pickup area determination unit is used to detect the contact operation of the target object with the display screen and determine the sound pickup area based on the set parameters; The sound processing unit is used to suppress sounds from outside the pickup area in the sound collected by the audio acquisition module to obtain the speech of the target object.

24. An electronic device, characterized in that, It includes a processor and a memory; the memory stores computer programs or instructions; the processor is used to execute the computer programs or instructions stored in the memory so that the data synchronization device performs the method of any one of claims 1 to 21.

25. A chip, characterized in that, It includes a processor and a power supply circuit; the power supply circuit is used to supply power to the processor, and the processor is used to execute a computer program to implement the method as described in any one of claims 1 to 21.

26. A sound pickup system, characterized in that, The sound pickup system includes the electronic device as described in claim 24 and an audio acquisition module connected to the electronic device.

27. The sound pickup system according to claim 26, characterized in that, The sound pickup system also includes a display screen connected to the electronic device.

28. A computer-readable storage medium, characterized in that, The storage medium stores a computer program or instructions, which, when executed on a computer, implement the method as described in any one of claims 1 to 21.

29. A computer program product, characterized in that, When the computer program product is run on a computer, it causes the computer to perform the method as described in any one of claims 1 to 21.