A speech separation method and device
By combining information from image and sound acquisition devices, the target sound source can be accurately located and separated, solving the problem of inaccurate sound source localization in multi-sound source scenarios and improving the accuracy of speech separation.
Patent Information
- Application Number
- CN202111362514.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-17
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2041-11-17
AI Technical Summary
When multiple sound sources are emitted simultaneously, the accuracy of sound source localization in existing speech separation technologies is low, which affects the accuracy of speech separation.
By acquiring target scene images captured by an image acquisition device within a preset time period and mixed sound signals captured by a sound acquisition device within the same time period, the sound source location information in the image and the device orientation information are used to enhance the sound signal of the target sound source and suppress other sound source signals, thereby achieving accurate location and separation of the sound source.
It improves the accuracy of sound source localization and speech separation, especially in multi-sound source scenarios, it can effectively avoid interference from sound source localization and improve the effect of speech separation.
Smart Images

Figure CN114038452B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer software technology, and in particular to a speech separation method and device. Background Technology
[0002] Currently, speech separation technology is applied in various life scenarios. For example, during a phone call, it separates the speaker's voice signal from background noise; in a multi-person conference, it separates each person's voice signal for easy recording. Speech separation can be based on a single microphone array or multiple distributed microphone arrays to locate sound sources, and then output the speech emitted by one or more sound sources to achieve the purpose of speech separation. However, when multiple sound sources are emitting sound simultaneously during the microphone array's sound acquisition process, it causes significant interference to the localization of multiple sound sources, resulting in low accuracy in sound source localization and severely affecting the accuracy of speech separation. Summary of the Invention
[0003] This application provides a speech separation method and apparatus to improve the accuracy of sound source localization and enhance the accuracy of speech separation.
[0004] To achieve the above objectives, the embodiments of this application provide the following technical solutions:
[0005] In a first aspect, a speech separation method is provided, comprising: acquiring any image to be processed of a target scene acquired by an image acquisition device within a preset time period, and a mixed sound signal of the target scene acquired by a sound acquisition device within a preset time period; the image to be processed includes an image of a first sound source, and the mixed sound signal is composed of the sound signal of the first sound source and other sound signals; determining a first orientation of the first sound source relative to the sound acquisition device based on the position information of the image of the first sound source in the image to be processed, and the orientation information of the image acquisition device relative to the sound acquisition device; enhancing the sound signal of the first orientation in the mixed sound signal, and suppressing the sound signals of other orientations besides the first orientation, to obtain the sound signal of the first sound source.
[0006] Because the sound signals encountered in current speech separation scenarios are often complex and may contain one or more sound sources, the location of the first sound source to be separated is usually interfered with by other sound signals, leading to inaccurate sound source localization and thus affecting the accuracy of speech separation. Therefore, this technical solution, which uses images to determine the sound source, helps improve the accuracy of sound source localization, and the accuracy of speech separation after image-based sound source localization is also significantly improved.
[0007] In one possible implementation, the image to be processed further includes an image of a second sound source, and other sound signals include sound signals from the second sound source; the method further includes: determining a second orientation of the second sound source relative to the sound acquisition device based on the position information of the image of the second sound source in the image to be processed and the orientation information of the image acquisition device relative to the sound acquisition device; enhancing the sound signal from the second orientation in the mixed sound signals and suppressing sound signals from other orientations besides the second orientation to obtain the sound signal from the second sound source.
[0008] This possible implementation provides a specific way to achieve speech separation in scenarios with multiple sound sources. Computer devices can use the above method to locate different sound sources. Since the sound source localization does not interfere with each other in this method, the speech separation accuracy based on this method is higher.
[0009] In one possible implementation, determining the first orientation of the first sound source relative to the sound acquisition device based on the position information of the image of the first sound source in the image to be processed and the orientation information of the image acquisition device relative to the sound acquisition device includes: determining the orientation information of the first sound source relative to the image acquisition device based on the position information of the image of the first sound source in the image to be processed; and determining the first orientation of the first sound source relative to the sound acquisition device based on the orientation information of the first sound source relative to the image acquisition device and the orientation information of the image acquisition device relative to the sound acquisition device.
[0010] This possible implementation provides a specific way for computer devices to locate sound sources based on images. The computer device determines the location information between the sound source and the sound acquisition device by determining the location information between the sound source and the image acquisition device, as well as the location information between the image acquisition device and the sound acquisition device, thereby accurately locating the sound source and helping to achieve speech separation.
[0011] In one possible implementation, the first sound source is a person, and the method further includes: determining the position information of the first sound source's image in the image to be processed by a head and shoulder detection algorithm.
[0012] This possible implementation provides a possible recognition method for computer devices to determine sound sources based on images. By using head and shoulder detection, the sound source in the image can be determined quickly, and the image recognition steps are simplified, making it simple and convenient to implement.
[0013] In one possible implementation, enhancing the first directional sound signal in the mixed sound signal includes: enhancing the first directional sound signal in the mixed sound signal based on a beamforming method.
[0014] This possible implementation provides a specific way for computer devices to achieve speech separation, by using beamforming to output the sound signal of the sound source based on the azimuth after the sound source is located.
[0015] In one possible implementation, the image acquisition device is integrated with the sound acquisition device.
[0016] This possible implementation method facilitates equipment assembly during the implementation process, while reducing the computational workload of coordinate transformation in computer equipment, thus achieving the effect of convenient management.
[0017] In a second aspect, a computer device is provided, comprising: functional units for performing any of the methods provided in the first aspect, wherein the actions performed by each functional unit are implemented by hardware or by hardware executing corresponding software. For example, the computer device may include: an acquisition unit, a determination unit, and a processing unit; the acquisition unit is used to acquire any one image to be processed of a target scene acquired by an image acquisition device within a preset time period, and a mixed sound signal of the target scene acquired by a sound acquisition device within a preset time period; the image to be processed includes an image of a first sound source, and the mixed sound signal is composed of the sound signal of the first sound source and other sound signals; the determination unit is used to determine a first position of the first sound source relative to the sound acquisition device based on the position information of the image of the first sound source in the image to be processed and the orientation information of the image acquisition device relative to the sound acquisition device; the processing unit is used to enhance the sound signal of the first position in the mixed sound signal and suppress the sound signals of other positions besides the first position to obtain the sound signal of the first sound source.
[0018] Thirdly, a computer device is provided, comprising: a processor and a memory. The processor is connected to the memory, the memory being used to store computer execution instructions, and the processor executing the computer execution instructions stored in the memory, thereby implementing any of the methods provided in the first aspect.
[0019] Fourthly, a chip is provided, comprising: a processor and an interface circuit; the interface circuit for receiving code instructions and transmitting them to the processor; and the processor for executing the code instructions to perform any of the methods provided in the first aspect.
[0020] Fifthly, a computer-readable storage medium is provided, including computer-executable instructions that, when executed on a computer, cause the computer to perform any of the methods provided in the first aspect.
[0021] In a sixth aspect, a computer program product is provided, including computer execution instructions that, when executed on a computer, cause the computer to perform any of the methods provided in the first aspect.
[0022] The technical effects of any of the implementation methods in aspects two through six can be found in the technical effects of the corresponding implementation methods in aspect one, and will not be repeated here. Attached Figure Description
[0023] Figure 1 This is a schematic diagram of the composition of a speech separation system provided in an embodiment of this application;
[0024] Figure 2 This is a schematic diagram of the structure of an image and sound acquisition device provided in an embodiment of this application;
[0025] Figure 3 This is a schematic diagram of an application scenario provided by an embodiment of this application;
[0026] Figure 4 A schematic diagram of the structure of a computer device provided in an embodiment of this application;
[0027] Figure 5 A schematic flowchart of a speech separation method provided in an embodiment of this application;
[0028] Figure 6 A schematic diagram illustrating sound source localization provided in an embodiment of this application;
[0029] Figure 7 A coordinate diagram illustrating the calculation of orientation information provided in this application embodiment;
[0030] Figure 8 This is a schematic diagram of the composition of a computer device provided in an embodiment of this application. Detailed Implementation
[0031] In the description of this application, unless otherwise stated, " / " means "or," for example, A / B can mean A or B. The "and / or" in this document is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, and B alone. Furthermore, "at least one" means one or more, and "multiple" means two or more. The terms "first," "second," etc., do not limit the quantity or order of execution, and "first," "second," etc., do not necessarily imply differences.
[0032] It should be noted that, in this application, the terms "exemplary" or "for example" are used to indicate that something is being described as an example, illustration, or illustration. Any embodiment or design described as "exemplary" or "for example" in this application should not be construed as being more preferred or advantageous than other embodiments or design solutions. Specifically, the use of terms such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.
[0033] like Figure 1 As shown, this application provides a speech separation system 10, which may include a computer device 11, a sound acquisition device 12, and an image acquisition device 13.
[0034] The computer device 11 in this embodiment can be a terminal device or a network device. The terminal device can be referred to as: terminal, user equipment (UE), terminal device, access terminal, user unit, user station, mobile station, remote station, remote terminal, mobile device, user terminal, wireless communication device, user agent, or user equipment, etc. Specifically, the terminal device can be a mobile phone, augmented reality (AR) device, virtual reality (VR) device, tablet computer, laptop computer, ultra-mobile personal computer (UMPC), netbook, personal digital assistant (PDA), etc. The network device can specifically be a server, etc. The server can be a single physical or logical server, or two or more physical or logical servers sharing different responsibilities and cooperating to achieve the various functions of the server.
[0035] The sound acquisition device 12 in this embodiment can be a microphone array, which refers to a system composed of multiple microphones used for sound acquisition. The number of microphones can be 8, 16, 32, etc., and different numbers of microphones result in different topologies. Each microphone converts the acquired sound signal into an electrical signal. As each microphone element constituting the array, the acquired sound signal can be used for audio processing, such as noise reduction, speech separation, etc.
[0036] The image acquisition device 13 in this embodiment can be a camera for capturing still images or videos. The voice separation system may include one or N cameras, where N is a positive integer greater than 1.
[0037] It should be noted that the embodiments of this application do not impose any restrictions on the specific forms of computer equipment, sound acquisition equipment and image acquisition equipment. The following will uniformly use computer equipment, microphone array and camera as examples for description.
[0038] The devices in the speech separation system of this application can be deployed in different ways. The specific deployment method is set based on the camera's ability to capture the first sound source, without limiting the specific location. One possible implementation is that the camera and microphone array are set up independently. With this deployment method, devices can be supplemented and adjusted based on the existing scenario, thus simplifying implementation and reducing the complexity of device configuration.
[0039] Another possible implementation is to integrate the camera with the microphone array, such as... Figure 2 As shown, camera M is deployed at the center of a 32-channel microphone array. This deployment method integrates the audio and image acquisition devices, resulting in a smaller footprint and facilitating coordinate transformation during subsequent sound source localization.
[0040] It should be noted that in this application, when the camera and microphone array are integrated together, the positional relationship between the camera and the microphone array is not limited. That is, the camera does not necessarily have to be deployed in the center of the microphone array, but can also be deployed in other positions of the microphone array.
[0041] The method provided in this application embodiment can be applied to, for example... Figure 3 The scenario shown is a multi-person conference, where the acquisition device is based on... Figure 2 The device, which integrates a camera and microphone array, represents two sound sources in a multi-person conference, A and B. Typically, in multi-person conferences, multiple sound sources may be emitting sound simultaneously or alternately. Figure 3 The acquisition device shown is capable of acquiring mixed sound signals and images from multiple sound sources, which can be applied to the speech separation method in this application.
[0042] like Figure 4 The diagram shown is a hardware structure schematic of a computer device 40 provided in an embodiment of this application. The computer device 40 can be used to implement the functions of the aforementioned computer device.
[0043] Figure 4 The computer device 40 shown may include a processor 401, a memory 402, a communication interface 403, and a bus 404. The processor 401, the memory 402, and the communication interface 403 can be connected via the bus 404.
[0044] The processor 401 is the control center of the computer device 40. It can be a general-purpose central processing unit (CPU) or other general-purpose processors. The general-purpose processor can be a microprocessor or any conventional processor.
[0045] As an example, processor 401 may include one or more CPUs, for example Figure 4 CPU 0 and CPU 1 are shown in the diagram.
[0046] The memory 402 may be a read-only memory (ROM) or other type of static storage device capable of storing static information and instructions, random access memory (RAM) or other type of dynamic storage device capable of storing information and instructions, or electrically erasable programmable read-only memory (EEPROM), disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but is not limited thereto.
[0047] In one possible implementation, the memory 402 can exist independently of the processor 401. The memory 402 can be connected to the processor 401 via a bus 404 and is used to store data, instructions, or program code. When the processor 401 calls and executes the instructions or program code stored in the memory 402, it can implement the speech separation method provided in the embodiments of this application.
[0048] In another possible implementation, the memory 402 can also be integrated with the processor 401.
[0049] Communication interface 403 is used for connecting computer device 40 to other devices via a communication network, which may be Ethernet, radio access network (RAN), wireless local area network (WLAN), etc. Communication interface 403 may include a receiving unit for receiving data and a transmitting unit for transmitting data.
[0050] Bus 404 can be an industry standard architecture (ISA) bus, a peripheral component interconnect (PCI) bus, or an extended industry standard architecture (EISA) bus. This bus can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 4 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.
[0051] It should be pointed out that, Figure 4 The structure shown does not constitute a limitation on computer device 40, except... Figure 4 In addition to the components shown, computer device 40 may include more or fewer components than illustrated, or combine certain components, or have different component arrangements.
[0052] To make the embodiments of this application clearer, the following is a brief introduction to the concepts and some contents related to the embodiments of this application.
[0053] 1. Speech separation
[0054] Speech separation refers to separating the desired target sound from the background sound, which can also be understood as removing the sound signals in the background sound that interfere with the target sound signal.
[0055] 2. Beamforming
[0056] Beamforming refers to the process of processing sound signals using appropriate algorithms to form a beam pointing towards a sound source. Specifically, this can include time delay or phase compensation, amplitude weighting, and other processing of the sound signals collected by each element in a microphone array. Beamforming algorithms are divided into adaptive beamforming algorithms and non-adaptive beamforming algorithms. Adaptive beamforming algorithms include minimum variance distortionless response (MVDR) adaptive beamforming algorithms, generalized side-lobe canceller (GSC) adaptive beamforming algorithms, and so on.
[0057] 3. Sound source localization
[0058] Sound source localization refers to determining the spatial location of a sound source, which is usually based on the sound signal emitted by the sound source.
[0059] 4. Speech recognition
[0060] Speech recognition refers to the process by which computer devices recognize sound signals and convert them into natural human language for display, or convert them into machine language that computers can recognize and execute the instructions carried in the sound signals.
[0061] like Figure 5 The diagram shown is a flowchart of a speech separation method provided in this application. The method includes:
[0062] S501. The computer device acquires any one image of the target scene captured by the camera within a preset time period, and a mixed sound signal of the target scene captured by the microphone array within the preset time period. The image to be processed includes an image of a first sound source, and the mixed sound signal is a mixture of the sound signal from the first sound source and other sound signals.
[0063] The target scene is the environment in which the camera and microphone array are located. Optionally, the target scene can be an enclosed space (such as a room), in which the camera and microphone array are installed. Of course, the target scene can also be an open space, such as a spacious area.
[0064] The target scene includes a first sound source, which can be any sound source in the target scene. The sound source can be a person or other objects, and this application embodiment does not limit this. In the specific embodiments of this application, the sound source is described as a person.
[0065] The mixed sound signal collected by the microphone array is composed of the sound signal from the first sound source (i.e., the sound signal emitted by the first sound source) and other sound signals.
[0066] Optionally, other sound signals include ambient sound (or noise) and / or sound signals from at least one second sound source.
[0067] For example, in the application scenario of single-person voice separation, the mixed sound signal collected by the microphone array can be composed of a person's voice signal and ambient sound.
[0068] For example, in multi-person speech separation application scenarios (such as multi-person dialogue application scenarios), the mixed sound signal collected by the microphone array can be a mixture of the sound signals of multiple people, or the mixed sound signal collected by the microphone array can be a mixture of the sound signals of multiple people and ambient sound.
[0069] The image captured by the camera is the image of the target scene. The image to be processed can be any image captured by the camera within a preset time period.
[0070] Optionally, the position of the first sound source in the images captured by the camera does not change within a preset time period, or the change is within an allowable range. In this way, by referring to any image of the target scene captured by the camera within the preset time period (i.e., the image to be processed), separating the sound signal of the first sound source from the mixed sound signal of the target scene captured by the microphone array within the preset time period helps to achieve better separation results.
[0071] In other words, in one example, the technical solution provided by the embodiments of this application can be applied to scenarios where the location of the sound source does not change or changes by a small amount (negligible) over a period of time.
[0072] S502, The computer device determines the position information of the first sound source image in the image to be processed.
[0073] This application does not limit the specific implementation of S502. For example, when the first sound source is a person, the computer device can determine the position information of the first sound source's image in the image to be processed based on a head and shoulder detection algorithm.
[0074] For example, such as Figure 6 As shown, Image 1 includes person A and person B. Images of person A and person B can be identified from Image 1 using a head and shoulder detection algorithm. The head and shoulder detection algorithm refers to the ability to identify the head and shoulder region of a human body in an image or video using deep learning technology. This head and shoulder region can be a geometric shape, such as... Figure 6 The rectangular dashed box shown encloses the heads and shoulders of personnel A and personnel B respectively.
[0075] Optionally, the computer device predicts the location of the human voice source, i.e., the specific location of the sound source, based on the identified head and shoulder area. For example, such as... Figure 6 The auxiliary lines shown indicate that the mouth of a person is located approximately one-third of the way down from the center of the head and shoulder area. The coordinates of this position are used as the positional information of person A and person B in image 1.
[0076] It should be noted that there is no limitation on the method for determining the specific location of the sound source in the head and shoulder region. If the head and shoulder detection algorithm is more accurate, a more precise method can be used to calculate the specific location of the sound source.
[0077] S503. The computer device determines the first orientation of the first sound source relative to the microphone array based on the position information of the first sound source image in the image to be processed and the orientation information of the camera relative to the microphone array.
[0078] Optionally, S503 may include:
[0079] S503A: The computer device determines the orientation information of the first sound source relative to the camera based on the position information of the image of the first sound source in the image to be processed.
[0080] It is understandable that, given a fixed camera position, the positional information of an object within the camera's field of view within the captured image can characterize the object's position relative to the camera. This positional information can include distance and orientation information.
[0081] For example, still using the above Figure 6 For example, such as Figure 7 As shown, a pixel coordinate system is constructed with the width of image 1 as the horizontal axis u, the height of image 1 as the vertical axis v, and Oc as the origin. Here, w and h are the width and height of image 1, respectively, and O is the pixel center of the image. The pixel coordinates of O in the pixel coordinate system are... Establish a planar coordinate system with O as the origin, consisting of a horizontal axis (x) parallel to u and a vertical axis (y) parallel to v. Assume person A is... Figure 7 In this context, p represents the pixel coordinates (u). p v p The pixel coordinates represent the position of person A in image 1. The angle θ (azimuth angle) between p and O can be used to represent the orientation of person A relative to the camera. The azimuth angle θ can be obtained from the pixel coordinates of p and O using the following formula:
[0082] Among them, l pp' l is the length of pp' Op' The length between Op' is
[0083] It should be noted that the above formula for calculating the azimuth angle θ can be transformed based on trigonometric functions, and this application does not impose any restrictions on this.
[0084] S503B: The computer device determines the first position of the first sound source relative to the microphone array based on the positional information of the first sound source relative to the camera and the positional information of the camera relative to the microphone array.
[0085] Based on the above Figure 2 In the deployment shown, when the camera is at the center of the microphone array, it can be understood that the position information of the first sound source relative to the camera is the same as the position information of the first sound source relative to the microphone array. Therefore, the first azimuth information of the first sound source relative to the microphone array is the azimuth angle θ in the example above.
[0086] It should be noted that when the camera and microphone array are integrated, the positional relationship between the devices can be ignored; that is, the azimuth angle of the camera relative to the microphone array, used to represent directional information, can be approximated as zero. Therefore, the positional information of the first sound source relative to the camera is the same as the positional information of the first sound source relative to the microphone array. Based on this deployment method, the computer device requires less computation when determining the first directional information, and the sound source localization speed is faster.
[0087] When the camera and microphone array are deployed independently, the computer device, in addition to determining the positional relationship between the first sound source and the camera, also needs to determine the positional relationship between the microphone array and the camera. Based on the positional relationship between the first sound source and the camera, and the positional relationship between the camera and the microphone array, the first azimuth information of the first sound source relative to the microphone array is obtained.
[0088] Specifically, the orientation information of the camera relative to the microphone array is fixed after these two devices are installed. Therefore, the orientation information of the camera relative to the microphone array can be preset in the computer device before implementing the technical solution provided in this embodiment.
[0089] Once the computer device obtains the orientation information of the first sound source relative to the camera and the orientation information of the camera relative to the microphone array, it can convert the information to obtain the orientation information of the first sound source relative to the microphone array. This orientation information is the first orientation information.
[0090] S504. The computer device enhances the sound signal from the first direction in the mixed sound signal and suppresses the sound signals from other directions besides the first direction to obtain the sound signal from the first sound source.
[0091] The azimuth angle obtained in step S504 is applied to beamforming technology;
[0092] The beam can be enhanced using the MVDR or GSC algorithm in beamforming, which are well-known techniques to those skilled in the art and will not be elaborated here.
[0093] Through the above steps, computer equipment can separate and locate individual sound sources from mixed sound signals using images, thus achieving speech separation and improving the accuracy of speech separation.
[0094] Optionally, after step S504 above, the computer device may also perform the following steps S505-S506.
[0095] S505. The computer device determines the second orientation of the second sound source relative to the microphone array based on the position information of the image of the second sound source in the image to be processed and the orientation information of the camera relative to the microphone array.
[0096] S506. The computer device enhances the second-position sound signal in the mixed sound signal and suppresses sound signals from other positions besides the second position to obtain the sound signal of the second sound source.
[0097] It is understood that after step S506, the computer device obtains the sound signal from the first sound source and the sound signal from the second sound source, respectively. The sound signals from directions other than the second direction should include the sound signal from the first sound source in step S504. The sound signals from directions other than the first direction in step S504 should include the sound signal from the second sound source.
[0098] For example, such as Figure 6 As shown, assuming person A is the first sound source and person B is the second sound source, after the computer device determines the sound signal of person A based on the above steps S501-S504, the computer device can also determine the sound signal of person B.
[0099] By identifying the sound signals from multiple sound sources using the above method, speech separation can be achieved in scenarios with multiple sound sources. Since the location information of the sound source and microphone array is determined by using images, the interference of mixed sound signals from multiple sound sources on the sound source localization can be avoided, thereby improving the accuracy of speech separation.
[0100] Optionally, after performing steps S501-S506, the separated audio signal can be subjected to speech recognition to automatically record information from the audio signal. For example, the audio signal can be converted into text and automatically saved through speech recognition, thereby realizing the function of automatically saving the audio source record from audio acquisition, which helps to improve the intelligence of computer equipment.
[0101] The foregoing primarily describes the solutions of the embodiments of this application from a methodological perspective. It is understood that, in order to achieve the above-described functions, the computer device includes at least one of the hardware structures and software modules corresponding to the execution of each function. Those skilled in the art should readily recognize that, based on the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed in hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0102] This application embodiment can divide a computer device into functional units based on the above method examples. For example, each function can be divided into its own functional units, or two or more functions can be integrated into one processing unit. The integrated unit can be implemented in hardware or as a software functional unit. It should be noted that the unit division in this application embodiment is illustrative and only represents one logical functional division; other division methods may be used in actual implementation.
[0103] For example, Figure 8 A possible structural schematic diagram of the computer device (denoted as computer device 80) involved in the above embodiments is shown. The computer device 80 includes an acquisition unit 801, a determination unit 802, and a processing unit 803. The acquisition unit 801 is used to acquire any one image to be processed from the target scene acquired by the image acquisition device within a preset time period, and a mixed sound signal of the target scene acquired by the sound acquisition device within the preset time period; the image to be processed includes an image of a first sound source, and the mixed sound signal is composed of the sound signal from the first sound source and other sound signals. For example, Figure 5 The step S501 shown is as follows. The determining unit 802 is used to determine the first orientation of the first sound source relative to the sound acquisition device based on the position information of the first sound source's image within the image to be processed, and the orientation information of the image acquisition device relative to the sound acquisition device. For example, Figure 5 The steps S502 and S503 are shown. Processing unit 803 is used to enhance the sound signal from the first direction in the mixed sound signal and suppress sound signals from other directions besides the first direction, to obtain the sound signal from the first sound source. For example, Figure 5 Step S504 is shown.
[0104] Optionally, the image to be processed also includes the image of the second sound source, and other sound signals include the sound signal of the second sound source; the determining unit 802 is further configured to determine the second orientation of the second sound source relative to the sound acquisition device based on the position information of the image of the second sound source in the image to be processed and the orientation information of the image acquisition device relative to the sound acquisition device; the processing unit 803 is further configured to enhance the sound signal of the second orientation in the mixed sound signal and suppress the sound signals of other orientations besides the second orientation to obtain the sound signal of the second sound source.
[0105] Optionally, the determining unit 802 is specifically used to determine the orientation information of the first sound source relative to the image acquisition device based on the position information of the image of the first sound source in the image to be processed; and to determine the first orientation of the first sound source relative to the sound acquisition device based on the orientation information of the first sound source relative to the image acquisition device and the orientation information of the image acquisition device relative to the sound acquisition device.
[0106] Optionally, if the first sound source is a person, the determining unit 802 is also used to determine the position information of the image of the first sound source in the image to be processed by using a head and shoulder detection algorithm.
[0107] Optionally, the processing unit 803 is specifically used to enhance the first directional sound signal in the mixed sound signal based on the beamforming method.
[0108] Optionally, the image acquisition device and the sound acquisition device are integrated together. Optionally, the computer device 80 also includes a storage unit 804. The storage unit 804 is used to store computer execution instructions, and other units in the computer device can perform corresponding actions according to the computer execution instructions stored in the storage unit 804.
[0109] For a detailed description of the above-mentioned optional methods, please refer to the foregoing method embodiments, which will not be repeated here. Furthermore, the explanation of any of the computer devices 80 provided above and the description of their beneficial effects can be found in the corresponding method embodiments described above, and will not be repeated here.
[0110] As an example, combined Figure 4 The functions implemented by some or all of the acquisition unit 801, determination unit 802, processing unit 803, and storage unit 804 in computer device 80 can be achieved through... Figure 4 Processor 401 in the middle executes Figure 4 The program code in memory 402 is used for implementation. The acquisition unit 801 can also be accessed via... Figure 4 The receiving unit in the communication interface 403 is implemented.
[0111] This application also provides a computer-readable storage medium storing a computer program that, when run on a computer, causes the computer to perform the methods executed by any of the computer devices described above.
[0112] For explanations of the relevant content and descriptions of the beneficial effects in any of the computer-readable storage media provided above, please refer to the corresponding embodiments described above, which will not be repeated here.
[0113] This application also provides a chip. The chip integrates a control circuit for implementing the functions of the aforementioned computer device 80 and one or more ports. Optionally, the functions supported by the chip can be referred to above, and will not be repeated here. Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium. The aforementioned storage medium can be a read-only memory, random access memory, etc. The aforementioned processing unit or processor can be a central processing unit, a general-purpose processor, an application-specific integrated circuit (ASIC), a microprocessor (digital signal processor, DSP), a field-programmable gate array (FPGA), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof.
[0114] This application also provides a computer program product containing instructions that, when executed on a computer, cause the computer to perform any of the methods described in the above embodiments. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the flow or function according to the embodiments of this application is generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer instructions may be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can access or may include one or more data storage devices such as servers or data centers that can be integrated with the medium. The available medium may be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., SSD), etc.
[0115] It should be noted that the devices for storing computer instructions or computer programs provided in the embodiments of this application, such as but not limited to the memory, computer-readable storage medium and communication chip, are all non-transitory.
[0116] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented using software programs, implementation can be, in whole or in part, in the form of a computer program product. This computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the flow or function according to the embodiments of this application is generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer instructions can be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer or a data storage device containing one or more servers, data centers, etc., that can be integrated with the medium. The available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state disks, SSDs).
[0117] Although this application has been described herein in conjunction with various embodiments, those skilled in the art, by reviewing the accompanying drawings, the disclosure, and the appended claims, will understand and implement other variations of the disclosed embodiments in carrying out the claimed application. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude multiple instances. A single processor or other unit can implement several functions listed in the claims. While different dependent claims may recite certain measures, this does not mean that these measures cannot be combined to produce good results.
[0118] Although this application has been described in conjunction with specific features and embodiments, it is obvious that various modifications and combinations can be made thereto without departing from the spirit and scope of this application. Accordingly, this specification and drawings are merely exemplary illustrations of this application as defined by the appended claims, and are considered to cover any and all modifications, variations, combinations, or equivalents within the scope of this application. Clearly, those skilled in the art can make various alterations and modifications to this application without departing from the spirit and scope of this application. Thus, if such modifications and modifications of this application fall within the scope of the claims of this application and their equivalents, this application is also intended to include such modifications and modifications.
Claims
1. A speech separation method, characterized in that, include: The system acquires any one image of a target scene to be processed, captured by an image acquisition device within a preset time period, and a mixed sound signal of the target scene captured by a sound acquisition device within the preset time period; the image to be processed includes an image of a first sound source, and the mixed sound signal is composed of the sound signal of the first sound source and other sound signals; the first sound source is a person. The position information of the first sound source in the image to be processed is determined based on the head and shoulder detection algorithm; Based on the position information of the first sound source image in the image to be processed, the azimuth angle of the first sound source relative to the image acquisition device is determined by the following formula; Where θ is the azimuth angle, w and h are the width and height of the image to be processed, respectively, and the pixel coordinates of the first sound source in the image to be processed are (u p v p ); Based on the azimuth angle of the first sound source relative to the image acquisition device and the azimuth information of the image acquisition device relative to the sound acquisition device, the first azimuth of the first sound source relative to the sound acquisition device is determined; The sound signal from the first direction in the mixed sound signal is enhanced, and the sound signals from other directions are suppressed to obtain the sound signal from the first sound source.
2. The method according to claim 1, characterized in that, The image to be processed further includes an image of a second sound source, and the other sound signals include sound signals from the second sound source; the method further includes: Based on the position information of the image of the second sound source in the image to be processed, and the orientation information of the image acquisition device relative to the sound acquisition device, the second orientation of the second sound source relative to the sound acquisition device is determined; The sound signal from the second direction in the mixed sound signal is enhanced, and the sound signals from other directions besides the second direction are suppressed to obtain the sound signal from the second sound source.
3. The method according to claim 1 or 2, characterized in that, The enhancement of the first directional sound signal in the mixed sound signal includes: Based on beamforming methods, the sound signal in the first direction of the mixed sound signal is enhanced.
4. The method according to claim 1 or 2, characterized in that, The image acquisition device and the sound acquisition device are integrated together.
5. A computer device, characterized in that, include: The acquisition unit is used to acquire any one image to be processed from the target scene acquired by the image acquisition device within a preset time period, and a mixed sound signal of the target scene acquired by the sound acquisition device within the preset time period; the image to be processed includes an image of a first sound source, and the mixed sound signal is composed of the sound signal of the first sound source and other sound signals; the first sound source is a person. The determining unit is used to determine the position information of the image of the first sound source in the image to be processed by using a head and shoulder detection algorithm; The determining unit is further configured to determine the azimuth angle of the first sound source relative to the image acquisition device based on the position information of the first sound source image in the image to be processed by the following formula; Where θ is the azimuth angle, w and h are the width and height of the image to be processed, respectively, and the pixel coordinates of the first sound source in the image to be processed are (u p v p ); Based on the azimuth angle of the first sound source relative to the image acquisition device and the azimuth information of the image acquisition device relative to the sound acquisition device, the first azimuth of the first sound source relative to the sound acquisition device is determined; The processing unit is used to enhance the sound signal from the first direction in the mixed sound signal and suppress the sound signals from other directions besides the first direction to obtain the sound signal from the first sound source.
6. The computer device according to claim 5, characterized in that, The image to be processed also includes an image of a second sound source, and the other sound signals include sound signals from the second sound source; The determining unit is further configured to determine the second orientation of the second sound source relative to the sound acquisition device based on the position information of the image of the second sound source in the image to be processed, and the orientation information of the image acquisition device relative to the sound acquisition device; The processing unit is further configured to enhance the sound signal from the second direction in the mixed sound signal and suppress sound signals from other directions besides the second direction to obtain the sound signal from the second sound source.
7. The computer device according to claim 5 or 6, characterized in that, The processing unit is specifically used to enhance the first directional sound signal in the mixed sound signal based on a beamforming method.
8. The computer device according to claim 5 or 6, characterized in that, The image acquisition device and the sound acquisition device are integrated together.
9. A computer device, characterized in that, include: processor; The processor is connected to a memory for storing computer execution instructions, and the processor executes the computer execution instructions stored in the memory to enable the computer device to implement the method as described in any one of claims 1-4.
10. A computer-readable storage medium, characterized in that, Used to store computer instructions that, when executed on a computer, cause the computer to perform the method of any one of claims 1-4.
Citation Information
Patent Citations
Audio signal processing device and method and electronic device
CN106782584A