Sound acquisition method and device, computer equipment and storage medium

Through image processing technology, the target sound source body is determined and combined with sound generation actions to collect sound, which solves the problem of inaccurate sound acquisition for headphones in noisy environments, and realizes accurate sound acquisition and translation in noisy environments.

CN120295599APending Publication Date: 2025-07-11HUIZHOU TCL MOBILE COMM CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510365514.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

In noisy environments, existing translation headphones find it difficult to accurately collect the voice of the target's attention, resulting in a decrease in translation accuracy.

Method used

The target sound source body in the scene is determined through image processing technology, the target sound matching the target sound source body is collected, and the sound is collected based on the sound generation action, and the image and sound information are combined for precise positioning and acquisition.

Benefits of technology

The accuracy and reliability of sound acquisition of the target sound source body in noisy environments is achieved, and the translation accuracy of the translation headset in multi-person conversation scenarios is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120295599A_ABST
    Figure CN120295599A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a sound acquisition method and device, electronic equipment and a computer readable storage medium, and relates to the technical field of sound processing, and the method comprises the steps: determining a target sound source body in a scene image; collecting target sound matched with the target sound source body in a current scene; based on the scene image, determining a sounding action when the target sound source body generates the target sound; and based on the sound production action and the target sound, carrying out sound collection on the target sound source body. Therefore, positioning of the target sound source body can be realized through image processing, the speaking sound of the target sound source body can be matched in the current scene, and the sound of the target sound source body can be accurately and reliably acquired.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the technical field of sound processing, and more particularly, to a method, apparatus, computer device, and computer-readable storage medium for sound acquisition. Background Art

[0002] With the widespread application of large models, the addition of translation functions to headphones has become increasingly popular among users. Current translation headphone products have a high accuracy rate in quiet environments. However, in noisy environments or scenarios with multiple people talking, the collected sound will also be very noisy, greatly reducing the accuracy rate of translating the sound of the target object of interest.

[0003] Therefore, how to enable an electronic device to accurately and reliably collect the sound emitted by the target object of interest is an urgent problem to be solved at present. Summary of the Invention

[0004] Embodiments of the present disclosure provide a method, apparatus, electronic device, and computer-readable storage medium for sound acquisition, aiming to at least solve one of the technical problems in the related art to a certain extent.

[0005] In a first aspect, embodiments of the present disclosure provide a method for sound acquisition, the method comprising:

[0006] Determine a target sound source in the scene image;

[0007] Collect a target sound in the current scene that matches the target sound source;

[0008] Based on the scene image, determine the vocalization action of the target sound source when generating the target sound;

[0009] Based on the vocalization action and the target sound, perform sound acquisition on the target sound source.

[0010] In a second aspect, embodiments of the present disclosure further provide a device for sound acquisition, the device comprising:

[0011] A first determination module, configured to determine a target sound source in the scene image;

[0012] A first acquisition module, configured to collect a target sound in the current scene that matches the target sound source;

[0013] A second determination module, configured to determine the vocalization action of the target sound source when generating the target sound based on the scene image;

[0014] A second acquisition module, configured to perform sound acquisition on the target sound source based on the vocalization action and the target sound.

[0015] In a third aspect, an embodiment of the present disclosure further provides an electronic device, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, the steps in the above-mentioned sound acquisition method are implemented.

[0016] In a fourth aspect, an embodiment of the present disclosure further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned sound acquisition method are implemented.

[0017] In a fifth aspect, an embodiment of the present disclosure further provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods provided in various alternative implementations of the embodiments of the present disclosure.

[0018] In the embodiments of the present disclosure, first, a target sound source in the scene image is determined, then the target sound matching the target sound source in the current scene is collected, then based on the scene image, the vocalization action when the target sound source produces the target sound is determined, and finally, based on the vocalization action and the target sound, sound collection is performed on the target sound source. Thus, the positioning of the target sound source can be realized through image processing, and then the voice of the target sound source when it speaks can be matched in the current scene, and thus the sound of the target sound source can be accurately and reliably collected.

[0019] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions in the present disclosure, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative efforts.

[0021] Figure 1 is a schematic flowchart of the sound acquisition method provided by the first embodiment of the present disclosure;

[0022] Figure 2 is a schematic flowchart of the sound acquisition method provided by the second embodiment of the present disclosure;

[0023] Figure 3 is a schematic structural diagram of the sound acquisition device provided by the embodiments of the present disclosure;

[0024] Figure 4 It is a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure. Detailed implementation manners

[0025] Here, some embodiments of the present disclosure will be described in detail, and their examples are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. Various changes, modifications, and equivalents of the methods, apparatuses, and / or systems described herein will become apparent after understanding the present disclosure. For example, the order of operations described herein is merely an example and is not limited to those set forth herein, but can be changed as will be apparent after understanding the present disclosure, except for operations that must be performed in a specific order. Additionally, descriptions of features known in the art may be omitted for the sake of clarity and conciseness.

[0026] The implementation manners described in some embodiments of the present disclosure below do not represent all implementation manners consistent with the present disclosure. On the contrary, they are merely examples of apparatuses and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0027] It should be noted that the execution subject of the sound acquisition method in this embodiment can be a sound acquisition device, and this device can be configured in any type of electronic device, which is not limited herein.

[0028] In an embodiment of the present disclosure, the "sound acquisition device" will be used as the execution subject to execute the "sound acquisition method" for illustration, which is not limited herein.

[0029] It should be noted that the description order of the following embodiments does not limit the priority order of the embodiments.

[0030] Figure 1 It is a schematic flowchart of a sound acquisition method provided according to the first embodiment of the present disclosure.

[0031] As Figure 1 shown, the method includes:

[0032] Step 101, determining a target sound source object in a scene image.

[0033] Among them, the sound source object refers to an object that is making a sound, and can also be called a sound source or a sounding body. For example, if Zhang San is speaking, then Zhang San is the sound source object.

[0034] Among them, the target sound source object can be a sound source object that needs to be concerned or monitored in the current scene. For example, if there are Zhang San, Li Si, Wang Wu, and Zhao Liu in a room, and if the sound that needs to be concerned in the current scene is the sound made by Li Si, then "Li Si" can be used as the target sound source object, which is not limited herein.

[0035] Among them, the scene image can be an image obtained by photographing the current scene.

[0036] As a possible implementation, the scene image can be obtained first, and then the scene image can be recognized to determine at least one sound source object, and then the sound source object that meets the preset conditions in the scene image is used as the target sound source object.

[0037] Among them, the preset condition can be the target condition for selecting the target sound source object.

[0038] Specifically, a multi-sensor device with a camera and a microphone (such as a smart headset, a smartphone, a surveillance camera, or a dedicated acoustic monitoring device) is used to capture the scene image in real time or at regular intervals through the camera of the device. Then, the captured image can be preprocessed, including denoising, enhancing contrast, adjusting brightness, etc., to improve the recognition effect. After that, image processing algorithms can be used to extract the potential sound source object features in the image, such as shape, size, color, texture, and the contrast between the object and the background.

[0039] Alternatively, deep learning models (such as YOLO, Faster R-CNN, etc.) or traditional computer vision algorithms (such as edge detection, template matching, etc.) can also be applied to detect the objects in the image. These models or algorithms should be trained in advance to be able to recognize common sound source object types, such as speakers, musical instruments, vehicles, etc. Finally, a candidate list containing potential sound source objects can be generated according to the detection results.

[0040] Furthermore, a series of preset conditions can be set according to the requirements of the application scenario. These conditions may be based on the type, position, and size of the sound source object. Each sound source object in the candidate list is matched with the preset conditions, and the sound source object that meets all the preset conditions is determined as the target sound source object.

[0041] For example, if the preset condition is to be located directly in front of the scene image, the sound source object located directly in front of the scene image can be used as the target sound source object.

[0042] As another possible implementation, the scene image can be obtained first, and then the target sound source object selected by the user in the scene image can be obtained.

[0043] Specifically, the scene image captured by the user can be displayed in the graphical user interface. The device can provide selection tools, such as rectangular selection boxes, circular selection boxes, polygon selection boxes, or click selection, etc., to allow the user to mark the areas or objects of their interest in the image. If the system is complex or the user is not familiar with the operation, a user guide or help document can be provided to guide the user on how to select the target sound source object.

[0044] It should be noted that when the user uses the selection tool to mark the area of interest in the image, the device can recognize and capture the information of this area. The device can provide a confirmation button or mechanism to allow the user to confirm their selection.

[0045] Alternatively, the device can also send the scene image to other electronic devices so that other electronic devices can display the scene image. The user can select the target sound source on other electronic devices. Then, the selection information of the target sound source is sent to the device. For example, the earphone can collect the scene image and then transmit it to the mobile phone. The user selects the target sound source on the mobile phone and then sends the selection information to the earphone.

[0046] Step 102: Collect the target sound in the current scene that matches the target sound source.

[0047] Optionally, the initial sound of the target type in the current scene can be collected first, and then based on the sound characteristics of the target sound source, the initial sound is filtered to obtain the target sound corresponding to the sound characteristics.

[0048] It should be noted that due to different types of sound sources, the types of sounds can also be different, such as human voices, animal sounds, children's voices, device sounds (such as TV sounds), wind sounds, etc., which are not limited here.

[0049] Among them, the target type can be a pre-selected specific type of sound, such as a human voice, which is not limited here.

[0050] Among them, the initial sound can be the sound collected in the current scene according to the target type, and it can be the mixed sound of multiple sound sources of the same target type.

[0051] For example, there are 3 people in a room. The sound collected by the device in the current scene may include the voices of these 3 people, and may also include the ambient sound of the room, such as TV sounds, wind sounds, etc. If the target type of sound is a human voice, only the voices of these 3 people in this room can be collected and used as the initial sound.

[0052] It is understandable that the microphone array technology can be utilized to localize the sound source in the room by combining parameters such as time difference of arrival (TDOA) and direction of arrival (DOA). This helps to determine the positions of human voices and ambient sounds. Subsequently, based on the results of sound source localization, blind source separation (BSS) algorithms (such as independent component analysis ICA, non-negative matrix factorization NMF, etc.) or deep learning techniques (such as neural networks) can be used to separate the human voice from the sounds in the current scene. After that, feature extraction can be performed on the human voice signal, including spectral features (such as Mel-frequency cepstral coefficients MFCC), time-domain features (such as short-time energy, short-time zero-crossing rate), etc. These features can reflect the uniqueness of the human voice. Further, the extracted human voice features can be matched with a preset human voice feature library (or a machine learning-based classifier) to screen out the sound signals that match the target type (human voice). At the same time, a certain threshold can be set to ensure that the screened sound signals have a high confidence level. Thus, accurate human voice can be obtained as the initial sound.

[0053] Further, the target sound can be extracted from the initial sound according to the sound characteristics of the target sound source body.

[0054] It is understandable that the sound characteristics when different target sound source bodies emit sounds are different. For example, the sound characteristics when different people speak are different. Therefore, the device can separate the target sound from other sounds in the initial sound based on the unique sound characteristics corresponding to the target sound source body, so as to screen out the target sound.

[0055] It should be noted that everyone's voice has uniqueness. The voiceprint feature refers to the sound feature in the sound signal that can characterize and identify the unique attributes of the speaker. The voiceprint feature is the voice feature contained in the speech that can characterize and identify the speaker. For example, if the scene image is an image of a meeting room with 7 people in it, the voiceprint features of these 7 people are different. For example, Zhang San is the main speaker among the 7 people. Taking Zhang San as the target sound source body, the device can screen from the human voices in the current scene according to the pre-recorded voiceprint features of Zhang San, so as to extract Zhang San's speaking voice, that is, the target sound.

[0056] Step 103, based on the scene image, determine the vocalization action of the target sound source body when generating the target sound.

[0057] It should be noted that the vocalization action can be the action that the target sound source body needs to perform when generating the target sound. Taking the target sound source body as a person as an example, the vocalization action can be the action of the mouth.

[0058] Optionally, if the type of the target sound is a human voice, the mouth shape of the target sound source body in the scene image can be recognized to obtain the mouth vocalization action.

[0059] As an example, first, preprocess the scene image to improve the accuracy of recognizing the lip shape of the target sound source body in the scene image. The preprocessing steps may include image denoising, grayscale conversion, binarization, etc., in order to better extract the face and mouth regions. After that, in the preprocessed image, use face detection technology to detect and locate the face region. This can be achieved through deep learning-based face detection algorithms, such as Haar features, HOG+SVM, or deep learning models, etc. Once the face is detected, the mouth region can be further located.

[0060] Optionally, lip movement recognition can be adopted to detect the lip movement for sound production. Lip movement recognition is a key step in determining the sound production action of the target sound source body. For example, it can be implemented based on the method of feature vectors, such as extracting the feature vectors of the mouth region and using the Hidden Markov Model (HMM) or other machine learning algorithms for state matching. Or, it can also be based on the method of mouth shape classification. Since the mouth shapes are basically the same when a person makes the same sound and there are also great similarities in mouth shapes when making similar sounds, the changing mouth shapes for pronunciation can be clustered. Then, based on the clustering, lip reading recognition is performed to improve the accuracy and speed of recognition.

[0061] Step 104, perform sound acquisition on the target sound source body based on the sound production action and the target sound.

[0062] Optionally, first, based on the lip movement for sound production corresponding to the target sound source body, determine the start sound production time and the end sound production time. Then, based on the start sound production time and the end sound production time, select the first sound from the target sound. Then, based on the first sound and the lip movement for sound production, perform sound acquisition on the target sound source body.

[0063] Specifically, according to the lip movement for sound production of the target sound source body, mark the moment when the target sound source body starts speaking as the start sound production time, and mark the moment when the target sound source body ends speaking as the end sound production time. Then, the sound of the target sound source body speaking within the time period from the start sound production time to the end sound production time can be intercepted as the first sound.

[0064] Among them, the first sound can be the target sound generated by the target sound source body within the time period from the start sound production time to the end sound production time.

[0065] Optionally, when the sound signal intensity of the first sound is greater than or equal to the preset threshold, determine that the first sound is normal.

[0066] Optionally, when the sound signal intensity of the first sound is less than the preset threshold, determine that the first sound is abnormal.

[0067] It should be noted that the preset threshold is a pre-determined standard value for comparing the strength of sound signals. This value is usually determined based on factors such as the normal sound characteristics of the target sound source, the ambient noise level, and the performance of the recording equipment. If the sound signal strength of the first sound is greater than or equal to the preset threshold, the first sound is determined to be normal. If the sound signal strength of the first sound is less than the preset threshold, the first sound can be selectively determined to be abnormal. By measuring the sound signal strength of the first sound and comparing it with the preset threshold, it can be determined whether the first sound is normal or not. If the sound signal strength meets or exceeds the preset threshold, the first sound is determined to be normal. If the sound signal strength is lower than the preset threshold, the first sound is determined to be abnormal. By setting a reasonable preset threshold, normal or abnormal sound signals can be effectively screened out, providing a reliable basis for subsequent processing and analysis.

[0068] Optionally, when the mouth sound-making action is a continuous action and the first sound is normal, sound is collected from the target sound source.

[0069] It should be noted that if the mouth sound-making movement is a continuous movement, it means that the person is speaking, and if the first sound is normal, it means that the first sound is stable and reliable, and thus the sound can be collected normally, that is, the sound of the target sound source is collected.

[0070] Optionally, when the mouth sound production action is not performed and the first sound is abnormal, the sound collection is stopped;

[0071] It should be noted that if the mouth sound-making action is not moving, it means that no speech is made. If the first sound is abnormal, it means that the first sound is unstable or no sound can be detected, so the sound collection of the target sound source can be stopped.

[0072] Optionally, when the mouth sound-making action is not in progress and the first sound is normal, sound is collected from the target sound source.

[0073] It should be noted that if the mouth sound movement is not moving, it means that no speech is made. If the first sound is normal, it means that the first sound is stable and reliable. In this case, it may be because the target sound source is not recognized normally or the recognition is wrong, but the sound of the target sound source can be retained, that is, the sound of the target sound source is collected.

[0074] In the embodiments of the present disclosure, first, the target sound source object in the scene image is determined, then the target sound matching the target sound source object in the current scene is collected. After that, based on the scene image, the vocalization action when the target sound source object generates the target sound is determined. Finally, based on the vocalization action and the target sound, the sound of the target sound source object is collected. Thus, the positioning of the target sound source object can be achieved through image processing, and then its speaking sound can be matched in the current scene, and thus the sound of the target sound source object can be accurately and reliably collected.

[0075] Figure 2 It is a schematic flowchart of a method for collecting sound according to the second embodiment of the present disclosure.

[0076] As Figure 2 shown, the method includes:

[0077] Step 201, determine the target sound source object in the scene image.

[0078] Step 202, collect the target sound matching the target sound source object in the current scene.

[0079] Step 203, based on the scene image, determine the vocalization action when the target sound source object generates the target sound.

[0080] It should be noted that the specific implementation methods of steps 201-203 can refer to the above embodiments and will not be elaborated here.

[0081] Step 204, when the vocalization action at the mouth is a continuous action and the sound signal intensity of the first sound is less than a preset threshold, amplify the first sound.

[0082] It should be noted that the preset threshold is a pre-determined standard value for comparing the sound signal intensity. This value is usually determined comprehensively based on factors such as the normal vocalization characteristics of the target sound source object, the ambient noise level, and the performance of the recording device. If the sound signal intensity of the first sound is greater than or equal to the preset threshold, the first sound is determined to be normal. If the sound signal intensity of the first sound is less than the preset threshold, the first sound can be selectively determined to be abnormal. By measuring the sound signal intensity of the first sound and comparing it with the preset threshold, the normality of the first sound can be determined. If the sound signal intensity meets or exceeds the preset threshold, the first sound is determined to be normal. If the sound signal intensity is lower than the preset threshold, the first sound is determined to be abnormal. By setting a reasonable preset threshold, normal or abnormal sound signals can be effectively screened out, providing a reliable basis for subsequent processing and analysis.

[0083] It can be understood that if the sound signal intensity of the first sound is less than the preset threshold, the sound signal intensity of the first sound can be increased through amplification processing to improve the clarity and usability of the sound.

[0084] Step 205: If the sound signal intensity of the amplified first sound is greater than the preset threshold, collect the sound of the target sound source body; otherwise, stop the collection.

[0085] It can be understood that if the sound signal intensity of the first sound is greater than the preset threshold after amplification processing, it means that the first sound is stable and reliable, and its usability is improved. Therefore, the sound of the target sound source body can be collected, otherwise the collection can be stopped.

[0086] In the embodiments of the present disclosure, first, the target sound source body in the scene image is determined, then the target sound matching the target sound source body in the current scene is collected, then based on the scene image, the vocalization action of the target sound source body when generating the target sound is determined, and finally, when the vocalization action at the mouth is a continuous action and the sound signal intensity of the first sound is less than the preset threshold, the first sound is amplified. After that, if the sound signal intensity of the amplified first sound is greater than the preset threshold, the sound of the target sound source body is collected, otherwise the collection is stopped. By accurately identifying the target sound source body in the scene and its vocalization action, combined with the intelligent judgment and amplification processing of the sound signal intensity, the accuracy and flexibility of sound collection are effectively improved. On the premise of ensuring that the sound quality meets the preset standard, accurate collection of the target sound source body is carried out, avoiding ineffective collection and resource waste, and at the same time improving the application value of sound data and the accuracy of subsequent analysis.

[0087] To facilitate better implementation of the sound collection method of the present disclosure, the present disclosure also provides a sound collection device based on the above sound collection method. The meanings of the nouns are the same as those in the above sound collection method, and the specific implementation details can refer to the description in the method embodiments.

[0088] Please refer to Figure 3 , Figure 3 FIG. is a schematic structural diagram of the sound collection device provided by the embodiments of the present disclosure. The sound collection device 300 includes:

[0089] A first determination module 310, configured to determine the target sound source body in the scene image;

[0090] A first collection module 320, configured to collect the target sound matching the target sound source body in the current scene;

[0091] A second determination module 330, configured to determine the vocalization action of the target sound source body when generating the target sound based on the scene image;

[0092] A second collection module 340, configured to collect the sound of the target sound source body based on the vocalization action and the target sound.

[0093] Optionally, the first determining module 310 includes:

[0094] Get scene image;

[0095] Identify the scene image to determine at least one sound source;

[0096] The sound source body that meets the preset conditions in the scene image is taken as the target sound source body.

[0097] Optionally, the first determining module 310 includes:

[0098] Get scene image;

[0099] Obtain a target sound source body selected by a user in the scene image.

[0100] Optionally, the first acquisition module 320 is specifically used to:

[0101] Collect the initial sound of the target type in the current scene;

[0102] Based on the sound characteristics of the target sound source, the initial sound is screened to obtain a target sound corresponding to the sound characteristics.

[0103] Optionally, the type of the target sound is human voice, and the second determination module 330 is specifically configured to:

[0104] The mouth shape of the target sound source in the scene image is recognized to obtain the mouth sound-making action.

[0105] Optionally, the second acquisition module 340 includes:

[0106] A first determining unit, configured to determine a start time and an end time of the sound production based on a mouth sound production action corresponding to the target sound source body;

[0107] A selection unit, configured to select a first sound from the target sound based on the start sounding time and the end sounding time;

[0108] A collection unit is used to collect sound from the target sound source based on the first sound and the mouth sound production action.

[0109] Optionally, the acquisition unit is specifically used to:

[0110] When the mouth sounding action is a continuous action and the first sound is normal, collecting sound from the target sound source;

[0111] or,

[0112] When the mouth sound-making action is not performed and the first sound is abnormal, stop sound collection;

[0113] Or,

[0114] When the mouth sound-making action is not performed and the first sound is normal, perform sound collection on the target sound source body.

[0115] Optionally, the collection unit is specifically configured to:

[0116] When the mouth sound-making action is a continuous action and the sound signal intensity of the first sound is less than a preset threshold, perform amplification processing on the first sound;

[0117] If the sound signal intensity of the amplified first sound is greater than the preset threshold, perform sound collection on the target sound source body, otherwise stop collection.

[0118] Optionally, the apparatus further includes:

[0119] A third determination module, configured to determine that the first sound is normal when the sound signal intensity of the first sound is greater than or equal to the preset threshold;

[0120] A fourth determination module, configured to determine that the first sound is abnormal when the sound signal intensity of the first sound is less than the preset threshold.

[0121] In the embodiments of the present disclosure, first, the target sound source body in the scene image is determined, then the target sound matching the target sound source body in the current scene is collected, then based on the scene image, the sound-making action when the target sound source body generates the target sound is determined, and finally, based on the sound-making action and the target sound, sound collection is performed on the target sound source body. Thus, the positioning of the target sound source body can be achieved through image processing, and then the voice of the target sound source body when it speaks can be matched in the current scene, and thus the sound of the target sound source body can be accurately and reliably collected.

[0122] In addition, the present disclosure also provides an electronic device, as Figure 4 shown, which shows a schematic structural diagram of the electronic device involved in the present disclosure. Specifically:

[0123] The electronic device may include a processor 401 with one or more processing cores, a memory 402 with one or more computer-readable storage media, a power supply 403, an input unit 404, and other components. Those skilled in the art can understand that Figure 4 the structural diagram of the electronic device shown in does not constitute a limitation on the electronic device, and may include more or fewer components than shown, or combine certain components, or have different component arrangements. Among them:

[0124] The processor 401 is the control center of the electronic device, connecting various parts of the entire electronic device through various interfaces and circuits. By running or executing software programs and / or modules stored in the memory 402, and invoking the data stored in the memory 402, it executes various functions of the electronic device and processes data, thereby monitoring the electronic device as a whole. Optionally, the processor 401 may include one or more processing cores; preferably, the processor 401 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor may not be integrated into the processor 401 either.

[0125] The memory 402 can be used to store software programs and modules. The processor 401 executes various functional applications and data processing by running the software programs and modules stored in the memory 402. The memory 402 mainly includes a program storage area and a data storage area. Among them, the program storage area can store the operating system, application programs required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area can store data created according to the use of the electronic device, etc. In addition, the memory 402 may include high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, flash memory device, or other volatile solid-state storage devices. Correspondingly, the memory 402 may also include a memory controller to provide the processor 401 with access to the memory 402.

[0126] The electronic device also includes a power supply 403 for supplying power to each component. Preferably, the power supply 403 can be logically connected to the processor 401 through a power management system, so as to realize functions such as management of charging, discharging, and power consumption management through the power management system. The power supply 403 may also include any components such as one or more DC or AC power supplies, a recharge system, a power device debugging circuit, a power converter or inverter, and a power status indicator.

[0127] The electronic device may further include an input unit 404, which can be used to receive input digital or character information, and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function controls.

[0128] Although not shown, the electronic device may further include a display unit and the like, which will not be elaborated herein. Specifically, in this embodiment, the processor 401 in the electronic device will load the executable files corresponding to the processes of one or more application programs into the memory 402 according to the following instructions, and the processor 401 will run the application programs stored in the memory 402, so as to implement the steps in any of the sound acquisition methods provided by the embodiments of the present disclosure.

[0129] In the embodiments of the present disclosure, first, the target sound source body in the scene image is determined, then the target sound matching the target sound source body in the current scene is collected, then based on the scene image, the vocalization action when the target sound source body generates the target sound is determined, and finally, based on the vocalization action and the target sound, the sound of the target sound source body is collected. Thus, the positioning of the target sound source body can be realized through image processing, and then the sound of its speech can be matched in the current scene, and then the sound of the target sound source body can be accurately and reliably collected.

[0130] For the specific implementation of each of the above operations, reference can be made to the previous embodiments, which will not be elaborated herein.

[0131] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructions, or by controlling relevant hardware through instructions. The instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.

[0132] Therefore, the present disclosure provides a computer-readable storage medium, on which a computer program is stored. The computer program can be loaded by a processor to execute the steps in any of the sound acquisition methods provided by the present disclosure.

[0133] For the specific implementation of each of the above operations, reference can be made to the previous embodiments, which will not be elaborated herein.

[0134] Among them, the computer-readable storage medium may include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disk, etc.

[0135] Since the instructions stored in the computer-readable storage medium can execute the steps in any of the sound acquisition methods provided by the present disclosure, the beneficial effects that can be achieved by any of the sound acquisition methods provided by the present disclosure can be realized. For details, reference can be made to the previous embodiments, which will not be elaborated herein.

[0136] The above has introduced in detail a method, apparatus, electronic device, and computer-readable storage medium for collecting sound. In this article, specific examples are used to elaborate on the principles and implementation manners of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those skilled in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. A method for collecting sound, characterized in that, including: determining a target sound source in a scene image; collecting a target sound matching the target sound source in the current scene; determining a vocalization action when the target sound source produces the target sound based on the scene image; collecting sound for the target sound source based on the vocalization action and the target sound.

2. The method according to claim 1, wherein The determining of the target sound source in the scene image includes: obtaining a scene image; identifying the scene image to determine at least one sound source; using the sound source in the scene image that meets a preset condition as the target sound source.

3. The method according to claim 1, wherein The determining of the target sound source in the scene image includes: obtaining a scene image; obtaining the target sound source selected by the user in the scene image.

4. The method according to claim 1, wherein The collecting of the target sound matching the target sound source in the current scene includes: collecting an initial sound of a target type in the current scene; screening the initial sound based on the sound characteristics of the target sound source to obtain a target sound corresponding to the sound characteristics.

5. The method according to claim 1, characterized in that, When the type of the target sound is a human voice, the determining of the vocalization action when the target sound source produces the target sound based on the scene image includes: identifying the mouth shape of the target sound source in the scene image to obtain a mouth vocalization action.

6. The method according to claim 5, wherein The collecting of sound for the target sound source based on the vocalization action and the target sound includes: determining a start vocalization time and an end vocalization time based on the mouth vocalization action corresponding to the target sound source; selecting a first sound from the target sound based on the start vocalization time and the end vocalization time; collecting sound for the target sound source based on the first sound and the mouth vocalization action.

7. The method according to claim 6, wherein The collecting of sound for the target sound source based on the first sound includes: collecting sound for the target sound source when the mouth vocalization action is a continuous action and the first sound is normal; or, stopping sound collection when the mouth vocalization action is no action and the first sound is abnormal; or, collecting sound for the target sound source when the mouth vocalization action is no action and the first sound is normal.

8. The method according to claim 6, wherein The collecting of sound for the target sound source based on the first sound includes: performing amplification processing on the first sound when the mouth vocalization action is a continuous action and the sound signal intensity of the first sound is less than a preset threshold; if the sound signal intensity of the amplified first sound is greater than the preset threshold, collecting sound for the target sound source, otherwise stopping collection.

9. The method according to claim 7, wherein It further includes: determining that the first sound is normal when the sound signal intensity of the first sound is greater than or equal to the preset threshold; determining that the first sound is abnormal when the sound signal intensity of the first sound is less than the preset threshold.

10. A sound collection device, characterized in that, including: a first determination module for determining a target sound source in a scene image; a first collection module for collecting a target sound matching the target sound source in the current scene; A second determination module, configured to determine a sounding action of the target sound source when generating the target sound based on the scene image; A second acquisition module, configured to perform sound acquisition on the target sound source based on the sounding action and the target sound.

11. A computer device, characterized in that, It includes a processor and a memory, and the memory stores multiple instructions; the processor loads the instructions from the memory to execute the steps in the sound acquisition method according to any one of claims 1-9.

12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores multiple instructions, and the instructions are suitable for being loaded by a processor to execute the steps in the sound acquisition method according to any one of claims 1-9.