Audio separation method, separation device, processor, and electronic device
By calculating the angular relationship between the camera and microphone array, separating and splicing audio tracks, and using a voice activity detection algorithm, the problem of audio data cannot be automatically separated in the prior art is solved, and accurate extraction of target audio and convenient playback of conference records is achieved.
Patent Information
- Application Number
- CN202210892100.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-27
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2042-07-27
AI Technical Summary
In the prior art, audio in videos acquired based on configurations of camera and microphone arrays cannot be quickly positioned to the audio of a specific speaker, making it difficult to play back the conference record.
By calculating multiple angles, the audio tracks are separated and spliced based on the coordinate system relationship of the camera and microphone array, and the target audio is extracted using a speech activity detection algorithm.
实现了对目标对象音频的准确自动分离,方便会议记录的快速回放。
Smart Images

Figure CN115295005B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of audio separation. Specifically, it relates to a method and apparatus for separating audio, a computer-readable storage medium, a processor, and an electronic device. Background Art
[0002] In the prior art, video and audio are usually collected separately based on the configuration of a camera and a microphone array. As a result, there will be only one audio track in a video obtained in this way. When using the above configuration of the camera and the microphone array for meeting recording, since the audio of multiple speakers is on the same audio track, it is impossible to quickly locate all the audio of a speaker from the beginning to the end of the meeting during the automatic playback of the meeting recording.
[0003] Therefore, there is an urgent need for a method that can automatically separate audio data so that each speaker corresponds to an audio track, facilitating the quick location of all the audio of a certain speaker from the beginning to the end of the meeting during the playback of the meeting recording.
[0004] The above information disclosed in the background art section is only used to enhance the understanding of the background art of the technology described in this article. Therefore, the background art may contain certain information that is not prior art known to those skilled in the art in their own country. Summary of the Invention
[0005] The main object of the present application is to provide a method and apparatus for separating audio, a computer-readable storage medium, a processor, and an electronic device to solve the problem in the prior art that it is difficult to automatically separate audio data.
[0006] According to one aspect of an embodiment of the present invention, a method for separating audio is provided. A camera includes a camera and a microphone array. The camera is used to collect an original video, and the microphone array is used to collect multiple original audios. The separation method includes: calculating multiple second angles based on at least multiple first angles, where the first angle is the included angle between a first straight line and a second straight line. The first straight line is the connection line between the position point of the center of the camera and the position point of the target object, and the second straight line is any straight line where one coordinate axis of the camera coordinate system is located. The second angle is the included angle between a third straight line and a fourth straight line. The third straight line is the connection line between the position point of the center of the microphone array and the position point of the target object, and the fourth straight line is any straight line where one coordinate axis of the microphone coordinate system is located. The microphone array is located in the space constructed by the microphone coordinate system; obtaining multiple audio tracks based on the multiple original audios and the multiple second angles, where one audio track corresponds to one second angle; splicing the multiple audio tracks belonging to the same target object to obtain a target audio track corresponding to the target object, and using a voice activity detection algorithm to detect the target audio track to obtain the target audio of the target object.
[0007] Optionally, calculating multiple second angles based on at least multiple first angles includes: determining a relationship matrix for pose mutual conversion between the camera coordinate system and the microphone coordinate system; multiplying the multiple first angles by the relationship matrix to obtain the multiple second angles.
[0008] Optionally, the original video includes multiple consecutive images to be detected. The process of determining multiple first angles includes: detecting each of the images to be detected to obtain multiple position information groups of the multiple target objects, where each target object corresponds to one position information group, and each image to be detected corresponds to at least one position information group. The position information group includes information representing the position of the smallest rectangular region; determining the center points of each of the smallest rectangular regions and the center point of the image to be detected; inputting the center points of each of the smallest rectangular regions and the center point of the image to be detected into a camera imaging model to obtain the multiple first angles.
[0009] Optionally, the separation method further includes: sending at least the target image and the target audio track to a terminal device, so that a display screen of the terminal device displays the target image and a target scroll bar, the target scroll bar being located on one side of the corresponding target image, the target scroll bar being an icon of the target audio track, and the terminal device playing a corresponding part of the target audio track in response to a first predetermined operation applied to the target scroll bar; sending the original video to the terminal device, so that the display screen of the terminal device displays a video icon of the original video, and the terminal device playing the original video in response to a second predetermined operation applied to the video icon of the original video.
[0010] Optionally, the process of determining the target image includes: determining a plurality of predetermined regions corresponding to a plurality of position information groups of a plurality of target objects in each image to be detected, and cropping the plurality of predetermined regions to obtain a plurality of predetermined images, where one position information group corresponds to one predetermined region; using an image quality evaluation algorithm to evaluate the quality of the plurality of predetermined images belonging to the same target object to obtain the target image, the target image being the image with the optimal image quality among the plurality of predetermined images.
[0011] Optionally, sending at least the target image and the target audio track to the terminal device further includes: sending time information of the target audio on the target audio track to the terminal device, so that the terminal device displays a target marker on the target scroll bar, the target marker being generated according to the time information.
[0012] According to another aspect of the embodiments of the present invention, there is also provided a device for separating audio, including: a camera comprising a camera and a microphone array, the camera being configured to collect an original video, the microphone array being configured to collect a plurality of original audio signals, the separating device including: a first calculation unit configured to calculate a plurality of second angles based on at least a plurality of first angles, wherein the first angle is an included angle between a first straight line and a second straight line, the first straight line being a line connecting a position point of the center of the camera and a position point of a target object, the second straight line being a straight line where any one coordinate axis of the coordinate system of the camera is located, the second angle being an included angle between a third straight line and a fourth straight line, the third straight line being a line connecting a position point of the center of the microphone array and the position point of the target object, the fourth straight line being a straight line where any one coordinate axis of the microphone coordinate system is located, and the microphone array being located within a space constructed by the microphone coordinate system; a second calculation unit configured to obtain a plurality of audio tracks based on the plurality of original audio signals and the plurality of second angles, one audio track corresponding to one second angle; a first detection unit configured to splice the plurality of audio tracks belonging to the same target object to obtain a target audio track corresponding to the target object, and to detect the target audio track using a voice activity detection algorithm to obtain the target audio of the target object.
[0013] According to yet another aspect of the embodiments of the present invention, there is also provided a computer-readable storage medium, the computer-readable storage medium including a stored program, wherein the program executes any one of the methods described above.
[0014] According to still another aspect of the embodiments of the present invention, there is also provided a processor, the processor being configured to run a program, wherein when the program runs, it executes any one of the methods described above.
[0015] According to one aspect of the embodiments of the present invention, there is also provided an electronic device, including: one or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the one or more programs include those for executing any one of the methods described above.
[0016] In an embodiment of the present invention, in the audio separation method, first, multiple second angles are calculated based on at least multiple first angles. Then, multiple audio signals are obtained based on the multiple original audio signals and the multiple second angles. Finally, multiple audio tracks of the same target object are spliced to obtain a target audio track corresponding to the target object, and a voice activity detection algorithm is used to detect the target audio track to obtain the target audio of the target object. In this separation method, multiple audio signals are obtained based on the multiple original audio signals and the calculated multiple second angles, and then the multiple audio signals corresponding to the same target object are spliced to obtain the target audio track of the target object. Finally, the voice activity detection algorithm is used to detect the target audio track of the target object to obtain the target audio of the target object on the target audio track. That is, this solution realizes the automatic separation of the audio of the target object in the entire original video, forming one target audio track corresponding to one target object, and the voice activity detection method is used to detect the audible audio of the target object on the target audio track to obtain the target audio. This solution can more accurately automatically separate the audio of the target object in the entire original video, ensuring that the obtained target audio track and target audio are relatively accurate, and solving the problem that it is difficult to automatically separate audio data in the prior art. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The accompanying drawings forming a part of this application are used to provide a further understanding of this application. The schematic embodiments and descriptions thereof of this application are used to explain this application and do not constitute an improper limitation to this application. In the drawings:
[0018] Figure 1 shows a flowchart of a method for separating audio according to an embodiment of this application;
[0019] Figure 2 shows a schematic diagram displayed by a terminal device according to an embodiment of this application;
[0020] Figure 3 shows a schematic structural diagram of an audio separation device according to an embodiment of this application.
[0021] Among them, the above-mentioned accompanying drawings include the following reference numerals:
[0022] 100, video icon; 200, target image; 300, target scroll bar; 400, target marker. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0023] It should be noted that, without conflict, the embodiments in this application and the features in the embodiments can be combined with each other. The following will refer to the accompanying drawings and combine the embodiments to describe this application in detail.
[0024] To enable those skilled in the art to better understand the solution of this application, the technical solution in the embodiments of this application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this application.
[0025] It should be noted that the terms "first", "second", etc. in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so as to implement the embodiments of this application described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0026] As described in the background art, it is difficult to automatically separate audio data in the prior art. To solve the above problems, in a typical embodiment of this application, a method for separating audio, a separation device, a computer-readable storage medium, a processor, and an electronic device are provided.
[0027] According to an embodiment of this application, a method for separating audio is provided.
[0028] Figure 1 is a flowchart of a method for separating audio according to an embodiment of this application. The camera includes a camera and a microphone array. The above-mentioned camera is used to collect the original video, and the above-mentioned microphone array is used to collect multiple original audio, such as Figure 1 As shown, the separation method includes the following steps:
[0029] Step S101, calculate multiple second angles at least according to multiple first angles, where the above-mentioned first angle is the included angle between a first straight line and a second straight line, the above-mentioned first straight line is the connection line between the position point of the center of the above-mentioned camera and the position point of the target object, the above-mentioned second straight line is any one of the coordinate axes of the coordinate system of the above-mentioned camera, the above-mentioned second angle is the included angle between a third straight line and a fourth straight line, the above-mentioned third straight line is the connection line between the position point of the center of the above-mentioned microphone array and the position point of the above-mentioned target object, the above-mentioned fourth straight line is any one of the coordinate axes of the microphone coordinate system, and the above-mentioned microphone array is located in the space constructed by the above-mentioned microphone coordinate system;
[0030] Step S102: Based on the multiple original audios and the multiple second angles, obtain multiple audio tracks, where one of the audio tracks corresponds to one of the second angles;
[0031] Step S103: splice the multiple audio tracks belonging to the same target object to obtain the target audio track corresponding to the target object, and use a voice activity detection algorithm to detect the target audio track to obtain the target audio of the target object.
[0032] In the above audio separation method, first, calculate multiple second angles based on at least multiple first angles. Then, perform calculations based on the multiple original audios and the multiple second angles to obtain multiple audios. Finally, splice the multiple audio tracks belonging to the same target object to obtain the target audio track corresponding to the target object, and use a voice activity detection algorithm to detect the target audio track to obtain the target audio of the target object. In this separation method, based on the multiple original audios and the calculated multiple second angles, multiple audios are obtained, and then the multiple audios corresponding to the same target object are spliced to obtain the target audio track of the target object. Finally, the voice activity detection algorithm is used to detect the target audio track of the target object to obtain the target audio of the target object on the target audio track. That is, this solution realizes the automatic separation of the audio of the target object in the entire original video, forms one target audio track corresponding to one target object, and uses the voice activity detection method to detect the audible audio of the target object on the target audio track to obtain the target audio. This solution can relatively accurately automatically separate the audio of the target object in the entire original video, ensuring that the obtained target audio track and target audio are relatively accurate, and solving the problem that it is difficult to automatically separate audio data in the prior art.
[0033] Specifically, splice the audio tracks belonging to the same target object to obtain the target audio track corresponding to the target object. This target audio track is the audio track of the target object in the entire original video, that is, this target audio track includes audible and inaudible audios. Use a voice activity detection algorithm to detect the target audio track of the target object to obtain the target audio of the target object. This target audio is an audible segment of the target object on the entire target audio track.
[0034] In the actual application process, after detecting the target audio of the target object, the user can directly play the target audio according to actual needs, which is convenient for the user to perform playback. For example, in a meeting scenario, this solution can automatically record the meeting.
[0035] Specifically, when the above target object moves in the original video, a tracking algorithm can also be used to track the target object.
[0036] In a specific embodiment of the present application, a beamforming algorithm is used to calculate multiple original audio and multiple second angles to obtain multiple audio tracks.
[0037] In an embodiment of the present application, in order to accurately determine the above-mentioned second angles, at least multiple first angles are used to calculate multiple second angles, including: determining a relationship matrix for the mutual transformation of poses between the coordinate system of the above-mentioned camera and the coordinate system of the above-mentioned microphone; multiplying the multiple above-mentioned first angles by the relationship matrix to obtain the multiple above-mentioned second angles.
[0038] In another embodiment of the present application, the above-mentioned original video includes multiple consecutive images to be detected. The process of determining the multiple above-mentioned first angles includes: detecting each of the above-mentioned images to be detected to obtain multiple position information groups of the multiple above-mentioned target objects, where each of the above-mentioned target objects corresponds to one of the above-mentioned position information groups, and each of the above-mentioned images to be detected corresponds to at least one of the above-mentioned position information groups. The above-mentioned position information group includes information representing the position of the smallest rectangular region; determining the center points of each of the above-mentioned smallest rectangular regions and the center point of the above-mentioned image to be detected; inputting the center points of each of the above-mentioned smallest rectangular regions and the center point of the above-mentioned image to be detected into a camera imaging model to obtain the multiple above-mentioned first angles. In this embodiment, each image to be detected is detected to obtain multiple position information groups of the target objects, and the position information group is information representing the position of the smallest rectangular region, that is, each image to be detected is detected to obtain the smallest rectangular region corresponding to each target object, and then the center point corresponding to the smallest rectangular region and the center point of the image to be detected are input into the camera imaging model to obtain multiple first angles, which ensures that the first angles can be obtained more accurately, and further ensures that the obtained second angles are more accurate.
[0039] Specifically, the above-mentioned position information group includes the position information of a first target point and the position information of a second target point. The above-mentioned first target point and the above-mentioned second target point are on the target diagonal line, and the above-mentioned target diagonal line is a diagonal line of the smallest rectangular region including the target object in the detected image to be detected.
[0040] In the actual application process, the above-mentioned target object can be a person. When the above-mentioned target object is a person, at least one person is included in the image to be detected. Therefore, face detection can be performed on each of the above-mentioned images to be detected, and a position information group of the face of the target object in the entire image to be detected can be obtained. The smallest rectangular area including the target object can be determined through the above-mentioned position information group. Of course, the above-mentioned target object is not limited to a person. The above-mentioned target object can also be any object that can emit sound. When the above-mentioned target object is an object that can emit sound, at least one object that can emit sound is included in the image to be detected. Therefore, object detection can be performed on the above-mentioned image to be detected, and a position information group of the target object in the entire image to be detected can be obtained. The smallest rectangular area including the target object can be determined through the above-mentioned position information group.
[0041] Specifically, in this application, whether it is face detection on each image to be detected or object detection on each image to be detected, it can be realized through detection algorithms in the prior art. For example, in the case of face detection, algorithms such as RetinaFace and MTCNN (Multi-task Convolutional Neural Network) can be used to realize it; in the case of object detection, algorithms such as YoLo algorithm and Faster R-CNN algorithm can be used to realize it.
[0042] To further ensure that the user can play back more conveniently, in another embodiment of this application, as Figure 2 shown, the above-mentioned separation method further includes: sending at least the target image 200 and the above-mentioned target audio track to the terminal device, so that the display screen of the terminal device displays the target image 200 and the target scroll bar 300. The target scroll bar 300 is located on one side of the corresponding target image 200. The target scroll bar 300 is an icon of the target audio track. When the terminal device responds to a first predetermined operation on the target scroll bar 300, it plays the corresponding part of the target audio track; sending the above-mentioned original video to the terminal device, so that the display screen of the terminal device displays the video icon 100 of the above-mentioned original video. When the terminal device responds to a second predetermined operation on the video icon 100 of the above-mentioned original video, it plays the above-mentioned original video.
[0043] In the actual application process, as Figure 2As shown, the target scroll bar 300 can be located on one side of the target image 200. Of course, it is not limited to setting the target scroll bar 300 on one side of the target image 200. The target scroll bar 300 can also be set above the target image 200 or below the target image 200. Specifically, when the target scroll bar 300 is located on one side of the target image 200, the video icon 100 of the original video can be located above the entire target scroll bar 300 and the target image 200, or can also be located below the entire target scroll bar 300 and the target image 200. That is, in this solution, there is no limitation on the setting methods of the video icon 100 of the original video, the target image 200, and the target scroll bar 300 mentioned above.
[0044] Specifically, when the above terminal device responds to a first predetermined operation on the above target scroll bar, the corresponding part of the above target audio track is played, that is, the target audio is played. In this case, the above original video can also automatically jump to the position corresponding to the target audio for playing. Of course, when the above terminal device responds to a second predetermined operation on the video icon of the above original video and plays the above original video, it can also automatically jump to the target audio on the corresponding target audio track according to the playing position of the original video, that is, this ensures that the user can conveniently perform playback.
[0045] In another embodiment of the present application, the process of determining the above target image includes: determining a plurality of predetermined regions corresponding to a plurality of position information groups of a plurality of the above target objects in each image to be detected, and cropping the plurality of the above predetermined regions to obtain a plurality of predetermined images, where one of the above position information groups corresponds to one of the above predetermined regions; using an image quality evaluation algorithm to evaluate the quality of a plurality of the above predetermined images belonging to the same above target object to obtain the above target image, and the above target image is the image with the best image quality among the plurality of the above predetermined images. In this embodiment, according to a plurality of position information groups corresponding to a plurality of target objects, a plurality of predetermined regions are determined in each image to be detected, the plurality of predetermined regions are cropped to obtain a plurality of predetermined images, and then an image quality evaluation algorithm is used to evaluate a plurality of predetermined images of the same target object to determine the target image. Subsequently, the determined target image is sent to the terminal device so that the target image is displayed, further ensuring a better display effect and further facilitating the user to distinguish each target object.
[0046] In an embodiment of the present application, in order to facilitate the user to more intuitively observe the target audio on the terminal device, as Figure 2As shown, at least the target image and the above-mentioned target audio track are sent to the terminal device, further including: sending the time information of the above-mentioned target audio on the above-mentioned target audio track to the above-mentioned terminal device, so that the above-mentioned terminal device displays a target marker 400 on the above-mentioned target scroll bar 300, and the above-mentioned target marker 400 is generated according to the above-mentioned time information.
[0047] In the actual application process, the above-mentioned target marker is any marker that displays the target audio on the target scroll bar. For example, the above-mentioned target marker can be a color or other graphic shapes. For example, when the above-mentioned target marker is a color, the above-mentioned target marker can be yellow, green, etc.
[0048] It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0049] The embodiment of the present application also provides a device for separating audio. It should be noted that the device for separating audio in the embodiment of the present application can be used to execute the method for separating audio provided by the embodiment of the present application. The following introduces the device for separating audio provided by the embodiment of the present application.
[0050] Figure 3 is a schematic structural diagram of a device for separating audio according to an embodiment of the present application. The camera includes a camera and a microphone array. The above-mentioned camera is used to collect the original video, and the above-mentioned microphone array is used to collect multiple original audio, such as Figure 3 As shown, the separation device includes:
[0051] A first calculation unit 10, configured to calculate multiple second angles at least according to multiple first angles, where the above-mentioned first angle is the included angle between a first straight line and a second straight line, the above-mentioned first straight line is the connection line between the position point of the center of the above-mentioned camera and the position point of the target object, the above-mentioned second straight line is any coordinate axis of the coordinate system of the above-mentioned camera, the above-mentioned second angle is the included angle between a third straight line and a fourth straight line, the above-mentioned third straight line is the connection line between the position point of the center of the above-mentioned microphone array and the position point of the target object, the above-mentioned fourth straight line is any coordinate axis of the microphone coordinate system, and the above-mentioned microphone array is located in the space constructed by the above-mentioned microphone coordinate system;
[0052] A second calculation unit 20, configured to obtain multiple audio tracks based on the multiple above-mentioned original audio and the multiple above-mentioned second angles, and one above-mentioned audio track corresponds to one above-mentioned second angle;
[0053] The first detection unit 30 is configured to splice multiple of the above-mentioned audio tracks belonging to the same above-mentioned target object to obtain a target audio track corresponding to the above-mentioned target object, and use a voice activity detection algorithm to detect the above-mentioned target audio track to obtain the target audio of the above-mentioned target object.
[0054] In the above-mentioned audio separation device, the first calculation unit is configured to calculate multiple second angles at least according to multiple first angles, where the above-mentioned first angle is the included angle between a first straight line and a second straight line, the above-mentioned first straight line is the connection line between the position point of the center of the above-mentioned camera and the position point of the target object, the above-mentioned second straight line is any straight line where an axis of the coordinate system of the above-mentioned camera is located, the above-mentioned second angle is the included angle between a third straight line and a fourth straight line, the above-mentioned third straight line is the connection line between the position point of the center of the above-mentioned microphone array and the position point of the above-mentioned target object, the above-mentioned fourth straight line is any straight line where an axis of the microphone coordinate system is located, and the above-mentioned microphone array is located in the space constructed by the above-mentioned microphone coordinate system; the second calculation unit is configured to obtain multiple audio tracks based on the multiple above-mentioned original audio and the multiple above-mentioned second angles, and one of the above-mentioned audio tracks corresponds to one of the above-mentioned second angles; the first detection unit is configured to splice multiple of the above-mentioned audio tracks belonging to the same above-mentioned target object to obtain a target audio track corresponding to the above-mentioned target object, and use a voice activity detection algorithm to detect the above-mentioned target audio track to obtain the target audio of the above-mentioned target object. In this separation device, multiple audios are obtained based on the multiple original audios and the calculated multiple second angles, and then the multiple audios corresponding to the same target object are spliced to obtain the target audio track of the target object. Finally, the voice activity detection algorithm is used to detect the target audio track of the target object to obtain the target audio of the target object on the target audio track. That is, this solution realizes the automatic separation of the audio of the target object in the entire original video, forms one target audio track corresponding to one target object, and detects the audible audio of the target object on the target audio track through the voice activity detection method to obtain the target audio. This solution can more accurately automatically separate the audio of the target object in the entire original video, ensure that the obtained target audio track and target audio are relatively accurate, and solve the problem that it is difficult to automatically separate audio data in the prior art.
[0055] Specifically, the audio tracks belonging to the same target object are spliced to obtain a target audio track corresponding to the target object. This target audio track is the audio track of the target object in the entire original video, that is, this target audio track includes audible and inaudible audio. The voice activity detection algorithm is used to detect the target audio track of the target object to obtain the target audio of the target object. This target audio is a section of audible audio of the target object on the entire target audio track.
[0056] In the actual application process, after detecting the target audio of the target object, the user can directly play the target audio according to actual needs, which is convenient for the user to play back. For example, in a meeting scenario, this solution can automatically record the meeting.
[0057] Specifically, when the above-mentioned target object moves in the original video, a tracking algorithm can also be used to track the target object.
[0058] In a specific embodiment of the present application, a beamforming algorithm is used to calculate multiple original audio and multiple second angles to obtain multiple audio tracks.
[0059] In order to accurately determine the above-mentioned second angle, in an embodiment of the present application, the above-mentioned first calculation unit includes a first determination module and a multiplication module. Among them, the above-mentioned first determination module is used to determine the relationship matrix for the mutual pose conversion between the coordinate system of the above-mentioned camera and the coordinate system of the above-mentioned microphone; the above-mentioned multiplication module is used to multiply multiple above-mentioned first angles by the above-mentioned relationship matrix to obtain multiple above-mentioned second angles.
[0060] In another embodiment of the present application, the above-mentioned original video includes multiple consecutive images to be detected, and the separation device further includes a second detection unit, a determination unit, and an input unit. Among them, the above-mentioned second detection unit is used to detect each of the above-mentioned images to be detected to obtain multiple position information groups of multiple above-mentioned target objects. Among them, each of the above-mentioned target objects corresponds to one of the above-mentioned position information groups, and each of the above-mentioned images to be detected corresponds to at least one of the above-mentioned position information groups. The above-mentioned position information group includes information representing the position of the smallest rectangular area; the above-mentioned determination unit is used to determine the center point of each of the above-mentioned smallest rectangular areas and the center point of the above-mentioned image to be detected; the above-mentioned input unit is used to input the center points of each of the above-mentioned smallest rectangular areas and the center point of the above-mentioned image to be detected into the camera imaging model to obtain multiple above-mentioned first angles. In this embodiment, each image to be detected is detected to obtain multiple position information groups of the target object, and the position information group is information representing the position of the smallest rectangular area, that is, each image to be detected is detected to obtain the smallest rectangular area corresponding to each target object, and then the center point corresponding to the smallest rectangular area and the center point of the image to be detected are input into the camera imaging model to obtain multiple first angles, which ensures that the first angles can be obtained more accurately, and further ensures that the obtained second angles are more accurate.
[0061] Specifically, the above-mentioned position information group includes the position information of the first target point and the position information of the second target point. The above-mentioned first target point and the above-mentioned second target point are on the target diagonal line, and the above-mentioned target diagonal line is a diagonal line of the smallest rectangular area including the target object in the detected image to be detected.
[0062] In the actual application process, the above-mentioned target object can be a person. When the above-mentioned target object is a person, at least one person is included in the image to be detected. Therefore, face detection can be performed on each of the above-mentioned images to be detected, and a position information group of the face of the target object in the entire image to be detected can be obtained. The smallest rectangular area including the target object can be determined through the above-mentioned position information group. Of course, the above-mentioned target object is not limited to a person. The above-mentioned target object can also be any object that can emit sound. When the above-mentioned target object is an object that can emit sound, at least one object that can emit sound is included in the image to be detected. Therefore, object detection can be performed on the above-mentioned image to be detected, and a position information group of the target object in the entire image to be detected can be obtained. The smallest rectangular area including the target object can be determined through the above-mentioned position information group.
[0063] Specifically, in the present application, whether it is face detection on each image to be detected or object detection on each image to be detected, it can be implemented by detection algorithms in the prior art. For example, in the case of face detection, algorithms such as RetinaFace and MTCNN (Multi-task Convolutional Neural Network) can be used; in the case of object detection, algorithms such as YoLo algorithm and Faster R-CNN algorithm can be used.
[0064] In order to further ensure that the user can play back more conveniently, in another embodiment of the present application, as Figure 2 shown, the above-mentioned separation device further includes a first sending unit and a second sending unit. The first sending unit is used to send at least the target image 200 and the above-mentioned target audio track to the terminal device, so that the display screen of the terminal device displays the target image 200 and the target scroll bar 300. The target scroll bar 300 is located on one side of the corresponding target image 200. The target scroll bar 300 is an icon of the above-mentioned target audio track. When the terminal device responds to a first predetermined operation on the target scroll bar 300, it plays the corresponding part of the above-mentioned target audio track; the second sending unit is used to send the above-mentioned original video to the terminal device, so that the display screen of the terminal device displays the video icon 100 of the above-mentioned original video. When the terminal device responds to a second predetermined operation on the video icon 100 of the above-mentioned original video, it plays the above-mentioned original video.
[0065] In the actual application process, as Figure 2As shown, the target scroll bar 300 can be located on one side of the target image 200. Of course, it is not limited to setting the target scroll bar 300 on one side of the target image 200. The target scroll bar 300 can also be set above the target image 200, or the target scroll bar 300 can be set below the target image 200. Specifically, when the target scroll bar 300 is located on one side of the target image 200, the video icon 100 of the original video can be located above the entire target scroll bar 300 and the target image 200, or can also be located below the entire target scroll bar 300 and the target image 200. That is, in this solution, there is no limitation on the setting method of the video icon 100 of the original video, the target image 200, and the target scroll bar 300 as described above.
[0066] Specifically, when the above terminal device responds to a first predetermined operation acting on the above target scroll bar, the corresponding part of the above target audio track is played, that is, the target audio is played. In this case, the above original video can also automatically jump to the position corresponding to the target audio for playing. Of course, when the above terminal device responds to a second predetermined operation acting on the video icon of the above original video and plays the above original video, it can also automatically jump to the target audio on the corresponding target audio track according to the playing position of the original video, that is, this ensures that the user can conveniently perform playback.
[0067] In another embodiment of the present application, the above first sending unit includes a second determination module and an evaluation module. Among them, the second determination module is used to determine a plurality of predetermined regions corresponding to a plurality of position information groups of a plurality of the above target objects in each image to be detected, and crop the plurality of the above predetermined regions to obtain a plurality of predetermined images, where one of the above position information groups corresponds to one of the above predetermined regions; the evaluation module is used to evaluate the quality of a plurality of the above predetermined images belonging to the same above target object by using an image quality evaluation algorithm to obtain the above target image, and the above target image is the image with the optimal image quality among the plurality of the above predetermined images. In this embodiment, according to a plurality of position information groups corresponding to a plurality of target objects, a plurality of predetermined regions are determined in each image to be detected, and the plurality of predetermined regions are cropped to obtain a plurality of predetermined images. Then, an image quality evaluation algorithm is used to evaluate a plurality of predetermined images of the same target object to determine the target image, and the determined target image is subsequently sent to the terminal device so that the target image is displayed, further ensuring a better display effect and further facilitating the user to distinguish each target object.
[0068] In order to facilitate the user to more intuitively observe the target audio on the terminal device, in one embodiment of the present application, as Figure 2As shown, the above-mentioned first sending unit further includes a sending module, configured to send the time information of the above-mentioned target audio on the above-mentioned target audio track to the above-mentioned terminal device, so that the above-mentioned terminal device displays a target marker 400 on the above-mentioned target scroll bar 300, and the above-mentioned target marker 400 is generated according to the above-mentioned time information.
[0069] In actual application, the above-mentioned target marker is any marker that displays the target audio on the target scroll bar. For example, the above-mentioned target marker can be a color or other graphic shapes. For example, when the above-mentioned target marker is a color, the above-mentioned target marker can be yellow, green, etc.
[0070] The separation device of the above-mentioned audio includes a processor and a memory. The above-mentioned first calculation unit, second calculation unit, first detection unit, etc. are all stored in the memory as program units, and the processor executes the above-mentioned program units stored in the memory to implement corresponding functions.
[0071] The processor contains a kernel, and the kernel retrieves the corresponding program units from the memory. One or more kernels can be set, and by adjusting the kernel parameters, the problem that it is difficult to automatically separate audio data in the prior art can be solved.
[0072] The memory may include non-permanent memory in a computer-readable medium, forms such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM), and the memory includes at least one storage chip.
[0073] An embodiment of the present invention provides a computer-readable storage medium, on which a program is stored, and when the program is executed by a processor, the above-mentioned audio separation method is implemented.
[0074] An embodiment of the present invention provides a processor, and the above-mentioned processor is used to run a program, wherein when the above-mentioned program runs, the above-mentioned audio separation method is executed.
[0075] An embodiment of the present invention provides an electronic device, which includes one or more processors, a memory, and one or more programs, wherein the above-mentioned one or more programs are stored in the above-mentioned memory and are configured to be executed by the above-mentioned one or more processors, and the above-mentioned one or more programs include those for executing any one of the above-mentioned methods.
[0076] An embodiment of the present invention provides a device, which includes a processor, a memory, and a program stored on the memory and executable on the processor. When the processor executes the program, at least the following steps are implemented:
[0077] Step S101, calculate a plurality of second angles based on at least a plurality of first angles, where the first angle is the included angle between a first straight line and a second straight line, the first straight line is the connection line between the position point of the center of the camera and the position point of the target object, the second straight line is any one of the coordinate axes of the coordinate system of the camera, the second angle is the included angle between a third straight line and a fourth straight line, the third straight line is the connection line between the position point of the center of the microphone array and the position point of the target object, the fourth straight line is any one of the coordinate axes of the microphone coordinate system, and the microphone array is located in the space constructed by the microphone coordinate system;
[0078] Step S102, obtain a plurality of audio tracks based on the plurality of original audio and the plurality of second angles, where one audio track corresponds to one second angle;
[0079] Step S103, splice the plurality of audio tracks belonging to the same target object to obtain the target audio track corresponding to the target object, and use a voice activity detection algorithm to detect the target audio track to obtain the target audio of the target object.
[0080] The device in this article can be a server, a PC, a PAD, a mobile phone, etc.
[0081] This application also provides a computer program product, which, when executed on a data processing device, is suitable for executing a program initialized with at least the following method steps:
[0082] Step S101, calculate a plurality of second angles based on at least a plurality of first angles, where the first angle is the included angle between a first straight line and a second straight line, the first straight line is the connection line between the position point of the center of the camera and the position point of the target object, the second straight line is any one of the coordinate axes of the coordinate system of the camera, the second angle is the included angle between a third straight line and a fourth straight line, the third straight line is the connection line between the position point of the center of the microphone array and the position point of the target object, the fourth straight line is any one of the coordinate axes of the microphone coordinate system, and the microphone array is located in the space constructed by the microphone coordinate system;
[0083] Step S102, obtain a plurality of audio tracks based on the plurality of original audio and the plurality of second angles, where one audio track corresponds to one second angle;
[0084] Step S103, splice the plurality of audio tracks belonging to the same target object to obtain the target audio track corresponding to the target object, and use a voice activity detection algorithm to detect the target audio track to obtain the target audio of the target object.
[0085] In the above embodiments of the present invention, the descriptions of the respective embodiments each have their own emphasis. For parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0086] In several embodiments provided by the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are merely illustrative. For example, the above division of units can be a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of units or modules can be in an electrical or other form.
[0087] The units described above as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0088] In addition, in each embodiment of the present invention, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0089] If the above integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the above methods in each embodiment of the present invention. The foregoing storage medium includes: USB flash drive, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, magnetic disk, or optical disk and other various media that can store program codes.
[0090] From the above description, it can be seen that the above embodiments of the present application achieve the following technical effects:
[0091] 1), in the audio separation method of the present application, first, at least according to a plurality of first angles, a plurality of second angles are calculated. Then, based on the plurality of original audios and the plurality of second angles, a plurality of audios are obtained. Finally, the plurality of audio tracks of the same target object are spliced to obtain a target audio track corresponding to the target object, and a voice activity detection algorithm is used to detect the target audio track to obtain the target audio of the target object. In this separation method, based on the plurality of original audios and the calculated plurality of second angles, a plurality of audios are obtained, and then the plurality of audios corresponding to the same target object are spliced to obtain the target audio track of the target object. Finally, the target audio track of the target object is detected by the voice activity detection algorithm to obtain the target audio of the target object on the target audio track. That is, this solution realizes the automatic separation of the audio of the target object in the entire original video, forms one target audio track corresponding to one target object, and detects the audible audio of the target object on the target audio track through the voice activity detection method to obtain the target audio. This solution can relatively accurately automatically separate the audio of the target object in the entire original video, ensuring that the obtained target audio track and target audio are relatively accurate, and solving the problem that it is difficult to automatically separate audio data in the prior art.
[0092] 2) In the audio separation device of the present application, the first calculation unit is used to calculate a plurality of second angles at least based on a plurality of first angles, where the first angle is the included angle between a first straight line and a second straight line, the first straight line is the connection line between the position point of the center of the camera and the position point of the target object, the second straight line is any straight line where an axis of the camera coordinate system is located, the second angle is the included angle between a third straight line and a fourth straight line, the third straight line is the connection line between the position point of the center of the microphone array and the position point of the target object, the fourth straight line is any straight line where an axis of the microphone coordinate system is located, and the microphone array is located in the space constructed by the microphone coordinate system; the second calculation unit is used to obtain a plurality of audio tracks based on the plurality of original audio and the plurality of second angles, and one audio track corresponds to one second angle; the first detection unit is used to splice the plurality of audio tracks belonging to the same target object to obtain the target audio track corresponding to the target object, and use a voice activity detection algorithm to detect the target audio track to obtain the target audio of the target object. In this separation device, calculations are performed based on a plurality of original audio and the calculated plurality of second angles to obtain a plurality of audio, then the plurality of audio corresponding to the same target object are spliced to obtain the target audio track of the target object, and finally the target audio track of the target object is detected by a voice activity detection algorithm to obtain the target audio of the target object on the target audio track. That is, this solution realizes the automatic separation of the audio of the target object in the entire original video, forms one target audio track corresponding to one target object, and detects the audible audio of the target object on the target audio track through a voice activity detection method to obtain the target audio. This solution can relatively accurately automatically separate the audio of the target object in the entire original video, ensuring that the obtained target audio track and target audio are relatively accurate, and solving the problem that it is difficult to automatically separate audio data in the prior art.
[0093] The above are only the preferred embodiments of the present application and are not used to limit the present application. For those skilled in the art, the present application can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A method for separating audio, characterized in that, The camera includes a camera and a microphone array. The camera is used to collect the original video, and the microphone array is used to collect multiple original audio signals. The separation method includes: Calculating multiple second angles based on at least multiple first angles. Wherein, the first angle is the included angle between a first straight line and a second straight line. The first straight line is the line connecting the position point of the center of the camera and the position point of the target object. The second straight line is any one of the coordinate axes of the camera's coordinate system. The second angle is the included angle between a third straight line and a fourth straight line. The third straight line is the line connecting the position point of the center of the microphone array and the position point of the target object. The fourth straight line is any one of the coordinate axes of the microphone coordinate system. The microphone array is located in the space constructed by the microphone coordinate system; Obtaining multiple audio tracks based on the multiple original audio signals and the multiple second angles. One audio track corresponds to one second angle; Splicing multiple audio tracks belonging to the same target object to obtain the target audio track corresponding to the target object, and using a voice activity detection algorithm to detect the target audio track to obtain the target audio of the target object; Wherein, the original video includes multiple consecutive images to be detected. The process of determining the multiple first angles includes: Detecting each of the images to be detected to obtain multiple position information groups of the multiple target objects. Wherein, each target object corresponds to one position information group, and each image to be detected corresponds to at least one position information group. The position information group includes information representing the position of the smallest rectangular region; Determining the center point of each of the smallest rectangular regions and the center point of the image to be detected; Inputting the center points of each of the smallest rectangular regions and the center point of the image to be detected into the camera imaging model to obtain the multiple first angles.
2. The separation method according to claim 1, characterized in that, Calculating multiple second angles based on at least multiple first angles includes: Determining the relationship matrix for the pose mutual conversion between the camera's coordinate system and the microphone coordinate system; Multiplying the multiple first angles by the relationship matrix to obtain the multiple second angles.
3. The separation method according to any one of claims 1 or 2, characterized in that The separation method further includes: Sending at least the target image and the target audio track to the terminal device so that the display screen of the terminal device displays the target image and a target scroll bar. The target scroll bar is located on one side of the corresponding target image. The target scroll bar is the icon of the target audio track. When the terminal device responds to a first predetermined operation on the target scroll bar, it plays the corresponding part of the target audio track; Sending the original video to the terminal device so that the display screen of the terminal device displays the video icon of the original video. When the terminal device responds to a second predetermined operation on the video icon of the original video, it plays the original video.
4. The method according to claim 3, characterized in that, The process of determining the target image includes: Determine multiple predetermined regions corresponding to multiple position information groups of multiple target objects in each image to be detected, and crop the multiple predetermined regions to obtain multiple predetermined images, where one position information group corresponds to one predetermined region; Use an image quality assessment algorithm to perform quality assessment on multiple predetermined images belonging to the same target object to obtain the target image, where the target image is the image with the optimal image quality among the multiple predetermined images.
5. The separation method according to claim 3, characterized in that, At least send the target image and the target audio track to the terminal device, and further include: Send the time information of the target audio on the target audio track to the terminal device, so that the terminal device displays a target marker on the target scroll bar, and the target marker is generated according to the time information.
6. An audio separation device, characterized in that, The camera includes a camera and a microphone array. The camera is used to collect the original video, and the microphone array is used to collect multiple original audios. The separation device includes: A first calculation unit, configured to calculate multiple second angles at least according to multiple first angles, where the first angle is the included angle between a first straight line and a second straight line. The first straight line is the connection line between the position point of the center of the camera and the position point of the target object, and the second straight line is any coordinate axis of the coordinate system of the camera. The second angle is the included angle between a third straight line and a fourth straight line. The third straight line is the connection line between the position point of the center of the microphone array and the position point of the target object, and the fourth straight line is any coordinate axis of the microphone coordinate system. The microphone array is located in the space constructed by the microphone coordinate system; A second calculation unit, configured to obtain multiple audio tracks based on the multiple original audios and the multiple second angles, where one audio track corresponds to one second angle; A first detection unit, configured to splice multiple audio tracks belonging to the same target object to obtain the target audio track corresponding to the target object, and use a voice activity detection algorithm to detect the target audio track to obtain the target audio of the target object; Wherein, the original video includes multiple consecutive images to be detected, and the separation device further includes a second detection unit, a determination unit, and an input unit; The second detection unit is configured to detect each image to be detected to obtain multiple position information groups of multiple target objects, where each target object corresponds to one position information group, and each image to be detected corresponds to at least one position information group. The position information group includes information representing the position of the smallest rectangular region; The determination unit is configured to determine the center points of each of the smallest rectangular regions and the center point of the image to be detected; The input unit is configured to input the center points of each of the smallest rectangular regions and the center point of the image to be detected into the camera imaging model to obtain multiple first angles.
7. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, where the program executes the method according to any one of claims 1 to 5.
8. A processor, characterized in that, The processor is used to run a program, wherein when the program runs, it executes the method described in any one of claims 1 to 5.
9. An electronic device, characterized in that, Comprising: One or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the one or more programs include those for executing the method described in any one of claims 1 to 5.
Citation Information
Patent Citations
Electrical device and associated operating method for displaying user interface related to a sound track
US20150363157A1
System and method for audio-visual multi-speaker speech separation with location-based selection
US20210312915A1