Voice tracking photographing method, system and device and nonvolatile storage medium
Patent Information
- Application Number
- CN202211718745.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-29
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2042-12-29
AI Technical Summary
[0004]本发明实施例提供了一种语音跟踪摄像方法、系统、装置及非易失性存储介质,以至少解决由于相关技术中的跟踪摄像方法无法根据说话主体进行摄像跟踪切换,造成的跟踪拍摄效果差,容易出现音视频缺失的技术问题
[0009]在本发明实施例中,通过基于麦克风阵列中包括的多个第一麦克风分别针对目标发声对象采集到的第一拾音幅度,确定摄像装置沿水平方向转动的初始水平方向角,以及沿垂直方向转动的初始俯仰角;控制上述摄像装置旋转至上述初始水平方向角和上述初始俯仰角所对应的第一拾音区域之后,获取主麦克风在多个不同的拾音方向针对上述目标发声对象分别采集到的第二拾音幅度;基于上述主麦克风在上述多个不同的拾音方向分别采集到的第二拾音幅度,确定上述摄像装置沿上述水平方向转动的目标水平方向角,以及沿上述垂直方向转动的目标俯仰角;控制上述摄像装置由上述初始水平方向角旋转至上述目标水平方向角,以及由上述初始俯仰角旋转至上述目标俯仰角进行视频图像采集,得到视频图像采集结果,达到了通过麦克风阵列和主麦克风结合的方式进行语音视频跟踪,准确针对目标发声对象进行位置跟踪和视频拍摄的目的,从而实现了提升视频跟踪效果和音视频流畅度的技术效果,进而解决了由于相关技术中的跟踪摄像方法无法根据说话主体进行摄像跟踪切换,造成的跟踪拍摄效果差,容易出现音视频缺失的技术问题。
Smart Images

Figure CN116389906B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and more specifically, to a voice tracking camera method, system, device, and non-volatile storage medium. Background Technology
[0002] In office settings, there is often a need to record meetings. Typically, only one speaker is speaking at a time. When interrupted or asked a question, the speaker pauses, and another speaker takes over. Meeting cameras are usually fixed and cannot follow the speaker in real-time. For example, some technologies offer a method based on facial recognition to track the speaker on stage, but this cannot keep up with the switching between different speakers, making accurate tracking and positioning difficult. This can lead to audio and video loss and poor recording quality.
[0003] There is currently no effective solution to the above problems. Summary of the Invention
[0004] This invention provides a voice tracking camera method, system, device, and non-volatile storage medium to at least solve the technical problem that the tracking camera method in the related art cannot switch the camera tracking according to the speaker, resulting in poor tracking and shooting effects and easy audio and video loss.
[0005] According to one aspect of the present invention, a voice tracking camera method is provided, comprising: determining an initial horizontal angle and an initial pitch angle of a camera device rotating in a horizontal direction based on a first pickup amplitude collected by a plurality of first microphones included in a microphone array for a target sound-emitting object; controlling the camera device to rotate to a first pickup area corresponding to the initial horizontal angle and the initial pitch angle, and then acquiring a second pickup amplitude collected by a main microphone for the target sound-emitting object in a plurality of different pickup directions; determining a target horizontal angle and a target pitch angle of the camera device rotating in the horizontal direction based on the second pickup amplitude collected by the main microphone in the plurality of different pickup directions; controlling the camera device to rotate from the initial horizontal angle to the target horizontal angle and from the initial pitch angle to the target pitch angle to acquire video images, thereby obtaining a video image acquisition result.
[0006] According to another aspect of the present invention, a voice tracking camera system is also provided, comprising: a microphone array; a multi-channel micro-motor for controlling a plurality of first microphones included in the microphone array; a camera device; a main microphone parallel to the central axis of the camera device; an azimuth motor for controlling the horizontal rotation of the camera device; a pitch motor for controlling the vertical rotation of the camera device; and a digital processing device, wherein the number of channels of the multi-channel micro-motor corresponds to the number of the plurality of first microphones, and is used to control the pitch angles corresponding to the plurality of first microphones respectively; wherein the microphone array is connected to the multi-channel micro-motor and the digital processing device; the multi-channel micro-motor is connected to the digital processing device; the main microphone is connected to the camera device and the digital processing device; the camera device is connected to the azimuth motor and the pitch motor; the azimuth motor and the pitch motor are respectively connected to the digital processing device; wherein the plurality of first microphones are used to collect a first pickup amplitude emitted by a target sound source, and transmit the first pickup amplitude collected for the target sound source... The signal is sent to the aforementioned digital processing device; the aforementioned digital processing device is used to determine the initial horizontal angle of rotation of the aforementioned camera device in the horizontal direction and the initial pitch angle in the vertical direction based on the first pickup amplitude collected by the aforementioned plurality of first microphones for the aforementioned target sound-emitting object; control the aforementioned camera device to rotate to the first pickup area corresponding to the aforementioned initial horizontal angle and the aforementioned initial pitch angle; the aforementioned main microphone is used to collect the second pickup amplitude of the aforementioned target sound-emitting object in the aforementioned plurality of different pickup directions, and send the second pickup amplitude collected in the aforementioned plurality of different pickup directions to the aforementioned digital processing device; the aforementioned digital processing device is used to determine the target horizontal angle of rotation of the aforementioned camera device in the aforementioned horizontal direction and the target pitch angle in the aforementioned vertical direction based on the second pickup amplitude collected by the aforementioned main microphone in the aforementioned plurality of different pickup directions, and control the aforementioned camera device to rotate from the aforementioned initial horizontal angle to the aforementioned target horizontal angle, and from the aforementioned initial pitch angle to the aforementioned target pitch angle; the aforementioned camera device performs video image acquisition to obtain video image acquisition results.
[0007] According to another aspect of the present invention, a non-volatile storage medium is also provided, characterized in that the non-volatile storage medium stores a plurality of instructions, the instructions being adapted to be loaded by a processor and executed any one of the above-described voice tracking camera methods.
[0008] According to another aspect of the present invention, an electronic device is also provided, characterized in that it includes one or more processors and a memory, wherein the memory is used to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement any one of the above-described voice tracking camera methods.
[0009] In this embodiment of the invention, based on the first pickup amplitudes collected by multiple first microphones in the microphone array targeting the target sound source, an initial horizontal angle and an initial pitch angle of the camera device rotating in the horizontal direction are determined. After controlling the camera device to rotate to the first pickup area corresponding to the initial horizontal angle and the initial pitch angle, a second pickup amplitude is obtained by the main microphone collecting data from multiple different pickup directions targeting the target sound source. Based on the second pickup amplitudes collected by the main microphone from the multiple different pickup directions, a target horizontal angle of the camera device rotating in the horizontal direction is determined. The target pitch angle is rotated along the aforementioned vertical direction; the camera device is controlled to rotate from the aforementioned initial horizontal direction angle to the aforementioned target horizontal direction angle, and from the aforementioned initial pitch angle to the aforementioned target pitch angle to acquire video images, thereby obtaining video image acquisition results. This achieves the purpose of accurately tracking the position and capturing video of the target speaker by combining a microphone array and a main microphone, thereby improving the technical effect of video tracking and audio-visual smoothness. It also solves the technical problem that the tracking camera method in related technologies cannot switch camera tracking according to the speaker, resulting in poor tracking and shooting effects and easy audio-visual loss. Attached Figure Description
[0010] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this application, illustrate exemplary embodiments of the invention and, together with their description, serve to explain the invention and do not constitute an undue limitation thereof. In the drawings:
[0011] Figure 1 This is a flowchart of a voice tracking camera method according to an embodiment of the present invention;
[0012] Figure 2 This is a schematic diagram of the structure of a voice tracking camera system according to an embodiment of the present invention;
[0013] Figure 3 This is a schematic diagram of the structure of an optional voice tracking camera system according to an embodiment of the present invention;
[0014] Figure 4 This is a flowchart of an optional voice tracking camera method according to an embodiment of the present invention;
[0015] Figure 5a This is an optional microphone array horizontal beam pattern according to an embodiment of the present invention;
[0016] Figure 5b This is an optional microphone array vertical beam pattern according to an embodiment of the present invention;
[0017] Figure 6 This is a simplified diagram of a single microphone in an optional microphone array according to an embodiment of the present invention;
[0018] Figure 7 This is a simplified directional diagram of an optional single microphone according to an embodiment of the present invention;
[0019] Figure 8 This is a schematic diagram of an optional pickup area covered by a single microphone according to an embodiment of the present invention;
[0020] Figure 9 This is a schematic diagram of the boundary line of an optional pickup area according to an embodiment of the present invention;
[0021] Figure 10 This is a three-dimensional schematic diagram of the structure of an optional voice tracking camera system according to an embodiment of the present invention;
[0022] Figure 11 This is a schematic diagram of the optional pickup area path length dimension according to an embodiment of the present invention;
[0023] Figure 12 This is a schematic diagram of the response curve of an optional human voice filter according to an embodiment of the present invention. Detailed Implementation
[0024] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0025] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0026] According to an embodiment of the present invention, a method embodiment for voice tracking camera is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0027] Figure 1 This is a flowchart of a voice tracking camera method according to an embodiment of the present invention, such as... Figure 1 As shown, the method includes the following steps:
[0028] Step S102: Based on the first sound pickup amplitude collected by the multiple first microphones included in the microphone array for the target sound-emitting object, determine the initial horizontal direction angle of the camera device rotating in the horizontal direction and the initial pitch angle of the camera device rotating in the vertical direction.
[0029] Optionally, the microphone array described above is a circular array composed of multiple first microphones, all of which are narrow-beam directional microphones of equal specifications. It is used to obtain the initial horizontal and initial pitch angles corresponding to the camera device based on the acquired sound pickup amplitude, and then to coarsely adjust the position of the camera device based on the obtained initial horizontal and initial pitch angles. The camera device described above may be, but is not limited to, a webcam.
[0030] In an optional embodiment, before determining the initial horizontal angle of the camera device rotating in the horizontal direction and the initial pitch angle rotating in the vertical direction based on the first pickup amplitudes collected by the plurality of first microphones included in the microphone array for the target sound-emitting object, the method further includes: acquiring first speech information collected by the plurality of first microphones for the target sound-emitting object; performing a first filtering process on the first speech information corresponding to the plurality of first microphones using a plurality of first voice filters to obtain a third pickup amplitude corresponding to the plurality of first microphones, wherein the plurality of first voice filters correspond one-to-one with the plurality of first microphones; performing a first noise reduction and amplification process on the third pickup amplitudes corresponding to the plurality of first microphones using a plurality of first low-noise amplifiers to obtain a fourth pickup amplitude corresponding to the plurality of first microphones, wherein the plurality of first low-noise amplifiers correspond one-to-one with the plurality of first microphones; and performing a first compensation process on the fourth pickup amplitudes corresponding to the plurality of first microphones using a pre-constructed path compensation matrix to obtain a first pickup amplitude corresponding to the plurality of first microphones.
[0031] Optionally, the aforementioned path compensation matrix is used to perform path compensation on the pickup information corresponding to the multiple first microphones. Through this method, the position and direction of the camera device are coarsely adjusted using the pickup amplitudes (i.e., the first pickup amplitudes) obtained by the multiple first microphones in the microphone array. Specifically: when someone speaks at the start of the meeting, the first voice information collected by the multiple first microphones in the microphone array targeting the target speaker is first processed by a corresponding first voice filter to obtain the third pickup amplitude corresponding to each of the multiple first microphones; further, a corresponding first low-noise amplifier is used to perform first noise reduction and amplification processing on the third pickup amplitudes corresponding to the multiple first microphones to obtain the fourth pickup amplitude corresponding to each of the multiple first microphones; the path compensation matrix is used to compensate for the fourth pickup amplitude corresponding to each of the multiple first microphones, resulting in the first pickup amplitude corresponding to each of the multiple first microphones. Based on the first pickup amplitude, coarse tracking of the camera device begins.
[0032] In an optional embodiment, the determination of the initial horizontal angle of the camera device rotating in the horizontal direction and the initial pitch angle rotating in the vertical direction based on the first pickup amplitude collected by the multiple first microphones in the microphone array for the target sound-emitting object includes: obtaining the first pickup area corresponding to the largest pickup amplitude among the first pickup amplitudes of the multiple first microphones; determining the initial horizontal angle according to the center line direction of the first pickup area; and using the pre-obtained initial pitch angle corresponding to the first pickup area as the initial pitch angle.
[0033] Optionally, the initial directional angle of the microphone array is the position of the center line (i.e., angle bisector) of each pickup area. α is taken as the horizontal half-beamwidth corresponding to each of the multiple first microphones. It is required that an integer multiple of α is 360° (or an integer multiple of α [0.9, 1.1] is 360°). The multiple is N. Then the horizontal direction of the target area (such as a conference room) is divided into N pickup areas Rn = [r1, r2, r3...rn]. The angle of the pickup area corresponding to each microphone is α (or β), and the angle bisector of each pickup area is the center line.
[0034] Using the above method, based on the principle of maximum pickup amplitude, the center line direction and initial pitch angle of the first pickup area corresponding to the largest pickup amplitude among the multiple first microphones are respectively used as the initial horizontal direction angle and initial pitch angle of the camera device (or the main microphone, the initial horizontal direction angle and initial pitch angle of the camera device and the main microphone are the same), thereby achieving coarse positioning of the target sound-emitting object.
[0035] Optionally, the main microphone mentioned above is a narrow-beam directional microphone with the same specifications as the microphone array elements.
[0036] In an optional embodiment, before using the pre-acquired initial pitch angle corresponding to the first pickup area as the initial pitch angle, the method further includes: acquiring multiple pickup areas corresponding to the multiple first microphones, and multiple pitch angles corresponding to the multiple pickup areas respectively, wherein there is a one-to-one correspondence between the multiple first microphones and the multiple pickup areas; acquiring example speech emitted by the main microphone to an example speaker within the multiple pickup areas, and collecting second speech information corresponding to the multiple pitch angles corresponding to the multiple pickup areas respectively; and using a second voice filter to perform a second filtering process on the second speech information corresponding to the multiple pitch angles corresponding to the multiple pickup areas respectively. The fifth pickup amplitude corresponding to each of the multiple pickup areas and each of the multiple pitch angles is obtained; a second low-noise amplifier is used to perform a third noise reduction and amplification process on the fifth pickup amplitude corresponding to each of the multiple pickup areas and each of the multiple pitch angles, to obtain a sixth pickup amplitude corresponding to each of the multiple pickup areas and each of the multiple pitch angles; based on the sixth pickup amplitude corresponding to each of the multiple pickup areas and each of the multiple pitch angles, the maximum pickup amplitude corresponding to each of the multiple pickup areas is determined; based on the maximum pickup amplitude corresponding to each of the multiple pickup areas, the initial pitch angle corresponding to each of the multiple pickup areas is determined; and based on the maximum pickup amplitude corresponding to each of the multiple pickup areas, the path compensation matrix is determined.
[0037] It should be noted that for the calibration of any one of the multiple pickup areas, considering the two uncontrollable factors—the non-linear change in the distance between the target speaker and the microphone, and the non-linear change in the amplitude and angle of the microphone within the half-beamwidth—an extreme value approach is adopted to simplify the design. Specifically, the calibration height is defined as the approximate height of a person's mouth on the conference chair, simulating the actual sound-emitting position of the target speaker during the meeting. The calibrator (a device capable of continuously emitting human voice) is used as the example speaker. The calibration user is required to hold the calibrator (a device capable of continuously emitting human voice) at the calibration height and slowly walk around the edge of the conference table as the example speech. In the slow-moving scenario, the error caused by the Doppler effect is negligible. Since the microphone has a narrow beam and strong attenuation outside the beam, each microphone will pick up sound individually during the calibration user's walk around the conference table, and the human voice collected by other microphones is negligible. Using the above method, each pickup area places the example speech (a short passage requiring stable human voice amplitude and fast speech rate) at a standard height and begins to emit sound. After recognizing the sound source in a certain pickup area, the main microphone aligns with the center line of that pickup area. By changing the pitch angle of the main microphone, the second speech information corresponding to multiple pitch angles of the main microphone in that pickup area is obtained. The second speech information corresponding to multiple pitch angles in that pickup area is then filtered by the first human voice filter and amplified by the first low-noise amplifier to obtain the sixth pickup amplitude. The main microphone finds the maximum pickup amplitude within the half-beamwidth corresponding to multiple pitch angles in that pickup area. The pitch angle corresponding to the maximum pickup amplitude is βn, which is used as the initial value of the pitch angle of the corresponding microphone. The initial values of the pitch angles of each microphone in the microphone array are obtained, forming a pitch angle processing matrix [b1,b2,b3…bn], in degrees.
[0038] In one optional embodiment, determining the path compensation matrix based on the maximum pickup amplitude corresponding to the plurality of pickup areas includes: determining the median pickup amplitude corresponding to the maximum pickup amplitude corresponding to the plurality of pickup areas; and determining the path compensation matrix based on the maximum pickup amplitude corresponding to the plurality of pickup areas and the median pickup amplitude.
[0039] Using the above method, the pickup result matrix An = [a1, a2, a3…an] (in dBm) is obtained based on the maximum pickup amplitude corresponding to each of the multiple pickup areas. The median pickup amplitude ax corresponding to the pickup result matrix is then calculated. The path compensation matrix Xn = [x1, x2, x3…xn] (in dB) is obtained by subtracting the median pickup amplitude ax from the maximum pickup amplitude an corresponding to each of the multiple pickup areas.
[0040] In an optional embodiment, after determining the initial pitch angle corresponding to each of the plurality of pickup areas based on the maximum pickup amplitude corresponding to each of the plurality of pickup areas, the method further includes: controlling multiple micro motors to adjust the plurality of first microphones to the corresponding initial pitch angle positions.
[0041] Using the above method, the pitch angle corresponding to the largest pickup amplitude is bn, which is taken as the initial pitch angle value of the corresponding microphone. The initial pitch angle value of each microphone in the microphone array is obtained, forming a pitch angle initial value matrix [b1,b2,b3…bn], in degrees. Then, according to this pitch angle initial value matrix, the multiple first microphones in the microphone array are set to correspond to micro motors respectively. After setting, the pitch angle corresponding to each microphone in each area of the microphone array is consistent with the pitch angle initial value matrix value.
[0042] Step S104: After controlling the camera device to rotate to the first pickup area corresponding to the initial horizontal angle and the initial pitch angle, the second pickup amplitude of the main microphone is acquired from the target sounding object in multiple different pickup directions. The main microphone is a microphone independent of the microphone array.
[0043] Optionally, the horizontal angle of the camera device is adjusted by the azimuth motor, and the pitch angle of the camera device is adjusted by the pitch motor. That is, the azimuth motor is adjusted to adjust the axis of the camera device to the center line position of the first sound pickup area corresponding to the maximum sound pickup amplitude, and the pitch motor is adjusted to adjust the pitch angle to the initial value of the pitch angle of the first sound pickup area.
[0044] In one optional embodiment, the acquisition of the second pickup amplitude from the main microphone in multiple different pickup directions relative to the target sound source includes: acquiring third speech information from the main microphone in multiple different pickup directions relative to the target sound source; using a second voice filter to perform a third filtering process on the third speech information corresponding to the multiple different pickup directions to obtain a seventh pickup amplitude corresponding to the multiple different pickup directions; and using a second low-noise amplifier to perform a third noise reduction and amplification process on the seventh pickup amplitude corresponding to the multiple different pickup directions to obtain the second pickup amplitude corresponding to the multiple different pickup directions. Through this method, the third speech information acquired by the main microphone of the camera device in multiple different pickup directions relative to the target sound source is sequentially filtered by the second voice filter and further amplified by the second low-noise amplifier to obtain the second pickup amplitude corresponding to the multiple different pickup directions. The filtered and amplified pickup amplitude has better sound clarity and intelligibility.
[0045] Optional features include first and second voice filters, first and second low-noise amplifiers, etc.
[0046] Step S106: Based on the second pickup amplitude collected by the main microphone in the multiple different pickup directions, determine the target horizontal angle of the camera device rotating in the horizontal direction and the target pitch angle of the camera device rotating in the vertical direction.
[0047] By using the above methods and based on the principle of maximizing the pickup amplitude, the second pickup amplitude collected by the main microphone from multiple different pickup directions for the target sound source is used as the basis for fine-tuning the horizontal and vertical angles of the camera device, thereby achieving precise positioning of the target sound source.
[0048] In an optional embodiment, when the plurality of different pickup directions include a first number of horizontal angles and a second number of pitch angles, determining the target horizontal angle of the camera device rotating along the horizontal direction and the target pitch angle of the camera device rotating along the vertical direction based on the second pickup amplitudes collected by the main microphone in the plurality of different pickup directions includes: taking the horizontal angle corresponding to the largest pickup amplitude among the second pickup amplitudes corresponding to the first number of horizontal angles as the target horizontal angle; and taking the pitch angle corresponding to the largest pickup amplitude among the second pickup amplitudes corresponding to the second number of pitch angles as the target pitch angle.
[0049] Optionally, the pickup amplitude of the main microphone after filtering and amplification is divided into K parts in the pitch angle range [0,90] as multiple different pickup directions. Then, the pitch angle of the camera device is adjusted by traversal to obtain the pitch angle βm corresponding to the maximum pickup amplitude as the target pitch angle. Similarly, the horizontal direction angle is adjusted by traversal to obtain the horizontal direction angle αm corresponding to the maximum pickup amplitude. After fine-tuning, it is assumed that the camera device is now aimed at the target sound source.
[0050] Step S108: Control the camera device to rotate from the initial horizontal angle to the target horizontal angle, and from the initial pitch angle to the target pitch angle to acquire video images and obtain video image acquisition results.
[0051] By finely adjusting the camera device from the initial horizontal and initial pitch angles to the positions corresponding to the target horizontal and pitch angles, the camera device is considered to be aligned with the target sound-emitting object, and video image acquisition begins, thus achieving precise positioning of the target sound-emitting object.
[0052] In an optional embodiment, after controlling the camera device to rotate from the initial horizontal angle to the target horizontal angle and from the initial pitch angle to the target pitch angle to acquire video images and obtain video image acquisition results, the method further includes: increasing the maximum pickup amplitude among the first pickup amplitudes corresponding to the plurality of first microphones by a preset emphasis coefficient; and decreasing the other pickup amplitudes (excluding the maximum pickup amplitude) among the first pickup amplitudes corresponding to the plurality of first microphones by a preset attenuation coefficient to obtain audio processing results; and obtaining video synthesis results based on the video image acquisition results and the audio processing results.
[0053] Optionally, during the final video recording process (i.e., during video synthesis), the method for synthesizing the audio processing result through the microphone array is as follows: If the camera device is coarsely adjusted, the first pickup area (i.e., the first pickup area corresponding to the largest pickup amplitude) corresponding to the initial horizontal angle and the initial pitch angle is rk. The pickup amplitude corresponding to the first pickup area is added with a preset emphasis coefficient WdB. Then, the pickup amplitudes of the other pickup areas r1, ...,rn in multiple pickup areas other than the first pickup area are subtracted from the preset attenuation coefficient ZdB and then superimposed to obtain the final audio track as the final audio processing result.
[0054] It should be noted that in practical applications, the pickup area of the main microphone continuously changes with the corresponding horizontal and pitch angles, inevitably leading to inaccuracies or noise in the acquired pickup results. However, the sum of the pickup areas of the microphone array always covers all speakers, resulting in relatively stable pickup results. Therefore, this embodiment of the invention synthesizes audio tracks using only the pickup results from multiple first microphones in the microphone array. The method for restoring the audio track processes the original microphone array pickup results, preserving or amplifying the pickup results from the pickup area where the target speaker is located, while attenuating other areas before superimposing them together to form the restored audio track. This preserves all the sound in the meeting room during the coarse adjustment stage, without losing the content of the speech, thereby accurately eliminating noise and preventing the loss of the speaker's voice during adjustment, thus improving the accuracy of audio track synthesis.
[0055] It is understood that the execution entity of the above steps S102 to S108 is a digital processing device. Through the above steps S102 to S108, coarse positioning is obtained by microphone array pickup statistics, and precise positioning is obtained by fine adjustment of the main microphone. The obtained pickup amplitude is used to determine the target sound-emitting object to be tracked, and to prompt the camera device to align with the target sound-emitting object. It can achieve the purpose of voice and video tracking by combining microphone array and main microphone, accurately tracking the position of the target sound-emitting object and capturing video, thereby improving the technical effect of video tracking and audio-visual smoothness. This solves the technical problem that the tracking camera method in related technologies cannot switch camera tracking according to the speaker, resulting in poor tracking and shooting effect and easy audio and video loss.
[0056] According to an embodiment of the present invention, a system embodiment for implementing the above-described voice tracking camera method is also provided. Figure 2 This is a schematic diagram of the structure of a voice tracking camera system according to an embodiment of the present invention, as shown below. Figure 2 As shown, the aforementioned voice tracking camera system includes: a microphone array 200; a multi-channel micro-motor 201 for controlling multiple first microphones included in the microphone array; a camera device 202; a main microphone 203 parallel to the central axis of the camera device; an azimuth motor 204 for controlling the horizontal rotation of the camera device; a pitch motor 205 for controlling the vertical rotation of the camera device; and a digital processing device 206. The number of channels of the multi-channel micro-motor corresponds to the number of the multiple first microphones, and is used to control... The aforementioned plurality of first microphones correspond to different pitch angles. The microphone array 200 is connected to the multi-channel micro-motor 201 and the digital processing device 206; the multi-channel micro-motor 201 is connected to the digital processing device 206; the main microphone 203 is connected to the camera device 202 and the digital processing device 206; the camera device 202 is connected to the azimuth motor 204 and the pitch motor 205; the azimuth motor 204 and the pitch motor 205 are respectively connected to the digital processing device.
[0057] The aforementioned plurality of first microphones are used to acquire a first pickup amplitude emitted by the target sound source, and send the first pickup amplitude acquired for the target sound source to the aforementioned digital processing device; the aforementioned digital processing device is used to determine, based on the first pickup amplitude acquired by the plurality of first microphones for the target sound source, an initial horizontal angle for the camera device to rotate in the horizontal direction and an initial pitch angle for the camera device to rotate in the vertical direction; control the camera device to rotate to the first pickup area corresponding to the initial horizontal angle and the initial pitch angle; the aforementioned main microphone is used to target the target sound source from multiple different pickup directions. The second pickup amplitude is collected separately, and the second pickup amplitudes collected in the multiple different pickup directions are sent to the digital processing device; the digital processing device is used to determine the target horizontal angle of the camera device rotating in the horizontal direction and the target pitch angle of the camera device rotating in the vertical direction based on the second pickup amplitudes collected by the main microphone in the multiple different pickup directions, and to control the camera device to rotate from the initial horizontal angle to the target horizontal angle and from the initial pitch angle to the target pitch angle; the camera device performs video image acquisition to obtain video image acquisition results.
[0058] With the above system setup, voice and video tracking can be achieved by combining a microphone array and a main microphone. This allows for accurate location tracking and video capture of the target speaker, thereby improving video tracking performance and audio-visual smoothness. It also solves the technical problem of poor tracking and capture results and audio-visual loss caused by the inability of tracking camera methods in related technologies to switch camera tracking based on the speaker.
[0059] Optionally, the aforementioned digital processing device can be a standalone device or a sub-device of other devices built into the voice tracking camera system, such as being built into a microphone array, a camera device, a main microphone, or a multi-channel micro motor, etc.
[0060] Optional features include a microphone array for defining the pickup area and marking the azimuth angle of the center line of each pickup area; the microphone array includes multiple first microphones. A voice filter, connected to the microphone array and main microphone, filters out frequencies other than human voice, eliminating interference from background noise on the coarse / fine adjustment criteria. A low-noise amplifier module, connected to the voice filter, increases the pickup amplitude with high gain, while improving the signal-to-noise ratio. A main camera and a main microphone parallel to its axis acquire video, serving as the criterion for fine-tuning the camera's position. Pitch and azimuth motors adjust the camera's pitch and horizontal angles. A digital processing unit acquires the filtered and amplified pickup amplitude signal in the digital domain to obtain the corresponding human voice amplitude value, calculates, analyzes, and compares the human voice amplitude values, controls the motor rotation, and synthesizes the video track provided by the camera and the audio track obtained by the microphone array into a single video stream for output.
[0061] Optionally, the above-mentioned voice tracking camera system includes three steps: camera device calibration, coarse positioning, and fine positioning. The calibration process determines the initial pitch angle of the microphone array, uses the microphone array to achieve coarse positioning of the camera device's initial horizontal and pitch angles, and uses a main microphone parallel to the camera device to acquire the fine horizontal and pitch angles. The tracking and recognition approach for the target voice source involves obtaining coarse positioning through microphone array pickup statistics, fine-tuning the main microphone to achieve precise positioning, and using the acquired pickup amplitude to determine the target voice source and prompt the camera device to align with it.
[0062] Optionally, the multiple first microphones in the microphone array and the main microphone are all directional microphones of the same specification, with narrow beam characteristics in both azimuth and elevation angles. The multiple first microphones in the microphone array are used with multi-channel micro-motors, each motor channel controlling the elevation angle of one of the multiple first microphones. The camera device and the main microphone are axially parallel, mounted on the same tray, and controlled by the same azimuth and elevation motors, respectively controlling the azimuth and elevation angles. The camera device can be, but is not limited to, a webcam. A digital processing device is used for processing the acquisition, analysis, and comparison of human voice signals, controlling the micro-motors and the elevation and azimuth motors, and finally synthesizing the video stream. The final video stream is generated in two parts: the camera device only generates the video track (i.e., the video image acquisition result), and the audio track is weighted and superimposed from the sound pickup results of the multiple first microphones in the microphone array to form the audio track (i.e., the audio processing result). The video track and the audio track are combined to form the video stream, thus obtaining the synthesized video result.
[0063] In an alternative embodiment, it remains as follows Figure 2As shown, the system further includes: a first voice filter 211 corresponding to each of the plurality of first microphones, and a first low-noise amplifier 212 corresponding to each of the plurality of first microphones, wherein the plurality of first microphones are connected to the digital processing device 206 in sequence through the corresponding first voice filter 211 and the first low-noise amplifier 212; a second voice filter 213 corresponding to the main microphone 203, and a second low-noise amplifier 214 corresponding to the main microphone 203, wherein the main microphone 203 is connected to the digital processing device 206 in sequence through the second voice filter 213 and the second low-noise amplifier 214.
[0064] It is understandable that, in order to improve the signal-to-noise ratio (SNR) of the channel, a voice filter and a low-noise amplifier are used after the microphone converts the signal into an electrical signal to improve the SNR and avoid noise interference. The specifications of the aforementioned first voice filter 211 and second voice filter 213, and the specifications of the aforementioned first low-noise amplifier 212 and second low-noise amplifier 214 are as follows.
[0065] Based on the above embodiments and optional embodiments, the present invention proposes an optional implementation method. Figure 3 This is a schematic diagram of the structure of an optional voice tracking camera system according to an embodiment of the present invention. Figure 4 This is a flowchart of an optional voice tracking camera method according to an embodiment of the present invention, such as... Figure 3 and Figure 4 As shown, the voice tracking camera system includes: a circular array composed of several narrow-beam directional microphones of equal specifications; a multi-channel micro-motor controlling the pitch angle of the microphone array; a camera device; a main microphone parallel to the central axis of the camera device; an azimuth motor controlling the horizontal rotation of the camera device; a pitch motor controlling the vertical rotation of the camera device; a low-noise amplifier for human voice (i.e., a first low-noise amplifier and a second low-noise amplifier for human voice); human voice filters (i.e., multiple first human voice filters and second human voice filters); and digital processing equipment. The specific method steps corresponding to this system are as follows:
[0066] Step S11: Use multiple narrow-beam directional microphones of the same specification to form a circular array, and divide the horizontal plane of the conference room into multiple sound pickup areas according to the number of microphones.
[0067] Step S12: Perform path loss calibration and initial pitch angle calibration on-site to obtain the path loss from the sound source to the microphone array for each pickup area, as well as the initial pitch angle of the main microphone of the camera device. Set the pitch angle of each microphone in the microphone array according to the initial pitch angle value using the corresponding micro motor.
[0068] Step S13, the real-time adjustment during use after calibration is divided into two steps: coarse adjustment and fine adjustment. The specific steps are as follows:
[0069] Step S131: Begin coarse adjustment and positioning of the camera device. First, compensate for the loss of the pickup amplitude corresponding to each of the multiple first microphones in the microphone array. Then, compare the compensated pickup amplitudes (i.e., the first pickup amplitudes) corresponding to the multiple first microphones in the microphone array and take the maximum value. The initial horizontal direction angle and initial pitch angle corresponding to the center line position of the area where the array is located are used as the coarse adjustment result of the camera device. Use a stepper motor (i.e., a multi-channel micro motor) to control the camera device to rotate horizontally to the center line of the area obtained by coarse adjustment, and vertically rotate to the initial pitch angle value of the pickup area corresponding to the maximum pickup amplitude.
[0070] Step S132: Fine-tuning and positioning of the camera device begins. Below the camera device is a directional microphone (i.e., the main microphone) parallel to the lens axis. The main microphone is used to perform local fine-tuning based on the sound pickup results (i.e., the second sound pickup amplitude) corresponding to multiple different sound pickup directions. The horizontal and pitch angles corresponding to the main microphone are further adjusted until the maximum sound pickup amplitude is obtained. When the main microphone is adjusted to the target horizontal and pitch angles corresponding to the maximum sound pickup amplitude, the adjustment is considered complete, and the target voice in the meeting is accurately tracked.
[0071] In step S14, the final recorded video consists of two parts: the video track represents the video image capture result obtained by the camera device, and the audio track represents the audio processing result synthesized by the microphone array. Specifically, the final video stream is generated in two parts: the camera device generates only the video track (i.e., the video image capture result), and the audio track (i.e., the audio processing result) is generated by weighted and superimposed from the sound pickup results of multiple first microphones in the microphone array. The video track and audio track are then combined to form the video stream, resulting in the final video synthesis result.
[0072] The embodiments of the present invention can achieve at least the following technical effects: (1) The microphone array is used to roughly define and analyze the position of the target sound-emitting object, and then the extreme value method is used to accurately locate it. The location is unique, fast and effective. (2) The original microphone array pickup results are processed in the method of restoring the audio track. The pickup results of the pickup area where the target sound-emitting object is located are retained, and the pickup results of other areas are attenuated and then superimposed to form the restored audio track. In this way, the main sound can be preserved in the coarse adjustment stage without losing the content of the meeting. The fine adjustment of the position is only to facilitate the camera device to better track the target sound-emitting object.
[0073] Based on the above embodiments and optional embodiments, the present invention proposes another optional implementation of the voice tracking camera system and corresponding method. Figure 5a and Figure 5bThese are, respectively, an optional horizontal beam pattern and a vertical beam pattern of a microphone array according to an embodiment of the present invention; Figure 6 This is a simplified diagram of a single microphone in an optional microphone array according to an embodiment of the present invention; Figure 7 This is a simplified directional diagram of an optional single microphone according to an embodiment of the present invention; Figure 8 This is a schematic diagram of an optional pickup area covered by a single microphone according to an embodiment of the present invention; Figure 9 This is a schematic diagram of the boundary line of an optional pickup area according to an embodiment of the present invention; Figure 10 This is a three-dimensional schematic diagram of the structure of an optional voice tracking camera system according to an embodiment of the present invention; Figure 11 This is a schematic diagram of the optional pickup area path length dimension according to an embodiment of the present invention; Figure 12 This is a schematic diagram of an optional human voice filter response curve according to an embodiment of the present invention. Taking a rectangular conference room as an example, 5 meters long and 3 meters wide, the camera device is placed in the center. A microphone array composed of multiple narrow-beam directional microphones with a horizontal beamwidth of 30° and a vertical beamwidth of 20° is used. The radiation pattern is shown in Figure 5. A simplified diagram of a single microphone in the microphone array is shown in Figure 5. Figure 6 The corresponding simplified diagram is as follows: Figure 7 The pickup area covered by a single microphone is as follows: Figure 8 Since the beamwidth is 30°, the conference table can be divided into 360 / 30 = 12 pickup zones, such as... Figure 9 The long dashed line is the dividing line of the sound pickup area. The microphone beam is narrow and the half-amplitude edge attenuation is very high. It can be assumed that each microphone in the microphone array is responsible for a 30° sound pickup area.
[0074] For ease of understanding Figure 10 A three-dimensional schematic diagram of the voice tracking camera system is shown. For simplicity, only four arrays are depicted. Figure 10In the diagram, 207 is the product base, 200 is the microphone array, 201 is the multi-channel micro motor, 208 is the array tray providing a fulcrum for the microphone array 200 and the multi-channel micro motor 201, 203 is the main microphone, 202 is the camera device, 209 is the housing housing the main microphone 203 and the camera device 202, and after fixing, the axes of the main microphone 203 and the camera device 202 are aligned, 205 is the pitch motor behind the housing 209, and 204 is the azimuth motor behind the housing 209. After dividing the sound pickup area, user A can hold a repeater as an example voice and continuously play "ah~" (the voice should be continuous and without fluctuation) as the example speech. The calibration height of the repeater is defined as the approximate height of a person's mouth on the conference chair. This simulates the actual voice position of the target voice during the meeting. The user A walks slowly around the edge of the conference room (areas 1-12) at a speed of 0.5 m / s. The perimeter of the entire conference room is 16 m, and this takes 32 seconds.
[0075] As user A walks, the microphone array collects the received sound information in real time. After passing through a voice filter, the sound enters the digital signal processing equipment, which converts it into sound pickup amplitude. The digital signal processing equipment collects the sound pickup amplitude every 0.2 seconds, that is, one sampling point is collected every 0.1m. Each pickup area corresponds to a different distance traveled by user A. The four pickup areas with the longest distances (2, 5, 8, and 11) are 1.69m long, and 16 sampling points can be collected for each of them. The four pickup areas with the shortest distances (3, 4, 9, and 10) are 0.87m long (e.g., ...). Figure 11 As shown in the image, it can collect up to 8 sampling points. After the microphone collects the analog signal and converts it into an electrical signal, it is filtered using a human voice filter. The pickup amplitude-frequency curve of the filter used is shown in the image. Figure 12 As shown, it exhibits good filtering performance within the pickup amplitude (DC-3kHz, DC-3kHz), with significant noise reduction outside the 3kHz band. The purpose of the first round of acquisition is to identify the area where the speaker is located, facilitating the rotation of the main microphone to the corresponding area for pitch angle calibration.
[0076] The initial value of the pitch angle is recorded simultaneously with the first round of acquisition. When a microphone in a certain pickup area acquires the first point and the difference between the first point and the previous acquisition time point is 40dB, it is considered that a sound source has appeared in the current pickup area. Conversely, if the next point is 40dB lower than the previous point, it is considered that the sound source has left the current pickup area. Before starting the second sampling point, the motor is adjusted so that the camera device is aligned with the center line of the newly reached sound pickup area. At this time, the main microphone, which is parallel to the axis of the camera device, is also aligned with the center line of the sound pickup area. The [0, 90°] is divided into the minimum number of sampling points minus 2. For example, in this embodiment of the invention, the number of sampling points in sound pickup area 3 is the minimum, which is 8 sampling points. Therefore, only 6 areas need to be taken. 90° / 6 = 15°. In the 8 sampling points of sound pickup area 3, the motor adjusts the pitch angle in sequence as [0°, 0°, 15°, 30°, 45°, 60°, 75°, 90°]. In sound pickup area 2, the motor adjusts the pitch angle in sequence as [0, 0, 15°, 30°, 45°, 60°, 75°, 90 ... Taking pickup area 3 as an example, after the main microphone samples at 8 time points in pickup area 3 are filtered by the human voice filter, the amplitudes obtained are [-8, -8.2, -7, -6, -8.5, -9, -20, -40], in dBm (decibels milliwatts). Taking the angle corresponding to the maximum point as 30°, the initial value of the pitch angle of this area is 30°. Similarly, the initial values of the pitch angles of the 12 areas are [0, 15°, 30°, 30°, 15°, 0°, 0°, 15°, 30°, 30°, 15°, 0°]. The maximum pickup amplitudes corresponding to the 12 pickup areas are [-13, -10.9, -8, -7.6, -10.5, -13.2, -12.8, -10.8, -8.2, -8.3, -11, -13.1], in dBm. The median is -10.85, so the path compensation matrix Xn = [-2.15, -0.05, 2.85, 3.25, 0.35, -2.35, -1.95, 0.05, 2.65, 2.55, -0.15, -2.25]. The calibration is now complete, and the path compensation matrix and initial pitch angle values for each pickup area are obtained. When it is determined that the sound source has left a certain pickup area, the horizontal and pitch angles of the camera device are then coarsely and finely adjusted for positioning in actual use. The specific steps are as follows:
[0077] Step S21, the camera device coarse adjustment positioning process, sets the amplitude difference for determining whether there is sound to Y = 30dB. At the beginning of the meeting, the main speaker is usually a single person. After the array 3 receives the sound pickup amplitude and it is higher than the sound pickup amplitude Y = 30dB corresponding to other sound pickup areas (already superimposed with the calibration matrix), the azimuth motor of the camera device is controlled to rotate to the center line (15°) of the first sound pickup area corresponding to the largest sound pickup amplitude in the horizontal direction of the camera device axis, which is the initial value of the horizontal direction angle. Then the pitch motor rotates to the initial value of the pitch angle of the first sound pickup area (30°).
[0078] Step S22: Fine-tuning of the camera device's positioning process. This fine-tuning is achieved by traversing the horizontal and vertical angles of the main microphone within the first pickup area. At this stage, the digital signal processing equipment is required to acquire the pickup amplitude of the main microphone. The frequency of adjustment of the azimuth and pitch motors, which separately adjust the horizontal and vertical angles of the main microphone, is increased from 0.2 seconds / time during calibration to 0.01 seconds / time for rapid fine-tuning. For the horizontal angle, since the initial beam angle is relatively large, 5° is taken as the minimum accuracy. This process is repeated throughout the entire... Only 7 sampling points are needed for the horizontal azimuth angle range of [0°, 30°]. Since the coarse adjustment has reached the center line of the pickup area, it is necessary to fine-tune it 6 more times, traversing the horizontal azimuth angles [0, 5, 10, 15, 20, 25, 30] (unit: °). After 0.06s, the corresponding pickup amplitudes are [-15.1, -15.0, -14.9, -14.5, -14.1, -14.6, -15]. The fine-tuned azimuth angle is 20°. Then, control the azimuth motor to adjust the azimuth angle to 20°.
[0079] For the pitch angle, since the initial beam angle is small, 3° is taken as the minimum precision. Traversing the entire pitch angle range of (0.5, 1.5) times [15°, 45°] requires 11 sampling points. Since the coarse adjustment has reached the 30° line, it needs to be fine-tuned 10 more times. Traversing the pitch angle [15, 18, 21, 24, 27, 30, 33, 36, 39, 42, 45] (unit: °) for 0.1s, the pickup amplitude corresponding to different pitch angles is obtained as follows: [-15.5, -15.1, -14.5, -14.1, -13.6, -14.1, -14.6, -15.3, -15.9, -16.5, -16.8]. The fine-tuned pitch angle is 27°. Then, the pitch angle motor is controlled to adjust to the 27° pitch angle position. Once the fine-tuning is complete, the storage time for the fine-tuning process should not exceed 0.5 seconds.
[0080] In step S23, if a change in the target sound source is detected, return to step S21, determine the target based on the amplitude difference Y = 30 between the pickup areas, and then execute step S22 for fine-tuning.
[0081] Step S24: Recording results. The video captured by the camera device is used as the video track of the final video stream. The audio track synthesis rule is as follows: for the coarsely located sound pickup area, the amplitude is increased by 2dB, and the sound pickup amplitude of other sound pickup areas is halved. Finally, they are superimposed and synthesized into an audio track. The video track and audio track are synthesized into a complete video stream output, which is the video synthesis result. The specific synthesis function is implemented in the digital processing device.
[0082] It should be noted that in this application Figures 2 to 3 The specific structure of the voice tracking camera system shown is merely illustrative. In practical applications, the voice tracking camera system in this application can be more advanced than... Figures 2 to 3 The voice tracking camera system shown has more or less structure.
[0083] It should be noted that any optional or preferred voice tracking camera method in the above method embodiments can be executed or implemented in the voice tracking camera system provided in this embodiment.
[0084] Furthermore, it should be noted that the optional or preferred implementation methods of this embodiment can be found in the relevant descriptions in the method embodiments, which will not be repeated here.
[0085] According to an embodiment of this application, an embodiment of a non-volatile storage medium is also provided. Optionally, in this embodiment, the non-volatile storage medium includes a stored program, wherein, when the program is running, it controls the device where the non-volatile storage medium is located to execute any of the above-mentioned voice tracking and video recording methods.
[0086] Optionally, in this embodiment, the non-volatile storage medium may be located in any computer terminal in a group of computer terminals in a computer network, or in any mobile terminal in a group of mobile terminals, and the non-volatile storage medium includes stored programs.
[0087] Optionally, during program execution, the device containing the non-volatile storage medium can be controlled to execute any of the above-mentioned voice tracking and video recording methods.
[0088] According to an embodiment of this application, an embodiment of a processor is also provided. Optionally, in this embodiment, the processor is used to run a program, wherein the program executes any of the above-described voice tracking camera methods.
[0089] According to an embodiment of this application, an embodiment of a computer program product is also provided, which, when executed on a data processing device, is adapted to execute a program that initializes the voice tracking camera method steps described above.
[0090] Optionally, when the above-mentioned computer program product is executed on a data processing device, it is suitable to execute an initialization program having the following method steps: based on the first pickup amplitude collected by the multiple first microphones included in the microphone array for the target sound-emitting object, determine the initial horizontal angle of the camera device rotating in the horizontal direction and the initial pitch angle rotating in the vertical direction; after controlling the camera device to rotate to the first pickup area corresponding to the initial horizontal angle and the initial pitch angle, acquire the second pickup amplitude collected by the main microphone for the target sound-emitting object in multiple different pickup directions; based on the second pickup amplitude collected by the main microphone in the multiple different pickup directions, determine the target horizontal angle of the camera device rotating in the horizontal direction and the target pitch angle rotating in the vertical direction; control the camera device to rotate from the initial horizontal angle to the target horizontal angle and from the initial pitch angle to the target pitch angle to acquire video images, and obtain video image acquisition results.
[0091] This invention provides an electronic device, which includes a processor, a memory, and a program stored in the memory and executable on the processor. When the processor executes the program, it implements the steps of any of the above-described voice tracking camera methods.
[0092] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0093] In the above embodiments of the present invention, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0094] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of modules described above can be a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, or indirect coupling or communication connection between modules, and may be electrical or other forms.
[0095] The modules described above as separate components may or may not be physically separate. Similarly, the components shown as modules may or may not be physical modules; they may be located in one place or distributed across multiple modules. Some or all of the modules can be selected to achieve the purpose of this embodiment, depending on actual needs.
[0096] Furthermore, the functional modules in the various embodiments of the present invention can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module. The integrated modules described above can be implemented in hardware or as software functional modules.
[0097] If the aforementioned integrated modules are implemented as software functional modules and sold or used as independent products, they can be stored in a computer-readable non-volatile storage medium. Based on this understanding, the technical solution of this invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a non-volatile storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this invention. The aforementioned non-volatile storage medium includes various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.
[0098] The above are merely preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A voice tracking camera method, characterized in that, include: Based on the first sound pickup amplitude collected by the multiple first microphones included in the microphone array for the target sound-emitting object, the initial horizontal angle of the camera device rotating in the horizontal direction and the initial pitch angle of the camera device rotating in the vertical direction are determined. After controlling the camera device to rotate to the first pickup area corresponding to the initial horizontal angle and the initial pitch angle, the second pickup amplitude of the main microphone is acquired from the target sound source in multiple different pickup directions. The main microphone is a microphone independent of the microphone array. Based on the second pickup amplitude collected by the main microphone from multiple different pickup directions, the target horizontal angle of the camera device rotating along the horizontal direction and the target pitch angle of the camera device rotating along the vertical direction are determined. The camera device is controlled to rotate from the initial horizontal angle to the target horizontal angle, and from the initial pitch angle to the target pitch angle to acquire video images, thereby obtaining video image acquisition results; The method further includes: constructing a path compensation matrix based on the following method, wherein the path compensation matrix is used to perform path compensation on the pickup information collected by the plurality of first microphones for the target vocal object to obtain the first pickup amplitude: obtaining a plurality of pickup areas corresponding to the plurality of first microphones, and a plurality of pitch angles corresponding to the plurality of pickup areas, wherein there is a one-to-one correspondence between the plurality of first microphones and the plurality of pickup areas; obtaining the example speech emitted by the main microphone for the example vocal object in the plurality of pickup areas, and the second speech information corresponding to the plurality of pitch angles corresponding to the plurality of pickup areas; using a second human voice filter to process the second speech information corresponding to the plurality of pitch angles corresponding to the plurality of pickup areas. The two voice information undergo a second filtering process to obtain the fifth pickup amplitude corresponding to the multiple pitch angles of the multiple pickup areas. A second low-noise amplifier is then used to perform a second noise reduction and amplification process on the fifth pickup amplitude corresponding to the multiple pitch angles of the multiple pickup areas, resulting in the sixth pickup amplitude corresponding to the multiple pitch angles of the multiple pickup areas. Based on the sixth pickup amplitude corresponding to the multiple pitch angles of the multiple pickup areas, the maximum pickup amplitude corresponding to each of the multiple pickup areas is determined. Based on the maximum pickup amplitude corresponding to each of the multiple pickup areas, the initial pitch angle corresponding to each of the multiple pickup areas is determined. Finally, based on the maximum pickup amplitude corresponding to each of the multiple pickup areas, a path compensation matrix is determined.
2. The method according to claim 1, characterized in that, Before determining the initial horizontal angle of the camera device's rotation in the horizontal direction and the initial pitch angle of its rotation in the vertical direction based on the first sound pickup amplitude collected by the multiple first microphones included in the microphone array from the target sound-emitting object, the method further includes: Acquire the first voice information collected by the plurality of first microphones respectively targeting the target vocal object; Multiple first voice filters are used to perform first filtering processing on the first speech information corresponding to the multiple first microphones respectively to obtain the third pickup amplitude corresponding to the multiple first microphones respectively, wherein the multiple first voice filters correspond one-to-one with the multiple first microphones; Multiple first low-noise amplifiers are used to perform first noise reduction and amplification processing on the third pickup amplitudes corresponding to the multiple first microphones respectively, to obtain the fourth pickup amplitudes corresponding to the multiple first microphones respectively; A first compensation process is performed on the fourth pickup amplitude corresponding to each of the plurality of first microphones using a pre-constructed path compensation matrix to obtain the first pickup amplitude corresponding to each of the plurality of first microphones.
3. The method according to claim 2, characterized in that, The determination of the initial horizontal angle of the camera device's rotation in the horizontal direction and the initial pitch angle of its rotation in the vertical direction, based on the first sound pickup amplitude collected by the multiple first microphones included in the microphone array from the target sound source, includes: Obtain the first pickup area corresponding to the largest pickup amplitude among the multiple first microphones; The initial horizontal direction angle is determined based on the centerline direction of the first pickup area; The initial pitch angle corresponding to the first pickup area obtained in advance is used as the initial pitch angle.
4. The method according to claim 1, characterized in that, Determining the path compensation matrix based on the maximum pickup amplitude corresponding to each of the multiple pickup areas includes: Based on the maximum pickup amplitude corresponding to the multiple pickup areas respectively, the median pickup amplitude corresponding to the maximum pickup amplitude corresponding to the multiple pickup areas is determined; The path compensation matrix is determined based on the maximum pickup amplitude corresponding to the multiple pickup areas and the median pickup amplitude.
5. The method according to claim 1, characterized in that, After determining the initial pitch angle corresponding to each of the plurality of pickup areas based on the maximum pickup amplitude corresponding to each of the plurality of pickup areas, the method further includes: The multiple micro motors are controlled to adjust the multiple first microphones to their respective initial pitch angle positions.
6. The method according to claim 1, characterized in that, When the plurality of different pickup directions include a first number of horizontal angles and a second number of pitch angles, determining the target horizontal angle of rotation of the camera device along the horizontal direction and the target pitch angle of rotation along the vertical direction based on the second pickup amplitude collected by the main microphone in the plurality of different pickup directions includes: The horizontal angle corresponding to the largest pickup amplitude among the second pickup amplitudes corresponding to the first number of horizontal angles is taken as the target horizontal angle; The pitch angle corresponding to the largest pickup amplitude among the second number of pitch angles is taken as the target pitch angle.
7. The method according to claim 1, characterized in that, The acquisition of the second pickup amplitude, obtained by the main microphone from multiple different pickup directions for the target sound source, includes: Acquire third speech information collected by the main microphone from multiple different pickup directions for the target sound-emitting object; A second human voice filter is used to perform a third filtering process on the third speech information corresponding to the multiple different pickup directions to obtain the seventh pickup amplitude corresponding to the multiple different pickup directions. A second low-noise amplifier is used to perform a third noise reduction and amplification process on the seventh pickup amplitude corresponding to the multiple different pickup directions to obtain the second pickup amplitude corresponding to the multiple different pickup directions.
8. The method according to claim 1, characterized in that, After controlling the camera device to rotate from the initial horizontal angle to the target horizontal angle, and from the initial pitch angle to the target pitch angle to acquire video images and obtain video image acquisition results, the method further includes: For the first pickup amplitudes corresponding to the plurality of first microphones, the largest pickup amplitude is increased by a preset amplification factor; and for the other pickup amplitudes corresponding to the plurality of first microphones, except for the largest pickup amplitude, the preset attenuation factor is decreased, to obtain the audio processing result. Based on the video image acquisition results and the audio processing results, a video synthesis result is obtained.
9. A voice tracking camera system, characterized in that, The system is used to perform the method according to any one of claims 1 to 8, the system comprising: a microphone array; a multi-channel micromotor for controlling a plurality of first microphones included in the microphone array; a camera device; a main microphone parallel to the central axis of the camera device; an azimuth motor for controlling the horizontal rotation of the camera device; a pitch motor for controlling the vertical rotation of the camera device; and a digital processing device, wherein the main microphone is a microphone independent of the microphone array, and the number of channels of the multi-channel micromotor corresponds to the number of the plurality of first microphones, and is used to control the pitch angle corresponding to each of the plurality of first microphones. The microphone array is connected to the multi-channel micro-motor and the digital processing device; the multi-channel micro-motor is connected to the digital processing device; the main microphone is connected to the camera device and the digital processing device; the camera device is connected to the azimuth motor and the pitch motor; the azimuth motor and the pitch motor are respectively connected to the digital processing device, wherein... The plurality of first microphones are used to collect the first pickup amplitude emitted by the target sound source, and send the first pickup amplitude collected for the target sound source to the digital processing device; The digital processing device is used to determine the initial horizontal angle of rotation of the camera device in the horizontal direction and the initial pitch angle in the vertical direction based on the first pickup amplitude collected by the plurality of first microphones for the target sound-emitting object; and to control the camera device to rotate to the first pickup area corresponding to the initial horizontal angle and the initial pitch angle. The main microphone is used to collect the second pickup amplitude from the target sound source in multiple different pickup directions, and to send the second pickup amplitude collected in the multiple different pickup directions to the digital processing device. The digital processing device is used to determine the target horizontal angle of the camera device rotating along the horizontal direction and the target pitch angle of the camera device rotating along the vertical direction based on the second pickup amplitude collected by the main microphone in the multiple different pickup directions, and to control the camera device to rotate from the initial horizontal angle to the target horizontal angle and from the initial pitch angle to the target pitch angle; The camera device acquires video images and obtains the video image acquisition results.
10. The system according to claim 9, characterized in that, The system further includes: a first voice filter corresponding to each of the plurality of first microphones, and a first low-noise amplifier corresponding to each of the plurality of first microphones, wherein the plurality of first microphones are connected to the digital processing device in sequence through the corresponding first voice filter and the first low-noise amplifier; a second voice filter corresponding to the main microphone, and a second low-noise amplifier corresponding to the main microphone, wherein the main microphone is connected to the digital processing device in sequence through the second voice filter and the second low-noise amplifier.
11. A non-volatile storage medium, characterized in that, The non-volatile storage medium stores multiple instructions, which are adapted to be loaded by a processor and executed by the voice tracking camera method according to any one of claims 1 to 8.
12. An electronic device, characterized in that, It includes one or more processors and a memory, the memory being used to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors cause the one or more processors to implement the voice tracking camera method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Microphone array-based wireless video tracking and monitoring system
CN203366132U