A recording and broadcasting device, method, apparatus and medium
The recording and broadcasting equipment, consisting of a controller and a microphone, uses audio information to determine the location of the sound source and adjust the angle of the pan-tilt unit. Combined with the lens to collect video information, it solves the problem of poor adaptability of the recording and broadcasting equipment in various scenarios and achieves flexible recording and broadcasting effects.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ZHEJIANG DAHUA TECH CO LTD
- Filing Date
- 2023-04-18
- Publication Date
- 2026-04-14
AI Technical Summary
Existing recording and broadcasting equipment has high requirements in various recording and broadcasting scenarios, makes it difficult to record targets in scattered locations, and has poor environmental adaptability.
The recording and broadcasting equipment consists of a controller that controls a pan-tilt unit and a microphone. The microphone collects audio information to determine the location of the sound source, and the pan-tilt unit is adjusted to accurately record the target person. Video information is collected in combination with a wide-angle lens and an auxiliary lens.
It enables accurate switching of perspectives in various recording and broadcasting scenarios, adapts to different environments, and improves the flexibility and recording effect of recording and broadcasting equipment.
Smart Images

Figure CN116546328B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of surveillance technology, and in particular to a recording and broadcasting device, method, apparatus and medium. Background Technology
[0002] Recorded videos are often used in classrooms, meetings, and other similar settings. Figure 1a This diagram illustrates a common recording and broadcasting scenario using existing technologies. Figure 1b This is a diagram illustrating another common recording and broadcasting scenario provided by existing technology. Figure 1a For meeting scenarios, Figure 1b For classroom settings. Figure 1a The participants and Figure 1b The number of students in the recording is relatively large and their locations are scattered; any participant or student could become the target of recording at any time. Therefore, the recording scene is characterized by a large screen and random perspectives.
[0003] Existing recording methods typically include the following: First, based on the captured panoramic video, moving targets are cut out to obtain close-up shots, and the classroom video is recorded by switching between panoramic and close-up views. Second, multiple pan-tilt cameras and recording devices are integrated onto the recording blackboard, and the positions of the captured target person are calculated and recorded. Third, the position of the speaker is monitored by a monitoring module, along with the position, angle, and distance of the microphone on the moving rail, to record the speaker. However, the first method has high requirements for the number, resolution, and installation location of camera lenses; the second method can only be used in classrooms with blackboards, limiting its application; and the third method requires the installation of a rail, placing high demands on the environment.
[0004] Therefore, how to provide a convenient recording and broadcasting device that can be applied to various recording and broadcasting scenarios and can accurately switch perspectives to record and broadcast the target person is a technical problem that urgently needs to be solved by those skilled in the art. Summary of the Invention
[0005] This application provides a recording and broadcasting device, method, apparatus, and medium to solve the problems in the prior art.
[0006] In a first aspect, this application provides a recording and broadcasting device, which includes: a controller, a base, a pan-tilt unit, a lens, three first microphones and a second microphone; wherein the pan-tilt unit and the three first microphones are all disposed under the base, and the three first microphones are arranged in a triangular pattern;
[0007] The lens is mounted on the bottom of the gimbal, and the second microphone is mounted on the bottom of the lens;
[0008] The controller is configured to determine the angular deviation value between the sound source and multiple preset positions based on the audio information collected by each microphone, wherein the preset position is the midpoint of the line connecting any two microphones; determine the target position of the sound source based on each angular deviation value, the distance between any two microphones stored in advance, and a first distance between the preset sound source and the preset position; determine the target angle value for the pan-tilt unit to rotate based on the target position of the sound source; and control the pan-tilt unit to rotate based on the target angle value, so that the rotated lens and the second microphone can collect video and audio information of the sound source.
[0009] In one possible implementation, the controller is specifically configured to, for any angular deviation value, determine the superposition value and average value of the sound intensities collected by the two microphones connected at a preset position to determine the angular deviation value, based on the sound intensity of the sound source collected by the two microphones, and determine the angular deviation value based on the superposition value and the average value.
[0010] In one possible implementation, the controller is specifically configured to determine the horizontal and vertical angles between the sound source and the gimbal based on the target position of the sound source, and to determine the horizontal and vertical angles as the target angle values that the gimbal needs to rotate.
[0011] In one possible implementation, the controller is specifically configured to, after controlling the pan-tilt unit to rotate according to the target angle value, determine whether the sound intensity in the audio information of the sound source collected by the second microphone is less than a preset intensity; if so, update the first distance using a preset second distance, and redetermine the target position of the sound source based on the updated first distance.
[0012] In one possible implementation, the controller is specifically configured to acquire audio information of the sound source collected by the second microphone; determine whether the sound intensity of the audio information is less than a preset intensity; if not, generate a video file based on the acquired video and audio information.
[0013] In one possible implementation, the controller is specifically configured to adjust the audio acquisition gain of the second microphone if it is determined that the sound intensity of the audio information is less than a preset intensity, until the sound intensity of the audio information acquired by the second microphone is not less than the preset intensity.
[0014] In one possible implementation, the recording and broadcasting device further includes: at least three wide-angle lenses, and the at least three wide-angle lenses are evenly arranged under the base;
[0015] The controller is further configured to, if it acquires sub-audio information of other sound sources collected by each of the first microphones, determine other locations of the other sound sources, and, based on the determined other locations of the other sound sources and the panoramic video information acquired by the at least three wide-angle lenses, determine the image position of the other sound sources in the panoramic video information, and respectively extract local video information corresponding to each image position.
[0016] In one possible implementation, the controller is further configured to generate a video file based on the local video information, the subaudio information, the main audio information of the sound source acquired by the second microphone, and the main video information of the sound source acquired by the lens.
[0017] In one possible implementation, the recording and broadcasting device further includes:
[0018] At least one auxiliary lens located elsewhere;
[0019] The controller is further configured to determine the target auxiliary lens corresponding to the shooting range of the target position of the sound source based on the determined target position of the sound source, the pre-saved position of each auxiliary lens and the shooting range, determine the third distance, horizontal offset angle and vertical offset angle between the target auxiliary lens and the sound source; control the target auxiliary lens to focus based on the third distance, horizontal offset angle and vertical offset angle, and collect the sub-video information of the sound source.
[0020] Secondly, this application provides a recording method, the method comprising:
[0021] Based on the audio information collected by each microphone, the angular deviation values between the sound source and multiple preset positions are determined, wherein the preset positions are the midpoints of the lines connecting any two microphones; based on each angular deviation value, the distance between any two microphones that are stored in advance, and the first distance between the preset sound source and the preset positions, the target position of the sound source is determined; based on the target position of the sound source, the target angle value that the pan-tilt unit needs to rotate is determined; based on the target angle value, the pan-tilt unit is controlled to rotate so that the rotated lens and the second microphone can collect the video and audio information of the sound source.
[0022] In one possible implementation, determining the angular deviation value between the sound source and multiple preset positions based on the audio information collected by each microphone includes:
[0023] For any given angular deviation value, based on the sound intensity of the sound source collected by two microphones connected at a preset position that determines the angular deviation value, the superposition value and average value of the sound intensity collected by the two microphones are determined, and the angular deviation value is determined based on the superposition value and average value.
[0024] In one possible implementation, determining the target angle value that the pan-tilt unit needs to rotate based on the target location of the sound source includes:
[0025] Based on the target location of the sound source, determine the horizontal and vertical angles between the sound source and the gimbal, and then determine the horizontal and vertical angles as the target angle values that the gimbal needs to rotate.
[0026] In one possible implementation, after controlling the gimbal rotation according to the target angle value, the method further includes:
[0027] Determine whether the sound intensity in the audio information of the sound source collected by the second microphone is less than a preset intensity; if so, update the first distance using a preset second distance, and redetermine the target position of the sound source based on the updated first distance.
[0028] In one possible implementation, after controlling the pan-tilt unit to rotate according to the target angle value so that the rotated lens and second microphone can acquire video and audio information of the sound source, the method further includes:
[0029] The system acquires the audio information of the sound source collected by the second microphone; it determines whether the sound intensity of the audio information is less than a preset intensity; if not, it generates a video file based on the acquired video and audio information.
[0030] In one possible implementation, if it is determined that the sound intensity of the audio information is less than a preset intensity, the audio acquisition gain of the second microphone is adjusted until the sound intensity of the audio information acquired by the second microphone is not less than the preset intensity.
[0031] In one possible implementation, the method further includes:
[0032] If the sub-audio information of other sound sources collected by each of the first microphones is obtained, the other positions of the other sound sources are determined. Based on the determined other positions of the other sound sources and the panoramic video information collected by the at least three wide-angle lenses, the position of the other sound sources in the panoramic video information is determined, and the local video information corresponding to each position is extracted.
[0033] In one possible implementation, after extracting the local video information corresponding to each frame position, the method further includes:
[0034] A video file is generated based on the local video information, the secondary audio information, the main audio information of the sound source collected by the second microphone, and the main video information of the sound source collected by the lens.
[0035] In one possible implementation, the method further includes:
[0036] Based on the determined target location of the sound source, the pre-saved position and shooting range of each auxiliary lens, the target auxiliary lens corresponding to the shooting range of the target location of the sound source is determined, and the third distance, horizontal offset angle and vertical offset angle between the target auxiliary lens and the sound source are determined; the target auxiliary lens is controlled to focus according to the third distance, horizontal offset angle and vertical offset angle, and the sub-video information of the sound source is collected.
[0037] Thirdly, this application also provides a recording and broadcasting device, the device comprising:
[0038] The determination module is used to determine the angular deviation value between the sound source and multiple preset positions based on the audio information collected by each microphone, wherein the preset position is the midpoint of the line connecting any two microphones; determine the target position of the sound source based on each angular deviation value, the distance between any two microphones that are stored in advance, and the first distance between the preset sound source and the preset position; and determine the target angle value that the pan-tilt unit needs to rotate based on the target position of the sound source.
[0039] The control module is used to control the rotation of the pan-tilt unit according to the target angle value, so that the rotated lens and the second microphone can collect video and audio information of the sound source.
[0040] In one possible implementation, the determining module is specifically used to determine, for any angle deviation value, the superposition value and average value of the sound intensity collected by the two microphones connected at a preset position to determine the angle deviation value, based on the sound intensity of the sound source collected by the two microphones. The angle deviation value is then determined based on the superposition value and the average value.
[0041] In one possible implementation, the determining module is specifically used to determine the horizontal and vertical angles between the sound source and the pan-tilt unit based on the target position of the sound source, and to determine the horizontal and vertical angles as the target angle values that the pan-tilt unit needs to rotate.
[0042] In one possible implementation, the determining module is further configured to determine whether the sound intensity in the audio information of the sound source collected by the second microphone is less than a preset intensity; if so, the first distance is updated using a preset second distance, and the target position of the sound source is re-determined based on the updated first distance.
[0043] In one possible implementation, the device further includes:
[0044] The first generation module is used to acquire the audio information of the sound source collected by the second microphone; determine whether the sound intensity of the audio information is less than a preset intensity; if not, generate a video file based on the acquired video information and audio information.
[0045] In one possible implementation, the determining module is further configured to, if it is determined that the sound intensity of the audio information is less than a preset intensity, adjust the audio acquisition gain of the second microphone until the sound intensity of the audio information acquired by the second microphone is not less than the preset intensity.
[0046] In one possible implementation, the determining module is further configured to, if the sub-audio information of other sound sources acquired by each of the first microphones is obtained, determine the other locations of the other sound sources, and, based on the determined other locations of the other sound sources and the panoramic video information acquired by the at least three wide-angle lenses, determine the image position of the other sound sources in the panoramic video information, and respectively extract the local video information corresponding to each image position.
[0047] In one possible implementation, the device further includes:
[0048] The second generation module is used to generate a video file based on the local video information, the sub-audio information, the main audio information of the sound source collected by the second microphone, and the main video information of the sound source collected by the lens.
[0049] In one possible implementation, the determining module is further configured to determine, based on the determined target location of the sound source, the pre-saved location of each auxiliary lens and the shooting range, the target auxiliary lens corresponding to the shooting range of the target location of the sound source, and the third distance, horizontal offset angle and vertical offset angle between the target auxiliary lens and the sound source.
[0050] The control module is also used to control the target auxiliary lens to focus based on the third distance, horizontal offset angle and vertical offset angle, and to collect sub-video information of the sound source.
[0051] Fourthly, this application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of any of the recording and broadcasting methods described above.
[0052] In this embodiment, the recording and broadcasting device includes: a controller, a base, a pan-tilt unit, a lens, three first microphones, and one second microphone. The pan-tilt unit and the three first microphones are all mounted under the base, and the three first microphones are arranged in a triangular pattern. The lens is mounted at the bottom of the pan-tilt unit, and the second microphone is mounted at the bottom of the lens. This allows for the acquisition of audio information from multiple directions. The controller is used to determine the angular deviation value between the sound source and multiple preset positions based on the audio information acquired by each microphone. Based on each angular deviation value, the pre-saved distance between any two microphones, and a preset first distance between the sound source and a preset position, the controller can determine the target angle value for the pan-tilt unit to rotate. The controller controls the pan-tilt unit to rotate according to the target angle value, so that the rotated lens and second microphone acquire video and audio information of the sound source. This allows for application in various recording and broadcasting scenarios and enables accurate switching of perspectives to record and broadcast the target person from the sound source. Attached Figure Description
[0053] To more clearly illustrate the implementation methods in the embodiments of this application or related technologies, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings.
[0054] Figure 1a A schematic diagram of a common recording and broadcasting scenario provided for existing technologies;
[0055] Figure 1b A schematic diagram of another common recording and broadcasting scenario provided by existing technology;
[0056] Figure 2a A structural side view of a recording and broadcasting device provided in an embodiment of this application;
[0057] Figure 2b A top view of the structure of a recording and broadcasting device provided in an embodiment of this application;
[0058] Figure 3 This is a schematic diagram of a directional sound pickup unit provided in an embodiment of this application;
[0059] Figure 4 A schematic diagram of sound source localization provided in an embodiment of this application;
[0060] Figure 5 A schematic diagram of a sound source angle provided in an embodiment of this application;
[0061] Figure 6 A schematic diagram illustrating the process of locating a sound source using a recording and broadcasting device provided in an embodiment of this application;
[0062] Figure 7This application provides a schematic diagram illustrating the process of optimizing recording effects using a recording and broadcasting device.
[0063] Figure 8a This application provides a schematic diagram of a meeting scenario in a silent state, as illustrated in an embodiment of the present application.
[0064] Figure 8b This application provides a schematic diagram of a meeting scenario where participants are speaking, as illustrated in an embodiment of the present application.
[0065] Figure 8c This is a schematic diagram of the rotation of a recording and broadcasting device provided in an embodiment of this application;
[0066] Figure 8d This is a schematic diagram of a recording and broadcasting device provided in an embodiment of this application;
[0067] Figure 9 A side view of the structure of another recording and broadcasting device provided in an embodiment of this application;
[0068] Figure 10 A schematic diagram of a multi-sound-source recording scene and video file provided in an embodiment of this application;
[0069] Figure 11 This is a schematic diagram illustrating the process of generating a video file, as provided in an embodiment of this application.
[0070] Figure 12 A schematic diagram of another multi-sound-source recording and broadcasting scenario and video information provided in an embodiment of this application;
[0071] Figure 13 This is a schematic diagram of the operation of a recording and broadcasting system provided in an embodiment of this application;
[0072] Figure 14 This is a schematic diagram of the recording process of a recording and broadcasting system provided in an embodiment of this application;
[0073] Figure 15 This is a schematic diagram of the structure of a recording and broadcasting device provided in an embodiment of this application. Detailed Implementation
[0074] To make the objectives, technical solutions, and advantages of this application clearer, a further detailed description of this application will be provided below with reference to the accompanying drawings. Obviously, the embodiments described in this application are merely some embodiments, not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0075] It should be noted that the brief descriptions of terms in this application are only for the convenience of understanding the embodiments described below, and are not intended to limit the embodiments of this application. Unless otherwise stated, these terms should be understood in their ordinary and common meaning.
[0076] The terms "first," "second," "third," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar or related objects or entities, and do not necessarily imply a specific order or sequence, unless otherwise specified. It should be understood that such terms are interchangeable where appropriate.
[0077] The terms “comprising” and “having”, and any variations thereof, are intended to cover but not exclude inclusion, for example, a product or device that includes a range of components is not necessarily limited to all of the components that are clearly listed, but may include other components that are not clearly listed or that are inherent to such product or device.
[0078] The term "module" refers to any known or subsequently developed hardware, software, firmware, artificial intelligence, fuzzy logic, or combination of hardware and / or software code that is capable of performing the functions associated with that element.
[0079] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.
[0080] This application provides a recording and broadcasting device, method, apparatus, and medium. The recording and broadcasting device includes: a controller, a base, a pan-tilt unit, a lens, three first microphones, and a second microphone. The pan-tilt unit and the three first microphones are all disposed under the base, and the three first microphones are arranged in a triangular pattern. The lens is mounted on the bottom of the pan-tilt unit, and the second microphone is mounted on the bottom of the lens. The controller is configured to: determine the angular deviation value between a sound source and multiple preset positions based on the audio information collected by each microphone, wherein the preset position is the midpoint of the line connecting any two microphones; determine the target position of the sound source based on each angular deviation value, a pre-stored distance between any two microphones, and a preset first distance between the sound source and the preset position; determine the target angle value by which the pan-tilt unit needs to rotate based on the target position of the sound source; and control the pan-tilt unit to rotate based on the target angle value, so that the rotated lens and the second microphone collect video and audio information of the sound source.
[0081] In order to adapt to a variety of different recording and broadcasting scenarios, this application provides a recording and broadcasting device, method, apparatus and medium.
[0082] Example 1:
[0083] Figure 2a This is a structural side view of a recording and broadcasting device provided in an embodiment of this application. Figure 2b This is a top view of the structure of a recording and broadcasting device provided in an embodiment of this application. Figure 2a and Figure 2b As shown, the recording and broadcasting equipment includes:
[0084] The system includes a controller 100, a base 101, a gimbal 102, a lens 103, three first microphones 104, and a second microphone 105; wherein the gimbal 102 and the three first microphones 104 are all located under the base 101, and the three first microphones 104 are arranged in a triangular pattern.
[0085] The lens 103 is mounted on the bottom of the gimbal 102, and the second microphone 105 is mounted on the bottom of the lens 103;
[0086] The controller 100 is configured to determine the angular deviation value between the sound source and multiple preset positions based on the audio information collected by each microphone, wherein the preset position is the midpoint of the line connecting any two microphones; determine the target position of the sound source based on each angular deviation value, the distance between any two microphones stored in advance, and a first distance between the preset sound source and the preset position; determine the target angle value that the pan-tilt unit 102 needs to rotate based on the target position of the sound source; and control the pan-tilt unit 102 to rotate based on the target angle value, so that the rotated lens 103 and the second microphone 105 can collect the video and audio information of the sound source.
[0087] The controller 100 can also be an electronic device such as a server located outside the recording and broadcasting equipment.
[0088] Based on directional sound pickup technology, any two microphones can form a directional sound pickup unit. The greater the angular deviation between the sound source and the directional sound pickup unit, the lower the sound intensity of the audio information collected by the directional sound pickup unit. When the angular deviation is greater than 45 degrees, the sound intensity is zero. Figure 3 This is a schematic diagram of a directional sound pickup unit provided in an embodiment of this application. Figure 3As shown, the directional pickup unit consists of two microphones, which can acquire the sound intensity of audio information located within the effective area, and then determine the angular deviation value between the sound source and the directional pickup unit, i.e., the value of α in the figure. The effective area is within 45 degrees of the direction of the line connecting the two microphones, and the area outside 45 degrees of the direction of the line connecting the two microphones is the ineffective area. The angular deviation value is the angle between the two connecting lines. The first connecting line is the line connecting the midpoint of the line connecting the sound source and the two microphones of the directional pickup unit, and the second connecting line is the line connecting the two microphones of the directional pickup unit.
[0089] Based on this, in this embodiment of the application, the recording and broadcasting device is equipped with three first microphones located on the same horizontal plane, namely the first microphone 104 located under the base 101, and a second microphone located on another horizontal plane, namely the second microphone 105 located at the bottom of the lens 103. Any two first microphones located on the same horizontal plane can form a directional pickup unit, and any one of the first and second microphones can also form a directional pickup unit, thereby enabling the detection of sound sources from multiple directions. The midpoint of the line connecting any two microphones is determined as a preset position, thereby determining the angular deviation values between the sound source and the multiple preset positions. Furthermore, based on each angular deviation value, the pre-stored distance between any two microphones, and the preset first distance between the sound source and the preset position, the target position of the sound source can be determined.
[0090] Based on the target location of the sound source, the angle of the sound source relative to the second microphone in the horizontal and vertical directions can be determined. This is the target angle value that the pan-tilt unit carrying the lens and the second microphone needs to rotate. Therefore, the pan-tilt unit can be controlled to rotate according to the target angle value so that the rotated lens and the second microphone can collect video and audio information of the sound source.
[0091] Specifically, the process of determining the target location of the sound source in the embodiments of this application will be described below with a specific example.
[0092] Figure 4 This is a schematic diagram illustrating sound source localization provided in an embodiment of this application. Figure 4As shown, Mic1, Mic2, and Mic3 are three first microphones arranged in a triangle under the base, and Mic4 is a second microphone mounted at the bottom of the lens. The speaker is the sound source D. Since Mic4 is closer to the lens, a Cartesian coordinate system can be established with Mic4 as the origin. Assume the coordinates of the target position of sound source D are (x, y, z), and the first microphones Mic1, Mic2, and Mic3 are arranged in an equilateral triangle on the same horizontal plane, with the distance between any two first microphones being L, and the distance between the second microphone Mic4 and the plane containing the first microphones Mic1, Mic2, and Mic3 being H. Therefore, the coordinates of the positions of Mic1, Mic2, Mic3, and Mic4 can be determined as follows: And (0, 0, 0). The preset positions are the midpoints of the lines connecting Mic1 and Mic4 (A), Mic2 and Mic4 (B), and Mic3 and Mic4 (C). The coordinates of A, B, and C can be determined as follows: The angle between the line connecting the sound source and the preset position A and the line connecting Mic1 and Mic4 is α; the angle between the line connecting the sound source and the preset position B and the line connecting Mic2 and Mic4 is β; and the angle between the line connecting the sound source and the preset position C and the line connecting Mic3 and Mic4 is γ.
[0093] Based on the dot product relationship of vectors, we can obtain the following equation (1):
[0094]
[0095] Where cosα is the cosine of the angle α between the sound source and the preset position A. It is a vector determined by the coordinates of the locations of Mic1 and Mic4; yes The magnitude of the vector is also the distance between Mic1 and Mic4; It is a vector determined by the coordinates of the preset position A and the position of the sound source D; yes The magnitude of the vector is also the distance between the preset position A and the sound source D; cosβ is the cosine of the angle β between the sound source and the preset position B. It is a vector determined by the coordinates of the locations of Mic2 and Mic4; yes The magnitude of the vector is also the distance between Mic2 and Mic4; It is a vector determined by the coordinates of the preset position B and the position of the sound source D; yes The magnitude of the vector is also the distance between the preset position B and the sound source D; cosγ is the cosine of the angle γ between the sound source and the preset position C. It is a vector determined by the coordinates of the locations of Mic3 and Mic4; yes The magnitude of the vector is also the distance between Mic3 and Mic4; It is a vector determined by the coordinates of the preset position C and the position of the sound source D; yes The magnitude of the vector is also the distance between the preset position C and the sound source D.
[0096] Since the recording equipment is installed at a relatively high position, the distances between the sound source and the preset positions A, B, and C can be considered approximately equal, and the first distance between the preset sound source and the preset position is L. D Therefore, we can obtain the following equation (2):
[0097]
[0098] Among them, L D The first distance between the preset sound source and the preset position. It is the magnitude of the vector determined by the coordinates of the preset position A and the position of the sound source D, and it is also the distance between the preset position A and the sound source D. It is the magnitude of the vector determined by the coordinates of the preset position B and the position of the sound source D, and it is also the distance between the preset position B and the sound source D; It is the magnitude of the vector determined by the coordinates of the preset position C and the position of the sound source D, and it is also the distance between the preset position C and the sound source D; Let x, y, z be the first distance between the location (x, y, z) of sound source D and the location (0, 0, 0) of Mic4. x, y, z are the coordinates of the location of sound source D in a Cartesian coordinate system with Mic4 as the origin.
[0099] Based on equations (1) and (2), we can obtain equation (3), and then determine the coordinates of the target location where the sound source D is located:
[0100]
[0101] in, Let L be the vector formed by the coordinates (x, y, z) of the location of sound source D in a rectangular coordinate system. D (cosβ-cosα) is the value of x. For the value corresponding to y, 2L D L(cosα+cosβ+cosγ)-L 2 -3H 2 / 6H is the value corresponding to z; cosα is the cosine of the angle α between the sound source and the preset position A; cosβ is the cosine of the angle β between the sound source and the preset position B; cosγ is the cosine of the angle γ between the sound source and the preset position C; the distance between any two first pickups is L; and the distance between the second pickup Mic4 and the plane containing the first pickups Mic1, Mic2, and Mic3 is H.
[0102] Since the coordinates of sound source D are determined using a Cartesian coordinate system with the second microphone Mic4 as the origin, the coordinates of sound source D in the spherical coordinate system with Mic4 as the origin can be determined based on the transformation relationship between the Cartesian and spherical coordinate systems and the coordinates of sound source D in the Cartesian coordinate system. This allows us to obtain the angle values of sound source D relative to Mic4 in the horizontal and vertical directions, which are also the target angle values that the pan-tilt unit supporting the lens and Mic4 needs to rotate. Therefore, the pan-tilt unit can be controlled to rotate based on the target angle values, so that the rotated lens and second microphone can acquire video and audio information from sound source D.
[0103] In this embodiment, the recording and broadcasting device includes: a controller, a base, a pan-tilt unit, a lens, three first microphones, and one second microphone. The pan-tilt unit and the three first microphones are all mounted under the base, and the three first microphones are arranged in a triangular pattern. The lens is mounted at the bottom of the pan-tilt unit, and the second microphone is mounted at the bottom of the lens. This allows for the acquisition of audio information from multiple directions. The controller is used to determine the angular deviation value between the sound source and multiple preset positions based on the audio information acquired by each microphone. Based on each angular deviation value, the pre-saved distance between any two microphones, and a preset first distance between the sound source and a preset position, the controller can determine the target angle value for the pan-tilt unit to rotate. The controller controls the pan-tilt unit to rotate according to the target angle value, so that the rotated lens and second microphone acquire video and audio information of the sound source. This allows for application in various recording and broadcasting scenarios and enables accurate switching of perspectives to record and broadcast the target person from the sound source.
[0104] Example 2:
[0105] In order to accurately determine the angular deviation value between the sound source and the preset position, based on the above embodiments, in this embodiment of the application, the controller is specifically used to determine the superposition value and average value of the sound intensity collected by the two microphones connected to the preset position where the angular deviation value is determined, and to determine the angular deviation value based on the superposition value and the average value.
[0106] In this embodiment, based on directional sound pickup technology, the midpoint of the line connecting any two microphones is determined as a preset position. Since the larger the angular deviation between the sound source and the preset position, the lower the sound intensity of the audio information collected by the two microphones, and the sound intensity is zero when the angular deviation is greater than 45 degrees, the correspondence between sound intensity and angular deviation can be determined. In this embodiment, by determining the midpoint of the line connecting any two microphones as the preset position, the angular deviation between the sound source and each preset position can be determined by the sound intensity of the audio information collected by the two microphones connected to each preset position.
[0107] Specifically, for any given angular deviation value, based on the sound intensities of the audio information collected by two microphones connected at a preset position that determines the angular deviation value, the superposition value and average value of the sound intensities collected by the two microphones can be determined. The angular deviation value can then be determined by the ratio of the superposition value to the average value.
[0108] The angle deviation value can be determined according to the following formula (4):
[0109] α=k α / k(4)
[0110] Wherein, the angle deviation value is α, and the summation value of the sound intensities of the audio information collected by the two microphones is k. α The average value is k.
[0111] In this embodiment of the application, the angle deviation value is determined based on the superposition and average value of the sound intensity of the sound source collected by two microphones connected at a preset position to determine the angle deviation value.
[0112] Example 3:
[0113] In order to record the target person of the sound source more accurately, based on the above embodiments, in this embodiment of the application, the controller is specifically used to determine the horizontal and vertical angles of the sound source relative to the gimbal based on the target position of the sound source, and to determine the horizontal and vertical angles as the target angle values that the gimbal needs to rotate.
[0114] Figure 5 This is a schematic diagram of a sound source angle provided for an embodiment of this application. For example... Figure 5 As shown, θ is the angle between the sound source and the horizontal plane of the pan-tilt unit. Let θ be the angle between the sound source and the vertical plane relative to the pan-tilt unit. Based on the coordinates of the target location of the sound source, the horizontal and vertical angles θ between the sound source and the pan-tilt unit can be determined. This allows us to determine the angle θ between the horizontal plane and the vertical plane. The target angle value for the gimbal to rotate is determined.
[0115] Specifically, the coordinates of the target location where the sound source D is located in the above equation (3) will be used as an example for explanation.
[0116] The coordinates of the target location where the sound source D is located in equation (3) above are determined based on a rectangular coordinate system established with the second microphone at the bottom of the lens on the bottom of the pan-tilt unit as the origin. Figure 2a As can be seen, the gimbal, lens, and second microphone are closely connected, so they can be considered as a single unit. Based on the conversion relationship between Cartesian and spherical coordinate systems, the coordinates of the target position where the sound source D is located can be directly converted from the Cartesian coordinate system to the spherical coordinate system.
[0117] Based on the above equation (3) and the transformation relationship between rectangular coordinate system and spherical coordinate system (5), the coordinates of the target position where the sound source D is located in the spherical coordinate system are determined by the following equation (6):
[0118]
[0119]
[0120] in, Let θ represent the coordinates of the target location of sound source D in a spherical coordinate system. r represents the distance between sound source D and the origin of the spherical coordinate system, and θ represents the horizontal angle between sound source D and the origin of the spherical coordinate system. This represents the angle between the sound source D and the origin of the spherical coordinate system. x, y, z are the coordinates of the sound source D in the rectangular coordinate system. cosα is the cosine of the angle deviation α between the sound source and the preset position A. cosβ is the cosine of the angle deviation β between the sound source and the preset position B. cosγ is the cosine of the angle deviation γ between the sound source and the preset position C. The distance between any two first pickups is L. The distance between the second pickup Mic4 and the plane containing the first pickups Mic1, Mic2, and Mic3 is H.
[0121] In this embodiment, the horizontal and vertical angles between the sound source and the pan-tilt unit are determined based on the target location of the sound source. These angles are then used as the target angle values for the pan-tilt unit to rotate, thereby enabling the rotated lens and the second microphone to accurately capture the video and audio information of the sound source.
[0122] Example 4:
[0123] To ensure recording quality, based on the above embodiments, in this embodiment, the controller is specifically used to control the pan-tilt unit to rotate according to the target angle value, and then determine whether the sound intensity in the audio information of the sound source collected by the second microphone is less than a preset intensity; if so, the first distance is updated using a preset second distance, and the target position of the sound source is re-determined based on the updated first distance.
[0124] In this embodiment, to ensure recording quality, a preferred sound intensity can be predetermined as a preset intensity. After controlling the pan-tilt unit to rotate according to the target angle value so that the rotated second microphone collects the audio information of the sound source, it can be determined whether the sound intensity in the audio information collected by the second microphone is less than the preset intensity. If so, it indicates that the sampling angle of the second microphone may not be perfectly aligned with the target position of the sound source. Since the target position is determined based on multiple angular deviation values between the sound source and the preset position, the distance between any two microphones that are stored in advance, and the first distance between the sound source and the preset position, the preset position is the midpoint of the line connecting any two microphones, including the midpoint of the line connecting any two first microphones and the midpoint of the line connecting any one first microphone and the second microphone. Since the recording equipment is installed at a high position and is relatively small, the distance between the sound source and the midpoint of the line connecting any two microphones in the recording equipment can be preset as the first distance between the sound source and the recording equipment. The distance between any two microphones is fixed in the pre-stored data, and the angular deviation between the sound source and the preset position determined by directional sound pickup technology is also relatively accurate. Therefore, the first distance between the preset sound source and the preset position may be inaccurate. Thus, the first distance can be updated using the second distance between the preset sound source and the preset position to re-determine the target position of the sound source.
[0125] The preset first and second distances can be determined based on the installation height of the recording equipment, the horizontal distance between the installation position and the speaker corresponding to the sound source, and the height of the speaker's mouth. For example, the installation height of the recording equipment, the horizontal distance between the positions of multiple speakers and the installation position of the recording equipment, and the height of the speaker's mouth in standing and sitting positions can be pre-saved. Then, based on each horizontal distance, mouth height, and installation height of the recording equipment, a set of distances between the sound source and the preset position can be determined. When determining the target position of the sound source for the first time, any distance can be selected from the distance set as the first distance. Subsequently, any distance can be selected from the remaining distances as the second distance. The first distance is updated using the second distance, and the target position of the sound source is re-determined based on the updated first distance. The sound intensity in the audio information of the sound source collected by the second microphone is not less than the preset intensity until the pan-tilt unit is rotated according to the target position.
[0126] In this embodiment, after controlling the pan-tilt unit to rotate according to the target angle value, if it is determined that the sound intensity in the audio information of the sound source collected by the second microphone is less than a preset intensity, then the first distance is updated using a preset second distance, and the target position of the sound source is re-determined based on the updated first distance. This enables the recording equipment to be aligned with the sound source, improving the recording effect.
[0127] Based on the above embodiments, the following specific example will be used to illustrate the overall process of determining the target location of a sound source.
[0128] Figure 6 This is a schematic diagram illustrating the process of locating a sound source using a recording and broadcasting device provided in an embodiment of this application. Figure 6 As shown, the process includes the following steps:
[0129] S601: The microphone collects audio information.
[0130] If the microphone of the recording equipment detects sound, it will collect the audio information of the sound.
[0131] S602: The controller determines the angular deviation value between the sound source and multiple preset positions based on the audio information collected by each microphone, wherein the preset position is the midpoint of the line connecting any two microphones.
[0132] S603: The controller determines the target position of the sound source based on each of the angle deviation values, the distance between any two microphones that are stored in advance, and the first distance between the preset sound source and the preset position.
[0133] S604: The controller determines the target angle value that the pan-tilt unit needs to rotate based on the target position of the sound source.
[0134] S605: The controller controls the pan-tilt unit to rotate according to the target angle value, so that the rotated lens and the second microphone can collect video and audio information of the sound source.
[0135] S606: The controller determines whether the sound intensity in the audio information of the sound source collected by the second microphone is less than a preset intensity.
[0136] S607a: If so, the first distance is updated using a preset second distance, and the target position of the sound source is re-determined based on the updated first distance. Then, 603 is executed.
[0137] S607b: If not, the positioning process ends.
[0138] Example 5:
[0139] To ensure the audio quality of the video file, based on the above embodiments, in this embodiment, the controller is specifically used to acquire the audio information of the sound source collected by the second microphone; determine whether the sound intensity of the audio information is less than a preset intensity; if not, generate a video file based on the acquired video information and audio information.
[0140] In this embodiment, to ensure the audio quality of the video file generated by the recording device, after acquiring the audio information of the sound source collected by the second microphone, it is first determined whether the sound intensity of the audio information is less than a preset intensity. If it is determined that the sound intensity of the audio information is not less than the preset sound intensity, then a video file is generated based on the acquired video information and audio information.
[0141] To further ensure the audio quality of the video file, based on the above embodiments, in this embodiment of the application, the controller is specifically used to adjust the audio acquisition gain of the second microphone if it is determined that the sound intensity of the audio information is less than a preset intensity, until the sound intensity of the audio information acquired by the second microphone is not less than the preset intensity.
[0142] To further improve recording quality and ensure high-quality audio in the video file, after controlling the pan-tilt-zoom (PTZ) rotation based on the target angle, the directional pickup function of the second microphone mounted under the lens can be activated to filter noise from directions other than the target angle. The magnification of the lens mounted on the rotated PTG is then adjusted to make the video and audio information from the sound source captured by the lens and the second microphone clearer. Based on this, it can be determined whether the sound intensity of the captured audio information is less than a preset intensity. If it is determined that the sound intensity of the captured audio information is less than the preset intensity, the audio acquisition gain of the second microphone can be adjusted until the sound intensity of the audio information captured by the second microphone is no less than the preset intensity.
[0143] In this embodiment of the application, by determining whether the sound intensity of the audio information collected by the second microphone is less than a preset intensity, if so, the audio acquisition gain of the second microphone is adjusted; if not, a video file is generated based on the collected video and audio information, thus ensuring the recording effect of the video file.
[0144] Based on the above embodiments, the following specific example will be used to illustrate the overall process of optimizing recording effects.
[0145] Figure 7 This is a schematic diagram illustrating the process of optimizing recording effects using a recording and broadcasting device, as provided in an embodiment of this application. Figure 7 As shown, the process includes the following steps:
[0146] S701: After the controller controls the pan-tilt unit to rotate according to the target angle value, the rotated lens will find the target image corresponding to the sound source and adjust the lens magnification.
[0147] S702: The controller adjusts the acquisition angle of the second microphone, enables directional pickup, and filters noise from non-target angles.
[0148] S703: The second pickup collects audio information.
[0149] S704: The controller acquires the audio information of the sound source collected by the second microphone; and determines whether the sound intensity of the audio information is less than the preset intensity.
[0150] S7045a: If the controller determines that the sound intensity of the audio information is less than the preset intensity, it adjusts the audio acquisition gain of the second microphone and executes S703.
[0151] S705b: If the controller determines that the sound intensity of the audio information is not less than the preset intensity, it will generate a recording file based on the collected video and audio information.
[0152] Based on the above explanation, taking a meeting scenario as an example, the common states of recording and broadcasting equipment are as follows:
[0153] (a) There was no obvious human sound in the environment. Figure 8a This application provides a schematic diagram of a meeting scenario in a silent state, as illustrated in the embodiments of this application. Figure 8a As shown, the recording equipment remains stationary. Alternatively, it may be performing other tasks, such as patrolling.
[0154] (b) Speeches by attendees, Figure 8b This application provides a schematic diagram of a meeting scenario where participants are speaking, as shown in the embodiments of this application. Figure 8b As shown, when the microphone of the recording equipment detects sound, it pauses the currently running service and collects audio information;
[0155] (c) The controller determines the target position of the sound source based on the audio information collected by each microphone; determines the target angle value of the pan-tilt unit to rotate based on the target position of the sound source; and controls the pan-tilt unit to rotate based on the target angle value. Figure 8c This is a schematic diagram of the rotation of a recording and broadcasting device provided in an embodiment of this application.
[0156] (d) Figure 8d This is a schematic diagram of a recording and broadcasting device provided in an embodiment of this application. Figure 8d As shown, the controller rotates the pan-tilt unit according to the target angle value, causing the rotated lens to locate the target image corresponding to the sound source and adjust the lens magnification. It also controls and adjusts the microphone's pickup angle, activating directional sound pickup and filtering noise from non-target angles until the recording effect reaches the desired level.
[0157] (e) After the participants finish speaking and the sound disappears for a few seconds, the system returns to state (a) and the recording equipment continues to perform its original function.
[0158] Example 6:
[0159] In order to acquire video information from multiple sound sources, based on the above embodiments, in this embodiment of the application, the recording and broadcasting device further includes: at least three wide-angle lenses, and the at least three wide-angle lenses are evenly arranged under the base;
[0160] The controller is further configured to, if it acquires sub-audio information of other sound sources collected by each of the first microphones, determine other locations of the other sound sources, and, based on the determined other locations of the other sound sources and the panoramic video information acquired by the at least three wide-angle lenses, determine the image position of the other sound sources in the panoramic video information, and respectively extract local video information corresponding to each image position.
[0161] For applications involving multiple people communicating, i.e., multiple sound sources, at least three wide-angle lenses can be installed under the base of the recording equipment. Figure 9 A side view of the structure of another recording and broadcasting device provided in an embodiment of this application. (In conjunction with...) Figure 2a and Figure 9 It can be seen that, Figure 9 The recording equipment shown is in Figure 2a Based on the recording and broadcasting equipment shown, three evenly distributed wide-angle lenses 106 were added under the base to collect panoramic video information of the application scene.
[0162] In this embodiment, the first sound source detected by the microphone can be identified as the primary sound source. Based on the target position of the primary sound source, the target angle value for the pan-tilt unit to rotate is determined. The pan-tilt unit is then controlled to rotate according to the target angle value, so that the rotated lens and the second microphone can capture video and audio information of the primary sound source. If sub-audio information of other sound sources captured by each first microphone is obtained, the other positions of the other sound sources are determined. Specifically, the other positions of the other sound sources can be determined using the method described in the above embodiment.
[0163] Based on the identified locations of other sound sources and panoramic video information captured by at least three wide-angle lenses, determine the position of each sound source within the panoramic video image. Extract local video information corresponding to each position. This local video information can be magnified to obtain close-up shots of the other sound sources.
[0164] In order to generate a video file containing multiple sound sources, based on the above embodiments, in this embodiment of the application, the controller is further configured to generate a video file according to the local video information, the sub-audio information, the main audio information of the sound source collected by the second microphone, and the main video information of the sound source collected by the lens.
[0165] In this embodiment of the application, the partial video information of each other sound source extracted from the panoramic video information can be stitched together with the main video information of the main sound source, i.e., the first sound source detected by the microphone, and the sub-audio information of each other sound source can be stitched together with the main audio information of the main sound source in chronological order to generate a video file.
[0166] Based on the above embodiments, the process of a recording device generating video files will be described below with a specific example.
[0167] Figure 10 This is a schematic diagram of a multi-sound-source recording scene and video file provided in an embodiment of this application. Figure 10 As shown, the panoramic view contains sub-views of multiple people. The partial view of person A corresponding to other sound sources is extracted and stitched together with the view of person B corresponding to the main sound source to obtain the output view and generate a video file.
[0168] Figure 11 This is a schematic diagram illustrating the process of generating a video file, as provided in an embodiment of this application. Figure 11 As shown, the process includes the following steps:
[0169] S1101: After the controller controls the pan-tilt unit to rotate according to the target angle value, the rotated lens and the second microphone will collect the main video and main audio information of the main sound source.
[0170] S1102: The first pickup detects and determines whether there are other sound sources.
[0171] If not, proceed to step S1105. If yes, collect sub-audio information from other sound sources and proceed to step S1103.
[0172] S1103: When the controller obtains the sub-audio information of other sound sources collected by each of the first microphones, it determines the other positions of the other sound sources. Based on the determined other positions of the other sound sources and the panoramic video information collected by the at least three wide-angle lenses, it determines the position of the other sound sources in the panoramic video information and extracts the local video information corresponding to each position.
[0173] S1104: The controller splices together the local video information, sub-audio information, the main audio information of the main sound source collected by the second microphone, and the main video information of the main sound source collected by the lens.
[0174] S1105: Generate video file.
[0175] Example 7:
[0176] In order to collect video information from multiple sound sources, based on the above embodiments, in this embodiment of the application, the recording and broadcasting device further includes:
[0177] At least one auxiliary lens located elsewhere;
[0178] The controller is further configured to determine the target auxiliary lens corresponding to the shooting range of the target position of the sound source based on the determined target position of the sound source, the pre-saved position of each auxiliary lens and the shooting range, determine the third distance, horizontal offset angle and vertical offset angle between the target auxiliary lens and the sound source; control the target auxiliary lens to focus based on the third distance, horizontal offset angle and vertical offset angle, and collect the sub-video information of the sound source.
[0179] For multi-sound-source applications, in order to accurately capture video information from each sound source, the recording and broadcasting equipment can also include at least one auxiliary lens positioned elsewhere. Based on the determined target location of the sound source and the pre-saved shooting range of each auxiliary lens, the target auxiliary lens corresponding to the shooting range of the sound source's target location is determined. Furthermore, based on the target location of the sound source and the position of the target auxiliary lens, a third distance, a horizontal offset angle, and a vertical offset angle between the sound source and the target auxiliary location are determined. Based on the third distance, the horizontal offset angle, and the vertical offset angle, the image position of the sound source within the shooting frame of the auxiliary lens is determined, thereby enabling the target auxiliary lens to focus on the image position and capture sub-video information of the sound source.
[0180] The following example illustrates the process of controlling the auxiliary lens to acquire sub-video information of the sound source in the embodiments of this application.
[0181] Figure 12 This is a schematic diagram illustrating another multi-sound-source recording and broadcasting scenario and video information provided in an embodiment of this application. For example... Figure 12As shown, the recording equipment F is installed directly above the circular conference table, and the auxiliary lens E is installed on the right side of the circular conference table (left and right in the figure), with the shooting range covering the area between angles α and γ in the horizontal plane. Based on the determined target position of the sound source G, and the pre-saved position and shooting range of the auxiliary lens E, the target auxiliary lens corresponding to the shooting range of the target position of the sound source G is determined, i.e., the auxiliary lens E. A rectangular coordinate system is established with the position of the recording equipment F as the origin (0, 0, 0), and the coordinates of the target position of the sound source G are determined as (x, y, z), and the coordinates of the pre-saved position of the auxiliary lens E are (x, y, z). E y E , z E The recording camera F, the sound source G, and the auxiliary lens E satisfy the following equation (7):
[0182]
[0183] in, The coordinates (x) of auxiliary lens E E y E , z E The vector determined by the coordinates (0, 0, 0) of the recording device F and the recording equipment F is specifically as follows: The vector that determines the coordinates (0, 0, 0) of the recording device F and the coordinates (x, y, z) of the sound source G, specifically: The coordinates (x) of auxiliary lens E E y E , z E The vector determined by the coordinates (x, y, z) of the sound source G and the sound source G is specifically...
[0184] From equation (7), we can obtain the following equation (8):
[0185]
[0186] Where, r, θ, Let x, y, and z be the third distance, horizontal offset angle, and vertical offset angle between the sound source G and the auxiliary lens E. Let x, y, and z be the coordinates of the target position of the sound source G. E y E , z E The coordinates are for the position of the auxiliary lens E.
[0187] The position of the sound source in the frame of the auxiliary lens can be determined by the horizontal and vertical offset angles. For example, an 8192 coordinate system can be used to determine the horizontal and vertical coordinates of the sound source's position in the frame of the auxiliary lens. The following explanation uses the determination of the horizontal coordinates of the position based on the horizontal offset angle as an example.
[0188] According to the following formula (9), the horizontal coordinates of the position of the sound source G in the shooting frame of the auxiliary lens E can be determined:
[0189] x 8192 =8192·θ / (γ-α)(9)
[0190] Where, x 8192 Let be the horizontal coordinate of the sound source's position in the image captured by the auxiliary lens E, and let α and γ be the minimum and maximum shooting angles of the auxiliary lens E in the horizontal plane. θ is the horizontal offset angle between the sound source G and the auxiliary lens E.
[0191] Similarly, the vertical coordinate of the sound source's position in the frame captured by the auxiliary lens E, y 8192 This can be achieved by shifting the vertical angle between the sound source G and the auxiliary lens E. The minimum and maximum values of the shooting angle of the auxiliary lens E in the vertical plane are obtained, and will not be elaborated here.
[0192] like Figure 12 As shown, the sub-video information of sound source G obtained by auxiliary lens E, corresponding to the image of person B, is spliced with the main video information of person A obtained by the lens set below the recording device F to obtain the output image.
[0193] Based on the above embodiments, a recording and broadcasting system can also be built using the above recording and broadcasting equipment and other cameras. Figure 13 This is a schematic diagram of a recording and broadcasting system provided in an embodiment of this application.
[0194] like Figure 13 As shown, the system includes recording equipment, camera 1, camera 2... camera n.
[0195] The recording equipment can send the main audio and main video information of person A corresponding to the main sound source (i.e., the audio stream and video stream of person A) to the encoder of the platform, send the secondary audio information of person B corresponding to other sound sources (i.e., the audio stream of person B) to the encoder, and send the coordinates of the determined location of person B to the device management center of the platform.
[0196] Based on the coordinates of Person B's location, the pre-saved positions of each camera, and the shooting range, the equipment management center determines the target camera (Camera 2) corresponding to the shooting range of Person B's location, and then sends the coordinates of Person B's location to Camera 2.
[0197] Camera 2 focuses on person B based on the received coordinates of person B's location and acquires the secondary video information of person B, i.e., the video stream of person B, and sends the acquired secondary video information of person B to the encoder.
[0198] The encoder encodes the main audio and video information of person A and the secondary video and audio information of person B, and then splices them together to output the video.
[0199] Figure 14 This is a schematic diagram illustrating the recording process of a recording and broadcasting system provided in an embodiment of this application. Figure 14 As shown, the process includes the following steps:
[0200] S1401: After the controller controls the pan-tilt unit to rotate according to the target angle value, the rotated lens and the second microphone will collect the main video and main audio information of the main sound source.
[0201] S1402: The first pickup collects audio information from other directions to determine whether there are other sound sources.
[0202] If not, execute S1407 directly; if yes, execute S1403.
[0203] S1403: The controller reports the main audio and main video information of the main sound source, as well as the secondary audio information of other sound sources, to the platform, and sends the coordinates of the determined locations of other sound sources to the platform.
[0204] S1404: Based on the coordinates of the received sub-audio location, the pre-saved positions and shooting ranges of each camera, the platform determines the target camera corresponding to the shooting range of the coordinates of other sound sources.
[0205] S1405: The platform calls the target camera, focuses on the target area, and collects secondary video information from other sound sources.
[0206] S1406: The platform encodes the received main audio and main video information, secondary video and secondary audio information, and completes the image stitching.
[0207] S1407: Generate video file.
[0208] Example 8:
[0209] Based on the same technical concept, this application provides a recording method applied to a recording device, the method comprising:
[0210] Based on the audio information collected by each microphone, the angular deviation values between the sound source and multiple preset positions are determined, wherein the preset positions are the midpoints of the lines connecting any two microphones; based on each angular deviation value, the distance between any two microphones that are stored in advance, and the first distance between the preset sound source and the preset positions, the target position of the sound source is determined; based on the target position of the sound source, the target angle value that the pan-tilt unit needs to rotate is determined; based on the target angle value, the pan-tilt unit is controlled to rotate so that the rotated lens and the second microphone can collect the video and audio information of the sound source.
[0211] In one possible implementation, determining the angular deviation value between the sound source and multiple preset positions based on the audio information collected by each microphone includes:
[0212] For any given angular deviation value, based on the sound intensity of the sound source collected by two microphones connected at a preset position that determines the angular deviation value, the superposition value and average value of the sound intensity collected by the two microphones are determined, and the angular deviation value is determined based on the superposition value and average value.
[0213] In one possible implementation, determining the target angle value that the pan-tilt unit needs to rotate based on the target location of the sound source includes:
[0214] Based on the target location of the sound source, determine the horizontal and vertical angles between the sound source and the gimbal, and then determine the horizontal and vertical angles as the target angle values that the gimbal needs to rotate.
[0215] In one possible implementation, after controlling the gimbal rotation according to the target angle value, the method further includes:
[0216] Determine whether the sound intensity in the audio information of the sound source collected by the second microphone is less than a preset intensity; if so, update the first distance using a preset second distance, and redetermine the target position of the sound source based on the updated first distance.
[0217] In one possible implementation, after controlling the pan-tilt unit to rotate according to the target angle value so that the rotated lens and second microphone can acquire video and audio information of the sound source, the method further includes:
[0218] The system acquires the audio information of the sound source collected by the second microphone; it determines whether the sound intensity of the audio information is less than a preset intensity; if not, it generates a video file based on the acquired video and audio information.
[0219] In one possible implementation, if it is determined that the sound intensity of the audio information is less than a preset intensity, the audio acquisition gain of the second microphone is adjusted until the sound intensity of the audio information acquired by the second microphone is not less than the preset intensity.
[0220] In one possible implementation, the method further includes:
[0221] If the sub-audio information of other sound sources collected by each of the first microphones is obtained, the other positions of the other sound sources are determined. Based on the determined other positions of the other sound sources and the panoramic video information collected by the at least three wide-angle lenses, the position of the other sound sources in the panoramic video information is determined, and the local video information corresponding to each position is extracted.
[0222] In one possible implementation, after extracting the local video information corresponding to each frame position, the method further includes:
[0223] A video file is generated based on the local video information, the secondary audio information, the main audio information of the sound source collected by the second microphone, and the main video information of the sound source collected by the lens.
[0224] In one possible implementation, the method further includes:
[0225] Based on the determined target location of the sound source, the pre-saved position and shooting range of each auxiliary lens, the target auxiliary lens corresponding to the shooting range of the target location of the sound source is determined, and the third distance, horizontal offset angle and vertical offset angle between the target auxiliary lens and the sound source are determined; the target auxiliary lens is controlled to focus according to the third distance, horizontal offset angle and vertical offset angle, and the sub-video information of the sound source is collected.
[0226] Example 9:
[0227] Based on the same technical concept, this application also provides a recording and playback device. Figure 15 This is a schematic diagram of a recording and broadcasting device provided in an embodiment of this application. Figure 15 As shown, the device includes:
[0228] The determining module 151 is used to determine the angle deviation value between the sound source and multiple preset positions based on the audio information collected by each microphone, wherein the preset position is the midpoint of the line connecting any two microphones; determine the target position of the sound source based on each angle deviation value, the distance between any two microphones that are stored in advance, and the first distance between the preset sound source and the preset position; and determine the target angle value that the pan-tilt unit needs to rotate based on the target position of the sound source.
[0229] Control module 152 is used to control the pan-tilt unit to rotate according to the target angle value, so that the rotated lens and the second microphone can collect video and audio information of the sound source.
[0230] In one possible implementation, the determining module 151 is specifically used to determine, for any angle deviation value, the superposition value and average value of the sound intensity collected by the two microphones connected at a preset position to determine the angle deviation value, based on the sound intensity of the sound source collected by the two microphones. The angle deviation value is then determined based on the superposition value and the average value.
[0231] In one possible implementation, the determining module 151 is specifically used to determine the horizontal and vertical angles between the sound source and the gimbal based on the target position of the sound source, and to determine the horizontal and vertical angles as the target angle values that the gimbal needs to rotate.
[0232] In one possible implementation, the determining module 151 is further configured to determine whether the sound intensity in the audio information of the sound source collected by the second microphone is less than a preset intensity; if so, the first distance is updated using a preset second distance, and the target position of the sound source is re-determined based on the updated first distance.
[0233] In one possible implementation, the device further includes:
[0234] The first generation module is used to acquire the audio information of the sound source collected by the second microphone; determine whether the sound intensity of the audio information is less than a preset intensity; if not, generate a video file based on the acquired video information and audio information.
[0235] In one possible implementation, the determining module 151 is further configured to adjust the audio acquisition gain of the second microphone if it is determined that the sound intensity of the audio information is less than a preset intensity, until the sound intensity of the audio information acquired by the second microphone is not less than the preset intensity.
[0236] In one possible implementation, the determining module 151 is further configured to, if the sub-audio information of other sound sources collected by each of the first microphones is obtained, determine the other positions of the other sound sources, and, based on the determined other positions of the other sound sources and the panoramic video information collected by the at least three wide-angle lenses, determine the image position of the other sound sources in the panoramic video information, and respectively extract the local video information corresponding to each image position.
[0237] In one possible implementation, the device further includes:
[0238] The second generation module is used to generate a video file based on the local video information, the sub-audio information, the main audio information of the sound source collected by the second microphone, and the main video information of the sound source collected by the lens.
[0239] In one possible implementation, the determining module 151 is further configured to determine the target auxiliary lens corresponding to the shooting range where the target position of the sound source is located, based on the determined target position of the sound source, the position of each auxiliary lens and the shooting range that are saved in advance, and to determine the third distance, horizontal offset angle and vertical offset angle between the target auxiliary lens and the sound source.
[0240] The control module 152 is also used to control the target auxiliary lens to focus according to the third distance, the horizontal offset angle and the vertical offset angle, and to collect the sub-video information of the sound source.
[0241] Based on the same technical concept, embodiments of this application also provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of any of the recording and broadcasting methods described above.
[0242] In this embodiment, the recording and broadcasting device includes: a controller, a base, a pan-tilt unit, a lens, three first microphones, and one second microphone. The pan-tilt unit and the three first microphones are all mounted under the base, and the three first microphones are arranged in a triangular pattern. The lens is mounted at the bottom of the pan-tilt unit, and the second microphone is mounted at the bottom of the lens. This allows for the acquisition of audio information from multiple directions. The controller is used to determine the angular deviation value between the sound source and multiple preset positions based on the audio information acquired by each microphone. Based on each angular deviation value, the pre-saved distance between any two microphones, and a preset first distance between the sound source and a preset position, the controller can determine the target angle value for the pan-tilt unit to rotate. The controller controls the pan-tilt unit to rotate according to the target angle value, so that the rotated lens and second microphone acquire video and audio information of the sound source. This allows for application in various recording and broadcasting scenarios and enables accurate switching of perspectives to record and broadcast the target person from the sound source.
[0243] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0244] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions specified in one or more blocks of the flowchart illustrations and / or one or more blocks of the block diagrams.
[0245] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means that implement the functions specified in one or more flowcharts and / or one or more block diagrams.
[0246] These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, such that the instructions, which execute on the computer or other programmable apparatus, provide steps for implementing the functions specified in one or more flowcharts and / or one or more block diagrams.
[0247] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.
Claims
1. A recording and broadcasting device, characterized in that, The recording and broadcasting equipment includes: a controller, a base, a pan-tilt unit, a lens, three first microphones, and one second microphone; wherein the pan-tilt unit and the three first microphones are all mounted under the base, and the three first microphones are arranged in a triangular pattern. The lens is mounted on the bottom of the gimbal, and the second microphone is mounted on the bottom of the lens; The controller is configured to: determine the angular deviation value between the sound source and multiple preset positions based on the audio information collected by each microphone, wherein the preset position is the midpoint of the line connecting any two microphones; determine the target position of the sound source based on each angular deviation value, the pre-saved distance between any two microphones, and a preset first distance between the sound source and the preset position; determine the target angle value for the pan-tilt unit to rotate based on the target position of the sound source; and control the pan-tilt unit to rotate based on the target angle value, so that the rotated lens and the second microphone collect video and audio information of the sound source. The controller is further configured to, after controlling the pan-tilt unit to rotate according to the target angle value, determine whether the sound intensity in the audio information of the sound source collected by the second microphone is less than a preset intensity; if so, update the first distance with a preset second distance, and redetermine the target position of the sound source based on the updated first distance.
2. The recording and broadcasting equipment according to claim 1, characterized in that, The controller is specifically used to determine the superposition and average value of the sound intensities collected by the two microphones connected at a preset position to determine the angle deviation value for any given angle deviation value, based on the sound intensity of the sound source collected by the two microphones. The controller then determines the angle deviation value based on the superposition and average value.
3. The recording and broadcasting equipment according to claim 1, characterized in that, The controller is specifically used to determine the horizontal and vertical angles between the sound source and the gimbal based on the target position of the sound source, and to determine the horizontal and vertical angles as the target angle values that the gimbal needs to rotate.
4. The recording and broadcasting equipment according to claim 1, characterized in that, The controller is specifically used to acquire the audio information of the sound source collected by the second microphone; determine whether the sound intensity of the audio information is less than a preset intensity; if not, generate a video file based on the acquired video information and audio information.
5. The recording and broadcasting equipment according to claim 4, characterized in that, The controller is specifically configured to adjust the audio acquisition gain of the second microphone if it is determined that the sound intensity of the audio information is less than a preset intensity, until the sound intensity of the audio information acquired by the second microphone is not less than the preset intensity.
6. The recording and broadcasting equipment according to claim 1, characterized in that, The recording and broadcasting equipment also includes at least three wide-angle lenses, and the at least three wide-angle lenses are evenly arranged under the base; The controller is further configured to, if it acquires sub-audio information of other sound sources collected by each of the first microphones, determine other locations of the other sound sources, and, based on the determined other locations of the other sound sources and the panoramic video information acquired by the at least three wide-angle lenses, determine the image position of the other sound sources in the panoramic video information, and respectively extract local video information corresponding to each image position.
7. The recording and broadcasting equipment according to claim 6, characterized in that, The controller is further configured to generate a video file based on the local video information, the sub-audio information, the main audio information of the sound source acquired by the second microphone, and the main video information of the sound source acquired by the lens.
8. The recording and broadcasting equipment according to claim 1, characterized in that, The recording and broadcasting equipment also includes: At least one auxiliary lens located elsewhere; The controller is further configured to determine the target auxiliary lens corresponding to the shooting range of the target position of the sound source based on the determined target position of the sound source, the pre-saved position of each auxiliary lens and the shooting range, determine the third distance, horizontal offset angle and vertical offset angle between the target auxiliary lens and the sound source; control the target auxiliary lens to focus based on the third distance, horizontal offset angle and vertical offset angle, and collect the sub-video information of the sound source.
9. A recording and broadcasting method, characterized in that, The method is applied to a recording and broadcasting equipment, which includes a base, a pan-tilt unit, a lens, three first microphones and one second microphone; the three first microphones are all located under the base and are arranged in a triangular pattern. The lens is mounted on the bottom of the gimbal, and the second microphone is mounted on the bottom of the lens; the method includes: Based on the audio information collected by each microphone, the angular deviation value between the sound source and multiple preset positions is determined, wherein the preset position is the midpoint of the line connecting any two microphones; Based on each of the aforementioned angle deviation values, the pre-saved distance between any two microphones, and the first distance between the preset sound source and the preset position, the target position of the sound source is determined; based on the target position of the sound source, the target angle value that the pan-tilt unit needs to rotate is determined; based on the target angle value, the pan-tilt unit is controlled to rotate so that the rotated lens and the second microphone can collect video and audio information of the sound source; After controlling the pan-tilt unit to rotate according to the target angle value, it is determined whether the sound intensity in the audio information of the sound source collected by the second microphone is less than a preset intensity; if so, the first distance is updated using a preset second distance, and the target position of the sound source is re-determined based on the updated first distance.
10. The method according to claim 9, characterized in that, The step of determining the angular deviation value between the sound source and multiple preset positions based on the audio information collected by each microphone includes: For any given angular deviation value, based on the sound intensity of the sound source collected by two microphones connected at a preset position that determines the angular deviation value, the superposition value and average value of the sound intensity collected by the two microphones are determined, and the angular deviation value is determined based on the superposition value and average value.
11. The method according to claim 9, characterized in that, Determining the target angle value that the pan-tilt unit needs to rotate based on the target location of the sound source includes: Based on the target location of the sound source, determine the horizontal and vertical angles between the sound source and the gimbal, and then determine the horizontal and vertical angles as the target angle values that the gimbal needs to rotate.
12. The method according to claim 9, characterized in that, After controlling the pan-tilt unit to rotate according to the target angle value so that the rotated lens and second microphone can acquire video and audio information of the sound source, the method further includes: The system acquires the audio information of the sound source collected by the second microphone; it determines whether the sound intensity of the audio information is less than a preset intensity; if not, it generates a video file based on the acquired video and audio information.
13. The method according to claim 12, characterized in that, If it is determined that the sound intensity of the audio information is less than the preset intensity, the audio acquisition gain of the second microphone is adjusted until the sound intensity of the audio information acquired by the second microphone is not less than the preset intensity.
14. The method according to claim 9, characterized in that, The recording and broadcasting equipment further includes: at least three wide-angle lenses, and the at least three wide-angle lenses are evenly arranged under the base; the method further includes: If the sub-audio information of other sound sources collected by each of the first microphones is obtained, the other positions of the other sound sources are determined. Based on the determined other positions of the other sound sources and the panoramic video information collected by the at least three wide-angle lenses, the position of the other sound sources in the panoramic video information is determined, and the local video information corresponding to each position is extracted.
15. The method according to claim 14, characterized in that, After extracting the local video information corresponding to each frame position, the method further includes: A video file is generated based on the local video information, the secondary audio information, the main audio information of the sound source collected by the second microphone, and the main video information of the sound source collected by the lens.
16. The method according to claim 9, characterized in that, The method further includes: Based on the determined target location of the sound source, the pre-saved position and shooting range of each auxiliary lens, the target auxiliary lens corresponding to the shooting range of the target location of the sound source is determined, and the third distance, horizontal offset angle and vertical offset angle between the target auxiliary lens and the sound source are determined; the target auxiliary lens is controlled to focus according to the third distance, horizontal offset angle and vertical offset angle, and the sub-video information of the sound source is collected.
17. A recording and playback device, characterized in that, The device includes: The determination module is used to determine the angular deviation value between the sound source and multiple preset positions based on the audio information collected by each microphone, wherein the preset position is the midpoint of the line connecting any two microphones; determine the target position of the sound source based on each angular deviation value, the pre-saved distance between any two microphones, and a preset first distance between the sound source and the preset position; and determine the target angle value for the pan-tilt unit to rotate based on the target position of the sound source; wherein the microphones include three first microphones and one second microphone; the three first microphones are all disposed under the base and are arranged in a triangle; the second microphone is installed at the bottom of the lens, and the lens is installed at the bottom of the pan-tilt unit; The control module is used to control the pan-tilt unit to rotate according to the target angle value, so that the rotated lens and the second microphone can collect video and audio information of the sound source. The determining module is further configured to determine whether the sound intensity in the audio information of the sound source collected by the second microphone is less than a preset intensity; if so, the first distance is updated using a preset second distance, and the target position of the sound source is re-determined based on the updated first distance.
18. The apparatus according to claim 17, characterized in that, The determining module is specifically used to determine the superposition value and average value of the sound intensity collected by the two microphones connected at the preset position of the determined angle deviation value for any angle deviation value, based on the sound intensity of the sound source collected by the two microphones. The angle deviation value is then determined based on the superposition value and the average value.
19. The apparatus according to claim 17, characterized in that, The determining module is specifically used to determine the horizontal and vertical angles between the sound source and the pan-tilt unit based on the target position of the sound source, and to determine the horizontal and vertical angles as the target angle values that the pan-tilt unit needs to rotate.
20. The apparatus according to claim 17, characterized in that, The device further includes: The first generation module is used to acquire the audio information of the sound source collected by the second microphone; determine whether the sound intensity of the audio information is less than a preset intensity; if not, generate a video file based on the acquired video information and audio information.
21. The apparatus according to claim 20, characterized in that, The determining module is further configured to, if it is determined that the sound intensity of the audio information is less than a preset intensity, adjust the audio acquisition gain of the second microphone until the sound intensity of the audio information acquired by the second microphone is not less than the preset intensity.
22. The apparatus according to claim 17, characterized in that, The determining module is further configured to, if it acquires sub-audio information of other sound sources collected by each of the first microphones, determine other locations of the other sound sources, and, based on the determined other locations of the other sound sources and the panoramic video information acquired by at least three wide-angle lenses, determine the image position of the other sound sources in the panoramic video information, and respectively extract local video information corresponding to each image position; wherein, the at least three wide-angle lenses are evenly arranged under the base.
23. The apparatus according to claim 22, characterized in that, The device further includes: The second generation module is used to generate a video file based on the local video information, the sub-audio information, the main audio information of the sound source collected by the second microphone, and the main video information of the sound source collected by the lens.
24. The apparatus according to claim 17, characterized in that, The determining module is further configured to determine the target auxiliary lens corresponding to the shooting range of the target position of the sound source based on the determined target position of the sound source, the pre-saved position of each auxiliary lens and the shooting range, and to determine the third distance, horizontal offset angle and vertical offset angle between the target auxiliary lens and the sound source. The control module is also used to control the target auxiliary lens to focus based on the third distance, horizontal offset angle and vertical offset angle, and to collect sub-video information of the sound source.
25. A computer-readable storage medium, characterized in that, It stores a computer program that, when executed by a processor, implements the steps of the recording method as described in any one of claims 9-16.
Citation Information
Patent Citations
Panoramic video recording method and device based on voice tracking
CN111163281A
Method for automatically capturing and tracking speaker
CN113163148A