A camera control method, system and apparatus
Patent Information
- Application Number
- CN202310628234.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-30
- Publication Date
- 2026-10-09
- Estimated Expiration
- 2043-05-30
AI Technical Summary
尤其面对多路口的监控场景和用户多样化的监控需求,普通的监控设备显然无法满足需求,无法兼顾多场景全面监控
Smart Images

Figure CN116567420B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of intelligent monitoring technology, and in particular to a camera control method, system and device. Background Technology
[0002] As the security market continues to expand, the demand for surveillance equipment is also increasing. Especially in multi-intersection monitoring scenarios and with diverse user needs, ordinary surveillance equipment is clearly insufficient to meet these demands and cannot provide comprehensive monitoring across multiple scenarios. To address this issue, current solutions often employ multiple cameras working in tandem, such as using binocular surveillance equipment. One camera monitors different directions within a scene, while the other assists by responding to alarms in real time and capturing corresponding footage. However, this increases the overall cost. Therefore, it is necessary to address the problem of ordinary surveillance equipment's inability to respond in real time to multi-directional monitoring needs in multi-intersection scenarios. Summary of the Invention
[0003] This application provides a camera control method, system, and device to enable comprehensive environmental control of cameras without the assistance of other cameras, thereby improving the adaptability of camera control.
[0004] This application provides a camera control method, including:
[0005] Based on the sound signal of the external target sound source received by the microphone array on the camera, determine the audio coordinates of the target sound source in the audio coordinate system of the microphone array and the sound type of the target sound source;
[0006] When the location of the audio coordinates is within the pre-set monitoring area for the camera, and the sound type is consistent with the preset sound type, the audio coordinates of the target sound source are converted into pan-tilt coordinates in the pan-tilt coordinate system according to the pre-determined mapping relationship between the audio coordinate system of the microphone array on the camera and the pan-tilt coordinate system of the camera. The parameters that the camera needs to adjust are determined according to the pan-tilt coordinates, and the camera is rotated to the corresponding position according to the parameters.
[0007] This method determines the audio coordinates and sound type of an external target sound source in the audio coordinate system of the microphone array, based on the sound signal received by the microphone array on the camera. When the location of the audio coordinates is within a pre-set monitoring area for the camera, and the sound type matches a preset sound type, the audio coordinates of the target sound source are converted into pan-tilt coordinates in the pan-tilt coordinate system according to the pre-determined mapping relationship between the audio coordinate system of the microphone array on the camera and the pan-tilt coordinate system of the camera. Based on the pan-tilt coordinates, the parameters that the camera needs to adjust are determined, and the camera is rotated to the corresponding position according to the parameters. This enables the camera to achieve comprehensive environmental control without the assistance of other cameras, improving the adaptability of camera deployment.
[0008] In some embodiments, the predetermined mapping relationship between the audio coordinate system of the microphone array on the camera and the pan-tilt coordinate system of the camera is specifically determined by the following methods:
[0009] The microphone array receives audio signals of a preset fixed frequency emitted by a sound-emitting device mounted on the horizontally rotating camera.
[0010] Calculate the audio energy of the audio signal received by each microphone in the microphone array, and record the pan-tilt coordinates of the camera when the audio energy of the audio signal received by each microphone is at its maximum.
[0011] By using the audio coordinates of the audio signal received by each microphone at its maximum audio energy and the corresponding pan-tilt coordinates of the camera, a mapping relationship between the audio coordinate system and the pan-tilt coordinate system is established.
[0012] This method establishes a mapping relationship between the audio coordinate system of the microphone array and the pan-tilt coordinate system of the camera, facilitating subsequent coordinate transformations using this mapping relationship.
[0013] In some embodiments, the method further includes:
[0014] The audio signal received by each microphone in the microphone array is subjected to frequency filtering.
[0015] This method filters out audio signals of other frequencies from the audio signal received by each microphone, thereby improving the accuracy of subsequent audio energy calculations.
[0016] In some embodiments, determining the audio coordinates of the target sound source in the audio coordinate system and the sound type of the target sound source based on the sound signal of the external target sound source received by the microphone array specifically includes:
[0017] The sound signal of the target sound source received by the microphone array is input into a multi-class neural network, which outputs the audio coordinates of the target sound source in the audio coordinate system and the sound type of the target sound source.
[0018] The multi-class neural network is obtained by dividing pre-collected audio samples from different directions into a preset number of directional categories, using the audio samples from the preset number of directional categories as training samples to train the convolutional neural network, and then using different target sound sources to replace the sound types in the training samples to train the trained convolutional neural network.
[0019] This method enables the determination of the direction and sound type of a target sound source using a multi-class neural network, thereby improving the accuracy of target sound source localization.
[0020] This application provides a camera control system, including:
[0021] The sound source localization module is used to determine the audio coordinates of the target sound source in the audio coordinate system of the microphone array and the sound type of the target sound source based on the sound signal of the external target sound source received by the microphone array on the camera.
[0022] The camera adjustment module is used to convert the audio coordinates of the target sound source into pan-tilt coordinates in the pan-tilt coordinate system according to the pre-determined mapping relationship between the audio coordinate system of the microphone array on the camera and the pan-tilt coordinate system of the camera, when the location of the audio coordinates is within the pre-set monitoring area for the camera and the sound type is consistent with the preset sound type. Based on the pan-tilt coordinates, the module determines the parameters that the camera needs to be adjusted, and rotates the camera to the corresponding position according to the parameters.
[0023] This system enables cameras to achieve comprehensive and coordinated environmental control without the assistance of other cameras, thus improving the adaptability of camera deployment.
[0024] In some embodiments, the system further includes a coordinate calibration module:
[0025] This is used to establish the mapping relationship between the audio coordinate system of the microphone array on the camera and the pan-tilt coordinate system of the camera.
[0026] This system enables the establishment of a mapping relationship between the audio coordinate system and the pan-tilt coordinate system through a coordinate calibration module.
[0027] In some embodiments, the coordinate calibration module is specifically used for:
[0028] The microphone array receives audio signals of a preset fixed frequency emitted by a sound-emitting device mounted on the horizontally rotating camera.
[0029] Calculate the audio energy of the audio signal received by each microphone in the microphone array, and record the pan-tilt coordinates of the camera when the audio energy of the audio signal received by each microphone is at its maximum.
[0030] By using the audio coordinates of the audio signal received by each microphone at its maximum audio energy and the corresponding pan-tilt coordinates of the camera, a mapping relationship between the audio coordinate system and the pan-tilt coordinate system is established.
[0031] This system establishes a mapping relationship between the audio coordinate system of the microphone array and the pan-tilt coordinate system of the camera, facilitating subsequent coordinate transformations using this mapping relationship.
[0032] In some embodiments, the coordinate calibration module is further configured to:
[0033] The audio signal received by each microphone in the microphone array is subjected to frequency filtering.
[0034] This system filters out other frequencies of audio signals from the audio signal received by each microphone, thereby improving the accuracy of subsequent audio energy calculations.
[0035] Another embodiment of this application provides a camera control device, which includes a memory and a processor, wherein the memory is used to store program instructions, and the processor is used to call the program instructions stored in the memory and execute any of the methods described above according to the obtained program.
[0036] Furthermore, according to embodiments, for example, a computer program product for a computer is provided, which includes software code portions that, when the product is run on the computer, perform the steps of the methods defined above. The computer program product may include a computer-readable medium on which the software code portions are stored. Furthermore, the computer program product may be directly loaded into the computer's internal memory and / or sent via a network through at least one of an upload process, a download process, and a push process.
[0037] Another embodiment of this application provides a computer-readable storage medium storing computer-executable instructions for causing the computer to perform any of the methods described above. Attached Figure Description
[0038] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0039] Figure 1 This is a schematic diagram of the overall process of a camera control method provided in an embodiment of this application;
[0040] Figure 2 A schematic diagram of a direction category in an audio coordinate system provided for an embodiment of this application;
[0041] Figure 3 A schematic diagram of a region of interest for a camera provided in an embodiment of this application;
[0042] Figure 4 This is a schematic diagram illustrating the overall process of establishing a mapping relationship between an audio coordinate system and a gimbal coordinate system, provided in an embodiment of this application.
[0043] Figure 5 This is a schematic diagram of the structure of a camera provided in an embodiment of this application;
[0044] Figure 6 This is a schematic diagram illustrating the overlap of two coordinate systems, as provided in an embodiment of this application.
[0045] Figure 7 This is a schematic flowchart illustrating a camera control method provided in an embodiment of this application.
[0046] Figure 8 A schematic diagram illustrating the specific process for establishing a mapping relationship between an audio coordinate system and a gimbal coordinate system, provided in an embodiment of this application;
[0047] Figure 9 This is a schematic diagram of the structure of a camera control system provided in an embodiment of this application;
[0048] Figure 10 This is a schematic diagram of the structure of a camera control device provided in an embodiment of this application. Detailed Implementation
[0049] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0050] This application provides a camera control method, system, and device to enable comprehensive environmental control of cameras without the assistance of other cameras, thereby improving the adaptability of camera control.
[0051] The method and apparatus are based on the same concept of the application. Since the methods and apparatus solve problems in similar ways, the implementation of the apparatus and methods can refer to each other, and the repeated parts will not be described again.
[0052] The terms "first," "second," etc. (if present) in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0053] The following examples and embodiments are to be understood as illustrative only. While this specification may refer to "a," "an," or "some" examples or embodiments in several places, this does not mean that every such reference relates to the same example or embodiment, nor does it mean that the feature applies only to a single example or embodiment. Individual features of different embodiments may also be combined to provide other embodiments. Furthermore, terms such as "comprising" and "including" should be understood not to limit the described embodiments to consisting only of those features mentioned; such examples and embodiments may also include features, structures, units, modules, etc., not specifically mentioned.
[0054] The various embodiments of this application will now be described in detail with reference to the accompanying drawings. It should be noted that the order in which the embodiments are presented in this application represents only a chronological order and does not represent the superiority or inferiority of the technical solutions provided by the embodiments.
[0055] It should be noted that the technical solution provided in this application embodiment receives the sound signal emitted by the external target sound source through the microphone array on the camera, inputs the received sound signal into a multi-classification neural network, outputs the audio coordinates of the target sound source in the audio coordinate system of the microphone array and the sound type, and compares the output result with the preset listening area and sound type. When they match, the camera is rotated to the location of the target sound source. This is an example of the method, but it is not limited to this.
[0056] See Figure 1This application provides a camera control method, including:
[0057] Step S101: Based on the sound signal of the external target sound source received by the microphone array on the camera, determine the audio coordinates of the target sound source in the audio coordinate system of the microphone array and the sound type of the target sound source;
[0058] The camera, for example, is a spherical camera or a camera with a pan-tilt-zoom (PTZ) unit; the sound signal of the external target sound source is, for example, the sound of a motor vehicle horn, an alarm, a dog barking, or other various types of sounds in the environment.
[0059] In this step, the sound signal of the external target sound source received by the microphone array is input into a multi-class neural network. The multi-class neural network predicts the sound signal and outputs the prediction result, which is the audio coordinates (specific direction category, such as due north) of the target sound source in the audio coordinate system of the microphone array and the sound type of the target sound source.
[0060] Multi-class neural networks are obtained by training convolutional neural networks in the following two stages:
[0061] Collect sound data from different directions and classify this sound data into, for example, N directions (the specific number depends on the actual training situation), to obtain audio samples of N direction categories, for example... Figure 2 The audio coordinate system shown has four directional categories (e.g., corresponding to the front, back, left, and right of the microphone array; directional 1 corresponds to due north to due east, directional 2 corresponds to due east to due south, directional 3 corresponds to due south to due west, and directional 4 corresponds to due west to due north). Audio samples from these N directional categories are used as training datasets to train a convolutional neural network, resulting in a directional classification neural network that can classify audio signals by direction.
[0062] By replacing the sound types in the training dataset with different types of sounds (for example, the sound type in the original training dataset is the sound of a motor vehicle horn, which can be replaced with the sound of children playing, alarms, dogs barking, or other various types of sounds in the environment), a new training dataset is obtained. The new training dataset is then used to train the above-mentioned directional classification neural network to obtain a multi-class neural network. The specific training method can be any method for training a neural network, which will not be elaborated in the embodiments of this application.
[0063] Step S102: When the location of the audio coordinates is within the monitoring area pre-set for the camera, and the sound type is consistent with the preset sound type, the audio coordinates of the target sound source are converted into pan-tilt coordinates in the pan-tilt coordinate system according to the pre-determined mapping relationship between the audio coordinate system of the microphone array on the camera and the pan-tilt coordinate system of the camera. The parameters that the camera needs to adjust are determined according to the pan-tilt coordinates, and the camera is rotated to the corresponding position according to the parameters.
[0064] The monitoring area (i.e., the region of interest) in this step is drawn in the audio coordinate system through the region of interest configuration interface of the upper-level application, such as a webpage on a computer. This monitoring area can be of any shape, such as a rectangle, sector, triangle, or circle. Three points can be determined first (audio coordinates, or other numbers of points), and then the shape is drawn using these three points. For example... Figure 3 As shown, the drawn sector-shaped region of interest (ROI) is defined, with the x-axis pointing south and the y-axis pointing east. The ROI is located between northwest and northeast, and can be represented as (45°, 135°). West corresponds to 0°. Assuming that step S101 determines the direction of the target sound source to be north (90°), the pre-defined ROI includes this direction; that is, the audio coordinates of the target sound source fall within this region. Figure 3 Within the region of interest shown; when the sound type of the target sound source determined by the multi-class neural network is, for example, the sound of a motor vehicle horn, and if the preset sound type is also the sound of a motor vehicle horn, it means that the sound type of the target sound source is consistent with the preset sound type, which meets the condition for linkage with the camera. Then, according to the preset mapping relationship between the audio coordinate system of the microphone array on the camera and the pan-tilt coordinate system of the camera, the audio coordinates of the target sound source are converted into pan-tilt coordinates, and the adjustment parameters of the camera are determined according to the pan-tilt coordinates of the target sound source. The camera is rotated according to the adjustment parameters so that the location of the target sound source falls within the monitoring range of the camera.
[0065] Step S102 enables the camera to respond in real time to the multi-directional deployment requirements of the environment, thereby improving the adaptability of camera deployment.
[0066] To establish the mapping relationship between the audio coordinate system of the microphone array on the camera and the pan-tilt coordinate system of the camera, see [reference needed]. Figure 4 The specific steps include:
[0067] For example Figure 5As shown, the camera includes a microphone array, a camera, and a sound-generating device. Directly above the camera C1, four microphones M1, M2, M3, and M4 are integrated, forming a microphone array. The sound-receiving surfaces of M1 and M3 are opposite, and the sound-receiving surfaces of M2 and M4 are opposite. A sound-generating device S1 is installed on the camera.
[0068] It should be noted that the microphone array described above can also be composed of other numbers of microphones, as long as the number of microphones is not less than four. The more microphones there are, the more accurate the final mapping relationship will be, but the corresponding cost will also be higher. This embodiment of the application takes into account both cost and the accuracy of the mapping relationship, and selects four microphones to form the microphone array; the microphones in the microphone array are not limited to using... Figure 5 The arrangement shown is not limited to any other form; other arrangements may also be used.
[0069] Step S201: Construct the audio coordinate system of the microphone array and the pan-tilt coordinate system of the camera;
[0070] In this step, an audio coordinate system is established on the plane where the microphone array is located (e.g., parallel to the horizontal rotation plane of the camera). For example, the origin is the intersection of the line connecting microphones M1 and M3 and the line connecting microphones M2 and M4, with the line connecting M1 and M3 as the x-axis and the line connecting M2 and M4 as the y-axis. The direction of M1 is the x-axis direction, and the direction of M2 is the y-axis direction. A pan-tilt-zoom (PTZ) coordinate system is then established on the horizontal rotation plane of the camera, with the pan-tilt-zoom (PTZ) tether as the origin. Figure 5 The gimbal coordinate system shown has the south direction as the X-axis and the east direction as the Y-axis.
[0071] It should be noted that the audio coordinate system of the microphone array and the pan-tilt coordinate system of the camera can also be constructed using other methods, and are not limited to the methods in the embodiments of this application.
[0072] Step S202: Turn on the sound-emitting device on the camera lens, use the sound-emitting device to emit an audio signal of a fixed frequency (e.g., frequency f), and at the same time rotate the camera horizontally.
[0073] The horizontal rotation of the camera in this step can be either counterclockwise or clockwise; no restrictions are imposed in this embodiment.
[0074] Step S203: Calculate the audio energy of the audio signal received by each microphone in the microphone array. When the audio energy of the audio signal received by each microphone reaches its maximum, record the position of the camera at this time, i.e., the pan-tilt coordinates of the camera (the coordinates corresponding to the camera rotating to a certain direction or angle).
[0075] In this step, the audio energy of the audio signal received by each microphone can be calculated using any existing method, and no restrictions are imposed on this embodiment.
[0076] Before calculating the audio energy, the audio signal received by each microphone is frequency filtered to remove audio signals of other frequencies, retaining only the audio signal with frequency f. The audio energy is then calculated from the frequency-filtered audio signal, resulting in a more accurate audio energy calculation.
[0077] Step S204: Using the camera's pan-tilt coordinates and the corresponding microphone coordinates recorded in step S203, establish a mapping relationship between the audio coordinate system of the microphone array and the camera's pan-tilt coordinate system.
[0078] For example Figure 5 As shown, through step S203, when the sound-emitting device is directly below M1, the audio signal (i.e., audio energy) received by M1 is the strongest. Assuming the audio coordinates of M1 are (x1, 0), the pan-tilt coordinates of camera C1 at this time are recorded as (X1, Y1); when the sound-emitting device is directly below M2, the audio signal received by M2 is the strongest. Assuming the audio coordinates of M2 are (0, y2), the pan-tilt coordinates of camera C1 at this time are recorded as (X2, Y2); when the sound-emitting device is directly below M3, the audio signal received by M3 is the strongest. Assuming the audio coordinates of M3 are (-x3, 0), the pan-tilt coordinates of camera C1 at this time are recorded as (X3, Y3); when the sound-emitting device is directly below M4, the audio signal received by M4 is the strongest. Assuming the audio coordinates of M4 are (0, -y4), the pan-tilt coordinates of camera C1 at this time are recorded as (X4, Y4).
[0079] The audio coordinates of each microphone and the corresponding pan-tilt coordinates of the camera are used to calculate the coordinate mapping rotation θ and translation (a, b) using the following formulas one and two:
[0080] x s =X p *cos(θ)+Y p Formula 1 for sin(θ) + a
[0081] y s =Y p *cos(θ)-X p Formula 2 for sin(θ) + b
[0082] Among them, (x s y s (X) represents the coordinates of a point in the audio coordinate system. p Y p () represents the coordinates of a point in the gimbal coordinate system;
[0083] For example Figure 6 As shown, to make coordinate system uv coincide with coordinate system xy, coordinate system uv needs to be rotated by θ degrees, and then moved horizontally by a and vertically by b on the plane.
[0084] To achieve the conversion between audio coordinates and pan-tilt coordinates, in some embodiments, the pre-determined mapping relationship between the audio coordinate system of the microphone array on the camera and the pan-tilt coordinate system of the camera is specifically determined by the following methods:
[0085] The microphone array receives an audio signal of a preset fixed frequency (e.g., frequency f) emitted by a sound-emitting device mounted on the horizontally rotating camera.
[0086] Calculate the audio energy of the audio signal received by each microphone in the microphone array. When the audio energy of the audio signal received by each microphone is at its maximum, record the pan-tilt coordinates of the camera at this time (e.g., (X1, Y1), (X2, Y2), (X3, Y3), (X4, Y4) as mentioned above).
[0087] Using the audio coordinates at which the audio energy of the audio signal received by each microphone is at its maximum and the corresponding pan-tilt coordinates of the camera, a mapping relationship between the audio coordinate system and the pan-tilt coordinate system is established (for example, using (x1, 0) and (X1, Y1), (0, y2) and (X2, Y2), (-x3, 0) and (X3, Y3), (0, -y4) and (X4, Y4) mentioned above, the coordinate mapping rotation θ and translation (a, b) are calculated using formulas one and two).
[0088] To improve the accuracy of subsequent audio energy calculations, in some embodiments, the method further includes:
[0089] The audio signal received by each microphone in the microphone array is subjected to frequency filtering (i.e., filtering out audio signals of other frequencies).
[0090] To improve the accuracy of target sound source localization, in some embodiments, determining the audio coordinates of the target sound source in the audio coordinate system and the sound type of the target sound source based on the sound signal of the external target sound source received by the microphone array specifically includes:
[0091] The sound signal of the target sound source received by the microphone array is input into a multi-class neural network, which outputs the audio coordinates of the target sound source in the audio coordinate system and the sound type of the target sound source.
[0092] The multi-class neural network is obtained by dividing pre-collected audio samples from different directions into a preset number of directional categories (e.g., the N directional categories mentioned above), using the audio samples from the preset number of directional categories as training samples to train the convolutional neural network, and then using different target sound sources to replace the sound types in the training samples to train the trained convolutional neural network.
[0093] The following are examples of specific methodologies and procedures.
[0094] Example 1:
[0095] See Figure 7 The camera control method provided in this application includes the following steps:
[0096] Step S301: Establish the mapping relationship between the audio coordinate system of the microphone array on the camera and the pan-tilt coordinate system of the camera, and store the mapping relationship.
[0097] Step S302: Set the region of interest and sound type for the camera to monitor;
[0098] The region of interest set in this step is, for example... Figure 3 The range of 45 to 135 degrees shown corresponds to the sound type of a motor vehicle horn.
[0099] Step S303: Monitor the target environment using a camera and acquire the sound source signal collected by the microphone array on the camera;
[0100] Step S304: Input the sound source signals collected by the microphone array into the multi-classification neural network, and output the audio coordinates and sound type of the sound source signals;
[0101] In this step, for example, the direction corresponding to the audio coordinates of the output sound source signal is 60 degrees, and the sound type is the horn of a motor vehicle.
[0102] Step S305: Compare the output results of the multi-class neural network with the preset region of interest and sound type. If they are consistent, convert the audio coordinates of the sound source signal into gimbal coordinates according to the mapping relationship established in step S301.
[0103] The direction corresponding to the audio coordinates of the sound source signal is 60 degrees. It is within the range of 45 degrees to 135 degrees, meaning that the location of the sound source signal is within the region of interest. The sound type of the sound source signal is the sound of a motor vehicle horn, which is consistent with the preset sound type. The linkage condition is met, and the camera needs to be switched.
[0104] Step S306: Adjust the camera parameters according to the pan-tilt coordinates of the sound source signal, and rotate the camera to the location of the sound source signal, i.e., the location of the sound source signal falls within the camera's monitoring range.
[0105] After the camera switches positions, repeat steps S304-S306 to enable the camera to respond in real time to the multi-directional deployment requirements of the environment.
[0106] Example 2:
[0107] See Figure 8 This application provides a coordinate system calibration method, assuming the camera's microphone array consists of four microphones, for example... Figure 5 As shown, the specific steps include:
[0108] Step S401: Construct the audio coordinate system of the microphone array on the camera and the pan-tilt coordinate system of the camera, and record the audio coordinates of each microphone.
[0109] Step S402: Turn on the sound device on the camera to emit an audio signal with a frequency of f, and rotate the camera horizontally clockwise.
[0110] Step S403: Acquire the audio signal collected by each microphone in the microphone array, and perform frequency filtering on the acquired audio signal to obtain the filtered audio signal.
[0111] Step S404: Calculate the audio energy of the audio signal after filtering for each microphone. When the audio energy of the microphone is at its maximum, record the pan-tilt coordinates of the camera at this time, which is the combination of the microphone's audio coordinates and the camera's pan-tilt coordinates.
[0112] Step S405: Calculate the coordinate mapping rotation and translation using the above formulas one and two, thus obtaining the mapping relationship between the audio coordinate system and the gimbal coordinate system.
[0113] See Figure 9 This application provides a camera control system, comprising:
[0114] The sound source localization module 100 is used to determine the audio coordinates (i.e., the direction of the target sound source) and the sound type of the target sound source in the audio coordinate system of the microphone array based on the sound signal of the external target sound source received by the microphone array on the camera.
[0115] The camera adjustment module 200 is used to convert the audio coordinates of the target sound source into pan-tilt coordinates in the pan-tilt coordinate system according to the pre-determined mapping relationship between the audio coordinate system of the microphone array on the camera and the pan-tilt coordinate system of the camera, when the location of the audio coordinates is within the monitoring area pre-set for the camera and the sound type is consistent with the preset sound type, determine the parameters that the camera needs to be adjusted according to the pan-tilt coordinates, and rotate the camera to the corresponding position according to the parameters.
[0116] To achieve the conversion between audio coordinates and gimbal coordinates, in some embodiments, see [reference needed]. Figure 9 The system also includes a coordinate calibration module 300:
[0117] This is used to establish the mapping relationship between the audio coordinate system of the microphone array on the camera and the pan-tilt coordinate system of the camera.
[0118] To establish a mapping relationship between the audio coordinate system of the microphone array and the pan-tilt coordinate system of the camera, facilitating subsequent coordinate transformations using this mapping relationship, in some embodiments, the coordinate calibration module 300 is specifically used for:
[0119] The microphone array receives audio signals of a preset fixed frequency emitted by a sound-emitting device mounted on the horizontally rotating camera.
[0120] Calculate the audio energy of the audio signal received by each microphone in the microphone array, and record the pan-tilt coordinates of the camera when the audio energy of the audio signal received by each microphone is at its maximum.
[0121] By using the audio coordinates of the audio signal received by each microphone at its maximum audio energy and the corresponding pan-tilt coordinates of the camera, a mapping relationship between the audio coordinate system and the pan-tilt coordinate system is established.
[0122] To improve the accuracy of subsequent audio energy calculations, in some embodiments, the coordinate calibration module 300 is further used for:
[0123] The audio signal received by each microphone in the microphone array is subjected to frequency filtering.
[0124] The following describes the device or apparatus provided in the embodiments of this application, and the explanations or examples of the same or corresponding technical features as those described in the above methods will not be repeated hereafter.
[0125] See Figure 10 This application provides a camera control device, comprising:
[0126] Processor 600 is used to read the program from memory 620 and execute the following procedures:
[0127] Based on the sound signal of the external target sound source received by the microphone array on the camera, determine the audio coordinates of the target sound source in the audio coordinate system of the microphone array and the sound type of the target sound source;
[0128] When the location of the audio coordinates is within the pre-set monitoring area for the camera, and the sound type is consistent with the preset sound type, the audio coordinates of the target sound source are converted into pan-tilt coordinates in the pan-tilt coordinate system according to the pre-determined mapping relationship between the audio coordinate system of the microphone array on the camera and the pan-tilt coordinate system of the camera. The parameters that the camera needs to adjust are determined according to the pan-tilt coordinates, and the camera is rotated to the corresponding position according to the parameters.
[0129] In some embodiments, the processor 600 is further configured to read the program in the memory 620 and execute the predetermined mapping relationship between the audio coordinate system of the microphone array on the camera and the pan-tilt coordinate system of the camera.
[0130] The microphone array receives audio signals of a preset fixed frequency emitted by a sound-emitting device mounted on the horizontally rotating camera.
[0131] Calculate the audio energy of the audio signal received by each microphone in the microphone array, and record the pan-tilt coordinates of the camera when the audio energy of the audio signal received by each microphone is at its maximum.
[0132] By using the audio coordinates of the audio signal received by each microphone at its maximum audio energy and the corresponding pan-tilt coordinates of the camera, a mapping relationship between the audio coordinate system and the pan-tilt coordinate system is established.
[0133] In some embodiments, the processor 600 is further configured to read a program from the memory 620 and execute it:
[0134] The audio signal received by each microphone in the microphone array is subjected to frequency filtering.
[0135] In some embodiments, determining the audio coordinates of the target sound source in the audio coordinate system and the sound type of the target sound source based on the sound signal of the external target sound source received by the microphone array specifically includes:
[0136] The sound signal of the target sound source received by the microphone array is input into a multi-class neural network, which outputs the audio coordinates of the target sound source in the audio coordinate system and the sound type of the target sound source.
[0137] The multi-class neural network is obtained by dividing pre-collected audio samples from different directions into a preset number of directional categories, using the audio samples from the preset number of directional categories as training samples to train the convolutional neural network, and then using different target sound sources to replace the sound types in the training samples to train the trained convolutional neural network.
[0138] In some embodiments, the camera control device provided in this application further includes a transceiver 610 for receiving and sending data under the control of the processor 600.
[0139] Among them, Figure 10 In this context, the bus architecture can include any number of interconnected buses and bridges, specifically linking various circuits together, represented by one or more processors (processor 600) and memory (memory 620). The bus architecture can also link together various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and therefore will not be described further herein. The bus interface provides an interface. The transceiver 610 can be multiple elements, including a transmitter and a receiver, providing a unit for communicating with various other devices over a transmission medium.
[0140] In some embodiments, the camera control device provided in this application further includes a user interface 630. The user interface 630 may be an interface that can connect to external or internal devices, including but not limited to keypads, displays, speakers, microphones, joysticks, etc.
[0141] The processor 600 is responsible for managing the bus architecture and general processing, while the memory 620 can store the data used by the processor 700 during operation.
[0142] In some embodiments, the processor 600 may be a CPU (Central Processing Unit), an ASIC (Application Specific Integrated Circuit), an FPGA (Field-Programmable Gate Array), or a CPLD (Complex Programmable Logic Device).
[0143] This application provides a computing device, which may specifically be a desktop computer, portable computer, smartphone, tablet computer, personal digital assistant (PDA), etc. The computing device may include a central processing unit (CPU), memory, input / output devices, etc. Input devices may include a keyboard, mouse, touchscreen, etc., and output devices may include display devices, such as a liquid crystal display (LCD) or a cathode ray tube (CRT).
[0144] The memory may include read-only memory (ROM) and random access memory (RAM), and provides the processor with program instructions and data stored in the memory. In the embodiments of this application, the memory may be used to store the program of any of the methods provided in the embodiments of this application.
[0145] The processor executes any of the methods described in the embodiments of this application according to the program instructions stored in the memory.
[0146] This application also provides a computer program product or computer program that includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform any of the methods described in the above embodiments. The program product may employ any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof.
[0147] This application provides a computer-readable storage medium for storing computer program instructions used in the apparatus provided in the above-described embodiments, including a program for performing any of the methods provided in the above-described embodiments. The computer-readable storage medium may be a non-transitory computer-readable medium.
[0148] The computer-readable storage medium can be any available medium or data storage device that a computer can access, including but not limited to magnetic storage (e.g., floppy disks, hard disks, magnetic tapes, magneto-optical disks (MOs), etc.), optical storage (e.g., CDs, DVDs, BDs, HVDs, etc.), and semiconductor storage (e.g., ROMs, EPROMs, EEPROMs, non-volatile memory (NAND flash), solid-state drives (SSDs)).
[0149] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0150] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0151] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0152] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0153] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.
Claims
1. A camera control method, characterized in that, The method includes: Based on the sound signal of the external target sound source received by the microphone array on the camera, determine the audio coordinates of the target sound source in the audio coordinate system of the microphone array and the sound type of the target sound source; When the location of the audio coordinates is within the monitoring area pre-set for the camera, and the sound type is consistent with the preset sound type, the audio coordinates of the target sound source are converted into pan-tilt coordinates in the pan-tilt coordinate system according to the pre-determined mapping relationship between the audio coordinate system of the microphone array on the camera and the pan-tilt coordinate system of the camera. The parameters that the camera needs to adjust are determined according to the pan-tilt coordinates, and the camera is rotated to the corresponding position according to the parameters. Specifically, determining the audio coordinates of the target sound source in the audio coordinate system and the sound type of the target sound source based on the sound signal received by the microphone array from the external target sound source includes: The sound signal of the target sound source received by the microphone array is input into a multi-class neural network, which outputs the audio coordinates of the target sound source in the audio coordinate system and the sound type of the target sound source. The multi-class neural network is obtained by dividing pre-collected audio samples from different directions into a preset number of directional categories, using the audio samples from the preset number of directional categories as training samples to train the convolutional neural network, and then using different target sound sources to replace the sound types in the training samples to train the trained convolutional neural network. The predetermined mapping relationship between the audio coordinate system of the microphone array on the camera and the pan-tilt coordinate system of the camera is specifically determined by the following methods: The microphone array receives audio signals of a preset fixed frequency emitted by a sound-emitting device mounted on the horizontally rotating camera. Calculate the audio energy of the audio signal received by each microphone in the microphone array, and record the pan-tilt coordinates of the camera when the audio energy of the audio signal received by each microphone is at its maximum. By using the audio coordinates of the audio signal received by each microphone at its maximum audio energy and the corresponding pan-tilt coordinates of the camera, a mapping relationship between the audio coordinate system and the pan-tilt coordinate system is established.
2. The method according to claim 1, characterized in that, The method further includes: The audio signal received by each microphone in the microphone array is subjected to frequency filtering.
3. A camera control system, characterized in that, The system includes: The sound source localization module is used to determine the audio coordinates of the target sound source in the audio coordinate system of the microphone array and the sound type of the target sound source based on the sound signal of the external target sound source received by the microphone array on the camera. The camera adjustment module is used to convert the audio coordinates of the target sound source into pan-tilt coordinates in the pan-tilt coordinate system according to the pre-determined mapping relationship between the audio coordinate system of the microphone array on the camera and the pan-tilt coordinate system of the camera, when the location of the audio coordinates is within the pre-set monitoring area for the camera and the sound type is consistent with the preset sound type. Based on the pan-tilt coordinates, the module determines the parameters that the camera needs to be adjusted, and rotates the camera to the corresponding position according to the parameters. Specifically, determining the audio coordinates of the target sound source in the audio coordinate system of the microphone array and the sound type of the target sound source based on the sound signal of the external target sound source received by the microphone array on the camera includes: The sound signal of the target sound source received by the microphone array is input into a multi-class neural network, which outputs the audio coordinates of the target sound source in the audio coordinate system and the sound type of the target sound source. The multi-class neural network is obtained by dividing pre-collected audio samples from different directions into a preset number of directional categories, using the audio samples from the preset number of directional categories as training samples to train the convolutional neural network, and then using different target sound sources to replace the sound types in the training samples to train the trained convolutional neural network. The system also includes a coordinate calibration module, specifically used for: The microphone array receives audio signals of a preset fixed frequency emitted by a sound-emitting device mounted on the horizontally rotating camera. Calculate the audio energy of the audio signal received by each microphone in the microphone array, and record the pan-tilt coordinates of the camera when the audio energy of the audio signal received by each microphone is at its maximum. By using the audio coordinates of the audio signal received by each microphone at its maximum audio energy and the corresponding pan-tilt coordinates of the camera, a mapping relationship between the audio coordinate system and the pan-tilt coordinate system is established.
4. The system according to claim 3, characterized in that, The coordinate calibration module is also used for: The audio signal received by each microphone in the microphone array is subjected to frequency filtering.
5. A camera control device, characterized in that, include: Memory, used to store program instructions; A processor is configured to invoke program instructions stored in the memory and execute the method of claim 1 or 2 according to the obtained program.
6. A computer program product for use in a computer, characterized in that, Includes a software code portion that, when the product is run on the computer, is used to perform the method according to claim 1 or 2.
7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions for causing the computer to perform the method of claim 1 or 2.
Citation Information
Patent Citations
Camera adjusting method, device and system based on sound source localization
CN110389597A
Security monitoring method, device, robot and storage medium
CN111601074A