Shooting position updating method, image acquisition equipment and storage medium
By combining conference audio and images, dynamically adjusting the camera angle and tracking speakers in real time, the problems of camera adjustment lag and inaccurate sound source positioning are solved, and the accuracy and completeness of conference records are improved.
Patent Information
- Application Number
- CN202510573452.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-07-18
AI Technical Summary
In the existing conference recording methods, the camera adjustment is lagging and the sound source positioning is inaccurate, resulting in inaccurate information, especially in multi-person discussion scenarios, which cannot efficiently capture the sound and images of each speaker.
By combining conference audio and images, the predicted shooting direction is determined using the sound source direction, the target image is collected and the location of the target speaker is determined, the camera angle is dynamically adjusted, and the speaker is tracked in real time.
Improve the accuracy and completeness of meeting minutes, ensure the audio and image matching of speakers, and achieve efficient capture of each speaker.
Smart Images

Figure CN120343405A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of shooting position adjustment, and particularly to a method for updating a shooting position, an image acquisition device, and a storage medium. Background Art
[0002] The current meeting recording method manually adjusts the microphone and camera to align with the speaking member.
[0003] However, in a meeting scenario based on manual adjustment, there are defects such as lag in camera adjustment and inaccurate sound source positioning, resulting in inaccurate meeting recording information.
[0004] The above content is only used to assist in understanding the technical solution of this application, and does not represent an admission that the above content is prior art. Summary of the Invention
[0005] The main purpose of this application is to provide a method for updating a shooting position, an image acquisition device, and a storage medium, aiming to solve the technical problem of inaccurate current meeting recording information.
[0006] To achieve the above purpose, this application proposes a method for updating a shooting position, and the method for updating a shooting position includes: Determine a predicted shooting direction according to the sound source direction of the meeting audio; Collect a target image corresponding to the predicted shooting direction, determine a target speaker according to the meeting mode corresponding to the meeting audio, and the target position of the target speaker in the target image; Adjust the current shooting direction according to the target position.
[0007] In one embodiment, the step of determining a predicted shooting direction according to the sound source direction of the meeting audio includes: Obtain historical meeting data corresponding to the meeting mode; Determine the predicted shooting direction based on a network model trained based on the historical meeting data and the sound source direction.
[0008] In one embodiment, the step of adjusting the current shooting direction according to the target position includes: Determine a position change amount of the target position according to the target position and inertial sensing data; Determine a motion parameter of the video module based on the position change amount, and adjust the current shooting direction of the video module based on the motion parameter.
[0009] In one embodiment, after the step of determining a position change amount of the target position according to the target position and inertial sensing data, the method for updating a shooting position further includes: Determine the shooting mode corresponding to the meeting mode; Determine the motion parameters of each video module according to the shooting mode and the position change amount; Adjust the current shooting direction of each video module based on the motion parameters.
[0010] In one embodiment, after the step of controlling the operation of the video module based on the motion parameters to adjust the current shooting direction, the method for updating the shooting position further includes: Obtain the offset angle corresponding to the inertial sensing data; If the offset angle is greater than a preset angle, execute the step of determining the position change amount of the target position according to the target position and the inertial sensing data.
[0011] In one embodiment, the steps of collecting the target image corresponding to the predicted shooting direction, determining the target speaker according to the meeting mode of the meeting audio, and the target position of the target speaker in the target image include: Collect the target image corresponding to the preset shooting direction, and determine the speaking weight of the speaker corresponding to the meeting audio according to the meeting mode; Determine the target speaker based on the speaking weight; Obtain the target position of the target speaker in the target image.
[0012] In one embodiment, the steps of collecting the target image corresponding to the preset shooting direction and determining the speaking weight of the speaker corresponding to the meeting audio according to the meeting mode include: Collect the target image corresponding to the preset shooting direction, and determine the historical meeting data and shooting direction adjustment data corresponding to the meeting mode; Fuse the current meeting data, the historical meeting data, and the shooting direction adjustment data to obtain a fused feature, and calculate the speaking weight of the speaker in the fused feature according to the multi-head attention algorithm.
[0013] In one embodiment, after the step of adjusting the current shooting direction according to the target position, the method for updating the shooting position further includes: Continuously collect the target audio information and target image information of the target speaker; Generate the meeting record information of the target shooting person according to the target audio information and the target image information.
[0014] In addition, to achieve the above object, the present application further provides an image acquisition device, which includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, and the computer program is configured to implement the steps of the shooting position updating method as described above.
[0015] In addition, to achieve the above object, the present application further provides a storage medium, which is a computer-readable storage medium, and a computer program is stored on the storage medium, and when the computer program is executed by a processor, the steps of the shooting position updating method as described above are implemented.
[0016] One or more technical solutions provided by the present application have at least the following technical effects: During the process of meeting recording, first, the predicted shooting direction to be photographed is determined through the sound source direction of the meeting audio, and then the target image in the predicted shooting direction is collected. Then, according to the meeting mode corresponding to the meeting audio, the target speaker in the target image and the target position of the target speaker on the target image are determined. Finally, the current shooting direction is adjusted according to the target position. Based on this, during the meeting acquisition process, by combining the meeting audio and the meeting image, the speaker is tracked in real time, and the camera angle is dynamically adjusted to ensure the matching of the audio and image of the speaker, efficiently capture the voice and image of each speaker, and improve the accuracy of the meeting record information. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The accompanying drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present application and used together with the specification to explain the principles of the present application.
[0018] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0019] Figure 1 It is a schematic flowchart provided for the first embodiment of the shooting position updating method of the present application; Figure 2 It is a schematic flowchart provided for the second embodiment of the shooting position updating method of the present application; Figure 3 It is a schematic flowchart of a brief process when the data is processed based on the Transformer model for the shooting position updating method provided by the present application; Figure 4 It is a schematic diagram of the device structure of the hardware operating environment involved in the shooting position updating method in the embodiments of the present application.
[0020] The realization of the purpose, functional features and advantages of this application will be further described in conjunction with the embodiments with reference to the accompanying drawings. Specific embodiments
[0021] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of this application and are not used to limit this application.
[0022] The current meeting recording method manually adjusts the microphone and camera to aim at the speaking member.
[0023] However, in the meeting scenario based on manual adjustment, there are defects such as lag in camera adjustment and inaccurate sound source localization, resulting in inaccurate meeting recording information.
[0024] At the same time, for the meeting scenario where multiple people discuss simultaneously, traditional devices cannot efficiently capture the voices and images of each speaker, resulting in omission or inaccuracy of meeting information. Therefore, the existing meeting minutes generation methods usually rely on manual collation, or can only generate simplified text records based on speech transcription, lacking comprehensive capture of meeting participants and speech content, and having problems such as lag in camera adjustment, inaccurate sound source localization, video frame jitter, and incomplete capture of meeting content.
[0025] Based on this, the main solution of the embodiments of this application is: determining a predicted shooting direction according to the sound source direction of the meeting audio; Collecting a target image corresponding to the predicted shooting direction, determining a target speaker according to the meeting mode corresponding to the meeting audio, and the target position of the target speaker in the target image; Adjusting the current shooting direction according to the target position.
[0026] Specifically, during the meeting collection process, by combining the meeting audio and the meeting image, the speaker is tracked in real time, and the camera angle is dynamically adjusted to ensure the matching of the audio and image of the speaking person, efficiently capturing the voices and images of each speaker, and improving the accuracy of the meeting recording information.
[0027] It should be noted that the execution subject of this embodiment can be a computing service device with data processing, network communication, and program running functions, or an image acquisition device that can implement the above functions. Among them, the image acquisition device includes headphones with cameras, such as TWS (True Wireless Stereo) headphones, OWS (Open-Ear Wireless Stereo) headphones, etc., and wearable devices such as smart home systems, online education systems, and security monitoring systems equipped with video modules that need to take images through the video module. This can improve the user experience and market competitiveness of these devices in actual use and promote the development of intelligent wearable technology. Hereinafter, taking headphones with cameras as an example, this embodiment and the following embodiments will be described. Among them, the headphones are equipped with two independent cameras (left and right ears), and based on its inherent stereo vision structure advantages, such as sufficient baseline length, it can effectively capture high-precision depth (distance) information. At the same time, the spatial positioning of feature points is more accurate and stable, significantly reducing the measurement error of monocular vision and improving the accuracy of pose angle estimation.
[0028] To better understand the technical solution of this application, the following will be described in detail in combination with the specification drawings and specific implementation manners.
[0029] An embodiment of this application provides a method for updating the shooting position, referring to Figure 1 , Figure 1 which is a schematic flowchart of the first embodiment of the method for updating the shooting position of this application.
[0030] In this embodiment, the method for updating the shooting position includes steps S10 to S30: Step S10, determine the predicted shooting direction according to the sound source direction of the conference audio.
[0031] In this embodiment, the predicted shooting direction can be the direction for shooting at the current moment or the shooting direction at a future moment.
[0032] Therefore, as an optional implementation manner for determining the predicted shooting direction, the headphones collect the conference audio through a microphone array, detect the voice characteristics of the speaker, and thus calculate the time difference (TDOA) or phase difference when the sound reaches different microphone arrays. Combining with a beamforming algorithm (such as the MVDR algorithm), the azimuth coordinates of the sound source in three-dimensional space are output. That is, by analyzing the multi-channel audio signal through the sound source localization technology, the azimuth angle θ and elevation angle φ of the speaker are calculated, and then the predicted shooting direction is calculated according to the angle parameters. The camera of the headphones can adjust the captured target image based on the real-time sound source direction when the speaker moves frequently, and at the same time, adjust the shooting direction based on the position change information of the speaker in the target image.
[0033] In another alternative embodiment for determining the predicted shooting direction, step S10 may include steps S11 to S12: Step S11, obtaining historical meeting data corresponding to the meeting mode.
[0034] In this embodiment, different meeting modes correspond to different historical meeting data. Therefore, to improve the accuracy of predicting the shooting direction in the future period, it is necessary to obtain the historical meeting data corresponding to the meeting mode. Among them, the historical meeting data can be stored in the cloud, and when the headset records a meeting, the data is obtained from the cloud. By obtaining the historical meeting data, the neural network model can be trained based on the historical meeting data, and then the next speaker can be predicted according to the trained model.
[0035] Step S12, determining the predicted shooting direction based on the network model trained with the historical meeting data and the sound source direction.
[0036] In this embodiment, when calculating the predicted shooting direction at a future moment, model training can be performed based on historical meeting data. The training input parameters include audio data and video data. The audio data includes the historical meeting speech transcription text, the speaking duration of the speaker, and the interaction frequency of the speakers. The video data includes the camera adjustment trajectory, the spatial position change of the speaker, etc. Among them, the above parameters are only for explanation and are not a limitation on the historical meeting data.
[0037] Model training is performed with these historical model data. After the model is trained, the sound source direction is used as an input parameter of the network model, so that the model determines the current user's speaking direction based on the sound source direction. Subsequently, the speaking probability corresponding to each shooting direction at a future moment is calculated based on this direction, and the predicted shooting direction is determined based on the speaking probability, so as to analyze and predict the possible speaking time and speaker through historical meeting data, and combine real-time sound source localization technology to accurately identify the position of the current speaker, that is, predict the predicted shooting direction at a future moment.
[0038] In this embodiment, in a normal meeting scenario, based on the predicted possible speaking time and the sound source direction corresponding to the main speaker, combined with real-time sound source localization technology, the direction to be shot is determined, so as to adjust the shooting direction of the headset camera in advance or in real time based on the predicted shooting direction, avoid the situation of out-of-sync audio and video in the meeting data, and improve the adjustment efficiency of the headset camera.
[0039] Step S20, collecting the target image corresponding to the predicted shooting direction, determining the target speaker according to the meeting mode corresponding to the meeting audio, and the target position of the target speaker in the target image.
[0040] It should be noted that the predicted shooting direction is only a general direction. In this general direction, the target speaker may be located in the edge area of the target image. At this time, to improve the effect of the meeting record, it is necessary to calculate the target position to adjust the shooting direction of the camera and shooting parameters based on this target position, so as to ensure that the speaker is always in the middle of the image and achieve dynamic tracking of the speaker.
[0041] In this embodiment, since the earphone is provided with two cameras, in the meeting scenario, the main and secondary cameras can be set. The main camera is used for tracking and positioning the speaker, and the secondary camera is used for shooting and recording meeting materials such as the meeting screen. Therefore, the main / secondary camera for image acquisition can also be determined according to the sound source direction. For example, the camera closer to the sound source direction is set as the main camera.
[0042] Specifically, after obtaining the predicted shooting direction, if the shooting direction at the current moment is the predicted shooting direction, determine the main camera for shooting the predicted shooting direction through the sound source direction, and control the main camera to rotate to the predicted shooting direction (if the deviation between the current shooting direction of the main camera and the predicted shooting direction is small, there is no need to rotate) and execute the image shooting action.
[0043] While shooting the target image, the earphone determines the meeting mode according to the audio recognition result of the meeting audio. For example, if there is one speaking voice in the meeting audio, the meeting mode is determined as the single-speaker mode. If there are multiple different voices in the meeting audio, the meeting mode is determined as the multi-person discussion mode. Then, after the main camera shoots the target image, determine the target speaker through the meeting mode corresponding to the meeting audio. For example, in the single-speaker mode, determine the target speaker through the timbre of the meeting audio, and calculate the target position of the target speaker on the target image through the image information. In the multi-person discussion mode, it is necessary to select the target speaker from multiple speakers based on the speaking weights of each speaker.
[0044] Optionally, if the predicted shooting direction is the predicted shooting direction for the next moment, when the adjustment condition of the shooting direction is met, such as the current target speaker finishes speaking, or after 5 seconds, control the main camera to rotate to this predicted shooting direction. In addition, the current speaker can be shot by one of the cameras on one side, and at the same time, the other camera on the other side is rotated in advance to the predicted shooting direction, so that after the current speaker finishes speaking, it can seamlessly switch to the predicted shooting direction for video shooting. In this way, through historical data modeling, predict possible future speakers and corresponding predicted directions, and at the same time, combine the sound source data to optimize the camera adjustment strategy and reduce the switching delay.
[0045] In this embodiment, by determining the position information of the target speaker on the target image, when the speaker is in a moving state, the target speaker can be tracked and located in real time for shooting. Based on sound source localization + visual recognition, it is ensured that the camera is always aimed at the correct speaker.
[0046] Step S30: Adjust the current shooting direction according to the target position.
[0047] In this embodiment, after obtaining the target position, the current shooting direction is adjusted according to the direction of the target position in the target image. For example, if the target position is directly to the right of the target image, at this time, it is necessary to control the camera to rotate to the right so that the target speaker is in the middle area of the image, ensuring that the camera is always aimed at the speaker. Therefore, the rotation parameters of the camera can be calculated based on the target position, and the camera operation can be controlled based on the rotation parameters to adjust the current shooting direction.
[0048] It should be noted that the above parameters in this embodiment are only for explanation and are not limitations to this application.
[0049] Furthermore, after adjusting the current shooting direction, the target audio information and target image information of the target speaker can be continuously collected. Subsequently, according to the target audio information and target image information, the meeting record information of the target shooting person can be generated, thereby improving the integrity and stability of the meeting record.
[0050] This embodiment provides a method for updating the shooting position. In a meeting scenario, by calculating the predicted shooting direction at the current moment or a future moment, then analyzing the image collected in the predicted shooting direction based on an image analysis method, and determining the speaker through the meeting mode, and finally automatically tracking and shooting the speaker, it is possible to dynamically adjust the headphone settings according to the meeting mode, automatically adjust the receiving direction of the microphone and the shooting angle of the camera, track the speaker in real time during the meeting, dynamically adjust the camera angle, and improve the accuracy of meeting information recording.
[0051] Based on the first embodiment of this application, in the second embodiment of this application, the same or similar content as the above first embodiment can be referred to the above introduction and will not be repeated hereinafter. On this basis, in addition to setting a camera, the headphone is also provided with an inertial sensor (Inertial Measurement Unit). The inertial sensor usually includes a gyroscope, an accelerometer, and a magnetometer, and can collect triaxial angular velocity, linear acceleration data, and geomagnetic data. Therefore, the collected inertial sensing data includes at least one of triaxial angular velocity, linear acceleration, and geomagnetic data. During the actual meeting recording process, the line of sight of the headphone wearer changes with the speaker, and the headphone generates inertial sensing data due to the wearer turning the head / twisting the head.
[0052] Therefore, please refer to Figure 2 , step S30 includes steps S31 to S32: Step S31, determine the position change amount of the target position according to the target position and inertial sensing data.
[0053] In this embodiment, in addition to adjusting the current shooting direction through the target position, the rotation parameters can also be calculated by combining the target position and the inertial sensing data of the earphone, that is, the data compensation is performed on the position change amount calculated by the target position through the inertial sensing data to improve the accuracy of the position change amount. Thus, when the user wearing the earphone moves the head to align the camera with the speaker, through the dynamic adjustment of the inertial sensing data and the intelligent camera, the automation, accuracy, and stability of the meeting record are realized.
[0054] Specifically, the position change amount of the target position on the target image can be calculated first. For example, data conversion is performed based on the camera internal parameters to calculate the angle that the main camera needs to rotate. At the same time, the inertial sensing data from the moment when the target image is collected to the present is calculated, and the movement amount of the main camera relative to the target image is calculated based on this inertial sensing data. Finally, the data of the two in the same dimension are compensated to obtain the position change amount.
[0055] Step S32, determine the motion parameters of the video module based on the position change amount, and adjust the current shooting direction of the video module based on the motion parameters.
[0056] In this embodiment, after obtaining the position change amount, it is necessary to convert the position change amount into the motion parameters of the video module, that is, the main camera. For example, it is necessary to move N° to the right, calculate the motion parameters of the camera associated with this N°, and then update the current shooting direction of the video module based on this motion parameter, so as to flexibly adjust the shooting angle of the camera and make it align with the person who is about to speak, avoiding the jitter caused by frequent switching of the camera.
[0057] Optionally, after adjusting the current shooting direction, if the head of the earphone wearer moves to change the shooting direction, at this time, it is necessary to calculate the offset angle corresponding to the inertial sensing data. If the offset angle is small, such as less than or equal to the preset angle, the current camera angle is maintained at this time. Otherwise, it is necessary to re-execute the processing action of step S31 to fine-tune the camera angle to ensure the centered position of the speaker in the image.
[0058] Optionally, in this embodiment, the shooting mode of the earphone is different under different meeting modes, that is, the focused picture of the camera is different. For example, in the single speaker mode, the main camera focuses on the speaker, and the secondary camera locks on the projection screen or the meeting whiteboard. While in the multi-person discussion mode, it is necessary to switch the two side cameras to the panoramic mode to ensure that the content of the entire meeting process is completely recorded. Therefore, after determining the position change amount of the target position, it is also necessary to determine the shooting mode corresponding to the meeting mode, and determine the motion parameters of each video module according to the shooting mode and the position change amount. Finally, adjust the current shooting direction of each video module based on the motion parameters. Thus, different shooting position update actions are performed under different meeting modes, improving the comprehensiveness of meeting records.
[0059] Further, in this embodiment, when it is necessary to adjust the shooting direction based on inertial sensing data, when training the network model and outputting the predicted shooting direction based on the network model, the historical inertial sensing data can be used as training data, so that when determining the predicted shooting direction, based on the probability of head movement during the speech, the predicted shooting direction can be optimized.
[0060] This embodiment provides a method for updating the shooting position. In a meeting scenario, based on historical meeting prediction, sound source localization, multi-camera dynamic adjustment, and inertial sensing data head tracking, the image acquisition direction of the camera is dynamically optimized, so as to combine audio signals, image information, and inertial sensing data to achieve accurate meeting scenario analysis and improve the accuracy of meeting information recording.
[0061] Based on the first embodiment of the present application, in the third embodiment of the present application, the same or similar content as that in the above second embodiment can be referred to the above introduction and will not be repeated hereinafter. On this basis, when determining the target position of the target speaker and the current meeting mode is the multi-person discussion mode, the position that needs to be accurately shot can be determined based on the weight ratio of different speakers. Therefore, step S20 includes steps S21 to S23: Step S21, collect the target image corresponding to the preset shooting direction, and determine the speech weight of the speaker corresponding to the meeting audio according to the meeting mode.
[0062] In this embodiment, after collecting the target image corresponding to the preset direction, the historical meeting data and shooting direction adjustment data corresponding to the meeting mode can be determined first. Subsequently, the current meeting data (the target image of the speaker and the meeting audio collected currently), the historical meeting data, and the shooting direction adjustment data are fused through the self-attention algorithm to obtain a fused feature, and the speech weight corresponding to the speaker in the fused feature is calculated according to the multi-head attention algorithm.
[0063] Exemplarily, the calculation formula of the fused feature is as follows: , Among them, Q (Query) represents the current conference data, i.e., the latest speaker information; K (Key) represents the historical conference data (past speakers, camera adjustment records); V (Value) represents the historical records of camera adjustment and microphone direction adjustment.
[0064] The formula for multi-head attention is as follows:
[0065] Among them, h is the number of multi-head attention. Each attention head learns different speaker weights to improve the accuracy of camera adjustment. W 0 is the output projection matrix.
[0066] In this embodiment, data fusion is performed through the self-attention mechanism and the multi-head attention mechanism, enabling the earphone to simultaneously focus on multiple speakers and calculate the importance of each speaker. Based on the importance of the speakers, the camera adjustment strategy is optimized to ensure that the main speaker is always at the center of the screen. At the same time, the microphone direction is optimized based on the main speaker to ensure that the earphone's sound collection focuses on the most important speaker.
[0067] It can be understood that by fusing the historical conference data and the shooting direction adjustment data, the speaking weights of each speaker in the conference scenario can be accurately analyzed.
[0068] Step S22: Determine the target speaker based on the speaking weight.
[0069] In this embodiment, the greater the speaking weight, the more necessary it is to take this speaker as the main shooting target in the current conference to improve the effectiveness of conference information recording.
[0070] It can be understood that if the conference mode is a single-speaker mode, the target speaker is the speaker corresponding to the single-speaker mode.
[0071] Step S23: Obtain the target position of the target speaker in the target image.
[0072] After determining the target speaker, the target position of the target speaker in the target image is determined through image recognition technology, so as to adjust the shooting direction based on the target position.
[0073] This embodiment provides a method for updating the shooting position. By obtaining the corresponding historical meeting data and shooting direction adjustment data in the meeting mode, and performing feature fusion based on the self-attention algorithm to improve the correlation between features. Subsequently, the speech weights of different speakers in the fused features are calculated through the multi-head attention algorithm. Finally, the target speaker with the highest weight value is selected through the speech weight, so as to select the key shooting target for tracking and positioning in the multi-person discussion mode, thereby improving the effectiveness and importance of the meeting record.
[0074] Based on any of the above embodiments of the present application, in the fourth embodiment of the present application, the same or similar content as any of the above embodiments can be referred to the above introduction and will not be repeated hereinafter. On this basis, it can be trained through a neural network model of the Transformer architecture, including training the historical meeting data based on the model, so as to predict the behavior of the speaker and optimize the camera adjustment strategy.
[0075] Exemplarily, in the data input stage, the input of the Transformer is historical meeting data, including: Audio data: The meeting speech transcription text (after ASR processing). Speech duration, speaker interaction frequency.
[0076] Video data: The meeting camera adjustment trajectory, recording how the camera is adjusted. The spatial position change of the speaker. And inertial sensing data: Head movement trend (Pitch, Yaw, Roll).
[0077] In the feature extraction stage, speech features are extracted through MFCC (Mel-Frequency Cepstral Coefficients), and at the same time, the speech energy distribution is calculated to judge the active speakers in the meeting through the meeting audio.
[0078] The extracted video features include the historical camera trajectory, which records how the camera was adjusted in the past, so as to statistically analyze the position distribution of the speakers in the picture. In the process of extracting the inertial sensing data, it includes calculating the change trend of the head movement direction to predict the possible head movement of the wearer.
[0079] Furthermore, the Transformer model also includes a multi-head self-attention mechanism, which is used to predict the meeting mode, calculate the speaking weights, and determine the adjustment strategies for cameras and microphones based on the multi-head attention mechanism. After the attention mechanism of the Transformer model, a feed-forward neural network (FFN) is used to further extract features, so as to learn the camera adjustment strategies in different meeting modes, identify the single-speaker mode / multi-person discussion mode, and adjust the camera movement method, and then optimize the dynamic adjustment of the microphone to ensure that the microphone can dynamically pick up sound and avoid recording interference noise.
[0080]
[0081] Where, W1 and W2 are the weight matrices of the feed-forward network; b1 and b2 are the bias terms.
[0082] Exemplarily, please refer to Figure 3 , Figure 3 A schematic flowchart of the process of processing data based on the Transformer model is provided during the process of providing a method for updating the shooting position. Specifically, after the earphone runs and captures the meeting audio and meeting images in the meeting scenario, the historical meeting data is loaded, and the Transformer model is used to predict the possible future speakers to obtain the predicted shooting direction. Subsequently, the sound source direction of the current meeting audio is calculated through TDOA to determine the position of the current speaker. At the same time, the camera angle is corrected based on the IMU head tracking, and the microphone direction is adjusted. Then, the side of the speaker is judged through the sound source direction. If it is on the left, the left camera is set as the main camera. If it is on the right, the right camera is set as the main camera. If it is in the middle, the default camera is set as the main camera and the target speaker is captured.
[0083] After shooting, the yaw angle θyaw of the user's head movement is calculated through inertial sensing data. If it is greater than the preset 5°, the camera angle is adjusted. Otherwise, the current shooting angle is maintained to ensure the dynamic fine-tuning of the camera. Subsequently, it is analyzed and judged whether the sound collection direction of the microphone needs to be adjusted, that is, if so, the sound source direction is calculated through TDOA, IMU head angle compensation, etc., and the microphone angle is adjusted. Otherwise, the sound collection angle of the microphone is maintained. Finally, the data is stored and the Transformer model is optimized.
[0084] It should be noted that the above examples are only for understanding the present application and do not constitute a limitation on the method for updating the shooting position of the present application. Based on this technical concept, more forms of simple transformations are within the protection scope of the present application.
[0085] The present application provides an image acquisition device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method for updating the shooting position in the above first embodiment.
[0086] Reference is made below Figure 4 to the structure schematic diagram of the image acquisition device suitable for implementing the embodiments of the present application. Figure 4 The illustrated image acquisition device is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present application.
[0087] As Figure 4 shown, the image acquisition device may include a processing device 1001 (such as a central processing unit, a graphics processing unit, etc.), which can execute various appropriate actions and processes according to the program stored in a read-only memory (ROM, Read Only Memory) 1002 or the program loaded from a storage device 1003 into a random access memory (RAM, Random Access Memory) 1004. In the random access memory 1004, various programs and data required for the operation of the image acquisition device are also stored. The processing device 1001, the read-only memory 1002, and the random access memory 1004 are connected to each other through a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems may be connected to the input / output interface 1006: an input device 1007 including, for example, a touch screen, a touch pad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD, Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 may allow the image acquisition device to communicate with other devices wirelessly or wiredly to exchange data. Although the figure shows an image acquisition device with various systems, it should be understood that it is not required to implement or have all the shown systems. More or fewer systems may be alternatively implemented or had.
[0088] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product that includes a computer program carried on a computer-readable medium, and the computer program contains program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device, or installed from a storage device 1003, or installed from a read-only memory 1002. When the computer program is executed by a processing device 1001, the above-mentioned functions defined in the methods of the embodiments disclosed in the present application are executed.
[0089] The image acquisition device provided by the present application adopts the method for updating the shooting position in the above-mentioned embodiment, and can solve the technical problem of inaccurate current meeting record information. Compared with the prior art, the beneficial effects of the image acquisition device provided by the present application are the same as those of the method for updating the shooting position provided by the above-mentioned embodiment, and other technical features in the image acquisition device are the same as the features disclosed in the method of the previous embodiment, which will not be elaborated here.
[0090] It should be understood that each part disclosed in the present application can be implemented by hardware, software, firmware or a combination thereof. In the description of the above embodiments, specific features, structures, materials or characteristics can be combined in a suitable manner in any one or more embodiments or examples.
[0091] As described above, the above are only the specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in the present application, and all of them should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
[0092] The present application provides a computer-readable storage medium having computer-readable program instructions (i.e., computer programs) stored thereon, and the computer-readable program instructions are used to execute the method for updating the shooting position in the above-mentioned embodiment.
[0093] The computer-readable storage medium provided by this application can be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or components, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memories, read-only memories, erasable programmable read-only memories (EPROMs), optical fibers, portable compact disk read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this embodiment, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, device, or component. The program code contained on the computer-readable storage medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, radio frequency (RF), etc., or any suitable combination of the above.
[0094] The above computer-readable storage medium can be included in an image acquisition device; or it can exist independently and not be assembled into the image acquisition device.
[0095] The above computer-readable storage medium carries one or more programs. When the above one or more programs are executed by the image acquisition device, the image acquisition device is caused to: determine a predicted shooting direction according to the sound source direction of the conference audio; acquire a target image corresponding to the predicted shooting direction, determine a target speaker according to the conference mode corresponding to the conference audio, and the target position of the target speaker in the target image; adjust the current shooting direction according to the target position.
[0096] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The above-mentioned programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (for example, by connecting through the Internet using an Internet service provider).
[0097] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutively represented blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0098] The modules described in the embodiments of this application can be implemented in software or in hardware. Among them, the name of the module does not constitute a limitation on the unit itself in some cases.
[0099] The readable storage medium provided by this application is a computer-readable storage medium. The computer-readable storage medium stores computer-readable program instructions (i.e., computer programs) for performing the above-mentioned method for updating the shooting position, and can solve the technical problem of inaccurate current meeting record information. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided by this application are the same as those of the method for updating the shooting position provided by the above embodiments, and will not be elaborated here.
[0100] The above are only some embodiments of the present application, and thus do not limit the patent scope of the present application. Any equivalent structural transformation made under the technical concept of the present application by using the content of the specification and drawings of the present application, or any direct / indirect application in other related technical fields is included in the patent protection scope of the present application.
Claims
1. A method for updating a shooting position, characterized in that, The method for updating the shooting position includes: Determine a predicted shooting direction according to the sound source direction of the conference audio; Collect a target image corresponding to the predicted shooting direction, determine a target speaker according to the conference mode corresponding to the conference audio, and determine a target position of the target speaker in the target image; Adjust the current shooting direction according to the target position.
2. The method for updating the shooting position according to claim 1, wherein The step of determining the predicted shooting direction according to the sound source direction of the conference audio includes: Obtain historical conference data corresponding to the conference mode; Determine the predicted shooting direction based on a network model trained based on the historical conference data and the sound source direction.
3. The method for updating the shooting position according to any one of claims 1 to 2, characterized in that The step of adjusting the current shooting direction according to the target position includes: Determine a position change amount of the target position according to the target position and inertial sensing data; Determine a motion parameter of the video module based on the position change amount, and adjust the current shooting direction of the video module based on the motion parameter.
4. The method for updating the shooting position according to claim 3, wherein After the step of determining the position change amount of the target position according to the target position and inertial sensing data, the method for updating the shooting position further includes: Determine a shooting mode corresponding to the conference mode; Determine the motion parameters of each video module according to the shooting mode and the position change amount; Adjust the current shooting direction of each video module based on the motion parameter.
5. The method for updating the shooting position according to claim 3, wherein After the step of controlling the operation of the video module based on the motion parameter to adjust the current shooting direction, the method for updating the shooting position further includes: Obtain an offset angle corresponding to the inertial sensing data; If the offset angle is greater than a preset angle, execute the step of determining the position change amount of the target position according to the target position and inertial sensing data.
6. The method for updating the shooting position according to claim 1, characterized in that, The step of collecting a target image corresponding to the predicted shooting direction, determining a target speaker according to the conference mode corresponding to the conference audio, and determining a target position of the target speaker in the target image includes: Collect the target image corresponding to the preset shooting direction, and determine the speech weight of the speaker corresponding to the conference audio according to the conference mode; Determine the target speaker based on the speech weight; Obtain the target position of the target speaker in the target image.
7. The method for updating the shooting position according to claim 6, wherein The step of collecting the target image corresponding to the preset shooting direction and determining the speech weight of the speaker corresponding to the conference audio according to the conference mode includes: Collect the target image corresponding to the preset shooting direction, and determine the historical conference data and shooting direction adjustment data corresponding to the conference mode; Fuse the current conference data, the historical conference data, and the shooting direction adjustment data to obtain a fused feature, and calculate the speech weight corresponding to the speaker in the fused feature according to the multi-head attention algorithm.
8. The method for updating the shooting position according to claim 1, characterized in that, After the step of adjusting the current shooting direction according to the target position, the method for updating the shooting position further includes: Continuously collect target audio information and target image information of the target speaker; Generate the meeting record information of the target shooting person according to the target audio information and the target image information.
9. An image acquisition device, characterized in that, The image acquisition device includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, and the computer program is configured to implement the steps of the shooting position updating method according to any one of claims 1 to 8.
10. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the shooting position updating method according to any one of claims 1 to 8 are implemented.