Sound-vision collaborative focusing method, device and equipment based on cross-modal distillation
The audio-visual data is obtained simultaneously by photographing equipment and the focus prediction is carried out using the cross-modal distillation model, which solves the shortcomings of the existing focus technology in dynamic target positioning, and achieves a more efficient focus effect, which is suitable for film and television shooting and conference records.
Patent Information
- Application Number
- CN202510566175.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-08-05
AI Technical Summary
The existing automatic focus technology mainly relies on visual data, lacks cross-modal collaboration capabilities, and it is difficult to accurately locate dynamic targets, especially in film and television shooting and conference record scenarios, which are prone to miss selection of non-subject targets.
The photography equipment is used to obtain the audio and visual data simultaneously, including sound signals and continuous image sequences, extract the audio and visual features and input the pre-trained cross-modal distillation model for focus prediction, and dynamically adjust the focal length of the photography equipment.
It significantly improves the intelligence of focus and scene relevance, can quickly and accurately locate dynamic targets, and improves the shooting effect of film and television shooting and conference records.
Smart Images

Figure CN120434504A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of photography technology, and in particular to a method, device and equipment for audio-visual collaborative focusing based on cross-modal distillation. Background Art
[0002] Autofocus technology is one of the key functions of modern imaging equipment. With breakthroughs in electronic technology and artificial intelligence, modern autofocus systems have integrated high-precision sensors, intelligent algorithms and high-speed communication interfaces, and can analyze shooting scenes in real time and dynamically adjust focus.
[0003] In film and television, autofocus technology has evolved from initially assisting with static scene capture to supporting high-speed motion capture and precise focusing in complex lighting conditions. In conference recording, from early fixed-focus cameras to today's conference systems supporting intelligent facial recognition, autofocus has become a core feature for improving recording efficiency. However, when shooting multiple people or recording meetings, the system may mistakenly select non-subjective subjects. Furthermore, the system lacks the ability to quickly respond to dynamically moving objects, resulting in low autofocus accuracy. Summary of the Invention
[0004] The main technical problem solved by the present invention is that the existing automatic focusing technology mainly relies on visual data and lacks cross-modal collaboration capabilities, making it difficult to accurately locate dynamic targets in film and television shooting and meeting recording scenarios.
[0005] According to the first aspect, an embodiment provides an audio-visual collaborative focusing method based on cross-modal distillation, including:
[0006] Synchronously acquiring audio and visual data of a target user using a photographic device; wherein the audio and visual data includes an audio signal and a sequence of continuous images; the photographic device includes a microphone array and a camera, wherein the microphone array is used to collect the audio signal, and the camera is used to capture the sequence of continuous images;
[0007] Extracting audio-visual features corresponding to the audio-visual data; wherein the audio-visual features include the sound source direction of the sound signal, the image saliency of the continuous image sequence, and the location of the target user;
[0008] Inputting the audio-visual features into a pre-trained cross-modal distillation model to obtain a focus prediction result of the target user;
[0009] The focal length of the photographic device is adjusted according to the focus prediction result.
[0010] In some embodiments, the pre-trained cross-modal distillation model is trained in the following manner:
[0011] Acquire a training data set; wherein the training data set includes training sound signals and training image sequences of sample users collected synchronously under different environmental conditions, wherein the environmental conditions include the type of sound source, the lighting conditions of the sample users, and the distance between the photographic device and the sample users;
[0012] Extracting a training feature set corresponding to the training data set; wherein the training feature set includes the sound source direction of the training sound signal, the image saliency of the training image sequence, and the position of the sample user;
[0013] The cross-modal distillation model to be trained is trained according to the training feature set to obtain a trained cross-modal distillation model.
[0014] In some embodiments, the cross-modal distillation model to be trained includes a teacher model, a student model, and a distillation network; and training the cross-modal distillation model to be trained according to the training feature set to obtain a trained cross-modal distillation model includes:
[0015] Inputting the training feature set into the teacher model to obtain a first focus prediction result of the teacher model, inputting the training feature set into the student model to obtain a second focus prediction result of the student model, and inputting the training feature set into the distillation network to obtain a third focus prediction result;
[0016] Calculating a supervision loss based on the second focus prediction result and a preset focus reference result, calculating a distillation loss based on the first focus prediction result and the second focus prediction result, calculating a consistency loss of the distillation network based on the first focus prediction result and the third focus prediction result, and constructing a final loss function based on the supervision loss, the distillation loss, and the consistency loss;
[0017] Calculate the gradient of the final loss function with respect to the model parameters of the student model, and update the model parameters of the student model according to the gradient and the optimizer until the student model converges or reaches a preset number of iterations, and use the corresponding cross-modal distillation model as the trained cross-modal distillation model.
[0018] In some embodiments, the distillation network adopts a cross-modal attention distillation mechanism based on the Transformer architecture.
[0019] In some embodiments, after synchronously acquiring the audio and visual data of the target user using a photographic device, the audio and visual collaborative focusing method further includes:
[0020] Determining the scene type of the target user based on the sound signal and the continuous image sequence; wherein the scene type includes a meeting scene, a film and television shooting scene, and an outdoor dynamic scene;
[0021] Before inputting the audio-visual features into a pre-trained cross-modal distillation model, the audio-visual collaborative focusing method further includes:
[0022] Adjust model parameters of the pre-trained cross-modal distillation model according to the scenario type.
[0023] In some embodiments, the sound source direction of the sound signal includes the angle and distance of the sound source located by the microphone array and the preset signal localization algorithm, the image saliency of the continuous image sequence includes edge features, texture features and salient features, and the position of the target user is determined based on the fusion of the sound source direction and the image saliency based on the attention mechanism.
[0024] In some embodiments, the photographic device includes a single camera or multiple cameras; and the audio-visual collaborative focusing method further includes:
[0025] Using a photographic device including a single camera to capture a sequence of continuous images of the target user; or;
[0026] A photographic device including multiple cameras is used to capture a sequence of continuous images of the target user at different viewing angles.
[0027] According to the second aspect, an embodiment provides an audio-visual collaborative focusing device based on cross-modal distillation, comprising:
[0028] a data acquisition module, configured to synchronously acquire audio and visual data of a target user using a photographic device; wherein the audio and visual data includes audio signals and a sequence of continuous images, and the photographic device includes a microphone array and a camera, wherein the microphone array is configured to acquire the audio signals, and the camera is configured to capture the sequence of continuous images;
[0029] a feature extraction module, configured to extract audio-visual features corresponding to the audio-visual data; wherein the audio-visual features include the direction of the sound source of the sound signal, the image saliency of the continuous image sequence, and the location of the target user;
[0030] A focus prediction module, configured to input the audio-visual features into a pre-trained cross-modal distillation model to obtain a focus prediction result of the target user;
[0031] A focal length adjustment module is used to adjust the focal length of the photographic device according to the focus prediction result.
[0032] According to a third aspect, an embodiment provides an audio-visual collaborative focusing device based on cross-modal distillation, including:
[0033] Memory, used to store programs;
[0034] A processor is used to implement the audio-visual collaborative focusing method by executing the program stored in the memory.
[0035] According to a fourth aspect, an embodiment provides a computer program product, comprising a computer program and / or instructions, which implement the audio-visual collaborative focusing method when executed by a processor.
[0036] According to the above-mentioned embodiments of the audio-visual collaborative focusing method, apparatus, device, and computer program product based on cross-modal distillation, the audio-visual data of the target user are synchronously acquired using a photographic device, and the audio-visual features corresponding to the audio-visual data are extracted. The focus prediction result of the target user can be predicted by combining the audio-visual features with the cross-modal distillation model. When the cross-modal distillation model is used for prediction, the dynamic target user can be located using sound information. Through the cross-modal collaborative capability of the sound signal and the continuous image sequence, the accurate association of the audio-visual data is achieved. This significantly improves the intelligence and scene relevance of focusing, and is suitable for the dynamic needs of film and television shooting and conference recording photography. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 This is a flowchart of an audio-visual collaborative focusing method based on cross-modal distillation according to an embodiment of the present application;
[0038] Figure 2 A flowchart of a method for training a pre-trained cross-modal distillation model according to an embodiment;
[0039] Figure 3 A flowchart of an embodiment of training a cross-modal distillation model to be trained according to a training feature set to obtain a trained cross-modal distillation model;
[0040] Figure 4 Schematic diagram of the structure of an audio-visual collaborative focusing device based on cross-modal distillation in one embodiment. DETAILED DESCRIPTION
[0041] The present invention will be further described in detail below by means of specific embodiments in conjunction with the accompanying drawings. Similar elements in different embodiments are numbered with associated similar elements. In the following embodiments, many detailed descriptions are provided to enable the present application to be better understood. However, those skilled in the art will readily appreciate that some of the features may be omitted in different circumstances, or may be replaced by other elements, materials, or methods. In some cases, some operations related to the present application are not shown or described in the specification. This is to avoid the core portion of the present application being overwhelmed by excessive descriptions, and for those skilled in the art, it is not necessary to describe these related operations in detail. They will fully understand the related operations based on the description in the specification and the general technical knowledge in the art.
[0042] In addition, the features, operations, or characteristics described in the specification may be combined in any appropriate manner to form various embodiments. Furthermore, the steps or actions in the method description may be reordered or adjusted in a manner readily apparent to those skilled in the art. Therefore, the various sequences in the specification and drawings are provided solely for the purpose of clearly describing a particular embodiment and are not intended to be mandatory, unless otherwise specified.
[0043] The serial numbers assigned to components herein, such as "first," "second," etc., are used solely to distinguish the objects being described and do not convey any sequential or technical meaning. References to "connection" and "coupling" herein, unless otherwise specified, include both direct and indirect connections (couplings).
[0044] Existing autofocus technologies rely primarily on visual data and struggle to leverage sound to locate dynamic targets. This is particularly true in scenarios requiring sound source correlation, such as filming or conference recording. Traditional methods lack cross-modal collaboration capabilities. For example, in multi-person filming or conference recording, the system may mistakenly select non-subjective targets. For example, when a reporter raises their hand, the camera may mistakenly focus on the reporter's arm instead of the speaker's face.
[0045] In order to improve the intelligence and scene relevance of focusing, and to be applicable to the dynamic needs in film and television shooting and conference record photography, an embodiment of the present application provides an audio-visual collaborative focusing method based on cross-modal distillation. In this audio-visual collaborative focusing method, a photographic device is used to synchronously obtain the audio-visual data of the target user; wherein the audio-visual data includes a sound signal and a continuous image sequence, and the photographic device includes a microphone array and a camera, the microphone array is used to collect the sound signal, and the camera is used to shoot the continuous image sequence; the audio-visual features corresponding to the audio-visual data are extracted; wherein the audio-visual features include the sound source direction of the sound signal, the image saliency of the continuous image sequence and the position of the target user; the audio-visual features are input into a pre-trained cross-modal distillation model to obtain the focus prediction result of the target user; and the focal length of the photographic device is adjusted according to the focus prediction result.
[0046] The following describes the audio-visual collaborative focusing method based on cross-modal distillation provided in an embodiment of the present application in conjunction with the accompanying drawings.
[0047] Figure 1 A flowchart of a method for audio-visual collaborative focusing based on cross-modal distillation provided in an embodiment of the present application is shown, and is described in detail below:
[0048] Step S10: synchronously obtain the audio and video data of the target user using a photographic device.
[0049] Specifically, the photographic device includes a microphone array and a camera. The microphone array is used to collect sound signals, and the camera is used to capture a continuous image sequence. Therefore, the microphone array and camera in the photographic device are used to synchronously obtain the target user's sound signal and the continuous image sequence.
[0050] For example, in a meeting recording scenario, when a speaker is speaking, the microphone array focuses on the speaker's voice signal through beamforming, and at the same time, a continuous image sequence containing the speaker is obtained through the camera.
[0051] In this embodiment, the target user's audio and visual data are acquired synchronously, ensuring a one-to-one correspondence between the recorded target user's voice signal and the sequence of continuously captured images. For example, in a meeting, when a speaker discusses a technical solution, their voice data is synchronized with the image data, making it easier to later reconstruct the specific scene of the voice discussion through the image. This improves information utilization in practical application scenarios and is suitable for the dynamic needs of filming, recording, and recording meetings.
[0052] Step S20: extracting audio-visual features corresponding to the audio-visual data.
[0053] Specifically, the audio-visual features corresponding to the audio-visual data are extracted, wherein the audio-visual features include the sound source direction of the sound signal, the image saliency of the continuous image sequence, and the position of the target user. The sound source direction of the sound signal refers to the time difference or phase difference calculated by the microphone array for the sound signal of the target user to reach each microphone to determine the azimuth and pitch angles and distance of the sound source relative to the array. Image saliency is the extraction of high-attention areas by calculating the visual saliency (such as color contrast, edge strength, motion characteristics) of pixels or areas in the continuous image sequence. The position of the target user refers to the three-dimensional coordinates of the target user in space.
[0054] For example, in a conference recording scenario, the microphone array detects that when a speaker speaks, the sound source direction is "30° in front, 1 meter away" (corresponding to the speaker's position relative to the camera). The speaker's position is x = 2.5m, y = 1.8m, z = 1.2m, and is dynamically updated based on the speaker's movement. When the speaker points to specific data, if the saliency score of a certain area in the continuous image sequence is higher than the background, the feature of the pointed area can be used as image saliency.
[0055] In an embodiment of the present application, by obtaining the direction of the sound source and combining it with the continuous image sequence and position of the speaker, the sound source and the speaker's position can be associated to achieve accurate positioning of the speaker. Combined with the saliency of the speaker's image, the speaker's behavior and the areas that need to be focused on can be clarified, thereby improving the accuracy of subsequent focusing.
[0056] Step S30: Input the audio-visual features into a pre-trained cross-modal distillation model to obtain the focus prediction result of the target user.
[0057] Specifically, the cross-modal distillation model is a model training technique that transfers knowledge from one modality to another, improving the performance of the target modality model by leveraging the complementary information between different modalities. The architecture of the cross-modal distillation model consists of a teacher model and a student model. The teacher model is typically a multimodal model that can process both acoustic and visual information, while the student model is a unimodal model (e.g., only processes acoustic or visual information). During training, the teacher model extracts high-level semantic features from bimodal data and transfers this knowledge to the student model via a distillation loss function.
[0058] For example, in a video classification task, the teacher model can learn a joint representation of scenes and events from video frames and audio, and then distill this knowledge into a student model that only processes audio, thereby improving the accuracy of video classification.
[0059] In an embodiment of the present application, a cross-modal distillation model can improve the performance of the target modality model by leveraging the complementary information between different modalities. Since data from different modalities have different sensitivities to noise and interference, the cross-modal distillation model can learn more robust feature representations from multiple different modalities, improving its performance in complex environments.
[0060] Step S40: adjusting the focal length of the photographic device according to the focus prediction result.
[0061] Specifically, the focus prediction result can be a focus position or a focus parameter. The focus position refers to the coordinates or area range of the area in the picture that needs to be clearly imaged, while the focus parameter is a parameter used to describe the imaging characteristics of the focus area and is used to optimize the focusing effect.
[0062] In the embodiments of this application, dynamically adjusting the focal length of a camera based on focus prediction results can significantly improve shooting quality and efficiency, enabling precise focusing. Focus prediction results allow for real-time adjustment of the focal length, eliminating the need for manual adjustment. This allows for rapid location of dynamic target users in scenarios requiring sound source correlation, such as filming or recording meetings.
[0063] In an embodiment of the present application, a camera is used to synchronously acquire the target user's audio and visual data, extract the corresponding audio and visual features, and combine these features with a cross-modal distillation model to predict the target user's focus. This allows the cross-modal distillation model to accurately locate the dynamic target user using sound information, and through the cross-modal collaboration between sound signals and continuous image sequences, accurately associate the audio and visual data. This significantly improves the intelligence and scene relevance of focusing, making it suitable for the dynamic needs of film and television shooting and conference recording photography.
[0064] In some embodiments, please refer to Figure 2 The pre-trained cross-modal distillation model is trained by following steps S31 to S33:
[0065] Step S31: Obtain a training data set.
[0066] Specifically, the training dataset includes training audio signals and training image sequences of sample users collected simultaneously under different environmental conditions. Environmental conditions include the type of sound source, the lighting conditions of the sample users, and the distance between the camera and the sample users. The training dataset is sourced from meeting recordings or film and television shooting scenes.
[0067] For example, in a conference scenario, the sound source types include direct human voice sources, device-assisted sound sources, and environmental interference sound sources. Direct human voice sources are the original voices actively produced by participants, device-assisted sound sources are human voices enhanced or converted by electronic devices, and environmental interference sound sources refer to non-target sound sources that may affect the quality of the meeting. The lighting conditions of the sample users can be strong light environments, such as direct sunlight or high-intensity light sources, or weak light environments, such as insufficient light or low-light scenes, or evenly lit environments, such as scenes with evenly distributed optical fibers and no obvious shadows. The distance between the photographic equipment and the sample users can be an extremely close distance of 1-10 cm for macro photography of insects or plant details, a moderate distance of 1-3 meters for standard photography of portraits or daily scenery, or a long distance of 10-100 meters for long-distance photography of scenery or wildlife.
[0068] In the embodiment of the present application, by obtaining a diverse training data set to train the model, the generalization ability and robustness of the model can be improved.
[0069] Step S32: extracting a training feature set corresponding to the training data set.
[0070] Specifically, the training feature set includes the sound source direction of the training sound signal, the image saliency of the training image sequence, and the position of the sample user.
[0071] Step S33: Train the cross-modal distillation model to be trained according to the training feature set to obtain a trained cross-modal distillation model.
[0072] Specifically, the cross-modal distillation model to be trained is trained according to the training feature set to obtain a trained cross-modal distillation model, and the trained cross-modal distillation model is used to perform focus prediction to obtain an accurate focus prediction result.
[0073] For example, the focus prediction results of the cross-modal distillation model to be trained are evaluated by preset evaluation indicators to obtain evaluation results, and the model is trained and optimized based on the evaluation results. The preset evaluation indicators include focus prediction error and focusing success rate. The focus prediction error includes but is not limited to two-dimensional focus prediction error, three-dimensional focus prediction error and normalized focus error. The two-dimensional focus prediction error refers to the deviation in two-dimensional space between the predicted focus position in the focus prediction result and the true focus position in the preset focus reference position, and the Euclidean distance between the two can be calculated. The three-dimensional focus prediction error refers to the deviation in three-dimensional space between the predicted focus position in the focus prediction result and the true focus position in the preset focus reference position. The normalized focus error refers to scaling the error between the predicted focus position in the focus prediction result and the true focus position in the preset focus reference position to between 0 and 1 or a specific range to eliminate the dimensional effect. The focusing success rate is used to evaluate the proportion of successfully predicted true focus positions in the focus prediction result.
[0074] In an embodiment of the present application, the cross-modal distillation model to be trained is optimized and trained through a training feature set, covering conference, film and television, and outdoor dynamic scenes, and is particularly suitable for dynamic requirements in film and television shooting and conference recording photography.
[0075] In some embodiments, please refer to Figure 3 , the cross-modal distillation model to be trained includes a teacher model, a student model and a distillation network; step S33: training the cross-modal distillation model to be trained according to the training feature set to obtain a trained cross-modal distillation model, including steps S331 to S333, which are described in detail below.
[0076] Step S331: Input the training feature set into the teacher model to obtain the first focus prediction result of the teacher model, input the training feature set into the student model to obtain the second focus prediction result of the student model, input the training feature set into the distillation network to obtain the third focus prediction result.
[0077] Specifically, the teacher model provides high-quality focus prediction results as the learning goal of the student model. The teacher model is usually a pre-trained high-performance model, the student model is a lightweight model, and the distillation network is a Transformer-based cross-modal attention network.
[0078] Step S332: Calculate the supervision loss based on the second focus prediction result and the preset focus reference result, calculate the distillation loss based on the first focus prediction result and the second focus prediction result, calculate the consistency loss of the distillation network based on the first focus prediction result and the third focus prediction result, and construct the final loss function based on the supervision loss, distillation loss and consistency loss.
[0079] Specifically, the cross-entropy loss is calculated based on the second focal prediction result and the preset focal reference result as the supervision loss. The KL divergence is used to measure the difference in the focal prediction results of the teacher model and the student model, that is, the distillation loss of the first and second focal prediction results is calculated. The consistency loss is used to measure the consistency of the focal prediction results of the teacher model and the distillation network, that is, the consistency loss of the distillation network is calculated based on the first and third focal prediction results. Combining the above losses and controlling the weight of each loss through different hyperparameters, the final loss function is obtained.
[0080] In the embodiments of the present application, the output of the teacher model is generally more stable and can guide the student model to converge quickly. The distillation network fuses sound and image features through a cross-modal attention mechanism to generate additional supervisory signals. At the same time, the output of the distillation model can serve as a supplementary knowledge source for the student model, further improving the performance of the student model. By minimizing the final loss function, the student model gradually approaches the performance of the teacher model and the distillation network. The lightweight design of the student model makes it more suitable for deployment in resource-constrained environments.
[0081] Step S333: Calculate the gradient of the final loss function with respect to the model parameters of the student model, and update the model parameters of the student model according to the gradient and the optimizer until the student model converges or reaches a preset number of iterations, and use the corresponding cross-modal distillation model as the trained cross-modal distillation model.
[0082] In this embodiment of the present application, by jointly training the teacher model, student model, and distillation network, the information in the training dataset can be fully utilized to improve the generalization ability of the model. The consistency loss of the distillation network helps to ensure that the focus prediction results of the teacher model and the distillation network are consistent, further improving the learning effect of the student model.
[0083] In some embodiments, the distillation network adopts a cross-modal attention distillation mechanism based on the Transformer architecture.
[0084] In the embodiment of the present application, the distillation network adopts a cross-modal attention distillation mechanism based on the Transformer architecture, which can improve the model performance through multimodal information fusion, especially in terms of robustness, generalization ability and reasoning efficiency in complex scenarios.
[0085] In some embodiments, after synchronously acquiring the audio and visual data of the target user using a photographic device, the audio and visual collaborative focusing method further includes:
[0086] The scene type of the scene where the target user is located is determined based on the sound signal and the continuous image sequence.
[0087] Specifically, by combining deep learning technology, the target user's voice signal and scene characteristics (such as sound source type or image content) in the continuous image sequence are analyzed to determine the scene type of the target user. Scene types include meeting scenes, film and television shooting scenes, and outdoor dynamic scenes.
[0088] In some embodiments, before inputting the audio-visual features into a pre-trained cross-modal distillation model, the audio-visual collaborative focusing method further includes:
[0089] Adjust the model parameters of the pre-trained cross-modal distillation model according to the scenario type.
[0090] Specifically, the model parameters are adaptively adjusted based on the scene type, so that the focus prediction results output by the cross-modal distillation model after the adjustment of the model parameters are more scene-adaptive and suitable for the dynamic requirements in film and television shooting and conference recording photography.
[0091] For example, in a meeting scene, the focus is prioritized on the speaker area, and in a film and television scene, the focus is adjusted to the dynamic subject according to the direction of the sound.
[0092] In an embodiment of the present application, the scene type obtained through analysis provides cross-modal support for focus prediction, thereby improving the focusing effect.
[0093] In some embodiments, the focus strategy is adjusted based on the basis of scene switching and focus adjustment, for example, sound source change: adjust the focus target by detecting changes in audio direction; image saliency change: combine visual analysis to optimize focus area selection; shooting task type: adjust the focus strategy according to preset parameters (such as sound source priority, picture priority).
[0094] In some embodiments, the sound source direction of a sound signal includes the angle and distance of the sound source as determined by a microphone array and a preset signal localization algorithm. The image saliency of a continuous image sequence includes edge features, texture features, and salient features. The location of the target user is determined based on an attention mechanism that fuses the sound source direction and image saliency. The preset signal localization algorithm may include a time difference calculation-based algorithm or a beamforming algorithm.
[0095] Specifically, the attention mechanism learns the credibility of the sound source direction and image saliency, analyzes the spatial distribution of the target user in the scene, and dynamically assigns weights to determine the target user's location. For example, in noisy environments, the weight of the sound source direction is reduced, while in low-light environments, the weight of image saliency is increased.
[0096] In the embodiment of the present application, the sound source direction and image saliency are fused through the attention mechanism, which can dynamically balance the reliability of the two modalities and significantly improve the positioning accuracy in complex scenarios.
[0097] In some embodiments, the photographic device may include a single camera or multiple cameras. When photographing a target user using a photographic device including multiple cameras, a continuous sequence of images of the target user from different perspectives may be captured. A single camera is suitable for simple meeting scenarios, while multiple cameras are suitable for complex filming or outdoor dynamic scenes to enhance the robustness of the image data.
[0098] For example, in film and television shooting scenarios, a camera system consisting of multiple cameras is used to capture dynamic sound sources, such as actors' dialogue. These cameras can be divided into a primary camera and a secondary camera. The primary camera adjusts focus based on the sound source, while the secondary camera provides perspective supplementation and saliency correction.
[0099] In an embodiment of the present application, focus prediction can be collaboratively optimized by synchronously capturing a multi-perspective continuous image sequence and multi-channel audio.
[0100] In some embodiments, to ensure focusing accuracy, the sound signal and the continuous image sequence need to meet the following conditions: (1) The audio signal-to-noise ratio is high enough to support sound source localization. (2) The image resolution meets a minimum threshold to ensure the accuracy of visual feature extraction. (3) The ambient lighting is maintained within a set range to avoid image quality degradation. Therefore, the imaging conditions for obtaining audio-visual data include: the audio signal-to-noise ratio, image resolution, and ambient lighting meet a minimum threshold to ensure focusing effect.
[0101] In some embodiments, a cross-modal distillation-based audio-visual collaborative focusing method is applicable to video cameras, smart conferencing equipment, and filming equipment. By embedding the cross-modal distillation model into the device chip and combining it with a microphone array and camera for real-time computation, this method increases focusing response speed by over 40%, significantly improving shooting quality in sound-source-dependent scenarios.
[0102] In some embodiments, the audio-visual collaborative focusing method based on cross-modal distillation provided in the embodiments of the present application can be applied in the following scenarios. Scenario 1: Speaker shooting for conference recording, shooting the speaker in the conference room, locating the sound direction through the microphone array, adjusting the focus to the speaker's face in combination with the image data, and generating a clear picture. Scenario 2: Dynamic scene tracking for film and television shooting, when shooting dialogue scenes, analyzing the actor's voice and movement, dynamically focusing on the sound source body, ensuring picture continuity and detail integrity. Scenario 3: Outdoor recording for multi-camera systems, using multi-camera equipment to shoot outdoor activities, collaboratively processing multi-channel audio and multi-perspective images, and optimizing the focus to the main sound source area.
[0103] Corresponding to the above-mentioned audio-visual collaborative focusing method based on cross-modal distillation in the above embodiment, Figure 4A structural schematic diagram of an audio-visual collaborative focusing device based on cross-modal distillation provided in an embodiment of the present application is shown. For ease of explanation, only the parts related to the embodiment of the present application are shown.
[0104] Please refer to Figure 4 The audio-visual collaborative focusing device based on cross-modal distillation includes a data acquisition module 10, a feature extraction module 20, a focus prediction module and a focal length adjustment module, which are described in detail below.
[0105] The data acquisition module 10 is used to synchronously acquire the audio and visual data of the target user using a photographic device; wherein the audio and visual data include sound signals and a continuous image sequence, and the photographic device includes a microphone array and a camera, the microphone array is used to collect sound signals, and the camera is used to shoot a continuous image sequence.
[0106] The feature extraction module 20 is used to extract audio-visual features corresponding to the audio-visual data; wherein the audio-visual features include the sound source direction of the sound signal, the image saliency of the continuous image sequence, and the location of the target user.
[0107] The focus prediction module 30 is used to input the audio-visual features into a pre-trained cross-modal distillation model to obtain the focus prediction result of the target user.
[0108] The focal length adjustment module 40 is used to adjust the focal length of the photographic device according to the focus prediction result.
[0109] In some embodiments, the focus prediction module 30 is further used to train a cross-modal distillation model to be trained:
[0110] Obtain a training data set; wherein the training data set includes training sound signals and training image sequences of sample users collected synchronously under different environmental conditions, and the environmental conditions include the type of sound source, the lighting conditions of the sample users, and the distance conditions between the photographic equipment and the sample users.
[0111] A training feature set corresponding to the training data set is extracted; wherein the training feature set includes the sound source direction of the training sound signal, the image saliency of the training image sequence, and the position of the sample user.
[0112] The cross-modal distillation model to be trained is trained according to the training feature set to obtain a trained cross-modal distillation model.
[0113] In some embodiments, the cross-modal distillation model to be trained includes a teacher model, a student model, and a distillation network; the focus prediction module 30 is further configured to train the cross-modal distillation model to be trained based on the training feature set to obtain a trained cross-modal distillation model, including:
[0114] The training feature set is input into the teacher model to obtain the first focus prediction result of the teacher model, the training feature set is input into the student model to obtain the second focus prediction result of the student model, and the training feature set is input into the distillation network to obtain the third focus prediction result.
[0115] The supervision loss is calculated based on the second focus prediction result and the preset focus reference result, the distillation loss is calculated based on the first focus prediction result and the second focus prediction result, the consistency loss of the distillation network is calculated based on the first focus prediction result and the third focus prediction result, and the final loss function is constructed based on the supervision loss, distillation loss and consistency loss.
[0116] Calculate the gradient of the final loss function with respect to the model parameters of the student model, and update the model parameters of the student model according to the gradient and the optimizer until the student model converges or reaches the preset number of iterations. Then, use the corresponding cross-modal distillation model as the trained cross-modal distillation model.
[0117] In some embodiments, the distillation network adopts a cross-modal attention distillation mechanism based on the Transformer architecture.
[0118] In some embodiments, after synchronously acquiring the audio and visual data of the target user using a photographic device, the data acquisition module 10 is further configured to:
[0119] The scene type of the target user's scene is determined based on the sound signal and the continuous image sequence; wherein the scene types include meeting scenes, film and television shooting scenes, and outdoor dynamic scenes.
[0120] In some embodiments, before inputting the audio-visual features into a pre-trained cross-modal distillation model, the focus prediction module 30 is further configured to:
[0121] Adjust the model parameters of the pre-trained cross-modal distillation model according to the scenario type.
[0122] In some embodiments, the sound source direction of the sound signal includes the angle and distance of the sound source located by a microphone array and a preset signal localization algorithm. The image saliency of the continuous image sequence includes edge features, texture features, and salient features. The position of the target user is determined based on the attention mechanism that fuses the sound source direction and image saliency.
[0123] In some embodiments, the photographic device includes a single camera or multiple cameras; the data acquisition module 10 is further configured to:
[0124] capturing a sequence of continuous images of the target user using a photographic device comprising a single camera; or
[0125] A photographic device including multiple cameras is used to capture a sequence of continuous images of a target user at different viewing angles.
[0126] In some embodiments, the present application further provides an audio-visual collaborative focusing device based on cross-modal distillation, comprising:
[0127] Memory, used to store programs;
[0128] The processor is used to implement the audio-visual collaborative focusing method by executing the program stored in the memory.
[0129] In some embodiments, the present application also provides a computer program product, including a computer program and / or instructions, which implements the audio-visual collaborative focusing method when the computer program and / or instructions are executed by a processor.
[0130] Those skilled in the art will appreciate that all or part of the functions of the various methods in the above embodiments can be implemented by hardware or by computer program. When all or part of the functions in the above embodiments are implemented by computer program, the program can be stored in a computer-readable storage medium, and the storage medium can include: read-only memory, random access memory, disk, optical disk, hard disk, etc., and the program is executed by a computer to implement the above functions. For example, the program is stored in the memory of the device, and when the program in the memory is executed by the processor, all or part of the above functions can be implemented. In addition, when all or part of the functions in the above embodiments are implemented by computer program, the program can also be stored in a storage medium such as a server, another computer, disk, optical disk, flash disk or mobile hard disk, and saved in the memory of the local device by downloading or copying, or the system of the local device is updated. When the program in the memory is executed by the processor, all or part of the functions in the above embodiments can be implemented.
[0131] The above examples are used to illustrate the present invention, which are only used to help understand the present invention and are not intended to limit the present invention. Those skilled in the art can make several simple deductions, modifications or substitutions based on the concept of the present invention.
Claims
1. A method for audio-visual collaborative focusing based on cross-modal distillation, characterized in that: include: Synchronously acquiring audio and visual data of a target user using a photographic device; wherein the audio and visual data includes an audio signal and a sequence of continuous images; the photographic device includes a microphone array and a camera, wherein the microphone array is used to collect the audio signal, and the camera is used to capture the sequence of continuous images; Extracting audio-visual features corresponding to the audio-visual data; wherein the audio-visual features include the sound source direction of the sound signal, the image saliency of the continuous image sequence, and the location of the target user; Inputting the audio-visual features into a pre-trained cross-modal distillation model to obtain a focus prediction result of the target user; The focal length of the photographic device is adjusted according to the focus prediction result.
2. The audio-visual collaborative focusing method according to claim 1, wherein: The pre-trained cross-modal distillation model is trained in the following way: Acquire a training data set; wherein the training data set includes training sound signals and training image sequences of sample users collected synchronously under different environmental conditions, wherein the environmental conditions include the type of sound source, the lighting conditions of the sample users, and the distance between the photographic device and the sample users; Extracting a training feature set corresponding to the training data set; wherein the training feature set includes the sound source direction of the training sound signal, the image saliency of the training image sequence, and the position of the sample user; The cross-modal distillation model to be trained is trained according to the training feature set to obtain a trained cross-modal distillation model.
3. The audio-visual collaborative focusing method according to claim 2, wherein: The cross-modal distillation model to be trained includes a teacher model, a student model, and a distillation network; and the cross-modal distillation model to be trained is trained according to the training feature set to obtain a trained cross-modal distillation model, including: Inputting the training feature set into the teacher model to obtain a first focus prediction result of the teacher model, inputting the training feature set into the student model to obtain a second focus prediction result of the student model, and inputting the training feature set into the distillation network to obtain a third focus prediction result; Calculating a supervision loss based on the second focus prediction result and a preset focus reference result, calculating a distillation loss based on the first focus prediction result and the second focus prediction result, calculating a consistency loss of the distillation network based on the first focus prediction result and the third focus prediction result, and constructing a final loss function based on the supervision loss, the distillation loss, and the consistency loss; Calculate the gradient of the final loss function with respect to the model parameters of the student model, and update the model parameters of the student model according to the gradient and the optimizer until the student model converges or reaches a preset number of iterations, and use the corresponding cross-modal distillation model as the trained cross-modal distillation model.
4. The audio-visual collaborative focusing method according to claim 3, wherein: The distillation network adopts a cross-modal attention distillation mechanism based on the Transformer architecture.
5. The audio-visual collaborative focusing method according to claim 1, wherein: After synchronously acquiring the audio and visual data of the target user using the photographic device, the audio and visual collaborative focusing method further includes: Determining the scene type of the target user based on the sound signal and the continuous image sequence; wherein the scene type includes a meeting scene, a film and television shooting scene, and an outdoor dynamic scene; Before inputting the audio-visual features into a pre-trained cross-modal distillation model, the audio-visual collaborative focusing method further includes: Adjust model parameters of the pre-trained cross-modal distillation model according to the scenario type.
6. The audio-visual collaborative focusing method according to claim 1, wherein: The sound source direction of the sound signal includes the angle and distance of the sound source located by the microphone array and the preset signal localization algorithm. The image saliency of the continuous image sequence includes edge features, texture features and salient features. The position of the target user is determined based on the attention mechanism by fusing the sound source direction and the image saliency.
7. The audio-visual collaborative focusing method according to claim 1, wherein: The photographic device includes a single camera or multiple cameras; The audio-visual collaborative focusing method further includes: Using a photographic device including a single camera to capture a sequence of continuous images of the target user; or; A photographic device including multiple cameras is used to capture a sequence of continuous images of the target user at different viewing angles.
8. An audio-visual collaborative focusing device based on cross-modal distillation, characterized in that: include: a data acquisition module, configured to synchronously acquire audio and visual data of a target user using a photographic device; wherein the audio and visual data includes audio signals and a sequence of continuous images, and the photographic device includes a microphone array and a camera, wherein the microphone array is configured to acquire the audio signals, and the camera is configured to capture the sequence of continuous images; a feature extraction module, configured to extract audio-visual features corresponding to the audio-visual data; wherein the audio-visual features include the direction of the sound source of the sound signal, the image saliency of the continuous image sequence, and the location of the target user; A focus prediction module, configured to input the audio-visual features into a pre-trained cross-modal distillation model to obtain a focus prediction result of the target user; A focal length adjustment module is used to adjust the focal length of the photographic device according to the focus prediction result.
9. An audio-visual collaborative focusing device based on cross-modal distillation, characterized in that: include: Memory, used to store programs; A processor, configured to implement the audio-visual collaborative focusing method as described in any one of claims 1 to 7 by executing the program stored in the memory.
10. A computer program product comprising a computer program and / or instructions, characterized in that When the computer program and / or instructions are executed by a processor, the audio-visual collaborative focusing method according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Image acquisition method and device of intelligent glasses, intelligent glasses and storage medium
CN120751256A