Shooting control method, device, equipment, medium and program product

By using voice interaction and TWS earphones for coordinated control, hands-free shooting on mobile terminals is achieved, solving the problem of complex operation in existing technologies and improving user convenience and shooting efficiency.

CN122002122APending Publication Date: 2026-05-08SHENZHEN PHICOUSTIC SYST DEV CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHENZHEN PHICOUSTIC SYST DEV CO LTD
Filing Date
2025-12-26
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

The operation of existing mobile terminal camera photography is complicated and cannot meet users' needs for convenient shooting, which affects the user experience, especially when the operation accuracy and response speed are limited in motion or complex environments.

Method used

The system allows users to control TWS earbuds to take photos via voice interaction. It uses a command parsing model to analyze user voice commands and executes image capture through the TWS earbuds. The system also combines inertial sensor data to stitch and optimize the images, enabling hands-free operation.

Benefits of technology

It simplifies the shooting process, improves the convenience and shooting efficiency for users in various scenarios, and enhances the user experience, especially the convenience of shooting in motion and the integrity of the image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122002122A_ABST
    Figure CN122002122A_ABST
Patent Text Reader

Abstract

The invention discloses a shooting control method and device, equipment, a medium and a program product, and relates to the technical field of computers. The method comprises the following steps: acquiring target voice information of a user; inputting the target voice information into the instruction analysis model to obtain a target voice command corresponding to the target voice information; the instruction analysis model is obtained by training a to-be-trained model based on at least one to-be-trained voice sample, the to-be-trained voice sample comprises voice information and a first instruction label and a second instruction label corresponding to the voice information, the first instruction label is used for indicating whether an image needs to be shot, and the second instruction label is used for indicating an image shooting type; in response to the target voice command indicating that an image needs to be shot, sending an image shooting instruction corresponding to the target voice command to the TWS earphone, so that the TWS earphone shoots a scene image; and receiving the scene image sent by the TWS earphone. According to the invention, convenient shooting requirements of the user can be met, and the use experience of the user on the shooting function is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of computer technology, and in particular relates to a shooting control method, device, equipment, medium and program product. Background Technology

[0002] In recent years, computer technology has developed rapidly, and mobile terminals such as mobile phones and tablets have become indispensable electronic products in people's lives. Among the usage scenarios of these mobile terminals, image shooting functions have received widespread attention, and users have a strong demand for convenient shooting methods.

[0003] Currently, mobile devices are typically equipped with camera functions, and users can trigger shooting through physical buttons or virtual buttons on the mobile device.

[0004] However, the existing methods for operating the camera on mobile terminals are relatively complex, which cannot meet users' needs for convenient shooting and affects the user's experience of using the shooting function. Summary of the Invention

[0005] This application provides a shooting control method, device, equipment, medium, and program product that can meet users' shooting needs for convenience and improve users' experience of using shooting functions.

[0006] A first aspect of this application provides a shooting control method applied to a mobile terminal, comprising: Obtain the user's target voice information; The target speech information is input into the instruction parsing model to obtain the target speech command corresponding to the target speech information. The instruction parsing model is trained based on at least one speech sample to be trained. The speech sample to be trained includes speech information and a first instruction label and a second instruction label corresponding to the speech information. The first instruction label is used to indicate whether it is necessary to capture an image, and the second instruction label is used to indicate the type of image capture. In response to a target voice command indicating that an image needs to be captured, an image capture instruction corresponding to the target voice command is sent to the TWS earphone so that the TWS earphone can capture the scene image. Receive scene images sent by TWS earphones.

[0007] A second aspect of this application provides a shooting control method applied to TWS earphones, including: Obtain image capture instructions; In response to an image capture command, perform the corresponding image capture action to obtain a scene image; Send scene images to mobile devices.

[0008] A third aspect of this application provides a shooting control device applied to a mobile terminal, comprising: The information acquisition module is used to acquire the user's target voice information; The command determination module is used to input the target speech information into the command parsing model to obtain the target speech command corresponding to the target speech information. The command parsing model is trained based on at least one training speech sample. The training speech sample includes speech information and a first command label and a second command label corresponding to the speech information. The first command label is used to indicate whether it is necessary to capture an image, and the second command label is used to indicate the type of image capture. The instruction sending module is used to respond to the target voice command indicating that an image needs to be captured, and send an image capture instruction corresponding to the target voice command to the TWS earphone so that the TWS earphone can capture scene images. The image receiving module is used to receive scene images sent by TWS earphones.

[0009] A fourth aspect of this application provides a shooting control device applied to TWS earphones, comprising: The instruction acquisition module is used to acquire image capture instructions; The image capture module is used to respond to image capture commands, execute image capture actions corresponding to the image capture commands, and obtain scene images; The image transmission module is used to send scene images to the mobile terminal.

[0010] A fifth aspect of the embodiments of this application provides an electronic device, the device comprising: a memory and a program or instructions stored in the memory and executable on a processor, wherein when the program or instructions are executed by the processor, they implement the shooting control method provided in any of the embodiments of this application described above.

[0011] A sixth aspect of the embodiments of this application provides a readable storage medium on which a program or instructions are stored, and when the program or instructions are executed by a processor, they implement the shooting control method provided by any aspect of the embodiments of this application described above.

[0012] A seventh aspect of the embodiments of this application provides a computer program product, wherein when the instructions in the computer program product are executed by the processor of an electronic device, the electronic device performs the shooting control method provided in any aspect of the embodiments of this application described above.

[0013] The shooting control method provided in this application acquires the user's target voice information and uses an instruction parsing model trained on a large number of training voice samples to accurately parse the corresponding target voice command. This method eliminates the need for manual button operation by the user; commands can be issued solely through voice, greatly improving operational convenience. Furthermore, when the target voice command indicates the need to capture an image, an image capture command can be sent to the TWS earphone to complete the capture and receive the scene image. This wireless TWS earphone shooting mode allows users to shoot with greater freedom, unrestricted by the operation of the mobile terminal itself. Thus, compared to the traditional complex button shooting operation, this application, with voice interaction at its core, simplifies the operation process, enabling users to complete image capture more easily and quickly, effectively meeting users' needs for convenient shooting and thereby improving the user experience of the shooting function. Attached Figure Description

[0014] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments of this application will be briefly introduced below. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0015] Figure 1 This is a schematic flowchart of a first shooting control method provided in an embodiment of this application; Figure 2 This is a schematic diagram of a TWS earphone photography process provided in one embodiment of this application; Figure 3 This is a flowchart illustrating a second shooting control method provided in one embodiment of this application; Figure 4 This is a schematic diagram of the structure of a TWS earphone provided in one embodiment of this application; Figure 5 This is a schematic diagram illustrating the process of image transmission in a TWS earphone according to one embodiment of this application; Figure 6 This is a schematic diagram of the structure of a first shooting control device provided in one embodiment of this application; Figure 7 This is a schematic diagram of the structure of a second shooting control device provided in one embodiment of this application; Figure 8 This is a schematic diagram of the structure of a shooting control device provided in one embodiment of this application. Detailed Implementation

[0016] The features and exemplary embodiments of various aspects of this application will be described in detail below. To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain this application and not to limit it. For those skilled in the art, this application can be implemented without some of these specific details. The following description of the embodiments is merely to provide a better understanding of this application by illustrating examples.

[0017] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus that includes said element.

[0018] It should be noted that the acquisition, storage, use, and processing of data in the technical solution of this application all comply with the relevant provisions of national laws and regulations.

[0019] It should be noted that in the embodiments of this application, certain software, components, models and other existing solutions in the industry may be mentioned. These should be regarded as exemplary and are only intended to illustrate the feasibility of implementing the technical solution of this application. However, it does not mean that the applicant has used or necessarily used the solution.

[0020] In traditional mobile terminal image capture control methods, users need to trigger the shooting function through physical buttons or virtual on-screen buttons, resulting in a multi-level interaction in the operation process. Physical button triggering requires the user to precisely press a specific area, while virtual on-screen button triggering requires the user to focus their gaze on the touch interface; both require the user to maintain direct contact with the device. When the user is in motion or in a complex environment, the operational accuracy and response speed are limited by the device's grip posture and environmental interference, significantly increasing the trigger failure rate and affecting the execution efficiency of image acquisition tasks.

[0021] For example, in smart cycling navigation scenarios on mobile devices, users need to capture images of the scenery along the way while maintaining vehicle control. Traditional operation requires users to remove one hand from the handlebars to perform touch operations, leading to decreased device stability and a 300-500 millisecond delay in triggering shooting commands. When the vehicle is traveling at 20 km / h, the 0.5-second interval between the user's gaze shifting to the screen creates a 2.8-meter blind spot, directly causing a target positioning deviation of more than 15 degrees. At this point, it becomes impossible to effectively distinguish whether the user intends to adjust the navigation path or trigger the shooting function, resulting in an erroneous command execution rate of up to 18%.

[0022] To address the aforementioned challenges, this application first explores how to establish a contactless command triggering mechanism to eliminate reliance on physical operations. Traditional button touch control methods are limited by eye movement and physical contact. This application considers using voice as a natural interaction medium, capturing user intent through acoustic signals. However, simple voice recognition suffers from environmental noise interference and command ambiguity, necessitating the construction of a model with multi-dimensional command parsing capabilities. Further analysis reveals that image capture requirements involve two layers of decision logic: whether to execute the capture and which capture mode to select. This requires the command recognition system to simultaneously handle classification tasks. This application compares single-stage classification with multi-task learning frameworks, finding that the latter effectively captures the correlation between commands, ultimately choosing to construct a dual-classification head model architecture. To solve the problem of training data labeling costs, a self-distillation pre-training method is used to improve the model's semantic understanding ability, reducing reliance on labeled samples through unsupervised learning. Simultaneously, recognizing the synergistic advantages of mobile terminals and wearable devices, this application uses TWS earphones as image acquisition terminals, leveraging their spatial characteristics to achieve hands-free operation and avoid direct interaction between the user and the mobile terminal.

[0023] In this regard, such as Figure 1 As shown, this application provides a shooting control method applied to a mobile terminal, including the following steps S101 to S104: S101, acquire the user's target voice information.

[0024] In this step, the target speech information refers to the user's voice content with a specific intent. For example, if a user says, "I want to take a picture," this is the target speech information, which contains the user's need to take a picture. Specifically, this can be achieved by using a microphone array to collect ambient sound and extracting clear human voice through a noise reduction algorithm, which is then used to trigger the subsequent command parsing process.

[0025] S102, the target speech information is input into the instruction parsing model to obtain the target speech command corresponding to the target speech information; the instruction parsing model is obtained by training the training model based on at least one training speech sample. The training speech sample includes speech information and a first instruction label and a second instruction label corresponding to the speech information. The first instruction label is used to indicate whether it is necessary to capture an image, and the second instruction label is used to indicate the type of image capture.

[0026] In this step, the instruction parsing model is a trained algorithm model whose function is to analyze and process the input speech information, converting it into commands that the computer can understand and execute. Specifically, it can be implemented using a natural language understanding model based on deep neural networks, which analyzes the semantics of the speech to determine whether filming is needed and the type of filming.

[0027] The target speech command is a specific instruction corresponding to the user's speech intent, obtained by the instruction parsing model after parsing the input target speech information. For example, if a user says "I want to take a picture," the target speech command obtained after processing by the instruction parsing model might be "Take a single picture."

[0028] The model to be trained is an initial algorithm model that has not been fully trained and whose performance needs to be improved. It has the potential to perform speech parsing, but its performance needs to be optimized through learning from a large amount of data.

[0029] The speech information is the raw speech data in the speech samples to be trained, such as various statements spoken by the user related to image capture; the "first instruction label" indicates whether image capture is required. For example, for the speech information "take a photo," the first instruction label could be "take a photo"; for the speech information "please check the current time," the first instruction label would be "do not take a photo"; the "second instruction label" indicates the type of image capture. For example, when the speech information is "take a photo," the second instruction label is "take a photo"; if the speech information is "record a video," the second instruction label would be "record a video."

[0030] S103, in response to the target voice command indicating that an image needs to be captured, sends an image capture instruction corresponding to the target voice command to the TWS earphone, so that the TWS earphone captures the scene image.

[0031] In this step, the image capture command is a specific instruction generated based on the target voice command, used to control the TWS earbuds to capture images. For example, if the target voice command is "take a single photo," then the image capture command will include relevant information such as the image capture type being a single photo. After receiving this image capture command, the TWS earbuds will execute the operation of taking a single photo. Specifically, the command data can be encapsulated using the Bluetooth protocol and transmitted via a wireless communication link to ensure the real-time performance and accuracy of the capture command.

[0032] TWS earphones refer to wireless ear-worn devices with integrated cameras. Specifically, they can be implemented using dual independent camera modules, which are used to execute shooting commands and return image data.

[0033] S104 receives scene images sent by TWS earphones.

[0034] In this step, the scene image is the actual scene captured by the TWS earbuds according to the image capture command, such as a photo or video. Specifically, it can be achieved by using an image sensor to capture the image and transmitting it back to the mobile terminal via Bluetooth, thus completing the synchronization and storage of the image data.

[0035] This application enables users to complete the shooting process without manually operating the mobile terminal by coordinating the instruction parsing model with the shooting function of TWS earphones. At the same time, it utilizes the wearable characteristics of TWS earphones to expand the shooting angle and scene coverage, solving the problems of cumbersome operation and limited field of view of traditional mobile terminal shooting.

[0036] The working process and principle of this application are as follows: the mobile terminal first acquires the user's target voice information. This step is completed through the mobile terminal's microphone or other audio input device, converting the user's voice into a digital signal.

[0037] Next, the mobile terminal inputs the target voice information into the command parsing model. The command parsing model is a pre-trained deep learning model capable of understanding and parsing voice commands. This model is trained on a large number of training voice samples, which contain voice information and two corresponding command labels. The first command label indicates whether video recording is required, and the second command label indicates the type of video recording. Through this multi-task learning approach, the model can simultaneously understand whether the user wants to record and what shooting mode they want to use.

[0038] Then, the instruction parsing model processes the target speech information and outputs the target speech command. If the target speech command indicates that an image needs to be captured, the mobile terminal will generate a corresponding image capture instruction. This instruction includes specific shooting parameters, such as image capture type and exposure settings.

[0039] Then, the mobile terminal sends the image capture command to the paired TWS earbuds via the Bluetooth communication protocol. After receiving the command, the TWS earbuds activate their built-in camera and capture the current scene image according to the shooting parameter settings in the image capture command.

[0040] Finally, the TWS earbuds transmit the captured scene images back to the mobile terminal via Bluetooth. The mobile terminal receives and stores this image data, completing the entire shooting process.

[0041] The core of this solution lies in the collaborative work of voice interaction and wearable devices, which enables hands-free, contactless shooting control, greatly improving the convenience of shooting in various scenarios.

[0042] like Figure 2 The diagram illustrates a process for taking photos using TWS earbuds. The user emits target voice information, which is captured by the microphone and transmitted to the TWS earbuds. The TWS earbuds then transmit the target voice information to a mobile terminal via Bluetooth Low Energy (BLE). The mobile terminal inputs the target voice information into a command parsing model, parses the target voice command, and if it indicates the need to capture an image, sends an image capture command to the TWS master earbud via Bluetooth (BT), while simultaneously synchronizing the image capture command to the slave earbud via the TWS protocol. Upon receiving the image capture command, both the master and slave earbuds take a picture and upload it to the mobile terminal.

[0043] As an example, the mobile terminal could be a smartphone or tablet, equipped with a high-performance processor, memory, Bluetooth module, and microphone array. The user wears TWS earbuds paired with the mobile terminal, which integrate a miniature camera module. When the user wants to take a picture, they speak a command such as "I want to take a picture" or "I want to record a video." The mobile terminal's microphone array captures this speech and performs initial noise reduction.

[0044] The processed target speech information is input into a pre-trained instruction parsing model. This model employs a Transformer architecture, comprising a backbone network and two classifiers. The backbone network extracts semantic features from the speech, the first classifier determines whether a photo needs to be taken, and the second classifier determines the specific photo type. The instruction parsing model then outputs a message indicating that a photo needs to be taken, and the photo type is "take a picture." Based on this, the mobile terminal generates an image capture instruction containing parameters such as the photo mode and autofocus.

[0045] Using Bluetooth 5.0, the mobile device sends an image capture command to the TWS earbuds. Upon receiving the command, the TWS earbuds activate their built-in 12-megapixel camera and set the appropriate shooting parameters. After capturing the image, the TWS earbuds transmit the JPEG image data back to the mobile device via Bluetooth. The mobile device receives and saves the photo, completing the process.

[0046] This embodiment acquires the user's target voice information and accurately parses the corresponding target voice command using a command parsing model trained on a large number of training voice samples. This method eliminates the need for manual button operation; commands can be issued solely through voice, greatly improving ease of use. Furthermore, when the target voice command indicates the need to capture an image, it can send an image capture command to the TWS earphone to complete the capture and receive the scene image. This wireless TWS earphone shooting mode allows users greater freedom in shooting, unrestricted by the mobile terminal's own operation. Thus, compared to traditional complex button-based shooting operations, this application, with voice interaction at its core, simplifies the operation process, enabling users to complete image capture more easily and quickly, effectively meeting users' needs for convenient shooting and improving the user experience of the shooting function.

[0047] In some of the solutions described above in this application, the mobile terminal controls the TWS earphone to capture scene images through image capture commands. However, due to differences in shooting angle or position, the field of view of the first and second scene images captured by the two sides of the TWS earphone is limited.

[0048] In this regard, this application further proposes that the TWS earphone includes a first earphone and a second earphone, and the scene image includes a first scene image transmitted by the first earphone and a second scene image transmitted by the second earphone. Following S104, this shooting control method also includes: Receive target inertial sensor data sent by TWS earphones; the target inertial sensor data is the inertial sensor data of TWS earphones when capturing the first side scene image and the second side scene image; Based on the target inertial sensor data, the relative pose between the first side scene image and the second side scene image is determined; Based on the relative pose, the first side scene image and the second side scene image are stitched together to obtain the target scene image.

[0049] In this embodiment, the first and second earpieces are each integrated with an inertial sensor to collect motion data during shooting; the target inertial sensor data includes acceleration, angular velocity, and orientation information; the relative pose is calculated using a six-degree-of-freedom pose estimation algorithm, which includes translation and rotation parameters; the image stitching uses a feature point matching algorithm, combined with the relative pose, to spatially align the two images.

[0050] Specifically, when the TWS earbuds perform a shooting action, the first and second earbuds simultaneously acquire scene images from their respective perspectives and record corresponding inertial sensor data. After receiving the scene images and inertial sensor data from both sides, the mobile terminal analyzes the acceleration and angular velocity information to calculate the spatial position difference and angular offset of the two earbuds at the moment of shooting, generating a relative pose matrix. Subsequently, this relative pose matrix is ​​used to perform geometric transformations on the first and second scene images to eliminate the perspective differences. Then, feature point matching is used to further optimize the alignment accuracy, and finally, the two scene images are fused into a seamlessly stitched target scene image. This process solves the problem of image stitching misalignment caused by inconsistent shooting angles, improving the accuracy and visual effect of multi-view image synthesis.

[0051] As an example, the first and second earbuds in a TWS earphone are equipped with a camera and an inertial sensor, respectively. During recording, both earbuds simultaneously capture scene images and record corresponding inertial sensor data. This data is then sent to a mobile terminal for processing.

[0052] After receiving scene images from both sides and inertial sensor data, the mobile terminal first analyzes the inertial sensor data. By comparing data such as acceleration and angular velocity from both sides, it calculates the relative position and orientation between the two cameras, i.e., the relative pose.

[0053] Then, the mobile terminal uses the calculated relative pose information to perform geometric transformations and registration on the scene images on both sides. This process includes operations such as image rotation, scaling, and translation, so that the two scene images can be accurately aligned in the overlapping area.

[0054] Finally, the mobile terminal fuses and stitches the registered scene images from both sides. Image fusion algorithms, such as weighted averaging or multi-band fusion, are used in overlapping areas to ensure a natural transition and no obvious seams in the stitched image. After stitching, a complete panoramic image of the target scene is obtained.

[0055] This embodiment achieves automatic stitching of scene images captured using dual headphones, increasing the field of view and enhancing the integrity of the scene images. Simultaneously, using inertial sensor data for pose estimation improves the accuracy and efficiency of image stitching, avoiding misalignment and distortion problems that may occur in traditional image stitching. Furthermore, this solution eliminates the need for users to manually adjust the shooting angle or perform post-processing, improving the user experience.

[0056] In some of the solutions described above in this application, when a mobile terminal controls TWS earphones to capture images via voice commands, the accuracy of the command parsing model directly affects the triggering precision of the shooting action. In the prior art, if an unoptimized general speech recognition model is directly used, insufficient training data or an unreasonable model structure may lead to a risk of misrecognition in the dual judgment of shooting requirements and shooting type in the voice commands.

[0057] In this regard, prior to S102, this application further proposes that the shooting control method also includes: At least one speech sample to be trained and a model to be trained are obtained; wherein, the model to be trained is constructed from a target self-distillation model, a first classification head and a second classification head. The target self-distillation model is a model with semantic feature extraction capability obtained by self-distillation pre-training using unlabeled speech samples. The first classification head is used to output the probability that the speech sample to be trained indicates whether or not an image needs to be captured. The second classification head is used to output the probability that the speech sample to be trained indicates various types of image capture. Based on the speech samples to be trained, the model parameters of the target self-distillation model, the first classification head parameters of the first classification head, and the second classification head parameters of the second classification head are updated to obtain the instruction parsing model.

[0058] In this embodiment, the target self-distillation model extracts general semantic features from unlabeled speech samples through self-distillation pre-training, enabling the model to possess basic semantic understanding capabilities. The first and second classification heads each employ independent fully connected layer structures. The first classification head outputs a binary classification result to determine whether shooting is required, while the second classification head outputs a multi-classification result to determine the shooting type. During model parameter updates, the first and second instruction labels in the training speech samples serve as supervision signals, and the parameters of the target self-distillation model, the first classification head, and the second classification head are simultaneously optimized using a backpropagation algorithm.

[0059] Specifically, during the model training phase, the target self-distillation model first encodes the features of the input speech samples to be trained, generating a high-dimensional semantic vector. The first classification head receives this vector, outputs the probability value of whether a shot is needed through an activation function, and calculates the cross-entropy loss with the first instruction label. The second classification head receives the same vector, outputs the probability distribution of each shooting type through a normalization function, and calculates the cross-entropy loss with the second instruction label. The weighted sum of the losses from the two classification heads drives the model parameter update. In this way, the target self-distillation model simultaneously learns the dual judgment task of speech commands during training. Self-distillation pre-training endows the model with generalization ability, while the independent parameter design of the two classification heads avoids interference between tasks. For example, when the speech sample to be trained contains "take a panoramic photo," the model determines that a shot is needed through the first classification head, and simultaneously identifies the shooting type as "panoramic photo" through the second classification head. The resulting instruction parsing model can accurately distinguish the complex intent in the speech command.

[0060] As an example, the speech samples to be trained can include multiple speech messages and their corresponding first instruction labels and second instruction labels. For instance, the first instruction label for the speech message "Please take a photo" is "Need to take a photo," and the second instruction label is "Take a photo." The first instruction label for the speech message "Please record a video" is "Need to take a photo," and the second instruction label is "Record a video." The first instruction label for the speech message "Exit camera" is "Do not need to take a photo," and the second instruction label is empty.

[0061] Therefore, the speech samples to be trained are input into the training model. The target self-distillation model extracts the semantic features of the speech samples. The first classification head outputs the probability of whether a shot is needed, and the second classification head outputs the probabilities of various shooting types. The output results are compared with the labels of the speech samples, the loss function is calculated, and the model parameters are updated through the backpropagation algorithm to obtain the trained instruction parsing model.

[0062] This embodiment employs a target self-distillation model as the base model, combined with fine-tuning using two classification heads, to effectively extract semantic features from voice commands and accurately identify the user's shooting intention and specific shooting type. This method fully leverages the advantages of pre-training with unlabeled speech data, while supervised fine-tuning improves the model's performance on specific tasks. Compared to traditional speech recognition methods, this approach better understands the semantics of user voice commands, improving the accuracy of voice-controlled shooting and enhancing the user experience.

[0063] In some of the solutions mentioned above in this application, after the user triggers the TWS earphone to capture scene images via voice, the mobile terminal receives the images but cannot automatically identify the elements in the images and provide relevant information, causing the user to have to manually query or not be able to understand the scene content in a timely manner, which affects the user experience.

[0064] In response, this application further proposes that, following S104, the shooting control method also includes: Input the scene image into the image recognition model to extract the iconic elements in the scene image; Search the background knowledge base for information about elements associated with the iconic elements; Send element description information to TWS earphones to introduce iconic elements to users.

[0065] In this embodiment, the image recognition model employs a convolutional neural network architecture, extracting image features through multiple convolutional and pooling layers, with fully connected layers outputting category labels for distinctive elements. A background knowledge base stores element names, attributes, and associated text data, using an inverted index structure to accelerate searching. Element description information is transmitted to the TWS earphones in text or voice format, and the earphones' built-in speech synthesis module converts the text into speech for playback.

[0066] Specifically, after the mobile terminal receives the scene image, the image recognition model performs region segmentation on the image, identifying elements such as buildings, landmarks, or text labels. For example, when the image contains the Eiffel Tower, the model outputs the "Eiffel Tower" tag. The background knowledge base matches the tag with pre-stored building history, height data, and opening hours. The searched information is sent to headphones via Bluetooth, and the headphones play the audio content through bone conduction speakers. Thus, users can obtain real-time explanations without manual operation during the shooting process, enhancing the intelligence of image interaction.

[0067] As an example, after receiving scene images sent by TWS earbuds, the scene images are input into an image recognition model. The image recognition model can be a pre-trained deep learning model, such as a convolutional neural network model. This image recognition model processes the input scene images and extracts representative elements from them. Representative elements can include representative objects or scenes such as buildings, sculptures, and natural landscapes.

[0068] Search a background knowledge base for descriptive information about elements associated with the iconic elements. The background knowledge base can be a pre-built database containing information about various iconic elements. The search process can employ methods such as keyword matching to find descriptive text related to the identified iconic elements.

[0069] The system sends information about the searched elements to the TWS earbuds. Upon receiving the information, the TWS earbuds use speech synthesis technology to convert the text into speech, which is then played back to the user through the speaker, introducing the iconic elements of the scene.

[0070] This embodiment achieves intelligent recognition and narration of the shooting scene. Users can obtain introductions to important elements in the scene without manually searching for relevant information, enhancing the shooting experience. Simultaneously, this solution utilizes TWS earphones as the output device, avoiding the inconvenience of users needing to look at their phone screen and improving the convenience of information acquisition. Furthermore, with the support of a background knowledge base, the system can provide rich and accurate descriptive information to help users better understand and appreciate the shooting scene.

[0071] In some of the solutions described above in this application, after the mobile terminal receives the scene images sent by the TWS earphones, the image processing process does not take into account the user's personalized needs, resulting in a deviation between the image processing results and the user's preferences, which affects the user experience.

[0072] In response, this application further proposes that, following S104, the shooting control method also includes: Obtain the user's historical image style preferences; Based on historical image preferences and styles, determine personalized optimization strategies corresponding to historical image preferences and styles; Based on a personalized optimization strategy, scene images are optimized to obtain optimized scene images.

[0073] In this embodiment, the historical image preference style is extracted by analyzing the user's adjustment operation records or saved records of historical images. The adjustment operation records include contrast adjustment parameters, saturation adjustment parameters, and filter type selection records, etc. The personalized optimization strategy includes the contrast optimization parameter range, saturation optimization parameter range, and filter type combination, etc. When optimizing the scene image based on the personalized optimization strategy, a preset algorithm is used to adjust the contrast parameter of the scene image to the contrast optimization parameter range, adjust the saturation parameter of the scene image to the saturation optimization parameter range, and overlay the scene image according to the filter type combination.

[0074] Specifically, after receiving the scene image, the mobile terminal retrieves the user's adjustment operation records for historical images from the local storage unit. It determines the range of contrast optimization parameters by statistically analyzing the average value of contrast adjustment parameters, the range of saturation optimization parameters by statistically analyzing the average value of saturation adjustment parameters, and the filter type selection frequency to determine the filter type combination. The scene image is then input into the image processing engine. The image processing engine performs a linear transformation on the contrast of the scene image based on the contrast optimization parameter range, and a non-linear enhancement on the saturation of the scene image based on the saturation optimization parameter range. Preset filter algorithms are then applied sequentially according to the filter type combination. The processed optimized scene image automatically overwrites the original scene image or generates a copy and stores it in a designated directory. Users can view the optimized image content through the mobile terminal.

[0075] As an example, after receiving scene images sent by TWS earbuds, the system obtains the user's historical image preference style. This historical image preference style can be obtained by analyzing image data that the user has previously saved or shared. For example, the frequency of style features such as high saturation, high contrast, and soft tones in the user's saved images can be statistically analyzed, and the most frequently occurring feature can be identified as the user's historical image preference style.

[0076] Based on historical image preferences, a personalized optimization strategy is then determined. Specifically, if the user's historical image preference is high saturation, the personalized optimization strategy may include increasing the color saturation of the scene image; if the user prefers high contrast, the optimization strategy may include enhancing the brightness and darkness contrast of the scene image.

[0077] Finally, based on the determined personalized optimization strategy, the scene image is optimized to obtain an optimized scene image. For example, image processing algorithms can be used to adjust parameters such as saturation, contrast, and hue of the scene image to match the user's preferred style. The optimized scene image is closer to the user's aesthetic preferences.

[0078] This embodiment enables automatic personalized optimization of scene images based on the user's historical image preferences, improving the visual appeal and user satisfaction. Furthermore, since the optimization strategy is automatically generated based on the user's historical preferences, manual adjustment of image parameters is unnecessary, simplifying user operation and enhancing the user experience. In addition, personalized optimization makes the captured scene images more in line with the user's aesthetic needs, enhancing the practicality of the shooting function.

[0079] Based on the shooting control method for mobile terminals provided in this application, this application further proposes a shooting control method for TWS earphones.

[0080] like Figure 3 As shown, this application provides a shooting control method applied to TWS earphones, including the following steps S301 to S303: S301, receive image capture command; S302, in response to the image capture command, executes the image capture behavior corresponding to the image capture command to obtain the scene image; S303 sends scene images to the mobile terminal.

[0081] In this embodiment, as Figure 4The diagram illustrates the structure of a TWS (True Wireless Stereo) earphone. The Bluetooth (BT) module, acting as the main control unit, is responsible for overall control and human-computer interaction. It connects to the Wi-Fi module and the Image Signal Processor (ISP) module via a Universal Asynchronous Receiver / Transmitter (UART), and also connects to buttons, a microphone, and a speaker to receive user input and output prompts. The Wi-Fi module and ISP module integrate Wi-Fi and image signal processing functions. They receive commands from the Bluetooth module via UART, control the camera to capture images, and store the data in a NOR flash memory module using Wi-Fi. Furthermore, the Wi-Fi and ISP modules interact with the Bluetooth module via an I2S interface to exchange audio data and can work in conjunction with an Inertial Measurement Unit (IMU) to support various earphone functions.

[0082] The image capture command can be acquired through various methods, including user interaction with the target button on the TWS earbuds, matching the user's target voice information with a preset voice command template, or direct transmission from the mobile terminal. The image capture process involves adjusting the TWS earbuds' camera parameters to meet preset shooting conditions, such as focal length range or exposure thresholds, and then triggering the camera to capture the scene image. The scene image is transmitted to the mobile terminal via a wireless communication protocol, and compression encoding techniques can be used during transmission to reduce the data volume.

[0083] Specifically, once the image capture command is received, the camera parameters of the TWS earbuds are first adjusted to preset shooting conditions. For example, the camera's aperture value is set to a specific range, and the shutter speed is dynamically adjusted according to the ambient light. After the parameters are adjusted, the camera initiates the shooting process to capture the scene image. After shooting, the scene image is encoded and compressed, and then transmitted to the mobile terminal via the Wi-Fi module. In this process, the TWS earbuds independently complete the shooting operation, without requiring the user to directly operate the mobile terminal, thus simplifying the shooting process. Through the constraints of preset shooting conditions, image quality is guaranteed, avoiding image blurring or overexposure issues caused by improper parameter settings.

[0084] like Figure 5The diagram illustrates a process for image transmission using TWS earbuds. After the TWS earbuds capture a full-size image or video, the image transmission process begins. At this point, the mobile terminal acts as the access point (AP), while the main and secondary earbuds of the TWS earbuds act as stations (STAs). The main and secondary earbuds transmit the captured full-size image or recorded video to the mobile terminal (AP) via Wi-Fi. In this process, Wi-Fi serves as the data transmission channel, ensuring stable and rapid transmission of image data from the TWS earbuds to the mobile terminal, allowing users to perform subsequent viewing and processing on the mobile terminal.

[0085] As an example, TWS earbuds first acquire a video recording command. This command can be acquired in several ways. For instance, the earbuds can respond to the user's interaction with a button on the earbuds and acquire the corresponding video recording command. Alternatively, the earbuds can respond to the user's voice input, matching the voice with a preset voice command template to determine the appropriate video recording command. Furthermore, TWS earbuds can also receive video recording commands sent from a mobile terminal.

[0086] Upon receiving an image capture command, the TWS earbuds respond by executing the corresponding image capture actions to obtain a scene image. Specifically, the TWS earbuds first adjust their camera parameters until they meet preset shooting conditions. These parameters may include exposure time, ISO sensitivity, white balance, etc. When the camera parameters meet the preset shooting conditions, the TWS earbuds execute the shooting actions corresponding to the image capture command to capture the scene image.

[0087] Finally, the TWS earbuds send the captured scene images to the mobile terminal. This can be achieved via wireless communication. After receiving the scene images, the mobile terminal can further process, store, or display them.

[0088] This embodiment demonstrates the ability to capture images using TWS earphones. Users can then directly shoot using the TWS earphones they are wearing, without needing to hold the mobile device. This significantly improves the convenience and flexibility of shooting, allowing users to quickly capture moments in various scenarios. Furthermore, the shooting angle of TWS earphones is closer to the human eye's perspective, resulting in more natural and visually pleasing images. This innovative shooting method provides users with a completely new image recording experience and expands the application scope of mobile terminal image shooting functions.

[0089] In some of the solutions mentioned above in this application, the way TWS earphones obtain image shooting commands is relatively simple, which cannot meet the diverse operational needs of users in different scenarios, resulting in users being unable to flexibly trigger the shooting function and affecting the user experience.

[0090] In this regard, this application further proposes that S201 includes any one of the following: In response to the interaction between the user and the target button on the TWS earphone, obtain the image capture command corresponding to the interaction action; In response to the user's target voice information, the target voice information is matched with a preset voice command template to determine the image capture instruction corresponding to the target voice information; Receive image capture commands sent by the mobile terminal.

[0091] In this embodiment, the target button is configured as a physical button or a virtual touch button, and the interaction actions include single-click, double-click, or long-press operations. Preset voice command templates are stored in the local storage unit of the TWS earphone, containing multiple predefined keyword sets, each keyword set corresponding to a type of image capture command. The process of the mobile terminal sending the image capture command is implemented through wireless communication protocols, including Bluetooth, Wi-Fi, or near-field communication technologies.

[0092] Specifically, when a user presses the physical button on the side of the TWS earbuds, a trigger electrical signal is transmitted to the processing unit. The processing unit identifies the press duration as a single click or a long press, and then generates a corresponding image capture command. When the user issues a voice command containing preset keywords, the voice acquisition module converts the voice signal into text data and performs a similarity match with locally stored voice templates. If a match is successful, the corresponding shooting mode is activated. When the mobile terminal detects a user's voice-triggered shooting request, it sends an image capture command containing shooting parameters to the earbuds via Bluetooth. By coordinating multiple command acquisition methods, the lack of flexibility caused by a single operation mode is solved, allowing users to conveniently trigger the shooting function in scenarios such as exercise, holding objects with both hands, or remote control.

[0093] As an example, when obtaining image capture commands based on button interaction, TWS earbuds are equipped with a target button, which can be a physical button or a virtual touch button. When the user interacts with the target button, such as by clicking, double-clicking, or long-pressing, the relevant sensing components of the TWS earbuds detect the interaction and transmit the signal to the processing unit. The processing unit identifies the type of interaction according to preset rules and then obtains the image capture command corresponding to that interaction. For example, a single click may correspond to a photo capture command, and a long press may correspond to a video recording command. This button interaction-based approach provides users with an intuitive and convenient operation method, especially suitable for users to quickly trigger the shooting function in specific scenarios.

[0094] As another example, when obtaining image shooting commands based on voice matching, the TWS earbuds' local storage unit stores a preset image shooting command template. This template contains multiple predefined keyword sets, each corresponding to a type of image shooting command. When the user emits the target voice information, the earbuds' voice acquisition module picks up the voice signal and converts it into text data. Then, the processing unit matches the text data corresponding to the target voice information with the preset voice command template, determining the image shooting command corresponding to the target voice information by calculating similarity and other methods. For example, when the user says a voice command containing the keyword "take a picture," a photo shooting command will be matched. This method allows users to flexibly trigger the shooting function simply by voice without manual operation, improving ease of use.

[0095] This embodiment provides multiple methods for acquiring image capture commands, allowing users to flexibly choose according to their actual usage scenarios. Triggering capture through various methods such as physical button interaction, voice control, or remote control via mobile terminal greatly improves the convenience and flexibility of the shooting operation, meeting users' shooting needs in different scenarios. At the same time, the combination of multiple control methods enhances the system's fault tolerance and reliability; even if one control method fails, users can still achieve shooting control through other methods, thereby improving the overall user experience.

[0096] In some of the solutions described above in this application, when the TWS earphones respond to the image capture command, the camera parameters may not be adjusted in time due to changes in ambient light or user movement, affecting the quality of the captured image.

[0097] In this regard, this application further proposes S302, which includes: In response to the image capture command, adjust the camera parameters of the TWS earphones until the camera parameters meet the preset shooting conditions; In response to the camera parameters meeting the preset shooting conditions, the system executes the image shooting behavior corresponding to the image shooting command to obtain the scene image.

[0098] In this embodiment, adjusting camera parameters includes at least one of autofocus, exposure compensation, or white balance adjustment; preset shooting conditions include at least one of image sharpness threshold, brightness threshold, or contrast threshold. During camera parameter adjustment, the parameter configuration is dynamically optimized by real-time detection of the matching degree between the current parameters and the preset conditions. For example, in low-light environments, ISO sensitivity and aperture size are adjusted first to bring the brightness to the preset threshold.

[0099] Specifically, after receiving an image capture command, the TWS earbuds activate the camera parameter adjustment module. The focus module automatically adjusts the focus based on changes in the distance to objects in the scene, while the exposure module adjusts the shutter speed and ISO based on the ambient light intensity. During parameter adjustment, the image sensor captures real-time images and extracts key metrics, comparing them with preset shooting conditions. When the image sharpness exceeds a set threshold and the brightness is within a reasonable range, the shooting action is triggered. This process ensures that the camera parameters are adapted to the current environment during shooting, avoiding blurry or overexposed images due to unsuitable parameters. Through phased adjustments and condition verification, shooting efficiency is improved while ensuring image quality.

[0100] As an example, when the TWS earbuds receive an image capture command, they first use the built-in gyroscope and accelerometer to detect the current device's posture stability. If the detected device shaking exceeds a preset threshold, the electronic image stabilization function is automatically activated, and the image sensor's operating mode is adjusted to motion compensation mode. When the ambient light intensity is below 50 lux, the image sensor's sensitivity is gradually increased to ISO 3200, while the shutter speed is reduced to 1 / 30 second. When the camera parameters meet the preset stable shooting conditions and the exposure parameters reach the optimal combination, the image sensor is triggered to perform a single-frame long exposure capture to obtain a high-quality scene image with low noise.

[0101] This embodiment effectively solves the technical challenge of inaccurate parameter adjustment for external mobile devices in complex shooting environments, achieving automatic optimization of shooting parameters in low-light conditions or during motion. Specifically, through the collaborative work of multiple sensors, it ensures clear and stable image data can be acquired even in scenarios with device shake or insufficient light, avoiding the shooting delays and improper parameter settings caused by traditional manual adjustment methods. This significantly improves the imaging quality and operational reliability of TWS earphones as independent shooting devices.

[0102] Based on the shooting control method provided in this application, correspondingly, this application also provides specific embodiments of the shooting control device.

[0103] like Figure 6 As shown in the diagram, this application provides a schematic diagram of a shooting control device. The shooting control device includes an information acquisition module 610, a command determination module 620, an instruction sending module 630, and an image receiving module 640.

[0104] Information acquisition module 610 is used to acquire the user's target voice information; The command determination module 620 is used to input the target speech information into the command parsing model to obtain the target speech command corresponding to the target speech information. The command parsing model is trained based on at least one training speech sample. The training speech sample includes speech information and a first command label and a second command label corresponding to the speech information. The first command label is used to indicate whether it is necessary to capture an image, and the second command label is used to indicate the type of image capture. The instruction sending module 630 is used to respond to a target voice command indicating that an image needs to be captured, and send an image capture instruction corresponding to the target voice command to the TWS earphone so that the TWS earphone captures scene images. The image receiving module 640 is used to receive scene images sent by TWS earphones.

[0105] The shooting control device provided in this application acquires the user's target voice information and accurately parses the corresponding target voice command using an instruction parsing model trained on a large number of training voice samples. This method eliminates the need for manual button operation by the user; commands can be issued solely through voice, greatly improving operational convenience. Simultaneously, when the target voice command indicates the need to capture an image, it can send an image capture command to the TWS earphone to complete the capture and receive the scene image. This wireless TWS earphone shooting mode allows users to shoot with greater freedom, unrestricted by the operation of the mobile terminal itself. Thus, compared to the traditional complex button shooting operation, this application, with voice interaction at its core, simplifies the operation process, enabling users to complete image capture more easily and quickly, effectively meeting users' needs for convenient shooting, and thereby improving the user experience of the shooting function.

[0106] Furthermore, this application also proposes that the TWS earphone includes a first earphone and a second earphone, and the scene image includes a first scene image transmitted by the first earphone and a second scene image transmitted by the second earphone. After receiving the scene image sent by the TWS earphones, the shooting control device may further include: The data receiving module is used to receive target inertial sensor data sent by the TWS earphones; the target inertial sensor data is the inertial sensor data of the TWS earphones when capturing the first side scene image and the second side scene image; The pose determination module is used to determine the relative pose between the first side scene image and the second side scene image based on the target inertial sensor data. The image stitching module is used to stitch together the first side scene image and the second side scene image based on relative pose to obtain the target scene image.

[0107] Furthermore, this application also proposes that before inputting the target speech information into the instruction parsing model to obtain the target speech command corresponding to the target speech information, the shooting control device may further include: The model acquisition module is used to acquire at least one speech sample to be trained and a model to be trained. The model to be trained is constructed from a target self-distillation model, a first classification head, and a second classification head. The target self-distillation model is a model with semantic feature extraction capability obtained by self-distillation pre-training using unlabeled speech samples. The first classification head is used to output the probability that the speech sample to be trained indicates whether or not an image needs to be captured. The second classification head is used to output the probability that the speech sample to be trained indicates various types of image capture. The model training module is used to update the model parameters of the target self-distillation model, the first classification head parameters of the first classification head, and the second classification head parameters of the second classification head based on the speech samples to be trained, so as to obtain the instruction parsing model.

[0108] Furthermore, this application also proposes that after receiving the scene image sent by the TWS earphone, the shooting control device may further include: The element extraction module is used to input scene images into the image recognition model and extract the iconic elements in the scene images; The information retrieval module is used to search for element description information associated with the iconic element from the background knowledge base; The information sending module is used to send element introduction information to the TWS earphones so as to introduce the iconic elements to the user.

[0109] Furthermore, this application also proposes that after receiving the scene image sent by the TWS earphone, the shooting control device may further include: The style acquisition module is used to acquire the user's historical image style preferences; The strategy determination module is used to determine personalized optimization strategies corresponding to historical image preference styles based on historical image preference styles. The image optimization module is used to optimize scene images based on personalized optimization strategies to obtain optimized scene images.

[0110] Furthermore, such as Figure 7 As shown, this application provides a schematic diagram of another shooting control device. This shooting control device includes: The instruction acquisition module 710 is used to acquire image capture instructions; The image capture module 720 is used to respond to the image capture command, execute the image capture behavior corresponding to the image capture command, and obtain the scene image; The image transmission module 730 is used to send scene images to a mobile terminal.

[0111] Furthermore, this application also proposes an instruction acquisition module 710, used for: In response to the interaction between the user and the target button on the TWS earphone, obtain the image capture command corresponding to the interaction action; In response to the user's target voice information, the target voice information is matched with a preset voice command template to determine the image capture instruction corresponding to the target voice information; Receive image capture commands sent by the mobile terminal.

[0112] Furthermore, this application also proposes an image capturing module 720 for: In response to the image capture command, adjust the camera parameters of the TWS earphones until the camera parameters meet the preset shooting conditions; In response to the camera parameters meeting the preset shooting conditions, the system executes the image shooting behavior corresponding to the image shooting command to obtain the scene image.

[0113] Based on the shooting control method provided in this application, correspondingly, this application also provides specific embodiments of the shooting control device.

[0114] Figure 8 A schematic diagram of the hardware structure of the shooting control device provided in an embodiment of this application is shown.

[0115] The device for verifying feature matching results may include a processor 801 and a memory 802 storing computer program instructions.

[0116] Specifically, the processor 801 may include a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits that can be configured to implement the embodiments of this application.

[0117] Memory 802 may include mass storage for data or instructions. For example, and not limitingly, memory 802 may include a hard disk drive (HDD), floppy disk drive, flash memory, optical disk, magneto-optical disk, magnetic tape, or Universal Serial Bus (USB) drive, or a combination of two or more of these. Where appropriate, memory 802 may include removable or non-removable (or fixed) media. Where appropriate, memory 802 may be internal to an integrated gateway disaster recovery device. In a particular embodiment, memory 802 is non-volatile solid-state memory.

[0118] The processor 801 reads and executes computer program instructions stored in the memory 802 to implement any of the shooting control methods in the above embodiments.

[0119] In one example, the shooting control device may also include a communication interface 803 and a bus 810. For example, Figure 8 As shown, the processor 801, memory 802, and communication interface 803 are connected through bus 810 and complete communication with each other.

[0120] The communication interface 803 is mainly used to realize communication between various modules, devices, units and / or equipment in the embodiments of this application.

[0121] Bus 810 includes hardware, software, or both, that couples components of the shooting control device together. For example, and not limitingly, the bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an Infinite Bandwidth Interconnect, a Low Pin Count (LPC) bus, a memory bus, a Microchannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or other suitable buses, or combinations of two or more of these. Where appropriate, bus 810 may include one or more buses. Although specific buses are described and illustrated in embodiments of this application, this application contemplates any suitable bus or interconnect.

[0122] Furthermore, in conjunction with the shooting control methods in the above embodiments, this application embodiment can provide a computer storage medium for implementation. The computer storage medium stores computer program instructions; when these computer program instructions are executed by a processor, they implement any of the shooting control methods in the above embodiments.

[0123] In addition, in conjunction with the shooting control method in the above embodiments, this application embodiment can provide a computer program product for implementation. When the instructions in the computer program product are executed by the processor of the electronic device, the electronic device performs the shooting control method provided by any aspect of the above embodiments of this application.

[0124] It should be clarified that this application is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of this application is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order of steps, after understanding the spirit of this application.

[0125] The functional blocks shown in the above-described structural diagram can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, they can be, for example, electronic circuits, application-specific integrated circuits (ASICs), appropriate firmware, plug-ins, function cards, etc. When implemented in software, the elements of this application are programs or code segments used to perform the required tasks. Programs or code segments can be stored on a machine-readable medium or transmitted over a transmission medium or communication link via data signals carried on a carrier wave. "Machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical disks, hard disks, fiber optic media, radio frequency (RF) links, etc. Code segments can be downloaded via computer networks such as the Internet, intranets, etc.

[0126] It should also be noted that the exemplary embodiments mentioned in this application describe methods or systems based on a series of steps or apparatus. However, this application is not limited to the order of the above steps; that is, the steps can be performed in the order mentioned in the embodiments, or in a different order, or several steps can be performed simultaneously.

[0127] The aspects of this disclosure have been described above with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It should be understood that each block in the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that these instructions, executable via the processor of the computer or other programmable data processing apparatus, enable the implementation of the functions / actions specified in one or more blocks of the flowchart illustrations and / or block diagrams. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor, or a field-programmable logic circuit. It is also understood that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can also be implemented by special-purpose hardware performing the specified functions or actions, or can be implemented by a combination of special-purpose hardware and computer instructions.

[0128] The above description is merely a specific implementation of this application. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, modules, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. It should be understood that the protection scope of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the protection scope of this application.

Claims

1. A shooting control method, characterized in that, The method is applied to a mobile terminal and includes: Obtain the user's target voice information; The target speech information is input into the instruction parsing model to obtain the target speech command corresponding to the target speech information; the instruction parsing model is trained based on at least one speech sample to be trained, the speech sample to be trained includes speech information and a first instruction label and a second instruction label corresponding to the speech information, the first instruction label is used to indicate whether it is necessary to capture an image, and the second instruction label is used to indicate the type of image capture. In response to the target voice command indicating that an image needs to be captured, an image capture instruction corresponding to the target voice command is sent to the TWS earphone so that the TWS earphone captures a scene image. Receive the scene image sent by the TWS earphone.

2. The method according to claim 1, characterized in that, The TWS earphones include a first earphone and a second earphone, and the scene image includes a first scene image sent by the first earphone and a second scene image sent by the second earphone. After receiving the scene image sent by the TWS earphone, the method further includes: Receive target inertial sensor data sent by the TWS earphone; the target inertial sensor data is the inertial sensor data of the TWS earphone when capturing the first side scene image and the second side scene image; Based on the target inertial sensor data, the relative pose between the first side scene image and the second side scene image is determined; Based on the relative pose, the first side scene image and the second side scene image are stitched together to obtain the target scene image.

3. The method according to claim 1, characterized in that, Before inputting the target speech information into the instruction parsing model to obtain the target speech command corresponding to the target speech information, the method further includes: At least one of the speech samples to be trained and the model to be trained are obtained; wherein, the model to be trained is constructed from a target self-distillation model, a first classification head and a second classification head, the target self-distillation model is a model with semantic feature extraction capability obtained by self-distillation pre-training using unlabeled speech samples, the first classification head is used to output the probability that the speech sample to be trained indicates whether an image needs to be captured, and the second classification head is used to output the probability that the speech sample to be trained indicates various types of image capture. Based on the speech samples to be trained, the model parameters of the target self-distillation model, the first classification head parameters of the first classification head, and the second classification head parameters of the second classification head are updated to obtain the instruction parsing model.

4. The method according to claim 1, characterized in that, After receiving the scene image sent by the TWS earphone, the method further includes: The scene image is input into the image recognition model to extract the iconic elements in the scene image; Search the background knowledge base for descriptive information about elements associated with the iconic element; The element description information is sent to the TWS earphones to introduce the iconic element to the user.

5. The method according to claim 1, characterized in that, After receiving the scene image sent by the TWS earphone, the method further includes: Obtain the user's historical image preference style; Based on the historical image preference style, determine the personalized optimization strategy corresponding to the historical image preference style; Based on the personalized optimization strategy, the scene image is optimized to obtain an optimized scene image.

6. A shooting control method, characterized in that, The method is applied to TWS earphones and includes: Obtain image capture instructions; In response to the image capture command, perform the image capture action corresponding to the image capture command to obtain the scene image; The scene image is sent to the mobile terminal.

7. The method according to claim 6, characterized in that, The image capture command includes any one of the following: In response to the interaction between the user and the target button of the TWS earphone, the image capture command corresponding to the interaction is obtained; In response to the user's target voice information, the target voice information is matched with a preset voice command template to determine the image capture instruction corresponding to the target voice information; Receive the image capture command sent by the mobile terminal.

8. The method according to claim 6, characterized in that, The step of responding to the image capture command by executing an image capture action corresponding to the image capture command to obtain a scene image includes: In response to the image capture command, the camera parameters of the TWS earphone are adjusted until the camera parameters meet the preset capture conditions; In response to the camera parameters meeting the preset shooting conditions, an image shooting action corresponding to the image shooting command is executed to obtain the scene image.

9. A shooting control device, characterized in that, The device is applied to a mobile terminal and includes: The information acquisition module is used to acquire the user's target voice information; The command determination module is used to input the target speech information into the command parsing model to obtain the target speech command corresponding to the target speech information; the command parsing model is obtained by training the training model based on at least one training speech sample, the training speech sample includes speech information and a first command label and a second command label corresponding to the speech information, the first command label is used to indicate whether it is necessary to capture an image, and the second command label is used to indicate the type of image capture. The instruction sending module is used to respond to the target voice command indicating that an image needs to be captured, and send an image capture instruction corresponding to the target voice command to the TWS earphone so that the TWS earphone captures scene images; The image receiving module is used to receive the scene images sent by the TWS earphones.

10. A shooting control device, characterized in that, The device is used in TWS earphones and includes: The instruction acquisition module is used to acquire image capture instructions; The image capture module is used to respond to the image capture command, execute the image capture behavior corresponding to the image capture command, and obtain the scene image; The image transmission module is used to send the scene image to the mobile terminal.