Snapshot method, model training method, electronic equipment and training equipment

By combining a lightweight visual feature extraction model with a visual language model, the accuracy and real-time performance issues of capturing exciting moments in dynamic scenes have been solved, achieving efficient capture results and improving user experience.

CN121908116APending Publication Date: 2026-04-21HONOR DEVICE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HONOR DEVICE CO LTD
Filing Date
2024-10-11
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately and promptly capture fleeting moments in dynamic scenes, leading to a degraded user experience.

Method used

A lightweight visual feature extraction model is trained using a model distillation technique based on a visual language model. By combining visual features and preset key text features, the captured image is determined by calculating the first and second key text scores, thus achieving real-time capture.

Benefits of technology

It improves the accuracy and real-time performance of image capture, reduces latency and power consumption, and enhances the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121908116A_ABST
    Figure CN121908116A_ABST
Patent Text Reader

Abstract

The invention provides a snapshot method, a model training method, electronic equipment and training equipment, relates to the field of photographing, and can snapshot wonderful moments in a preset dynamic scene in real time. The method comprises the following steps: acquiring a to-be-determined frame set from a preview image displayed in a preview interface by the electronic equipment under the condition that the preview interface is displayed and a snapshot function is started; the electronic equipment obtains visual features corresponding to the first preview image based on a visual feature extraction model; the electronic equipment fuses the visual features of all the first preview images to obtain visual fusion features; the electronic device determines a first wonderful score and a second wonderful score of the target preview image based on the visual fusion feature and a preset wonderful text feature; and the electronic device determines a snapshot image based on the first wonderful score and the second wonderful score.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The application relates to the field of photography, and in particular to a snapshot method, a model training method, electronic devices, and training equipment. Background Technology

[0002] As the imaging capabilities of mobile phones and other electronic devices with camera functions continue to improve, users can usually easily capture photos of posed portraits, landscapes, architecture, and other wonderful moments in daily life. However, when faced with fleeting moments, users often cannot control their electronic devices to take pictures in time, resulting in the devices failing to capture photos that capture the perfect moment and thus reducing the user experience. Summary of the Invention

[0003] This application provides a snapshot method, a model training method, an electronic device, and a training device, which can capture exciting moments in a preset dynamic scene in real time.

[0004] To achieve the above objectives, the embodiments of this application adopt the following technical solutions:

[0005] In a first aspect, embodiments of this application provide a snapshot method applied to an electronic device. The method includes: when the electronic device displays a preview interface and activates the snapshot function, acquiring a set of pending frames from the preview images displayed in the preview interface; the set of pending frames includes a latest preset number of first preview images;

[0006] The electronic device obtains the visual features corresponding to the first preview image based on a visual feature extraction model. The visual feature extraction model is obtained by model distillation based on a visual language model. The visual language model has the ability to process image and text information simultaneously and obtain output information corresponding to both image and text information. The difference between the number of parameters of the visual language model and the number of parameters of the visual feature extraction model is greater than a first threshold.

[0007] The electronic device fuses the visual features of all the first preview images to obtain visual fusion features;

[0008] The electronic device determines a first highlight score and a second highlight score for the target preview image based on visual fusion features and preset highlight text features. The preset highlight text features are the text features of the highlight text corresponding to the highlight images of highlight moments in a preset dynamic scene. The first highlight score is used to characterize the similarity between the target preview image and the highlight text. The second highlight score is used to characterize the matching degree between the target preview image and the highlight text. The target preview image is the first preview image generated in the target order from the set of undetermined frames.

[0009] The electronic device determines the image to be captured based on the first and second highlight scores.

[0010] Based on the technical solution provided in this application, visual features of multiple preview images can be extracted first, and then combined with preset brilliant text features to obtain a first brilliant score and a second brilliant score for the target preview image. Subsequently, brilliant frames that can be used as capture images can be determined based on the first and second brilliant scores. Then, upon receiving a capture command, the corresponding brilliant frames can be output as capture images. In this technical solution, the visual feature extraction model for extracting visual features of the preview image is obtained through model distillation of a large visual language model with multimodal understanding capabilities; the text and image spaces of the two are consistent. Through model distillation of the large visual language model, the representational ability of the visual feature extraction model can be improved, enhancing its understanding of multimodal data. This allows the visual feature extraction model to better extract visual features from the preview image, and ensures that the visual features extracted by the model can smoothly perform similarity and matching calculations with the preset brilliant visual features and preset brilliant text features generated by the large visual language model. In this way, the first and second "highlight scores," derived from the visual features extracted by the visual feature extraction model and the preset "highlight text" features generated by the visual language big data model, can better characterize the correlation between the target preview image and the "highlight images" and "highlight text." Therefore, by referencing both visual and textual features—two multimodal information sources—the captured image can be accurately determined. Furthermore, because the visual feature extraction model is small in size, has low latency, and low power consumption, this solution can capture "highlight moments" in preset dynamic scenes in real time. Compared to existing technologies, it can improve capture accuracy while reducing latency and power consumption, thus enhancing the user experience.

[0011] Furthermore, because this technical solution employs a lightweight visual feature extraction model, it can process multiple frames of images in real time to obtain the visual features of those frames, and then determine the capture image based on these visual features. This effectively avoids false detection problems (i.e., obtaining inaccurate capture images) caused by unclear or unstable visual features in a single frame.

[0012] In one possible implementation of the first aspect, the electronic device obtains a set of frames to be determined from the preview images displayed in the preview interface, including: the electronic device extracts a preview image from the preview images displayed in the preview interface as a first preview image according to a preset frame interval, so as to obtain a set of frames to be determined.

[0013] Based on the technical solution corresponding to the above implementation method, since the multiple first preview images in the undetermined frame set are extracted from the preview images displayed on the preview interface according to a preset frame interval, the number of images processed by the final capture method can be reduced, thus reducing the computational load. Furthermore, since the preset frame interval is not large, it does not affect the determination of the final captured image.

[0014] In one possible implementation of the first aspect, the electronic device obtains visual features corresponding to the first preview image based on a visual feature extraction model, including: when the electronic device determines that a first object exists in all the first preview images in the set of undetermined frames, and the second object in each of the first preview images in the set of undetermined frames is the same object, it obtains visual features corresponding to the first preview image based on the visual feature extraction model; wherein the first object includes any one of the following: dog, cat, or human; and the second object is the first object that occupies the largest area.

[0015] Based on the technical solutions corresponding to the above implementation methods, since the purpose of the technical solutions provided in this application embodiment is to obtain photos (i.e., captured images) of exciting moments in a specific dynamic scene, and the main objects in this specific dynamic scene are fixed first objects (or captured objects), such as people, cats, dogs, etc. Therefore, for preview images that do not contain target objects, there is no need to perform subsequent processing of the capture algorithm. This also avoids processing preview images that cannot be used as captured images, reducing the waste of computing resources. In addition, if it is determined that the first objects exist in all the first preview images in the undetermined frame set, it can be considered that it is possible to determine the captured image through the first preview images. Furthermore, the captured image generally contains only a single main object. Since the first preview images in the undetermined frame set are generated at very close times, under normal circumstances, if the first preview image can be used to determine the captured image, the same first object is likely to exist in the first preview image, and the area occupied by the target object (the area in the image) is the largest in all the first preview images. If the above situation does not exist in the first preview image, it can be considered that there is no unified main target in the first preview images of the undetermined frame set, and therefore, the undetermined frame set cannot be used to determine the capture image. Therefore, to avoid wasting computational resources by still processing the first preview images in the undetermined frame set even when there is no unified main target (i.e., the capture object), it is necessary to further determine whether there is a unified capture object in the first preview images, even if the first object is determined to exist in all the first preview images in the undetermined frame set. Only when the first object is determined to exist in all the first preview images in the undetermined frame set, and the second object in each first preview image in the undetermined frame set is the same object, can it be considered that there is a unified capture object in the first preview images in the undetermined frame set, and therefore, these first preview images can be used to determine the capture image.

[0016] Therefore, the technical solution based on the above implementation method can avoid further feature extraction when the first preview image in the undetermined frame set does not meet the capture conditions, thus reducing the resource consumption during the capture method's operation.

[0017] In one possible implementation of the first aspect, the electronic device obtains visual features corresponding to the first preview image based on a visual feature extraction model, including: the electronic device cropping the first preview image based on the category of the second object to obtain a second preview image; the electronic device inputting the second preview image into the visual feature extraction model to obtain visual features of a second preset image; the visual features of the second preset image are the visual features corresponding to the first preview image to which the second preset image belongs.

[0018] To avoid the influence of redundant background content in the first preview image on the subsequent judgment of the captured image, the first preset image can be cropped based on the second object to reduce redundant background content. Since the purpose of capturing is primarily to capture a specific moment of the subject's action, the determination of the subject's action may require the assistance of other objects in the first preview image. For example, besides a person jumping in the air, at least a basketball is needed to determine if it's an aerial shooting motion. Therefore, assuming the second object in each first preview image in the undetermined frame set is the same object, the cropping of the first preset image should not only include the second object, but also other content within a certain range of the second object. Therefore, based on the above implementation method, redundant background content in the first preview image can be reduced, the computational load of the subsequent visual feature extraction model can be decreased, the execution efficiency of the capture method can be improved, and the adverse effects of background content on the capture method's execution process can be avoided.

[0019] In one possible implementation of the first aspect, the visual features include first sub-visual features and second sub-visual features, and the visual fusion features include first sub-visual fusion features and second sub-visual fusion features; the electronic device fuses the visual features of all the first preview images to obtain the visual fusion features, including: the electronic device determining the average value of the first sub-visual features of a preset number of second preview images as the first sub-visual fusion features; the electronic device using a preset feature fusion model to fuse the second sub-visual features of the preset number of second preview images to obtain the second sub-visual fusion features.

[0020] Based on the technical solution corresponding to the above implementation method, the mobile phone can use a suitable fusion method to obtain the fused first sub-visual fusion feature and second sub-visual fusion feature of the first preview image (i.e., the first sub-visual feature and second sub-visual feature of the second preview image), which provides data support for the subsequent calculation of the first and second highlights score, and enables the entire capture method to be implemented smoothly.

[0021] In one possible implementation of the first aspect, before the electronic device fuses the second sub-visual features of a preset number of second preview images using a preset feature fusion model to obtain the second sub-visual fusion features, the method further includes: the electronic device fusing preset temporal features with the second sub-visual features of the second preview images to update the second sub-visual features of the second preview images; the preset temporal features are used to characterize the temporal change characteristics of the image content in the video, which includes all the exciting moments of the preset dynamic scene.

[0022] In this way, by incorporating preset temporal features into the second sub-visual features, the second sub-visual features can reflect the temporal changes in the actions of the main object in the preset dynamic scene. Consequently, the second sub-visual fusion feature at the point of fusion can also better reflect this. Based on this, the second highlight score calculated based on the second sub-visual fusion feature can better determine the captured image, allowing the phone to obtain more accurate captured images and improving the user experience.

[0023] In one possible implementation of the first aspect, the preset brilliant text features include a first sub-brilliant text feature and a second sub-brilliant text feature; the electronic device determines a first brilliant score and a second brilliant score of the target preview image based on the visual fusion features and the preset brilliant text features, including: the electronic device calculates a first brilliant score based on the first sub-visual fusion features and the first sub-brilliant text features; the electronic device processes the second sub-visual fusion features and the second sub-brilliant text features using a multimodal fusion model to obtain a second brilliant score.

[0024] Based on the technical solutions corresponding to the above implementation methods, the mobile phone can use appropriate methods to obtain the first and second highlights scores, providing data support for the subsequent determination of the captured image.

[0025] In one possible implementation of the first aspect, the electronic device calculates a first brilliant score based on the first sub-visual fusion feature and the first sub-brilliant text feature, including: the electronic device calculates the image-text contrast loss ITC based on the first sub-visual fusion feature and the first sub-brilliant text feature; the electronic device calculates the image-image contrast loss IIC based on the first sub-visual fusion feature and the preset brilliant image feature; and the electronic device determines the weighted average of the ITC and IIC as the first brilliant score.

[0026] Based on the technical solution corresponding to the above implementation method, the first highlight score determined by IIC and ITC can not only characterize the similarity between the target preview image and the highlight text corresponding to the first sub-highlight text feature, but also characterize the similarity between the target preview image and the highlight image corresponding to the preset highlight visual feature. Therefore, it can also characterize the correlation between the target preview image and the highlight moments in the preset dynamic scene. Subsequently, when determining the capture image based on this first highlight score, the accuracy can be improved.

[0027] In one possible implementation of the first aspect, the electronic device determines the captured image based on a first highlight score and a second highlight score, including: the electronic device determines a highlight frame based on the first highlight score and the second highlight score; and the electronic device, in response to a capture command, determines the highlight frame corresponding to the capture command as the captured image.

[0028] Based on the technical solution corresponding to the above implementation method, the electronic device can first determine the highlights, and then determine the corresponding highlights as the captured image only when it receives the capture command, so that the final captured image better meets the user's needs.

[0029] In one possible implementation of the first aspect, the electronic device determines a highlight frame based on a first highlight score and a second highlight score, including: if the second highlight score is greater than a probability threshold, the electronic device determines a highlight window based on the maximum number of delayed frames and obtains the first highlight score of the first preview image in the highlight window; the highlight window is a combination of several consecutive preview images of the maximum delay frame, and the latest preview image in the highlight window is the latest first preview image in the set of undetermined frames; the maximum number of delayed frames is greater than a preset number; if the first highlight score of the first target preview image is the highest among the historical first preview images in the highlight window, the first target preview image is determined as a highlight frame; the historical first preview image is the first preview image in the highlight window whose generation time is before the target preview image.

[0030] Based on the technical solution described above, the mobile phone can determine the best frames that meet the latency requirements and can be used as captured images, using a pre-set maximum latency frame count. Furthermore, since this method for determining best frames combines a first best frame score (determined by ITC or IIC and ITC) and a second best frame score (i.e., ITM), and the second best frame score is less affected by the background in the image, the final determined best frames are more accurate.

[0031] Secondly, a model training method is provided, applied to a training device. The method includes: the training device acquiring sample data and corresponding label information; wherein, the sample data pair includes multiple first sample data, multiple second sample data, and second sub-highlight text features; the second sub-highlight text features are text features of highlight text corresponding to highlight images of highlight moments in a preset dynamic scene; the first sample data pair includes sample visual fusion features of a sample frame set, and the second sample data includes sample visual features of sample images; the sample frame set includes a preset number of sample video frames, which are video frames from videos in any scene, and the generation order of the preset number of sample video frames is arranged in an arithmetic sequence, with the common difference being the same as the preset frame interval; the sample images are images from any scene.

[0032] The training device uses sample data as training data and the label information corresponding to the sample data as supervision information. Iterative training is used to obtain an initial multimodal fusion model, and then a multimodal fusion model is obtained.

[0033] Based on the above technical solution, a multimodal fusion model can be trained through supervised learning. This multimodal fusion model has the ability to obtain the matching degree between visual features and second-sub-highlight text features (i.e., the matching degree between the image corresponding to the visual feature and the highlight text corresponding to the second-sub-highlight text feature). This provides data support for subsequently determining the second highlight score. Furthermore, this training method uses two types of training data—sample frame sets and sample images—to train the multimodal fusion model, reducing the difficulty of obtaining sample data and effectively improving the model's convergence efficiency.

[0034] Thirdly, this application provides an electronic device including a display, a memory, and one or more processors; the display and the memory are both coupled to the processors; wherein the memory stores computer program code, the computer program code including computer commands, and when the computer commands are executed by the processor, the electronic device performs the snapshot method provided by the first aspect and any possible design thereof.

[0035] Fourthly, this application provides a computer-readable storage medium including computer commands that, when executed on an electronic device, cause the electronic device to perform the snapshot method provided by the first aspect and any possible design thereof.

[0036] Fifthly, this application provides a computer program product that, when run on a training device, causes the training device to perform a model training method as provided in the first aspect and any of its possible design methods.

[0037] Sixthly, this application provides a training apparatus including a processor and a memory for storing processor-executable instructions. The processor is configured to execute the executable instructions to implement the model training method provided in the second aspect.

[0038] In a seventh aspect, this application provides a computer-readable storage medium including computer commands that, when executed on a training device, cause the training device to perform a model training method as provided in the first aspect and any of its possible design embodiments.

[0039] Eighthly, this application provides a computer program product that, when run on a training device, causes the training device to perform a model training method as provided in the first aspect and any possible design thereof.

[0040] Understandably, the beneficial effects that the technical solutions provided in the second to fourth aspects described above can be achieved can be referred to the beneficial effects of the first aspect and any of its possible design methods, which will not be repeated here. Attached Figure Description

[0041] Figure 1 A schematic diagram of the highlights provided in the embodiments of this application;

[0042] Figure 2 A schematic diagram illustrating the principle of a snapshot method provided in this application embodiment;

[0043] Figure 3 A schematic diagram of an implementation environment provided for an embodiment of this application;

[0044] Figure 4 A schematic diagram of the hardware architecture of an electronic device provided in an embodiment of this application;

[0045] Figure 5 A schematic diagram of the software architecture of an electronic device provided in an embodiment of this application;

[0046] Figure 6 This is a schematic diagram of the structure of a training device provided in an embodiment of this application;

[0047] Figure 7 A flowchart illustrating a snapshot method provided in this application embodiment. Figure 1 ;

[0048] Figure 8 This application provides an example of enabling a snapshot function. Figure 1 ;

[0049] Figure 9 This application provides an example of enabling a snapshot function. Figure 2 ;

[0050] Figure 10 This application provides an example of enabling a snapshot function. Figure 3 ;

[0051] Figure 11 A flowchart illustrating a snapshot method provided in this application embodiment. Figure 2 ;

[0052] Figure 12 A cropping diagram of a first preview image provided in an embodiment of this application;

[0053] Figure 13 A comparative schematic diagram of a second preview image and a first preview image provided for an embodiment of this application;

[0054] Figure 14A flowchart illustrating a snapshot method provided in this application embodiment. Figure 2 ;

[0055] Figure 15 A flowchart illustrating a snapshot method provided in this application embodiment. Figure 3 ;

[0056] Figure 16 A flowchart illustrating a snapshot method provided in this application embodiment. Figure 4 ;

[0057] Figure 17 A flowchart illustrating a snapshot method provided in this application embodiment. Figure 5 ;

[0058] Figure 18 A schematic diagram illustrating a training method for a multimodal fusion model provided in an embodiment of this application;

[0059] Figure 19 A flowchart illustrating a snapshot method provided in this application embodiment. Figure 6 ;

[0060] Figure 20 This is a schematic diagram illustrating the effect of determining a highlight frame in an embodiment of this application.

[0061] Figure 21 A flowchart illustrating a snapshot method provided in this application embodiment. Figure 7 . Detailed Implementation

[0062] The terminology used in the following embodiments of this application is for the purpose of describing particular embodiments only and is not intended to be limiting of this application. As used in the specification and appended claims of this application, the singular expressions “a,” “an,” “the,” “the,” “the,” and “this” are intended to include the plural expressions as well, unless the context clearly indicates otherwise. It should also be understood that “ / ” means “or,” for example, A / B can mean A or B; “and / or” in the text is merely a description of the relationship between related objects, indicating that three relationships can exist, for example, A and / or B can mean: A alone, A and B simultaneously, and B alone.

[0063] In this application, the reference to "embodiment" means that a specific feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a mutually exclusive, independent, or alternative embodiment. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described in this application can be combined with other embodiments.

[0064] The terms "first" and "second" in the following embodiments of this application are for descriptive purposes only and should not be construed as implying relative importance or implicitly indicating the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature, and in the description of the embodiments of this application, unless otherwise stated, "multiple" means two or more.

[0065] First, the terms used in this application are explained as follows:

[0066] (1) Image-text contrastive loss (ITC): ITC is a loss function used to measure the similarity between an image and text. ITC can be used to calculate the image-text contrastive loss ITC value, which can be used to characterize the degree of similarity between an image and text. The greater the similarity, the higher the ITC value. That is, the closer the content of a text is to the scene depicted by the image, the greater the ITC value between the text and the image.

[0067] (2) Image-Text Contrastive Loss (IIC): IIC can be understood as a loss function used to measure the similarity between two images. IIC can be used to calculate the image-text contrastive loss IIC value, which can be used to characterize the degree of similarity between images. The greater the similarity, the higher the IIC value. That is, the closer the content of two images is, the greater the IIC value between the two images.

[0068] (3) Image-text matching loss (ITM): ITM is a loss function used to determine whether input text matches an image; essentially, it's a binary classification problem. The purpose of ITM is to learn multimodal image-text labels, capturing fine-grained alignment between visual and linguistic elements. ITM is a binary classification task, and the corresponding model uses an ITM head (which can be a fully connected layer) to predict whether an image-text pair is positive (match) or negative (mismatch) given corresponding multimodal features (e.g., image features and text features). If an image and a text match perfectly, their ITM values ​​can be represented by "1"; if they don't match at all, their ITM values ​​can be represented by "0". The ITM loss function minimizes the prediction error of mismatched pairs (i.e., mismatched image-text pairs) while maximizing the prediction accuracy of matched pairs (i.e., matched image-text pairs).

[0069] (4) Object Detection Algorithms: Object detection algorithms are an important technology in the field of computer vision, designed to identify and locate objects of interest (i.e., targets) from images or videos. These objects can be any entity, such as people, vehicles, animals, buildings, etc. Object detection algorithms are generally divided into two categories: those based on traditional methods and those based on deep learning.

[0070] Traditional methods typically rely on hand-designed feature extractors and classifiers. Specifically, they may include the following steps: candidate region selection, feature extraction, classifier, and post-processing. Candidate region selection involves generating a series of candidate regions using methods such as sliding windows. Feature extraction can utilize hand-designed feature extractors such as scale-invariant feature transform (SIFT), speedup robust features (SURF), or histogram of oriented gradients (HOG) to extract features from the candidate regions. The classifier can use machine learning algorithms such as support vector machine (SVM) or AdaBoost to classify the extracted features and determine whether the candidate regions contain the target. Post-processing involves using methods such as non-maximum suppression (NMS) to remove duplicate detection boxes and improve detection accuracy.

[0071] Deep learning-based methods typically use neural network models such as convolutional neural networks (CNNs) to automatically extract features and directly perform target localization and classification. For example, the following series of neural network models can be used: R-CNN (region) series, YOLO (you only look once) series, SSD (Single Shot MultiBox Detector), EfficientDet, etc. Among them, the R-CNN series can include R-CNN, SPP-Net, Fast R-CNN, Faster R-CNN, etc.; the YOLO series can include YOLO, YOLOv2, YOLOv3, etc.

[0072] In R-CNN, selective search is first used to generate candidate regions, then a CNN is used to extract features, and an SVM is used for classification to determine the target region. SPP-Net, compared to R-CNN, further introduces spatial pyramid pooling to solve the problem of redundant feature computation in R-CNN. Fast R-CNN further accelerates feature extraction through ROI (Region of Interest) pooling layers, while simultaneously performing classification and bounding box regression, improving the efficiency of object detection. Faster R-CNN introduces a Region Proposal Network (RPN) to achieve end-to-end object detection, further improving object detection efficiency. Overall, the R-CNN series of neural network models employ a two-stage processing mode for object detection: first, candidate regions are selected, and then objects within these candidate regions are identified and classified to determine the target.

[0073] YOLO treats object detection as a regression problem, directly predicting bounding boxes and class probabilities on the image, achieving real-time object detection. Compared to the R-CNN series, it is more efficient in object detection. YOLOv2, building upon YOLO, introduced batch normalization and anchor boxes, improving both accuracy and speed. YOLOv3 further utilizes deeper (more layers) neural network structures and introduces multi-scale prediction, further enhancing object detection performance.

[0074] In SSD, the regression idea of ​​YOLO and the anchor box mechanism of Faster R-CNN are combined to realize the detection of multi-scale feature maps, which improves the detection accuracy and speed.

[0075] In EfficientDet, based on the backbone network of EfficientNet, combined with composite scaling methods and a weighted bi-directional feature pyramid network (BiFPN), efficient and high-precision target detection is achieved.

[0076] (5) Multi-object tracking (MOT / MTT) algorithms: Multi-object tracking is a computer vision task designed to analyze video to identify and track objects belonging to one or more categories, such as pedestrians, cars, animals, and inanimate objects, without any prior knowledge of the object's appearance or number. In single-object tracking, the appearance of the object is known in advance. In multi-object tracking, a detection step is needed to identify objects entering or leaving the scene. Specifically, multi-object tracking can include the following steps: object detection, object tracking, and object association. Object detection can use object detection algorithms to locate objects in images and videos. Object tracking is the process of tracking objects in consecutive image frames. Object tracking algorithms need to use the object's appearance features and motion information to infer the object's position in subsequent frames. Common object tracking algorithms include correlation filter-based methods (such as mean filters, kernel correlation filters, etc.), particle filter-based methods (such as Kalman filters, particle filters, etc.), and deep learning-based methods (such as Siamese networks, MDNet, etc.). Finally, object association can associate the tracking results of the object in different frames to maintain the consistency of the object's identity.

[0077] (6) Model Distillation: Model distillation is a model compression technique. Its basic principle is to transfer knowledge from a complex, large model (teacher model) to a smaller model (student model), thereby compressing and accelerating the model. The teacher model is usually a pre-trained complex network with high performance and generalization ability, while the student model is a network with fewer parameters and lower computational complexity. By guiding the student model's learning through the output of the teacher model, the student model can learn the teacher model's "dark knowledge," thus significantly reducing the model's complexity while maintaining high performance.

[0078] Model distillation is widely used in various machine learning tasks, especially in scenarios requiring model compression and acceleration. For example, in image classification tasks, a teacher model can learn rich image features and classification knowledge by extracting and classifying features from a large number of images. A small student model can then learn this knowledge from the output of the teacher model, thus achieving better performance in classification tasks.

[0079] Model distillation generally includes the following steps:

[0080] First, a complex, large model (i.e., the teacher model) is trained to learn on a large amount of data, thereby achieving high performance and generalization ability.

[0081] Secondly, prepare the training dataset for training the teacher model and the student model.

[0082] Next, the training data from the training dataset is input into the teacher model and the student model, and then the output signal of the teacher model is used as supervision information to train the student model.

[0083] In addition, during the training of the student model, some hyperparameters, such as the temperature parameter, can be adjusted to optimize the distillation effect.

[0084] Of course, in practice, model distillation can also employ any other possible methods, and this application does not impose any specific restrictions on this.

[0085] (7) Vision-language model (VLM): VLM is a model that combines visual and textual information to understand and generate content involving images and text. By fusing image and text data, VLM enables cross-modal interaction and reasoning. It can not only understand visual information in images, but also combine this information with textual descriptions to generate more accurate and richer outputs.

[0086] A Virtual Model (VLM) typically consists of an image encoder and a text encoder, which are responsible for extracting feature representations of the image and text, respectively. These feature representations are then interacted and fused in a fusion layer to generate the final output. During training, the VLM learns how to better fuse image and text information and generate a satisfactory output by optimizing the loss function.

[0087] VLM can process image and text information simultaneously, enabling cross-modal interaction and reasoning. This allows it to excel in a variety of tasks, such as image captioning and visual question answering.

[0088] By combining image and text information, VLM can generate more accurate and richer output. For example, in image (or video) captioning tasks, VLM can generate descriptive text for images that not only accurately reflects the content of the image but also contains rich semantic information. As another example, in visual question answering, VLM can obtain the answer text in the image that matches the question text based on the input image / video and the question text. Furthermore, in text-to-image tasks, VLM can generate an image that matches the input text.

[0089] As a multimodal large model, VLM has an extremely large number of parameters, ranging from billions to tens of billions; the number of model layers may be dozens or even more.

[0090] In this embodiment of the application, the visual language model used can be the VAST (vision-audio-subtitle-text) model.

[0091] (8) Moment of Brilliance: This refers to the moment within a given period when the subject's state and / or movement are at their best. In some examples, the subject can be stationary or in motion. When the subject is stationary, a moment of brilliance is the moment when the image of the subject is sharp and without blur. For example, when shooting a close-up portrait, a moment of brilliance is the moment when the person is smiling with their eyes open and the image is sharp. When the subject is in motion, a moment of brilliance is the image of the subject in motion. For example, when shooting a person jumping, a moment of brilliance is the moment when the person is in the air after jumping and the image is sharp. Images of moments of brilliance can be considered masterpieces.

[0092] Currently, with the continuous improvement of imaging capabilities in electronic devices such as smartphones, users can usually easily capture photos of wonderful moments in static scenes or scenes with slow changes. However, for example... Figure 1 The moment the dog / cat jumps in the air, as shown in (a), Figure 1 The pet playing as shown in (b) Figure 1 The moment the dog high-fives the person shown in (c) Figure 1 The moment of the basketball shot in the air shown in (d) Figure 1 The timing of the shuttlecock reception shown in (e) is as follows: Figure 1 In the dynamic scenes (or dynamic scenes) shown in (f), such as the key dance moves, users often cannot control their electronic devices to take pictures in time, thus failing to capture photos of the exciting moments that meet their needs.

[0093] To enable electronic devices to capture exciting moments in dynamic scenes in a timely manner, current camera applications on electronic devices can include a snapshot function. When the snapshot function is enabled, the electronic device can capture and save images of exciting moments from the preview screen in the camera preview interface using a specific snapshot method.

[0094] In related technologies, there are three main methods for image capture:

[0095] The first capture method is based on a human keypoint detection algorithm. This method extracts key points of the human body from the preview image and then judges the range of motion, opening and closing angles, etc., to identify exciting images (images of exciting moments) that match specific dynamic scenes.

[0096] The second capture algorithm is a visual feature matching method. This method extracts visual features from the preview image and calculates the similarity between these features and the visual features of a preset reference image. The frame with the highest similarity is then selected as the featured image. The preset reference image can be a featured image from a specific dynamic scene obtained through various possible means.

[0097] The third type of capture method can be based on a multimodal large model. This type of capture method can extract the visual features of the preview image and the text features of the text (used to describe the content of the pre-set reference image) using models such as CLIP (contrastive language-image pre-training) and Q-former. Then, it calculates the image-text similarity using methods such as ITC loss and cosine similarity, and then determines whether the preview image is a wonderful picture in the dynamic scene.

[0098] Of the three image capture methods mentioned above, the first method requires the identification of all key human body points. However, in reality, the shooting environment is highly variable, which can lead to insufficient key points being identified from the preview image, resulting in detection failure. Furthermore, the diversity of human postures also causes bias in key point detection. Therefore, this method has poor generalization and significant limitations, making it unsuitable for effectively capturing dynamic and exciting images.

[0099] In the second capture method, firstly, the extraction of visual features from the preview image is affected by the shooting environment, leading to unstable or inaccurate features. Secondly, since the preset reference image is relatively limited, if it fails to encompass the current shooting scene, it will result in failure to capture the image or obtaining an inaccurate captured image. Therefore, this capture method has poor generalization and significant limitations, and cannot effectively capture exciting images in dynamic scenes.

[0100] Furthermore, the neural network models used in the first and second capture methods are generally small in size, have insufficient training data, and lack multimodal understanding capabilities (simultaneously combining visual and textual features to recognize images). This further leads to poor generalization of the first and second capture methods, which may fail to accurately capture images once the interfering factors change.

[0101] The third capture method achieves excellent capture results due to the massive training data and complex structure of the multimodal large model, which possesses multimodal understanding capabilities that combine visual and textual features. However, the large multimodal large model also suffers from significant latency and power consumption due to its large number of parameters and complex structure, making it unsuitable for real-time capture requirements of mobile devices.

[0102] Based on the above description, it is clear that existing capture methods cannot adequately meet the accuracy and real-time requirements of mobile phones and other edge electronic devices when performing capture operations.

[0103] Based on this, and to address the aforementioned technical problems, this application provides a snapshot method that can be applied to scenarios where electronic devices perform real-time snapshots. (Refer to...) Figure 2 As shown, in this technical solution, when the camera application is open and the preview interface is displayed, the electronic device can acquire the latest multi-frame preview images and input them into the visual feature extraction model to obtain the visual features of the multi-frame preview images. Figure 2 This example uses only three preview images—Preview Image 1, Preview Image 2, and Preview Image 3—and their visual features (Visual Feature 1 of Preview Image 1, Visual Feature 2 of Preview Image 2, and Visual Feature 3 of Preview Image 3) as an example, and does not represent a specific limitation on the actual implementation. The visual feature extraction model is based on a visual language model and obtained using model distillation. This visual language model is a multimodal large model with a much larger number of parameters and a longer latency than the visual feature extraction model. The visual feature extraction model, on the other hand, has fewer parameters and a shorter latency than a preset threshold.

[0104] This visual language model has the ability to process both image and text information simultaneously, obtaining output information corresponding to both image and text information. For example, this visual language model can at least have the following capabilities: generating text based on images, generating images based on text, and generating answer text corresponding to the question text and image based on the question text and image.

[0105] After obtaining the visual features of multiple preview images, the visual features of the multiple preview images can be fused to obtain fused visual features. Then, based on the preset brilliant text features and preset brilliant visual features, the first brilliant score and the second brilliant score of the target preview image in the multiple preview images can be determined.

[0106] Among them, the preset visual features are the visual features of the highlights of the images in the preset dynamic scene, and the preset text features are the text features of the corresponding text in the images. The preset visual features and preset text features can be obtained by inputting the images and corresponding text of the highlights of the preset dynamic scene into a visual language model.

[0107] The target preview image is one of a series of preview images, for example, it can be a preview image whose acquisition time is approximately the middle of the acquisition period. The acquisition period is the time interval during which the multi-frame preview images are acquired, with the start time being the earliest acquisition time among the multi-frame preview images and the end time being the latest acquisition time among the multi-frame preview images. The first "highlight score" is used to characterize at least the similarity between the target preview image and the preset "highlight text," or the first "highlight score" is used to characterize at least the similarity between the target preview image and the preset "highlight image" and the preset "highlight text"; the second "highlight score" is used to characterize the matching degree between the target preview image and the preset "highlight text."

[0108] After obtaining the first and second highlight scores of the target preview image, a specific frame output strategy can be used to determine the highlight frames that can be used as capture images in the current target preview image and previous target preview images, and output these highlight frames. Alternatively, upon receiving a capture command, the highlight frame closest to the time the capture command was generated can be output as the capture image (specifically, it can be stored in a gallery).

[0109] Based on the technical solution provided in this application, the visual feature extraction model for extracting visual features from preview images is obtained through model distillation of a large visual language model with multimodal understanding capabilities, and the image and text spaces of the two are consistent. By distilling the large visual language model, the representational ability of the visual feature extraction model can be improved, enhancing its understanding of multimodal data. This allows the visual feature extraction model to better extract visual features from preview images, and ensures that the visual features extracted by this model can be smoothly compared to the preset high-quality visual features and preset high-quality text features generated by the large visual language model for similarity and matching calculations. Therefore, the first and second high-quality scores derived from the visual features extracted by this model, the preset high-quality visual features and preset high-quality text features generated by the large visual language model, can better characterize the correlation between the target preview image and the high-quality images and high-quality text. Consequently, by referencing both visual and textual features—two types of multimodal information—the captured image in the target preview image can be accurately identified. Furthermore, because the visual feature extraction model is small in size, has low latency, and low power consumption, this solution can improve the accuracy of image capture while reducing latency and power consumption, thus enhancing the user experience compared to existing technologies.

[0110] Furthermore, because this technical solution employs a lightweight visual feature extraction model, it can process multiple frames of images in real time to obtain the visual features of those frames, and then determine the capture image based on these visual features. This effectively avoids false detection problems (i.e., obtaining inaccurate capture images) caused by unclear or unstable visual features in a single frame.

[0111] The technical solutions provided in the embodiments of this application will be described in detail below with reference to the accompanying drawings.

[0112] The technical solution provided in this application can be applied to, for example... Figure 3 The implementation environment shown. (Refer to...) Figure 3 As shown, the real-time environment may include training device 01 and electronic device 02. The training device 01 and electronic device 02 can establish a communication connection via wired or wireless communication.

[0113] The training device 01 is mainly used to train a large visual language model and to train a visual feature extraction model through model distillation. The electronic device 02 is used to implement the image capture method provided in this application when the user needs to capture images, after obtaining the visual feature extraction model from the training device and the preset excellent visual features and preset excellent text features from the large visual language model.

[0114] Of course, in the embodiments of this application, if the computing and storage resources of the electronic device 02 are sufficient, all actions or processes performed by the training device 01 can also be implemented by the electronic device 02, and this application does not impose any specific restrictions on this.

[0115] It is understood that the aforementioned electronic device 02 and training device 01 can be two separate devices or the same device. This application does not impose any specific restrictions in this regard.

[0116] For example, the electronic devices in the above wireless charging system can be mobile phones, tablets, handheld computers, personal computers (PCs), ultra-mobile personal computers (UMPCs), netbooks, as well as cellular phones, personal digital assistants (PDAs), augmented reality (AR) devices, virtual reality (VR) devices, artificial intelligence (AI) devices, wearable devices, in-vehicle devices, smart home devices, and / or smart city devices, etc., any electronic devices with wireless charging capabilities. This application embodiment does not impose any special restrictions on the specific type of electronic device.

[0117] For example, taking a mobile phone as an electronic device, Figure 4 A schematic diagram of the structure of an electronic device provided in an embodiment of this application is shown.

[0118] Reference Figure 4 As shown, the electronic device may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, antenna 1, antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone jack 170D, a sensor module 180, buttons 190, a motor 191, an indicator 192, a display screen 193, a subscriber identification module (SIM) card interface 194, and a camera 195, etc. The sensor module 180 may include pressure sensors, gyroscope sensors, barometric pressure sensors, magnetic sensors, accelerometers, gravity sensors, distance sensors, proximity sensors, fingerprint sensors, temperature sensors, touch sensors, ambient light sensors, bone conduction sensors, etc.

[0119] Processor 110 may include one or more processing units, such as: application processor (AP), modem processor, graphics processing unit (GPU), image signal processor (ISP), controller, memory, video codec, digital signal processor (DSP), baseband processor, and / or neural network processing unit (NPU), etc. Different processing units may be independent devices or integrated into one or more processors.

[0120] A controller can be the nerve center and command center of an electronic device. Based on command opcodes and timing signals, the controller generates operation control signals to complete the control of command retrieval and execution.

[0121] The processor 110 may also include a memory for storing commands and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can store commands or data that the processor 110 has just used or that are used repeatedly. If the processor 110 needs to use the command or data again, it can retrieve it directly from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.

[0122] In some embodiments, the processor 110 may include one or more interfaces. Interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.

[0123] The charging management module 140 is used to receive charging input from wireless power supply devices (such as chargers, laptop batteries, etc.). The charger can be a wireless charger or a wired charger. In some wired charging embodiments, the charging management module 140 can receive charging input from the wired charger via the USB interface 130. In some wireless charging embodiments, the charging management module 140 can receive wireless charging input via the receiving coil in the wireless charging chip of the electronic device. Alternatively, the charging management module 140 includes a wireless charging chip, and can then receive wireless charging input via the receiving coil in the wireless charging chip. Of course, in some embodiments, the receiving coil can be set separately from the charging management module 140, in which case the charging management module can use the receiving coil to receive wireless charging input via the wireless charging chip. Furthermore, in other embodiments, the charging management module 140 can also wirelessly charge other electronic devices via the receiving coil.

[0124] While charging the battery 142, the charging management module 140 can also supply power to the electronic device through the power management module 141. Specifically, the battery 142 can be composed of multiple batteries connected in series. The power management module 141 connects the battery 142, the charging management module 140, and the processor 110. In some other embodiments, the charging management module 140 may also be located within the processor 110.

[0125] The power management module 141 connects the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140, providing power to the processor 110, internal memory 121, display screen 193, camera 195, and wireless communication module 160, etc. The power management module 141 can also monitor parameters such as battery voltage, current, battery cycle count, and battery health status (leakage current, impedance). In some other embodiments, the power management module 141 may also be located within the processor 110. In some embodiments, the charging management module 140 and the power management module 141 may be located in the same device.

[0126] The external memory interface 120 can be used to connect to external non-volatile memory, thereby expanding the storage capacity of the electronic device. The external non-volatile memory communicates with the processor 110 through the external memory interface 120 to perform data storage functions. For example, music, video, and other files can be stored in the external non-volatile memory.

[0127] Internal memory 121 may include one or more random access memory (RAM) and one or more non-volatile memory (NVM). The RAM can be directly read and written by the processor 110 and can be used to store executable programs (e.g., machine commands) of the operating system or other running programs, as well as user and application data. The NVM can also store executable programs and user and application data, and can be pre-loaded into the RAM for direct read and write operations by the processor 110.

[0128] A touch sensor, also known as a "touch device," can be located on the display screen 193. The touch sensor and the display screen 193 together form a touchscreen, also called a "touchscreen." The touch sensor detects touch operations applied to or near it. The touch sensor can transmit the detected touch operation to the application processor to determine the type of touch event. Visual output related to the touch operation can be provided through the display screen 193. In other embodiments, the touch sensor may also be located on the surface of the electronic device, in a different position than the display screen 193.

[0129] An ambient light sensor is used to detect ambient light intensity. A pressure sensor is used to sense pressure signals and can convert these signals into electrical signals. In some embodiments, the pressure sensor may be located on the display screen 193. There are many types of pressure sensors, such as resistive pressure sensors, inductive pressure sensors, and capacitive pressure sensors.

[0130] An accelerometer (G-sensor), also called a gravity sensor, is a device that can sense acceleration in any direction. A triaxial accelerometer works based on the fundamental principle of acceleration. Acceleration is a spatial vector; on the one hand, to accurately understand the motion of an object, its components on its three coordinate axes must be measured; on the other hand, in situations where the direction of the object's motion is unknown beforehand, only a triaxial accelerometer can detect the acceleration signal.

[0131] A gyroscope (GYRO-sensor), also known as a ground sensor, traditionally contains an internal gyroscope. A three-axis gyroscope can simultaneously measure position, trajectory, and acceleration in six directions. A single-axis gyroscope can only measure quantities in two directions, meaning a system typically requires three gyroscopes. A single three-axis gyroscope can replace three single-axis gyroscopes. The working principle of a three-axis gyroscope is to measure the angle between the vertical axis of the gyroscope rotor and the device in a three-dimensional coordinate system, and calculate the angular velocity. The angle and angular velocity are used to determine the object's motion state in three-dimensional space. A three-axis gyroscope can simultaneously measure six directions: up, down, left, right, forward, and backward (the composite direction can also be decomposed into three-axis coordinates), ultimately determining the device's trajectory and acceleration. In other words, by measuring its own rotation, the three-axis gyroscope determines the device's current motion state, such as forward, backward, up, down, left, or right; and whether it is accelerating (angular velocity) or decelerating (angular velocity).

[0132] In some embodiments, an electronic device may include one or N cameras 195, where N is a positive integer greater than 1. In this application embodiment, the type of camera 195 can be distinguished based on hardware configuration and physical location. For example, a camera located on the side of the electronic device's display screen 193 can be called a front-facing camera, and a camera located on the side of the electronic device's back cover can be called a rear-facing camera; another example is that a camera with a short focal length and a wide field of view can be called a wide-angle camera, while a camera with a long focal length and a narrow field of view can be called a regular camera. Here, focal length and field of view are relative concepts and are not specifically limited by parameters. Therefore, wide-angle cameras and regular cameras are also relative concepts, and can be specifically distinguished based on physical parameters such as focal length and field of view.

[0133] The electronic device implements display functions through a GPU, a display screen 193, and an application processor. The GPU is a microprocessor for image editing, connected to the display screen 193 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 110 may include one or more GPUs, which execute program commands to generate or modify display information.

[0134] Electronic devices can achieve shooting functions through ISP, camera 195, video codec, GPU, display 193, and application processor. The GPU is used to perform mathematical and geometric calculations and for graphics rendering. Processor 110 may include one or more GPUs, which execute program commands to generate or modify display information.

[0135] The Information Service Provider (ISP) is used to process data fed back from the camera 195. For example, when taking a picture, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element. The light signal is converted into an electrical signal, and the camera's photosensitive element transmits the electrical signal to the ISP for processing, transforming it into an image visible to the naked eye. The ISP can also perform algorithmic optimization on image noise and brightness. The ISP can also optimize parameters such as exposure and color temperature of the shooting scene. In some embodiments, the ISP can be integrated into the camera 195. The camera 195 is used to capture still images or videos.

[0136] Digital signal processors (DSPs) are used to process digital signals. Besides digital image signals, they can also process other digital signals. For example, when an electronic device is selecting a frequency, a DSP can perform a Fourier transform on the frequency energy.

[0137] Video codecs are used to compress or decompress digital video. Electronic devices can support one or more video codecs. This allows the electronic device to play or record video in various encoded formats, such as Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, MPEG4, etc.

[0138] Display screen 193 is used to display images, videos, etc. Display screen 193 includes a display panel. The display panel may be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a Mini LED, a MicroLED, a Micro-OLED, a quantum dot light-emitting diode (QLED), etc. In some embodiments, the electronic device may include one or N displays 193, where N is a positive integer greater than 1.

[0139] In this embodiment of the application, the display screen 193 can be used to display pages required by the electronic device (e.g., a page displaying captured images, etc.), and to display images captured by any one or more cameras 195 in the interface.

[0140] The wireless communication function of electronic devices can be realized through antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, modem, and baseband processor.

[0141] Antenna 1 and Antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in an electronic device can be used to cover one or more communication frequency bands. Different antennas can also be reused to improve antenna utilization.

[0142] The mobile communication module 150 can provide wireless communication solutions, including 2G / 3G / 4G / 5G, for use in electronic devices. The mobile communication module 150 can receive electromagnetic waves via antenna 1, and perform filtering, amplification, and other processing on the received electromagnetic waves before transmitting them to a modem processor for demodulation. The mobile communication module 150 can also amplify the signal modulated by the modem processor and convert it into electromagnetic waves for radiation via antenna 1. In some embodiments, at least some functional modules of the mobile communication module 150 can be housed in the processor 110. In some embodiments, at least some functional modules of the mobile communication module 150 and at least some modules of the processor 110 can be housed in the same device.

[0143] The modem processor may include a modulator and a demodulator. The modulator modulates the low-frequency baseband signal to be transmitted into a mid-to-high frequency signal. The demodulator demodulates the received electromagnetic wave signal into a low-frequency baseband signal. The demodulator then transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After processing by the baseband processor, the low-frequency baseband signal is transmitted to the application processor. The application processor outputs sound signals through audio devices (not limited to speaker 170A, receiver 170B, etc.) or displays images or videos through the display screen 193. In some embodiments, the modem processor may be a separate device. In other embodiments, the modem processor may be independent of the processor 110 and may be housed in the same device as the mobile communication module 150 or other functional modules.

[0144] The wireless communication module 160 can provide solutions for wireless communication applications in electronic devices, including wireless local area networks (WLANs) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), and infrared (IR) technologies. The wireless communication module 160 can be one or more devices integrating at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via antenna 2, performs frequency modulation and filtering of the electromagnetic wave signals, and sends the processed signal to processor 110. The wireless communication module 160 can also receive signals to be transmitted from processor 110, perform frequency modulation and amplification, and convert them into electromagnetic waves for radiation via antenna 2.

[0145] In some embodiments, antenna 1 of the electronic device is coupled to mobile communication module 150, and antenna 2 is coupled to wireless communication module 160, enabling the electronic device to communicate with networks and other devices via wireless communication technology. The wireless communication technology may include Global System for Mobile Communications (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Time-Division Code Division Multiple Access (TDSCDMA), Long Term Evolution (LTE), BT, GNSS, WLAN, NFC, FM, and / or IR technologies, etc. The GNSS may include the Global Positioning System (GPS), the Global Navigation Satellite System (GLONASS), the BeiDou Navigation Satellite System (BDS), the Quasi-Zenith Satellite System (QZSS), and / or satellite-based augmentation systems (SBAS).

[0146] The SIM card interface 194 is used to connect a SIM card. The SIM card can be inserted into or removed from the SIM card interface 194 to make contact with and detach from the electronic device. The electronic device can support one or more SIM card interfaces. The SIM card interface 194 supports Nano SIM cards, Micro SIM cards, and other SIM cards. Multiple cards can be inserted into the same SIM card interface 194 simultaneously. The SIM card interface 194 is also compatible with external memory cards. The electronic device interacts with the network through the SIM card to achieve functions such as calls and data communication. One SIM card corresponds to one user number.

[0147] It is understood that the interface connection relationships between the modules illustrated in the embodiments of the present invention are merely illustrative and do not constitute a limitation on the structure of the electronic device. In other embodiments of this application, the electronic device may also employ different interface connection methods or combinations of multiple interface connection methods as described in the above embodiments.

[0148] Of course, it is understandable that the above... Figure 4 The illustration shown is merely an example when the electronic device is in the form of a mobile phone. If the electronic device is in the form of a tablet, handheld computer, PC, PDA, wearable device (such as a smartwatch, smart bracelet), or other similar device, the structure of the electronic device may include more advanced features. Figure 4 The fewer structures shown can also include more than Figure 4 The structures shown are not limited here.

[0149] It is understandable that, generally speaking, the implementation of various functions in electronic devices requires not only hardware support but also software cooperation. The software system of electronic devices can adopt a layered architecture, event-driven architecture, microkernel architecture, microservice architecture, or cloud architecture. This application's embodiment uses a layered architecture... Taking the system as an example, the software structure of the electronic device is illustrated.

[0150] Figure 5 This is a schematic diagram of the layered architecture of the software system of the electronic device provided in the embodiments of this application. The layered architecture divides the software into several layers, each with a clear role and division of labor. The layers communicate with each other through software interfaces (e.g., APIs).

[0151] In some examples, refer to Figure 5 As shown in this embodiment, the software system located on the application processor (AP) in the system-on-a-chip (SOC) of the electronic device is divided into five layers, from top to bottom: application layer, framework layer (or application framework layer), system library and Android runtime, HAL layer (hardware abstraction layer), and kernel layer (or driver layer). The system library and Android runtime can also be referred to as the native framework layer or native layer.

[0152] The application layer can include a series of applications. For example... Figure 5 As shown, the application layer can include applications (APPs) such as camera, gallery, calendar, map, WLAN, Bluetooth, news, music, video, SMS, call, navigation, and instant messaging.

[0153] The framework layer provides application programming interfaces (APIs) and programming frameworks for applications in the application layer. The application framework layer includes predefined functions or services. For example, the application framework layer may include an activity manager, window manager, content provider, audio service, view system, phone manager, resource manager, notification manager, package manager, etc., but this embodiment does not impose any limitations on these.

[0154] The window manager is used to manage windowed applications. It can retrieve screen size, determine the presence of a status bar, lock the screen, and capture screenshots, among other things.

[0155] Content providers store and retrieve data, making that data accessible to applications. This data can include videos, images, audio, phone calls made and received, browsing history and bookmarks, phone books, etc.

[0156] A view system includes visual controls, such as controls for displaying text and controls for displaying images. View systems can be used to build applications. A display interface can consist of one or more views. For example, a display interface including a text notification icon could include views for displaying text and views for displaying images.

[0157] A phone manager is used to provide communication functionality for electronic devices. For example, a phone manager can manage the call status of a calling application (including initiation, connection, and termination).

[0158] The primary function of camera services is to provide applications with a unified interface and functionality for accessing and operating camera devices. Here are some of the main functions of camera services:

[0159] Camera Access: The camera service provides an interface to access camera hardware, enabling applications to communicate with camera devices. It abstracts away the underlying camera drivers and hardware details, hiding the differences in the underlying implementation, allowing applications to use camera functionality uniformly across different devices.

[0160] Camera Control: The camera service provides camera settings and control functions, such as adjusting parameters like exposure, focus, and flash, as well as switching between front and rear cameras. Applications can use the camera service to flexibly control and configure the camera to meet different photographic needs.

[0161] In this embodiment, the camera service may include a snapshot service, which provides image capture functionality, allowing the application to capture key frames from images acquired by the camera. Specifically, the snapshot service can call the snapshot algorithm in the HAL layer based on the camera HAL layer to implement the snapshot method provided in this embodiment.

[0162] The camera service also provides a callback mechanism for camera (i.e., camera module) status and events, enabling applications to promptly obtain information about changes in the camera module's status or events such as completing a shot. Through callbacks, applications can respond to camera operations in a timely manner, such as updating the UI or executing other logic.

[0163] In summary, the camera service plays a bridging role in the framework layer of the software system architecture. It encapsulates the underlying camera hardware details, provides a unified interface and functions, and offers applications convenient capabilities for camera access, control, image capture, and processing.

[0164] The file explorer provides applications with various resources, such as localized strings, icons, images, layout files, video files, and more.

[0165] The notification manager allows applications to display notifications in the status bar. These notifications can be used to deliver informational messages and can disappear automatically after a short pause, requiring no user interaction. For example, the notification manager can be used to notify users of completed downloads or message alerts. The notification manager can also display notifications as icons or scrolling text in the top status bar, such as notifications from background applications, or as dialog boxes on the screen. Examples include displaying text messages in the status bar, emitting sounds, vibrating electronic devices, and flashing indicator lights.

[0166] Package manager in The package manager is used to manage application packages. It allows applications to obtain detailed information about installed applications and their services, permissions, etc. The package manager is also used to manage events such as application installation, uninstallation, and upgrades.

[0167] Of course, in practice, the technical solution provided in this application can be implemented by any possible module in the application processor. The above is only an example and is not intended to impose specific limitations on the actual implementation.

[0168] The system library can include multiple functional modules. For example: a surface manager, a display compositing system, media libraries, open graphics library embedded systems (OpenGL ES), SGL, etc. The surface manager manages the display subsystem and provides 2D and 3D layer blending for multiple applications. The media libraries support playback and recording of various common audio and video formats, as well as still image files. The media libraries support various audio and video encoding formats, such as MPEG4, H.264, MP3, AAC, AMR, JPG, PNG, etc. OpenGL ES is used for 3D graphics drawing, image rendering, compositing, and layer processing. SGL is a 2D graphics engine. The display compositing system can specifically manage the display subsystem (e.g., controlling the display of an electronic device's screen to show a charging icon or adjust screen brightness), and provides layer generation or blending for multiple applications.

[0169] The Android runtime consists of the core libraries and the ART virtual machine. The Android runtime is responsible for scheduling and managing the Android system. The core libraries comprise two parts: one part contains the functionalities that Java code needs to call, and the other part consists of the Android core libraries. The application layer and application framework layer run in the ART virtual machine. The ART virtual machine executes the Java files of the application layer and application framework layer into binary files. The ART virtual machine is used for managing object lifecycles, stack management, thread management, security and exception management, and garbage collection.

[0170] The Hardware Abstraction Layer (HAL) is the interface layer between the operating system kernel and the hardware circuitry, designed to abstract the hardware. It hides the platform-specific hardware interface details, providing the operating system with a virtual hardware platform that is hardware-independent and portable across multiple platforms. The HAL provides a standard interface that exposes device hardware functionality to the higher-level Java API framework (i.e., the framework layer). The HAL contains multiple library modules, each implementing an interface for a specific type of hardware component, such as: audio HAL, Bluetooth HAL, camera HAL (also known as camera HAL or camera hardware abstraction module), sensors HAL (or i-sensor service), and display HAL (display module), etc.

[0171] In this embodiment of the application, in order to achieve image capture, the HAL layer may further include an image capture algorithm for implementing the image capture method of this embodiment. This image capture algorithm may include runtime code and data that implement the image capture method provided in this embodiment, such as a visual feature extraction model, a fusion algorithm for fusing visual features of multiple preview images, preset excellent text features, preset excellent visual features, methods for calculating the first excellent score and the second excellent score, and specific frame output strategies, etc.

[0172] In this embodiment, when a user opens the camera application, the camera application calls the camera service in the framework layer to start the application. Then, it calls the camera HAL in the HAL layer to instruct the camera driver to start the camera device (i.e., the camera). For example, the camera HAL can send a start command to the camera driver. Upon receiving the start command, the camera driver can drive the camera to collect external light signals to obtain a preview image stream (an image sequence composed of multiple preview images), and then send the preview image stream back to the HAL layer for display. Simultaneously, if the user has enabled the snapshot function, the camera HAL can process the preview images in the preview image stream in real time based on camera algorithms to determine the best frames that can be used as snapshot images. Then, upon receiving a snapshot command, the corresponding best frames can be stored in the image library.

[0173] The kernel layer is the layer between hardware and software. The kernel layer contains at least various drivers and a TCP / IP protocol stack. These drivers can include display drivers, camera drivers, audio drivers, sensor drivers, wireless charging drivers, etc., but this application does not limit the scope.

[0174] It should be noted that although the embodiments in this application use the Android system as an example for illustration, the basic principles are equally applicable to systems based on... Electronic devices using operating systems such as iOS and Windows.

[0175] For example, the wireless power supply device in the above-described wireless charging system can be a device capable of providing charging services to electronic devices, such as tablet computers, mobile phones, in-vehicle wireless charging docks, laptops, super mobile personal computers, netbooks or PDAs, and other devices capable of wirelessly charging other devices. The wireless power supply device has a charging area; when an electronic device is located within the charging area of ​​the wireless power supply device, the wireless power supply device can wirelessly charge the electronic device. This application embodiment does not impose any special limitations on the specific type of wireless power supply device.

[0176] For example, the training device 01 provided in this application can be a server, which can be a single server, a server cluster composed of multiple servers, or a cloud computing service center. This application does not impose any specific restrictions on this.

[0177] For example, taking the training device as the server, Figure 6 A schematic diagram of a server structure is shown. (Refer to...) Figure 6 As shown, the server includes one or more processors 601, a communication line 602, and at least one communication interface. Figure 6 (This is merely an example illustration of a communication interface 603 and a processor 601; optionally, a memory 604 may also be included.)

[0178] The processor 601 may be a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits used to control the execution of the program of the present application.

[0179] Communication line 602 may include a communication bus for communication between different components.

[0180] The communication interface 603 can be a transceiver module used to communicate with other devices or communication networks, such as Ethernet, RAN, and wireless local area networks (WLAN). For example, the transceiver module can be a transceiver or similar device. Optionally, the communication interface 603 can also be a transceiver circuit located within the processor 601, used to implement the processor's signal input and signal output.

[0181] The memory 604 can be a device with storage functionality. For example, it can be read-only memory (ROM) or other types of static storage devices capable of storing static information and instructions; random access memory (RAM) or other types of dynamic storage devices capable of storing information and instructions; electrically erasable programmable read-only memory (EEPROM); compact disc read-only memory (CD-ROM) or other optical disc storage; optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.); magnetic disk storage media or other magnetic storage devices; or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but not limited thereto. The memory can exist independently and be connected to the processor via communication line 602. The memory can also be integrated with the processor.

[0182] The memory 604 stores computer execution instructions for implementing the scheme of this application, and the processor 601 controls the execution. The processor 601 executes the computer execution instructions stored in the memory 604 to obtain the visual feature extraction model required in the scheme of this application or to execute the corresponding training method.

[0183] Alternatively, in this embodiment, the processor 601 may execute the processing-related functions in the behavior recognition model generation method provided in the following embodiments of this application, and the communication interface 603 may be responsible for communicating with other devices (e.g., electronic devices) or communication networks. This embodiment does not specifically limit this.

[0184] Optionally, the computer execution instructions in the embodiments of this application may also be referred to as application code, and the embodiments of this application do not specifically limit this.

[0185] In a specific implementation, as one embodiment, the processor 601 may include one or more CPUs, for example... Figure 6 CPU0 and CPU1 in the CPU.

[0186] In a specific implementation, as one example, the server may include multiple processors, for example... Figure 6The processors 601 and 607 are described herein. Each of these processors may be a single-core processor or a multi-core processor. The processors herein may include, but are not limited to, at least one of the following: a central processing unit (CPU), a microprocessor, a digital signal processor (DSP), a microcontroller unit (MCU), or an artificial intelligence processor, and other computing devices that run software. Each computing device may include one or more cores for executing software instructions to perform calculations or processing.

[0187] In a specific implementation, as one embodiment, the server may further include an output device 605 and an input device 606. The output device 605 communicates with the processor 601 and can display information in various ways. For example, the output device 605 may be a liquid crystal display (LCD), a light-emitting diode (LED) display device, a cathode ray tube (CRT) display device, or a projector, etc. The input device 606 communicates with the processor 601 and can receive user input in various ways. For example, the input device 606 may be a mouse, keyboard, touchscreen device, or sensing device, etc.

[0188] The server described above can be a general-purpose device or a dedicated device. For example, the server can be a desktop computer, laptop, network server, PDA (personal digital assistant), mobile phone, tablet computer, wireless terminal device, embedded device, the aforementioned terminal device, the aforementioned network device, or something else with... Figure 6 Devices with similar structures. This application does not limit the type of server to specific embodiments.

[0189] The technical solutions provided in this application embodiment can be implemented in an electronic device having the above-described hardware and software architecture, or in a training device having the above-described hardware architecture.

[0190] The following combination Figure 7 As shown, the processing flow of the snapshot method provided in the embodiments of this application is introduced. Figure 7 This is a flowchart illustrating a snapshot method provided in an embodiment of this application. For example, a mobile phone is used as an example. Figure 7 As shown, this snapshot method may include S701-S711:

[0191] S701: The phone responds to the user's shooting operation by displaying a shooting preview interface.

[0192] In this embodiment of the application, the shooting preview interface can be a photo preview interface or a video preview interface.

[0193] In one possible implementation, if the shooting preview interface is a photo preview interface, then the shooting operation can be a user's operation to open the target application. This opening operation can be a touch operation (e.g., a tap) on the phone's interface, or a voice command input by the user to open the target application. This application does not impose specific limitations on the specific implementation of this shooting operation.

[0194] The target application can be any application on the phone that has shooting and snapshot functions, such as a camera app. When the target application is launched, it will default to using the normal photo-taking function and displaying the photo preview interface.

[0195] For example, taking the target application as a camera application, and the operation of opening the target application as a user's touch operation on the phone screen, such as... Figure 8 As shown in (a), the mobile phone can display a desktop 801. This desktop 801 includes a camera application icon 802. When the user needs to open the camera application, they can click on the camera application icon 802 (i.e., open it). The mobile phone can respond to the user's click on the camera application icon 802, launch the camera application, and display as shown in (a). Figure 8 The preview interface 803 shown in (b) can be used as a camera preview interface to display the preview image captured by the camera.

[0196] For example, after a user launches the camera application, the camera application can send a preview request to the camera service in the framework layer. The camera service then sends the preview request to the camera module in the HAL layer, and the camera module sends the preview request to the camera driver in the driver layer to start the camera. After the camera starts, the image sensor in the camera can acquire image light signals and transmit these signals to the image signal processor for preprocessing to obtain a raw image. The raw image is then transmitted back to the HAL layer through the camera driver. The raw image is the image obtained after the photosensitive element in the camera converts the captured light source signal into an electrical signal; it can also be called the original image.

[0197] In some examples, after the camera application is launched, it can continuously generate multiple preview requests and send each preview request to the image-related functional modules in the electronic device in sequence. The image-related functional modules in the electronic device respond to each preview request, capture raw images, generate preview images corresponding to each preview request based on the raw images, and display them in the photo preview interface.

[0198] In another possible implementation, if the shooting preview interface is a video recording preview interface, then the shooting operation can be a user's operation to enable the video recording function of the target application. This enabling operation can be a touch operation (e.g., a click) on the phone interface, or a voice command input by the user to the phone to enable the target application. This application does not impose specific limitations on the specific implementation of this shooting operation.

[0199] For example, taking the target application as a camera application, the operation to enable the recording function of the target application is a user's touch operation on the mobile phone interface, such as... Figure 8 As shown in (b), the preview interface 803 may include a recording option 802. When a user needs to enable the recording function, the user can trigger the recording option, such as by clicking it. In response to the user's triggering action on the recording option 802, refer to... Figure 8 As shown in (c), the mobile phone can display the recording control 805 in the preview interface 803. The user can then click on the recording control 805 (i.e., activate the recording function), causing the phone to start recording video using the camera. During video recording, the phone can display... Figure 8 The video preview interface 806 shown in (d) is a video recording preview interface. This video preview interface can display a preview image corresponding to each video frame recorded by the camera. The preview image can be the image frame itself, or it can be the image frame before image optimization.

[0200] In some embodiments, after the recording function of the target application is activated, multiple recording requests can be generated continuously and sent sequentially to the image-related functional modules in the electronic device. In response to each recording request, the image-related functional modules in the electronic device can continuously acquire raw images and generate preview images and video frames corresponding to each recording request based on the raw images. The preview images are displayed in real-time on the recording preview interface, while the video frames are encoded and stored.

[0201] S702: The mobile phone responds to the user's operation to enable the snapshot function and activates the snapshot function.

[0202] Since the capture method provided in this application is intended to capture exciting photos of exciting moments in dynamic scenes, the capture function needs to be enabled before actually determining the exciting frames to be captured based on the preview image.

[0203] In some embodiments, the snapshot function may specifically include a manual snapshot function and an automatic snapshot function.

[0204] With automatic capture enabled, the phone can automatically determine and save the captured image based on the preview image in the preview interface, without requiring user intervention (such as triggering the camera controls, like clicking). This reduces user interaction and prevents missing exciting moments. With manual capture enabled, the phone continuously identifies frames suitable for capture based on the preview image in the preview interface, but doesn't save them directly. When the user performs a capture operation in the preview interface, the phone saves the latest identified frame as the captured image. In other words, automatic capture generates and saves images without requiring a user-triggered capture command (it's generated by the capture operation itself), while manual capture requires a user-triggered capture command to generate and save the image.

[0205] Based on this, in the embodiments of this application, the user's operation to enable the capture function can be either the user's operation to enable the automatic capture function or the user's operation to enable the manual capture function.

[0206] The following is combined with Figure 9 and Figure 10 Taking the preview interface as an example, this paper provides an exemplary description of how users can enable the automatic capture function and the manual capture function.

[0207] In some embodiments, such as Figure 9 As shown in (a), the photo preview interface 901 may include a manual capture control 91. Initially (e.g., when the camera application is first opened), the manual capture function is disabled by default. At this time, the manual capture control 91 is unselected, and its display format can be a first display format. For example, the first display format can be an icon or pattern without color fill.

[0208] When a user needs to enable the manual capture function, they can trigger the manual capture control 91 in the photo preview interface 901 (this triggering operation is the operation to enable the manual capture function). In response to the user's triggering operation on the manual capture control 91, the phone can enable the manual capture function, and as shown... Figure 9 As shown in (b), the display format of the manual capture control 91 is adjusted to a second display format. For example, the second display format can be an icon or pattern filled with a specific color, such as black.

[0209] The different display formats of the manual capture control 91 when the manual capture function is enabled or disabled allow users to easily and intuitively understand whether the manual capture function is on or off. Of course, in practice, the specific implementation of the first and second display formats can be any other possible implementation method, and this application does not impose any specific restrictions on them.

[0210] When the user needs to disable the manual capture function later, they can click the manual capture control 91 again in the photo preview interface 901 to disable the manual capture function. After the manual capture function is disabled, the display format of the manual capture control 91 will revert to the first display format.

[0211] In other embodiments, such as Figure 9 As shown in (a), the photo preview interface 901 may also include an automatic capture control 92. In the initial state (e.g., when the camera application has just been opened), the automatic capture function is off by default. At this time, the automatic capture control 92 is unselected, and its display format can be the first display format.

[0212] When a user needs to enable the automatic snapshot function, they can trigger the automatic snapshot control 92 on the photo preview interface 901 (this triggering operation is the operation to enable the automatic snapshot function). In response to the user's triggering operation on the automatic snapshot control 92, the phone can enable the automatic snapshot function. And as... Figure 9 As shown in (c), the display format of the automatic capture control 91 is adjusted to the second display format. The different display formats of the automatic capture control 91 when the automatic capture function is on and off allow users to intuitively understand whether the automatic capture function is enabled or disabled.

[0213] When the user needs to turn off the automatic capture function later, they can click the automatic capture control 92 again in the photo preview interface 901 to turn off the automatic capture function. After the automatic capture function is turned off, the display format of the automatic capture control 92 will revert to the first display format.

[0214] This application embodiment does not limit the setting position of the manual capture control 91 and the automatic capture control 92 in the preview interface. The automatic capture control 92 can be set in... Figure 9 The preview interface 901 shown in (a) can be positioned at the lower right side, or it can be set in any other possible position.

[0215] In some other embodiments, the preview interface of the mobile phone may not include the automatic capture control. In this case, the automatic capture function can be enabled through the settings control in the preview interface.

[0216] For example, such as Figure 10 As shown in (a), the photo preview interface 1001 may include settings controls. To enter the capture mode, the user can perform a trigger operation (e.g., a click operation) on the settings control 101 in the preview interface 1001. In response to the user's trigger operation on the settings control 101, the phone can display as shown in [image / description]. Figure 10 The settings interface shown in (b) is as follows. Figure 10As shown in (b), the settings interface includes several controls related to taking photos (such as photo ratio control, smart photo control, and filter photo control) and several controls related to video (such as video resolution control, video frame rate control, high-efficiency video format control, and one-record-multiple-views control). Users can... Figure 10 In the settings interface shown in (b), the smart camera control 102 is triggered.

[0217] In response to the user's triggering operation on the smart camera control 102, the phone can display... Figure 10 The intelligent camera settings interface shown in (c) includes voice-activated camera controls, gesture camera controls, smile capture controls, and automatic capture controls. The phone's automatic capture function is disabled by default. In this case, the automatic capture control 103 can be displayed in a third display format, such as... Figure 10 As shown in (c), the slider in the automatic capture control 103 is on the left, and the right side of the automatic capture control 103 is not filled with color. When the user needs to enable the automatic capture function, they can... Figure 10 In the intelligent photo-taking settings interface shown in (c), a trigger operation (e.g., a click or swipe) is performed on the automatic capture control 103. In response to the user's trigger operation on the automatic capture control 103, such as... Figure 10 As shown in (d), the mobile phone can enable the automatic capture function and change the display format of the automatic capture control 103 to a fourth display format. For example, the slider can be moved to the right, and the left area of ​​the automatic capture control 103 can be set to a predetermined color (e.g., blue).

[0218] The different display formats of the automatic capture control 103 in the smart photo settings interface when the automatic capture function is enabled or disabled allow users to easily and intuitively understand whether the automatic capture function is on or off.

[0219] The actions described above, including the user's gradual opening of the settings interface and the smart camera settings interface, and the combination of triggering the automatic capture control, can be considered as the activation of the automatic capture function. Alternatively, the final triggering of the automatic capture control can also be considered the activation of the automatic capture function.

[0220] Subsequently, if the user needs to disable the automatic capture function, they can re-trigger the automatic capture control 103 (displayed in the fourth display format) in the smart camera settings interface to disable the automatic capture function. After the automatic capture function is disabled, the display format of the automatic capture control 103 will revert to the third display format.

[0221] With the phone displaying a preview interface and the capture function enabled, the phone can begin to determine the best frames that can be captured based on the preview image, i.e., execute the subsequent S703-S710.

[0222] S703: The mobile phone obtains the set of frames to be determined from the preview image displayed in the preview interface.

[0223] With the phone displaying a preview interface and the snapshot function enabled, the phone's camera continuously captures preview images, generating a preview image frame stream, until it receives a snapshot command. The preview interface then displays the preview images from this frame stream sequentially. To promptly identify the best frames from the preview image stream that can be used as snapshot images, the phone continuously processes the latest preview images.

[0224] Furthermore, since the photos needed to capture key moments in dynamic scenes are generally snapshots of an object in motion, relying solely on a preview image may not be accurate enough. For example, in a scene where a dog jumps from the ground, the most exciting moment might be when the dog is at its highest point. However, if the dog is also in the air during its jump and fall, using only a preview image to determine the capture image could result in multiple photos of the dog in the air being mistakenly identified as capture images. Similarly, the visual features of a single frame cannot distinguish between lifting and putting down a child in a parent-child activity, making it impossible to accurately identify the highest point or other exciting moments. In these cases, the captured images are likely to not meet the user's needs.

[0225] Based on this, the mobile phone can determine the capture image based on multiple preview images generated at different but similar times. The combination of these multiple preview images constitutes the set of frames to be determined. In this embodiment, the set of frames to be determined includes a preset number of the latest first preview images. These first preview images are preview images, and the preset number of first preview images includes the latest preview image. For example, the preset number can be 3.

[0226] In the snapshot method provided in this application embodiment, after obtaining the set of frames to be determined each time, a visual feature extraction model is first used to process each first preview image in the set of frames to be determined (specifically, feature extraction) to obtain the visual features corresponding to each first preview image. If the mobile phone uses every preview image as the first preview image in a certain set of frames to be determined, it will cause the visual feature extraction model to have an excessive amount of computation.

[0227] Based on this, to reduce the computational burden on the visual feature extraction model, in some embodiments, the mobile phone can extract preview images from the preview image frame stream (or preview images) according to a preset frame interval to obtain a set of undetermined frames. That is, the frame interval between two adjacent first preview images extracted from the undetermined frame set is the preset frame interval. Here, the frame interval refers to the number of preview image frames between two adjacent first preview images. Furthermore, in practice, the mobile phone can assign consecutive frame numbers to all preview images in the preview image frame stream according to their generation order (or display order), with earlier generation resulting in smaller frame numbers. The frame numbers of two adjacent preview images are consecutive. Based on this, the frame interval can also refer to the difference between the frame numbers of two adjacent first preview images minus 1.

[0228] For example, taking a preset quantity of 3 and a preset frame interval of 2 as an example, refer to... Figure 11 As shown, for the preview image frame stream, the mobile phone can extract a preview image as the first preview image every two frames. When three new first preview images are obtained, a set of frames to be determined can be considered. The strategy of extracting a preview image as the first preview image every two frames to obtain the set of frames to be determined can be called a frame extraction strategy.

[0229] For example, if the phone receives three first preview images with frame numbers t-3, t, and t+3, it can be considered to have obtained a set of undetermined frames, which includes the three first preview images with frame numbers t-3, t, and t+3. Subsequently, if the phone receives a first preview image with frame number t+6, it can be considered to have obtained a new set of undetermined frames, which includes the three first preview images with frame numbers t, t+3, and t+6.

[0230] In this way, by using a specific frame extraction strategy, the mobile phone can avoid processing all preview images in the subsequent steps of the capture algorithm, thereby reducing the computational power consumption and computational burden of the entire technical solution.

[0231] After obtaining the set of undetermined frames, in order to determine the captured image, it is necessary to first acquire the visual features of each first preview image in the set of undetermined frames, and then determine the captured image based on the visual features. However, since the purpose of the technical solution provided in this application embodiment is to obtain photos (i.e., captured images) of exciting moments in a specific dynamic scene, and the main objects in this specific dynamic scene are fixed first objects (or captured objects), such as people, cats, dogs, etc., it is not necessary to perform subsequent processing of the capture algorithm for preview images that do not contain target objects. This also avoids processing preview images that are impossible to be captured images, reducing the waste of computing resources.

[0232] Based on this, after obtaining the set of frames to be determined, the mobile phone can first determine whether the target object exists in all frames in the set, and then decide whether to extract visual features. That is, S703 is followed by S704.

[0233] S704, The mobile phone determines whether the first object exists in all the first preview images in the set of pending frames.

[0234] The first object may include, but is not limited to: dogs, cats, and people.

[0235] If it is determined that the first object exists in all the first preview images in the undetermined frame set, it can be considered that it is possible to determine the captured image from the first preview images. Furthermore, the captured image generally contains only a single main object. Since the first preview images in the undetermined frame set are generated at very close times, under normal circumstances, if the first preview images can be used to determine the captured image, there is a high probability that the same first object exists in the first preview images, and the area occupied by the target object (the area in the image) is the largest among all the first preview images. If the above situation does not exist in the first preview images, it can be considered that there is no unified main target in the first preview images of the undetermined frame set, and therefore it can be considered that the undetermined frame set cannot be used to determine the captured image. Therefore, to avoid wasting computational resources by still processing the first preview images in the undetermined frame set even when there is no unified main target (i.e., the captured object) in the first preview images of the undetermined frame set, if it is determined that the first object exists in all the first preview images of the undetermined frame set, a further judgment can be made as to whether there is a unified captured object in the first preview images, i.e., S705 is executed.

[0236] If it is determined that the first object does not exist in any of the first preview images in the pending frame set, it can be assumed that there is no specific dynamic scene in the scene captured by the user with the mobile phone within the time period corresponding to the pending frame set. Therefore, it is not necessary to determine the capture image based on the first preview images in the current pending frame set. At this time, a new pending frame set can be obtained and the corresponding judgment process can be repeated, that is, S703 can be re-executed.

[0237] In some embodiments, S704 can specifically be referred to as object detection, and the mobile phone can use an object detection algorithm to determine whether the first object exists in all the first preview images in the set of undetermined frames. For example, the object detection algorithm used in this embodiment can be the YOLO algorithm. Of course, any other possible object detection algorithm can also be used.

[0238] Taking the first category, which includes cats, dogs, and people, as an example, refer to... Figure 11As shown, after obtaining the set of frames to be determined, the mobile phone can use a target detection algorithm to sequentially determine whether a cat / dog and a person are present in each first preview image. If it is determined that one or more of the objects—cat, dog, or person—are present in all the first preview images, then a further determination can be made as to whether a unified main target exists in the first preview images, i.e., step S705 is executed. If it is determined that one or more of the objects—cat, dog, or person—are not present in all the first preview images, then it can be determined that there is no target to be captured in the first preview images of the set of frames to be determined, and therefore the subsequent process is not executed.

[0239] S705: The mobile phone determines whether the second object in each first preview image in the undetermined frame set is the same object.

[0240] The second object refers to the first object that occupies the largest area in the first preview image.

[0241] If the second object in each first preview image in the set of undetermined frames is determined to be the same object, it can be assumed that there is a unified capture object in the first preview images in the set of undetermined frames, and thus it can be assumed that these first preview images can be used to determine the capture image.

[0242] Furthermore, to avoid the influence of redundant background content in the first preview image on the subsequent judgment of the captured image, the first preset image can be cropped based on the second object to reduce redundant background content. Since the purpose of capturing is primarily to capture a specific moment of the subject's action, and the determination of the subject's action may require the assistance of other objects in the first preview image. For example, besides a person jumping in the air, at least a basketball is needed to determine if it's an aerial shooting motion. Therefore, assuming that the second object in each first preview image in the undetermined frame set is the same object, the cropping of the first preset image should not only include the second object, but also other content within a certain range of the second object. Different types of second objects perform different ranges of motion, and the corresponding ranges can differ. For example, when the second object is a person, the cropped image area can be a first multiple of the person's area; when the second object is a dog, the cropped image area can be a second multiple of the person's area; the second multiple is different from the first multiple.

[0243] Based on this, if it is determined that the second object in each first preview image in the set of pending frames is the same object, the first preview image can be cropped based on the type of the second object, i.e., S706 is executed.

[0244] If it is determined that the second object in each first preview image of the pending frame set is not the same object, it can be assumed that there is no consistent capture object in the first preview images of the pending frame set. In this case, it can be assumed that there is no consistent capture object in the scene captured by the user with the mobile phone within the time period corresponding to the pending frame set, so there is no need to determine the capture image based on the first preview images in the current pending frame set. At this time, a new pending frame set can be obtained and the corresponding judgment process can be repeated, that is, S703 can be re-executed.

[0245] In some embodiments, when the mobile phone performs S704 and employs an object detection method, a corresponding detection box is generated in the first preview image for each detected first object. For example, refer to... Figure 12 As shown in (a), during the object detection algorithm, the mobile phone generates corresponding rectangular detection boxes around each person (i.e., the first object) in the first preview image, such as detection boxes 1201, 1202, and 1203. The detection box of the first object completely encompasses the first object, and the detection box only completely includes the first object without including any other areas besides the first object. The first object with the largest detection box in the first preview image can be identified as the second object. For example... Figure 12 As shown in (a), the detection box 1201 is the largest detection box, so the first object included in the detection box 1201 can be identified as the second object.

[0246] In some embodiments, S704 may specifically be referred to as multi-target tracking. When S705 is executed, the mobile phone may employ a multi-target tracking algorithm to determine whether the second object is the same object. (See also...) Figure 11 As shown, if it is determined that the first object exists in all first preview images, a multi-object tracking algorithm can be used to track the second object in each first preview image, and then determine whether the same second object exists in all first preview images (i.e., whether the same target is detected in all first preview images). If it is determined that the second object exists, the subsequent cropping step, i.e., S706, is executed. If the second object does not exist, it can be considered that there is no target to capture in the first preview images of the undetermined frame set, and the subsequent process of determining the capture image is not required.

[0247] Furthermore, in practice, for a given dynamic scene, the second object that can be captured may not occupy the largest area in all first preview images due to its movement. However, it is highly likely that the average area occupied by the second object in all first preview images will be the largest among all first objects. Therefore, the second object mentioned in S705 can also refer to the first object that occupies the largest average area in all first preview images.

[0248] S706. The mobile phone crops the first preview image based on the category of the second object to obtain the second preview image.

[0249] The second preview image includes the second object, and the size of the second preview image is a preset multiple of the size of the area occupied by the second object. This size includes both length and width. The area occupied by the second object can be a rectangular region.

[0250] In some embodiments, when the mobile phone performs S704 and employs an object detection method, a corresponding detection box is generated in the first preview image for each detected first object. For example... Figure 12 As shown.

[0251] In this case, refer to Figure 11 As shown, when the phone executes S706, it first determines the detection frame of the second object. Then, it determines the expansion ratio (or preset multiple) based on the category of the second object. For example, if the category of the second object is "human," the expansion ratio can be 2.8 times. After that, the detection frame of the second object can be resized (length and width) according to the expansion ratio.

[0252] Next, the portion of the expanded detection box included in the first preview image is determined as the cropped second preview image. If the expanded detection box extends beyond the first preview image, the portion of the expanded detection box that extends beyond the first preset image is ignored. Alternatively, when expanding the detection box of the second object, if the length or width reaches the edge of the first preview image before being expanded to 2.8 times its original size, the corresponding length or width is no longer expanded.

[0253] For example, the detection box of the second object is... Figure 12 Taking the detection box 1201 shown in (a) as an example, the enlarged detection box 1204 is as follows: Figure 12 As shown in (b). The second preview image obtained after cropping can be viewed as follows. Figure 12 As shown in (c).

[0254] After cropping the first preview image in the frame set to obtain the second preview image, the second preview image can be further input into the visual feature extraction model to obtain the visual features in the second preview image, i.e., S707 is executed.

[0255] In this embodiment of the application, S707 is executed for each first preset image in the set of undetermined frames, so after S707 is executed, a preset number of second preview images can be obtained.

[0256] In some embodiments, the visual feature extraction model has a requirement for the resolution of the input image, which is generally greater than the resolution of the second preset image. Based on this, after cropping to obtain the second preset image, refer to... Figure 11 As shown, the second preset image also needs to be updated using zero-element padding to ensure its resolution matches the input resolution required by the visual feature extraction model. Zero-element padding involves using pixels with a value of 0 to augment the second preset image, increasing its resolution. A pixel value of 0 refers to a pixel whose value is 0 for all color channels (e.g., red R, green G, and blue B channels). For example, with a preset number of 3, the three first preset images in the undetermined frame set are as follows: Figure 13 As shown in (a), the three second preset images, after being cropped and filled with zero elements, can be displayed as follows: Figure 13 As shown in (b).

[0257] Based on the technical solutions corresponding to S704-S706 mentioned above, to avoid problems such as complex backgrounds and small targets interfering with the visual feature extraction model's capabilities, this solution integrates algorithms such as object detection and multi-object tracking, along with data preprocessing strategies (i.e., cropping and zero-element padding). This allows the visual feature extraction model to focus more on the target region containing the captured object, reducing the pressure on the accuracy of the visual feature extraction model's representation from the overall capture scheme.

[0258] S707 The mobile phone inputs the second preview image into the visual feature extraction model to obtain the visual features of the second preset image.

[0259] The visual features of the second preset image are the same as the visual features of the first preset image corresponding to the second preset image. In this embodiment, the visual feature extraction model can also be called a visual encoder.

[0260] In this embodiment, the visual feature extraction model can be obtained using model distillation techniques based on a visual language model. This visual language model is a multimodal large model with a significantly larger number of parameters and a longer latency than the visual feature extraction model. The difference between the number of parameters in the visual language model and the number of parameters in the visual feature extraction model is greater than a first threshold.

[0261] This visual language model has the ability to simultaneously process image and text information, obtaining output information corresponding to both image and text information. In this embodiment, the visual language model needs to be trained using sample pairs consisting of exciting images of exciting moments in a preset dynamic scene and the text features of the corresponding exciting text (e.g., an image of a dog jumping in the air and the text "dog jumping in the air"). This allows the visual language model, after obtaining a visual feature extraction model through model distillation, to better extract visual features that are compatible with the exciting images and exciting text. The training method of the visual language model can be any possible training method, and this application does not impose any specific limitations on it. For example, the visual language model can be a VAST model or any other possible neural network model.

[0262] In some embodiments, model distillation of the visual language model to obtain the visual feature extraction model can specifically involve using multiple sample images positively or negatively correlated with the preview image as training data, simultaneously inputting them into both the visual language model and the initial visual feature extraction model. This yields first sample visual features obtained by the visual language model processing the sample images, and second sample visual features obtained by the initial visual feature extraction model processing the sample images. Subsequently, the first sample visual features can be used as supervision information, and the difference between the first and second sample visual features can be used as the loss value to iteratively update the initial visual feature extraction model, resulting in the final visual feature extraction model. The final visual feature extraction model's image-text space will be approximately consistent with the visual language model's image-text space. Consequently, its ability to extract visual features from images will be approximately the same as the visual language model's ability to extract image features. It will also possess a multimodal data understanding capability similar to the visual language model, enabling the visual feature extraction model to better extract visual features from preview images.

[0263] In some embodiments, the visual feature extraction model may include any possible neural network model, such as a lightweight convolutional neural network (CNN) model (e.g., MobileNetV2) or a ViT (vision transformer) model with a transformer architecture, depending on the actual needs. This application does not impose any specific restrictions on this.

[0264] After obtaining the visual features corresponding to the first preset image, the mobile phone can further determine a first "highlight score" based on the visual features, which is used to characterize the similarity between the target preview image, the featured image, and the corresponding featured text, and a second "highlight score" based on the visual features, which is used to characterize the matching degree between the target preview image and the featured text. The level of detail of the visual features required for the calculation of the first "highlight score" and the second "highlight score" is different.

[0265] Therefore, in some embodiments, the visual features obtained by the visual feature extraction model may include a first sub-visual feature and a second sub-visual feature. The first and second sub-visual features have different dimensions, or in other words, the first and second sub-visual features have different numbers of features. The first sub-visual feature has a smaller dimension than the second sub-visual feature.

[0266] For example, taking a visual feature extraction model that uses a CNN as the backbone network for feature extraction as an example, refer to... Figure 14 As shown, after receiving a frame of image (specifically, a second preview image), the CNN first extracts a preliminary visual feature. For example, the dimension of this visual feature can be 1280, meaning that the visual feature can have 1280 features, and this visual feature can be represented by 1×1280.

[0267] Subsequently, the total visual features can be mapped to first sub-visual features through a first feature projection layer, and the total visual features can be mapped to second sub-visual features through a second feature projection layer. For example, the dimension of the first sub-visual feature can be 512, meaning it can contain 512 features, and can be represented by 1×512. The dimension of the second sub-visual feature can be 768, meaning it can contain 768 features, and can be represented by 1×768. The first sub-visual feature can be used to characterize coarse-grained main features in the second preset image, while the second sub-visual feature can be used to characterize finer-grained detail features in the second preset image.

[0268] In some embodiments, refer to Figure 14 As shown, the first feature projection layer can be a linear layer, and the second feature projection layer can include a linear layer and a normalization layer. Of course, the normalization layer doesn't have to be a separate layer; it can be a normalization operation performed on the output of the linear layer. In practice, the first and second feature projection layers may or may not belong to a visual feature extraction model; this application does not impose specific restrictions on this.

[0269] After performing the above processing on each second preset image, the first sub-visual feature and the second sub-visual feature of each second preset image can be obtained.

[0270] In this way, the mobile phone obtains first and second sub-visual features with different dimensions through the visual feature extraction model, which can then be used to calculate the first and second highlights scores respectively, thereby enabling the subsequent determination of the captured image based on the first and second highlights scores.

[0271] After obtaining the visual features corresponding to the first preset image, the mobile phone can determine the first and second highlights scores for the captured image based on these visual features. If the first and second highlights scores are calculated directly using the visual features corresponding to a preset number of first preset images, the computational load would be large, and the accuracy would be reduced. Therefore, to reduce the computational load and facilitate accurate calculation of the first and second highlights scores, all first preview images can be considered comprehensively. The visual features corresponding to all first preview images in the undetermined frame set (i.e., the visual features of the second preview image after cropping from all first preview images) can be fused to obtain the visual fusion features. This is followed by step S708 after step S707.

[0272] S708: The mobile phone fuses the visual features of a preset number of second preview images to obtain visual fusion features.

[0273] In this embodiment of the application, when fusing the visual features of a preset number of second preview images, any possible fusion method can be adopted, and this application does not impose any specific restrictions on it.

[0274] In some embodiments, the visual features include a first sub-visual feature and a second sub-visual feature, and the first and second sub-visual features have different uses, one for calculating a first highlight score and the other for calculating a second highlight score. In this case, the visual fusion features may include the first sub-visual fusion feature and the second sub-visual fusion feature.

[0275] Based on this, combined Figure 7 , refer to Figure 15 As shown, S708 may specifically include S1501 and S1502:

[0276] S1501. The mobile phone determines the average value of the first sub-visual features of a preset number of second preview images as the first sub-visual fusion feature.

[0277] Specifically, the mobile phone can use the mean operation to obtain the average value of the first sub-visual features of a preset number of second preview images. The average value of the first sub-visual features of the preset number of second preview images is obtained by averaging the first sub-visual features element-wise. For example, with a preset number of 3, and the first sub-visual features of the preset number of second preview images being (0,2,2), (2,4,6), and (4,6,7), the final first sub-visual fusion feature is (2,4,5). In the first sub-visual fusion feature, 2 is (0+2+4) / 3, 4 is (2+4+6) / 3, and 5 is (2+6+7) / 3.

[0278] In addition, to facilitate calculation, after using the mean operation, the average value of the first sub-visual features of a preset number of second preview images can be normalized to obtain the first sub-visual fusion features that are easier to calculate.

[0279] For example, taking a preset quantity of 3 and a first sub-visual feature dimension of 512 as an example, refer to... Figure 14 As shown, after the visual feature extraction model obtains the first sub-visual features (identified by 3×512) of all the second preview images, the first sub-visual fusion feature can be obtained through mean operation and normalization layer (or normalization operation). For example, if the dimension of the first sub-visual feature is 512, then the dimension of the first sub-visual fusion feature is also 512, which can also be represented by 1×512.

[0280] S1502. The mobile phone uses a preset feature fusion model to fuse the second sub-visual features of a preset number of second preview images to obtain the second sub-visual fusion features.

[0281] Since the second brilliant score, calculated based on the second sub-visual fusion features, is primarily used to determine whether the first preview image in the undetermined frame set is an image within the preset dynamic scene, the fusion of the second sub-visual features needs to more fully consider the differences between the second sub-visual features of different second preview images. Therefore, a preset feature fusion model with better fusion performance can be used to fuse the second sub-visual features of a preset number of second preview images to obtain second sub-visual fusion features that better represent the differences among all second sub-visual features.

[0282] In this embodiment, the preset feature fusion model can be any lightweight neural network model with feature fusion capabilities, such as a multilayer perceptron (MLP) or a more lightweight neural network consisting of several sequentially connected convolutional layers, or a more complex transformer architecture model. The specific neural network architecture chosen depends on actual requirements (power consumption, latency, and performance, etc.), and this application does not impose specific restrictions on it.

[0283] For example, taking a preset quantity of 3 and a dimension of 768 for the second sub-visual feature as an example, refer to... Figure 14 As shown, after the visual feature extraction model obtains the second sub-visual features (identified by 3×768) of all the second preview images, the second sub-visual fusion feature can be obtained through an MLP. The MLP can include linear layers and normalization layers (or normalization operations on the output of the linear layers). For example, if the dimension of the first sub-visual feature is 768, then the dimension of the second sub-visual fusion feature is also 768, or it can be represented as 1×768. In some embodiments, refer to... Figure 14As shown, the process of obtaining the first sub-visual fusion feature and the second sub-visual fusion feature by mobile phone fusion can be called temporal information fusion.

[0284] In this embodiment, S1501 and S1502 can be executed simultaneously or sequentially, depending on the actual needs. This application does not impose any specific restrictions on this.

[0285] Based on the above technical solution, the mobile phone can use a suitable fusion method to obtain the fused first sub-visual features and second sub-visual features corresponding to the first preview image (i.e., the first sub-visual features and second sub-visual features of the second preview image), which provides data support for the subsequent calculation of the first and second highlights scores, and enables the entire capture method to be implemented smoothly.

[0286] In some embodiments, the images capturing key moments in certain preset dynamic scenes are highly dependent on spatiotemporal information. For example, in a scene with a jumping dog, the temporal sequence of multiple images in that scene needs to be considered to determine the key moment when the dog is at its highest point, thus identifying the key image. Therefore, to better determine whether the first preview image in the set of undetermined frames is an image from the preset dynamic scene, the temporal changes of the main object in the preset dynamic scene also need to be considered. Based on this, combined with... Figure 15 , refer to Figure 16 As shown, before S1502 is executed, the method further includes S1601:

[0287] S1601. The mobile phone fuses the preset timing features with the second sub-visual features of the second preview image to update the second sub-visual features of the second preview image.

[0288] In this embodiment, the preset temporal feature can be obtained by inputting a video containing exciting images of all possible preset dynamic scenes into a visual language model, or it can be obtained when the visual language model is trained to convergence. This preset temporal feature can be used to characterize the temporal variation characteristics of image content in videos containing exciting images of all possible preset dynamic scenes. In some embodiments, the preset temporal feature can be a temporal information vector (frame embedding) generated by the visual language model. The preset temporal feature and the second sub-visual feature can have the same dimension.

[0289] In one possible implementation, the mobile phone can use an add operation to fuse the preset temporal features with the second sub-visual features of the second preview image (that is, to add the preset temporal features and the second sub-visual features element by element) in order to update the second sub-visual features of the second preview image.

[0290] S1601 is implemented for each second preview image, so after S1601 is executed, the updated second sub-visual features of the second preview images obtained by cropping each first preview image in the undetermined frame set can be obtained. Then, the mobile phone can use the updated second sub-visual features to execute S1502.

[0291] For example, refer to Figure 14 As shown, each time the second feature projection layer generates a second sub-visual feature, it uses an add operation to fuse the temporal features into the second sub-visual feature. Subsequently, the second sub-visual fused feature can be obtained based on the second sub-visual feature that has been fused with the temporal features.

[0292] In this way, by incorporating preset temporal features into the second sub-visual features, the second sub-visual features can reflect the temporal changes in the actions of the main object in the preset dynamic scene. Consequently, the second sub-visual fusion feature at the point of fusion can also better reflect this. Based on this, the second highlight score calculated based on the second sub-visual fusion feature can better determine the captured image, allowing the phone to obtain more accurate captured images and improving the user experience.

[0293] After S708 is executed, and the visual fusion features are obtained, the mobile phone can further determine the first and second highlight scores based on the visual fusion features, so as to further determine the captured image based on the first and second highlight scores. That is, S709 is executed after S708.

[0294] S709: The mobile phone determines the first and second highlights scores of the target preview image based on visual fusion features and preset highlights features.

[0295] The preset highlighted text features are the text features of highlighted images corresponding to highlighted moments in a preset dynamic scene. The target preview image is the first preview image generated in the target order within the set of pending frames. For example, if the preset quantity is 3, the target order can be the second order, meaning the target preview image is the first preview image generated in the second order within the set of pending frames.

[0296] Since there are multiple preset dynamic scenes, there can be multiple preset exciting text features in this embodiment. Each pair of preset exciting text features corresponds to one preset dynamic scene.

[0297] In some embodiments, visual features include a first sub-visual feature and a second sub-visual feature, and visual fusion features include a first sub-visual fusion feature and a second sub-visual fusion feature. Preset compelling text features may include a first sub-compassionate text feature with the same dimension as the first sub-visual fusion feature and a second compelling text feature with the same dimension as the second sub-visual fusion feature.

[0298] Specifically, the first sub-visual fusion feature can be used to determine the first highlight score, and the second sub-visual fusion feature can be used to determine the second highlight score. Based on this, combined with... Figure 15 , refer to Figure 17 As shown, S709 may specifically include S1701 and S1702:

[0299] S1701: The mobile phone calculates the first brilliant score based on the first sub-visual fusion feature and the first sub-brilliant text feature.

[0300] The first highlight score is used to characterize the similarity between the target preview image and the highlight text corresponding to the first sub-highlight text feature (or preset highlight text feature).

[0301] In this embodiment, the first highlight score can be ITC (specifically, an ITC value or an ITC score). When the first highlight score is ITC, the mobile phone can use any possible ITC loss function to calculate the ITC in S1701, and this application does not impose any specific restrictions on this.

[0302] In some embodiments, the first highlight score can be either ITC or a weighted average of IIC and ITC. Therefore, S1701 specifically includes: the mobile phone calculating ITC based on the first sub-visual fusion feature and the first sub-highlight text feature, and calculating IIC based on the first sub-visual fusion feature and preset highlight image features; the mobile phone determining the weighted average of IIC and IIC as the first highlight score. The preset highlight image feature is the visual feature of the highlight image corresponding to the highlight text to which the first sub-highlight text feature (or preset highlight text feature) belongs. The preset highlight image is an image of a highlight moment in a preset dynamic scene. The preset highlight visual feature and preset highlight text feature can be obtained by inputting the image of the highlight moment in the preset dynamic scene and the corresponding highlight text into a visual language model.

[0303] The mobile phone can use any possible IIC loss function to calculate the IIC, and this application does not impose any specific restrictions on this. The IIC and the corresponding weights can be determined according to the actual situation, and this application does not impose any specific restrictions on this.

[0304] In this way, the first highlight score determined by the mobile phone through IIC and ITC can not only characterize the similarity between the target preview image and the highlighted text corresponding to the first sub-highlight text feature, but also the similarity between the target preview image and the highlighted image corresponding to the preset highlighted visual feature. Furthermore, it can characterize the correlation between the target preview image and the highlighted moments in the preset dynamic scene. Therefore, determining the capture image based on this first highlight score can be more accurate.

[0305] S1702. The mobile phone uses a multimodal fusion model to process the second sub-visual fusion features and the second sub-highlighted text features to obtain the second highlight score.

[0306] The second highlight score characterizes the degree of matching between the target preview image and the highlighted text corresponding to the second highlight text feature. The multimodal fusion model can be a pre-trained model capable of obtaining the second highlight score based on visual features and the second highlight text feature. For example, the second highlight score can be an ITM (Intense Phenomenon Measure).

[0307] In the embodiments of this application, S1702 and S1701 can be executed simultaneously or sequentially, and this application does not impose specific restrictions on this.

[0308] To obtain a suitable multimodal fusion model, embodiments of this application also provide a training method for the multimodal fusion model. This method can be applied in a training device and may include the following steps:

[0309] S1. The training device acquires sample data and the corresponding label information.

[0310] The sample data pair may include multiple first sample data, multiple second sample data, and second sub-features of the text. The first sample data pair includes the visual fusion features of the sample frame set, and the second sample data includes the visual features of the sample images.

[0311] The sample frame set includes a predetermined number of sample video frames. These sample video frames are video frames from any scene. The predetermined number of sample video frames are generated in an arithmetic sequence, and the common difference is the same as the predetermined frame interval. The scene corresponding to the first sample data includes a predetermined dynamic scene, and the video frames included in the multiple first sample data include highlights from the predetermined dynamic scene. The visual fusion feature of the sample frame set is the fusion feature of the visual features of the predetermined number of sample video frames.

[0312] The sample images are images from any scene. The scene corresponding to the second sample data includes a preset dynamic scene, and the sample images included in the multiple second sample data include wonderful pictures of wonderful moments in the preset dynamic scene.

[0313] For example, refer to Figure 18 As shown, with a preset quantity of 3, the visual features of three sample video frames in the sample frame set are represented by 3×768. The sample visual fusion features of the sample frame set can then be achieved using an MLP. Specifically, this MLP can be the preset feature fusion model described in the aforementioned embodiment. The final sample visual fusion features of the sample frame set can then be represented by 1×768. Similarly, sample visual features can also be represented by 1×768. In this embodiment, S1 can be referred to as the temporal fusion part of the training process.

[0314] The label information corresponding to the first sample data includes whether the sample frame set matches or does not match the second sub-feature of the featured text. If the sample frame set includes featured images corresponding to the featured text to which the second sub-feature of the featured text belongs, then the label information of the first sample data is a match; if the sample frame set does not include featured images corresponding to the featured text to which the second sub-feature of the featured text belongs, then the label information of the first sample data is a mismatch.

[0315] The label information corresponding to the second sample data includes whether the sample image matches or does not match the second sub-feature of the featured text. If the sample image is a featured image corresponding to the featured text to which the second sub-feature of the featured text belongs, then the label information of the second sample data is a match; if the sample image is not a featured image corresponding to the featured text to which the second sub-feature of the featured text belongs, then the label information of the second sample data is a mismatch.

[0316] In this embodiment, the purpose of adding second sample data to the sample data is to improve the convergence speed of model training. If only the first sample data is included during training, the visual features input to the multimodal fusion model are fused features of the visual features of multiple sample video frames in the first sample data. Since judging fused features is more complex than judging the visual features of a single image, the training efficiency will be insufficient. Furthermore, obtaining a set of sample frames is more difficult than obtaining sample images. If the sample data only includes the first sample data besides the second sub-feature text features, obtaining the sample data will be more challenging.

[0317] To better obtain a multimodal fusion model capable of determining the matching degree between multiple images and the featured text corresponding to the second sub-feature, the aforementioned sample frame set and sample images should include hard negative examples corresponding to the featured text. For example, if the featured text is a dog jumping, then the sample frame set and sample images could include images of a dog running. This allows the finally trained multimodal fusion model to accurately determine the low matching degree between the target preview image corresponding to the second sub-visual fusion feature and the featured text corresponding to the second sub-feature.

[0318] Furthermore, since the second sub-visual fusion feature processed by the multimodal fusion model is obtained by fusing multiple second sub-visual features with preset temporal features, in order to better train a multimodal fusion model that meets the processing requirements, refer to... Figure 18 As shown, before generating the visual fusion features of the sample frame set, the visual features of each sample video frame in the sample frame set need to be fused and updated with the preset temporal features using the add operation; similarly, the sample data features of the sample image also need to be fused with the preset temporal features using the add operation.

[0319] S2. The training device uses sample data as training data and the label information corresponding to the sample data as supervision information. Iterative training is used to obtain an initial multimodal fusion model, and then a multimodal fusion model is obtained.

[0320] The initial multimodal fusion model can be any feasible lightweight neural network model. For example, refer to... Figure 18 As shown, the initial multimodal fusion model may include a linear layer 1, a normalization layer, an activation layer (GELU, Gaussian error linear units), and a linear layer 2. The output of this initial or final multimodal fusion model is binary data (or matching data). Specifically, this binary data may include the probability of a match (or semantic similarity) between the sample data and the corresponding label information, and the probability of a mismatch (or semantic opposition). The probability of a match can be represented by the ITM (Intense Second Brilliance Score).

[0321] In some embodiments, S2 may specifically include the following steps:

[0322] S21. Initialize the training equipment to create the initial multimodal fusion model.

[0323] The initialization of the initial multimodal fusion model can specifically involve initializing the weight parameters and bias parameters of the initial multimodal fusion model using any feasible initialization method. Four commonly used initialization methods are Gaussian initialization, Xavier initialization, MSRA initialization, and He initialization. Generally, the bias parameters are initialized to 0, and the weight parameters are randomly initialized. The specific initialization process will not be described in detail in this application.

[0324] S22. The training device inputs the sample data into the initial multimodal fusion model to obtain the predicted matching data.

[0325] The predicted matching data can be the binary data mentioned in the previous embodiments.

[0326] S23. The training device determines the matching loss value based on the label information of the predicted matching data and the sample data.

[0327] The matching loss can be obtained based on any feasible loss function, and this application does not impose any specific restrictions on it.

[0328] S24. The training device iteratively updates the initial multimodal fusion model based on the matching loss value to obtain the multimodal fusion model.

[0329] Specifically, step S24 can involve tuning the initial multimodal fusion model based on the matching loss value (adjusting weight parameters and bias parameters, etc.). After each tuning, steps S22-S24 are repeated until the matching loss value is less than a preset matching loss value. At this point, the current initial multimodal fusion model is determined as the final multimodal fusion model. The sample data input to the initial multimodal fusion model can be different each time steps S22-S24 are repeated. The preset matching loss value can be derived empirically; it can be assumed that if the matching loss value is less than the preset matching loss value, the error of the multimodal fusion model based on the matching data obtained from the visual features and the second sub-text features is within an acceptable range.

[0330] Based on the technical solutions corresponding to SS1 and S2 mentioned above, a multimodal fusion model can be trained through supervised learning. This multimodal fusion model has the ability to obtain the matching degree between visual features and second-sub-highlighted text features (i.e., the matching degree between the image corresponding to the visual feature and the highlighted text corresponding to the second-sub-highlighted text feature) using visual features and second-sub-highlighted text features. This provides data support for subsequently determining the second-highlighted text score. Furthermore, this training method uses two types of training data—sample frame sets and sample images—to train the multimodal fusion model, reducing the difficulty of obtaining sample data and effectively improving the model convergence efficiency.

[0331] Based on the technical solutions corresponding to S1702 and S1701, the mobile phone can obtain the first and second highlights scores in an appropriate manner, providing data support for the subsequent determination of the captured image.

[0332] After S709 is executed, the phone can obtain the first and second highlights scores, and then determine the captured image based on the first and second highlights scores, that is, execute S710.

[0333] The S710 mobile phone determines the highlight frames based on the first and second highlight scores.

[0334] In some embodiments, the first highlight score can be a weighted average of IIC and ITC, and the second highlight score can be ITM. In this case, refer to Figure 19 As shown, S710 may specifically include S1901-S1902:

[0335] S1901, The mobile phone determines whether the ITM is greater than the probability threshold.

[0336] Since ITM can accurately determine the degree of matching between the target preview image and the preset featured text features, it can determine whether the target preview image is an image within the preset dynamic scene before or after a featured moment (i.e., the featured interval). Based on this, the mobile phone can first determine whether the ITM is greater than a probability threshold. For example, the probability threshold can be 50%.

[0337] If ITM is greater than the probability threshold, the target preview image can be considered to be in the excellent range, i.e., S1902 is executed.

[0338] If the ITM is less than or equal to the probability threshold, the target preview image can be considered not to be in the exciting range, and the subsequent processes of the capture method will not be performed.

[0339] S1902, The mobile phone confirms that the target preview image is in the excellent range.

[0340] Furthermore, to ensure a low time interval between the final determined highlight frame and the current moment or the latest preview image, the latency between the captured image and the current moment will also be low when the corresponding capture command is received and the highlight frame is determined as the captured image. Therefore, in this embodiment, a maximum latency frame number can be set, that is, the number of frames between the highlight frame and the latest preview image. The mobile phone can then determine the highlight frame based on the maximum latency frame number. This leads to the execution of steps S1903-S1905.

[0341] In this embodiment, since there are multiple preset dynamic scenes, there are also multiple pairs of preset exciting image features and preset exciting text features. Accordingly, in the aforementioned embodiments, the first exciting score and the second exciting score determined in S709 (or S1701 and S1702) correspond to multiple preset dynamic scenes, that is, multiple first exciting scores and multiple second exciting scores are obtained, each first exciting score corresponds to one preset dynamic scene, and each second exciting score corresponds to one preset dynamic scene.

[0342] Based on this, when executing step 1901, the phone can determine whether the largest target ITM among multiple ITMs is greater than a probability threshold. If it is determined that the target ITM is greater than the probability threshold, then the target preview image is determined to be within the exciting range of the preset dynamic scene corresponding to the target ITM. Subsequently, when using the first exciting score, the first exciting score of the preset dynamic scene corresponding to the target ITM is also adopted.

[0343] Of course, in practice, the determination of which preset dynamic scene corresponds to the first and second highlights can be achieved in any other possible way, as long as the final determined preset dynamic scene is the preset dynamic scene that the target preview image is most likely to match, as reflected by its corresponding first and second highlights.

[0344] S1903. The mobile phone determines the highlight window based on the maximum latency frame rate and obtains the first highlight score of the first preview image in the highlight window.

[0345] The "Highlights Window" refers to a combination of several consecutive preview images with the maximum latency. "Consecutive" here means continuously displayed in the preview interface when the capture function is enabled; that is, the Highlights Window includes several consecutive preview images with the maximum latency displayed in the preview interface when the capture function is enabled. The latest preview image in the Highlights Window is the newest first preview image in the set of pending frames. The maximum latency frame count is greater than a preset number.

[0346] Since the method provided in this application embodiment is continuously executed by the mobile phone after the capture function is enabled, a corresponding first wonderful score will be determined after each preset number of first preset images are obtained. Therefore, it is convenient to obtain the first wonderful score of each preview image in the wonderful window.

[0347] It should be noted that, since this application identifies the first and second highlights scores, derived from the visual features of multiple first preview images, as the target preview images within those multiple first preview images, after the phone activates its capture function, the first preview images preceding the target preview images in the initial set of undetermined frames do not possess first and second highlights scores; these scores can be set to 0 or empty. The first preview images following the target preview images in the initial set of undetermined frames will then be used as target preview images, and attempts will be made to obtain the first and second highlights scores. If, due to the absence of a first object or a uniform capture object in some undetermined frames, subsequent calculations of the first and second highlights scores are not performed, the corresponding target preview images' first and second highlights scores will also be set to 0 or empty.

[0348] After obtaining the first "wonder" score of the first preview image in the "wonder" window, the higher the score, the more similar the first preview image is to the "wonder" image, and therefore the more likely it is to be a "wonder" frame suitable for capture. Based on this, the phone determines whether a first preview image with a maximum (local maximum) score exists in the "wonder" window. If it does, the first preview image with the maximum score is identified as the "wonder" frame. If it does not exist, the entire capture method is re-executed. That is, S1904 and S1905 are executed after S1903.

[0349] S1904. The mobile phone determines whether the first target preview image has the highest score among the first preview images in the history of the first preview image in the wonderful window.

[0350] Among them, the first historical preview image is the first preview image generated in the highlight window before the target preview images that have obtained the first and second highlight scores.

[0351] For example, taking a preset quantity of 3, a maximum latency of 6 frames, and the target preview image being the second-ranked preview image in the pending frame set, the "Highlights" window can include six first preview images, including all the first preview images in the pending frame set. The first preview images in the "Highlights" window, arranged in the order of generation, can be image 1, image 2, image 3, image 4, image 5, and image 6. Here, images 4, 5, and 6 are three first preview images from the pending frame set, and image 5 is the current target preview image. If image 5 has the highest first "Highlights" score among images 1, 2, 3, 4, and 5, it indicates that there may be subsequent first preview images with even higher first "Highlights" scores. In other words, there is no first preview image with a maximum first "Highlights" score in the historical first preview images of the "Highlights" window, and the entire capture method needs to be re-executed.

[0352] If, among images 1, 2, 3, 4, and 5, image 4 has the highest first highlight score, it indicates that image 4, the first target preview image, has the highest first highlight score among the historical first preview images in the highlight window. In other words, image 4 is the first preview image in the highlight window with the maximum first highlight score. In this case, image 4 is determined as the highlight frame. That is, S1905 is executed.

[0353] S1905, the mobile phone identifies the first target preview image as the best frame.

[0354] In the embodiments of this application, the strategy for determining the exciting frames in the schemes corresponding to S1903-S1905 can be a low-latency frame output strategy.

[0355] Once S1905 is completed, the highlight frame is determined. If the highlight frame is determined as the captured image, the capture delay is the frame interval between the first target preview image and the latest preview image in the undetermined frame set.

[0356] For example, taking a probability threshold of 50% as an example, the effect of this low-latency frame output strategy is illustrated in the diagram below. Figure 20 As shown, within multiple frame ranges after the capture function is enabled, there are preview images with ITM values ​​greater than the probability threshold. At this point, it can be determined that these preview images are within the "highlight" range. Furthermore, preview images within the "highlight" range where IIC and ITC (i.e., based on the first highlight score) are local maxima can be captured as "highlight frames".

[0357] Based on the above technical solution, the mobile phone can determine the best frames that meet the latency requirements and can be used as captured images, using a pre-set maximum latency frame count. Furthermore, because this method for determining best frames combines a first best frame score (determined by ITC or IIC and ITC) and a second best frame score (i.e., ITM), and the second best frame score is less affected by the background in the image, the final determined best frames are more accurate.

[0358] S711: The mobile phone responds to the capture command and determines the best frame corresponding to the capture command as the captured image.

[0359] In this embodiment, the capture command can be triggered by the user or automatically by the mobile phone. After the captured image is determined, the mobile phone can store the captured image in the gallery.

[0360] In some examples, when the manual capture function is enabled, users can trigger a capture command by inputting a photo-taking action (such as triggering the camera control) in the preview interface. In this case, the phone can determine the most recently identified frame before the capture command is triggered as the captured image.

[0361] In other examples, when the automatic snapshot function is enabled, the phone can generate a snapshot command when a highlight frame is identified. The highlight frame corresponding to this snapshot command will then be the one identified at the time the snapshot command was issued. In response to this snapshot command, the phone will identify that highlight frame as the captured image.

[0362] Based on S709-S711, the mobile phone can promptly determine the captured image based on the first and second highlights scores.

[0363] Based on the technical solution provided in this application, visual features of multiple preview images can be extracted first, and then combined with preset brilliant text features to obtain a first brilliant score and a second brilliant score for the target preview image. Subsequently, brilliant frames that can be used as capture images can be determined based on the first and second brilliant scores. Then, upon receiving a capture command, the corresponding brilliant frames can be output as capture images. In this technical solution, the visual feature extraction model for extracting visual features of the preview image is obtained through model distillation of a large visual language model with multimodal understanding capabilities; the text and image spaces of the two are consistent. Through model distillation of the large visual language model, the representational ability of the visual feature extraction model can be improved, enhancing its understanding of multimodal data. This allows the visual feature extraction model to better extract visual features from the preview image, and ensures that the visual features extracted by the model can smoothly perform similarity and matching calculations with the preset brilliant visual features and preset brilliant text features generated by the large visual language model. In this way, the first and second "highlight scores," derived from the visual features extracted by the visual feature extraction model and the preset "highlight text" features generated by the visual language big data model, can better characterize the correlation between the target preview image and the "highlight images" and "highlight text." Therefore, by referencing both visual and textual features—two multimodal information sources—the captured image within the target preview image can be accurately identified. Furthermore, because the visual feature extraction model is small in size, has low latency, and low power consumption, this solution can capture "highlight moments" in preset dynamic scenes in real time. Compared to existing technologies, it can improve capture accuracy while reducing latency and power consumption, thus enhancing the user experience.

[0364] Furthermore, because this technical solution employs a lightweight visual feature extraction model, it can process multiple frames of images in real time to obtain the visual features of those frames, and then determine the capture image based on these visual features. This effectively avoids false detection problems (i.e., obtaining inaccurate capture images) caused by unclear or unstable visual features in a single frame.

[0365] In some embodiments, the preset excellent text features (first sub-excellent text features and second sub-excellent text features) in the technical solution provided in this application remain unchanged after being extracted from the visual language model. However, to make the text-image spaces of the visual feature extraction model and the visual language model more similar, the visual feature extraction model can fully inherit the text representation learning module of the teacher model, such as the BERT (bidirectional encoder representations from transformers) structure; that is, the visual feature extraction model can include this text representation learning module. Then, during model distillation, the final visual feature extraction model can be obtained by distilling and training with text-image pairs (i.e., data pairs consisting of video / images and corresponding text, refer to the sample data in the aforementioned embodiments). In this case, the obtained visual feature extraction model can also have the ability to extract text features, and its ability is similar to that of the visual language model. Therefore, the preset excellent text features can also be generated by the visual feature extraction model, and the generation method is similar to that of the large visual language model.

[0366] It should be noted that the capture method provided in this application supports online real-time capture. However, by expanding the cache window and other methods, this method can be easily applied to scenarios such as video motion detection and video highlight frame retrieval. In these scenarios, a complete video can be input at once, and the mobile phone can obtain all the images at once, and highlight frames can be identified from all video frames. In this case, the highlight interval is still determined using ITM, and the highlight frame is determined by the video frame with the highest or highest highlight score among all frames. The specific implementation details can be found in the aforementioned embodiments and will not be repeated here.

[0367] For ease of understanding, the following example illustrates visual features including a first sub-visual feature and a second sub-visual feature, visual fusion features including a first sub-visual fusion feature and a second sub-visual fusion feature, and preset excellent text features that can include a first sub-excellent text feature with the same dimension as the first sub-visual fusion feature and a second sub-excellent text feature with the same dimension as the second sub-visual fusion feature. The first sub-visual feature has a dimension of 512, the second sub-visual feature has a dimension of 768, the preset feature fusion model is MLP, the preset number is 3, the visual feature extraction model is a visual encoder, and the first excellent score is generated by IIC and ITC, as an example. Figure 21 The process for determining the best frames in the technical solution provided in this application is explained as follows:

[0368] First, the mobile phone can acquire a set of frames to be determined. The specific acquisition process can be referred to the relevant description in S703 of the aforementioned embodiment, and will not be repeated here.

[0369] Subsequently, the mobile phone can preprocess the set of frames to be obtained to obtain three second preview images. Specifically, the preprocessing may include detecting a first object, multi-object tracking, and cropping based on the detection bounding box region. The specific implementation of the preprocessing can be referred to the relevant descriptions in S704-S706 of the aforementioned embodiments, and will not be repeated here.

[0370] Next, the mobile phone can use a visual encoder and a first feature projection layer to extract visual features from the second preview image, obtaining a first sub-visual feature (1, 512) and a second sub-visual feature (1, 768). The first feature projection layer can include fully connected linear layers, while the normalization layer within it is not shown. The specific implementation of this part can be referred to the relevant description in S707 of the aforementioned embodiment, and will not be repeated here.

[0371] Then, the phone can perform the following two tasks in parallel:

[0372] Firstly, the mobile phone uses the mean operation to fuse the first sub-visual features of the three second preview images to obtain the first sub-visual fusion feature (1,512); then, based on the first sub-visual fusion feature, the first sub-highlight text feature (m,1,512) and the preset highlight visual feature (m,1,512), the first highlight score of the target preview image is obtained.

[0373] Where m is the number of preset dynamic scenes. The number of preset dynamic scenes determines the number of pairs of first sub-text features and preset visual features.

[0374] The specific implementation of the first aspect can be referred to the relevant descriptions of S1501 and S1701 in the foregoing embodiments, which will not be repeated here.

[0375] Secondly, the mobile phone can first fuse and update the second sub-visual features with the preset temporal features using an add method. Then, it uses a preset feature fusion model to fuse the second sub-visual features of the three second preview images to obtain the second sub-visual fusion feature (m, 1, 768). Furthermore, a multimodal fusion model can be used to process the second sub-visual fusion feature and the second sub-highlight text feature (m, 1, 768) to obtain the second highlight score of the target preview image, thereby determining whether the target preview image is in the highlight range.

[0376] The specific implementation of the second aspect can refer to the relevant descriptions of S1601, S1502, S1702, S1901 and S1902 in the aforementioned embodiments, which will not be repeated here.

[0377] After the first and second aspects are executed, the mobile phone can determine the highlight frames based on the low-latency frame output strategy. The specific implementation of the mobile phone determining the highlight frames based on the low-latency frame output strategy can be referred to the relevant descriptions in S1903-S1905 of the aforementioned embodiments, and will not be repeated here.

[0378] The implementation and beneficial effects of the technical solutions provided in the above embodiments can be referred to the relevant content of the capture method provided in the foregoing embodiments, and will not be repeated here.

[0379] It is understood that, in order to achieve the aforementioned functions, the electronic device includes corresponding hardware structures and / or software modules for performing each function. Those skilled in the art should readily recognize that, based on the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein, the embodiments of the present invention can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed in a hardware-driven or software-driven manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of the embodiments of this application.

[0380] This application embodiment can divide the above-described electronic device into functional modules based on the method example described above. For example, each function can be divided into its own functional modules, or two or more functions can be integrated into one processing module. The integrated modules can be implemented in hardware or as software functional modules. It should be noted that the module division in this embodiment is illustrative and only represents one logical functional division; in actual implementation, there may be other division methods.

[0381] This application also provides an electronic device, which includes a memory and one or more processors; the memory is coupled to the processors; wherein the memory stores computer program code, which includes computer instructions, and when the computer instructions are executed by the processor, the electronic device performs the image capture method provided in the foregoing embodiments. The specific structure of this electronic device can be referred to... Figure 4 The structure of the electronic device shown is illustrated.

[0382] This application also provides a computer-readable storage medium that includes computer instructions that, when executed on an electronic device, cause the electronic device to perform the snapshot method provided in the foregoing embodiments.

[0383] This application also provides a computer program product containing executable instructions that, when run on an electronic device, cause the electronic device to perform the snapshot method provided in the foregoing embodiments.

[0384] This application also provides a training device, which includes a processor and a memory for storing executable instructions of the processor. The processor is configured to execute the executable instructions to implement the model training method provided in the foregoing embodiments. The specific structure of this training device can be found in [reference needed]. Figure 6 The structure of the training device shown is illustrated.

[0385] This application also provides a computer-readable storage medium including computer instructions that, when executed on a training device, cause the training device to perform the model training method provided in the foregoing embodiments.

[0386] This application also provides a computer program product containing executable instructions that, when run on a training device, cause the training device to perform the model training method provided in the foregoing embodiments.

[0387] Through the above description of the embodiments, those skilled in the art can clearly understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.

[0388] In the several embodiments provided in this application, it should be understood that the disclosed apparatus / device and method can be implemented in other ways. For example, the apparatus / device embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0389] The units described as separate components may or may not be physically separate. A component shown as a unit can be one or more physical units; that is, it can be located in one place or distributed in multiple different locations. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0390] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0391] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, essentially or in other words, the parts that contribute to the prior art, or all or part of the technical solutions, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0392] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for capturing images, characterized in that, Applied to electronic devices, the method includes: When the electronic device displays a preview interface and activates the snapshot function, it obtains a set of pending frames from the preview images displayed on the preview interface; the set of pending frames includes the latest preset number of first preview images; The electronic device obtains visual features corresponding to the first preview image based on a visual feature extraction model. The visual feature extraction model is obtained by model distillation based on a visual language model. The visual language model has the ability to process image and text information simultaneously and obtain output information corresponding to both image and text information. The difference between the number of parameters of the visual language model and the number of parameters of the visual feature extraction model is greater than a first threshold. The electronic device fuses the visual features of all the first preview images to obtain visual fusion features; The electronic device determines a first highlight score and a second highlight score for the target preview image based on the visual fusion features and preset highlight text features; wherein, the preset highlight text features are the text features of the highlight text corresponding to the highlight images of highlight moments in a preset dynamic scene; the first highlight score is used to characterize the similarity between the target preview image and the highlight text; the second highlight score is used to characterize the matching degree between the target preview image and the highlight text; the target preview image is the first preview image generated in the target order from the set of undetermined frames; The electronic device determines the captured image based on the first and second "wonderful" scores.

2. The method according to claim 1, characterized in that, The electronic device obtains a set of frames to be determined from the preview image displayed on the preview interface, including: The electronic device extracts a preview image from the preview image displayed on the preview interface at a preset frame interval to obtain the set of frames to be determined.

3. The method according to claim 1 or 2, characterized in that, The electronic device obtains visual features corresponding to the first preview image based on a visual feature extraction model, including: When the electronic device determines that a first object exists in all the first preview images in the set of undetermined frames, and that the second object in each of the first preview images in the set of undetermined frames is the same object, it obtains the visual features corresponding to the first preview image based on a visual feature extraction model; wherein, the first object includes any one of the following: dog, cat, or human; and the second object is the first object that occupies the largest area.

4. The method according to claim 3, characterized in that, The electronic device obtains visual features corresponding to the first preview image based on a visual feature extraction model, including: The electronic device crops the first preview image based on the category of the second object to obtain the second preview image; The electronic device inputs the second preview image into the visual feature extraction model to obtain the visual features of the second preset image; the visual features of the second preset image are the visual features corresponding to the first preview image to which the second preset image belongs.

5. The method according to claim 4, characterized in that, The visual features include a first sub-visual feature and a second sub-visual feature, and the visual fusion features include a first sub-visual fusion feature and a second sub-visual fusion feature; The electronic device fuses the visual features of all the first preview images to obtain visual fusion features, including: The electronic device determines the average value of the first sub-visual features of a preset number of second preview images as the first sub-visual fusion feature; The electronic device uses a preset feature fusion model to fuse a preset number of second sub-visual features of the second preview image to obtain the second sub-visual fusion feature.

6. The method according to claim 5, characterized in that, Before the electronic device fuses the second sub-visual features of a preset number of the second preview images using a preset feature fusion model to obtain the second sub-visual fusion features, the method further includes: The electronic device fuses the preset temporal features with the second sub-visual features of the second preview image to update the second sub-visual features of the second preview image; the preset temporal features are used to characterize the temporal change characteristics of the image content in the video, which includes all the exciting moments of the preset dynamic scene.

7. The method according to claim 5, characterized in that, The preset excellent text features include a first sub-excellent text feature and a second sub-excellent text feature; The electronic device determines a first and second "highlight score" for the target preview image based on the visual fusion features and preset "highlight text" features, including: The electronic device calculates a first brilliant score based on the first sub-visual fusion feature and the first sub-brilliant text feature; The electronic device uses a multimodal fusion model to process the second sub-visual fusion features and the second sub-highlighted text features to obtain the second highlight score.

8. The method according to claim 7, characterized in that, The electronic device calculates a first "highlight score" based on the first sub-visual fusion feature and the first sub-highlight text feature, including: The electronic device calculates the image-text contrast loss ITC based on the first sub-visual fusion feature and the first sub-highlighted text feature. The electronic device calculates the image-to-image contrast loss (IIC) based on the first sub-visual fusion feature and the preset brilliant image features. The electronic device determines the first excellent score by taking the weighted average of the ITC and the IIC.

9. The method according to any one of claims 1-7, characterized in that, The electronic device determines the captured image based on the first and second highlight scores, including: The electronic device determines the highlight frames based on the first highlight score and the second highlight score; The electronic device responds to the capture command and determines the best frame corresponding to the capture command as the captured image.

10. The method according to claim 8, characterized in that, The electronic device determines highlight frames based on the first highlight score and the second highlight score, including: If the second highlight score is greater than the probability threshold, the electronic device determines the highlight window based on the maximum latency frame number and obtains the first highlight score of the first preview image in the highlight window; the highlight window is a combination of several consecutive preview images with a maximum latency frame number, and the latest preview image in the highlight window is the latest first preview image in the undetermined frame set; the maximum latency frame number is greater than the preset number; If it is determined that the first target preview image has the highest score among the historical first preview images in the highlight window, the first target preview image is determined as the highlight frame; the historical first preview image is the first preview image in the highlight window that was generated before the target preview image.

11. A model training method, characterized in that, Applied to training devices, the method includes: The training device acquires sample data and corresponding label information; wherein, the sample data pair includes multiple first sample data, multiple second sample data, and a second sub-highlight text feature; the second sub-highlight text feature is the text feature of the highlight text corresponding to the highlight image of a highlight moment in a preset dynamic scene; the first sample data pair includes sample visual fusion features of a sample frame set, and the second sample data includes sample visual features of sample images; the sample frame set includes a preset number of sample video frames, the sample video frames are video frames in a video of any scene, the preset number of sample video frames are generated in an arithmetic sequence, and the common difference is the same as the preset frame interval; the sample image is an image in any scene. The training device uses sample data as training data and the label information corresponding to the sample data as supervision information to iteratively train and obtain an initial multimodal fusion model, thus obtaining a multimodal fusion model.

12. An electronic device, characterized in that, include: The device includes a display, a memory, and one or more processors; the display and the memory are both coupled to the processors; wherein the memory stores computer program code, the computer program code including computer instructions, which, when executed by the processor, cause the electronic device to perform the snapshot method as described in any one of claims 1-10.

13. A computer-readable storage medium, characterized in that, It includes computer instructions that, when executed on an electronic device, cause the electronic device to perform the snapshot method as described in any one of claims 1-10.

14. A training device, characterized in that, include: processor; Memory used to store the processor's executable instructions; The processor is configured to execute the executable instructions to implement the model training method as described in claim 11.

15. A computer-readable storage medium, characterized in that, Includes computer instructions that, when executed on a training device, cause the electronic device to perform the model training method as described in claim 11.