Image processing method and device and storage medium
By implementing the image processing method in the terminal device, the image of exciting actions in the shooting content is recognized, and the problem of single action recognition scheme in the prior art is solved, and the exciting image recognition with high efficiency and low power consumption is achieved.
Patent Information
- Application Number
- CN202311626921.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-29
- Publication Date
- 2025-06-06
AI Technical Summary
The existing action recognition scheme is relatively single and cannot meet the diverse needs of users, especially when identifying images of exciting actions in the shooting content.
By implementing the image processing method in the terminal device, an image stream collected by the camera is obtained, and action recognition is performed based on text description statements of preset actions. The method includes acquiring a first image stream, identifying a preset action in the first image, acquiring a continuous second image, calculating a wonderfulness score of each image, and determining a target image.
It realizes the accurate identification of wonderful images with exciting actions during the action recognition process, improves recognition efficiency, reduces equipment power consumption, and meets the diverse needs of users.
Smart Images

Figure CN120108029A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of terminals, and in particular to an image processing method, device and storage medium. Background Art
[0002] With the continuous development of terminal technology, users can use terminal devices such as mobile phones and tablets to record their lives by shooting videos or images. In order to automatically understand the content of the shooting, shooting content understanding technologies such as motion recognition have also been rapidly developed. However, existing motion recognition solutions are often relatively simple and cannot meet the needs of users. Summary of the invention
[0003] In order to solve the above technical problems, the present application provides an image processing method, device and storage medium, which can identify wonderful images with wonderful actions during the process of action recognition in the captured content, thereby meeting the user's usage needs.
[0004] In a first aspect, the present application provides an image processing method. The method includes: a terminal device obtains a first image stream captured by a camera in response to a first operation of a user, wherein the first image stream includes a first image. The terminal device performs action recognition of a preset action on the first image based on a text description sentence of the preset action. When the preset action is recognized in the first image, at least one second image is obtained. Wherein, at least one second image is continuous with the first image in the first video stream. The terminal device obtains a wonderfulness score of the first image and a wonderfulness score of each second image in at least one second image. The terminal device determines a target image based on the wonderfulness score of the first image and the wonderfulness score of the second image.
[0005] Through this embodiment, after acquiring the first image stream, the preset action in the first image can be accurately recognized based on the text description statement of the preset action and the multimodal information of text and image. Also, for the first image with the preset action, after at least one second image in the first image stream, since the at least one second image and the first image are continuous and the action of the photographed object is continuous, the second image adjacent to the first image will also have the preset action. Thus, after the wonderfulness score of at least one second image and the wonderfulness score of the first image are obtained, an image with a higher degree of wonderfulness can be selected as the target image from among the multiple images with the preset action, so that the wonderful images with wonderful actions can be accurately identified in the process of performing action recognition on the image stream, thereby meeting the user's usage needs. Furthermore, it is not necessary to perform action recognition on each image to accurately find images with wonderful actions, thereby improving recognition efficiency and reducing device power consumption.
[0006] Among them, the user's first operation may be a shooting operation performed by the user on the target shooting interface. Among them, the target shooting interface may be a shooting interface or a recording interface related to capturing wonderful photos. Exemplarily, it may be an operation performed by the user on a shooting interface with a wonderful photo capturing function. In one example, the first operation may be a user's triggering operation of a shooting control on a shooting interface in a "wonderful capture" mode. In another example, when the "one shot, multiple results" shooting function is turned on, the first operation may be a user's triggering operation of a shooting control on a shooting interface in a "recording" mode. In yet another example, the first operation may be a user's triggering operation of a wonderful capture control on a shooting interface in a "photographing" mode.
[0007] Among them, the first image stream can be a video stream to be processed or an image stream to be processed for image processing, wherein the video stream to be processed can be composed of multiple video frames to be processed collected at different times, and the image stream to be processed can be composed of multiple image frames to be processed collected at different times. Exemplarily, after the terminal device obtains the image light signal through the camera, the image light signal can be converted into a Bayer image through the image sensor. Then, the Bayer image is processed using ISP, etc. to obtain the original video frame or the original image frame. Then, after downsampling the original video frame or the original image frame, the video frame to be processed or the image frame to be processed is obtained.
[0008] Among them, the terminal device can be a mobile phone, computer, tablet computer and other devices.
[0009] The first image may be an image in the first image stream. For example, it may be a video frame to be processed or an image frame to be processed. Exemplarily, the first image may be an image frame or a video frame collected in real time, so that a wonderful image can be determined during recording or taking a photo.
[0010] The second image and the first image are continuous. For example, if 10 images are collected continuously, the 10th image may be the first image, and the 4th to 9th images may be the second images. Exemplarily, the first image may be located between at least one second image, or the first image may be collected later than all the second images, or the first image may be collected earlier than all the second images.
[0011] The preset action may be an action that can be action-recognized. For example, the preset action may be jumping, running, walking, sitting, and the like.
[0012] The text description sentence of the preset action may be a sentence in a fixed text format for describing the preset action. For example, it may be generated according to the action classification label (object) and prompt word (prompt) template of the preset action.
[0013] Among them, action recognition can be used to identify whether the subject in the image has made a preset action. Among them, the subject can be a person, an animal, a cartoon character, etc.
[0014] The wonderfulness score is used to measure the wonderfulness of the image. For example, it can be used to measure the wonderfulness of the action of the subject in the image. For example, the action height, the stretch of the action, etc. In some embodiments, the wonderfulness score can be determined using a wonderfulness scoring model.
[0015] The target image may contain a wonderful video frame or a wonderful image frame of a wonderful action, for example, an image frame or a video frame acquired near a position with the highest degree of action completion (highest action, most relaxed action).
[0016] Among them, the terminal device can determine the first target image according to the maximum score of the wonderfulness score of the first image and the wonderfulness score of each second image. Exemplarily, when other images with preset actions are also identified, the maximum score of the wonderfulness score of the first image and the wonderfulness score of each second image, the wonderfulness scores of other images, and the wonderfulness scores of images continuous with other images can also be determined, and the image corresponding to the maximum score is determined as the first target image. In another exemplary embodiment, when other images with preset actions are not identified, the image corresponding to the maximum score of the wonderfulness score of the first image and the wonderfulness score of each second image can be determined as the first target image.
[0017] According to the first aspect, or any implementation of the first aspect above, the first image stream also includes a third image; based on the wonderfulness score of the first image and the wonderfulness score of each second image, determining the first target image includes: based on the text description sentence of the preset action, performing action recognition of the preset action on the third image; in the case where the preset action is recognized in the third image, obtaining at least one fourth image, wherein the at least one fourth image is continuous with the third image in the first image stream; obtaining the wonderfulness score of the third image and the wonderfulness score of each fourth image in the at least one fourth image; determining the first target image based on the wonderfulness score of the first image, the wonderfulness score of each second image, the wonderfulness score of the third image, and the wonderfulness score of each fourth image. In this way, when the third image also includes the preset action, the fourth image continuous with the third image also includes the preset action. When the first target image is selected according to the wonderfulness scores of the first image, the second image, the third image, and the fourth image, since the acquisition time of the third image and the first image is different, the target image with the highest action wonderfulness can be selected from the images acquired at different action times, thereby further improving the wonderfulness of the acquired target image.
[0018] The third image may be acquired at a third moment, the first image may be acquired at a first moment, and the third moment is later than the first moment.
[0019] Exemplarily, the image with the highest wonderful score among the wonderful scores of the first image, the second image, the third image, and the fourth image may be determined as the first target image.
[0020] Exemplarily, when other images with preset actions are identified, the first target image may be determined according to the wonderfulness scores of the other images and the wonderfulness scores of images continuous with the other images.
[0021] According to the first aspect, or any implementation of the first aspect above, based on the wonderfulness score of the first image, the wonderfulness score of each second image, the wonderfulness score of the third image, and the wonderfulness score of each fourth image, determining the first target image includes: obtaining the maximum value of the wonderfulness score of the first image and the wonderfulness score of each second image to obtain a first score; storing the first score in the target cache; obtaining the maximum value of the wonderfulness score of the third image and the wonderfulness score of each fourth image to obtain a second score; storing the second score in the target cache; when the number of scores stored in the target cache reaches a preset number threshold, obtaining multiple scores including the first score and the second score from the target cache; obtaining the maximum score among the multiple scores, and determining the image corresponding to the maximum score as the first target image. In this way, since the first score and the second score can be collected at different stages of the action, after a sufficient number of scores are cached in the target cache, the image with the highest wonderfulness can be selected as the target image according to the score, so that the target image can be selected near the highest degree of action completion, thereby improving the wonderfulness of the target image.
[0022] For example, the first score and the image number of the image corresponding to the first score may be stored in the target storage area in a corresponding manner. For example, if the fifth image among the 1st to 7th images has the highest score of p1, the image number 5 and p1 may be stored in a corresponding manner.
[0023] Exemplarily, after determining the maximum score among the multiple scores, the image sequence number of the image corresponding to the maximum score may be output. If the image is a video frame, the image sequence number may be a frame index.
[0024] Exemplarily, the preset number may be 4. For example, when the scores stored in the target cache include a first score p1, a second score p2, and a third score p3, if the target cache acquires a fourth score p4, the maximum value among p1-p4 may be determined, and the image corresponding to the maximum value is the first target image.
[0025] According to the first aspect, or any implementation of the first aspect above, based on the wonderfulness score of the first image and the wonderfulness score of each second image, determining the first target image includes: obtaining the maximum value of the wonderfulness score of the first image and the wonderfulness score of each second image to obtain the first score; storing the first score in the target cache area; after acquiring the fifth image through the camera, when the difference between the image sequence number of the fifth image and the reference image sequence number is greater than or equal to the preset difference threshold, obtaining the stored score in the target cache area, the stored score includes the first score, and the reference image sequence number is the image sequence number of the image corresponding to the first score cached in the target cache area; obtaining the maximum score in the stored score, and determining the image corresponding to the maximum score as the first target image. In this way, the image with the highest wonderfulness can be selected as the target image according to the score every time a certain number of images are collected, so that the target image can be selected near the highest action completion degree, thereby improving the wonderfulness of the target image.
[0026] Exemplarily, the preset difference threshold may be set according to actual conditions and specific scenarios, for example, may be set to 9.
[0027] Exemplarily, if after the fifth image is acquired, the cache scores of the target cache area only include the first score p1, then the image corresponding to the first score p1 may be determined as the first target image.
[0028] Exemplarily, if after acquiring the fifth image, the cache scores of the target cache area include the first score p1 and other scores, a maximum value may be determined between the first score p1 and the other scores, and the image corresponding to the maximum value may be determined as the first target image.
[0029] Exemplarily, the fifth image may be acquired at a fifth moment, which is later than the first moment.
[0030] Exemplarily, when the first score is the first score stored in the target cache, the reference image sequence number may be the image sequence number of the image corresponding to the first score.
[0031] According to the first aspect, or any implementation of the first aspect above, after determining the first target image, the method further includes: clearing the target buffer area. In this way, the terminal device can determine the next target image according to the new score in the target buffer area, thereby improving the recognition accuracy of the wonderful image.
[0032] According to the first aspect, or any implementation of the first aspect above, the method further includes: acquiring a sixth image through a camera; in the case where the sixth image is separated from the reference image by a preset number of images, based on the text description statement of the preset action, performing action recognition of the preset action on the sixth image, the reference image is the last image collected among the third image and at least one fourth image, or the fifth image; in the case where the preset action is recognized in the sixth image, acquiring at least one seventh image in the first image stream containing the sixth image, wherein at least one sixth image and the seventh image are continuous in the first image stream; acquiring the wonderfulness score of the sixth image and the wonderfulness score of each seventh image in at least one seventh image; determining the second target image based on the wonderfulness score of the sixth image and the wonderfulness score of each seventh image. Through this embodiment, after each target image is recognized, the next target image can be recognized after a cooling interval, thereby avoiding misrecognition of the cooling device and improving the recognition accuracy of the wonderful image.
[0033] Exemplarily, the predetermined number may be the interval length of the cooling interval, for example, it may be set to 30.
[0034] Exemplarily, the second target image is determined in a similar manner to the first target image. For example, the maximum score of the sixth image and the maximum score of each seventh image can be cached in the target cache area, so that the number of scores cached in the target cache area reaches a preset number threshold, or when the difference between the image sequence number of the acquired image and the reference image sequence number is greater than or equal to a preset difference threshold, the scores cached in the target cache area are acquired, and then the second target image is determined based on the maximum value of the cached scores.
[0035] For example, if the sixth image is not separated from the reference image by a preset number of images, that is, the sixth image is in the cooling interval, the action recognition of the preset action may not be performed on the sixth image. Alternatively, the action recognition may be performed, but the wonderfulness score is not calculated. Alternatively, the wonderfulness scores of the sixth image and the seventh image may be calculated, but the maximum score of the wonderfulness scores of the sixth image and the seventh image is not placed in the target buffer area.
[0036] According to the first aspect, or any implementation of the first aspect above, when the sixth image is separated from the reference image by a preset number of images, based on the text description of the preset action, the sixth image is subjected to action recognition of the preset action, including: when the difference between the image sequence number of the sixth image and the image sequence number of the reference image is greater than or equal to the interval length of the cooling interval, the sixth image is subjected to action recognition of the preset action based on the text description of the preset action. In this way, after acquiring an image (such as the sixth image), it is possible to accurately determine whether it is in the cooling interval based on the image sequence number of the acquired image, and if it is not in the cooling interval, the next target image is determined, thereby effectively avoiding the misidentification of wonderful images and improving the recognition accuracy.
[0037] According to the first aspect, or any implementation of the first aspect above, the terminal device includes a first flag, and the method further includes: after acquiring the first target image, setting the first flag to a first value; and acquiring an eighth image through a camera; taking the difference between the image sequence number of the eighth image and the image sequence number of the reference image; when the difference is equal to a preset number, setting the first flag to a second value. In this way, the first flag can be used to accurately indicate whether the terminal device has recognized the target image (highlight frame), so that the terminal device can quickly confirm the target image recognition status of the terminal device according to the first flag.
[0038] Exemplarily, the first flag bit may be a wonderful frame flag bit. The first value is in a "TRUE" state, indicating that a wonderful frame has been identified. The second value is in a "FALSE" state, indicating that a wonderful frame has not been identified. Optionally, the terminal device may set the default value of the first flag bit to the second value.
[0039] According to the first aspect, or any implementation of the first aspect above, when the sixth image is separated from the reference image by a preset number of images, based on the text description of the preset action, the sixth image is subjected to action recognition of the preset action, including: when the first flag is set to a second value, based on the text description of the preset action, the sixth image is subjected to action recognition of the preset action. In this way, it is possible to quickly determine whether it is in the cooling interval according to the state of the first flag, and to quickly determine whether the received sixth image needs to be processed, thereby improving the processing rate of the sixth image, and thereby improving the recognition efficiency of the target image.
[0040] According to the first aspect, or any implementation of the first aspect above, based on the text description sentence of the preset action, the action recognition of the preset action is performed on the first image, including: obtaining image features of the first image; obtaining text features of the text description sentence; obtaining a similarity score between the image features of the first image and the text features of the text description sentence; based on the similarity score, determining whether the preset action is recognized in the first image. In this way, through this embodiment, the preset action can be recognized on the first image based on the multimodal information of image-text, thereby improving the recognition accuracy.
[0041] For example, the similarity score may be calculated by calculating the cosine similarity between the image feature and the text feature. The similarity score may also be referred to as a similarity score.
[0042] According to the first aspect, or any implementation of the first aspect above, obtaining the image features of the first image includes: inputting the first image into the lightweight model to obtain the image features of the first image. In this way, since the lightweight model can be easily deployed in the terminal device, the image features of the first image can be accurately extracted in the terminal device, and the power consumption of the terminal device is saved.
[0043] Exemplarily, the lightweight model may be MobileNetV1, MobileNetV2, MobileNetV3, ShuffleNet, ShuffleNetV2, SqueezeNet, Xception, etc.
[0044] According to the first aspect, or any implementation of the first aspect above, the lightweight model is obtained by distilling the image encoder of the multimodal large model by the server. In this way, by distillation, the extraction accuracy of image features is taken into account while reducing the number of model parameters, so that high-precision feature extraction can be achieved on the terminal device, and energy consumption of the terminal device is saved.
[0045] Exemplarily, the multimodal large model may be a large language model such as a CLIP model, an Action CLIP model, a UNITER model, a ViLT model, or a BEiT model.
[0046] Exemplarily, the image encoder may be a ResNet or Vit model.
[0047] According to the first aspect, or any implementation of the first aspect above, obtaining text features of a text description sentence includes: obtaining text features of a text description sentence sent by a server, wherein the text features of the text description sentence are obtained by the server extracting text features from the text description sentence using a text encoder, and the text encoder is a text encoder of a multimodal large model. In this way, since the preset action is relatively fixed, the extraction of features of the text description sentence is implemented on the server side, which saves the function of the terminal device while ensuring the extraction accuracy of the text features.
[0048] Exemplarily, the text encoder and the image encoder may belong to the same multimodal macro model.
[0049] Exemplarily, when the preset action changes, the server may extract text features of the text description statement of the new preset action and send it to the terminal device.
[0050] Exemplarily, the text encoder may be a CBOW model or a text Transformer model.
[0051] According to the first aspect, or any implementation of the first aspect above, the number of preset actions is multiple, and obtaining a similarity score between the image features of the first image and the text features of the text description sentence includes: obtaining a similarity score between the image features of the first image and the text features of the text description sentence of each preset action in multiple preset actions; based on the similarity score, determining whether the preset action is recognized in the first image, including: determining the maximum similarity score among the multiple similarity scores obtained; based on the maximum similarity score and the preset score threshold, determining whether the preset action is recognized in the first image. In this way, when there are multiple preset actions, when the maximum similarity score is greater than or equal to the preset score threshold, it can be determined that at least one of the multiple preset actions exists in the first image, without determining the specific category of the action in the image, so that the preset action can be accurately recognized in the first image, thereby improving recognition efficiency and taking into account recognition accuracy.
[0052] Exemplarily, the preset action may be a pre-set action such as running, jumping, rolling, sitting, etc.
[0053] Exemplarily, the motion recognition method of the third image and the sixth image is the same as the motion recognition method of the first image.
[0054] According to the first aspect, or any implementation of the first aspect above, after the first image is input into the lightweight model and the image features of the first image are obtained, the method further includes: obtaining the image features of at least one ninth image in the image feature sequence including the image features of the first image, the image features of at least one ninth image and the image features of the first image are continuous in the image feature sequence, wherein the image feature sequence is a sequence composed of image features of images in the second image stream, and the second image stream is obtained by extracting frames from the first image stream; smoothing the image features of the first image using the image features of at least one ninth image; wherein obtaining the similarity score between the image features of the first image and the text features of the text description sentence includes: obtaining the similarity score between the image features of the first image after smoothing and the text features of the text description sentence. In this way, the outliers in the image features can be eliminated by smoothing, the accuracy of the image features is improved, and the accuracy of subsequent action recognition and the accuracy of wonderful frame recognition are thereby improved.
[0055] Exemplarily, the smoothing process may include averaging and normalization processes.
[0056] Exemplarily, the first image stream may extract an image frame every interval of a preset number of images to obtain the second image stream.
[0057] Exemplarily, the first image is also an image in the second image stream. In other words, the first image can be obtained by extracting frames.
[0058] Exemplarily, a first image stream may be acquired, and frames of the first image stream may be extracted to obtain a second image stream. Then, image features of each image in the second image stream may be acquired to obtain an image feature sequence corresponding to the second image stream. Also exemplarily, image features of each image in the first image stream may be acquired to obtain an image feature sequence corresponding to the first image stream. Then, a second image value may be extracted at intervals of a preset number of features in the image feature sequence to obtain a feature sequence of the second image stream.
[0059] Exemplarily, the first image may be located between the ninth images, or the first image may be acquired later than all the ninth images, or the first image may be acquired earlier than all the ninth images.
[0060] According to the first aspect, or any implementation of the first aspect above, in an image feature sequence including the image feature of the first image, obtaining at least one image feature of a ninth image includes: using a first window to slide in the image feature sequence to a first window position corresponding to the image feature of the first image; in the first window located at the first window position, obtaining other features except the image feature of the first image, to obtain at least one image feature of the ninth image. In this way, the image feature of the ninth image can be obtained by drawing a sliding window.
[0061] Exemplarily, the window size of the first window may be 3.
[0062] Exemplarily, within the first window at the first window position, the image feature of the first image may be located at the rightmost end, or at the middle position, or at the leftmost end.
[0063] According to the first aspect, or any implementation of the first aspect above, obtaining at least one second image in a first image stream containing a first image includes: using a second window to slide in the first image stream to a second window position corresponding to the first image; in the second window located at the second window position, obtaining other images except the first image to obtain at least one second image. In this way, the second image adjacent to the first image can be quickly determined by sliding the window, and since the action is often continuous, the second image also often contains similar actions, so that the image with more exciting actions can be selected from the first image and the second image as the target image, thereby improving the action excitement of the target image.
[0064] Exemplarily, the window size of the second window may be 7.
[0065] Exemplarily, in the second window at the second window position, the first image may be located at the rightmost end, or at the middle position, or at the leftmost end.
[0066] According to the first aspect, or any implementation of the first aspect above,
[0067] Exemplarily, the third target image may be a wonderful photo.
[0068] For example, one may be selected from the first target image and the second target image, and the selected target image may be directly used as the third target image. Alternatively, the selected target image may be processed (such as fused with adjacent images) and used as the third target image. Alternatively, the third target image may be determined based on the fourth target image after the fourth target image is determined by other wonderful frame recognition algorithms.
[0069] Exemplarily, a third target image may be displayed.
[0070] In a second aspect, the present application provides an electronic device. The electronic device includes: one or more processors; a memory; and one or more computer programs, wherein the one or more computer programs are stored in the memory, and when the computer programs are executed by the one or more processors, the electronic device performs the following steps: in response to a first operation of a user, obtaining a first image stream captured by a camera, wherein the first image stream includes a first image; based on a text description statement of a preset action, performing action recognition of a preset action on the first image; when a preset action is recognized in the first image, obtaining at least one second image, wherein the at least one second image is continuous with the first image in the first image stream; obtaining a wonderfulness score of the first image and a wonderfulness score of each second image in at least one second image; and determining a first target image based on the wonderfulness score of the first image and the wonderfulness score of each second image.
[0071] According to the second aspect, or any implementation of the second aspect above, the first image stream also includes a third image; when the computer program is executed by one or more processors, the electronic device executes the following steps: based on the text description sentence of the preset action, performing action recognition of the preset action on the third image; when the preset action is recognized in the third image, obtaining at least one fourth image, wherein the at least one fourth image is continuous with the third image in the first image stream; obtaining a wonderfulness score of the third image and a wonderfulness score of each fourth image in at least one fourth image; and determining the first target image based on the wonderfulness score of the first image, the wonderfulness score of each second image, the wonderfulness score of the third image, and the wonderfulness score of each fourth image.
[0072] According to the second aspect, or any implementation of the second aspect above, when the computer program is executed by one or more processors, the electronic device performs the following steps: obtain the maximum value of the wonderfulness score of the first image and the wonderfulness score of each second image to obtain a first score; store the first score in a target cache area; after acquiring the fifth image through the camera, when the difference between the image sequence number of the fifth image and the reference image sequence number is greater than or equal to a preset difference threshold, obtain the stored score in the target cache area, the stored score includes the first score, and the reference image sequence number is the image sequence number of the image corresponding to the first score cached in the target cache area; obtain the maximum score in the stored scores, and determine the image corresponding to the maximum score as the first target image.
[0073] According to the second aspect, or any implementation of the second aspect above, the first image stream also includes a third image; when the computer program is executed by one or more processors, the electronic device performs the following steps: acquiring a sixth image through a camera; when the sixth image is separated from the reference image by a preset number of images, based on a text description statement of the preset action, performing action recognition of a preset action on the sixth image, the reference image being the last image captured among the third image and at least one fourth image, or the fifth image; when the preset action is recognized in the sixth image, acquiring at least one seventh image in the first image stream containing the sixth image, wherein at least one sixth image and the seventh image are continuous in the first image stream; acquiring a wonderfulness score of the sixth image and a wonderfulness score of each of the at least one seventh image; and determining a second target image based on the wonderfulness score of the sixth image and the wonderfulness score of each of the seventh images.
[0074] The second aspect and any implementation of the second aspect correspond to the first aspect and any implementation of the first aspect respectively. The technical effects corresponding to the second aspect and any implementation of the second aspect can refer to the technical effects corresponding to the above-mentioned first aspect and any implementation of the first aspect, which will not be repeated here.
[0075] In a third aspect, the present application provides a computer-readable medium for storing a computer program, wherein the computer program includes instructions for executing the method in the first aspect or any possible implementation of the first aspect.
[0076] The third aspect and any implementation of the third aspect correspond to the first aspect and any implementation of the first aspect, respectively. The technical effects corresponding to the third aspect and any implementation of the third aspect can refer to the technical effects corresponding to the first aspect and any implementation of the first aspect, which will not be repeated here.
[0077] In a fourth aspect, the present application provides a computer program, comprising instructions for executing the method in the first aspect or any possible implementation of the first aspect.
[0078] The fourth aspect and any implementation of the fourth aspect correspond to the first aspect and any implementation of the first aspect, respectively. The technical effects corresponding to the fourth aspect and any implementation of the fourth aspect can refer to the technical effects corresponding to the above-mentioned first aspect and any implementation of the first aspect, which will not be repeated here.
[0079] In a fifth aspect, the present application provides a chip, the chip comprising a processing circuit and a transceiver pin, wherein the transceiver pin and the processing circuit communicate with each other through an internal connection path, and the processing circuit executes the method in the first aspect or any possible implementation of the first aspect to control the receiving pin to receive a signal and control the sending pin to send a signal.
[0080] The fifth aspect and any implementation of the fifth aspect correspond to the first aspect and any implementation of the first aspect, respectively. The technical effects corresponding to the fifth aspect and any implementation of the fifth aspect can refer to the technical effects corresponding to the first aspect and any implementation of the first aspect, which will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0081] Figure 1A-1B A group of exemplary video frames provided by an embodiment of the present application are respectively shown;
[0082] Figure 2A-2L A group of user interface schematic diagrams of the image processing method provided by the embodiment of the present application are exemplarily shown;
[0083] Figure 3A-3D Another group of user interface schematic diagrams of the image processing method provided by the embodiment of the present application is exemplarily shown;
[0084] Figure 4A-4D Another group of user interface schematic diagrams of the image processing method provided by the embodiment of the present application is exemplarily shown;
[0085] Figure 5 A system architecture diagram of an image processing system provided by an embodiment of the present application is shown;
[0086] Figure 6 A schematic diagram of a process flow of an image processing method provided in an embodiment of the present application is shown;
[0087] Figure 7 A schematic diagram of the structure of an electronic device is shown;
[0088] Figure 8 is a software structure block diagram of the electronic device of the embodiment of the present application;
[0089] Fig. 9 A schematic diagram showing a flow chart of a video frame processing part of an image processing method provided in an embodiment of the present application;
[0090] Fig.10 A schematic diagram of an exemplary video frame processing process is shown;
[0091] Fig.11 A schematic diagram of an exemplary process of generating an original video stream is shown;
[0092] Fig.12 A schematic diagram of the structure of an exemplary multi-modal large model provided in an embodiment of the present application is shown;
[0093] Fig.13 A schematic diagram of an exemplary training process of a multimodal large model is shown;
[0094] Fig.14 A schematic diagram showing a flow chart of an action recognition part of another image processing method provided in an embodiment of the present application;
[0095] Fig.15 A schematic diagram showing an exemplary process of action recognition provided by an embodiment of the present application is shown;
[0096] Fig.16 A schematic diagram of the structure of an exemplary lightweight model provided in an embodiment of the present application is shown;
[0097] Fig.17 A schematic diagram of a knowledge distillation process provided by an embodiment of the present application is shown;
[0098] Fig.18 A schematic diagram of another knowledge distillation provided by an embodiment of the present application is shown;
[0099] Fig.19 A schematic diagram showing another exemplary action recognition process provided by an embodiment of the present application is shown;
[0100] Fig. 20 A schematic diagram of an exemplary smoothing process provided by an embodiment of the present application is shown;
[0101] Fig.21 Another exemplary schematic diagram of the training process of a multimodal large model is shown;
[0102] Fig. 22 A schematic diagram showing an exemplary method of determining a maximum score provided in an embodiment of the present application is shown;
[0103] Fig.23 A schematic diagram showing an exemplary action recognition process provided by an embodiment of the present application;
[0104] Fig.24 A schematic flow chart of a wonderful frame recognition part of an image processing method provided in an embodiment of the present application is shown;
[0105] Fig.25 A schematic diagram of a flow chart of an exemplary wonderful frame recognition process provided by an embodiment of the present application is shown;
[0106] Fig.26A schematic diagram of an exemplary wonderful frame recognition process provided by an embodiment of the present application is shown;
[0107] Fig. 27 A schematic diagram of training an exemplary wonderfulness scoring model provided in an embodiment of the present application is shown;
[0108] Fig.28 A schematic diagram of an exemplary process of wonderful frame recognition provided by an embodiment of the present application is shown;
[0109] Fig.29 An exemplary curve diagram showing the change of the action completion degree of the photographed object over time is shown;
[0110] Fig.30 A schematic block diagram of a device according to an embodiment of the present application is shown. DETAILED DESCRIPTION
[0111] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0112] The term "and / or" in this article is merely a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone.
[0113] The terms "first" and "second" in the description and claims of the embodiments of the present application are used to distinguish different objects rather than to describe a specific order of objects. For example, a first target object and a second target object are used to distinguish different target objects rather than to describe a specific order of target objects.
[0114] In the embodiments of the present application, words such as "exemplary" or "for example" are used to indicate examples, illustrations or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a specific way.
[0115] In the description of the embodiments of the present application, unless otherwise specified, the meaning of "multiple" refers to two or more than two. For example, multiple processing units refer to two or more processing units; multiple systems refer to two or more systems.
[0116] In daily life, users of mobile phones and other terminal devices often record their lives by taking photos and videos. Therefore, image understanding technologies such as video have become one of the research directions of terminal devices. In image understanding technologies, image action recognition is an important research branch.
[0117] In the existing action recognition process, most of them focus on identifying the category of the action occurring in the entire video or preview image, which cannot meet the usage needs of users.
[0118] Based on this, an embodiment of the present application provides an image processing method, which can, after recognizing a preset action in a captured first image, determine a target image where a wonderful action has occurred based on the wonderfulness scores of the first image and a second image continuous with the first image, thereby being able to identify wonderful images with wonderful actions during the process of action recognition in the captured content, thereby meeting the user's usage needs.
[0119] It should be noted that the image processing method provided in the embodiments of the present application can be applied to terminal devices with image processing functions such as mobile phones and tablet computers.
[0120] In some embodiments, the image processing method of the embodiment of the present application can be applied to a terminal device with a camera function. Specifically, the terminal device with a camera function can use the image processing method to process a video being shot (i.e., being shot), a video that has been shot, a video downloaded from the network, or a video sent by an external device.
[0121] In other embodiments, the image processing method of the embodiment of the present application can be applied to a terminal device that does not have a camera function. Specifically, the terminal device that does not have a camera function can download or receive a captured video or image stream sent by an external device from the network, and use the image processing method to process the captured video or image stream. For ease of understanding, the technical solution of the embodiment of the present application is described below using a video stream as an example. However, it should be understood that the image processing method of the embodiment of the present application can also be applied to an image stream.
[0122] To facilitate understanding, before introducing the image processing method provided in the embodiment of the present application, the wonderful capture process involved in the embodiment of the present application is first explained through wonderful video frames and wonderful photos.
[0123] Wonderful video frames are video frames with meaningful or representative picture content in the captured video content (a group of continuous video frames). For example, in the process of capturing the actions of the captured object (such as a person or an animal), the video frames captured at the moment (referred to as the wonderful moment) when the action is most complete (such as the highest jump or the most relaxed action) or the video frames captured near the wonderful moment can be determined as wonderful video frames.
[0124] For example, taking a person's jumping action as an example, Figure 1A A set of exemplary video frames provided by an embodiment of the present application is shown. Among them, Figure 1A The horizontal axis in the middle is the time axis, and the arrow direction of the time axis points to the direction of time growth. Figure 1A The vertical direction (vertical direction of the time axis) represents the jumping height of the person being filmed in the video frame. The farther away from the time axis, the higher the jumping height of the person.
[0125] like Figure 1A As shown, the video shot for the jumping action of the person being photographed may include multiple continuous video frames A1 to A7. In the video frame A1 captured at the first moment t, the person being photographed starts to jump, and in the video frames A1 to A3, the jumping height of the person being photographed gradually increases until the jumping height of the person being photographed in the video frame A4 captured at the second moment t+k (i.e., the wonderful moment) reaches the highest, and in the video frames A5 to A7 captured after the second moment t+k, the jumping height of the person being photographed gradually decreases until the person being photographed in A7 in the video captured at the third moment t+n ends the jump. In the embodiment of the present application, for the jumping action, the jumping height can be used to measure the action completion. Accordingly, when the user jumps to the highest point, it can be considered that the jumping action is the most complete. At this time, the terminal device can capture a wonderful video frame near the highest point of the jump (the second moment t+k).
[0126] It should be noted that during the shooting process, the terminal device can obtain the original video stream (including a group of continuous original video frames) through the camera. Since the resolution of the original video stream is relatively high, its data volume is often large and is not suitable for complex image processing. Therefore, in order to reduce the image processing cost and image processing time, the terminal device can perform a resolution reduction operation on the original video stream to obtain a video stream to be processed (i.e., a video stream for subsequent action recognition and wonderful frame recognition, including multiple video frames to be processed, and the resolution of the video frames to be processed is lower than that of the original video frames), and perform image processing operations such as action recognition on the video stream to be processed. Accordingly, in the image processing method provided in the embodiment of the present application, the terminal device can perform action recognition on multiple video frames to be processed, and then use the action recognition results of the multiple video frames to be processed as the basis for selecting wonderful video frames. For example, Figure 1A The video frames A1 to A7 shown may all be video frames to be processed, and one of the video frames to be processed (such as video frame A4) may be selected as a wonderful video frame to be processed.
[0127] After determining the wonderful video frame to be processed in the video stream to be processed, the terminal device can determine the original video frame corresponding to the wonderful video frame to be processed in the original video stream, and determine the corresponding original video frame as the target wonderful video frame, and then obtain a wonderful photo based on the target wonderful video frame, and display the wonderful photo on the terminal device. The terminal device can perform image processing such as image optimization on the target wonderful video frame to obtain the wonderful photo. Alternatively, the terminal device can directly use the target wonderful video frame as the wonderful photo.
[0128] As another example, taking the running action of an animal as an example, Figure 1B A set of exemplary video frames provided by an embodiment of the present application is shown. Figure 1B and Figure 1A The difference is that Figure 1B The video frames B1-B7 in FIG. 1 show the running action of the deer. Figure 1B The vertical direction of represents the degree of spreading of the deer’s limbs in the video frame. Figure 1B As shown, in video frames B1-B3, as the deer's hind legs accumulate strength, the deer's limbs gradually open, and in video frame B4 captured at the second moment t+k, the deer's limbs are opened to the maximum, and in video frames B5-B7, the deer's limbs gradually close. In the embodiment of the present application, the terminal device can measure the degree of action completion by the degree of opening and closing of the deer's limbs. Accordingly, since the deer's limbs are opened to the maximum in video frame B4, the terminal device can capture a wonderful video frame near the second moment t+k (the moment when the deer's limbs are opened to the maximum).
[0129] After introducing the wonderful capture process, for ease of understanding, the user interfaces of multiple feasible shooting scenes for wonderful capture are described below.
[0130] In some embodiments, in the shooting scene of the "Great Capture" mode, Figure 2A-2L A group of user interface schematic diagrams of the image processing method provided by the embodiment of the present application are exemplarily shown. It can be understood that Figure 2A The user interface described subsequently merely illustrates a possible user interface style of the terminal device 100 taking a mobile phone as an example, and should not constitute a limitation on the embodiments of the present application.
[0131] first, Figure 2A The main interface (homepage) of the terminal device 100 is shown as an example. Figure 2A As shown, the main interface may include a status bar 111 , a page indicator 112 , and a plurality of application icons 113 .
[0132] Among them, the status bar 111 may include one or more signal strength indicators of a mobile communication signal (also referred to as a cellular signal), a wireless fidelity (Wi-Fi) signal strength indicator, a battery status indicator, a time indicator, etc.
[0133] The page indicator 112 may be used to indicate the positional relationship between the currently displayed page and other pages. Specifically, the page indicator 112 may be used to indicate which page the user is currently browsing among the multiple pages carrying the multiple application icons 113. The user may swipe left and right to browse other pages.
[0134] For multiple application icons 113, each application icon 113 corresponds to an application. It should be noted that the application icons can be distributed on multiple pages, and the user can browse the application icons on other pages by sliding left and right. For the camera application icon 113A, it can be located in the bottom fixed bar 114 at the bottom of the main interface, or it can be directly located in the application icon display area 115 in the middle of the main interface, or in the application folder of the application icon display area 115, and its display position is not specifically limited.
[0135] Next, the terminal device 100 detects that the user acts on Figure 2A After the application trigger operation (such as a click operation) of the camera application icon 113A on the user interface shown in FIG. 1 is performed, the terminal device 100 may respond to the above operation by displaying a Figure 2B The user interface shown. Figure 2B The user interface for shooting (hereinafter referred to as shooting interface) provided by the terminal device 100 is exemplarily shown. It should be noted that in this embodiment and subsequent embodiments, in addition to opening the camera application by clicking the camera application icon 113A on the main interface, the camera application can also be opened by clicking the camera control in a specific application when the terminal device 100 runs a specific application, such as opening the camera application by clicking the camera control in the instant messaging system, and there is no specific limitation on this.
[0136] like Figure 2B As shown, the shooting interface may include a menu bar 121, a shooting control 122A, a preview window 123, and a review control 124. The following will describe each of them.
[0137] The menu bar 121 may include multiple shooting mode options, such as "capture", "take photo", "record video", etc. Figure 2B As shown, the terminal device 100 can detect the mode switching operation (such as left swipe, right swipe, click, etc.) performed by the user on the menu bar 121 to switch the shooting mode. Specifically, the user can switch to the shooting mode as shown in FIG. Figure 2C Highlight Shot mode shown.
[0138] exist Figure 2C In the interface shown, the "wonderful capture" mode can also be called the "AI capture" mode. In the "wonderful capture" mode, the terminal device 100 can capture and generate wonderful photos.
[0139] In the "wonderful capture" mode, when the terminal device 100 detects a trigger operation of the user on the shooting control 122A (such as a user operation such as clicking to trigger the wonderful capture function), the terminal device 100 can respond to the shooting operation and start recording video (including a group of video frames) through the camera, and display the video as shown in FIG. Figure 2D The recording interface shown. It should be noted that when the terminal device 100 starts recording, the preview window 123 displays a video frame 1231, and the video frame 1231 may be the first frame image of the recorded video. It should be noted that during the recording process, the preview video stream displayed in the preview window 123 may be a video stream obtained by processing the original video stream through a processing algorithm for the user to preview the video. Among them, the preview resolution may be lower than the original video stream, and there is no restriction on this.
[0140] Specifically, Figure 2D As shown, Figure 2D The recording interface shown includes a recording window and a shooting control 122B. The recording interface can display the video frames collected by the camera in real time, wherein the video frame 1232 can be a frame of video in the middle of the recorded video. The shooting control 122B and Figure 2B The style of the shooting control 122A of the shooting interface can be different, so as to prompt the user that the video recording has started through the shooting control 122B. Also, during the process of the terminal device 100 recording the video, the terminal device 100 can also generate a to-be-processed video stream based on the recorded video, and then select one of the to-be-processed video frames in the to-be-processed video stream as a wonderful video frame. Also, determine a wonderful photo based on the wonderful video frame.
[0141] pass Figure 2D In the recording interface shown, the terminal device 100 can automatically select wonderful video frames and wonderful photos. Figure 2E Another recording interface is shown. Figure 2E The recording interface shown is the same as Figure 2D The recording interface shown is different in that Figure 2E The recording interface shown may also include a snapshot control 125, so that the terminal device 100 can select wonderful videos and wonderful photos according to the user's triggering operation (such as clicking) on the snapshot control 125. Specifically,.
[0142] And, at a certain moment after starting recording (hereinafter referred to as the stop recording moment), the recording interface is as follows Figure 2F As shown, after the terminal device 100 detects a stop recording operation (such as a user operation such as clicking to stop video recording) on the shooting control 122B, the terminal device 100 can stop video recording in response to the stop recording operation. At the moment of stopping recording, the video frame 1233 displayed in the preview window 123 can be the last frame of the recorded video.
[0143] After stopping recording, the interface of the terminal device 100 may change to the following: Figure 2G The interface shown, that is, the user interface can be changed back to Figure 2C Style shown. Figure 2G The user interface shown is similar to Figure 2C The difference between the user interface of the present invention and the present invention is that the thumbnail of the wonderful photo can be displayed in the review control 124, or the thumbnail of the video cover of the recorded video can be displayed, and there is no specific limitation on this. As for the video cover of the recorded video, in some embodiments, in order to facilitate the user to intuitively understand the video content and quickly capture the theme of the video, the recorded video cover can be a wonderful photo. In other embodiments, the video cover of the recorded video can also be the first video frame of the recorded video.
[0144] And, when the user Figure 2G After the user interface shown in the figure performs a trigger operation (such as a user operation such as clicking) on the review control 124, the terminal device 100 can respond to the trigger operation and display the Figure 2H The user interface shown (i.e., the shooting content display interface). Figure 2H The shooting content display interface shown may include a recorded video file 131 and a gallery control 126 .
[0145] In some embodiments, Figure 2H As shown, the terminal device 100 can display the recorded video file 131 according to the viewing needs of the user (the recorded video file 131 can be a video stored on the terminal device and browsed or viewed by the user from the terminal device 100. Exemplarily, the recorded video file 131 can be obtained based on the original video stream, and the resolution (or clarity) of the recorded video file 131 is higher than or equal to the resolution (or clarity) of the original video stream). Optionally, in order to facilitate the user to quickly capture the video shooting subject, the wonderful photos (or the video frames corresponding to the wonderful photos in the recorded video file 131) can be used as the cover of the recorded video file 131. Alternatively, the first frame image of the recorded video file 131 can be used as the cover of the recorded video file 131. Optionally, in order to facilitate user viewing, the display of the recorded video file 131 and the wonderful photos 132 can be switched by sliding left and right.
[0146] In other embodiments, Fig.2I As shown, the video file 131 and the wonderful photos 132 obtained by shooting can be displayed in the shooting content display area. It should be noted that Fig.2I Only one display style of the recorded video file 131 and the wonderful photos 132 is exemplarily shown. In the embodiment of the present application, the recorded video file 131 and the wonderful photos 132 can be displayed in other styles, which is not limited.
[0147] And, continue to see Figure 2H , when the user is Figure 2H After the user interface shown in the figure performs a play operation (such as a user operation such as a click) on the video file 131, the terminal device 100 can respond to the trigger operation and display the video file 131. Figure 2J The user interface shown is used to play the recorded video file 131. Specifically, Figure 2J The user interface shown may include a play progress bar 127 and a wonderful content display area 128. The play progress bar 127 is used to display the play progress of the video file 131, and the wonderful content display area 128 may include a video display frame 1281 and a wonderful photo display frame 1282. The video display frame 1281 can synchronously display the real-time image of the video file, and the wonderful photo display frame 1282 can display the wonderful photos of the video file 131. And after the video file 131 is played, the terminal device 100 can display Fig.2I or Figure 2H The user interface shown in FIG. 1 (or the user interface shown in FIG. J , wherein, after the playback is completed, the video frames corresponding to the wonderful photos in the video file 131 or the last frame of the video file 131 are displayed at the display position of the video file 131).
[0148] And, when the user Fig.2I or Figure 2H The user interface shown performs a trigger operation (such as a user operation such as a click) on the gallery control 125, or Figure 2A After a trigger operation (such as a click or other user operation) is performed on the gallery icon 113B on the interface shown, a Figure 2K The user interface (gallery display interface) shown in FIG. Figure 2K In the gallery display interface shown, users can view stored videos and photos. Figure 2KIn the gallery display interface shown, the thumbnail 141a of the recorded video file 131 and the thumbnail 142a of the wonderful photo 132 can be displayed accordingly, wherein the display styles of the thumbnail 141a of the recorded video file 131 and the thumbnail 142a of the wonderful photo 1322 can be different from the thumbnails of other videos and photos, or can be the same, and there is no limitation on this. Alternatively, the thumbnail 141a of the recorded video file 131 and the thumbnail 142a of the wonderful photo 132 can be displayed in multiple albums, and there is no limitation on this. It should be noted that in this embodiment and in subsequent embodiments, in addition to the above-mentioned opening of the gallery, when the terminal device 100 runs a specific application, the camera application can be opened by clicking the gallery control in the specific application, such as opening the camera application by clicking the gallery control in the instant messaging tool, and there is no specific limitation on this.
[0149] Optionally, the "wonderful shooting" mode can also be set in the "more" mode. Accordingly, the terminal device 100 can also detect the user's action on Figure 2B After switching to the "more" mode, the terminal device 100 displays the following Figure 2L The user interface shown in Figure 1 is as follows. Figure 2L The interface shown in FIG. 1 shows multiple shooting mode controls, such as a wonderful capture control 129 (i.e., a control for controlling the terminal device 100 to enter the "wonderful capture" mode). When the terminal device 100 detects that the user has performed a trigger operation (such as a click or other user operation) of the wonderful capture control 129, the terminal device 100 may display the following: Figure 2C In addition, the other interfaces of the wonderful capture through the wonderful capture control in the "More" mode are similar to the interface of the wonderful capture directly through the "Great Capture" mode, which can be seen as follows Figure 2C-2K The relevant description will not be repeated here.
[0150] It should be noted that Figure 2A-2L The process of triggering corresponding controls on the user interface to perform user operations is shown. Users can also Figure 2A-2L , and the user operations on the subsequent user interfaces are performed by pressing specific physical keys or key combinations, inputting voice, air gestures, etc., which are not specifically limited. And, it should also be noted that other controls can be set on each user interface according to actual scenarios and specific needs, which are not specifically limited.
[0151] In passing Figure 2A-2L After explaining the wonderful capture scene corresponding to the "wonderful capture" mode by way of example, the wonderful capture scene of the "recording" mode will be explained next.
[0152] In other embodiments, the terminal device 100 can capture wonderful photos through the "one-record-multiple-photos" function in the shooting scene of the "record" mode. The "one-record-multiple-photos" mode is a recording mode that can analyze the video stream, automatically identify and generate wonderful photos when recording a video.
[0153] Accordingly, Figure 3A-3D Another set of user interface diagrams of the image processing method provided by the embodiment of the present application is exemplarily shown. First, the user triggers Figure 2A After the camera application icon 113A on the main interface is opened, the terminal device 100 may display the following Figure 2B After the user switches to the "recording" mode through the mode switching operation on the menu bar 121, the terminal device 100 may display the following Figure 3A Specifically, Figure 3A As shown, the video recording interface may include a setting control 141, a shooting control 142, and a replay control 143. After the user detects a trigger operation (such as a user operation such as a click) performed by the user on the setting control 141, the following may be displayed: Figure 3B The video recording settings interface is shown.
[0154] like Figure 3B As shown, the video recording setting interface may include setting controls for multiple video recording functions, such as a switch control 144 for a "record one, get multiple results" function. The switch control 144 is used to control the opening and closing of the "record one, get multiple results" function. For example, when the switch control 144 is in the closed state, the terminal device 100 detects the user's right swipe, click, and other operations on the switch control 144, and can control the switch 144 to be in the open state, and turn on "record one, get multiple results". And, when the switch control 144 is in the open state, the switch 144 is controlled to be in the closed state by swiping left or clicking. It should be noted that the video recording setting interface may also include other setting controls for "record one, get multiple results", such as a setting control for the minimum recording time, a setting control for the video resolution, etc., and there is no specific limitation on this. It should be noted that in the embodiment of the present application, the opening setting of the "record one, get multiple results" mode may also be performed in other ways. For example, when the user is in Figure 2A After clicking the setting icon on the main interface shown, the terminal device 100 can display the camera setting page, and the camera setting page can display a switch control of the "one record, multiple get" function. The user can operate the switch control of the "one record, multiple get" function to turn the "one record, multiple get" mode on and off.
[0155] Continue to see Figure 3BWhen the terminal device 100 detects that the user has turned on the "record one, take multiple shots" function through the switch control 144, the terminal device 100 may return to Figure 3A The video recording interface is shown. Figure 3A After a trigger operation (such as a user operation such as clicking to trigger the recording function) is performed on the shooting control 142 on the recording interface shown, the terminal device 100 may display Figure 3C The one-shot-multiple-recording interface shown.
[0156] like Figure 3C As shown, the one-shot-multiple-recording interface may include a preview window 145, a recording stop control 146, a recording pause control 147, and a photo control 148. Specifically, the user may click the photo control 148 to manually capture a wonderful photo, or, if the user does not click the photo control 148, the terminal device 100 may automatically generate a wonderful photo. Optionally, the one-shot-multiple-recording interface may not include the photo control 148, and there is no specific limitation on this.
[0157] And, continue to see Figure 3C When the terminal device 100 detects that the user clicks the recording stop control 146, the terminal device 100 can end the recording and generate a recording video file 131 and wonderful photos 132. And after the recording is finished, the terminal device 100 can switch to the recording interface. It should be noted that the switched recording interface is different from the Figure 3A The video interface shown in FIG. 1 is similar in style, except that the review control 143 of the switched video interface can display thumbnails of wonderful photos or cover thumbnails of the video file 131, and the difference also includes that the switched video interface is different from the video interface 131. Figure 3A The real-time preview screen displayed in the video recording interface shown is different.
[0158] When the user Figure 3A After clicking the replay control 143 on the video interface shown, the display page of the video file 131 can be displayed (its specific display style is similar to Figure 2H or Fig.2I The styles of the two are similar. When the recording objects are different, the video files 131 displayed by the two are different). After the user plays the video file 131 on the display page of the video file 131, the terminal device 100 can display Figure 3D The video playback page shown in the figure. Figure 2J The user interfaces shown are similar and will not be described in detail.
[0159] Also, it should be noted that how to click to enter the gallery, and how to display the video files and wonderful photos in the gallery can be referred to in the above combination Figure 2H-2KThe relevant instructions are not repeated here.
[0160] It should be noted that in the present application, the terminal device 100 can also enable the "one record, multiple results" mode in other ways, for example, Figure 2L The "more" mode user interface shown in the figure may display a "record one, get multiple" mode, where the user Figure 2B The menu bar 121 of the shooting interface shown in the figure switches the mode. After switching to the "more" mode, the following information can be displayed: Figure 2L After the terminal device 100 detects the user's selection operation (click or other user operation) for the "one record, multiple results" mode, it can display the "one record, multiple results" shooting interface (not shown in the figure, the shooting interface is the same as the above Figure 2C The shooting interface of the “Great Snap” mode shown in the figure is similar and no specific limitation is made to this).
[0161] In passing Figure 3A-3D After exemplifying the wonderful capture scenes in the “video recording” mode, the wonderful capture scenes in the “photographing” mode will be described next.
[0162] In some other embodiments, in the shooting scene of the "photographing" mode, the terminal device 100 can also capture wonderful photos through the "photographing" mode. Figure 4A-4D Provide explanation.
[0163] First, the user triggers Figure 2A After the camera application icon 113A on the main interface is opened, the terminal device 100 may display the following Figure 2B After the user switches to the "shooting" mode through the mode switching operation on the menu bar 121 (if the default "photographing" mode is set after opening the camera application, no mode switching operation is required), the terminal device 100 may display the following Figure 4A The photo taking interface is shown.
[0164] like Figure 4A As shown, a wonderful capture control 151A is displayed on the photo taking interface. After the user performs a trigger operation (such as a click or other user operation) on the wonderful capture control 151A, the terminal device 100 can display Figure 4B Optionally, a wonderful capture control 151A may be directly set on the photo taking interface, or the user may click a setting control on the photo taking interface to enter a setting interface (the setting interface is similar to Figure 3B The difference is that the setting interface entered through the photo taking interface includes various photo taking function options, such as a switch control for the wonderful capture function. The user can turn on the wonderful capture function by operating the switch control of the wonderful capture function on the setting interface.
[0165] Figure 4B The interface shown includes a wonderful capture control 151B and a wonderful capture icon 152. Among them, the wonderful capture control 151B and the wonderful capture control 151A are different in style to distinguish whether the wonderful capture function is turned on. Specifically, when the wonderful capture control 151A is displayed on the shooting interface, it means that the wonderful capture function is not turned on, and the photo is taken in the normal shooting mode. And, when the wonderful capture control 151B is displayed on the shooting interface, it means that the wonderful capture function is turned on, and a wonderful capture is taken at this time. Optionally, when the user uses the wonderful capture function for the first time, a guide frame 153 (or prompt frame) can also be displayed. The guide frame 153 is used to prompt the user with the following information: the wonderful capture is turned on.
[0166] When the terminal device 100 displays Figure 4B When the user interface shown in the figure is displayed, the terminal device 100 can automatically take a photo of a wonderful photo. For example, after clicking the wonderful snapshot control 151A, the terminal device 100 takes a wonderful snapshot of the pony's jumping action (jumping-flying-landing).
[0167] For example, when the pony starts jumping, the terminal device 100 collects Figure 4B In the picture 161, the terminal device 100 does not take pictures at this time, and after a period of time, when the pony flies to the highest point, the terminal device 100 collects Figure 4C In the screen 162 shown in FIG. 160 , the terminal device 100 determines that a wonderful video frame has been captured, and can automatically capture and generate a wonderful photo. The terminal device 100 ends the wonderful shooting and displays Figure 4D After that, the pony starts to land. At this time, the terminal device 100 is not Figure 4D Alternatively, continue to refer to Figure 4D After the terminal device 100 captures a wonderful photo, the wonderful photo can be displayed in the review control 154. And after the user opens the gallery, the terminal device 100 can display the wonderful photo in the gallery display interface. Among them, the opening method of the gallery can refer to the relevant description of the above part of the embodiment of the present application, which will not be repeated here.
[0168] And, it should be noted that in the present application, the terminal device 100 can also enable the "wonderful capture" mode in other ways, for example, Figure 2L The "More" mode user interface shown in FIG. 1 may display a "Great Snapshot" mode. On the "More" mode user interface, if the terminal device 100 detects a user selection operation (click or other user operation) for the "Great Snapshot" mode, a "Great Snapshot" shooting interface may be displayed (not shown in the figure, and its shooting interface is similar to the above-mentioned "Great Snapshot" mode). Figure 4B and Figure 4CThe shooting interface of the "wonderful capture" mode shown in FIG. 1 is similar to that of the "wonderful capture" mode shown in FIG. 1 , and no specific limitation is made to this). Alternatively, you can also Figure 2B The menu bar 121 shown in FIG. 1 is set to a "wonderful capture" mode to automatically capture the subject's action when the subject's action is completed to the highest degree. Figure 2A-2L The difference of the wonderful snapshot shown is that the wonderful snapshot in this embodiment does not perform video recording, but is only used for photo shooting.
[0169] It can be seen from the above content that wonderful capture can be performed through recording or shooting mode. Among them, in the recording mode, it involves the processing of video streams and video frames in the video streams, and in the shooting mode, it involves the processing of image streams and image frames in the image streams. For the sake of convenience, the video stream is used as an example for explanation. However, it should be understood that the image processing method provided in the embodiment of the present application can also be adapted to the processing of image streams.
[0170] After introducing feasible ways to take wonderful photos through the above-mentioned multiple shooting modes, the image processing system involved in the embodiment of the present application is briefly described.
[0171] Figure 5 FIG. 1 shows a system architecture diagram of an image processing system provided by an embodiment of the present application. Figure 5 As shown, the image processing system may include a terminal device 100 and a server 200. The terminal device 100 and the server 200 may establish a direct or indirect communication connection through wired or wireless communication.
[0172] The terminal device 100 may be a mobile phone, a tablet computer, a desktop computer, a laptop computer, a handheld computer, a notebook computer, an ultra mobile personal computer (UMPC), a netbook, a cellular phone, a personal digital assistant (PDA), an augmented reality (AR) device, a virtual reality (VR) device, an artificial intelligence (AI) device, a wearable device, a vehicle-mounted device, a smart home device and / or a smart city device, and the present application embodiment does not impose any special restrictions on this. The user interface of the terminal device 100 may refer to the above-mentioned part of the present application embodiment in combination with the above-mentioned part. Figures 2A-4D The relevant description is omitted here.
[0173] The server 200 may be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. It should be noted that in actual applications, the number of servers 200 may be one or more. Figure 5 The number of terminal devices 100 and servers 200 in the application scenario shown is only an adaptive example and is not limited in this application.
[0174] In the embodiments of the present application, Figure 6 A schematic diagram of a process flow of an image processing method provided by an embodiment of the present application is shown. Figure 6 , the server 200 can distill the multimodal large model to obtain a lightweight model, and send the lightweight model to the terminal device 100 (for example, by sending the lightweight model by sending model parameters). And, the server 200 can also use the multimodal large model to extract features of the preset action description statement (i.e., a description statement in a fixed text format for describing the preset action), and send the extracted text features of the preset action to the terminal device 100. The terminal device 100 can shoot the action of the photographed object P1, and obtain a video stream to be processed after processing. The terminal device 100 extracts image features of multiple video frames to be processed in the video stream to be processed based on the lightweight model to obtain image features of multiple video frames. The terminal device 100 performs similarity calculation based on the extracted image features and the text features of the preset action. And, according to the similarity calculation result, it is determined whether there is a preset action in each video frame to be processed. For the video frames to be processed with the preset action, the wonderful frame recognition is performed according to the preset wonderful frame selection logic, the wonderful video frame is selected based on the wonderful frame to be recognized, and the wonderful photos corresponding to the wonderful video frame are generated. Optionally, if there is no preset action in the video frames to be processed, the video processing is terminated.
[0175] It should be noted that, in the embodiment of the present application, the multimodal information can be used to improve the recognition effect of the action category. Also, by combining the action category information to recognize the wonderful frame, the function of capturing the wonderful action moment can be realized.
[0176] After introducing the image processing system, the hardware structure of the terminal device involved in this application is described next.
[0177] Figure 7 Schematic diagram of the structure of the electronic device 300 is shown. For example, Figure 7 The structure of the electronic device 300 in the embodiment can be applied to the terminal device 100 shown in the above embodiment. It should be understood that Figure 7 The electronic device 300 shown is only one example of an electronic device, and the electronic device 300 may have more or fewer components than shown in the figure, may combine two or more components, or may have a different configuration of components. Figure 7 The various components shown in the EMBODIMENTS 2000 may be implemented in hardware, software, or a combination of hardware and software, including one or more signal processing and / or application specific integrated circuits.
[0178] The electronic device 300 may include: a processor 310, an external memory interface 320, an internal memory 321, a Universal Serial Bus (USB) interface 330, a charging management module 340, a power management module 341, a battery 342, an antenna 1, an antenna 2, a mobile communication module 350, a wireless communication module 360, an audio module 370, a speaker 370A, a receiver 370B, a microphone 370C, an earphone interface 370D, a sensor module 380, a button 390, a motor 391, an indicator 392, a camera 393, a display screen 394, and a Subscriber Identification Module (SIM) card interface 395, etc. The sensor module 380 may include a pressure sensor 380A, a gyroscope sensor 380B, an air pressure sensor 380C, a magnetic sensor 380D, an acceleration sensor 380E, a distance sensor 380F, a proximity light sensor 380G, a fingerprint sensor 380H, a temperature sensor 380J, a touch sensor 380K, an ambient light sensor 380L, a bone conduction sensor 380M, etc.
[0179] The processor 310 may include one or more processing units, for example, the processor 310 may include an application processor (AP), a modem processor, a graphics processor (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Different processing units may be independent devices or integrated into one or more processors.
[0180] The controller may be the nerve center and command center of the electronic device 300. The controller may generate an operation control signal according to the instruction operation code and the timing signal to complete the control of fetching and executing instructions.
[0181] A memory may also be provided in the processor 310 for storing instructions and data. In some embodiments, the memory in the processor 310 is a high-speed cache memory. The USB interface 330 is an interface that complies with USB standard specifications, and may specifically be a Mini USB interface, a Micro USB interface, a USB Type C interface, etc. The USB interface 330 may be used to connect a charger to charge the electronic device 300, or may be used to transfer data between the electronic device 300 and a peripheral device. It may also be used to connect headphones to play audio through the headphones. The interface may also be used to connect other electronic devices, such as AR devices, etc.
[0182] The charging management module 340 is used to receive charging input from a charger. The charger may be a wireless charger or a wired charger. In some wired charging embodiments, the charging management module 340 may receive charging input from a wired charger through the USB interface 330. In some wireless charging embodiments, the charging management module 340 may receive wireless charging input through a wireless charging coil of the electronic device 300. While the charging management module 340 is charging the battery 342, it may also power the electronic device through the power management module 341.
[0183] The power management module 341 is used to connect the battery 342, the charging management module 340 and the processor 310. The power management module 341 receives input from the battery 342 and / or the charging management module 340, and supplies power to the processor 310, the internal memory 321, the external memory, the display screen 394, the camera 393, and the wireless communication module 360. The wireless communication function of the electronic device 300 can be implemented through the antenna 1, the antenna 2, the mobile communication module 350, the wireless communication module 360, the modem processor and the baseband processor.
[0184] Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in the electronic device 300 can be used to cover a single or multiple communication frequency bands. Different antennas can also be reused to improve the utilization of the antennas. For example, antenna 1 can be reused as a diversity antenna for a wireless local area network. In some other embodiments, the antenna can be used in combination with a tuning switch.
[0185] The mobile communication module 350 can provide solutions for wireless communication including 2G / 3G / 4G / 5G etc. applied to the electronic device 300. The mobile communication module 350 can include at least one filter, switch, power amplifier, low noise amplifier (Low Noise Amplifier, LNA) etc. The wireless communication module 360 can provide solutions for wireless communication including wireless local area network (WLAN) (such as wireless fidelity (Wireless Fidelity, Wi-Fi) network), Bluetooth (Bluetooth, BT) (), Global Navigation Satellite System (Global Navigation Satellite System, GNSS), Frequency Modulation (Frequency Modulation, FM), Near Field Communication (NFC), Infrared (IR) () applied to the electronic device 300. In some embodiments, the antenna 1 of the electronic device 300 is coupled to the mobile communication module 350, and the antenna 2 is coupled to the wireless communication module 360, so that the electronic device 300 can communicate with the network and other devices through wireless communication technology.
[0186] In an embodiment of the present application, the server 200 may send model parameters of the lightweight model and text features of a preset action description statement to the electronic device 300 via the mobile communication module 350 and / or the wireless communication module 360 .
[0187] The electronic device 300 implements the display function through a GPU, a display screen 394, and an application processor. The GPU is a microprocessor for image processing, which connects the display screen 394 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 310 may include one or more GPUs, which execute program instructions to generate or change display information.
[0188] The display screen 394 is used to display images, videos, etc. The display screen 394 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), etc. In some embodiments, the electronic device 300 may include 1 or N display screens 394, where N is a positive integer greater than 1.
[0189] In the embodiment of the present application, the electronic device 300 can display the display capability provided by the display screen 194. Figure 2A-2L , Figure 3A-3D , Figure 4A-4DThe user interface shown includes image resources such as photos and videos displayed in the interface.
[0190] The electronic device 300 can realize the shooting function through ISP, camera 393, video codec, GPU, display screen 394 and application processor.
[0191] The ISP is used to process the data fed back by the camera 393. For example, when taking a photo, the shutter is opened, and the light is transmitted to the camera photosensitive element through the lens. The light signal is converted into an electrical signal, and the camera photosensitive element transmits the electrical signal to the ISP for processing and converts it into an image visible to the naked eye. The ISP can also perform algorithm optimization on the noise, brightness, and skin color of the image. The ISP can also optimize the exposure, color temperature and other parameters of the shooting scene. In some embodiments, the ISP can be set in the camera 393.
[0192] The camera 393 is used to capture still images or videos. The object generates an optical image through the lens and projects it onto the photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal oxide semiconductor (CMOS) phototransistor. The photosensitive element converts the optical signal into an electrical signal, and then transmits the electrical signal to the ISP to be converted into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in a standard RGB, YUV or other format. In some embodiments, the electronic device 300 may include 1 or N cameras 393, where N is a positive integer greater than 1.
[0193] The digital signal processor is used to process digital signals, and can process not only digital image signals but also other digital signals. For example, when the electronic device 300 is selecting a frequency point, the digital signal processor is used to perform Fourier transform on the frequency point energy.
[0194] The video codec is used to compress or decompress digital video. The electronic device 300 may support one or more video codecs. Thus, the electronic device 300 may play or record videos in a variety of coding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, MPEG4, etc.
[0195] In the embodiment of the present application, the electronic device 300 acquires images through the capabilities provided by the camera 193, ISP, and digital signal processor to obtain the original video stream. When shooting a video, the electronic device 300 can generate a video in a specific format through a video codec. The electronic device 300 can perform image feature extraction, action recognition, and wonderful frame selection on the video frame through the intelligent cognitive capabilities provided by the NPU.
[0196] The external memory interface 320 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 300. The external memory card communicates with the processor 310 through the external memory interface 320 to implement a data storage function, such as storing music, video and other files in the external memory card.
[0197] The internal memory 321 can be used to store computer executable program codes, which include instructions. The processor 310 executes various functional applications and data processing of the electronic device 300 by running the instructions stored in the internal memory 321. The internal memory 321 may include a program storage area and a data storage area. Among them, the program storage area may store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc. The data storage area may store data created during the use of the electronic device 300 (such as audio data, a phone book, etc.), etc. In addition, the internal memory 321 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (Universal Flash Storage, UFS), etc.
[0198] The electronic device 300 can implement audio functions such as music playing and recording through the audio module 370, the speaker 370A, the receiver 370B, the microphone 370C, the headphone jack 370D, and the application processor.
[0199] The audio module 370 is used to convert digital audio information into analog audio signal output, and is also used to convert analog audio input into digital audio signal. The audio module 370 can also be used to encode and decode audio signals. In some embodiments, the audio module 370 can be arranged in the processor 310, or some functional modules of the audio module 370 can be arranged in the processor 310.
[0200] The speaker 370A, also called a "speaker", is used to convert an audio electrical signal into a sound signal. In the embodiment of the present application, when playing a video, the electronic device 300 can play the audio signal in the video through the speaker 370A.
[0201] Microphone 370C, also called "microphone" or "microphone", is used to convert sound signals into electrical signals. In an embodiment of the present application, during video recording, the electronic device 300 can collect sound signals through the microphone 370C. The electronic device 300 can be provided with at least one microphone 370C. In other embodiments, the electronic device 300 can be provided with two microphones 370C, which can not only collect sound signals but also realize noise reduction function. In other embodiments, the electronic device 300 can also be provided with three, four or more microphones 370C to realize the collection of sound signals, noise reduction, identification of sound sources, and realization of directional recording functions, etc.
[0202] The headphone jack 370D is used to connect a wired headphone.
[0203] The pressure sensor 380A is used to sense the pressure signal and can convert the pressure signal into an electrical signal. In some embodiments, the pressure sensor 380A can be set on the display screen 394. There are many types of pressure sensors 380A, such as resistive pressure sensors, inductive pressure sensors, capacitive pressure sensors, etc. The capacitive pressure sensor can be a parallel plate including at least two conductive materials. When a force acts on the pressure sensor 380A, the capacitance between the electrodes changes. The electronic device 300 determines the intensity of the pressure according to the change in capacitance. When a touch operation acts on the display screen 394, the electronic device 300 detects the touch operation intensity according to the pressure sensor 380A. The electronic device 300 can also calculate the touch position according to the detection signal of the pressure sensor 380A. In some embodiments, touch operations acting on the same touch position but with different touch operation intensities can correspond to different operation instructions. For example: when a touch operation with a touch operation intensity less than the first pressure threshold acts on the short message application icon, an instruction to view the short message is executed. When a touch operation with a touch operation intensity greater than or equal to the first pressure threshold acts on the short message application icon, an instruction to create a new short message is executed.
[0204] The gyro sensor 380B can be used to determine the motion posture of the electronic device 300. In some embodiments, the angular velocity of the electronic device 300 around three axes (i.e., x, y, and z axes) can be determined by the gyro sensor 380B. The gyro sensor 380B can be used for anti-shake shooting. For example, when the shutter is pressed, the gyro sensor 380B detects the angle of the electronic device 300 shaking, calculates the distance that the lens module needs to compensate based on the angle, and allows the lens to offset the shaking of the electronic device 300 through reverse movement to achieve anti-shake. The gyro sensor 380B can also be used for navigation and somatosensory game scenes.
[0205] The acceleration sensor 380E can detect the magnitude of the acceleration of the electronic device 300 in all directions (generally three axes). When the electronic device 300 is stationary, the magnitude and direction of gravity can be detected. It can also be used to identify the posture of the electronic device and is applied to applications such as horizontal and vertical screen switching and pedometers.
[0206] The distance sensor 380F is used to measure the distance. The electronic device 300 can measure the distance by infrared or laser. In some embodiments, when shooting a scene, the electronic device 300 can use the distance sensor 380F to measure the distance to achieve fast focusing.
[0207] The fingerprint sensor 180H is used to collect fingerprints. The electronic device 100 can use the collected fingerprint characteristics to implement fingerprint unlocking, access application locks, fingerprint photography, fingerprint call answering, etc.
[0208] The touch sensor 380K is also called a "touch panel". The touch sensor 380K can be set on the display screen 394, and the touch sensor 380K and the display screen 394 form a touch screen, also called a "touch screen". The touch sensor 380K is used to detect touch operations acting on or near it. The touch sensor can pass the detected touch operation to the application processor to determine the type of touch event. Visual output related to the touch operation can be provided through the display screen 394. In other embodiments, the touch sensor 380K can also be set on the surface of the electronic device 300, which is different from the position of the display screen 394.
[0209] In the embodiment of the present application, the electronic device 300 can detect user operations such as clicking and sliding on the display screen 394 through the touch detection capability provided by the touch sensor 380K, thereby controlling the activation and deactivation of applications and devices and realizing jump switching between different user interfaces.
[0210] The key 390 includes a power button, a volume button, etc. The key 390 may be a mechanical key. It may also be a touch key. The electronic device 300 may receive key input and generate key signal input related to user settings and function control of the electronic device 300. Optionally, in the embodiment of the present application, the electronic device 300 may control the shooting of photos or videos through the key 390.
[0211] Motor 391 can generate vibration prompts. Motor 391 can be used for incoming call vibration prompts, and can also be used for touch vibration feedback. For example, touch operations acting on different applications (such as taking pictures, audio playback, etc.) can correspond to different vibration feedback effects. Indicator 392 can be an indicator light, which can be used to indicate charging status, power changes, messages, missed calls, notifications, etc.
[0212] In passing Figure 7After introducing the hardware structure of the terminal device, the software structure of the terminal device will be described next.
[0213] The software system of the electronic device 300 may adopt a layered architecture, an event-driven architecture, a micro-core architecture, a micro-service architecture, or a cloud architecture. The embodiment of the present application takes the Android system of the layered architecture as an example to exemplify the software structure of the electronic device 300.
[0214] Figure 8 It is a software structure block diagram of the electronic device 300 according to an embodiment of the present application.
[0215] The layered architecture of the electronic device 300 divides the software into several layers, each layer has a clear role and division of labor. The layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into four layers, from top to bottom, namely, the application layer, the framework layer, the hardware abstraction layer (HAL) and the system library, and the kernel layer.
[0216] The application layer can include a series of application packages.
[0217] like Figure 8 As shown, the application package may include applications such as camera, gallery, etc.
[0218] The framework layer provides an application programming interface (API) and a programming framework for the applications in the application layer. The application framework layer includes some predefined functions.
[0219] like Figure 8 As shown, the application framework layer may include a window manager, a content provider, a view system, a resource manager, a notification manager, etc. In the embodiment of the present application, a camera access interface may also be included. The camera access interface is used to provide an application programming interface and a programming framework for camera applications.
[0220] The window manager is used to manage window programs. The window manager can obtain the display screen size, determine whether there is a status bar, lock the screen, capture the screen, etc.
[0221] Content providers are used to store and retrieve data and make it accessible to applications. The data can include videos, images, etc.
[0222] The view system includes visual controls, such as controls for displaying text, controls for displaying images, etc. The view system can be used to build applications. A display interface can be composed of one or more views. For example, a display interface including a text notification icon can include a view for displaying text and a view for displaying images.
[0223] The resource manager provides various resources for applications, such as localized strings, icons, images, layout files, video files, and so on.
[0224] The notification manager enables applications to display notification information in the status bar. It can be used to convey notification-type messages and can disappear automatically after a short stay without user interaction. For example, the notification manager is used to notify download completion, message reminders, etc. The notification manager can also be a notification that appears in the system top status bar in the form of a chart or scroll bar text, such as notifications of applications running in the background, or a notification that appears on the screen in the form of a dialog window. For example, a text message is displayed in the status bar, a prompt sound is emitted, an electronic device vibrates, an indicator light flashes, etc.
[0225] The hardware abstraction layer is an interface layer between the application framework layer and the kernel layer, which is used to provide a virtual hardware platform for the operating system. The hardware abstraction layer may include a camera HAL, an audio HAL, a WiFi HAL, etc. In an embodiment of the present application, the camera HAL includes a video frame processing module, a motion recognition module, a wonderful frame recognition module, etc. It should be noted that the video frame processing module, the motion recognition module, and the wonderful frame recognition module can also be set in other HALs except the camera HAL, without limitation.
[0226] The kernel layer is the layer between hardware and software. The kernel layer includes at least display driver, camera driver, audio driver, and sensor driver. Among them, the camera driver is used to drive the camera sensor to collect images and drive the image signal sensor to pre-process the image. The digital signal processor driver is used to drive the digital signal processor to process images. The image processor driver is used to drive the graphics processor to process images.
[0227] Understandably, Figure 8 The layers in the software structure shown and the components included in each layer do not constitute a specific limitation on the electronic device 300. In other embodiments of the present application, the electronic device 300 may include more or fewer layers than shown, and each layer may include more or fewer components, which is not limited in the present application.
[0228] Exemplarily, the camera application calls the camera access interface in the application framework layer to start the camera service, and when the user inputs the camera control instruction (such as preview, zoom, take pictures, record, capture, etc.), the camera application can send the camera control instruction to the camera hardware abstraction layer (Camera HAL) in the hardware abstraction layer through the camera access interface. The camera hardware abstraction layer can call the camera device driver in the kernel layer according to the received camera control instruction, and the camera device driver drives the camera to capture images, and drives the image sensor (Sensor) to convert the image light signal into an image electrical signal (Bayer image). In addition, the camera device driver drives the ISP to convert the image electrical signal into a raw video stream, and transmits the raw video stream to the camera HAL through the camera device driver.
[0229] The video frame processing module of the camera HAL can determine the video stream to be processed based on the capabilities of the digital signal processor and the image processor and based on the original video stream. The action recognition module can perform action recognition on the video frames in the video stream to be processed to determine whether the photographed object in each video frame has made a preset action. The video frames to be recognized that have made the preset action are transmitted to the wonderful frame recognition module. The wonderful frame recognition module selects the target wonderful frames that meet the preset wonderful frame selection conditions from the wonderful frames to be recognized, and generates wonderful photos based on the target wonderful frames.
[0230] And, after determining the wonderful photos, the camera HAL can report the wonderful photos to the camera application or the gallery application through the camera access interface, and the camera application or the gallery application can display the wonderful photos in the display interface, or save the wonderful photos in the mobile phone.
[0231] After introducing the software structure of the electronic device, the image processing method of the embodiment of the present application is described in conjunction with the flowchart. It should be noted that the methods in the following embodiments can all be implemented by the above-mentioned image processing system. Figure 9-Figure 29 , the specific implementation method of the embodiment of the present application is described in detail. Specifically, the image processing process in the embodiment of the present application can be divided into three parts. The first part is the video frame processing part, the second part is the action recognition part, and the third part is the wonderful frame selection part. Next, each part will be described one by one.
[0232] For the video frame processing part, next combine Figure 9-11 The video frame processing portion is described. Fig. 9 A flowchart of a video frame processing part of an image processing method provided in an embodiment of the present application is shown.
[0233] like Fig. 9 As shown, the video frame processing part may include the following steps S101 to S103.
[0234] S101, the terminal device obtains the original video stream.
[0235] In S101, after the terminal device performs a shooting operation in the target shooting interface (shooting or recording interface related to capturing wonderful photos), the camera can continuously capture images of the object being photographed to obtain multiple image light signals. Then, the image sensor can perform photoelectric conversion on the captured image light signals to obtain a Bayer image. Then, the ISP is used to process the Bayer image to obtain an original video stream.
[0236] The original video stream may include multiple original video frames. Fig.10 FIG. 1 is a flow chart showing an exemplary video frame processing process. Fig.10 As shown, the terminal device may include a video frame processing module 410, an action recognition module 420, a wonderful frame recognition module 430, a first storage module 440, etc. Specifically, the video frame processing module 410 may obtain an original video stream. The original video stream may include k original video frames C1 to Ck. Wherein k is any positive integer greater than or equal to 2.
[0237] In one example, taking a shooting scene in the "wonderful capture" mode as an example, Fig.11 FIG. 1 shows a schematic diagram of an exemplary process of generating an original video stream. Fig.11 As shown in (1), the user Figure 2C After the trigger operation is performed on the shooting control 122A on the user interface shown, the camera application of the terminal device generates a snapshot instruction in response to the trigger operation, and sends the snapshot instruction to the camera HAL through the camera access interface. The camera HAL calls the camera device driver according to the snapshot instruction to drive the camera, image sensor, ISP, etc. through the camera device driver to continuously collect and start generating raw image frames (start generating raw video streams). And when the user Figure 2F After the trigger operation is performed on the shooting control 122B on the user interface shown, the camera application can generate a stop snap command and send it to the camera HAL through the camera access interface, so that the camera HAL responds to the stop snap command, stops the device driver driving the camera, image sensor, and ISP, and stops generating raw image frames (stops generating raw video streams). Among them, between detecting the snap command (starting to generate raw image frames) and detecting the stop snap command (stopping to generate raw video frames), the terminal device acquires a total of k raw video frames.
[0238] In another example, taking the shooting scene in the "video recording" mode as an example, Fig.11 As shown in (2), the user Figure 3AThe user interface shown in FIG. 140 performs a trigger operation on the shooting control 142, and the camera HAL can generate a one-shot multiple shooting instruction. Figure 3C After clicking the recording stop control 146 on the user interface shown, the camera HAL can generate a recording and multiple stop shooting instruction, and in response to the recording and multiple stop shooting instruction, the terminal device stops the continued generation of the original video stream, during which the terminal device obtains a total of k original video frames. It should be noted that the specific generation process of the original video stream based on the "recording" mode is similar to the specific generation process of the original video stream in the above-mentioned "wonderful capture" mode, and can refer to the relevant description of the previous example, which will not be repeated here.
[0239] In another example, taking the shooting scene in the "shooting" mode as an example, Fig.11 As shown in (3), the user Figure 4A After the trigger operation is performed on the wonderful capture control 151A on the user interface shown, the terminal device can obtain the original video stream through the camera, image sensor, ISP and other devices. And after the moment when the action of the photographed object is most completed, if it is detected that the wonderful video frame has been captured, the camera HAL controls the camera device driver to stop driving the camera, image sensor, and ISP, thereby stopping the continued acquisition of the original image frame. During this period, the terminal device obtains a total of k original video frames. Alternatively, the camera HAL can also generate a stop capture instruction in response to the user's trigger operation on the user interface to stop the continued generation of the original video stream.
[0240] After introducing the generation process of the original video stream through S101, S102 will be described next.
[0241] S102, the terminal device performs a downsampling operation on the original video stream to obtain a video stream to be processed. Exemplarily, after each original video frame is extracted, the terminal device may perform a downsampling operation on the extracted original video frame to obtain a video stream to be processed.
[0242] For example, see Fig.10 , each of the original video frames can be downsampled to obtain a video frame to be processed. Accordingly, after obtaining the original video frame C1, it can be downsampled to obtain the video frame to be processed E1;...; after obtaining the original video frame Ck, it can be downsampled to obtain k video frames to be processed En. It should be noted that in the embodiment of the present application, in addition to downsampling, other methods that can reduce the image resolution can also be used to process the original video frame to obtain the video frame to be processed. And the n video frames to be processed are sent to the action recognition module 420.
[0243] S103: The terminal device extracts frames from the video stream to be processed, and obtains video frames to be processed.
[0244] In S102, the terminal device may extract a frame of video frame to be processed for subsequent calculation every preset number of video frames. The preset number a may be set according to the actual image processing situation and specific image processing requirements, such as 3. For example, it may be set to a larger value when the calculation speed is increased; and to a smaller value when the calculation accuracy is increased. For another example, it may be set to a larger value when the subject is moving faster; and to a smaller value when the subject is moving slower, without specific restrictions.
[0245] Exemplarily, in order to ensure real-time computing, the terminal device can perform frame extraction operations in real time. Specifically, after each acquisition of a to-be-processed video frame, the terminal device can determine whether the currently acquired to-be-processed video frame satisfies a preset frame extraction condition. If so, the currently acquired original video frame is extracted to continue subsequent operations. Among them, the preset frame extraction condition can be whether the frame index of the currently acquired to-be-processed video frame satisfies a*x+b, where a is a preset number, b is the frame index of the first extracted video frame, and x is any integer greater than or equal to 0. Alternatively, the preset frame extraction condition can be that the distance from the last extracted to-be-processed video frame is a preset number of video frames. In one example, continue to refer to Fig.10 , a to-be-processed video frame can be extracted every 3 video frames. For example, to-be-processed video frame E3, to-be-processed video frame E6, ..., to-be-processed video frame E3n can be extracted. Accordingly, when the kth to-be-processed video frame is received, a total of n to-be-processed video frames are extracted. Wherein n is any integer greater than or equal to k / 3. For ease of illustration, the following takes the extraction of to-be-processed video frame E3, to-be-processed video frame E6, ..., to-be-processed video frame E3n as an example for illustration. It should be noted that other preset numbers a and other extraction positions can also be selected according to actual conditions, and there is no specific limitation on this.
[0246] Optionally, in this embodiment, in addition to the frame extraction operation, other operations may be performed to reduce the amount of image processing operations. For example, each preset number of original video frames may be averaged and normalized to obtain a video frame to be processed, and no specific limitation is made to this.
[0247] In one embodiment, the difference from the above S102 and S103 is that after the original video stream is obtained through S101, the original video stream can be downsampled to obtain the video stream to be processed, and the video stream to be processed is output to the action recognition module 420.
[0248] In another embodiment, different from the above S102 and S103, after the original video stream is obtained through S101, a frame extraction operation may be performed on the original video stream, and the original video stream after frame extraction may be downsampled to obtain a video stream to be processed.
[0249] And, it should be noted that, continue to refer to Fig.10 In order to facilitate the subsequent generation of wonderful photos, after obtaining the original video stream, the video frame processing module 410 can use one original video stream for the generation of the video stream to be processed, and can also store another original video stream in the first storage module 440 of the terminal device. For example, the first storage module 440 can be a buffer queue, that is, multiple original video frames C1-Ck can be stored in the buffer queue. Optionally, the extracted original video frames D1-Dn can be stored in the first storage module 440 for the subsequent generation of wonderful photos.
[0250] After introducing the video frame processing part through the above content, let's continue with Figure 12-Figure 23 The action recognition part is described.
[0251] It should be noted that, since the action recognition part of the embodiment of the present application is implemented based on a multimodal large model, the multimodal large model involved in the embodiment of the present application will be described below.
[0252] The multimodal large model can be a multimodal model that can combine information of different modes such as text information and image information for training and use. In some embodiments, the multimodal large model in the embodiment of the present application can be a large language model such as a contrastive language-image pre-training (CLIP) model, ActionCLIP (i.e., a multimodal pre-training model for video action recognition), a unified reconstruction model for multimodal image text understanding (UNiversal Image-TExt Representation Learning, UNITER), (Vision-and-Language Transformer, ViLT), a pre-trained language representation model (Bidirectional Encoder representation from Image Transformers, BEiT), etc., without specific restrictions.
[0253] In order to facilitate understanding of the image feature extraction model in the embodiment of the present application, CLIP will be used as an example to illustrate the multimodal large model.
[0254] For example, Fig.12 FIG. 1 shows a schematic diagram of the structure of an exemplary multi-modal large model provided in an embodiment of the present application. Fig.12 As shown, the multimodal large model can be CLIP, which is a cross-modal pre-training model for contrast-based image-text learning. In an example of the present application, the CLIP model can be used by a server. Specifically, the server can obtain CLIP 51 through the clip.load function.
[0255] In the embodiment of the present application, CLIP51 may include an image encoder 511 , a text encoder 512 , a normalization processing module 513 , a similarity calculation module 514 and a classification module 515 .
[0256] Among them, the image encoder 511 can extract image features from the input image data to obtain image features Fv (i.e., image encoding results). Exemplarily, the image encoder 511 can be a residual neural network (ResNet) or a visual Transformer (Vision Transformer, Vit) model. Among them, Transformer is an attention-based encoder-decoder architecture. In an example of the present application, the image encoder 511 can first process the image using a preprocessing operation, and then use the model.encode_image function to extract image features from the input image.
[0257] Among them, the text encoder 512 can extract text features from the text data to obtain text features Ft (i.e., text encoding results). Exemplarily, the text encoder 512 can be a continuous bag of words (CBOW) model or a text transformer model. In an example of the present application, the text encoder 512 can first call clip.tokenize to vectorize the text, and then use the model.encode_text function to extract text features from the input text.
[0258] The standardization processing module 513 is used to standardize the image feature Fv and the text feature Ft respectively so that they can be mapped to a common multimodal vector space, so that the standardized image feature Fv and the standardized text feature Ft can be aligned in the vector space. Exemplarily, the standardization processing module 513 can linearly map the image feature Fv and the text feature Ft to the same multimodal embedding representation space (Multi-modal EmbeddingSpace). In an example of the present application, the standardization processing module 513 can first linearly map the image feature Fv and the text feature Ft, and then process them using the l2_normalize function.
[0259] The similarity calculation module 514 is used to calculate the similarity between the two modalities of text and image (vision). Exemplarily, the similarity calculation module 514 can calculate the cosine similarity of the standardized image feature Fv and the standardized text feature Ft to obtain a similarity score. In one example of the present application, the similarity calculation module 514 can perform similarity calculation using the np.dot formula.
[0260] The identification module 515 is used to determine whether the input video frame to be processed is a video frame in which the subject performs a preset action based on the similarity score. In the embodiment of the present application, the wonderful frame to be identified in which the subject performs a preset action can be determined based on the similarity score between the input video frame to be processed and the preset action description sentences of m different preset actions (m input texts).
[0261] For the training process of CLIP51, CLIP51 can use 400 million pairs of image and text data from the Internet as a data set and perform model training through comparative learning. For example, Fig.13 FIG. 2 shows a schematic diagram of an exemplary training process of a multi-modal large model. Fig.13 As shown, p image-text pairs are taken as an example, where in each image-text pair, the text prompt sentence is used as the image label of the corresponding image. Specifically, p images are input into the image encoder 511, and p image features I can be obtained. 1 to I p ; Inputting p text prompt sentences into the text encoder 512, p text features T can be obtained 1 To T p. Based on p image features and p text features, a p*p dimensional similarity matrix 521A can be obtained. Among them, the elements on the diagonal of the similarity matrix 521A are positive samples, and those on the off-diagonal are negative samples. Among them, each element on the similarity matrix 521A is the cosine similarity of the corresponding image feature and text feature (i.e., logits as the output). The cross entropy loss (loss) is calculated using the similarity matrix 521A and the p*p dimensional label matrix (ground truth) 522A, and the CLIP model can be optimized according to the cross entropy loss. Among them, the cross entropy loss of the image and the cross entropy loss of the text can be calculated separately first, and then the average of the cross entropy loss of the image and the cross entropy loss of the text can be used as the cross entropy loss of CLIP51.
[0262] As another example, the multimodal large model in the embodiment of the present application can be Action CLIP. Action CLIP is a new paradigm for video action recognition. Action CLIP can be based on the structural framework of CLIP, and its structure is different in that: first, textual prompts and visual prompts are provided. Second, the label matrix is improved.
[0263] In an example of the present application, for a text prompt, after obtaining the action label of a preset action, the server can use the filling function f fill Obtain a preset action description sentence of the preset action, and input the preset action description sentence into the text encoder of CLIP to obtain text features. Among them, according to the filling position of the action label in the prompt word (prompt) template, fill function f fill It can be divided into the following three categories: prefix prompt, cloze prompt and suffix prompt.
[0264] In one example of the present application, for visual prompts, after obtaining the input image, the server can use the prompt function f tem Process the input image. Prompt function f tem It can be divided into the following three categories: pre-network prompt, in-network prompt and post-network prompt. Among them, the pre-network prompt is used to add position encoding to the image blocks in time and space of the video frame and input them into the image encoder. The in-network prompt is used to insert a time shift module between the feature layers of the image encoder to increase. The post-network prompt is used to use the spatial encoder and the temporal encoder to sequentially encode spatial features and temporal features.
[0265] And, Action CLIP is different in that: in the label matrix of Action CLIP, elements that are not on the diagonal may also be positive samples. Optionally, the loss function can be replaced by KL divergence.
[0266] After introducing the multimodal large model through the above CLIP and Action CLIP, we will continue to explain the action recognition part.
[0267] Fig.14 FIG. 2 is a flow chart showing the action recognition part of another image processing method provided by an embodiment of the present application. Fig.14 As shown, the action recognition part may include the following steps S201 to S205.
[0268] S201, the terminal device extracts image features from a video frame to be processed to obtain image features of the video frame to be processed. Exemplarily, after obtaining each video frame to be processed, the terminal device may extract image features from the video frame to be processed to obtain image features of the video frame to be processed.
[0269] In S201, the terminal device may extract image features from the video frame to be processed based on the image feature extraction model. For example, the current video frame to be processed may be input into the image feature extraction model to obtain the image features of the current video frame to be processed. For example, Fig.15 FIG. 1 is a flow chart showing an exemplary action recognition process provided by an embodiment of the present application. Fig.15 As shown, after obtaining the video frame E3 to be processed, the image feature F can be obtained through the lightweight model 532. V3 After obtaining the video frame E6 to be processed, the image feature F can be obtained through the lightweight model 532 V6 , ..., after obtaining the video frame to be processed E3n, the image feature F can be obtained through the lightweight model 532 V3n .
[0270] In some embodiments, in order to facilitate image feature extraction in the terminal device, the image feature extraction model can be a lightweight model. For example, it can be Fig.15. Exemplarily, the lightweight model 532 may be a lightweight neural network model such as MobileNetV1, MobileNetV2, MobileNetV3, ShuffleNet, ShuffleNetV2, SqueezeNet, Xception, etc., which is applied to the terminal device and can realize the image feature extraction function. The embodiment of the present application does not impose any specific restrictions on it. In this embodiment, by selecting a lightweight model, accurate image feature extraction can be realized in the terminal device, so that fast and accurate action recognition can be realized in the terminal device.
[0271] Next, the structure of the lightweight model is explained.
[0272] In one example, the image feature extraction model can use a MobileNet model. For example, in order to facilitate action recognition on a terminal device, a MobileNetV2 model can be used as an image feature extraction model. Exemplarily, the image data of the video frame to be processed can be input into the MobileNetV2 model, and the image features of the image frame to be processed can be output. The video frame to be processed can be pre-processed and then input into the MobileNetV2 model.
[0273] It should be noted that MobileNetV2 is a lightweight convolutional neural network that can be applied to mobile terminals. Compared with traditional convolutional neural networks, it greatly reduces the amount of calculation and parameters, has higher computing efficiency and smaller model size, can achieve fast and accurate image feature extraction on mobile terminals, and reduces the power consumption of terminal devices during action recognition.
[0274] Next, the image feature extraction model of the embodiment of the present application is described by taking MobileNetV2 as an example. It should be noted that other lightweight models can also be selected to perform image feature extraction on the terminal device, and there is no specific limitation on this.
[0275] In one example, Fig.16 A schematic structural diagram of an exemplary lightweight model provided in an embodiment of the present application is shown.
[0276] For example, Fig.16 As shown, MobileNetV2 adopts an inverted residual structure (InvertedResiduals), which may include the first 1*1 point convolution layer 5311, a 3*3 depth convolution layer 5312 and a 1*1 point convolution layer 5313.
[0277] Among them, after the image data of the video frame to be processed is input into the 1*1 pointwise convolution (also known as pointwise convolution) layer 5311, the input image data is increased in dimension (6 times) by the combination of the 1*1 point convolution layer 5311 and the Relu6 function, so as to expand the channels of the feature map through the 1*1 gradual convolution operation and enrich the number of features. Then, the features are filtered using the combination of the 3*3 deep convolution layer 5312 and the Relu6 function, and then the features are reduced in dimension by the combination of the 1*1 point convolution layer 5313 and the linear activation function. And, for comparison Fig.16 (1) and Fig.16 (2) It can be seen that when the stride of the 3*3 deep convolution layer 5312 is 1 and the image data matrix composed of the input video frame data to be processed and the output image feature matrix have the same shape, a shortcut connection is made between the input and output of MobileNetV2. And, when the stride of the 3*3 deep convolution layer 5312 is 2, no shortcut connection is made between the input and output of MobileNetV2.
[0278] It should be noted that depthwise separable convolution can reduce the number of model parameters and the amount of computation. Specifically, depthwise separable convolution can decompose the traditional convolution operation into two independent operations: depthwise convolution and pointwise convolution. Depthwise convolution can perform convolution operations on the channel dimension, and pointwise convolution can perform convolution operations on the spatial dimension. This decomposition operation can reduce the computational complexity while ensuring the accuracy of image processing.
[0279] Among them, the Relu6 activation function is used after the 1*1 point convolution layer 5311 and the 3*3 depth convolution layer 5312, which can achieve good numerical resolution even when the terminal device has low-precision floating point numbers. The linear activation function is used after the 1*1 point convolution layer 5313, which can avoid the loss of low-dimensional information caused by the Relu6 activation function when the dimension is greatly reduced.
[0280] In another example, the image feature extraction model can use the ShuffleNet model.
[0281] Exemplarily, in the first ShuffleNet model, the difference from the above-mentioned MobileNetV2 model is that the first 1*1 point convolution layer can be replaced by a combination of 1*1 grouped convolution layer + channel shuffle (channelshuffleShuffleNet) operations so that the information between groups can interact. In addition, a batch normalization (BatchNorm, BN) and Relu6 combination operation is performed after the first 1*1 grouped convolution layer. In addition, a BN operation is performed after the 3*3 deep convolution layer and the second 1*1 grouped convolution layer. It should be noted that for other contents of the ShuffleNet model, please refer to the relevant description of the lightweight model in the above part of the embodiment of the present application, which will not be repeated here.
[0282] As another example, in the second ShuffleNet model, the difference from the first ShuffleNet model is that an average pooling layer can be added to the shortcut connection, and the last feature addition operation (element-wise addition, represented by ⊕) is replaced by a channel concatenation operation to increase the output dimension without causing too much computation. It should be noted that the other contents of the ShuffleNet model can refer to the relevant description of the lightweight model in the above part of the embodiment of the present application, which will not be repeated here.
[0283] In another example, the image feature extraction model can use the ShuffleNetV2 model.
[0284] Exemplarily, in the first ShuffleNetV2 model, channel split can be performed to divide the feature map of the video frame to be processed into two branches, one branch will be directly passed backward, and the other branch will complete the convolution operation through 1*1 point convolution + BN + Rule6 operation, 3*3 depth convolution + BN operation, 1*1 point convolution + BN + Rule6 operation, and then the features of the two branches will be concatenated to restore to the size of the input feature, and then the channel dimension concatenation operation will be performed. It should be noted that the other contents of the ShuffleNetV2 model can refer to the relevant description of the lightweight model in the above part of the embodiment of this application, which will not be repeated here.
[0285] As another example, in the second ShuffleNetV2 model, the difference from the first ShuffleNetV2 model is that the channel separation operation is removed, and a 3*3 depth convolution + BN operation and a 1*1 point convolution + BN + Rule6 operation are added to the first branch to complete the convolution operation. It should be noted that for other contents of the ShuffleNetV2 model, please refer to the relevant description of the lightweight model in the above part of the embodiment of the present application, which will not be repeated here.
[0286] After introducing the specific structure of the lightweight model through the MobileNetV2 model and the ShuffleNet model, we will continue to explain how to obtain the lightweight model.
[0287] In one embodiment, Fig.15 As shown, the lightweight model 531 can be obtained by distilling the image encoder 511 of the multimodal large model. The multimodal large module can refer to the above-mentioned part of the embodiment of the present application. Fig.12 and Fig.13 The relevant description of , will not be repeated here. It should be noted that the server can also obtain a lightweight model through model compression methods such as weight quantization, pruning, attention migration, etc., and there is no restriction on this. Alternatively, the lightweight model can also be obtained by training the terminal device using training data, and there is no restriction on this.
[0288] In one example, Fig.17 FIG. 1 shows a schematic diagram of a knowledge distillation process provided by an embodiment of the present application. Fig.17 As shown, the pre-trained image encoder 511 (Teacher Model) can obtain image features of image samples based on the obtained image samples. At this time, the image features output by the pre-trained image encoder 511 can be used as soft labels (Soft targets) for knowledge distillation. In addition, the lightweight module 532 (Student Model) to be trained can obtain the same image samples and obtain the first output result when the distillation temperature T is equal to t. In addition, the lightweight module 532 to be trained can also obtain the second output result when the distillation temperature is 1.
[0289] And, the server can calculate the cross entropy loss between the first output result and the soft label to obtain loss 1 (also known as distillation loss). And, the server can also calculate the cross entropy loss between the second output result and the true label (hard targets) to obtain loss 2 (also known as student model loss). And, the weighted sum of loss 1 and loss 2 is obtained to obtain the total loss.
[0290] And, if Fig.17 As shown by the dotted line in , the server can optimize the lightweight model 532 according to the total loss. For example, the server can fine-tune the model parameters of the lightweight module 532 until the total loss function reaches the convergence condition, thereby obtaining the trained lightweight module 532. Exemplarily, the convergence condition can be a condition for measuring whether the loss function converges. For example, the convergence condition can include: the total loss function converges to a preset value. It should be noted that the convergence condition can also be other conditions that can characterize the convergence of the model, and there is no specific limitation on this.
[0291] It should be noted that, since soft labels contain more knowledge and information than real labels, the soft label training method can further improve the representation ability of the lightweight model 532. In addition, the distillation of the lightweight model using soft labels can avoid overfitting of the lightweight model 532, thereby improving the reliability of the lightweight model 532 training process.
[0292] Also, it should be noted that, since the image encoder is a relatively large model, the lightweight model 532 obtained by distillation is relatively light, so the lightweight model obtained by knowledge distillation can be deployed in the terminal device. Also, through the method of knowledge distillation, the lightweight model 532 can achieve a high image feature extraction accuracy without using a large amount of sample data for training, and realizes zero-sample and few-sample model training. Also, the lightweight model obtained by distillation can be deployed on the mobile terminal, while ensuring the accuracy of feature extraction, and meeting the performance and power consumption requirements of the terminal device.
[0293] In another example, Fig.18 FIG. 2 shows another schematic diagram of knowledge distillation provided by an embodiment of the present application. Fig.18 As shown, in the knowledge distillation of this example, the training data set can be divided into a base class data set and a novel class data set. Then, in an embodiment of the present application, samples can be extracted from the base class data set and the novel class data set to obtain image samples (or, image samples can be selected in other ways, such as obtaining from the network, or shooting images containing preset actions as image samples). Then, the image samples are sent to the pre-trained image encoder 511 (teacher model) after cropping and scaling to obtain a first feature embedding, that is, the feature embedding output by the image encoder 511). And, the same image sample can be input into the lightweight model 532 to be trained to obtain a second feature embedding, that is, the feature embedding output by the lightweight model 532.
[0294] Then, knowledge distillation can be performed based on the first feature embedding and the second feature embedding. Specifically, the image loss (ViLD-image) can be calculated using the first feature embedding and the second feature embedding, and then the lightweight model 532 can be optimized according to the image loss.
[0295] Among them, the L1 loss function can be used to calculate the image loss. Specifically, the image loss Satisfies the following formula (1):
[0296]
[0297] in, represents the first feature embedding, represents the second feature embedding. Optionally, the image samples can be processed using two scaling methods, 1x and 1.5x, and the sum of the feature embedding corresponding to 1x and the feature embedding corresponding to 1.5x can be normalized to obtain
[0298] Optionally, in order to save distillation time, before the lightweight model 532 is distilled, the computing device can use the trained image encoder 511 to pre-acquire the first image embedding, and send the first image embedding to the server in the embodiment of the present application. Thus, when the lightweight model 532 is distilled, the server can directly use the prepared first image embedding for knowledge distillation, saving the time of knowledge distillation and reducing the amount of computing of the server.
[0299] In passing Fig.17 and Fig.18 After introducing the distillation process of lightweight models, we will now combine Fig.15 and Fig.19 The implementation method of S201 is described.
[0300] For example, Fig.19 FIG. 2 is a flow chart showing another exemplary action recognition process provided by an embodiment of the present application. Fig.19 As shown, the action recognition module 420 of the terminal device may include an image feature extraction module 421 , a text feature acquisition module 422 , a standardization processing module 423 , a similarity calculation module 424 and a judgment module 425 .
[0301] In order to enable the terminal device to use the lightweight model to extract image features, after the server uses the image encoder 511 to distill to obtain the lightweight model 532, the model parameters of the lightweight model 532 can be sent to the terminal device. After receiving the model parameters, the terminal device can store them in the second storage module 450. Among them, the second storage module 450 can be a hardware result, functional module, etc. that can store data, such as a memory, storage space, storage queue, etc., and there is no specific limitation on this.
[0302] When image feature extraction is required for the video frame to be processed, the image feature extraction module 421 can obtain the model parameters of the lightweight module 532 from the second storage module 450, and generate the lightweight model 532 locally on the terminal device based on the model parameters. Fig.15 After obtaining the first video frame E3 to be processed, the image feature Fv3 can be obtained through the lightweight model 532; ...; after obtaining the 3nth video frame E3n to be processed, it is input into the lightweight module 532, and the image feature Fv3n can be obtained.
[0303] Alternatively, if Fig.19 As shown, in order to improve the recognition accuracy of the wonderful frame, the action recognition module 420 may further include a feature smoothing module 426. Specifically, after the image features are obtained, the feature smoothing module 426 may be used to smooth the obtained image features. Fig. 20 A schematic diagram of an exemplary smoothing process provided in an embodiment of the present application is shown.
[0304] like Fig. 20 As shown, after obtaining the image feature Fv 12 After (smoothing image features), the image feature Fv can be determined 12 The corresponding current window, where the window size (windowsize) can be set to r1, and accordingly, the current window with the window size of r1 can be selected to include the image feature Fv 12 Then, the image features in the current window are averaged and normalized to obtain the smoothed image feature Fv 12′ Then the smoothed image feature Fv 12′ It occurs to the standardization processing module so that it can be standardized and then similarity calculated. It should be noted that through smoothing, outliers in image features can be eliminated, the accuracy of image features can be improved, and the accuracy of subsequent action recognition and the accuracy of wonderful frame recognition can be improved. Among them, r1 can be any positive integer, which can be set to other values according to actual calculation conditions and calculation requirements, and there is no specific limitation on this.
[0305] In one example, if Fig. 20 As shown, the window size r1 can be set to 3. Accordingly, the feature smoothing module 426 can be based on the image feature Fv 6 , image feature Fv 9 , image feature Fv 12 Perform averaging and normalization to obtain the smoothed Fv 12′ . In this example, when determining the window corresponding to a certain image feature to be smoothed, the image feature to be smoothed can be located at the rightmost end of the window, so that after obtaining the image feature to be smoothed, the feature smoothing module 426 can perform real-time calculations based on other image features obtained before the image feature to be smoothed, thereby improving the real-time performance and recognition efficiency of wonderful frame recognition. In addition, this example can also effectively solve the problem of calculation boundaries and reduce the error probability of image processing. Optionally, the image feature to be smoothed can be located at other positions of the corresponding window, such as the leftmost, middle, etc., according to actual calculation conditions and calculation requirements, and there is no specific limitation on this.
[0306] Optionally, in order to improve the reliability of calculation, the feature smoothing module 426 may determine whether the image feature to be smoothed satisfies a preset smoothing condition after obtaining the image feature to be smoothed. The preset smoothing condition may be a condition for determining whether the number of existing image features can be smoothed. In one example, the preset smoothing condition may include: whether the video frame index corresponding to the current image feature is greater than or equal to a first preset index, and the first preset index may be the position of the image feature to be smoothed in the first window, such as Fig.19 For example, the first preset index may be the index corresponding to the last image feature of the first window, for example, the first preset index may be 9. That is to say, when the video index corresponding to the acquired image feature is greater than or equal to 9, the acquired image feature may be smoothed and subsequently subjected to action recognition and other operations. If the video index corresponding to the acquired image feature is less than 9, for example, the image feature Fv 3 , image feature Fv 6 , it is not necessary to perform smoothing and subsequent processing on it. Alternatively, the image feature Fv 3 , image feature Fv 6 For image features that do not meet the preset smoothing conditions, the feature smoothing module 426 can directly use the image features for subsequent calculations, or use the image features and image features before the image features for averaging and normalization processing, and continue with subsequent calculations after processing.
[0307] After introducing S201, S202 will be described next.
[0308] S202: The terminal device obtains text features of a preset action.
[0309] In S202 , the text feature acquisition module 422 may acquire text features of m preset actions.
[0310] The text feature of each preset action can be obtained by extracting the text feature of the preset action description sentence of the preset action by a text encoder. The text encoder can be a text encoder of a multimodal large model, such as Fig.15 The text encoder 512 in S202 may be used to uniformly express visual and text information across modalities. The text encoder in S202 and the image encoder used to distill the lightweight model may belong to the same multimodal large model, such as the same CLIP model.
[0311] For ease of understanding, the preset actions and preset action description statements are explained below.
[0312] The preset action, i.e., the action that can be action recognized, can be predefined. For example, the preset action can include jumping, running, walking, sitting, etc. It should be noted that other actions can also be preset according to the actual action recognition scene and specific action recognition requirements, and there is no specific limitation on this.
[0313] For the preset action description sentence, it can be generated according to the action classification label (object) and prompt word (prompt) template of the preset action.
[0314] Among them, the action classification label can be used to represent the action category information, which can be used as a specific category identifier of the preset action. Exemplarily, as shown in Table 1, the action classification label can include "jump", "run", "walk", "sit", etc. It should be noted that the action classification label can also be set to other language forms, such as Chinese, according to the actual scenario and specific needs, and there is no limitation on this.
[0315] The prompt word template, that is, the sentence template, can be set to a specific format. For example, as shown in Table 1 below, the prompt word template can be "The video of a{object}". Among them, "{object}" is the filling position of the action classification label. It should be noted that the specific format and content of the prompt word template can also be set to other styles according to the actual action recognition scene and action recognition requirements, and there is no specific limitation on this.
[0316] The preset action description sentence is used to describe the action that appears in the video. By filling in the action classification labels of different preset actions in the prompt word template, the preset action description sentences of different preset actions can be generated. For example, referring to Table 1, the preset action description sentence of the jumping action can be "The video of a jump"; the preset action description sentence of the running action can be "The video of a run", etc.
[0317] Table 1
[0318]
[0319] In some embodiments, the text features of the preset action may be pre-extracted by the server and sent to the terminal device. Fig.15 as well as Fig.18 It can be seen that after obtaining the preset action description sentences G1-Gn of the m preset actions, the server can input the preset action description sentences G1-Gn of the m preset actions into the text encoder 512 to obtain m text features F t1 To F tm Then, the server sends m text features F t1 To F tm The terminal device receives m text features F t1 To F tm Afterwards, it can be stored in the third storage module 460. Among them, the third storage module 460 and the second storage module 450 can be the same storage module or different storage modules. The third storage module 460 is similar to the second storage module 450. Please refer to the relevant description of the second storage module 450 in the above part of the embodiment of the present application, which will not be repeated here.
[0320] And, when action recognition is required, the text feature acquisition module 422 can acquire m text features F from the third storage module 460 t1 To F tm .
[0321] It should be noted that, since the preset action can be pre-set in advance, its content and analogy will not change easily, and the preset actions of different terminal devices are relatively consistent. Accordingly, in this embodiment, the text features can be extracted in advance by the server, and text features with high feature accuracy can be used without arranging a text feature extraction model on the terminal device, thereby saving the computing power of the terminal device.
[0322] Optionally, the few-shot capability of the multimodal large model can be utilized to quickly expand new actions. For example, when a new preset action needs to be added, the server can update the multimodal large model with a small number of image-text pairs of the newly added action.
[0323] For example, Fig.21 Another exemplary multi-modal large model training process diagram is shown. Fig.13 and Fig.21 It can be seen that, taking the newly added action "rolling" as an example, a small number of image-text pairs of rolling actions can be newly added on the basis of the original p image-text pairs (p text prompt sentences 541A and p training images 542A), such as text description sentences 541B ("The video of a roll") and rolling pictures 542B. It should be noted that, in order to simplify the representation, Fig.21 In the example, one image-text pair of the rolling action is used, but it should be understood that the number of image-text pairs of the rolling action can be multiple, such as 10 pairs, 100 pairs, etc., and there is no specific limitation on this.
[0324] And, continue to see Fig.13 and Fig.21 , the text description statement 541B can be input into the text encoder 512 to obtain the text feature T p+1 And the tumbling picture 542B is input into the image encoder 511 to obtain the image feature I p+1 Then, based on the p+1 image features and the p+1 text features, a similarity matrix 521B may be obtained. And after determining a new label matrix 522B, a cross entropy loss may be calculated using the similarity matrix 521B and the label matrix 522B to adjust the CLIP model according to the calculated cross entropy loss.
[0325] And, after adjusting the CLIP model, the lightweight model can be redistilled using the image encoder of the adjusted CLIP model, so as to update the lightweight model using the redistillation method. Then, the model parameters of the re-updated lightweight model are sent to the terminal device, so that the terminal device uses the updated lightweight model to extract image features. It should be noted that in the embodiment of the present application, other methods can also be used to update the lightweight model, such as using the image of the tumbling action to train and adjust the lightweight model, etc., which is not limited to this.
[0326] Furthermore, after the CLIP model is adjusted, the text encoder of the adjusted CLIP model can be used to extract text features of the text description statement of the rolling action, and the extracted text features are sent to the third storage module 460 for storage.
[0327] In some other embodiments, the server may also distill the text encoder 512 to obtain a lightweight model, and then deploy the lightweight model on the terminal device. Accordingly, in S201, the text feature acquisition module 422 may generate a lightweight model obtained by distilling the text encoder 512, and then use the lightweight model to extract text features.
[0328] After introducing S202, S203 will be described next.
[0329] S203, the terminal device performs standardization processing on the image features of the video frame to be processed and the text features of the preset action.
[0330] In S203, continue to refer to Fig.18 , the normalization processing module 423 can align the image features of the video frame to be processed and the text features of the preset action in the multimodal vector space. For example, the normalization processing module 423 can first perform linear mapping on the image features Fv and the text features Ft, and then use the l2_normalize function to process (this processing process is not described in Fig.15 ).
[0331] S204, the terminal device calculates the similarity between the image features of the video frame to be processed and the text features of the preset action.
[0332] In S204, continue to see Fig.18 , the similarity calculation module 424 can calculate the cosine similarity of the standardized image feature Fv and the standardized text feature Ft to obtain a similarity score. Among them, the similarity score is used to measure the similarity between the image feature and the text feature, and the high or low similarity score can indicate the probability that the subject in the video frame to be processed has performed a preset action (i.e., the action corresponding to the preset action). Optionally, in an embodiment of the present application, for the image feature Fv, a smoothing process can be performed before the standardization process, or after the standardization process. Among them, the content of the smoothing process can be referred to in the above part of the embodiment of the present application in combination with Fig. 20 The relevant description will not be repeated here.
[0333] For example, a cosine similarity calculation may be performed on each normalized image feature Fv and each normalized text feature Ft to obtain a similarity score. V3 To F V3n , and m text features F t1 To F tm , m*n similarity scores can be calculated.
[0334] In a specific example, see Fig.15 , we can use the i-th image feature F Vi and the jth text feature F tj Perform similarity calculation and obtain the similarity score S ij . Similarity score S ij It can represent the probability of the jth preset action appearing in the i-th image to be processed. Wherein, i can be an integer multiple of 3 and less than or equal to n, and j can be any positive integer less than or equal to m.
[0335] S205: The terminal device determines whether the subject in the video frame to be processed has performed the preset action based on the similarity between the image features of the video frame to be processed and the text features of the preset action.
[0336] For example, see Fig.15 For the i-th image frame to be processed Ei, the decision module 425 may calculate the similarity score S between the image feature Fvi of the image frame to be processed and the text features of the m preset actions. 1i -S mi In the example, the maximum score Smaxi corresponding to the image frame to be processed is determined. Fig. 22 FIG. 2 shows an exemplary schematic diagram of determining the maximum score provided in an embodiment of the present application. Fig. 22 As shown, for the image features of the video frame to be processed, the similarity between it and the text features of the jumping action is 80%, the similarity between it and the text features of the running action is 36%, the similarity between it and the text features of the walking action is 36%, and the similarity between it and the text features of the sitting action is 8%. Since 86% is the maximum value among the four scores, the maximum score corresponding to the video frame to be processed is 86%. Among them, the maximum score is used to represent the probability that the subject of the image frame to be processed has made m preset actions. In other words, if the subject performs any one of the m preset actions, theoretically, the maximum score of the corresponding processed image frame is higher.
[0337] Next, after determining the maximum score corresponding to each image frame to be processed, the decision module 425 can determine that the subject in the video frame to be processed has performed the preset action if the maximum score corresponding to each video frame to be processed is greater than or equal to the preset score threshold. Similarly, if the maximum score corresponding to the video frame to be processed is less than the preset score threshold, it is determined that the subject in the video frame to be processed has not performed the preset action (including the case where there is no subject, or there is a subject but the preset action is not performed).
[0338] In one example, Fig.23FIG. 1 is a schematic diagram showing an exemplary action recognition process provided by an embodiment of the present application. Fig.23 As shown, after the image features of the five to-be-processed video frames E1-E5 are extracted using a lightweight model, the similarity calculation module can calculate the similarity between the image features of each to-be-processed video frame and the text features of m preset actions. And, the judgment module can determine the maximum score corresponding to each to-be-processed feature frame according to the similarity between the image features of each to-be-processed video frame and the text features of m preset actions, such as 51%, 86%, 97%, 89%, and 64% respectively. If the preset score threshold is 75%, since the maximum score of the to-be-processed image frames E2-E4 exceeds 75%, it can be determined that the subject in the to-be-processed image frames E2-E4 has performed the preset action (i.e., there is a random action in the to-be-processed video). And the image features of the to-be-processed image frames are sent to the wonderful frame recognition module to continue the wonderful frame recognition module.
[0339] It should be noted that in the embodiment of the present application, since the lightweight model is obtained by distilling the multimodal large model, it can reduce the amount of model parameters while ensuring the feature extraction accuracy of the multimodal large model, thereby enabling accurate image feature extraction on the terminal device side.
[0340] After introducing the action recognition part in the above section, the wonderful frame recognition part will be explained next.
[0341] For the wonderful frame recognition part, we will combine Figure 24-Figure 29 The video frame processing portion is described. Fig.24 A flow chart of a wonderful frame recognition part of an image processing method provided in an embodiment of the present application is shown.
[0342] like Fig.24 As shown, the wonderful frame identification part may include steps S301-S305.
[0343] S301, when the subject in the video frame to be processed performs a preset action, the terminal device determines a plurality of wonderful frames to be identified in the video stream to be processed. The plurality of wonderful frames to be identified include the video frame to be processed and adjacent video frames to be processed of the video frame to be processed. For example, in order to facilitate real-time calculation, the video frame to be processed and a plurality of video frames to be processed before the video frame to be processed can be selected as video frames to be identified. Exemplarily, the wonderful frame identification module 430 can include a video frame selection module, which can determine a plurality of wonderful frames to be identified in the video stream to be processed.
[0344] In some embodiments, the terminal device may determine a plurality of to-be-identified wonderful frames in the to-be-processed video stream when the maximum score corresponding to the to-be-processed video frame is greater than or equal to a preset score threshold.
[0345] For example, Fig.25 FIG. 2 is a flow chart showing an exemplary wonderful frame recognition process provided by an embodiment of the present application. Fig.25 As shown, when the subject in the to-be-processed video frame with index 12 (hereinafter referred to as video frame 12) performs a preset action, the terminal device can determine the current window corresponding to video frame 12. The window size (windowsize) can be set to r2, and accordingly, the current window with window size r2 can select r2 to-be-processed video frames including the to-be-processed video frame. For example, the window size r2 can be set to 7. Accordingly, in Fig.25 In the example, when determining the window corresponding to the video frame to be processed, the video frame to be processed can be located at the rightmost end of the window, so that after obtaining the video frame to be processed, the terminal device can perform real-time wonderful frame recognition based on other video frames to be processed obtained before the video frame to be processed, thereby improving the real-time and recognition efficiency of wonderful frame recognition. In addition, this example can also effectively solve the problem of calculation boundaries and reduce the error probability of image processing. Optionally, the video frame to be processed can be located at other positions of the corresponding window, such as the leftmost, the middle, etc., according to the actual calculation situation and calculation requirements, and no specific restrictions are made on this. It should be noted that the window size r2 can be set to other values according to the actual situation and specific requirements, for example, the window size r2 can be set according to the power consumption of the terminal device. Exemplarily, in order to improve the recognition accuracy, the window size r2 can be set to a value greater than 7 to obtain a larger global field of view. For example, in order to reduce the power consumption of the terminal device or to improve the calculation efficiency, the window size r2 can be set to a value less than 7.
[0346] Optionally, in order to improve the reliability of calculation, the terminal device may determine whether the video frame to be processed satisfies a preset window selection condition after acquiring the video frame to be processed. The preset window selection condition may be a condition for determining that the number of existing video frames to be processed is sufficient for window selection.
[0347] In one example, the preset window selection may include: whether the video frame index corresponding to the to-be-processed video frame is greater than or equal to the second preset index, and the preset index may be the position of the to-be-processed video frame in the first window, such as Fig.25For example, the second preset index may be the index corresponding to the last video frame to be processed of the first window, for example, the second preset index may be 7. That is to say, when the video index corresponding to the acquired video frame to be processed is greater than or equal to 7, window selection and other operations may be performed on the acquired video frame to be processed. If the video index corresponding to the acquired video frame to be processed is less than 7, for example, there is a preset action in video frame 3 (the video frame to be processed with index 3), no subsequent processing may be performed on it. Alternatively, according to actual needs, for the video frame to be processed that does not meet the preset window selection conditions, the terminal device may directly use the video frame to be processed for subsequent calculations, or use the video frame to be processed and the video frame to be processed before the video frame to be processed for subsequent calculations. For example, if there is a preset action in video frame 3, video frames 1-video frames 3 may be determined as video frames to be identified.
[0348] Since similar actions exist in frames near the to-be-processed video frame where the preset action is recognized, through window selection, the frame with the highest degree of excitement can be selected from the video frames near the to-be-processed video frame where the preset action exists for subsequent calculation, thereby further improving the accuracy of exciting frame recognition.
[0349] In other embodiments, since subsequent calculations require the use of image features of the wonderful frames to be identified, the terminal device can use a lightweight model to extract the image features of the video stream to be processed to obtain an image feature sequence. Among them, the image feature sequence can be a sequence composed of image features of each video frame to be processed in the video stream to be processed. Then, the terminal device performs window selection in the image feature sequence to select image features of multiple wonderful frames to be identified. It should be noted that the window selection in the image feature sequence can refer to the window selection in the video stream to be processed in the above embodiment, which will not be repeated here.
[0350] Furthermore, it should be noted that, in the above embodiment, after obtaining the video stream to be processed, the frame extraction operation may not be performed on it. After the image feature sequence of the video stream to be processed is determined by using a lightweight model, the image features of the video stream to be processed can be extracted from the image feature sequence in a frame extraction manner to perform smoothing, action recognition and other related operations. For example, the image features are extracted every three image features in the image feature sequence.
[0351] S302, the terminal device determines the wonderfulness scores of multiple wonderful frames to be identified. Exemplarily, the color frame identification module 430 may include a wonderfulness scoring module, and the wonderful frame scoring module may determine the wonderfulness score of each wonderful frame to be identified.
[0352] As for the wonderfulness score, it is used to measure the wonderfulness of the wonderful frame to be identified. For example, it can be the wonderfulness of the action of the subject in the video frame to be identified. Among them, the wonderfulness of the action can be related to the degree of completion of the action, and the degree of completion of the action can be the jumping height, the degree of stretching of the action, etc., which will not be specifically described. Exemplarily, the wonderfulness score includes a positive score and a negative score. Among them, the positive score is used to positively evaluate the degree of completion of the action in the wonderful frame to be identified, that is, the higher the degree of completion of the action, the higher the positive score. The negative score is used to reversely evaluate the degree of completion of the action in the wonderful frame to be identified. In other words, the higher the degree of completion of the action, the lower the negative score. Alternatively, the wonderfulness score can also include a score, such as a positive score, which is not specifically limited. It should be noted that the wonderfulness score can also be related to the image quality such as the clarity of the picture quality, which will not be specifically described.
[0353] In some embodiments, the terminal device may use a wonderfulness scoring model to determine the wonderfulness score of the wonderful frame to be identified.
[0354] For the excitement scoring model, it can be a model that can calculate the excitement score. Exemplarily, the excitement scoring model may include at least two fully connected layers. In one example, the number of channels of the first fully connected layer is the same as the dimension of the image features of the exciting frame to be identified, so as to be able to align with the image features of the exciting frame to be identified. For example, the number of channels of the first fully connected layer may be 512. The number of channels of the last fully connected layer may be 2 to output a positive class score and a negative class score. Exemplarily, in order to reduce the power consumption of the terminal device, the number of fully connected layers may be 2. Alternatively, in order to improve the accuracy of the excitement score, the number of fully connected layers may be greater than or equal to 3.
[0355] It should be noted that in the embodiments of the present application, other network models or functions may also be used to calculate the wonderfulness score, and there is no specific limitation on this. For example, the image quality score may be determined using an image quality scoring function, and the completion score may be determined using an action completion scoring function or a scoring model. The wonderfulness score is determined based on the image quality score and the completion score. The embodiments of the present application do not specifically limit this.
[0356] In some embodiments, the terminal device can obtain image features of multiple wonderful frames to be identified. Then, the image features of the multiple wonderful frames to be identified are input into the wonderfulness scoring model to obtain an output result. After the output result is activated by an activation function, the wonderfulness scores of the multiple wonderful frames to be identified are obtained. Among them, the activation function can be a sigmoid function. It should be noted that other activation functions can also be selected according to actual conditions and specific needs, and there is no specific limitation on this.
[0357] For example, Fig.26FIG. 2 is a flow chart showing an exemplary wonderful frame recognition process provided by an embodiment of the present application. Fig.26 As shown, when video frames 6 to 12 are selected as the wonderful frames to be identified, the wonderfulness scoring model can be used to determine the wonderfulness scores of each of video frames 6 to 12.
[0358] It should be noted that the image features of the remaining wonderful frames to be identified except the video frame to be processed may be obtained in step S201. Alternatively, the image features of the remaining wonderful frames to be identified may be input into the lightweight model by the terminal device after S301 for determination. The embodiment of the present application does not specifically limit the method for obtaining the image features of the wonderful frames to be identified.
[0359] In some embodiments, the wonderfulness scoring model may be trained using training data. For example, it may be trained by a server or other external device and sent to the terminal device. Alternatively, it may be trained by the terminal device itself.
[0360] For example, Fig. 27 FIG. 2 shows a schematic diagram of training an exemplary wonderfulness scoring model provided in an embodiment of the present application. Fig. 27 As shown, the training data of the wonderfulness scoring model may include image samples and wonderfulness scores corresponding to the image samples. During the training process, the image samples as samples may be input into the wonderfulness scoring model to be trained to obtain an output result. Based on the output result and the wonderfulness scores corresponding to the image samples, a loss function is determined. The wonderfulness scoring model is trained using the loss function, and when the training stop condition is met, a trained wonderfulness scoring model is obtained.
[0361] S303, the terminal device determines the maximum wonderful score among the wonderful scores of the multiple wonderful frames to be identified. Exemplarily, the wonderful frame identification module 420 may include a score processing module. The score processing module may determine the maximum wonderful score. Exemplarily, continue to refer to Fig.26 If video frame 9 has the highest wonderfulness score among video frames 6 to 12, the wonderfulness score of video frame 9, 89%, can be determined as the maximum wonderfulness score.
[0362] S304: When the preset score cache condition is met, the maximum excitement score and the video frame index are cached in the target cache area. The target cache area may be a cache process such as a cache queue. Alternatively, the target cache area may also refer to a cache device, a cache module, or other device or functional module capable of data cache in a terminal device, which is not specifically limited.
[0363] The preset score cache condition refers to the condition that must be met in order to put the maximum excitement score into the target cache area. In one example, the preset score cache condition may include that the maximum excitement score is greater than or equal to a preset excitement score threshold. The preset excitement score threshold can be set according to actual conditions and specific needs, and its specific value is not limited. Fig.26 If the cache score 89% is greater than the preset wonderful score threshold, the score 89% and the frame index 9 can be cached in the target cache area. Alternatively, the wonderful score and the video frame index can be cached in the target cache area without determining whether the maximum wonderful score is greater than or equal to the preset wonderful score threshold.
[0364] In some embodiments, the terminal device is also provided with a data cache flag, which defaults to a "FALSE" state (indicating a state where no data is stored). After the first maximum excitement rating is stored, the data cache flag is set to a "TRUE" state (indicating a state where data is stored).
[0365] Optionally, in order to improve the calculation efficiency, the difference from S301-S303 is that, when the subject in the video frame to be processed performs a preset action, the terminal device can directly calculate the wonderfulness score using the image features of the video frame to be processed. When the wonderfulness score meets the preset score caching condition, the wonderfulness score is cached in the target cache area.
[0366] S305, when the preset wonderful frame identification condition is met, the terminal device determines the maximum value of the maximum wonderfulness score in the buffer area, and the wonderful frame index corresponding to the maximum value. Exemplarily, the wonderful frame identification module 420 may include a wonderful frame determination module. Accordingly, the wonderful frame index may be determined by the wonderful frame determination module.
[0367] The preset wonderful frame recognition condition may refer to a condition that needs to be satisfied in order to perform wonderful frame recognition using the maximum wonderful score in the target buffer.
[0368] In one embodiment, the wonderful frame recognition condition may include that the number of maximum wonderful scores in the target buffer area reaches a preset number threshold. In consideration of factors such as power consumption of the terminal device or wonderful frame recognition rate, the preset number threshold may be set to 4, or may be set to other values according to actual conditions and specific needs, and there is no specific limitation on this.
[0369] In a specific example, Fig.28 FIG. 2 shows a schematic diagram of an exemplary wonderful frame recognition process provided by an embodiment of the present application. Fig.28As shown, for the video stream to be processed, the currently acquired video frame to be processed can be subjected to frame extraction and smoothing and then action recognition can be performed. If the currently acquired video frame to be processed includes a preset action, a plurality of outstanding frames to be identified are determined by selecting frames through a window. Then a maximum outstanding score p1 is determined based on the plurality of outstanding frames to be identified, and is cached. The maximum outstanding scores p2, p3, and p4 are determined by the same method. And, after the score p4 is cached, the terminal device determines that the cache length of the target cache area is equal to 4, then the cached 4 maximum outstanding scores can be compared to determine the maximum value of the 4 maximum outstanding scores (i.e., the number of the maximum outstanding scores cached in the target cache area). If p4 is the largest, the frame index v (outstanding frame index) corresponding to the maximum outstanding score p4 can be determined. And, after determining the outstanding frame index v, the outstanding frame index v can be output and the target cache area can be cleared. Optionally, after clearing the target cache area, the data cache flag can be optionally changed from the "TRUE" state to the "FALSE" state. Among them, the specific content of the data cache flag bit can be found in the relevant description of the above part of the embodiment of this application, and will not be repeated here.
[0370] In other embodiments, the wonderful frame identification condition may include that the difference between the frame index of the currently acquired video frame to be processed and the frame index corresponding to the first maximum wonderful score in the target buffer area is greater than or equal to a third preset index. Exemplarily, the third preset index can be set according to the power consumption of the terminal device, for example, it can be set to 9. Exemplarily, if the frame index corresponding to the first maximum wonderful score in the target buffer area is 9, then after the terminal device receives the frame indexes of the 10th to 17th video frames to be processed, the wonderful frame identification condition will not be met. If the wonderful score of the received video frame to be processed is the maximum wonderful score through S303, (when the preset score cache condition is met) the maximum wonderful score is cached to the target buffer area. When the terminal device receives the 18th video frame to be processed, or a video frame to be processed with a frame index greater than 18, the maximum wonderful score in the target buffer area can be processed in S305, and the target buffer area is cleared.
[0371] Optionally, when the data cache flag is changed from the "FALSE" state to the "TRUE" state, the stored video frame index of the maximum excitement score (the frame index corresponding to the first maximum excitement score) may be recorded.
[0372] In some embodiments, after S305 , S306 (not shown in the figure) may also be included.
[0373] S306, the terminal device determines the wonderful video frame based on the wonderful frame index, and generates a wonderful photo according to the wonderful video frame. For example, in order to meet the user's wonderful frame selection needs, the number of wonderful video frames determined can be 1. It should be noted that the terminal device can also select multiple wonderful video frames according to actual conditions and specific needs. For example, the number of wonderful photos can be 1. Similarly, the terminal device can also generate multiple wonderful photos, which is not limited.
[0374] The wonderful frame identification module 420 of the terminal device can send a wonderful frame index to the wonderful frame decision module. The wonderful frame decision module determines the wonderful video frame in the original video stream based on the wonderful frame index. And, based on the wonderful video frame, generates a wonderful photo. Exemplarily, the wonderful frame decision module can use the wonderful video frame as a wonderful photo. Alternatively, the wonderful frame decision module can obtain a wonderful photo after performing image processing on the wonderful video frame. Among them, image processing may include image processing operations such as video frame fusion. In the embodiment of the present application, the image processing operation can be selected according to the actual situation and specific scene, and its specific type is not limited.
[0375] In some embodiments, the wonderful frame decision module can determine the video frame corresponding to the wonderful frame index as the wonderful video frame. Exemplarily, the terminal device can determine the original video frame in the original video stream based on the wonderful frame index, and determine the original video frame as the wonderful video frame. Then, the original video frame is image processed to generate a wonderful photo.
[0376] In other embodiments, the terminal device can select a wonderful video frame from the video frames corresponding to the wonderful frame index and other candidate video frames determined by other wonderful frame recognition algorithms. Exemplarily, the wonderful frame recognition module 420 of the terminal device can send the wonderful frame index to the wonderful frame decision module. After obtaining the wonderful frame index, the wonderful frame decision module can select a wonderful video frame from the original video frame corresponding to the wonderful frame index and other candidate video frames according to the wonderful frame selection logic. Among them, other wonderful frame recognition algorithms can be algorithms for determining candidate video frames based on image quality scores, portrait scores, etc., and relevant technical personnel can set other wonderful frame algorithms according to actual conditions and specific scenarios, and no specific restrictions are made to this. In addition, relevant technical personnel can set the wonderful frame selection logic according to actual wonderful frame selection requirements and selection scenarios, and no specific restrictions are made to this.
[0377] In some embodiments, in order to avoid false positives, a cooling interval is also provided in the embodiments of the present application.
[0378] In one embodiment, when the wonderful frame identification condition may include that the number of maximum wonderful scores in the target buffer reaches a preset number threshold, the terminal device may determine the frame index of the last maximum wonderful score cached to the target, for example, Fig.28 , the frame index of the last maximum excitement score cached to the target is v. Enter the cooling interval after the vth video frame to be processed (or it can be the last video frame to be processed that has been obtained, or it can be the last exciting frame to be identified that has been scored for excitement), and do not process the video frame after obtaining the video frame to be processed. Until the v+30th (cooling interval length) video frame to be processed is received, continue to process and cache the video frame. It should be noted that the cooling interval length can also be set to other values according to actual conditions and specific scenarios. For example, it can be set according to the power consumption of the terminal device. For another example, the terminal device can determine the movement speed of the captured target. If the movement speed is large, the cooling interval can be set to a smaller value. If the movement speed is slow, the cooling interval can be set to a larger value. It should be noted that the cooling interval can also be set according to other factors, and there is no specific restriction on this.
[0379] It should be noted that in an embodiment of the present application, the terminal device may not perform operations such as image feature extraction, action recognition, and wonderful frame recognition on the video frames to be processed in the cooling interval. Alternatively, the terminal device may not perform score caching operations on the video frames to be processed in the cooling interval. For example, the terminal device may perform image feature extraction and action recognition operations on the video frames to be processed in the cooling interval. When a preset action is recognized in the video to be processed in the cooling interval, multiple wonderful frames to be recognized corresponding to the video frames to be processed can be determined, and after the maximum wonderful score is determined based on the multiple wonderful frames to be recognized, the determined maximum wonderful score is not cached.
[0380] In one example, a wonderful frame flag may be set in the terminal device, which may default to the "FALSE" state (indicating that no wonderful frame has been identified). Also, after determining the frame index v corresponding to the last wonderful score (or after determining the wonderful frame index or clearing the cache), the terminal device switches the wonderful frame flag from the "FALSE" state to the "TRUE" state (indicating that a wonderful frame has been identified), at which point it indicates that the cooling interval has been entered. Also, if the difference between the frame index of the currently acquired video frame to be processed and v is equal to the length of the cooling interval (for example, 30), that is, when the frame index of the currently acquired video frame to be processed is v+30, the terminal device switches the wonderful frame flag from the "TRUE" state to the "FALSE" state, at which point it indicates that the cooling interval has ended.
[0381] Exemplarily, if the terminal device detects that the state of the wonderful frame flag is "FALSE", the maximum wonderful score can be cached. Correspondingly, if the terminal device detects that the state of the wonderful frame flag is "TRUE", the maximum wonderful score is stopped.
[0382] In another exemplary embodiment, after detecting the output wonderful frame index, or detecting that the target buffer area is cleared, or detecting that the data cache flag is changed from the "TRUE" state to the "FALSE" state, the terminal device stops caching the maximum wonderful score (entering the cooling interval). And, when the difference between the frame index of the currently acquired video frame to be processed and v is equal to the cooling interval length, the terminal device continues caching the maximum wonderful score.
[0383] In another embodiment, when the wonderful frame identification condition may include that the difference between the frame index u of the currently acquired video frame to be processed and the frame index corresponding to the first maximum wonderful score in the target buffer is greater than or equal to the third preset index, the relevant content of the cooling interval is similar to the previous embodiment, except that the cooling interval is entered after the wonderful frame index is output (i.e., u video frames to be processed), and the cooling interval is ended after the u+30th (cooling interval length) video frame to be processed is acquired. It should be noted that the other contents of the cooling interval can refer to the relevant description of the previous embodiment, which will not be repeated here.
[0384] For example, Fig.29 FIG. 1 shows an exemplary curve diagram of the change of the action completion degree of the photographed object over time. Fig.29 As shown, when the subject performs action 1, the image processing solution of the embodiment of the present application can be used to identify wonderful frames at the stage when the action completion of action 1 gradually increases, and the video frame corresponding to the wonderful frame index is identified near the highest point of the action completion (that is, the candidate wonderful video frame determined for action 1). Also, in the stage when the action completion of action 1 decreases, it enters the cooling-down interval, and the candidate wonderful video frame is no longer identified, thereby avoiding misidentification at this stage. Also, after the cooling-down interval ends, if the subject continues to perform action 2, the wonderful frame identification of action 2 can be entered. Similarly, the identification of candidate wonderful video frames of action 2 can be completed during the stage when the completion of action 2 increases. The setting of the cooling-down interval avoids repeated misidentification of the same action during the stage when the completion decreases, avoids false positives in identification, and reduces false recall. Also, it does not affect the identification of the next action.
[0385] It should be noted that, in the embodiment of the present application, after the candidate wonderful video frames of action 1 and the candidate wonderful video frames of action 2 are determined, the wonderful frame decision module can determine a wonderful video frame based on the candidate wonderful video frames of action 1 and the candidate wonderful video frames of action 2. Alternatively, the wonderful frame decision module can determine a wonderful video frame corresponding to the candidate wonderful video frame of action 1, and another wonderful video frame corresponding to the candidate wonderful video frame of action 2. It should be noted that the embodiment of the present application does not specifically limit how the wonderful frame decision module selects wonderful video frames.
[0386] It is understandable that, in order to realize the above functions, the electronic device includes hardware and / or software modules corresponding to the execution of each function. In combination with the algorithm steps of each example described in the embodiments disclosed herein, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application in combination with the embodiments, but such implementation should not be considered to be beyond the scope of the present application.
[0387] In one example, Fig.30 A schematic block diagram of a device 400 according to an embodiment of the present application is shown. The device 400 may include: a processor 401 and a transceiver / transceiver pin 402 , and optionally, a memory 403 .
[0388] The components of the device 400 are coupled together via a bus 404, wherein the bus 404 includes a power bus, a control bus, and a status signal bus in addition to a data bus. However, for the sake of clarity, all buses are referred to as bus 404 in the figure.
[0389] Optionally, the memory 403 may be used for the instructions in the aforementioned method embodiment. The processor 401 may be used to execute the instructions in the memory 403, and control the receiving pin to receive a signal, and control the sending pin to send a signal.
[0390] The apparatus 400 may be the electronic device or a chip of the electronic device in the above method embodiment.
[0391] Among them, all relevant contents of each step involved in the above method embodiment can be referred to the functional description of the corresponding functional module, and will not be repeated here.
[0392] The steps performed by the terminal device 100 in the image processing method provided in the above embodiment of the present application may also be performed by a chip system included in the terminal device 100, wherein the chip system may include a processor and a Bluetooth chip. The chip system may be coupled to a memory so that the chip system calls a computer program stored in the memory when it is running to implement the steps performed by the above terminal 100. The processor in the chip system may be an application processor or a processor other than an application processor.
[0393] It should be noted that, as used in the specification and the appended claims of the present application, the singular expressions "a", "a", "said", "above", "the" and "this" are intended to also include plural expressions, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in the present application refers to and includes any or all possible combinations of one or more listed items. As used in the above embodiments, the term "when..." may be interpreted to mean "if..." or "after..." or "in response to determining..." or "in response to detecting...", depending on the context. Similarly, the phrase "when determining..." or "if (stated condition or event) is detected" may be interpreted to mean "if determining..." or "in response to determining..." or "when (stated condition or event) is detected" or "in response to detecting (stated condition or event)", depending on the context.
[0394] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.
Claims
1. An image processing method, It is characterized in that include: In response to a first operation of the user, acquiring a first image stream captured by a camera, wherein the first image stream includes a first image; Based on the text description sentence of the preset action, performing action recognition of the preset action on the first image; When the preset action is recognized in the first image, acquiring at least one second image, wherein the at least one second image and the first image are continuous in the first image stream; Obtaining a wonderfulness score of the first image and a wonderfulness score of each second image in the at least one second image; A first target image is determined based on the wonderfulness score of the first image and the wonderfulness score of each of the second images.
2. The image processing method according to claim 1, It is characterized in that The first image stream also includes a third image; Determining a first target image based on the wonderfulness score of the first image and the wonderfulness score of each of the second images includes: Based on the text description sentence of the preset action, performing action recognition of the preset action on the third image; When the preset action is recognized in the third image, acquiring at least one fourth image, wherein the at least one fourth image is continuous with the third image in the first image stream; Obtaining a wonderfulness score of the third image and a wonderfulness score of each fourth image in the at least one fourth image; The first target image is determined based on the wonderfulness score of the first image, the wonderfulness score of each of the second images, the wonderfulness score of the third image, and the wonderfulness score of each of the fourth images.
3. The image processing method according to claim 2, It is characterized in that Determining the first target image based on the wonderfulness score of the first image, the wonderfulness score of each of the second images, the wonderfulness score of the third image, and the wonderfulness score of each of the fourth images includes: Obtaining a maximum value of the wonderfulness score of the first image and the wonderfulness scores of each of the second images to obtain a first score; Storing the first score in a target buffer; Obtaining a maximum value of the wonderfulness score of the third image and the wonderfulness scores of each of the fourth images to obtain a second score; storing the second score in a target buffer; When the number of scores stored in the target cache reaches a preset number threshold, obtaining a plurality of scores including the first score and the second score from the target cache; A maximum score among the multiple scores is obtained, and an image corresponding to the maximum score is determined as the first target image.
4. The image processing method according to claim 1, It is characterized in that The determining of the first target image based on the wonderfulness score of the first image and the wonderfulness score of each of the second images comprises: Obtaining a maximum value of the wonderfulness score of the first image and the wonderfulness scores of each of the second images to obtain a first score; Storing the first score in a target buffer; After acquiring a fifth image through the camera, if the difference between the image sequence number of the fifth image and the reference image sequence number is greater than or equal to a preset difference threshold, acquiring a stored score in the target cache area, the stored score including the first score, and the reference image sequence number is the image sequence number of the image corresponding to the first score cached in the target cache area; The maximum score among the stored scores is obtained, and an image corresponding to the maximum score is determined as the first target image.
5. The image processing method according to claim 3 or 4, It is characterized in that After determining the first target image, the method further includes: Clear the target buffer area.
6. The image processing method according to claim 2 or 4, It is characterized in that The method further comprises: Acquire a sixth image through the camera; In a case where the sixth image is separated from the reference image by a preset number of images, based on a text description sentence of the preset action, performing action recognition of the preset action on the sixth image, the reference image being the last image collected among the third image and at least one fourth image, or the fifth image; When the preset action is recognized in the sixth image, acquiring at least one seventh image in a first image stream including the sixth image, wherein the at least one sixth image and the seventh image are continuous in the first image stream; Obtaining a wonderfulness score of the sixth image and a wonderfulness score of each seventh image in the at least one seventh image; A second target image is determined based on the wonderfulness score of the sixth image and the wonderfulness scores of each of the seventh images.
7. The image processing method according to claim 6, It is characterized in that When the sixth image is separated from the reference image by a preset number of images, performing action recognition of the preset action on the sixth image based on the text description of the preset action includes: When the difference between the image sequence number of the sixth image and the image sequence number of the reference image is greater than or equal to the preset number, action recognition of the preset action is performed on the sixth image based on the text description sentence of the preset action.
8. The image processing method according to claim 6, It is characterized in that The terminal device includes a first flag bit, The method further comprises: After acquiring the first target image, setting the first flag bit to a first value; and, acquiring an eighth image through the camera; Acquire a difference between an image sequence number of the eighth image and an image sequence number of the reference image; When the difference is equal to the preset number, the first flag bit is set to a second value.
9. The image processing method according to claim 8, It is characterized in that When the sixth image is separated from the reference image by a preset number of images, performing action recognition of the preset action on the sixth image based on the text description of the preset action includes: In a case where the first flag is set to the second value, based on the text description sentence of the preset action, action recognition of the preset action is performed on the sixth image.
10. The image processing method according to claim 1, It is characterized in that The step of performing action recognition of a preset action on the first image based on the text description sentence of the preset action includes: Acquire image features of the first image; Acquire text features of the text description sentence; Obtaining a similarity score between an image feature of the first image and a text feature of the text description sentence; Based on the similarity score, it is determined whether the preset action is recognized in the first image.
11. The image processing method according to claim 10, It is characterized in that The acquiring the image feature of the first image includes: The first image is input into a lightweight model to obtain image features of the first image.
12. The image processing method according to claim 11, It is characterized in that The lightweight model is obtained by distilling the image encoder of the multimodal large model by the server.
13. The image processing method according to claim 10, It is characterized in that The obtaining of the text features of the text description sentence includes: Obtaining text features of the text description sentence sent by the server, The text features of the text description sentence are obtained by the server using a text encoder to extract text features from the text description sentence, and the text encoder is a text encoder of a multimodal large model.
14. The image processing method according to claim 10, It is characterized in that The number of the preset actions is multiple, The obtaining of a similarity score between the image feature of the first image and the text feature of the text description sentence includes: Obtaining a similarity score between an image feature of the first image and a text feature of a text description sentence of each preset action in a plurality of preset actions; The determining, based on the similarity score, whether the preset action is recognized in the first image includes: Determining a maximum similarity score among the obtained multiple similarity scores; Based on the maximum similarity score and a preset score threshold, it is determined whether the preset action is recognized in the first image.
15. The image processing method according to claim 11, It is characterized in that After inputting the first image into the lightweight model to obtain the image features of the first image, the method further includes: Acquire image features of at least one ninth image in an image feature sequence including image features of the first image, wherein the image features of the at least one ninth image and the image features of the first image are continuous in the image feature sequence, wherein the image feature sequence is a sequence composed of image features of images in a second image stream, and the second image stream is obtained by extracting frames from the first image stream; performing smoothing processing on image features of the first image using image features of the at least one ninth image; The step of obtaining a similarity score between the image feature of the first image and the text feature of the text description sentence includes: A similarity score between the image features of the smoothed first image and the text features of the text description sentence is obtained.
16. The image processing method according to claim 15, It is characterized in that The acquiring, from the image feature sequence including the image feature of the first image, at least one image feature of a ninth image comprises: Sliding a first window in the image feature sequence to a first position corresponding to the image feature of the first image; Other features except the image feature of the first image in the first window are acquired to obtain at least one image feature of the ninth image.
17. The image processing method according to claim 1, It is characterized in that The acquiring at least one second image from a first image stream containing the first image comprises: Sliding a second window in the first image stream to a second position corresponding to the first image; Acquire other images in the second window except the first image to obtain the at least one second image.
18. The image processing method according to claim 1, It is characterized in that Determining a third target image based on the first target image and the second target image; The third target image is saved.
19. An electronic device, It is characterized in that include: one or more processors; Memory; and one or more computer programs, wherein the one or more computer programs are stored in the memory, and when the computer programs are executed by the one or more processors, the electronic device performs the following steps: In response to a first operation of the user, acquiring a first image stream captured by a camera, wherein the first image stream includes a first image; Based on the text description sentence of the preset action, performing action recognition of the preset action on the first image; When the preset action is recognized in the first image, acquiring at least one second image, wherein the at least one second image and the first image are continuous in the first image stream; Obtaining a wonderfulness score of the first image and a wonderfulness score of each second image in the at least one second image; A first target image is determined based on the wonderfulness score of the first image and the wonderfulness score of each of the second images.
20. The electronic device according to claim 19, It is characterized in that The first image stream also includes a third image; When the computer program is executed by the one or more processors, the electronic device performs the following steps: Based on the text description sentence of the preset action, performing action recognition of the preset action on the third image; When the preset action is recognized in the third image, acquiring at least one fourth image, wherein the at least one fourth image is continuous with the third image in the first image stream; Obtaining a wonderfulness score of the third image and a wonderfulness score of each fourth image in the at least one fourth image; The first target image is determined based on the wonderfulness score of the first image, the wonderfulness score of each of the second images, the wonderfulness score of the third image, and the wonderfulness score of each of the fourth images.
21. The electronic device according to claim 19, It is characterized in that When the computer program is executed by the one or more processors, the electronic device performs the following steps: Obtaining a maximum value of the wonderfulness score of the first image and the wonderfulness scores of each of the second images to obtain a first score; Storing the first score in a target buffer; After acquiring a fifth image through the camera, if the difference between the image sequence number of the fifth image and the reference image sequence number is greater than or equal to a preset difference threshold, acquiring a stored score in the target cache area, the stored score including the first score, and the reference image sequence number is the image sequence number of the image corresponding to the first score cached in the target cache area; The maximum score among the stored scores is obtained, and an image corresponding to the maximum score is determined as the first target image.
22. The electronic device according to claim 20 or 21, It is characterized in that The first image stream further includes a third image; when the computer program is executed by the one or more processors, the electronic device performs the following steps: Acquire a sixth image through the camera; In a case where the sixth image is separated from the reference image by a preset number of images, based on a text description sentence of the preset action, performing action recognition of the preset action on the sixth image, the reference image being the last image collected among the third image and at least one fourth image, or the fifth image; When the preset action is recognized in the sixth image, acquiring at least one seventh image in a first image stream including the sixth image, wherein the at least one sixth image and the seventh image are continuous in the first image stream; Obtaining a wonderfulness score of the sixth image and a wonderfulness score of each seventh image in the at least one seventh image; A second target image is determined based on the wonderfulness score of the sixth image and the wonderfulness scores of each of the seventh images.
23. A chip, It is characterized in that including one or more interface circuits and one or more processors; Wherein, the interface circuit is used to receive a signal from a memory of an electronic device and send the signal to the processor, wherein the signal includes computer instructions stored in the memory; when the processor executes the computer instructions, the electronic device executes the image processing method described in any one of claims 1 to 18.
24. A computer-readable storage medium, It is characterized in that It comprises a computer program, which, when executed on an electronic device, enables the electronic device to execute the image processing method according to any one of claims 1 to 18.
Citation Information
Patent Citations
Campus safety monitoring and early warning management system based on artificial intelligence
CN111967400A
Image optimization method and device, mobile terminal and storage medium
CN113256503A
Shooting method and electronic equipment
CN115525188A
Image recognition method and device, electronic equipment and storage medium
CN116977707A
Action recognition method and device, electronic equipment and storage medium
CN116994188A