Motion focusing method, electronic equipment and storage medium

By acquiring and processing multi-frame images in electronic devices, building a trusted trajectory of moving objects and predicting their locations, the focus delay problem of electronic devices when shooting moving objects is solved, and image clarity is improved.

CN120302152APending Publication Date: 2025-07-11HONOR DEVICE CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202410009866.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-02
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

When shooting moving objects, existing electronic devices are affected by focus delay, resulting in low image clarity and unable to meet user needs.

Method used

By acquiring the continuous N-frame images acquired by the camera, using motion detection and prediction models to construct a trusted trajectory of the moving object, predict its position in subsequent image frames, and focus based on the predicted position to improve focus accuracy.

Benefits of technology

The image clarity of electronic devices when shooting moving objects is improved, the problem of inaccurate prediction trajectory caused by false detection of motion detection algorithms is reduced, and the accuracy of focus is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120302152A_ABST
    Figure CN120302152A_ABST
Patent Text Reader

Abstract

The invention provides a motion focusing method, electronic equipment and a storage medium, and relates to the technical field of terminals.The method comprises the steps that under the condition that a camera application of the electronic equipment starts a motion focusing mode, the electronic equipment can build a credible track of a moving object, and the motion focusing mode of the moving object in continuous image frames can be obtained based on the credible position of the moving object in the continuous image frames; and predicting the motion trail of the motion object in the subsequent image frame, and controlling the camera to perform motion focusing based on the predicted motion trail of the motion object in the subsequent image frame. According to the method, the camera focusing accuracy can be improved, and then the image definition of the moving object shot by the electronic equipment is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of terminals, and in particular, to a motion focusing method, an electronic device, and a storage medium. Background Art

[0002] With the popularization of intelligent terminal devices, it has become a normal state for users to use electronic devices such as mobile phones and tablet computers to take pictures, and users' requirements for the image quality of the pictures taken are also getting higher and higher. For example, when users use electronic devices to take pictures of people in a moving state or objects moving at high speed, etc., users expect to capture a clear image of the moving object.

[0003] However, in practical applications, affected by the focusing delay of the electronic device, the image clarity of the moving object captured by the electronic device is low and cannot meet the shooting needs of users. Summary of the Invention

[0004] Embodiments of this application provide a motion focusing method, an electronic device, and a storage medium, which can improve the accuracy of camera focusing and further improve the image clarity of the moving object captured by the electronic device.

[0005] In a first aspect, an embodiment of this application proposes a motion focusing method. In response to an operation of turning on the motion focusing mode, the electronic device acquires N consecutive frames of images collected by the camera. Each of the N frames of images includes the same moving object, and N is a positive integer greater than 1. The electronic device acquires the reliable position information of the moving object in each of the N frames of images. The electronic device determines the predicted position information of the moving object in the M frames of images after the N frames of images according to the reliable position information of the moving object in the N frames of images, and M is a positive integer greater than 1. The electronic device controls the camera to perform motion focusing according to the predicted position information of the moving object in the P-th frame of the M frames of images, and P is a positive integer less than or equal to M.

[0006] Exemplarily, as shown in Figure 1 c, in response to an operation on the motion focusing control 1032 on the window 1031 in the interface 103, the electronic device acquires N consecutive frames of images collected by the camera to execute the subsequent motion focusing method. Taking a certain frame of the N frames of images as an example, the reliable position information of the moving object in this frame of image is determined according to the detected position information and the predicted position information of the moving object in this frame of image. Exemplarily, as shown in Figure 7 , the position information of the reliable box 2 of the moving object in the second frame of image is obtained by weighted summation of the position information of the detected box 2 of the moving object in the second frame of image and the position information of the predicted box 2 of the moving object in the second frame of image.

[0007] In the above embodiments, when the camera application of the electronic device enables the motion focus mode, the electronic device can construct a reliable trajectory of the moving object (i.e., the trajectory corresponding to the reliable positions of the moving object in consecutive image frames), predict the motion trajectory of the moving object in subsequent image frames (i.e., the predicted positions of the moving object in subsequent image frames) based on the reliable trajectory of the moving object in consecutive image frames, and control the camera to perform motion focus based on the predicted motion trajectory of the moving object in subsequent image frames. Compared with the electronic device predicting the motion trajectory of the moving object in subsequent image frames based on the detected trajectory of the moving object in consecutive image frames (i.e., the trajectory corresponding to the detected positions of the moving object in consecutive image frames), the above method can reduce the problem that the inaccuracy of the detected trajectory caused by the misdetection of the motion detection algorithm affects the accuracy of the predicted trajectory of the moving object, improve the accuracy of camera focus, and thus improve the image clarity of the moving object captured by the electronic device.

[0008] In an optional embodiment of the first aspect, obtaining the reliable position information of the moving object in each of the N frames of images includes: obtaining the detected position information of the moving object in the i-th frame of the N frames of images and the predicted position information of the moving object in the i-th frame of images, where the i-th frame of images is any image frame other than the first frame of the N frames of images; determining the reliable position information of the moving object in the i-th frame of images according to the detected position information and the predicted position information of the moving object in the i-th frame of images. Where i is an integer greater than 1 and less than or equal to N.

[0009] The above embodiments illustrate the method of constructing the reliable positions of the moving object in consecutive image frames. The electronic device can determine the reliable position of the moving object in each frame of images according to the detected position and the predicted position of the moving object in each frame of images. The reliability of the reliable position of the moving object in this frame of images is higher than the detected position of the moving object in this frame of images. The reliable positions of the moving object in the constructed consecutive image frames can be used by the electronic device to predict the motion trajectory of the moving object in subsequent image frames.

[0010] In an optional embodiment of the first aspect, the detected position information of the moving object in the i-th frame of images includes the coordinate position of the first detection box of the moving object in the i-th frame of images, and the predicted position information of the moving object in the i-th frame of images includes the coordinate position of the first prediction box of the moving object in the i-th frame of images; determining the reliable position information of the moving object in the i-th frame of images according to the detected position information and the predicted position information of the moving object in the i-th frame of images includes: obtaining the coordinate position of the first reliable box of the moving object in the i-th frame of images through a fusion operation on the coordinate positions of the first detection box and the first prediction box, and the reliable position information of the moving object in the i-th frame of images includes the coordinate position of the first reliable box.

[0011] In the above embodiments, the electronic device performs coordinate position fusion based on the coordinate positions of the upper left corner in the detection box and the prediction box of the i-th frame image to obtain the coordinate position of the upper left corner in the reliable box of the i-th frame image, and performs coordinate position fusion based on the coordinate positions of the lower right corner in the detection box and the prediction box of the i-th frame image to obtain the coordinate position of the lower right corner in the reliable box of the i-th frame image. The coordinate positions of the upper left corner and the lower right corner of the reliable box of the i-th frame image can uniquely determine the position of the reliable box in the i-th frame image, and this position can be used for subsequent prediction of the motion trajectory of the moving object.

[0012] In an alternative embodiment of the first aspect, the coordinate position of the first detection box includes the coordinate position of the upper left corner in the first detection box and the coordinate position of the lower right corner in the first detection box, and the coordinate position of the first prediction box includes the coordinate position of the upper left corner in the first prediction box and the coordinate position of the lower right corner in the first prediction box; by performing a fusion operation on the coordinate positions of the first detection box and the first prediction box, the coordinate position of the first reliable box of the moving object in the i-th frame image is obtained, including: by performing a weighted summation operation on the coordinate position of the upper left corner in the first detection box and the coordinate position of the upper left corner in the first prediction box, the coordinate position of the upper left corner in the first reliable box is obtained; and by performing a weighted summation operation on the coordinate position of the lower right corner in the first detection box and the coordinate position of the lower right corner in the first prediction box, the coordinate position of the lower right corner in the first reliable box is obtained.

[0013] In the above embodiments, the electronic device respectively performs a weighted summation operation on the coordinate positions of the upper left corner of the detection box and the prediction box in the same frame image, and performs a weighted summation operation on the coordinate positions of the lower right corner of the detection box and the prediction box in the same frame image, to obtain the coordinate positions of the upper left corner and the lower right corner of the reliable box in this frame image, which are used for subsequent prediction of the motion trajectory of the moving object.

[0014] In an alternative embodiment of the first aspect, obtaining the coordinate position of the upper left corner in the first reliable box by performing a weighted summation operation on the coordinate position of the upper left corner in the first detection box and the coordinate position of the upper left corner in the first prediction box includes:

[0015] Determining the coordinate position (x, y) of the upper left corner in the first reliable box through the following formula:

[0016] X = w1 × x predict + w2 × x detect

[0017] y = w1 × y predict + w2 × y detect

[0018] In the formula, (x detect , y detect ) represents the coordinate position of the upper left corner in the first detection box, (xpredict , y predict ) represents the coordinate position of the upper left corner in the first prediction box, w1 represents the weight value of the first prediction box, and w2 represents the weight value of the first detection box.

[0019] In the above embodiments, the electronic device performs weighted summation operations on the x-axis coordinate positions of the upper left corners of the detection box and the prediction box in the same frame of image respectively, and performs weighted summation operations on the y-axis coordinate positions of the upper left corners of the detection box and the prediction box in the same frame of image, so as to obtain the coordinate position of the upper left corner of the reliable box in this frame of image, which is used for subsequent prediction of the motion trajectory of the moving object. It should be understood that based on the principle of the above formula, the coordinate positions of the lower right corners of the reliable box in the x-axis and y-axis in this frame of image can be obtained respectively.

[0020] In an alternative embodiment of the first aspect, the sum of the weight value of the first prediction box and the weight value of the first detection box is 1, and the weight value of the first prediction box is determined according to the first distance value and the first moving speed of the moving object; the first distance value is the distance value between the detection box and the upper left corner in the prediction box of the moving object in the i-th frame of image, and the first moving speed is the average moving speed of the upper left corner of the reliable box of the moving object from the (i - N)-th frame of image to the (i - 1)-th frame of image.

[0021] It should be understood that the sum of the weight value of the first prediction box and the weight value of the first detection box is 1. After the electronic device determines the weight value of the first prediction box, the weight value of the first detection box can be determined. The calculation of the weight value of the first prediction box can refer to Formula 3, and the calculation of the weight value of the first detection box can refer to Formula 4.

[0022] It should be understood that when the moving speed of the reliable box of the moving object (which can correspond to the above first moving speed) remains unchanged, as the position offset value between the detection box and the prediction box (which can correspond to the above first distance value) increases, the weight value of the prediction box gradually increases, and the weight value of the detection box gradually decreases. When the position offset value between the detection box and the prediction box remains unchanged, as the moving speed of the reliable box of the moving object increases, the weight value of the detection box gradually increases, and the weight value of the prediction box gradually decreases.

[0023] In an alternative embodiment of the first aspect, determining the predicted position information of the moving object in the M frames of images after the N frames of images according to the reliable position information of the moving object in the N frames of images includes: inputting the reliable position information of the moving object in the N frames of images into a motion prediction model, and after being processed by the motion prediction model, obtaining the predicted position information of the moving object in the M frames of images after the N frames of images; the motion prediction model is trained by using a neural network model.

[0024] In the above embodiments, the electronic device inputs the coordinate positions of the bounding boxes of the moving objects in N frames of images into a preset motion prediction model. After model operations, the predicted coordinate positions of the moving objects in the M frames of images after the N frames of images are obtained. The model input is not the N frames of images, but the position data of the moving objects in the images, which can improve the processing speed of the motion prediction model, thereby shortening the focusing time of the device and improving the motion focusing speed of the device.

[0025] In some embodiments, the motion prediction model can be trained using a lightweight fully connected neural network model. Using a lightweight model structure can, on the one hand, save the storage space of the device, and on the other hand, improve the execution speed of the model. In addition, due to the simple model structure and small computational amount of the model, the power consumption of the device can be reduced.

[0026] In an optional embodiment of the first aspect, after the electronic device controls the camera to perform motion focusing according to the predicted position information of the moving object in the P-th frame of the M frames of images, the method further includes: the electronic device acquires the P-th frame of the M frames of images collected by the camera; the electronic device displays the P-th frame of the image and the predicted position information of the moving object in the P-th frame of the image, and the image clarity of the moving object in the P-th frame of the image is greater than a preset clarity.

[0027] In the above embodiments, when the electronic device displays the P-th frame of the image, it simultaneously displays the predicted bounding box position of the moving object in the P-th frame of the image. This predicted bounding box position is the focusing position when the P-th frame of the image is captured. Since this predicted bounding box position is determined based on the foregoing method, the deviation between the predicted bounding box position and the actual position of the moving object in the P-th frame of the image is less than a preset value, making the image clarity of the moving object in the P-th frame of the image greater than the preset clarity. It should be understood that subsequent image frames are all focused according to the foregoing method, which can improve the image clarity of the moving objects captured by the electronic device as a whole.

[0028] In a second aspect, an embodiment of the present application provides an electronic device, which includes: a memory and a processor. The processor is used to call the computer program in the memory to execute the method according to any one of the first aspects.

[0029] In a third aspect, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program runs on an electronic device, the electronic device is caused to execute the method according to any one of the first aspects.

[0030] In a fourth aspect, an embodiment of the present application provides a chip, which includes a processor. The processor is used to call the computer program in the memory to cause the chip to execute the method according to any one of the first aspects.

[0031] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program, which, when run, causes a computer to execute the method described in any item of the first aspect.

[0032] It should be understood that the second to fifth aspects of the present application correspond to the technical solutions of the first aspect of the present application. The beneficial effects obtained by each aspect and the corresponding feasible implementation manners are similar and will not be repeated. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 It is a schematic diagram of the change of the mobile phone interface provided by the embodiment of the present application;

[0034] Figure 2 It is a schematic diagram of the principle of the motion focusing method provided by the embodiment of the present application Figure 1 ;

[0035] Figure 3 It is a schematic diagram of the process of the motion focusing method provided by the embodiment of the present application Figure 1 ;

[0036] Figure 4 It is a schematic diagram of the structure of the motion prediction model provided by the embodiment of the present application;

[0037] Figure 5 It is a schematic diagram of the principle of the motion focusing method provided by the embodiment of the present application Figure 2 ;

[0038] Figure 6 It is a schematic diagram of the process of the motion focusing method provided by the embodiment of the present application Figure 2 ;

[0039] Figure 7 It is a schematic diagram of the principle of the motion prediction provided by the embodiment of the present application;

[0040] Figure 8 It is a schematic diagram of the structure of an electronic device provided by the embodiment of the present application;

[0041] Figure 9 It is a schematic diagram of the software architecture and internal interaction of an electronic device provided by the embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0042] In order to clearly describe the technical solutions of the embodiments of the present application, in the embodiments of the present application, terms such as "first" and "second" are used to distinguish identical items or similar items with basically the same functions and effects. Those skilled in the art can understand that the terms "first" and "second" do not limit the quantity and execution order, and the terms "first" and "second" do not necessarily limit to being different.

[0043] It should be noted that in the embodiments of the present application, words such as "exemplary" or "for example" are used to represent examples, illustrations or explanations. Any embodiment or design described as "exemplary" or "for example" in the present application should not be construed as being more preferred or having more advantages than other embodiments or designs. Rather, the use of words such as "exemplary" or "for example" is intended to present relevant concepts in a specific manner.

[0044] In the embodiments of the present application, "at least one" means one or more, and "a plurality" means two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone, where A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after. "At least one (kind / individual) of the following" or its similar expressions refer to any combination of these items, including any combination of single items (kinds / individuals) or plural items (kinds / individuals). For example, at least one (kind / individual) of a, b or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, c can be single or multiple.

[0045] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc., and the data includes images or videos, etc.) involved in the present application are all information and data that have been authorized by the user or fully authorized by all parties. And the collection, use and processing of the relevant data need to comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entrances are provided for the user to choose to authorize or refuse.

[0046] For the sake of easy understanding, the following first explains the professional terms involved in the embodiments of the present application.

[0047] First, motion focus is a focusing mode of an electronic device, which is mainly applicable to shooting dynamic objects, such as shooting athletes on a sports field or flying birds, etc. This focusing mode can quickly and accurately capture the moving object and maintain focus. When the shooting object moves in the frame, the electronic device will automatically adjust the camera focusing parameters, such as the focusing position, focal length, etc., to ensure that the object always remains clear. The embodiments of the present application optimize the algorithm in the motion focus mode of the electronic device.

[0048] Second, subject focus is another focusing mode of the electronic device, mainly applicable to shooting portraits, animals or objects with specific features. This focusing mode can automatically identify the subject in the picture and prioritize focusing on it. For example, when shooting a portrait, if the person's face is facing the device camera, the subject focus will automatically align the focus on the face; if the person is moving or turning, the focus will automatically track and remain aligned with the subject.

[0049] Third, the motion detection algorithm is an important field in computer vision, mainly used to detect and track moving objects, such as people, vehicles, and animals in motion. Common algorithms include: detection algorithms based on bounding boxes (or detection boxes), single-stage detection algorithms, key-point based detection algorithms, etc., which are not limited in the embodiments of this application.

[0050] Fourth, the lightweight neural network is a neural network model that performs deep learning tasks under resource-constrained conditions. Its design aims to reduce the number of model parameters and computational complexity while maintaining sufficient performance to meet the requirements of specific application scenarios. In the embodiments of this application, the motion prediction model can be trained using a lightweight neural network, for example, trained using a lightweight fully connected neural network. The structure of the lightweight neural network is not limited in the embodiments of this application.

[0051] Fifth, the linear layer is one of the basic layer types in the neural network and is also known as the fully connected layer. The purpose of the linear layer is to multiply the input vector by the weight matrix and add the bias vector to generate a new output vector.

[0052] Sixth, ResBlock is the basic structural unit in the residual network ResNet and is usually called the residual block. It alleviates the vanishing gradient problem by introducing skip connections to pass the information of the previous layer to the subsequent layers. For example, Figure 4 In the shown ResBlock module, the input of the ResBlock module is summed with the output of the last LeakyReLu in the ResBlock module, so as to pass the information of the previous layer along the depth of the network, which can improve the network performance. LeakyReLu is a neural network activation function. Figure 4 The shown ResBlock module also includes Dropout. Dropout is a regularization method for neural network models. During the training process, by randomly ignoring a part of the neurons and making a part of the neurons stop working with a certain probability, it can prevent the model from overfitting to the training data and improve the generalization ability of the model to new data.

[0053] In the scenario of shooting a moving object, the user can turn on the motion focus mode in the camera application of the electronic device to trigger the electronic device to shoot the moving object in the motion focus mode.

[0054] Exemplarily, Figure 1 is a schematic diagram of the change of the mobile phone interface provided by the embodiment of the present application. As Figure 1 shown, in response to an operation by the user on the icon 101 of the camera application on the main interface of the mobile phone, the mobile phone displays an interface 102, and the interface 102 includes a shooting mode selection area 1021, a control 1022 for viewing the album, a shooting control 1023, a control 1024 for switching the camera, etc. The shooting mode selection area 1021 displays a night scene mode, a portrait mode, a photo shooting mode, a video recording mode, etc., and the user can select a shooting mode in the area 1021 for shooting. As Figure 1 shown in b, the camera application can default to turn on the photo shooting mode.

[0055] In some embodiments, in response to an operation by the user on the photo shooting mode 1025 in the area 1021, the mobile phone displays an interface 103, and the interface 103 includes a window 1031, and the window 1031 displays a control 1032 for motion focus. In one example, the camera application of the mobile phone defaults to preferentially turn on the motion focus mode. Therefore, when the mobile phone displays the window 1031, the control 1032 for motion focus can be defaultly displayed, and at this time, the motion focus mode has not been turned on.

[0056] In one example, in response to an operation by the user on the control 1032, the mobile phone defaultly turns on the motion focus mode. As Figure 1 shown in d, the mobile phone displays an interface 104, and the interface 104 includes a window 1051, and the window 1051 displays a control 1052 for motion focus priority and a control 1053 for subject focus priority. The control 1052 for motion focus priority is displayed in bold to indicate that the camera application of the mobile phone has turned on the motion focus mode.

[0057] The user can switch the focus mode in the window 1051. In one example, in response to an operation by the user on the control 1053 on the window 1051, the camera application of the mobile phone can be switched from the motion focus mode to the subject focus mode. Correspondingly, the control 1053 for subject focus is displayed in bold (not shown) to indicate that the camera application of the mobile phone has turned on the subject focus mode.

[0058] In some embodiments, after the electronic device acquires an image collected by the camera, it detects the position of the moving object in the image, and adjusts the focus position of the camera based on the detected position to guide the camera to collect subsequent images.

[0059] Exemplarily, Figure 2 is a schematic diagram of the principle of the motion focus method provided by the embodiment of the present application Figure 1 . As Figure 2As shown in a, in the motion focus mode, the electronic device acquires the i-th frame image collected by the camera, and the i-th frame image includes a moving vehicle. Based on the motion detection algorithm, the electronic device can obtain the position information of the vehicle in the i-th frame image. The position information of the vehicle includes the coordinate position of the detection frame 21 of the vehicle in the i-th frame image. For example, the coordinate positions of the upper left corner and the lower right corner of the detection frame 21 in the i-th frame image are used as the coordinate position of the detection frame 21. When the shooting angle remains unchanged, the electronic device can adjust the focus position of the camera based on the position information of the vehicle in the i-th frame image. When the moving speed of the vehicle is less than the first threshold, the camera of the electronic device can focus based on the coordinate position of the detection frame 21 in the i-th frame image to acquire the (i + 1)-th frame image.

[0060] However, when the moving speed of the vehicle is greater than the second threshold, the actual position of the vehicle in the (i + 1)-th frame image collected by the camera of the electronic device will deviate from the detected position of the vehicle in the i-th frame image. As Figure 2 shown in b, the detection frame 22 can be regarded as the position where the vehicle position in the i-th frame image is projected onto the (i + 1)-th frame image, and the detection frame 23 is the actual position of the vehicle in the (i + 1)-th frame image. The electronic device focuses based on the detection frame 22 to acquire the (i + 1)-th frame image. Due to the inaccurate focus position of the camera, the clarity of the vehicle in the (i + 1)-th frame image is low, such as the clarity of the vehicle being less than the preset clarity.

[0061] In view of the above problems, an embodiment of the present application provides a motion focus method. After the electronic device acquires multiple frames of images collected by the camera, it detects the positions of moving objects in the multiple frames of images, and uses a preset motion prediction model based on the multiple detected positions of the moving objects to obtain the predicted positions of the moving objects in the images after the multiple frames of images. The electronic device can adjust the focus position based on the predicted positions of the moving objects to guide the camera to collect subsequent images.

[0062] In this embodiment, the electronic device can learn the historical motion trajectory of the moving object according to the multiple detected positions of the moving object, and based on this historical motion trajectory and combined with the motion prediction model, obtain the motion trajectory of the moving object in the subsequent images. Compared with the foregoing embodiments, since the focus position is more accurate, the image clarity of the moving object in the subsequent images can be improved.

[0063] Next, in combination with Figures 3 to 5 a detailed description of the motion focus method provided by the embodiment of the present application will be given.

[0064] Exemplarily, Figure 3 is a schematic flow chart of the motion focus method provided by the embodiment of the present application Figure 1 . This method can be applied to an electronic device with shooting and image processing functions. Such as Figure 3As shown in the figure, the motion focusing method of this embodiment includes:

[0065] S301. In response to an operation to enable the motion focusing mode, the electronic device acquires N consecutive frames of images captured by the camera. N is an integer greater than or equal to 2.

[0066] Exemplarily, as shown in c of Figure 1 , in response to an operation by the user on the motion focusing control 1032, such as the user clicking on the control 1032, the electronic device enables the motion focusing mode. After the electronic device enables the motion focusing mode, the motion focusing method provided in this embodiment can be executed.

[0067] S302. The electronic device performs motion detection on the N frames of images to obtain the detection box position information of the moving object in each of the N frames of images.

[0068] In some embodiments, the electronic device can perform motion detection on the N frames of images based on a detection algorithm for bounding boxes to obtain the detection box position information of the moving object in each of the N frames of images.

[0069] The detection box position information of the moving object may include the coordinate positions of the diagonal points of the detection box of the moving object in the image, such as the coordinate positions of the upper left corner and the lower right corner of the detection box in the image, or the coordinate positions of the upper right corner and the lower left corner of the detection box in the image.

[0070] The embodiments of the present application do not limit the algorithm for motion detection.

[0071] S303. The electronic device predicts the predicted box position information of the moving object in the M frames of images after the N frames of images according to the detection box position information of the moving object in the N frames of images.

[0072] In some embodiments, the electronic device can predict the predicted box position information of the moving object in the M frames of images after the N frames of images through a motion prediction model. In this embodiment, a trained motion prediction model is preset in the electronic device, and the electronic device can input the detection box position information of the moving object in the N frames of images into the motion prediction model to obtain the predicted box position information of the moving object in the M frames of images after the N frames of images.

[0073] The predicted box position information of the moving object may include the coordinate positions of the diagonal points of the predicted box of the moving object in the image, such as the coordinate positions of the upper left corner and the lower right corner of the predicted box in the image, or the coordinate positions of the upper right corner and the lower left corner of the predicted box in the image.

[0074] In some embodiments, the motion prediction model can be trained using a lightweight neural network structure, for example, trained using a lightweight fully-connected neural network structure. By adopting a lightweight model structure, on the one hand, it can save the storage space of the device, on the other hand, it can also improve the execution speed of the model. In addition, due to the simple model structure and small computational amount of the model, the device power consumption can be reduced.

[0075] Exemplarily, Figure 4 is a schematic structural diagram of the motion prediction model provided by the embodiments of the present application. As Figure 4 shown, the input of the motion prediction model includes the coordinate positions of the detection frames of the moving objects in N consecutive frames of images, that is, Figure 4 the coordinate positions of N detection frames in. The output of the motion prediction model includes the coordinate positions of the predicted frames of the moving objects in M frames of images after the N frames of images, that is, Figure 4 the coordinate positions of M predicted frames in. N is a positive integer greater than 1, and M is a positive integer greater than 1. Exemplarily, M is less than N. For example, N is 8 and M is 4. Another example is that N is 3 and M is 2. The embodiments of the present application do not limit N and M.

[0076] In this embodiment, the input of the motion prediction model is the coordinate positions of the detection frames of the moving objects in N frames of images. The coordinate positions of the detection frames of the moving objects in each frame of image can be denoted as (x1, y1, x2, y2). For example, (x1, y1) can be the coordinate position of the upper left corner of the detection frame in the image, and (x2, y2) can be the coordinate position of the upper left corner of the detection frame in the image. Of course, in some embodiments, the coordinate positions of the detection frame can also be marked by using the other two diagonal points of the detection frame, such as the coordinate positions of the upper right corner and the lower left corner in the image. It can be seen that the input of the motion prediction model includes N×4 numbers. Correspondingly, the output of the motion prediction model includes M×4 numbers. The simple model input and lightweight model structure can improve the processing speed of the motion prediction model, thereby shortening the focusing time and improving the motion focusing speed of the device.

[0077] In some embodiments, such as Figure 4As shown in the figure, the motion prediction model includes two linear layers and Q ResBlock modules. The two linear layers are respectively denoted as linear layer 1 and linear layer 2. The output of linear layer 1 is connected to the first ResBlock module among the Q ResBlock modules, and the output of the last ResBlock module among the Q ResBlock modules is connected to linear layer 2. Where Q is a positive integer, for example, Q can be 3. Exemplarily, a ResBlock module may include a linear layer 3, a Dropout layer 1, a LeakyReLu layer 1, a linear layer 4, a Dropout layer 2, and a LeakyReLu layer 2 connected in sequence. The output of this ResBlock module is the sum value of the output of LeakyReLu3 and the input of linear layer 3.

[0078] In the motion prediction model shown in this embodiment, there are ResBlock modules. The ResBlock modules can transfer the information of the previous layer of the model along the depth of the network, which can improve the motion prediction ability of the model.

[0079] The motion prediction model shown in this embodiment is only an example. The motion prediction model may include more or fewer layers than shown in the figure to implement the motion prediction function. The structure of the motion prediction model is not limited in this application embodiment.

[0080] S304. The electronic device adjusts the focus position of the camera according to the predicted box position information of the moving object in the P-th frame image among the M frame images. M is a positive integer, and P is a positive integer less than or equal to M.

[0081] Exemplarily, if N is 8 and M is 4, the electronic device can obtain the predicted box position of the moving object in each of the last 4 frame images according to the detection box positions of the moving object in the previous 8 frame images. For example, the electronic device adjusts the focus position of the camera according to the predicted box position of the moving object in the first frame image among the last 4 frame images. Another example is that the electronic device adjusts the focus position of the camera according to the predicted box position of the moving object in the second frame image among the last 4 frame images.

[0082] The following combines Figure 5 to elaborate in detail how the electronic device adjusts the camera focus position according to the predicted box position information. Exemplarily, Figure 5 is a scenario schematic of the motion focus method provided by this application embodiment Figure 2 . For ease of understanding, Figure 5 the solution is described by taking N as 3 and M as 2 as an example.

[0083] In the motion focus mode, the electronic device acquires three consecutive frame images collected by the camera, such as Figure 5For the (i-2)th frame image, the (i-1)th frame image, and the ith frame image in [reference], the electronic device first obtains the coordinate positions of the detection frames of the vehicle in these three frames of images based on a motion detection algorithm, such as Figure 5 the detection frames 51, 52, and 53 in [reference]. Based on the dashed arrows in the figure, it can be seen that the vehicle's motion trajectory is moving from right to left along the lane. After the electronic device obtains the coordinate positions of the detection frames of the vehicle in the three frames of images, it inputs the coordinate positions of the detection frames 51, 52, and 53 into a motion prediction model. After being operated by the motion prediction model, the model outputs the coordinate positions of the predicted frames of the vehicle in the next two frames of images, such as Figure 5 the predicted frames 54' and 55' in [reference].

[0084] In some embodiments, P is configured to be 1. The electronic device can adjust the focusing position of the camera according to the position information of the predicted frame of the moving object in the first frame of the M-frame images, so as to capture the first frame of the images after the first N frames. Exemplarily, referring to Figure 5 , the electronic device can focus based on the coordinate position of the predicted frame 54', and capture the (i + 1)th frame image, which can improve the image clarity of the vehicle in the (i + 1)th frame image.

[0085] In the embodiments of the present application, the camera of the electronic device includes an image sensor. The image output speed of the image sensor is usually measured by the frame rate, that is, the number of frames per second (frames per second, fps). The higher the frame rate, the faster the image output speed of the image sensor. Exemplarily, the frame rate of the image sensor is 30 frames per second, that is, the image sensor outputs one frame of image approximately every 33 milliseconds (ms).

[0086] Exemplarily, ideally, the electronic device adjusts the camera focusing position to the focusing position 1 according to the position information of the predicted frame of the moving object in the first frame of the M-frame images, and the image sensor outputs the first frame of the M-frame images based on the focusing position 1. However, if the image output speed of the image sensor is greater than the preset speed, after the electronic device adjusts the camera focusing position to the focusing position 1 according to the position information of the predicted frame of the moving object in the first frame of the M-frame images, the image sensor may have already output the first frame of the M-frame images based on the previous focusing position, then the image sensor will output the second frame of the M-frame images based on the focusing position 1. Ideally, the electronic device adjusts the camera focusing position to the focusing position 2 according to the position information of the predicted frame of the moving object in the second frame of the M-frame images, and the image sensor outputs the second frame of the M-frame images based on the focusing position 2. Since the focusing position is inaccurate when capturing the second frame of the M-frame images, it will affect the image clarity of the moving object in the second frame of the M-frame images.

[0087] To address the above problems, an appropriate P value can be configured for the electronic device. For example, P is configured to 2. After the electronic device obtains the predicted bounding box positions of the moving objects in the M frames of images following the N frames of images based on S303, instead of using the predicted bounding box position of the moving object in the first frame of the M frames of images for focusing, the electronic device can use the predicted bounding box position of the moving object in the second frame of the M frames of images for focusing. This can avoid the problem of inaccurate focusing positions caused by the latency of running the prediction, and improve the image clarity of the moving objects in the images output by the image sensor.

[0088] Exemplarily, still referring to Figure 5 , when the electronic device obtains the predicted bounding box 54', the image sensor may have output the (i + 1)-th frame of image. When the electronic device obtains the predicted bounding box 54' and the predicted bounding box 55', it can focus based on the coordinate position of the predicted bounding box 55' to capture the (i + 2)-th frame of image, which can improve the image clarity of the vehicle in the (i + 2)-th frame of image.

[0089] In the motion focusing method shown in the above embodiments, the electronic device predicts the predicted bounding box positions of the moving objects in the images following the N frames of images based on the detected bounding box positions of the moving objects in the N frames of images collected by the camera, and adjusts the focusing position of the camera based on the predicted bounding box positions of the moving objects in the images following the N frames of images. Compared with adjusting the camera focusing position through the detected bounding box positions, adjusting the camera focusing position through the predicted bounding box positions can improve the image clarity of the moving objects in the images captured by the camera because the focusing position is more accurate.

[0090] Based on the foregoing embodiments, the electronic device can determine the predicted bounding box positions of the moving objects in subsequent multiple frames of images according to the detected bounding box positions of the moving objects in multiple frames of historical images. Considering the capabilities of the motion detection algorithm, in some cases, the detected bounding box positions obtained by the electronic device based on the motion detection algorithm may be inaccurate, such as false detections, which will affect the accuracy of the predicted bounding box positions and further affect the accuracy of camera focusing, resulting in a decrease in the image clarity of the moving objects in the images captured by the camera.

[0091] To further improve the accuracy of the predicted bounding box positions, an embodiment of the present application proposes a motion focusing method. The electronic device can fuse the coordinate positions of the detected bounding box of the moving object in the i-th frame of image with the predicted bounding box position of the moving object in the i-th frame of image to obtain the reliable bounding box position of the moving object in the i-th frame of image. Based on this method, the electronic device can obtain the reliable bounding box positions of the moving objects in multiple frames of images. The electronic device can predict the predicted bounding box positions of the moving objects in the images following the multiple frames of images based on the reliable bounding box positions of the moving objects in the multiple frames of images, and perform focusing based on the predicted bounding box positions.

[0092] In the above method, the position of the reliable box is obtained by fusing the position of the detection box and the position of the prediction box. The electronic device can predict the motion trajectory of the moving object according to the position of the reliable box, which can reduce the problem of inaccurate position of the prediction box caused by misdetection of the motion detection algorithm.

[0093] It should be noted that, based on the position of the detection box of the moving object in multiple frames of images, the electronic device can determine the detection box trajectory of the moving object in consecutive image frames. The detection box trajectory can be simply referred to as the detection trajectory. Based on the position of the prediction box of the moving object in multiple frames of images, the electronic device can determine the prediction box trajectory of the moving object in consecutive image frames. The prediction box trajectory can be simply referred to as the prediction trajectory. Based on the position of the reliable box of the moving object in multiple frames of images, the electronic device can determine the reliable box trajectory of the moving object in consecutive image frames. The reliable box trajectory can be simply referred to as the reliable trajectory. The electronic device can store the detection trajectory, prediction trajectory, and reliable trajectory of the moving object in consecutive image frames.

[0094] Next, in combination with Figure 6 the motion focusing method provided in this embodiment will be described in detail. Exemplarily, Figure 6 is a schematic flowchart of the motion focusing method provided in an embodiment of the present application Figure 2 . This method can be applied to an electronic device with shooting and image processing functions. As Figure 6 shown, the motion focusing method in this embodiment includes:

[0095] S601. In response to an operation to turn on the motion focusing mode, the electronic device acquires the i-th frame of image collected by the camera.

[0096] In this embodiment, the i-th frame of image can be any frame of image collected by the camera after the electronic device turns on the motion focusing mode.

[0097] S602. The electronic device performs motion detection on the i-th frame of image to obtain the position information of the detection box of the moving object in the i-th frame of image. S602 in this embodiment can refer to S302 in the foregoing embodiment. For the sake of brevity, it will not be elaborated here.

[0098] S603. The electronic device determines whether there is position information of the prediction box of the moving object in the i-th frame of image.

[0099] In some embodiments, if the i-th frame of image is the first frame of image collected by the camera after the electronic device turns on the motion diagonal mode, since the electronic device has not performed motion detection and motion prediction on the images collected by the camera before turning on the motion focusing mode, there is no position information of the prediction box of the moving object in the first frame of image in the electronic device.

[0100] In some embodiments, if the i-th frame image is an image after the first frame image captured by the camera after the electronic device turns on the motion focus mode, the electronic device stores the predicted frame position information of the moving object in the i-th frame image. For example, after the electronic device turns on the motion focus mode, the camera captures the second frame image, and the electronic device stores the predicted frame position information of the moving object in the second frame image. The predicted frame position information of the moving object in the second frame image can be determined by the electronic device based on the detection frame position information of the moving object in the first frame image. For details, please refer to Figure 7 The first example in the embodiment.

[0101] Combine the following Figure 7 The process of obtaining the predicted frame position information of the moving object in the image after the first frame image captured by the camera is illustrated by way of example. Figure 7 The schematic diagram of the principle of motion prediction provided by the embodiment of the present application. Take the first four frames of images captured by the camera after the electronic device turns on the motion focus mode as an example. Figure 7 As shown, the electronic device obtains the detection frame of the moving object in each of the first four frames of images, such as detection frame 1 to detection frame 4, based on the motion detection algorithm, and the detection trajectory of the moving object includes detection frame 1 to detection frame 4. The electronic device is pre-installed with a motion prediction model. The input of the motion prediction model of this embodiment includes the credible frame position of the moving object in three consecutive frames of images, and the output of the motion prediction model includes the predicted frame position of the moving object in the last two frames of images, that is, N is 3 and M is 2. Figure 7 The trusted trajectory includes trusted box 1 to trusted box 3. Figure 7 The predicted trajectory includes prediction box 2 to prediction box 5.

[0102] In one example, Figure 7 As shown, the electronic device splices the position information of the detection frame 1 of the moving object in the three copies of the first frame image, and inputs it into the motion prediction model. After being processed by the motion prediction model, the position information of the prediction frame 2 of the moving object in the second frame image and the position information of the prediction frame 3 of the moving object in the third frame image are obtained. Then, the electronic device can determine the position information of the credible frame 2 according to the position information of the prediction frame 2 and the position information of the detection frame 2. For example, the electronic device obtains the position information of the credible frame 2 by performing a weighted sum operation on the position information of the prediction frame 2 and the position information of the detection frame 2. The operation principle can refer to S604 below. It should be understood that in this example, since there is no credible frame position information initially, the electronic device can construct the input of the model according to the position information of the detection frame 1. Exemplarily, after obtaining the position information of the prediction frame 3, the electronic device can adjust the camera focus position according to the position information of the prediction frame 3 to capture the third frame image.

[0103] In one example, Figure 7As shown, the electronic device splices the position information of two detection frames 1 and the position information of a trusted frame 2, and inputs the spliced information into a motion prediction model. After being processed by the motion prediction model, the position information of a predicted frame 3' of the moving object in the third frame image and the position information of a predicted frame 4 of the moving object in the fourth frame image are obtained. Further, the electronic device can determine the position information of a trusted frame 3 based on the position information of the predicted frame 3' and the position information of a detection frame 3. For example, the electronic device performs a weighted summation operation on the position information of the predicted frame 3' and the position information of the detection frame 3 to obtain the position information of the trusted frame 3. It should be understood that in this example, since there is less than three pieces of trusted frame position information, the electronic device can construct the input of the model based on the position information of the detection frame 1 and the position information of the trusted frame 2. Exemplarily, after obtaining the position information of the predicted frame 4, the electronic device can adjust the camera focus position according to the position information of the predicted frame 4 to capture the fourth frame image.

[0104] In one example, as Figure 7 shown, the electronic device splices the position information of a detection frame 1, the position information of a trusted frame 2, and the position information of a trusted frame 3, and inputs the spliced information into a motion prediction model. After being processed by the motion prediction model, the position information of a predicted frame 4' of the moving object in the fourth frame image and the position information of a predicted frame 5 of the moving object in the fifth frame image are obtained. Further, the electronic device can determine the position information of a trusted frame 4 based on the position information of the predicted frame 4' and the position information of a detection frame 4. For example, the electronic device performs a weighted summation operation on the position information of the predicted frame 4' and the position information of the detection frame 4 to obtain the position information of the trusted frame 4. It should be understood that in this example, since there is less than three pieces of trusted frame position information, the electronic device can construct the input of the model based on the position information of the detection frame 1, the position information of the trusted frame 2, and the position information of the trusted frame 3. Exemplarily, after obtaining the position information of the predicted frame 5, the electronic device can adjust the camera focus position according to the position information of the predicted frame 5 to capture the fifth frame image.

[0105] In one example, as Figure 7 shown, the electronic device splices the position information of trusted frames 2 to 4, and inputs the spliced information into a motion prediction model. After being processed by the motion prediction model, the position information of a predicted frame 5' of the moving object in the fifth frame image and the position information of a predicted frame 6 of the moving object in the sixth frame image are obtained. Further, the electronic device can determine the position information of a trusted frame 5 based on the position information of the predicted frame 5' and the position information of a detection frame 5. For example, the electronic device performs a weighted summation operation on the position information of the predicted frame 5' and the position information of the detection frame 5 to obtain the position information of the trusted frame 5. Exemplarily, after obtaining the position information of the predicted frame 6, the electronic device can adjust the camera focus position according to the position information of the predicted frame 6 to capture the sixth frame image.

[0106] In some embodiments, if the electronic device determines the position information of the prediction box of the moving object in the i-th frame image, the following operations can be performed:

[0107] S604. The electronic device determines the position information of the credible box of the moving object in the i-th frame image according to the position information of the detection box of the moving object in the i-th frame image and the position information of the prediction box of the moving object in the i-th frame image.

[0108] In some embodiments, the position information of the detection box of the moving object in the i-th frame image includes: the coordinate positions of the upper left corner and the lower right corner of the detection box of the moving object in the i-th frame image. The position information of the prediction box of the moving object in the i-th frame image includes: the coordinate positions of the upper left corner and the lower right corner of the prediction box of the moving object in the i-th frame image. The position information of the credible box of the moving object in the i-th frame image includes: the coordinate positions of the upper left corner and the lower right corner of the credible box of the moving object in the i-th frame image.

[0109] It should be understood that the representation methods of the position information of the detection box, the prediction box, and the credible box should be consistent. For example, they are all represented by the coordinate positions of the upper left corner and the lower right corner of the box, or they are all represented by the coordinate positions of the upper right corner and the lower left corner of the box.

[0110] For convenience, the following uses the coordinate positions of the upper left corner and the lower right corner of the box to represent the position information of the box as an example to illustrate the solution.

[0111] In some embodiments, the electronic device determines the coordinate position of the upper left corner of the credible box of the moving object in the i-th frame image according to the coordinate position of the upper left corner of the detection box of the moving object in the i-th frame image and the coordinate position of the upper left corner of the prediction box of the moving object in the i-th frame image. Exemplarily, the electronic device performs a weighted sum operation on the coordinate position of the upper left corner of the detection box of the moving object in the i-th frame image and the coordinate position of the upper left corner of the prediction box of the moving object in the i-th frame image to obtain the coordinate position of the upper left corner of the credible box of the moving object in the i-th frame image.

[0112] In one example, the electronic device can determine the coordinate position (x, y) of the upper left corner of the credible box of the moving object in the i-th frame image through Formula 1 and Formula 2.

[0113] x = w1 × x predict + w2 × x detect Formula -

[0114] y = w1 × y predict + w2 × y detect Formula 2

[0115] In the formula, (x detect , y detect ) represents the coordinate position of the upper left corner of the detection box of the moving object in the i-th frame image, (xpredict , y predict ) represents the coordinate position of the upper left corner of the prediction box of the moving object in the i-th frame image, w1 represents the weight value of the prediction box, and w2 represents the weight value of the detection box.

[0116] In one example, the electronic device can determine the weight value w1 of the prediction box through Formula 3 and determine the weight value w2 of the detection box through Formula 4.

[0117]

[0118]

[0119] In the formula, Dist represents the distance value between the upper left corners of the detection box and the prediction box of the moving object in the i-th frame image, Speed represents the average moving speed of the upper left corner of the reliable box of the moving object in the N frames of images before the i-th frame image (i.e., the i-N-th frame image to the i-1-th frame image), and e is a constant.

[0120] Based on Formula 3 and Formula 4, it can be seen that when Speed remains unchanged, as Dist increases, the weight w1 of the prediction box gradually increases. When Dist remains unchanged, as Speed increases, the weight w2 of the detection box becomes larger.

[0121] In one example, Speed can be determined through the following Formula 5.

[0122]

[0123] In the formula, S1 represents the average value of the diagonal lengths of the reliable boxes of the moving object in the i-N-th frame image and the i-1-th frame image, that is, the average value of the diagonal lengths of the reliable boxes of the moving object in the first and last frame images among the N frames of images before the i-th frame image, and S2 represents the total moving distance value of the upper left corner of the reliable box of the moving object from the i-N-th frame image to the end of the i-1-th frame image. It should be noted that in the above example, i is a positive integer greater than N.

[0124] Exemplarily, referring to Figure 7 , taking N as 3 as an example, when determining the coordinate position of the upper left corner of the reliable box 5, the electronic device can determine w1 and w2 used to calculate the coordinate position of the upper left corner of the reliable box 5 based on the average moving speed of the upper left corners of the reliable boxes of the moving object in the previous 3 frames of images (i.e., the reliable boxes 2 to 4) and the distance value between the upper left corners of the detection box 5 and the prediction box 5' of the moving object in the 5th frame image.

[0125] In some embodiments, the electronic device determines the coordinate position of the lower right corner of the reliable bounding box of the moving object in the i-th frame image based on the coordinate position of the lower right corner of the detection bounding box of the moving object in the i-th frame image and the coordinate position of the lower right corner of the predicted bounding box of the moving object in the i-th frame image. Exemplarily, the electronic device performs a weighted summation operation on the coordinate position of the lower right corner of the detection bounding box of the moving object in the i-th frame image and the coordinate position of the lower right corner of the predicted bounding box of the moving object in the i-th frame image to obtain the coordinate position of the lower right corner of the reliable bounding box of the moving object in the i-th frame image. The implementation principle can refer to Formulas 1 to 5 of the previous embodiment and will not be elaborated here.

[0126] It should be noted that based on the distance value between the upper left corner of the detection bounding box of the moving object in the i-th frame image and the upper left corner of the predicted bounding box, and the average moving speed of the upper left corner of the reliable bounding box of the moving object in the N frames of images before the i-th frame image, the weight of the predicted bounding box (denoted as weight 1) and the weight of the detection bounding box (denoted as weight 2) are determined. The sum of weight 1 and weight 2 is 1. Based on the distance value between the lower right corner of the detection bounding box of the moving object in the i-th frame image and the lower right corner of the predicted bounding box, and the average moving speed of the lower right corner of the reliable bounding box of the moving object in the N frames of images before the i-th frame image, the weight of the predicted bounding box (denoted as weight 3) and the weight of the detection bounding box (denoted as weight 4) are determined, and the sum of weight 3 and weight 4 is 1. Generally, weight 1 is different from weight 3, and weight 2 is different from weight 4, that is, the weights determined based on the upper left corner are different from the weights determined based on the lower right corner.

[0127] It should be understood that as the moving speed of the moving object changes, the weights of the above-mentioned predicted bounding box and the detection bounding box can be dynamically changed. The electronic device determines the position information of the reliable bounding box of the moving object in the i-th frame image according to the dynamic weight value, providing data support for subsequent motion focusing.

[0128] S605. The electronic device stores the position information of the reliable bounding box of the moving object in the i-th frame image into the reliable trajectory set of the moving object.

[0129] Based on the above S601 to S605, the electronic device constructs a reliable trajectory set of the moving object. Exemplarily, the reliable trajectory set includes the position information of multiple reliable bounding boxes, such as Figure 7 reliable bounding box 2 to reliable bounding box 5. It should be understood that as time goes by, the camera continuously captures image frames, and the electronic device can store the position information of the reliable bounding boxes of the moving object in the continuously captured image frames into the reliable trajectory set.

[0130] In some embodiments, after S605, the following can be executed:

[0131] S606. The electronic device obtains the position information of the reliable bounding boxes of the moving object in N consecutive frames of images from the reliable trajectory set.

[0132] S607. The electronic device predicts the predicted box position information of the moving object in the M frames of images after the N frames of images according to the trusted box position information of the moving object in the continuous N frames of images in the trusted trajectory set.

[0133] In some embodiments, a trained motion prediction model is pre - installed in the electronic device. The electronic device can input the trusted box position information of the moving object in the continuous N frames of images in the trusted trajectory set into the motion prediction model to obtain the predicted box position information of the moving object in the M frames of images after the N frames of images.

[0134] Exemplarily, when the trusted trajectory set includes the trusted box position information of the moving object in three consecutive frames of images, such as Figure 7 the position information of trusted box 2 to trusted box 4, the electronic device can splice the position information of trusted box 2 to trusted box 4 and input it into the motion prediction model. After being processed by the motion prediction model, the position information of predicted box 5' of the moving object in the 5th frame of image and the position information of predicted box 6 of the moving object in the 6th frame of image are obtained. That is, the electronic device predicts the predicted box position information of the moving object in the latter two frames according to the trusted box position information of the moving object in three consecutive frames of images.

[0135] In this embodiment, the electronic device can predict the position of the moving object in the subsequent image based on the trusted box position of the moving object in multiple frames of images, rather than predicting the position of the moving object in the subsequent image based on the detection box position of the moving object in multiple frames of images. This is because there may be false detections in motion detection, that is, motion detection may be inaccurate. If the detection box position of the moving object in multiple frames of images is directly input into the motion prediction model, the predicted position of the moving object output by the model may have a large deviation from the actual position of the moving object, resulting in inaccurate prediction results.

[0136] By adopting the method of this embodiment, the trusted box position is obtained by fusing the detection box position and the predicted box position, which can play an anti - false - detection effect, improve the accuracy of motion prediction, guide the camera to accurately focus, and improve the image clarity of the moving object in the image. The motion prediction result is the predicted box position information.

[0137] S608. The electronic device adjusts the focusing position of the camera according to the predicted box position information of the moving object in the P - th frame of the M frames of images. S608 of this embodiment can refer to S304 of the foregoing embodiment and will not be elaborated here.

[0138] In the motion focusing method shown in the above embodiments, the electronic device determines the reliable bounding box position of the moving object in the i-th frame image based on the detection bounding box position of the moving object in the i-th frame image collected by the camera and the predicted bounding box position of the moving object in the i-th frame image. The electronic device can obtain the reliable bounding box positions of the moving object in a series of consecutive frames in this way. The electronic device can predict the predicted bounding box position of the moving object in the image after the N frames based on the reliable bounding box positions of the moving object in the consecutive N frames, and then adjust the focusing position of the camera based on the predicted bounding box position, which can further improve the accuracy of the predicted bounding box position, and thus improve the accuracy of camera focusing, enabling the camera to capture a clear moving object.

[0139] The motion focusing method shown in the above several embodiments can be applied to an electronic device with a photographing function. The electronic device can also be referred to as a terminal, a user equipment (UE), a mobile station (MS), a mobile terminal (MT), etc. The electronic device can be a mobile phone with a touch screen, a smart TV, a wearable device, a tablet computer (Pad), a computer with a wireless transceiver function, a virtual reality (VR) terminal device, an augmented reality (AR) terminal device, a wireless terminal in industrial control, a wireless terminal in self-driving, a wireless terminal in remote medical surgery, a wireless terminal in a smart grid, a wireless terminal in transportation safety, a wireless terminal in a smart city, a wireless terminal in a smart home, etc. The embodiments of the present application do not limit the specific technologies and specific device forms adopted by the electronic device.

[0140] Exemplarily, Figure 8 FIG. is a schematic structural diagram of an electronic device provided by an embodiment of the present application. As Figure 8 shown, the electronic device 100 includes: a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, a sensor 180, a button 190, a camera 193, and a display screen 194.

[0141] It can be understood that the structure illustrated in this embodiment does not constitute a specific limitation on the electronic device 100. In some embodiments, the electronic device 100 may include more or fewer components than those shown in the figure, or combine certain components, or split certain components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0142] It can be understood that the interface connection relationships between the modules illustrated in the embodiments are only illustrative and do not constitute a structural limitation on the electronic device 100. In some embodiments, the electronic device 100 may also adopt different interface connection methods or a combination of multiple interface connection methods in the above embodiments.

[0143] The processor 110 may include one or more processing units. Among them, different processing units may be independent devices or integrated in one or more processors. A memory may also be provided in the processor 110 for storing instructions and data. In the embodiments of the present application, the processor 110 may be used to call the computer program in the memory so that the electronic device executes the steps of the foregoing method embodiments to achieve motion focusing and improve the image clarity of the moving object captured by the electronic device.

[0144] The USB interface 130 is an interface that conforms to the USB standard specification, and may specifically be a Mini USB interface, a Micro USB interface, a USB Type C interface, etc. The USB interface 130 can be used to connect a charger to charge the terminal device, can also be used for data transmission between the terminal device and peripheral devices, and can also be used to connect headphones to play audio through the headphones.

[0145] The charging management module 140 is used to receive a charging input from a charger. The power management module 141 is used to connect the battery 142, the charging management module 140, and the processor 110.

[0146] The wireless communication function of the electronic device 100 can be implemented through antenna 1, antenna 2, the mobile communication module 150, the wireless communication module 160, the modulation and demodulation processor, and the baseband processor, etc. The mobile communication module 150 can provide solutions for wireless communications including 2G / 3G / 4G / 5G, etc. applied to the electronic device 100. The wireless communication module 160 can provide solutions for wireless communications including wireless local area networks (WLAN), Bluetooth, global navigation satellite system (GNSS), frequency modulation (FM), NFC, infrared technology (IR), etc. applied to the electronic device 100.

[0147] The electronic device 100 can implement a display function through a GPU, a display screen 194, an application processor, etc. The GPU is a microprocessor for image processing, connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 110 may include one or more GPUs, which execute instructions to generate or change display information.

[0148] The display screen 194 is used to display images, videos, etc. The display screen 194 includes a display panel. In some embodiments, the electronic device 100 may include one or N display screens 194, where N is a positive integer greater than 1.

[0149] The electronic device 100 can implement a shooting function through an image signal process (ISP) module, one or more cameras 193, a video codec, a GPU, one or more display screens 194, and an application processor, etc.

[0150] The camera 193 is used to capture still images or videos. In some embodiments, the electronic device 100 may include one or more cameras 193. The camera 193 includes a lens, an image sensor (such as a CMOS image sensor (complementary metal oxide semiconductor image sensor, abbreviated as CIS)), a motor, etc.

[0151] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 through the external memory interface 120 to implement a data storage function. For example, data files such as music, photos, and videos are saved in the external memory card.

[0152] The internal memory 121 can be used to store one or more computer programs, and the one or more computer programs include instructions. The processor 110 can execute the above instructions stored in the internal memory 121, so that the electronic device 100 performs various functional applications and data processing, etc.

[0153] The sensor 180 may include one or more of the following, for example: a pressure sensor, a gyroscope sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a distance sensor, a proximity light sensor, a fingerprint sensor, a temperature sensor, a touch sensor, an ambient light sensor, or a bone conduction sensor, etc.

[0154] The button 190 includes a power-on button, volume buttons, etc. The button 190 can be a mechanical button or a touch button. The electronic device 100 can receive button inputs and generate key signal inputs related to the user settings and function controls of the electronic device 100. For example, when the camera application is launched, the user can trigger the camera to take pictures or record videos by pressing the power-on button.

[0155] In addition, on top of the above components, an operating system also runs on the electronic device. For example, iOS operating system, Android operating system, or Windows operating system, etc. Application programs can be installed and run on the operating system.

[0156] The software system of the electronic device can adopt a layered architecture, event-driven architecture, microkernel architecture, microservices architecture, or cloud architecture. In the embodiments of the present application, taking the software system with a layered architecture as the Android system as an example, the software structure of the electronic device is exemplarily described. Figure 9 It is a schematic diagram of the software architecture and internal interaction of an electronic device provided by the embodiments of the present application. The layered architecture divides the software system of the electronic device into several layers, and each layer has a clear role and division of labor. The layers communicate with each other through software interfaces. Refer to Figure 9 , the electronic device includes: an application layer, an application framework layer, a hardware abstraction layer, and a driver layer.

[0157] The application layer may include a camera application, and the user can use the camera application to take pictures or videos. In the embodiments of the present application, the camera application provides various shooting modes, for example, a sports focus mode, etc. In some embodiments, the application package may also include applications such as a gallery, a calendar, a call, a map, a navigation, Bluetooth, music, and a video.

[0158] The application framework layer can provide application programming interfaces (APIs) and programming frameworks for the applications in the application layer. In the embodiments of the present application, the application framework layer may include a camera management module and a window manager.

[0159] The camera management module is responsible for managing camera device information. The camera application can obtain camera characteristics, such as parameters like the number of cameras and shooting capabilities, through the camera management module. The camera management module can also be used to transfer data between the camera application and the camera hardware abstraction layer. Exemplarily, refer to Figure 9 , in response to an operation to activate the sports focus mode, the camera application transmits a notification message (S901) to the camera hardware abstraction layer through the camera management module, so that after receiving the image captured by the camera, the camera hardware abstraction layer can invoke the motion detection module and the motion prediction module to guide the AF module to perform sports focus.

[0160] The window management module is responsible for managing the windows in the application and interacting with the user interface. For example, it is responsible for managing the camera application window and sending the window content to the display driver to drive the display screen to display. The window content may include the image captured by the camera and the focus frame (i.e., the prediction frame) for capturing the image.

[0161] The hardware abstraction layer is an interface layer located between the kernel layer and the hardware circuit. In the embodiments of the present application, the hardware abstraction layer includes a camera hardware abstraction layer, which may include a motion detection module, a motion prediction module, and an AF module. A motion detection algorithm is preset in the motion detection module for detecting the position of a moving object in the image. A motion prediction model is preset in the motion prediction module, which can be used to determine the predicted box position of the moving object in the next M frames of images according to the position of the moving object in the previous N frames of images, and can also be used to send the predicted box position of the moving object in the next M frames of images to the AF module. The AF module can be used to control the camera to focus according to the predicted box position of the moving object sent by the motion prediction module.

[0162] The driver layer provides drivers for different hardware devices. In the embodiments of the present application, the driver layer may include a camera driver and a display driver. The camera driver can be used to drive the camera of the electronic device to work. The display driver is used to drive the display screen of the electronic device to work.

[0163] The hardware layer includes hardware devices such as a camera and a display screen. The camera may include a lens, a lens motor, an image sensor, and an image signal processing (ISP), etc. Among them, the ISP can be used to perform image processing on the raw image output by the image sensor, including, for example, linear correction, noise removal, dead pixel removal, white balance, automatic exposure control, etc.

[0164] Exemplarily, continue to refer to Figure 9, after the camera application enables the motion focus mode, the camera hardware abstraction layer sends an exposure control instruction to the image sensor through the camera driver (S902) to instruct the image sensor to capture continuous image frames. The image sensor sends the continuous image frames to the ISP (S903). After the ISP processes the continuous image frames, the continuous image frames are sent to the camera hardware abstraction layer through the ISP driver (S904). The camera hardware abstraction layer performs motion detection on the continuous image frames by invoking the motion detection module to obtain the detection box positions of the moving objects in the continuous image frames. The motion detection module sends the detection box positions of the moving objects in the continuous image frames to the motion prediction module (S905). The motion prediction module can input the received detection box positions of the moving objects in the continuous image frames into a preset motion prediction model to obtain the predicted box positions of the moving objects in multiple frames of images after the continuous image frames. Optionally, in some embodiments, the input of the motion prediction model in the motion prediction module includes the reliable box positions of the moving objects in the continuous image frames.

[0165] The motion prediction module sends the predicted box position of the moving object in the Pth frame image (e.g., P = 2) after the continuous image frames to the AF module (S906), and the predicted box position of the moving object in the Pth frame image is the focus position. After receiving the focus position, the AF module sends focus parameters to the motor driver (S907), and the focus parameters include the focus position. The motor driver sends a focus control instruction to the lens motor according to the received focus parameters (S908) to drive the lens motor to move and achieve motion focus. Exemplarily, referring to Figure 7 , the motion prediction model in the motion detection module outputs the position of the predicted box 5' of the moving object in the 5th frame image and the position of the predicted box 6 of the moving object in the 6th frame image. The motion detection module sends the position of the predicted box 6 to the AF module, and the AF module can send the position of the predicted box 6 to the motor driver. The position of the predicted box 6 is the focus position.

[0166] After receiving the continuous image frames from the ISP, the camera hardware abstraction layer can send the continuous image frames to the camera application through the camera management module (S909). It should be understood that the camera application continuously receives the continuous image frames from the ISP, including the image frames captured by the camera after motion focus. After receiving the predicted box position of the moving object in the Pth frame image after the continuous image frames sent by the motion detection module, the AF module can send the predicted box position of the moving object in the Pth frame image to the camera application through the camera management module (S910). In some embodiments, the AF module can send the predicted box position of the moving object in the Pth frame image to the window manager (not shown) so that the window manager can perform content merging for display.

[0167] The camera application may send the P-th frame image and the predicted bounding box position of the moving object in the P-th frame image to the window manager (S911). After merging the P-th frame image and the predicted bounding box position of the moving object in the P-th frame image, the window manager sends the display content to the display driver. The display content includes the P-th frame image of the predicted bounding box of the moving object (S912), so as to drive the display to display the P-th frame image including the predicted bounding box of the moving object.

[0168] It can be understood that Figure 9 The modules included in each layer shown are the modules involved in the embodiments of the present application. The modules included in each layer do not constitute a limitation on the structure and module deployment hierarchy of the electronic device. In some embodiments, the electronic device may include more or fewer layers than shown, and each layer may include more or fewer components, which are not limited in the present application.

[0169] It should be noted that in the above embodiments, the "module" may be a software program, a hardware circuit, or a combination of the two to implement the above functions. The hardware circuit may include an application specific integrated circuit (ASIC), an electronic circuit, a processor (such as a shared processor, a dedicated processor, or a group of processors, etc.) for executing one or more software or firmware programs, a memory, a merging logic circuit, and / or other suitable components to support the described functions.

[0170] Therefore, the modules of each example described in the embodiments of the present application can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software form depends on the specific application and design constraints of the technical solution. A person skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0171] Based on the foregoing several embodiments, an embodiment of the present application proposes a motion focusing method. In response to an operation of turning on the motion focusing mode, the electronic device acquires N consecutive frames of images collected by the camera. Each frame of the N frames of images includes the same moving object, and N is a positive integer greater than 1; the electronic device acquires the reliable position information of the moving object in each frame of the N frames of images; the electronic device determines the predicted position information of the moving object in the M frames of images after the N frames of images according to the reliable position information of the moving object in the N frames of images, and M is a positive integer greater than 1; the electronic device controls the camera to perform motion focusing according to the predicted position information of the moving object in the P-th frame image of the M frames of images, and P is a positive integer less than or equal to M.

[0172] Exemplarily, referring to Figure 1As shown in Figure C, in response to an operation on the motion focus control 1032 acting on the window 1031 in the interface 103, the electronic device acquires N consecutive frames of images captured by the camera to perform subsequent motion focus methods. Taking a certain frame of the N frames of images as an example, the reliable position information of the moving object in this frame is determined based on the detected position information and predicted position information of the moving object in this frame. Exemplarily, referring to Figure 7 As shown, the position information of the reliable bounding box 2 of the moving object in the second frame of image is obtained by weighted summation of the position information of the detected bounding box 2 of the moving object in the second frame of image and the position information of the predicted bounding box 2 of the moving object in the second frame of image.

[0173] In the above embodiment, when the camera application of the electronic device enables the motion focus mode, the electronic device can construct a reliable trajectory of the moving object (i.e., the trajectory corresponding to the reliable positions of the moving object in consecutive image frames), predict the motion trajectory of the moving object in subsequent image frames (i.e., the predicted positions of the moving object in subsequent image frames) based on the reliable trajectory of the moving object in consecutive image frames, and control the camera to perform motion focus based on the predicted motion trajectory of the moving object in subsequent image frames. Compared with the electronic device predicting the motion trajectory of the moving object in subsequent image frames based on the detected trajectory of the moving object in consecutive image frames (i.e., the trajectory corresponding to the detected positions of the moving object in consecutive image frames), the above method can reduce the problem that the inaccuracy of the detected trajectory caused by misdetection of the motion detection algorithm affects the accuracy of the predicted trajectory of the moving object, improve the accuracy of camera focus, and thus improve the clarity of the images of the moving objects captured by the electronic device.

[0174] In an alternative embodiment, obtaining the reliable position information of the moving object in each frame of the N frames of images includes: obtaining the detected position information of the moving object in the i-th frame of the N frames of images, and the predicted position information of the moving object in the i-th frame of images, where the i-th frame of images is any image frame other than the first frame of the N frames of images; determining the reliable position information of the moving object in the i-th frame of images according to the detected position information and predicted position information of the moving object in the i-th frame of images. Where i is an integer greater than 1 and less than or equal to N.

[0175] The above embodiment shows a way to construct the reliable positions of the moving object in consecutive image frames. The electronic device can determine the reliable position of the moving object in this frame according to the detected position and predicted position of the moving object in each frame of image. The reliability of the reliable position of the moving object in this frame is higher than the detected position of the moving object in this frame. The constructed reliable positions of the moving object in consecutive image frames can be used by the electronic device to predict the motion trajectory of the moving object in subsequent image frames.

[0176] In an optional embodiment, the detection position information of the moving object in the i-th frame image includes the coordinate position of the first detection box of the moving object in the i-th frame image, and the predicted position information of the moving object in the i-th frame image includes the coordinate position of the first prediction box of the moving object in the i-th frame image; according to the detection position information and the predicted position information of the moving object in the i-th frame image, the credible position information of the moving object in the i-th frame image is determined, including: by performing a fusion operation on the coordinate position of the first detection box and the coordinate position of the first prediction box, the coordinate position of the first credible box of the moving object in the i-th frame image is obtained, and the credible position information of the moving object in the i-th frame image includes the coordinate position of the first credible box.

[0177] In the above embodiment, the electronic device performs coordinate position fusion based on the coordinate positions of the upper left corners in the detection box and the prediction box of the i-th frame image to obtain the coordinate position of the upper left corner in the credible box of the i-th frame image, and performs coordinate position fusion based on the coordinate positions of the lower right corners in the detection box and the prediction box of the i-th frame image to obtain the coordinate position of the lower right corner in the credible box of the i-th frame image. The coordinate positions of the upper left corner and the lower right corner of the credible box of the i-th frame image can uniquely determine the position of the credible box in the i-th frame image, and this position can be used for subsequent prediction of the motion trajectory of the moving object.

[0178] In an optional embodiment, the coordinate position of the first detection box includes the coordinate position of the upper left corner in the first detection box and the coordinate position of the lower right corner in the first detection box, and the coordinate position of the first prediction box includes the coordinate position of the upper left corner in the first prediction box and the coordinate position of the lower right corner in the first prediction box; by performing a fusion operation on the coordinate position of the first detection box and the coordinate position of the first prediction box, the coordinate position of the first credible box of the moving object in the i-th frame image is obtained, including: by performing a weighted summation operation on the coordinate position of the upper left corner in the first detection box and the coordinate position of the upper left corner in the first prediction box, the coordinate position of the upper left corner in the first credible box is obtained; and by performing a weighted summation operation on the coordinate position of the lower right corner in the first detection box and the coordinate position of the lower right corner in the first prediction box, the coordinate position of the lower right corner in the first credible box is obtained.

[0179] In the above embodiment, the electronic device performs a weighted summation operation on the coordinate positions of the upper left corners of the detection box and the prediction box in the same frame image respectively, and performs a weighted summation operation on the coordinate positions of the lower right corners of the detection box and the prediction box in the same frame image respectively, to obtain the coordinate positions of the upper left corner and the lower right corner of the credible box in this frame image, for subsequent prediction of the motion trajectory of the moving object.

[0180] In an optional embodiment, by performing a weighted summation operation on the coordinate position of the upper left corner in the first detection box and the coordinate position of the upper left corner in the first prediction box, the coordinate position of the upper left corner in the first credible box is obtained, including:

[0181] Determine the coordinate position (x, y) of the upper left corner in the first credible box through the following formula:

[0182] x = w1 × x predict + w2 × x detect

[0183] y = w1 × y predict + w2 × y detect

[0184] In the formula, (x detect , y detect ) represents the coordinate position of the upper left corner in the first detection box, (x predict , y predict ) represents the coordinate position of the upper left corner in the first prediction box, w1 represents the weight value of the first prediction box, and w2 represents the weight value of the first detection box.

[0185] In the above embodiment, the electronic device performs weighted summation operations on the coordinate positions of the upper left corners of the detection box and the prediction box in the x-axis of the same frame of image respectively, and performs weighted summation operations on the coordinate positions of the upper left corners of the detection box and the prediction box in the y-axis of the same frame of image, so as to obtain the coordinate position of the upper left corner of the credible box in this frame of image, which is used for subsequent prediction of the motion trajectory of the moving object. It should be understood that based on the principle of the above formula, the coordinate positions of the lower right corners of the credible box in the x-axis and y-axis of this frame of image can be obtained respectively.

[0186] In an alternative embodiment, the sum of the weight value of the first prediction box and the weight value of the first detection box is 1, and the weight value of the first prediction box is determined according to the first distance value and the first moving speed of the moving object; the first distance value is the distance value between the detection box of the moving object and the upper left corner in the prediction box in the i-th frame of image, and the first moving speed is the average moving speed of the upper left corner of the credible box of the moving object from the (i - N)-th frame of image to the (i - 1)-th frame of image.

[0187] It should be understood that the sum of the weight value of the first prediction box and the weight value of the first detection box is 1. After the electronic device determines the weight value of the first prediction box, the weight value of the first detection box can be determined. The calculation of the weight value of the first prediction box can refer to Formula 3, and the calculation of the weight value of the first detection box can refer to Formula 4.

[0188] It should be understood that when the moving speed of the credible box of the moving object (which can correspond to the above first moving speed) remains unchanged, as the position offset value between the detection box and the prediction box (which can correspond to the above first distance value) increases, the weight value of the prediction box gradually increases, and the weight value of the detection box gradually decreases. When the position offset value between the detection box and the prediction box remains unchanged, as the moving speed of the credible box of the moving object increases, the weight value of the detection box gradually increases, and the weight value of the prediction box gradually decreases.

[0189] In an alternative embodiment, determining predicted position information of a moving object in M frames of images after N frames of images according to the reliable position information of the moving object in the N frames of images includes: inputting the reliable position information of the moving object in the N frames of images into a motion prediction model, and after being processed by the motion prediction model, obtaining the predicted position information of the moving object in M frames of images after the N frames of images; the motion prediction model is trained by using a neural network model.

[0190] In the above embodiment, the electronic device inputs the reliable frame coordinate positions of the moving object in the N frames of images into a preset motion prediction model, and after model operation, obtains the predicted coordinate positions of the moving object in M frames of images after the N frames of images. The model input is not the N frames of images, but the position data of the moving object in the images, which can improve the processing speed of the motion prediction model, and thus can shorten the focusing time of the device and improve the motion focusing speed of the device.

[0191] In some embodiments, the motion prediction model can be trained by using a lightweight fully connected neural network model. By adopting a lightweight model structure, on the one hand, it can save the storage space of the device, and on the other hand, it can also improve the execution speed of the model. In addition, due to the simple model structure and small amount of model operations, the device power consumption can be reduced.

[0192] In an alternative embodiment, after the electronic device controls the camera to perform motion focusing according to the predicted position information of the moving object in the P-th frame of the M frames of images, the method further includes: the electronic device acquires the P-th frame of the M frames of images collected by the camera; the electronic device displays the P-th frame of the image and the predicted position information of the moving object in the P-th frame of the image, and the image clarity of the moving object in the P-th frame of the image is greater than a preset clarity.

[0193] In the above embodiment, when the electronic device displays the P-th frame of the image, it simultaneously displays the predicted frame position of the moving object in the P-th frame of the image, and this predicted frame position is the focusing position when the P-th frame of the image is acquired. Since this predicted frame position is determined based on the foregoing method, the deviation between the predicted frame position and the actual position of the moving object in the P-th frame of the image is less than a preset value, so that the image clarity of the moving object in the P-th frame of the image is greater than the preset clarity. It should be understood that subsequent image frames are focused according to the foregoing method, which can improve the image clarity of the moving object captured by the electronic device as a whole.

[0194] An embodiment of this application further provides an electronic device, which includes: one or more processors and a memory. The memory is coupled to the one or more processors. The memory is used to store computer program code, and the computer program code includes computer instructions. The one or more processors call the computer instructions to cause the electronic device to execute the steps in the foregoing method embodiments. The implementation principle and technical effects are similar to those of the foregoing related embodiments, and will not be elaborated herein.

[0195] An embodiment of this application further provides a chip system, which is applied to an electronic device. The chip system includes one or more processors, and the one or more processors are used to call computer instructions to cause the electronic device to execute the steps in the foregoing method embodiments. The implementation principle and technical effects are similar to those of the foregoing related embodiments, and will not be elaborated herein.

[0196] An embodiment of this application further provides a computer-readable storage medium, which includes computer instructions. When the computer instructions run on an electronic device, the electronic device is caused to execute the steps in the foregoing method embodiments. The implementation principle and technical effects are similar to those of the foregoing related embodiments, and will not be elaborated herein.

[0197] An embodiment of this application further provides a computer program product, which includes computer program code. When the computer program code runs on an electronic device, the electronic device is caused to execute the steps in the foregoing method embodiments. The implementation principle and technical effects are similar to those of the foregoing related embodiments, and will not be elaborated herein.

[0198] The method described in the foregoing embodiments can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. If implemented in software, the functions can be stored on or transmitted over a computer-readable medium as one or more instructions or code. The computer-readable medium can include a computer storage medium and a communication medium, and can also include any medium that can transfer a computer program from one place to another. The storage medium can be any target medium accessible by a computer.

[0199] In some embodiments, a computer-readable medium may include RAM, ROM, a compact disc read-only memory (CD-ROM), or other optical disc storage, magnetic disk storage, or any other medium that is intended to store the desired program code in the form of instructions or data structures and that can be accessed by a computer. Also, any connection is properly termed a computer-readable medium. For example, if software is transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, Digital Subscriber Line (DSL), or wireless technologies such as infrared, radio, and microwave, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of the medium. As used herein, disk and disc include optical discs, laser discs, optical discs, Digital Versatile Discs (DVDs), floppy disks, and Blu-ray discs, where disks typically reproduce data magnetically, while discs reproduce data optically using lasers. Combinations of the above should also be included within the scope of computer-readable media.

[0200] Embodiments of the present application are described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and combinations of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processing unit of a general-purpose computer, special-purpose computer, embedded processor, or other programmable device to generate a machine, such that the instructions executed by the processing unit of the computer or other programmable data processing device produce means for implementing the functions specified in Figure 1 one or more of the flows or multiple flows and / or blocks Figure 1 one or more of the blocks or multiple blocks.

[0201] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data that have been authorized by the user or fully authorized by all parties. Moreover, the collection, use, and processing of the relevant data need to comply with relevant laws, regulations, and standards, and corresponding operation entrances are provided for users to select authorization or rejection.

[0202] In the above specific embodiments, the purpose, technical solution and beneficial effects of the present invention have been further described in detail. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent substitutions, improvements, etc. made on the basis of the technical solution of the present invention should be included in the protection scope of the present invention.

Claims

1. A method for motion focusing, characterized in that, Including: In response to an operation of enabling the motion focus mode, the electronic device acquires N consecutive images captured by the camera, each of the N images including the same moving object, and N is a positive integer greater than 1; The electronic device acquires the reliable position information of the moving object in each of the N images; The electronic device determines the predicted position information of the moving object in M images after the N images according to the reliable position information of the moving object in the N images, and M is a positive integer greater than 1; The electronic device controls the camera to perform motion focus according to the predicted position information of the moving object in the P-th image of the M images, and P is a positive integer less than or equal to M.

2. The method according to claim 1, characterized in that, Acquiring the reliable position information of the moving object in each of the N images includes: Acquiring the detected position information of the moving object in the i-th image of the N images, and the predicted position information of the moving object in the i-th image, where the i-th image is any image frame other than the first image frame in the N images; According to the detected position information and the predicted position information of the moving object in the i-th image, determining the reliable position information of the moving object in the i-th image.

3. The method according to claim 2, wherein The detected position information of the moving object in the i-th image includes the coordinate position of the first detection frame of the moving object in the i-th image, and the predicted position information of the moving object in the i-th image includes the coordinate position of the first prediction frame of the moving object in the i-th image; According to the detected position information and the predicted position information of the moving object in the i-th image, determining the reliable position information of the moving object in the i-th image includes: By performing a fusion operation on the coordinate positions of the first detection frame and the first prediction frame, obtaining the coordinate position of the first reliable frame of the moving object in the i-th image, and the reliable position information of the moving object in the i-th image includes the coordinate position of the first reliable frame.

4. The method according to claim 3, wherein The coordinate position of the first detection frame includes the coordinate position of the upper left corner and the coordinate position of the lower right corner in the first detection frame, and the coordinate position of the first prediction frame includes the coordinate position of the upper left corner and the coordinate position of the lower right corner in the first prediction frame; By performing a fusion operation on the coordinate positions of the first detection frame and the first prediction frame, obtaining the coordinate position of the first reliable frame of the moving object in the i-th image includes: By performing a weighted summation operation on the coordinate position of the upper left corner in the first detection frame and the coordinate position of the upper left corner in the first prediction frame, obtaining the coordinate position of the upper left corner in the first reliable frame; and By performing a weighted summation operation on the coordinate position of the lower right corner in the first detection frame and the coordinate position of the lower right corner in the first prediction frame, obtaining the coordinate position of the lower right corner in the first reliable frame.

5. The method according to claim 4, characterized in that By performing a weighted summation operation on the coordinate positions of the upper left corner in the first detection box and the coordinate positions of the upper left corner in the first prediction box, the coordinate positions of the upper left corner in the first confidence box are obtained, including: Determine the coordinate positions (x, y) of the upper left corner in the first confidence box through the following formula: x = w1 × x predict + w2 × x detect y = w1×y predict +w2×y detect where (x detect , y detect ) represents the coordinate position of the upper left corner in the first detection box, (x predict , y predict ) represents the coordinate position of the upper left corner in the first prediction box, w1 represents the weight value of the first prediction box, and w2 represents the weight value of the first detection box.

6. The method according to claim 5, wherein: The sum of the weight value of the first prediction box and the weight value of the first detection box is 1, and the weight value of the first prediction box is determined according to the first distance value and the first moving speed of the moving object; The first distance value is the distance value between the detection box of the moving object in the i-th frame image and the upper left corner in the prediction box, and the first moving speed is the average moving speed of the upper left corner of the confidence box of the moving object from the (i - N)-th frame image to the (i - 1)-th frame image.

7. The method according to any one of claims 1 to 6, wherein: Determine the predicted position information of the moving object in the M frame images after the N frame images according to the reliable position information of the moving object in the N frame images, including: Input the reliable position information of the moving object in the N frame images into a motion prediction model, and after being processed by the motion prediction model, obtain the predicted position information of the moving object in the M frame images after the N frame images; The motion prediction model is trained by using a neural network model.

8. The method according to any one of claims 1 to 7, wherein: After the electronic device controls the camera to perform motion focusing according to the predicted position information of the moving object in the P-th frame image of the M frame images, the method further includes: The electronic device acquires the P-th frame image in the M frame images collected by the camera; The electronic device displays the P-th frame image and the predicted position information of the moving object in the P-th frame image, and the image clarity of the moving object in the P-th frame image is greater than a preset clarity.

9. An electronic device, characterized in that, The electronic device includes: one or more processors and a memory; The memory is coupled to the one or more processors, and the memory is used to store computer program code, the computer program code includes computer instructions, and the one or more processors call the computer instructions to cause the electronic device to execute the method according to any one of claims 1 to 8.

10. A chip system, characterized in that, The chip system is applied to an electronic device, the chip system includes one or more processors, and the one or more processors are used to call computer instructions to cause the electronic device to execute the method according to any one of claims 1 to 8.

11. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes computer instructions, and when the computer instructions run on the electronic device, the electronic device is caused to execute the method according to any one of claims 1 to 8.

12. A computer program product, characterized in that, The computer program product includes computer program code, and when the computer program code runs on the electronic device, the electronic device is caused to execute the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Focusing method, mobile terminal, and computer readable storage medium

    CN109167910A

  • Automatic focusing method and device, electronic equipment and computer readable storage medium

    CN115037869A

  • Track determination method and device, computer readable storage medium and electronic equipment

    CN116645396A

  • Focusing processing method, electronic equipment and storage medium

    CN117135451A