Motion focusing method, and electronic device and storage medium

By obtaining the trusted position information of moving objects in continuous image frames in electronic devices, constructing trusted trajectories and predicting subsequent motion trajectories, the focus delay problem of electronic devices when shooting moving objects is solved, and image clarity is improved.

WO2025145891A1PCT designated stage expired Publication Date: 2025-07-10HONOR DEVICE CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/140092
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-02
Filing Date
2024-12-17
Publication Date
2025-07-10

AI Technical Summary

Technical Problem

When shooting moving objects, existing electronic devices are limited by focus delays, resulting in low image clarity and cannot meet users' shooting needs.

Method used

By acquiring the continuous N-frame images collected by the camera, using motion detection algorithms and prediction models, the trusted position information of the moving object is determined, the trusted trajectory of the moving object is constructed, its motion trajectory in subsequent image frames is predicted, and the camera is controlled to perform motion focusing based on this.

Benefits of technology

It improves the accuracy of camera focus, improves the image clarity of moving objects captured by electronic devices, and reduces the problem of inaccurate prediction trajectory caused by false detection of motion detection algorithms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024140092_10072025_PF_FP_ABST
    Figure CN2024140092_10072025_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of terminals. Provided are a motion focusing method, and an electronic device and a storage medium. In the method, when a camera application on an electronic device enables a motion focusing mode, the electronic device may predict a motion trajectory of a moving object in subsequent image frames by means of constructing a credible trajectory of the moving object and on the basis of credible positions of the moving object in consecutive image frames, and on the basis of the predicted motion trajectory of the moving object in the subsequent image frames, control a camera to perform motion focusing. By means of the method, the accuracy of camera focusing can be improved, and thus the definition of images of a moving object photographed by means of an electronic device is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Motion focusing method, electronic device and storage medium

[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on January 2, 2024, with application number 202410009866.2 and application name “Motion Focusing Method, Electronic Device and Storage Medium”, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The present application relates to the field of terminal technology, and in particular to a motion focusing method, an electronic device, and a storage medium. Background Art

[0003] With the widespread use of smart devices, it has become common for users to use electronic devices such as mobile phones and tablets to take photos, and users have increasingly higher requirements for image quality. For example, users use electronic devices to shoot people in motion or objects moving at high speeds, and they expect to capture the moving objects in clear images.

[0004] However, in actual applications, due to the focus delay of the electronic device, the image clarity of the moving object captured by the electronic device is low and cannot meet the user's shooting needs. Summary of the Invention

[0005] The embodiments of the present application provide a motion focusing method, an electronic device, and a storage medium, which can improve the accuracy of camera focusing, thereby improving the image clarity of moving objects captured by the electronic device.

[0006] In a first aspect, an embodiment of the present application proposes a motion focus method. In response to an operation of turning on a motion focus mode, the electronic device obtains N consecutive frames of images captured by a camera, each of the N frames of images includes the same moving object, and N is a positive integer greater than 1; the electronic device obtains trusted position information of the moving object in each of the N frames of images; the electronic device determines the predicted position information of the moving object in M ​​frames after the N frames of images based on the trusted position information of the moving object in the N frames of images, where M is a positive integer greater than 1; the electronic device controls the camera to perform motion focus based on the predicted position information of the moving object in the P-th frame of the M frames of images, where P is a positive integer less than or equal to M.

[0007] Exemplarily, as shown in c of FIG1 , in response to the operation of the motion focus control 1032 acting on the window 1031 in the interface 103, the electronic device obtains N consecutive frames of images captured by the camera to perform a subsequent motion focus method. Taking a certain frame image among the N frames of images as an example, the credible position information of the moving object in the frame image is determined based on the detected position information and the predicted position information of the moving object in the frame image. Exemplarily, as shown in FIG7 , the position information of the credible frame 2 of the moving object in the second frame image is obtained by weighted summation based on the position information of the detected frame 2 of the moving object in the second frame image and the position information of the predicted frame 2 of the moving object in the second frame image.

[0008] In the above embodiment, when the motion focus mode is turned on in the camera application of the electronic device, the electronic device can construct a credible trajectory of the moving object (i.e., the trajectory corresponding to the credible position of the moving object in consecutive image frames), predict the motion trajectory of the moving object in subsequent image frames (i.e., the predicted position of the moving object in subsequent image frames) based on the credible trajectory of the moving object in the consecutive image frames, and control the camera to perform motion focus based on the predicted motion trajectory of the moving object in the subsequent image frames. Compared with the electronic device predicting the motion trajectory of the moving object in subsequent image frames based on the detection trajectory of the moving object in consecutive image frames (i.e., the trajectory corresponding to the detected position of the moving object in consecutive image frames), the above method can reduce the problem of inaccurate detection trajectory due to misdetection of the motion detection algorithm, which affects the accuracy of the predicted trajectory of the moving object, improve the accuracy of camera focus, and thus improve the image clarity of the moving object captured by the electronic device.

[0009] In an optional embodiment of the first aspect, obtaining credible position information of a moving object in each of N frames of images includes: obtaining detected position information of the moving object in an i-th frame of the N frames of images, and predicted position information of the moving object in the i-th frame of the N frames of images, where the i-th frame of the N frames of images is any frame other than the first frame of the N frames of images; and determining credible position information of the moving object in the i-th frame of the image based on the detected position information and the predicted position information of the moving object in the i-th frame of the image, where i is an integer greater than 1 and less than or equal to N.

[0010] The above embodiment shows a method for constructing a trusted position of a moving object in continuous image frames. The electronic device can determine the trusted position of the moving object in each frame image based on the detected position and predicted position of the moving object in the frame image. The credibility of the trusted position of the moving object in the frame image is higher than the detected position of the moving object in the frame image. The constructed trusted position of the moving object in the continuous image frames can be used by the electronic device to predict the motion trajectory of the moving object in subsequent image frames.

[0011] In an optional embodiment of the first aspect, the detection position information of the moving object in the i-th frame image includes the coordinate position of the first detection frame of the moving object in the i-th frame image, and the predicted position information of the moving object in the i-th frame image includes the coordinate position of the first prediction frame of the moving object in the i-th frame image; based on the detection position information and the predicted position information of the moving object in the i-th frame image, the credible position information of the moving object in the i-th frame image is determined, including: obtaining the coordinate position of the first credible frame of the moving object in the i-th frame image by performing a fusion operation on the coordinate position of the first detection frame and the coordinate position of the first prediction frame, and the credible position information of the moving object in the i-th frame image includes the coordinate position of the first credible frame.

[0012] In the above embodiment, the electronic device performs coordinate fusion based on the coordinate positions of the upper left corners of the detection and prediction frames of the i-th image frame to obtain the coordinate position of the upper left corner of the credible frame of the i-th image frame, and performs coordinate fusion based on the coordinate positions of the lower right corners of the detection and prediction frames of the i-th image frame to obtain the coordinate position of the lower right corner of the credible frame of the i-th image frame. The coordinate positions of the upper left corner and lower right corner of the credible frame of the i-th image frame uniquely determine the position of the credible frame in the i-th image frame, and this position can be used to predict the motion trajectory of subsequent moving objects.

[0013] In an optional embodiment of the first aspect, the coordinate position of the first detection frame includes the coordinate position of the upper left corner of the first detection frame and the coordinate position of the lower right corner of the first detection frame, and the coordinate position of the first prediction frame includes the coordinate position of the upper left corner of the first prediction frame and the coordinate position of the lower right corner of the first prediction frame; by performing a fusion operation on the coordinate position of the first detection frame and the coordinate position of the first prediction frame, the coordinate position of the first credible frame of the moving object in the i-th frame image is obtained, including: performing a weighted sum operation on the coordinate position of the upper left corner of the first detection frame and the coordinate position of the upper left corner of the first prediction frame to obtain the coordinate position of the upper left corner of the first credible frame; and performing a weighted sum operation on the coordinate position of the lower right corner of the first detection frame and the coordinate position of the lower right corner of the first prediction frame to obtain the coordinate position of the lower right corner of the first credible frame.

[0014] In the above embodiment, the electronic device performs a weighted sum operation on the coordinate positions of the upper left corner of the detection box and the prediction box in the same frame image, and performs a weighted sum operation on the coordinate positions of the lower right corner of the detection box and the prediction box in the same frame image, so as to obtain the coordinate positions of the upper left corner and the lower right corner of the credible box in the frame image, which are used for subsequent motion trajectory prediction of the moving object.

[0015] In an optional embodiment of the first aspect, obtaining the coordinate position of the upper left corner of the first credible frame by performing a weighted sum operation on the coordinate position of the upper left corner of the first detection frame and the coordinate position of the upper left corner of the first prediction frame includes:

[0016] The coordinate position (x, y) of the upper left corner of the first credible frame is determined by the following formula: x = w1 × x predict +w2×x detect y=w1×y predict +w2×y detect

[0017] In the formula, (x detect ,y detect ) represents the coordinate position of the upper left corner of the first detection frame, (x predict ,y predict ) represents the coordinate position of the upper left corner of the first prediction box, w1 represents the weight value of the first prediction box, and w2 represents the weight value of the first detection box.

[0018] In the above embodiment, the electronic device performs a weighted summation of the x-axis coordinate positions of the upper left corners of the detection frame and the prediction frame in the same frame image, as well as a weighted summation of the y-axis coordinate positions of the upper left corners of the detection frame and the prediction frame in the same frame image, to obtain the coordinate position of the upper left corner of the credible frame in the frame image for subsequent motion trajectory prediction of the moving object. It should be understood that based on the principles of the above formula, the x-axis and y-axis coordinate positions of the lower right corner of the credible frame in the frame image can be obtained.

[0019] In an optional embodiment of the first aspect, the sum of the weight value of the first prediction box and the weight value of the first detection box is 1, and the weight value of the first prediction box is determined based on the first distance value and the first moving speed of the moving object; the first distance value is the distance value between the detection box of the moving object in the i-th frame image and the upper left corner of the prediction box, and the first moving speed is the average moving speed of the upper left corner of the credible box of the moving object from the iN-th frame image to the i-1-th frame image.

[0020] It should be understood that the sum of the weight value of the first prediction frame and the weight value of the first detection frame is 1. After determining the weight value of the first prediction frame, the electronic device can determine the weight value of the first detection frame. The weight value of the first prediction frame can be calculated with reference to Formula 3, and the weight value of the first detection frame can be calculated with reference to Formula 4.

[0021] It should be understood that, when the moving speed of the credible frame of the moving object remains unchanged (which may correspond to the first moving speed described above), as the position offset value between the detection frame and the prediction frame increases (which may correspond to the first distance value described above), the weight value of the prediction frame gradually increases, while the weight value of the detection frame gradually decreases. When the position offset value between the detection frame and the prediction frame remains unchanged, as the moving speed of the credible frame of the moving object increases, the weight value of the detection frame gradually increases, while the weight value of the prediction frame gradually decreases.

[0022] In an optional embodiment of the first aspect, based on the trusted position information of the moving object in N frames of images, the predicted position information of the moving object in M ​​frames of images after the N frames of images is determined, including: inputting the trusted position information of the moving object in the N frames of images into a motion prediction model, and obtaining the predicted position information of the moving object in M ​​frames of images after processing by the motion prediction model; the motion prediction model is obtained by training using a neural network model.

[0023] In the above embodiment, the electronic device inputs the credible frame coordinate position of the moving object in N frames of images into a preset motion prediction model, and after model calculation, obtains the predicted coordinate position of the moving object in M ​​frames of images after the N frames of images. The model input is not N frames of images, but the position data of the moving object in the image, which can improve the processing speed of the motion prediction model, thereby shortening the focusing time of the device and improving the motion focusing speed of the device.

[0024] In some embodiments, the motion prediction model can be trained using a lightweight, fully connected neural network model. This lightweight model structure not only saves device storage space but also improves model execution speed. Furthermore, due to its simple structure and low computational complexity, it reduces device power consumption.

[0025] In an optional embodiment of the first aspect, after the electronic device controls the camera to perform motion focus based on the predicted position information of the moving object in the P-th frame image in the M-frame image, the method also includes: the electronic device obtains the P-th frame image in the M-frame image captured by the camera; the electronic device displays the P-th frame image and the predicted position information of the moving object in the P-th frame image, and the image clarity of the moving object in the P-th frame image is greater than the preset clarity.

[0026] In the above embodiment, when the electronic device displays the P-th frame image, it also displays the predicted frame position of the moving object in the P-th frame image. This predicted frame position is the focus position when the P-th frame image is captured. Because this predicted frame position is determined based on the aforementioned method, the deviation between this predicted frame position and the actual position of the moving object in the P-th frame image is less than a preset value, resulting in the image clarity of the moving object in the P-th frame image being greater than a preset clarity. It should be understood that subsequent image frames are all focused according to the aforementioned method, which can generally improve the image clarity of the moving object captured by the electronic device.

[0027] In a second aspect, an embodiment of the present application provides an electronic device, comprising: a memory and a processor, wherein the processor is configured to call a computer program in the memory to execute a method as described in any one of the first aspects.

[0028] In a third aspect, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program runs on an electronic device, the electronic device executes the method described in any one of the first aspects.

[0029] In a fourth aspect, an embodiment of the present application provides a chip, the chip including a processor, the processor being used to call a computer program in a memory so that the chip executes the method as described in any one of the first aspects.

[0030] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program, which, when executed, enables a computer to execute the method as described in any one of the first aspects.

[0031] It should be understood that the second to fifth aspects of the present application correspond to the technical solutions of the first aspect of the present application, and the beneficial effects achieved by each aspect and the corresponding feasible implementation methods are similar and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] FIG1 is a schematic diagram of changes in a mobile phone interface provided by an embodiment of the present application;

[0033] FIG2 is a first schematic diagram of the principle of the motion focusing method provided in an embodiment of the present application;

[0034] FIG3 is a first flow chart of a motion focusing method according to an embodiment of the present application;

[0035] FIG4 is a schematic diagram of the structure of a motion prediction model provided in an embodiment of the present application;

[0036] FIG5 is a second schematic diagram of the principle of the motion focusing method provided in an embodiment of the present application;

[0037] FIG6 is a second flow chart of the motion focusing method provided in an embodiment of the present application;

[0038] FIG7 is a schematic diagram showing the principle of motion prediction provided by an embodiment of the present application;

[0039] FIG8 is a schematic structural diagram of an electronic device provided in an embodiment of the present application;

[0040] FIG9 is a schematic diagram of the software architecture and internal interactions of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0041] To facilitate the clear description of the technical solutions of the embodiments of the present application, in the embodiments of the present application, the words "first" and "second" are used to distinguish between identical or similar items with substantially the same functions and effects. Those skilled in the art will understand that the words "first" and "second" do not limit the quantity or execution order, and the words "first" and "second" do not necessarily mean different.

[0042] It should be noted that in the embodiments of this application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described in this application as "exemplary" or "for example" should not be construed as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.

[0043] In the embodiments of the present application, "at least one" refers to one or more, and "more" refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent: the existence of A alone, the existence of A and B at the same time, and the existence of B alone, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items (kind / individual)" or similar expressions refers to any combination of these items, including any combination of single items (kind / individual) or plural items (kind / individual). For example, at least one of a, b or c (kind / individual) can represent: a, b, c, ab, ac, bc, or abc, where a, b, c can be single or multiple.

[0044] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc., data including images or videos, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0045] To facilitate understanding, the professional terms involved in the embodiments of this application are first explained below.

[0046] First, motion focus is a focus mode for electronic devices, primarily suitable for capturing dynamic subjects, such as athletes on a sports field or birds in flight. This focus mode can quickly and accurately capture moving subjects and maintain focus. As the subject moves within the frame, the electronic device automatically adjusts the camera's focus parameters, such as focus position and focal length, to ensure the subject remains sharp. The embodiments of this application optimize the algorithm for the motion focus mode of electronic devices.

[0047] Second, subject focus is another focus mode on electronic devices, primarily suitable for photographing portraits, animals, or objects with specific features. This focus mode automatically identifies the subject in the frame and prioritizes focus on it. For example, when photographing a person, if the person's face is facing the device's camera, subject focus priority will automatically focus on the face; if the person is moving or turning, the focus will automatically track and maintain focus on the subject.

[0048] Third, motion detection algorithms are an important area in computer vision, primarily used to detect and track moving objects, such as people, vehicles, and animals. Commonly used algorithms include bounding box (or detection box)-based detection algorithms, single-stage detection algorithms, and keypoint-based detection algorithms, which are not limited to these in the present embodiment.

[0049] Fourth, a lightweight neural network is a neural network model that performs deep learning tasks under resource-constrained conditions. Its design aims to reduce the number of model parameters and computational complexity while maintaining sufficient performance to meet the needs of specific application scenarios. In the embodiment of the present application, the motion prediction model can be obtained by training using a lightweight neural network, for example, a lightweight fully connected neural network. The embodiment of the present application does not limit the structure of the lightweight neural network.

[0050] Fifth, the linear layer is one of the basic layer types in a neural network, also known as a fully connected layer. The purpose of the linear layer is to multiply the input vector by the weight matrix and add the bias vector to generate a new output vector.

[0051] Sixth, ResBlock is the basic structural unit in the residual network ResNet, commonly known as the residual block. It passes the information of the previous layer to the subsequent layer by introducing skip connections to alleviate the gradient vanishing problem. For example, in the ResBlock module shown in Figure 4, the input of the ResBlock module is summed with the output of the last LeakyReLu in the ResBlock module, thereby passing the information of the previous layer along the depth of the network, which can improve network performance. LeakyReLu is a neural network activation function. The ResBlock module shown in Figure 4 also includes Dropout, which is a regularization method for neural network models. During the training process, by randomly ignoring some neurons and allowing some neurons to stop working with a certain probability, it can prevent the model from overfitting the training data and improve the model's generalization ability to new data.

[0052] In the scenario of photographing a moving object, the user can turn on the motion focus mode in the camera application of the electronic device, triggering the electronic device to photograph the moving object in the motion focus mode.

[0053] For example, Figure 1 is a schematic diagram of the changes in the mobile phone interface provided by an embodiment of the present application. As shown in Figure 1, in response to the user's operation on the camera application icon 101 on the main interface of the mobile phone, the mobile phone displays interface 102. Interface 102 includes a shooting mode selection area 1021, a control 1022 for viewing an album, a shooting control 1023, a control 1024 for switching cameras, etc. The shooting mode selection area 1021 displays night scene mode, portrait mode, photo mode, video mode, etc. The user can select a shooting mode in area 1021 to shoot. As shown in Figure 1 b, the camera application can turn on the photo mode by default.

[0054] In some embodiments, in response to a user's operation on camera mode 1025 in area 1021, the mobile phone displays interface 103, which includes window 1031. Window 1031 displays motion focus controls 1032. In one example, the mobile phone's camera application prioritizes motion focus mode by default. Therefore, when the mobile phone displays window 1031, it may display motion focus controls 1032 by default, even though motion focus mode is not yet enabled.

[0055] In one example, in response to a user's operation on control 1032, the mobile phone turns on motion focus mode by default. As shown in FIG1(d), the mobile phone displays interface 104, which includes window 1051. Window 1051 displays a control 1052 for prioritizing motion focus and a control 1053 for prioritizing subject focus. Control 1052 for prioritizing motion focus is displayed in bold to indicate to the user that the camera application has turned on motion focus mode.

[0056] The user can switch focus modes in window 1051. In one example, in response to a user operating control 1053 on window 1051, the camera application of the mobile phone can switch from motion focus mode to subject focus mode. Accordingly, the subject focus control 1053 is displayed in bold (not shown) to indicate to the user that the camera application has turned on subject focus mode.

[0057] In some embodiments, after acquiring an image captured by a camera, the electronic device detects the position of a moving object in the image and adjusts the focus position of the camera based on the detected position to guide the camera to capture subsequent images.

[0058] Exemplarily, FIG2 is a schematic diagram of the principle of the motion focus method provided in an embodiment of the present application. As shown in FIG2 a, in the motion focus mode, the electronic device obtains the i-th frame image captured by the camera, and the i-th frame image includes a moving vehicle. Based on the motion detection algorithm, the electronic device can obtain the position information of the vehicle in the i-th frame image. The position information of the vehicle includes the coordinate position of the detection frame 21 of the vehicle in the i-th frame image. For example, the coordinate positions of the two points of the upper left corner and the lower right corner of the detection frame 21 in the i-th frame image are used as the coordinate positions of the detection frame 21. When the shooting angle remains unchanged, the electronic device can adjust the focus position of the camera based on the position information of the vehicle in the i-th frame image. When the vehicle moving speed is less than the first threshold, the camera of the electronic device can focus based on the coordinate position of the detection frame 21 in the i-th frame image to capture the i+1-th frame image.

[0059] However, if the vehicle's moving speed exceeds the second threshold, the actual position of the vehicle in the i+1th frame image captured by the electronic device's camera will deviate from the detected position of the vehicle in the i+1th frame image. As shown in FIG2(b), detection frame 22 can be viewed as the projection of the vehicle's position in the i+1th frame image, and detection frame 23 represents the actual position of the vehicle in the i+1th frame image. The electronic device focuses based on detection frame 22 to capture the i+1th frame image. Due to the inaccurate camera focus position, the clarity of the vehicle in the i+1th frame image is low, for example, the vehicle clarity is less than the preset clarity.

[0060] To address the above issues, embodiments of the present application provide a motion focus method. After acquiring multiple frames of images captured by a camera, an electronic device detects the position of a moving object in the multiple frames. Based on the multiple detected positions of the moving object, a preset motion prediction model is used to obtain a predicted position of the moving object in images subsequent to the multiple frames. The electronic device can adjust the focus position based on the predicted position of the moving object to guide the camera in capturing subsequent images.

[0061] In this embodiment, the electronic device can obtain the historical motion trajectory of the moving object based on multiple detected positions of the moving object, and obtain the motion trajectory of the moving object in subsequent images based on the historical motion trajectory combined with the motion prediction model. Compared with the previous embodiment, since the focus position is more accurate, the image clarity of the moving object in subsequent images can be improved.

[0062] The motion focusing method provided in the embodiment of the present application is described in detail below with reference to FIG. 3 to FIG. 5 .

[0063] For example, FIG3 is a flow chart of a motion focusing method provided in an embodiment of the present application. The method can be applied to electronic devices with shooting and image processing functions. As shown in FIG3, the motion focusing method of this embodiment includes:

[0064] S301. In response to an operation of starting a motion focus mode, the electronic device acquires N consecutive frames of images captured by a camera, where N is an integer greater than or equal to 2.

[0065] 1 c, in response to a user's operation on the motion focus control 1032, such as a user clicking the control 1032, the electronic device turns on the motion focus mode. After the electronic device turns on the motion focus mode, the motion focus method provided in this embodiment can be executed.

[0066] S302. The electronic device performs motion detection on N frames of images, and obtains detection frame position information of the moving object in each of the N frames of images.

[0067] In some embodiments, the electronic device may perform motion detection on N frames of images based on a bounding box detection algorithm to obtain detection box position information of the moving object in each of the N frames of images.

[0068] The detection frame position information of the moving object may include the coordinate positions of the diagonal points of the detection frame of the moving object in the image, such as the coordinate positions of the upper left and lower right corners of the detection frame in the image, or the coordinate positions of the upper right and lower left corners of the detection frame in the image.

[0069] The embodiment of the present application does not limit the motion detection algorithm.

[0070] S303. The electronic device predicts the predicted frame position information of the moving object in the M frames following the N frames based on the detection frame position information of the moving object in the N frames.

[0071] In some embodiments, the electronic device may use a motion prediction model to predict the predicted box position information of the moving object in M ​​frames of imagery following N frames of imagery. In this embodiment, the electronic device is pre-installed with a trained motion prediction model. The electronic device may input the detection box position information of the moving object in N frames of imagery into the motion prediction model to obtain the predicted box position information of the moving object in M ​​frames of imagery following N frames of imagery.

[0072] The predicted box position information of the moving object may include the coordinate positions of the diagonal points of the predicted box of the moving object in the image, such as the coordinate positions of the upper left and lower right corners of the predicted box in the image, or the coordinate positions of the upper right and lower left corners of the predicted box in the image.

[0073] In some embodiments, the motion prediction model can be trained using a lightweight neural network structure, such as a lightweight fully connected neural network structure. This lightweight model structure can save device storage space and improve model execution speed. Furthermore, due to its simple structure and low computational complexity, it can reduce device power consumption.

[0074] Exemplarily, Figure 4 is a structural diagram of the motion prediction model provided in an embodiment of the present application. As shown in Figure 4, the input of the motion prediction model includes the coordinate positions of the detection frames of the moving objects in N consecutive frames of images, that is, the coordinate positions of the N detection frames in Figure 4. The output of the motion prediction model includes the coordinate positions of the prediction frames of the moving objects in M ​​frames of images after N frames of images, that is, the coordinate positions of the M prediction frames in Figure 4. N is a positive integer greater than 1, and M is a positive integer greater than 1. Exemplarily, M is less than N. For example, N is 8 and M is 4. For another example, N is 3 and M is 2. The embodiment of the present application does not limit N and M.

[0075] In this embodiment, the input of the motion prediction model is the coordinate position of the detection frame of the moving object in N frames of image. The coordinate position of the detection frame of the moving object in each frame of image can be recorded as (x1, y1, x2, y2). For example, (x1, y1) can be the coordinate position of the upper left corner of the detection frame in the image, and (x2, y2) can be the coordinate position of the upper left corner of the detection frame in the image. Of course, in some embodiments, the coordinate positions of the other two diagonal points of the detection frame, such as the upper right corner and the lower left corner in the image, can also be used to mark the coordinate position of the detection frame. It can be seen that the input of the motion prediction model includes N×4 numbers. Correspondingly, the output of the motion prediction model includes M×4 numbers. Simple model input and lightweight model structure can improve the processing speed of the motion prediction model, thereby shortening the focusing time and improving the motion focusing speed of the device.

[0076] In some embodiments, as shown in FIG4 , the motion prediction model includes two linear layers and Q ResBlock modules, where the two linear layers are respectively denoted as Linear Layer 1 and Linear Layer 2. The output of Linear Layer 1 is connected to the first ResBlock module of the Q ResBlock modules, and the output of the last ResBlock module of the Q ResBlock modules is connected to Linear Layer 2. Where Q is a positive integer, for example, Q can be 3. Exemplarily, a ResBlock module may include a Linear Layer 3, a Dropout Layer 1, a LeakyReLu Layer 1, a Linear Layer 4, a Dropout Layer 2, and a LeakyReLu Layer 2, which are connected in sequence. The output of the ResBlock module is the sum of the output of LeakyReLu 3 and the input of Linear Layer 3.

[0077] The motion prediction model shown in this embodiment includes a ResBlock module. The ResBlock module can transmit the front-layer information of the model along the depth of the network, thereby improving the motion prediction capability of the model.

[0078] The motion prediction model shown in this embodiment is only an example. The motion prediction model may include more or fewer layers than shown in the figure to implement the motion prediction function. The structure of the motion prediction model is not limited in this embodiment of the application.

[0079] S304. The electronic device adjusts the focus position of the camera according to the predicted frame position information of the moving object in the P-th frame image among the M frames of image. M is a positive integer, and P is a positive integer less than or equal to M.

[0080] For example, if N is 8 and M is 4, the electronic device may obtain the predicted frame position of the moving object in each of the next four frames based on the detection frame position of the moving object in the first eight frames. For example, the electronic device may adjust the camera's focus position based on the predicted frame position of the moving object in the first of the next four frames. For another example, the electronic device may adjust the camera's focus position based on the predicted frame position of the moving object in the second of the next four frames.

[0081] The following describes in detail how the electronic device adjusts the camera focus position based on the predicted frame position information, in conjunction with Figure 5. For example, Figure 5 is a second schematic diagram of a scenario of the motion focus method provided in an embodiment of the present application. For ease of understanding, Figure 5 uses N as 3 and M as 2 as an example to illustrate the solution.

[0082] In motion focus mode, the electronic device obtains three consecutive frames of images captured by the camera, such as the i-2th frame image, the i-1th frame image, and the i-th frame image in Figure 5. The electronic device first obtains the coordinate positions of the detection frames of the vehicle in these three frames based on the motion detection algorithm, such as detection frames 51, 52, and 53 in Figure 5. Based on the dotted arrows in the figure, it can be seen that the vehicle's motion trajectory moves from right to left along the lane. After obtaining the coordinate positions of the detection frames of the vehicle in the three frames, the electronic device inputs the coordinate positions of detection frames 51, 52, and 53 into the motion prediction model. After the motion prediction model is calculated, the model outputs the coordinate positions of the prediction frames of the vehicle in the next two frames, such as prediction frames 54' and 55' in Figure 5.

[0083] In some embodiments, if P is set to 1, the electronic device may adjust the camera's focus position based on the predicted frame position information of the moving object in the first frame of M frames to capture the first frame after the previous N frames. For example, referring to FIG. 5 , the electronic device may focus based on the coordinate position of the predicted frame 54 ′ to capture the (i+1)th frame, thereby improving the image clarity of the vehicle in the (i+1)th frame.

[0084] In the embodiments of the present application, the camera of the electronic device includes an image sensor. The image sensor's image output speed is typically measured in terms of frame rate, i.e., frames per second (fps). The higher the frame rate, the faster the image sensor outputs images. For example, the image sensor's frame rate is 30 fps, meaning the image sensor outputs one frame approximately every 33 milliseconds (ms).

[0085] For example, ideally, the electronic device adjusts the camera focus position to focus position 1 based on the predicted frame position information of the moving object in the first frame of the M-frame image, and the image sensor outputs the first frame of the M-frame image based on focus position 1. However, if the image sensor's image output speed is greater than a preset speed, after the electronic device adjusts the camera focus position to focus position 1 based on the predicted frame position information of the moving object in the first frame of the M-frame image, the image sensor may have already output the first frame of the M-frame image based on the previous focus position. In this case, the image sensor will output the second frame of the M-frame image based on focus position 1. Ideally, the electronic device adjusts the camera focus position to focus position 2 based on the predicted frame position information of the moving object in the second frame of the M-frame image, and the image sensor outputs the second frame of the M-frame image based on focus position 2. Since the focus position is inaccurate when capturing the second frame of the M-frame image, the image clarity of the moving object in the second frame of the M-frame image will be affected.

[0086] To address the above issue, a suitable P value can be configured for the electronic device, for example, P is configured to be 2. After the electronic device obtains the predicted frame position of the moving object in the M frames following the N frames based on S303, the electronic device may focus on the predicted frame position of the moving object in the second frame of the M frames instead of using the predicted frame position of the moving object in the first frame of the M frames. This can avoid the problem of inaccurate focus position caused by the delay in running the prediction, and improve the image clarity of the moving object in the image output by the image sensor.

[0087] For example, referring again to FIG. 5 , when the electronic device acquires prediction frame 54 ′, the image sensor may have already output the (i+1)th frame. Upon acquiring prediction frame 54 ′ and prediction frame 55 ′, the electronic device can focus based on the coordinates of prediction frame 55 ′ to capture the (i+2)th frame, thereby improving the image clarity of the vehicle in the (i+2)th frame.

[0088] In the motion focus method described in the above embodiment, the electronic device predicts the predicted frame position of the moving object in images subsequent to the N frames captured by the camera based on the detection frame position of the moving object. Based on the predicted frame position of the moving object in the N frames subsequent to the N frames, the electronic device adjusts the camera's focus position. Compared to adjusting the camera's focus position based on the detection frame position, adjusting the camera's focus position based on the predicted frame position improves the image clarity of the moving object in the images captured by the camera due to its more accurate focus position.

[0089] Based on the aforementioned embodiments, the electronic device can determine the predicted frame position of the moving object in multiple subsequent frames of images based on the detection frame position of the moving object in multiple frames of historical imagery. Given the capabilities of the motion detection algorithm, in some cases, the detection frame position obtained by the electronic device based on the motion detection algorithm may be inaccurate. For example, if there are false detections, this will affect the accuracy of the predicted frame position, and thus affect the accuracy of camera focus, resulting in reduced image clarity of the moving object in the images captured by the camera.

[0090] To further improve the accuracy of the predicted frame position, an embodiment of the present application proposes a motion focusing method, in which the electronic device can fuse the coordinate positions of the detection frame position of the moving object in the i-th frame image with the predicted frame position of the moving object in the i-th frame image to obtain the credible frame position of the moving object in the i-th frame image. Based on this method, the electronic device can obtain the credible frame position of the moving object in multiple frames. Based on the credible frame position of the moving object in the multiple frames, the electronic device can predict the predicted frame position of the moving object in the image after the multiple frames, and focus based on the predicted frame position.

[0091] In the above method, the credible frame position is obtained by fusing the detection frame position and the prediction frame position. The electronic device can predict the motion trajectory of the moving object based on the credible frame position, which can reduce the problem of inaccurate prediction frame position caused by false detection of the motion detection algorithm.

[0092] It should be noted that the electronic device can determine the detection frame trajectory of the moving object in consecutive image frames based on the detection frame position of the moving object in multiple image frames. The detection frame trajectory can be simply referred to as the detection trajectory. The electronic device can determine the predicted frame trajectory of the moving object in consecutive image frames based on the predicted frame position of the moving object in multiple image frames. The predicted frame trajectory can be simply referred to as the predicted trajectory. The electronic device can determine the credible frame trajectory of the moving object in consecutive image frames based on the credible frame position of the moving object in multiple image frames. The credible frame trajectory can be simply referred to as the credible trajectory. The electronic device can store the detection trajectory, predicted trajectory, and credible trajectory of the moving object in consecutive image frames.

[0093] The motion focusing method provided by this embodiment is described in detail below in conjunction with FIG6 . For example, FIG6 is a second flow chart of the motion focusing method provided by this embodiment of the application. This method can be applied to electronic devices with shooting and image processing functions. As shown in FIG6 , the motion focusing method of this embodiment includes:

[0094] S601. In response to the operation of starting the motion focus mode, the electronic device obtains the i-th frame image captured by the camera.

[0095] In this embodiment, the i-th frame image may be any frame image captured by a camera after the electronic device turns on the motion focus mode.

[0096] S602. The electronic device performs motion detection on the i-th frame image and obtains the detection frame position information of the moving object in the i-th frame image. S602 of this embodiment can refer to S302 of the above embodiment and will not be described here for brevity.

[0097] S603. The electronic device determines whether there is predicted frame position information of a moving object in the i-th frame image.

[0098] In some embodiments, if the i-th frame image is the first frame image captured by the camera after the electronic device turns on the motion diagonal mode, since the electronic device did not perform motion detection and motion prediction on the image captured by the camera before turning on the motion focus mode, the electronic device does not have the predicted frame position information of the moving object in the first frame image.

[0099] In some embodiments, if the i-th frame image is an image subsequent to the first frame image captured by the camera after the electronic device turns on motion focus mode, the electronic device stores the predicted frame position information of the moving object in the i-th frame image. For example, after the electronic device turns on motion focus mode, the camera captures the second frame image, and the electronic device stores the predicted frame position information of the moving object in the second frame image. The predicted frame position information of the moving object in the second frame image can be determined by the electronic device based on the detection frame position information of the moving object in the first frame image. For details, please refer to the first example in the embodiment of Figure 7.

[0100] The following is an example of the process of obtaining the predicted frame position information of the moving object in the image after the first frame image captured by the camera, with reference to Figure 7. For example, Figure 7 is a schematic diagram of the principle of motion prediction provided by an embodiment of the present application. Taking the first four frames of images captured by the camera after the electronic device turns on the motion focus mode as an example, as shown in Figure 7, the electronic device obtains the detection frame of the moving object in each of the first four frames of images based on the motion detection algorithm, such as detection frame 1 to detection frame 4, and the detection trajectory of the moving object includes detection frame 1 to detection frame 4. A motion prediction model is pre-installed in the electronic device. The input of the motion prediction model of this embodiment includes the credible frame position of the moving object in 3 consecutive frames of images, and the output of the motion prediction model includes the predicted frame position of the moving object in the last 2 frames of images, that is, N is 3 and M is 2. The credible trajectory in Figure 7 includes credible frames 1 to credible frames 3, and the predicted trajectory in Figure 7 includes prediction frames 2 to prediction frames 5.

[0101] In one example, as shown in Figure 7, the electronic device concatenates the three pieces of position information for the detection frame 1 of the moving object in the first image frame and inputs them into a motion prediction model. After processing by the motion prediction model, the electronic device obtains the position information for the prediction frame 2 of the moving object in the second image frame and the position information for the prediction frame 3 of the moving object in the third image frame. Furthermore, the electronic device can determine the position information for the credible frame 2 based on the position information for the prediction frame 2 and the position information for the detection frame 2. For example, the electronic device performs a weighted sum operation on the position information for the prediction frame 2 and the position information for the detection frame 2 to obtain the position information for the credible frame 2. The operation principle is described in S604 below. It should be understood that in this example, since the credible frame position information does not initially exist, the electronic device can construct the model input based on the position information for the detection frame 1. For example, after obtaining the position information for the prediction frame 3, the electronic device can adjust the camera focus position based on the position information for the prediction frame 3 to capture the third image frame.

[0102] In one example, as shown in Figure 7, the electronic device concatenates the two pieces of position information for detection frame 1 and the position information for trusted frame 2, and then inputs the concatenated information into a motion prediction model. After processing by the motion prediction model, the electronic device obtains the position information for predicted frame 3' of the moving object in the third image frame, and the position information for predicted frame 4 of the moving object in the fourth image frame. Furthermore, the electronic device can determine the position information for trusted frame 3 based on the position information for predicted frame 3' and the position information for detection frame 3. For example, the electronic device can perform a weighted sum operation on the position information for predicted frame 3' and the position information for detection frame 3 to obtain the position information for trusted frame 3. It should be understood that in this example, since there are fewer than three pieces of trusted frame position information, the electronic device can construct the model input based on the position information for detection frame 1 and the position information for trusted frame 2. For example, after obtaining the position information for predicted frame 4, the electronic device can adjust the camera focus position based on the position information for predicted frame 4 to capture the fourth image frame.

[0103] In one example, as shown in Figure 7, the electronic device concatenates the position information of detection frame 1, the position information of trusted frame 2, and the position information of trusted frame 3, and then inputs the concatenated information into a motion prediction model. After processing by the motion prediction model, the position information of predicted frame 4' for the moving object in the fourth image frame and the position information of predicted frame 5 for the moving object in the fifth image frame are obtained. Furthermore, the electronic device can determine the position information of trusted frame 4 based on the position information of predicted frame 4' and the position information of detection frame 4. For example, the electronic device can perform a weighted sum operation on the position information of predicted frame 4' and the position information of detection frame 4 to obtain the position information of trusted frame 4. It should be understood that in this example, since there are fewer than three pieces of trusted frame position information, the electronic device can construct the model input based on the position information of detection frame 1, trusted frame 2, and trusted frame 3. For example, after obtaining the position information of predicted frame 5, the electronic device can adjust the camera focus position based on the position information of predicted frame 5 to capture the fifth image frame.

[0104] In one example, as shown in FIG7 , the electronic device concatenates the position information of credible frames 2 to 4 and inputs the information into a motion prediction model. After processing by the motion prediction model, the position information of the prediction frame 5' of the moving object in the fifth frame image and the position information of the prediction frame 6 of the moving object in the sixth frame image are obtained. Furthermore, the electronic device can determine the position information of the credible frame 5 based on the position information of the prediction frame 5' and the position information of the detection frame 5. For example, the electronic device obtains the position information of the credible frame 5 by performing a weighted sum operation on the position information of the prediction frame 5' and the position information of the detection frame 5. Exemplarily, after obtaining the position information of the prediction frame 6, the electronic device can adjust the camera focus position based on the position information of the prediction frame 6 to capture the sixth frame image.

[0105] In some embodiments, if the electronic device determines that there is predicted frame position information of a moving object in the i-th frame image, the electronic device may execute:

[0106] S604. The electronic device determines the credible frame position information of the moving object in the i-th frame image according to the detection frame position information of the moving object in the i-th frame image and the prediction frame position information of the moving object in the i-th frame image.

[0107] In some embodiments, the detection frame position information of the moving object in the i-th frame image includes: the coordinate positions of the upper left corner and the lower right corner of the detection frame of the moving object in the i-th frame image. The prediction frame position information of the moving object in the i-th frame image includes: the coordinate positions of the upper left corner and the lower right corner of the prediction frame of the moving object in the i-th frame image. The credible frame position information of the moving object in the i-th frame image includes: the coordinate positions of the upper left corner and the lower right corner of the credible frame of the moving object in the i-th frame image.

[0108] It should be understood that the position information of the detection box, prediction box and credible box should be represented in the same way, for example, they should all be represented by the coordinate positions of the upper left corner and lower right corner of the box, or they should all be represented by the coordinate positions of the upper right corner and lower left corner of the box.

[0109] For convenience, the following solution is explained using the coordinate positions of the upper left corner and lower right corner of the box as an example to represent the position information of the box.

[0110] In some embodiments, the electronic device determines the coordinate position of the upper left corner of the credible frame of the moving object in the i-th frame image based on the coordinate position of the upper left corner of the detection frame of the moving object in the i-th frame image and the coordinate position of the upper left corner of the prediction frame of the moving object in the i-th frame image. Exemplarily, the electronic device obtains the coordinate position of the upper left corner of the credible frame of the moving object in the i-th frame image by performing a weighted sum operation on the coordinate position of the upper left corner of the detection frame of the moving object in the i-th frame image and the coordinate position of the upper left corner of the prediction frame of the moving object in the i-th frame image.

[0111] In one example, the electronic device can determine the coordinate position (x, y) of the upper left corner of the credible frame of the moving object in the i-th frame image by using Formula 1 and Formula 2. predict +w2×x detect Formula 1

[0112] y=w1×y predict +w2×y detect Formula 2

[0113] In the formula, (x detect ,y detect ) represents the coordinate position of the upper left corner of the detection frame of the moving object in the i-th frame image, (xpredict ,y predict ) represents the coordinate position of the upper left corner of the prediction box of the moving object in the i-th frame image, w1 represents the weight value of the prediction box, and w2 represents the weight value of the detection box.

[0114] In one example, the electronic device may determine the weight value w1 of the prediction box using Formula 3, and determine the weight value w2 of the detection box using Formula 4.

[0115] Where Dist represents the distance between the upper left corner of the detection box and the upper left corner of the prediction box of the moving object in the i-th frame image, Speed ​​represents the average moving speed of the upper left corner of the credible box of the moving object in the N frames before the i-th frame image (i.e., the iN-th frame image to the i-1-th frame image), and e is a constant.

[0116] Based on Formula 3 and Formula 4, we can see that when Speed ​​remains unchanged, as Dist increases, the weight w1 of the prediction box gradually increases. When Dist remains unchanged, as Speed ​​increases, the weight w2 of the detection box increases.

[0117] In one example, Speed ​​can be determined by the following formula 5.

[0118] Where S1 represents the average of the diagonal lengths of the credible box of the moving object in the iNth frame and the diagonal lengths of the credible box of the moving object in the i-1th frame, that is, the average of the diagonal lengths of the credible box of the moving object in the first and last frames of the N frames before the i-th frame. S2 represents the total movement distance of the upper left corner of the credible box of the moving object from the iNth frame to the i-1th frame. It should be noted that in the above example, i is a positive integer greater than N.

[0119] Exemplarily, referring to Figure 7, taking N as 3 as an example, when determining the coordinate position of the upper left corner of the credible box 5, the electronic device can determine w1 and w2 for calculating the coordinate position of the upper left corner of the credible box 5 based on the average moving speed of the upper left corner of the credible boxes of the moving object in the first three frames (i.e., credible boxes 2 to credible boxes 4), and the distance value between the upper left corner of the detection box 5 of the moving object in the fifth frame image and the upper left corner of the prediction box 5'.

[0120] In some embodiments, the electronic device determines the coordinate position of the lower right corner of the credible frame of the moving object in the i-th frame image based on the coordinate position of the lower right corner of the detection frame of the moving object in the i-th frame image and the coordinate position of the lower right corner of the prediction frame of the moving object in the i-th frame image. Exemplarily, the electronic device obtains the coordinate position of the lower right corner of the credible frame of the moving object in the i-th frame image by performing a weighted sum operation on the coordinate position of the lower right corner of the detection frame of the moving object in the i-th frame image and the coordinate position of the lower right corner of the prediction frame of the moving object in the i-th frame image. The implementation principle can be referred to Formulas 1 to 5 of the previous embodiment and will not be expanded here.

[0121] It should be noted that the weight of the prediction frame (denoted as weight 1) and the weight of the detection frame (denoted as weight 2) are determined based on the distance between the upper left corner of the detection frame of the moving object in the i-th frame image and the upper left corner of the prediction frame, and the average movement speed of the upper left corner of the credible frame of the moving object in the N frames before the i-th frame image. The sum of weight 1 and weight 2 is 1. The weight of the prediction frame (denoted as weight 3) and the weight of the detection frame (denoted as weight 4) are determined based on the distance between the lower right corner of the detection frame of the moving object in the i-th frame image and the lower right corner of the prediction frame, and the average movement speed of the lower right corner of the credible frame of the moving object in the N frames before the i-th frame image. The sum of weight 3 and weight 4 is 1. Typically, weight 1 is different from weight 3, and weight 2 is different from weight 4. That is, the weight determined based on the upper left corner is different from the weight determined based on the lower right corner.

[0122] It should be understood that as the moving speed of the moving object changes, the weight of the above-mentioned prediction frame and the weight of the detection frame can change dynamically. The electronic device determines the credible frame position information of the moving object in the i-th frame image based on the dynamic weight value to provide data support for subsequent motion focusing.

[0123] S605. The electronic device stores the credible frame position information of the moving object in the i-th frame image into the credible trajectory set of the moving object.

[0124] Based on steps S601 to S605 above, the electronic device constructs a trusted trajectory set for the moving object. Exemplarily, the trusted trajectory set includes the location information of multiple trusted boxes, such as trusted boxes 2 to 5 in Figure 7. It should be understood that as the camera continuously captures image frames over time, the electronic device may store the location information of the trusted boxes of the moving object in the continuously captured image frames in the trusted trajectory set.

[0125] In some embodiments, after S605 , the following operations may be performed:

[0126] S606. The electronic device obtains the credible frame position information of the moving object in N consecutive frames of images from the credible trajectory set.

[0127] S607. The electronic device predicts the predicted frame position information of the moving object in M ​​frames of image following the N frames of image based on the credible frame position information of the moving object in the consecutive N frames of image in the credible trajectory set.

[0128] In some embodiments, a trained motion prediction model is pre-installed in the electronic device. The electronic device can input the credible box position information of the moving object in N consecutive frames of images in the credible trajectory set into the motion prediction model to obtain the predicted box position information of the moving object in M ​​frames of images after the N frames of images.

[0129] Exemplarily, when the credible trajectory set includes the credible box position information of the moving object in three consecutive frames of images, such as the position information of credible boxes 2 to credible boxes 4 in Figure 7, the electronic device can splice the position information of credible boxes 2 to credible boxes 4 and input the information into the motion prediction model. After processing by the motion prediction model, the position information of the predicted box 5' of the moving object in the fifth frame image and the position information of the predicted box 6 of the moving object in the sixth frame image are obtained. That is, the electronic device predicts the predicted box position information of the moving object in the next two frames based on the credible box position information of the moving object in the three consecutive frames of images.

[0130] In this embodiment, the electronic device can predict the position of the moving object in subsequent images based on the trusted frame position of the moving object in multiple frames, rather than predicting the position of the moving object in subsequent images based on the detection frame position of the moving object in multiple frames. This is because motion detection may have false detections, that is, motion detection may be inaccurate. If the detection frame position of the moving object in multiple frames is directly input into the motion prediction model, the predicted position of the moving object output by the model may deviate significantly from the actual position of the moving object, resulting in inaccurate prediction results.

[0131] Using the method of this embodiment, the trusted frame position is derived by fusing the detection frame position with the predicted frame position. This can mitigate false detections and improve the accuracy of motion prediction, guiding the camera to focus accurately and enhancing the clarity of moving objects in the image. The motion prediction result is the predicted frame position information.

[0132] S608. The electronic device adjusts the focus position of the camera according to the predicted frame position information of the moving object in the P-th frame image among the M frames. S608 of this embodiment can refer to S304 of the above embodiment and will not be repeated here.

[0133] In the motion focus method illustrated in the above embodiment, the electronic device determines the credible frame position of the moving object in the i-th frame image captured by the camera based on the detection frame position of the moving object in the i-th frame image and the predicted frame position of the moving object in the i-th frame image. In this way, the electronic device can obtain the credible frame position of the moving object in multiple consecutive frames of images. Based on the credible frame position of the moving object in N consecutive frames of images, the electronic device can predict the predicted frame position of the moving object in images subsequent to N frames of images, and then adjust the camera's focus position based on the predicted frame position. This can further improve the accuracy of the predicted frame position, thereby improving the accuracy of the camera's focus, allowing the camera to capture clear images of moving objects.

[0134] The motion focus methods shown in the above embodiments can be applied to electronic devices with shooting functions, which can also be called terminals, user equipment (UE), mobile stations (MS), mobile terminals (MT), etc. The electronic devices can be mobile phones with touch screens, smart TVs, wearable devices, tablet computers, computers with wireless transceiver functions, virtual reality (VR) terminal devices, augmented reality (AR) terminal devices, wireless terminals in industrial control, wireless terminals in self-driving, wireless terminals in remote medical surgery, wireless terminals in smart grids, wireless terminals in transportation safety, wireless terminals in smart cities, wireless terminals in smart homes, etc. The embodiments of the present application do not limit the specific technologies and specific device forms used by the electronic devices.

[0135] For example, Figure 8 is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. As shown in Figure 8, the electronic device 100 includes: a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, a sensor 180, a button 190, a camera 193, and a display 194.

[0136] It should be understood that the structure illustrated in this embodiment does not constitute a specific limitation on the electronic device 100. In some embodiments, the electronic device 100 may include more or fewer components than shown, or may combine or separate certain components, or arrange the components differently. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0137] It is understood that the interface connection relationship between the modules shown in the embodiment is only for illustrative purposes and does not limit the structure of the electronic device 100. In some embodiments, the electronic device 100 may also adopt different interface connection methods from the above embodiments, or a combination of multiple interface connection methods.

[0138] The processor 110 may include one or more processing units. The different processing units may be independent devices or integrated into one or more processors. The processor 110 may also be provided with a memory for storing instructions and data. In an embodiment of the present application, the processor 110 may be configured to call a computer program in the memory so that the electronic device executes the steps of the aforementioned method embodiment, implements motion focus, and improves the image clarity of a moving object captured by the electronic device.

[0139] The USB interface 130 is an interface that complies with USB standards and may be a Mini USB interface, a Micro USB interface, a USB Type-C interface, etc. The USB interface 130 can be used to connect a charger to charge the terminal device, to transfer data between the terminal device and peripheral devices, or to connect headphones to play audio through the headphones.

[0140] The charging management module 140 is configured to receive charging input from a charger. The power management module 141 is configured to connect the battery 142 , the charging management module 140 and the processor 110 .

[0141] The wireless communication function of the electronic device 100 can be implemented through the antenna 1, the antenna 2, the mobile communication module 150, the wireless communication module 160, the modem processor, and the baseband processor. The mobile communication module 150 can provide solutions for wireless communications such as 2G / 3G / 4G / 5G applied to the electronic device 100. The wireless communication module 160 can provide solutions for wireless communications such as wireless local area networks (WLAN), Bluetooth, global navigation satellite systems (GNSS), frequency modulation (FM), NFC, and infrared technology (IR) applied to the electronic device 100.

[0142] Electronic device 100 implements display functionality through a GPU, display screen 194, and an application processor. A GPU is a microprocessor for image processing that connects display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. Processor 110 may include one or more GPUs that execute instructions to generate or modify display information.

[0143] The display screen 194 is used to display images, videos, etc. The display screen 194 includes a display panel. In some embodiments, the electronic device 100 may include one or N display screens 194, where N is a positive integer greater than one.

[0144] The electronic device 100 can implement a shooting function through an image signal processing (ISP) module, one or more cameras 193, a video codec, a GPU, one or more display screens 194, and an application processor.

[0145] The camera 193 is used to capture still images or videos. In some embodiments, the electronic device 100 may include one or more cameras 193. The camera 193 includes a lens, an image sensor (such as a complementary metal oxide semiconductor image sensor (CIS)), a motor, and the like.

[0146] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 via the external memory interface 120 to implement data storage functions. For example, data files such as music, photos, and videos can be stored on the external memory card.

[0147] The internal memory 121 may be used to store one or more computer programs, which include instructions. The processor 110 may execute the instructions stored in the internal memory 121 to enable the electronic device 100 to perform various functional applications and data processing.

[0148] The sensor 180 may include one or more of the following, for example: a pressure sensor, a gyroscope sensor, an air pressure sensor, a magnetic sensor, an acceleration sensor, a distance sensor, a proximity light sensor, a fingerprint sensor, a temperature sensor, a touch sensor, an ambient light sensor, or a bone conduction sensor.

[0149] Keys 190 include a power button, a volume button, and the like. Keys 190 can be mechanical or touch-sensitive. Electronic device 100 can receive key inputs and generate key signal inputs related to user settings and function control of electronic device 100. For example, when the camera application is enabled, the user can trigger the camera to take photos or record videos by pressing the power button.

[0150] In addition, on top of the above components, the electronic device also runs an operating system, such as the iOS operating system, the Android operating system, or the Windows operating system. Applications can be installed and run on the operating system.

[0151] The software system of the electronic device can adopt a layered architecture, an event-driven architecture, a microkernel architecture, a microservice architecture, or a cloud architecture. The embodiment of the present application takes the Android system as an example of a software system with a layered architecture to illustrate the software structure of the electronic device. Figure 9 is a schematic diagram of the software architecture and internal interactions of an electronic device provided in an embodiment of the present application. The layered architecture divides the software system of the electronic device into several layers, and each layer has a clear role and division of labor. The layers communicate with each other through software interfaces. Referring to Figure 9, the electronic device includes: an application layer, an application framework layer, a hardware abstraction layer, and a driver layer.

[0152] The application layer may include a camera application, which users can use to capture images or videos. In the embodiment of the present application, the camera application provides multiple shooting modes, such as motion focus mode. In some embodiments, the application package may also include applications such as gallery, calendar, call, map, navigation, Bluetooth, music, and video.

[0153] The application framework layer can provide an application programming interface (API) and a programming framework for the application programs in the application layer. In the embodiment of the present application, the application framework layer can include a camera management module and a window manager.

[0154] The camera management module is responsible for managing camera device information. Camera applications can use this module to obtain camera characteristics, such as the number of cameras and shooting capabilities. The camera management module is also used to transfer data between the camera application and the camera hardware abstraction layer. For example, referring to FIG9 , in response to an operation to enable motion focus mode, the camera application transmits a notification message ( S901 ) to the camera hardware abstraction layer via the camera management module. Upon receiving an image captured by the camera, the camera hardware abstraction layer accesses the motion detection module and the motion prediction module to guide the AF module to perform motion focus.

[0155] The window management module is responsible for managing application windows and their interaction with the user interface. For example, it manages the camera application window and sends its contents to the display driver for display. Window contents can include the image captured by the camera and the focus frame (i.e., the predicted frame) of the image.

[0156] The hardware abstraction layer is an interface layer located between the kernel layer and the hardware circuit. In an embodiment of the present application, the hardware abstraction layer includes a camera hardware abstraction layer, and the camera hardware abstraction layer may include a motion detection module, a motion prediction module, and an AF module. The motion detection module is pre-installed with a motion detection algorithm for detecting the position of a moving object in an image. The motion prediction module is pre-installed with a motion prediction model, which can be used to determine the predicted frame position of the moving object in the subsequent M frames of image based on the position of the moving object in the previous N frames of image, and can also be used to send the predicted frame position of the moving object in the subsequent M frames of image to the AF module. The AF module can be used to control the camera to focus based on the predicted frame position of the moving object sent by the motion prediction module.

[0157] The driver layer provides drivers for various hardware devices. In embodiments of the present application, the driver layer may include a camera driver and a display driver. The camera driver can be used to drive the camera of an electronic device. The display driver is used to drive the display screen of an electronic device.

[0158] The hardware layer includes hardware devices such as cameras and displays. A camera may include a lens, lens motor, image sensor, and image signal processing (ISP). The ISP can be used to process the raw image output by the image sensor, including linearity correction, noise removal, bad pixel removal, white balancing, and automatic exposure control.

[0159] For example, referring to Figure 9, after the camera application turns on the motion focus mode, the camera hardware abstraction layer sends an exposure control instruction to the image sensor through the camera driver (S902) to instruct the image sensor to capture continuous image frames. The image sensor sends the continuous image frames to the ISP (S903). After the ISP performs image processing on the continuous image frames, the continuous image frames are sent to the camera hardware abstraction layer through the ISP driver (S904). The camera hardware abstraction layer performs motion detection on the continuous image frames by calling the motion detection module to obtain the detection frame position of the moving object in the continuous image frames. The motion detection module sends the detection frame position of the moving object in the continuous image frames to the motion prediction module (S905). The motion prediction module can input the detection frame position of the moving object in the received continuous image frames into a preset motion prediction model to obtain the predicted frame position of the moving object in multiple frames of images after the continuous image frames. Optionally, in some embodiments, the input of the motion prediction model in the motion prediction module includes the credible frame position of the moving object in the continuous image frames.

[0160] The motion prediction module sends the predicted frame position of the moving object in the P-th frame image (e.g., P is 2) after the continuous image frames to the AF module (S906). The predicted frame position of the moving object in the P-th frame image is the focus position. After receiving the focus position, the AF module sends a focus parameter to the motor driver (S907), and the focus parameter includes the focus position. The motor driver sends a focus control instruction to the lens motor according to the received focus parameter (S908) to drive the lens motor to move and realize motion focus. For example, referring to Figure 7, the motion prediction model in the motion detection module outputs the position of the predicted frame 5' of the moving object in the 5th frame image and the position of the predicted frame 6 of the moving object in the 6th frame image. The motion detection module sends the position of the predicted frame 6 to the AF module. The AF module can send the position of the predicted frame 6 to the motor driver. The position of the predicted frame 6 is the focus position.

[0161] After receiving continuous image frames from the ISP, the camera hardware abstraction layer can send continuous image frames to the camera application through the camera management module (S909). It should be understood that the camera application continues to receive continuous image frames from the ISP, including image frames captured by the camera after motion focus. After the AF module receives the predicted frame position of the moving object in the P-th frame image after the continuous image frames sent by the motion detection module, the camera management module can send the predicted frame position of the moving object in the P-th frame image to the camera application (S910). In some embodiments, the AF module can send the predicted frame position of the moving object in the P-th frame image to the window manager (not shown) so that the window manager can merge the display content.

[0162] The camera application may send the P-th frame image and the predicted frame position of the moving object in the P-th frame image to the window manager (S911). After merging the P-th frame image and the predicted frame position of the moving object in the P-th frame image, the window manager sends display content including the P-th frame image with the predicted frame of the moving object to the display driver (S912), thereby driving the display to display the P-th frame image including the predicted frame of the moving object.

[0163] It is understood that the modules included in each layer shown in Figure 9 are modules involved in the embodiments of this application, and the modules included in each layer do not constitute a limitation on the structure of the electronic device and the hierarchy of module deployment. In some embodiments, the electronic device may include more or fewer layers than shown, and each layer may include more or fewer components, which is not limited in this application.

[0164] It should be noted that in the above embodiments, a "module" may be a software program, a hardware circuit, or a combination of the two that implements the above functions. The hardware circuit may include an application-specific integrated circuit (ASIC), an electronic circuit, a processor (e.g., a shared processor, a dedicated processor, or a group processor) and memory for executing one or more software or firmware programs, combined logic circuits, and / or other suitable components that support the described functions.

[0165] Therefore, the modules of each example described in the embodiments of this application can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0166] Based on the aforementioned embodiments, an embodiment of the present application proposes a motion focus method. In response to an operation of turning on a motion focus mode, the electronic device obtains N consecutive frames of images captured by a camera, each of the N frames of images includes the same moving object, and N is a positive integer greater than 1; the electronic device obtains the trusted position information of the moving object in each of the N frames of images; the electronic device determines the predicted position information of the moving object in the M frames of images after the N frames of images based on the trusted position information of the moving object in the N frames of images, where M is a positive integer greater than 1; the electronic device controls the camera to perform motion focus based on the predicted position information of the moving object in the P-th frame of the M frames of images, where P is a positive integer less than or equal to M.

[0167] Exemplarily, as shown in c of FIG1 , in response to the operation of the motion focus control 1032 acting on the window 1031 in the interface 103, the electronic device obtains N consecutive frames of images captured by the camera to perform a subsequent motion focus method. Taking a certain frame image among the N frames of images as an example, the credible position information of the moving object in the frame image is determined based on the detected position information and the predicted position information of the moving object in the frame image. Exemplarily, as shown in FIG7 , the position information of the credible frame 2 of the moving object in the second frame image is obtained by weighted summation based on the position information of the detected frame 2 of the moving object in the second frame image and the position information of the predicted frame 2 of the moving object in the second frame image.

[0168] In the above embodiment, when the motion focus mode is turned on in the camera application of the electronic device, the electronic device can construct a credible trajectory of the moving object (i.e., the trajectory corresponding to the credible position of the moving object in consecutive image frames), predict the motion trajectory of the moving object in subsequent image frames (i.e., the predicted position of the moving object in subsequent image frames) based on the credible trajectory of the moving object in the consecutive image frames, and control the camera to perform motion focus based on the predicted motion trajectory of the moving object in the subsequent image frames. Compared with the electronic device predicting the motion trajectory of the moving object in subsequent image frames based on the detection trajectory of the moving object in consecutive image frames (i.e., the trajectory corresponding to the detected position of the moving object in consecutive image frames), the above method can reduce the problem of inaccurate detection trajectory due to misdetection of the motion detection algorithm, which affects the accuracy of the predicted trajectory of the moving object, improve the accuracy of camera focus, and thus improve the image clarity of the moving object captured by the electronic device.

[0169] In an optional embodiment, obtaining the credible position information of the moving object in each of N frames of images includes: obtaining detected position information of the moving object in an i-th frame of the N frames of images, and predicted position information of the moving object in the i-th frame of the N frames of images, where the i-th frame of the N frames of images is any frame other than the first frame of the N frames of images; and determining the credible position information of the moving object in the i-th frame of the image based on the detected position information and the predicted position information of the moving object in the i-th frame of the image, where i is an integer greater than 1 and less than or equal to N.

[0170] The above embodiment shows a method for constructing a trusted position of a moving object in continuous image frames. The electronic device can determine the trusted position of the moving object in each frame image based on the detected position and predicted position of the moving object in the frame image. The credibility of the trusted position of the moving object in the frame image is higher than the detected position of the moving object in the frame image. The constructed trusted position of the moving object in the continuous image frames can be used by the electronic device to predict the motion trajectory of the moving object in subsequent image frames.

[0171] In an optional embodiment, the detected position information of the moving object in the i-th frame image includes the coordinate position of the first detection frame of the moving object in the i-th frame image, and the predicted position information of the moving object in the i-th frame image includes the coordinate position of the first prediction frame of the moving object in the i-th frame image; based on the detected position information and the predicted position information of the moving object in the i-th frame image, the credible position information of the moving object in the i-th frame image is determined, including: obtaining the coordinate position of the first credible frame of the moving object in the i-th frame image by performing a fusion operation on the coordinate position of the first detection frame and the coordinate position of the first prediction frame, and the credible position information of the moving object in the i-th frame image includes the coordinate position of the first credible frame.

[0172] In the above embodiment, the electronic device performs coordinate fusion based on the coordinate positions of the upper left corners of the detection and prediction frames of the i-th image frame to obtain the coordinate position of the upper left corner of the credible frame of the i-th image frame, and performs coordinate fusion based on the coordinate positions of the lower right corners of the detection and prediction frames of the i-th image frame to obtain the coordinate position of the lower right corner of the credible frame of the i-th image frame. The coordinate positions of the upper left corner and lower right corner of the credible frame of the i-th image frame uniquely determine the position of the credible frame in the i-th image frame, and this position can be used to predict the motion trajectory of subsequent moving objects.

[0173] In an optional embodiment, the coordinate position of the first detection frame includes the coordinate position of the upper left corner of the first detection frame and the coordinate position of the lower right corner of the first detection frame, and the coordinate position of the first prediction frame includes the coordinate position of the upper left corner of the first prediction frame and the coordinate position of the lower right corner of the first prediction frame; by performing a fusion operation on the coordinate position of the first detection frame and the coordinate position of the first prediction frame, the coordinate position of the first credible frame of the moving object in the i-th frame image is obtained, including: obtaining the coordinate position of the upper left corner of the first credible frame by performing a weighted sum operation on the coordinate position of the upper left corner of the first detection frame and the coordinate position of the upper left corner of the first prediction frame; and obtaining the coordinate position of the lower right corner of the first credible frame by performing a weighted sum operation on the coordinate position of the lower right corner of the first detection frame and the coordinate position of the lower right corner of the first prediction frame.

[0174] In the above embodiment, the electronic device performs a weighted sum operation on the coordinate positions of the upper left corner of the detection box and the prediction box in the same frame image, and performs a weighted sum operation on the coordinate positions of the lower right corner of the detection box and the prediction box in the same frame image, so as to obtain the coordinate positions of the upper left corner and the lower right corner of the credible box in the frame image, which are used for subsequent motion trajectory prediction of the moving object.

[0175] In an optional embodiment, obtaining the coordinate position of the upper left corner of the first credible frame by performing a weighted sum operation on the coordinate position of the upper left corner of the first detection frame and the coordinate position of the upper left corner of the first prediction frame includes:

[0176] The coordinate position (x, y) of the upper left corner of the first credible frame is determined by the following formula: x = w1 × x predict +w2×x detect y=w1×y predict +w2×y detect

[0177] In the formula, (x detect ,y detect ) represents the coordinate position of the upper left corner of the first detection frame, (x predict ,y predict) represents the coordinate position of the upper left corner of the first prediction box, w1 represents the weight value of the first prediction box, and w2 represents the weight value of the first detection box.

[0178] In the above embodiment, the electronic device performs a weighted summation of the x-axis coordinate positions of the upper left corners of the detection frame and the prediction frame in the same frame image, as well as a weighted summation of the y-axis coordinate positions of the upper left corners of the detection frame and the prediction frame in the same frame image, to obtain the coordinate position of the upper left corner of the credible frame in the frame image for subsequent motion trajectory prediction of the moving object. It should be understood that based on the principles of the above formula, the x-axis and y-axis coordinate positions of the lower right corner of the credible frame in the frame image can be obtained.

[0179] In an optional embodiment, the sum of the weight value of the first prediction frame and the weight value of the first detection frame is 1, and the weight value of the first prediction frame is determined based on the first distance value and the first moving speed of the moving object; the first distance value is the distance value between the detection frame of the moving object in the i-th frame image and the upper left corner of the prediction frame, and the first moving speed is the average moving speed of the upper left corner of the credible frame of the moving object from the iN-th frame image to the i-1-th frame image.

[0180] It should be understood that the sum of the weight value of the first prediction frame and the weight value of the first detection frame is 1. After determining the weight value of the first prediction frame, the electronic device can determine the weight value of the first detection frame. The weight value of the first prediction frame can be calculated with reference to Formula 3, and the weight value of the first detection frame can be calculated with reference to Formula 4.

[0181] It should be understood that, when the moving speed of the credible frame of the moving object remains unchanged (which may correspond to the first moving speed described above), as the position offset value between the detection frame and the prediction frame increases (which may correspond to the first distance value described above), the weight value of the prediction frame gradually increases, while the weight value of the detection frame gradually decreases. When the position offset value between the detection frame and the prediction frame remains unchanged, as the moving speed of the credible frame of the moving object increases, the weight value of the detection frame gradually increases, while the weight value of the prediction frame gradually decreases.

[0182] In an optional embodiment, based on the credible position information of the moving object in N frames of images, the predicted position information of the moving object in M ​​frames of images after the N frames of images is determined, including: inputting the credible position information of the moving object in the N frames of images into a motion prediction model, and obtaining the predicted position information of the moving object in M ​​frames of images after processing by the motion prediction model; the motion prediction model is obtained by training using a neural network model.

[0183] In the above embodiment, the electronic device inputs the credible frame coordinate position of the moving object in N frames of images into a preset motion prediction model, and after model calculation, obtains the predicted coordinate position of the moving object in M ​​frames of images after the N frames of images. The model input is not N frames of images, but the position data of the moving object in the image, which can improve the processing speed of the motion prediction model, thereby shortening the focusing time of the device and improving the motion focusing speed of the device.

[0184] In some embodiments, the motion prediction model can be trained using a lightweight, fully connected neural network model. This lightweight model structure not only saves device storage space but also improves model execution speed. Furthermore, due to its simple structure and low computational complexity, it reduces device power consumption.

[0185] In an optional embodiment, after the electronic device controls the camera to perform motion focus based on the predicted position information of the moving object in the P-th frame image among the M-frame images, the method further includes: the electronic device obtains the P-th frame image among the M-frame images captured by the camera; the electronic device displays the P-th frame image and the predicted position information of the moving object in the P-th frame image, and the image clarity of the moving object in the P-th frame image is greater than the preset clarity.

[0186] In the above embodiment, when the electronic device displays the P-th frame image, it also displays the predicted frame position of the moving object in the P-th frame image. This predicted frame position is the focus position when the P-th frame image is captured. Because this predicted frame position is determined based on the aforementioned method, the deviation between this predicted frame position and the actual position of the moving object in the P-th frame image is less than a preset value, resulting in the image clarity of the moving object in the P-th frame image being greater than a preset clarity. It should be understood that subsequent image frames are all focused according to the aforementioned method, which can generally improve the image clarity of the moving object captured by the electronic device.

[0187] An embodiment of the present application also provides an electronic device, which includes: one or more processors and a memory, the memory is coupled to the one or more processors, the memory is used to store computer program code, the computer program code includes computer instructions, and the one or more processors call the computer instructions to enable the electronic device to execute the steps in the aforementioned method embodiment. Its implementation principle and technical effects are similar to those of the aforementioned related embodiments and will not be repeated here.

[0188] An embodiment of the present application also provides a chip system, which is applied to an electronic device. The chip system includes one or more processors, and the one or more processors are used to call computer instructions to enable the electronic device to execute the steps in the aforementioned method embodiment. Its implementation principle and technical effects are similar to those of the aforementioned related embodiments and will not be repeated here.

[0189] An embodiment of the present application also provides a computer-readable storage medium, which includes computer instructions. When the computer instructions are executed on an electronic device, the electronic device executes the steps in the aforementioned method embodiment. The implementation principle and technical effects are similar to those of the aforementioned related embodiments and will not be repeated here.

[0190] An embodiment of the present application also provides a computer program product, which includes computer program code. When the computer program code runs on an electronic device, the electronic device executes the steps in the aforementioned method embodiment. Its implementation principle and technical effects are similar to those of the aforementioned related embodiments and will not be repeated here.

[0191] The methods described in the above embodiments can be implemented in whole or in part through software, hardware, firmware, or any combination thereof. If implemented in software, the functions can be stored as one or more instructions or codes on a computer-readable medium or transmitted on a computer-readable medium. Computer-readable media can include computer storage media and communication media, and can also include any medium that can transfer a computer program from one place to another. The storage medium can be any target medium that can be accessed by a computer.

[0192] In some embodiments, computer-readable media may include RAM, ROM, compact disc read-only memory (CD-ROM) or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium designed to carry or store the desired program code in the form of instructions or data structures and accessible by a computer. Moreover, any connection is appropriately referred to as a computer-readable medium. For example, if a coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL) or wireless technology (such as infrared, radio and microwave) is used to transmit software from a website, server or other remote source, the coaxial cable, fiber optic cable, twisted pair, DSL or wireless technology such as infrared, radio and microwave are included in the definition of medium. Disk and optical disk as used herein include optical disk, laser disk, optical disk, digital versatile disk (DVD), floppy disk and Blu-ray disk, where disks generally reproduce data magnetically, while optical disks reproduce data optically using lasers. Combinations of the above should also be included within the scope of computer-readable media.

[0193] The present application embodiment is described with reference to the flow chart and / or block diagram of the method, device (system) and computer program product according to the embodiment of the present application.It should be understood that each flow process and / or box in the flow chart and / or block diagram and the combination of the flow process and / or box in the flow chart and / or block diagram can be realized by computer program instructions.These computer program instructions can be provided to the processing unit of general-purpose computer, special-purpose computer, embedded processing machine or other programmable device to produce a machine, so that the instruction executed by the processing unit of computer or other programmable data processing device produces the device for realizing the function specified in one flow chart flow or multiple flows and / or one block or multiple blocks of block diagram.

[0194] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0195] The above specific implementation methods further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above are only specific implementation methods of the present invention and are not intended to limit the scope of protection of the present invention. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solutions of the present invention should be included in the scope of protection of the present invention.

Claims

1. A motion focusing method, characterized in that, including: In response to an operation to turn on the motion focus mode, the electronic device acquires N consecutive images collected by the camera, each of the N images including the same moving object, and N is a positive integer greater than 1; The electronic device acquires reliable position information of the moving object in each of the N images; The electronic device determines predicted position information of the moving object in M images after the N images according to the reliable position information of the moving object in the N images, and M is a positive integer greater than 1; The electronic device controls the camera to perform motion focus according to the predicted position information of the moving object in the P-th image among the M images, and P is a positive integer less than or equal to M.

2. The method according to claim 1, wherein Acquiring reliable position information of the moving object in each of the N images includes: Acquiring detection position information of the moving object in the i-th image among the N images, and predicted position information of the moving object in the i-th image, where the i-th image is any image frame other than the first image frame among the N images; According to the detection position information and the predicted position information of the moving object in the i-th image, determining the reliable position information of the moving object in the i-th image.

3. The method according to claim 2, characterized in that The detection position information of the moving object in the i-th image includes the coordinate position of the first detection box of the moving object in the i-th image, and the predicted position information of the moving object in the i-th image includes the coordinate position of the first prediction box of the moving object in the i-th image; According to the detection position information and the predicted position information of the moving object in the i-th image, determining the reliable position information of the moving object in the i-th image includes: By performing a fusion operation on the coordinate positions of the first detection box and the first prediction box, obtaining the coordinate position of the first reliable box of the moving object in the i-th image, and the reliable position information of the moving object in the i-th image includes the coordinate position of the first reliable box.

4. The method according to claim 3, characterized in that, The coordinate position of the first detection box includes the coordinate position of the upper left corner and the coordinate position of the lower right corner in the first detection box, and the coordinate position of the first prediction box includes the coordinate position of the upper left corner and the coordinate position of the lower right corner in the first prediction box; By performing a fusion operation on the coordinate positions of the first detection box and the first prediction box, obtaining the coordinate position of the first reliable box of the moving object in the i-th image includes: By performing a weighted summation operation on the coordinate position of the upper left corner in the first detection box and the coordinate position of the upper left corner in the first prediction box, obtaining the coordinate position of the upper left corner in the first reliable box; and By performing a weighted summation operation on the coordinate position of the lower right corner in the first detection box and the coordinate position of the lower right corner in the first prediction box, obtaining the coordinate position of the lower right corner in the first reliable box.

5. The method according to claim 4, wherein Performing a weighted summation operation on the coordinate position of the upper left corner in the first detection box and the coordinate position of the upper left corner in the first prediction box to obtain the coordinate position of the upper left corner in the first confidence box, including: Determining the coordinate position (x, y) of the upper left corner in the first confidence box through the following formula: x = w1 × x predict + w2 × x detect y = w1 × y predict + w2 × y detect Wherein, (x detect , y detect ) represents the coordinate position of the upper left corner in the first detection box, (x predict , y predict ) represents the coordinate position of the upper left corner in the first prediction box, w1 represents the weight value of the first prediction box, and w2 represents the weight value of the first detection box.

6. The method according to claim 5, wherein: The sum of the weight value of the first prediction box and the weight value of the first detection box is 1, and the weight value of the first prediction box is determined according to the first distance value and the first moving speed of the moving object; The first distance value is the distance value between the detection box of the moving object and the upper left corner in the prediction box in the i-th frame image, and the first moving speed is the average moving speed of the upper left corner of the confidence box of the moving object from the (i - N)-th frame image to the (i - 1)-th frame image.

7. The method according to any one of claims 1 to 6, wherein: Determining the predicted position information of the moving object in the M frames of images after the N frames of images according to the reliable position information of the moving object in the N frames of images, including: Inputting the reliable position information of the moving object in the N frames of images into a motion prediction model, and after being processed by the motion prediction model, obtaining the predicted position information of the moving object in the M frames of images after the N frames of images; The motion prediction model is trained by using a neural network model.

8. The method according to any one of claims 1 to 7, wherein: After the electronic device controls the camera to perform motion focusing according to the predicted position information of the moving object in the P-th frame image among the M frames of images, the method further includes: The electronic device acquires the P-th frame image among the M frames of images collected by the camera; The electronic device displays the P-th frame image and the predicted position information of the moving object in the P-th frame image, and the image clarity of the moving object in the P-th frame image is greater than a preset clarity.

9. An electronic device, characterized in that, The electronic device includes: one or more processors and a memory; The memory is coupled to the one or more processors, and the memory is used to store computer program code, and the computer program code includes computer instructions. The one or more processors call the computer instructions to cause the electronic device to execute the method according to any one of claims 1 to 8.

10. A chip system, characterized in that, The chip system is applied to an electronic device, and the chip system includes one or more processors, and the one or more processors are used to call computer instructions to cause the electronic device to execute the method according to any one of claims 1 to 8.

11. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes computer instructions, and when the computer instructions run on an electronic device, the electronic device is caused to execute the method according to any one of claims 1 to 8.

12. A computer program product, characterized in that, The computer program product includes computer program code, and when the computer program code runs on an electronic device, the electronic device is caused to execute the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Target tracking method and related device

    CN103259962A

  • A target tracking method and device

    CN109712188A

  • Motion track identification method and device, equipment and medium

    CN115018886A

  • Tracking focusing method, electronic equipment and computer readable storage medium

    CN116055844A

  • Method and device for detecting the movement state of objects

    EP1962245A2