Position determination method and device, image prediction model training method and device

By acquiring and predicting image frames at each moment, combined with the image prediction model training method, the problem of inaccurate determination of image surfaces of multiple vehicles is solved, and the accuracy of vehicle route planning is achieved.

CN115937303BActive Publication Date: 2025-08-26BEIJING HORIZON INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211739030.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-30
Publication Date
2025-08-26
Estimated Expiration
2042-12-30

AI Technical Summary

Technical Problem

When multiple vehicles are included in the historical image frame, the prior art cannot accurately determine the image plane position of the vehicle in the predicted image frame, resulting in the inability to route planning for each vehicle.

Method used

By acquiring the first image frame of each moment and a second image frame of each moment, a second predicted image frame of each detection object is predicted, and using the image prediction model training method, the image prediction model is iteratively updated to accurately determine the image plane position of each detection object.

Benefits of technology

The accurate positioning of each detection object in the predicted image frame is realized, and effective route planning for each vehicle can be carried out to avoid congestion and accidental road sections.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115937303B_ABST
    Figure CN115937303B_ABST
Patent Text Reader

Abstract

Disclosed are a method and device for determining a position, and a method and device for training an image prediction model, relating to the field of intelligent driving technology. The method includes: determining the position information of N detection objects at each of multiple moments; determining a first image frame for each of the N detection objects at each moment, and a second image frame for each of the N detection objects at each moment, based on the position information of the N detection objects at each moment; predicting a first predicted image frame for the N detection objects, and a second predicted image frame for each of the N detection objects, based on the first image frames at the multiple moments and the N second image frames at each of the multiple moments; and determining the image plane position of each detection object in the first predicted image frame based on the first predicted image frame and the N second predicted image frames. The present disclosure accurately determines the image plane position of each detection object from the first predicted image frame comprising the N detection objects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of intelligent driving technology, and specifically to a method and device for determining a position, and a method and device for training an image prediction model. Background Art

[0002] Currently, predicting a vehicle's trajectory can bring a lot of convenience to users. For example, based on the predicted trajectory, the vehicle can be route-planned to avoid congested sections, accident-prone sections, and so on.

[0003] However, when the historical image frame includes multiple vehicles, the predicted image frame predicted based on the historical image frame also includes multiple vehicles. Therefore, it may be impossible to determine the image plane positions of the multiple vehicles in the predicted image frame, and thus route planning cannot be performed for each vehicle. Summary of the Invention

[0004] Existing prediction image frames obtained based on historical image frames including multiple vehicles also include multiple vehicles. It is impossible to determine the image plane positions of the multiple vehicles in the prediction image frames, and thus it is impossible to perform route planning for each vehicle.

[0005] In order to solve the above technical problems, the present application is proposed. The embodiments of the present application provide a method and device for determining a position, and a method and device for training an image prediction model. In the position determination scheme provided by the present application, the first image frame includes N detection objects (such as vehicles), so the first predicted image frame including the N detection objects can be predicted based on the first image frames at multiple moments. Each second image frame only includes one detection object, so for each detection object, the second predicted image frame of the detection object can be predicted based on the second image frames at multiple moments that only include the detection object. Then, the second predicted image frame of each detection object in the N detection objects can be predicted. The image plane position of each detection object can be accurately known based on the second predicted image frame of each detection object, and then the predicted image plane position of each of the N detection objects can be accurately known based on the second predicted image frame of each of the N detection objects. Then, the predicted image plane position of each of the N detection objects can be accurately determined from the first predicted image frame including the N detection objects by taking the predicted image plane positions of each of the N detection objects. Thus, route planning can be performed for each detection object.

[0006] According to one aspect of the present application, a position determination method is provided, comprising: first determining position information of N detection objects at each of a plurality of moments, where the plurality of moments include a current moment and at least two historical moments, and N is a positive integer greater than 1; then, based on the position information of the N detection objects at each moment, determining a first image frame of the N detection objects at each moment, and a second image frame of each of the N detection objects at each moment; thereafter, predicting a first predicted image frame of the N detection objects and a second predicted image frame of each of the N detection objects based on the first image frames at the plurality of moments and the N second image frames at each of the plurality of moments; finally, determining an image plane position of each detection object in the first predicted image frame based on the first predicted image frame and the N second predicted image frames.

[0007] Based on the above scheme, not only the first image frame at each moment (such as the current moment, historical moment) is obtained, but also N second image frames at each moment are obtained. The first image frame includes N detection objects (such as vehicles). The N second image frames each include N detection objects, that is, each second image frame includes one detection object among the N detection objects. Since the first image frame includes N detection objects, a first predicted image frame including N detection objects can be predicted based on the first image frames at multiple moments. Since each second image frame only includes one detection object, for each detection object, a second predicted image frame of the detection object can be predicted based on the second image frames at multiple moments that only include the detection object. Furthermore, a second predicted image frame of each of the N detection objects can be predicted. Based on the second predicted image frame of each detection object, the image plane position of each detection object can be accurately determined, and further, based on the second predicted image frames of each of the N detection objects, the predicted image plane position of each of the N detection objects can be accurately determined. By taking the predicted image plane positions of each of the N detection objects, the image plane position of each detection object can be accurately determined from the first predicted image frame including the N detection objects.

[0008] Furthermore, based on the image plane position of each detected object in the first predicted image frame, as well as the image plane positional relationship between each detected object and other detected objects in the first predicted image frame, a route can be planned for the detected object. For example, it can be determined whether the road ahead of the detected object is congested or has an accident section, so that the detected object can be reminded to avoid congested and accident sections.

[0009] According to one aspect of the present application, a method for training an image prediction model is provided, comprising: first, obtaining multiple groups of input samples and output samples corresponding to each group of input samples. Each group of input samples includes: a first input image frame of each of the M detection objects at a plurality of first moments, and a second input image frame of each of the M detection objects at the first moment; the output samples include: a first output image frame of the M detection objects at the second moment, and a second output image frame of each of the M detection objects at the second moment; the second moment is after the first moment, and M is a positive integer greater than 1. Then, each group of input samples is input into the initial image prediction model to obtain a training output image, which includes the first predicted output image frame of the M detection objects at the second moment, and the second predicted output image frame of each of the M detection objects at the second moment. Finally, the initial image prediction model is iteratively trained using the M second predicted output image frames and the M second output image frames as supervisory information to obtain a trained image prediction model.

[0010] Based on the above technical solution, during the training of the initial image prediction model, since each second predicted output image frame includes only one detection object, the image plane position of the detection object in the second predicted output image can be clearly represented. Therefore, the image plane position of the detection object in the second predicted output image can be compared with the image plane position of the detection object in a second output image frame (i.e., M second predicted output image frames and M second output image frames are used as supervision information), which is used to iteratively update the initial image detection module, thereby obtaining a trained image detection model with more accurate image plane position prediction.

[0011] According to one aspect of the present application, a position determination device is provided, including: a position acquisition module for determining position information of N detection objects at each moment in a plurality of moments, where the plurality of moments include a current moment and at least two historical moments, and N is a positive integer greater than 1; an image acquisition module for determining a first image frame of the N detection objects at each moment, and a second image frame of each of the N detection objects at each moment based on the position information of the N detection objects at each moment; a prediction module for predicting a first predicted image frame of the N detection objects and a second predicted image frame of each of the N detection objects based on the first image frames at the plurality of moments and the N second image frames at each moment in the plurality of moments; and a position determination module for determining an image plane position of each detection object in the first predicted image frame based on the first predicted image frame and the N second predicted image frames.

[0012] According to one aspect of the present application, a training device for an image prediction model is provided, comprising: a sample acquisition module for acquiring multiple groups of input samples and output samples corresponding to each group of input samples; wherein each group of input samples comprises: a first input image frame of each of the M detection objects at a plurality of first moments, and a second input image frame of each of the M detection objects at the first moment; the output samples comprise: a first output image frame of the M detection objects at the second moment, and a second output image frame of each of the M detection objects at the second moment; the second moment is after the first moment, and M is a positive integer greater than 1; an image prediction module for inputting each group of input samples into an initial image prediction model to obtain a training output image, the training output image comprising the first predicted output image frame of the M detection objects at the second moment, and the second predicted output image frame of each of the M detection objects at the second moment; an iterative training module for iteratively training the initial image prediction model using the M second predicted output image frames and the M second output image frames as supervisory information to obtain a trained image prediction model.

[0013] According to one aspect of the present application, a computer-readable storage medium is provided, which stores a computer program, and the computer program is used to execute the method provided in any of the above aspects.

[0014] According to one aspect of the present application, an electronic device is provided, which includes: a processor; a memory for storing processor-executable instructions; and a processor for reading executable instructions from the memory and executing the instructions to implement the method provided in any of the above aspects. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] The above and other purposes, features, and advantages of the present application will become more apparent through a more detailed description of the embodiments of the present application in conjunction with the accompanying drawings. The accompanying drawings are intended to provide a further understanding of the embodiments of the present application and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present application and do not constitute a limitation of the present application. In the drawings, the same reference numerals generally represent the same components or steps.

[0016] Figure 1 It is a schematic diagram of the implementation environment of the technical solution provided by this application.

[0017] Figure 2 This is a flow diagram of a location determination method provided by this application. Figure 1 .

[0018] Figure 3 This is a schematic diagram of a first image frame and a second image frame provided by the present application.

[0019] Figure 4This is a flow diagram of a location determination method provided by this application. Figure 2 .

[0020] Figure 5 This is a flow chart of an image prediction model provided in this application for predicting image plane position.

[0021] Figure 6 This is a structural diagram of a training method for an image prediction model provided by this application. Figure 1 .

[0022] Figure 7 This is a structural diagram of a training method for an image prediction model provided by this application. Figure 2 .

[0023] Figure 8 This is a schematic diagram of the structure of a position determination device provided by this application. Figure 1 .

[0024] Figure 9 This is a schematic diagram of the structure of a position determination device provided by this application. Figure 2 .

[0025] Figure 10 This is a schematic diagram of the structure of a position determination device provided by this application. Figure 3 .

[0026] Figure 11 This is a schematic diagram of the structure of a training device for an image prediction model provided by this application. Figure 1 .

[0027] Figure 12 This is a schematic diagram of the structure of a training device for an image prediction model provided by this application. Figure 2 .

[0028] Figure 13 This is a schematic diagram of the structure of a training device for an image prediction model provided by this application. Figure 3 .

[0029] Figure 14 This is a structural diagram of an electronic device provided by this application. DETAILED DESCRIPTION

[0030] Below, the exemplary embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application, and it should be understood that the present application is not limited to the exemplary embodiments described herein.

[0031] Hereinafter, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of the technical features being referred to. Thus, a feature designated "first" or "second" may explicitly or implicitly include one or more of such features. In the description of this application, unless otherwise specified, "plurality" means two or more. "A and / or B" includes the following three combinations: A only, B only, and a combination of A and B.

[0032] Currently, to predict vehicle trajectories, related solutions capture historical image frames from multiple vehicles traveling on the road and then predict a predicted image frame based on these historical image frames. Because the historical image frames include multiple vehicles, the predicted image frame also includes multiple vehicles. The predicted image frame includes the image positions of multiple vehicles, making it impossible to distinguish the image position of each vehicle in the predicted image frame, making it impossible to plan a route for each vehicle.

[0033] In response to the above problems, an embodiment of the present application provides a position determination method, which obtains a first image frame including N detection objects (such as vehicles) at each moment (such as the current moment, a historical moment), and a second image frame including one detection object at each moment. Since the first image frame includes N detection objects, a first predicted image frame including N detection objects can be predicted based on the first image frames at multiple moments. Since each second image frame only includes one detection object, for each detection object, a second predicted image frame of the detection object can be predicted based on the second image frames at multiple moments that only include the detection object. The image plane position of each detection object can be accurately known based on the second predicted image frame of each detection object, and then the predicted image plane position of each of the N detection objects can be accurately known based on the second predicted image frame of each of the N detection objects. By taking the predicted image plane positions of each of the N detection objects, the image plane position of each detection object can be accurately determined from the first predicted image frame including the N detection objects.

[0034] Furthermore, based on the image plane position of each detected object in the first predicted image frame, as well as the image plane positional relationship between the detected object and other detected objects in the first predicted image frame, a route can be planned for the detected object. For example, it can be determined whether the road ahead of the detected object is congested or has an accident section, so that the detected object can be reminded to avoid congested and accident sections.

[0035] Figure 1 This is a structural diagram of an electronic device according to an exemplary embodiment. The location determination method provided in the embodiment of the present application can be applied to Figure 1 The electronic device 100 shown in FIG. Figure 1As shown, the electronic device 100 includes one or more processors 101 and a memory 102 .

[0036] The processor 101 may be a central processing unit (CPU) or other forms of processing units having data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 100 to perform desired functions.

[0037] The memory 102 may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory. The non-volatile memory may include, for example, read-only memory (ROM), a hard disk, a flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 101 may execute the program instructions to implement the position determination method or image prediction model training method of the various embodiments of the present application described above and / or other desired functions.

[0038] In one example, the electronic device 100 may further include an input device 103 and an output device 104 , and these components are interconnected via a bus system and / or other forms of connection mechanisms (not shown).

[0039] Of course, to simplify, Figure 1 Only some of the components related to the present application in the electronic device 100 are shown, and components such as buses, input / output interfaces, etc. are omitted. In addition, the electronic device 100 may further include any other appropriate components according to specific application scenarios.

[0040] In some embodiments, Figure 1 The electronic device 100 shown can be a terminal or a server. The terminal can be used to obtain the position information of N vehicles at multiple times. The server can be used to train the initial image prediction model based on the position information of the N vehicles at multiple times to obtain an image prediction model. The server can also send the image prediction model to the terminal so that the terminal predicts the image plane positions of the N vehicles in the first predicted image frame based on the image prediction model, and performs route planning for each of the N vehicles based on the image plane positions of the N vehicles in the first predicted image frame to determine a driving route with better road conditions. Of course, the terminal can also train the initial image prediction model based on the position information of the N vehicles at multiple times, and the embodiment of the present application does not limit the specific device for training the image prediction model.

[0041] The location determination method provided in the embodiments of the present application is described below with reference to the accompanying drawings.

[0042] The embodiment of the present application provides a location determination method, which can be executed by a location determination device, which can be located at Figure 1 The electronic device 100 shown in FIG. 1 is a part of the electronic device 100 (eg, a CPU, a GPU, or a hardware acceleration chip). Figure 2 As shown, the method may include S201-S204.

[0043] S201. Determine location information of N detection objects at each of a plurality of moments, where the plurality of moments include a current moment and at least two historical moments, and N is a positive integer greater than 1.

[0044] The position determination device may determine the position information of each of the N detection objects at each of the multiple moments. In other words, the position information of the N detection objects at each moment may include: the position information of each of the N detection objects at each moment.

[0045] In some embodiments, the interval between two adjacent moments in the multiple moments may be the same or different. The following embodiments are illustrative examples of the same interval between two adjacent moments in the multiple moments. For example, the interval between two adjacent historical moments in the multiple moments may be the same, and the interval is relatively short, such as 1 second.

[0046] In some embodiments, the N detected objects may include all detected objects that are within the target area at any one of the multiple moments in time. Because the intervals between the multiple moments in time are relatively short, any one of the N detected objects may be within the target area at multiple moments in time. In other words, the detected objects within the target area at multiple moments in time may be the same.

[0047] In some embodiments, the N detected objects may include all detected objects located within the target area at multiple moments in time.

[0048] For example, if the N detection objects are three vehicles and the multiple time points include t1, t2, t3, and t4, each vehicle is located within a target area at at least one of t1, t2, t3, and t4. The position determination device can determine the position information of the three vehicles at t1, the position information of the three vehicles at t2, the position information of the three vehicles at t3, and the position information of the three vehicles at t4.

[0049] The target area may be an area defined by at least one designated geographic location. For example, a circular area with a designated geographic location as its center and a designated radius may be the target area. Another example may be a square area with four designated geographic locations as its vertices. Furthermore, the vehicle's geographic location may be used to determine whether it is within the target area.

[0050] In an embodiment of the present application, taking the detection object as a vehicle as an example, the position information of each detection object can characterize the position of each detection object on the road. The position information of each detection object can be the geographical location (for example, longitude and latitude coordinate values) of each detection object, or the image plane position (for example, two-dimensional coordinate values ​​in the image) of each detection object in the image. Wherein, the types of the position information of the detection object obtained using different acquisition methods (such as, positioning system, image frame) are different. Different acquisition methods and the type of position information obtained by each acquisition method are introduced below.

[0051] Illustratively, the position determination device may obtain the geographical location of each detection object at each of multiple moments through a positioning system.

[0052] For example, the position determination device may obtain the position information of each detection object at each moment through a global navigation satellite system (GNSS). GNSS may include a global positioning system (GPS), a global navigation satellite system (GLONASS), a Beidou navigation satellite system (BDS), a quasi-zenith satellite system (QZSS) and / or a satellite-based augmentation system (SBAS), etc. For example, the geographical location of each detection object at each moment may be the latitude and longitude coordinate values ​​obtained using GPS.

[0053] Exemplarily, the position determination device may also use a light detection and ranging (Lidar) laser radar to obtain the geographic location of each detected object at each moment.

[0054] Exemplarily, the position determining device may also obtain the image plane position of each detection object at each of multiple moments through image frames.

[0055] For example, the position determination device can obtain an initial image frame of each detection object at each of multiple moments (e.g., a historical moment, a current moment); and then obtain the image plane position of each detection object at each moment based on the initial image frame of each detection object at each moment, where the image plane position is the image plane position of each detection object in the initial image frame. The initial image frame of each detection object at each moment can be obtained by capturing it through an image acquisition module (e.g., a camera), or can be predicted using the position determination method provided in an embodiment of the present application.

[0056] Exemplarily, the initial image frame of each detection object at each moment may be a top view of the detection object at each moment.

[0057] S202 : Determine, based on position information of the N detection objects at each moment, a first image frame of the N detection objects at each moment, and a second image frame of each of the N detection objects at each moment.

[0058] The position determination device can determine, based on the position information of the N detection objects at each moment, a first image frame that includes the N detection objects at each moment. The position determination device can also determine, based on the position information of each detection object at each moment, a second image frame that only includes the detection object at each moment, thereby determining N second image frames. The N second image frames each include a different detection object.

[0059] Exemplarily, all image frames at each moment determined by the position determination apparatus can be expressed as h*w*c. h is the length of each image frame (e.g., the first image frame, the second image frame). w is the width of each image frame (e.g., the first image frame, the second image frame). c is the total number of all image frames at each moment, where c=N+1. That is, all image frames at each moment include one first image frame and N second image frames, wherein the first image frame includes N detection objects, and each of the N second image frames includes one detection object.

[0060] Exemplarily, if the position information of each detection object at each moment is the latitude and longitude coordinate value of each detection object in space, or the two-dimensional coordinate value of each detection object in the overhead view, then the first image frame determined by the position determination device based on the position information of N detection objects at each moment is an overhead view including N detection objects, and the second image frame determined based on the position information of each detection object at each moment also includes an overhead view of the detection object.

[0061] In some embodiments, the position determination device may render a first image frame of the N detected objects at each moment based on the position information of the N detected objects at each moment. Similarly, the position determination device may also render a second image frame of each detected object at each moment based on the position information of each detected object at each moment.

[0062] For example, if the position information of the N detected objects at each moment is image plane positions, the position determination device may render the first image frame based on the image plane positions of the N detected objects at each moment. If the position information of the N detected objects at each moment is geographic locations, the position determination device may render the first image frame based on the geographic locations of the N detected objects at each moment and the image scale. Image scale refers to the ratio of the length of a line segment on the image to the horizontal projection length of the corresponding line segment in the field.

[0063] It should be noted that, for the process of the position determination apparatus drawing and rendering the second image frame, reference may be made to the detailed introduction of the position determination apparatus drawing and rendering the first image frame, which will not be described in detail in the embodiment of the present application.

[0064] In some embodiments, because the image plane position of the detected object at the next moment after multiple moments (or the next moment after the current moment) is affected by the surrounding environment, the position determination device may obtain not only the position information of the N detected objects at each of the multiple moments, but also the environmental information of the N detected objects. For example, information about the road on which the N vehicles are located (which may be referred to as road information) may include road coordinate information. The next moment may also be referred to as a future moment adjacent to the current moment.

[0065] Furthermore, the position determination device can generate a first image frame of the N detection objects at each moment based on the position information of the N detection objects at each moment and the environmental information of the N detection objects. The first image frame includes not only the N detection objects, but also other objects in the environment where the N detection objects are located (for example, the road where the vehicle is located). The position determination device can also generate a second image frame of each detection object at each moment based on the position information of each detection object at each moment and the environmental information where each detection object is located. The second image frame includes not only one detection object, but also other objects in the environment where the detection object is located (for example, the road where the vehicle is located). The environmental information where the N detection objects are located and the environmental information where each detection object is located can be the same.

[0066] For example, continuing to take the N detection objects as 3 vehicles as an example, the position determination device can draw and render a first image frame including the 3 vehicles and 3 second image frames including the 3 vehicles respectively based on the position information of the 3 vehicles at t1 and the road information around the 3 vehicles. Figure 3 The first image frame shown in (a) is a top view of the road where three vehicles are located. Figure 3 (b) Figure 3 (c) and Figure 3 The second image frame shown in (d) is a top view of a road including one vehicle and three vehicles, and Figure 3 (b) Figure 3 (c) and Figure 3 The second image frame shown in (d) includes a different vehicle. Figure 3 The white areas in the figure are vehicles. Figure 3 The gray area in the figure is the road where the three vehicles are located.

[0067] S203. Predict first predicted image frames of N detection objects and second predicted image frames of each of the N detection objects based on the first image frames at multiple moments and N second image frames at each of the multiple moments.

[0068] The position determination device can predict the first predicted image frame of the N detection objects at the next moment based on the first image frames at multiple moments and the N second image frames at each of the multiple moments, and predict the second predicted image frame of each detection object at the next moment, thereby predicting N second predicted image frames. The first predicted image frame includes the N detection objects and is used to represent the image plane positions of the N detection objects at the next moment. The second predicted image frame includes only one detection object and is used to represent the image plane position of the one detection object at the next moment. The N second predicted image frames each include different detection objects.

[0069] The duration between each two adjacent moments in the multiple moments (ie, the interval duration) may be equal, and the duration between the next moment and the current moment in the multiple moments may be equal to the interval duration.

[0070] In an embodiment of the present application, other objects in the environment where the N detection objects are located may be objects that affect the image plane position of the detection objects at the next moment. That is to say, in addition to the N detection objects, the first image frame may only include other objects that affect the image plane position of the detection object at the next moment, and in addition to one detection object, the second image frame may only include other objects that affect the image plane position of the detection object at the next moment. Since the first image frame and the second image frame do not include information that is irrelevant to the image plane position of the detection object at the next moment, the influence of information that is irrelevant to the image plane position of the detection object at the next moment on the prediction of the image plane position of the detection object at the next moment can be avoided, thereby improving the accuracy of the prediction of the image plane position of the detection object at the next moment.

[0071] In an embodiment of the present application, to predict the image plane position of the detection object at the next moment, the motion state of the detection object at the current moment can be obtained based on the image plane position of the detection object at at least two historical moments before the current moment (i.e., the current moment before the next moment); the image plane position of the detection object at the current moment is also obtained. Then, based on the image plane position of the detection object at the current moment and the motion state of the detection object at the current moment, the image plane position of the detection object at the next moment after the current moment can be predicted.

[0072] Specifically, the position determination device can obtain the optical flow field Flow of the N detection objects at the current moment (which can be simply referred to as the optical flow field at the current moment). The optical flow field can represent the motion state of the N detection objects at the current moment. Then, the position determination device can use the optical flow field of the N detection objects at the current moment to process the first image frame of the N detection objects at the current moment to obtain the first predicted image frame of the N detection objects. The first predicted image frame can represent the image plane position of the N detection objects at the next moment. Similarly, the position determination device can also use the optical flow field of the N detection objects at the current moment to process the second image frame of each detection object at the current moment to obtain the second predicted image frame of each detection object. The second predicted image frame of each detection object can represent the image plane position of the detection object at the next moment.

[0073] The processing operation performed by the position determination device on each image frame (eg, the first image frame, the second image frame) at the current moment using the optical flow field may be referred to as a remapping operation (ie, a warp operation).

[0074] In an embodiment of the present application, the position determination device can determine the motion state of N detection objects at the current moment (i.e., the optical flow field of the N detection objects at the current moment) based on the image plane positions of the N detection objects at at least two historical moments before the current moment. The first image frames of at least two historical moments can represent the image plane positions of the N detection objects at at least two historical moments, and the N second image frames of at least two historical moments can also represent the image plane positions of the N detection objects at at least two historical moments. Therefore, the position determination device can obtain the optical flow field at the current moment based on the first image frames of at least two historical moments, or the N second image frames of at least two historical moments.

[0075] In some embodiments, the position determination device may use an image prediction model to predict a first predicted image frame of N detected objects and a second predicted image frame of each of the N detected objects based on a first image frame at multiple moments and N second image frames at each of the multiple moments.

[0076] Exemplarily, the image prediction model may include an optical flow prediction layer and an image prediction layer. Figure 4 As shown, S203 in the location determination method provided in the embodiment of the present application may include S401-S402.

[0077] S401 : Based on the optical flow field prediction layer in the image prediction model, process the first image frames of at least two historical moments or N second image frames of at least two historical moments to obtain the optical flow field at the current moment.

[0078] The position determination device may input the first image frames at multiple time points and N second image frames at each of the multiple time points into an image prediction model. The optical flow field prediction layer in the image prediction model first learns the optical flow field of the N detected objects at the current time point based on the first image frames at at least two historical time points or the N second image frames at at least two historical time points.

[0079] In the embodiments of the present application, the image prediction model may be a neural network model, such as a Long Short-Term Memory (LSTM) network, a Convolutional Neural Network (CNN), or a Convolutional LSTM. A Convolutional LSTM is a neural network composed of an LSTM and a CNN.

[0080] For example, take the N detection objects as 3 vehicles, the multiple time points include t1, t2, t3 and t4, and the image prediction model is an LSTM model, as shown in the following example: Figure 5As shown, the position determination device can determine a first image frame including three vehicles and three second image frames including three vehicles at each time t1, t2, t3, and t4, and input all the determined first image frames and all the second image frames into the LSTM model. The optical flow field prediction layer in the LSTM model can predict the optical flow field of the three vehicles at t4 based on the first image frames at t1, t2, and t3, or the second image frames at t1, t2, and t3.

[0081] It should be noted that the LSTM model can predict the optical flow field at a time point at least two historical moments later based on the first image frames of at least two historical moments, or the second image frames of at least two historical moments. Therefore, the LSTM model can predict the optical flow field of three vehicles at t3 based on the first image frames of t1 and t2, or the second image frames of t1 and t2. The LSTM model can also predict the optical flow field of three vehicles at t4 based on the first image frames of t2 and t3, or the second image frames of t2 and t3.

[0082] S402 : Based on the image prediction layer in the image prediction model and the optical flow field at the current moment, the first image frame at the current moment and the N second image frames at the current moment are processed respectively to obtain a first predicted image frame and N second predicted image frames.

[0083] The optical flow prediction layer in the image prediction model is connected to the image prediction layer in the image prediction model. The optical flow prediction layer in the image prediction model sends the learned optical flow field at the current moment to the image prediction layer. The image prediction layer processes the first image frame at the current moment based on the optical flow field to obtain a first predicted image frame, and processes the N second image frames at the current moment to obtain N second predicted image frames.

[0084] In some embodiments, the image prediction layer in the image prediction model may include a pre-processing layer and a convolution layer. The position determination device may first process the optical flow field at the current moment using the pre-processing layer in the image prediction model to obtain a convolution kernel. Then, based on the convolution layer in the image prediction model and the convolution kernel, the device may process the first image frame at the current moment and the N second image frames at the current moment, respectively, to obtain a first predicted image frame and N second predicted image frames.

[0085] Specifically, the convolution layer in the image prediction model can use the convolution kernel to perform a convolution operation on the first image frames of the N detection objects at the current moment to obtain a first predicted image frame including the N detection objects. The convolution layer in the image prediction model can also use the convolution kernel to process the second image frame of each detection object at the current moment to obtain a second predicted image frame including the detection object, thereby obtaining N second predicted image frames.

[0086] Exemplarily, the convolution operation may be a conventional convolution operation, or a depth-wise convolution operation, and the like.

[0087] For example, let's take the LSTM model as an example. Figure 5 As shown in the figure, after the optical flow prediction layer in the LSTM model obtains the optical flow fields of the three vehicles at t4, the pre-processing layer in the LSTM model processes the optical flow fields of the three vehicles at t4 to obtain the convolution kernel for t4. The convolution layer in the LSTM model then uses this convolution kernel to convolve the first image frame of the three vehicles at t4 to obtain the first predicted image frame of the three vehicles at t5. The convolution kernel is also used to convolve the second image frame of each vehicle at t4 to obtain the second predicted image frame of the vehicle at t5. Here, t5 is the time immediately after t4.

[0088] It should be noted that, in addition to the method described in S201-S202 above, the position determination device can also determine the first image frame of the N detection objects at each moment and the second image frame of each of the N detection objects at each moment by referring to the method described in S201-S203 above to obtain the first predicted image frame of the N detection objects and the second predicted image frame of each of the N detection objects.

[0089] S204 : Determine an image plane position of each detection object in the first predicted image frame based on the first predicted image frame and the N second predicted image frames.

[0090] In the embodiment of the present application, the second predicted image frame of each detection object only includes the detection object. Therefore, the image plane position of the detection object at the next moment can be determined based on the second predicted image frame of the detection object. However, the first predicted image frame of N detection objects includes N detection objects, and the image plane positions of each of the N detection objects cannot be regioned in the first predicted image frame. Therefore, the position determination device can use the second predicted image frame of each detection object in the N detection objects to determine the image plane position of the detection object in the first predicted image frame.

[0091] Specifically, the position determination device can determine that the image plane position of a detection object in the first predicted image frame that is the same as the first image plane position is the image plane position of the detection object based on the image plane position (which can be called the first image plane position) of each detection object in the second predicted image frame.

[0092] It is understandable that the position determination device can accurately determine the image plane position of each detection object from the first predicted image frame including N detection objects. Furthermore, the position determination device can also determine the image plane position relationship between each detection object and other detection objects other than itself in the first predicted image frame. Based on the image plane position of each detection object in the first predicted image frame, and the image plane position relationship between each detection object and other detection objects other than itself in the first predicted image frame, a route planning can be performed for the detection object. For example, it is determined whether the road ahead on which the detection object is currently traveling is congested, whether there is an accident section, etc., so as to remind the detection object to avoid congested sections and accident sections.

[0093] In an embodiment of the present application, in order for the position determination device to make accurate predictions, it is necessary to train in advance (at least before S203) the image prediction model used in the aforementioned embodiment. Based on this, an embodiment of the present application also provides a method for training an image prediction model, which can be applied to a training device for an image prediction model, and the training device for the image prediction model can be an electronic device or a part of an electronic device. The electronic device can be the terminal or server in the aforementioned embodiment. Figure 6 As shown, the training method of an image prediction model provided by the present application may include S601-S603.

[0094] S601. Obtain multiple groups of input samples and output samples corresponding to each group of input samples; wherein each group of input samples includes: a first input image frame of M detection objects at each first moment in multiple first moments, and a second input image frame of each detection object among the M detection objects at the first moment; the output samples include: a first output image frame of the M detection objects at the second moment, and a second output image frame of each detection object among the M detection objects at the second moment; the second moment is after the first moment.

[0095] The training device can obtain multiple input samples and output samples corresponding to each group of input samples. Each group of input samples and a group of output samples corresponding to each group of input samples can be used to train an initial image prediction model. The first input image frame in each group of input samples includes M detection objects. The M second input image frames in each group of input samples only include one detection object among the M detection objects, and the detection objects included in the M second input image frames are different. The first output image frame in each group of output samples includes M detection objects. The M second output image frames in each group of output samples only include one detection object among the M detection objects, and the detection objects included in the M second output image frames are different. M is a positive integer greater than 1.

[0096] It should be noted that the details of the multiple first moments can refer to the detailed description of the multiple moments in the aforementioned embodiment. The details of the second moment can refer to the detailed description of the next moment in the aforementioned embodiment. The details of the first input image frame can refer to the detailed description of the first image frame at multiple moments in the aforementioned embodiment. The details of the first output image frame can refer to the detailed description of the second image frame at multiple moments in the aforementioned embodiment. The embodiments of this application will not be repeated here.

[0097] It should also be noted that the specific process of the training device obtaining the first input image frame of the M detection objects at each first moment can be referred to the detailed description of the training device determining the first image frame of the N detection objects at each moment in the aforementioned embodiment. The specific process of the training device obtaining the second input image frame of each of the M detection objects at each first moment can be referred to the detailed description of the training device determining the second image frame of each of the N detection objects at each moment in the aforementioned embodiment. The embodiments of the present application will not be repeated here.

[0098] S602. Input each group of input samples into the initial image prediction model to obtain a training output image, where the training output image includes a first predicted output image frame of the M detection objects at the second moment, and a second predicted output image frame of each of the M detection objects at the second moment.

[0099] After inputting each set of input samples into the initial image prediction model, the training device obtains a training output image corresponding to each set of input samples. The training output image includes: a first predicted output image frame including M detection objects at a second moment, and a second predicted output image frame including one detection object at a second moment.

[0100] In some embodiments, an initial image prediction model may be designed based on experience, and may be a model of any neural network, such as LSTM, CNN, or convolutional LSTM.

[0101] In an embodiment of the present application, the multiple first moments include a third moment and at least two fourth moments before the third moment. The training device can determine the motion state of the M detection objects at the third moment (i.e., the optical flow field of the M detection objects at the third moment) based on the image plane positions of the M detection objects at at least two fourth moments before the third moment. Among them, the first input image frames of at least two fourth moments can represent the image plane positions of the M detection objects at at least two fourth moments, and the M second input image frames of at least two fourth moments can also represent the image plane positions of the M detection objects at at least two fourth moments. Therefore, the position determination device can obtain the optical flow field of the M detection objects at the third moment based on the first input image frames of at least two fourth moments, or the M second input image frames of at least two fourth moments.

[0102] Furthermore, after the training device determines the motion state of the detection object at the third moment (i.e., the optical flow field of the detection object at the third moment), it can predict the image plane position of the detection object at the second moment after the third moment in combination with the image plane position of the detection object at the third moment. Specifically, the training device can use the optical flow field of the M detection objects at the third moment to process the first input image frame of the M detection objects at the third moment (used to represent the image plane position of the M detection objects at the third moment) to obtain the first predicted input image frame of the M detection objects at the second moment after the third moment. The first predicted input image frame can represent the image plane position of the M detection objects at the second moment. The training device can also use the optical flow field of the M detection objects at the third moment to process the second input image frame of each detection object in the M detection objects at the third moment (used to represent the image plane position of each detection object at the third moment) to obtain the second predicted input image frame of each detection object at the second moment after the third moment. The second predicted input image frame can represent the image plane position of each detection object at the second moment.

[0103] Exemplarily, the initial image prediction model may include an optical flow prediction layer and an image prediction layer. Figure 7 As shown, S602 in the image prediction model training method provided in the embodiment of the present application may include S701-S702.

[0104] S701. Based on the optical flow field prediction layer in the initial image prediction model, process at least two first input image frames at the fourth moment or at least two M second input image frames at the fourth moment to obtain a predicted optical flow field at the third moment.

[0105] The training device may input multiple first input image frames at a first moment and M second input image frames at each of the multiple first moments into an initial image prediction model. The optical flow field prediction layer in the initial image prediction model first learns to obtain the optical flow field of the M detection objects at a third moment based on at least two first input image frames at a fourth moment or at least two M second input image frames at a fourth moment.

[0106] It should be noted that the specific process of S701 can refer to the detailed description of S401 in the aforementioned embodiment, and will not be repeated here in the embodiment of this application.

[0107] S702 : Based on the image prediction layer in the initial image prediction model and the predicted optical flow field at the third moment, the first input image frame at the third moment and the M second input image frames at the third moment are processed respectively to obtain a training output image.

[0108] The optical flow prediction layer in the initial image prediction model is connected to the image prediction layer in the initial image prediction model. The optical flow prediction layer in the initial image prediction model sends the learned optical flow field at the third moment to the image prediction layer. Based on the optical flow field at the third moment, the image prediction layer processes the first input image frame at the third moment to obtain a first predicted output image frame at the second moment in the training output image. The image prediction layer then processes the M second input image frames at the third moment to obtain M second predicted output image frames at the second moment in the training output image.

[0109] In some embodiments, the image prediction layer in the initial image prediction model may include a pre-processing layer and a convolution layer. The training device may process the predicted optical flow field at the third moment based on the pre-processing layer in the initial image prediction model to obtain a predicted convolution kernel. The training device may then process the first input image frame at the third moment and the M second input image frames at the third moment based on the convolution layer and the predicted convolution kernel in the initial image prediction model to obtain a training output image.

[0110] Specifically, the convolution layer in the initial image prediction model can use the convolution kernel to perform a convolution operation on the first input image frame of the M detection objects at the third moment to obtain a first output predicted image frame including the M detection objects. The convolution layer in the initial image prediction model can also use the convolution kernel to process the second input image frame of each detection object at the third moment to obtain a second predicted output image frame including the detection object, thereby obtaining M second predicted output image frames.

[0111] It should be noted that the details of the convolution operation can be referred to the detailed introduction of the convolution operation in the aforementioned S402, and will not be repeated here in the embodiment of the present application.

[0112] S603 : Using the M second predicted output image frames and the M second output image frames as supervisory information, iteratively train the initial image prediction model to obtain a trained image prediction model.

[0113] The training device can use the M second output images in each set of input samples and the M second predicted output image frames in the training output images corresponding to each set of input samples as supervision information to iteratively train the initial image prediction model to obtain a trained image prediction model. The trained image prediction model is the image prediction model in the aforementioned embodiment.

[0114] It is understandable that in the process of training the initial image prediction model, since each second predicted output image frame only includes one detection object, the image plane position of the detection object in the second predicted output image can be clearly represented. Therefore, the image plane position of the detection object in the second predicted output image can be compared with the image plane position of the detection object in a second output image frame (i.e., M second predicted output image frames and M second output image frames are used as supervision information) to iteratively update the initial image detection module, thereby obtaining a trained image detection model with more accurate image plane position prediction.

[0115] In some embodiments, reference Figure 7 As shown, S603 may specifically include S703 and S704.

[0116] S703 : Determine a loss value according to the M second predicted output image frames and the M second output image frames.

[0117] For example, S703 may specifically use a loss function to determine the loss value. For example, the loss function may be specifically represented by the following formula (1):

[0118]

[0119] Among them, loss is the loss value. i (x) is the image plane position of the i-th detection object among the M detection objects in the second predicted output image frame, for example, the coordinate value of the i-th detection object in the second predicted output image frame. i is the image plane position of the i-th detection object in the second output image frame, for example, the coordinate value of the i-th detection object in the second output image frame.

[0120] S704: Iteratively update the initial image prediction model according to the loss value to obtain a trained image prediction model.

[0121] By continuously iteratively optimizing the loss value, an image prediction model can be obtained that can accurately determine the image plane position of the predicted detection object.

[0122] It can be understood that since each second predicted output image frame and each second output image frame only includes one detection object, a second predicted output image frame and a second output image frame including the same detection object can be determined from the M second predicted output image frames and the M second output image frames, and the loss value calculated based on the second predicted output image frame and the second output image frame can characterize the accuracy of the initial image prediction model in predicting the image plane position of the detection object. By iteratively updating the initial image prediction model using the loss value, the trained image prediction model can make the image plane position prediction of the detection object more accurate. It can be further seen that by calculating the loss value using the M second predicted output image frames and the M second output image frames, and iteratively updating the initial image prediction model based on the calculated loss value, the trained image prediction model can make the image plane position prediction of the M detection objects more accurate. It can also be said that the image plane position of the trained image prediction model has a higher accuracy in predicting the image plane position.

[0123] It is understandable that, in order to realize the above functions, the above electronic device includes hardware structures and / or software modules corresponding to the execution of each function. Those skilled in the art should easily realize that, in combination with the units and algorithm steps of each example described in the embodiments disclosed herein, the embodiments of the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the embodiments of the present application.

[0124] In the case of dividing each functional module into corresponding functional modules, the embodiment of the present application also provides a position determination device. Figure 8 FIG. 8 is a schematic diagram of a position determination device according to an embodiment of the present invention. The device may include: a position acquisition module 81 , an image acquisition module 82 , a prediction module 83 and a position determination module 84 .

[0125] Among them, the position acquisition module 81 is used to determine the position information of N detection objects at each of multiple moments, where the multiple moments include the current moment and at least two historical moments, and N is a positive integer greater than 1. The image acquisition module 82 is used to determine the first image frame of the N detection objects at each moment, and the second image frame of each of the N detection objects at each moment based on the position information of the N detection objects at each moment. The prediction module 83 is used to predict the first predicted image frame of the N detection objects and the second predicted image frame of each of the N detection objects based on the first image frames at multiple moments and the N second image frames at each of the multiple moments. The position determination module 84 is used to determine the image plane position of each detection object in the first predicted image frame based on the first predicted image frame and the N second predicted image frames.

[0126] In some embodiments, combined Figure 8 , refer to Figure 9 As shown, the prediction module 83 may include an optical flow field acquisition unit 831 and an image processing unit 832. The optical flow field acquisition unit 831 is configured to process the first image frame of at least two historical moments or the N second image frames of at least two historical moments based on the optical flow field prediction layer in the image prediction model to obtain the optical flow field at the current moment. The image processing unit 832 is configured to process the first image frame at the current moment and the N second image frames at the current moment based on the image prediction layer in the image prediction model and the optical flow field at the current moment to obtain the first predicted image frame and the N second predicted image frames.

[0127] In some embodiments, combined Figure 9 , refer to Figure 10 As shown, image processing unit 832 includes a pre-processing unit 8321 and a convolution unit 8322. The image prediction layer includes a pre-processing layer and a convolution layer. Pre-processing unit 8321 is configured to process the optical flow field at the current moment based on the pre-processing layer in the image prediction model to obtain a convolution kernel. Convolution unit 8322 is configured to process the first image frame at the current moment and the N second image frames at the current moment based on the convolution layer and convolution kernel in the image prediction model, respectively, to obtain a first predicted image frame and N second predicted image frames.

[0128] In some embodiments, combined Figure 8 , refer to Figure 10As shown, the image acquisition module 82 includes a first image acquisition unit 821 and a second image acquisition unit 822. The first image acquisition unit 821 is configured to generate a first image frame of the N detection objects at each moment based on the position information of the N detection objects at each moment and the environmental information of the N detection objects. The second image acquisition unit 822 is configured to generate a second image frame of each detection object at each moment based on the position information of each detection object at each moment and the environmental information of the detection object.

[0129] Regarding the position determination device in the above embodiment, the specific manner in which each module performs operations and the corresponding beneficial effects have been described in detail in the embodiment of the position determination method mentioned above, and will not be repeated here.

[0130] In the case of dividing each functional module into corresponding functional modules, the embodiment of the present application also provides a training device for an image prediction model. Figure 11 FIG. 1 is a schematic diagram of a training device for an image prediction model according to an embodiment of the present invention. The training device may include: a sample acquisition module 91 , an image prediction module 92 , and an iterative training module 93 .

[0131] Specifically, the sample acquisition module 91 is used to obtain multiple groups of input samples and output samples corresponding to each group of input samples; wherein each group of input samples includes: the first input image frame of each first moment of M detection objects in multiple first moments, and the second input image frame of each detection object in the M detection objects at the first moment; the output samples include: the first output image frame of the M detection objects at the second moment, and the second output image frame of each detection object in the M detection objects at the second moment; the second moment is after the first moment, and M is a positive integer greater than 1.

[0132] The image prediction module 92 is used to input each group of input samples into the initial image prediction model to obtain a training output image, where the training output image includes a first predicted output image frame of the M detection objects at the second moment, and a second predicted output image frame of each detection object in the M detection objects at the second moment.

[0133] The iterative training module 93 is configured to iteratively train the initial image prediction model using the M second predicted output image frames and the M second output image frames as supervisory information to obtain a trained image prediction model.

[0134] In some embodiments, combined Figure 11 , refer to Figure 12As shown, iterative training module 93 includes a loss unit 931 and an iteration unit 932. Loss unit 931 is configured to determine a loss value based on the M second predicted output image frames and the M second output image frames. Iterative unit 932 is configured to iteratively update the initial image prediction model based on the loss value to obtain a trained image prediction model.

[0135] In some embodiments, combined Figure 11 , refer to Figure 12 As shown, the image prediction module 92 includes an optical flow prediction unit 921 and an image processing unit 922. The multiple first moments include a third moment and at least two fourth moments before the third moment.

[0136] Specifically, the optical flow prediction unit 921 is configured to process at least two first input image frames at the fourth moment or at least two M second input image frames at the fourth moment based on the optical flow field prediction layer in the initial image prediction model to obtain a predicted optical flow field at the third moment. The image processing unit 922 is configured to process the first input image frame at the third moment and the M second input image frames at the third moment based on the image prediction layer in the initial image prediction model and the predicted optical flow field at the third moment to obtain a training output image.

[0137] In some embodiments, combined Figure 12 , refer to Figure 13 As shown, image processing unit 922 includes a pre-processing unit 9221 and a convolution unit 9222. The image prediction layer includes a pre-processing layer and a convolution layer. Pre-processing unit 9221 is configured to process the predicted optical flow field at the third moment based on the pre-processing layer in the initial image prediction model to obtain a predicted convolution kernel. Convolution unit 9222 is configured to process the first input image frame at the third moment and the M second input image frames at the third moment based on the convolution layer and predicted convolution kernel in the initial image prediction model to obtain a training output image.

[0138] The trained image prediction model here is the image prediction model used in the position determination method in the aforementioned embodiment.

[0139] Figure 14 This is a possible structural diagram of an electronic device according to an exemplary embodiment. The electronic device may be the above-mentioned position determination device or image prediction model training device, or may be a terminal or server including the position determination device and / or image prediction model training device. Figure 14As shown, the electronic device includes a processor 1001 and a memory 1002. The memory 1002 is used to store instructions executable by the processor 1001, and the processor 1001 can implement the functions of each module in the position determination device and / or the image prediction model training device in the above-mentioned embodiments. The memory 1002 stores at least one instruction, which is loaded and executed by the processor 1001 to implement the position determination method and / or image prediction model training method provided in the above-mentioned various method embodiments.

[0140] In a specific implementation, as an embodiment, the processor 1001 (eg, processor 1001-1 and processor 1001-2) may include one or more CPUs, for example Figure 14 As an embodiment, the electronic device may include multiple processors 1001, such as CPU0 and CPU1. Figure 14 1 and 1001 - 2 are shown in FIG. Each CPU in processors 1001 may be a single-core processor (Single-CPU) or a multi-core processor (Multi-CPU). Processor 1001 herein may refer to one or more devices, circuits, and / or processing cores for processing data (e.g., computer program instructions).

[0141] The memory 1002 may be a read-only memory (ROM) or other type of static storage device that can store static information and instructions, a random access memory (RAM) or other type of dynamic storage device that can store information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, an optical disc storage (including a compact disc, laser disc, optical disc, digital versatile disc, Blu-ray disc, etc.), a magnetic disk computer storage medium or other magnetic storage device, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 1002 may exist independently and be connected to the processor 1001 via the communication bus 1003. The memory 1002 may also be integrated with the processor 1001.

[0142] The communication bus 1003 can be an industry standard architecture (ISA) bus, a peripheral component interconnect (PCI) bus, or an extended industry standard architecture (EISA) bus. The communication bus 1003 can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 14 Only one thick line is used in the diagram, but this does not mean that there is only one bus or one type of bus.

[0143] In addition, in order to facilitate information interaction between the electronic device and other devices (for example, when the electronic device is a terminal, it interacts with the server, or when the electronic device is a server, it interacts with the terminal), the electronic device includes a communication interface 1004. The communication interface 1004 uses any transceiver-like device to communicate with other devices or communication networks, such as control systems, radio access networks (RAN), wireless local area networks (WLAN), etc. The communication interface 1004 may include a receiving unit to implement a receiving function, and a sending unit to implement a sending function. The communication interface 1004, the processor 1001, and the memory 1002 are connected through a communication bus 1003 to complete mutual communication.

[0144] The present application also provides a computer-readable storage medium storing computer instructions, which, when executed on an electronic device, cause the electronic device to execute the position determination method and / or the image prediction model training method in the above method embodiments.

[0145] For example, the computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a compact disc (CD-ROM), a magnetic tape, a floppy disk, an optical data storage device, and the like.

[0146] An embodiment of the present application also provides a computer program product comprising computer instructions, which, when executed on an electronic device, enables the electronic device to execute the position determination method and / or the image prediction model training method in the above-mentioned method embodiment.

[0147] Among them, the electronic device, computer-readable storage medium or computer program product provided in the embodiments of the present application are all used to execute the corresponding methods provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding methods provided above, and will not be repeated here.

[0148] Through the description of the above implementation methods, those skilled in the art can clearly understand that for the convenience and brevity of description, only the division of the above-mentioned functional modules is used as an example. In actual applications, the above-mentioned functions can be assigned to different functional modules as needed, that is, the internal structure of the device (such as an electronic device) is divided into different functional modules to complete all or part of the functions described above. The specific working process of the system, device (such as an electronic device) and unit described above can refer to the corresponding process in the aforementioned method embodiment and will not be repeated here.

[0149] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices (such as electronic devices) and methods can be implemented in other ways. For example, the device (such as electronic devices) embodiments described above are merely schematic. For example, the division of the modules or units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.

[0150] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0151] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0152] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as flash memory, mobile hard disk, read-only memory, random access memory, magnetic disk or optical disk.

[0153] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

Claims

1. A method for determining a position, comprising: Determine location information of N detection objects at each of a plurality of moments, where the plurality of moments include a current moment and at least two historical moments, and N is a positive integer greater than 1; Determining, based on the position information of the N detection objects at each moment, a first image frame of the N detection objects at each moment, and a second image frame of each of the N detection objects at each moment; Predicting a first predicted image frame of the N detected objects and a second predicted image frame of each of the N detected objects based on the first image frames at the multiple moments and N second image frames at each of the multiple moments; Based on the first predicted image frame and N second predicted image frames, an image plane position of each of the detection objects in the first predicted image frame is determined.

2. The method according to claim 1, wherein The predicting, based on the first image frames at the multiple moments and N second image frames at each of the multiple moments, first predicted image frames of the N detected objects and second predicted image frames of each of the N detected objects includes: Processing the first image frames of the at least two historical moments or the N second image frames of the at least two historical moments based on the optical flow field prediction layer in the image prediction model to obtain the optical flow field at the current moment; Based on the image prediction layer in the image prediction model and the optical flow field at the current moment, the first image frame at the current moment and the N second image frames at the current moment are processed respectively to obtain the first predicted image frame and N second predicted image frames.

3. The method according to claim 2, wherein: The image prediction layer includes a pre-processing layer and a convolutional layer; The processing of the first image frame at the current moment and the N second image frames at the current moment based on the image prediction layer in the image prediction model and the optical flow field at the current moment to obtain the first predicted image frame and the N second predicted image frames includes: Processing the optical flow field at the current moment based on the pre-processing layer in the image prediction model to obtain a convolution kernel; Based on the convolution layer and the convolution kernel in the image prediction model, the first image frame at the current moment and the N second image frames at the current moment are processed respectively to obtain the first predicted image frame and N second predicted image frames.

4. The method according to any one of claims 1 to 3, wherein The determining, based on the position information of the N detection objects at each moment, a first image frame of the N detection objects at each moment, and a second image frame of each of the N detection objects at each moment, comprises: generating a first image frame of the N detection objects at each moment based on position information of the N detection objects at each moment and information about the environment in which the N detection objects are located; Based on the position information of each of the detection objects at each moment and the environment information where the detection objects are located, a second image frame of each of the detection objects at each moment is generated.

5. A method for training an image prediction model, comprising: Acquire multiple groups of input samples and output samples corresponding to each group of the input samples; wherein each group of the input samples includes: a first input image frame of M detection objects at each first moment among multiple first moments, and a second input image frame of each of the M detection objects at the first moment; the output samples include: a first output image frame of the M detection objects at a second moment, and a second output image frame of each of the M detection objects at the second moment; the second moment is after the first moment, and M is a positive integer greater than 1; Inputting each group of input samples into an initial image prediction model to obtain a training output image, wherein the training output image includes a first predicted output image frame of the M detection objects at the second moment and a second predicted output image frame of each detection object in the M detection objects at the second moment; The initial image prediction model is iteratively trained using the M second predicted output image frames and the M second output image frames as supervisory information to obtain a trained image prediction model.

6. The method according to claim 5, wherein: The method of iteratively training the initial image prediction model using the M second predicted output image frames and the M second output image frames as supervisory information to obtain a trained image prediction model includes: determining a loss value according to the M second predicted output image frames and the M second output image frames; The initial image prediction model is iteratively updated according to the loss value to obtain the trained image prediction model.

7. The method according to claim 5 or 6, wherein: The multiple first moments include a third moment and at least two fourth moments before the third moment, and inputting each group of input samples into an initial image prediction model to obtain a training output image includes: Processing the at least two first input image frames at the fourth moments or the M second input image frames at the at least two fourth moments based on the optical flow field prediction layer in the initial image prediction model to obtain a predicted optical flow field at the third moment; Based on the image prediction layer in the initial image prediction model and the predicted optical flow field at the third moment, the first input image frame at the third moment and the M second input image frames at the third moment are processed respectively to obtain the training output image.

8. The method according to claim 7, wherein: The image prediction layer includes a pre-processing layer and a convolutional layer; The step of processing the first input image frame at the third moment and the M second input image frames at the third moment based on the image prediction layer in the initial image prediction model and the predicted optical flow field at the third moment to obtain the training output image includes: Processing the predicted optical flow field at the third moment based on the pre-processing layer in the initial image prediction model to obtain a predicted convolution kernel; Based on the convolution layer and the prediction convolution kernel in the initial image prediction model, the first input image frame at the third moment and the M second input image frames at the third moment are processed respectively to obtain the training output image.

9. A position determination device comprising: a position acquisition module, configured to determine the position information of N detection objects at each of a plurality of moments, wherein the plurality of moments include a current moment and at least two historical moments, and N is a positive integer greater than 1; an image acquisition module, configured to determine, based on position information of the N detection objects at each moment, a first image frame of the N detection objects at each moment, and a second image frame of each of the N detection objects at each moment; a prediction module, configured to predict, based on the first image frames at the multiple moments and the N second image frames at each of the multiple moments, a first predicted image frame of the N detected objects and a second predicted image frame of each of the N detected objects; A position determination module is used to determine the image plane position of each of the detection objects in the first prediction image frame based on the first prediction image frame and N second prediction image frames.

10. A training device for an image prediction model, comprising: a sample acquisition module, configured to acquire multiple groups of input samples and output samples corresponding to each group of the input samples; wherein each group of the input samples includes: a first input image frame of M detection objects at each first moment among multiple first moments, and a second input image frame of each of the M detection objects at the first moment; the output samples include: a first output image frame of the M detection objects at a second moment, and a second output image frame of each of the M detection objects at the second moment; the second moment is after the first moment, and M is a positive integer greater than 1; an image prediction module, configured to input each group of input samples into an initial image prediction model to obtain a training output image, wherein the training output image includes a first predicted output image frame of the M detection objects at the second moment and a second predicted output image frame of each detection object in the M detection objects at the second moment; An iterative training module is used to iteratively train the initial image prediction model using M second predicted output image frames and M second output image frames as supervision information to obtain a trained image prediction model.

11. A computer-readable storage medium storing a computer program, wherein the computer program is used to execute the position determination method described in any one of claims 1 to 4, or the image prediction model training method described in any one of claims 5 to 8.

12. An electronic device, comprising: processor; a memory for storing instructions executable by the processor; The processor is used to read the executable instructions from the memory and execute the instructions to implement the position determination method described in any one of claims 1-4 above, or to execute the image prediction model training method described in any one of claims 5-8 above.

Citation Information

Patent Citations

  • Behavior recognition lightweight method, system and equipment based on multi-target tracking

    CN113158909A

  • Image similarity calculation method and device, equipment and medium

    CN115375926A