Image-based Speed Determination Method, Apparatus, Device, and Storage Medium

By extracting and fusing features from the two-frame images, combining convolution operation and spatial position, the problem of large velocity prediction error caused by unstable image brightness is solved, and accurate and efficient determination of motion speed is achieved.

CN115830065BActive Publication Date: 2025-07-29HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO LTD

Patent Information

Application Number
CN202211399309.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-09
Publication Date
2025-07-29
Estimated Expiration
2042-11-09

AI Technical Summary

Technical Problem

In the prior art, there is a problem that the motion speed prediction error and low accuracy due to unstable brightness of the input image.

Method used

By extracting different features from the two frames of images and fusion, the displacement information feature map of the target object is determined using convolution operations, and the motion speed is obtained based on the spatial position of the target object.

Benefits of technology

It realizes the accurate determination of the movement speed of the target object under different image brightness conditions, and improves the accuracy and efficiency of velocity prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115830065B_ABST
    Figure CN115830065B_ABST
Patent Text Reader

Abstract

The present application provides an image-based speed determination method, apparatus, device, and storage medium, which relate to the technical field of image recognition. The method includes: obtaining a first image and a second image during the movement of a target object in an environmental space, where the acquisition time of the first image is later than that of the second image; fusing a first feature extracted from the first image and a second feature extracted from the second image to obtain a fused feature map, where the fused feature map includes first pixel position information and second pixel position information of the target object; performing a first convolution operation on the fused feature map to obtain a displacement information feature map, where the displacement information feature map is used to represent the correspondence between pixel positions and displacements; extracting a target displacement corresponding to a target pixel position from the displacement information feature map; and determining the movement speed of the target object at the acquisition time of the first image according to the target displacement.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image recognition technology, and particularly to an image-based speed determination method, apparatus, device, and storage medium. Background Art

[0002] Currently, it has become a general trend to estimate the motion speed of a target object by using the video of the target object taken. For example, if the target object is a target vehicle moving relative to the present vehicle, the current driving speed of the target vehicle can be determined by taking the video of the target vehicle.

[0003] In the related art, the optical flow method is usually adopted to analyze the moving image to predict the motion speed of the target object in the image. The implementation of this method is based on the premise that "the image brightness of the input image is constant", but in practical applications, it is very difficult to ensure that the image brightness of the target object is constant. It can be seen that when using the above method to predict the speed, due to the unstable image brightness of the input image, problems such as large error and low accuracy of the predicted running speed will occur. Therefore, how to avoid the influence of the characteristics of the input image and improve the accuracy of the predicted running speed has become an urgent technical problem to be solved at present. Summary of the Invention

[0004] Based on the above technical problems, this application provides an image-based speed determination method, apparatus, device, and storage medium, which are used to solve the problems such as large error and low accuracy of the predicted running speed caused by the unstable image brightness of the input image.

[0005] In a first aspect, an image-based speed determination method is provided. The method includes: acquiring a first image and a second image of a target object during the movement in the environmental space, where the acquisition time of the first image is later than that of the second image; fusing a first feature extracted from the first image and a second feature extracted from the second image to obtain a fused feature map, where the fused feature map includes first pixel position information and second pixel position information of the target object, the first pixel position information represents the pixel position of the target object in the first image, and the second pixel position information represents the pixel position of the target object in the second image; performing a first convolution operation on the fused feature map to obtain a displacement information feature map, where the displacement information feature map is used to represent the corresponding relationship between the pixel position and the displacement; extracting a target displacement corresponding to a target pixel position from the displacement information feature map, where the target pixel position is the pixel position in the displacement information feature map corresponding to the first pixel position information of the target object; and determining the motion speed of the target object at the acquisition time of the first image according to the target displacement.

[0006] It should be understood that the above first feature and second feature include underlying features and multi-degree features at different environmental space angles.

[0007] In the speed determination method provided by this application, different features of the target object at different positions in the environmental space are extracted from two frames of images respectively, that is, the first feature and the second feature. Then, by fusing the first feature and the second feature corresponding to the two frames of images into a fused feature, the first feature and the second feature are associated. Since the first feature and the second feature have different feature angles for representing the environmental space, fusing the first feature and the second feature into a single fused feature is more conducive to identifying and extracting the feature differences of the target object in the environmental space in the two frames of images. Therefore, the displacement information feature map obtained based on the fused feature can more accurately determine the displacement of the target object. Based on this, by extracting the information of two frames of images of the target object at different acquisition times and performing corresponding convolution operations, the displacement of the target object can be directly determined, and thus the motion speed of the target object at the acquisition time of the first image can be determined. This speed determination method is not limited by the image characteristics (such as image brightness) of the input image and can be used for various input images with different frames. It has a wide range of application scenarios and a simple determination process, avoiding the problem of low accuracy of the determined motion speed of the target object caused by the limitation of the image characteristics of the input image in the related art.

[0008] In a possible implementation, the method further includes: obtaining the target space position of the target object at the acquisition time of the first image, where the target space position is the three-dimensional position of the target object in the environmental space; determining the two-dimensional projection position of the target space position projected onto the first image; and converting the two-dimensional projection position into a target pixel position according to the ratio of the size of the first image to the size of the displacement information feature map.

[0009] It should be understood that at least one pixel position on the first image includes the two-dimensional projection position of the target space position projected onto the first image, and there is a projection conversion relationship between the space position and the first image, that is, when the space position is known, the pixel position projected onto the first image is known. Therefore, this implementation provides a method for obtaining the target pixel position based on the target space position of the target object according to the conversion relationship between the two-dimensional pixels in the first image and the space position of the target object in the environmental space. Compared with determining the target pixel position by extracting the pixel position of the target object from the first image, this implementation is more in line with the environmental space scenario based on the space position of the target object, and thus the obtained target pixel position is more accurate.

[0010] In another possible implementation, the speed determination method further includes: performing a second convolution operation on the fused feature map to obtain a displacement uncertainty feature map; the displacement uncertainty feature map is used to represent the correspondence between the pixel position and the displacement uncertainty, and the displacement uncertainty represents the accuracy of the displacement corresponding to any pixel position in the displacement information feature map; and extracting the displacement uncertainty corresponding to the target displacement from the displacement uncertainty feature map according to the target pixel position.

[0011] It is understandable that the displacement corresponds one-to-one with the displacement uncertainty.

[0012] In this implementation manner, the magnitude of the displacement uncertainty reflects the accuracy of the determined target displacement. That is, the larger the displacement uncertainty, the lower the accuracy of the obtained target displacement. Correspondingly, the smaller the displacement uncertainty, the higher the accuracy of the obtained target displacement. Therefore, in practical applications, multiple sets of two frames of images can be collected to obtain multiple target displacements and multiple displacement uncertainties. Then, from the multiple target displacements, the target displacement with the smallest corresponding displacement uncertainty is selected, and the target displacement with the smallest displacement uncertainty is used to determine the motion speed of the target object, so as to make the determination of the motion speed of the target object more accurate.

[0013] In another possible implementation manner, the first feature extracted from the first image is fused with the second feature extracted from the second image to obtain a fused feature map, and a first convolution operation is performed on the fused feature map to obtain a displacement information feature map, including: inputting the first image and the second image into a displacement determination model to obtain the displacement information feature map output by the displacement determination model.

[0014] It can be understood that the above displacement determination model is a model trained based on training samples.

[0015] In this implementation manner, by inputting two frames of images of the target object at different acquisition times into the displacement determination model, the displacement of the target object can be directly determined, and thus the motion speed of the target object at the acquisition time of the first image can be determined. The above input of the two frames of images into the displacement determination model replaces the steps of feature extraction, fusion, and execution of the first convolution operation in one step, and can determine the displacement more quickly, thereby improving the efficiency of determining the motion speed of the target object.

[0016] In another possible implementation manner, the displacement determination model includes a feature extraction module, a fusion module, and a displacement information feature map determination module; the feature extraction module is used to extract the first feature from the first image and the second feature from the second image; the fusion module is used to fuse the first feature and the second feature to output a fused feature map; the displacement information feature map determination module is used to output a displacement information feature map according to the fused feature map.

[0017] In another possible implementation manner, the displacement determination model further includes a displacement uncertainty feature map determination module; the displacement uncertainty feature map determination module is used to output a displacement uncertainty feature map according to the fused feature map.

[0018] In this implementation method, by inputting two frames of images of the target object at different acquisition times into the displacement determination model, the displacement uncertainty of the target object can be directly determined. That is, by inputting the fused feature map obtained from the two frames of images into the displacement uncertainty feature map determination module, the displacement uncertainty of the target object can be obtained. This method is simple and direct, facilitating practical applications.

[0019] In a second aspect, a training method for a displacement determination model is provided. The method includes: obtaining a training sample set, where each training sample in the training sample set includes a first training image and a second training image during the movement of the sample object, and a label corresponding to the training sample; the first sampling time of the first training image is later than the second sampling time of the second training image; the label includes the first true pixel position of the sample object in the first training image and the true displacement generated by the sample object between the first sampling time and the second sampling time; using the training sample set to train a preset initial model to obtain a displacement determination model; where the initial model includes a feature extraction module, a fusion module, and a displacement information feature map determination module; the feature extraction module is used to extract a first feature from the first training image and a second feature from the second training image; the fusion module is used to fuse the first feature and the second feature to obtain a fused feature map, the fused feature map includes the first pixel position information and the second pixel position information of the sample object, the first pixel position information represents the pixel position of the sample object in the first training image, and the second pixel position information represents the pixel position of the sample object in the second image; the displacement information feature map determination module is used to output a displacement information feature map according to the fused feature map, and the displacement information feature map represents the corresponding relationship between the pixel position and the displacement; the displacement corresponding to the first true pixel position in the displacement information feature map is the predicted displacement output by the initial model.

[0020] In a possible implementation method, using the training sample set to train a preset initial model to obtain a displacement determination model includes: inputting the training sample into the initial model to obtain the predicted displacement output by the initial model; determining a first training loss according to the predicted displacement and the true displacement; adjusting the parameters of the initial model according to the training loss to obtain a displacement determination model; the training loss includes the first training loss.

[0021] In another possible implementation, the tag further includes the true pixel displacement and the first true position of the sample object, and the training method further includes: predicting the true position of the sample object at the second sampling time based on the predicted displacement and the first true position to obtain the predicted second true position; predicting the true pixel position of the sample object in the second training image based on the predicted second true position to obtain the predicted second true pixel position; obtaining the predicted pixel displacement based on the predicted second true pixel position and the first true pixel position; determining the second training loss based on the predicted pixel displacement and the true pixel displacement; and the training loss further includes the second training loss.

[0022] In another possible implementation, the initial model further includes a displacement uncertainty feature map determination module; the displacement uncertainty feature map determination module is configured to output a displacement uncertainty feature map according to the fused feature map, and the displacement uncertainty feature map is used to represent the correspondence between the pixel position and the displacement uncertainty, and the displacement uncertainty represents the accuracy of the displacement corresponding to any pixel position in the displacement information feature map.

[0023] In another possible implementation, using a training sample set to train a preset initial model to obtain a displacement determination model further includes: determining a third training loss based on the displacement uncertainty and the true displacement; and the training loss further includes the third training loss.

[0024] In a third aspect, an image-based speed determination device is provided. The speed determination device includes: a first acquisition unit configured to acquire a first image and a second image during the movement of a target object in an environmental space, where the acquisition time of the first image is later than that of the second image; a fusion unit configured to fuse a first feature extracted from the first image and a second feature extracted from the second image to obtain a fused feature map, where the fused feature map includes first pixel position information and second pixel position information of the target object, the first pixel position information represents the pixel position of the target object in the first image, and the second pixel position information represents the pixel position of the target object in the second image; a convolution unit configured to perform a first convolution operation on the fused feature map to obtain a displacement information feature map, where the displacement information feature map is used to represent the correspondence between the pixel position and the displacement; a displacement determination unit configured to extract a target displacement corresponding to a target pixel position from the displacement information feature map, where the target pixel position is the pixel position in the displacement information feature map corresponding to the first pixel position information of the target object; and a speed determination unit configured to determine the movement speed of the target object at the acquisition time of the first image according to the target displacement.

[0025] In a possible implementation, the displacement determination unit is further configured to specifically perform: obtaining the target spatial position of the target object at the acquisition time of the first image, where the target spatial position is the three-dimensional position of the target object in the environmental space; determining the two-dimensional projection position of the target spatial position projected on the first image; and converting the two-dimensional projection position into a target pixel position according to the ratio of the size of the first image to the size of the displacement information feature map.

[0026] In another possible implementation, the speed determination device further includes a displacement uncertainty determination unit, which is configured to specifically perform a second convolution operation on the fusion feature map to obtain a displacement uncertainty feature map; the displacement uncertainty feature map is used to characterize the correspondence between the pixel position and the displacement uncertainty, and the displacement uncertainty represents the accuracy of the displacement corresponding to any pixel position in the displacement information feature map; and extracting the displacement uncertainty corresponding to the target displacement from the displacement uncertainty feature map according to the target pixel position.

[0027] In another possible implementation, the speed determination device is further configured to perform: inputting the first image and the second image into a displacement determination model to obtain a displacement information feature map output by the displacement determination model.

[0028] In another possible implementation, the displacement determination model includes a feature extraction module, a fusion module, and a displacement information feature map determination module; the feature extraction module is used to extract a first feature from the first image and a second feature from the second image; the fusion module is used to fuse the first feature and the second feature to output a fusion feature map; and the displacement information feature map determination module is used to output a displacement information feature map according to the fusion feature map.

[0029] In another possible implementation, the displacement determination model further includes a displacement uncertainty feature map determination module; the displacement uncertainty feature map determination module is used to output a displacement uncertainty feature map according to the fusion feature map.

[0030] Fourthly, a training device for a displacement determination model is provided. The training device includes: a second acquisition unit configured to acquire a training sample set, where each training sample in the training sample set includes a first training image and a second training image during the movement of a sample object, and a label corresponding to the training sample; the first sampling time of the first training image is later than the second sampling time of the second training image; the label includes the first true pixel position of the sample object in the first training image and the true displacement generated by the sample object between the first sampling time and the second sampling time; a feature map training unit configured to use the training sample set to train a preset initial model to obtain a displacement determination model; where the initial model includes a feature extraction module, a fusion module, and a displacement information feature map determination module; the feature extraction module is used to extract a first feature from the first training image and a second feature from the second training image; the fusion module is used to fuse the first feature and the second feature to obtain a fused feature map, and the fused feature map includes the first pixel position information and the second pixel position information of the sample object, the first pixel position information represents the pixel position of the sample object in the first training image, and the second pixel position information represents the pixel position of the sample object in the second image; the displacement information feature map determination module is used to output a displacement information feature map according to the fused feature map, and the displacement information feature map represents the corresponding relationship between the pixel position and the displacement; the displacement corresponding to the first true pixel position in the displacement information feature map is the predicted displacement output by the initial model.

[0031] In a possible implementation, the feature map training unit is configured to specifically execute inputting the training sample into the initial model to obtain the predicted displacement output by the initial model; determining a first training loss according to the predicted displacement and the true displacement; adjusting the parameters of the initial model according to the training loss to obtain a displacement determination model; the training loss includes the first training loss.

[0032] In another possible implementation, the feature map training unit is further configured to specifically execute predicting the true position of the sample object at the second sampling time according to the predicted displacement and the first true position to obtain a predicted second true position; predicting the true pixel position of the sample object in the second training image according to the predicted second true position to obtain a predicted second true pixel position; obtaining a predicted pixel displacement according to the predicted second true pixel position and the first true pixel position; determining a second training loss according to the predicted pixel displacement and the true pixel displacement; the training loss further includes the second training loss.

[0033] In another possible implementation, a displacement uncertainty feature map determination module is configured to output a displacement uncertainty feature map based on the fused feature map. The displacement uncertainty feature map is used to represent the correspondence between pixel positions and displacement uncertainties, and the displacement uncertainty represents the accuracy of the displacement corresponding to any pixel position in the displacement information feature map.

[0034] In another possible implementation, the feature map training unit is further configured to specifically execute determining a third training loss based on the displacement uncertainty and the true displacement; the training loss further includes the third training loss.

[0035] In a fifth aspect, an image-based speed determination device is provided, including: a memory and a processor; the memory and the processor are coupled; the memory is used to store computer program code, and the computer program code includes computer instructions; wherein, when the processor executes the computer instructions, the image-based speed determination device is caused to execute the image-based speed determination method according to the first aspect.

[0036] In a sixth aspect, a computer-readable storage medium is provided, and the computer-readable storage medium includes computer instructions; wherein, when the computer instructions run on the image-based speed determination device, the image-based speed determination device is caused to execute the image-based speed determination method according to the first aspect.

[0037] In a seventh aspect, an image-based speed determination device is provided, including: a memory and a processor; the memory and the processor are coupled; the memory is used to store computer program code, and the computer program code includes computer instructions; wherein, when the processor executes the computer instructions, the training device of the displacement determination model is caused to execute the training method of the displacement determination model according to the second aspect.

[0038] In an eighth aspect, a computer-readable storage medium is provided, and the computer-readable storage medium includes computer instructions; wherein, when the computer instructions run on the training device of the displacement determination model, the training device of the displacement determination model is caused to execute the training method of the displacement determination model according to the second aspect.

[0039] In a ninth aspect, a computer device is provided, including: a processor and a memory. The processor is connected to the memory, and the memory is used to store computer execution instructions. The processor executes the computer execution instructions stored in the memory, thereby implementing any one of the methods provided in the first aspect and the second aspect.

[0040] In a tenth aspect, a computer program product is provided, including computer execution instructions. When the computer execution instructions run on a computer, the computer is caused to execute any one of the methods provided in the first aspect and the second aspect.

[0041] For the technical effects brought by any of the implementation manners in the second aspect to the tenth aspect, reference may be made to the technical effects brought by the corresponding implementation manner in the first aspect, which will not be elaborated herein. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0043] Figure 1 Schematic diagram of the composition of the speed prediction system provided by the embodiment of the present application;

[0044] Figure 2 Schematic diagram of the application scenario of the method for determining speed based on an image provided by the embodiment of the present application;

[0045] Figure 3 Schematic diagram of the composition of the speed determination device provided by the embodiment of the present application;

[0046] Figure 4 Schematic flowchart of the method for determining speed based on an image provided by the embodiment of the present application;

[0047] Figure 5 Schematic diagram of the feature fusion process provided by the embodiment of the present application;

[0048] Figure 6 Schematic diagram of the network structure of the displacement determination model provided by the embodiment of the present application;

[0049] Figure 7 Another schematic flowchart of the method for determining speed based on an image provided by the embodiment of the present application;

[0050] Figure 8 Schematic diagram of a speed determination device provided by the embodiment of the present application;

[0051] Figure 9 Schematic diagram of a training device provided by the embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0052] In the description of this application, unless otherwise specified, " / " means "or". For example, A / B can mean A or B. "And / or" in this article is just a relationship describing associated objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist simultaneously, and B exists alone. In addition, "at least one" means one or more, and "a plurality" means two or more. The words such as "first" and "second" do not limit the quantity and execution order, and the words such as "first" and "second" do not necessarily limit to be different.

[0053] It should be noted that in this application, words such as "exemplary" or "for example" are used to give examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in this application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Exactly, the use of words such as "exemplary" or "for example" aims to present relevant concepts in a specific way.

[0054] First, for the convenience of understanding this application, the relevant elements involved in this application are described below.

[0055] Low-level features of an image: contours, edges, colors, textures, and shape features. Edges and contours can reflect the content of the image; if edges and key points can be reliably extracted, many vision problems can be basically solved. The semantic information of the low-level features of the image is relatively less, but the target position is accurate.

[0056] Multi-scale features of an image: Extract high-level semantic features in both spatial and channel directions by using convolution kernels of various different sizes to extract image features. Among them, the high-level semantic features of an image refer to the features that can be seen. For example, when extracting low-level features of a vehicle, features such as the contour and color of the vehicle can be extracted, then the high-level features are displayed as a vehicle. The semantic information of the high-level features is relatively rich, but the target position is relatively rough.

[0057] Process of extracting fused features of an image: Directly connect the low-level and high-level feature maps to transfer low-level information to the high level.

[0058] Secondly, a simple introduction to the application scenarios involved in this application is given.

[0059] Currently, it has become a common trend to estimate the motion speed of a target object by using the video of the captured target object. For example, if the target object is a target vehicle moving relative to the host vehicle, the current driving speed of the target vehicle can be determined by capturing the video of the target vehicle. In related technologies, the optical flow method is usually used to analyze the moving image to predict the motion speed of the target object in the image. The implementation of this method is based on the premise that "the image brightness of the input image is constant". However, in actual applications, it is very difficult to ensure that the image brightness of the target object is constant. It can be seen that when using the above method to predict the speed, due to the unstable image brightness of the input image, problems such as large errors and low accuracy in the predicted running speed will occur. Therefore, how to avoid the influence of the characteristics of the input image and improve the accuracy of the predicted running speed has become an urgent technical problem to be solved at present.

[0060] Based on this, the present application provides a speed determination method, which extracts different features of the target object at different positions in the environmental space from two frames of images, namely, the first feature and the second feature. Then, by fusing the first feature and the second feature corresponding to the two frames of images into a fusion feature, the first feature and the second feature are associated. Since the first feature and the second feature have different feature angles for representing the environmental space, fusing the first feature and the second feature into a single fusion feature is more conducive to identifying and extracting the feature differences of the target object in the environmental space in the two frames of images. Therefore, the displacement information feature map obtained based on the fusion feature can more accurately determine the displacement of the target object. Based on this, by extracting the information of two frames of images of the target object at different acquisition times and performing corresponding convolution operations, the displacement of the target object can be directly determined, thereby determining the motion speed of the target object at the acquisition time of the first image. This speed determination method is not limited by the characteristics of the input image (such as image brightness), and can be used for various input images of different frames. It has a wide range of usage scenarios and a simple determination process, avoiding the problem of low accuracy of the determined motion speed of the target object due to the limitation of the characteristics of the input image in related technologies.

[0061] Again, a brief introduction to the implementation environment (implementation architecture) involved in the present application is given.

[0062] The embodiment of the present application provides a schematic diagram of a speed prediction system as Figure 1 shown. As Figure 1 shown, the speed prediction system includes a server 101, a speed determination device 102, and a first moving device 103 (i.e., the target object). Among them, the server 101 can establish connections with the speed determination device 102 and the first moving device 103 through a wired network or a wireless network. The speed determination device 102 and the first moving device 103 can communicate with each other through a wired network or a wireless network.

[0063] The speed determination device 102 is asFigure 2 As shown, the server of the present application stores a pre-trained displacement determination model. The speed determination device acquires two frames of images of the first moving device at different times during the movement and the current spatial coordinates of the first moving device, and sends the two frames of images and the current spatial coordinates to the server. The server inputs the two frames of images into the displacement determination model to obtain a displacement information feature map. The server determines the pixel coordinates of the first moving device based on the current spatial coordinates, and determines the current moving speed of the first moving device at the position of the displacement information feature map through the pixel coordinates of the first moving device. The server then sends the current moving speed to the speed determination device so that the speed determination device can display the current moving speed.

[0064] It should be noted that each of the two frames of images includes multiple objects, that is, multiple moving devices, and the target object is the first moving device. Therefore, the displacement information feature map includes the displacements of multiple moving devices.

[0065] Exemplarily, the speed determination device can be set on the second moving device. The speed determination device acquires two frames of images of the first moving device at different times (t1 and t2, where t1 is later than t2) during the movement, and determines the moving speed of the first moving device at time t2 based on the two different frames of images. The speed determination device can also be set at a certain fixed position to determine the moving speed of the first moving device. The position where the second moving device is set can be determined according to the specific application scenario. For example, when applied to the scenario of an autonomous vehicle, the speed determination device is set on the vehicle. Another example is that the speed determination device is set on a highway to predict the speed of vehicles running on the highway. Therefore, the present application does not specifically limit the usage scenario and installation method of the speed determination device.

[0066] In some embodiments, the server can be a single server, or alternatively, it can also be a server cluster composed of multiple servers. In some embodiments, the server cluster can also be a distributed cluster. The present disclosure also does not limit the specific implementation manner of the server.

[0067] As another implementation, the speed determination method of the present application can also be applied to a computer device. The specific form of the computer device in the embodiments of the present application is not limited in any way. For example, the computer device can specifically be a terminal device or a network device. Among them, the terminal device can be referred to as: terminal, user equipment (UE), terminal device, access terminal, user unit, user station, mobile station, remote station, remote terminal, mobile device, user terminal, wireless communication device, user agent or user device, etc. The terminal device can specifically be a mobile phone, augmented reality (AR) device, virtual reality (VR) device, tablet computer, laptop computer, ultra-mobile personal computer (UMPC), netbook, personal digital assistant (PDA), etc. The network device can specifically be a server, etc. Among them, the server can be a physical or logical server, or two or more physical or logical servers sharing different responsibilities and cooperating with each other to implement the various functions of the server.

[0068] In terms of hardware implementation, the above computer device can be implemented by a computer device as Figure 3 shown. As Figure 3 shown, it is a schematic diagram of the hardware structure of a computer device provided by an embodiment of the present application. The computer device can be used to implement the functions of the above computer device.

[0069] Figure 3 It is a schematic diagram of the composition of the image-based speed determination device 102 provided by an embodiment of the present application. As Figure 3 shown, the speed determination device 102 can include a processor 10, a memory 20, a communication line 30, a communication interface 40, and an input / output interface 50.

[0070] Among them, the processor 10, the memory 20, the communication interface 40, and the input / output interface 50 can be connected through the communication line 30.

[0071] A processor 10, configured to execute instructions stored in a memory 20 to implement the gesture recognition method provided in the following embodiments of the present application. The processor 10 may be a central processing unit (CPU), a general-purpose processor, a network processor (NP), a digital signal processor (DSP), a microprocessor, a microcontroller, a programmable logic device (PLD), or any combination thereof. The processor 10 may also be any other device with processing capabilities, such as a circuit, a device, or a software module, and the embodiments of the present application do not limit this. In one example, the processor 10 may include one or more CPUs, such as Figure 3 CPU0 and CPU1 in Figure 3 shown by way of a dashed line. As an alternative implementation, the electronic device may include multiple processors. For example, in addition to the processor 10, it may further include a processor 60 (

[0072] The memory 20 is configured to store instructions. For example, the instructions may be a computer program. Optionally, the memory 20 may be a read-only memory (ROM) or other type of static storage device that can store static information and / or instructions, or it may be a random access memory (RAM) or other type of dynamic storage device that can store information and / or instructions. It may also be an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices, and the embodiments of the present application do not limit this.

[0073] It should be noted that the memory 20 may exist independently of the processor 10 or may be integrated with the processor 10. The memory 20 may be located inside the speed determination device 102 or outside the speed determination device 102, and the embodiments of the present application do not limit this.

[0074] A communication line 30 is used to transmit information between various components included in an electronic device. The communication line 30 can be an Industry Standard Architecture (ISA) line, a Peripheral Component Interconnect (PCI) line, an Extended Industry Standard Architecture (EISA) line, etc. The communication line 30 can be divided into an address line, a data line, a control line, etc. For ease of representation, Figure 3 only one solid line is shown in [description], but it does not mean that there is only one line or one type of line.

[0075] A communication interface 40 is used to communicate with other devices (such as the above-mentioned image acquisition device 200) or other communication networks. The other communication network can be an Ethernet, a Radio Access Network (RAN), a Wireless Local Area Networks (WLAN), etc. The communication interface � can be a module, a circuit, a transceiver, or any device capable of implementing communication.

[0076] An input / output interface 50 is used to implement human-machine interaction between a user and an electronic device. For example, it can implement action interaction, text interaction, or voice interaction between a user and the speed determination device 102.

[0077] Exemplarily, the input / output interface 50 can be a central control screen in the first motion device 103, or a physical button in the operation area, or a physical button in the operation area, etc. Through the central control screen or physical button, etc., action interaction (such as gesture interaction) or text interaction between a vehicle occupant and the speed determination device 102 can be achieved.

[0078] Exemplarily, the input / output interface 50 can also be an audio module, which can include a speaker and a microphone, etc. Through this audio module, voice interaction between a user and the speed determination device 102 can be achieved.

[0079] It should be noted that Figure 3 the structure shown in [description] does not constitute a limitation on the speed determination device 102 (electronic device). Except Figure 3 for the components shown, the speed determination device 102 (electronic device) can include more or fewer components than those shown in the figure, or a combination of certain components, or different component arrangements.

[0080] The method for determining speed based on an image provided in an embodiment of the present application will be described below with reference to the accompanying drawings. For ease of understanding, "coordinates" are used to describe "positions" in the following text.

[0081] Figure 4 It is a schematic flowchart of the speed determination method provided in an embodiment of the present application. Optionally, this method can be executed by a speed determination device having the above Figure 3 shown hardware structure, as Figure 4 shown, this method includes S401 to S405.

[0082] S401, acquire a first image and a second image during the movement of a target object in an environmental space.

[0083] Among them, the acquisition time of the first image is later than that of the second image; for ease of understanding, the acquisition time of the first image is referred to as the first acquisition time in the following text; the acquisition time of the second image is referred to as the second acquisition time, then the first acquisition time of the first image is later than the second acquisition time of the second image.

[0084] S402, fuse the first feature extracted from the first image with the second feature extracted from the second image to obtain a fused feature map.

[0085] It should be understood that the above first feature and second feature include underlying features and multi-degree features at different environmental space angles.

[0086] Exemplarily, as Figure 5 shown, taking the first feature as feature (t - 1) and the second feature t as an example, the fusion process of the first feature and the second feature is described as follows. First, two 4-dimensional 1xCxHxW first original feature maps (t) and second original feature maps (t - 1) are respectively subjected to two downsamplings and two upsamplings (i.e., 2-fold downsampling, 2-fold downsampling, 2-fold upsampling, 2-fold upsampling), corresponding to obtaining 4-dimensional 1xCxHxW first feature maps (t) and second feature maps (t - 1). Then, multi-scale feature extraction is performed on the first feature map (t) and the second feature map (t - 1) to respectively obtain 4-dimensional 1xCxHxW multi-scale feature maps (t - 1) and multi-scale feature maps (t). Finally, the above multi-scale feature maps (t - 1) and multi-scale feature maps (t) are fused to obtain a 4-dimensional 1x2CxHxW fused feature map (t). Among them, in 4-dimensional 1xCxHxW, 1 represents the current map, C represents the number of channels, H represents the length, and W represents the width.

[0087] S403, perform a first convolution operation on the fused feature map to obtain a displacement information feature map.

[0088] Among them, the displacement information feature map is used to characterize the correspondence between pixel positions and displacements. That is, the displacement information feature map characterizes the correspondence between pixel coordinates and the displacements of the object during the time interval between the first acquisition time and the second acquisition time.

[0089] As an implementation manner, the above steps S402 and S403 can be implemented in the following way: obtaining a displacement information feature map according to the first image, the second image, and the trained displacement determination model.

[0090] It can be understood that the above displacement determination model is a model trained based on training samples.

[0091] In some embodiments, the displacement determination model includes a feature extraction module, a fusion module, and a displacement information feature map determination module; the feature extraction module is used to extract a first feature from the first image and a second feature from the second image; the fusion module is used to fuse the first feature and the second feature to output a fused feature map; the displacement information feature map determination module is used to output a displacement information feature map according to the fused feature map.

[0092] In this implementation manner, by inputting two frames of images of the target object at different acquisition times into the displacement determination model, the displacement of the target object can be directly determined, and thus the movement speed of the target object at the acquisition time of the first image can be determined. The above-mentioned input of the two frames of images into the displacement determination model replaces the execution steps of feature extraction, fusion, and the first convolution operation in one step, and can determine the displacement more quickly, thereby improving the efficiency of determining the movement speed of the target object.

[0093] S404, extracting a target displacement corresponding to a target pixel position from the displacement information feature map, where the target pixel position is the pixel position corresponding to the first pixel position information of the target object in the displacement information feature map.

[0094] As a possible implementation manner, this step S404 can be specifically implemented in the following way: obtaining the target spatial coordinates of the target object at the first acquisition time; the target spatial coordinates are the three-dimensional coordinates of the target object in the environmental space; determining the two-dimensional projection coordinates of the target spatial coordinates projected onto the first image; and converting the two-dimensional projection coordinates into target pixel coordinates according to the size ratio of the first image to the displacement information feature map.

[0095] It should be understood that at least one pixel coordinate on the first image includes the two-dimensional projection coordinates of the target spatial coordinates projected onto the first image, and there is a projection conversion relationship between the spatial coordinates and the first image, that is, when the spatial coordinates are known, the pixel coordinates projected onto the first image are known.

[0096] Based on this, this implementation provides a method for obtaining target pixel coordinates based on the target spatial coordinates of a target object according to the conversion relationship between two-dimensional pixels in the first image and the spatial coordinates of the target object in the environmental space. Compared with determining the target pixel coordinates by extracting the pixel coordinates of the target object from the first image, this implementation is more in line with the environmental space scenario based on the spatial coordinates of the target object, so that the obtained target pixel coordinates are more accurate.

[0097] In some examples, a 3D (3 dimensions, three-dimensional) target detection method is used to obtain the target spatial coordinates. Exemplarily, taking the target object as a vehicle, an autonomous vehicle uses a 3D target detection method to obtain the spatial coordinates of an obstacle to identify the exact position and orientation of the obstacle.

[0098] As an implementation, as shown in the following formula (1), the above three-dimensional target spatial coordinates (x, y, z) p can obtain the two-dimensional projection coordinates (u, v) of the target object in the first image through the following matrix conversion method p . The two-dimensional projection coordinates (u, v) p are scaled to obtain target pixel coordinates that match the displacement information feature map. Among them, after scaling, rounding up or down to the nearest integer is performed. For example, 4.3 is taken as 4, and 4.9 is taken as 5. Based on the target pixel coordinates (u, v) ps , the target displacement (dx, dy, dz) is obtained from the displacement information feature map p .

[0099]

[0100] Among them, represents the three-dimensional spatial coordinate vector of the target object in the x direction, y direction, and z direction at the first acquisition time, represents the coordinate vector formed by the pixel coordinates of the target object and the real number 1 at the first acquisition time, t represents time, p represents the target object, fx represents the horizontal equivalent focal length, fy represents the vertical equivalent focal length, cx represents the vertical camera internal parameter center, cy represents the vertical camera internal parameter center, u represents the pixel coordinate in the horizontal direction on the image, and v represents the pixel coordinate in the vertical direction on the image.

[0101] S405. Determine the motion speed of the target object at the acquisition time of the first image according to the target displacement.

[0102] As shown in the following formula (2), the above three-dimensional target spatial coordinates can obtain the motion speed of the target object through the following matrix conversion method.

[0103]

[0104] Among them, represents a velocity vector, a target displacement vector representing the target displacement, dx represents the displacement in the x direction, dy represents the displacement in the y direction, dz represents the displacement in the z direction, vx represents the velocity in the x direction, vy represents the velocity in the y direction, vz represents the velocity in the z direction. time t is the first acquisition time, time t-1 is the second acquisition time.

[0105] Based on the above implementation of determining velocity based on images, different features of the target object at different positions in the environmental space can be extracted from two frames of images respectively, that is, the first feature and the second feature. Then, by fusing the first feature and the second feature corresponding to the two frames of images into a fused feature, the first feature and the second feature are associated. Since the first feature and the second feature have different feature angles for representing the environmental space, fusing the first feature and the second feature into a single fused feature is more conducive to identifying and extracting the feature differences of the target object in the environmental space in the two frames of images. Therefore, the displacement information feature map obtained based on the fused feature can more accurately determine the displacement of the target object.

[0106] Based on this, by extracting the information of two frames of images of the target object at different acquisition times and performing corresponding convolution operations, the displacement of the target object can be directly determined, and thus the motion velocity of the target object at the acquisition time of the first image can be determined. This velocity determination method is not limited by the image characteristics (such as image brightness) of the input image and can be used for various input images of different frames. It has a wide range of application scenarios and a simple determination process, avoiding the problem of low accuracy of the determined motion velocity of the target object due to the limitation of the image characteristics of the input image in the related technology.

[0107] Based on the displacement determined by the above implementation, the following provides two methods for determining the displacement uncertainty of the target displacement to evaluate the accuracy of the target displacement. It can be understood that displacement and displacement uncertainty are in one-to-one correspondence.

[0108] As an implementation: perform a second convolution operation on the above determined fused feature map to obtain a displacement uncertainty feature map. Then, according to the target pixel position, extract the displacement uncertainty corresponding to the target displacement from the displacement uncertainty feature map.

[0109] Among them, the displacement uncertainty feature map is used to represent the correspondence between the pixel position and the displacement uncertainty, and the displacement uncertainty represents the accuracy of the displacement corresponding to any pixel position in the displacement information feature map.

[0110] In this implementation, the magnitude of the displacement uncertainty reflects the accuracy of the determined target displacement. That is, the larger the displacement uncertainty, the lower the accuracy of the obtained target displacement. Correspondingly, the smaller the displacement uncertainty, the higher the accuracy of the obtained target displacement. Therefore, in practical applications, multiple sets of two-frame images can be collected and input into the displacement determination model multiple times to obtain multiple target displacements and multiple displacement uncertainties. Then, from the multiple target displacements, the target displacement with the smallest corresponding displacement uncertainty is selected, and the target displacement with the smallest displacement uncertainty is used to determine the motion speed of the target object, so as to make the determination of the motion speed of the target object more accurate.

[0111] As another implementation, the displacement determination model further includes a displacement uncertainty feature map determination module; the displacement uncertainty feature map determination module is configured to output a displacement uncertainty feature map according to the fused feature map.

[0112] Specifically, according to the first image, the second image, and the displacement determination model, a displacement uncertainty feature map is obtained; the displacement determination model is further configured to obtain the displacement uncertainty feature map according to the fused feature. The displacement uncertainty feature map includes the displacement uncertainty corresponding to the target displacement, and the displacement uncertainty is used to indicate the degree of uncertainty of the target displacement.

[0113] In this implementation, by inputting two-frame images of the target object at different acquisition times into the displacement determination model, the displacement uncertainty of the target object can be directly determined. That is, by inputting the fused feature map obtained based on the two-frame images into the displacement uncertainty feature map determination module, the displacement uncertainty of the target object can be obtained. This method is simple and direct, facilitating practical applications.

[0114] Exemplarily, if multiple target displacements are obtained for the same target object using multiple-frame images. Taking three target displacements as an example, the displacement uncertainty corresponding to the first target displacement is 0.9; the displacement uncertainty corresponding to the second target displacement is 0.5, and the displacement uncertainty corresponding to the third target displacement is 0.7; then the second target displacement has a higher accuracy than the third target displacement; the third target displacement has a higher accuracy than the first target displacement.

[0115] The above-mentioned feature extraction module is used to extract the first feature and the second feature. The first feature and the second feature include underlying features and multi-degree features in different environmental space angles. The first feature layer extracts the underlying features of the images from the first image and the second image respectively, that is, the first underlying feature and the second underlying feature. The second feature layer extracts the multi-scale features of the images from the first underlying feature and the second underlying feature respectively, that is, the first feature and the second feature. The fusion module is used to fuse the first feature and the second feature to obtain a fused feature. The displacement information feature map determination module is used to perform a convolution operation on the fused feature to obtain a displacement information feature map; the target displacement output module is used to obtain the target displacement according to the target pixel coordinates and the displacement information feature map.

[0116] Exemplarily, as Figure 6 shown, the displacement determination model includes an underlying feature extraction module, a multi-scale feature extraction module, a front and rear frame feature fusion module, a displacement dx / dy / dz prediction module, and a displacement uncertainty σx / σy / σz prediction module. In the first step, for the images Img t at the time stamp time t and the images Img t-1 at the time stamp time t-1 respectively, underlying features are extracted to obtain underlying features. In the second step, 4-dimensional 1xCxHxW FPN multi-scale features are extracted from the underlying features respectively. The front and rear frame feature fusion module fuses the features extracted from the front and rear frames respectively, and fuses them in a superposition manner in the second dimension to obtain a fused feature map of 4 dimensions 1x2CxHxW. Through the first convolution operation, the high-order feature information of the fused feature map is extracted and fused to obtain a high-order feature map of 4 dimensions 1x256xHxW. Then, the high-order feature information of the front and rear frames is associated, and the second convolution operation is performed to extract the inter-frame features, obtaining a displacement information feature map and a displacement uncertainty feature map of 4 dimensions 1x3xHxW to predict the target displacement.

[0117] Based on the above steps S402 and S403, it can be completed by the displacement determination model. In order to ensure the accuracy of obtaining the target, this model is a trained model. The following will make a detailed description of the training process of this model in conjunction with the accompanying drawings.

[0118] Figure 7 It is a schematic flowchart of the training method of the displacement determination model provided by the embodiment of the present application. The method includes the following steps:

[0119] S701, obtain a training sample set.

[0120] Each training sample in the training sample set includes a first training image and a second training image during the movement of the sample object, as well as a label corresponding to the training sample; the first sampling time of the first training image is later than the second sampling time of the second training image; the label includes the first true pixel position of the sample object in the first training image and the true displacement generated by the sample object between the first sampling time and the second sampling time.

[0121] S702. Use the training sample set to train a preset initial model to obtain a displacement determination model.

[0122] The displacement determination model is used to obtain a displacement information feature map based on images at two different times, and the displacement information feature map represents the corresponding relationship between pixel coordinates and the displacement of the sample object at two different time intervals.

[0123] The initial model includes a feature extraction module, a fusion module, and a displacement information feature map determination module. The feature extraction module is used to extract a first feature from the first training image and a second feature from the second training image; the fusion module is used to fuse the first feature and the second feature to obtain a fused feature map, and the fused feature map includes the first pixel position information and the second pixel position information of the sample object. The first pixel position information represents the pixel position of the sample object in the first training image, and the second pixel position information represents the pixel position of the sample object in the second image; the displacement information feature map determination module is used to output a displacement information feature map according to the fused feature map, and the displacement information feature map represents the corresponding relationship between pixel positions and displacements; the displacement corresponding to the first true pixel position in the displacement information feature map is the predicted displacement output by the initial model.

[0124] In one implementation, the training loss needs to be considered during the training process in step S702, and the following steps will be looped multiple times: after one training is completed, calculate the training loss of this time, and adjust the parameters of the initial model based on the training loss of this time. After multiple trainings, a displacement determination model is finally obtained.

[0125] As a way of implementation, the training loss includes a first training loss, which is also called a spatial displacement loss. Step S702 can be specifically implemented in the following way: First, input the training sample into the initial model to obtain the predicted displacement output by the initial model. Second, determine the first training loss according to the predicted displacement and the true displacement. Third, adjust the parameters of the initial model according to the training loss to obtain a displacement determination model.

[0126] As another implementation, the training loss includes a second training loss, which is also referred to as the pixel displacement loss. Among them, the label also includes the true pixel displacement and the first true position of the sample object. The above step S602 can be specifically implemented in the following manner: First, according to the predicted displacement and the first true position, predict the true position of the sample object at the second sampling time to obtain the predicted second true position. Second, according to the predicted second true position, predict the true pixel position of the sample object in the second training image to obtain the predicted second true pixel position. Third, according to the predicted second true pixel position and the first true pixel position, obtain the predicted pixel displacement. Fourth, determine the second training loss according to the predicted pixel displacement and the true pixel displacement.

[0127] Exemplarily, during training, the loss value includes a spatial displacement loss and a pixel displacement loss, where the pixel displacement loss is the difference between the value output by the pixel displacement prediction module and the true pixel displacement.

[0128] In some embodiments, the displacement determination model is obtained after training an initial model using a training sample set and then deleting the pixel displacement prediction module in the initial model; the network structure of the initial model includes a feature extraction module, a fusion module, a displacement information feature map determination module, a target displacement output module, and a pixel displacement prediction module; the pixel displacement prediction module is used to obtain the predicted pixel displacement according to the target displacement output by the target displacement output module.

[0129] It should be understood that during training, multi-task learning is performed to share the feature extraction module and the fusion module, and the shared feature extraction module and fusion module can learn more underlying, more general, and more relevant features. The pixel displacement prediction module uses the output of the target displacement module as an input to determine the predicted pixel displacement. That is, the fusion feature obtained through the feature extraction module and the fusion module is independently transmitted to the target displacement module. That is, the target displacement module is not affected by the pixel displacement prediction module, and deleting the pixel displacement prediction module will not affect the target displacement module. However, during training, the predicted pixel displacement output by the pixel displacement prediction module is used to help the feature extraction module and the fusion module improve the learning emphasis on pixel displacement features during the regression process, improve the accuracy of the training process, and thus improve the accuracy of the target displacement determination model. And in this implementation, the pixel displacement prediction module is deleted in the displacement determination model, saving computational overhead on the basis of not affecting the accuracy of the displacement determination model and realizing the miniaturization of the displacement determination model.

[0130] In this embodiment, the displacement loss function includes displacement errors in two dimensions, namely, the error of three-dimensional space coordinate displacement and the pixel displacement error on the image. In view of the projection conversion relationship between the image pixel coordinates and the environmental space coordinates, the pixel coordinates of the image will affect the displacement of the three-dimensional space coordinates. Therefore, considering the pixel displacement deviation and the displacement deviation together during the training process will make the trained displacement determination model more accurate.

[0131] As another implementation, the training loss includes a third training loss, which is also called the displacement uncertainty loss.

[0132] Exemplarily, the displacement uncertainty loss value is the expected variance of the uncertainty output by the displacement uncertainty output module and the true displacement. Specifically, the network structure of the initial model further includes a displacement uncertainty output module; during the training process, the loss value further includes an uncertainty loss value, where the displacement uncertainty loss value is the difference between the uncertainty output by the displacement uncertainty output module and the expected value of the true displacement.

[0133] In this implementation, considering the influence of displacement uncertainty and using the displacement uncertainty deviation as the loss function will make the determination degree of the displacement output by the displacement determination model higher, thus ensuring the accuracy of the target displacement.

[0134] As a specific implementation, the training loss is called the loss value. The first training loss is the spatial displacement loss. The second training loss is the pixel displacement loss. The third training loss is the displacement uncertainty loss. During the training process, the loss value includes the spatial displacement loss, the pixel displacement loss, and the displacement uncertainty loss, where the pixel displacement loss is the difference between the predicted pixel displacement output by the pixel displacement prediction module and the true pixel displacement.

[0135] After performing a training task on each training sample once, based on the initial model, a training displacement information feature map corresponding to each training sample is obtained. According to the first pixel coordinates of the sample object in the first training image, the predicted displacement of the sample object is obtained from the training displacement information feature map; the difference between the predicted displacement and the true displacement corresponding to each training sample is determined as the displacement error; according to the predicted displacement, the predicted pixel coordinates of the sample object in the second training image are predicted; the difference between the predicted pixel coordinates and the second pixel coordinates of the sample object in the second training image is determined as the error of the pixel displacement.

[0136] Exemplarily, according to the following formula (3), based on the true coordinates (x, y, z) of the sample object gt , the true first pixel coordinates (u, v) on the first training image are obtained gt .

[0137]

[0138] Among them, represents the true three-dimensional spatial coordinate vector in the x, y, and z directions of the sample object, is the coordinate vector composed of the true pixel coordinates of the sample object on the first training image and the real number 1. t represents time, and gt represents the sample object of the first training image.

[0139] Furthermore, as shown in formula (4) below, according to the true first pixel coordinates (u, v) gt obtain the predicted displacement (dx, dy, dz) of the sample object from the training displacement information feature map gt . Add the predicted displacement (dx, dy, dz) gt and the true coordinates (x, y, z) gt to predict the predicted spatial coordinates of the sample object in the second training image, and thus obtain the predicted pixel coordinates (u, v) according to the predicted spatial coordinates gt′ .

[0140]

[0141] Among them, represents the coordinate vector composed of the predicted pixel coordinates of the sample object and the real number 1, represents the predicted displacement vector of the predicted displacement, and gt′ represents the sample object of the second training image.

[0142] Furthermore, obtain the predicted pixel displacement according to formula (5) below.

[0143]

[0144] Among them, is the two-dimensional vector of the predicted pixel displacement, is the two-dimensional vector of the true pixel coordinates of the sample object on the first training image, is the two-dimensional vector of the predicted pixel coordinates of the sample object on the second training image; du represents the pixel displacement in the horizontal direction on the first training image, and dv represents the pixel displacement in the vertical direction on the first training image.

[0145] Furthermore, obtain the true pixel displacement according to formula (6) below. The pixel displacement loss is the difference between the true pixel displacement and the predicted pixel displacement.

[0146]

[0147] Among them, is the two-dimensional vector of the true pixel displacement, is the two-dimensional vector of the true pixel coordinates of the sample object on the second training image, Δu is the true pixel displacement in the horizontal direction, and Δv is the true pixel displacement in the vertical direction.

[0148] Further, calculate the following displacement losses according to formula (7), formula (8), and formula (9) respectively: the spatial displacement loss in the x horizontal direction is |dx - Δx|, the spatial displacement loss in the y horizontal direction is |dy - Δy|, the spatial displacement loss in the z horizontal direction is |dz - Δz|, the horizontal pixel displacement loss is |du - Δu|, and the vertical displacement loss is |dv - Δv|.

[0149] L(dx) = |dx - Δx| + |du - Δu| (7)

[0150] L(dy) = |dy - Δy| + |dv - Δv| (8)

[0151] L(dz) = |dz - Δz| (9)

[0152] Among them, the displacement loss in the x direction is L(dx), the displacement loss in the y direction is L(dy), the horizontal displacement loss in the z direction is L(dz), Δx is the true displacement in the x direction, Δy is the true displacement in the y direction, and Δz is the true displacement in the z direction.

[0153] In this implementation manner, the displacement loss function includes displacement errors in two dimensions, that is, the error of the three-dimensional space coordinate displacement and the pixel displacement error on the image. Considering the projection conversion relationship between the image pixel coordinates and the environmental space coordinates, the pixel coordinates of the image will affect the displacement of the three-dimensional space coordinates. Therefore, considering the pixel displacement deviation and the displacement deviation together during the training process will make the trained displacement determination model more accurate.

[0154] As a possible implementation manner, the network structure of the initial model further includes a displacement uncertainty output module; during the training process, the loss value further includes an uncertainty loss value, where the displacement uncertainty loss value is the difference between the uncertainty output by the displacement uncertainty output module and the expected value of the true displacement.

[0155] Specifically, after each training sample executes a training task, a training uncertainty feature map corresponding to each training sample is obtained. According to the first pixel coordinates of the target sample object in the first training image, the uncertainty of the predicted displacement is obtained from the training uncertainty feature map; according to the true displacement and uncertainty corresponding to each training sample, the error of the uncertainty corresponding to the predicted displacement is determined.

[0156] Exemplarily, it is calculated as the following formula (10).

[0157]

[0158]

[0159]

[0160] Among them, L(σx) is the uncertainty loss in the x direction, L(σy) is the uncertainty loss in the y direction, L(σz) is the uncertainty loss in the z direction, σx is the uncertainty in the x direction, σy is the uncertainty in the y direction, and σz is the uncertainty in the z direction.

[0161] In this implementation manner, the influence of displacement uncertainty is considered, and the displacement uncertainty deviation is used as the loss function, which can make the certainty of the displacement output by the displacement determination model higher, thus ensuring the accuracy of the target displacement.

[0162] Therefore, the total loss obtained is the sum of each loss, as shown in formula (13).

[0163] L = L(dx) + L(dy) + L(dz) + L(σx) + L(σy) + L(σz) (13)

[0164] It should be noted that the above-mentioned various modules introduce the initial model of multi-task learning provided by the embodiments of the present application. The multiple tasks may also include tasks such as feature segmentation that can enable the feature extraction network module to extract underlying and general features. The embodiments of the present application do not limit this.

[0165] The above mainly introduces the solution provided by the embodiments of the present application from the perspective of the method. To implement the above functions, it includes the corresponding hardware structure and / or software module for executing each function. Those skilled in the art should easily realize that, in combination with the units and algorithm steps of each example described in the embodiments disclosed in this article, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the way of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0166] The present application also provides an image-based speed determination device, as Figure 8 shown. The speed determination device includes: a first acquisition unit 81, a fusion unit 82, a convolution unit 83, a displacement determination unit 84, and a speed determination unit 85.

[0167] The first acquisition unit 81 is configured to acquire a first image and a second image during the movement of the target object in the environmental space, where the acquisition time of the first image is later than that of the second image; the fusion unit 82 is configured to fuse the first feature extracted from the first image and the second feature extracted from the second image to obtain a fused feature map, where the fused feature map includes the first pixel position information and the second pixel position information of the target object, the first pixel position information represents the pixel position of the target object in the first image, and the second pixel position information represents the pixel position of the target object in the second image; the convolution unit 83 is configured to perform a first convolution operation on the fused feature map to obtain a displacement information feature map, where the displacement information feature map is used to represent the correspondence between the pixel position and the displacement; the displacement determination unit 84 is configured to extract a target displacement corresponding to the target pixel position from the displacement information feature map, where the target pixel position is the pixel position corresponding to the first pixel position information of the target object in the displacement information feature map; the speed determination unit 85 is configured to determine the movement speed of the target object at the acquisition time of the first image according to the target displacement.

[0168] In a possible implementation manner, the displacement determination unit 84 is further configured to specifically acquire the target spatial position of the target object at the acquisition time of the first image, where the target spatial position is the three-dimensional position of the target object in the environmental space; determine the two-dimensional projection position of the target spatial position projected on the first image; and convert the two-dimensional projection position into the target pixel position according to the ratio of the size of the first image to the size of the displacement information feature map.

[0169] In another possible implementation manner, the speed determination device further includes a displacement uncertainty determination unit, which is configured to specifically perform a second convolution operation on the fused feature map to obtain a displacement uncertainty feature map; the displacement uncertainty feature map is used to represent the correspondence between the pixel position and the displacement uncertainty, where the displacement uncertainty represents the accuracy of the displacement corresponding to any pixel position in the displacement information feature map; and extract the displacement uncertainty corresponding to the target displacement from the displacement uncertainty feature map according to the target pixel position.

[0170] In another possible implementation manner, the speed determination device is further configured to: input the first image and the second image into a displacement determination model to obtain a displacement information feature map output by the displacement determination model.

[0171] In another possible implementation manner, the displacement determination model includes a feature extraction module, a fusion module, and a displacement information feature map determination module; the feature extraction module is used to extract the first feature from the first image and the second feature from the second image; the fusion module is used to fuse the first feature and the second feature to output a fused feature map; and the displacement information feature map determination module is used to output a displacement information feature map according to the fused feature map.

[0172] In another possible implementation, the displacement determination model further includes a displacement uncertainty feature map determination module; the displacement uncertainty feature map determination module is configured to output a displacement uncertainty feature map according to the fused feature map.

[0173] The present application also provides a training device for a displacement determination model, as Figure 9 shown, the training device includes: a second acquisition unit 91 and a feature map training unit 92.

[0174] The second acquisition unit 91 is configured to acquire a training sample set, each training sample in the training sample set includes a first training image and a second training image during the movement of the sample object, and a label corresponding to the training sample; the first sampling time of the first training image is later than the second sampling time of the second training image; the label includes the first true pixel position of the sample object in the first training image and the true displacement generated by the sample object between the first sampling time and the second sampling time; the feature map training unit 92 is configured to use the training sample set to train a preset initial model to obtain a displacement determination model; wherein, the initial model includes a feature extraction module, a fusion module and a displacement information feature map determination module; the feature extraction module is configured to extract a first feature from the first training image and a second feature from the second training image; the fusion module is configured to fuse the first feature and the second feature to obtain a fused feature map, the fused feature map includes the first pixel position information and the second pixel position information of the sample object, the first pixel position information represents the pixel position of the sample object in the first training image, and the second pixel position information represents the pixel position of the sample object in the second image; the displacement information feature map determination module is configured to output a displacement information feature map according to the fused feature map, the displacement information feature map represents the correspondence between the pixel position and the displacement; the displacement corresponding to the first true pixel position in the displacement information feature map is the predicted displacement output by the initial model.

[0175] In a possible implementation, the feature map training unit 92 is configured to specifically execute inputting the training sample into the initial model to obtain the predicted displacement output by the initial model; determining a first training loss according to the predicted displacement and the true displacement; adjusting the parameters of the initial model according to the training loss to obtain a displacement determination model; the training loss includes the first training loss.

[0176] In another possible implementation, the feature map training unit 92 is further configured to specifically perform predicting the true position of the sample object at the second sampling time according to the predicted displacement and the first true position to obtain the predicted second true position; predicting the true pixel position of the sample object in the second training image according to the predicted second true position to obtain the predicted second true pixel position; obtaining the predicted pixel displacement according to the predicted second true pixel position and the first true pixel position; determining the second training loss according to the predicted pixel displacement and the true pixel displacement; and the training loss further includes the second training loss.

[0177] In another possible implementation, the displacement uncertainty feature map determination module is configured to output a displacement uncertainty feature map according to the fused feature map, and the displacement uncertainty feature map is used to represent the correspondence between the pixel position and the displacement uncertainty, and the displacement uncertainty represents the accuracy of the displacement corresponding to any pixel position in the displacement information feature map.

[0178] In another possible implementation, the feature map training unit 92 is further configured to specifically perform determining the third training loss according to the displacement uncertainty and the true displacement; and the training loss further includes the third training loss.

[0179] It should be noted that the above division of the units is illustrative, and is only a logical function division. In actual implementation, there may be other division methods. For example, two or more functions may also be integrated into one processing unit. The above integrated units may be implemented in the form of hardware or in the form of software functional units.

[0180] In an exemplary embodiment, the embodiment of the present application further provides a readable storage medium, including software instructions, which when running on an electronic device (computing processing device), cause the electronic device (computing processing device) to execute any one of the methods provided in the above embodiments.

[0181] In an exemplary embodiment, the embodiment of the present application further provides a computer program product including computer execution instructions, which when running on an electronic device (computing processing device), cause the electronic device (computing processing device) to execute any one of the methods provided in the above embodiments.

[0182] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using a software program, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer-executable instructions. When the computer-executable instructions are loaded and executed on a computer, the processes or functions according to the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer-executable instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer-executable instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that includes one or more integrated media. The available medium can be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)), etc.

[0183] Although the present application has been described in conjunction with various embodiments, however, in the process of implementing the claimed present application, those skilled in the art can understand and implement other variations of the disclosed embodiments by viewing the drawings, the disclosure, and the appended claims. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "one" does not exclude a plurality. A single processor or other unit can implement several functions recited in the claims. Certain measures are recited in mutually different dependent claims, but this does not mean that these measures cannot be combined to produce good results.

[0184] Although the present application has been described in conjunction with specific features and their embodiments, it is obvious that various modifications and combinations can be made without departing from the spirit and scope of the present application. Accordingly, the present specification and the drawings are only exemplary descriptions of the present application defined by the appended claims, and are considered to have covered any and all modifications, variations, combinations, or equivalents within the scope of the present application. Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application also intends to include these changes and modifications.

[0185] As described above, it is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions within the technical scope disclosed in the present application should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. An image-based speed determination method, characterized in that, The method includes: Obtaining a first image and a second image during the movement of a target object in an environmental space, wherein the acquisition time of the first image is later than that of the second image; Fusing a first feature extracted from the first image and a second feature extracted from the second image to obtain a fused feature map, where the fused feature map includes first pixel position information and second pixel position information of the target object, the first pixel position information represents the pixel position of the target object in the first image, and the second pixel position information represents the pixel position of the target object in the second image; Performing a first convolution operation on the fused feature map to obtain a displacement information feature map, where the displacement information feature map is used to represent the correspondence between pixel positions and displacements; Extracting a target displacement corresponding to a target pixel position from the displacement information feature map, where the target pixel position is the pixel position in the displacement information feature map corresponding to the first pixel position information of the target object; Determining the movement speed of the target object at the acquisition time of the first image according to the target displacement.

2. The speed determination method according to claim 1, characterized in that The method further includes: Obtaining a target spatial position of the target object at the acquisition time of the first image, where the target spatial position is the three-dimensional position of the target object in the environmental space; Determining the two-dimensional projection position of the target spatial position projected onto the first image; Converting the two-dimensional projection position into the target pixel position according to the ratio of the size of the first image to the size of the displacement information feature map.

3. The method for determining speed based on an image according to claim 2, wherein The method further includes: Performing a second convolution operation on the fused feature map to obtain a displacement uncertainty feature map; the displacement uncertainty feature map is used to represent the correspondence between pixel positions and displacement uncertainties, and the displacement uncertainty represents the accuracy of the displacement corresponding to any pixel position in the displacement information feature map; Extracting the displacement uncertainty corresponding to the target displacement from the displacement uncertainty feature map according to the target pixel position.

4. The method for determining speed based on an image according to claim 3, wherein The fusing the first feature extracted from the first image and the second feature extracted from the second image to obtain a fused feature map, and performing a first convolution operation on the fused feature map to obtain a displacement information feature map includes: Inputting the first image and the second image into a displacement determination model to obtain the displacement information feature map output by the displacement determination model.

5. The method for image-based speed determination according to claim 4, characterized in that, The displacement determination model includes a feature extraction module, a fusion module, and a displacement information feature map determination module; The feature extraction module is used to extract the first feature from the first image and the second feature from the second image; The fusion module is used to fuse the first feature and the second feature to output a fused feature map; The displacement information feature map determination module is used to output the displacement information feature map according to the fused feature map.

6. The method for determining speed based on an image according to claim 5, wherein The displacement determination model further includes a displacement uncertainty feature map determination module; The displacement uncertainty feature map determination module is used to output the displacement uncertainty feature map according to the fused feature map.

7. A training method for a displacement determination model, characterized in that, The training method includes: Obtain a training sample set, where each training sample in the training sample set includes a first training image and a second training image during the movement of the sample object, and a label corresponding to the training sample; the first sampling time of the first training image is later than the second sampling time of the second training image; the label includes the first true pixel position of the sample object in the first training image and the true displacement generated by the sample object between the first sampling time and the second sampling time; Use the training sample set to train a preset initial model to obtain a displacement determination model; wherein, the initial model includes a feature extraction module, a fusion module, and a displacement information feature map determination module; The feature extraction module is used to extract a first feature from the first training image and a second feature from the second training image; The fusion module is used to fuse the first feature and the second feature to obtain a fused feature map, the fused feature map includes the first pixel position information and the second pixel position information of the sample object, the first pixel position information represents the pixel position of the sample object in the first training image, and the second pixel position information represents the pixel position of the sample object in the second training image; The displacement information feature map determination module is used to output a displacement information feature map according to the fused feature map, and the displacement information feature map represents the corresponding relationship between the pixel position and the displacement; The displacement corresponding to the first true pixel position in the displacement information feature map is the predicted displacement output by the initial model.

8. The training method according to claim 7, characterized in that, The using the training sample set to train a preset initial model to obtain a displacement determination model includes: Input the training sample into the initial model to obtain the predicted displacement output by the initial model; Determine a first training loss according to the predicted displacement and the true displacement; Adjust the parameters of the initial model according to the training loss to obtain the displacement determination model; the training loss includes the first training loss.

9. The training method according to claim 8, wherein The label further includes the true pixel displacement and the first true position of the sample object, and the training method further includes: Predict the true position of the sample object at the second sampling time according to the predicted displacement and the first true position to obtain a predicted second true position; Predict the true pixel position of the sample object in the second training image according to the predicted second true position to obtain a predicted second true pixel position; Obtain a predicted pixel displacement according to the predicted second true pixel position and the first true pixel position; Determine a second training loss according to the predicted pixel displacement and the true pixel displacement; The training loss further includes the second training loss.

10. The training method according to any one of claims 8 to 9, characterized in that, The initial model further includes a displacement uncertainty feature map determination module; The displacement uncertainty feature map determination module is configured to output the displacement uncertainty feature map according to the fused feature map. The displacement uncertainty feature map is used to characterize the correspondence between pixel positions and displacement uncertainties, and the displacement uncertainty represents the accuracy of the displacement corresponding to any pixel position in the displacement information feature map.

11. The training method according to claim 10, wherein Using the training sample set to train a preset initial model to obtain a displacement determination model further includes: Determining a third training loss according to the displacement uncertainty and the true displacement; The training loss further includes the third training loss.

12. An image-based speed determination device, characterized in that, The speed determination device includes: A first acquisition unit configured to acquire a first image and a second image during the movement of a target object in an environmental space, where the acquisition time of the first image is later than that of the second image; A fusion unit configured to fuse a first feature extracted from the first image and a second feature extracted from the second image to obtain a fused feature map. The fused feature map includes first pixel position information and second pixel position information of the target object. The first pixel position information characterizes the pixel position of the target object in the first image, and the second pixel position information characterizes the pixel position of the target object in the second image; A convolution unit configured to perform a first convolution operation on the fused feature map to obtain a displacement information feature map. The displacement information feature map is used to characterize the correspondence between pixel positions and displacements; A displacement determination unit configured to extract a target displacement corresponding to a target pixel position from the displacement information feature map. The target pixel position is the pixel position in the displacement information feature map corresponding to the first pixel position information of the target object; A speed determination unit configured to determine the movement speed of the target object at the acquisition time of the first image according to the target displacement.

13. The speed determination device according to claim 12, characterized in that, The speed determination device further includes a displacement uncertainty determination unit; The displacement determination unit is specifically configured to acquire the target spatial position of the target object at the acquisition time of the first image. The target spatial position is the three-dimensional position of the target object in the environmental space; Determining the two-dimensional projection position of the target spatial position projected on the first image; converting the two-dimensional projection position into the target pixel position according to the ratio of the size of the first image to the size of the displacement information feature map; The displacement uncertainty determination unit is specifically configured to perform a second convolution operation on the fused feature map to obtain a displacement uncertainty feature map. The displacement uncertainty feature map is used to characterize the correspondence between pixel positions and displacement uncertainties. The displacement uncertainty represents the accuracy of the displacement corresponding to any pixel position in the displacement information feature map; extracting the displacement uncertainty corresponding to the target displacement from the displacement uncertainty feature map according to the target pixel position; The speed determination device is further configured to perform: inputting the first image and the second image into a displacement determination model to obtain a displacement information feature map output by the displacement determination model; Wherein, the displacement determination model includes a feature extraction module, a fusion module, and a displacement information feature map determination module; the feature extraction module is configured to extract the first feature from the first image and the second feature from the second image; the fusion module is configured to fuse the first feature and the second feature to output a fused feature map; the displacement information feature map determination module is configured to output the displacement information feature map according to the fused feature map; the displacement determination model further includes a displacement uncertainty feature map determination module; the displacement uncertainty feature map determination module is configured to output the displacement uncertainty feature map according to the fused feature map.

14. A training device for a displacement determination model, characterized in that, The training device includes: A second acquisition unit configured to acquire a training sample set, each training sample in the training sample set including a first training image and a second training image during the movement of a sample object, and a label corresponding to the training sample; the first sampling time of the first training image is later than the second sampling time of the second training image; the label includes a first true pixel position of the sample object in the first training image and a true displacement generated by the sample object between the first sampling time and the second sampling time; A feature map training unit configured to perform training on a preset initial model using the training sample set to obtain a displacement determination model; wherein, the initial model includes a feature extraction module, a fusion module, and a displacement information feature map determination module; The feature extraction module is configured to extract a first feature from the first training image and a second feature from the second training image; The fusion module is configured to fuse the first feature and the second feature to obtain a fused feature map, the fused feature map including first pixel position information and second pixel position information of the sample object, the first pixel position information representing the pixel position of the sample object in the first training image, and the second pixel position information representing the pixel position of the sample object in the second training image; The displacement information feature map determination module is configured to output a displacement information feature map according to the fused feature map, the displacement information feature map representing the correspondence between pixel positions and displacements; The displacement corresponding to the first true pixel position in the displacement information feature map is the predicted displacement output by the initial model.

15. The training device according to claim 14, characterized in that, The label further includes the true pixel displacement and the first true position of the sample object: the initial model further includes a displacement uncertainty feature map determination module; the training device is further configured to; The feature map training unit is configured to specifically execute: inputting the training sample into the initial model to obtain the predicted displacement output by the initial model; determining a first training loss according to the predicted displacement and the true displacement; adjusting the parameters of the initial model according to the training loss to obtain the displacement determination model; the training loss includes the first training loss; The feature map training unit is further configured to specifically execute: predicting the true position of the sample object at the second sampling time according to the predicted displacement and the first true position to obtain a predicted second true position; Predicting the true pixel position of the sample object in the second training image according to the predicted second true position to obtain a predicted second true pixel position; Obtaining a predicted pixel displacement according to the predicted second true pixel position and the first true pixel position; determining a second training loss according to the predicted pixel displacement and the true pixel displacement; the training loss further includes the second training loss; The displacement uncertainty feature map determination module is configured to output the displacement uncertainty feature map according to the fused feature map, where the displacement uncertainty feature map is used to represent the correspondence between the pixel position and the displacement uncertainty, and the displacement uncertainty represents the accuracy of the displacement corresponding to any pixel position in the displacement information feature map; The feature map training unit is further configured to specifically execute: determining a third training loss according to the displacement uncertainty and the true displacement; the training loss further includes the third training loss.

16. An image-based speed determination device, characterized in that, Comprising: A memory and a processor; the memory and the processor are coupled; the memory is used to store computer program code, and the computer program code includes computer instructions; Wherein, when the processor executes the computer instructions, the image-based speed determination device is caused to execute the image-based speed determination method according to any one of claims 1 to 6.

17. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes computer instructions; Wherein, when the computer instructions run on the image-based speed determination device, the image-based speed determination device is caused to execute the image-based speed determination method according to any one of claims 1 to 6.

18. An image-based speed determination device, characterized in that, Comprising: A memory and a processor; the memory and the processor are coupled; the memory is used to store computer program code, and the computer program code includes computer instructions; Wherein, when the processor executes the computer instructions, the training device of the displacement determination model is caused to execute the training method of the displacement determination model according to any one of claims 7 to 11.

19. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes computer instructions; Wherein, when the computer instructions run on the training device of the displacement determination model, the training device of the displacement determination model is caused to execute the training method of the displacement determination model according to any one of claims 7 to 11.

Citation Information

Patent Citations

  • Remote sensing image super-resolution reconstruction method based on multi-dictionary learning and non-local information fusion

    CN105825477A

  • Integrated navigation algorithm based on fusion of optical flow position and velocity information

    CN109916394A

Cited By

  • Image-based speed determination method and apparatus, and device and storage medium

    EP4618016A1

  • Image-based speed determination method and apparatus, and device and storage medium

    WO2024099068A1