Data processing method, device, storage medium, and processor
By converting from a two-dimensional detection box to a three-dimensional detection box and determining the projection information of the target object, the problem of low two-dimensional detection accuracy is solved, and the precise positioning of the object and physical spatial mapping are achieved.
Patent Information
- Application Number
- CN202010246855.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-03-31
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2040-03-31
AI Technical Summary
When detecting objects, the prior art, based on two-dimensional detection and tracking will lead to low accuracy in determining the position, positioning and speed calculation of the object, especially because the plane where the object is located is not on the same plane as the picture.
By acquiring an image containing the target object, acquiring its two-dimensional detection box, and obtaining at least one surface of the three-dimensional detection box of the target object based on the two-dimensional detection box, the projection information of the target object in the target plane is determined.
The precise positioning of objects is achieved, the accuracy problem of two-dimensional detection is avoided, and the mapping accuracy of objects in physical space is improved.
Smart Images

Figure CN113470067B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computers, and in particular, to a data processing method, apparatus, storage medium, and processor. Background Art
[0002] Currently, when detecting an object, mainly two-dimensional detection (2D) and tracking of the object are performed to determine the behavior of the object. However, since the plane where the object is located and the picture are usually not in the same plane, the two-dimensional detection and tracking will fail in the following aspects:
[0003] (1) In terms of determining the location of the object in the service area: The two-dimensional detection box of the object detected in the image is very likely to span multiple parking spaces, so that the parking area of the object cannot be accurately located;
[0004] (2) In terms of object positioning: The two-dimensional detection box of the object detected in the image is very likely to span multiple regions, so that misjudgment is easily caused for regional-level positioning;
[0005] (3) In terms of speed calculation, the two-dimensional detection box of the object based on the picture cannot use the known specification information of the object to perform physical modeling of the true length corresponding to the pixels on the road surface, so that the true physical space distance difference cannot be determined to accurately calculate parameters such as the speed of the object.
[0006] Since the two-dimensional detection and tracking will fail in the above aspects, it leads to the technical problem of low accuracy in positioning the object.
[0007] For the above problems, no effective solution has been proposed yet. Summary of the Invention
[0008] Embodiments of the present invention provide a data processing method, apparatus, storage medium, and processor to at least solve the technical problem of low accuracy in positioning an object.
[0009] According to one aspect of the embodiments of the present invention, a data processing method is provided. The method may include: obtaining an image including a target object; obtaining a two-dimensional detection box of the target object; obtaining at least one surface of a three-dimensional detection box of the target object based on the two-dimensional detection box; and determining projection information of the target object on a target plane through at least one surface of the three-dimensional detection box of the target object.
[0010] According to another aspect of the embodiments of the present invention, another data processing method is further provided. The method may include: displaying an image including a target object on a target interface, and displaying a two-dimensional detection frame of the target object; displaying at least one surface of a three-dimensional detection frame of the target object on the target interface, where at least one surface of the three-dimensional detection frame of the target object is obtained based on the two-dimensional detection frame; displaying projection information of the target object on a target plane on the target interface, where the projection information is determined by at least one surface of the three-dimensional detection frame of the target object.
[0011] According to another aspect of the embodiments of the present invention, another data processing method is further provided. The method may include: obtaining an image including a target object; determining that there is a preset position in the image; in the case where no transformation matrix is associated with the preset position, performing three-dimensional detection on the image to obtain at least one surface of a three-dimensional detection frame, where the transformation matrix is used to transform a two-dimensional detection frame of an object corresponding to any image with a preset position into at least one surface of a three-dimensional detection frame of the object; determining projection information of the target object on a target plane through at least one surface of the three-dimensional detection frame of the target object.
[0012] According to another aspect of the embodiments of the present invention, another data processing method is further provided. The method may include: obtaining a target request, where the target request carries an image to be processed input on a target interface, and the image includes a target object; in response to the target request, obtaining a two-dimensional detection frame of the target object; based on the two-dimensional detection frame, obtaining at least one surface of a three-dimensional detection frame of the target object; determining projection information of the target object on a target plane through at least one surface of the three-dimensional detection frame of the target object, and sending the projection information to the target interface for display.
[0013] According to another aspect of the embodiments of the present invention, a data processing apparatus is further provided. The apparatus may include: a first obtaining unit, configured to obtain an image including a target object; a second obtaining unit, configured to obtain a two-dimensional detection frame of the target object; a third obtaining unit, configured to obtain at least one surface of a three-dimensional detection frame of the target object based on the two-dimensional detection frame; a first determining unit, configured to determine projection information of the target object on a target plane through at least one surface of the three-dimensional detection frame of the target object.
[0014] According to another aspect of the embodiments of the present invention, another data processing device is also provided. The device may include: a first display unit, configured to display an image including a target object on a target interface and display a two-dimensional detection frame of the target object; a second display unit, configured to display at least one surface of a three-dimensional detection frame of the target object on the target interface, where at least one surface of the three-dimensional detection frame of the target object is obtained based on the two-dimensional detection frame; a third display unit, configured to display projection information of the target object on a target plane on the target interface, where the projection information is determined by at least one surface of the three-dimensional detection frame of the target object.
[0015] According to another aspect of the embodiments of the present invention, another data processing device is also provided. The device may include: a fourth acquisition unit, configured to acquire an image including a target object; a second determination unit, configured to determine that there is a preset position in the image; a detection unit, configured to perform three-dimensional detection on the image to obtain at least one surface of a three-dimensional detection frame of the target object when the preset position is not associated with a conversion matrix, where the conversion matrix is used to convert a two-dimensional detection frame of an object corresponding to any image with a preset position into at least one surface of a three-dimensional detection frame of the object; a third determination unit, configured to determine projection information of the target object on a target plane through at least one surface of the three-dimensional detection frame of the target object.
[0016] According to another aspect of the embodiments of the present invention, a storage medium is also provided. The storage medium includes a stored program, where when the program runs, it controls the device where the storage medium is located to perform the following steps: acquire an image including a target object; acquire a two-dimensional detection frame of the target object; based on the two-dimensional detection frame, acquire at least one surface of a three-dimensional detection frame of the target object; determine projection information of the target object on a target plane through at least one surface of the three-dimensional detection frame of the target object.
[0017] According to another aspect of the embodiments of the present invention, a processor is also provided. The processor is used to run a program, where when the program runs, it performs the following steps: acquire an image including a target object; acquire a two-dimensional detection frame of the target object; based on the two-dimensional detection frame, acquire at least one surface of a three-dimensional detection frame of the target object; determine projection information of the target object on a target plane through at least one surface of the three-dimensional detection frame of the target object.
[0018] According to another aspect of the embodiments of the present invention, a mobile terminal is also provided. The mobile terminal may include: a processor; a transmission device, configured to transmit an image including a target object; and a memory, connected to the transmission device, configured to provide instructions for the processor to perform the following processing steps: acquire a two-dimensional detection frame of the target object; based on the two-dimensional detection frame, acquire at least one surface of a three-dimensional detection frame of the target object; determine projection information of the target object on a target plane through at least one surface of the three-dimensional detection frame of the target object.
[0019] In an embodiment of the present invention, an image including a target object is acquired; a two-dimensional detection frame of the target object is acquired; at least one surface of a three-dimensional detection frame of the target object is acquired based on the two-dimensional detection frame; and projection information of the target object on a target plane is determined through at least one surface of the three-dimensional detection frame of the target object. This application is a three-dimensional object detection technology based on a two-dimensional detection frame. The three-dimensional detection frame is obtained by using the two-dimensional detection frame of the object, and then the projection information of the target object on the plane is obtained through the three-dimensional detection frame, achieving the purpose of physically mapping the object. It avoids directly performing two-dimensional detection on the object to judge the object's behavior. Since the plane and the picture are not in the same plane, the detection of the object fails in some cases, thereby solving the technical problem of low accuracy in positioning the object and achieving the technical effect of accurately positioning the object. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The drawings described herein are used to provide a further understanding of the present invention, and constitute a part of this application. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings:
[0021] Figure 1 is a hardware structure block diagram of a computer terminal (or mobile device) for implementing a data processing method according to an embodiment of the present invention;
[0022] Figure 2 is a flowchart of a data processing method according to an embodiment of the present invention;
[0023] Figure 3 is a flowchart of another data processing method according to an embodiment of the present invention;
[0024] Figure 4 is a flowchart of another data processing method according to an embodiment of the present invention;
[0025] Figure 5 is a flowchart of another method for detecting an object according to an embodiment of the present invention;
[0026] Figure 6 is a schematic diagram of projection plane calculation according to an embodiment of the present invention;
[0027] Figure 7 is a schematic diagram of an architecture diagram of a traffic video object physical mapping solution according to an embodiment of the present invention;
[0028] Figure 8 is a schematic diagram of an interaction interface of a data processing method according to an embodiment of the present invention;
[0029] Figure 9Schematic diagram of a data processing device according to an embodiment of the present invention;
[0030] Figure 10 Schematic diagram of another data processing device according to an embodiment of the present invention;
[0031] Figure 11 Schematic diagram of another data processing device according to an embodiment of the present invention; and
[0032] Figure 12 Block diagram of a mobile terminal according to an embodiment of the present invention. Detailed implementation manners
[0033] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0034] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily need to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0035] First, some nouns or terms that appear in the process of describing the embodiments of the present application are applicable to the following explanations:
[0036] Three-dimensional object target detection, an algorithm for three-dimensional target object detection and three-dimensional position estimation by combining deep neural network regression learning and geometric constraints;
[0037] Preset position, a way of associating the key area to be monitored with the operating status of the PTZ camera;
[0038] Linear regression, a linear model for prediction by linear combination, the purpose of which is to find a straight line, a plane or a hyperplane of a higher dimension such that the error between the predicted value and the true value is minimized;
[0039] Pseudo inverse matrix, a generalized form of the inverse matrix.
[0040] Example 1
[0041] According to an embodiment of the present invention, an embodiment of a data processing method is further provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0042] The method embodiment provided in the first embodiment of this application can be executed on a mobile terminal, a computer terminal, or a similar computing device. Figure 1 It is a hardware structure block diagram of a computer terminal (or mobile device) for implementing a data processing method according to an embodiment of the present invention. As Figure 1 shown, the computer terminal 10 (or mobile device 10) may include one or more (shown as 102a, 102b,..., 102n in the figure) processors 102 (the processor 102 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission module 106 for communication functions. In addition, it may further include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the I / O interface), a network interface, a power supply, and / or a camera. Those of ordinary skill in the art can understand that Figure 1 the structure shown is only schematic and does not limit the structure of the above-mentioned electronic device. For example, the computer terminal 10 may further include more or fewer components than Figure 1 shown in, or have a different configuration from Figure 1 shown.
[0043] It should be noted that the above one or more processors 102 and / or other data processing circuits are generally referred to as "data processing circuits" in this article. The data processing circuit may be embodied in whole or in part as software, hardware, firmware, or any other combination. In addition, the data processing circuit may be a single independent processing module, or be incorporated in whole or in part into any one of the other elements in the computer terminal 10 (or mobile device). As involved in the embodiments of this application, the data processing circuit is a kind of processor control (such as the selection of a variable resistance terminal path connected to an interface).
[0044] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the data processing method in the embodiments of the present invention. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, implements the data processing method of the above application program. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely disposed relative to the processor 102, and these remote memories can be connected to the computer terminal 10 through a network. Examples of the above network include but are not limited to the Internet, intranet, local area network, mobile communication network, and combinations thereof.
[0045] The transmission device 106 is used to receive or send data via a network. Specific examples of the above network may include a wireless network provided by a communication provider of the computer terminal 10. In one instance, the transmission device 106 includes a network adapter (Network Interface Controller, abbreviated as NIC), which can be connected to other network devices through a base station and thus can communicate with the Internet. In one instance, the transmission device 106 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0046] The display can be, for example, a touch-screen liquid crystal display (LCD), and the liquid crystal display enables a user to interact with the user interface of the computer terminal 10 (or mobile device).
[0047] It should be noted here that in some alternative embodiments, the above Figure 1 shown computer device (or mobile device) may include hardware elements (including circuits), software elements (including computer code stored on a computer-readable medium), or a combination of both hardware elements and software elements. It should be pointed out that Figure 1 is only an example of a specific specific instance and is intended to show the types of components that may exist in the above computer device (or mobile device).
[0048] In Figure 1 the shown operating environment, the present application provides a data processing method as Figure 2 shown. It should be noted that the data processing method of this embodiment can be executed by Figure 1 the mobile terminal shown in the embodiment.
[0049] Figure 2 is a flowchart of a data processing method according to an embodiment of the present invention. As Figure 2As shown, the method includes the following steps:
[0050] Step S202, obtain an image including a target object.
[0051] In the technical solution provided in step S202 of the present invention, the target object is an object to be detected on a target plane. For example, if the target plane is a traffic road, the target object can be a target vehicle to be detected on the traffic road. An image obtained by real-time monitoring of the target plane can be acquired. This image can be a monocular color (RGB) image and can form a video stream of traffic video. Then, an image of the target object is obtained from the video stream. This image is also a single frame in the video stream and is the current video frame of the target object.
[0052] Step S204, obtain a two-dimensional detection frame of the target object.
[0053] In the technical solution provided in step S204 of the present invention, after obtaining the image including the target object, a two-dimensional detection frame of the target object is obtained.
[0054] After obtaining the image of the target object in this embodiment, the image can be subjected to two-dimensional detection. For example, two-dimensional detection of traffic targets is performed on the image to obtain a two-dimensional detection frame of the target object. This two-dimensional detection frame can be represented by a vector X*(x, y, w, h) and can be a bottom vector. Among them, x is used to represent the abscissa of the upper left point of the two-dimensional detection frame, y is used to represent the ordinate of the upper left point of the two-dimensional detection frame, w is used to represent the width of the two-dimensional detection frame, and h is used to represent the height of the two-dimensional detection frame.
[0055] Step S206, based on the two-dimensional detection frame, obtain at least one surface of the three-dimensional detection frame of the target object.
[0056] In the technical solution provided in step S206 of the present invention, after obtaining the two-dimensional detection frame of the target object, based on the two-dimensional detection frame, at least one surface of the three-dimensional detection frame (3D) of the target object is obtained.
[0057] This embodiment can obtain a transformation matrix based on the image and perform transformation processing on the two-dimensional detection frame of the target object through the transformation matrix to obtain at least one surface of the three-dimensional detection frame of the target object. This at least one surface can be the bottom surface of the three-dimensional detection frame.
[0058] In this embodiment, based on the image of the target object, it can be determined whether there is a transformation matrix. This transformation matrix can be used to convert the two-dimensional detection frame of the target object into at least one surface of the three-dimensional detection frame of the target object. This at least one surface of the three-dimensional detection frame can be a parallelogram and can be represented by a vector y*(x 1 , y 1 , x 2, y 2 , w), where (x 1 , y 1 ) can be used to represent the abscissa and ordinate of the upper left point of the parallelogram respectively, and (x 2 , y 2 ) can be used to represent the abscissa and ordinate of the lower left point of the parallelogram respectively, and w can be used to represent the width of the parallelogram.
[0059] Optionally, the above conversion matrix of this embodiment is a linear conversion matrix, which can linearly and quickly transform the two-dimensional detection frame of the target object to at least one surface of the three-dimensional detection frame of the target object accurately.
[0060] Step S208, determine the projection information of the target object on the target plane through at least one surface of the three-dimensional detection frame of the target object.
[0061] In the technical solution provided in step S208 of the present invention above, after obtaining at least one surface of the three-dimensional detection frame of the target object based on the two-dimensional detection frame, determine the projection information of the target object on the target plane through at least one surface of the three-dimensional detection frame of the target object.
[0062] In this embodiment, after performing conversion processing on the two-dimensional detection frame of the target object through the conversion matrix to obtain at least one surface of the three-dimensional detection frame of the target object, the projection information of the target object on the target plane can be determined through at least one surface of the three-dimensional detection frame of the target object.
[0063] In this embodiment, the projection information of the target object on the target plane indicates the projection of the target object on the target plane, thereby achieving the purpose of physically mapping the target object. Among them, the projection information is also the physical space mapping result of the target object, and output this physical space mapping result.
[0064] This embodiment can convert the two-dimensional detection frames of all objects included in the image of the target object into at least one surface of the three-dimensional detection frame through the conversion matrix, and determine the physical space mapping results of all objects on the target plane, and then output the physical space mapping results of all objects on the target plane to complete the physical space mapping process of the objects.
[0065] Through the above steps S202 to S208, an image containing the target object is obtained; a two-dimensional detection frame of the target object is obtained; based on the two-dimensional detection frame, at least one surface of the three-dimensional detection frame of the target object is obtained; and the projection information of the target object on the target plane is determined through at least one surface of the three-dimensional detection frame of the target object. That is to say, in this embodiment, the three-dimensional detection frame is obtained by using the two-dimensional detection frame of the object, and then the projection information of the target object on the road is obtained through the three-dimensional detection frame, achieving the purpose of physically mapping the object, avoiding directly performing two-dimensional detection on the object to judge the object's behavior. Since the plane and the picture are not in the same plane, the detection of the object fails in some cases, thus solving the technical problem of low accuracy in positioning the object and achieving the technical effect of accurately positioning the object.
[0066] Further, in terms of vehicle positioning, this embodiment uses the two-dimensional detection frame of the vehicle and quickly and accurately converts it using a transformation matrix to obtain the projection information of the vehicle on the road, achieving the purpose of physically mapping the vehicle, avoiding directly performing two-dimensional detection on the vehicle to judge the vehicle's behavior. Since the road plane and the picture are not in the same plane, the detection of the vehicle fails in some cases. Therefore, in aspects such as vehicle positioning in a lane, for example, in determining the parking space where the vehicle is located, in vehicle positioning based on a high-speed parallel narrow lane, and in calculating the vehicle speed, this embodiment can also solve the problem of low accuracy in positioning the vehicle through the above method and achieve the technical effect of accurately positioning the vehicle.
[0067] The above method of this embodiment will be further introduced below.
[0068] As an optional implementation manner, obtaining a transformation matrix based on an image includes: determining that there is a preset position in the image; obtaining the transformation matrix associated with the preset position, where the transformation matrix is used to convert the two-dimensional detection frame of the object corresponding to any image with the preset position into at least one surface of the three-dimensional detection frame of the object.
[0069] Optionally, if the image in the above method is the current video frame, any image is any video frame, and the object is a vehicle, then obtaining a transformation matrix based on the current video frame includes: determining that there is a preset position in the current video frame of the target vehicle; obtaining the transformation matrix associated with the preset position, where the transformation matrix is used to convert the two-dimensional detection frame of the vehicle corresponding to any video frame with the preset position into the bottom surface of the three-dimensional detection frame of the vehicle.
[0070] In this embodiment, when obtaining the transformation matrix based on the current video frame, a pre-position judgment can be first performed on the current video frame to determine whether there is a pre-position in the current video frame. The pre-position is a way to associate the key area to be monitored with the operating status of the PTZ camera. When the PTZ runs to the key area to be monitored on the traffic road, a command to set the pre-position can be sent to the PTZ camera, and the PTZ camera will record the orientation of the PTZ and the state of the camera at this time, and associate them with the number of the pre-position. When a recall command is sent to the PTZ camera, the PTZ immediately runs to the pre-position at the fastest speed, and the camera also returns to the state remembered at that time, so as to facilitate the monitoring personnel to quickly view the area to be monitored.
[0071] In this embodiment, after determining that there is a pre-position in the current video frame, it can be determined whether there is already a transformation matrix for converting the two-dimensional detection frame of the vehicle indicated by any video frame to the bottom surface of the three-dimensional detection frame. If it is determined that the above transformation matrix already exists for the pre-position, the transformation matrix can be obtained, and the two-dimensional detection frame of the target vehicle can be transformed using the transformation matrix to obtain the bottom surface of the three-dimensional detection frame of the target vehicle.
[0072] This embodiment can utilize the characteristic of the scene space invariance of the pre-position to perform physical space mapping on the two-dimensional detection frame of the vehicle through the existing transformation matrix, so as to obtain the projection of the vehicle on the road plane. And this embodiment performs physical mapping transformation on the entire vehicle for a single pre-position, which can effectively avoid the problems that are easily affected by light, occlusion, and small targets in three-dimensional detection, so that the scene can be generalized and the recall rate is higher.
[0073] As an alternative embodiment, the method further includes: when the pre-position is not associated with a transformation matrix, performing three-dimensional detection on the image to obtain at least one surface of the three-dimensional detection frame of the target object.
[0074] In this embodiment, the image in the above method can be the current video frame, the target object can be the target vehicle, and at least one surface of the three-dimensional detection frame can be the bottom surface of the three-dimensional detection frame. Then, when the pre-position is not associated with a transformation matrix, three-dimensional detection is performed on the current video frame of the target vehicle to obtain the bottom surface of the three-dimensional detection frame of the target vehicle.
[0075] In this embodiment, when it is determined that the preset position is not associated with a conversion matrix, that is, it is determined that there is no conversion matrix for converting the two-dimensional detection frame into the bottom surface of the three-dimensional detection frame at this preset position, the three-dimensional detection is directly performed on the current video frame. That is, the current video frame is detected according to the algorithm for three-dimensional traffic target detection, and the bottom surface of the three-dimensional detection frame of the target vehicle is obtained. At the same time, the projection information of the target vehicle on the traffic road at this time can be output, that is, the physical space mapping result at this time is output. Among them, the three-dimensional traffic target detection, that is, the three-dimensional vehicle target detection, is an algorithm for three-dimensional target vehicle detection and three-dimensional position estimation by combining deep neural network regression learning and geometric constraints.
[0076] The method for determining the above conversion matrix in this embodiment will be introduced below.
[0077] As an alternative embodiment, the method further includes: determining a conversion matrix based on at least one surface of the three-dimensional detection frame of the target object and the two-dimensional detection frame of the target object.
[0078] Optionally, the above target object may be a target vehicle, and at least one surface of the three-dimensional detection frame may be the bottom surface of the three-dimensional detection frame. When the preset position is not associated with a conversion matrix, the above method may be to determine the conversion matrix based on the bottom surface of the three-dimensional detection frame of the target vehicle and the two-dimensional detection frame of the target vehicle.
[0079] In this embodiment, when the preset position is not associated with a conversion matrix, it is necessary to determine the conversion matrix for converting from the two-dimensional detection frame to the three-dimensional detection frame at this preset position, and the conversion matrix can be determined based on the bottom surface of the three-dimensional detection frame of the target vehicle that has been detected and the two-dimensional detection frame of the target vehicle.
[0080] As an alternative embodiment, determining the conversion matrix based on at least one surface of the three-dimensional detection frame of the target object and the two-dimensional detection frame of the target object includes: performing one-to-one matching between the two-dimensional detection frame of the target object and at least one surface of the three-dimensional detection frame of the target object to obtain a first matching result; performing at least linear regression calculation on the first matching result to obtain a linear regression model; and determining the weight of the linear regression model as the conversion matrix.
[0081] In this embodiment, the target object in the above method may be a target vehicle, and at least one surface of the three-dimensional detection frame may be the bottom surface of the three-dimensional detection frame. Then the above method may be to determine the conversion matrix based on the bottom surface of the three-dimensional detection frame of the target vehicle and the two-dimensional detection frame of the target vehicle, including: performing one-to-one matching between the two-dimensional detection frame of the target vehicle and the bottom surface of the three-dimensional detection frame of the target vehicle to obtain a first matching result; performing at least linear regression calculation on the first matching result to obtain a linear regression model; and determining the weight of the linear regression model as the conversion matrix.
[0082] In this embodiment, when determining the transformation matrix based on the bottom surface of the three-dimensional detection box of the target vehicle and the two-dimensional detection box of the target vehicle, the bottom surface of the two-dimensional detection box of the target vehicle and the three-dimensional detection box of the target vehicle can be matched one by one to obtain a matching result, that is, a set of two-dimensional to three-dimensional detection box pairs corresponding to the current video frame image of the target vehicle can be obtained, and linear regression calculation can be performed on it to obtain a linear regression model. Among them, linear regression calculation refers to an algorithm that makes predictions through linear combinations, and its purpose is to find a straight line, a plane, or a hyperplane in a higher dimension, so that the error between the predicted value and the true value is minimized. This embodiment can apply the pseudo-inverse matrix in the vehicle physical space mapping, and the pseudo-inverse matrix of the linear regression model can be obtained. After the pseudo-inverse matrix remains stable after multiple online linear regression calculations, the linear regression calculation is stopped. At this time, the weight of the pseudo-linear regression model can be determined as the transformation matrix for converting the two-dimensional detection box into a three-dimensional detection box.
[0083] As an alternative implementation, one-to-one matching is performed on at least one surface of the two-dimensional detection box of the target object and the three-dimensional detection box of the target object to obtain a first matching result, including: obtaining a parallelogram corresponding to at least one surface of the three-dimensional detection box of the target object; performing one-to-one matching on the parallelogram and the two-dimensional detection box of the target object to obtain a first matching result.
[0084] In this embodiment, the target object in the above method can be the target vehicle, and at least one surface of the three-dimensional detection box can be the bottom surface of the three-dimensional detection box. Then the above method can be to perform one-to-one matching on the bottom surface of the two-dimensional detection box of the target vehicle and the three-dimensional detection box of the target vehicle to obtain a first matching result, including: obtaining a parallelogram corresponding to the bottom surface of the three-dimensional detection box of the target vehicle; performing one-to-one matching on the parallelogram and the two-dimensional detection box of the target vehicle to obtain a first matching result.
[0085] In this embodiment, when implementing one-to-one matching between the two-dimensional detection box of the target vehicle and the bottom surface of the three-dimensional detection box of the target vehicle to obtain the first matching result, the bottom surface of the obtained three-dimensional detection box can be subjected to two-dimensional transformation, and the parallelogram corresponding to the bottom surface of the three-dimensional detection box can be calculated. This parallelogram can be the minimum two-dimensional circumscribed rectangle of the bottom surface of the three-dimensional detection box. After calculating the parallelogram corresponding to the bottom surface of the three-dimensional detection box, the two-dimensional detection box is then subjected to improvement (refining) processing such as inflation (infate method), and the Hungarian algorithm is used based on the Intersection over Union (IOU) to perform one-to-one matching between the above parallelogram and the processed two-dimensional detection box, thereby obtaining the first matching result.
[0086] It should be noted that in this embodiment, the bottom edge of the bottom quadrilateral of the above three-dimensional detection box is parallel to the horizontal line, that is, it can be ideally assumed that the edges of the bottom surface of the three-dimensional detection box are parallel to the horizontal line, and the offset angle can be ignored.
[0087] As an alternative embodiment, the method further includes: obtaining at least one second matching result stored in the memory, where each second matching result is obtained by one-to-one matching between the two-dimensional detection box of the object corresponding to each historical image with a preset bit and at least one surface of the corresponding three-dimensional detection box, and each historical image is generated before the current image; at least performing linear regression calculation on the first matching result to obtain a linear regression model, including: performing linear regression calculation on the first matching result and at least one second matching result to obtain a linear regression model.
[0088] In this embodiment, the above historical image can be a historical video frame, and at least one surface of the three-dimensional detection box can be the bottom surface of the three-dimensional detection box. Then the method can be to obtain at least one second matching result stored in the memory, where each second matching result is obtained by one-to-one matching between the two-dimensional detection box of the vehicle corresponding to each historical video frame with a preset bit and the bottom surface of the corresponding three-dimensional detection box, and each historical video frame occurs before the current video frame; at least performing linear regression calculation on the first matching result to obtain a linear regression model, including: performing linear regression calculation on the first matching result and at least one second matching result to obtain a linear regression model.
[0089] In this embodiment, at least one second matching result stored in the memory is obtained. Each of the second matching results is obtained by one-to-one matching of the two-dimensional detection frame of the vehicle corresponding to each historical video frame with a preset bit and the bottom surface of the corresponding three-dimensional detection frame. The occurrence time of each historical video frame is earlier than the occurrence time of the current video frame. Then, a linear regression calculation is performed by combining the first matching result and at least one second matching result stored in the memory to obtain a linear regression model.
[0090] Optionally, the linear regression model is obtained by the following method in this embodiment:
[0091] The weight W of the linear regression model 权 = X + Y, where X + is used to represent the pseudo-inverse matrix of the linear regression model, and Y is used to represent the bottom surface of the three-dimensional detection frame. The W 权 can be the weight obtained during the calculation of the linear regression model. That is, the weight of the linear regression model in this embodiment can be obtained by solving the pseudo-inverse matrix of the linear regression model.
[0092] The linear regression model can be represented by Y = W T 权 X, or Y = XW 权 where Y is used to represent the bottom surface of the three-dimensional detection frame, W 权 is used to represent the weight during the calculation of the linear regression model, and X is used to represent the two-dimensional detection frame.
[0093] As an alternative implementation, after at least performing a linear regression calculation on the first matching result to obtain a linear regression model, the method further includes: obtaining at least one third matching result without changing the preset position, where each third matching result is obtained by matching the two-dimensional detection frame of the object corresponding to each new image with a preset position and at least one surface of the corresponding three-dimensional detection frame, and each new image is generated after the image; updating the linear regression model with at least one third matching result.
[0094] In this embodiment, the above new image can be a new video frame, and at least one surface of the three-dimensional detection frame can be the bottom surface of the three-dimensional detection frame. After at least performing a linear regression calculation on the first matching result to obtain a linear regression model, the method further includes: obtaining at least one third matching result without changing the preset position, where each third matching result is obtained by matching the two-dimensional detection frame of the vehicle corresponding to each new video frame with a preset position and the bottom surface of the corresponding three-dimensional detection frame, and each new video frame occurs after the current video frame; updating the linear regression model with at least one third matching result.
[0095] In this embodiment, each time a new video frame is acquired, while ensuring that the preset position remains unchanged, the linear regression model is updated online with the new video frame. The bottom surfaces of the two-dimensional detection box and the corresponding three-dimensional detection box of the vehicle corresponding to each new video frame with a preset position can be matched to obtain a third matching result. Furthermore, the linear regression model is updated through at least one third matching result, so that Y = W T 权 X, or Y = XW 权 In Y = WX or Y = XW, Y can represent the bottom surface of the new three-dimensional detection box, and X can represent the new two-dimensional detection box.
[0096] As an optional implementation manner, the weights of the updated linear regression model are obtained; determining the weights of the linear regression model as the transformation matrix includes: when the change between the weights of the updated linear regression model and the weights of the linear regression model before update is within the target threshold, determining the weights of the updated linear regression model as the transformation matrix.
[0097] In this embodiment, after the linear regression model is updated while the preset position remains unchanged, the weights of the updated linear regression model and the weights of the linear regression model before update can be obtained, and it is determined whether the change between the weights of the updated linear regression model and the weights of the linear regression model before update is within the target threshold. If it is determined that the change between the weights of the updated linear regression model and the weights of the linear regression model before update is within the target threshold, it can be determined that the weights of the linear regression model tend to be stable, and the weights of the updated linear regression model can be determined as the transformation matrix.
[0098] Optionally, this embodiment can also obtain the similarity between the weights of the updated linear regression model and the weights of the linear regression model before update, determine whether the similarity is within the preset threshold. If it is determined that the similarity is within the preset threshold, then it is determined whether the change between the weights of the updated linear regression model and the weights of the linear regression model before update is within the target threshold. If so, it is determined that the weights of the linear regression model tend to be stable, and the weights of the updated linear regression model can be determined as the transformation matrix.
[0099] Optionally, this embodiment can also obtain the pseudo-inverse matrix of the linear regression model before update and the pseudo-inverse matrix of the updated linear regression model, and determine whether the change or similarity between the pseudo-inverse matrix of the linear regression model before update and the pseudo-inverse matrix of the updated linear regression model is within the corresponding threshold. If it is determined that the change or similarity between the pseudo-inverse matrix of the linear regression model before update and the pseudo-inverse matrix of the updated linear regression model is within the corresponding threshold, it is determined that the weights of the linear regression model tend to be stable, and the weights of the updated linear regression model can be determined as the transformation matrix.
[0100] As an alternative implementation, performing three-dimensional detection on an image to obtain at least one surface of a three-dimensional detection box of a target object includes: performing three-dimensional detection on the image through a three-dimensional detection model to obtain at least one surface of the three-dimensional detection box of the target object, where the three-dimensional detection model is trained through pre-collected image samples and at least one surface of the corresponding three-dimensional detection box.
[0101] In this embodiment, the image may be the current video frame of the target vehicle, at least one surface of the three-dimensional detection box may be the bottom surface of the three-dimensional detection box, and the image sample may be a video frame sample. Then the above method may be to perform three-dimensional detection on the current video frame to obtain the bottom surface of the three-dimensional detection box of the target vehicle, including: performing three-dimensional detection on the current video frame through a three-dimensional detection model to obtain the bottom surface of the three-dimensional detection box of the target vehicle, where the three-dimensional detection model is trained through pre-collected video frame samples and the bottom surface of the corresponding three-dimensional detection box.
[0102] In this embodiment, in the case where there is no conversion matrix for converting a two-dimensional detection box into a three-dimensional detection box at the preset position, it is possible to directly enter the three-dimensional detection model for model inference, perform three-dimensional detection on the current video frame through the three-dimensional detection model, so as to obtain the bottom surface of the three-dimensional detection box of the target vehicle, where the three-dimensional detection model, that is, the short-term traffic target three-dimensional detection module, is trained through pre-collected video frame samples and the bottom surface of the corresponding three-dimensional detection box. The video frame sample and the corresponding bottom surface of the three-dimensional detection box may be a public data set and a small amount of labeled traffic data sets, so as to train them to obtain a three-dimensional detection module for giving accurate results of the three-dimensional detection box of the vehicle for real-time video frames in a short period of time, that is, performing short-term three-dimensional detection. Once the conversion matrix from the two-dimensional detection box based on a single preset position to the bottom surface of the three-dimensional detection box is calculated online subsequently, it is no longer necessary to perform model inference through this three-dimensional detection module.
[0103] As an alternative example, different from the above idea of short-term three-dimensional detection and linear transformation using a two-dimensional detection box in this solution, this embodiment may directly perform three-dimensional detection on a real-time video stream and output a physical space mapping result, and this method strongly depends on the accuracy of the three-dimensional detection box.
[0104] Compared with other technical solutions with relatively high costs, this embodiment does not rely on additional information such as radar and depth. It trains a 3D detection model with fewer samples in the early stage, detects the 3D detection box of the target object through this simple secondary model, determines the transformation matrix by combining the 2D detection box of the target object, and directly uses the transformation matrix and the 2D detection box of the object to determine the 3D detection box of the object in the later stage. This is more effective than directly predicting the 3D detection box of the object in scenarios such as occlusion, night, and distance, thus achieving the effect of accurately positioning the object.
[0105] In addition, related technologies based on physical mapping solutions such as radar, depth, and binocular all pose higher requirements for the amount of calculation. However, this embodiment only performs short-term inference on the 3D detection model, uses the transformation matrix with small computational complexity to quickly and accurately transform the 2D detection box for long-term calculation, so the real-time performance is higher, achieving the effect of accurately positioning the object.
[0106] Since this embodiment performs physical mapping transformation on all objects for a single preset position, it can effectively avoid problems that are easily affected by light, occlusion, and small targets in 3D detection. The scene is more generalizable and the recall rate is higher; for object positioning (such as parking spaces in the service area, lanes on the highway main road) and lane occupancy rate, the information of at least one surface of the 3D detection box of the object can be directly used for quick calculation, thus achieving the effect of accurately positioning the object.
[0107] Through the above method, this embodiment can focus on the object positioning problem in the traffic scene. For example, simplifying the occupancy of parking spaces in the service area into the physical space mapping of the object, that is, determining the projection information of the object on the road plane, and solving the object positioning problem by combining multiple dimensions of 2D detection and 3D detection of the object.
[0108] The embodiment of the present invention also provides another data processing method from the perspective of user interaction.
[0109] Figure 3 It is a flowchart of another data processing method according to the embodiment of the present invention. As Figure 3 shown, this method may include the following steps:
[0110] Step S302, display an image including the target object on the target interface, and display the 2D detection box of the target object.
[0111] In the technical solution provided in step S302 of the present invention above, the target object may be a target vehicle, and the image may be the current video frame, which may be the current video frame of the target vehicle on the traffic road displayed on the target interface. Optionally, obtain the video stream obtained by real-time monitoring of the traffic road, and then obtain the current video frame of the target vehicle from the video stream, and display the current video frame on the target interface.
[0112] After obtaining an image containing a target object, this embodiment can perform two-dimensional detection on the image to obtain a two-dimensional detection box of the target object, and can display the two-dimensional detection box of the target object on the target interface. For example, display the two-dimensional detection box of the target vehicle on the target interface.
[0113] Step S304: Display at least one surface of the three-dimensional detection box of the target object on the target interface, where at least one surface of the three-dimensional detection box of the target object is obtained based on the two-dimensional detection box.
[0114] In the technical solution provided in step S304 of the present invention, after displaying the two-dimensional detection box of the target object on the target interface, display at least one surface of the three-dimensional detection box of the target object on the target interface, where at least one surface of the three-dimensional detection box of the target object is obtained by performing transformation processing on the two-dimensional detection box of the target object through a transformation matrix, and the transformation matrix is obtained based on the image.
[0115] In this embodiment, based on the image of the target object, it can be determined whether there is a transformation matrix, which can be used to transform the two-dimensional detection box of the target object into at least one surface of the three-dimensional detection box of the target object. At least one surface of the three-dimensional detection box can be the bottom surface and can be a parallelogram.
[0116] Optionally, the transformation matrix of this embodiment is a linear transformation matrix, which can linearly, quickly, and accurately transform the two-dimensional detection box of the target object to at least one surface of the three-dimensional detection box of the target object, and then display at least one surface of the three-dimensional detection box of the target object on the target interface.
[0117] Step S306: Display the projection information of the target object on the target plane on the target interface, where the projection information is determined by at least one surface of the three-dimensional detection box of the target object.
[0118] In the technical solution provided in step S306 of the present invention, after displaying at least one surface of the three-dimensional detection box of the target object on the target interface, display the projection information of the target object on the target plane on the target interface, and the projection information is determined by at least one surface of the three-dimensional detection box of the target object.
[0119] In this embodiment, the projection information displayed on the target interface indicates the projection of the target object on the target plane, thereby achieving the purpose of physically mapping the target object. Among them, the projection information is also the physical space mapping result of the target object, and output this physical space mapping result.
[0120] In this embodiment, the two-dimensional detection boxes of all objects included in the image can be converted into at least one surface of a three-dimensional detection box through a transformation matrix, and the physical space mapping results of all objects on the target plane can be determined. Furthermore, the physical space mapping results of all objects on the target plane are displayed on the target interface, completing the physical space mapping process of the objects.
[0121] Through the above steps S302 to S306, an image containing the target object is displayed on the target interface, and the two-dimensional detection box of the target object is displayed; at least one surface of the three-dimensional detection box of the target object is displayed on the target interface, where at least one surface of the three-dimensional detection box of the target object is obtained based on the two-dimensional detection box; the projection information of the target object on the target plane is displayed on the target interface, where the projection information is determined by at least one surface of the three-dimensional detection box of the target object. That is to say, in this embodiment, the three-dimensional detection box is obtained by using the two-dimensional detection box of the object, and then the projection information of the target object on the plane is obtained through the three-dimensional detection box, achieving the purpose of physical space mapping of the object, avoiding directly performing two-dimensional detection on the object to judge the object behavior. Since the plane and the picture are not in the same plane, the data processing fails in some cases, thus solving the technical problem of low accuracy in positioning the object and achieving the technical effect of accurately positioning the object.
[0122] The embodiment of the present invention also provides another data processing method.
[0123] Figure 4 It is a flowchart of another data processing method according to the embodiment of the present invention. As Figure 4 shown, the method may include the following steps:
[0124] Step S402, obtain an image containing the target object.
[0125] In the technical solution provided in step S402 of the present invention above, the target object is an object that needs to be detected on the target plane, and an image obtained by real-time monitoring of the target plane can be obtained, which is also the real-time video frame in the video stream. Optionally, obtain the current video frame of the target vehicle on the traffic road.
[0126] Step S404, determine that there is a preset position in the image.
[0127] In the technical solution provided in step S404 of the present invention above, after obtaining the image of the target object on the target plane, it is determined that there is a preset position in the image. Optionally, in this embodiment, it is determined that there is a preset position in the current video frame.
[0128] This embodiment can judge whether there is a preset position in the image, and the preset position is a way to associate the key area to be monitored with the operating condition of the pan-tilt camera.
[0129] Step S406: When no transformation matrix is associated with the preset position, perform three-dimensional detection on the image to obtain at least one surface of the three-dimensional detection frame of the target object.
[0130] In the technical solution provided in step S406 of the present invention, it may be that when no transformation matrix is associated with the preset position, perform three-dimensional detection on the current video frame to obtain the bottom surface of the three-dimensional detection frame of the target vehicle. The transformation matrix is used to convert the two-dimensional detection frame of the object corresponding to any image with a preset position into at least one surface of the three-dimensional detection frame of the object. For example, the transformation matrix is used to convert the two-dimensional detection frame of the vehicle corresponding to any video frame with a preset position into the bottom surface of the three-dimensional detection frame of the vehicle.
[0131] Step S408: Determine the projection information of the target object on the target plane through at least one surface of the three-dimensional detection frame of the target object.
[0132] In the technical solution provided in step S408 of the present invention, the projection information of the target object on the target plane indicates the projection of the target object on the target plane. It may be to determine the projection information of the target vehicle on the traffic road through the bottom surface of the three-dimensional detection frame of the target vehicle, so as to achieve the purpose of physically mapping the target object. Among them, the projection information is also the physical space mapping result of the target object, and output this physical space mapping result.
[0133] This embodiment can convert the two-dimensional detection frames of all objects included in the image into at least one surface of the three-dimensional detection frame through the transformation matrix, determine the physical space mapping results of all objects on the target plane, and then output the physical space mapping results of all objects on the target plane to complete the physical space mapping process of the objects.
[0134] As an alternative embodiment, in step S402, after obtaining the image containing the target object, the method further includes: performing two-dimensional detection on the image to obtain the two-dimensional detection frame of the target object; when a transformation matrix is associated with the preset position, perform transformation processing on the two-dimensional detection frame of the target object through the transformation matrix to obtain at least one surface of the three-dimensional detection frame of the target object.
[0135] In this embodiment, the above image may be a video frame, the target object may be a target vehicle, and at least one surface of the three-dimensional detection frame may be the bottom surface of the three-dimensional detection frame. The above method may be to perform two-dimensional detection on the current video frame of the target vehicle on the traffic road after obtaining it, to obtain the two-dimensional detection frame of the target vehicle; when a transformation matrix is associated with the preset position, perform transformation processing on the two-dimensional detection frame of the target vehicle through the transformation matrix to obtain the bottom surface of the three-dimensional detection frame of the target vehicle.
[0136] After obtaining the current video frame of the target vehicle, two-dimensional detection can also be performed on the current video frame, that is, two-dimensional traffic target detection is performed on the current video frame to obtain the two-dimensional detection box of the target vehicle. The two-dimensional detection box can be represented by the vector X*(x, y, w, h), where x is used to represent the abscissa of the upper left point of the two-dimensional detection box, y is used to represent the ordinate of the upper left point of the two-dimensional detection box, w is used to represent the width of the two-dimensional detection box, and h is used to represent the height of the two-dimensional detection box. If it is determined that the above conversion matrix already exists in the preset position, the conversion matrix can be obtained, and the two-dimensional detection box of the target vehicle can be converted by using the conversion matrix, so as to obtain the bottom surface of the three-dimensional detection box of the target vehicle, and further determine the projection information of the target vehicle on the traffic road through the bottom surface of the three-dimensional detection box of the target vehicle.
[0137] Through the above steps S402 to S408, an image containing the target object is obtained; it is determined that there is a preset position in the image; in the case where the preset position is not associated with a conversion matrix, three-dimensional detection is performed on the image to obtain at least one surface of the three-dimensional detection box of the target object, where the conversion matrix is used to convert the two-dimensional detection box of the object corresponding to any image with a preset position into at least one surface of the three-dimensional detection box of the object; the projection information of the target object on the target plane is determined through at least one surface of the three-dimensional detection box of the target object. That is to say, this embodiment avoids directly performing two-dimensional detection on the object to judge the object behavior. Since the plane and the picture are not in the same plane, the data processing fails in some cases, thus solving the technical problem of low positioning accuracy of the object and achieving the technical effect of accurately positioning the object.
[0138] The data processing method of the embodiment of the present invention will be further introduced from the perspective of cloud services below.
[0139] As an optional example, a target request is obtained, where the target request carries an image to be processed input on the target interface, and the image contains the target object; in response to the target request, a two-dimensional detection box of the target object is obtained; based on the two-dimensional detection box, at least one surface of the three-dimensional detection box of the target object is obtained; the projection information of the target object on the target plane is determined through at least one surface of the three-dimensional detection box of the target object, and the projection information is sent to the target interface for display.
[0140] In this embodiment, the server may obtain a target request, which may be a request entered by a user on a target interface to indicate personalized requirements, and may carry an image to be processed. The image contains a target object, and the target object may be an object to be detected on a target plane. For example, if the target plane is a traffic road, the target object may be a target vehicle to be detected on the traffic road. Optionally, the image to be processed by the server in this embodiment may be an image obtained by real-time monitoring of the target plane. The image may be a monocular RGB image and may form a video stream of traffic video. Furthermore, the server may obtain an image of the target object from the video stream.
[0141] After the server obtains the target request, in response to the target request, it performs two-dimensional detection on the image. For example, the server performs two-dimensional traffic target detection on the current video frame of the target vehicle to obtain a two-dimensional detection box of the target vehicle. The two-dimensional detection box may be represented by a vector X*(x, y, w, h) and may be a bottom vector. Among them, x is used to represent the abscissa of the upper left point of the two-dimensional detection box, y is used to represent the ordinate of the upper left point of the two-dimensional detection box, w is used to represent the width of the two-dimensional detection box, and h is used to represent the height of the two-dimensional detection box.
[0142] After the server obtains the two-dimensional detection box of the target object, it may obtain at least one surface of the three-dimensional detection box of the target object based on the two-dimensional detection box. Optionally, the server may obtain a transformation matrix based on the image of the target object and perform transformation processing on the two-dimensional detection box of the target object through the transformation matrix to obtain at least one surface of the three-dimensional detection box of the target object. The at least one surface may be the bottom surface of the three-dimensional detection box. Optionally, the server may determine whether there is a transformation matrix based on the image of the target object. The transformation matrix may be used to transform the two-dimensional detection box of the target object into at least one surface of the three-dimensional detection box of the target object. The at least one surface of the three-dimensional detection box may be a parallelogram and may be represented by a vector y*(x 1 , y 1 , x 2 , y 2 , w), where (x 1 , y 1 ) may be used to represent the abscissa and ordinate of the upper left point of the parallelogram respectively, (x 2 , y 2 ) may be used to represent the abscissa and ordinate of the lower left point of the parallelogram respectively, and w may be used to represent the width of the parallelogram.
[0143] After the server processes the two-dimensional detection box of the target object through a transformation matrix to obtain at least one surface of the three-dimensional detection box of the target object, the projection information of the target object on the target plane can be determined through at least one surface of the three-dimensional detection box of the target object, and the projection information is sent to the target interface for display. Among them, the projection information of the target object on the target plane indicates the projection of the target object on the target plane, thereby achieving the purpose of physically mapping the target object. Among them, the projection information is also the physical space mapping result of the target object, and this physical space mapping result is output.
[0144] Optionally, the server in this embodiment can convert the two-dimensional detection boxes of all objects included in the image of the target object into at least one surface of the three-dimensional detection box through the transformation matrix, determine the physical space mapping results of all objects on the target plane, and then output the physical space mapping results of all objects on the target plane on the target interface, completing the physical space mapping process of the objects, thereby achieving the purpose of meeting the personalized needs of users using cloud services.
[0145] In the above method of this embodiment, the server uses the two-dimensional detection box of the object to obtain the three-dimensional detection box, and then through the three-dimensional detection box, the projection information of the target object on the road is obtained, achieving the purpose of physically mapping the object, avoiding directly performing two-dimensional detection on the object to judge the object's behavior. Since the plane and the picture are not in the same plane, the detection of the object fails in some cases, thus solving the technical problem of low accuracy in positioning the object and achieving the technical effect of accurately positioning the object. Further, in terms of vehicle positioning in the lane, for example, in determining the parking space where the vehicle is located, in vehicle positioning based on a high-speed parallel narrow lane, in calculating the vehicle speed, etc., this embodiment can also solve the problem of low accuracy in positioning the vehicle through the above method and achieve the technical effect of accurately positioning the vehicle.
[0146] Embodiment 2
[0147] The technical solution of the present invention will be further described below in conjunction with the preferred embodiments. Specifically, taking the object in Embodiment 1 as a vehicle, the image as a video frame, and the plane as the road, the technical solution of this embodiment will be further described.
[0148] In the related art, the vehicle behavior is judged by directly performing two-dimensional detection on the vehicle. Since the road plane and the picture are not in the same plane, the detection of the vehicle fails in some cases. Three-dimensional traffic target detection is currently mainly applied in the field of autonomous driving. However, the common methods are to combine radar, depth information, and neural networks, mainly including the following several:
[0149] 1) Radar (LiDAR) + 3D point cloud. The three-dimensional point cloud data is represented as a set composed of unordered data points and is obtained relying on LiDAR, etc. The methods for three-dimensional object detection based on the point cloud data obtained by radar and combined with deep learning can include: (1) Projecting the point cloud to certain specific perspectives (such as the front view perspective and the bird's-eye view perspective) for processing, and at the same time fusing and using the image information from the camera; (2) Dividing the point cloud data into voxels with spatial dependence relationships, introducing spatial dependence relationships into the point cloud data, and then using methods such as three-dimensional convolution for processing. However, the accuracy of this method depends on the fineness of the segmentation in the three-dimensional space, and the computational complexity is also relatively high; (3) The method of directly applying a deep learning model to the point cloud data.
[0150] 2) Binocular (multi-view) stereo vision. Using the parallax principle, learning the automatic alignment of the left and right images, and optimizing the final detection result through dense matching. However, the disadvantage is that it requires an additional binocular camera.
[0151] 3) Depth image (RGB-D). Conducting three-dimensional detection on ordinary RGB three-channel color images and depth maps. However, the latter has a relatively high acquisition cost and requires additional equipment.
[0152] And this embodiment mainly conducts three-dimensional object detection on vehicles based on a monocular RGB image, and based on the technology of two-dimensional detection of traffic vehicles, uses the scene space invariance of a single preset position to perform a physical space mapping on the vehicle to obtain the projection of the vehicle on the road plane.
[0153] Figure 5 It is a flowchart of another method for detecting vehicles according to an embodiment of the present invention. As Figure 5 shown, this method may include the following steps:
[0154] Step S501, input a real-time video stream and obtain a video frame.
[0155] Step S502, perform a preset position determination on the video frame.
[0156] Step S503, perform two-dimensional detection on the video frame.
[0157] Step S504, determine whether there is already a transformation matrix for converting the two-dimensional detection box to the bottom surface of the three-dimensional detection box at the preset position.
[0158] If it is determined that there is already a transformation matrix for converting the two-dimensional detection box to the bottom surface of the three-dimensional detection box at the preset position, then execute step S505; otherwise, execute step S506.
[0159] Step S505: Linearly transform the two-dimensional detection box through a transformation matrix to obtain the bottom surface of the three-dimensional detection box.
[0160] Step S506: Perform three-dimensional detection on the video frame through a three-dimensional detection module to obtain the bottom surface of the three-dimensional detection box.
[0161] If there is no transformation matrix for converting the two-dimensional detection box to the bottom surface of the three-dimensional detection box in the preset position, then enter the three-dimensional detection module for model inference.
[0162] This embodiment can also, after performing two-dimensional detection on the video frame, utilize the two-dimensional-three-dimensional detection box pair to perform three-dimensional detection on the video frame through a three-dimensional detection module to obtain the bottom surface of the three-dimensional detection box.
[0163] Step S507: Output the physical space mapping result through the bottom surface of the three-dimensional detection box.
[0164] After linearly transforming the two-dimensional detection box through a transformation matrix to obtain the bottom surface of the three-dimensional detection box, or performing three-dimensional detection on the video frame through a three-dimensional detection module to obtain the bottom surface of the three-dimensional detection box, output the physical space mapping result through the bottom surface of the three-dimensional detection box.
[0165] Step S508: Calculate the transformation matrix for converting from the two-dimensional detection box to the bottom surface of the three-dimensional detection box under the preset position.
[0166] In this embodiment, the obtained bottom surface of the three-dimensional detection box can be one-to-one matched with the two-dimensional detection box through two-dimensional transformation to obtain a set of two-dimensional-three-dimensional detection box pairs, and calculate the transformation matrix for converting from the two-dimensional detection box to the bottom surface of the three-dimensional detection box under this preset position.
[0167] This embodiment can perform linear regression calculation on a set of 2D-3D detection box pairs of the obtained video frame and the two-dimensional-three-dimensional detection box pairs in the previous memory to obtain a linear regression model.
[0168] Figure 6 It is a schematic diagram of a projection plane calculation according to an embodiment of the present invention. As Figure 6 shown, A is the two-dimensional detection box, B is the bottom surface of the corresponding three-dimensional detection box, and C is the bottom side of B parallel to the horizontal line. The two-dimensional detection box is represented by X*(x, y, w, h), where x is used to represent the abscissa of the upper left point of the two-dimensional detection box, y is used to represent the ordinate of the upper left point of the two-dimensional detection box, w is used to represent the width of the two-dimensional detection box, h is used to represent the height of the two-dimensional detection box, and the bottom surface of the three-dimensional detection box is represented by the vector y*(x 1 , y 1 , x 2 , y 2 , w), where (x1 , y 1 ) can be respectively used to represent the abscissa and ordinate of the upper left point of the parallelogram. (x 2 , y 2 ) can be respectively used to represent the abscissa and ordinate of the lower left point of the parallelogram. w can be used to represent the width of the parallelogram. Calculate the pseudo-inverse matrix X + , and the linear regression model can be obtained through the following formula.
[0169] The weight W of the linear regression model 权 = X + Y, where X + is used to represent the pseudo-inverse matrix of the linear regression model, and Y is used to represent the bottom surface of the three-dimensional detection frame. This W 权 can be the weight obtained during the calculation of the linear regression model. That is, the weight of the linear regression model in the embodiment can be obtained by solving the pseudo-inverse matrix of the linear regression model.
[0170] The linear regression model of this embodiment can be represented by Y = W T 权 X, or Y = XW 权 , where Y is used to represent the bottom surface of the three-dimensional detection frame, and W 权 is used to represent the weight during the calculation of the linear regression model, and X is used to represent the two-dimensional detection frame.
[0171] Step S509, determine whether the conversion matrix of the preset position has been stabilized.
[0172] The W of the linear conversion model in this embodiment can be 权 determined as the conversion matrix. Each new video frame entering can be used to update the linear regression model online. When W 权 changes relatively stably, that is, the conversion matrix has been stabilized.
[0173] Step S510, the calculation of the conversion matrix under the preset position is completed.
[0174] In the case where the conversion matrix has been stabilized, after the calculation of the conversion matrix under the preset position is completed, the linear regression calculation can be stopped. At this time, take Y = W T 权 X, or Y = XW 权 The W in 权 as the final conversion matrix from the two-dimensional detection frame to the bottom surface of the three-dimensional detection frame for this preset position.
[0175] Figure 7 is a schematic diagram of the architecture of a traffic video vehicle physical mapping scheme according to an embodiment of the present invention. As Figure 7As shown, the traffic video vehicle physical mapping architecture of this embodiment may include a short-term traffic object 3D detection module 71 (short-term 3D object detection), a 2D-3D box matching module 72 (2D-3D box matching), a transformation matrix calculation module 73 (online matrix computing), and an online linear transformation module 74 (online linear transformation) in the dotted part.
[0176] The short-term traffic object 3D detection module 71 is used to obtain real-time video frames, perform calculations on the real-time video frames, and perform 2D object detection. Then, within a short period of time, it gives a 3D accurate detection result of traffic objects for the real-time image, completing the physical space mapping of the video frame. Once the subsequent online linear transformation module 74 based on a single preset for the 2D detection box to the bottom surface of the 3D detection is calculated, model inference is no longer performed through this module.
[0177] The 2D-3D box matching module 72 is used to provide a one-to-one matching result between the 2D detection box and the 3D detection box. First, the 3D detection box is subjected to a 2D transformation to calculate the corresponding parallelogram of the 3D detection box, which can be the minimum 2D circumscribed rectangle of the bottom surface of the 3D detection box (assuming the bottom quadrilateral has a bottom side parallel to the horizontal line). After calculating the corresponding parallelogram of the 3D detection box, operations such as inflating (infate) and refining are performed on the 2D detection box, and the matching between the 2D detection box and the 3D detection box can be achieved based on IOU using the hungarian algorithm.
[0178] The transformation matrix calculation module 73 is used to obtain a linear regression model of the 3D detection box y*(x 1 , y 1 , x 2 , y 2 , w) through linear regression using the X*(x, y, w, h) of the 2D detection box. This assumes that the bottom side of the 3D detection box is parallel to the horizontal line and ignores the offset angle. Among them, the weights of the linear regression model can be obtained by solving the pseudo-inverse matrix of the linear regression model. When the pseudo-inverse matrix remains stable after multiple online calculations, the linear regression calculation can be stopped, and at this time, the conversion matrix calculation based on a single preset is completed.
[0179] The online linear transformation module 74 uses the transformation matrix based on a single preset to transform the 2D detection box to obtain a 3D detection box, and completes the physical space mapping of the vehicle through this 3D detection box.
[0180] This embodiment can also achieve two-dimensional target tracking and process traffic events and traffic parameters.
[0181] As an alternative example, different from the idea of short-term three-dimensional detection and linear transformation using two-dimensional detection boxes in this solution, for real-time video streams, three-dimensional detection can be directly performed on each video frame to output the physical space mapping result. This method strongly depends on the accuracy of the three-dimensional detection box.
[0182] Figure 8 It is a schematic diagram of the interaction interface of a vehicle detection method according to an embodiment of the present invention. As Figure 8 shown, the user can drag the video frame of the image of the vehicle on the traffic road into the "Add" text box, and by clicking the "Vehicle Detection" button, detect the video frame of the vehicle, and finally generate the projection information of the target vehicle on the traffic road, so as to achieve the purpose of accurately and quickly detecting the vehicle, thus solving the technical problem of low positioning accuracy of the vehicle and achieving the technical effect of accurately positioning the vehicle.
[0183] Through the above vehicle detection method of this embodiment, compared with other solutions with relatively high costs, this embodiment does not rely on additional information such as radar and depth, uses fewer labeled samples to train the three-dimensional detection model in the early stage, determines the transformation matrix for converting the two-dimensional detection box into a three-dimensional detection box by using the three-dimensional detection box and the two-dimensional detection box, and has better effects than directly predicting 3D detection boxes in scenarios such as occlusion, night, and distance; this embodiment has high real-time performance. Physical mapping solutions based on radar, depth, and binoculars all pose higher requirements for computational complexity, while this embodiment only performs short-term inference on the three-dimensional detection model and uses a transformation matrix with low computational complexity to perform long-term calculations quickly and accurately through linear transformation; this embodiment has a high recall rate. Since the physical mapping transformation of the entire vehicle is performed for a single preset position, it can effectively avoid the problems that are greatly affected by light, occlusion, and small targets in three-dimensional detection, so the scene can be generalized and the recall rate is higher; this embodiment has rich application scenarios. Vehicle positioning (service area parking spaces, highway main road lanes) and lane occupancy rate can directly use the bottom surface information of the vehicle three-dimensional detection box for rapid calculation.
[0184] This embodiment is based on a vehicle 3D detection technology using 2D detection boxes, which does not require additional information such as complex radar (point cloud), depth annotation, and binocular conditions. It uses a publicly available dataset and a small amount of annotated traffic datasets for training to obtain a 3D detection model, and combines more accurate 2D detection boxes to determine the projection information of the vehicle on the traffic road. This embodiment uses the 2D detection box of the vehicle and performs a fast and accurate linear transformation using a transformation matrix to obtain the physical space mapping result, that is, the projection information of the vehicle on the road plane. Focusing on the vehicle positioning problem in the traffic scenario, such as the occupancy of parking spaces in the service area, etc., can be simplified to the physical space mapping of the vehicle, that is, the projection calculation of the vehicle on the road plane, and solves the technical problem of low positioning accuracy of the vehicle by combining multiple dimensions such as 2D detection and 3D detection.
[0185] Embodiment 3
[0186] According to an embodiment of the present invention, there is also provided a data processing apparatus for implementing the above Figure 2 shown data processing method.
[0187] Figure 9 is a schematic diagram of a data processing apparatus according to an embodiment of the present invention. As Figure 9 shown, the data processing apparatus 90 may include: a first acquisition unit 91, a second acquisition unit 92, a third acquisition unit 93, and a first determination unit 94.
[0188] The first acquisition unit 91 is configured to acquire an image including a target object.
[0189] The second acquisition unit 92 is configured to acquire a 2D detection box of the target object.
[0190] The third acquisition unit 93 is configured to acquire at least one surface of the 3D detection box of the target object based on the 2D detection box.
[0191] The first determination unit 94 is configured to determine the projection information of the target object on the target plane through at least one surface of the 3D detection box of the target object.
[0192] It should be noted here that the above first acquisition unit 91, second acquisition unit 92, third acquisition unit 93, and first determination unit 94 respectively correspond to steps S202 to S208 of Embodiment 1. The instances and application scenarios implemented by the four units and the corresponding steps are the same, but are not limited to the content disclosed in the above Embodiment 1. It should be noted that the above units, as part of the apparatus, can run in the computer terminal 10 provided in Embodiment 1.
[0193] According to an embodiment of the present invention, there is also provided a data processing apparatus for implementing the above Figure 3 shown data processing method.
[0194] Figure 10 It is a schematic diagram of another data processing device according to an embodiment of the present invention. As Figure 10 shown, the data processing device 100 may include: a first display unit 101, a second display unit 102, and a third display unit 103.
[0195] The first display unit 101 is configured to display an image including a target object on a target interface, and display a two-dimensional detection frame of the target object.
[0196] The second display unit 102 is configured to display at least one surface of a three-dimensional detection frame of the target object on the target interface, where at least one surface of the three-dimensional detection frame of the target object is obtained based on the two-dimensional detection frame.
[0197] The third display unit 103 is configured to display projection information of the target object on a target plane on the target interface, where the projection information is determined by at least one surface of the three-dimensional detection frame of the target object.
[0198] It should be noted here that the above-mentioned first display unit 91, second display unit 92, and third display unit 93 respectively correspond to steps S302 to S306 of Embodiment 1. The three units have the same implementation examples and application scenarios as the corresponding steps, but are not limited to the content disclosed in the above-mentioned Embodiment 1. It should be noted that the above units, as part of the device, can run in the computer terminal 10 provided in Embodiment 1.
[0199] According to an embodiment of the present invention, there is also provided a data processing device for implementing the above-mentioned Figure 4 shown data processing method.
[0200] Figure 11 It is a schematic diagram of another data processing device according to an embodiment of the present invention. As Figure 11 shown, the data processing device 110 may include: a fourth acquisition unit 111, a second determination unit 112, a detection unit 113, and a third determination unit 114.
[0201] The fourth acquisition unit 111 is configured to acquire an image including a target object.
[0202] The second determination unit 112 is configured to determine that there is a preset position in the image.
[0203] The detection unit 113 is configured to perform three-dimensional detection on the image when the preset position is not associated with a transformation matrix, and obtain at least one surface of a three-dimensional detection frame of the target object, where the transformation matrix is used to convert a two-dimensional detection frame of an object corresponding to any image with a preset position into at least one surface of a three-dimensional detection frame of the object.
[0204] A third determination unit 114, configured to determine projection information of the target object on the target plane through at least one surface of the three-dimensional detection frame of the target object.
[0205] It should be noted here that the above-mentioned third acquisition unit 111, second determination unit 112, detection unit 113, and third determination unit 114 respectively correspond to steps S402 to S408 of Embodiment 1. The functions of the three units and the corresponding steps are the same in terms of the implemented examples and application scenarios, but are not limited to the content disclosed in the above-mentioned Embodiment 1. It should be noted that the above units, as part of the device, can run in the computer terminal 10 provided in Embodiment 1.
[0206] The data processing method of this embodiment uses an object three-dimensional detection technology based on a two-dimensional detection frame, obtains a three-dimensional detection frame by using the two-dimensional detection frame of the object, and further obtains the projection information of the target object on the plane through the three-dimensional detection frame, achieving the purpose of physically mapping the object, avoiding directly performing two-dimensional detection on the object to judge the object's behavior. Since the plane and the picture are not in the same plane, the data processing fails in some cases, thus solving the technical problem of low positioning accuracy of the object and achieving the technical effect of accurately positioning the object.
[0207] Embodiment 4
[0208] An embodiment of the present invention can provide a mobile terminal, which can be any computer terminal device in a computer terminal group.
[0209] Optionally, in this embodiment, the above-mentioned mobile terminal can be at least one network device among multiple network devices in a computer network.
[0210] In this embodiment, the above-mentioned mobile terminal can execute the program code of the following steps in the data processing method of the application program: acquire an image containing the target object; acquire a two-dimensional detection frame of the target object; based on the two-dimensional detection frame, acquire at least one surface of the three-dimensional detection frame of the target object; determine the projection information of the target object on the target plane through at least one surface of the three-dimensional detection frame of the target object.
[0211] Optionally, Figure 12 is a structural block diagram of a mobile terminal according to an embodiment of the present invention. As Figure 12 shown, the mobile terminal A may include: one or more (only one is shown in the figure) processors 122, a memory 124, and a transmission device 126.
[0212] Among them, a transmission device is configured to transmit an image of a target object on a target plane; and a memory is connected to the transmission device and is configured to provide instructions for a processor to perform the following processing steps: obtaining a two-dimensional detection frame of the target object; based on the two-dimensional detection frame, obtaining at least one surface of a three-dimensional detection frame of the target object; and determining projection information of the target object on the target plane through at least one surface of the three-dimensional detection frame of the target object.
[0213] Among them, the memory can be used to store software programs and modules, such as program instructions / modules corresponding to the data processing method and device in the embodiments of the present invention. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, implements the above-mentioned data processing method. The memory may include a high-speed random access memory, and may further include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory may further include a memory remotely disposed relative to the processor, and these remote memories may be connected to the mobile terminal A through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0214] The processor can also call the information and application programs stored in the memory through the transmission device to perform the following steps: obtaining an image including a target object; obtaining a two-dimensional detection frame of the target object; based on the two-dimensional detection frame, obtaining at least one surface of a three-dimensional detection frame of the target object; and determining projection information of the target object on the target plane through at least one surface of the three-dimensional detection frame of the target object.
[0215] Optionally, the above processor further executes program code for the following steps: obtaining a transformation matrix based on the image, and performing transformation processing on the two-dimensional detection frame through the transformation matrix to obtain at least one surface of the three-dimensional detection frame of the target object.
[0216] Optionally, the above processor further executes program code for the following steps: obtaining a transformation matrix based on the image, including: determining that there is a preset position in the image; obtaining a transformation matrix associated with the preset position, where the transformation matrix is used to transform a two-dimensional detection frame of an object corresponding to any image with a preset position into at least one surface of a three-dimensional detection frame of the object.
[0217] Optionally, the above processor further executes program code for the following steps: performing three-dimensional detection on the image in the case where the preset position is not associated with a transformation matrix to obtain at least one surface of the three-dimensional detection frame of the target object.
[0218] Optionally, the above processor further executes program code for the following steps: determining a transformation matrix based on at least one surface of the three-dimensional detection frame of the target object and the two-dimensional detection frame of the target object.
[0219] Optionally, the above-mentioned processor also executes program code for the following steps: one-to-one match at least one surface of the two-dimensional detection box of the target object and the three-dimensional detection box of the target object to obtain a first matching result; perform at least linear regression calculation on the first matching result to obtain a linear regression model; determine the weight of the linear regression model as the transformation matrix.
[0220] Optionally, the above-mentioned processor also executes program code for the following steps: obtain a parallelogram corresponding to at least one surface of the three-dimensional detection box of the target object; perform one-to-one match between the parallelogram and the two-dimensional detection box of the target object to obtain a first matching result.
[0221] Optionally, the above-mentioned processor also executes program code for the following steps: obtain at least one second matching result stored in the memory, where each second matching result is obtained by one-to-one matching at least one surface of the two-dimensional detection box and the corresponding three-dimensional detection box of the object corresponding to each historical image with a preset bit, and each historical image is generated before the image; perform linear regression calculation on the first matching result and the at least one second matching result to obtain a linear regression model.
[0222] Optionally, the above-mentioned processor also executes program code for the following steps: after performing at least linear regression calculation on the first matching result to obtain a linear regression model, without changing the preset bit, obtain at least one third matching result, where each third matching result is obtained by matching at least one surface of the two-dimensional detection box and the corresponding three-dimensional detection box of the object corresponding to each new image with a preset bit, and each new image is generated after the image; update the linear regression model through the at least one third matching result.
[0223] Optionally, the above-mentioned processor also executes program code for the following steps: obtain the weight of the updated linear regression model; determine the weight of the linear regression model as the transformation matrix, including: when the change between the weight of the updated linear regression model and the weight of the linear regression model before update is within the target threshold, determine the weight of the updated linear regression model as the transformation matrix.
[0224] Optionally, the above-mentioned processor also executes program code for the following steps: perform linear transformation on the two-dimensional detection box of the target object through the transformation matrix to obtain at least one surface of the three-dimensional detection box of the target object.
[0225] Optionally, the above-mentioned processor also executes the program code of the following steps: performing three-dimensional detection on an image through a three-dimensional detection model to obtain at least one surface of a three-dimensional detection frame of a target object, where the three-dimensional detection model is trained through pre-collected image samples and at least one surface of the corresponding three-dimensional detection frame.
[0226] Optionally, the above-mentioned processor also executes the program code of the following steps: obtaining a current video frame from a video stream of a target plane.
[0227] Optionally, the above-mentioned processor also executes the program code of the following steps: performing two-dimensional detection on the current video frame to obtain a two-dimensional detection frame of a target vehicle.
[0228] Optionally, the above-mentioned processor also executes the program code of the following steps: obtaining a transformation matrix based on the current video frame, and performing transformation processing on the two-dimensional detection frame of the target vehicle through the transformation matrix to obtain the bottom surface of the three-dimensional detection frame of the target vehicle.
[0229] Optionally, the above-mentioned processor also executes the program code of the following steps: determining projection information of the target vehicle on a traffic road through the bottom surface of the three-dimensional detection frame of the target vehicle.
[0230] As an optional implementation manner, the processor can also call the information and application programs stored in the memory through a transmission device to execute the following steps: displaying an image including the target object on a target interface, and displaying a two-dimensional detection frame of the target object; displaying at least one surface of the three-dimensional detection frame of the target object on the target interface, where at least one surface of the three-dimensional detection frame of the target object is obtained based on the two-dimensional detection frame; displaying projection information of the target object on the target plane on the target interface, where the projection information is determined through at least one surface of the three-dimensional detection frame of the target object.
[0231] As an optional implementation manner, the processor can also call the information and application programs stored in the memory through a transmission device to execute the following steps: obtaining an image including the target object; determining that there is a preset position in the image; in the case where the preset position is not associated with a transformation matrix, performing three-dimensional detection on the image to obtain at least one surface of the three-dimensional detection frame of the target object, where the transformation matrix is used to convert the two-dimensional detection frame of the object corresponding to any image with a preset position into at least one surface of the three-dimensional detection frame of the object; determining projection information of the target object on the target plane through at least one surface of the three-dimensional detection frame of the target object.
[0232] Optionally, the above-mentioned processor also executes the program code of the following steps: after obtaining the image of the target object on the target plane, performing two-dimensional detection on the image to obtain the two-dimensional detection frame of the target object; in the case where a conversion matrix is associated with the preset position, performing conversion processing on the two-dimensional detection frame of the target object through the conversion matrix to obtain at least one surface of the three-dimensional detection frame of the target object.
[0233] As an alternative implementation, the processor can also call the information and application programs stored in the memory through the transmission device to execute the following steps: obtaining a target request, where the target request carries the image to be processed input on the target interface, and the image includes the target object; responding to the target request to obtain the two-dimensional detection frame of the target object; based on the two-dimensional detection frame, obtaining at least one surface of the three-dimensional detection frame of the target object; determining the projection information of the target object on the target plane through at least one surface of the three-dimensional detection frame of the target object, and sending the projection information to the target interface for display.
[0234] By adopting the embodiment of the present invention, a data processing method is provided, which includes obtaining an image containing a target object; obtaining the two-dimensional detection frame of the target object; based on the two-dimensional detection frame, obtaining at least one surface of the three-dimensional detection frame of the target object; and determining the projection information of the target object on the target plane through at least one surface of the three-dimensional detection frame of the target object. That is to say, this application is a three-dimensional object detection technology based on a two-dimensional detection frame, which uses the two-dimensional detection frame of the object to obtain a three-dimensional detection frame, and then obtains the projection information of the target object on the plane through the three-dimensional detection frame, achieving the purpose of physically mapping the object, avoiding directly performing two-dimensional detection on the object to judge the object's behavior. Since the plane and the picture are not in the same plane, the detection of the object fails in some cases, thus solving the technical problem of low accuracy in positioning the object and achieving the technical effect of accurately positioning the object.
[0235] Those of ordinary skill in the art can understand that Figure 12 the structure shown is only schematic, and the mobile terminal A can also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a palm computer, and terminal devices such as Mobile Internet Devices (MID), PAD, etc. Figure 12 It does not limit the structure of the above-mentioned mobile terminal. For example, the mobile terminal A may further include more or fewer components (such as a network interface, a display device, etc.) than those shown Figure 12 in the figure, or have a different configuration from that shown Figure 12 in the figure.
[0236] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by a program instructing the hardware related to the terminal device. The program can be stored in a computer-readable storage medium, and the storage medium can include: a flash drive, a read-only memory (ROM for short), a random access memory (RAM for short), a magnetic disk, an optical disk, etc.
[0237] Embodiment 5
[0238] An embodiment of the present invention also provides a storage medium. Optionally, in this embodiment, the above storage medium can be used to store the program code executed by the data processing method provided in the above Embodiment 1.
[0239] Optionally, in this embodiment, the above storage medium can be located in any one of the computer terminals in the computer terminal group in the computer network, or in any one of the mobile terminals in the mobile terminal group.
[0240] Optionally, in this embodiment, the storage medium is set to store program code for performing the following steps: obtaining an image containing a target object; obtaining a two-dimensional detection box of the target object; based on the two-dimensional detection box of the image, obtaining at least one surface of the three-dimensional detection box of the target object; and determining the projection information of the target object on the target plane through at least one surface of the three-dimensional detection box of the target object.
[0241] Optionally, the storage medium is further set to store program code for performing the following steps: obtaining a transformation matrix based on the image, and performing transformation processing on the two-dimensional detection box through the transformation matrix to obtain at least one surface of the three-dimensional detection box of the target object.
[0242] Optionally, the storage medium is further set to store program code for performing the following steps: determining that there is a preset position in the image; obtaining a transformation matrix associated with the preset position, where the transformation matrix is used to transform the two-dimensional detection box of the object corresponding to any image with a preset position into at least one surface of the three-dimensional detection box of the object.
[0243] Optionally, the storage medium is further set to store program code for performing the following steps: when there is no transformation matrix associated with the preset position, performing three-dimensional detection on the image to obtain at least one surface of the three-dimensional detection box of the target object.
[0244] Optionally, the storage medium is further set to store program code for performing the following steps: determining a transformation matrix based on at least one surface of the three-dimensional detection box of the target object and the two-dimensional detection box of the target object.
[0245] Optionally, the storage medium is further configured to store program code for performing the following steps: one-to-one match at least one surface of the two-dimensional detection box of the target object and the three-dimensional detection box of the target object to obtain a first matching result; perform at least linear regression calculation on the first matching result to obtain a linear regression model; determine the weight of the linear regression model as the transformation matrix.
[0246] Optionally, the storage medium is further configured to store program code for performing the following steps: obtain a parallelogram corresponding to at least one surface of the three-dimensional detection box of the target object; perform one-to-one match between the parallelogram and the two-dimensional detection box of the target object to obtain a first matching result.
[0247] Optionally, the storage medium is further configured to store program code for performing the following steps: obtain at least one second matching result stored in the memory, where each second matching result is obtained by one-to-one matching at least one surface of the two-dimensional detection box and the corresponding three-dimensional detection box of the object corresponding to each historical image with a preset bit, and each historical image is generated before the image; perform linear regression calculation on the first matching result and the at least one second matching result to obtain a linear regression model.
[0248] Optionally, the storage medium is further configured to store program code for performing the following steps: after performing at least linear regression calculation on the first matching result to obtain a linear regression model, without changing the preset bit, obtain at least one third matching result, where each third matching result is obtained by matching at least one surface of the two-dimensional detection box and the corresponding three-dimensional detection box of the object corresponding to each new image with a preset bit, and each new image is generated after the image; update the linear regression model through the at least one third matching result.
[0249] Optionally, the storage medium is further configured to store program code for performing the following steps: obtain the pseudo-inverse matrix of the updated linear regression model; when the change between the pseudo-inverse matrix of the updated linear regression model and the pseudo-inverse matrix of the linear regression model before update is within the target threshold, determine the pseudo-inverse matrix of the updated linear regression model as the transformation matrix.
[0250] Optionally, the storage medium is further configured to store program code for performing the following steps: perform linear transformation on the two-dimensional detection box of the target object through the transformation matrix to obtain at least one surface of the three-dimensional detection box of the target object.
[0251] Optionally, the storage medium is further configured to store program code for performing the following steps: performing three-dimensional detection on an image through a three-dimensional detection model to obtain at least one surface of a three-dimensional detection frame of a target object, where the three-dimensional detection model is trained through pre-collected image samples and at least one surface of the corresponding three-dimensional detection frame.
[0252] Optionally, the storage medium is further configured to store program code for performing the following steps: obtaining a current video frame from a video stream of a target plane.
[0253] Optionally, the storage medium is further configured to store program code for performing the following steps: performing two-dimensional detection on the current video frame to obtain a two-dimensional detection frame of a target vehicle.
[0254] Optionally, the storage medium is further configured to store program code for performing the following steps: obtaining a transformation matrix based on the current video frame, and performing transformation processing on the two-dimensional detection frame of the target vehicle through the transformation matrix to obtain the bottom surface of the three-dimensional detection frame of the target vehicle.
[0255] Optionally, the storage medium is further configured to store program code for performing the following steps: determining projection information of the target vehicle on a traffic road through the bottom surface of the three-dimensional detection frame of the target vehicle.
[0256] Optionally, in this embodiment, the storage medium is configured to store program code for performing the following steps: displaying an image including the target object on a target interface, and displaying a two-dimensional detection frame of the target object; displaying at least one surface of the three-dimensional detection frame of the target object on the target interface, where at least one surface of the three-dimensional detection frame of the target object is obtained based on the two-dimensional detection frame; displaying projection information of the target object on the target plane on the target interface, where the projection information is determined through at least one surface of the three-dimensional detection frame of the target object.
[0257] Optionally, in this embodiment, the storage medium is configured to store program code for performing the following steps: obtaining an image including the target object; determining that there is a preset position in the image; in the case where no transformation matrix is associated with the preset position, performing three-dimensional detection on the image to obtain at least one surface of the three-dimensional detection frame of the target object, where the transformation matrix is used to transform the two-dimensional detection frame of the object corresponding to any image with a preset position into at least one surface of the three-dimensional detection frame of the object; determining projection information of the target object on the target plane through at least one surface of the three-dimensional detection frame of the target object.
[0258] Optionally, the storage medium is further configured to store program code for performing the following steps: after obtaining an image of a target object on a target plane, performing two-dimensional detection on the image to obtain a two-dimensional detection box of the target object; when a conversion matrix is associated with a preset position, performing conversion processing on the two-dimensional detection box of the target object through the conversion matrix to obtain at least one surface of a three-dimensional detection box of the target object.
[0259] Optionally, in this embodiment, the storage medium is configured to store program code for performing the following steps: obtaining a target request, where the target request carries an image to be processed input on a target interface, and the image includes a target object; in response to the target request, obtaining a two-dimensional detection box of the target object; based on the two-dimensional detection box, obtaining at least one surface of a three-dimensional detection box of the target object; determining projection information of the target object on the target plane through at least one surface of the three-dimensional detection box of the target object, and sending the projection information to the target interface for display.
[0260] Embodiment 6
[0261] An embodiment of the present invention further provides a processor. Optionally, in this embodiment, the above-mentioned processor runs the program of the data processing method provided in the above-mentioned Embodiment 1.
[0262] Optionally, in this embodiment, the processor is configured to run program code for performing the following steps: obtaining an image including a target object; obtaining a two-dimensional detection box of the target object; based on the two-dimensional detection box, obtaining at least one surface of a three-dimensional detection box of the target object; determining projection information of the target object on the target plane through at least one surface of the three-dimensional detection box of the target object.
[0263] Optionally, in this embodiment, the processor is configured to run program code for performing the following steps: displaying an image including a target object on the target interface, and displaying a two-dimensional detection box of the target object; displaying at least one surface of a three-dimensional detection box of the target object on the target interface, where at least one surface of the three-dimensional detection box of the target object is obtained based on the two-dimensional detection box; displaying projection information of the target object on the target plane on the target interface, where the projection information is determined through at least one surface of the three-dimensional detection box of the target object.
[0264] Optionally, in this embodiment, the processor is configured to run program code for the following steps: obtaining an image containing a target object; determining that there is a preset position in the image; in the case where no transformation matrix is associated with the preset position, performing three-dimensional detection on the image to obtain at least one surface of the three-dimensional detection frame of the target object, where the transformation matrix is used to transform the two-dimensional detection frame of the object corresponding to any image with a preset position into at least one surface of the three-dimensional detection frame of the object; determining projection information of the target object on the target plane through at least one surface of the three-dimensional detection frame of the target object.
[0265] Optionally, in this embodiment, the processor is configured to run program code for the following steps: obtaining a target request, where the target request carries an image to be processed input on the target interface, and the image contains a target object; in response to the target request, obtaining a two-dimensional detection frame of the target object; based on the two-dimensional detection frame, obtaining at least one surface of the three-dimensional detection frame of the target object; determining projection information of the target object on the target plane through at least one surface of the three-dimensional detection frame of the target object, and sending the projection information to the target interface for display.
[0266] The serial numbers of the above embodiments of the present invention are only for description and do not represent the superiority or inferiority of the embodiments.
[0267] In the above embodiments of the present invention, the descriptions of the various embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0268] In the several embodiments provided by the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection can be through some interfaces, and the indirect coupling or communication connection of units or modules can be in an electrical or other form.
[0269] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0270] In addition, in each embodiment of the present invention, each functional unit can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0271] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks, or optical discs.
[0272] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. A data processing method, characterized in that, comprising: obtaining an image including a target object; obtaining a two-dimensional detection frame of the target object; based on the two-dimensional detection frame, obtaining at least one surface of a three-dimensional detection frame of the target object; determining projection information of the target object on a target plane through at least one surface of the three-dimensional detection frame of the target object; wherein, based on the two-dimensional detection frame, obtaining at least one surface of the three-dimensional detection frame of the target object includes: determining whether a preset position of the image is associated with a transformation matrix, wherein the transformation matrix is obtained by solving a pseudo-inverse matrix of a linear regression model, and the linear regression model is obtained by performing one-to-one matching on the two-dimensional detection frame of the target object and at least one surface of the three-dimensional detection frame of the target object, and performing linear regression calculation on at least the obtained first matching result; in the case where the preset position is associated with the transformation matrix, performing transformation processing on the two-dimensional detection frame to obtain at least one surface of the three-dimensional detection frame of the target object.
2. The method according to claim 1, characterized in that, in the case where the preset position is associated with the transformation matrix, performing transformation processing on the two-dimensional detection frame to obtain at least one surface of the three-dimensional detection frame of the target object, including: performing transformation processing on the two-dimensional detection frame through the transformation matrix to obtain at least one surface of the three-dimensional detection frame of the target object.
3. The method according to claim 2, characterized in that, obtaining the transformation matrix based on the image includes: determining that the preset position exists in the image; obtaining the transformation matrix associated with the preset position, wherein the transformation matrix is used to transform a two-dimensional detection frame of an object corresponding to any image in which the preset position exists into at least one surface of a three-dimensional detection frame of the object.
4. The method according to claim 3, characterized in that, the method further includes: in the case where the preset position is not associated with the transformation matrix, performing three-dimensional detection on the image to obtain at least one surface of the three-dimensional detection frame of the target object.
5. The method according to claim 1, characterized in that, the method further includes: determining weights of the linear regression model based on the pseudo-inverse matrix and the bottom surface of the three-dimensional detection frame; determining the weights of the linear regression model as the transformation matrix.
6. The method according to claim 5, characterized in that, performing one-to-one matching on the two-dimensional detection frame of the target object and at least one surface of the three-dimensional detection frame of the target object to obtain a first matching result, including: obtaining a parallelogram corresponding to at least one surface of the three-dimensional detection frame of the target object; performing one-to-one matching on the parallelogram and the two-dimensional detection frame of the target object to obtain the first matching result.
7. The method according to claim 5, characterized in that, edges of at least one surface of the three-dimensional detection frame are parallel to a horizontal line.
8. The method according to claim 5, characterized in that, The method further includes: obtaining at least one second matching result stored in the memory, where each of the second matching results is obtained by one-to-one matching of at least one surface of the two-dimensional detection box and the corresponding three-dimensional detection box of the object corresponding to each historical image having the preset bit, and each of the historical images is generated before the image; Performing at least linear regression calculation on the first matching result to obtain a linear regression model, including: performing linear regression calculation on the first matching result and at least one of the second matching results to obtain the linear regression model.
9. The method according to claim 5, wherein, after performing at least linear regression calculation on the first matching result to obtain a linear regression model, the method further includes: when the preset bit remains unchanged, obtaining at least one third matching result, where each of the third matching results is obtained by matching at least one surface of the two-dimensional detection box and the corresponding three-dimensional detection box of the object corresponding to each new image having the preset bit, and each of the new images is generated after the image; updating the linear regression model through the at least one third matching result.
10. The method according to claim 9, wherein, the method further includes: obtaining the weight of the updated linear regression model; determining the weight of the linear regression model as the transformation matrix, including: when the change between the weight of the updated linear regression model and the weight of the linear regression model before update is within the target threshold, determining the weight of the updated linear regression model as the transformation matrix.
11. The method according to claim 5, wherein, performing transformation processing on the two-dimensional detection box of the target object through the transformation matrix to obtain at least one surface of the three-dimensional detection box of the target object, including: performing a linear transformation on the two-dimensional detection box of the target object through the transformation matrix to obtain at least one surface of the three-dimensional detection box of the target object.
12. The method according to claim 4, wherein, performing three-dimensional detection on the image to obtain at least one surface of the three-dimensional detection box of the target object, including: performing three-dimensional detection on the image through a three-dimensional detection model to obtain at least one surface of the three-dimensional detection box of the target object, where the three-dimensional detection model is trained through pre-collected image samples and at least one surface of the corresponding three-dimensional detection box.
13. The method according to any one of claims 1 to 12, wherein, the target object is a target vehicle, and the image is the current video frame.
14. The method according to claim 13, wherein, obtaining an image containing the target object, including: obtaining the current video frame from the video stream of the target plane.
15. The method according to claim 13, wherein, obtaining the two-dimensional detection box of the target object, including: performing two-dimensional detection on the current video frame to obtain the two-dimensional detection box of the target vehicle.
16. The method according to claim 13, wherein, obtaining at least one surface of the three-dimensional detection frame of the target object based on the two-dimensional detection frame includes: obtaining a transformation matrix based on the current video frame, and performing a transformation process on the two-dimensional detection frame of the target vehicle through the transformation matrix to obtain the bottom surface of the three-dimensional detection frame of the target vehicle.
17. The method according to claim 16, wherein, determining the projection information of the target object on the target plane through at least one surface of the three-dimensional detection frame of the target object includes: determining the projection information of the target vehicle on the traffic road through the bottom surface of the three-dimensional detection frame of the target vehicle.
18. A data processing method, wherein, comprises: displaying an image including a target object on a target interface, and displaying a two-dimensional detection frame of the target object; displaying at least one surface of the three-dimensional detection frame of the target object on the target interface, wherein at least one surface of the three-dimensional detection frame of the target object is obtained by performing a transformation process on the two-dimensional detection frame in the case where a transformation matrix is associated with a preset position of the image, the transformation matrix is obtained by solving the pseudo-inverse matrix of a linear regression model, and the linear regression model is obtained by performing one-to-one matching on the two-dimensional detection frame of the target object and at least one surface of the three-dimensional detection frame of the target object, and performing at least linear regression calculation on the obtained first matching result; displaying the projection information of the target object on the target plane on the target interface, wherein the projection information is determined through at least one surface of the three-dimensional detection frame of the target object.
19. A data processing method, wherein, comprises: obtaining an image including a target object; determining that there is a preset position in the image; determining whether the preset position is associated with a transformation matrix, wherein the transformation matrix is obtained by solving the pseudo-inverse matrix of a linear regression model, and the linear regression model is obtained by performing one-to-one matching on the two-dimensional detection frame of the target object and at least one surface of the three-dimensional detection frame of the target object, and performing at least linear regression calculation on the obtained first matching result; in the case where the preset position is not associated with the transformation matrix, performing three-dimensional detection on the image to obtain at least one surface of the three-dimensional detection frame of the target object, wherein the transformation matrix is used to transform the two-dimensional detection frame of the object corresponding to any image with the preset position into at least one surface of the three-dimensional detection frame of the object; in the case where the preset position is associated with the transformation matrix, performing a transformation process on the two-dimensional detection frame to obtain at least one surface of the three-dimensional detection frame of the target object; determining the projection information of the target object on the target plane through at least one surface of the three-dimensional detection frame of the target object.
20. The method according to claim 19, wherein, When the conversion matrix is associated with the preset position, performing a conversion process on the two-dimensional detection frame to obtain at least one surface of the three-dimensional detection frame of the target object, including: Performing two-dimensional detection on the image to obtain a two-dimensional detection frame of the target object; When the conversion matrix is associated with the preset position, performing a conversion process on the two-dimensional detection frame of the target object through the conversion matrix to obtain at least one surface of the three-dimensional detection frame of the target object.
21. A data processing method, characterized in that, it includes: Obtaining a target request, where the target request carries an image to be processed input on a target interface, and the image includes a target object; Responding to the target request to obtain a two-dimensional detection frame of the target object; Based on the two-dimensional detection frame, obtaining at least one surface of the three-dimensional detection frame of the target object; Determining projection information of the target object on a target plane through at least one surface of the three-dimensional detection frame of the target object, and sending the projection information to the target interface for display; Among them, based on the two-dimensional detection frame, obtaining at least one surface of the three-dimensional detection frame of the target object includes: determining whether a conversion matrix is associated with the preset position of the image, where the conversion matrix is obtained by solving the pseudo-inverse matrix of a linear regression model, and the linear regression model is obtained by performing one-to-one matching on the two-dimensional detection frame of the target object and at least one surface of the three-dimensional detection frame of the target object, and performing at least linear regression calculation on the obtained first matching result; when the conversion matrix is associated with the preset position, performing a conversion process on the two-dimensional detection frame to obtain at least one surface of the three-dimensional detection frame of the target object.
22. A data processing device, characterized in that, it includes: A first acquisition unit for acquiring an image including a target object; A second acquisition unit for acquiring a two-dimensional detection frame of the target object; A third acquisition unit for obtaining at least one surface of the three-dimensional detection frame of the target object based on the two-dimensional detection frame; A first determination unit for determining projection information of the target object on a target plane through at least one surface of the three-dimensional detection frame of the target object; Among them, the third acquisition unit is used to perform the following steps to obtain at least one surface of the three-dimensional detection frame of the target object based on the two-dimensional detection frame, including: determining whether a conversion matrix is associated with the preset position of the image, where the conversion matrix is obtained by solving the pseudo-inverse matrix of a linear regression model, and the linear regression model is obtained by performing one-to-one matching on the two-dimensional detection frame of the target object and at least one surface of the three-dimensional detection frame of the target object, and performing at least linear regression calculation on the obtained first matching result; when the conversion matrix is associated with the preset position, performing a conversion process on the two-dimensional detection frame to obtain at least one surface of the three-dimensional detection frame of the target object.
23. A data processing device, characterized in that, it includes: The first display unit is configured to display an image including a target object on a target interface and display a two-dimensional detection frame of the target object; The second display unit is configured to display at least one surface of a three-dimensional detection frame of the target object on the target interface, where at least one surface of the three-dimensional detection frame of the target object is obtained by performing a conversion process on the two-dimensional detection frame when it is determined that a preset position in the image is associated with a conversion matrix, the conversion matrix is obtained by solving a pseudo-inverse matrix of a linear regression model, and the linear regression model is obtained by performing a one-to-one matching on the two-dimensional detection frame of the target object and at least one surface of the three-dimensional detection frame of the target object and performing at least a linear regression calculation on the obtained first matching result; The third display unit is configured to display projection information of the target object on a target plane on the target interface, where the projection information is determined by at least one surface of the three-dimensional detection frame of the target object.
24. A data processing device Characterized in that It includes: A fourth acquisition unit for acquiring an image including a target object; A second determination unit for determining that there is a preset position in the image; A detection unit for performing three-dimensional detection on the image to obtain at least one surface of a three-dimensional detection frame of the target object when the preset position is not associated with a conversion matrix, where the conversion matrix is used to convert a two-dimensional detection frame of an object corresponding to any image with the preset position into at least one surface of a three-dimensional detection frame of the object; A third determination unit for determining projection information of the target object on a target plane through at least one surface of the three-dimensional detection frame of the target object; Wherein, the device is further configured to perform the following steps: determining whether the preset position is associated with a conversion matrix, where the conversion matrix is obtained by solving a pseudo-inverse matrix of a linear regression model, and the linear regression model is obtained by performing a one-to-one matching on the two-dimensional detection frame of the target object and at least one surface of the three-dimensional detection frame of the target object and performing at least a linear regression calculation on the obtained first matching result; and performing a conversion process on the two-dimensional detection frame to obtain at least one surface of a three-dimensional detection frame of the target object when the preset position is associated with the conversion matrix.
25. A storage medium Characterized in that The storage medium includes a stored program, where when the program runs, it controls the device where the storage medium is located to perform the following steps: Acquiring an image including a target object; Acquiring a two-dimensional detection frame of the target object; Based on the two-dimensional detection frame, acquiring at least one surface of a three-dimensional detection frame of the target object; Determining projection information of the target object on a target plane through at least one surface of the three-dimensional detection frame of the target object; Among them, obtaining at least one surface of the three-dimensional detection frame of the target object based on the two-dimensional detection frame includes: determining whether a preset position of the image is associated with a transformation matrix, where the transformation matrix is obtained by solving the pseudo-inverse matrix of a linear regression model, and the linear regression model is obtained by performing one-to-one matching on the two-dimensional detection frame of the target object and at least one surface of the three-dimensional detection frame of the target object, and performing linear regression calculation on at least the obtained first matching result; in the case where the preset position is associated with the transformation matrix, performing transformation processing on the two-dimensional detection frame to obtain at least one surface of the three-dimensional detection frame of the target object.
26. A processor, characterized in that, the processor is used to run a program, and when the program runs, the following steps are executed: obtaining an image including a target object; obtaining a two-dimensional detection frame of the target object; obtaining at least one surface of the three-dimensional detection frame of the target object based on the two-dimensional detection frame; determining projection information of the target object on a target plane through at least one surface of the three-dimensional detection frame of the target object; Among them, obtaining at least one surface of the three-dimensional detection frame of the target object based on the two-dimensional detection frame includes: determining whether a preset position of the image is associated with a transformation matrix, where the transformation matrix is obtained by solving the pseudo-inverse matrix of a linear regression model, and the linear regression model is obtained by performing one-to-one matching on the two-dimensional detection frame of the target object and at least one surface of the three-dimensional detection frame of the target object, and performing linear regression calculation on at least the obtained first matching result; in the case where the preset position is associated with the transformation matrix, performing transformation processing on the two-dimensional detection frame to obtain at least one surface of the three-dimensional detection frame of the target object.
27. A mobile terminal, characterized in that, it includes: a processor; a transmission device for transmitting an image including a target object; and a memory connected to the transmission device for providing instructions for the processor to perform the following processing steps: obtaining a two-dimensional detection frame of the target object; obtaining at least one surface of the three-dimensional detection frame of the target object based on the two-dimensional detection frame; determining projection information of the target object on a target plane through at least one surface of the three-dimensional detection frame of the target object; Among them, obtaining at least one surface of the three-dimensional detection frame of the target object based on the two-dimensional detection frame includes: determining whether a preset position of the image is associated with a transformation matrix, where the transformation matrix is obtained by solving the pseudo-inverse matrix of a linear regression model, and the linear regression model is obtained by performing one-to-one matching on the two-dimensional detection frame of the target object and at least one surface of the three-dimensional detection frame of the target object, and performing linear regression calculation on at least the obtained first matching result; in the case where the preset position is associated with the transformation matrix, performing transformation processing on the two-dimensional detection frame to obtain at least one surface of the three-dimensional detection frame of the target object.
Citation Information
Patent Citations
Object three-dimensional position detection method and device based on a depth fitting degree evaluation network
CN109872366A
Object three-dimensional detection and intelligent driving control method and device, medium and equipment
CN110826357A