A 3D video-based target detection method, system, device and medium

Through a 3D video-based target detection method, the length, width, and height information of the vehicle is accurately obtained using the 3D detection frame and the longitude and latitude coordinates of the acquisition device, which solves the problem of insufficient detection accuracy in existing technologies and achieves higher data credibility and traffic safety.

CN119229088BActive Publication Date: 2025-09-19BEIJING SINOITS TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411188492.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-28
Publication Date
2025-09-19
Estimated Expiration
2044-08-28

AI Technical Summary

Technical Problem

Existing millimeter-wave radar and 2D video target detection algorithms cannot accurately obtain the length, width, and height dimensions of a vehicle, and there are deviations in geographic location mapping, especially when detecting trucks.

Method used

A target detection method based on 3D video is adopted. The latitude and longitude coordinates of the reference vertex of the target detection object are determined through the 3D detection frame. The height information of the target detection object is obtained by combining the latitude and longitude coordinates of the acquisition device. The accurate reference fixed point is obtained through coordinate conversion to realize the determination of the basic attribute information of the target detection object.

Benefits of technology

The accuracy of target detection is improved, detailed feature information of the vehicle is obtained, the credibility of the data is improved, the problem of the inability to obtain the actual height in the existing technology is solved, and the accuracy of traffic events and the safety of autonomous driving are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119229088B_ABST
    Figure CN119229088B_ABST
Patent Text Reader

Abstract

The present invention discloses a 3D video-based target detection method, system, device, and medium, relating to the field of image processing technology. The method includes: determining a 3D detection frame corresponding to a target detection object within the to-be-detected area based on any frame of image data in the area to be detected; determining the latitude and longitude coordinates of a reference vertex corresponding to any target detection object based on the 3D detection frame, and determining basic attribute information of the target detection object in combination with the latitude and longitude coordinates of an acquisition device, wherein the basic attribute information includes the height information of the target detection object. In the process of determining the basic attribute information of the target detection object, this solution incorporates the latitude and longitude coordinates of the acquisition device to ensure that the determination result of the basic attribute information of the target detection object is more consistent with the actual situation and the data is more credible.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing technology, and in particular to a 3D video-based target detection method, system, device and medium. Background Art

[0002] With the launch of intelligent connected vehicle-road-cloud integrated urban pilot projects, traditional millimeter-wave radars and cameras, as the primary road test hardware, have the advantage of relatively low cost and are widely used. Typically, millimeter-wave radars are unable to detect the specific length, width, and height scale information of a target, and 2D video target detection algorithms are also unable to obtain the target's scale information. This paper proposes a method based on 3D video target detection. Through certain algorithmic calculations, it can accurately obtain the vehicle's length, width, and height dimensions in real time. In addition, the pixel coordinates after 3D target detection and the geographic location coordinate information mapped by the algorithm are more accurate than the target geographic information obtained by millimeter-wave radar algorithms and 2D video target detection algorithms. This can then be used in digital twin displays or reported to cloud servers as traffic event information, allowing terminals or monitoring platforms to obtain accurate vehicle dynamic information in real time, providing safer protection for road traffic and promoting the increasingly mature development of autonomous driving technology.

[0003] By detecting targets in the video image, the vehicle's size information can be obtained. Existing technologies usually use geodetic meter coordinates as a benchmark, measure the meter coordinates of ground marking lines, and associate the pixel points on the vehicle edge with the meter coordinate points of the marking lines to calculate the length and width of the vehicle target. Obtaining vehicle height information is more difficult. In addition, when calculating the vehicle target's geographic location information, the center pixel coordinates or the bottom center pixel coordinates of the 2D target detection frame are usually used as a benchmark. There is a certain deviation between the geographic coordinates mapped by the algorithm and the actual geographic coordinates of the vehicle, especially when detecting trucks, the deviation is relatively large. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to address the deficiencies of the existing technology and specifically provide a 3D video-based target detection method, system, device and medium, as follows:

[0005] 1) In a first aspect, the present invention provides a method for object detection based on 3D video, the specific technical solution of which is as follows:

[0006] Determine a 3D detection frame corresponding to a target detection object within the area to be detected based on any frame image data of the area to be detected;

[0007] Based on the 3D detection frame, the latitude and longitude coordinates of the reference vertex corresponding to any target detection object are determined, and the basic attribute information of the target detection object is determined in combination with the latitude and longitude coordinates of the acquisition device, where the basic attribute information includes the height information of the target detection object.

[0008] The beneficial effects of the 3D video-based object detection method provided by the present invention are as follows:

[0009] This solution uses a 3D detection frame to determine the latitude and longitude of subsequent benchmark points, which is more accurate than the 2D detection frame method and can reflect the characteristics of the target detection object in more detail. In addition, this solution combines the latitude and longitude coordinates of the acquisition device in the process of determining the basic attribute information of the target detection object, so that the determination results of the basic attribute information of the target detection object are more consistent with the actual situation and the data is more credible. In addition, the basic attribute information also includes height information, which solves the problem that the current video target detection cannot obtain the actual height.

[0010] Based on the above solution, the present invention can also be improved as follows.

[0011] Furthermore, the process of determining the latitude and longitude coordinates of the reference vertex corresponding to any target detection object is as follows:

[0012] Determine the pixel coordinates of a reference vertex corresponding to any target detection object based on the 3D detection frame;

[0013] According to the coordinate conversion principle, the pixel coordinates of any reference vertex are converted into latitude and longitude coordinates.

[0014] Furthermore, the process of determining the basic attribute information of the target detection object is as follows:

[0015] Based on the correspondence formula of the longitude and latitude coordinates of all reference vertices, the longitude and latitude coordinates of the acquisition device and the center coordinates of the target detection object, the longitude and latitude coordinates of the reference vertices and the target center coordinates corresponding to the longitude and latitude coordinates of the acquisition device are determined, and the basic attribute information is determined according to the target center coordinates.

[0016] Furthermore, it also includes:

[0017] The basic attribute information is sent to the cloud server for storage.

[0018] 2) In a second aspect, the present invention further provides a 3D video-based object detection system, the specific technical solution of which is as follows:

[0019] Determine a 3D detection frame corresponding to a target detection object within the area to be detected based on any frame image data of the area to be detected;

[0020] Based on the 3D detection frame, the latitude and longitude coordinates of the reference vertex corresponding to any target detection object are determined, and the basic attribute information of the target detection object is determined in combination with the latitude and longitude coordinates of the acquisition device, where the basic attribute information includes the height information of the target detection object.

[0021] Based on the above solution, the present invention can also be improved as follows.

[0022] Furthermore, the process of determining the latitude and longitude coordinates of the reference vertex corresponding to any target detection object is as follows:

[0023] Determine the pixel coordinates of a reference vertex corresponding to any target detection object based on the 3D detection frame;

[0024] According to the coordinate conversion principle, the pixel coordinates of any reference vertex are converted into latitude and longitude coordinates.

[0025] Furthermore, the process of determining the basic attribute information of the target detection object is as follows:

[0026] Based on the correspondence formula of the longitude and latitude coordinates of all reference vertices, the longitude and latitude coordinates of the acquisition device and the center coordinates of the target detection object, the longitude and latitude coordinates of the reference vertices and the target center coordinates corresponding to the longitude and latitude coordinates of the acquisition device are determined, and the basic attribute information is determined according to the target center coordinates.

[0027] Furthermore, it also includes:

[0028] The basic attribute information is sent to the cloud server for storage.

[0029] 3) In a third aspect, the present invention further provides an electronic device, comprising a processor, wherein the processor is coupled to a memory, wherein at least one computer program is stored in the memory, and the at least one computer program is loaded and executed by the processor so that the electronic device implements any of the above methods.

[0030] 4) In a fourth aspect, the present invention further provides a computer-readable storage medium, wherein the computer-readable storage medium stores at least one computer program, and the at least one computer program is loaded and executed by a processor to enable a computer to implement any of the above methods.

[0031] It should be noted that the beneficial effects achieved by the technical solutions of the second to fourth aspects of the present invention and the corresponding possible implementation methods can be found in the above-mentioned technical effects of the first aspect and its corresponding possible implementation methods, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Other features, objects and advantages of the present invention will become more apparent upon reading the detailed description of non-limiting embodiments made with reference to the following drawings:

[0033] Figure 1 This is a flow chart of a method for object detection based on 3D video according to an embodiment of the present invention;

[0034] Figure 2 This is a second flow chart of a method for object detection based on 3D video according to an embodiment of the present invention;

[0035] Figure 3 This is a structural framework diagram of an electronic device in this solution;

[0036] Figure 4 Schematic diagram of reference vertices of a 3D video-based object detection method according to an embodiment of the present invention. DETAILED DESCRIPTION

[0037] To make the objectives, technical solutions and advantages of the present invention more clear, the embodiments of the present invention will be described in further detail below with reference to the accompanying drawings.

[0038] like Figure 1 as well as Figure 2 As shown, a 3D video-based target detection method according to an embodiment of the present invention includes the following steps:

[0039] S1, determining a 3D detection frame corresponding to a target detection object in the area to be detected based on any frame image data of the area to be detected;

[0040] S2, based on the 3D detection frame, determine the latitude and longitude coordinates of the reference vertex corresponding to any target detection object, and determine the basic attribute information of the target detection object in combination with the latitude and longitude coordinates of the acquisition device, wherein the basic attribute information includes the height information of the target detection object.

[0041] The beneficial effects of the 3D video-based object detection method provided by the present invention are as follows:

[0042] This solution uses a 3D detection frame to determine the latitude and longitude of subsequent benchmark points, which is more accurate than the 2D detection frame method and can reflect the characteristics of the target detection object in more detail. In addition, this solution combines the latitude and longitude coordinates of the acquisition device in the process of determining the basic attribute information of the target detection object, so that the determination results of the basic attribute information of the target detection object are more consistent with the actual situation and the data is more credible. In addition, the basic attribute information also includes height information, which solves the problem that the current video target detection cannot obtain the actual height.

[0043] S1, based on any frame image data of the area to be detected, determine the 3D detection frame corresponding to the target detection object in the area to be detected.

[0044] The area to be detected refers to the road area that can be collected by the collection equipment and allows vehicles and pedestrians to pass normally. Depending on the location of the collection equipment, the corresponding location of the area to be detected is also different, including but not limited to: intersections, highway intersections, etc.

[0045] The image data is obtained as follows:

[0046] By decoding the video stream captured by the acquisition device, each frame of image data is obtained. The decoding process includes:

[0047] The decoder decodes the compressed video stream that complies with a specific coding standard to obtain image data.

[0048] The process of determining the 3D detection frame corresponding to the target detection object according to the image data is as follows: wherein the target detection object is a vehicle.

[0049] By presetting the vehicle detection model, the image data is processed to obtain the initial 3D detection frame of all target detection objects corresponding to any image data;

[0050] All initial 3D detection frames on any image data are re-checked. The initial 3D detection frames corresponding to non-motor vehicles are first eliminated, and then the initial 3D detection frames with obvious detection errors are adjusted to obtain the 3D detection frames.

[0051] An initial 3D detection frame with obvious detection errors refers to: determining the actual length of the vehicle and comparing it with the length of the corresponding detection edge in the initial 3D detection frame. If the difference between the detection edge length and the actual length exceeds the error range, the initial 3D detection frame corresponding to the detection edge is determined as the initial 3D detection frame with obvious detection errors. At this time, feature extraction is performed on the area corresponding to the initial 3D detection frame to obtain an image that only contains the area corresponding to the initial 3D detection frame, or partial highlighting is performed on the image data corresponding to the initial 3D detection frame to obtain a new image, so that the features of the initial 3D detection frame with obvious detection errors are more prominent. The processed image data is re-input into the preset vehicle detection model for detection, and the results are verified again. If there are still initial 3D detection frames with obvious detection errors, the 3D detection frame is re-determined through manual intervention.

[0052] The preset vehicle detection model is trained as follows:

[0053] After constructing the training sample set, the sample images from the training sample set are input into the vehicle detection model and passed through the input layer to the hidden layer for feature extraction. The sample images are then processed through the pooling layer for dimensionality reduction. The reduced sample images are then input into three convolutional layers for vehicle feature extraction. The extracted results are then integrated and classified through the fully connected layer and passed to the output layer for output. The entire classification result is compared with the label category corresponding to each image in the validation set. Combined with the loss function, it is determined whether the current vehicle detection model has achieved the preset accuracy. If so, the vehicle model is output.

[0054] During the training process, a feature extraction layer can be added after any convolution layer. The feature extraction layer is not connected to the fully connected layer. The feature extraction layer directly extracts the features of the image processed by a certain convolution layer according to preset rules, and analyzes the extracted features (this analysis process can be done manually or by comparing features with labels, such as performing semantic analysis on labels, extracting keywords, and determining the similarity between features and keywords. When the similarity is higher than the preset similarity, it is determined that the requirements are met or that the feature extraction is correct). It is determined whether the processing of the current convolution layer is accurate, and if the processing result is inaccurate, the convolution layer of the model is replaced or adjusted to make the convolution layer more accurate in extracting features.

[0055] S2, based on the 3D detection frame, determine the latitude and longitude coordinates of the reference vertex corresponding to any target detection object, and combine the latitude and longitude coordinates of the acquisition device to determine the basic attribute information of the target detection object, wherein the basic attribute information includes the height information of the target detection object.

[0056] There are four reference vertices. Specifically, the central axis of the 2D image corresponding to any frame of image data (the axis parallel to the 2D image boundary line that divides the entire 2D image into two equal parts along the shooting direction of the acquisition device) is used as the reference, and the 2D image is divided into a left half and a right half. In the 2D image of the left half, the 3D detection frame corresponding to any target detection object is composed of six 2D detection frames, where the six 2D detection frames are three groups of two-by-two parallel detection frames. The 2D detection frame that is parallel to the ground and closer to the ground is defined as the reference 2D detection frame. Within the reference 2D detection frame, the vertex with the smallest distance from the acquisition device among the two vertices closer to the central axis is determined and defined as the reference point. The reference point, the two vertices adjacent to the reference point, and the vertex perpendicular to the reference point in the 2D detection frame that is parallel to the ground and farther away from the ground (also called the target point) are defined as the reference vertices. Wherein, the vertical correspondence means that the line connecting the target point and the reference point is perpendicular to the ground.

[0057] For the convenience of subsequent explanation, Figure 4As shown in the figure, the four reference vertices are represented by a, b, c, and d, where d is the target point and b is the reference point. f1 is the installation location of the acquisition device, and f2 is the point perpendicular to the ground where the acquisition device is located. The line connecting f1 and point d in the figure is line 1, and the line connecting f2 and point b is line 2. When the vertical error is small, lines 1 and 2 should be approximately equal to ground point e. The acquisition device is installed at a geographic location, Pos0 (lon0, lat0), and a height of h meters above the ground (i.e., the length of the sides of f1 and f2 is h). The pixels in each frame are mapped using an algorithm to obtain their corresponding geographic longitude and latitude coordinates. The geographic longitude and latitude coordinates of the reference vertices a, b, c, d, and e are: reference vertex a: Possa (aLon, aLot), reference vertex b: Posb (bLon, bLot), reference vertex c: Posc (cLon, cLot), reference vertex d: Posd (dLon, dLot), and reference vertex e: Pose (eLon, eLot). The center coordinates of the target object are centerPos (centerLon, centerLot), where centerLon = (aLon + cLon) / 2; centerLat = (aLat + cLat) / 2. The center point coordinates are the midpoint of the line connecting the two vertices adjacent to the reference point.

[0058] Since the pixel coordinates and longitude and latitude algorithm mapping in this description are based on the geodetic plane, from the perspective of the camera (acquisition device), point d is projected to point e on the ground through the camera, and the pixel coordinates of point d reach point e on the geodetic plane after algorithm mapping, and then dLon == eLon, dLot == eLot; the distance (posA, posB) function is a function for calculating the distance in meters between the longitude and latitude of two points, based on the theorem of parallel line properties of one side of a triangle.

[0059] The target height high = h × distance (Pose, Posb) / distance (Pose, Pos0). After taking the average of the heights over multiple frames, the obtained height is more accurate.

[0060] By obtaining the precise location and size information of the target, the judgment logic of traffic events is simplified. For example, events such as trucks exceeding the height limit, trucks not driving in the designated lane, and motor vehicles changing lanes continuously can be detected. This can greatly improve the accuracy of event reporting and solve the defects of 2D target detection.

[0061] According to the protocol, the algorithm sends the geographic coordinate information and size information of the vehicle target, as well as the event messages after logical processing, to the cloud server. The terminal or monitoring platform accurately displays the twin in real time, providing safer protection for pedestrians or vehicles on the surrounding roads. At the same time, it can make predictions in advance and plan and issue reasonable driving routes for autonomous vehicles.

[0062] Furthermore, the process of determining the latitude and longitude coordinates of the reference vertex corresponding to any target detection object is as follows:

[0063] Determine the pixel coordinates of a reference vertex corresponding to any target detection object based on the 3D detection frame;

[0064] According to the coordinate conversion principle, the pixel coordinates of any reference vertex are converted into latitude and longitude coordinates.

[0065] Furthermore, the process of determining the basic attribute information of the target detection object is as follows:

[0066] Based on the correspondence formula of the longitude and latitude coordinates of all reference vertices, the longitude and latitude coordinates of the acquisition device and the center coordinates of the target detection object, the longitude and latitude coordinates of the reference vertices and the target center coordinates corresponding to the longitude and latitude coordinates of the acquisition device are determined, and the basic attribute information is determined according to the target center coordinates.

[0067] Furthermore, it also includes:

[0068] The basic attribute information is sent to the cloud server for storage.

[0069] In the above embodiments, although the steps are numbered S1, S2, etc., these are only specific embodiments given by the present invention. Those skilled in the art may adjust the execution order of S1, S2, etc. according to actual conditions, which is also within the scope of protection of the present invention. It can be understood that in some embodiments, some or all of the above embodiments may be included.

[0070] The bounding box of 3D object detection represents the edge information of the target vehicle in more detail than that of 2D object detection. This paper proposes a method that, after a simple five-point calibration configuration, uses image pixels and geographic coordinate mapping to obtain the target's precise latitude and longitude location information, as well as its length, width, and height dimensions. The obtained center latitude and longitude coordinates are more accurate than those calculated by millimeter-wave radar algorithms and 2D video object detection algorithms. In particular, the method for obtaining height information solves the problem that current video object detection cannot obtain the target vehicle's actual height. When sent to the terminal or monitoring platform, a 1:1 display of real-time traffic targets is achieved. In addition, because the bottom bounding box of 3D object detection has little height information on the ground, it is easier to determine the vehicle's driving trajectory. These characteristics simplify the judgment logic of traffic events such as trucks exceeding the height limit, trucks driving outside the designated lane, and motor vehicles changing lanes continuously, greatly improving the accuracy of event reporting. It provides greater safety for pedestrians and vehicles on surrounding roads, and can also make early predictions to plan and issue reasonable driving routes for autonomous vehicles.

[0071] The present invention also provides a 3D video-based target detection system, the specific technical solution is as follows:

[0072] Determine a 3D detection frame corresponding to a target detection object within the area to be detected based on any frame image data of the area to be detected;

[0073] Based on the 3D detection frame, the latitude and longitude coordinates of the reference vertex corresponding to any target detection object are determined, and the basic attribute information of the target detection object is determined in combination with the latitude and longitude coordinates of the acquisition device, where the basic attribute information includes the height information of the target detection object.

[0074] Based on the above solution, the present invention can also be improved as follows.

[0075] Furthermore, the process of determining the latitude and longitude coordinates of the reference vertex corresponding to any target detection object is as follows:

[0076] Determine the pixel coordinates of a reference vertex corresponding to any target detection object based on the 3D detection frame;

[0077] According to the coordinate conversion principle, the pixel coordinates of any reference vertex are converted into latitude and longitude coordinates.

[0078] Furthermore, the process of determining the basic attribute information of the target detection object is as follows:

[0079] Based on the correspondence formula of the longitude and latitude coordinates of all reference vertices, the longitude and latitude coordinates of the acquisition device and the center coordinates of the target detection object, the longitude and latitude coordinates of the reference vertices and the target center coordinates corresponding to the longitude and latitude coordinates of the acquisition device are determined, and the basic attribute information is determined according to the target center coordinates.

[0080] Furthermore, it also includes:

[0081] The basic attribute information is sent to the cloud server for storage.

[0082] It should be noted that the beneficial effects of the target detection system based on 3D video provided by the above embodiment are the same as the beneficial effects of the target detection method based on 3D video, which will not be described in detail here. In addition, when implementing its functions, the system provided by the above embodiment is only illustrated by the division of the above functional modules. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the system can be divided into different functional modules according to actual conditions to complete all or part of the functions described above. In addition, the system and method embodiments provided by the above embodiment belong to the same concept. The specific implementation process is detailed in the method embodiment, which will not be described in detail here.

[0083] like Figure 3As shown, an electronic device 300 according to an embodiment of the present invention includes a processor 320, which is coupled to a memory 310. The memory 310 stores at least one computer program 330. The at least one computer program 330 is loaded and executed by the processor 320 to enable the electronic device 300 to implement any of the above methods. Specifically:

[0084] The electronic device 300 may vary significantly due to different configurations or performance, and may include one or more processors 320 (Central Processing Units, CPUs) and one or more memories 310, wherein the one or more memories 310 store at least one computer program 330, which is loaded and executed by the one or more processors 320 to enable the electronic device 300 to implement a 3D video-based target detection method provided in the above embodiment. Of course, the electronic device 300 may also have components such as a wired or wireless network interface, a keyboard, and an input / output interface for input and output. The electronic device 300 may also include other components for implementing device functions, which will not be described in detail here.

[0085] A computer-readable storage medium according to an embodiment of the present invention stores at least one computer program, and the at least one computer program is loaded and executed by a processor to enable a computer to implement any of the above methods.

[0086] Alternatively, the computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a compact disc (CD-ROM), a magnetic tape, a floppy disk, an optical data storage device, or the like.

[0087] In an exemplary embodiment, a computer program product or computer program is also provided, the computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the electronic device to perform any of the above methods.

[0088] It should be noted that the terms "first," "second," and the like in the specification and claims of this application are used to distinguish similar objects and to define a specific order or precedence. Where appropriate, the order used for similar objects may be interchanged, such that the embodiments of the present application described herein can be implemented in an order other than the order shown or described.

[0089] Those skilled in the art will appreciate that the present invention may be implemented as a system, method, or computer program product. Therefore, the present disclosure may be implemented in the following forms: entirely in hardware, entirely in software (including firmware, resident software, microcode, etc.), or in a combination of hardware and software, generally referred to herein as a "circuit," "module," or "system." Furthermore, in some embodiments, the present invention may be implemented in the form of a computer program product embodied in one or more computer-readable media containing computer-readable program code.

[0090] Any combination of one or more computer-readable media can be used. A computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or component, or any combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device or device.

[0091] Although the embodiments of the present invention have been shown and described above, it will be understood that the above embodiments are illustrative and are not to be construed as limitations on the present invention. A person skilled in the art may change, modify, replace and modify the above embodiments within the scope of the present invention.

Claims

1. A target detection method based on 3D video, characterized in that: include: Determine a 3D detection frame corresponding to a target detection object within the area to be detected based on any frame image data of the area to be detected; Determine the latitude and longitude coordinates of a reference vertex corresponding to any target detection object based on the 3D detection frame, and determine basic attribute information of the target detection object in combination with the latitude and longitude coordinates of the acquisition device, wherein the basic attribute information includes height information of the target detection object; The image data is obtained as follows: By decoding the video stream captured by the acquisition device, each frame of image data is obtained; the decoding process includes: Through the decoder, the compressed video stream that meets the specific coding standard is decoded to obtain image data; The process of determining a 3D detection frame corresponding to a target detection object according to image data is as follows: wherein the target detection object is a vehicle; By presetting the vehicle detection model, the image data is processed to obtain the initial 3D detection frame of all target detection objects corresponding to any image data; Recheck all initial 3D detection frames on any image data, first remove the initial 3D detection frames corresponding to non-motor vehicles, then adjust the initial 3D detection frames with obvious detection errors to obtain the 3D detection frames; An initial 3D detection frame with obvious detection errors refers to: determining the actual length of the vehicle and comparing it with the length of the corresponding detection edge in the initial 3D detection frame. If the difference between the detection edge length and the actual length exceeds the error range, the initial 3D detection frame corresponding to the detection edge is determined as the initial 3D detection frame with obvious detection errors. At this time, feature extraction is performed on the area corresponding to the initial 3D detection frame to obtain an image containing only the area corresponding to the initial 3D detection frame, or partial highlighting is performed on the image data corresponding to the initial 3D detection frame to obtain a new image, so that the features of the initial 3D detection frame with obvious detection errors are more prominent. The processed image data is re-input into the preset vehicle detection model for detection, and the results are verified again. If there are still initial 3D detection frames with obvious detection errors, the 3D detection frame is re-determined through manual intervention. The preset vehicle detection model is trained as follows: After constructing the training sample set, the sample images in the training sample set are input into the vehicle detection model and passed through the input layer to the hidden layer for feature extraction. The sample images are then reduced in dimension through the pooling layer and input into the three-layer convolution layer for vehicle feature extraction. The extracted results are integrated and classified through the fully connected layer and passed to the output layer for output. The entire classification result is compared with the label category corresponding to each image in the validation set. Combined with the loss function, it is determined whether the current vehicle detection model has achieved the preset accuracy. If so, the vehicle model is output. During the training process, a feature extraction layer is added after any convolutional layer. This feature extraction layer is not connected to the fully connected layer. The feature extraction layer directly extracts the features of the image processed by any convolutional layer according to preset rules, analyzes the extracted features, and determines whether the processing of the current convolutional layer is accurate. If the processing result is inaccurate, the convolutional layer of the model is replaced or adjusted.

2. The object detection method based on 3D video according to claim 1, characterized in that: The process of determining the latitude and longitude coordinates of the reference vertex corresponding to any target detection object is as follows: Determine the pixel coordinates of a reference vertex corresponding to any target detection object based on the 3D detection frame; According to the coordinate conversion principle, the pixel coordinates of any reference vertex are converted into latitude and longitude coordinates.

3. The object detection method based on 3D video according to claim 1, characterized in that: The process of determining the basic attribute information of the target detection object is as follows: Based on the correspondence formula of the longitude and latitude coordinates of all reference vertices, the longitude and latitude coordinates of the acquisition device and the center coordinates of the target detection object, the longitude and latitude coordinates of the reference vertices and the target center coordinates corresponding to the longitude and latitude coordinates of the acquisition device are determined, and the basic attribute information is determined according to the target center coordinates.

4. The object detection method based on 3D video according to claim 1, characterized in that: Also includes: The basic attribute information is sent to the cloud server for storage.

5. A 3D video-based target detection system, using the 3D video-based target detection method according to claim 1, characterized in that: The system includes: Determine a 3D detection frame corresponding to a target detection object within the area to be detected based on any frame image data of the area to be detected; Based on the 3D detection frame, the latitude and longitude coordinates of the reference vertex corresponding to any target detection object are determined, and the basic attribute information of the target detection object is determined in combination with the latitude and longitude coordinates of the acquisition device, where the basic attribute information includes the height information of the target detection object.

6. The 3D video-based target detection system according to claim 5, characterized in that: The process of determining the latitude and longitude coordinates of the reference vertex corresponding to any target detection object is as follows: Determine the pixel coordinates of a reference vertex corresponding to any target detection object based on the 3D detection frame; According to the coordinate conversion principle, the pixel coordinates of any reference vertex are converted into latitude and longitude coordinates.

7. The 3D video-based target detection system according to claim 5, characterized in that: The process of determining the basic attribute information of the target detection object is as follows: Based on the correspondence formula of the longitude and latitude coordinates of all reference vertices, the longitude and latitude coordinates of the acquisition device and the center coordinates of the target detection object, the longitude and latitude coordinates of the reference vertices and the target center coordinates corresponding to the longitude and latitude coordinates of the acquisition device are determined, and the basic attribute information is determined according to the target center coordinates.

8. The 3D video-based target detection system according to claim 5, characterized in that: Also includes: The basic attribute information is sent to the cloud server for storage.

9. An electronic device, characterized in that: The electronic device includes a processor coupled to a memory, wherein the memory stores at least one computer program, and the at least one computer program is loaded and executed by the processor so that the electronic device implements the method according to any one of claims 1 to 4.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores at least one computer program, which is loaded and executed by a processor to enable a computer to implement the method according to any one of claims 1 to 4.