A method and system for laser radar depth information and video target detection tracking
By combining lidar and machine vision, and utilizing laser beam detection and video information processing, the problem of low accuracy in real-time object position detection has been solved, achieving higher detection precision.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- INTELLIGENT INTER CONNECTION TECH CO LTD
- Filing Date
- 2023-04-13
- Publication Date
- 2026-04-24
AI Technical Summary
Existing methods for real-time object position detection are limited, resulting in low detection accuracy.
By combining LiDAR and machine vision, the system detects the initial position information of the target object using a laser beam, collects target video information, performs keyframe analysis and feature extraction, verifies the real-time position information of the target object, and tracks it based on its movement trajectory.
It improves the accuracy of real-time object position detection.
Smart Images

Figure CN116520345B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, specifically to a method and system for detecting and tracking video targets using lidar depth information. Background Technology
[0002] Currently, commercially available real-time object position detection methods include video target detection and tracking, which determines the real-time position of an object by capturing its movement trajectory via video. LiDAR detection systems use laser beams to detect the position, speed, and other characteristics of a target. Their working principle involves emitting a detection signal (laser beam) towards the target, then comparing the received signal reflected back from the target (target echo) with the emitted signal. After appropriate processing, information about the target, such as distance, azimuth, altitude, speed, attitude, and even shape, can be obtained.
[0003] However, due to the inherent defects of each of the aforementioned object position detection methods, current real-time object position detection technology still suffers from the technical problem of low detection accuracy caused by the reliance on a single target object position detection method. Summary of the Invention
[0004] This application provides a method and system for laser radar depth information and video target detection and tracking, which solves the technical problem in the prior art where the protection effect is poor due to inaccurate protection circuit settings.
[0005] The first aspect of this application provides a method for laser radar depth information and video target detection and tracking. The method includes: emitting a laser beam at a target object using a laser radar to detect the first position information of the target object; acquiring video information of the target object using machine vision, wherein the target video information includes the target object; performing keyframe analysis on the target video to obtain A keyframes, where A is a positive integer greater than 1; extracting features from the A keyframes to determine the second position information of the target object based on the target object feature information; verifying the position of the first position information and the second position information to determine the real-time position information of the target object; and tracking the target object based on the movement trajectory of the real-time position information of the target object.
[0006] A second aspect of this application provides a system for laser radar depth information and video target detection and tracking. The system includes: a first position information acquisition module, used to emit a laser beam from a laser radar onto a target object to detect the first position information of the target object; a target video information acquisition module, used to acquire target video information by machine vision, wherein the target video information includes the target object; a keyframe analysis module, used to perform keyframe analysis on the target video to acquire A keyframes, where A is a positive integer greater than 1; a second position information acquisition module, used to extract features from the A keyframes and determine the second position information of the target object based on the target object's feature information; a position verification module, used to verify the first and second position information to determine the real-time position information of the target object; and a target object tracking module, used to track the target object based on the movement trajectory of the real-time position information of the target object.
[0007] One or more technical solutions provided in this application have at least the following technical effects or advantages:
[0008] This application provides a method for laser radar depth information and video target detection and tracking, relating to the field of data processing technology. It involves using a laser radar to emit a laser beam at a target object to detect its first position information, acquiring target video information using machine vision, performing keyframe analysis on the target video to obtain A keyframes and extracting their features, determining the target object's second position information based on these features, verifying the first and second position information to determine the target object's real-time position information, and tracking the target object using its real-time position trajectory. This method solves the technical problem of low accuracy in real-time object position detection due to the reliance on a single method in existing technologies, thus improving the accuracy of real-time object position detection. Attached Figure Description
[0009] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0010] Figure 1 This application provides a schematic flowchart of a method for detecting and tracking video targets using lidar depth information, as provided in an embodiment of the present application.
[0011] Figure 2 This is a flowchart illustrating the process of detecting the first position information of a target object in a method for detecting and tracking a target object using lidar depth information and video, as provided in an embodiment of this application.
[0012] Figure 3 This is a flowchart illustrating the process of determining the second position information of a target object in a method for detecting and tracking a target object using lidar depth information and video, as provided in an embodiment of this application.
[0013] Figure 4 This is a schematic diagram of a system structure for lidar depth information and video target detection and tracking provided in an embodiment of this application.
[0014] Explanation of reference numerals in the attached figures: First position information acquisition module 11, target video information acquisition module 12, keyframe analysis module 13, second position information acquisition module 14, position verification module 15, target object tracking module 16. Detailed Implementation
[0015] This application provides a method for detecting and tracking video targets using lidar depth information, which addresses the technical problem in the prior art where inaccurate protection circuit settings lead to poor protection performance.
[0016] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.
[0017] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or server that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or modules not explicitly listed or inherent to such processes, methods, products, or devices.
[0018] Example 1
[0019] like Figure 1As shown, this application provides a method for detecting and tracking video targets using lidar depth information, the method comprising:
[0020] S100: The laser beam is emitted from the laser radar to detect the first position information of the target object;
[0021] Specifically, the aforementioned lidar refers to a radar system that detects the position, velocity, and other characteristics of a target by emitting a laser beam. It consists of a transmitting system, a receiving system, and information processing components. This application transmits a laser beam (detection signal) to the target object through the transmitting system. Then, it compares the received signal reflected back from the target with the transmitted signal and performs appropriate processing to obtain relevant information about the target object, such as its distance, azimuth, height, velocity, attitude, and even shape. This information serves as the target object's initial position information and can be used as a source for subsequently determining the target object's real-time position.
[0022] Furthermore, such as Figure 2 As shown, step S100 in this embodiment further includes:
[0023] S110: Determine the detectable range based on the emission limit of the laser beam detected by the lidar;
[0024] S120: Based on the detectable range, obtain the starting angle and ending angle of the lidar detection;
[0025] S130: Based on the starting angle and the ending angle, the laser beam is emitted and detected to determine the first position information of the target object.
[0026] Specifically, the maximum detection range of a lidar can be determined by the emission limit of its laser beam, i.e., the farthest emission distance of the laser beam. From this detection range, the starting and ending angles of the lidar's detection can be obtained; these are the starting and ending angles for the lidar to scan the object. The lidar's emission system emits a laser beam towards the target object, and then compares the received signal reflected from the target with the emitted signal. After appropriate processing, information about the target object, such as distance, orientation, height, speed, attitude, and even shape, can be obtained. This serves as the target object's initial position information, which can be used as a basis for subsequently determining the target object's real-time position.
[0027] S200: The target object is captured by machine vision to obtain target video information, wherein the target video information includes the target object;
[0028] Specifically, machine vision uses machines to replace human eyes for measurement and judgment. A camera captures video of the target object, which is then transmitted to a dedicated image processing system. Based on pixel distribution and information such as brightness and color, the video is converted into a digital signal and transmitted to a computer system. This digital signal serves as the target video information, including the target object's volume, size, position, and movement. This target video information can then be used as the basis for subsequently obtaining secondary positional information of the target object.
[0029] S300: Perform keyframe analysis on the target video to obtain A keyframes, where A is a positive integer greater than 1;
[0030] Specifically, video is a hybrid medium in which multiple still image frames and continuous audio information move synchronously along the timeline. A keyframe refers to the frame in which a character or object performs a key action in the movement of several frames that constitute a video. Analyzing the keyframes of the target video is to find and extract the keyframes that embody the significant features of each shot in the target video. This application assumes that a total of A keyframes are extracted, where A is a positive integer greater than 1, meaning that the number of keyframes is two or more, which can be used as the data source for subsequent feature extraction.
[0031] Furthermore, step S300 in this embodiment of the application also includes:
[0032] S310: Perform video segmentation on the target video based on the lens boundary detection algorithm to obtain information on multiple video segments;
[0033] S320: Extract key frames from the multiple video segment information to obtain the A key frames.
[0034] Specifically, the lens boundary detection algorithm is one of the key technologies in video retrieval. When a shot changes, the video image undergoes significant changes, such as an increase in the difference between corresponding pixels in consecutive frames, a significant change in color distribution, or the sudden appearance or disappearance of object edges. Lens boundary detection essentially detects these changes in obvious features. This application utilizes the lens boundary detection algorithm to segment the target video into multiple segments based on different obvious features, using these segments as the information for multiple video segments. Then, A keyframes with different obvious features are extracted sequentially from these multiple video segments, which can serve as the data source for subsequent feature extraction.
[0035] S400: Perform feature extraction on the A keyframes, and determine the second position information of the target object based on the target object feature information;
[0036] Specifically, based on the image information of the A key frames, different features of the target object contained in the A key frames are extracted sequentially, including the position, size, quantity and other feature information of the target object in the image. These target object feature information are used as the second position information of the target object, which can serve as another source of information for subsequently determining the real-time position information of the target object.
[0037] Furthermore, step S400 in this embodiment of the application also includes:
[0038] S410: Extract I-frame image;
[0039] S420: Segment the I-frame image according to the preset segmentation standard to determine the target I-frame image block;
[0040] S430: Calculate the discrete cosine transformation coefficients based on the target I-frame image block to obtain the first feature value and the second feature value;
[0041] S440: Calculate the feature value of the I-frame image based on the first feature value and the second feature value using the feature value calculation formula;
[0042] S450: Add the feature values of the I-frame image to the feature information of the target object.
[0043] S431: The formula for calculating the first eigenvalue is as follows:
[0044] S432:
[0045] (a+b=1, and a>b)
[0046] S433:T n This refers to the first eigenvalue, DC. n (x′,y′) refers to the first DC coefficient, AC n (x′,y′) refers to the first AC coefficient, n refers to the first I-frame image, (x′,y′) refers to the (x′,y′)th sub-block of the first I-frame image, a refers to the influence factor of the first DC coefficient on the first eigenvalue, and b refers to the influence factor of the first AC coefficient on the first eigenvalue.
[0047] S434: The formula for calculating the second eigenvalue is as follows:
[0048] (c+d=1, and c>d)
[0049] T n+1 This refers to the second eigenvalue, DC. n+1 (x″, y″) refers to the second DC coefficient, AC n+1(x″,y″) refers to the second AC coefficient, n+1 refers to the second I-frame image, (x″,y″) refers to the (x″,y″)th sub-block of the second I-frame image, c refers to the influence factor of the second DC coefficient on the second eigenvalue, and d refers to the influence factor of the second AC coefficient on the second eigenvalue.
[0050] Specifically, the preset segmentation standard is preset by relevant technicians based on pixel blocks in the I-frame image. The I-frame image in the keyframe is extracted, and then the extracted I-frame image is further segmented according to the preset segmentation standard to obtain target I-frame image blocks. Discrete cosine transformation coefficients are then calculated based on these target I-frame image blocks. The calculation of discrete cosine transformation coefficients is a transformation based on the cosine function, commonly used in image processing and image recognition. The formula for calculating the first feature value in this application is as follows:
[0051] (a+b=1, and a>b)
[0052] T n This refers to the first eigenvalue, DC. n (x′,y′) refers to the first DC coefficient, AC n (x′,y′) refers to the first AC coefficient, n refers to the first I-frame image, (x′,y′) refers to the (x′,y′)th sub-block of the first I-frame image, a refers to the influence factor of the first DC coefficient on the first eigenvalue, and b refers to the influence factor of the first AC coefficient on the first eigenvalue.
[0053] Substituting the first DC coefficient, first AC coefficient, etc. of the above I-frame image into the formula, the first characteristic value is calculated;
[0054] The formula for calculating the second eigenvalue in this application is as follows:
[0055] (c+d=1, and c>d)
[0056] T n+1 This refers to the second eigenvalue, DC. n+1 (x″, y″) refers to the second DC coefficient, AC n+1 (x″,y″) refers to the second AC coefficient, n+1 refers to the second I-frame image, (x″,y″) refers to the (x″,y″)th sub-block of the second I-frame image, c refers to the influence factor of the second DC coefficient on the second eigenvalue, and d refers to the influence factor of the second AC coefficient on the second eigenvalue.
[0057] Substituting the second DC coefficient and second AC coefficient of the above I-frame image into the formula, the second feature value is calculated. The first feature value and the second feature value are the feature values of the I-frame image. The feature values of the I-frame image are added to the feature information of the target object, and used as the feature information of the target object. This can be used as the basic data for obtaining the second position information of the target object in the future.
[0058] Furthermore, such as Figure 3 As shown, step S400 in this embodiment further includes:
[0059] S460: Based on the feature values of the I-frame image in the target object feature information, perform grid division on the images in the A key frames to obtain the divided image information;
[0060] S470: Based on the convolution kernel, the segmented image information is traversed and recognized to obtain grid image recognition information;
[0061] S480: Determine the second position information of the target object based on the grid image recognition information.
[0062] Specifically, based on the I-frame image feature values obtained from the target object feature information, the image information in the A keyframes is divided into a grid, resulting in image information for each grid. A convolution kernel, in image processing, is a weighted average of pixels in a small region of an input image, which becomes the corresponding pixel in the output image. The weights are defined by a function called the convolution kernel. Based on the convolution kernel, the divided image information is traversed and identified to pinpoint the grid positions where the target object appears in the image. This grid position is used as grid image recognition information, which is then used to determine the object's location. This serves as the target object's second position information, providing another source of information for subsequently determining the target object's real-time position.
[0063] S500: Verify the position of the first position information and the second position information to determine the real-time position information of the target object;
[0064] Specifically, in chronological order, the obtained first location information is matched one by one with the target object's location information in the second location information. Based on the change path pattern of the target object's location and the accuracy of location information from different sources, the location information from the two sources is filtered and verified. The verified location information is used as the real-time location information of the target object, which can determine the real-time information and movement trajectory of the target object.
[0065] Furthermore, step S500 in this embodiment of the application also includes:
[0066] S510: Perform location matching between different sources for the location sources corresponding to the first location information and the second location information to obtain a set of overlapping locations and a set of single locations;
[0067] S520: Determine the accurate value range based on the location and source characteristics of the single location set;
[0068] S530: Based on the accurate value range, the single location set is filtered, and the location verification is completed according to the filtered single location set and the overlapping location set to determine the real-time location information of the target object.
[0069] Specifically, following a chronological order, the positions of the target objects in the first location information from lidar detection and the second location information from video detection are matched sequentially. All cases where the target object in the first and second location information is in the same location are considered as overlapping locations, while all cases where the target object in the first and second location information is in different locations are considered as single location sets. The location feature of each single location set represents the location information of the target object from two different sources at the same time. The source feature of each single location set refers to the accuracy of the target object's location information obtained using lidar or video detection. For example, because lidar detection uses a laser beam for location detection, and light travels very quickly, lidar detection has high accuracy. Video detection, on the other hand, transmits video from a camera to a computer system via a network, which may have network latency, resulting in lower accuracy of the detected location information. By comparing the location features and source features of each single location set, the location information of the target object from two different sources at the same time can be compared, and the more accurate lidar detection location information can be referenced to determine the range of the target object's accurate location, which serves as the accuracy range. Based on the position range of the accurate value range, all position information in the single position set is filtered, and single positions within the accurate value range are retained. The filtered single position set is added to the overlapping position set for verification. If the overall change path of the fused position set matches the change path of the overlapping position set, the fused position set can be used as the real-time position information of the target object, and the real-time position of the target object can be obtained.
[0070] S600: Track the target object based on its movement trajectory according to the real-time location information of the target object.
[0071] Specifically, the target object's movement trajectory and real-time position can be determined from the target object's real-time position information. By acquiring the target object's real-time position, the target object can be tracked to achieve the effect of detecting the object's real-time position.
[0072] Furthermore, the embodiments of this application also include step S700, which further includes:
[0073] S710: Perform position deviation correction on the real-time position information of the target object to obtain the deviation correction result;
[0074] S720: Based on the deviation correction result, perform position correction on the real-time position information of the target object to obtain the corrected real-time position information of the target object;
[0075] S730: Optimize the real-time position information of the target object based on the real-time position information of the target object.
[0076] Specifically, based on the detection characteristics and accuracy of the aforementioned lidar and video detection methods, position deviation correction is performed on the real-time position information of the target object. The deviation information, caused by network latency and other factors, is compared with the correct position to obtain the information indicating the target object's deviation that needs correction, which serves as the deviation correction result. Based on the deviation information from the correction result, the real-time position information of the target object is corrected to obtain the corrected real-time position information. This corrected real-time position information replaces the uncorrected real-time position information of the target object, thereby optimizing the real-time position information of the target object and improving the accuracy of real-time object position detection.
[0077] In summary, the embodiments of this application have at least the following technical effects:
[0078] This application uses a lidar to emit a laser beam at a target object to detect the first position information of the target object, acquires target video information of the target object through machine vision, performs keyframe analysis on the target video, obtains A key frames and extracts their features, determines the second position information of the target object based on the target object feature information, verifies the position information of the first position information and the second position information to determine the real-time position information of the target object, and tracks the target object based on the movement trajectory of the real-time position information of the target object.
[0079] This achieves the technical effect of improving the accuracy of real-time object position detection.
[0080] Example 2
[0081] Based on the same inventive concept as the method for detecting and tracking video targets using lidar depth information in the foregoing embodiments, such as Figure 4 As shown, this application provides a system for laser radar depth information and video target detection and tracking. The system and method embodiments in this application are based on the same inventive concept. The system includes:
[0082] The first location information acquisition module is used to emit a laser beam at the target object using a lidar to detect the first location information of the target object.
[0083] A target video information acquisition module is used to acquire target video information by capturing video of the target object through machine vision, wherein the target video information includes the target object;
[0084] A keyframe analysis module is used to perform keyframe analysis on the target video to obtain A keyframes, where A is a positive integer greater than 1.
[0085] The second location information acquisition module is used to extract features from the A key frames and determine the second location information of the target object based on the target object feature information.
[0086] A location verification module is used to verify the location information of the first location information and the second location information to determine the real-time location information of the target object.
[0087] A target object tracking module is used to track the target object based on its movement trajectory according to its real-time location information.
[0088] Furthermore, the system also includes:
[0089] A detectable range determination module, wherein the detectable range determination module is used to determine the detectable range based on the emission limit of the laser beam detected by the lidar;
[0090] An angle acquisition module is used to acquire the start angle and end angle of the lidar detection based on the detectable range.
[0091] A detection module is used to detect the laser beam based on the starting angle and the ending angle, and to determine the first position information of the target object.
[0092] Furthermore, the system also includes:
[0093] A video segmentation module is used to segment the target video based on a shot boundary detection algorithm to obtain information of multiple video segments.
[0094] A keyframe extraction module is used to extract keyframes from the multiple video segment information to obtain the A keyframes.
[0095] Furthermore, the system also includes:
[0096] I-frame image extraction module, the I-frame image extraction module is used to extract I-frame images;
[0097] An image segmentation module is used to segment the I-frame image according to a preset segmentation standard to determine the target I-frame image block;
[0098] A discrete cosine transformation coefficient calculation module is used to calculate the discrete cosine transformation coefficient based on the target I-frame image block to obtain a first feature value and a second feature value.
[0099] The I-frame image feature value calculation module is used to calculate the I-frame image feature value based on the first feature value and the second feature value using the feature value calculation formula.
[0100] I-frame image feature value addition module, the I-frame image feature value addition module is used to add the I-frame image feature values to the target object feature information;
[0101] Furthermore, the system also includes:
[0102] The first eigenvalue calculation module uses the following formula to calculate the first eigenvalue:
[0103] (a+b=1, and a>b)
[0104] T n This refers to the first eigenvalue, DC. n (x′,y′) refers to the first DC coefficient, AC n (x′,y′) refers to the first AC coefficient, n refers to the first I-frame image, (x′,y′) refers to the (x′,y′)th sub-block of the first I-frame image, a refers to the influence factor of the first DC coefficient on the first eigenvalue, and b refers to the influence factor of the first AC coefficient on the first eigenvalue.
[0105] The second eigenvalue calculation module uses the following formula to calculate the second eigenvalue:
[0106] (c+d=1, and c>d)
[0107] T n+1 This refers to the second eigenvalue, DC. n+1 (x″, y″) refers to the second DC coefficient, AC n+1(x″,y″) refers to the second AC coefficient, n+1 refers to the second I-frame image, (x″,y″) refers to the (x″,y″)th sub-block of the second I-frame image, c refers to the influence factor of the second DC coefficient on the second eigenvalue, and d refers to the influence factor of the second AC coefficient on the second eigenvalue.
[0108] Furthermore, the system also includes:
[0109] A grid partitioning module is used to partition the images in the A key frames into grids based on the feature values of the I-frame image in the target object feature information, and to obtain partitioned image information.
[0110] A grid image recognition information acquisition module is used to perform traversal recognition on the divided image information based on a convolution kernel to acquire grid image recognition information.
[0111] The second location information acquisition module is used to determine the second location information of the target object based on the grid image recognition information.
[0112] Furthermore, the system also includes:
[0113] A location matching module is used to perform location matching between different sources for the location sources corresponding to the first location information and the second location information to obtain a set of overlapping locations and a single location set.
[0114] An accurate value range determination module is used to determine the accurate value range based on the location features and source features of the single location set.
[0115] A real-time location information determination module is used to filter the single location set based on the accurate value range, perform location verification based on the filtered single location set and the overlapping location set, and determine the real-time location information of the target object.
[0116] Furthermore, the system also includes:
[0117] A position deviation correction module is used to correct the position deviation of the target object's real-time position information to obtain a deviation correction result.
[0118] A position correction module is used to correct the real-time position information of the target object based on the deviation correction result, so as to obtain the corrected real-time position information of the target object.
[0119] A position optimization module is used to optimize the real-time position information of the target object based on the real-time position information of the target object.
[0120] It should be noted that the order of the embodiments described above is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. Furthermore, specific embodiments have been described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps described in the claims can be performed in a different order than that shown in the embodiments and still achieve the desired result. Additionally, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0121] The above description is only a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
[0122] This specification and accompanying drawings are merely illustrative examples of this application and are intended to cover any and all modifications, variations, combinations, or equivalents within the scope of this application. Clearly, those skilled in the art can make various alterations and modifications to this application without departing from its scope. Therefore, if such modifications and modifications fall within the scope of this application and its equivalents, this application intends to include such modifications and modifications.
Claims
1. A method for detecting and tracking video targets using lidar depth information, characterized in that, The method includes: The laser radar emits a laser beam at the target object to detect the first position information of the target object; The target object is captured by machine vision to obtain target video information, wherein the target video information includes the target object; Perform keyframe analysis on the target video to obtain A keyframes, where A is a positive integer greater than 1; Feature extraction is performed on the A keyframes, and the second position information of the target object is determined based on the target object feature information; The first location information and the second location information are verified to determine the real-time location information of the target object. The target object is tracked based on its movement trajectory according to its real-time location information; Detecting the first position information of the target object includes: The detectable range is determined based on the emission limit of the laser beam detected by the lidar; Based on the detectable range, the starting angle and ending angle of the lidar detection are obtained; Based on the starting angle and the ending angle, the laser beam is emitted and detected to determine the first position information of the target object; Feature extraction is performed on the A keyframes, including: Extract I-frame images; The I-frame image is segmented according to a preset segmentation standard to determine the target I-frame image block; Based on the target I-frame image block, the discrete cosine transformation coefficient is calculated to obtain the first feature value and the second feature value; Based on the first feature value and the second feature value, the feature value of the I-frame image is calculated using the feature value calculation formula; The feature values of the I-frame image are added to the feature information of the target object; Determining the second position information of the target object based on the target object's feature information includes: Based on the feature values of the I-frame image in the target object feature information, the images in the A key frames are divided into grids to obtain the divided image information; Based on the convolution kernel, the segmented image information is traversed and recognized to obtain grid image recognition information; Based on the grid image recognition information, the second position information of the target object is determined; Determining the real-time position information of the target object includes: The location sources corresponding to the first location information and the second location information are matched between different sources to obtain a set of overlapping locations and a set of single locations. The accurate value range is determined by the location and source characteristics of the single location set; The single location set is filtered based on the accurate value range, and the location verification is completed based on the filtered single location set and the overlapping location set to determine the real-time location information of the target object.
2. The method as described in claim 1, characterized in that, Obtaining the A keyframes includes: The target video is segmented based on a lens boundary detection algorithm to obtain information on multiple video segments. Keyframes are extracted from the multiple video segments to obtain the A keyframes.
3. The method as described in claim 1, characterized in that, Obtaining the first feature value and the second feature value includes: The formula for calculating the first eigenvalue is as follows: It refers to the first eigenvalue. This refers to the first DC coefficient. This refers to the first AC coefficient, and n refers to the first I-frame image. It refers to the first I-frame image. In the sub-block, 'a' refers to the influence factor of the first DC coefficient on the first eigenvalue, and 'b' refers to the influence factor of the first AC coefficient on the first eigenvalue. The formula for calculating the second eigenvalue is as follows: It refers to the second eigenvalue. This refers to the second DC coefficient. This refers to the second AC coefficient, and n+1 refers to the second I-frame image. It refers to the first I-frame image. In the sub-block, c refers to the influence factor of the second DC coefficient on the second eigenvalue, and d refers to the influence factor of the second AC coefficient on the second eigenvalue.
4. The method as described in claim 1, characterized in that, include: The real-time position information of the target object is corrected for position deviation to obtain the deviation correction result. Based on the deviation correction result, the real-time position information of the target object is corrected to obtain the corrected real-time position information of the target object. The real-time position information of the target object is optimized based on the real-time position information of the target object.
5. A system for laser radar depth information and video target detection and tracking, used to implement the method of claim 1, characterized in that, The system includes: The first location information acquisition module is used to emit a laser beam at the target object using a lidar to detect the first location information of the target object. A target video information acquisition module is used to acquire target video information by capturing video of the target object through machine vision, wherein the target video information includes the target object; A keyframe analysis module is used to perform keyframe analysis on the target video to obtain A keyframes, where A is a positive integer greater than 1. The second location information acquisition module is used to extract features from the A key frames and determine the second location information of the target object based on the target object feature information. A location verification module is used to verify the location information of the first location information and the second location information to determine the real-time location information of the target object. A target object tracking module is used to track the target object based on its movement trajectory according to its real-time location information.
Citation Information
Patent Citations
Target detection and identification device and method based on multi-fusion sensor
CN110428008A
SAR image change detection method based on unsupervised space-frequency representation learning fusion
CN115393706A