Target recognition and tracking method and system
By combining a multi-sensor fusion method with binocular cameras and UWB sensors, the problem of inaccurate identification and tracking of workers by intelligent working vehicles in orchards is solved, high-precision target tracking and safe following are achieved, and work efficiency and safety are improved.
Patent Information
- Application Number
- CN202210152153.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-18
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2042-02-18
AI Technical Summary
In orchards, intelligent work vehicles have inaccuracies in target recognition and tracking of workers, especially under the influence of complex environmental factors such as light intensity and tree branch obstruction. The algorithm processing results of existing technologies are unreliable, leading to driving safety hazards.
A method combining binocular cameras and ultra-wideband (UWB) sensors is adopted to obtain the first position information of the target through the instance segmentation model, and the second position information is obtained by determining the relative distance between nodes using UWB. The two information are then fused through the Kalman filter algorithm to achieve accurate tracking position calculation.
It improves the recognition accuracy and tracking stability of intelligent work vehicles for workers, ensures the safety of vehicles and personnel, and improves work efficiency.
Smart Images

Figure CN114693773B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of target detection, and in particular to a target recognition and tracking method and system. Background Art
[0002] In the current agricultural orchard management, it is necessary to uniformly manage the orchard's intelligent transport vehicles and operators, and usually it is necessary to grasp the whereabouts of the operators in real time, which is of great significance to the transportation management of the smart orchard.
[0003] Currently, the mainstream target tracking solution for intelligent orchard vehicles uses algorithms to process real-time camera images to provide data support for the vehicles' operations. However, orchards are plagued by numerous factors, such as light intensity and tree branch obstruction, which can affect the quality of camera images. This can lead to unreliable algorithmic processing results and pose a risk to subsequent driving safety. The complex orchard environment places high demands on the reliability of target recognition and tracking for intelligent transport vehicles. The distance between the transport vehicle and the operator, as well as positioning accuracy, are crucial for achieving stable and accurate tracking.
[0004] Therefore, how to achieve accurate identification and stable tracking of operating personnel by operating vehicles has become an urgent problem that needs to be solved in intelligent orchard management. Summary of the Invention
[0005] The present invention provides a target recognition and tracking method and system, which are used to solve the defect in the prior art that conventional cameras are used to acquire images and perform calculations in real time for intelligent orchard operation vehicles to achieve target tracking of operators, while ignoring environmental factors, resulting in inaccurate calculation results.
[0006] In a first aspect, the present invention provides a target recognition and tracking method, comprising:
[0007] Acquire a target image of a detection target using a binocular camera, and process the target image using an instance segmentation model to obtain first position information of the detection target;
[0008] Determining a relative distance between nodes of the detection target using ultra-wideband, and acquiring second position information of the detection target based on the relative distance between nodes;
[0009] The first position information and the second position information are fused and calculated to obtain tracking position information of the detection target.
[0010] According to a target recognition and tracking method provided by the present invention, the method uses a binocular camera to obtain a target image of a detection target, and uses an instance segmentation model to process the target image to obtain first position information of the detection target, including:
[0011] acquiring a target image of the detection target at regular intervals, and constructing a target classification image set based on the target image;
[0012] Obtaining an initial neural network model, simplifying and training the initial neural network model to obtain an optimized neural network model;
[0013] Training the optimized neural network model based on the target classification image set to obtain an instance segmentation model;
[0014] The target image is input into the instance segmentation model to complete target detection and tracking, to obtain a target frame of the target image, and the first position information is obtained based on the center point of the target frame.
[0015] According to a target recognition and tracking method provided by the present invention, the target image of the detected target is acquired at regular intervals, and a target classification image set is constructed based on the target image, including:
[0016] Using the binocular camera to collect multiple environmental images, annotating the multiple environmental images to obtain a target annotated image set;
[0017] Determine the classification of the target annotated image set, and obtain the target classified image set corresponding to the target annotated image set.
[0018] According to a target recognition and tracking method provided by the present invention, the obtaining of an initial neural network model, simplifying and training the initial neural network model to obtain the instance segmentation model includes:
[0019] Using preset coarse-grained parameter structured pruning to simplify the convolution kernel parameters of the initial neural network model to obtain an optimized neural network model;
[0020] The instance segmentation model is obtained by quantifying the preset network parameters and training the optimized neural network model.
[0021] According to a target recognition and tracking method provided by the present invention, the preset coarse-grained parameter structured pruning is used to simplify the convolution kernel parameters of the initial neural network model to obtain an optimized neural network model, including:
[0022] An evaluation factor is set for the convolution kernel parameters or feature map parameters, or a preset part of the convolution kernel and any channel of the feature map is deleted to obtain the optimized neural network model.
[0023] According to a target recognition and tracking method provided by the present invention, the method of training the optimized neural network model by quantizing preset network parameters to obtain the instance segmentation model includes:
[0024] The network parameters of the optimized neural network model are compressed using a uniform bit width or a combined bit width, and the instance segmentation model is obtained by training.
[0025] According to a target recognition and tracking method provided by the present invention, the method of determining the relative distance between nodes of the detection target by using ultra-wideband, and obtaining the second position information of the detection target based on the relative distance between nodes, includes:
[0026] Setting an ultra-wideband mobile node at the detection target;
[0027] The inter-node relative distance between the ultra-wideband mobile node and the ultra-wideband base station node is acquired, and the second location information is obtained from the inter-node relative distance.
[0028] According to a target recognition and tracking method provided by the present invention, the binocular camera and the ultra-wideband base station are located on the same horizontal plane and the same center line on the same mobile platform.
[0029] In a second aspect, the present invention further provides a target recognition and tracking system, comprising:
[0030] A first recognition module is configured to obtain a target image of a detection target using a binocular camera, and to obtain first position information of the detection target after processing the target image using an instance segmentation model;
[0031] a second identification module, configured to determine a relative distance between nodes of the detection target using ultra-wideband, and obtain second position information of the detection target based on the relative distance between nodes;
[0032] The integrated recognition module is used to fuse and calculate the first position information and the second position information to obtain the tracking position information of the detection target.
[0033] In a third aspect, the present invention also provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps of any of the target identification and tracking methods described above are implemented.
[0034] In a fourth aspect, the present invention further provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any of the target recognition and tracking methods described above.
[0035] In a fifth aspect, the present invention further provides a computer program product, comprising a computer program, which, when executed by a processor, implements the steps of any of the target recognition and tracking methods described above.
[0036] The target recognition and tracking method and system provided by the present invention use multiple types of sensors to perform target recognition and tracking on the same detection target, so as to ensure that the intelligent working vehicle can accurately identify the working personnel and stably track them in the orchard environment. Based on the precise spatial position information obtained, reliable data is provided for the vehicle's autonomous following, effectively improving the working efficiency of the intelligent working vehicle and ensuring the safety of the vehicle and the working personnel. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0038] Figure 1 This is one of the flow charts of the target recognition and tracking method provided by the present invention;
[0039] Figure 2 This is the second flow chart of the target recognition and tracking method provided by the present invention;
[0040] Figure 3 This is a working diagram of the intelligent working vehicle provided by the present invention;
[0041] Figure 4 This is a schematic diagram of the target operator identification and following results provided by the present invention;
[0042] Figure 5 It is a structural diagram of the target recognition and tracking system provided by the present invention;
[0043] Figure 6 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0044] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0045] The mainstream target tracking method for intelligent working vehicles on workers in smart orchards is to use algorithms to process images obtained by cameras in real time. The input data is relatively simple and does not fully consider the influencing factors of the orchard's complex environment, which easily causes significant interference to the camera and leads to low accuracy of the algorithm processing results. This paper proposes a new target recognition and tracking method based on multi-sensor.
[0046] Figure 1 This is one of the flow charts of the target recognition and tracking method provided by the present invention, such as Figure 1 Shown, including:
[0047] Step S1, using a binocular camera to obtain a target image of a detection target, and using an instance segmentation model to process the target image to obtain first position information of the detection target;
[0048] Step S2, using ultra-wideband to determine the relative distance between nodes of the detection target, and obtaining second position information of the detection target based on the relative distance between nodes;
[0049] Step S3: Fusion-calculate the first position information and the second position information to obtain tracking position information of the detection target.
[0050] Specifically, the present invention obtains the position information of the operating personnel, that is, the detection target, from multiple angles through the fusion of multiple sensors. First, a binocular camera is used to obtain the image of the detection target in real time, and the first position information in the current environment is calculated. Second, ultra-wideband (UWB) technology is used to utilize the UWB signal between the base station node and the mobile node to obtain the second position information in the current environment. Finally, the first position information and the second position information obtained respectively are fused and calculated to remove errors, thereby obtaining more accurate tracking position information of the detection target, so that the intelligent working vehicle can locate and follow the operating personnel according to the tracking position information.
[0051] Figure 2This is the second flow chart of the target recognition and tracking method provided by the present invention. The specific scheme adopted by the present invention includes: first, using a binocular camera to collect data images under different light intensities and different angles, and annotating the collected images, wherein only the workers are annotated, and then the annotated images are divided into a training data set and a verification data set according to a certain ratio. The purpose is to improve the generalization ability of the model. After obtaining a certain amount of training data, it is necessary to select and optimize the neural network to ensure high-quality instance segmentation accuracy and target tracking effect. By taking the midpoint of the target box as the tracking point, the actual coordinates of the tracking point are calculated according to the image coordinates of the tracking point, that is, the first position information; then using UWB ranging technology to accurately measure the distance between the worker and the intelligent working vehicle, and finally using the Kalman filter algorithm to fuse the two sets of data to achieve continuous target tracking of the intelligent working vehicle.
[0052] The present invention ensures the recognition accuracy of the intelligent transport vehicle for the detection target through the designed multi-sensor fusion target recognition and tracking method, improves the accuracy of the distance between the intelligent transport vehicle and the operator, indirectly improves the work efficiency of the operator and ensures the safety of the vehicle and personnel.
[0053] Based on the above embodiment, step S1 includes:
[0054] acquiring a target image of the detection target at regular intervals, and constructing a target classification image set based on the target image;
[0055] Obtaining an initial neural network model, simplifying and training the initial neural network model to obtain an optimized neural network model;
[0056] Training the optimized neural network model based on the target classification image set to obtain an instance segmentation model;
[0057] The target image is input into the instance segmentation model to complete target detection and tracking, to obtain a target frame of the target image, and the first position information is obtained based on the center point of the target frame.
[0058] The step of periodically acquiring the target image of the detection target and constructing the target classification image set based on the target image includes:
[0059] Using the binocular camera to collect multiple environmental images, annotating the multiple environmental images to obtain a target annotated image set;
[0060] Determine the classification of the target annotated image set, and obtain the target classified image set corresponding to the target annotated image set.
[0061] The obtaining of the initial neural network model, simplifying and training the initial neural network model to obtain the instance segmentation model includes:
[0062] Using preset coarse-grained parameter structured pruning to simplify the convolution kernel parameters of the initial neural network model to obtain an optimized neural network model;
[0063] The instance segmentation model is obtained by quantifying the preset network parameters and training the optimized neural network model.
[0064] The method of adopting preset coarse-grained parameter structured pruning to simplify the convolution kernel parameters of the initial neural network model to obtain an optimized neural network model includes:
[0065] An evaluation factor is set for the convolution kernel parameters or feature map parameters, or the preset part of the convolution kernel and any channel of the feature map are deleted to obtain the optimized neural network model.
[0066] The step of training the optimized neural network model by quantizing preset network parameters to obtain the instance segmentation model includes:
[0067] The network parameters of the neural network model are compressed using a uniform bit width or a combined bit width, and the instance segmentation model is obtained by training.
[0068] Specifically, intelligent working vehicles are usually in a state of constant speed. The binocular camera installed on the intelligent working vehicle reads the front environment information in real time, and identifies the working personnel and the corresponding position information through the trained target detection and tracking model. The model is deployed in the main controller of the binocular camera. Since the box regression mechanism has been added to the model during training, the identified detection target is segmented with a box. Here, considering the requirement of real-time segmentation of environmental images of intelligent working vehicles to avoid running jams and large consumption of main controller memory, the segmented image data is extracted every N frames.
[0069] After the center coordinates of the target frame are selected as the tracking point, the actual coordinates of the tracking point are converted according to the image pixel coordinates of the tracking point. Here, the midpoint of the binocular camera is calibrated as the coordinate origin, and the image pixel coordinates of the tracking point are obtained by eliminating the error between the left and right eye images.
[0070] Another example Figure 2As shown, before the binocular camera obtains the position information of the actual detection target, it is first necessary to prepare a data set for model training, including collecting images. Here, the image collection needs to comprehensively consider the influence of various factors in the orchard environment. Therefore, it is necessary to collect multiple images under different environments, such as strong light, weak light, and normal light images. Then, the collected images are annotated using the Lableme software, and classified according to the annotation results to obtain image data that only contains workers. Here, when processing the annotation file, the annotation file is named consistent with the original image. The Lableme software used in the present invention is a Javascript annotation tool for online image annotation. Compared with traditional image annotation tools, its advantage is that the tool can be used anywhere. In addition, it can also annotate images without installing or copying large datasets on the computer.
[0071] After selecting a suitable initial neural network, the instance segmentation model is compressed and accelerated to a certain extent.
[0072] Due to the complex structure and large number of parameters of deep neural networks, the trained models have long inference time and large memory requirements, which poses huge difficulties and challenges for their deployment on in-vehicle computing platforms with limited computing power.
[0073] The present invention uses pruning and quantization methods to achieve acceleration and compression of the model, including:
[0074] The model is accelerated using a coarse-grained parameter structured pruning method, where the smallest unit is the combination of parameters within the convolution kernel (filter). By setting evaluation factors for the filter or feature map, even entire filters or certain channels can be deleted to simplify the network and achieve acceleration on existing software and hardware.
[0075] Model compression is achieved through quantization of network parameters. The core idea is to use low-bit-width data instead of typical 32-bit floating-point network parameters such as weights, activation values, gradients, and errors. By using a uniform bit width (16-bit, 8-bit, etc.) or freely combining different bit widths to adjust network parameters, parameter storage and memory usage can be reduced, device energy consumption can be lowered, and computing speed can be accelerated.
[0076] After the model training is completed, the model is deployed to realize the detection and tracking of the detection target, calculate the image coordinates of the tracking point and convert them into actual coordinates, that is, obtain the first position information.
[0077] The present invention locates and tracks workers through a binocular camera installed on an intelligent work vehicle, selects a suitable neural network to train image data, and ensures high-quality instance segmentation accuracy and target tracking effect.
[0078] Based on any of the above embodiments, step S2 includes:
[0079] Setting an ultra-wideband mobile node at the detection target;
[0080] The inter-node relative distance between the ultra-wideband mobile node and the ultra-wideband base station node is acquired, and the second location information is obtained from the inter-node relative distance.
[0081] Specifically, to overcome the problem of target loss or reduced accuracy caused by lighting, vehicle bumps and occlusion in the orchard environment, UWB sensors are used as auxiliary relative posture perception devices to obtain the relative posture between the vehicle and the operator. UWB has the characteristics of low power consumption, insensitivity to channel fading (such as multipath, non-line-of-sight channels), strong anti-interference ability, no interference to other devices in the same environment, and strong penetration, which can effectively complement the vision of binocular cameras. Figure 3 As shown in the figure, the operator wears a UWB mobile node and the intelligent operation vehicle is equipped with a UWB base station node, which can obtain the relative position of the operator and the intelligent operation vehicle.
[0082] It's understandable that binocular cameras can accurately measure the relative position between the vehicle and the tracked person, but are significantly affected by lighting, road conditions, and occlusion. UWB, on the other hand, is unaffected by lighting, but the relative position accuracy it measures is relatively low. To build a stable and accurate target recognition and tracking system, the present invention utilizes a Kalman filter algorithm to fuse binocular camera and UWB data to obtain stable and accurate relative position information between the intelligent working vehicle and the tracked person. This fused relative position information serves as the input to the tracking controller, achieving stable tracking control.
[0083] The present invention adopts a multi-algorithm fusion and multi-hardware combination method to achieve high-precision target tracking of orchard transport vehicles.
[0084] Based on any of the foregoing embodiments, the binocular camera and the ultra-wideband base station are located on the same horizontal plane and the same center line on the same mobile platform.
[0085] The binocular camera and UWB base station node in the present invention are installed on the same mobile platform, that is, the intelligent working vehicle. At the same time, the UWB base station node and the binocular camera should be located on the same horizontal plane and the same center line, such as Figure 4 As shown in the figure, the center point of the UWB base station node and the midpoint of the binocular camera are both on the midpoint line of the dotted line shown in the figure. This setting is to facilitate the binocular camera and the UWB base station node to uniformly perform coordinate transformation when converting the acquired image of the detection target, which is convenient for subsequent fusion calculation.
[0086] In order to illustrate the practical effect of the technical solution of the present invention, a specific example is used for illustration and verification.
[0087] The main controller uses the NVIDIA JETSON AGX XAVIER with a high visual processing speed, and the binocular camera uses the ZED2 camera, which has the advantages of wide field of view and low distortion, and can feedback parameters such as object distance and angle; the UWB sensor uses the D-DWM-PG3.6 from Guangzhou Network Technology, which has the characteristics of small size, wide ranging range and small error, and can be used as a base station or tag.
[0088] Assume that the intelligent working vehicle is in an orchard environment with a driving speed of V. The ZED2 camera is used to read the orchard environment information in front. After the deployed instance segmentation model (Mask-RCNN neural network) is used to segment the image, the coordinate position (X cam ,Y cam ), the coordinates of the UWB node assembled by the operator relative to the UWB base station are (X uwb ,Y uwb ), and finally fused through the Kalman filter algorithm (X cam ,Y cam ) and (X uwb ,Y uwb ), and obtain a more accurate relative position between the intelligent vehicle and the tracked operator (X fuse ,Y fuse ) information to achieve accurate identification and stable tracking. The specific implementation results are as follows Figure 4 shown.
[0089] The target recognition and tracking system provided by the present invention is described below. The target recognition and tracking system described below and the target recognition and tracking method described above can be referenced to each other.
[0090] Figure 5 Schematic diagram of the target recognition and tracking system provided by the present invention. Figure 5 As shown, it includes: a first recognition module 51, a second recognition module 52 and a comprehensive recognition module 53, wherein:
[0091] The first recognition module 51 is used to obtain the target image of the detection target using a binocular camera, and obtain the first position information of the detection target after processing the target image using an instance segmentation model; the second recognition module 52 is used to determine the relative distance between the nodes of the detection target using ultra-wideband, and obtain the second position information of the detection target based on the relative distance between the nodes; the comprehensive recognition module 53 is used to fuse and calculate the first position information and the second position information to obtain the tracking position information of the detection target.
[0092] The present invention ensures the recognition accuracy of the intelligent transport vehicle for the detected target through the designed multi-sensor fusion target recognition and tracking system, improves the accuracy of the distance between the intelligent transport vehicle and the operator, indirectly improves the work efficiency of the operator and ensures the safety of the vehicle and personnel.
[0093] Figure 6 An example of a physical structure diagram of an electronic device is shown below. Figure 6 As shown, the electronic device may include: a processor 610, a communication interface 620, a memory 630, and a communication bus 640, wherein the processor 610, the communication interface 620, and the memory 630 communicate with each other via the communication bus 640. The processor 610 may call the logic instructions in the memory 630 to execute a target recognition and tracking method, which includes: using a binocular camera to acquire a target image of a detection target, using an instance segmentation model to process the target image to acquire first position information of the detection target; using ultra-wideband to determine the relative distance between nodes of the detection target, and acquiring second position information of the detection target based on the relative distance between nodes; and fusing and calculating the first position information and the second position information to obtain tracking position information of the detection target.
[0094] In addition, the logic instructions in the above-mentioned memory 630 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0095] On the other hand, the present invention also provides a computer program product, which includes a computer program, which can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the target recognition and tracking methods provided by the above methods, which include: using a binocular camera to obtain a target image of the detection target, and using an instance segmentation model to process the target image to obtain first position information of the detection target; using ultra-wideband to determine the relative distance between nodes of the detection target, and obtaining second position information of the detection target based on the relative distance between nodes; and fusing and calculating the first position information and the second position information to obtain tracking position information of the detection target.
[0096] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the target recognition and tracking method provided by the above-mentioned methods, the method comprising: using a binocular camera to obtain a target image of the detection target, and using an instance segmentation model to process the target image to obtain first position information of the detection target; using ultra-wideband to determine the relative distance between nodes of the detection target, and obtaining second position information of the detection target based on the relative distance between the nodes; and fusing and calculating the first position information and the second position information to obtain tracking position information of the detection target.
[0097] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.
[0098] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, or of course, by hardware. Based on this understanding, the essence of the above technical solution or the part that contributes to the existing technology can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or certain parts of the embodiments.
[0099] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A target recognition and tracking method, characterized in that: include: Acquire a target image of a detection target using a binocular camera, and process the target image using an instance segmentation model to obtain first position information of the detection target; Determining a relative distance between nodes of the detection target using ultra-wideband, and acquiring second position information of the detection target based on the relative distance between nodes; fusing and calculating the first position information and the second position information to obtain tracking position information of the detection target; The determining the relative distance between nodes of the detection target by using ultra-wideband, and acquiring the second position information of the detection target based on the relative distance between nodes, includes: Setting an ultra-wideband mobile node at the detection target; Acquire a relative distance between the ultra-wideband mobile node and an ultra-wideband base station node, and obtain the second location information based on the relative distance between the nodes; The binocular camera and the ultra-wideband base station are located on the same horizontal plane and the same center line on the same mobile platform; The method of acquiring a target image of a detection target by using a binocular camera and obtaining first position information of the detection target after processing the target image by using an instance segmentation model includes: acquiring a target image of the detection target at regular intervals, and constructing a target classification image set based on the target image; Obtaining an initial neural network model, simplifying and training the initial neural network model to obtain an optimized neural network model; Training the optimized neural network model based on the target classification image set to obtain an instance segmentation model; The target image is input into the instance segmentation model to complete target detection and tracking, to obtain a target frame of the target image, and the first position information is obtained based on the center point of the target frame.
2. The target recognition and tracking method according to claim 1, characterized in that: The step of regularly acquiring a target image of the detection target and constructing a target classification image set based on the target image includes: Using the binocular camera to collect multiple environmental images, annotating the multiple environmental images to obtain a target annotated image set; Determine the classification of the target annotated image set, and obtain the target classified image set corresponding to the target annotated image set.
3. The target recognition and tracking method according to claim 1, characterized in that: The obtaining of the initial neural network model, simplifying and training the initial neural network model to obtain the instance segmentation model includes: Using preset coarse-grained parameter structured pruning to simplify the convolution kernel parameters of the initial neural network model to obtain an optimized neural network model; The instance segmentation model is obtained by quantifying the preset network parameters and training the optimized neural network model.
4. The target recognition and tracking method according to claim 3, characterized in that: The method adopts preset coarse-grained parameter structured pruning to simplify the convolution kernel parameters of the initial neural network model to obtain an optimized neural network model, including: An evaluation factor is set for the convolution kernel parameters or feature map parameters, or a preset part of the convolution kernel and any channel of the feature map is deleted to obtain the optimized neural network model.
5. The target recognition and tracking method according to claim 3, characterized in that: The step of training the optimized neural network model by quantizing preset network parameters to obtain the instance segmentation model includes: The network parameters of the optimized neural network model are compressed using a uniform bit width or a combined bit width, and the instance segmentation model is obtained by training.
6. A target recognition and tracking system, characterized in that: include: A first recognition module is configured to obtain a target image of a detection target using a binocular camera, and to obtain first position information of the detection target after processing the target image using an instance segmentation model; a second identification module, configured to determine a relative distance between nodes of the detection target using ultra-wideband, and obtain second position information of the detection target based on the relative distance between nodes; a comprehensive identification module, configured to fuse and calculate the first position information and the second position information to obtain tracking position information of the detection target; The second identification module is specifically configured to: Setting an ultra-wideband mobile node at the detection target; Acquire a relative distance between the ultra-wideband mobile node and an ultra-wideband base station node, and obtain the second location information based on the relative distance between the nodes; The binocular camera and the ultra-wideband base station are located on the same horizontal plane and the same center line on the same mobile platform; The first identification module is specifically configured to: acquiring a target image of the detection target at regular intervals, and constructing a target classification image set based on the target image; Obtaining an initial neural network model, simplifying and training the initial neural network model to obtain an optimized neural network model; Training the optimized neural network model based on the target classification image set to obtain an instance segmentation model; The target image is input into the instance segmentation model to complete target detection and tracking, to obtain a target frame of the target image, and the first position information is obtained based on the center point of the target frame.
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the target recognition and tracking method according to any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Target positioning method and device, electronic equipment and storage medium
CN110889873A