Robot target identifying and positioning method and system based on deep learning

By analyzing multi-scale boundary gradients and channel response features, filtering attention migration paths, and optimizing spatial density fusion, the problem of overlapping and fuzzy boundary responses in multi-target scenarios in traditional robot target recognition and localization is solved, achieving higher accuracy and consistency in target localization and trajectory fusion.

CN121505403AInactive Publication Date: 2026-02-10HANGZHOU ELECTRONIC INFORMATION VOCATIONAL SCHOOL
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511582422.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-31
Publication Date
2026-02-10
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional robot target recognition and localization technologies suffer from overlapping and ambiguous boundary responses in multi-target scenarios. They lack fusion control over the spatial density and aggregation of multi-view feature points, leading to unstable recognition of overlapping target areas in complex environments and interference during multi-target trajectory output.

Method used

By analyzing multi-scale boundary gradient changes, selecting dominant channels, fusing boundary response parameters, locating the activation pixel cluster center, calculating the direction vector set, combining multi-view point cluster coordinate alignment, optimizing spatial density fusion, and adjusting target trajectory weights, target localization and trajectory fusion are achieved.

Benefits of technology

It enhances the accuracy of boundary extraction, optimizes the stability of attention transfer paths, improves the accuracy of spatial coordinate fusion, enhances the continuity of multi-target localization and the consistency of trajectory fusion, and strengthens the consistency of localization results in complex dynamic scenes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121505403A_ABST
    Figure CN121505403A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of electronic technology application specialty and computer specialty, in particular to a robot target identifying and positioning method and system based on deep learning, comprising the following steps: acquiring an input image and analyzing boundary gradient, extracting a dominant channel and fusing boundary response, positioning an aggregation center and expanding an attention area; the method comprises the following steps: collecting a multi-view point cluster and adjusting coordinates, positioning a stable base point to construct an interpolation region, extracting a center sequence and calculating trajectory change, generating a target trajectory fusion result, fusing multi-scale boundary gradient information and channel response characteristics, enhancing boundary extraction precision, and utilizing a space consistency direction screening mechanism to obtain a target trajectory fusion result. The stability of an attention migration path is optimized, multi-view point cluster distribution density and direction consistency are combined, space coordinate fusion accuracy is improved, an interpolation region fitting and target trajectory direction offset regulation and control method is adopted, multi-target positioning continuity and trajectory fusion consistency are improved, and multi-target positioning accuracy is improved. And the consistency and continuity of positioning results in a complex dynamic scene are enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of electronic technology application and computer technology, and in particular to a robot target recognition and localization method and system based on deep learning. Background Technology

[0002] Computer vision technology encompasses the techniques of acquiring, analyzing, and understanding image or video information using computers and sensing devices. It focuses on enabling computers to perceive and understand external visual information through processing steps such as image acquisition, feature extraction, object detection, image segmentation, and object recognition and localization. It broadly involves image processing, visual perception, pattern recognition, 3D reconstruction, and motion analysis. Its overall system includes multiple stages such as image acquisition, feature analysis, and result output. It acquires raw image data through visual sensors, extracts key feature information through algorithmic models, and achieves the recognition of target objects and the inference of spatial relationships, providing visual basis for robot control and environmental perception. Among these, deep learning-based robot target recognition and localization methods... This method refers to using deep neural network structures to perform target detection and spatial location estimation on image data collected by robots. Addressing the need for robots to identify and locate multiple objects in complex environments, the method covers technical aspects such as image data preprocessing, feature map generation, target detection network inference, and spatial coordinate calculation. Specifically, it involves using a convolutional neural network to perform multi-layer convolutional feature extraction on the input image to obtain the target's category information and bounding box position. Then, a 3D coordinate transformation algorithm combined with camera intrinsic and extrinsic parameters is used to map the detection results between pixels and world coordinates, thereby obtaining the target's spatial location information. Multi-view image matching and feature point correspondence calculation are combined to improve positioning accuracy, and a unified output of multi-target localization is achieved through coordinate fitting and geometric constraint methods.

[0003] Traditional robot target recognition and localization technologies rely on the overall gradient distribution of the image to extract target regions during the boundary feature processing stage. However, they do not conduct detailed analysis on the differences in boundary pixel distribution between channels, resulting in overlapping and blurred boundary responses in multi-target scenarios. This leads to interference issues in attention allocation and target separation. The 3D coordinate processing relies on basic projection transformation for viewpoint alignment and lacks fusion control for the spatial density and aggregation of feature points under multiple viewpoints. In complex environments, the recognition effect in overlapping target areas is unstable, and the output process of multi-target trajectories fails to consider the interference factors caused by changes in directional consistency, resulting in trajectory misalignment or recognition interruption when targets approach or intersect. Summary of the Invention

[0004] To address the technical problems existing in the prior art, embodiments of the present invention provide a robot target recognition and localization method based on deep learning, comprising the following steps: To achieve the above objectives, the present invention adopts the following technical solution: a robot target recognition and localization method based on deep learning, comprising the following steps: S1: Using image acquisition equipment, acquire input images, analyze gradient changes in multi-scale boundary regions, compare the distribution of boundary pixels in channels, filter dominant channels, determine gradient direction differences, match channel responses and fuse boundary regions to generate boundary gradient density response parameters. S2: Based on the boundary gradient density response parameters, locate the active pixel clustering center, calculate the set of direction vectors from the clustering center to each pixel in the neighborhood, compare the direction consistency features, filter the attention migration path, expand the attention coverage area and update the pixel response weights, and generate the focus attention migration configuration. S3: Based on the focus attention migration configuration, collect the coordinates of multi-viewpoint clusters, align the reference axis, compare the density and aggregation of point clusters, filter the spatial alignment reference, and perform coordinate correction and fusion adjustment in combination with the directional consistency between adjacent viewpoints to generate spatial density fusion parameters. S4: Based on the spatial density fusion parameters, analyze the spatial distribution of the three-dimensional coordinates of the point clusters in each local area, calculate the base point spacing, compare the inter-frame offset and density features, locate stable base points and construct the interpolation region, set the fitting benchmark, and generate target positioning base point information. S5: Based on the target positioning base point information, extract the center position sequence of each target in the image coordinate system in the continuous frame sequence, calculate the displacement direction and construct a vector sequence, analyze the trajectory change direction of each target direction vector, compare the directional offset of multiple target trajectory paths, adjust the weight distribution ratio of each target in the fusion stage, and generate the target trajectory fusion result.

[0005] As a further embodiment of the present invention, the boundary gradient density response parameters include channel gradient direction difference distribution, boundary pixel concentration intensity, and channel response matching coefficient; the focus attention migration configuration includes attention path direction set, pixel response weight distribution information, and spatial coverage information; the spatial density fusion parameters include point cluster coordinate alignment relationship data, spatial aggregation center information, and multi-view direction consistency features; the target positioning base point information includes base point spatial distribution structure, point cloud stability feature parameters, and interpolation region fitting benchmark; and the target trajectory fusion result includes trajectory direction change trend, target coordinate fusion weight ratio, and multi-target trajectory offset relationship.

[0006] As a further aspect of the present invention, the step of obtaining the boundary gradient density response parameters specifically includes: S101: Using an image acquisition device, acquire the input image and analyze the edge pixels of the multi-layer boundary region, calculate the gradient magnitude change trend of pixels under multiple scales, compare the concentration of boundary pixels in each channel, filter dense regions, obtain the dominant channel output, and generate the dominant channel output value. S102: Based on the output value of the dominant channel, determine the degree of difference in gradient direction between channels, and dynamically adjust the responses between channels by matching the responses between channels to generate feature mapping output parameters; S103: Based on the feature mapping output parameters, call the weight of each channel to perform response fusion on the boundary region, establish the response parameter set of the boundary region, and obtain the boundary gradient density response parameters.

[0007] As a further aspect of the present invention, the process of filtering dense regions specifically involves: constructing an equally spaced grid set within the boundary region of each channel of the input image, counting the number of edge pixels in each grid and normalizing it to a boundary pixel density value, sorting the boundary pixel density values ​​from largest to smallest to construct a boundary pixel density sorting sequence, comparing each grid density value with a boundary pixel aggregation determination threshold, filtering grids whose boundary pixel density exceeds the threshold, and merging them into a dense region set based on spatial adjacency. The process of obtaining the boundary pixel aggregation determination threshold is as follows: calculate the difference sequence of adjacent items in the boundary pixel density sorting sequence, locate the maximum transition point, and use the corresponding boundary pixel density as the boundary pixel aggregation determination threshold.

[0008] As a further aspect of the present invention, the step of obtaining the focus attention migration configuration specifically includes: S201: Based on the boundary gradient density response parameters, locate and analyze the boundary response concentration area, identify the spatial clustering center of continuously activated pixels, calculate the set of direction vectors from the clustering center to each pixel in the neighborhood, and obtain the pixel direction distribution parameters. S202: Based on the pixel orientation distribution parameters, compare the spatial consistency of orientation vectors, filter stable orientation regions as attention migration candidate paths, and generate migration path recognition results. S203: Based on the migration path identification result, update the original attention distribution along the continuous range of the path and expand the spatial coverage area, compare the pixel response intensity of the expanded area and reallocate the weights to establish the focus attention migration configuration.

[0009] As a further aspect of the present invention, the step of obtaining the spatial density fusion parameters specifically includes: S301: Based on the focus attention migration configuration, collect the feature point cluster coordinate set of the target area under multiple camera views, analyze the spatial angle between the camera projection direction of each group of point clusters and the target surface normal, calculate the rotation transformation direction of each point cluster and align it to the reference axis, and establish multi-view alignment coordinate parameters. S302: Based on the multi-view alignment coordinate parameters, compare the spatial distribution density and central aggregation degree between point clusters, analyze the density fluctuation of each point cluster, locate the spatial alignment reference based on the density difference fluctuation, and obtain the spatial alignment reference parameters. S303: Based on the aforementioned spatial alignment reference parameters, the consistency of point cluster orientation between adjacent viewpoints is judged, coordinates are corrected and fused, the spatial consistency of point clusters under multiple viewpoints is optimized, and spatial density fusion parameters are established.

[0010] As a further aspect of the present invention, the step of obtaining the target positioning base point information specifically includes: S401: Based on the spatial density fusion parameters, analyze the spatial arrangement of the three-dimensional coordinates of the point clusters in each local area, calculate the average spatial distance between candidate base points and neighboring points, identify the coverage relationship of the aggregation center point, and obtain the spatial distribution parameters of the base points; S402: Based on the spatial distribution parameters of the base points, compare the coordinate offset trends of the same points in consecutive frames, analyze the compactness of the point cloud distribution in multiple regions, determine the distribution stability of the region and locate the stable base points based on the combined characteristics of displacement change and density distribution, and obtain the stable base point positioning parameters. S403: Based on the stable base point positioning parameters, analyze the spatial coverage relationship between stable base points, construct a continuous interpolation region, set a fitting benchmark for the base points within the interpolation region, and establish target positioning base point information.

[0011] As a further aspect of the present invention, the step of obtaining the target trajectory fusion result specifically includes: S501: Based on the target positioning base point information, extract the center position of each target in the image coordinate system in the continuous frame sequence, establish the arrangement relationship of the center point of each target in each frame, and obtain the center position sequence parameters; S502: Based on the center position sequence parameters, calculate the displacement direction of each target in adjacent frames, construct the target frame direction vector sequence, analyze the rotation trend of each direction vector, determine the trajectory change direction, and obtain the trajectory change direction parameters; S503: Based on the trajectory change direction parameters, compare the directional offsets of multiple target trajectory paths within the same time period, sort the order of coordinate fusion participation according to the degree of directional difference, adjust the weight distribution ratio of each target in the fusion stage, and generate the target trajectory fusion result.

[0012] A deep learning-based robot target recognition and localization system includes: The boundary response construction module uses image acquisition equipment to acquire input images, analyze gradient changes in multi-scale boundary regions, compare the distribution of boundary pixels in channels, filter dominant channels, determine gradient direction differences, match channel responses and fuse boundary regions to generate boundary gradient density response parameters. The attention migration module, based on the boundary gradient density response parameters, locates the active pixel clustering center, calculates the set of direction vectors from the clustering center to each pixel in the neighborhood, compares the direction consistency features, filters the attention migration path, expands the attention coverage area and updates the pixel response weights, and generates the attention migration configuration. Based on the focal attention migration configuration, the spatial density calibration module collects multi-viewpoint cluster coordinates, aligns with the reference axis, compares the point cluster density and aggregation, filters the spatial alignment reference, and performs coordinate correction and fusion adjustment in combination with the directional consistency between adjacent viewpoints to generate spatial density fusion parameters. Based on the spatial density fusion parameters, the stable base point fitting module analyzes the spatial distribution of the three-dimensional coordinates of the point clusters in each local region, calculates the base point spacing, compares the inter-frame offset and density features, locates stable base points and constructs interpolation regions, sets the fitting benchmark, and generates target positioning base point information. Based on the target positioning base point information, the trajectory fusion calculation module extracts the center position sequence of each target in the image coordinate system in the continuous frame sequence, calculates the displacement direction and constructs a vector sequence, analyzes the trajectory change direction of each target direction vector, compares the directional offset of multiple target trajectory paths, adjusts the weight distribution ratio of each target in the fusion stage, and generates the target trajectory fusion result.

[0013] Compared with the prior art, the advantages and positive effects of the present invention are as follows: In this invention, the accuracy of boundary extraction is enhanced by fusing multi-scale boundary gradient information and channel response features. The stability of the attention migration path is optimized by using a spatial consistency direction filtering mechanism. The accuracy of spatial coordinate fusion is improved by combining the distribution density and direction consistency of multi-view point clusters. The continuity of multi-target positioning and the consistency of trajectory fusion are improved by using interpolation region fitting and target trajectory direction offset control methods. This enhances the consistency and continuity of positioning results in complex dynamic scenes. Attached Figure Description

[0014] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0015] Figure 1 This is a schematic diagram of the steps of the present invention; Figure 2 This is a detailed schematic diagram of S1 of the present invention; Figure 3 This is a detailed schematic diagram of S2 of the present invention; Figure 4 This is a detailed schematic diagram of S3 of the present invention; Figure 5 This is a detailed schematic diagram of S4 of the present invention; Figure 6 This is a detailed schematic diagram of S5 of the present invention; Figure 7 This is a system module diagram of the present invention. Detailed Implementation

[0016] The technical solution of the present invention will now be described with reference to the accompanying drawings.

[0017] In embodiments of the present invention, words such as "exemplarily," "for example," etc., are used to indicate that something is an example, illustration, or description. Any embodiment or design described as "exemplary" in the present invention should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of the word "exemplary" is intended to present the concept in a concrete manner. Furthermore, in embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one.

[0018] In the embodiments of this invention, the terms "image" and "picture" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, they convey the same meaning. Similarly, the terms "of," "corresponding (relevant)," and "corresponding" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, they convey the same meaning.

[0019] In this embodiment of the invention, sometimes a subscript such as W1 may be written in a non-subscript form such as W1. When the difference is not emphasized, the meaning they express is the same.

[0020] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.

[0021] Please see Figure 1 This invention provides a robot target recognition and localization method based on deep learning, comprising the following steps: S1: Using image acquisition equipment, acquire input images, analyze gradient changes in multi-scale boundary regions, compare the distribution of boundary pixels in channels, filter dominant channels, determine gradient direction differences, match channel responses and fuse boundary regions to generate boundary gradient density response parameters. S2: Based on the boundary gradient density response parameters, locate the cluster center of the active pixels, calculate the set of direction vectors from the cluster center to each pixel in the neighborhood, compare the direction consistency features, filter the attention migration path, expand the attention coverage area and update the pixel response weights, and generate the focus attention migration configuration. S3: Based on the focus attention migration configuration, collect the coordinates of multi-viewpoint clusters, align the reference axis, compare the density and aggregation of point clusters, filter the spatial alignment benchmark, and perform coordinate correction and fusion adjustment in combination with the directional consistency between adjacent viewpoints to generate spatial density fusion parameters. S4: Based on spatial density fusion parameters, analyze the spatial distribution of point cluster 3D coordinates in each local region, calculate the base point spacing, compare inter-frame offset and density features, locate stable base points and construct interpolation regions, set fitting benchmarks, and generate target positioning base point information; S5: Based on the target positioning base point information, extract the center position sequence of each target in the image coordinate system in the continuous frame sequence, calculate the displacement direction and construct a vector sequence, analyze the trajectory change direction of each target direction vector, compare the directional offset of multiple target trajectory paths, adjust the weight distribution ratio of each target in the fusion stage, and generate the target trajectory fusion result.

[0022] Boundary gradient density response parameters include channel gradient direction difference distribution, boundary pixel concentration intensity, and channel response matching coefficient. Focus attention migration configuration includes attention path direction set, pixel response weight distribution information, and spatial coverage information. Spatial density fusion parameters include point cluster coordinate alignment relationship data, spatial aggregation center information, and multi-view direction consistency features. Target localization base point information includes base point spatial distribution structure, point cloud stability feature parameters, and interpolation region fitting benchmark. Target trajectory fusion results include trajectory direction change trend, target coordinate fusion weight ratio, and multi-target trajectory offset relationship.

[0023] Please see Figure 2 The specific steps for obtaining the boundary gradient density response parameters are as follows: S101: Using an image acquisition device, acquire the input image and analyze the edge pixels of the multi-layer boundary region, calculate the gradient magnitude change trend of pixels under multiple scales, compare the concentration of boundary pixels in each channel, filter dense regions, obtain the dominant channel output, and generate the dominant channel output value. Using image acquisition equipment, an input image is acquired and the edge pixels of multiple boundary regions are analyzed. Specifically, an image with a resolution of [resolution value missing] is acquired. of Three-channel images, for , , The three channels are used respectively , , The Sobel operator at three scales calculates the gradient magnitude, with aisle Taking scale as an example, the gradient magnitude map is calculated. Subsequently Above the construction A grid of equally spaced pixels, totaling Count the number of grid cells, and count the gradient magnitude within each grid cell. Number of edge pixels, grid The number of pixels is Then the normalized boundary pixel density pass Divide by total grid pixels Get, get all Value (total) Sort the pixels from largest to smallest to obtain the boundary pixel density sorting sequence. Next, calculate the difference sequence of adjacent terms in the sequence. For example, the difference between the second term and the first term is The difference between the third term and the second term is The difference between the fourth and third terms is ,get The maximum transition point is located as follows: The boundary pixel density value corresponding to this transition point is Therefore, a threshold for determining boundary pixel aggregation is set. All grid densities and Compare and filter Grids, such as grids , , The densities are respectively , , All of them exceeded the threshold, and they were spatially adjacent, so they were merged into a dense region. Repeat the above process for all channels and all scales of this image, and finally compare all channels at... The number of grid cells exceeding the threshold in the region contribution, if The channel contributed indivual, The channel contributed indivual, The channel contributed One, then The channel is selected as the dominant channel for this area, and the dominant channel output value is generated.

[0024] S102: Based on the output value of the dominant channel, determine the degree of difference in gradient direction between channels, and dynamically adjust the responses between channels by matching the responses between channels to generate feature mapping output parameters; Based on the output value of the dominant channel, i.e., the region The dominant channel is Channels and their corresponding grid information, in Each pixel within the region To determine the degree of difference in gradient direction between channels, for example at a pixel. At the location, calculate gradient direction of the channel , channel , channel Set a fixed inter-channel directional matching threshold. This threshold is set based on the signal-to-noise ratio characteristics of the image sensor. For currently used high-definition CMOS sensors, the directional noise fluctuation range is within... Therefore, it is set inside. ,calculate Channel and , Differences in the direction of the channels and The absolute difference is , and The absolute difference is The difference value is compared with Compare, ,determination Channels and Dominant Channels match, ,determination Channels and Dominant Channels Mismatches are identified by matching responses between channels. , , The original gradient magnitude response of the three channels at this pixel. Dynamically adjust the response between channels for matching. The channel, its adjusted response The calculation process is as follows: using Subtract directional differences Divide by the base value The quotient, multiplied by the original response. ,get For mismatches The channel, its adjusted response The calculation process is as follows: using Subtract directional differences Divide by the base value The quotient, multiplied by the original response. ,get Dominant channel The response remains unchanged. ,right This adjustment process is repeated for all pixels within the region to generate feature map output parameters.

[0025] S103: Based on the feature mapping output parameters, call the weight of each channel to perform response fusion on the boundary region, establish the response parameter set of the boundary region, and obtain the boundary gradient density response parameters; Based on the feature mapping output parameters, i.e., the region Channel response after adjusting all pixels For example, pixels The response value at that location Call the weight of each channel, that channel weight It is not a fixed value, but rather dynamically calculated based on the contribution of each channel in S101 to the dense area. The specific calculation process is as follows: In all dense areas of S101, those determined to be... Channel-dominated total number of grids , Channel-dominated total number of grids , Channel-dominated total number of grids Assuming the statistical results are Calculate the total number of grid cells The fusion weights for each channel are then set to... for Divide by ,Right now , for Divide by ,Right now , for Divide by ,Right now Using these weights, response fusion is performed on the boundary region to calculate pixels. The final fusion response at the location Its calculation method is as follows Multiply , plus Multiply In addition Multiply Substitute the values: The products of each term are respectively , and The sum of the three is: This weighted summation operation is performed on all adjusted pixel responses to establish a set of response parameters for the boundary region, which is a fused 2D response map. The boundary gradient density response parameters are obtained.

[0026] Please see Figure 3 The specific steps to obtain the focus migration configuration are as follows: S201: Based on the boundary gradient density response parameters, locate and analyze the concentrated region of the boundary response, identify the spatial clustering center of continuously activated pixels, calculate the set of direction vectors from the clustering center to each pixel in the neighborhood, and obtain the pixel direction distribution parameters. Based on the boundary gradient density response parameters, i.e., the fused response map To locate and analyze the concentrated area of ​​boundary response, an activation threshold is first set. This threshold is used to filter pixels with strong responses. Settings reference Maximum response value in the figure Query get , Set as A fixed ratio, for example ,Right now ,get traversal Filter all response values These pixels constitute "continuously active pixels," and the spatial clustering center of these active pixels is identified by calculating these active pixels. This is achieved by finding the coordinate mean (centroid), assuming a centroid containing... The centroid (center of clustering) of a cluster of active pixels. The calculation method is as follows: Equal to all Sum of coordinates divided by , Equal to all Sum of coordinates divided by ,get Next, define a... Centered Pixel neighborhood, calculating cluster center To each active pixel in that neighborhood (common A set of direction vectors ,in pass coordinates minus The coordinates are obtained, for example, the coordinates of an active pixel in the neighborhood. Its response Then calculate its direction vector. The x-component is The y-component is ,Right now For all This calculation is repeated for each active pixel to obtain the pixel orientation distribution parameters.

[0027] S202: Based on the pixel orientation distribution parameters, compare the spatial consistency of orientation vectors, filter stable orientation regions as candidate paths for attention migration, and generate migration path recognition results. Based on pixel orientation distribution parameters, i.e., including A set of vectors (For example To compare the spatial consistency of direction vectors, first calculate this... The average vector of vectors Its x-component is all Sum divided by y component is all Sum divided by Assuming the calculation yields Its average direction (angle) Through calculation and The reverse tangent Then calculate each vector direction ,For example direction Through calculation and The reverse tangent , direction Through calculation and The reverse tangent Then, to filter stable directional regions, an angular consistency threshold needs to be set. The threshold Based on all directions in the set Angular standard deviation To dynamically set, assuming the calculation is obtained ,but Set as A specific proportion, for example times, that is ,get Compare each and The absolute difference and The absolute difference is , ,but It belongs to the stable direction region. and The absolute difference (after processing the angle surround) is , ,but Pixels that do not belong to the stable direction region The path is retained and used as a candidate path for attention migration, generating migration path identification results.

[0028] S203: Based on the migration path recognition results, update the original attention distribution along the continuous range of the path and expand the spatial coverage area, compare the pixel response intensity of the expanded area and reallocate the weights to establish the focus attention migration configuration; Based on the migration path identification results, i.e., the set of stable direction pixels (e.g., containing...) (pixels) and average vector , will this The set of pixels is used as the "original attention distribution". The original attention distribution is updated along a continuous range of the path, which is the path formed by... The defined direction, from the center of aggregation Begin, along Directional step, step size is iteration Next, a series of new center points are generated. , equal Plus Multiply The product, for example ,get , , , In each Define a surrounding area The neighborhood will Each new pixel is added to the attention distribution, forming an "extended spatial coverage area". (Include And newly added (pixels), compare extended area The response intensity of all pixels within, i.e., from S103. Query these pixels Value, found Maximum response within Assuming And reassign weights to establish the final focus attention transfer configuration (an attention graph). ),for any pixel within Its weight equal Divide by ,For example ( )of Value Then its weight is , while pixel weights outside the region for Establish a focus attention migration configuration.

[0029] Please see Figure 4 The specific steps for obtaining spatial density fusion parameters are as follows: S301: Based on the focus attention transfer configuration, collect the coordinate set of feature point clusters of the target area under multiple camera views, analyze the spatial angle between the camera projection direction of each point cluster and the normal of the target surface, calculate the rotation transformation direction of each point cluster and align it to the reference axis, and establish multi-view alignment coordinate parameters. Based on focus-attention transfer configuration, i.e., 2D attention graph The figure indicates the highlighted areas (weighted) of a target (e.g., a "T-joint") in the image. ,For example The weight is The method involves using three cameras (Cam1, Cam2, Cam3) to simultaneously acquire the coordinate set of feature point clusters in the target area. Specifically, on the 2D image from each camera, only... ORB feature points are extracted within the region, and the three-dimensional coordinates of these feature points (based on their respective camera coordinate systems) are calculated using triangulation to obtain three point cluster sets. Analyze the spatial angle between the camera projection direction and the target surface normal for each point cluster. First, in (Assuming the region is from Cam1) A plane is fitted to obtain the normal to the target surface. (In the Cam1 coordinate system), obtain the projection direction (Z-axis) of Cam1. The (calibrated) relative projection direction of Cam2 Calculate the included angle in space Through calculation and The dot product (the result is) Arccos is obtained ,calculate Through calculation and The dot product (the result is) Arccos is obtained Calculate the rotational transformation direction of each point cluster and align it to the reference axis, then set the robot's base coordinate system Z-axis. For the reference axis, the calculation will Rotate to Parallel rotation matrix (Through the axis-angle method), and Applied to All points in the array, obtain Simultaneously calculate the transformation matrix from Cam2 to Cam1. ,Will Transform the points in the image to the Cam1 coordinate system, and then apply... ,Right now equal Multiply Multiply The matrix product, for Perform a similar operation to establish multi-view alignment coordinate parameters.

[0030] S302: Based on the multi-view alignment coordinate parameters, compare the spatial distribution density and central aggregation degree between point clusters, analyze the density fluctuation of each point cluster, locate the spatial alignment reference based on the density difference fluctuation, and obtain the spatial alignment reference parameters. Based on the multi-view alignment coordinate parameters, the three point clusters already aligned to the robot's base coordinate system are used. Comparing the spatial distribution density and central aggregation degree among point clusters, firstly in The density of each point cluster in the voxel mesh is statistically analyzed, and the calculation is performed. Standard deviation of all voxel densities Assuming ,calculate center of mass And all the points to average distance Assuming ,right and Performing the same calculation yields... as well as Analyze the density fluctuation of each point cluster (i.e. Value) and degree of central aggregation (i.e. (Value), based on density difference fluctuations, a spatial alignment reference is established, and a quality assessment function is set. , equal Divide by The merchants, plus Divide by The merchants, among which and As a weighting factor, according to the calibration of this system, the reliability of density fluctuations takes precedence over the degree of aggregation, and is set as follows: Calculate the quality score for each point cluster: The calculation is as follows (Right now ) plus (Right now The sum is , The calculation is as follows (Right now ) plus (Right now The sum is , The calculation is as follows (Right now ) plus (Right now The sum is Compare the three scores. To get the highest score, therefore choose As a spatial alignment reference, the spatial alignment reference parameters are obtained.

[0031] S303: Combining spatial alignment reference parameters, the consistency of point cluster orientation between adjacent viewpoints is judged, coordinates are corrected and fused, the spatial consistency of point clusters under multiple viewpoints is optimized, and spatial density fusion parameters are established. Combining spatial alignment reference parameters, based on the reference point cluster For adjacent viewpoints (i.e. and , and The consistency of the direction of the point clusters is judged. and In the overlapping region, find the nearest neighbor pair, for example... and Calculate separately Local normal vector (using its) (nearest points) and Local normal vector Assuming , Calculate its dot product ,Right now ,get Set a normal vector consistency threshold. This threshold is set based on the angle of view analysis in S301. The cosine value, i.e. Compare the dot product results. If the directions of the two pairs of points are consistent, then the other pair of points... If the coordinates are not consistent, then the coordinates are corrected and merged, and all point clusters are merged. To generate the final point cloud The quality score calculated in S302 is used during fusion. As weights, normalized weights equal Divide by Calculations yielded , , For consistent points in overlapping regions (such as...) ), its fused coordinates The calculation method is as follows Multiply , plus Multiply In addition Multiply For points of inconsistency (such as...) ), its weight The value at this point is set to 0; for points in non-overlapping regions, they are directly merged. To optimize the spatial consistency of point clusters under multiple perspectives, spatial density fusion parameters are established.

[0032] Please see Figure 5 The specific steps for obtaining target positioning base point information are as follows: S401: Based on spatial density fusion parameters, analyze the spatial arrangement of the three-dimensional coordinates of point clusters in each local area, calculate the average spatial distance between candidate base points and neighboring points, identify the coverage relationship of aggregation center points, and obtain the spatial distribution parameters of base points; Based on spatial density fusion parameters, extract frames The final fused point cloud Analyze the spatial arrangement of the three-dimensional coordinates of point clusters within each local region, starting from... Candidate base points are selected, for example, by voxel lattice downsampling, selecting the center point of each voxel as a candidate base point. ,by (unit: For example, let's define a radius. Calculate the spherical local region. With all in this region Neighboring points The average spatial spacing is calculated. To each Euclidean distance , obtain the distance set Calculate the average spacing That is, all The sum (assuming it is) Divide by ,get Identify the coverage relationship of the aggregation center point and calculate this. Neighboring points geometric centroid That is, all average value average value The average value, assuming the calculation yields ,calculate Distance from the center of mass ,Right now and The Euclidean distance between them is calculated as follows: squared plus squared plus The square root of the square is... ,get ,right Repeat this process for all candidate base points to obtain the spatial distribution parameters of the base points.

[0033] S402: Based on the spatial distribution parameters of the base points, compare the coordinate offset trends of the same points in consecutive frames, analyze the compactness of the point cloud distribution in multiple regions, determine the distribution stability of the region and locate the stable base points based on the combined characteristics of displacement change and density distribution, and obtain the stable base point positioning parameters. Based on the spatial distribution parameters of the base points, i.e., the frame All candidate base points and their parameters and retrieve the previous frame. Corresponding parameters By comparing the coordinate offset trends of the same points within consecutive frames, and finding the nearest neighbor match... exist Corresponding points of frames Calculate its coordinate offset ,Right now and The Euclidean distance is calculated as follows: squared plus squared plus The square root of the square is... ,get Analyze the compactness of point cloud distribution in multiple regions, i.e., compare Changes, , Change in compactness It is the absolute value of the difference between the two, that is According to displacement change With density distribution (changes in compactness) The combined characteristics of the region are used to determine the distribution stability of the region and locate the stable base point, and a displacement change threshold is set. and compact change threshold , The settings are based on the motion estimation accuracy of the robot itself. If the motion estimation error is... Then set , The settings are based on the sensor's depth resolution. ,judge : and Both conditions are met, therefore It is positioned as a stable base point, if another point of ,but For unstable points, obtain the positioning parameters of stable reference points.

[0034] S403: Based on the stable base point positioning parameters, analyze the spatial coverage relationship between stable base points, construct a continuous interpolation region, set a fitting benchmark for the base points within the interpolation region, and establish target positioning base point information; Extracting a set of stable base points based on the location parameters of stable base points Analyze the spatial coverage relationship between stable reference points, and... The dataset undergoes 3DDelaunay triangulation to construct a tetrahedral mesh with these stable base points as vertices. This mesh defines the core stable volume of the target (T-joint). A continuous interpolation region is constructed, which is the total volume enclosed by the aforementioned tetrahedral mesh. A fitting reference is set for the base points within the interpolation region, and the RANSAC algorithm is called. Find the best-fit geometric primitive (plane or cylinder) in the set and set the RANSAC distance threshold. This threshold is referenced from the displacement threshold in S402. Settings, take ,exist In the nth iteration, the 1st In the next iteration, the three randomly sampled points form a plane. Its plane equation is ,calculate The distance from all points to the plane was found to be... The point distance is less than (i.e., "interior point"), in other iterations (e.g., fitting a cylinder), did not exceed The interior point scale is such that a plane is chosen. As a fitting benchmark, target positioning base point information is established.

[0035] Please see Figure 6The specific steps for obtaining the target trajectory fusion result are as follows: S501: Based on the target positioning base point information, extract the center position of each target in the image coordinate system in the continuous frame sequence, establish the arrangement relationship of the center point of each target in each frame, and obtain the center position sequence parameters; Based on the target positioning reference point information, using the fitting reference (plane) ) and the components constituting this benchmark Set of stable base points (interior points) This is target A (T-connector) in frame. The information is used to extract the center position of each target in the image coordinate system in a continuous frame sequence. First, the center position of target A in the frame is calculated. 3D Center ,Right now The mean of the 3D coordinates is obtained. Then, the camera intrinsic parameter matrix of Cam1 was used. and extrinsic parameter matrix Project the 3D center onto the 2D image coordinate system of Cam1. The calculation method is to use Multiply by left in sequence and Calculations yielded (Unit: pixels), for consecutive Perform this operation over three frames to obtain the center position sequence of target A. Meanwhile, there is a target B (a bolt) in the scene. Repeat all the steps from S1 to S501 on it to obtain the center position sequence of target B. Establish the arrangement relationship of each target center point in each frame to obtain the center position sequence parameters.

[0036] S502: Based on the center position sequence parameters, calculate the displacement direction of each target in adjacent frames, construct the target frame direction vector sequence, analyze the rotation trend of each direction vector, determine the trajectory change direction, and obtain the trajectory change direction parameters; Based on the center position sequence parameters, obtain two target sequences. and Calculate the displacement direction of each target in adjacent frames, and construct a sequence of target orientation vectors between frames. For target A: calculate frame... and Displacement vector between ,pass coordinates minus The coordinates are obtained, that is Calculate frames and Displacement vector between ,get For target B: , ,get Analyze the rotation trend of each direction vector and calculate the angle of each vector. , Through calculation and The reverse tangent , Through calculation and The reverse tangent , Through calculation and The reverse tangent , Through calculation and The reverse tangent Calculate the change in angle , equal minus ,Right now , equal minus ,Right now Determine the direction of trajectory change and set a trajectory rotation threshold. This threshold is calculated based on pixel jitter in the 2D projection of S501 and is set to... ,Compare The absolute value, The trajectory of target A is determined to be a "straight line". The trajectory of target B is determined to be "turning", and the trajectory change direction parameters are obtained.

[0037] S503: Based on the trajectory change direction parameter, compare the directional offset of multiple target trajectory paths within the same time period, sort the order of coordinate fusion participation according to the degree of directional difference, adjust the weight distribution ratio of each target in the fusion stage, and generate the target trajectory fusion result; Based on the trajectory change direction parameter, according to the trajectory change of target A (Straight line) and target B (Turn around), compare the same time period (frames) arrive The directional offset of multiple target trajectory paths within a given area, i.e., comparison. and Based on the degree of directional difference (i.e. The order in which the coordinates of the sorting (size) are integrated is determined by the following sorting rule: The smaller the value (the more stable the trajectory), the higher the priority, because... Therefore, the order of participation in the fusion process is determined as: [Target A (Priority1), Target B (Priority2)]. The weight distribution ratio of each target in the fusion stage is adjusted, and this weight is used in the next frame. In the state prediction model, weights The settings refer to its instability (i.e. To avoid division by zero, a minimal stable quantity is introduced. Weight The calculation method is as follows: First, calculate the inverse instability. , equal Divide by and The sum, then Calculate the inverse instability of target A by dividing by the sum of the inverse instabilities of all targets. ,get Calculate the inverse instability of target B. ,get Calculate the sum Normalization yields the weights: , These weights will be used to update the state covariance of the multi-target tracking filter and generate the target trajectory fusion result.

[0038] Please see Figure 7 A deep learning-based robot target recognition and localization system includes: The boundary response construction module uses image acquisition equipment to acquire input images, analyze gradient changes in multi-scale boundary regions, compare the distribution of boundary pixels in channels, filter dominant channels, determine gradient direction differences, match channel responses and fuse boundary regions to generate boundary gradient density response parameters. The attention migration module locates the cluster center of active pixels based on the boundary gradient density response parameters, calculates the set of direction vectors from the cluster center to each pixel in the neighborhood, compares the direction consistency features, filters the attention migration path, expands the attention coverage area and updates the pixel response weights, and generates the attention migration configuration. The spatial density calibration module is based on the focus attention migration configuration, collects the coordinates of point clusters from multiple perspectives, aligns them with the reference axis, compares the density and aggregation of point clusters, filters the spatial alignment reference, and performs coordinate correction and fusion adjustment in combination with the directional consistency between adjacent perspectives to generate spatial density fusion parameters. The stable base point fitting module analyzes the spatial distribution of the three-dimensional coordinates of point clusters in each local region based on spatial density fusion parameters, calculates the base point spacing, compares inter-frame offset and density features, locates stable base points and constructs interpolation regions, sets fitting benchmarks, and generates target positioning base point information. The trajectory fusion calculation module extracts the center position sequence of each target in the image coordinate system based on the target localization base point information, calculates the displacement direction and constructs a vector sequence, analyzes the trajectory change direction of each target direction vector, compares the directional offset of multiple target trajectory paths, adjusts the weight distribution ratio of each target in the fusion stage, and generates the target trajectory fusion result.

[0039] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A robot target recognition and localization method based on deep learning, characterized in that, Includes the following steps: S1: Using image acquisition equipment, acquire input images, analyze gradient changes in multi-scale boundary regions, compare the distribution of boundary pixels in channels, filter dominant channels, determine gradient direction differences, match channel responses and fuse boundary regions to generate boundary gradient density response parameters. S2: Based on the boundary gradient density response parameters, locate the active pixel clustering center, calculate the set of direction vectors from the clustering center to each pixel in the neighborhood, compare the direction consistency features, filter the attention migration path, expand the attention coverage area and update the pixel response weights, and generate the focus attention migration configuration. S3: Based on the focus attention migration configuration, collect the coordinates of multi-viewpoint clusters, align the reference axis, compare the density and aggregation of point clusters, filter the spatial alignment reference, and perform coordinate correction and fusion adjustment in combination with the directional consistency between adjacent viewpoints to generate spatial density fusion parameters. S4: Based on the spatial density fusion parameters, analyze the spatial distribution of the three-dimensional coordinates of the point cluster in each local area, calculate the base point spacing, compare the inter-frame offset and density features, locate stable base points and construct the interpolation region, set the fitting benchmark, and generate target positioning base point information.

2. The robot target recognition and localization method based on deep learning according to claim 1, characterized in that, The boundary gradient density response parameters include channel gradient direction difference distribution, boundary pixel concentration intensity, and channel response matching coefficient. The focus attention migration configuration includes attention path direction set, pixel response weight distribution information, and spatial coverage information. The spatial density fusion parameters include point cluster coordinate alignment relationship data, spatial aggregation center information, and multi-view direction consistency features. The target positioning base point information includes base point spatial distribution structure, point cloud stability feature parameters, and interpolation region fitting benchmark.

3. The robot target recognition and localization method based on deep learning according to claim 1, characterized in that, The specific steps for obtaining the boundary gradient density response parameters are as follows: S101: Using an image acquisition device, acquire the input image and analyze the edge pixels of the multi-layer boundary region, calculate the gradient magnitude change trend of pixels under multiple scales, compare the concentration of boundary pixels in each channel, filter dense regions, obtain the dominant channel output, and generate the dominant channel output value. S102: Based on the output value of the dominant channel, determine the degree of difference in gradient direction between channels, and dynamically adjust the responses between channels by matching the responses between channels to generate feature mapping output parameters; S103: Based on the feature mapping output parameters, call the weight of each channel to perform response fusion on the boundary region, establish the response parameter set of the boundary region, and obtain the boundary gradient density response parameters.

4. The robot target recognition and localization method based on deep learning according to claim 3, characterized in that, The process of filtering dense regions is as follows: construct an equally spaced grid set within the boundary region of each channel of the input image, count the number of edge pixels in each grid and normalize it to a boundary pixel density value, sort the boundary pixel density values ​​from largest to smallest, construct a boundary pixel density sorting sequence, compare the density values ​​of each grid with the boundary pixel aggregation judgment threshold, filter grids whose boundary pixel density exceeds the threshold, and merge them into a dense region set according to spatial adjacency. The process of obtaining the boundary pixel aggregation determination threshold is as follows: calculate the difference sequence of adjacent items in the boundary pixel density sorting sequence, locate the maximum transition point, and use the corresponding boundary pixel density as the boundary pixel aggregation determination threshold.

5. The robot target recognition and localization method based on deep learning according to claim 3, characterized in that, The specific steps for obtaining the focus attention migration configuration are as follows: S201: Based on the boundary gradient density response parameters, locate and analyze the boundary response concentration area, identify the spatial clustering center of continuously activated pixels, calculate the set of direction vectors from the clustering center to each pixel in the neighborhood, and obtain the pixel direction distribution parameters. S202: Based on the pixel orientation distribution parameters, compare the spatial consistency of orientation vectors, filter stable orientation regions as attention migration candidate paths, and generate migration path recognition results. S203: Based on the migration path identification result, update the original attention distribution along the continuous range of the path and expand the spatial coverage area, compare the pixel response intensity of the expanded area and reallocate the weights to establish the focus attention migration configuration.

6. The robot target recognition and localization method based on deep learning according to claim 5, characterized in that, The specific steps for obtaining the spatial density fusion parameters are as follows: S301: Based on the focus attention migration configuration, collect the feature point cluster coordinate set of the target area under multiple camera views, analyze the spatial angle between the camera projection direction of each group of point clusters and the target surface normal, calculate the rotation transformation direction of each point cluster and align it to the reference axis, and establish multi-view alignment coordinate parameters. S302: Based on the multi-view alignment coordinate parameters, compare the spatial distribution density and central aggregation degree between point clusters, analyze the density fluctuation of each point cluster, locate the spatial alignment reference based on the density difference fluctuation, and obtain the spatial alignment reference parameters. S303: Based on the aforementioned spatial alignment reference parameters, the consistency of point cluster orientation between adjacent viewpoints is judged, coordinates are corrected and fused, the spatial consistency of point clusters under multiple viewpoints is optimized, and spatial density fusion parameters are established.

7. The robot target recognition and localization method based on deep learning according to claim 6, characterized in that, The specific steps for obtaining the target positioning base point information are as follows: S401: Based on the spatial density fusion parameters, analyze the spatial arrangement of the three-dimensional coordinates of the point clusters in each local area, calculate the average spatial distance between candidate base points and neighboring points, identify the coverage relationship of the aggregation center point, and obtain the spatial distribution parameters of the base points; S402: Based on the spatial distribution parameters of the base points, compare the coordinate offset trends of the same points in consecutive frames, analyze the compactness of the point cloud distribution in multiple regions, determine the distribution stability of the region and locate the stable base points based on the combined characteristics of displacement change and density distribution, and obtain the stable base point positioning parameters. S403: Based on the stable base point positioning parameters, analyze the spatial coverage relationship between stable base points, construct a continuous interpolation region, set a fitting benchmark for the base points within the interpolation region, and establish target positioning base point information.

8. The robot target recognition and localization method based on deep learning according to claim 1, characterized in that, The method further includes: S5: Based on the target positioning base point information, extract the center position sequence of each target in the image coordinate system in the continuous frame sequence, calculate the displacement direction and construct a vector sequence, analyze the trajectory change direction of each target direction vector, compare the directional offset of multiple target trajectory paths, adjust the weight distribution ratio of each target in the fusion stage, and generate the target trajectory fusion result. The target trajectory fusion result includes the trajectory direction change trend, target coordinate fusion weight ratio, and multi-target trajectory offset relationship.

9. The robot target recognition and localization method based on deep learning according to claim 8, characterized in that, The specific steps for obtaining the target trajectory fusion result are as follows: S501: Based on the target positioning base point information, extract the center position of each target in the image coordinate system in the continuous frame sequence, establish the arrangement relationship of the center point of each target in each frame, and obtain the center position sequence parameters; S502: Based on the center position sequence parameters, calculate the displacement direction of each target in adjacent frames, construct the target frame direction vector sequence, analyze the rotation trend of each direction vector, determine the trajectory change direction, and obtain the trajectory change direction parameters; S503: Based on the trajectory change direction parameters, compare the directional offsets of multiple target trajectory paths within the same time period, sort the order of coordinate fusion participation according to the degree of directional difference, adjust the weight distribution ratio of each target in the fusion stage, and generate the target trajectory fusion result.

10. A robot target recognition and localization system based on deep learning, characterized in that, The system is used to implement the deep learning-based robot target recognition and localization method according to any one of claims 1-9, and the system comprises: The boundary response construction module uses image acquisition equipment to acquire input images, analyze gradient changes in multi-scale boundary regions, compare the distribution of boundary pixels in channels, filter dominant channels, determine gradient direction differences, match channel responses and fuse boundary regions to generate boundary gradient density response parameters. The attention migration module, based on the boundary gradient density response parameters, locates the active pixel clustering center, calculates the set of direction vectors from the clustering center to each pixel in the neighborhood, compares the direction consistency features, filters the attention migration path, expands the attention coverage area and updates the pixel response weights, and generates the attention migration configuration. Based on the focal attention migration configuration, the spatial density calibration module collects multi-viewpoint cluster coordinates, aligns with the reference axis, compares the point cluster density and aggregation, filters the spatial alignment reference, and performs coordinate correction and fusion adjustment in combination with the directional consistency between adjacent viewpoints to generate spatial density fusion parameters. Based on the spatial density fusion parameters, the stable base point fitting module analyzes the spatial distribution of the three-dimensional coordinates of the point clusters in each local region, calculates the base point spacing, compares the inter-frame offset and density features, locates stable base points and constructs interpolation regions, sets the fitting benchmark, and generates target positioning base point information. Based on the target positioning base point information, the trajectory fusion calculation module extracts the center position sequence of each target in the image coordinate system in the continuous frame sequence, calculates the displacement direction and constructs a vector sequence, analyzes the trajectory change direction of each target direction vector, compares the directional offset of multiple target trajectory paths, adjusts the weight distribution ratio of each target in the fusion stage, and generates the target trajectory fusion result.