Target association method, device and equipment for compact shelving complex background and medium

By combining visual and semantic feature extraction, feature fusion, and dynamic time warping algorithms with Kalman filtering in the complex background of mobile shelving, the problem of low target association accuracy in mobile shelving environments is solved, the continuity and consistency of target trajectories are achieved, and the stability of target recognition and tracking is improved.

CN121010777AInactive Publication Date: 2025-11-25ZHEJIANG BEITAI INTELLIGENT TECH CO LTD

Patent Information

Application Number
CN202511525237.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-24
Publication Date
2025-11-25
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the complex context of mobile shelving, existing technologies struggle to maintain the continuity of target trajectories, resulting in low target association accuracy. In particular, when a target is obscured or briefly leaves the field of vision, identity switching and trajectory breakage are likely to occur.

Method used

Visual and semantic feature vectors are obtained through feature extraction algorithms. Combined with long short-term memory networks and large visual-language models, feature fusion networks are used for feature fusion, dynamic time warping algorithms are used for temporal alignment, and Kalman filtering is used to predict target positions to ensure the continuity and consistency of the trajectory.

Benefits of technology

It improves the accuracy and robustness of target recognition, ensures the continuity and consistency of trajectories, reduces trajectory breakage issues, and enhances the accuracy of target association.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121010777A_ABST
    Figure CN121010777A_ABST
Patent Text Reader

Abstract

The invention provides a compact shelving complex background-oriented target association method, device and equipment and a medium, and relates to the technical field of computers, the method comprises the following steps: carrying out feature extraction on environment image data to obtain a visual feature vector, inputting the environment image data into a visual language large model to obtain a semantic feature vector, and carrying out feature extraction on the semantic feature vector; feature extraction is carried out on the motion data of the robot relative to the target through a long-short-term memory network, a motion feature vector is obtained, and the environment image data comprises the target space position of the target; performing feature fusion on the visual feature vector, the semantic feature vector and the motion feature vector through a feature fusion network to obtain a semantic enhanced feature vector; obtaining a target matching track of the target by using a dynamic time warping algorithm, and predicting a target prediction position of the target through Kalman filtering according to the target matching track; and obtaining a target association trajectory according to the target spatial position and the target prediction position. According to the invention, the association precision of target association can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer, in particular to a target association method and device for dense shelf complex background, equipment and medium. BACKGROUND

[0002] In the field of computer vision and automation, vision-based target recognition and tracking is a key foundation for realizing intelligent warehousing, archive management, automatic driving and other applications. The existing technology relies on traditional image feature matching, deep learning model, multi-sensor fusion and graph model association algorithm to realize the association of target trajectory.

[0003] However, in the application scene of dense shelf with complex structure, dense target and dynamic change, objects such as file boxes, shelves and mobile carts often have highly similar appearance features, combined with uneven lighting in the warehouse, personnel movement shielding and other interference factors. After the target is shielded or temporarily leaves the field and reappears, the existing technology has difficulty in maintaining trajectory continuity, and is prone to problems such as identity switching and trajectory breakage, resulting in low association accuracy of target association. SUMMARY

[0004] The problem solved by the present application is how to improve the association accuracy of target association in dense shelf complex background.

[0005] To solve the above problems, the present application provides a target association method and device for dense shelf complex background, equipment and medium.

[0006] In a first aspect, the present application provides a target association method for dense shelf complex background, comprising: Obtaining environment image data of a target and motion data of the target relative to a robot, the robot being used to obtain the environment image data, the environment image data including target spatial position; Performing feature extraction on the environment image data through a feature extraction algorithm to obtain a visual feature vector of the target, and inputting the environment image data into a visual language large model to obtain a semantic feature vector of the target, and performing feature extraction on the motion data through a long short-term memory network to obtain a motion feature vector; Performing feature fusion on the visual feature vector, the semantic feature vector and the motion feature vector through a feature fusion network to obtain a semantic enhanced feature vector of the target; Performing time sequence alignment on the semantic enhanced feature vector of the target and a historical trajectory feature vector sequence through a dynamic time warping algorithm to obtain a target matching trajectory of the target, and predicting a target predicted position of the target through Kalman filtering according to the target matching trajectory; Obtaining a target association trajectory according to the target spatial position and the target predicted position.

[0007] Optionally, the feature fusion of the visual feature vector, the semantic feature vector and the motion feature vector through the feature fusion network to obtain the semantic enhanced feature vector of the target comprises: feature fusion of the visual feature vector and the semantic feature vector through an attention mechanism to obtain a visual semantic fusion feature vector; feature fusion of the visual semantic fusion feature vector and the motion feature vector through an attention mechanism to obtain the semantic enhanced feature vector.

[0008] Optionally, the feature fusion of the visual feature vector and the semantic feature vector through an attention mechanism to obtain a visual semantic fusion feature vector comprises: mapping processing of the visual feature vector to obtain a visual mapping vector, and mapping processing of the semantic feature vector to obtain a semantic mapping vector; obtaining, according to the visual mapping vector and the semantic mapping vector, a first attention weight between the visual feature vector and the semantic feature vector through an attention mechanism; fusion of the visual feature vector and the semantic feature vector according to the first attention weight to obtain the visual semantic fusion feature vector.

[0009] Optionally, the feature fusion of the visual semantic fusion feature vector and the motion feature vector through an attention mechanism to obtain the semantic enhanced feature vector comprises: mapping processing of the visual semantic fusion feature vector to obtain a visual semantic mapping vector, and mapping processing of the motion feature vector to obtain a motion mapping vector; obtaining, according to the visual semantic mapping vector and the motion mapping vector, a second attention weight between the visual semantic fusion feature vector and the motion feature vector through an attention mechanism; fusion of the visual semantic fusion feature vector and the motion feature vector according to the second attention weight to obtain the semantic enhanced feature vector.

[0010] Optionally, the time sequence alignment of the semantic enhanced feature vector of the target and the historical trajectory feature vector sequence through the dynamic time warping algorithm to obtain the target matching trajectory of the target comprises: constructing a distance matrix of the semantic enhanced feature vector of the target and the historical trajectory feature vector sequence; obtaining, according to the distance matrix, the target matching trajectory of the target through a state transition equation.

[0011] Optionally, predicting the target's predicted location using Kalman filtering based on the target matching trajectory includes: The historical spatial location of the target is obtained based on the target matching trajectory; The target's predicted location is predicted based on the historical spatial location and state transition matrix.

[0012] Optionally, obtaining the target-related trajectory based on the target spatial location and the target predicted location includes: Calculate the spatial deviation between the target's spatial location and the predicted target location; The target matching trajectory is determined to be less than or equal to the preset deviation and matched with the target to obtain the target associated trajectory.

[0013] Secondly, the present invention provides a target association device for complex backgrounds of mobile shelving units, comprising: A data acquisition module is used to acquire environmental image data of the target and motion data of the target relative to the robot. The robot is used to acquire the environmental image data, which includes the spatial position of the target. The feature extraction module is used to extract features from the environmental image data using a feature extraction algorithm to obtain the visual feature vector of the target, input the environmental image data into a large visual language model to obtain the semantic feature vector of the target, and extract features from the motion data through a long short-term memory network to obtain the motion feature vector. The feature fusion module is used to fuse the visual feature vector, the semantic feature vector, and the motion feature vector through a feature fusion network to obtain the semantically enhanced feature vector of the target. The location prediction module is used to perform temporal alignment of the semantic enhancement feature vector and the historical trajectory feature vector sequence of the target using a dynamic time warping algorithm to obtain the target matching trajectory of the target, and predict the target prediction position of the target by Kalman filtering based on the target matching trajectory. The trajectory association module is used to obtain the target associated trajectory based on the target spatial location and the target predicted location.

[0014] Thirdly, the present invention provides an electronic device, including a memory and a processor; The memory is used to store computer programs; The processor is configured to, when executing the computer program, implement the target association method for complex backgrounds of mobile shelving units as described in the first aspect.

[0015] Fourthly, the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the target association method for complex backgrounds of mobile shelving units as described in the first aspect.

[0016] The beneficial effects of the target association method for complex backgrounds of mobile shelving units in this invention are as follows: Feature extraction algorithms effectively capture the spatial structure information of environmental image data, improving the accuracy of target recognition. Visual language models extract semantic features from environmental image data, acquiring semantic information such as "file box" and "shelf," effectively eliminating visual ambiguity and improving target recognition accuracy. Motion data features extracted through long short-term memory networks help solve the trajectory ID switching problem when the target is occluded or briefly leaves the field of vision, improving trajectory tracking stability. Feature fusion networks fuse visual feature vectors, semantic feature vectors, and motion feature vectors, improving the accuracy and robustness of target recognition. Dynamic time warping algorithms align the target's semantic enhancement feature vector with historical trajectory feature vector sequences to determine the target matching trajectory, ensuring the continuity of trajectory IDs. Kalman filtering predicts the target's predicted position based on motion state data in the target matching trajectory, improving trajectory tracking stability. Combining dynamic time warping with Kalman filtering achieves spatiotemporal dual constraints, ensuring the continuity and consistency of the target trajectory, reducing trajectory breakage problems, and thus improving the association accuracy of target association. Attached Figure Description

[0017] Figure 1 This is a flowchart illustrating a target association method for complex backgrounds of mobile shelving units in an embodiment of the present invention; Figure 2 This is a structural block diagram of the target association device for complex backgrounds of mobile shelving units in an embodiment of the present invention; Figure 3 This is a schematic diagram of the structure of an electronic device according to an embodiment of the present invention. Detailed Implementation

[0018] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Although some embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present invention. It should be understood that the accompanying drawings and embodiments of the present invention are for illustrative purposes only and are not intended to limit the scope of protection of the present invention.

[0019] It should be understood that the various steps described in the method embodiments of the present invention may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present invention is not limited in this respect.

[0020] The term "comprising" and its variations as used herein are open-ended, meaning "including but not limited to"; the term "based on" means "at least partially based on"; the term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments"; and the term "optionally" means "optional embodiments". Definitions of other terms will be given in the following description. It should be noted that the concepts of "first," "second," etc., mentioned in this invention are used only to distinguish different devices, modules, or units, and are not intended to limit the order of functions performed by these devices, modules, or units or their interdependencies.

[0021] It should be noted that the terms "a" and "a plurality of" used in this invention are illustrative rather than restrictive. Those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0022] The names of the messages or information exchanged between the multiple devices in the embodiments of the present invention are for illustrative purposes only and are not intended to limit the scope of these messages or information.

[0023] To address the problems existing in the aforementioned related technologies, this embodiment provides a target association method, apparatus, equipment, and medium for complex backgrounds of mobile shelving units.

[0024] like Figure 1 As shown, an embodiment of the present invention provides a target association method for complex backgrounds of mobile shelving units, comprising: S100, acquire environmental image data of the target and motion data of the target relative to the robot, wherein the robot is used to acquire the environmental image data, and the environmental image data includes the spatial position of the target; Specifically, RGB cameras mounted on warehouse robots collect environmental image data of the targets. Since the robot is continuously moving, it can acquire scene data from different perspectives of the targets. Most targets are stationary, and IMU sensors mounted on the warehouse robot synchronously record the robot's acceleration and angular velocity data as the target's motion data relative to the robot. In the archive storage room, the warehouse robot travels along tracks between mobile shelving units. RGB cameras mounted on the robot continuously capture images of targets including file boxes, shelf shelves, and handling trolleys. IMU sensors record the robot's forward, backward, turning, acceleration, and deceleration movements in real time.

[0025] S200: The environmental image data is processed by a feature extraction algorithm to obtain the visual feature vector of the target. The environmental image data is then input into a large visual language model to obtain the semantic feature vector of the target. The motion data is processed by a long short-term memory network to obtain the motion feature vector. Specifically, the environmental image data of the target is preprocessed, including grayscale conversion and noise reduction. Feature extraction is then performed on the preprocessed environmental image data, such as extracting local gradient features using the Histogram of Oriented Gradients (HOG) algorithm to generate visual feature vectors. The preprocessed environmental image data is then input into a large-scale visual language model, such as the CLIP model (Contrastive Language–Image Pre-training). The CLIP model can be pre-trained using a large amount of paired image and text data to learn the alignment relationships between images and text, enabling it to simultaneously understand information from two different modalities. This model achieves cross-modal alignment between images and text through contrastive learning and can output high-dimensional vectors related to the semantic content of the image, i.e., semantic feature vectors. These vectors can effectively capture key semantic information such as text labels and category attributes on the file box. Robot motion data collected by the IMU sensor is input into a Long Short-Term Memory (LSTM) network. The LSTM effectively captures long-term dependencies in the motion data through its gating mechanism, outputting motion feature vectors.

[0026] S300, the visual feature vector, the semantic feature vector, and the motion feature vector are fused through a feature fusion network to obtain the semantically enhanced feature vector of the target; Specifically, through the attention mechanism in the Transformer architecture, the visual feature vector searches for the most relevant semantic context in the semantic feature vector space, realizing the cross-guidance of visual and semantic features, and obtaining the fused features of visual feature vector and semantic feature vector. The fused features are then fused with the motion feature vector extracted by LSTM again through the attention mechanism. The motion features are used as contextual information to further fuse the features, and finally, the semantically enhanced feature vector is obtained.

[0027] S400, the semantic enhancement feature vector and the historical trajectory feature vector sequence of the target are temporally aligned using the dynamic time warping algorithm to obtain the target matching trajectory of the target, and the target prediction position of the target is predicted by Kalman filtering based on the target matching trajectory; Specifically, the Dynamic Time Warping (DTW) algorithm is used to temporally align the target's semantically enhanced feature vector with the historical trajectory feature vector sequence. The historical trajectory includes filtered, valid historical data associated with the target, and the historical trajectory feature vector sequence comprises the temporal set of semantically enhanced feature vectors generated by the target in previous consecutive frames. The DTW algorithm calculates the minimum cumulative distance between the target's semantically enhanced feature vector and the historical trajectory feature vector sequence to determine the matched historical trajectory, i.e., the target matching trajectory. A Kalman filter is then used to establish a state prediction model based on the historical motion states (such as position and velocity) in the target matching trajectory to predict the target's predicted position.

[0028] S500, the target-related trajectory is obtained based on the target spatial location and the target predicted location.

[0029] Specifically, by comparing the obtained target spatial location with the target predicted location predicted by the Kalman filter, the Euclidean distance between the target spatial location and the target predicted location can be calculated. If the Euclidean distance between the target spatial location and the target predicted location is less than the set threshold, the target matching trajectory matches the target, and the target matching trajectory is the target associated trajectory.

[0030] In this embodiment, feature extraction algorithms are used to extract features from environmental image data, effectively capturing the spatial structure information of the environmental image data and improving the accuracy of target recognition. A large visual language model extracts semantic features from the environmental image data, acquiring semantic information such as "file box" and "shelf," effectively eliminating visual ambiguity and improving target recognition accuracy. Motion data features extracted through a Long Short-Term Memory (LSTM) network help solve the trajectory ID switching problem when the target is occluded or briefly leaves the field of view, improving trajectory tracking stability. A feature fusion network integrates visual feature vectors, semantic feature vectors, and motion feature vectors, improving the accuracy and robustness of target recognition. A dynamic time warping algorithm temporally aligns the target's semantic enhancement feature vector with the historical trajectory feature vector sequence to determine the target matching trajectory, ensuring the continuity of trajectory IDs. Kalman filtering predicts the target's predicted position based on the motion state data in the target matching trajectory, improving trajectory tracking stability. By combining dynamic time warping with Kalman filtering to achieve spatiotemporal dual constraints, the continuity and consistency of the target trajectory are ensured, reducing trajectory breakage problems and thus improving the accuracy of target association.

[0031] Optionally, the step of fusing the visual feature vector, the semantic feature vector, and the motion feature vector through a feature fusion network to obtain the semantically enhanced feature vector of the target includes: The visual feature vector and the semantic feature vector are fused using an attention mechanism to obtain a visual-semantic fusion feature vector. The visual semantic fusion feature vector and the motion feature vector are fused using an attention mechanism to obtain the semantically enhanced feature vector.

[0032] Specifically, the visual feature vector extracted using the HOG algorithm is first used as the query vector, and the semantic feature vector extracted using the CLIP model is used as the key and value vectors. These are then fused using a multi-head attention mechanism to generate a visual-semantic fusion feature vector containing semantic guidance information. After fusing the visual and semantic feature vectors, the resulting visual-semantic fusion feature vector is used as the query vector, and the motion feature vector extracted from IMU data using an LSTM network is used as the key and value vectors. These are then fused using a multi-head attention mechanism to generate a semantically enhanced feature vector containing motion information.

[0033] In this optional embodiment, a multi-head attention mechanism can effectively capture the relationship between visual feature vectors and semantic feature vectors, achieving cross-guidance between them, effectively eliminating visual ambiguity, and improving the accuracy of target recognition. Based on the visual-semantic fusion feature vector, motion feature vectors are introduced for secondary feature fusion. By concatenating the outputs of multiple attention heads, the information captured by different attention heads is retained, thus ensuring that the fused feature vector not only contains rich semantic and visual information but also retains motion information, enabling more accurate target identification and tracking in complex environments.

[0034] Optionally, the step of fusing the visual feature vector and the semantic feature vector through an attention mechanism to obtain a visual-semantic fusion feature vector includes: The visual feature vector is mapped to obtain a visual mapping vector, and the semantic feature vector is mapped to obtain a semantic mapping vector. Based on the visual mapping vector and the semantic mapping vector, the first attention weight between the visual feature vector and the semantic feature vector is obtained through an attention mechanism; The visual feature vector and the semantic feature vector are fused according to the first attention weight to obtain the visual-semantic fusion feature vector.

[0035] Specifically, visual and semantic feature vectors from different modalities are mapped to the same dimension for attention computation. The visual feature vectors are mapped to obtain an environment mapping vector, which serves as the query vector. The semantic feature vectors are mapped to obtain semantic mapping vectors, which serve as the key and value vectors. The similarity between the query vector and the key vector for each attention head is calculated, and normalized first attention weights are generated using a softmax function. The weighted value vectors are then summed using these first attention weights to obtain the weighted output of each attention head. Finally, the features output from all attention heads are concatenated to obtain the visual-semantic fusion feature vector.

[0036] In this optional embodiment, in a mobile shelving environment, a large number of file boxes have highly similar appearances and are easily confused. By using semantic feature vectors to guide visual feature vectors, fusion features rich in contextual semantics are generated, which realizes the enhancement of vision by semantics, solves the problem of mismatch caused by visual similarity, and thus improves the association accuracy of target association.

[0037] Optionally, the step of fusing the visual semantic fusion feature vector and the motion feature vector through an attention mechanism to obtain the semantically enhanced feature vector includes: The visual semantic fusion feature vector is mapped to obtain a visual semantic mapping vector, and the motion feature vector is mapped to obtain a motion mapping vector. Based on the visual semantic mapping vector and the motion mapping vector, a second attention weight between the visual semantic fusion feature vector and the motion feature vector is obtained through an attention mechanism; The semantic enhancement feature vector is obtained by fusing the visual semantic fusion feature vector and the motion feature vector according to the second attention weight.

[0038] Specifically, the visual semantic fusion feature vectors of different modalities are ( The motion feature vector (M) and the motion feature vector (M) are mapped to the same intermediate dimension through independent linear projection layers. , to obtain the visual semantic mapping vector ( ) and motion mapping vector ( ); ; ; Use the visual semantic mapping vector as the query vector (Q= ), using the motion mapping vector as the key vector (K= ) and value vector (V= Similarity is calculated using a 4-head attention mechanism:

[0039] in = / 4, weights are assigned using softmax to highlight associated motion features.

[0040] The four attention output features are concatenated, mapped to the target dimension through a linear layer to obtain preliminary features, and then fused into a visual-semantic fusion feature vector through residual connections. Then perform Layer Normalization to obtain semantically enhanced feature vectors.

[0041] In this optional embodiment, visual semantic features may be affected by lighting conditions and occlusion, leading to a decrease in recognition accuracy. For example, under strong light or when an object is partially occluded, visual semantic features alone may not be sufficient to accurately identify the object, even though the object's motion state will not change due to lighting or slight occlusion in a short period of time. By introducing motion feature vectors for secondary feature fusion based on the visual semantic fusion feature vector, the fused feature vector not only contains rich semantic and visual information but also retains motion information. When the target reappears after being briefly occluded, its motion features can be used to assist in judgment and trajectory association, effectively mitigating the trajectory breakage problem caused by occlusion.

[0042] Optionally, the semantically enhanced feature vector of the target and the sequence of historical trajectory feature vectors are temporally aligned using a dynamic time warping algorithm to obtain the target matching trajectory of the target, including: Construct a distance matrix between the semantically enhanced feature vector of the target and the sequence of historical trajectory feature vectors; The target matching trajectory is obtained by using the state transition equation based on the distance matrix.

[0043] Specifically, construct the semantically enhanced feature vector of the target ( ) and historical trajectory feature vector sequence (H=[ , ,..., The distance matrix D of ]), where D[1,i]=distance( , Let represent the distance between the semantically enhanced feature vector of the target and the i-th feature vector in the sequence of historical trajectory feature vectors. The cumulative distance is calculated using the state transition equation in the dynamic time warping algorithm, and then the trajectory corresponding to the minimum cumulative distance is found, which is the target matching trajectory.

[0044] In this optional embodiment, by constructing a distance matrix between the semantically enhanced feature vector of the target and the sequence of historical trajectory feature vectors, the association between the target and the historical trajectory can be accurately identified. By using a dynamic programming algorithm to match paths, continuous tracking of the target can be maintained even in scenarios where the target is occluded or moving rapidly. This effectively solves the ID switching problem when the target is occluded or briefly out of sight, ensuring the continuity and consistency of the target trajectory, and thus improving the association accuracy of the target.

[0045] Optionally, predicting the target's predicted location using Kalman filtering based on the target matching trajectory includes: The historical spatial location of the target is obtained based on the target matching trajectory; The target's predicted location is predicted based on the historical spatial location and state transition matrix.

[0046] Specifically, the Kalman filter is a recursive filter used to predict the state of a target. The historical spatial position of the target can be obtained from historical motion data in the target's matching trajectory. A state prediction model using the Kalman filter is then constructed to build a state transition matrix, which is used to predict the target's predicted position. For example, if a cart disappears from view in frame 100 and reappears in frame 120, a cart trajectory similar to the one in frame 120 can be obtained using a dynamic time warping algorithm. Based on this trajectory, the Kalman filter can be used to predict the cart's predicted position in frame 120.

[0047] In this optional embodiment, the target prediction position is predicted by Kalman filtering, which improves the stability of trajectory tracking. Combined with the dynamic time warping algorithm, it forms a spatiotemporal dual constraint on the target trajectory association, ensuring the continuity and consistency of the target trajectory, reducing trajectory breakage problems, improving the accuracy and robustness of target tracking, and thus improving the association accuracy of target association.

[0048] Optionally, obtaining the target-related trajectory based on the target spatial location and the target predicted location includes: Calculate the spatial deviation between the target's spatial location and the predicted target location; The target matching trajectory is determined to be less than or equal to the preset deviation and matched with the target to obtain the target associated trajectory.

[0049] Specifically, the spatial deviation can be calculated by using the Euclidean distance between the target's spatial location and the predicted target location predicted by Kalman filtering based on historical motion data in the target's matching trajectory. When the calculated spatial deviation is less than or equal to a preset deviation, the target matching trajectory is considered to match the target, and this target matching trajectory is the target-associated trajectory. Multiple target matching trajectories that may be associated with the target can be obtained through dynamic time warping algorithms. By imposing spatial dimension constraints on the target matching trajectories, the uniqueness and accuracy of the final target identification are ensured, misidentification of the target is avoided, and the final target identification and tracking results are obtained.

[0050] In this optional embodiment, by calculating the spatial deviation between the target's spatial location and the predicted target location, a dual spatiotemporal constraint on the target association result is achieved, reducing the probability of target misalignment due to trajectory ID switching and trajectory breakage, ensuring the uniqueness and accuracy of the final target identification, and avoiding target misidentification.

[0051] like Figure 2 As shown, an embodiment of the present invention provides a target association device 200 for complex backgrounds of mobile shelving units, comprising: Data acquisition module 210 is used to acquire environmental image data of the target and motion data of the target relative to the robot. The robot is used to acquire the environmental image data, which includes the spatial position of the target. The feature extraction module 220 is used to extract features from the environmental image data using a feature extraction algorithm to obtain the visual feature vector of the target, input the environmental image data into a large visual language model to obtain the semantic feature vector of the target, and extract features from the motion data using a long short-term memory network to obtain the motion feature vector. The feature fusion module 230 is used to fuse the visual feature vector, the semantic feature vector and the motion feature vector through a feature fusion network to obtain the semantically enhanced feature vector of the target; The location prediction module 240 is used to perform temporal alignment of the semantic enhancement feature vector and the historical trajectory feature vector sequence of the target using a dynamic time warping algorithm to obtain the target matching trajectory of the target, and predict the target prediction position of the target by Kalman filtering based on the target matching trajectory. The trajectory association module 250 is used to obtain the target associated trajectory based on the target spatial location and the target predicted location.

[0052] like Figure 3As shown, an electronic device 300 provided in this embodiment of the invention includes a memory 310 and a processor 320; the memory 310 is used to store a computer program; the processor 320 is used to implement the target association method for complex backgrounds of mobile shelving as described above when the computer program is executed.

[0053] This invention provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it implements the target association method for complex backgrounds of mobile shelving units as described above.

[0054] The present invention will now be described an electronic device 300 that can serve as a server or client of the present invention, which is an example of a hardware device that can be applied to various aspects of the present invention. Electronic device 300 is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. Electronic device 300 can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0055] Electronic device 300 includes a computing unit that can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) or a computer program loaded from a storage unit into random access memory (RAM). The RAM may also store various programs and data required for device operation. The computing unit, ROM, and RAM are interconnected via a bus. Input / output (I / O) interfaces are also connected to the bus.

[0056] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc. In this application, the units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of the embodiments of the present invention according to actual needs. Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated units can be implemented in hardware or as software functional units.

[0057] While the present invention has been disclosed above, its scope of protection is not limited thereto. Those skilled in the art can make various changes and modifications without departing from the spirit and scope of the present invention, and all such changes and modifications will fall within the scope of protection of the present invention.

Claims

1. A target association method for complex backgrounds of mobile shelving units, characterized in that, include: The robot acquires environmental image data of the target and motion data of the target relative to the robot, wherein the robot is used to acquire the environmental image data, and the environmental image data includes the spatial position of the target. The environmental image data is subjected to feature extraction by a feature extraction algorithm to obtain the visual feature vector of the target. The environmental image data is then input into a large visual language model to obtain the semantic feature vector of the target. The motion data is subjected to feature extraction by a long short-term memory network to obtain the motion feature vector. The visual feature vector, the semantic feature vector, and the motion feature vector are fused using a feature fusion network to obtain the semantically enhanced feature vector of the target. The semantic enhancement feature vector and historical trajectory feature vector sequence of the target are temporally aligned using the dynamic time warping algorithm to obtain the target matching trajectory of the target. The target prediction position of the target is then predicted by Kalman filtering based on the target matching trajectory. The target-related trajectory is obtained based on the target's spatial location and the target's predicted location.

2. The target association method for complex backgrounds of mobile shelving units according to claim 1, characterized in that, The step of fusing the visual feature vector, the semantic feature vector, and the motion feature vector through a feature fusion network to obtain the semantically enhanced feature vector of the target includes: The visual feature vector and the semantic feature vector are fused using an attention mechanism to obtain a visual-semantic fusion feature vector. The visual semantic fusion feature vector and the motion feature vector are fused using an attention mechanism to obtain the semantically enhanced feature vector.

3. The target association method for complex backgrounds of mobile shelving units according to claim 2, characterized in that, The step of fusing the visual feature vector and the semantic feature vector through an attention mechanism to obtain a visual-semantic fusion feature vector includes: The visual feature vector is mapped to obtain a visual mapping vector, and the semantic feature vector is mapped to obtain a semantic mapping vector. Based on the visual mapping vector and the semantic mapping vector, the first attention weight between the visual feature vector and the semantic feature vector is obtained through an attention mechanism; The visual feature vector and the semantic feature vector are fused according to the first attention weight to obtain the visual-semantic fusion feature vector.

4. The target association method for complex backgrounds of mobile shelving units according to claim 3, characterized in that, The step of fusing the visual semantic fusion feature vector and the motion feature vector through an attention mechanism to obtain the semantically enhanced feature vector includes: The visual semantic fusion feature vector is mapped to obtain a visual semantic mapping vector, and the motion feature vector is mapped to obtain a motion mapping vector. Based on the visual semantic mapping vector and the motion mapping vector, a second attention weight between the visual semantic fusion feature vector and the motion feature vector is obtained through an attention mechanism; The semantic enhancement feature vector is obtained by fusing the visual semantic fusion feature vector and the motion feature vector according to the second attention weight.

5. The target association method for complex backgrounds of mobile shelving units according to claim 1, characterized in that, The step of using a dynamic time warping algorithm to temporally align the semantically enhanced feature vector of the target and the historical trajectory feature vector sequence to obtain the target matching trajectory includes: Construct a distance matrix between the semantically enhanced feature vector of the target and the sequence of historical trajectory feature vectors; The target matching trajectory is obtained by using the state transition equation based on the distance matrix.

6. The target association method for complex backgrounds of mobile shelving units according to claim 1, characterized in that, The step of predicting the target's predicted position using Kalman filtering based on the target matching trajectory includes: The historical spatial location of the target is obtained based on the target matching trajectory; The target's predicted location is predicted based on the historical spatial location and state transition matrix.

7. The target association method for complex backgrounds of mobile shelving units according to claim 1, characterized in that, The step of obtaining the target-related trajectory based on the target spatial location and the target predicted location includes: Calculate the spatial deviation between the target's spatial location and the predicted target location; The target matching trajectory is determined to be less than or equal to the preset deviation and matched with the target to obtain the target associated trajectory.

8. A target association device for complex backgrounds of mobile shelving units, characterized in that, include: A data acquisition module is used to acquire environmental image data of the target and motion data of the target relative to the robot. The robot is used to acquire the environmental image data, which includes the spatial position of the target. The feature extraction module is used to extract features from the environmental image data using a feature extraction algorithm to obtain the visual feature vector of the target, input the environmental image data into a large visual language model to obtain the semantic feature vector of the target, and extract features from the motion data through a long short-term memory network to obtain the motion feature vector. The feature fusion module is used to fuse the visual feature vector, the semantic feature vector, and the motion feature vector through a feature fusion network to obtain the semantically enhanced feature vector of the target. The location prediction module is used to perform temporal alignment of the semantic enhancement feature vector and the historical trajectory feature vector sequence of the target using a dynamic time warping algorithm to obtain the target matching trajectory of the target, and predict the target prediction position of the target by Kalman filtering based on the target matching trajectory. The trajectory association module is used to obtain the target associated trajectory based on the target spatial location and the target predicted location.

9. An electronic device, characterized in that, Including memory and processor; The memory is used to store computer programs; The processor is configured to, when executing the computer program, implement the target association method for complex backgrounds of mobile shelving as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The storage medium stores a computer program that, when executed by a processor, implements the target association method for complex backgrounds of mobile shelving as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Multi-modal middle school experiment step detection method and system based on moving target semantic enhancement

    CN118411652A

  • Logistics robot path planning method based on multi-modal perception

    CN120628104A

Cited By

  • Robot motion control method, robot motion control device and quadruped robot

    CN122151926A