Physical object traversal recognition detection method and device based on unmanned aerial vehicle equipment

By using drone equipment to collect and identify cargo information in warehouse management and compare it with the database, the shortcomings of traditional manual inspections are solved, and efficient and accurate inventory management and real-time information updates are achieved.

CN120688985AActive Publication Date: 2025-09-23ALADDIN UAV (SHENZHEN) CO LTD

Patent Information

Application Number
CN202510862397.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-25
Publication Date
2025-09-23
Estimated Expiration
2045-06-25

AI Technical Summary

Technical Problem

In warehouse management, traditional manual inspection methods have high labor costs, limited coverage, high operational risks, and unstable recognition accuracy, resulting in delayed inventory information updates, long inventory cycles, and frequent errors. Especially in large logistics warehouses or high-bay shelf areas, it is difficult to achieve comprehensive and timely recognition.

Method used

A physical object traversal recognition and detection method based on drone equipment is adopted. The drone equipment collects cargo information under a preset flight trajectory, uses the cargo recognition algorithm to identify cargo feature information, and compares it with the preset database to achieve inventory analysis.

Benefits of technology

It improves the accuracy of inventory management and the real-time nature of information updates, ensures the coverage and sequence of the collection process, improves the accuracy and adaptability of feature recognition, and outputs quantifiable inventory analysis results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120688985A_ABST
    Figure CN120688985A_ABST
Patent Text Reader

Abstract

The invention relates to an entity object traversal recognition detection method and device based on unmanned aerial vehicle equipment. The method comprises the following steps: acquiring cargo information acquired by unmanned aerial vehicle equipment in a cargo storage area, wherein the cargo information represents cargo storage related information acquired by the unmanned aerial vehicle equipment running according to a preset flight path and sequentially acquiring each storage position in the cargo storage area; according to the entity feature category in the cargo information, a cargo recognition algorithm matched with the cargo information is determined, cargo feature information corresponding to the cargo information is recognized based on the cargo recognition algorithm, and the cargo feature information comprises the product model, the storage number and the storage position of the corresponding cargo; and comparing the cargo feature information with inventory information in a preset database to obtain an inventory analysis result of the cargo storage area. By adopting the method, the systematic analysis of the cargo storage state can be realized, and the accuracy of inventory management and the real-time performance of information updating can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of warehouse management technology, and in particular to a method and device for detecting entity object traversal and identification based on drone equipment. Background Art

[0002] In the field of warehouse management technology, inventory counting is usually achieved through manual inspections. However, this method has disadvantages such as high labor costs, limited coverage, high operational risks, and unstable recognition accuracy, which leads to delayed inventory information updates, lengthy inventory counting cycles, and frequent errors. Especially in large logistics warehouses or high-bay racking areas, traditional methods cannot achieve comprehensive and timely identification, often requiring high-altitude operations or the coordinated operation of multiple trades, which seriously restricts the accuracy and efficiency of inventory management. Summary of the Invention

[0003] Based on this, it is necessary to provide a method, device, computer equipment and computer-readable storage medium for entity object traversal recognition and detection based on drone equipment to address the above technical problems, so as to realize systematic analysis of the storage status of goods, which will help improve the accuracy of inventory management and the real-time nature of information updates.

[0004] In a first aspect, the present application provides a method for detecting entity object traversal and recognition based on a drone device, comprising: Obtaining cargo information collected by the drone device in the cargo storage area, wherein the cargo information represents cargo storage-related information collected sequentially from each storage location in the cargo storage area by the drone device operating along a preset flight trajectory; determining, based on the entity feature category in the cargo information, a cargo identification algorithm that matches the cargo information, and identifying cargo feature information corresponding to the cargo information based on the cargo identification algorithm, the cargo feature information including a product model, storage quantity, and storage location of the corresponding cargo; The cargo characteristic information is compared with the inventory information in a preset database to obtain an inventory analysis result of the cargo storage area.

[0005] In a second aspect, the present application also provides a device for detecting entity object traversal and recognition based on a drone device, comprising: an acquisition module, configured to acquire cargo information collected by the drone device in the cargo storage area, wherein the cargo information represents cargo storage-related information collected sequentially from each storage location in the cargo storage area by the drone device operating along a preset flight trajectory; an identification module, configured to determine a cargo identification algorithm that matches the cargo information based on the entity feature category in the cargo information, and identify cargo feature information corresponding to the cargo information based on the cargo identification algorithm, wherein the cargo feature information includes a product model, storage quantity, and storage location of the corresponding cargo; The comparison module is used to compare the cargo characteristic information with the inventory information in a preset database to obtain an inventory analysis result of the cargo storage area.

[0006] In a third aspect, the present application further provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the above steps when executing the computer program.

[0007] In a fourth aspect, the present application further provides a computer-readable storage medium on which a computer program is stored, and the computer program implements the above steps when executed by a processor.

[0008] The above-mentioned entity object traversal recognition and detection method, device, computer device and computer-readable storage medium based on drone equipment first collects cargo information of each storage location in the cargo storage area in sequence according to a preset flight trajectory of the drone equipment, thereby ensuring that the collection process has coverage and sequentiality, and establishing a stable data foundation for subsequent identification; secondly, a matching cargo identification algorithm is determined according to the entity feature category in the cargo information, thereby realizing targeted information extraction to obtain cargo feature information, thereby improving the accuracy and adaptability of feature recognition; thirdly, the cargo feature information is compared with the inventory information in the preset database, thereby realizing the corresponding verification of the recognition result and the inventory registration data, and outputting a quantifiable inventory analysis result; based on this, through the whole process of data collection, feature recognition and inventory comparison processing, a systematic analysis of the storage status of the cargo is realized, which helps to improve the accuracy of inventory management and the real-time nature of information updating. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0010] Figure 1 1 is a flow chart of a method for traversal, identification and detection of entity objects based on a drone device in one embodiment; Figure 2 Schematic diagram of the structure of a drone device in one embodiment; Figure 3A schematic diagram of a process for identifying cargo feature information using a text detection and recognition algorithm based on relational modeling in one embodiment; Figure 4 A flowchart of a text detection and recognition algorithm based on relational modeling in one embodiment is shown; Figure 5 A schematic diagram of a process for obtaining text detection results using a text detection and recognition algorithm based on relational modeling in one embodiment; Figure 6 Schematic diagram of the structure of a text detection network with Mamba fusion and multi-channel sampling in one embodiment; Figure 7 Schematic diagram of the structure of multiple upsampling layers in a text detection network in one embodiment; Figure 8 Schematic diagram of the structure of the feature fusion layer in a text detection network in one embodiment; Figure 9 A schematic diagram of a process for obtaining text recognition results using a text detection and recognition algorithm based on relational modeling in one embodiment; Figure 10 A schematic diagram of the structure of a convolution-deconvolution text recognition network based on global relationship modeling in one embodiment; Figure 11 1 is a flow chart of identifying cargo feature information based on a hybrid attention and temporal fusion enhancement algorithm in one embodiment; Figure 12 1. A schematic diagram of a process for obtaining feature information of different scales based on a hybrid attention and temporal fusion enhancement algorithm in one embodiment; Figure 13 Schematic diagram of the structure of a cargo recognition and detection algorithm based on hybrid attention and Transformer fusion enhancement in one embodiment; Figure 14 Schematic diagram of the structure of the ACK module in a cargo recognition and detection algorithm based on hybrid attention and Transformer fusion enhancement in one embodiment; Figure 15 1. A flow chart of obtaining fusion features of different scales based on a hybrid attention and temporal fusion enhancement algorithm in one embodiment; Figure 16 Schematic diagram of the structure of a multi-scale temporal feature fusion layer in a cargo recognition and detection algorithm based on hybrid attention and Transformer fusion enhancement in one embodiment; Figure 17 1 is a schematic diagram of a process for identifying characteristic information of goods based on an identification recognition algorithm in one embodiment; Figure 18This is a structural block diagram of an entity object traversal recognition and detection device based on a drone device in one embodiment. DETAILED DESCRIPTION

[0011] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0012] In one embodiment, Figure 1 As shown, a method for detecting entity object traversal based on a drone device is provided. This embodiment uses the method applied to a server as an example. It is understood that the method can also be applied to a terminal, or to a system including a terminal and a server, and implemented through interaction between the terminal and the server. In this embodiment, the method includes the following steps S101 to S103.

[0013] Step S101: obtaining cargo information collected by the drone device in the cargo storage area. The cargo information represents cargo storage-related information collected by the drone device in sequence from each storage location in the cargo storage area according to a preset flight trajectory.

[0014] in, Figure 2 A schematic diagram of the structure of a drone device is shown, which includes a camera 1 installed on the bottom of the drone device body and an RFID reader 2. The camera 1 can represent a common RGB camera for capturing RGB images, or an RGB-D camera for simultaneously capturing RGB images and depth images to collect cargo information in the form of text, images, QR code identification, etc. The RFID reader 2 is used to read the radio frequency signal from the RFID tag to collect cargo information in the form of RFID tag identification.

[0015] Furthermore, Figure 2 The UAV device in the figure is a quad-rotor UAV, in which four sets of propeller arms 3 extend to the four corners of the body respectively, and a set of motors 4 and propellers 5 are installed at the end of each set of propeller arms 3, so that the UAV device has the ability to hover, accurately control the position and turn in space. It can fly at low altitude above or to the side of the cargo, and can move and shoot around the cargo from multiple angles.

[0016] Among them, the cargo storage area refers to the physical space range used to store different types of goods, such as a warehouse or shopping mall with multi-layer shelves; the storage location refers to a specific cargo storage unit in the cargo storage area with a unique space number or geometric coordinate identifier, such as the storage location corresponding to the 5th grid of the 3rd layer of a shelf.

[0017] Among them, cargo information refers to a set of data directly related to the cargo collected by the drone equipment through shooting, scanning or other sensing methods on the flight trajectory. The cargo information can represent information of data types such as text, images, identification codes, empty locations, etc.

[0018] Among them, the flight trajectory represents the pre-planned three-dimensional path followed by the drone equipment when operating in the cargo storage area. It is used to guide the drone equipment to complete the sequential access and data collection of all storage locations without duplication or omission. For example, a layer-by-layer flight path from the entrance of the area passes through each row of shelves from the upper layer to the lower layer, so as to realize real-time and orderly information acquisition and information modeling for the product model, location and count of each cargo based on the flight trajectory.

[0019] For example, first, the cargo storage area can be spatially divided and calibrated through static mapping or existing layout drawings, thereby forming basic environmental data for drone equipment flight trajectory planning to clarify the spatial location of each cargo storage unit. Subsequently, a set of preset flight trajectories is constructed for the above cargo storage area. These flight trajectories are required to cover all cargo storage units, ensuring that the drone equipment can sequentially collect data from each storage location during its operation. Furthermore, during actual operation, the drone equipment operates stably according to the preset flight trajectory, while controlling its flight altitude, steering, and hovering time to fully scan each storage location, ensuring that the cargo status from each perspective can be fully collected and recorded as cargo information.

[0020] Step S102: determining a cargo identification algorithm that matches the cargo information based on the entity feature category in the cargo information, and identifying cargo feature information corresponding to the cargo information based on the cargo identification algorithm. The cargo feature information includes the product model, storage quantity, and storage location of the corresponding cargo.

[0021] Among them, the entity feature category represents the classification identifier reflecting the data type in the cargo information, which is used to determine which type of cargo identification algorithm should be used to identify and process the cargo information. It may include the text feature category corresponding to the text data type, the image feature category corresponding to the image data type, the identification feature category corresponding to the identification code data type, etc.

[0022] Among them, the cargo identification algorithm represents a set of regularized processing logic or operation processes used to perform identification operations on different entity feature categories, so as to adaptively extract cargo feature fields with business significance from the cargo information according to the data type of the cargo information, namely, cargo feature information such as product model, storage quantity and storage location.

[0023] For example, first, based on the collected cargo information, the entity feature categories contained in the cargo information are identified. Specifically, the format of the cargo information is analyzed to determine whether it contains some form of content that can be used for feature extraction, such as structural features such as images, text, or identification codes. The entity feature categories corresponding to the cargo information are thereby identified. Furthermore, after identifying the entity feature categories corresponding to the cargo information, a corresponding cargo identification algorithm is employed based on the entity feature categories to identify the cargo feature information. Specifically, different cargo identification algorithms have different processing logic for different entity feature categories. Furthermore, during the identification process, the corresponding cargo identification algorithm processes the feature content in the cargo information according to the corresponding processing logic, thereby extracting feature information such as the product model, corresponding storage quantity, and storage location of each cargo. This feature information is expressed as structured fields and organized into cargo feature information for the corresponding cargo.

[0024] Step S103 : Compare the cargo characteristic information with the inventory information in the preset database to obtain the inventory analysis result of the cargo storage area.

[0025] Among them, the preset database refers to a data storage space established in advance and continuously maintained for storing a structured data set of the storage status of various types of goods. For example, it can represent a product information database in a warehouse management system, which includes inventory information such as product model, packaging specifications, registered storage location, inventory registration time, etc.

[0026] Among them, the inventory analysis result refers to the data analysis output result formed after comparing the identified cargo feature information with the existing inventory information in the database. It is used to display the difference, change or consistency relationship between the current actual recognition result and the system registration data. For example, it can indicate "Product A inventory is consistent, no difference", "Product B has a position deviation", "Product C is a newly added unregistered item" and other comparison results, and present them in the form of a table or list.

[0027] For example, before comparing the goods characteristic information with the inventory information in the database, it is necessary to ensure that each parameter field in the goods characteristic information is comparable, that is, its format, unit, and numbering coding system must all comply with the consistency standards defined in the database. Furthermore, during the comparison process, the product model field is used as the primary index field and matched with the inventory information registered in the database in sequence. Through the one-to-one correspondence between the field values, it is determined whether the corresponding goods have matching inventory items in the inventory information; if a certain product has a matching inventory item in the inventory information, the difference between the identified storage quantity and the storage quantity registered in the database is compared to further determine whether the current goods inventory is newly added, missing, or has changed in quantity; if a certain product does not have a matching inventory item in the inventory information, it is recorded as a newly added inventory item and marked separately. Furthermore, it is necessary to check the identified storage location with the storage location registered in the database to determine whether there is misplaced storage, inconsistent information, or erroneous recording. Ultimately, all comparison results will be organized into a unified inventory analysis result, which includes the inventory item matching status and quantity difference information of the goods. It also marks the accuracy of goods information recognition, inventory information completeness, and suspicious and abnormal storage locations. This will serve as the basis for subsequent inventory management operations and be used to trigger inventory corrections, replenishment warnings or manual review operations.

[0028] In the above-mentioned entity object traversal recognition and detection method based on drone equipment, first, the cargo information of each storage location in the cargo storage area is collected in sequence according to the preset flight trajectory of the drone equipment, so as to ensure the coverage and sequentiality of the collection process and establish a stable data foundation for subsequent identification; secondly, the matching cargo recognition algorithm is determined according to the entity feature category in the cargo information, so as to realize targeted information extraction to obtain cargo feature information, thereby improving the accuracy and adaptability of feature recognition; thirdly, the cargo feature information is compared with the inventory information in the preset database, so as to realize the corresponding verification of the recognition result and the inventory registration data, and output a quantifiable inventory analysis result; based on this, through the whole process of data collection, feature recognition and inventory comparison processing, a systematic analysis of the storage status of the cargo is realized, which helps to improve the accuracy of inventory management and the real-time nature of information update.

[0029] In an exemplary embodiment, Figure 3 As shown, according to the entity feature category in the cargo information, a cargo identification algorithm matching the cargo information is determined, and cargo feature information corresponding to the cargo information is identified based on the cargo identification algorithm, including steps S201 to S203.

[0030] In step S201 , if the entity feature category in the cargo information is a text feature category, a preset text detection and recognition algorithm based on relational modeling is used as a cargo recognition algorithm for cargo information matching.

[0031] The text feature category refers to the type of visual data composed of characters, words or symbols identified from the cargo information, such as printed text or handwritten marks appearing on cargo labels or the surface of cargo.

[0032] Among them, the text detection and recognition algorithm based on relational modeling represents a joint recognition processing flow that simultaneously considers the spatial arrangement and semantic structure association between characters, and is used to improve the recognition accuracy in complex backgrounds, curved arrangements or multilingual text scenarios.

[0033] For example, by analyzing factors such as pixel arrangement, boundary structure, and color density in the photographed cargo information, it is determined whether the cargo information contains character components. If it is determined that the cargo information contains character components, it is considered that the cargo information belongs to the text feature category, that is, there are recognizable character sequences or label texts. The text detection and recognition algorithm based on relational modeling can be called as a matching cargo recognition algorithm to subsequently perform more structured recognition of the data in the text area. Furthermore, if Figure 4 As shown in the figure, this type of algorithm includes two interrelated processing flows: on the one hand, there is the text area detection flow, which locates the text area in the input image based on the text detection network to obtain the text detection result; on the other hand, there is the text content recognition flow, which recognizes the character sequence or label text in the located text area based on the text recognition network to obtain the text recognition result.

[0034] In step S202 , based on the text detection network in the cargo identification algorithm, multi-scale feature extraction and fusion processing are performed on the cargo information to obtain text detection results corresponding to the cargo information.

[0035] Among them, the text detection network represents the neural network structure used to identify text areas in cargo information, so as to accurately locate and separate the area containing text content from the complex background, that is, output text detection results, which may include structured fields such as detection boxes, tilt angles, and detection confidence of several text areas.

[0036] For example, in a text detection network, image data from cargo information is used as input, and a multi-scale feature extraction and fusion mechanism is used to analyze and process possible text regions within the image data. This process first involves performing multi-layer feature extraction on the original image data through convolution operations. Each convolutional layer extracts information such as texture, edges, and local contrast at different scales, forming a set of feature maps with varying receptive fields and resolutions. In practice, because the size, arrangement, image quality, and format of text corresponding to cargo information vary, a comprehensive multi-scale analysis of the image data is necessary to enhance the perception of small characters, oblique text, or text regions against complex backgrounds.

[0037] After acquiring these multi-scale feature maps, the feature maps at each scale are further aligned through upsampling operations to ensure uniform spatial distribution of features from different convolutional layers, preserving the key structures in each layer and improving overall expressiveness. Feature fusion is then performed based on the upsampling process, integrating multiple layers of features in a unified scale space. The fusion process simultaneously considers local structural information and overall spatial associations to form a unified fused feature map. Subsequently, by detecting the spatial distribution and feature intensity information of the fused feature maps, possible text regions are identified in the cargo information. These text regions are output as detection boxes, indicating the specific location and range of the text regions in the cargo information, forming the final text detection results.

[0038] Step S203 , based on the text recognition network in the cargo identification algorithm, the text detection result is parsed to obtain a text recognition result corresponding to the text detection result, and the text recognition result is used as cargo feature information corresponding to the cargo information.

[0039] Among them, the text recognition network represents a neural network structure that performs content analysis on text detection results to convert the information in the text area into a specific text sequence, that is, outputs text recognition results, which may include text fields such as product model, storage quantity, and storage location.

[0040] For example, in a text recognition network, the text detection results are used as input, and a multi-stage recognition and parsing process is performed on the text regions contained therein to convert the text regions in image form into text sequences with clear semantics. During this process, the text regions are first uniformly formatted to meet the input requirements of the subsequent recognition structure. Subsequently, a processing path for feature extraction and feature modeling is constructed to achieve in-depth analysis of the content of the text regions. While extracting local image features, this processing path also introduces structural logic based on sequence modeling to capture the sequential relationships and semantic coherence between characters. This not only captures low-level image details such as character edges and strokes, but also establishes global associations between characters, thereby improving the ability to parse complex text arrangements. Furthermore, to further enhance recognition stability, a processing path for feature restoration and feature modeling is constructed. While maintaining structural modeling capabilities, some original image features in the text regions are restored to ensure consistency between character morphology and spatial structure. Based on this, the feature results generated by the two processing paths are fused in a unified space to form the final output features, which contain the image morphology, arrangement structure, and contextual semantic relationships of the characters. Finally, the output feature is subjected to character-by-character recognition, and the recognition result of each character is output through a structured classification and decoding process. These characters are then combined into a complete text string in the recognition order as the text recognition result contained in the text area.

[0041] In this embodiment, first, if the entity feature category in the cargo information is a text feature category, a text detection and recognition algorithm based on relational modeling is matched, thereby achieving targeted adaptation of the recognition algorithm to the information type and improving the accuracy and efficiency of subsequent processing; secondly, multi-scale feature extraction and fusion processing are performed based on the text detection network to obtain text detection results, thereby enhancing the detection ability of text areas of different sizes and in complex backgrounds, and improving the completeness and accuracy of text area extraction; thirdly, the text detection results are parsed and processed based on the text recognition network to obtain text recognition results, thereby achieving structured restoration of the text content and generating cargo feature information that can be directly used for business comparison; based on this, through the coordinated processing of entity feature category judgment, text area positioning and text content recognition, high-precision recognition of text content in cargo information is achieved.

[0042] In an exemplary embodiment, the text detection network includes a convolutional layer, a multi-sampling layer, a feature fusion layer, and a detection head structure; Figure 5 As shown, based on the text detection network in the cargo identification algorithm, multi-scale feature extraction and fusion processing are performed on the cargo information to obtain the text detection results corresponding to the cargo information, including steps S301 to S303.

[0043] Among them, the convolution layer represents the network structure unit for extracting local features of the input image to establish the feature map expression of the image at different scales; the multi-sampling layer represents the network structure unit for resizing the feature maps of different scales to increase the spatial resolution of the deep feature map to the same level as the shallow feature map. Figure 1 The feature fusion layer represents the network structure unit used to fuse feature maps from different scales and semantic levels, so as to enhance the overall semantic expression ability of the image while retaining local details; the detection head structure represents the network structure module used to perform region discrimination and boundary regression tasks based on the fused feature maps, so as to locate, classify or predict the confidence of text areas that may exist in the image.

[0044] In step S301, based on the convolutional layer, multi-scale feature extraction processing is performed on the cargo information to obtain feature maps of different scales.

[0045] In step S302 , the linear interpolation and deconvolution combined processing logic corresponding to the multiple upsampling layers and the global information modeling and splicing combined processing logic corresponding to the feature fusion layer are combined to perform upsampling and fusion processing on the feature maps of different scales to obtain a fused feature map.

[0046] Among them, the linear interpolation and deconvolution processing logic corresponding to multiple upsampling layers represents a processing mechanism that combines two different image enhancement methods to achieve size recovery in the upsampling process of multi-scale feature maps; among them, linear interpolation is used to smoothly fill the values ​​of the newly added pixels in the feature map to make them consistent with the original pixel distribution, and deconvolution is used to expand the spatial structure of the feature map and enhance the expression ability of edges and patterns, so as to retain the key structural information in the image while improving the spatial resolution of the feature map.

[0047] Among them, the global information modeling and splicing combined processing logic corresponding to the feature fusion layer represents a processing mechanism that performs both information splicing and contextual relationship modeling in the fusion processing of multi-scale feature maps; the splicing operation is used to retain the original feature content from different scales, and the global information modeling is used to establish spatial dependencies and cross-regional connections between features, so as to generate a more comprehensive and spatially consistent fused feature map.

[0048] Step S303 : Based on the detection head structure, the fused feature map is detected and analyzed to obtain text region detection information based on the cargo information, and the text region detection information is used as the text detection result corresponding to the cargo information.

[0049] Among them, the text area detection information represents a set of spatial positioning data output by the detection head structure after detecting and analyzing the fused feature map, which is used to identify the specific area location and structural features of the input image that may contain text. For example, each detection information may include the coordinates, size, direction angle corresponding to a rectangular bounding box that calibrates a certain area, as well as a confidence score of whether the area is a text area.

[0050] For example, in order to solve multiple problems such as inaccurate text area positioning and complex background interference in the text detection process, and to ensure accurate extraction of text information, the text detection and recognition algorithm based on relationship modeling in this embodiment is a text detection and recognition algorithm based on Mamba fusion and global relationship modeling, and the text detection network therein is a Mamba fusion and multi-path sampling text detection network.

[0051] like Figure 6 As shown, in the text detection network, the input image The images are sequentially input into Conv1 (i.e. the first convolutional layer), Conv2 (i.e. the second convolutional layer), Conv3 (i.e. the third convolutional layer), and Conv4 (i.e. the fourth convolutional layer) for feature extraction layer by layer to obtain feature maps at different levels. The low-level feature maps mainly correspond to edge, texture and other features, while the high-level feature maps mainly correspond to semantic features.

[0052] The feature map output by Conv4 is upsampled by Multi-upsample1 (i.e., the first multi-upsample layer) to obtain the feature ; Through Mamba Fusion1 (the first feature fusion layer) the features Fuse it with the feature map output by Conv4 to obtain the feature ; Through Multi-upsample2 (the second multi-upsample layer), the features Perform upsampling to obtain features ; Through Mamba Fusion2 (the second feature fusion layer) the features Fuse it with the feature map output by Conv2 to obtain the feature ; Through Multi-upsample3 (the third multi-upsample layer), the features Perform upsampling and then use Mamba Fusion3 (the third feature fusion layer) to combine the features output by Multi-upsample3 with the features Perform fusion processing to obtain features ; Through the detection head structure to detect features Perform detection analysis, that is, generate a probability map and a threshold map during the detection and analysis process, and fuse the two to calculate an approximate binary map, thereby achieving accurate detection of text areas and obtaining text detection results .

[0053] like Figure 7 As shown, in the multi-sampling layer, two different upsampling strategies are used to achieve the upsampling effect, that is, two different upsampling operations are designed, including an upsampling operation based on linear interpolation. and deconvolution-based upsampling operations ; Input features Pass separately and Perform upsampling, and then perform addition operation on the features obtained by the two upsampling operations. Operate with average value Perform averaging to obtain the final output upsampling result .

[0054] Furthermore, the processing logic of the above-mentioned multiple upsampling layers can refer to formula (1): (1) In formula (1), the upsampling operation based on linear interpolation Can be achieved through linear interpolation or bilinear interpolation; upsampling operation based on deconvolution This can be achieved through deconvolution with optimized parameters; It represents a function composite symbol, which is used to indicate that functions are executed sequentially. The processing logic of formula (1) corresponds to the linear interpolation and deconvolution processing logic corresponding to multiple sampling layers.

[0055] like Figure 8 As shown in the figure, in the feature fusion layer, a new MambaFusion network structure layer is designed by using Mamba and Non-local technology; on the one hand, an input feature With another input feature Input into Mamba for global information modeling of their respective features , and obtain the corresponding global information enhancement features; then input the corresponding global information enhancement features into Concat1 (i.e. the first splicing layer) for feature splicing operation , get the first splicing feature, and then input the first splicing feature into Non-local to perform global information fusion modeling operation , and obtain the first fusion feature.

[0056] On the other hand, the input features and Input to Concat2 (the second concatenation layer) for feature concatenation , get the second splicing feature; then add the second splicing feature to the first fusion feature based on the addition operation Operate with average value Perform average processing to obtain the final output fusion result .

[0057] Furthermore, the processing logic of the feature fusion layer can refer to equations (2) to (4): (2) (3) (4) In formulas (2) to (4), Represents the first fusion feature of Non-local output, Represents the second concatenated feature output by Concat2; Represents a function composite symbol, which is used to indicate that functions are executed sequentially. The processing logic of equations (2) to (4) corresponds to the global information modeling and splicing combined processing logic corresponding to the feature fusion layer.

[0058] In this embodiment, first, multi-scale feature extraction processing is performed on the cargo information according to the convolutional layer, so as to obtain feature maps that express the text region structure and contour features under different receptive fields, thereby enhancing the recognition support for characters of different sizes; secondly, upsampling and fusion processing are performed on the multi-scale feature map according to the multiple upsampling layers and feature fusion layers to obtain a fused feature map, thereby improving the uniformity of the feature map in terms of spatial structure and semantic information; thirdly, the fused feature map is detected and analyzed according to the detection head structure to obtain text region detection information, thereby accurately extracting the spatial position of the text region in the image. Based on this, accurate and efficient recognition and positioning of the text region in the image is achieved.

[0059] In an exemplary embodiment, the text recognition network includes a convolution layer, a deconvolution layer, a temporal modeling layer, and a detection head structure; Figure 9 As shown, based on the text recognition network in the cargo identification algorithm, the text detection result is parsed to obtain the text recognition result corresponding to the text detection result, including steps S401 to S404.

[0060] Among them, the convolution layer represents the network structure unit used to extract local features of the input data, so as to establish the feature expression of the original data at different scales; the deconvolution layer represents the network structure unit used to restore the spatial resolution of the features and enhance the expression of data details, so as to gradually restore the low-resolution features to a higher-resolution form; the temporal modeling layer represents the network structure unit used to capture the contextual relationship of characters in the sequence, so as to establish the semantic and sequential association between characters; the detection head structure represents the network structure module that performs content recognition and sequence generation on the received features, so as to recognize and sequence the text content in the text area and output the corresponding structured text recognition results.

[0061] Step S401 : Determine a first comprehensive processing logic combining feature extraction and feature modeling by combining the convolution processing logic corresponding to the convolution layer and the global relationship modeling processing logic corresponding to the temporal modeling layer.

[0062] Among them, the convolution processing logic corresponding to the convolution layer represents the processing mechanism for extracting local image features of the input text area, that is, capturing the character edges, contours and texture structures through a sliding window method to construct the local spatial expression of the characters.

[0063] Among them, the global relationship modeling processing logic corresponding to the temporal modeling layer represents a processing mechanism for establishing contextual semantic associations formed by characters in the text area according to their arrangement order, so as to express the global sequence relationship and logical order between characters.

[0064] Among them, the first comprehensive processing logic represents a feature processing path jointly constructed based on the convolution processing logic and the global relationship modeling logic, which is used to simultaneously complete the spatial feature extraction and sequence semantic modeling of the text area, thereby outputting character features with a temporal structure.

[0065] In step S402 , the second comprehensive processing logic combining feature restoration and feature modeling is determined by combining the deconvolution processing logic corresponding to the deconvolution layer and the global relationship modeling processing logic corresponding to the temporal modeling layer.

[0066] Among them, the deconvolution processing logic corresponding to the deconvolution layer represents the processing mechanism used to restore abstract features to higher-resolution expressions, emphasizing the reconstruction of data details and the recovery of character structure, thereby supplementing character edge information and repairing visual details lost during the compression process.

[0067] Among them, the second comprehensive processing logic represents a feature reconstruction path constructed based on the deconvolution processing logic and the global relationship modeling logic, which is used to maintain the semantic coherence between character sequences while restoring the data spatial structure, thereby outputting restored features with high resolution and semantic structure.

[0068] In step S403 , the first comprehensive processing logic and the second comprehensive processing logic are combined to perform comprehensive processing on the text detection result by combining feature extraction, modeling and restoration to obtain output features.

[0069] Step S404: Based on the detection head structure, the output features are detected and analyzed to obtain a text recognition result corresponding to the text detection result.

[0070] For example, Figure 10 As shown in the figure, the text recognition network is a convolution-deconvolution text recognition network based on global relationship modeling, which can accurately recognize text in complex backgrounds, distortions, and multi-language mixed scenes. Specifically, in the text recognition network, the text detection results are Input into Conv Block1 (i.e. the first convolution block) and Conv Block2 (i.e. the second convolution block) in sequence for convolution processing to obtain the feature ; The feature Input to Transformer Block1 (i.e. the first time series building module) for feature modeling processing to obtain the feature ;Will Input to Conv Block3 (the third convolution block) for convolution processing to obtain the feature ;Will Input to DeconvBlock1 (the first deconvolution block) for deconvolution processing to obtain the feature ;Will Input to TransformerBlock2 (the second time series building module) for feature modeling processing, and obtain ;Will Sequentially input to DeconvBlock2 (i.e. the second deconvolution block) and Deconv Block3 (i.e. the third deconvolution block) for deconvolution processing to obtain the feature ;Will Input into the detection head structure for detection and analysis to obtain the final text recognition result .

[0071] Furthermore, the processing logic of the above text recognition network can refer to formula (5): (5) In formula (5), Represents the convolution operations performed sequentially by Conv Block1 and Conv Block2; Represents the global relationship modeling operation performed by Transformer Block1; Represents the convolution operation performed by Conv Block3; Represents the deconvolution operation performed by Deconv Block1; Represents the global relationship modeling operation performed by Transformer Block2; Denotes the convolution operations performed sequentially by Deconv Block2 and Deconv Block3; Represents a function composite symbol, used to indicate that functions are executed sequentially.

[0072] Furthermore, 、 、 The combined processing logic corresponds to the first comprehensive processing logic in this embodiment; 、 、 The combination processing logic corresponds to the second comprehensive processing logic in this embodiment.

[0073] In this embodiment, first, according to the first comprehensive processing logic constructed by the convolution layer and the temporal modeling layer, the local structural features of the characters are extracted and the semantic associations between the characters are established, thereby enhancing the sequence consistency of the recognition basic expression; secondly, according to the second comprehensive processing logic constructed by the deconvolution layer and the temporal modeling layer, the detail restoration and semantic continuity of the characters are enhanced, thereby improving the resolution ability of the restored features; thirdly, according to the joint application of the first and second comprehensive processing logics, the extraction, modeling and restoration of features are jointly realized, and a unified output feature with spatial details and semantic expression is obtained; thirdly, the output features are detected and analyzed according to the detection head structure, thereby restoring the character information in the image into a structured text recognition result; based on this, accurate restoration and sequential recognition of the character content in the text image are achieved to ensure the integrity and reliability of the text recognition results.

[0074] In an exemplary embodiment, Figure 11 As shown, according to the entity feature category in the cargo information, a cargo identification algorithm matching the cargo information is determined, and cargo feature information corresponding to the cargo information is identified based on the cargo identification algorithm, including steps S501 to S504.

[0075] In step S501, if the entity feature category in the cargo information is an image feature category, a preset algorithm based on hybrid attention and temporal fusion enhancement is used as a cargo recognition algorithm for cargo information matching.

[0076] Among them, the image feature category represents the type of visual data identified from the cargo information and composed of visual content such as image structure, texture, and contour, such as the cargo outer packaging image, appearance, structural contour and other visual forms of content.

[0077] Among them, the algorithm based on hybrid attention and temporal fusion enhancement represents a joint recognition processing flow that combines the spatial attention mechanism with the multi-scale temporal information modeling method, which is used to simultaneously enhance the responsiveness of key areas in the image and the structural consistency between multi-layer features.

[0078] For example, if cargo information is detected to highlight features such as object structure, edges, and color distribution, the cargo information is considered to belong to the image feature category, indicating the presence of a recognizable object in the image. An algorithm based on hybrid attention and temporal fusion enhancement can be used as a matching cargo recognition algorithm to facilitate subsequent more structured recognition of the image data. This type of algorithm builds on deep visual perception mechanisms and incorporates attention mechanisms and temporal modeling logic to enhance the perception of the linkage between the spatial position and local details of target objects in an image.

[0079] Step S502 : Based on the backbone structure of the cargo identification algorithm, feature extraction processing is performed on the cargo information to obtain feature information of different scales.

[0080] Among them, the backbone structure represents the network structure module used to extract multi-scale semantic features, so as to extract low-level and high-level semantic features such as texture, edge, and shape from the input image.

[0081] Exemplarily, based on the backbone network in the cargo identification algorithm structure, multi-stage feature extraction processing is performed on the input cargo information; specifically, the cargo information enters the convolutional units in the backbone structure layer by layer, and the resolution, receptive field and response distribution of the feature map output by each layer are retained and recorded, thereby extracting feature information of different semantic levels such as structure, texture, edge, shape, etc. from the original image layer by layer to form a stable feature expression under multi-scale distribution; wherein, the backbone structure extracts and compresses local responses in the image by utilizing continuous feature transformations, thereby forming a layer-by-layer abstract expression of the characteristics of the object in the image.

[0082] Step S503: Based on the neck structure in the cargo identification algorithm, feature information of different scales is fused to obtain fused features of different scales.

[0083] Among them, the neck structure represents the network structure module used to integrate feature information of different scales, so as to fuse low-level detail information with high-level semantic information and generate unified and more expressive fusion features of different scales.

[0084] For example, based on the neck network in the cargo identification algorithm structure, feature information of different scales is fused. That is, feature information of different scales is uniformly modeled and fused according to dimensions such as feature hierarchy, spatial size, and response strength to form a fused feature that contains both low-level details and high-level semantics. Furthermore, during the fusion process, a structural enhancement mechanism can be introduced to strengthen the ability to distinguish between the target area and the background, thereby improving the accuracy of the fused feature's response to the target area. Ultimately, the fused features of different scales obtained after the fusion are completed have stronger target expression capabilities at each scale, higher context modeling capabilities, and semantic coverage.

[0085] In step S504, based on the detection head structure in the cargo recognition algorithm, the fusion features of different scales are detected and analyzed to obtain the image recognition results corresponding to the cargo information, and the image recognition results are used as the cargo feature information corresponding to the cargo information.

[0086] Among them, the detection head structure represents the network structure module used to perform region discrimination and boundary regression tasks based on fusion features at different scales, so as to perform target positioning, category recognition or confidence prediction of cargo targets in images.

[0087] Among them, the image recognition result identifier is the image description data output by the image recognition result after detecting and analyzing the fusion features of different scales, which is used to represent the specific recognition object in the image and its position, category, confidence level and other information in the image.

[0088] Exemplarily, based on the detection head structure within the cargo recognition algorithm, fused features at different scales are detected and analyzed to identify and determine the boundaries of target areas, thereby generating a complete image recognition result and extracting cargo feature information. Specifically, a sliding region analysis is first performed on the fused features, calculating the response strength of each region in the image along the feature dimension, block by block, to determine whether the corresponding region contains the target to be identified. Subsequently, the feature distribution within the target region containing the target to be identified is classified and positionally regressed, and identification and positioning information for the cargo target in each target region is output, such as the corresponding product type, confidence score, and bounding box position, thereby forming the image recognition result.

[0089] Based on this, in actual warehousing scenarios, goods are often found in complex environments such as dense stacking, partial occlusion, and structural overlap. Traditional cargo identification and counting methods often cannot accurately distinguish between individual cargo entities when encountering occlusion, which can easily lead to missed or redundant counts. This embodiment aims to improve the ability to accurately identify and count goods in the presence of occlusion. Specifically, by constructing a recognition framework with global context understanding capabilities, even if there are partially invisible areas on the surface of the goods, it can be effectively restored and inferred based on their edges, texture extensions, or spatial position relationships, achieving complete recovery of the structure in the invisible areas, thereby accurately identifying the category and quantity of obscured goods, and having stronger occlusion robustness and counting accuracy.

[0090] In this embodiment, first, if the entity feature category in the cargo information is an image feature category, an algorithm based on hybrid attention and temporal fusion enhancement is matched, thereby achieving targeted adaptation of the recognition algorithm to the information type and improving the accuracy and efficiency of subsequent processing; secondly, feature information of different scales is extracted based on the backbone structure, thereby enhancing the expression ability of multi-level visual details and semantic structures; thirdly, feature information of different scales is fused based on the neck structure to obtain feature information of different scales, thereby improving the consistency of feature expression and the response integrity of the target area; thirdly, the image recognition result is output based on the detection head structure, thereby achieving structured detection and positioning expression of cargo targets in the image; based on this, stable recognition and standardized result generation of image-based cargo information are achieved, improving the accuracy of image recognition and the versatility of the system.

[0091] In an exemplary embodiment, the backbone structure includes a convolutional optimization layer and an attention enhancement layer; Figure 12 As shown, based on the backbone structure in the cargo identification algorithm, feature extraction processing is performed on cargo information to obtain feature information of different scales, including steps S601 to S604.

[0092] Among them, the convolution optimization layer represents the network structure unit used to efficiently extract features from image information, so as to improve the extraction accuracy and computational efficiency through convolution structure adjustment or convolution parameter design; the attention enhancement layer represents the network structure unit used to improve the expression quality of feature map channels, which combines multiple attention modeling mechanisms to adjust the channel response intensity to strengthen key feature channels and suppress redundant information, thereby improving the expression stability of the target area.

[0093] Step S601: Based on the current convolution optimization layer, optimized feature extraction processing is performed on the cargo information to obtain feature information output by the current convolution optimization layer.

[0094] Step S602: Based on the channel attention modeling module, self-attention modeling module and C3K modeling module in the current attention enhancement layer, the feature information output by the current convolution optimization layer is comprehensively processed by combining data connection, channel dimension transformation and channel dimension segmentation to obtain the enhanced feature information output by the current attention enhancement layer.

[0095] Among them, the channel attention modeling module represents the modeling submodule in the attention enhancement layer that constructs attention weights based on the statistical characteristics and response distribution between feature map channels, so as to selectively enhance channels with discriminative capabilities; the self-attention modeling module represents the modeling submodule in the attention enhancement layer that is used to model the contextual dependencies between channels, so as to realize the dynamic interaction of information between different channels through internal feature similarity calculation; the C3K modeling module represents the modeling submodule in the attention enhancement layer that performs feature alignment and completion by fusing the convolutional feature path and the context modeling path, so as to comprehensively improve the spatial details of the feature map and the consistency of the structural expression between channels.

[0096] Step S603: Input the enhanced feature information output by the current attention enhancement layer to the next convolution optimization layer for processing to obtain the feature information output by the next convolution optimization layer; input the feature information output by the next convolution optimization layer to the next attention enhancement layer for processing to obtain the enhanced feature information output by the next attention enhancement layer, until the feature information output by all convolution optimization layers is traversed.

[0097] Step S604: The feature information output by different convolution optimization layers is used as feature information of different scales output by the backbone structure.

[0098] For example, Figure 13 As shown, the algorithm based on hybrid attention and temporal fusion enhancement in this embodiment is a cargo recognition and detection algorithm based on hybrid attention and Transformer fusion enhancement. The input of the algorithm can be a single image or a panoramic image synthesized from multiple perspectives to reduce or eliminate the influence of occlusion. Input to the Backbone part (i.e. backbone structure) for data processing, specifically: The output features of Layer 2 are obtained by sequentially inputting them into Layer 1 (i.e. the first convolution optimization layer), ACK1 (i.e. the first attention enhancement layer), and Layer 2 (i.e. the second convolution optimization layer) for convolution processing, attention enhancement processing, and convolution processing respectively. ;Will The input is sequentially sent to ACK2 (the second attention enhancement layer) and Layer3 (the third convolution optimization layer) for attention enhancement processing and convolution processing respectively, and the output features of Layer3 are obtained. ;Will The data is sequentially input to ACK3 (the third attention enhancement layer) and Layer4 (the fourth convolution optimization layer) for attention enhancement and convolution processing respectively, and the output features of Layer4 are obtained. ; Finally, 、 、 They are respectively used as input of the Neck part (i.e. neck structure).

[0099] Furthermore, in the Neck part, the high-level features input from the Backbone part are fused with the low-level features, thereby utilizing the high resolution of the low-level features and the rich semantic information of the high-level features, and supporting the back-end detection head part (i.e., the detection head structure) to independently predict multi-scale features; further, as Figure 13 As shown in the figure, in the detection head part, the One2Many Head (one-to-many detection head) and One2One Head (one-to-one detection head) in YoLov11 (an "end-to-end" single-stage target detection model) are used to detect features of different scales respectively, obtain the detection output of the input image, and obtain the image recognition result. .

[0100] Furthermore, if Figure 14 As shown in the figure, the ACK module (i.e., attention enhancement layer) in the cargo recognition and detection algorithm is a C3K network structure with enhanced channel attention and self-attention implemented on the basis of the original C3K module (i.e., a convolutional network with an attention mechanism introduced in YoLov11), and includes a channel attention modeling module. , self-attention modeling module and C3K modeling module Specifically, in the attention enhancement layer, the input features are processed by 1x1 Conv (i.e., 1x1 size convolution layer). Perform feature channel dimension transformation, and then split the output features of 1x1 Conv through Split to obtain 、 and .

[0101] In the self-attention modeling module In Sequentially input to Multi-head Attention1 (i.e. the first multi-head attention layer) and Multi-head Attention2 (i.e. the second multi-head attention layer) for self-attention enhancement processing to obtain the self-attention modeling module Output features .

[0102] In the channel attention modeling module In The output features of Multi-head Attention1 are concatenated through Concat1 (the first connection layer), and then the output features of Concat1 are input into ChannelAttention (the channel attention layer) for channel attention enhancement processing to obtain the channel attention modeling module. Output features .

[0103] In the C3K modeling module In The output features of the C3K modules at different layers are sequentially input into the C3K modules at different layers for attention enhancement processing, so that the output features of the C3K modules at different layers are used as the C3K modeling modules. Output features .

[0104] Based on this, 、 、 The concatenation is performed through Concat2 (the second connection layer), and then the feature channel dimension is restored through 1x1 Conv to obtain the output features of the ACK module. .

[0105] Furthermore, the processing logic of the ACK module can refer to equations (5) to (6): (5) (6) In formulas (5) to (6), Represents the channel attention modeling module Output features , Represents the self-attention modeling module Output features , Represents the C3K modeling module Output features ; Indicates the connection operation performed by Concat2. Represents the channel dimension transformation operation performed by 1x1 Conv, Indicates the splitting operation along the channel dimension performed by Split.

[0106] Furthermore, the convolution optimization layer in the cargo identification and detection algorithm is a convolutional network layer composed of deformable convolution and depthwise separable convolution. It can not only extract the shape features of irregular objects, but also reduce the time overhead of the model in the inference stage based on its characteristics of few parameters and small computational complexity.

[0107] In this embodiment, first, optimized feature extraction processing is performed on the cargo information according to the convolution optimization layer, thereby improving the clarity and stability of the feature information in local expressions such as edges and textures; secondly, the feature information is enhanced according to the multi-module modeling mechanism in the attention enhancement layer, thereby strengthening the responsiveness of key channels and improving the discriminability of features and the semantic coordination between channels; thirdly, according to the layer-by-layer iterative processing mechanism of the convolution optimization layer and the attention enhancement layer, multi-scale construction and semantic layered expression of the feature extraction process are realized; thirdly, according to the unified summary processing of the feature information output by each convolution optimization layer, multi-scale feature information covering different receptive fields and semantic levels is obtained; based on this, through the structural design combining multi-layer convolution optimization extraction and attention enhancement, deep feature modeling and multi-scale expression of cargo images are realized.

[0108] In an exemplary embodiment, the neck structure includes a multi-scale temporal feature fusion layer and an intermediate layer, wherein the intermediate layer includes at least one of an upsampling layer, a convolution optimization layer, and an attention enhancement layer; Figure 15 As shown, based on the neck structure in the cargo identification algorithm, feature information of different scales is fused to obtain fused features of different scales, including steps S701 to S704.

[0109] Among them, the multi-scale temporal feature fusion layer represents the network structure unit used to model the temporal relationship and jointly fuse feature information of different scales, so as to establish cross-scale and cross-time contextual relationships between multiple groups of feature information with inconsistent spatial scales, and improve the continuity and semantic expression ability of the fused features.

[0110] Among them, the intermediate layer represents a network structure unit or a collection of multiple network structure units used to perform intermediate processing of feature information of different scales, so as to perform structural adjustment, dimension reconstruction or response enhancement on the input feature information before multi-scale feature fusion, thereby improving the adaptability and expression quality of subsequent fusion processing.

[0111] Among them, the upsampling layer represents the network structure unit used to restore the size of the feature information, so as to increase the spatial resolution of the deep feature map to the same level as the shallow feature map. Figure 1 The convolution optimization layer represents the network structure unit for efficient feature extraction of image information, so as to improve the extraction accuracy and computational efficiency by adjusting the convolution structure or designing the convolution parameters. The attention enhancement layer represents the network structure unit for improving the expression quality of the feature map channel, which combines multiple attention modeling mechanisms to adjust the channel response strength to strengthen the key feature channels and suppress redundant information, thereby improving the expression stability of the target area.

[0112] Step S701 : determining feature information of a first target scale from feature information of different scales, and performing parsing processing on the feature information of the first target scale based on the current intermediate layer to obtain intermediate feature information output by the current intermediate layer.

[0113] In step S702, feature information of a second target scale is determined from feature information of different scales, and the intermediate feature information and the feature information of the second target scale are input into the current multi-scale temporal feature fusion layer for comprehensive processing combining feature splicing and feature similarity calculation to obtain fused features output by the current multi-scale temporal feature fusion layer.

[0114] Step S703: Input the fused features output by the current multi-scale temporal feature fusion layer to the next intermediate layer for processing to obtain the intermediate feature information output by the next intermediate layer; input the intermediate feature information output by the next intermediate layer to the next multi-scale temporal feature fusion layer for processing to obtain the fused features output by the next multi-scale temporal feature fusion layer, until all the fused features output by the multi-scale temporal feature fusion layers are traversed.

[0115] Step S704 : determining fusion features of different scales output by the neck structure from the fusion features output by all multi-scale temporal feature fusion layers.

[0116] For example, Figure 13 As shown, 、 、 As the input of the Neck part, specifically: The data is sequentially input to ACK4 (i.e. the fourth attention enhancement layer), Upsample1 (i.e. the first upsampling layer), and ACK5 (i.e. the fifth attention enhancement layer) for attention enhancement processing, upsampling processing, and attention enhancement processing respectively, and then The output features of ACK5 are fused with the scale time series features through Fusion1 (i.e. the first multi-scale time series feature fusion layer) to obtain the output features of Fusion1; the output features of Fusion1 are input to Upsample2 (i.e. the second upsampling layer) for upsampling, and then the output features of Upsample2 are combined with The feature fusion process is performed through Fusion2 (the second multi-scale temporal feature fusion layer) to obtain the output features of Fusion2 ;Will Input to Layer5 (the fifth convolution optimization layer) for convolution processing, and then the output features of Layer5 and the output features of Fusion1 are fused through Fusion3 (the third multi-scale temporal feature fusion layer) to obtain the output features of Fusion3. ;Will The output features of ACK6 and ACK4 are fused through Fusion4 (the fourth multi-scale temporal feature fusion layer) to obtain the output features of Fusion4. ; Finally, 、 、 They are respectively used as inputs of the detection head.

[0117] Furthermore, if Figure 16 As shown in Figure 2, the multi-scale temporal feature fusion layer in the cargo recognition and detection algorithm is a multi-scale feature fusion network structure based on Transformer. Specifically, the input features of two different levels are combined. and The splicing feature is obtained by splicing through Concat (i.e., the connection layer), and then the splicing feature is subjected to channel dimension reduction processing through 1x1 Conv (i.e., 1x1 size convolution layer) to obtain the converted feature; the similarity of a pair of converted features (i.e., the similarity between two different levels of input features) is calculated through multiplication operation, and then the similarity is multiplied with the original converted feature through multiplication operation to obtain the calculation result; the calculation result is subjected to channel dimension restoration processing through 1x1 Conv to obtain the restored feature, and then the restored feature is fused with the original splicing feature through addition operation to obtain the output feature of the multi-scale temporal feature fusion layer .

[0118] Furthermore, the processing logic of the multi-scale temporal feature fusion layer can refer to formula (7): (7) In formula (7), Indicates the concatenation operation performed by Concat; Represents the channel dimension transformation operation performed by 1x1 Conv, and in order to enhance the feature representation capability of the multi-scale temporal feature fusion layer, Figure 15 The two types of 1x1 Conv in do not share parameters with each other; represents the multiplication operation, Represents the addition operation.

[0119] In this embodiment, first, a first target scale is selected from the feature information of different scales and input into the intermediate layer for analysis and processing to obtain intermediate feature information, thereby realizing preprocessing and structural optimization of the original feature information; secondly, a second target scale is selected from the feature information of different scales, and feature splicing and similarity analysis are performed on the intermediate feature information and the second target scale feature information according to the multi-scale temporal feature fusion layer to obtain fused features, thereby improving the temporal consistency expression capability between cross-scale features; thirdly, according to the iterative processing logic of the multi-scale temporal feature fusion layer and the intermediate layer at each level, step-by-step enhancement and temporal depth expansion of multi-layer fusion features are realized; thirdly, according to the fusion features output by all multi-scale temporal feature fusion layers, the coordinated output and unified expression of multi-scale information are guaranteed; based on this, the structural consistency enhancement and temporal expression optimization of the feature information of different scales are realized, thereby improving the integrity and robustness of the image feature fusion processing.

[0120] In an exemplary embodiment, Figure 17 As shown, according to the entity feature category in the cargo information, a cargo identification algorithm matching the cargo information is determined, and cargo feature information corresponding to the cargo information is identified based on the cargo identification algorithm, including steps S801 to S803.

[0121] Step S801: If the entity feature category in the cargo information is an identification feature category, a preset identification recognition algorithm is used as a cargo recognition algorithm for cargo information matching.

[0122] Among them, the identification feature category represents the data type category that uniquely identifies the goods identified from the goods information, and is used to quickly locate and identify the identity information of the goods. For example, visual coding data such as barcodes and QR codes presented in the form of images, or radio frequency data presented in the form of radio frequency responses, can represent the information belonging to the identification feature category.

[0123] Among them, the identification recognition algorithm represents the logical processing flow designed for identification feature categories, used to identify, extract and parse identification information, and is used to convert visual coding data or radio frequency data into structured identification fields with business significance.

[0124] For example, first, if identification data with a coding structure, identification tag, or other unique identification function is detected in the cargo information, the entity feature category in the cargo information is classified as an identification feature category. Furthermore, after determining the identification feature category, to ensure consistency between subsequent processing methods and data types, the processing logic designed for identification data is selected as the cargo identification algorithm to parse various types of identification information, including but not limited to visual graphic codes or electronic radio frequency data. Furthermore, this matching process is not a simple static rule search, but rather a combined judgment based on preliminary analysis of the data structure and device configuration status, ensuring the identification algorithm has sufficient adaptability and processing coverage.

[0125] Step S802: Based on the identification recognition module configured in the drone device and matching the identification feature category, the identification information in the cargo information is identified. When the identification recognition module represents a visual coding recognition module, the identification information represents one of a barcode identification and a QR code identification. When the identification recognition module represents an RFID module, the identification information represents an RFID identification.

[0126] Among them, the identification recognition module refers to a set of functional components configured in the drone equipment for obtaining and identifying the identification information carried by the goods; the visual coding recognition module refers to the acquisition component used to collect the visual coding data attached to the surface of the goods, such as a camera device, image perception device, etc.; the radio frequency identification module refers to the wireless communication component used to receive the radio frequency signal emitted by the goods and extract the radio frequency data therein, such as a radio frequency receiving device, etc.

[0127] Among them, the identification information represents the coded data content extracted by the identification module to represent the identity of the goods; the barcode identification and QR code identification represent the coded data content that expresses the identity of the goods in the form of visual coding, and the radio frequency identification identification represents the coded data content that expresses the identity of the goods in the form of radio frequency response. Exemplarily, based on the category of identification data in the cargo information, the identification recognition module configured in the drone equipment is adaptively called to realize the acquisition operation of the identification information carried by the cargo; specifically, based on the presentation method of the identification characteristics, the adaptive collection channel is automatically selected, so that the identification information can be effectively extracted without human intervention. On the one hand, when the cargo information carries identification content in the form of visual coding, the visual coding recognition module scans the surface of the cargo or the collected cargo information and extracts the corresponding shape code identification or QR code identification; on the other hand, when the cargo information carries identification content in the form of radio frequency response, the radio frequency identification module receives radio frequency data from the radio frequency electronic tag set on the surface of the cargo or the collected cargo information to obtain the corresponding radio frequency identification identification.

[0128] Step S803: Based on a cargo identification algorithm, the identification information is parsed to obtain cargo characteristic information corresponding to the cargo information.

[0129] For example, for identification information in the form of visual coding, the identification information is input into the image decoding logic. By locating the coding area in the image, restoring the pattern structure and performing format correction, the content carried in the barcode identification or QR code identification is restored. For identification information in the form of radio frequency response, the received radio frequency data is subjected to protocol identification, displacement verification and field extraction based on the radio frequency decoding logic, and the electronic coding value stored in the radio frequency electronic tag is extracted, and the field is disassembled according to the preset protocol structure. Based on this, the decoded data content is further mapped to fields according to a unified format template, thereby obtaining cargo feature information with a standardized and comparable structure. Based on this, the structural restoration of the original identification content is completed through the image decoding logic or the radio frequency decoding logic, so that the coded information in the physical carrier is converted into logical fields that can be recognized by the system, realizing the key processing link of the transition from the perception layer to the data layer.

[0130] In this embodiment, first, if the entity feature category in the cargo information is an identification feature category, the identification recognition algorithm is matched to ensure that the information recognition path is compatible with the data type; secondly, the visual coding recognition module or the radio frequency identification module is called according to the identification feature category to identify the identification information, thereby realizing targeted collection of various identification forms and enhancing the adaptability to diversified cargo identification; thirdly, the identification information is parsed according to the image decoding logic or the radio frequency decoding logic, thereby realizing accurate conversion of the coded data into cargo feature information. Based on this, a stable recognition process for different identification features is constructed, realizing efficient recognition and standardized parsing of cargo identification information.

[0131] It should be understood that, although the various steps in the flowcharts involved in the various embodiments described above are displayed in sequence according to the instructions of the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts involved in the various embodiments described above can include multiple steps or multiple stages, and these steps or stages are not necessarily executed and completed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of steps or stages in other steps.

[0132] Based on the same inventive concept, the embodiments of the present application also provide a drone-based entity object traversal identification and detection device for implementing the aforementioned drone-based entity object traversal identification and detection method. The implementation solution provided by this device is similar to the implementation solution described in the aforementioned method. Therefore, the specific limitations of one or more drone-based entity object traversal identification and detection device embodiments provided below can be found in the above-mentioned limitations of the drone-based entity object traversal identification and detection method, and will not be repeated here.

[0133] In an exemplary embodiment, Figure 18 As shown, a device for detecting entity object traversal recognition based on a drone device is provided, comprising: an acquisition module 101, an identification module 102, and a comparison module 103, wherein: An acquisition module 101 is configured to acquire cargo information collected by a drone in a cargo storage area. The cargo information represents cargo storage-related information collected by the drone in sequence from each storage location in the cargo storage area while operating along a preset flight trajectory. Identification module 102, configured to determine a cargo identification algorithm that matches the cargo information based on the entity feature category in the cargo information, and identify cargo feature information corresponding to the cargo information based on the cargo identification algorithm, the cargo feature information including the product model, storage quantity, and storage location of the corresponding cargo; The comparison module 103 is used to compare the cargo feature information with the inventory information in the preset database to obtain the inventory analysis results of the cargo storage area.

[0134] In an exemplary embodiment, the recognition module 102 includes a first recognition unit, which is configured to: if the entity feature category in the cargo information is a text feature category, use a preset text detection and recognition algorithm based on relational modeling as a cargo recognition algorithm for matching the cargo information; perform multi-scale feature extraction and fusion processing on the cargo information based on the text detection network in the cargo recognition algorithm to obtain a text detection result corresponding to the cargo information; and parse the text detection result based on the text recognition network in the cargo recognition algorithm to obtain a text recognition result corresponding to the text detection result, and use the text recognition result as cargo feature information corresponding to the cargo information.

[0135] In an exemplary embodiment, the first recognition unit is further used to: perform multi-scale feature extraction processing on the cargo information based on the convolution layer to obtain feature maps of different scales; combine the linear interpolation and deconvolution combined processing logic corresponding to the multiple upsampling layers and the global information modeling and splicing combined processing logic corresponding to the feature fusion layer to jointly upsample and fuse the feature maps of different scales to obtain a fused feature map; based on the detection head structure, perform detection and analysis on the fused feature map to obtain text area detection information based on the cargo information, and use the text area detection information as the text detection result corresponding to the cargo information.

[0136] In an exemplary embodiment, the first recognition unit is also used to: determine a first comprehensive processing logic combining feature extraction and feature modeling in combination with the convolution processing logic corresponding to the convolution layer and the global relationship modeling processing logic corresponding to the temporal modeling layer; determine a second comprehensive processing logic combining feature restoration and feature modeling in combination with the deconvolution processing logic corresponding to the deconvolution layer and the global relationship modeling processing logic corresponding to the temporal modeling layer; combine the first comprehensive processing logic and the second comprehensive processing logic to perform comprehensive processing on the text detection results combining feature extraction, modeling and restoration to obtain output features; based on the detection head structure, detect and analyze the output features to obtain a text recognition result corresponding to the text detection result.

[0137] In an exemplary embodiment, the recognition module 102 includes a second recognition unit, which is configured to: if the entity feature category in the cargo information is an image feature category, use a preset algorithm based on hybrid attention and temporal fusion enhancement as a cargo recognition algorithm for matching the cargo information; based on the backbone structure in the cargo recognition algorithm, perform feature extraction processing on the cargo information to obtain feature information of different scales; based on the neck structure in the cargo recognition algorithm, perform fusion processing on the feature information of different scales to obtain fusion features of different scales; based on the detection head structure in the cargo recognition algorithm, perform detection and analysis on the fusion features of different scales to obtain an image recognition result corresponding to the cargo information, and use the image recognition result as the cargo feature information corresponding to the cargo information.

[0138] In an exemplary embodiment, the second recognition unit is also used to: based on the current convolution optimization layer, optimize the feature extraction processing of the cargo information to obtain the feature information output by the current convolution optimization layer; based on the channel attention modeling module, the self-attention modeling module and the C3K modeling module in the current attention enhancement layer, perform a comprehensive processing of the feature information output by the current convolution optimization layer by combining data connection, channel dimension transformation, and channel dimension segmentation to obtain the enhanced feature information output by the current attention enhancement layer; input the enhanced feature information output by the current attention enhancement layer to the next convolution optimization layer for processing to obtain the feature information output by the next convolution optimization layer, input the feature information output by the next convolution optimization layer to the next attention enhancement layer for processing to obtain the enhanced feature information output by the next attention enhancement layer, until the feature information output by all convolution optimization layers is traversed; and use the feature information output by different convolution optimization layers as the feature information of different scales output by the backbone structure.

[0139] In an exemplary embodiment, the second identification unit is also used to: determine the feature information of the first target scale in the feature information of different scales, analyze and process the feature information of the first target scale based on the current intermediate layer, and obtain the intermediate feature information output by the current intermediate layer; determine the feature information of the second target scale in the feature information of different scales, input the intermediate feature information and the feature information of the second target scale into the current multi-scale temporal feature fusion layer for comprehensive processing combining feature splicing and feature similarity calculation, and obtain the fusion features output by the current multi-scale temporal feature fusion layer; input the fusion features output by the current multi-scale temporal feature fusion layer into the next intermediate layer for processing to obtain the intermediate feature information output by the next intermediate layer, input the intermediate feature information output by the next intermediate layer into the next multi-scale temporal feature fusion layer for processing to obtain the fusion features output by the next multi-scale temporal feature fusion layer, until the fusion features output by all multi-scale temporal feature fusion layers are traversed; determine the fusion features of different scales output by the neck structure in the fusion features output by all multi-scale temporal feature fusion layers.

[0140] In an exemplary embodiment, the identification module 102 includes a third identification unit, which is used to: if the entity feature category in the cargo information is an identification feature category, use a preset identification recognition algorithm as a cargo identification algorithm for matching the cargo information; based on an identification recognition module configured in the drone device that matches the identification feature category, identify the identification information in the cargo information, when the identification recognition module represents a visual coding recognition module, the identification information represents one of a barcode identification and a QR code identification, and when the identification recognition module represents a radio frequency identification module, the identification information represents a radio frequency identification identification; based on the cargo identification algorithm, parse the identification information to obtain cargo feature information corresponding to the cargo information.

[0141] Each module in the aforementioned drone-based entity object traversal identification and detection device can be implemented in whole or in part through software, hardware, or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in hardware form, or stored in a computer device memory in software form, so that the processor can call and execute the corresponding operations of each module.

[0142] In an exemplary embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps in any of the above embodiments when executing the computer program.

[0143] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in any of the above embodiments are implemented.

[0144] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above-mentioned embodiments. In particular, any reference to memory, database, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processors involved in the various embodiments provided herein may be, but are not limited to, general-purpose processors, central processing units (CPUs), graphics processing units (GPUs), digital signal processors (DSPs), programmable logic devices (PLDs), data processing logic devices based on quantum computing, and the like.

[0145] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0146] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.

Claims

1. A method for traversal recognition and detection of entity objects based on drone equipment, characterized in that: The method comprises: Obtaining cargo information collected by the drone device in the cargo storage area, wherein the cargo information represents cargo storage-related information collected sequentially from each storage location in the cargo storage area by the drone device operating along a preset flight trajectory; determining, based on the entity feature category in the cargo information, a cargo identification algorithm that matches the cargo information, and identifying cargo feature information corresponding to the cargo information based on the cargo identification algorithm, the cargo feature information including a product model, storage quantity, and storage location of the corresponding cargo; The cargo characteristic information is compared with the inventory information in a preset database to obtain an inventory analysis result of the cargo storage area.

2. The method according to claim 1, characterized in that The determining, based on the entity feature category in the cargo information, a cargo identification algorithm that matches the cargo information, and identifying cargo feature information corresponding to the cargo information based on the cargo identification algorithm, includes: If the entity feature category in the cargo information is a text feature category, a preset text detection and recognition algorithm based on relational modeling is used as a cargo recognition algorithm for matching the cargo information; Based on the text detection network in the cargo identification algorithm, multi-scale feature extraction and fusion processing are performed on the cargo information to obtain text detection results corresponding to the cargo information; Based on the text recognition network in the cargo identification algorithm, the text detection result is parsed to obtain a text recognition result corresponding to the text detection result, and the text recognition result is used as cargo feature information corresponding to the cargo information.

3. The method according to claim 2, characterized in that The text detection network includes a convolutional layer, a multi-sampling layer, a feature fusion layer and a detection head structure; The text detection network in the cargo identification algorithm is used to perform multi-scale feature extraction and fusion processing on the cargo information to obtain text detection results corresponding to the cargo information, including: Based on the convolutional layer, multi-scale feature extraction processing is performed on the cargo information to obtain feature maps of different scales; Combined with the linear interpolation and deconvolution combined processing logic corresponding to the multiple upsampling layers and the global information modeling and splicing combined processing logic corresponding to the feature fusion layer, feature maps of different scales are jointly upsampled and fused to obtain a fused feature map; Based on the detection head structure, the fused feature map is detected and analyzed to obtain text region detection information based on the cargo information, and the text region detection information is used as the text detection result corresponding to the cargo information.

4. The method according to claim 2, characterized in that The text recognition network includes a convolution layer, a deconvolution layer, a temporal modeling layer and a detection head structure; The text recognition network in the cargo identification algorithm is used to parse the text detection result to obtain a text recognition result corresponding to the text detection result, including: Determine a first comprehensive processing logic combining feature extraction and feature modeling by combining the convolution processing logic corresponding to the convolution layer and the global relationship modeling processing logic corresponding to the temporal modeling layer; Determine a second comprehensive processing logic combining feature restoration and feature modeling by combining the deconvolution processing logic corresponding to the deconvolution layer and the global relationship modeling processing logic corresponding to the temporal modeling layer; Combining the first comprehensive processing logic and the second comprehensive processing logic, performing comprehensive processing combining feature extraction, modeling, and restoration on the text detection result to obtain output features; Based on the detection head structure, the output features are detected and analyzed to obtain a text recognition result corresponding to the text detection result.

5. The method according to claim 1, wherein The determining, based on the entity feature category in the cargo information, a cargo identification algorithm that matches the cargo information, and identifying cargo feature information corresponding to the cargo information based on the cargo identification algorithm, includes: If the entity feature category in the cargo information is an image feature category, a preset algorithm based on hybrid attention and temporal fusion enhancement is used as the cargo recognition algorithm for matching the cargo information; Based on the backbone structure of the cargo identification algorithm, feature extraction processing is performed on the cargo information to obtain feature information of different scales; Based on the neck structure in the cargo identification algorithm, feature information of different scales is fused to obtain fused features of different scales; Based on the detection head structure in the cargo identification algorithm, fusion features of different scales are detected and analyzed to obtain image recognition results corresponding to the cargo information, and the image recognition results are used as cargo feature information corresponding to the cargo information.

6. The method according to claim 5, characterized in that The backbone structure includes a convolution optimization layer and an attention enhancement layer; The backbone structure in the cargo identification algorithm is used to extract features from the cargo information to obtain feature information of different scales, including: Based on the current convolution optimization layer, performing optimized feature extraction processing on the cargo information to obtain feature information output by the current convolution optimization layer; Based on the channel attention modeling module, the self-attention modeling module, and the C3K modeling module in the current attention enhancement layer, the feature information output by the current convolution optimization layer is comprehensively processed by combining data connection, channel dimension transformation, and channel dimension segmentation to obtain enhanced feature information output by the current attention enhancement layer; Inputting the enhanced feature information output by the current attention enhancement layer to the next convolution optimization layer for processing to obtain the feature information output by the next convolution optimization layer, inputting the feature information output by the next convolution optimization layer to the next attention enhancement layer for processing to obtain the enhanced feature information output by the next attention enhancement layer, until the feature information output by all convolution optimization layers is traversed; The feature information output by different convolution optimization layers is used as the feature information of different scales output by the backbone structure.

7. The method according to claim 5, characterized in that The neck structure includes a multi-scale temporal feature fusion layer and an intermediate layer, wherein the intermediate layer includes at least one of an upsampling layer, a convolution optimization layer, and an attention enhancement layer; Based on the neck structure in the cargo identification algorithm, feature information of different scales is fused to obtain fusion features of different scales, including: Determining feature information of a first target scale from feature information of different scales, and performing parsing processing on the feature information of the first target scale based on the current intermediate layer to obtain intermediate feature information output by the current intermediate layer; Determining feature information of a second target scale from feature information of different scales, inputting the intermediate feature information and the feature information of the second target scale into the current multi-scale temporal feature fusion layer for comprehensive processing combining feature splicing and feature similarity calculation, to obtain fused features output by the current multi-scale temporal feature fusion layer; Input the fused features output by the current multi-scale temporal feature fusion layer to the next intermediate layer for processing to obtain intermediate feature information output by the next intermediate layer, input the intermediate feature information output by the next intermediate layer to the next multi-scale temporal feature fusion layer for processing to obtain the fused features output by the next multi-scale temporal feature fusion layer, until traversing the fused features output by all multi-scale temporal feature fusion layers; The fusion features of different scales output by the neck structure are determined from the fusion features output by all multi-scale temporal feature fusion layers.

8. The method according to claim 1, characterized in that The determining, based on the entity feature category in the cargo information, a cargo identification algorithm that matches the cargo information, and identifying cargo feature information corresponding to the cargo information based on the cargo identification algorithm, includes: If the entity feature category in the cargo information is an identification feature category, a preset identification recognition algorithm is used as a cargo recognition algorithm for matching the cargo information; Identify identification information in the cargo information based on an identification recognition module configured in the drone device and matching the identification feature category, where the identification information represents one of a barcode identification and a QR code identification when the identification recognition module represents a visual coding identification module, and represents an RFID identification when the identification recognition module represents an RFID module; Based on the cargo identification algorithm, the identification information is parsed to obtain cargo feature information corresponding to the cargo information.

9. A device for detecting entity object traversal and recognition based on drone equipment, characterized in that: The device comprises: an acquisition module, configured to acquire cargo information collected by the drone device in the cargo storage area, wherein the cargo information represents cargo storage-related information collected sequentially from each storage location in the cargo storage area by the drone device operating along a preset flight trajectory; an identification module, configured to determine a cargo identification algorithm that matches the cargo information based on the entity feature category in the cargo information, and identify cargo feature information corresponding to the cargo information based on the cargo identification algorithm, wherein the cargo feature information includes a product model, storage quantity, and storage location of the corresponding cargo; The comparison module is used to compare the cargo characteristic information with the inventory information in a preset database to obtain an inventory analysis result of the cargo storage area.

10. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • Intelligent logistics warehouse cargo checking method and system based on unmanned aerial vehicle

    CN109934318A

  • Warehouse management method, device, equipment and storage medium

    CN110532978A

  • Goods allocation identification system and method of mobile robot

    CN119273971A

  • Irregular text recognition method based on deep learning and Raspberry Pi

    CN119741713A

  • Goods warehouse-in and warehouse-out management method and system, computer equipment and storage medium

    CN119887055A

Cited By

  • Goods homing method, device and equipment based on goods missing

    CN120996066A