A dynamic modal warehouse in-out detection method and a warehouse device
By combining environmental and image sensors to dynamically adjust the antenna power and image sensor parameters of the RFID reader, the signal crosstalk problem in the case of multiple tags is solved, achieving high-precision and efficient warehouse inbound and outbound detection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHENZHEN WINIT TECH CO LTD
- Filing Date
- 2025-08-13
- Publication Date
- 2026-08-04
AI Technical Summary
In existing warehouse management, RFID readers are prone to signal crosstalk when reading multiple overlapping tags, which affects the accuracy and efficiency of inbound and outbound detection.
By collecting the motion direction and speed information of objects through environmental sensors, dynamically adjusting the antenna power of the RFID reader and the frame rate and field of view of the image sensor, and combining semantic segmentation and neural networks to identify the quantity and type of objects entering and leaving the warehouse, dynamic modal warehouse entry and exit detection is achieved.
It improves the accuracy and efficiency of warehousing equipment in detecting inbound and outbound items, adapts to the detection needs of different scenarios, and reduces computing power and computational load.
Smart Images

Figure CN120912119B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of warehouse management technology, specifically to a dynamic modal warehouse inbound / outbound detection method and warehouse equipment. Background Technology
[0002] In modern warehouse management, RFID tags are being used to detect the entry and exit of items. However, RFID readers have many drawbacks. For example, when reading multiple overlapping tags, crosstalk can occur between the signals received by the reader, affecting the reading results. Summary of the Invention
[0003] This invention addresses the problem of how to improve the accuracy and / or efficiency of detecting incoming and outgoing items in warehousing equipment.
[0004] To address the aforementioned problems, in a first aspect, this application provides a dynamic modal warehouse entry / exit detection method, applicable to warehouse equipment. The warehouse equipment includes: an environmental sensor for collecting environmental information, wherein the environmental information characterizes the movement direction and speed of the items entering / exiting; an RFID reader for reading the RFID tags of the items entering / exiting; and a first image sensor for acquiring a first image of the items entering / exiting. The detection method includes:
[0005] Based on the environmental information collected by the environmental sensors, the entry and exit methods of items are detected.
[0006] Based on the environmental information collected by the environmental sensor and the storage method, the antenna power of the radio frequency reader, the frame rate of the first image collected by the first image sensor, and / or the field of view are controlled to detect the quantity and type of the stored items.
[0007] In some embodiments, the environmental sensor includes a second image sensor, which is used to acquire multiple frames of second images of items entering and leaving the warehouse; the method further includes:
[0008] Based on the semantic segmentation module, the items entering and leaving the warehouse in the second image are identified; and
[0009] Based on the reading results from the RFID reader, the number of palletized items entering and leaving the warehouse is confirmed.
[0010] In some embodiments, the method further includes:
[0011] In the multiple frames of the second image after the semantic segmentation module identifies the items entering and leaving the warehouse in the second image, the position information of the items entering and leaving the warehouse and the changes in the area ratio of the items entering and leaving the warehouse in the multiple frames of the second image are identified, and the movement speed and direction of the items entering and leaving the warehouse are confirmed.
[0012] In some embodiments, the environmental sensor includes: an audio sensor, the audio sensor being used to collect audio information during the entry and exit of items; the method further includes:
[0013] The audio information after the semantic segmentation module identifies the items entering and leaving the warehouse in the second image is recorded as valid audio information. The valid audio information and multiple frames of the second image at the same time are fed into a pre-trained neural network. The neural network model outputs the movement speed and / or movement direction of the items entering and leaving the warehouse.
[0014] In some embodiments, the effective audio information is fed together with multiple frames of second images at the same time into a pre-trained neural network. The neural network model outputs the motion speed and direction of the objects entering and leaving the warehouse, including:
[0015] Convert the audio information into an audio image;
[0016] The first processor is used to stitch the audio image and the second image together to obtain a stitched image, and then the stitched image is sent to the second processor.
[0017] The second processor segments an audio image and a second image from the stitched image, and sends the audio image and the second image into their respective feature extraction networks to extract the audio feature information and image feature information;
[0018] The fused features are obtained by fusing the audio and image feature information;
[0019] Based on the fusion features, the speed and direction of movement of objects entering and leaving the warehouse are output.
[0020] In some embodiments, the dynamic modal warehouse inbound / outbound detection method further includes:
[0021] Construct a mapping relationship between the movement speed range and pallet quantity of items entering and leaving the warehouse, and the power of the radio frequency antenna, the frame rate of the first image acquired by the first image sensor, and / or the field of view size;
[0022] Based on the movement speed range of items entering and leaving the warehouse, the number of pallets, and the mapping relationship, the power of the radio frequency antenna, the frame rate of the first image acquired by the first image sensor, and / or the field of view size are determined; wherein,
[0023] In the mapping relationship, the power of the radio frequency antenna is positively correlated with the speed range of the items entering and leaving the warehouse and the number of stacks, and the frame rate and / or field of view of the first image acquired by the first image sensor are positively correlated with the speed range of the items entering and leaving the warehouse and the number of stacks.
[0024] In some embodiments, in the mapping relationship, the movement speed of items entering and leaving the warehouse is divided into at least three speed ranges; wherein...
[0025] The power of the radio frequency antenna in each motion zone is different, and the frame rate of the first image acquisition is different, but the field of view is the same.
[0026] Alternatively, the power of the radio frequency antennas in each motion zone may be different, and the field of view for acquiring the first image may be different while the video frame rate may be the same.
[0027] In some embodiments, the radio frequency reader has multiple radio frequency antennas, all of which are directional antennas, referred to as the first radio frequency antenna and the second radio frequency antenna, wherein the first radio frequency antenna faces the inbound direction and the second radio frequency antenna faces the outbound direction; the method further includes:
[0028] When the movement direction of the item entering or leaving the warehouse is the same as the direction of entry, the first radio frequency antenna is activated to read the radio frequency tag of the item entering or leaving the warehouse.
[0029] When the movement direction of the items entering or leaving the warehouse is the outbound direction, the second radio frequency antenna is activated to read the radio frequency tag of the items entering or leaving the warehouse.
[0030] In some embodiments, the first image sensor is a gimbal camera, which is configured to adjust the gimbal camera toward the inbound / outbound object based on the movement direction of the inbound / outbound object.
[0031] Alternatively, the first image sensor includes at least two cameras, referred to as the first camera and the second camera; wherein the first camera faces the outbound direction and the second camera faces the inbound direction; the method further includes:
[0032] When the direction of movement of the item entering or leaving the warehouse is the direction of entry, the first camera is invoked to capture the first image of the item entering or leaving the warehouse.
[0033] When the movement direction of the items entering or leaving the warehouse is the outbound direction, the second camera is invoked to capture the first image of the items entering or leaving the warehouse.
[0034] Secondly, this application provides a storage device, characterized in that it includes:
[0035] An environmental sensor is used to collect environmental information, which represents the direction and speed of movement of items entering and leaving the warehouse.
[0036] Radio frequency (RFID) readers are used to read the RFID tags on items entering and leaving the warehouse.
[0037] The first image sensor, and the first image processor, are used to acquire images of items entering and leaving the warehouse.
[0038] One or more processors are connected to the environmental sensor, the radio frequency reader, and the image sensor. The one or more processors are configured to execute a warehouse entry and exit detection program to implement the dynamic modal warehouse entry and exit detection method described above.
[0039] This application uses environmental sensors to detect environmental information, confirm the entry and exit methods, palletization quantities, and entry and exit speeds of items. Based on these factors, it adjusts the antenna power of the RFID reader and the frame rate and / or field of view of the first image acquired by the auxiliary image sensor. This dynamic adjustment of the RFID reader and the first image sensor meets the detection needs of items entering and leaving the warehouse in different scenarios, improving the accuracy and / or efficiency of warehouse equipment in detecting items entering and leaving the warehouse. Attached Figure Description
[0040] Figure 1 This is a flowchart of an embodiment of the dynamic modal warehouse inbound / outbound detection method of this application;
[0041] Figure 2 This is a flowchart of another embodiment of the dynamic modal warehouse inbound / outbound detection method of this application;
[0042] Figure 3 This is a flowchart of another embodiment of the dynamic modal warehouse entry and exit detection method of this application. Detailed Implementation
[0043] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0044] In the description of this invention, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the stated features. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.
[0045] In the description of this invention, the term "for example" is used to mean "used as an example, illustration, or description." Any embodiment described as "for example" in this invention is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is provided to enable any person skilled in the art to make and use the invention. Details are set forth in the following description for purposes of explanation. It should be understood that those skilled in the art will recognize that the invention can be made without using these specific details. In other instances, well-known structures and processes will not be described in detail to avoid obscuring the description of the invention with unnecessary detail. Therefore, the invention is not intended to be limited to the embodiments shown, but is consistent with the broadest scope of the principles and features disclosed herein.
[0046] Example 1
[0047] This application provides a dynamic modal method for detecting warehouse entry and exit, applicable to warehouse equipment.
[0048] In the embodiments of this application, the storage equipment can be installed at the warehouse entry and exit window or the entry and exit access control point to realize the detection function of the entry and exit equipment.
[0049] Scenario 1: The operator affixes RFID tags to the items entering and leaving the warehouse and sends them in one by one via a conveyor belt. This is one implementation method. In this case, the speed of the conveyor belt can be controlled to not be too fast, and then a reader is used to read the RFID tags one by one.
[0050] Scenario 2: Items are already neatly stacked and not suitable for being transported one by one into the warehouse via conveyor belt. In this case, the reader needs to read a large number of labels simultaneously and requires supplementary information from the user, such as affixing labels to the neatly stacked items indicating their quantity and type. The warehouse equipment uses image recognition to identify the quantity and type of the neatly stacked items and interpret the semantics of the labels. Then, based on the quantity and type of the items obtained from image recognition, the semantics of the labels, and the reader's reading results, it confirms the type and quantity of the items.
[0051] To adapt to the two scenarios mentioned above, the warehousing equipment in this application includes:
[0052] An environmental sensor is used to collect environmental information, which represents the direction and speed of movement of items entering and leaving the warehouse.
[0053] Environmental sensors can be one or more of cameras, microphones, light sensors, and vibration sensors to collect environmental information such as images, audio, and light when items enter or leave the warehouse. Furthermore, a recognition model can be used to determine the direction and speed of movement of the items based on the above information.
[0054] For example, confirming the movement direction and speed of items entering and leaving the warehouse includes: collecting video information of items entering and leaving the warehouse, identifying items entering and leaving the warehouse through a recognition model, calculating the position change of items entering and leaving the warehouse in different video frames and the relative position of the camera to obtain the movement direction of items entering and leaving the warehouse, and further, dividing the position change of items entering and leaving the warehouse in the video frame by the time change of the timestamp of the video frame to obtain the movement speed of items entering and leaving the warehouse.
[0055] For example, confirming the direction and speed of movement of items entering and leaving the warehouse includes: acquiring video information of items entering and leaving the warehouse, identifying the area where the items are located from the video, then using optical flow to calculate the speed of pixels in that area, and calculating their average speed.
[0056] It should be noted that there are many existing methods for calculating the velocity of objects in videos, and these are not limited to the distance & time and optical flow methods mentioned above. Other methods for calculating the velocity of objects in videos may also be included within the scope of this application.
[0057] Radio frequency (RFID) readers are used to read the RFID tags on items entering and leaving the warehouse.
[0058] The antenna power of the RF reader is adjustable.
[0059] In scenario 1, the number of RFID tags that need to be identified at a single moment is small, and the direction is clear (that is, the direction of movement of the conveyor belt). Therefore, the RFID reader does not need to be very powerful; it only needs to adjust its own power according to the speed of the conveyor belt.
[0060] In Scenario 2, the number of RFID tags to be identified at any given time is very large, and they overlap. In this case, the antenna power of the RFID reader needs to be increased to accurately read the RSSI (Received Signal Strength Indication) and phase value of the RFID signals fed back by the tags. This allows for the identification of the tag feedback information, thereby accurately determining the number of tags and the corresponding inbound / outbound item types.
[0061] The first image sensor is used to capture the first image of items entering and leaving the warehouse;
[0062] In scenario 1, the first image sensor can be turned off, as the RF reader can handle scenario 1 well and only needs to dynamically adjust the power of the RF antenna according to the conveyor belt speed.
[0063] In scenario 2, the first sensor needs to collect images of items entering and leaving the warehouse so that the processor can identify the quantity and type of neatly stacked items and the semantics of the labels based on image recognition. Then, based on the quantity and type of items entering and leaving the warehouse obtained from image recognition, the semantics of the labels, and the reading results of the reader, the type and quantity of items entering and leaving the warehouse are confirmed.
[0064] Reference Figure 1 In the embodiments of this application, the dynamic modal warehouse inbound / outbound detection method includes:
[0065] S100. Based on the environmental information collected by the environmental sensor, detect the entry and exit methods of the items entering and leaving the warehouse;
[0066] The methods for entering and leaving the warehouse include "sending items one by one into the warehouse via conveyor belt" in scenario 1, and "attaching labels to neatly stacked items, indicating the quantity and type of the items, and then carrying out the entire stacked group of items in and out of the warehouse" in scenario 2.
[0067] S200: Based on the environmental information collected by the environmental sensor and the storage method, control the antenna power of the radio frequency reader, the frame rate and / or field of view of the first image collected by the first image sensor, so as to detect the quantity and type of the stored items.
[0068] For example, when neatly stacked items enter or leave the warehouse via trolleys, the speed of the trolleys is controlled by the warehouse manager or unloading workers, and the speed of the trolleys may vary. In this case, it is necessary to collect environmental information to confirm the speed of the trolleys. When the speed is high, the antenna power and the frame rate of the first image should be increased, and the field of view of the first image should be reduced to provide more useful information. When the speed is low, the antenna power and the frame rate of the first image can be appropriately reduced to reduce the number of images that the processor of the warehousing equipment needs to recognize and reduce computing power consumption.
[0069] For example, the higher the inbound and outbound speed, the higher the antenna power of the RF reader, the higher the frame rate of the first image, and the smaller the field of view of the first image (in order to reduce the amount of information in the image and save computing power).
[0070] For example, after identifying the items entering or leaving the warehouse in the first image, this application can maintain the field of view size, but when sending the first image into the model, the first image is cropped, and only the image of the area where the items entering or leaving the warehouse are located is sent into the model to reduce computing power overhead. This is also a way to reduce the field of view.
[0071] Alternatively, after identifying items entering or leaving the warehouse, the focal length when capturing the first image can be adjusted to narrow the field of view, allowing for the capture of more useful information for the items entering or leaving the warehouse.
[0072] For example, this application also adjusts the acquisition frame rate of the first image based on the number of items stacked in the first image that are identified as entering or leaving the warehouse. For instance, the number of items stacked in or leaving the warehouse is confirmed by an RFID reader; the higher the number, the higher the acquisition frame rate of the first image, in order to provide more information.
[0073] This application uses environmental sensors to detect environmental information, confirm the entry and exit methods, palletization quantities, and entry and exit speeds of items entering and leaving the warehouse, and then adjusts the antenna power of the RFID reader and the frame rate and / or field of view of the first image acquired by the auxiliary detection first image sensor based on the entry and exit methods, palletization quantities, and entry and exit speeds. This achieves dynamic adjustment of the RFID reader and the first image sensor to meet the detection needs of items entering and leaving the warehouse in different scenarios.
[0074] It should be noted that in scenario 2, there may be discrepancies between the reading results of the reader and the image recognition results. Therefore, in the embodiments of this application, when the types and quantities of items entering and leaving the warehouse obtained by image recognition are different from those obtained by the radio frequency reader, the warehouse manager can be prompted to manually calibrate.
[0075] In some embodiments, the environmental sensor includes a second image sensor, which is used to acquire multiple frames of second images of items entering and leaving the warehouse; the method further includes:
[0076] Based on the semantic segmentation module, the items entering and leaving the warehouse in the second image are identified; and...
[0077] Based on the reading results from the RFID reader, the number of palletized items entering and leaving the warehouse is confirmed.
[0078] In some embodiments, the method further includes:
[0079] In the multiple frames of the second image after the semantic segmentation module identifies the items entering and leaving the warehouse in the second image, the position information of the items entering and leaving the warehouse and the changes in the area ratio of the items entering and leaving the warehouse in the multiple frames of the second image are identified, and the movement speed and direction of the items entering and leaving the warehouse are confirmed.
[0080] In this embodiment, the faster the movement speed of the items entering and leaving the warehouse, the greater the power of the radio frequency reader and the higher the frame rate of the first image. Conversely, the slower the movement speed of the items entering and leaving the warehouse, the lower the power of the radio frequency reader and the lower the frame rate of the first image, until the minimum power and frame rate required for the reader to work and the processor to recognize are met.
[0081] Similarly, the more items are stacked in and out of the warehouse, that is, the more items are entered and exited in a single transaction, the higher the power of the RFID reader and the higher the frame rate of the first image. Conversely, the fewer items are entered and exited in a single transaction, the lower the power of the RFID reader and the lower the frame rate of the first image.
[0082] Reference Figure 2 In some embodiments, the environmental sensor includes: an audio sensor, which is used to collect audio information during the entry and exit of items; the method further includes:
[0083] S300: The audio information after the semantic segmentation module identifies the items entering and leaving the warehouse in the second image is recorded as valid audio information. The valid audio information and multiple frames of the second image at the same time are fed into a pre-trained neural network. The neural network model outputs the movement speed and / or movement direction of the items entering and leaving the warehouse.
[0084] In this application, the entry and exit items in the second image are identified by the semantic segmentation module as prompt information, thereby reducing the amount of invalid audio and reducing computing power consumption.
[0085] Furthermore, the effective audio information, along with multiple frames of second images taken at the same time, is fed into a pre-trained neural network. The neural network model outputs the motion speed and direction of the objects entering and leaving the warehouse, including:
[0086] S301. Convert the audio information into an audio image;
[0087] S302. Use the first processor to stitch the audio image and the second image together to obtain a stitched image, and then send the stitched image to the second processor;
[0088] S303, the second processor segments the audio image and the second image from the stitched image, and sends the audio image and the second image into their respective feature extraction networks to extract the audio feature information and the image feature information;
[0089] S304. The audio feature information and the image feature information are fused to obtain the fused feature;
[0090] S305. Output the movement speed and direction of objects entering and leaving the warehouse based on the fusion features.
[0091] In this embodiment, in order to improve the calculation accuracy of motion direction and motion speed, this application can also fuse audio feature information and video feature information and then feed them into a neural network model. The neural network model outputs the motion speed and motion direction of the objects entering and leaving the warehouse.
[0092] The neural network model used here includes models for identifying items entering and leaving the warehouse, and models for calculating the speed and direction of movement of these items. For example, the optical flow calculation model used in the aforementioned embodiment for calculating the speed and direction of movement of items entering and leaving the warehouse.
[0093] In this embodiment, audio information is converted into an audio image. The conversion method is not limited, as long as it converts time-series data into matrix data. For example, time-frequency transforms such as short-time Fourier transform and wavelet transform can be used to calculate the time-frequency information (two-dimensional matrix) corresponding to the audio information, and then the corresponding heatmap (contour map) can be calculated to obtain the corresponding audio image. Alternatively, Gram angle field transform and Markov transform can be used to calculate the two-dimensional matrix corresponding to the audio information, and then the corresponding heatmap (contour map) can be calculated to obtain the corresponding audio image.
[0094] It is important to note that when generating the audio image, the size of the video image should be taken into account. The width or length of the audio image should be the same as that of the video image, and then the two should be stitched together into a single image.
[0095] By feeding the image into a neural network model, the velocity and direction of motion of the target object can be calculated. Specifically, the neural network model can identify the optical flow and audio features of the stitched image, and then calculate the velocity and direction of motion. It should be noted that the audio features are not extracted manually, but rather annotated by annotators, as detailed below:
[0096] Given the speed and direction of movement of the items entering and leaving the warehouse corresponding to the audio, and labeling them with these speed and direction, a regression model can be trained using these audio labels. Subsequently, inputting video into the model will yield the speed and direction of movement.
[0097] The model can be executed in the following steps:
[0098] When the stitched image is obtained, it is segmented to obtain matching video and audio images. The area where the items entering and leaving the warehouse are located in the video image and the optical flow in that area are identified. The movement speed and direction of the items entering and leaving the warehouse are calculated.
[0099] The audio image is fed into the pre-trained regression model, which outputs the motion speed and direction.
[0100] The final speed and direction of motion can be obtained by weighted summation of the two speed and direction results (e.g., by dividing the mean by 2).
[0101] The second processor will segment the audio image and the second image from the stitched image. At this time, the second processor will segment the audio image and the second image from a single stitched image.
[0102] In this device, the first processor and the second processor are different processors, and their distributed layout helps to improve the device's computing power. For example, the first processor is located at the image and audio acquisition device, while the second processor is a separate image processing device. The two are connected via a data cable or wireless communication. Unlike separate transmission methods, in this embodiment, the first processor sends the stitched image to the second processor.
[0103] In this embodiment of the application, the processor used for image acquisition and the processor running the pre-trained neural network are not the same processor.
[0104] Reference Figure 3 In some embodiments, the dynamic modal warehouse inbound / outbound detection method further includes:
[0105] S400: Construct the mapping relationship between the movement speed range and pallet quantity of items entering and leaving the warehouse, and the power of the radio frequency antenna, the frame rate of the first image acquired by the first image sensor, and / or the field of view.
[0106] S500, Based on the movement speed range of the items entering and leaving the warehouse, the number of stacks, and the mapping relationship, determine the power of the radio frequency antenna, the frame rate of the first image acquired by the first image sensor, and / or the field of view; wherein, in the mapping relationship, the power of the radio frequency antenna is positively correlated with the movement speed range of the items entering and leaving the warehouse and the number of stacks, and the frame rate of the first image acquired by the first image sensor and / or the field of view is positively correlated with the movement speed range of the items entering and leaving the warehouse and the number of stacks.
[0107] To reduce the computational load of the model, this application does not use a single frame rate for every speed. This reduces the quality requirements of the image sensor, which only needs to operate at a fixed number of frame rates, thus avoiding flickering. For example, the operating frame rates of the image sensor include 30 frames per second, 60 frames per second, and 120 frames per second.
[0108] In some embodiments, in the mapping relationship, the movement speed of items entering and leaving the warehouse is divided into at least three speed ranges; wherein...
[0109] The power of the radio frequency antenna in each motion zone is different, and the frame rate of the first image acquisition is different, but the field of view is the same.
[0110] Alternatively, the power of the radio frequency antennas in each motion zone may be different, and the field of view for acquiring the first image may be different while the video frame rate may be the same.
[0111] In some embodiments, the radio frequency reader has multiple radio frequency antennas, all of which are directional antennas, referred to as the first radio frequency antenna and the second radio frequency antenna, wherein the first radio frequency antenna faces the inbound direction and the second radio frequency antenna faces the outbound direction; the method further includes:
[0112] When the movement direction of the item entering or leaving the warehouse is the same as the direction of entry, the first radio frequency antenna is activated to read the radio frequency tag of the item entering or leaving the warehouse.
[0113] When the movement direction of the items entering or leaving the warehouse is the outbound direction, the second radio frequency antenna is activated to read the radio frequency tag of the items entering or leaving the warehouse.
[0114] In some embodiments, the first image sensor is a gimbal camera, which is configured to adjust the gimbal camera toward the inbound / outbound object based on the movement direction of the inbound / outbound object.
[0115] Alternatively, the first image sensor includes at least two cameras, referred to as the first camera and the second camera; wherein the first camera faces the outbound direction and the second camera faces the inbound direction; the method further includes:
[0116] When the direction of movement of the item entering or leaving the warehouse is the direction of entry, the first camera is invoked to capture the first image of the item entering or leaving the warehouse.
[0117] When the movement direction of the items entering or leaving the warehouse is the outbound direction, the second camera is invoked to capture the first image of the items entering or leaving the warehouse.
[0118] Secondly, this application provides a storage device, characterized in that it includes:
[0119] An environmental sensor is used to collect environmental information, which represents the direction and speed of movement of items entering and leaving the warehouse.
[0120] Radio frequency (RFID) readers are used to read the RFID tags on items entering and leaving the warehouse.
[0121] The first image sensor, and the first image processor, are used to acquire images of items entering and leaving the warehouse.
[0122] One or more processors are connected to the environmental sensor, the radio frequency reader, and the image sensor. The one or more processors are configured to execute a warehouse entry and exit detection program to implement the dynamic modal warehouse entry and exit detection method described above.
[0123] It should be noted that the descriptions of each embodiment in the above embodiments have different focuses. For parts that are not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.
[0124] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0125] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, as well as combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0126] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0127] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0128] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the invention.
[0129] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.
Claims
1. A dynamic modal warehouse in-out detection method, characterized in that, Applicable to warehousing equipment, the warehousing equipment includes: an environmental sensor for collecting environmental information, the environmental information representing the direction and speed of movement of items entering and leaving the warehouse; an RFID reader for reading the RFID tags of items entering and leaving the warehouse; a first image sensor for acquiring a first image of the items entering and leaving the warehouse; the detection method includes: Based on the environmental information collected by the environmental sensors, the entry and exit methods of items are detected. Based on the environmental information collected by the environmental sensors and the storage method, the antenna power of the RF reader, the frame rate and / or field of view of the first image acquired by the first image sensor are controlled to detect the quantity and type of the stored items; including: Construct a mapping relationship between the movement speed range and stacking quantity of items entering and leaving the warehouse, and the power of the radio frequency antenna, the frame rate of the first image acquired by the first image sensor, and the field of view. Based on the movement speed range of items entering and leaving the warehouse, the number of stacks, and the mapping relationship, the power of the radio frequency antenna, the frame rate of the first image acquired by the first image sensor, and the field of view size are determined. In the mapping relationship, the power of the radio frequency antenna is positively correlated with the speed range of the items entering and leaving the warehouse and the number of stacks, and the frame rate of the first image acquired by the first image sensor is positively correlated with the speed range of the items entering and leaving the warehouse and the number of stacks; the faster the items enter and leave the warehouse and the more stacks, the smaller the field of view of the first image acquired by the first image sensor. In the mapping relationship, the movement speed of items entering and leaving the warehouse is divided into at least three movement speed intervals; wherein, the power of the radio frequency antenna in each movement interval is different, and the frame rate of the first image is different while the field of view is the same; or, the power of the radio frequency antenna in each movement interval is different, and the field of view of the first image is different while the video frame rate is the same. The radio frequency reader has multiple radio frequency antennas, all of which are directional antennas, referred to as the first radio frequency antenna and the second radio frequency antenna. The first radio frequency antenna faces the inbound direction, and the second radio frequency antenna faces the outbound direction. The method further includes: when the movement direction of the item is inbound, the first radio frequency antenna is used to read the radio frequency tag of the item; when the movement direction of the item is outbound, the second radio frequency antenna is used to read the radio frequency tag of the item. Based on the environmental information collected by the environmental sensor, the entry and exit method, the movement speed range, the number of pallets, and the mapping relationship, the antenna power of the radio frequency reader, the frame rate and field of view of the first image collected by the first image sensor are controlled to detect the quantity and type of items entering and leaving the warehouse. The environmental sensor includes a second image sensor and an audio sensor. The second image sensor is used to acquire multiple frames of second images of items entering and leaving the warehouse. The audio sensor is used to acquire audio information of the items entering and leaving the warehouse. The detection of the entry and exit method of the items based on the acquired environmental information includes: Based on the semantic segmentation module, the items entering and leaving the warehouse in the second image are identified; In the multiple frames of the second image after the semantic segmentation module identifies the items entering and leaving the warehouse in the second image, the position information of the items entering and leaving the warehouse and the changes in the area ratio of the items entering and leaving the warehouse in the multiple frames of the second image are identified, and the movement speed and direction of the items entering and leaving the warehouse are confirmed; and: Based on the reading results of the radio frequency reader, the number of palletized items entering and leaving the warehouse is confirmed; Furthermore, based on the entry and exit items in the second image identified by the semantic segmentation module as prompt information, the audio information after the semantic segmentation module identifies the entry and exit items in the second image is recorded as valid audio information, so as to reduce the amount of invalid audio information to be processed. Convert the audio information into an audio image; After stitching the audio image with the second image, the audio image and the second image are segmented and sent to their respective feature extraction networks to extract audio feature information and image feature information. The audio feature information and image feature information are fused to obtain fused features. Based on the fused features, the movement speed and / or movement direction of the objects entering and leaving the warehouse are output.
2. The dynamic modal warehouse entry and exit detection method as described in claim 1, characterized in that, The first image sensor is a gimbal camera, which is configured to adjust its orientation toward the inbound / outbound object based on the movement direction of the object. Alternatively, the first image sensor includes at least two cameras, referred to as the first camera and the second camera; wherein the first camera faces the outbound direction and the second camera faces the inbound direction; the method further includes: When the direction of movement of the item entering or leaving the warehouse is the direction of entry, the first camera is invoked to capture the first image of the item entering or leaving the warehouse. When the movement direction of the items entering or leaving the warehouse is the outbound direction, the second camera is invoked to capture the first image of the items entering or leaving the warehouse.
3. A warehousing apparatus, characterised in that, include: An environmental sensor is used to collect environmental information, which represents the direction and speed of movement of items entering and leaving the warehouse. Radio frequency (RFID) readers are used to read the RFID tags on items entering and leaving the warehouse. The first image sensor, and the first image processor, are used to acquire images of items entering and leaving the warehouse. One or more processors are connected to the environmental sensor, the radio frequency reader, and the image sensor, and the one or more processors are configured to execute a warehouse entry and exit detection program to implement the dynamic modal warehouse entry and exit detection method as described in any one of claims 1 or 2.