A method and device for processing access deviation events of a multi-modal intelligent storage cabinet

By capturing video footage of stored items and extracting behavioral trajectories from smart lockers, and combining this with opening and closing records and cloud-based identity verification, the problem of time-consuming item tracing in smart lockers has been solved, enabling fast and accurate item location and secure locker opening.

CN121236849BActive Publication Date: 2026-02-03SHENZHEN MOTERN TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511798287.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-02
Publication Date
2026-02-03
Estimated Expiration
2045-12-02

AI Technical Summary

Technical Problem

During peak hours when there are many people, smart locker systems are prone to errors such as overlapping locker openings, misplacement, and mis-taking when processing multiple sets of verification and locker instructions. This leads to longer tracking times for items and increases the risk of loss.

Method used

By responding to user tracing requests, determining the initial time point, capturing video of stored items and extracting the trajectory of human behavior, and combining opening and closing records to estimate the location of items, the target cabinet is opened after cloud-based identity verification.

Benefits of technology

Quickly locate the cabinet where the item is stored, shorten the tracing time, reduce the risk of loss, and ensure the security of storage and retrieval.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121236849B_ABST
    Figure CN121236849B_ABST
Patent Text Reader

Abstract

The application relates to the field of intelligent storage cabinets, in particular to a method and device for processing intelligent storage cabinet access deviation events in a multi-modal manner, the method comprising the following steps: in response to a first article tracing request, determining an initial time node, presetting a storage time period, intercepting a target storage video in a video storage library, extracting a target person behavior track, and determining a target cabinet; determining an opening and closing record, and estimating whether the target article is in the target cabinet; if yes, initiating identity recognition verification on the first user through a cloud server; when the identity verification passes, the target cabinet is opened, the target storage video is intercepted through an edge node, and the target person behavior track is extracted to quickly locate the target cabinet storing the target article, without waiting for an operation and maintenance personnel to arrive, so that the tracing time is greatly shortened; the target cabinet is opened after the cloud server identity recognition verification passes, which guarantees the storage and access safety and efficiently solves the article misplacement tracing and use problems.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of smart lockers, and in particular to a method and apparatus for multimodal processing of smart locker access deviation events. Background Technology

[0002] With the development of modern society, people's demand for storage is increasing day by day, and traditional storage methods can no longer meet the requirements of efficiency, safety and convenience. Intelligent lockers have emerged to meet this need, integrating a variety of advanced technologies to bring a brand-new solution to the storage of items.

[0003] Smart lockers are storage devices with self-service operation and automatic management functions. They utilize intelligent technology, employing methods such as passwords, fingerprint recognition, and facial recognition for identity verification, enabling secure and convenient storage and retrieval of items. Simultaneously, they possess intelligent management functions, allowing for intelligent categorization, retrieval, and reminders for stored items, significantly improving storage efficiency and user experience.

[0004] During peak hours (8 AM to 6 PM), when multiple users access their devices simultaneously, the system processes multiple verifications and locker commands concurrently, easily leading to issues like overlapping locker openings and misplacement / retrieval. For example, users might mistakenly store items in a different locker that opened at the same time as their assigned locker. Subsequent retrieval of the items typically requires notifying maintenance personnel, who then review historical surveillance footage to determine the actual locker location. However, waiting for maintenance personnel to arrive and reviewing the video footage consumes significant time, extending the retrieval process and increasing the risk of lost items. Summary of the Invention

[0005] In view of the aforementioned problems, this application is proposed to provide a method and apparatus for multimodal processing of smart locker access deviation events to overcome or at least partially solve the aforementioned problems, comprising:

[0006] A method for multimodal processing of smart locker access deviation events, the method comprising:

[0007] In response to the first item traceability request initiated by the first user, determine the initial time node of the first user's last opening of the cabinet, and preset the storage period based on the initial time node;

[0008] Based on the storage time period, the target storage video is extracted from the video storage library. The target person's behavior trajectory in the target storage video is extracted through edge node calculation. Based on the target person's behavior trajectory, the target cabinet where the target item is stored is determined.

[0009] The system determines the opening and closing records of the target cabinet from the initial time point to the current time point when the first user opened the cabinet. Based on the opening and closing records, it estimates whether the target item is in the target cabinet at the current time point. If so, it initiates identity verification for the first user through the cloud server. When the cloud server returns a message indicating that the identity verification is successful, the target cabinet is opened.

[0010] Preferably, the step of extracting the target stored video from the video repository based on the stored time period, and extracting the target person's behavioral trajectory in the target stored video through edge node calculation, includes:

[0011] Based on the initial time node, the target access image is determined in the target access video, and at least one human figure outline is extracted from the target access image by calculating edge nodes;

[0012] Generate a character outline selection interface based on at least one character outline and display it to the first user;

[0013] In response to the target person outline selection result selected by the first user in the person outline selection interface, the target person's behavior trajectory in the target video is extracted based on the target person outline and by calculating edge nodes.

[0014] Preferably, determining whether the target item is in the target compartment at the current time point based on the opening and closing record includes:

[0015] The number of times the target cabinet has been opened and closed between the initial time point and the current time point is determined based on the opening and closing records.

[0016] When the number of opening and closing operations is one, the target cabinet is marked as not restarted.

[0017] When the number of opening and closing operations is twice, the target cabinet is marked as being in a state of only restarting once.

[0018] If the number of opening and closing operations is greater than two, the target cabinet is marked as having been restarted multiple times.

[0019] If the target cabinet is in a non-restarted state or has only been restarted once, it is determined that the target item is still in the target cabinet at the current time point.

[0020] Preferably, the step of initiating identity verification for the first user through a cloud server includes:

[0021] The identity capture camera and the environment capture camera of the smart locker are activated simultaneously. The identity capture camera is used to capture real-time images and video data of the upper body of the first user, and the environment capture camera is used to capture real-time environmental video including the first user.

[0022] Based on the upper body image and video data of the first user, facial detail features and clothing detail features are extracted, and the facial detail features are transmitted to the cloud server, while the clothing detail features are transmitted to the edge node;

[0023] The environmental video is transmitted to the edge node, and the human body area is located by calculating the edge node. The clothing area associated with the first user is locked in the human body area based on the clothing details. The edge node is used to calculate and determine the current behavior trajectory of the first user based on the clothing area.

[0024] Preferably, the step of extracting facial detail features and clothing detail features based on the upper body image and video data of the first user, transmitting the facial detail features to the cloud server, and transmitting the clothing detail features to the edge node includes:

[0025] Based on the upper body image and video data of the first user, facial feature vectors are extracted using a CNN model and transmitted to a cloud server. The cloud server is used to compare the facial feature vectors with pre-stored facial features.

[0026] Based on the upper body image and video data of the first user, the color feature values ​​of the clothing are extracted using the HSV color histogram algorithm, and the clothing color feature values ​​are transmitted to the edge nodes.

[0027] Preferably, the edge node is used to calculate and determine the real-time distance between the clothing area and the target cabinet area based on the current behavior trajectory of the first user. When the real-time distance is less than or equal to a preset interval threshold, an arrival confirmation message indicating that the first user is located at the corresponding position of the target cabinet is generated. The step of opening the target cabinet upon receiving authentication verification information returned by the cloud server includes:

[0028] When the cloud server returns a successful authentication message and the edge node simultaneously returns a location confirmation message, the target cabinet is opened.

[0029] Preferably, the method further includes:

[0030] Based on the opening and closing records, determine whether the target item is in the target compartment at the current time.

[0031] If the target cabinet is in a state of multiple restarts, or responds to the first user's request for no items found after the target cabinet is opened, then search for the client information of the second user associated with the target cabinet when it is opened and closed for the second time.

[0032] Determine the client information of the first user, generate communication connection channel information based on the client information of the first user and the identity information of the second user, and send it to the client of the first user.

[0033] A device for multimodal processing of access deviation events in smart lockers, the device comprising:

[0034] The response module is used to respond to the first item traceability request initiated by the first user, determine the initial time node of the first user's last opening of the cabinet, and preset the storage period based on the initial time node;

[0035] The calculation module is used to extract the target storage video from the video storage library based on the storage time period, extract the target person's behavior trajectory in the target storage video through edge node calculation, and determine the target cabinet where the target item is stored based on the target person's behavior trajectory.

[0036] The execution module is used to determine the opening and closing records of the target cabinet from the initial time point to the current time point when the first user opened the cabinet. Based on the opening and closing records, it estimates whether the target item is in the target cabinet at the current time point. If so, it initiates identity verification for the first user through the cloud server. When it receives the identity verification information returned by the cloud server, it opens the target cabinet.

[0037] An apparatus includes a processor, a memory, and a computer program stored in the memory and capable of running on the processor, wherein the computer program, when executed by the processor, implements the steps of a method for multimodal processing of smart locker access deviation events as described above.

[0038] A computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of a method for multimodal processing of smart locker access deviation events as described above.

[0039] This application has the following advantages:

[0040] In the embodiments of this application, in contrast to the time-consuming problem of item tracing in existing smart lockers, this application provides a solution using multimodal processing for such events. Specifically, in response to a first item tracing request initiated by a first user, the initial time node of the first user's last locker opening is determined, and a storage period is preset based on the initial time node; a target storage video is captured from a video storage repository based on the storage period, and the target person's behavior trajectory in the target storage video is extracted through edge node calculation; the target locker compartment for storing the target item is determined based on the target person's behavior trajectory; the opening and closing records of the target locker compartment from the initial time node to the current time node of the first user's current locker opening are determined, and the target item is estimated to be in the target locker compartment at the current time node based on the opening and closing records; if so, the first user's identity verification is initiated through a cloud server; when the identity verification information returned by the cloud server is received, the target locker compartment is opened. By responding to the first user's traceability request and determining the initial time node and preset storage time period, the edge node captures the target storage video and extracts the target person's behavior trajectory to quickly locate the target cabinet where the target item is stored, without waiting for maintenance personnel to arrive, significantly shortening the traceability time. At the same time, by querying the opening and closing records of the target cabinet from the initial time node to the current time node, it is possible to estimate in advance whether the target item is still in the cabinet, reducing the risk of item loss. Subsequently, after identity verification by the cloud server, the target cabinet is opened, which not only ensures the security of storage and retrieval but also efficiently solves the problem of tracing and retrieving misplaced items. Attached Figure Description

[0041] To more clearly illustrate the technical solution of this application, the drawings used in the description of this application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0042] Figure 1 This is a flowchart illustrating the steps of a method for multimodal processing of smart locker access deviation events according to an embodiment of this application;

[0043] Figure 2 This is a structural block diagram of the steps for multimodal processing of smart locker access deviation events according to an embodiment of this application;

[0044] Figure 3 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation

[0045] To make the objectives, features, and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.

[0046] Reference Figure 1 The diagram shows a flowchart of the steps of a method for multimodal processing of smart locker access deviation events according to an embodiment of this application.

[0047] The method includes:

[0048] S110, in response to the first item traceability request initiated by the first user, determine the initial time node of the first user's last opening of the cabinet, and preset the storage period according to the initial time node;

[0049] S120: Based on the storage time period, extract the target storage video from the video storage library, extract the target person's behavior trajectory in the target storage video through edge node calculation, and determine the target cabinet where the target item is stored based on the target person's behavior trajectory.

[0050] S130, determine the opening and closing records of the target cabinet from the initial time point to the current time point when the first user opened the cabinet, and estimate whether the target item is in the target cabinet at the current time point based on the opening and closing records; if so, initiate identity verification for the first user through the cloud server; when the cloud server returns a message indicating that the identity verification is successful, open the target cabinet.

[0051] In the embodiments of this application, in contrast to the time-consuming problem of item tracing in existing smart lockers, this application provides a solution using multimodal processing for such events. Specifically, in response to a first item tracing request initiated by a first user, the initial time node of the first user's last locker opening is determined, and a storage period is preset based on the initial time node; a target storage video is captured from a video storage repository based on the storage period, and the target person's behavior trajectory in the target storage video is extracted through edge node calculation; the target locker compartment for storing the target item is determined based on the target person's behavior trajectory; the opening and closing records of the target locker compartment from the initial time node to the current time node of the first user's current locker opening are determined, and the target item is estimated to be in the target locker compartment at the current time node based on the opening and closing records; if so, the first user's identity verification is initiated through a cloud server; when the identity verification information returned by the cloud server is received, the target locker compartment is opened. By responding to the first user's traceability request and determining the initial time node and preset storage time period, the edge node captures the target storage video and extracts the target person's behavior trajectory to quickly locate the target cabinet where the target item is stored, without waiting for maintenance personnel to arrive, significantly shortening the traceability time. At the same time, by querying the opening and closing records of the target cabinet from the initial time node to the current time node, it is possible to estimate in advance whether the target item is still in the cabinet, reducing the risk of item loss. Subsequently, after identity verification by the cloud server, the target cabinet is opened, which not only ensures the security of storage and retrieval but also efficiently solves the problem of tracing and retrieving misplaced items.

[0052] The method for multimodal processing of smart locker access deviation events in this exemplary embodiment will be further described below.

[0053] As described in step S110, in response to the first item traceability request initiated by the first user, the initial time node of the first user's last opening of the cabinet is determined, and the storage period is preset according to the initial time node.

[0054] In this embodiment, physical detection of items is eliminated. For example, ultra-thin, ultra-light, and transparent items are easy to evade detection. No sensors for detecting items are installed in the compartments of the smart locker in this embodiment. This saves the cost and installation of various sensors such as infrared and pressure sensors, and reduces the data load on the server.

[0055] It should be noted that when it is necessary to trace items, the smart locker does not need to know the specifications of the items. It directly uses the time node as the starting point and presets the storage time period based on the initial time node when the first user puts the items in. This reduces the amount of calculation required to capture the corresponding video data of the user and extract the behavior trajectory of the person in the subsequent process, thereby shortening the tracing time.

[0056] As an example, it usually takes a maximum of about one minute for a user to open the cabinet, store the item, and then close the cabinet. That is, if the initial time point is t0, the preset storage time period is (t0, t0+1).

[0057] Specifically, the industrial cameras configured in smart lockers generally employ an event-triggered mechanism for image / video acquisition. In this embodiment, the event-triggered mechanism for video acquisition, for example, presets the camera's operating time as x (generally greater than the average storage time for users). When user A, the first to open the locker, issues an opening command at time node a, the camera is activated to capture video of the locker environment, and the camera's closing time node is set to a+x. At a+x, the video image of the time period (a, a+x) is saved. If, during the time period (a, a+x) (before the video image is saved), a second user B issues an opening command at time node b, the camera's closing time node is updated to b+x. At a+x, no image is saved, until b+x when no new opening command is received, at which point the video image of the time period (a, b+x) is saved. Furthermore, if the initial time node t0 of the first user is within the time period (a, b+x) and before the second user B, then the target storage video of the first user is located within a segment of the video image of the time period (a, b+x).

[0058] As described in step S120, the target storage video is extracted from the video storage library according to the storage time period, the target person's behavior trajectory in the target storage video is extracted by edge node calculation, and the target cabinet for storing the target item is determined based on the target person's behavior trajectory.

[0059] In one embodiment of the present invention, the specific process of step S120, "to extract the target stored video from the video storage repository based on the stored time period and to extract the target person's behavior trajectory in the target stored video by calculating edge nodes," can be further explained in conjunction with the following description.

[0060] As described in the following steps

[0061] Based on the initial time node, the target access image is determined in the target access video, and at least one human figure outline is extracted from the target access image by calculating edge nodes;

[0062] Generate a character outline selection interface based on at least one character outline and display it to the first user;

[0063] In response to the target person outline selection result selected by the first user in the person outline selection interface, the target person's behavior trajectory in the target video is extracted based on the target person outline and by calculating edge nodes.

[0064] It should be noted that, in conjunction with the previous embodiment, the storage time period (t0, t0+1) for the first user is preset based on the initial time node t0. The video at (t0, t0+1) can then be directly extracted from the video image within the (a, b+x) time period. In this embodiment, a lightweight embedded database, such as SQLite, can also be deployed to construct a mapping relationship between each stored video image and associated user and metadata such as the opening time, facilitating more convenient indexing conditions during queries. The specific construction of the database tables can be determined by those skilled in the art based on the required data format, attributes, and fields, combined with the actual operation method, and will not be elaborated upon further.

[0065] In this embodiment, edge nodes calculate and identify the behavioral trajectories of people in a smart locker scenario, especially when multiple people are present at the same time, where automatic identification is prone to errors. By generating a person outline selection interface, users can select their own person outline, eliminating interference from irrelevant people. Edge nodes do not need to perform trajectory calculations on other irrelevant outlines, but only process the single outline confirmed by the user, effectively reducing unnecessary calculation overhead, further shortening the time for trajectory extraction, avoiding tracing deviations, and further shortening the tracing time.

[0066] In this scenario, even if the first user selects a different profile than themselves, and the server cannot determine whether the profile corresponds to the user, the subsequent steps involving user authentication before opening the locker still leave evidence that the user has taken an item that does not belong to them. This is because the user has left evidence for judicial intervention. This embodiment only addresses situations where normal users can reasonably trace their belongings; solutions for special individual cases as described above are not covered in this embodiment.

[0067] On the other hand, the general process of extracting the target person's behavior trajectory through edge computing nodes is as follows: video input, person detection, contour extraction, edge node calculation, trajectory tracking, behavior analysis, trajectory output, and target cabinet is determined based on the trajectory output results. In each process step, those skilled in the art can select the corresponding model tools from the currently mature open source projects according to their needs and the elements mentioned in the embodiments of this application. The specific parameter adjustments are not elaborated here.

[0068] As described in step S120, the opening and closing records of the target cabinet from the initial time node to the current time node when the first user opened the cabinet are determined. Based on the opening and closing records, it is estimated whether the target item is in the target cabinet at the current time node. If so, the first user is identified and verified through the cloud server. When the cloud server returns a message indicating that the identity verification is successful, the target cabinet is opened.

[0069] In one embodiment of the present invention, the specific process of "determining whether the target item is in the target cabinet at the current time point based on the opening and closing record" in step S130 can be further explained in conjunction with the following description.

[0070] As described in the following steps, determine the number of times the target cabinet has been opened and closed between the initial time node and the current time node based on the opening and closing record;

[0071] When the number of opening and closing operations is one, the target cabinet is marked as not restarted.

[0072] When the number of opening and closing operations is twice, the target cabinet is marked as being in a state of only restarting once.

[0073] If the number of opening and closing operations is greater than two, the target cabinet is marked as having been restarted multiple times.

[0074] If the target cabinet is in a non-restarted state or has only been restarted once, it is determined that the target item is still in the target cabinet at the current time point.

[0075] It should be noted that, under normal logic, for example, if user A uses user B's cabinet 1 to store item W, and as of the current time, cabinet 1 has not been opened since user B opened and closed it once, then the target cabinet is in a non-restarted state; if user B opens cabinet 1 and finds that the item W inside is not their own, they will generally report that the item W in cabinet 1 is not their own and trace their own items.

[0076] At this point, follow the steps below:

[0077] When the number of opening and closing operations has been completed twice, and a second item tracing request is received from a second user associated with the target cabinet, a command to close the target cabinet is generated and the second user is notified.

[0078] In response to the second user's closing operation on the target cabinet, the target cabinet is marked as a temporary storage cabinet.

[0079] It should be noted that the smart lockers have three categories of compartments: occupied compartments, unoccupied compartments, and temporary storage compartments. Occupied and unoccupied compartments are opened and closed in rotation to be available to users, while temporary storage compartments are not available to users and are only accessible to the person who claims the items inside.

[0080] In another embodiment, if user B opens cabinet 1 and finds that the item W inside is not theirs, takes the item W directly, closes cabinet 1, and does not report it, cabinet 1 may be opened and closed multiple times subsequently. In this case, the following steps shall apply:

[0081] Based on the opening and closing records, determine whether the target item is in the target compartment at the current time.

[0082] If the target cabinet is in a state of multiple restarts, or responds to the first user's request for no items found after the target cabinet is opened, then search for the client information of the second user associated with the target cabinet when it is opened and closed for the second time.

[0083] Determine the client information of the first user, generate communication connection channel information based on the client information of the first user and the identity information of the second user, and send it to the client of the first user.

[0084] It should be noted that when the first user's items are taken by the second user, the first user can contact the second user through the communication connection channel information. This communication connection channel information refers to the communication channel between the two parties, including but not limited to contact information push notifications and instant messaging access. Further, in serious cases, assistance from a third-party judicial authority may be sought; however, this is unrelated to the technical solution of this application and will not be elaborated upon.

[0085] This embodiment does not consider the handling method for malicious incidents where individual users open the cabinet without storing anything, waiting for other users to store the wrong items and then stealing other people's belongings.

[0086] In one embodiment of the present invention, the specific process of "initiating identity verification for the first user through the cloud server" in step S130 can be further described in conjunction with the following description.

[0087] As described in the following steps, the identity acquisition camera and the environment acquisition camera of the smart locker are started synchronously. The identity acquisition camera is used to collect the upper body image and video data of the first user in real time, and the environment acquisition camera is used to collect the environmental video containing the first user in real time.

[0088] Based on the upper body image and video data of the first user, facial detail features and clothing detail features are extracted, and the facial detail features are transmitted to the cloud server, while the clothing detail features are transmitted to the edge node;

[0089] The environmental video is transmitted to the edge node, and the human body area is located by calculating the edge node. The clothing area associated with the first user is locked in the human body area based on the clothing details. The edge node is used to calculate and determine the current behavior trajectory of the first user based on the clothing area.

[0090] It should be noted that the cloud server and edge nodes can process data in parallel without waiting for each other's data. By using multimodal processing of the first user's facial and clothing details, the system can quickly and accurately identify the first user, solving the problem of insufficient accuracy that may exist in single identity verification. It also takes into account the security of identity verification and the real-time nature of behavioral trajectory positioning, thus preventing others from claiming the property.

[0091] Specifically, the identity capture camera, which collects images and video data of the first user's upper body, requires the user's passive cooperation. The identity capture camera can be placed on the side of the smart locker in the area where the user interacts, and the user walks to this area to cooperate with data collection. Due to the installation position and angle of the environmental capture camera, the user needs to turn around during data collection in order to capture detailed features of the clothing from the front, back, and sides, which is convenient for subsequent combination with environmental video to determine the clothing area.

[0092] Furthermore, based on the upper body image and video data of the first user, facial feature vectors are extracted using a CNN model, and the facial feature vectors are transmitted to a cloud server. The cloud server is used to compare the facial feature vectors with pre-stored facial features.

[0093] Based on the upper body image and video data of the first user, the color feature values ​​of the clothing are extracted using the HSV color histogram algorithm, and the clothing color feature values ​​are transmitted to the edge nodes.

[0094] It's worth noting that CNN (Convolutional Neural Network) models excel at capturing subtle yet crucial facial features (such as facial contours and texture details). The facial feature vectors they extract can more accurately reflect a user's unique identity information. HSV (Histogram of Colors) can convert clothing colors from the RGB space to a "hue-saturation-brightness" space that better matches human visual perception, and it is more resistant to changes in ambient light.

[0095] In this embodiment, the edge node is further configured to calculate and determine the real-time distance between the clothing area and the target cabinet area based on the current behavioral trajectory of the first user. When the real-time distance is less than or equal to a preset interval threshold, an arrival confirmation message indicating that the first user is located at the corresponding position of the target cabinet is generated. The step of opening the target cabinet upon receiving authentication verification information returned by the cloud server includes:

[0096] When the cloud server returns a successful authentication message and the edge node simultaneously returns a location confirmation message, the target cabinet is opened.

[0097] It's important to note that relying solely on cloud-based identity verification to open a locker could lead to situations where someone impersonates the primary user (e.g., by stealing facial information) or the user's identity is verified but they are not actually at the locker. This could result in the target locker being opened incorrectly or items being fraudulently claimed. Edge nodes use real-time distance calculations to confirm the user's location (position information), creating a dual condition with cloud-based identity verification—the locker is only opened when both conditions are met. This locker access is secured based on both "identity legitimacy" and "physical location legitimacy." Furthermore, confirming the user's proximity to the target locker by measuring "real-time distance less than or equal to a preset interval" ensures the user can immediately retrieve their item upon opening, preventing invalid operations due to "locker opening and user arrival not being synchronized," and improving the accuracy of the retrieval process.

[0098] In another embodiment, before opening the target cabinet, the current cabinet opened by the first user is detected to have been opened and closed twice, and at the same time, a first item traceability request is initiated, generating a command to close the current cabinet and prompting the first user; in response to the result of the first user's closing operation of the current cabinet, the current cabinet is also marked as a temporary storage cabinet.

[0099] It should be noted that while generating the command to close the relevant cabinet, the smart locker's identity capture camera and environmental capture camera are also controlled to start collecting and storing video for subsequent data processing.

[0100] This application embodiment uses a cloud database and edge nodes to multimodally process smart locker access deviation events, without the need for other users or maintenance personnel to participate or cooperate. The cloud database and edge node data are used collaboratively and processed in parallel, making the item traceability process safe and fast.

[0101] As the apparatus embodiment is basically similar to the method embodiment, it is described in a relatively simple manner. For relevant details, please refer to the description of the method embodiment.

[0102] Reference Figure 2 The diagram shows a structural block diagram of a device for multimodal processing of smart locker access deviation events according to an embodiment of this application.

[0103] Specifically, it includes:

[0104] A device for multimodal processing of access deviation events in smart lockers, the device comprising:

[0105] Response module 110 is used to respond to the first item traceability request initiated by the first user, determine the initial time node of the first user's last opening of the cabinet, and preset the storage period according to the initial time node;

[0106] The calculation module 120 is used to extract the target storage video from the video storage library according to the storage time period, extract the target person's behavior trajectory in the target storage video through edge node calculation, and determine the target cabinet where the target item is stored based on the target person's behavior trajectory.

[0107] The execution module 130 is used to determine the opening and closing records of the target cabinet from the initial time node to the current time node of the first user's opening of the cabinet, and estimate whether the target item is in the target cabinet at the current time node based on the opening and closing records; if so, the first user is identified and verified through the cloud server, and the target cabinet is opened when the cloud server returns a message indicating that the identity verification is successful.

[0108] Reference Figure 3 The computer device illustrating a method for multimodal processing of access deviation events in smart lockers according to the present invention may specifically include the following:

[0109] The aforementioned computer device 12 is manifested in the form of a general-purpose computing device. The components of the computer device 12 may include, but are not limited to: one or more processors or processing units 16, memory 28, and a bus 18 connecting different system components (including memory 28 and processing unit 16).

[0110] Bus 18 refers to one or more of several types of bus 18 architectures, including memory bus 18 or memory controller, peripheral bus 18, graphics acceleration port, processor, or local bus 18 using any of the various bus 18 architectures. For example, these architectures include, but are not limited to, Industry Standard Architecture (ISA) bus 18, Micro Channel Architecture (MAC) bus 18, Enhanced ISA bus 18, Audio / Video Electronics Standards Association (VESA) local bus 18, and Peripheral Component Interconnect (PCI) bus 18.

[0111] Computer device 12 typically includes a variety of computer system readable media. These media can be any available media that can be accessed by computer device 12, including volatile and non-volatile media, removable and non-removable media.

[0112] Memory 28 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache 32. Computer device 12 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, storage system 34 may be used to read and write non-removable, non-volatile magnetic media (commonly referred to as a "hard disk drive"). Figure 3Not shown, a disk drive for reading and writing to a removable non-volatile disk (such as a "floppy disk") and an optical disk drive for reading and writing to a removable non-volatile optical disk (such as a CD-ROM, DVD-ROM, or other optical media) may be provided. In these cases, each drive may be connected to bus 18 via one or more data media interfaces. The memory may include at least one program product having a set (e.g., at least one) of program modules 42 configured to perform the functions of the embodiments of the present invention.

[0113] A program / utility 40 having a set (at least one) of program modules 42 may be stored, for example, in memory. Such program modules 42 include—but are not limited to—an operating system, one or more application programs, other program modules 42, and program data. Each or some combination of these examples may include an implementation of a network environment. Program modules 42 typically perform the functions and / or methods described in the embodiments of the present invention.

[0114] Computer device 12 can also communicate with one or more external devices 14 (e.g., keyboard, pointing device, display 24, camera, etc.), and with one or more devices that enable a user to interact with the computer device 12, and / or with any device that enables the computer device 12 to communicate with one or more other computing devices (e.g., network card, modem, etc.). This communication can be performed via input / output (I / O) interface 22. Furthermore, computer device 12 can also communicate with one or more networks (e.g., local area network (LAN)), wide area network (WAN), and / or public networks (e.g., the Internet) via network adapter 20. As shown, network adapter 20 communicates with other modules of computer device 12 via bus 18. It should be understood that, although... Figure 3 Not shown, it can be combined with computer device 12 to use other hardware and / or software modules, including but not limited to: microcode, device drivers, redundant processing unit 16, external disk drive array, RAID system, tape drive and data backup storage system 34, etc.

[0115] The processing unit 16 executes various functional applications and data processing by running programs stored in memory 28, such as implementing the method for multimodal processing of smart locker access deviation events provided in the embodiments of the present invention.

[0116] That is, when the processing unit 16 executes the above procedure, it performs the following: in response to the first item tracing request initiated by the first user, it determines the initial time node of the first user's last opening of the cabinet, and presets the storage time period based on the initial time node; it extracts the target storage video from the video storage library based on the storage time period, extracts the target person's behavior trajectory in the target storage video through edge node calculation, and determines the target cabinet where the target item is stored based on the target person's behavior trajectory; it determines the opening and closing records of the target cabinet from the initial time node to the current time node of the first user's current opening of the cabinet, and estimates whether the target item is in the target cabinet at the current time node based on the opening and closing records; if so, it initiates identity verification for the first user through the cloud server; when it receives the identity verification information returned by the cloud server, it opens the target cabinet.

[0117] In this embodiment of the invention, the present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method for multimodal processing of smart locker access deviation events as provided in all embodiments of this application:

[0118] That is, when the program is executed by the processor, it implements the following: In response to the first item tracing request initiated by the first user, it determines the initial time node of the first user's last opening of the cabinet, and presets the storage time period based on the initial time node; Based on the storage time period, it extracts the target storage video from the video storage library, extracts the target person's behavior trajectory in the target storage video through edge node calculation, and determines the target cabinet where the target item is stored based on the target person's behavior trajectory; It determines the opening and closing records of the target cabinet from the initial time node to the current time node of the first user's current opening of the cabinet, and estimates whether the target item is in the target cabinet at the current time node based on the opening and closing records; If so, it initiates identity verification for the first user through the cloud server; When it receives the identity verification information returned by the cloud server, it opens the target cabinet.

[0119] Any combination of one or more computer-readable media may be used. A computer-readable medium may be a computer-to-signal medium or a computer-readable storage medium. A computer-readable storage medium may be, for example—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium may be any tangible medium that contains or stores a program that may be used by or in connection with an instruction execution system, apparatus, or device.

[0120] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including—but not limited to—electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, capable of transmitting, propagating, or transmitting programs for use by or in connection with an instruction execution system, apparatus, or device.

[0121] Computer program code for performing the operations of this invention can be written in one or more programming languages ​​or a combination thereof. These programming languages ​​include object-oriented programming languages—such as Java, Smalltalk, and C++—and conventional procedural programming languages—such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider). The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably.

[0122] Although preferred embodiments of the present application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of the embodiments of the present application.

[0123] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal device. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal device that includes said element.

[0124] The above provides a detailed description of the method and apparatus for multimodal processing of smart locker access deviation events provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A method for multimodal processing of access deviation events in smart lockers, characterized in that, The method includes: In response to the first item traceability request initiated by the first user, determine the initial time node of the first user's last opening of the cabinet, and preset the storage period based on the initial time node; Based on the storage time period, the target storage video is extracted from the video storage library. The target person's behavior trajectory in the target storage video is extracted through edge node calculation. Based on the target person's behavior trajectory, the target cabinet where the target item is stored is determined. The system determines the opening and closing records of the target cabinet from the initial time point to the current time point when the first user opened the cabinet. Based on the opening and closing records, it estimates whether the target item is in the target cabinet at the current time point. If so, it initiates identity verification for the first user through the cloud server. When the cloud server returns a message indicating that the identity verification is successful, the target cabinet is opened.

2. The method for multimodal processing of smart locker access deviation events according to claim 1, characterized in that, The process of extracting target stored videos from the video repository based on the storage time period and extracting the target person's behavioral trajectory from the target stored video through edge node calculation includes: Based on the initial time node, the target access image is determined in the target access video, and at least one human figure outline is extracted from the target access image by calculating edge nodes; Generate a character outline selection interface based on at least one character outline and display it to the first user; In response to the target person outline selection result selected by the first user in the person outline selection interface, the target person's behavior trajectory in the target video is extracted based on the target person outline and by calculating edge nodes.

3. The method for multimodal processing of smart locker access deviation events according to claim 1, characterized in that, Determining whether the target item is in the target compartment at the current time point based on the opening and closing record includes: The number of times the target cabinet has been opened and closed between the initial time point and the current time point is determined based on the opening and closing records. When the number of opening and closing operations is one, the target cabinet is marked as not restarted. When the number of opening and closing operations is twice, the target cabinet is marked as being in a state of only restarting once. If the number of opening and closing operations is greater than two, the target cabinet is marked as having been restarted multiple times. If the target cabinet is in a non-restarted state or has only been restarted once, it is determined that the target item is still in the target cabinet at the current time point.

4. The method for multimodal processing of smart locker access deviation events according to claim 1, characterized in that, The step of initiating identity verification for the first user via a cloud server includes: The identity capture camera and the environment capture camera of the smart locker are activated simultaneously. The identity capture camera is used to capture real-time images and video data of the upper body of the first user, and the environment capture camera is used to capture real-time environmental video including the first user. Based on the upper body image and video data of the first user, facial detail features and clothing detail features are extracted, and the facial detail features are transmitted to the cloud server, while the clothing detail features are transmitted to the edge node; The environmental video is transmitted to the edge node, and the human body area is located by calculating the edge node. The clothing area associated with the first user is locked in the human body area based on the clothing details. The edge node is used to calculate and determine the current behavior trajectory of the first user based on the clothing area.

5. The method for multimodal processing of smart locker access deviation events according to claim 4, characterized in that, The step of extracting facial detail features and clothing detail features based on the upper body image and video data of the first user, transmitting the facial detail features to the cloud server, and transmitting the clothing detail features to the edge nodes includes: Based on the upper body image and video data of the first user, facial feature vectors are extracted using a CNN model and transmitted to a cloud server. The cloud server is used to compare the facial feature vectors with pre-stored facial features. Based on the upper body image and video data of the first user, the color feature values ​​of the clothing are extracted using the HSV color histogram algorithm, and the color feature values ​​of the clothing are transmitted to the edge nodes.

6. The method for multimodal processing of smart locker access deviation events according to claim 4, characterized in that, The edge node is used to calculate and determine the real-time distance between the clothing area and the target cabinet area based on the current behavior trajectory of the first user. When the real-time distance is less than or equal to a preset interval threshold, a confirmation message that the first user has been located at the corresponding position of the target cabinet is generated. The step of opening the target cabinet upon receiving authentication confirmation from the cloud server includes: When the cloud server returns a successful authentication message and the edge node simultaneously returns a location confirmation message, the target cabinet is opened.

7. The method for multimodal processing of smart locker access deviation events according to claim 3, characterized in that, The method further includes: Based on the opening and closing records, determine whether the target item is in the target compartment at the current time. If the target cabinet is in a state of multiple restarts, or responds to the first user's request for no items found after the target cabinet is opened, then search for the client information of the second user associated with the target cabinet when it is opened and closed for the second time. Determine the client information of the first user, generate communication connection channel information based on the client information of the first user and the identity information of the second user, and send it to the client of the first user.

8. A device for multimodal processing of access deviation events in smart lockers, characterized in that, The device includes: The response module is used to respond to the first item traceability request initiated by the first user, determine the initial time node of the first user's last opening of the cabinet, and preset the storage period based on the initial time node; The calculation module is used to extract the target storage video from the video storage library based on the storage time period, extract the target person's behavior trajectory in the target storage video through edge node calculation, and determine the target cabinet where the target item is stored based on the target person's behavior trajectory. The execution module is used to determine the opening and closing records of the target cabinet from the initial time point to the current time point when the first user opened the cabinet. Based on the opening and closing records, it estimates whether the target item is in the target cabinet at the current time point. If so, it initiates identity verification for the first user through the cloud server. When it receives the identity verification information returned by the cloud server, it opens the target cabinet.

9. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, which, when executed by a processor, implements the method for multimodal processing of smart locker access deviation events as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • An intelligent article storing and taking method based on a face recognition technology

    CN109948966A

  • Cargo video tracing system

    CN212160738U