An article tracking control method and system for an intelligent storage cabinet

Through the combination of binocular camera and emotional prediction model, high-accuracy item tracking of smart storage cabinets is achieved, solving the problems of high cost or low accuracy in the existing technology, reducing hardware costs and improving tracking efficiency.

CN119763087BActive Publication Date: 2025-06-27SHENZHEN MOTERN TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510260640.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-06-27
Estimated Expiration
2045-03-06

AI Technical Summary

Technical Problem

The item tracking method of existing storage cabinets is high or the accuracy is low, especially through RFID and sensors, which has a higher hardware cost; while the cost of using cameras is low, the accuracy of item tracking needs to be improved.

Method used

A binocular camera is used to obtain feature images, combine the preset emotion prediction model to obtain the user's emotional trend data, obtain the initial intention characteristics of the item through the emotional trend data, and combine the depth image of the item identification to calculate the item prediction trajectory information.

Benefits of technology

A high-accuracy and fast item tracking solution is achieved, reducing the cost of item tracking, no need to set up complex and expensive hardware, and saving human resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119763087B_ABST
    Figure CN119763087B_ABST
Patent Text Reader

Abstract

An object tracking control method and system for an intelligent storage cabinet according to an embodiment of the present invention include obtaining a feature image captured by an upper binocular camera of the intelligent storage cabinet; calling a preset emotion prediction model according to the feature image to obtain emotion trend data of a user, optimizing and iterating the weight parameters through a loss function, and then inputting new face and body features and voice features into the trained emotion prediction model; obtaining preliminary object intention features through the emotion trend data; performing object recognition on the feature image to obtain initial object position information; combining the preliminary object intention features and the initial object position information, calculating to obtain object prediction trajectory information, taking the result of the user's body movement as an auxiliary consideration factor, and combining the depth image of object recognition, realizing a high-accuracy and fast object tracking solution, reducing the cost of object tracking, not requiring the setting of complex and expensive hardware, and saving human resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular, to an article tracking control method for an intelligent storage cabinet, an article tracking control system for an intelligent storage cabinet, a computer device, and a storage medium. Background Art

[0002] The existing article tracking methods for various storage cabinets generally set RFID, weight / pressure sensors, cameras, and Internet of Things (IoT) sensors on articles or cabinet spaces. For example, vending storage cabinets can track the articles in the storage cabinet through RFID and weight / pressure sensors. However, the cost of RFID and sensor hardware is relatively high. In addition, the common article tracking method for existing vending storage cabinets is to identify the payment of articles through cameras, which has a lower cost, but the accuracy of article tracking needs to be improved. Summary of the Invention

[0003] In view of the above problems, embodiments of the present invention are proposed to provide an article tracking control method for an intelligent storage cabinet, an article tracking control system for an intelligent storage cabinet, a computer device, and a storage medium that overcome the above problems or at least partially solve the above problems.

[0004] To solve the above problems, embodiments of the present invention disclose an article tracking control method for an intelligent storage cabinet, including:

[0005] Obtaining a feature image captured by an upper binocular camera of the intelligent storage cabinet;

[0006] Invoking a preset emotion prediction model according to the feature image to obtain user emotion trend data;

[0007] Obtaining preliminary article intention features through the emotion trend data;

[0008] Performing article recognition on the feature image to obtain initial article position information;

[0009] Combining the preliminary article intention features and the initial article position information, and calculating to obtain predicted article trajectory information.

[0010] Preferably, the step of invoking a preset emotion prediction model according to the feature image to obtain user emotion trend data includes

[0011] Performing voice, face, or limb recognition on the feature image to obtain face and limb features;

[0012] Obtaining audio information corresponding to the feature image, and extracting voice features from the audio information,

[0013] Input the facial and limb features and voice features into the emotion prediction model to obtain the emotion trend data of the user.

[0014] Preferably, the emotion trend data includes first emotion trend data and second emotion trend data; obtaining the preliminary item intention features from the emotion trend data includes:

[0015] Determine the first side subtree and the final side subtree of the decision tree;

[0016] Set the first emotion trend data and the second emotion trend data as the feature attribute values of the first side subtree and the final side subtree of the decision tree respectively to obtain the classification result output by the decision tree, and determine the classification result as the preliminary item intention features.

[0017] Preferably, the step of performing item recognition on the feature image to obtain the initial item position information includes:

[0018] Input the feature image into the depth image recognition model to obtain the output initial item position information.

[0019] Preferably, the step of combining the preliminary item intention features and the initial item position information to calculate the predicted item trajectory information includes:

[0020] If the preliminary item intention features show a positive trend, adjust the first probability of the item moving to the first preset location;

[0021] If the preliminary item intention features show a negative trend, adjust the second probability of the item moving to the second preset location;

[0022] Generate the predicted item trajectory information according to the first probability and the second probability.

[0023] Preferably, the step of inputting the facial and limb features and voice features into the emotion prediction model to obtain the emotion trend data of the user includes

[0024] Combine the facial and limb features and voice features to form a feature training sample;

[0025] Initialize the learning rate and weight parameters of the emotion prediction model;

[0026] Calculate the aspect ratio between each image pixel in the feature training sample, and use an activation function to perform normalization to obtain the spatial relationship similarity;

[0027] Calculate the feature similarity between each node through the image pixels after linear transformation in each node;

[0028] Calculate the feature similarity between each voice feature point through the voice syllables after linear transformation in each voice feature;

[0029] Calculate the specific target similarity between each node included in the specific target area through the image pixels of the specific target area;

[0030] Construct a matrix of node edge weights through the spatial relationship similarity, specific target similarity, and feature similarity between each node in the emotion prediction model;

[0031] Construct a fully connected graph of each node composed of the feature map nodes of the feature training samples, the connecting node edges, and the matrix of node edge weights;

[0032] Optimize and iterate the weight parameters through the loss function to obtain a trained emotion prediction model, and then input new facial limb features and voice features into the trained emotion prediction model to obtain the output emotion trend data.

[0033] An embodiment of the present invention discloses an item tracking control system for an intelligent storage cabinet, including:

[0034] A first acquisition module, configured to acquire a feature image captured by the upper binocular camera of the intelligent storage cabinet;

[0035] An emotion trend data module, configured to call a preset emotion prediction model according to the feature image to obtain the emotion trend data of the user;

[0036] A first recognition module, configured to obtain preliminary item intention features through the emotion trend data;

[0037] A second recognition module, configured to perform item recognition on the feature image to obtain item initial position information;

[0038] An item prediction trajectory information module, configured to combine the preliminary item intention features and the item initial position information to calculate the item prediction trajectory information.

[0039] Preferably, the emotion trend data module includes

[0040] A face and limb recognition sub-module, configured to perform voice face or limb recognition on the feature image to obtain face and limb features;

[0041] An audio sub-module, configured to acquire audio information corresponding to the feature image and extract voice features from the audio information,

[0042] An input sub-module, configured to input the face and limb features and the voice features into the emotion prediction model to obtain the emotion trend data of the user.

[0043] An embodiment of the present invention also discloses a computer device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the item tracking control of the intelligent storage cabinet described above are implemented.

[0044] An embodiment of the present invention also discloses a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the item tracking control of the intelligent storage cabinet described above are implemented.

[0045] The embodiments of the present invention have the following advantages:

[0046] In the embodiment of the present invention, the item tracking control method of the intelligent storage cabinet includes: obtaining a feature image captured by the upper binocular camera of the intelligent storage cabinet; calling a preset emotion prediction model according to the feature image to obtain the emotion trend data of the user; obtaining the preliminary item intention features through the emotion trend data; performing item recognition on the feature image to obtain the initial item position information; combining the preliminary item intention features and the initial item position information to calculate and obtain the item prediction trajectory information, taking the result of the user's body movement as an auxiliary consideration factor, and combining the depth image of the item recognition, realizing a high-accuracy and fast item tracking solution, reducing the cost of item tracking, not requiring the setting of complex and expensive hardware, and saving human resources. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0048] Figure 1 is a flowchart of the steps of an embodiment of the item tracking control method of an intelligent storage cabinet according to an embodiment of the present invention;

[0049] Figure 2 is a block diagram of the structure of an embodiment of the item tracking control system of an intelligent storage cabinet according to an embodiment of the present invention;

[0050] Figure 3 is an internal structure diagram of a computer device in an embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0051] In order to make the technical problems, technical solutions and beneficial effects solved by the embodiments of the present invention clearer, the following further details the embodiments of the present invention in conjunction with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0052] In an embodiment of the present invention, a binocular camera is provided on an intelligent storage cabinet. The binocular camera can obtain the depth images of the item to be tracked and the user. The intended action of the user is obtained based on the depth image of the user, and the category information such as the name of the item is obtained based on the depth image of the item. By combining the intended action of the user and the category information such as the name of the item, the position of the item is predicted. Taking the result of the user's body movement as an auxiliary consideration factor and combining the depth image of item recognition, a high-accuracy and fast item tracking solution is achieved, reducing the cost of item tracking, without the need to set up complex and expensive hardware, and saving human resources.

[0053] Refer to Figure 1 , which shows a flowchart of the steps of an embodiment of an item tracking control method for an intelligent storage cabinet according to an embodiment of the present invention. Specifically, it may include the following steps:

[0054] Step 101, obtain the feature image captured by the binocular camera on the intelligent storage cabinet;

[0055] In an embodiment of the present invention, a terminal and a binocular camera may be provided on the intelligent storage cabinet; the terminal may be, but is not limited to, various in-vehicle computers, personal computers, laptop computers, smart phones, tablet computers, and portable wearable devices. The specific type of the terminal is not limited in the embodiment of the present invention. The operating system of the terminal may include Android, Harmony OS, IOS, Windows Phone, Windows, etc. The present invention does not impose too many restrictions on this.

[0056] In a preferred embodiment, the item tracking control method for the intelligent storage cabinet provided by the embodiment of the present invention may be applied to an application environment including a terminal and a server. Among them, the terminal communicates with the server through a network. The terminal may be, but is not limited to, various personal computers, laptop computers, and tablet computers, and the server may be implemented by an independent server or a server cluster composed of multiple servers.

[0057] It should be noted that the intelligent storage cabinet may refer to a vending cabinet or a locker. The embodiment of the present invention does not impose too many restrictions on this. In addition, for the structure of the intelligent storage cabinet, its main cabinet structure may include side plates, top plates, bottom plates, back plates, support feet, bases, adjustable casters, etc.; the internal structure may include adjustable shelves or fixed shelves, vertical or horizontal partition dividers, slide rail drawers, classification grids, pull-out metal wire baskets, swing doors or sliding doors and corresponding mechanical locks, etc. The embodiment of the present invention does not impose too many restrictions on the structure and function of the intelligent storage cabinet.

[0058] Specifically, in the embodiments of the present invention, a binocular camera refers to collecting images at two different positions simultaneously to obtain a fused depth image, realizing the function of binocular imaging, and solving the problem that users and items may block each other;

[0059] When the user opens the intelligent storage cabinet to pick up an item, the terminal can trigger the control of the binocular camera to perform an image acquisition operation to obtain a feature image. Specifically, the intelligent storage cabinet can be provided with an opening and closing sensor between the cabinet door and the main cabinet body. When the user opens the cabinet door, the opening and closing sensor is triggered to send a signal to the terminal on the intelligent storage cabinet, and the terminal triggers the control of the binocular camera to perform an image acquisition operation. That is, the terminal can be connected to the opening and closing sensor and the binocular camera. It should be noted that the opening and closing sensor can be a Hall sensor, and the embodiments of the present invention do not limit this too much;

[0060] Specifically, the feature image refers to an image that includes both a user image and an item image; after obtaining the feature image, the terminal can call an internal intelligent model to process the feature image data, and the specific steps are as follows.

[0061] Step 102, call a preset emotion prediction model according to the feature image to obtain the user's emotion trend data;

[0062] In the embodiments of the present invention, the terminal can include a preset emotion prediction model. After obtaining the feature image, the preset emotion prediction model is called to obtain the user's emotion trend data.

[0063] The emotion prediction model can be composed of multiple neural networks. For example, the emotion prediction model can include three network parts. The first network part can include a plurality of hidden layers and activation layers connected in sequence, and is used to identify the face features and limb features in the feature image; the second network part is used for the extraction of voice features, and it includes a voice conversion layer, a hidden layer, a fully connected layer, and a max pooling layer. Specifically, the voice conversion layer can include an encoder (Bi-LSTM) module, an attention module, and a decoder (LSTM) module. The embodiments of the present invention do not limit the specific structure of the voice conversion layer too much; in addition, the third network part can include a fully connected graph module, a hidden layer, a fully connected layer, a normalization layer, etc. for each node.

[0064] It should be noted that the intelligent storage cabinet can be provided with a voice acquisition device, and the voice acquisition device can be used to collect the user's voice data; specifically, the voice acquisition device can also be connected to the terminal, and the terminal triggers the control of the voice acquisition device to perform voice acquisition.

[0065] Specifically applied to the embodiments of the present invention, the calling a preset emotion prediction model according to the feature image to obtain the user's emotion trend data includes

[0066] Perform face or limb recognition on the feature image to obtain face and limb features;

[0067] Obtain the audio information corresponding to the feature image, and extract voice features from the audio information.

[0068] Input the face and limb features and voice features into the emotion prediction model to obtain the emotion trend data of the user.

[0069] First, perform face or limb recognition on the feature image to obtain face and limb features. Specifically, input the feature image into the first network part to obtain the output face and limb features. Further, when audio data information of the user is detected, that is, after the voice acquisition device acquires the voice data when the user picks up an object, voice features can also be extracted from the voice data. Specifically, input the voice data into the second network part, and perform recognition through an encoder (Bi-LSTM) module, an attention module, a decoder (LSTM) module, a hidden layer, a fully connected layer, and a max pooling layer to obtain the output voice features. Then input the face and limb features and voice features into the emotion prediction model to obtain the emotion trend data of the user.

[0070] Specifically, the emotion trend data can be score data. For example, within the score range of [1 - 100], 1 represents the most positive trend, 100 represents the most negative trend, and the score range from 1 to 100 represents the transition distribution process from a positive trend to a negative trend. Negative trends can be trends expressed by face or limb features such as "panic" and "alarm", while positive trends can be trends expressed by face or limb features such as "smile" and "nod".

[0071] In the embodiments of the present invention, by combining voice, face or limb movements, and through the method of multi-modal fusion, training samples can be obtained, which can enhance the robustness of model training, improve the accuracy of model recognition, and improve the real-time performance of model recognition.

[0072] Specifically applied to the embodiments of the present invention, the step of inputting the face and limb features and voice features into the emotion prediction model to obtain the emotion trend data of the user includes

[0073] Combine the face and limb features and voice features to form a feature training sample;

[0074] Initialize the learning rate and weight parameters of the emotion prediction model;

[0075] Calculate the aspect ratio between image pixels in the feature training sample, and perform normalization using an activation function to obtain the spatial relationship similarity;

[0076] It should be noted that the aspect ratio refers to the ratio of the width to the height of the feature map in the feature training samples; the spatial relationship similarity is obtained by normalizing with the aspect ratio.

[0077] Specifically, the spatial relationship similarity S1 may include:

[0078] ;

[0079] Among them, represents the spatial relationship similarity, represents the aspect ratio of the image pixels, i = 1, 2, 3 ······ n, where n is the number of image pixels, represents the first learnable parameter optimized by the first network part, represents the second learnable parameter optimized by the first network part.

[0080] Calculate the feature similarity between each voice feature point through the voice syllables after linear transformation in each voice feature;

[0081] Among them, the feature similarity is: ;

[0082] Among them, represents the feature similarity, represents the first voice syllable component after linear transformation, i = 1, 2, 3 ······ N, where N is the number of voice syllables, represents the second voice syllable component after linear transformation, i = 1, 2, 3 ······ N, where N is the number of voice syllables.

[0083] Calculate the specific target similarity between each node included in the specific target area through the image pixels of the specific target area; among them, the node refers to the pixel point in the feature map; the specific target area refers to the intersecting edge area between the user's body part and the item;

[0084] Furthermore, the specific target similarity can be expressed as follows:

[0085] ;

[0086] Among them, represents the specific target similarity, represents the proportion parameter of the specific target area in the feature map, represents the width data of the specific target area, i = 1, 2, 3 ······ o, where o refers to the number of image pixels in the specific target area, represents the height data of the specific target area, i = 1, 2, 3 ······ o, where o refers to the number of image pixels in the specific target area.

[0087] Construct a matrix of node edge weights based on the spatial relationship similarity, specific target similarity, and feature similarity between each pair of nodes in the emotion prediction model;

[0088] Then the matrix of node edge weights ;

[0089] Construct a fully connected graph for each node composed of the feature map nodes of the feature training samples, the connecting node edges, and the matrix of node edge weights;

[0090] Optimize and iterate the weight parameters through the loss function to obtain a trained emotion prediction model, and then input new facial limb features and speech features into the trained emotion prediction model to obtain the output emotion trend data.

[0091] Specifically, the loss function includes the following expression,

[0092] ;

[0093] Wherein, represents the loss function, w represents the node edge weight, represents the output emotion trend data, i = 1, 2, 3 ······ n, n is the number of image pixels, represents the label data of the sample, i = 1, 2, 3 ······ n, n is the number of image pixels, represents the proportion parameter of the specific target area in the feature map.

[0094] Step 103, obtain the preliminary intention features of the item through the emotion trend data;

[0095] In the embodiment of the present invention, after obtaining the emotion trend data, the preliminary intention features of the item can be obtained through the emotion trend data. Specifically, the emotion trend data includes the first emotion trend data and the second emotion trend data; the obtaining of the preliminary intention features of the item through the emotion trend data includes:

[0096] Determine the first sub - tree of the decision tree and the final sub - tree of the decision tree;

[0097] Set the first emotion trend data and the second emotion trend data as the feature attribute values of the first sub - tree of the decision tree and the final sub - tree of the decision tree respectively, obtain the classification result output by the decision tree, and determine the classification result as the preliminary intention features of the item.

[0098] Specifically, the first sentiment trend data may refer to positive trend data, and the second sentiment trend data may refer to negative trend data. By setting the positive trend data and the negative trend data in the subtrees on both sides of the decision tree structure, the combination of the sentiment prediction model and the decision tree can improve the accuracy of user intention prediction, reduce the low requirements for data preprocessing, and solve the problem of easy overfitting of the model;

[0099] Step 104: Perform item recognition on the feature image to obtain the initial item position information;

[0100] In a preferred embodiment of the embodiment of the present invention, the performing item recognition on the feature image to obtain the initial item position information includes: inputting the feature image into a depth image recognition model to obtain the output initial item position information.

[0101] The initial item position information may refer to the three-dimensional coordinates of the x-axis, y-axis, and z-axis, or may be coordinate data including depth distance information. The embodiment of the present invention does not impose excessive restrictions on this.

[0102] It should be noted that obtaining the initial item position information through the binocular camera and the depth image recognition model can be obtained by those skilled in the art through conventional technology, and the output initial item position information will not be elaborated one by one in the embodiment of the present invention.

[0103] Step 105: Combine the preliminary item intention feature and the initial item position information, and calculate to obtain the item prediction trajectory information.

[0104] Specifically applied to the embodiment of the present invention, after obtaining the preliminary item intention feature and the initial item position information, the preliminary item intention feature and the initial item position information can be combined to generate the item prediction trajectory information; specifically, the combining the preliminary item intention feature and the initial item position information and calculating to obtain the item prediction trajectory information includes: if the preliminary item intention feature is a positive trend, adjusting the first probability of the item moving to the first preset location; if the preliminary item intention feature is a negative trend, adjusting the second probability of the item moving to the second preset location; generating the item prediction trajectory information starting from the initial item position information according to the first probability and the second probability.

[0105] Specifically, the first preset location is the payment counter, the first probability refers to the probability that the user carries the item to the payment counter "i.e., checks out", the second preset location is the door, and the second probability refers to the probability that the user carries the item to the door "i.e., tries to escape without paying", and then generating the item prediction trajectory through the probability adjustment of the two.

[0106] In an embodiment of the present invention, the method for controlling item tracking of the intelligent storage cabinet includes: obtaining a feature image captured by the upper binocular camera of the intelligent storage cabinet; calling a preset emotion prediction model according to the feature image to obtain the emotion trend data of the user; obtaining the preliminary intention features of the item through the emotion trend data; performing item recognition on the feature image to obtain the initial position information of the item; combining the preliminary intention features of the item and the initial position information of the item, calculating to obtain the predicted trajectory information of the item, taking the result of the user's body movement as an auxiliary consideration factor, and combining the depth image of item recognition, realizing a high-accuracy and fast item tracking solution, reducing the cost of item tracking, not requiring the setting of complex and expensive hardware, and saving human resources.

[0107] It should be noted that for the method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the embodiments of the present invention are not limited by the described action sequence, because according to the embodiments of the present invention, some steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily essential for the embodiments of the present invention.

[0108] Refer to Figure 2 , which shows a structural block diagram of an embodiment of an item tracking control system of an intelligent storage cabinet according to an embodiment of the present invention, and specifically may include the following modules:

[0109] The first acquisition module 301 is used to obtain a feature image captured by the upper binocular camera of the intelligent storage cabinet;

[0110] The emotion trend data module 302 is used to call a preset emotion prediction model according to the feature image to obtain the emotion trend data of the user;

[0111] The first recognition module 303 is used to obtain the preliminary intention features of the item through the emotion trend data;

[0112] The second recognition module 304 is used to perform item recognition on the feature image to obtain the initial position information of the item;

[0113] The item predicted trajectory information module 305 is used to combine the preliminary intention features of the item and the initial position information of the item, and calculate to obtain the predicted trajectory information of the item.

[0114] Preferably, the emotion trend data module includes

[0115] The face and body recognition sub-module is used to perform voice face or body recognition on the feature image to obtain face and body features;

[0116] An audio sub-module, configured to obtain audio information corresponding to the feature image and extract speech features from the audio information.

[0117] An input sub-module, configured to input the face and limb features and the speech features into the emotion prediction model to obtain emotion trend data of the user.

[0118] Preferably, the emotion trend data includes first emotion trend data and second emotion trend data; the first recognition module obtaining the preliminary item intention features through the emotion trend data includes:

[0119] A determination sub-module, configured to determine the first side subtree and the final side subtree of the decision tree.

[0120] A setting sub-module, configured to respectively set the first emotion trend data and the second emotion trend data as the feature attribute values of the first side subtree and the final side subtree of the decision tree, obtain the classification result output by the decision tree, and determine the classification result as the preliminary item intention features.

[0121] Preferably, the second recognition module includes:

[0122] An item initial position sub-module, configured to input the feature image into a depth image recognition model to obtain the output item initial position information.

[0123] Preferably, the item prediction trajectory information module includes:

[0124] A first probability sub-module, configured to adjust the first probability of the item moving to a first preset location if the preliminary item intention features are in a positive trend.

[0125] A second probability sub-module, configured to adjust the second probability of the item moving to a second preset location if the preliminary item intention features are in a negative trend.

[0126] A generation sub-module, configured to generate item prediction trajectory information according to the first probability and the second probability.

[0127] Preferably, the input sub-module includes

[0128] A composition unit, configured to compose the face and limb features and the speech features into a feature training sample.

[0129] An initialization unit, configured to initialize the learning rate and weight parameters of the emotion prediction model.

[0130] An aspect ratio unit, configured to calculate the aspect ratio between each image pixel in the feature training sample, and perform normalization using an activation function to obtain the spatial relationship similarity.

[0131] A first calculation unit for calculating the feature similarity between each node by the linearly transformed image pixels in each node;

[0132] A second calculation unit for calculating the feature similarity between each voice feature point by the linearly transformed voice syllables in each voice feature;

[0133] A third calculation unit for calculating the specific target similarity between each node included in a specific target area by the image pixels of the specific target area;

[0134] A construction unit for constructing a matrix of node edge weights through the spatial relationship similarity, specific target similarity, and feature similarity between each node in the emotion prediction model;

[0135] A fully connected graph unit for constructing a fully connected graph of each node composed of the feature graph nodes of the feature training samples, the connecting node edges, and the matrix of node edge weights;

[0136] An iterative unit for optimizing and iterating the weight parameters through a loss function to obtain a trained emotion prediction model, and then inputting new face limb features and voice features into the trained emotion prediction model to obtain output emotion trend data.

[0137] For the apparatus embodiment, since it is basically similar to the method embodiment, the description is relatively simple. For related parts, refer to the partial description of the method embodiment.

[0138] For the specific limitations of the item tracking control system of the intelligent storage cabinet, reference can be made to the limitations on the item tracking control method of the intelligent storage cabinet in the above text, which will not be elaborated here. Each module in the above item tracking control system of the intelligent storage cabinet can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above respective modules.

[0139] The above provided item tracking control system of the intelligent storage cabinet can be used to execute the item tracking control method provided in any of the above embodiments, and has corresponding functions and beneficial effects.

[0140] In one embodiment, a computer device is provided. The computer device can be a terminal, and its internal structure diagram can be as Figure 3As shown in the figure. The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected by a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it realizes a method for tracking and controlling items in an intelligent storage cabinet. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covered on the display screen, or a button, a trackball, or a touchpad set on the shell of the computer device, or an external keyboard, touchpad, or mouse, etc.

[0141] Those skilled in the art can understand that Figure 3 the structure shown in the figure is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0142] In one embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, the following steps are realized:

[0143] Obtain a feature image captured by the upper binocular camera of the intelligent storage cabinet;

[0144] Call a pre-set emotion prediction model according to the feature image to obtain the emotion trend data of the user;

[0145] Obtain the preliminary item intention features through the emotion trend data;

[0146] Perform item recognition on the feature image to obtain the initial position information of the item;

[0147] Combine the preliminary item intention features and the initial position information of the item, and calculate to obtain the predicted trajectory information of the item.

[0148] Preferably, the step of calling a pre-set emotion prediction model according to the feature image to obtain the emotion trend data of the user includes

[0149] Perform voice, face, or limb recognition on the feature image to obtain face and limb features;

[0150] Obtain the audio information corresponding to the feature image, and extract voice features from the audio information,

[0151] Input the facial and limb features and the voice features into the emotion prediction model to obtain the emotion trend data of the user.

[0152] Preferably, the emotion trend data includes first emotion trend data and second emotion trend data; obtaining the preliminary item intention features from the emotion trend data includes:

[0153] Determine the first sub-tree of the decision tree and the final sub-tree of the decision tree;

[0154] Set the first emotion trend data and the second emotion trend data as the characteristic attribute values of the first sub-tree and the final sub-tree of the decision tree respectively, obtain the classification result output by the decision tree, and determine the classification result as the preliminary item intention features.

[0155] Preferably, the identifying the item from the feature image to obtain the initial item position information includes:

[0156] Input the feature image into the depth image recognition model to obtain the output initial item position information.

[0157] Preferably, the combining the preliminary item intention features and the initial item position information to calculate and obtain the item prediction trajectory information includes:

[0158] If the preliminary item intention features are in a positive trend, adjust the first probability of the item moving to the first preset location;

[0159] If the preliminary item intention features are in a negative trend, adjust the second probability of the item moving to the second preset location;

[0160] Generate the item prediction trajectory information according to the first probability and the second probability.

[0161] Preferably, the inputting the facial and limb features and the voice features into the emotion prediction model to obtain the emotion trend data of the user includes

[0162] Combine the facial and limb features and the voice features to form a feature training sample;

[0163] Initialize the learning rate and weight parameters of the emotion prediction model;

[0164] Calculate the aspect ratio between each image pixel in the feature training sample, and use an activation function to normalize it to obtain the spatial relationship similarity;

[0165] Calculate the feature similarity between each node through the image pixels after linear transformation in each node;

[0166] Calculate the feature similarity between each voice feature point through the voice syllables after linear transformation in each voice feature;

[0167] Calculate the specific target similarity between each node included in the specific target area through the image pixels of the specific target area;

[0168] Construct a matrix of node edge weights through the spatial relationship similarity, specific target similarity, and feature similarity between each node in the emotion prediction model;

[0169] Construct a fully connected graph of each node composed of the feature map nodes of the feature training samples, the connecting node edges, and the matrix of node edge weights;

[0170] Optimize and iterate the weight parameters through the loss function to obtain a trained emotion prediction model, and then input the new face limb features and voice features into the trained emotion prediction model to obtain the output emotion trend data.

[0171] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0172] Obtain the feature image captured by the upper binocular camera of the intelligent storage cabinet;

[0173] Call the preset emotion prediction model according to the feature image to obtain the emotion trend data of the user;

[0174] Obtain the preliminary item intention features through the emotion trend data;

[0175] Perform item recognition on the feature image to obtain the initial item position information;

[0176] Combine the preliminary item intention features and the initial item position information to calculate the predicted item trajectory information.

[0177] Preferably, the step of calling the preset emotion prediction model according to the feature image to obtain the emotion trend data of the user includes

[0178] Perform voice, face, or limb recognition on the feature image to obtain face limb features;

[0179] Obtain the audio information corresponding to the feature image, and extract the voice features from the audio information,

[0180] Input the face limb features and voice features into the emotion prediction model to obtain the emotion trend data of the user.

[0181] Preferably, the sentiment trend data includes first sentiment trend data and second sentiment trend data; obtaining the preliminary item intention features from the sentiment trend data includes:

[0182] Determine the first sub-tree and the final sub-tree of the decision tree;

[0183] Set the first sentiment trend data and the second sentiment trend data as the feature attribute values of the first sub-tree and the final sub-tree of the decision tree respectively, obtain the classification result output by the decision tree, and determine the classification result as the preliminary item intention features.

[0184] Preferably, the item recognition of the feature image to obtain the initial item position information includes:

[0185] Input the feature image into the depth image recognition model to obtain the output initial item position information.

[0186] Preferably, combining the preliminary item intention features and the initial item position information to calculate and obtain the item prediction trajectory information includes:

[0187] If the preliminary item intention feature is a positive trend, adjust the first probability of the item moving to the first preset location;

[0188] If the preliminary item intention feature is a negative trend, adjust the second probability of the item moving to the second preset location;

[0189] Generate the item prediction trajectory information according to the first probability and the second probability.

[0190] Preferably, inputting the face and limb features and the voice features into the emotion prediction model to obtain the sentiment trend data of the user includes

[0191] Combining the face and limb features and the voice features to form a feature training sample;

[0192] Initialize the learning rate and weight parameters of the emotion prediction model;

[0193] Calculate the aspect ratio between each image pixel in the feature training sample, and use the activation function to normalize to obtain the spatial relationship similarity;

[0194] Calculate the feature similarity between each node through the linearly transformed image pixels in each node;

[0195] Calculate the feature similarity between each voice feature point through the linearly transformed voice syllables in each voice feature;

[0196] Calculate the specific target similarity between each node included in the specific target area through the image pixels of the specific target area;

[0197] Construct a matrix of node edge weights through the spatial relationship similarity, specific target similarity, and feature similarity between each node in the emotion prediction model;

[0198] Construct a fully connected graph of each node composed of the feature map nodes of the feature training samples, the connecting node edges, and the matrix of node edge weights;

[0199] Optimize and iterate the weight parameters through the loss function to obtain the trained emotion prediction model, and then input the new facial limb features and speech features into the trained emotion prediction model to obtain the output emotion trend data.

[0200] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above method embodiments. Among them, any reference to memory, storage, database, or other media used in the various embodiments provided in the present application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0201] Each embodiment in this specification is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other.

[0202] Those skilled in the art should understand that the embodiments of the present invention can be provided as methods, apparatuses, or computer program products. Therefore, the embodiments of the present invention can take the form of all-hardware embodiments, all-software embodiments, or embodiments combining software and hardware aspects. Moreover, the embodiments of the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) that contain computer-usable program code.

[0203] The embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of apparatuses, terminal devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing terminal devices generate a device for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0204] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing terminal devices to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction method that implements the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0205] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal devices, such that a series of operation steps are executed on the computer or other programmable terminal devices to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable terminal devices provide steps for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0206] Although the preferred embodiments of the embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications once they know the basic creative concepts. Therefore, the appended claims are intended to be construed to include the preferred embodiments and all changes and modifications falling within the scope of the embodiments of the present invention.

[0207] Finally, it should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, device, article or terminal device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, device, article or terminal device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, device, article or terminal device comprising the said element.

[0208] The above has introduced in detail an article tracking control method for an intelligent storage cabinet, an article tracking control system for an intelligent storage cabinet, a computer device and a storage medium provided by the present invention. Specific examples are used in this text to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. A method for tracking and controlling items in a smart storage cabinet, characterized in that: include: Acquire the feature image captured by the binocular camera on the smart storage cabinet; Calling a preset emotion prediction model according to the feature image to obtain the user's emotion trend data; Obtaining preliminary intention characteristics of the item through the sentiment trend data; Performing object recognition on the feature image to obtain initial position information of the object; The preliminary intention feature of the item and the initial position information of the item are combined to calculate the predicted trajectory information of the item; The step of calling a preset emotion prediction model according to the feature image to obtain the user's emotion trend data includes: Inputting facial and body features and voice features into the emotion prediction model to obtain the user's emotion trend data includes the following steps: Combining the facial and body features and voice features into feature training samples; Initializing the learning rate and weight parameters of the emotion prediction model; The aspect ratio between each image pixel in the feature training sample is calculated, and the activation function is used for normalization to obtain the spatial relationship similarity; the spatial relationship similarity S1 may include: Among them, S1 represents the spatial relationship similarity, t i represents the aspect ratio of the image pixels, i=1,2,3······n, n is the number of image pixels, m represents the first learnable parameter optimized by the first network part, and d represents the second learnable parameter optimized by the first network part; The feature similarity between each speech feature point is calculated by the speech syllable after linear transformation in each speech feature; The specific target similarity between each node in the specific target area is obtained by calculating the image pixels of the specific target area; the specific target area refers to the intersection edge area of ​​the user's body part and the object; The matrix of node edge weights is constructed through the spatial relationship similarity, specific target similarity and feature similarity between each node in the sentiment prediction model; Construct a fully connected graph of each node consisting of feature graph nodes of feature training samples, edges connecting nodes, and a matrix of node edge weights; The weight parameters are iterated through loss function optimization to obtain a trained emotion prediction model, and then new facial and body features and voice features are input into the trained emotion prediction model to obtain output emotion trend data.

2. The method according to claim 1, characterized in that The step of calling a preset emotion prediction model according to the feature image to obtain the user's emotion trend data includes: Perform voice, face or body recognition on the feature image to obtain face and body features; The audio information corresponding to the feature image is obtained, and the speech feature is extracted from the audio information.

3. The method according to claim 1, characterized in that The emotional trend data includes first emotional trend data and second emotional trend data; The obtaining of preliminary intention features of the item through the sentiment trend data includes: Determine the first bypass subtree and the final bypass subtree of the decision tree; The first sentiment trend data and the second sentiment trend data are respectively set as the characteristic attribute values ​​of the first bypass subtree and the final bypass subtree of the decision tree, to obtain the classification result of the output of the decision tree, and the classification result is determined as the preliminary intention feature of the item.

4. The method according to claim 1, characterized in that The step of performing object recognition on the feature image to obtain initial position information of the object includes: The feature image is input into a deep image recognition model to obtain output information of the initial position of the object.

5. The method according to claim 1, characterized in that: The step of combining the preliminary intention feature of the item and the initial location information of the item to calculate the predicted trajectory information of the item includes: If the preliminary intention characteristic of the item is a positive trend, the first probability of the item moving to the first preset location is adjusted; If the preliminary intention characteristic of the item is a negative trend, then adjusting the second probability of the item moving to the second preset location; Generate object prediction trajectory information according to the first probability and the second probability.

6. An item tracking control system for a smart storage cabinet, characterized in that: include: A first acquisition module is used to acquire a feature image captured by a binocular camera on the smart storage cabinet; An emotion trend data module, used to call a preset emotion prediction model according to the feature image to obtain the user's emotion trend data; A first recognition module, used to obtain preliminary intention features of an item through the sentiment trend data; A second recognition module is used to perform object recognition on the feature image to obtain initial position information of the object; An item prediction trajectory information module is used to combine the item's preliminary intention features and the item's initial position information to calculate the item's prediction trajectory information; The step of calling a preset emotion prediction model according to the feature image to obtain the user's emotion trend data includes: Inputting facial and body features and voice features into the emotion prediction model to obtain the user's emotion trend data includes the following steps: Combining the facial and body features and voice features into feature training samples; Initializing the learning rate and weight parameters of the emotion prediction model; The aspect ratio between each image pixel in the feature training sample is calculated, and the activation function is used for normalization to obtain the spatial relationship similarity; the spatial relationship similarity S1 may include: Among them, S1 represents the spatial relationship similarity, t i represents the aspect ratio of the image pixels, i=1,2,3······n, n is the number of image pixels, m represents the first learnable parameter optimized by the first network part, and d represents the second learnable parameter optimized by the first network part; The feature similarity between each speech feature point is calculated by the speech syllable after linear transformation in each speech feature; The specific target similarity between each node in the specific target area is obtained by calculating the image pixels of the specific target area; the specific target area refers to the intersection edge area of ​​the user's body part and the object; The matrix of node edge weights is constructed through the spatial relationship similarity, specific target similarity and feature similarity between each node in the sentiment prediction model; Construct a fully connected graph of each node consisting of feature graph nodes of feature training samples, edges connecting nodes, and a matrix of node edge weights; The weight parameters are iterated through loss function optimization to obtain a trained emotion prediction model, and then new facial and body features and voice features are input into the trained emotion prediction model to obtain output emotion trend data.

7. The system according to claim 6, characterized in that The emotional trend data module includes A face and body recognition submodule, used for performing voice, face or body recognition on the feature image to obtain face and body features; The audio submodule is used to obtain the audio information corresponding to the feature image and extract the speech features from the audio information.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the item tracking control method of the smart storage cabinet described in any one of claims 1 to 5 are implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the item tracking control method of the smart storage cabinet described in any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Artificial intelligence vending method without fixed cabinet

    CN115482563A