User identification method and device for canteen intelligent settlement

By using semantic segmentation technology and consistency matching method in the canteen intelligent settlement system, the user identity identification problem caused by crowd occlusion is solved, and the accuracy of settlement is improved.

CN119442205BActive Publication Date: 2025-05-06STATE GRID SIJI FEITIAN (LANZHOU) CLOUD TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510042976.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-10
Publication Date
2025-05-06
Estimated Expiration
2045-01-10

AI Technical Summary

Technical Problem

In the existing canteen intelligent settlement system, due to crowd occlusion, the user identity of the meal plate is difficult to identify, affecting the accuracy of settlement.

Method used

The user identification device including a perception layer, a subject layer and a client is adopted to collect video data streams through the camera perception nodes, and the overall character, local unit and dining tray unit video streams are divided using semantic segmentation technology, and the correlation between dining tray and user is determined through clothing consistency and motion consistency matching.

Benefits of technology

It improves the accuracy of user identity identification of meal plates and solves the problem that user identity is difficult to identify in traditional unmanned settlement methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119442205B_ABST
    Figure CN119442205B_ABST
Patent Text Reader

Abstract

The present invention relates to a user identification method and device for intelligent settlement in a cafeteria, and to the technical field of identity identification, including: when receiving biometric data to be settled, receiving a video data stream from a camera perception node, and receiving a dish distribution data stream from a background configuration node; performing unit segmentation to obtain an overall unit video stream, a local unit video stream, and a plate unit video stream; performing clothing consistency matching to obtain a first compensated overall unit video stream and a clothing blind area local unit video stream; performing motion consistency matching to obtain a second compensated overall unit video stream, performing identity identification through the above data, and obtaining an identity identification plate unit video stream, then receiving the dish distribution data stream, associating it with the identity identification plate unit video stream and the biometric data to be settled, and constructing an identity identification document, which is sent to a client. This solves the technical problem in the prior art that the identity of a user with a plate is difficult to identify due to the occlusion of the crowd.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of identity authentication, and in particular to a user authentication method and device for intelligent settlement in a cafeteria. Background Art

[0002] The canteen's unmanned settlement platform can realize unmanned food selection and settlement. There are currently two ways to implement it. One is to use an image analysis device deployed at the settlement terminal to identify the food on the plate and then execute the bill settlement. The disadvantage is that when users select dishes, the dishes may be mixed, resulting in poor image recognition accuracy.

[0003] Secondly, in order to solve the shortcomings of the first method, it is proposed to deploy an image acquisition device in the food selection area to analyze the food selection process image, follow the plate to identify the dishes and execute the bill settlement. However, due to the crowd occlusion problem, the correlation between the plate and the user cannot be distinguished, resulting in the user identity of the plate being difficult to identify, so it is almost not used. Summary of the invention

[0004] The present invention aims to solve the technical problem in the prior art that the identity of a user with a meal plate is difficult to identify due to the obstruction of the crowd, and provides a user identification method and device for intelligent settlement in a cafeteria to solve the problem.

[0005] The technical solution of the present invention to solve the above technical problems is as follows:

[0006] In the first aspect, the present invention provides a user identification method for intelligent settlement in a cafeteria, and a user identification device for intelligent settlement in a cafeteria, the device comprising a perception layer, a main layer and a client, the perception layer comprising a camera perception node and a background configuration node, including: when receiving biometric data to be settled from the client, receiving a video data stream from the camera perception node through the perception layer, and receiving a dish distribution data stream from the background configuration node; receiving the video data stream from the perception layer through a semantic segmentation main network of the main layer for unit segmentation, and obtaining an overall unit video stream, a local unit video stream and a plate unit video stream, wherein the local unit video stream represents a video stream having at most two unit types of a plate unit, an arm unit, a clothing unit and a head unit as one, and the overall unit video stream represents a video stream having at least two unit types of a plate unit, an arm unit, a clothing unit and a head unit as one; receiving the video data stream from the perception layer through a primary semantic segmentation sub-network of the main layer The semantic segmentation main network receives the local unit video stream for clothing consistency matching to obtain a first compensated overall unit video stream and a clothing blind spot local unit video stream; receives the clothing blind spot local unit video stream from the semantic segmentation main network through the secondary semantic segmentation sub-network of the main layer for motion consistency matching to obtain a second compensated overall unit video stream; receives the overall unit video stream, the first compensated overall unit video stream and the second compensated overall unit video stream from the semantic segmentation main network, the primary semantic segmentation sub-network and the secondary semantic segmentation sub-network through the identity authentication network of the main layer, receives the biometric data to be calculated from the perception layer to authenticate the identity of the plate unit video stream, and after obtaining the identity authentication plate unit video stream, receives the dish distribution data stream from the perception layer, associates it with the identity authentication plate unit video stream and the biometric data to be calculated for storage, constructs an identity authentication document, and sends it to the client.

[0007] In a second aspect, the present invention provides a user authentication device for intelligent settlement in a cafeteria, including a perception layer, a main layer and a client, wherein the perception layer includes a camera perception node and a background configuration node, including: a data acquisition module, which is used to receive the biometric data to be settled from the client, receive the video data stream from the camera perception node through the perception layer, and receive the dish distribution data stream from the background configuration node; a unit segmentation module, which is used to receive the video data stream from the perception layer through the semantic segmentation main network of the main layer for unit segmentation, and obtain the overall unit video stream, the local unit video stream and the plate unit video stream, wherein the local unit video stream represents a video stream that has at most two unit types of a plate unit, an arm unit, a clothing unit and a head unit as one, and the overall unit video stream represents a video stream that has at least two unit types of a plate unit, an arm unit, a clothing unit and a head unit as one; a clothing segmentation module, which is used to receive the video data stream from the semantic segmentation sub-network of the main layer from the semantic segmentation sub-network; The main network receives the local unit video stream for clothing consistency matching to obtain a first compensated overall unit video stream and a clothing blind spot local unit video stream; the motion segmentation module is used to receive the clothing blind spot local unit video stream from the semantic segmentation main network through the secondary semantic segmentation sub-network of the main layer for motion consistency matching to obtain a second compensated overall unit video stream; the identity authentication module is used to receive the overall unit video stream, the first compensated overall unit video stream and the second compensated overall unit video stream from the semantic segmentation main network, the primary semantic segmentation sub-network and the secondary semantic segmentation sub-network through the identity authentication network of the main layer, receive the biometric data to be calculated from the perception layer to authenticate the identity of the plate unit video stream, and after obtaining the identity authentication plate unit video stream, receive the dish distribution data stream from the perception layer, associate it with the identity authentication plate unit video stream and the biometric data to be calculated for storage, construct an identity authentication document, and send it to the client.

[0008] The beneficial effects of the present invention are as follows: by collecting the video stream of the dish selection process and using advanced semantic segmentation technology to segment the video streams of the whole person, local units and plate units, and by matching the clothing consistency and motion consistency, the associated identity information is matched for each plate, thereby solving the problem of user identity being difficult to identify in traditional unmanned settlement methods, and achieving the technical effect of improving the accuracy of plate user identity identification. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] Figure 1 A schematic diagram of the process flow of the user identification method for canteen intelligent settlement provided by the present invention;

[0010] Figure 2 A schematic diagram of the structure of a user identification device for intelligent canteen settlement provided by the present invention.

[0011] In the accompanying drawings, the components represented by the reference numerals are described as follows:

[0012] Perception layer 001, main layer 002, client 003, camera perception node 0031, background configuration node 0032, data acquisition module 100, unit segmentation module 200, clothing segmentation module 300, motion segmentation module 400, identity authentication module 500. DETAILED DESCRIPTION

[0013] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.

[0014] In the description of the present invention, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. In the description of the present invention, the meaning of "plurality" is two or more, unless otherwise clearly and specifically defined.

[0015] In the description of the present invention, the term "for example" is used to mean "used as an example, illustration or explanation". Any embodiment described as "for example" in the present invention is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is given to enable any technician in the field to implement and use the present invention. In the following description, details are listed for the purpose of explanation. It should be understood that a person of ordinary skill in the art can recognize that the present invention can be implemented without using these specific details. In other examples, well-known structures and processes will not be elaborated in detail to avoid obscuring the description of the present invention with unnecessary details. Therefore, the present invention is not intended to be limited to the embodiments shown, but is consistent with the widest scope consistent with the principles and features disclosed in the present invention.

[0016] Embodiment 1:

[0017] like Figure 1 As shown, the embodiment of the present invention provides a user identification method for canteen intelligent settlement, which is applied to a user identification device for canteen intelligent settlement. The device includes a perception layer, a main layer and a client. The perception layer includes a camera perception node and a background configuration node, including the steps of:

[0018] S10: When receiving the biometric data to be calculated from the client, the perception layer receives the video data stream from the camera perception node and receives the dish distribution data stream from the background configuration node;

[0019] Specifically, the biometric data to be calculated refers to the user's authentication information, such as fingerprints, facial recognition data, etc., which can be used to determine the user's identity; the camera perception nodes refer to the cameras installed in the cafeteria, which are responsible for capturing video data when users choose dishes; the background configuration nodes refer to the server or computer system, which stores information about the distribution of dishes, including the location and identification of each dish; the video data stream and dish distribution data stream refer to the real-time stream of continuous video frame sequences and dish distribution information, which will be used for subsequent processing and analysis.

[0020] The workflow is as follows: First, the biometric data to be settled is received from the client for subsequent user identity verification. At the same time, the video data stream is received through the camera perception node. These video data streams contain dynamic information about the user's selection of dishes in the cafeteria. In addition, the dish distribution data stream is received from the background configuration node. These data streams contain the specific distribution information of the dishes, which is crucial for subsequent dish identification and settlement.

[0021] S20: receiving the video data stream from the perception layer through the semantic segmentation main network of the main layer to perform unit segmentation, and obtaining an overall unit video stream, a local unit video stream and a dinner plate unit video stream, wherein the local unit video stream represents a video stream having at most two unit types of a dinner plate unit, an arm unit, a clothing unit and a head unit as one, and the overall unit video stream represents a video stream having at least two unit types of a dinner plate unit, an arm unit, a clothing unit and a head unit as one;

[0022] Specifically, the semantic segmentation main network refers to a deep learning model that can understand the content in the video stream and segment it into different units or objects; unit segmentation refers to the process of separating different parts of the video stream (such as plates, arms, clothes, heads, etc.); the whole unit video stream, the local unit video stream and the plate unit video stream refer to different types of video streams obtained according to the segmentation results, which contain different parts of the video. The local unit video stream specifically refers to a video stream containing at most two of the plate unit, arm unit, clothing unit and head unit; while the whole unit video stream contains the plate unit and at least two other units (arm, clothing, head); the plate unit video stream represents the video stream of each plate.

[0023] In detail, the semantic segmentation main network of the main layer is used to process the video data stream received from the perception layer. The network analyzes and processes the video stream to identify and segment different units, such as dinner plates, arms, clothes, and heads. The specific process is as follows: the semantic segmentation main network first analyzes each frame in the video stream and identifies different objects and backgrounds in the frame; then the frames are segmented into different unit video streams according to different objects and backgrounds, and each unit video stream corresponds to a specific part of the video. For example, the dinner plate unit video stream only contains the image of the dinner plate, while the local unit video stream and the overall unit video stream may contain the dinner plate, the user's arm or clothes at the same time.

[0024] By segmenting the video stream into units, the system can more accurately track user behavior, identify the dishes selected by users, and ultimately achieve accurate settlement. This segmentation helps solve problems caused by crowd occlusion because it focuses on the video parts that are directly related to the user's identity and dish selection.

[0025] S30: receiving the local unit video stream from the semantic segmentation main network through the primary semantic segmentation sub-network of the main layer to perform clothing consistency matching, and obtaining a first compensated overall unit video stream and a clothing blind area local unit video stream;

[0026] Specifically, the first-level semantic segmentation sub-network is a deep learning network in the main layer used to process local unit video streams, focusing on clothing consistency matching; clothing consistency matching is to analyze the characteristics of clothing in the local unit video stream, identify and match the consistency of clothing in different video frames, to help determine the identity of the person in the video stream; the first compensation overall unit video stream refers to the video stream whose identity information is well matched with the plate unit video stream after clothing consistency matching. Clothing blind area local unit video streams refer to local unit video streams that fail to successfully match clothing features in clothing consistency matching. These areas may have unclear or missing clothing features due to occlusion or other factors.

[0027] The specific process is as follows: The first-level semantic segmentation sub-network receives local unit video streams from the semantic segmentation main network. These video streams contain at most two unit types: plate units, arm units, clothing units, and head units. The sub-network extracts key clothing features in the local unit video stream through the clothing feature extraction channel, and then uses the clothing feature comparison channel to match these features to identify the consistency of clothing in consecutive frames; for local unit video streams with high clothing feature matching, they are matched with the plate unit video stream to form the first compensated overall unit video stream, which helps to improve the accuracy of identity authentication. At the same time, the network will also identify areas where clothing features have low matching or cannot be matched, that is, clothing blind area local unit video streams, which may require additional processing or compensation mechanisms.

[0028] The role of this step in the whole scheme is to improve the accuracy of user identification, especially when clothing features are critical to identification. By matching clothing consistency, it is possible to better associate users with their selected plates, especially in crowded environments or restricted viewing angles.

[0029] S40: receiving the clothing blind area local unit video stream from the semantic segmentation main network through the secondary semantic segmentation sub-network of the main layer to perform motion consistency matching, and obtaining a second compensated overall unit video stream;

[0030] Specifically, the second-level semantic segmentation sub-network is another deep learning network in the main layer, which specializes in processing those local unit video streams of clothing blind spots that failed to pass clothing consistency matching in the first-level semantic segmentation sub-network; motion consistency matching refers to analyzing the motion trajectory and speed of objects in the video to identify and match the motion patterns in the video, which is crucial to solving the problem of unclear or obscured clothing features; the second compensatory overall unit video stream refers to the video stream whose motion features of the local unit video stream of clothing blind spots obtained after motion consistency matching are better matched with the dinner plate unit video stream, which compensates for the recognition error caused by the clothing blind spot problem.

[0031] The specific process is as follows: the secondary semantic segmentation sub-network receives the clothing blind area local unit video streams from the semantic segmentation main network; these video streams have not been successfully matched in the clothing consistency matching before, probably because the clothing features are not obvious or are blocked; the secondary network extracts the motion trajectory information in the video through the motion feature extraction channel, and then uses the motion trajectory comparison channel to match these motion features to identify the motion pattern in the video; for the video streams with high motion feature matching, they are matched with the dinner plate unit video stream to form a second compensated overall unit video stream, which helps to improve the accuracy of identity authentication.

[0032] The role of this step in the entire solution is to solve the identity authentication problem caused by unclear or obscured clothing features. Through motion consistency matching, it is possible to better associate users with their selected plates, especially in complex environments.

[0033] S50: Receive the whole unit video stream, the first compensated whole unit video stream and the second compensated whole unit video stream from the semantic segmentation main network, the first-level semantic segmentation sub-network and the second-level semantic segmentation sub-network through the identity authentication network of the main layer, receive the biometric data to be settled from the perception layer to perform identity authentication on the plate unit video stream, and after obtaining the identity authentication plate unit video stream, receive the dish distribution data stream from the perception layer, associate it with the identity authentication plate unit video stream and the biometric data to be settled, and construct an identity authentication document, which is sent to the client.

[0034] Specifically, the identity authentication network is a key component in the main layer. It is responsible for integrating the video stream data from the semantic segmentation main network, the first-level semantic segmentation sub-network and the second-level semantic segmentation sub-network, as well as the biometric data to be settled received from the perception layer, and performing comprehensive analysis to achieve identity authentication of the plate unit video stream; the biometric data to be settled refers to the user's biometric data, such as fingerprint or facial recognition information, which is used to match the user identity identified in the video; the identity authentication plate unit video stream refers to the plate video stream associated with a specific user identity after being processed by the identity authentication network, that is, each plate video stream has a matching identity information; the dish distribution data stream is a data stream received from the background configuration node of the perception layer, containing information such as the location and type of the dish; the identity authentication document is a document formed by integrating the identity authentication plate unit video stream, the biometric data to be settled and the dish distribution data stream, which is used for final settlement and recording, and then sent to the client.

[0035] The specific process is as follows: the identity authentication network first receives the whole unit video stream from the semantic segmentation main network, the first compensated whole unit video stream from the first-level semantic segmentation sub-network, and the second compensated whole unit video stream from the second-level semantic segmentation sub-network; then, the identity authentication network uses the biometric data to be calculated provided by the perception layer to perform identity matching on the user identified in the plate unit video stream. The biometric data to be calculated is the identity information entered by the user at the settlement terminal, and the settlement terminal also belongs to a frame image of the video stream. Therefore, according to the whole unit video stream, the first compensated whole unit video stream, and the second compensated whole unit video stream, the plate unit video stream associated with the entered user can be determined, and then the biometric data to be calculated is matched with the plate unit video stream to obtain the identity authentication plate unit video stream; finally, the identity authentication network receives the dish distribution data stream from the perception layer, associates and stores the identity authentication plate unit video stream, the biometric data to be calculated, and the dish distribution data stream, to construct a complete identity authentication document. This document contains the user's identity information and detailed information of the selected dishes. After this document is sent to the client, the client can calculate the dishes selected for each user's plate based on the identity authentication document, and then settle the account based on the selected dishes.

[0036] Since dishes can be identified in the video stream of the dish selection process, the problem that the terminal directly calculates the dishes as mixed and cannot be identified is solved. Furthermore, since the application constructs an identity authentication document that can effectively identify the user identity information of the plate, the problem of the plate being unable to be identified due to occlusion by the crowd is solved.

[0037] Furthermore, the perception layer further includes a weight perception node, which corresponds to the dish display cavity one by one and is deployed at the bottom of the dish display cavity, and further includes:

[0038] Receive the weight data stream of the dish display cavity through the weight sensing node of the sensing layer;

[0039] Performing time alignment on the identity authentication plate unit video stream and the dish distribution data stream through the main layer to obtain a selected dish data stream, wherein the selected dish data stream has a selected display cavity identification data stream;

[0040] Through the main layer, based on the display cavity identification data stream and the dish display cavity, the selected dish data stream and the weight data stream are aligned in timing and cavity to obtain a reduced weight data stream;

[0041] The weight reduction data stream and the dish distribution data stream are associated with the identity authentication plate unit video stream and the biometric data to be calculated and stored to construct an identity authentication document, which is then sent to the client.

[0042] Specifically, the weight sensing node refers to the sensor installed at the bottom of each dish display cavity, which is used to monitor and record the weight changes of the dishes in the cavity in real time, and track the dishes selected by the user and their quantity through the weight changes; the weight data stream refers to the data sequence about the weight changes of the dishes collected by the weight sensing node; the timing alignment refers to matching the data in different data streams according to the time sequence to ensure the synchronization of the data; the dish selection data stream refers to the data stream containing the information of the dishes selected by the user, and the display cavity selection identification data stream is the data stream that identifies the cavity corresponding to the dish selected by the user; the weight reduction data stream refers to the data stream that, after processing, reflects the weight reduction caused by the user's selection of dishes.

[0043] The specific process is as follows: the weight data stream from each dish display cavity is received through the weight perception node of the perception layer; the main layer then performs timing alignment on the identity authentication plate unit video stream and the dish distribution data stream to ensure that the dish selection event in the video stream matches the actual dish distribution data, and generates a selected dish data stream; then, the main layer performs further timing and cavity alignment on the selected dish data stream and the weight data stream based on the selected display cavity identification data stream and the corresponding dish display cavity to obtain a reduced weight data stream, which reflects the actual reduction in cavity weight after the user selects a dish; finally, the reduced weight data stream and the dish distribution data stream are associated with the identity authentication plate unit video stream and the biometric data to be settled and stored to construct a complete identity authentication document, which contains the user identity, the selected dish and its weight information, and then sent to the client for subsequent processing.

[0044] Furthermore, the video data stream is received from the perception layer through the semantic segmentation main network of the main layer for unit segmentation to obtain the overall unit video stream, the local unit video stream and the plate unit video stream, including:

[0045] Through the segmentation channel of the semantic segmentation main network, the video data stream is segmented to obtain the plate unit video stream, the arm unit video stream, the clothing unit video stream and the head unit video stream;

[0046] Through the occlusion identification channel of the semantic segmentation main network, the video data stream is subjected to occlusion identification to obtain an occlusion identification area;

[0047] The dinner plate unit video stream, the arm unit video stream, the clothing unit video stream and the head unit video stream are processed through the feature extraction channel of the semantic segmentation main network to extract posture features, shape features and distribution features;

[0048] The plate unit video stream, the arm unit video stream, the clothing unit video stream, the head unit video stream, the occlusion mark area, the posture features, the shape features and the distribution features are processed through the aggregation channel of the semantic segmentation main network to output the overall unit video stream and the local unit video stream.

[0049] Specifically, the main semantic segmentation network is a deep learning model responsible for in-depth analysis and processing of video data streams. It performs specific tasks through different channels; the segmentation channel refers to the functional layer in the network used to identify and separate different targets in the video (such as plates, arms, clothing, and heads); the occlusion identification channel is used to identify occluded areas in the video, which is crucial for subsequent analysis because the occluded areas may contain important information; the feature extraction channel is responsible for extracting key features from the identified units, such as posture, shape, and distribution features, which help further analysis and identification; the aggregation channel is responsible for integrating all the extracted information and features to output the final overall unit video stream and local unit video stream.

[0050] The specific process is as follows: First, the semantic segmentation main network identifies and segments the targets in the video data stream through the segmentation channel, and obtains the plate unit video stream, arm unit video stream, clothing unit video stream and head unit video stream respectively. Next, the occlusion identification channel identifies the occluded areas in the video and marks these areas. Then, the feature extraction channel processes the segmented unit video stream to extract posture features, shape features and distribution features. Finally, the aggregation channel combines all this information and outputs the overall unit video stream and the local unit video stream, providing basic data for subsequent identity authentication and dish recognition.

[0051] Furthermore, the semantic segmentation main network construction step includes:

[0052] Construct the segmentation channel loss function:

[0053] ,

[0054] in, Characterize the segmentation channel loss value, Represents the i-th unit supervision area of ​​the j-th frame video image, Represents the i-th unit prediction area of ​​the j-th frame video image, represents a small constant, N represents the number of unit supervision regions of the j-th frame video image, Characterizes the number of video stream frames;

[0055] Construct the occlusion identification channel loss function:

[0056] ,

[0057] in, Characterizes the loss value of the occlusion identification channel, The true label that represents whether the kth pixel of the jth frame video image is blocked. Representation is occluded, when The representation is not blocked, Characterizes the probability that the kth pixel of the jth frame video image predicted by the occlusion identification channel is occluded, The total number of pixels representing the j-th frame of the video image;

[0058] Construct the aggregate channel loss function:

[0059] ,

[0060] in, Characterizes the aggregate channel loss value, Unit grouping supervision data representing the j-th frame of the video image, Unit grouping aggregation data representing the j-th frame of the video image, Represents the intersection symbol, Represents the union symbol, Characterizes natural constants;

[0061] Construct the semantic segmentation main network loss function:

[0062] ,

[0063] in, Represents the loss value of the main network for semantic segmentation, , and Characterize loss weights;

[0064] The semantic segmentation main network is trained according to the segmentation channel loss function, the occlusion identification channel loss function, the aggregation channel loss function and the semantic segmentation main network loss function.

[0065] Specifically, first, according to the segmentation channel loss function: ,Training the segmentation channel, the training method is as follows: collect the video data stream record dataset and the segmentation unit supervision area dataset, set it as the segmentation channel construction dataset; divide the segmentation channel construction dataset according to 8:2 to generate the segmentation channel training dataset and the segmentation channel verification dataset; based on the segmentation channel loss function, call the segmentation channel training dataset to perform supervised training on the convolutional neural network. When the segmentation channel loss function is trained for 500 times in a row and the loss is less than or equal to the segmentation channel loss threshold for at least 480 times, use the segmentation channel verification dataset to perform supervised training on the convolutional neural network. When the segmentation channel loss function is trained for 100 times in a row and the loss is less than or equal to the segmentation channel loss threshold for at least 98 times, the segmentation channel is considered to have converged separately.

[0066] Using the same process, we identify the channel loss function based on occlusion: , train the occlusion identification channel; according to the aggregate channel loss function: , train the aggregation channel.

[0067] Furthermore, when the segmentation channel, occlusion identification channel, and aggregation channel have all converged individually, the loss function of the main semantic segmentation network is: , perform overall training. When the loss function of the semantic segmentation main network is less than or equal to the loss threshold of the semantic segmentation main network for at least 480 consecutive times out of 500 trainings, supervised training is performed using the validation dataset. When the loss function of the semantic segmentation main network is less than or equal to the loss threshold of the semantic segmentation main network for at least 98 consecutive times out of 100 trainings, the semantic segmentation main network is considered to have converged.

[0068] Furthermore, the first-level semantic segmentation sub-network of the main layer receives the local unit video stream from the semantic segmentation main network for clothing consistency matching to obtain a first compensated overall unit video stream and a clothing blind area local unit video stream, including:

[0069] Extracting a first frame local unit of the local unit video stream;

[0070] Extracting a first group of local units and a second group of local units whose spatial distribution distance is less than or equal to a consistent distance threshold from the first frame local units, wherein the first group of local units and the second group of local units do not have an intersection unit type;

[0071] Extracting first clothing features of the first group of local units according to the clothing feature extraction channel of the first-level semantic segmentation sub-network, and extracting second clothing features of the second group of local units;

[0072] Comparing the first clothing feature and the second clothing feature according to the clothing feature comparison channel of the first-level semantic segmentation sub-network to obtain clothing similarity;

[0073] When the clothing similarity is greater than or equal to the clothing similarity threshold, the first group of local units and the second group of local units are merged into the same group of units; when the clothing similarity is less than the clothing similarity threshold, no operation is performed;

[0074] Cyclic analysis is performed. When the clothing similarity of any two groups of local units is less than the clothing similarity threshold, a video stream that has at least two unit types, namely, a dinner plate unit, an arm unit, a clothing unit and a head unit, is added to the first compensated overall unit video stream. Otherwise, it is added to the clothing blind spot local unit video stream.

[0075] Specifically, the first-level semantic segmentation sub-network is a network in the main layer, which is specifically responsible for processing local unit video streams and performing clothing consistency matching; the local unit video stream refers to a video stream containing at most two unit types among the plate unit, arm unit, clothing unit and head unit; the clothing feature extraction channel and the clothing feature comparison channel are two key functional layers in the first-level semantic segmentation sub-network, which are used to extract clothing features and compare these features to determine the consistency of clothing, respectively, and both can be constructed based on convolutional neural networks; clothing similarity refers to the similarity score obtained by comparing the clothing features of two groups of local units; the clothing similarity threshold is a preset value used to determine whether two groups of clothing features are similar enough.

[0076] The specific process is as follows: The first-level semantic segmentation sub-network first receives the local unit video stream from the semantic segmentation main network. Then, the network extracts the local unit of the first frame of the local unit video stream, and identifies the first and second groups of local units whose spatial distribution distance is less than or equal to the consistent distance threshold preset by the manager from the first frame. These two groups of local units do not share any unit type. Next, the network extracts the clothing features of these two groups of local units respectively through the clothing feature extraction channel. Preferably, the clothing features include but are not limited to color features, texture features, shape features, etc., and uses the clothing feature comparison channel to compare these features to obtain clothing similarity:

[0077] If the clothing similarity is greater than or equal to a preset clothing similarity threshold, the clothing of the two groups of local units are considered to be consistent, and they are merged into the same group of units;

[0078] If the similarity is below the threshold, no merging is done.

[0079] This process is repeated until all local units have been analyzed. Finally, based on the results of clothing similarity, the video stream is classified into the first compensated overall unit video stream (clothing consistency matching is successful) or the clothing blind spot local unit video stream (clothing consistency matching fails).

[0080] Further, the secondary semantic segmentation sub-network of the main layer receives the clothing blind area local unit video stream from the semantic segmentation main network for motion consistency matching to obtain a second compensated overall unit video stream, including:

[0081] Extracting a first group of clothing blind spot local unit video streams and a second group of clothing blind spot local unit video streams of the clothing blind spot local unit video stream, wherein any frame spatial distribution distance between the first group of clothing blind spot local units and the second group of clothing blind spot local units is less than or equal to a consistent distance threshold, and does not have an intersection unit type;

[0082] Extracting first motion trajectory information of the first group of clothing blind area local unit video streams through the motion feature extraction channel of the secondary semantic segmentation sub-network, and extracting second motion trajectory information of the second group of clothing blind area local unit video streams;

[0083] Performing trajectory similarity evaluation on the first motion trajectory information and the second motion trajectory information through the motion trajectory comparison channel of the secondary semantic segmentation sub-network to obtain trajectory similarity;

[0084] When the trajectory similarity is greater than or equal to the trajectory similarity threshold, the first group of clothing blind area local unit video streams and the second group of clothing blind area local unit video streams are merged into the same group of unit video streams; when the trajectory similarity is less than the trajectory similarity threshold, no operation is performed;

[0085] Cyclic analysis is performed, and when the trajectory similarity of any two groups of local unit video streams is less than the trajectory similarity threshold, a video stream having at least two unit types of a dinner plate unit, an arm unit, a clothing unit and a head unit is added to the second compensated overall unit video stream.

[0086] Specifically, the second-level semantic segmentation sub-network is a network in the main layer, which is specially used to process those local unit video streams in clothing blind spots that failed to pass clothing consistency matching in the first-level semantic segmentation sub-network; local unit video streams in clothing blind spots refer to those video streams whose clothing features are not obvious or are blocked, and therefore need to be identified and matched by other methods (such as motion consistency matching); the motion feature extraction channel and the motion trajectory comparison channel are two key parts in the second-level semantic segmentation sub-network, which are used to extract motion trajectory information in the video and compare these trajectory information to determine the consistency of motion, respectively, and can be constructed through machine learning models such as neural networks and support vector machines; trajectory similarity refers to the similarity score obtained by comparing the motion trajectory information of two groups of local units; the trajectory similarity threshold is a preset value used to determine whether two sets of motion trajectories are similar enough.

[0087] The specific process is as follows: the second-level semantic segmentation sub-network first receives the clothing blind spot local unit video stream from the semantic segmentation main network, and extracts the first and second groups of clothing blind spot local unit video streams from them. The spatial distribution distance of any frame of these two groups of video streams is less than or equal to the consistent distance threshold pre-set by the manager, and they do not share any unit type; then, the network extracts the motion trajectory information of these two groups of local unit video streams respectively through the motion feature extraction channel, and uses the motion trajectory comparison channel to compare these trajectory information to obtain the trajectory similarity.

[0088] If the trajectory similarity is greater than or equal to a preset trajectory similarity threshold, the motion trajectories of the two groups of local unit video streams are considered to be consistent, and they are merged into the same group of unit video streams.

[0089] If the similarity is below the threshold, no merging is done.

[0090] This process is repeated until all local unit video streams have been analyzed. Finally, based on the results of trajectory similarity, the video stream is classified as the second compensated overall unit video stream (motion consistency matching is successful) or continues to be a clothing blind spot local unit video stream (motion consistency matching fails).

[0091] Furthermore, when the trajectory similarity of any two groups of local unit video streams is less than the trajectory similarity threshold, a video stream having at least two unit types of a dinner plate unit, an arm unit, a clothing unit, and a head unit as one is added to the second compensated overall unit video stream, further comprising:

[0092] When a video stream having a dinner plate unit and at least two unit types of an arm unit, a clothing unit and a head unit as one is added to the second compensated overall unit video stream, a residual local unit video stream is obtained, wherein the residual local unit video stream includes a local dinner plate unit video stream;

[0093] When the number of the local plate unit video streams is equal to 0, outputting the second compensated overall unit video stream;

[0094] When the number of the local dinner plate unit video streams is not equal to 0, trajectory unit hierarchical clustering is performed on each of the local dinner plate unit video streams until a video stream having at least two unit types of dinner plate unit, arm unit, clothing unit and head unit is found, and the video stream is added into the second compensated overall unit video stream.

[0095] Specifically, the trajectory similarity threshold refers to a preset value used to determine whether the motion trajectories of two groups of local unit video streams are similar enough to determine whether to merge them into the second compensated overall unit video stream; the residual local unit video streams refer to those local unit video streams that have not been merged with other video streams after motion consistency matching. They may be left alone because their trajectory similarity is lower than the threshold; the local plate unit video stream refers to the video stream containing the plate unit in the residual local unit video stream. These video streams need further processing to determine their association with the user identity.

[0096] The specific process is as follows: when the trajectory similarity of any two groups of local unit video streams is lower than the trajectory similarity threshold, those video streams that contain both the dinner plate unit and at least two other unit types (arm unit, clothing unit and head unit) are added to the second compensated overall unit video stream; after adding, check the remaining local unit video streams, especially those containing the local dinner plate unit video streams.

[0097] If the number of partial tray unit video streams is zero, that is, there are no remaining partial tray unit video streams to be processed, the constructed second compensated overall unit video stream will be output.

[0098] If the number of local plate unit video streams is non-zero, these local plate unit video streams are subjected to trajectory unit hierarchical clustering, which is a method of grouping unit video streams with similar motion trajectories. The clustering process continues until video streams containing both plate units and at least two other unit types are found, and then these video streams are added to the second compensated overall unit video stream. This process helps to further improve the accuracy of identification and ensure that all related video streams can be correctly associated and identified.

[0099] Further, performing trajectory unit hierarchical clustering on each of the local dinner plate unit video streams until a video stream having a dinner plate unit and at least two unit types of an arm unit, a clothing unit, and a head unit is obtained, comprising:

[0100] Extracting a first partial dining tray unit motion trajectory of a first partial dining tray unit video stream of the partial dining tray unit video stream;

[0101] Extracting a first residual local unit motion trajectory of a first residual local unit video stream from the residual local unit video stream, wherein the first residual local unit video stream does not have a dinner plate unit, and a spatial distribution distance between a first residual local unit and a first local dinner plate unit in any frame of the first residual local unit video stream is less than or equal to a sorting distance threshold, and the sorting distance threshold represents a maximum distribution distance that may belong to one body;

[0102] until an Nth residual local unit motion trajectory of the Nth residual local unit video stream is extracted from the residual local unit video stream, wherein the Nth residual local unit video stream does not have a dinner plate unit, and a spatial distribution distance between the Nth residual local unit and the first local dinner plate unit of any frame of the Nth residual local unit video stream is less than or equal to a sorting distance threshold, and N represents the number of all residual local unit video streams that meet the conditions of distribution distance and unit composition type;

[0103] Extracting the residual local unit video stream having the maximum trajectory similarity with the first local plate unit motion trajectory from the first residual local unit motion trajectory to the Nth residual local unit motion trajectory, and merging it with the first local plate unit video stream to generate an updated video stream;

[0104] When the update video stream has a dinner plate unit and at least two unit types of an arm unit, a clothing unit and a head unit as one, the update video stream is set as a second compensated whole unit video stream of the first local dinner plate unit video stream;

[0105] Otherwise, the trajectory unit hierarchical clustering is continued based on the updated video stream until a video stream having at least two unit types of a dinner plate unit, an arm unit, a clothing unit, and a head unit is found.

[0106] Specifically, trajectory unit hierarchical clustering is an algorithmic process that groups different video stream units together according to the similarity of motion trajectories; the first local plate unit video stream refers to the video stream that contains the plate unit in the residual local unit video stream; the first residual local unit video refers to other types of local unit video streams except the plate unit; the sorting distance threshold is a preset maximum distance value used to determine which unit video streams are close enough in spatial distribution and may belong to the same person; the maximum trajectory similarity value refers to the trajectory that is most similar to the first local plate unit motion trajectory among all compared trajectories.

[0107] The specific process is as follows: first, the motion trajectory of the first local plate unit in the local plate unit video stream is extracted; then, the motion trajectory of the first residual local unit that does not contain the plate unit but is close to the first local plate unit in spatial distribution (the distance is less than or equal to the sorting distance threshold) is extracted from the residual local unit video stream; this process will continue until all qualified residual local unit video streams are extracted; then, the residual local unit video stream with the highest similarity to the motion trajectory of the first local plate unit is found, and it is merged with the first local plate unit video stream to generate an updated video stream and then the process is executed based on the following two situations:

[0108] Process 1: If the updated video stream contains both the dinner plate unit and at least two other unit types (arm unit, clothing unit and head unit), set it as the second compensated whole unit video stream of the first local dinner plate unit video stream;

[0109] Process 2: If this condition is not met, trajectory unit hierarchical clustering will continue based on the updated video stream until a video stream that meets the condition is found. This process helps to improve the accuracy of identity authentication and ensures that all related video streams can be correctly associated and identified.

[0110] The user identification method for canteen intelligent settlement provided by the embodiment of the present invention has at least the following technical effects:

[0111] By collecting the video stream of the dish selection process and using advanced semantic segmentation technology, we can segment the video streams of the whole person, local units and plate units, and match the associated identity information for each plate through clothing consistency and motion consistency matching, thus solving the problem of user identity difficulty in traditional unmanned settlement methods and achieving the technical effect of improving the accuracy of plate user identity identification.

[0112] Embodiment 2:

[0113] like Figure 2 As shown, based on the same inventive concept as the user identification method for canteen intelligent settlement provided in Embodiment 1, the embodiment of the present invention also provides a user identification device for canteen intelligent settlement, including a perception layer 001, a main layer 002 and a client 003, wherein the perception layer includes a camera perception node 0031 and a background configuration node 0032, including:

[0114] The data acquisition module 100 is used to receive the biometric data to be calculated from the client, receive the video data stream from the camera perception node through the perception layer, and receive the dish distribution data stream from the background configuration node;

[0115] The unit segmentation module 200 is used to receive the video data stream from the perception layer through the semantic segmentation main network of the main layer to perform unit segmentation, and obtain an overall unit video stream, a local unit video stream and a plate unit video stream, wherein the local unit video stream represents a video stream having at most two unit types of a plate unit, an arm unit, a clothing unit and a head unit as one, and the overall unit video stream represents a video stream having at least two unit types of a plate unit, an arm unit, a clothing unit and a head unit as one;

[0116] A clothing segmentation module 300 is used to receive the local unit video stream from the semantic segmentation main network through the primary semantic segmentation sub-network of the main layer to perform clothing consistency matching, and obtain a first compensated overall unit video stream and a clothing blind area local unit video stream;

[0117] A motion segmentation module 400 is used to receive the clothing blind area local unit video stream from the semantic segmentation main network through the secondary semantic segmentation sub-network of the main layer to perform motion consistency matching to obtain a second compensated overall unit video stream;

[0118] The identity authentication module 500 is used to receive the whole unit video stream, the first compensated whole unit video stream and the second compensated whole unit video stream from the semantic segmentation main network, the first-level semantic segmentation sub-network and the second-level semantic segmentation sub-network through the identity authentication network of the main layer, receive the biometric data to be settled from the perception layer to perform identity authentication on the plate unit video stream, and after obtaining the identity authentication plate unit video stream, receive the dish distribution data stream from the perception layer, associate it with the identity authentication plate unit video stream and the biometric data to be settled, and construct an identity authentication document, which is sent to the client.

[0119] Furthermore, the perception layer further includes a weight perception node, which corresponds to the dish display cavity one by one and is deployed at the bottom of the dish display cavity, and further includes:

[0120] Receive the weight data stream of the dish display cavity through the weight sensing node of the sensing layer;

[0121] Performing time alignment on the identity authentication plate unit video stream and the dish distribution data stream through the main layer to obtain a selected dish data stream, wherein the selected dish data stream has a selected display cavity identification data stream;

[0122] Through the main layer, based on the display cavity identification data stream and the dish display cavity, the selected dish data stream and the weight data stream are aligned in timing and cavity to obtain a reduced weight data stream;

[0123] The weight reduction data stream and the dish distribution data stream are associated with the identity authentication plate unit video stream and the biometric data to be calculated and stored to construct an identity authentication document, which is then sent to the client.

[0124] Furthermore, the video data stream is received from the perception layer through the semantic segmentation main network of the main layer for unit segmentation to obtain the overall unit video stream, the local unit video stream and the plate unit video stream, including:

[0125] Through the segmentation channel of the semantic segmentation main network, the video data stream is segmented to obtain the plate unit video stream, the arm unit video stream, the clothing unit video stream and the head unit video stream;

[0126] Through the occlusion identification channel of the semantic segmentation main network, the video data stream is subjected to occlusion identification to obtain an occlusion identification area;

[0127] The dinner plate unit video stream, the arm unit video stream, the clothing unit video stream and the head unit video stream are processed through the feature extraction channel of the semantic segmentation main network to extract posture features, shape features and distribution features;

[0128] The plate unit video stream, the arm unit video stream, the clothing unit video stream, the head unit video stream, the occlusion mark area, the posture features, the shape features and the distribution features are processed through the aggregation channel of the semantic segmentation main network to output the overall unit video stream and the local unit video stream.

[0129] Furthermore, the semantic segmentation main network construction step includes:

[0130] Construct the segmentation channel loss function:

[0131] ,

[0132] in, Characterize the segmentation channel loss value, Represents the i-th unit supervision area of ​​the j-th frame video image, Represents the i-th unit prediction area of ​​the j-th frame video image, represents a small constant, N represents the number of unit supervision regions of the j-th frame video image, Characterizes the number of video stream frames;

[0133] Construct the occlusion identification channel loss function:

[0134] ,

[0135] in, Characterizes the loss value of the occlusion identification channel, The true label that represents whether the kth pixel of the jth frame video image is blocked. Representation is occluded, when The representation is not blocked, Characterizes the probability that the kth pixel of the jth frame video image predicted by the occlusion identification channel is occluded, The total number of pixels representing the j-th frame of the video image;

[0136] Construct the aggregate channel loss function:

[0137] ,

[0138] in, Characterizes the aggregate channel loss value, Unit grouping supervision data representing the j-th frame of the video image, Unit grouping aggregation data representing the j-th frame of the video image, Represents the intersection symbol, Represents the union symbol, Characterizes natural constants;

[0139] Construct the semantic segmentation main network loss function:

[0140] ,

[0141] in, Represents the loss value of the main network for semantic segmentation, , and Characterize loss weights;

[0142] The semantic segmentation main network is trained according to the segmentation channel loss function, the occlusion identification channel loss function, the aggregation channel loss function and the semantic segmentation main network loss function.

[0143] Furthermore, the first-level semantic segmentation sub-network of the main layer receives the local unit video stream from the semantic segmentation main network for clothing consistency matching to obtain a first compensated overall unit video stream and a clothing blind area local unit video stream, including:

[0144] Extracting a first frame local unit of the local unit video stream;

[0145] Extracting a first group of local units and a second group of local units whose spatial distribution distance is less than or equal to a consistent distance threshold from the first frame local units, wherein the first group of local units and the second group of local units do not have an intersection unit type;

[0146] Extracting first clothing features of the first group of local units according to the clothing feature extraction channel of the first-level semantic segmentation sub-network, and extracting second clothing features of the second group of local units;

[0147] Comparing the first clothing feature and the second clothing feature according to the clothing feature comparison channel of the first-level semantic segmentation sub-network to obtain clothing similarity;

[0148] When the clothing similarity is greater than or equal to the clothing similarity threshold, the first group of local units and the second group of local units are merged into the same group of units; when the clothing similarity is less than the clothing similarity threshold, no operation is performed;

[0149] Cyclic analysis is performed. When the clothing similarity of any two groups of local units is less than the clothing similarity threshold, a video stream that has at least two unit types, namely, a dinner plate unit, an arm unit, a clothing unit and a head unit, is added to the first compensated overall unit video stream. Otherwise, it is added to the clothing blind spot local unit video stream.

[0150] Further, the secondary semantic segmentation sub-network of the main layer receives the clothing blind area local unit video stream from the semantic segmentation main network for motion consistency matching to obtain a second compensated overall unit video stream, including:

[0151] Extracting a first group of clothing blind spot local unit video streams and a second group of clothing blind spot local unit video streams of the clothing blind spot local unit video stream, wherein any frame spatial distribution distance between the first group of clothing blind spot local units and the second group of clothing blind spot local units is less than or equal to a consistent distance threshold, and does not have an intersection unit type;

[0152] Extracting first motion trajectory information of the first group of clothing blind area local unit video streams through the motion feature extraction channel of the secondary semantic segmentation sub-network, and extracting second motion trajectory information of the second group of clothing blind area local unit video streams;

[0153] Performing trajectory similarity evaluation on the first motion trajectory information and the second motion trajectory information through the motion trajectory comparison channel of the secondary semantic segmentation sub-network to obtain trajectory similarity;

[0154] When the trajectory similarity is greater than or equal to the trajectory similarity threshold, the first group of clothing blind area local unit video streams and the second group of clothing blind area local unit video streams are merged into the same group of unit video streams; when the trajectory similarity is less than the trajectory similarity threshold, no operation is performed;

[0155] Cyclic analysis is performed, and when the trajectory similarity of any two groups of local unit video streams is less than the trajectory similarity threshold, a video stream having at least two unit types of a dinner plate unit, an arm unit, a clothing unit and a head unit is added to the second compensated overall unit video stream.

[0156] Furthermore, when the trajectory similarity of any two groups of local unit video streams is less than the trajectory similarity threshold, a video stream having at least two unit types of a dinner plate unit, an arm unit, a clothing unit, and a head unit as one is added to the second compensated overall unit video stream, further comprising:

[0157] When a video stream having a dinner plate unit and at least two unit types of an arm unit, a clothing unit and a head unit as one is added to the second compensated overall unit video stream, a residual local unit video stream is obtained, wherein the residual local unit video stream includes a local dinner plate unit video stream;

[0158] When the number of the local plate unit video streams is equal to 0, outputting the second compensated overall unit video stream;

[0159] When the number of the local dinner plate unit video streams is not equal to 0, trajectory unit hierarchical clustering is performed on each of the local dinner plate unit video streams until a video stream having at least two unit types of dinner plate unit, arm unit, clothing unit and head unit is found, and the video stream is added into the second compensated overall unit video stream.

[0160] Further, performing trajectory unit hierarchical clustering on each of the local dinner plate unit video streams until a video stream having a dinner plate unit and at least two unit types of an arm unit, a clothing unit, and a head unit is obtained, comprising:

[0161] Extracting a first partial dining tray unit motion trajectory of a first partial dining tray unit video stream of the partial dining tray unit video stream;

[0162] Extracting a first residual local unit motion trajectory of a first residual local unit video stream from the residual local unit video stream, wherein the first residual local unit video stream does not have a dinner plate unit, and a spatial distribution distance between a first residual local unit and a first local dinner plate unit in any frame of the first residual local unit video stream is less than or equal to a sorting distance threshold, and the sorting distance threshold represents a maximum distribution distance that may belong to one body;

[0163] until an Nth residual local unit motion trajectory of the Nth residual local unit video stream is extracted from the residual local unit video stream, wherein the Nth residual local unit video stream does not have a dinner plate unit, and a spatial distribution distance between the Nth residual local unit and the first local dinner plate unit of any frame of the Nth residual local unit video stream is less than or equal to a sorting distance threshold, and N represents the number of all residual local unit video streams that meet the conditions of distribution distance and unit composition type;

[0164] Extracting the residual local unit video stream having the maximum trajectory similarity with the first local plate unit motion trajectory from the first residual local unit motion trajectory to the Nth residual local unit motion trajectory, and merging it with the first local plate unit video stream to generate an updated video stream;

[0165] When the update video stream has a dinner plate unit and at least two unit types of an arm unit, a clothing unit and a head unit as one, the update video stream is set as a second compensated whole unit video stream of the first local dinner plate unit video stream;

[0166] Otherwise, the trajectory unit hierarchical clustering is continued based on the updated video stream until a video stream having at least two unit types of a dinner plate unit, an arm unit, a clothing unit, and a head unit is found.

[0167] It should be noted that in the above embodiments, the description of each embodiment has its own emphasis, and for parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0168] It should be understood by those skilled in the art that embodiments of the present invention may be provided as methods, devices, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0169] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (apparatus), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded computer, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0170] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0171] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0172] Although preferred embodiments of the present invention have been described, additional changes and modifications may occur to these embodiments once those skilled in the art understand the basic inventive concepts.

[0173] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention belong to the scope of the present invention and its equivalent technologies, the present invention is also intended to include these changes and variations.

Claims

1. A user identification method for canteen intelligent settlement, characterized in that: A user identification device applied to canteen intelligent settlement, the device includes a perception layer, a main layer and a client, the perception layer includes a camera perception node and a background configuration node, including: When receiving the biometric data to be calculated from the client, the perception layer receives the video data stream from the camera perception node and the dish distribution data stream from the background configuration node; The video data stream is received from the perception layer through the semantic segmentation main network of the main layer for unit segmentation to obtain an overall unit video stream, a local unit video stream and a dinner plate unit video stream, wherein the local unit video stream represents a video stream having at most two unit types of a dinner plate unit, an arm unit, a clothing unit and a head unit as one, and the overall unit video stream represents a video stream having at least two unit types of a dinner plate unit, an arm unit, a clothing unit and a head unit as one; The first-level semantic segmentation sub-network of the main layer receives the local unit video stream from the semantic segmentation main network for clothing consistency matching, and obtains a first compensated overall unit video stream and a clothing blind area local unit video stream; The secondary semantic segmentation sub-network of the main layer receives the clothing blind area local unit video stream from the semantic segmentation main network for motion consistency matching to obtain a second compensated overall unit video stream; The overall unit video stream, the first compensated overall unit video stream and the second compensated overall unit video stream are received from the semantic segmentation main network, the first-level semantic segmentation sub-network and the second-level semantic segmentation sub-network through the identity authentication network of the main layer, and the biometric data to be calculated is received from the perception layer to perform identity authentication on the plate unit video stream. After obtaining the identity authentication plate unit video stream, the dish distribution data stream is received from the perception layer, associated with the identity authentication plate unit video stream and the biometric data to be calculated for storage, and an identity authentication document is constructed and sent to the client.

2. The method according to claim 1, characterized in that The perception layer further includes a weight perception node, which corresponds to the dish display cavity one by one and is deployed at the bottom of the dish display cavity, and further includes: Receive the weight data stream of the dish display cavity through the weight sensing node of the sensing layer; Performing time alignment on the identity authentication plate unit video stream and the dish distribution data stream through the main layer to obtain a selected dish data stream, wherein the selected dish data stream has a selected display cavity identification data stream; Through the main layer, based on the display cavity identification data stream and the dish display cavity, the selected dish data stream and the weight data stream are aligned in timing and cavity to obtain a reduced weight data stream; The weight reduction data stream and the dish distribution data stream are associated with the identity authentication plate unit video stream and the biometric data to be calculated and stored to construct an identity authentication document, which is then sent to the client.

3. The method according to claim 1, characterized in that The video data stream is received from the perception layer through the semantic segmentation main network of the main layer for unit segmentation to obtain the overall unit video stream, the local unit video stream and the plate unit video stream, including: Through the segmentation channel of the semantic segmentation main network, the video data stream is segmented to obtain the plate unit video stream, the arm unit video stream, the clothing unit video stream and the head unit video stream; Through the occlusion identification channel of the semantic segmentation main network, the video data stream is subjected to occlusion identification to obtain an occlusion identification area; The dinner plate unit video stream, the arm unit video stream, the clothing unit video stream and the head unit video stream are processed through the feature extraction channel of the semantic segmentation main network to extract posture features, shape features and distribution features; The plate unit video stream, the arm unit video stream, the clothing unit video stream, the head unit video stream, the occlusion mark area, the posture features, the shape features and the distribution features are processed through the aggregation channel of the semantic segmentation main network to output the overall unit video stream and the local unit video stream.

4. The method according to claim 3, characterized in that The semantic segmentation main network construction step includes: Construct the segmentation channel loss function: , in, Characterize the segmentation channel loss value, Represents the i-th unit supervision area of ​​the j-th frame video image, Represents the i-th unit prediction area of ​​the j-th frame video image, represents a small constant, N represents the number of unit supervision regions of the j-th frame video image, Characterizes the number of video stream frames; Construct the occlusion identification channel loss function: , in, Characterizes the loss value of the occlusion identification channel, The true label that represents whether the kth pixel of the jth frame video image is occluded. Representation is occluded, when The representation is not blocked, Characterizes the probability that the kth pixel of the jth frame video image predicted by the occlusion identification channel is occluded, The total number of pixels representing the j-th frame of the video image; Construct the aggregate channel loss function: , in, Characterizes the aggregate channel loss value, Unit grouping supervision data representing the j-th frame of the video image, Unit grouping aggregation data representing the j-th frame of the video image, Represents the intersection symbol, Represents the union symbol, Characterizes natural constants; Construct the semantic segmentation main network loss function: , in, Represents the loss value of the main network for semantic segmentation, , and Characterize loss weights; The semantic segmentation main network is trained according to the segmentation channel loss function, the occlusion identification channel loss function, the aggregation channel loss function and the semantic segmentation main network loss function.

5. The method according to claim 1, characterized in that The first-level semantic segmentation sub-network of the main layer receives the local unit video stream from the semantic segmentation main network for clothing consistency matching, and obtains a first compensated overall unit video stream and a clothing blind area local unit video stream, including: Extracting a first frame local unit of the local unit video stream; Extracting a first group of local units and a second group of local units whose spatial distribution distance is less than or equal to a consistent distance threshold from the first frame local units, wherein the first group of local units and the second group of local units do not have an intersection unit type; Extracting first clothing features of the first group of local units according to the clothing feature extraction channel of the first-level semantic segmentation sub-network, and extracting second clothing features of the second group of local units; Comparing the first clothing feature and the second clothing feature according to the clothing feature comparison channel of the first-level semantic segmentation sub-network to obtain clothing similarity; When the clothing similarity is greater than or equal to the clothing similarity threshold, the first group of local units and the second group of local units are merged into the same group of units; when the clothing similarity is less than the clothing similarity threshold, no operation is performed; Cyclic analysis is performed. When the clothing similarity of any two groups of local units is less than the clothing similarity threshold, a video stream that has at least two unit types, namely, a dinner plate unit, an arm unit, a clothing unit and a head unit, is added to the first compensated overall unit video stream. Otherwise, it is added to the clothing blind spot local unit video stream.

6. The method according to claim 1, characterized in that The second-level semantic segmentation sub-network of the main layer receives the clothing blind area local unit video stream from the semantic segmentation main network for motion consistency matching to obtain a second compensated overall unit video stream, including: Extracting a first group of clothing blind spot local unit video streams and a second group of clothing blind spot local unit video streams of the clothing blind spot local unit video stream, wherein any frame spatial distribution distance between the first group of clothing blind spot local units and the second group of clothing blind spot local units is less than or equal to a consistent distance threshold, and does not have an intersection unit type; Extracting first motion trajectory information of the first group of clothing blind area local unit video streams through the motion feature extraction channel of the secondary semantic segmentation sub-network, and extracting second motion trajectory information of the second group of clothing blind area local unit video streams; Performing trajectory similarity evaluation on the first motion trajectory information and the second motion trajectory information through a motion trajectory comparison channel of the secondary semantic segmentation sub-network to obtain trajectory similarity; When the trajectory similarity is greater than or equal to the trajectory similarity threshold, the first group of clothing blind area local unit video streams and the second group of clothing blind area local unit video streams are merged into the same group of unit video streams; when the trajectory similarity is less than the trajectory similarity threshold, no operation is performed; Cyclic analysis is performed, and when the trajectory similarity of any two groups of local unit video streams is less than the trajectory similarity threshold, a video stream having at least two unit types of a dinner plate unit, an arm unit, a clothing unit and a head unit is added to the second compensated overall unit video stream.

7. The method according to claim 6, characterized in that When the trajectory similarity of any two groups of local unit video streams is less than the trajectory similarity threshold, a video stream having at least two unit types of a plate unit, an arm unit, a clothing unit and a head unit as one is added to the second compensated overall unit video stream, further comprising: When a video stream having a dinner plate unit and at least two unit types of an arm unit, a clothing unit and a head unit as one is added to the second compensated overall unit video stream, a residual local unit video stream is obtained, wherein the residual local unit video stream includes a local dinner plate unit video stream; When the number of the local plate unit video streams is equal to 0, outputting the second compensated overall unit video stream; When the number of the local dinner plate unit video streams is not equal to 0, trajectory unit hierarchical clustering is performed on each of the local dinner plate unit video streams until a video stream having at least two unit types of dinner plate unit, arm unit, clothing unit and head unit is found, and the video stream is added into the second compensated overall unit video stream.

8. The method according to claim 7, characterized in that Performing trajectory unit hierarchical clustering on each of the local dinner plate unit video streams until a video stream having a dinner plate unit and at least two unit types of an arm unit, a clothing unit, and a head unit is obtained, comprising: Extracting a first partial dining tray unit motion trajectory of a first partial dining tray unit video stream of the partial dining tray unit video stream; Extracting a first residual local unit motion trajectory of a first residual local unit video stream from the residual local unit video stream, wherein the first residual local unit video stream does not have a dinner plate unit, and a spatial distribution distance between a first residual local unit and a first local dinner plate unit in any frame of the first residual local unit video stream is less than or equal to a sorting distance threshold, and the sorting distance threshold represents a maximum distribution distance that may belong to one body; until an Nth residual local unit motion trajectory of the Nth residual local unit video stream is extracted from the residual local unit video stream, wherein the Nth residual local unit video stream does not have a dinner plate unit, and a spatial distribution distance between the Nth residual local unit and the first local dinner plate unit of any frame of the Nth residual local unit video stream is less than or equal to a sorting distance threshold, and N represents the number of all residual local unit video streams that meet the conditions of distribution distance and unit composition type; Extracting the residual local unit video stream having the maximum trajectory similarity with the first local plate unit motion trajectory from the first residual local unit motion trajectory to the Nth residual local unit motion trajectory, and merging it with the first local plate unit video stream to generate an updated video stream; When the update video stream has a dinner plate unit and at least two unit types of an arm unit, a clothing unit and a head unit as one, the update video stream is set as a second compensated whole unit video stream of the first local dinner plate unit video stream; Otherwise, the trajectory unit hierarchical clustering is continued based on the updated video stream until a video stream having at least two unit types of a dinner plate unit, an arm unit, a clothing unit, and a head unit is found.

9. A user identification device for canteen intelligent settlement, characterized in that: It includes a perception layer, a main layer and a client. The perception layer includes a camera perception node and a background configuration node, including: The data acquisition module is used to receive the biometric data to be calculated from the client, receive the video data stream from the camera perception node through the perception layer, and receive the dish distribution data stream from the background configuration node; A unit segmentation module, configured to receive the video data stream from the perception layer through a semantic segmentation main network of the main layer to perform unit segmentation, and obtain an overall unit video stream, a local unit video stream, and a plate unit video stream, wherein the local unit video stream represents a video stream having at most two unit types of a plate unit, an arm unit, a clothing unit, and a head unit as one, and the overall unit video stream represents a video stream having at least two unit types of a plate unit, an arm unit, a clothing unit, and a head unit as one; A clothing segmentation module, configured to receive the local unit video stream from the semantic segmentation main network through the primary semantic segmentation sub-network of the main layer to perform clothing consistency matching, and obtain a first compensated overall unit video stream and a clothing blind area local unit video stream; A motion segmentation module, configured to receive the clothing blind area local unit video stream from the semantic segmentation main network through the secondary semantic segmentation sub-network of the main layer to perform motion consistency matching, and obtain a second compensated overall unit video stream; An identity authentication module is used to receive the whole unit video stream, the first compensated whole unit video stream and the second compensated whole unit video stream from the semantic segmentation main network, the first-level semantic segmentation sub-network and the second-level semantic segmentation sub-network through the identity authentication network of the main layer, receive the biometric data to be settled from the perception layer to perform identity authentication on the plate unit video stream, and after obtaining the identity authentication plate unit video stream, receive the dish distribution data stream from the perception layer, associate it with the identity authentication plate unit video stream and the biometric data to be settled, store it, construct an identity authentication document, and send it to the client.

Citation Information

Patent Citations

  • Self-selection restaurant automatic charging method

    CN107122730A

  • Food safety data analysis system and method based on multi-modal retrieval

    CN118916498A