Self-checkout method, system and storage medium based on multimodal interaction
By adopting a multimodal interaction method in the self-service cashier system, combining deep recognition model and weight information for joint reasoning, and conducting shelf verification, the problem of low product recognition accuracy in the self-service cashier system is solved, and settlement efficiency and accuracy are improved.
Patent Information
- Application Number
- CN202510348316.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-03-24
AI Technical Summary
The existing self-service cashier system has low accuracy when identifying goods, and cannot improve the recognition accuracy from the source, resulting in insufficiency in settlement.
The self-service cashier method based on multimodal interaction is adopted. By taking initial images and identifying them based on the depth recognition model, combining weight information for joint reasoning, positioning and selling shelves for verification, ensuring the accuracy of product information.
It improves the accuracy and efficiency of product identification, solves the problems of low identification accuracy and insufficient settlement efficiency, and ensures the accuracy and speed of the settlement process.
Smart Images

Figure CN119863875B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data processing, and particularly relates to a self-service cashier method, system and storage medium based on multimodal interaction. Background Art
[0002] Self-service cashier systems have become the core facilities for improving operation efficiency in scenarios such as supermarkets and convenience stores. Current self-service cashier mainly relies on barcode recognition. However, in this process, users need to find the barcodes on the product packages by themselves. When the product packages are relatively complex or the settlement staff is not proficient, there may be problems such as a long settlement time and an impact on settlement efficiency.
[0003] To solve this problem, the following methods have been proposed in the prior art. For example, the Chinese patent document with the publication number CN109214806A discloses a self-service settlement method, device and storage medium. This method uses a variety of image recognition technologies to recognize the surveillance images, identifies the categories and quantities of the products to be settled from the surveillance images, then combines with a weight sensor to obtain the actual weight of the products to be settled, calculates the standard weight of the products in combination with the recognition results, and compares the actual weight with the standard weight to verify the accuracy of the recognition results. This method can complete the settlement recognition of products without scanning barcodes, thus greatly improving the settlement efficiency.
[0004] However, only by recognizing the product images in the settlement area in the above manner, there may be a situation where the product recognition accuracy is relatively low. Although the above technical solution introduces obtaining the product weight to verify the recognition results, this method can only determine whether the recognition results are correct. In case of incorrect results, manual confirmation is still required, and the recognition accuracy cannot be improved from the source. Summary of the Invention
[0005] To solve the above problems, the present invention provides a self-service cashier method, device and storage medium based on multimodal interaction to solve the problems existing in the prior art.
[0006] To achieve the above invention purpose, the present invention proposes a self-service cashier method based on multimodal interaction, including:
[0007] After receiving the start instruction, take an initial image of a predetermined area, split the initial image into multiple product images, and recognize each product image based on a depth recognition model to obtain the product information and the first confidence level included therein;
[0008] Generate a first product list based on the product information with the first confidence level greater than the first threshold, and count the first quantity and the second quantity of the recognized products and unrecognized products in the initial image;
[0009] If the first quantity and the second quantity meet a preset condition, obtain the weight information of the predetermined area, and perform joint inference based on all the commodity images and the weight information to obtain the commodity information of the unrecognized commodity and the second confidence level;
[0010] If the second confidence level is less than the first threshold, locate the corresponding sales shelf based on the commodity information of the unrecognized commodity, and verify the commodity information based on the picking and placing information of the sales shelf;
[0011] If the verification is successful, add the commodity information of the unrecognized commodity to the first commodity list to obtain a second commodity list. If the verification fails, generate a manual entry reminder;
[0012] After receiving the completion information, generate amount information based on the second commodity list and display a payment interface until a completion settlement information or a cancellation payment information is received.
[0013] Further, the recognition of the commodity image includes the following steps:
[0014] If a barcode area is detected in the commodity image, generate the commodity information based on the barcode and set the corresponding first confidence level to 100%. If the barcode area is not recognized, recognize the commodity image based on a depth recognition model to obtain the commodity information and the first confidence level. Determine the commodity type based on the commodity information. If the commodity type belongs to a preset type and the first confidence level is greater than the first threshold, continue to recognize the commodity quality. If a damaged commodity is recognized, generate a commodity quality reminder.
[0015] Further, the recognition of the commodity quality includes the following steps:
[0016] Perform grid segmentation on the commodity image to obtain a plurality of grid areas, obtain the original average pixel value of each grid area, set multiple interval ranges, and each interval range corresponds to a mapped pixel. Based on the interval range where the original average pixel value is located, map it to the corresponding mapped pixel;
[0017] Divide the grid area into multiple types of sub-areas, obtain the statistical features of the sub-areas. The statistical features include the mapped pixels existing in the sub-areas and the occupancy ratio. Establish a standard feature library. The standard feature library includes the standard features of each sub-area of each commodity type in a damaged state. Compare the statistical features of the sub-areas in the commodity image with the standard features, and determine the commodity quality based on the comparison result.
[0018] Further, verifying the product information based on the picking and placing information of the sales shelf includes the following steps:
[0019] Obtain the shelf image of the sales shelf. When a product is removed from the shelf in the shelf image, determine the product information of the removed product based on the shelf data, perform target tracking on the user who removes the product to obtain the user movement image, intercept and pre-store the facial image and body image in the user movement image, and label the product information of the sales shelf in the facial image and the body image as image tags. When there is a verification requirement, retrieve based on the product information of the unrecognized product and the image tags. If facial images and body images containing the product information of the unrecognized product are retrieved, compare them with the current user image. If the comparison passes, determine that the currently recognized product information is correct.
[0020] Further, jointly reasoning based on all the product images and the weight information includes the following steps:
[0021] Establish a weight database, where the weight database includes the standard weight and floating range of each product, set the error distribution of the weight sensor, obtain the first weight of the recognized product based on the weight database, calculate the second weight of the unrecognized product based on the total weight of the products in the predetermined area and the first weight, and correct the second weight based on the floating range and the error distribution to obtain its weight distribution range;
[0022] Obtain candidate information, where the candidate information is the product information in the recognition result of a single product image with the first confidence level greater than the second threshold. Obtain the standard weight and the floating range of the candidate information based on the weight database, generate candidate combinations that meet the weight distribution range based on the standard weight and the floating range, and calculate the combination score of each candidate combination based on the weight distribution range of the candidate combination;
[0023] Calculate the candidate score of each candidate combination based on the first confidence level and the combination score, and screen the candidate combination corresponding to the maximum candidate score as the target combination, where the candidate information included is used as the reasoning result of the unrecognized product, and the corresponding candidate score is used as the second confidence level.
[0024] Further, calculating the candidate score of the candidate combination includes the following steps:
[0025] Obtain the purchase probability of each candidate information based on historical purchase records, perform a weighted sum of the first confidence level, the combined score, and the purchase probability based on a preset weight to obtain the joint score of the candidate combination, normalize the joint score to obtain a normalized score, generate a penalty value based on the current number of existing candidate combinations, and correct all the normalized scores based on the penalty value to obtain the candidate scores.
[0026] Further, the preset conditions include that both the first quantity and the second quantity are less than a third threshold.
[0027] Further, if it is detected that the user moves the commodity back to the shelf or the user leaves the mall area in the shelf image, delete the user's facial image and body image.
[0028] The present invention also provides a self-checkout system based on multimodal interaction, which is used to implement the above-mentioned self-checkout method based on multimodal interaction. The system includes:
[0029] An identification module, after receiving a start instruction, captures an initial image of a predetermined area, splits the initial image into multiple commodity images, identifies each commodity image based on a depth recognition model to obtain the commodity information and the first confidence level included therein, and generates a first commodity list based on the commodity information with the first confidence level greater than a first threshold.
[0030] A first verification module, which counts the first quantity and the second quantity of the identified commodities and the unidentified commodities in the initial image. If the first quantity and the second quantity meet the preset conditions, obtains the weight information of the predetermined area, and performs joint inference based on all the commodity images and the weight information to obtain the commodity information and the second confidence level of the unidentified commodities.
[0031] A second verification module, if the second confidence level is less than the first threshold, locates the corresponding selling shelf based on the commodity information of the unidentified commodities, verifies the commodity information based on the picking and placing information of the selling shelf. If the verification is successful, adds the commodity information of the unidentified commodities to the first commodity list to obtain a second commodity list. If the verification fails, generates a manual entry reminder.
[0032] A settlement module, after receiving a completion information, generates amount information based on the second commodity list, and displays a payment interface until a completion settlement information or a cancellation payment information is received.
[0033] The present invention also discloses a computer storage medium, which stores program instructions. When the program instructions run, they control the device where the computer storage medium is located to execute the above-mentioned method.
[0034] Beneficial effects:
[0035] After receiving the start instruction, the present invention captures an initial image of a predetermined area, splits it into multiple product images based on an edge detection algorithm, and then uses a depth recognition model for recognition to quickly and accurately obtain product information in each product image. This method effectively solves the problem of low efficiency in manually scanning product barcodes one by one and improves the cash register efficiency.
[0036] By counting the number of recognized and unrecognized products in the initial image, after the system initially recognizes the products, it can clearly sort out a list of products to be settled. If the number of recognized and unrecognized products meets the preset conditions, the weight information of the predetermined area is continuously obtained, and joint reasoning is performed based on the weight information to obtain the product information of the unrecognized products. By adding weight information for joint reasoning, the accuracy and comprehensiveness of product recognition are improved. The present invention also locates the sales shelves based on the product information of the unrecognized products and verifies the product information according to the picking and placing information of the shelves. Through verification, the correctness of the inferred product information is ensured, and the problem that the product information is still inaccurate after joint reasoning or joint reasoning cannot be performed is solved; In summary, the present invention combines multiple methods to obtain the product information to be settled, thereby greatly improving the accuracy of product recognition. Brief description of the drawings
[0037] Figure 1 It is a flowchart of the steps of the self-service cash register method based on multi-modal interaction of the present invention;
[0038] Figure 2 It is a schematic structural diagram of the self-service cash register device of the present invention;
[0039] Figure 3 It is a schematic diagram of the principle of dividing product images of the present invention;
[0040] Figure 4 It is a schematic structural diagram of the self-service cash register system based on multi-modal interaction of the present invention.
[0041] In the figure, 1 is the placement area; 2 is the display area; 3 is the wireless transmission module; 4 is the scanning area. Detailed implementation manners
[0042] In order to make the purpose, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0043] It should be noted that all acquisitions and processing of information such as images and data in the present invention are carried out on the premise of complying with relevant laws, regulations and policies.
[0044] As Figure 1 shown, a self-service cashiering method based on multimodal interaction according to the present invention includes:
[0045] S1: After receiving a start instruction, capture an initial image of a predetermined area, split the initial image into multiple product images, and identify each product image based on a depth recognition model to obtain product information and a first confidence level included therein.
[0046] As Figure 2 shown, a self-service cashiering device for implementing the present invention includes a placement area 1 and a display area 2. The placement area 1 is used to place products. A weight sensor is provided in the placement area 1. The display area 2 is used to display a product list and amount information. The device further includes a wireless transmission module 3 for realizing the sending and receiving of data of the self-service cashiering device. When the products cannot be automatically identified according to the method of the present invention, the barcode scanning area 4 can also be used to identify the products based on barcodes. After the self-service cashiering device receives a start instruction, the start instruction is input by clicking on the screen, and the camera is called to capture an initial image of a predetermined area, that is, the initial image within the placement area 1. A fixed camera can be arranged above the cashiering device for shooting. Particularly, when placing items, the device reminds the user through voice, and the reminder information includes the following content: "Do not stack products when placing", "Please place fresh fruits and vegetables separately for identification".
[0047] After the shooting is completed, the initial image is obtained. The object contour in the initial image is detected based on the Canny edge detection algorithm, and then the initial image is cropped into multiple product images each containing only one object contour according to the object contour. Then, a depth recognition model based on CNN is used to identify each product image. The depth recognition model outputs the product information and the corresponding first confidence level included in the product image. For example, if the depth recognition model outputs the product information as potato chips of brand A and the first confidence level is 98%, the higher the first confidence level, the greater the possibility that the depth recognition model believes that the product image includes potato chips of brand A. Particularly, if the image includes a whole bag of fruits, each recognized contour therein is split into a product image based on the contour edge. If all the product images are recognized as the same kind of fruit, the product price is calculated based on the unit price and weight of the fruit.
[0048] S2: Generate a first product list based on the product information with the first confidence level greater than a first threshold, and count the first quantity and the second quantity of the recognized products and unrecognized products in the initial image.
[0049] S3: If the first quantity and the second quantity meet a preset condition, obtain the weight information of the predetermined area, and perform joint reasoning based on all the product images and the weight information to obtain the product information and a second confidence level of the unrecognized products.
[0050] In this embodiment, the preset conditions include that both the first quantity and the second quantity are less than the third threshold.
[0051] Here, the first threshold is set to 95%. If the first confidence level of the commodity information is greater than the first threshold, it is directly considered that the judgment result is correct, and it is added to the first commodity list. At the same time, the corresponding selling price is obtained from the price database. For fruits, the unit price of the fruit is multiplied by the weight to obtain the selling price. The first commodity list is the list of commodities to be settled. Then, the first quantity and the second quantity in the initial image are counted. For example, if multiple commodities are placed in a predetermined area, 5 of them are recognized and added to the first commodity list, and 3 are not recognized, then the first quantity is 5 and the second quantity is 4. The third threshold is set to 3. When the number of recognized commodities and unrecognized commodities both exceed 3, the error of combined reasoning by weight will increase greatly. In this case, the subsequent steps are used to determine the commodity information.
[0052] When it is determined that combined reasoning can be performed, combined reasoning is carried out according to the weight information of the recognized commodities and unrecognized commodities. The specific reasoning method will be introduced later. After the reasoning is completed, the commodity information and the second confidence level of the unrecognized commodities can be obtained.
[0053] S4: If the second confidence level is less than the first threshold, based on the commodity information of the unrecognized commodities, the corresponding selling shelf is located, and the commodity information is verified based on the picking and placing information of the selling shelf.
[0054] S5: If the verification is successful, the commodity information of the unrecognized commodities is added to the first commodity list to obtain the second commodity list. If the verification fails, a manual entry reminder is generated.
[0055] S6: After receiving the completion information, the amount information is generated based on the second commodity list, and the payment interface is displayed until the completion settlement information or the cancellation of payment information is received.
[0056] If the second confidence level is still less than the first threshold, or when combined reasoning cannot be performed, it indicates that the commodity information still cannot be determined accurately. At this time, according to the currently speculated commodity information, the shelf monitoring image of the commodity for sale is obtained, and it is determined whether the user currently settling accounts has purchased the speculated commodity based on the monitoring image, that is, the commodity information is verified. If the verification is passed, the speculated commodity information is added to the first commodity list and updated to the second commodity list. If the verification fails, it is recommended that the user manually scan the barcode to enter the unentered commodity.
[0057] After the user clicks to complete the entry on the display interface, the system receives the completion information, generates the amount information based on the second commodity list, and displays the payment interface after confirmation until it detects that the user has completed the payment or cancelled the payment.
[0058] After receiving the start instruction, the present invention captures an initial image of a predetermined area, splits it into multiple product images based on an edge detection algorithm, and then uses a depth recognition model for recognition to quickly and accurately obtain the product information in each product image. This method effectively solves the problem of low efficiency in manually scanning product barcodes one by one and improves the cash register efficiency.
[0059] By counting the number of recognized and unrecognized products in the initial image, after the system initially recognizes the products, it can clearly sort out the list of products to be settled. If the number of recognized and unrecognized products meets the preset conditions, it continues to obtain the weight information of the predetermined area and performs joint reasoning based on the weight information to obtain the product information of the unrecognized products. By adding weight information for joint reasoning, the accuracy and comprehensiveness of product recognition are improved. The present invention also locates the sales shelves based on the product information of the unrecognized products and verifies the product information according to the picking and placing information of the shelves. Through verification, the correctness of the inferred product information is ensured, and the problems that the product information is still inaccurate after joint reasoning or joint reasoning cannot be performed are solved; In summary, the present invention combines multiple methods to obtain the product information to be settled, thereby greatly improving the accuracy of product recognition.
[0060] Specifically, the recognition of the product image by the present invention includes the following steps:
[0061] If a barcode area is detected in the product image, product information is generated based on the barcode, and the corresponding first confidence level is set to 100%. If the barcode area is not recognized, the product image is recognized based on the depth recognition model to obtain the product information and the first confidence level. The product type is determined based on the product information. If the product type belongs to a preset type and the first confidence level is greater than the first threshold, the product quality is continuously recognized. If a damaged product is recognized, a product quality reminder is generated.
[0062] First, the barcode area in the product image is detected. Since the barcode area usually shows a high gradient change in the horizontal direction (strip edge) and a small gradient change in the vertical direction, a detection method based on the gradient direction can be used. Specifically, the Sobel operator is used to calculate the X and Y direction gradients of the image, and then the gradient magnitude is calculated. A threshold is set to filter the low-gradient areas and retain the high-gradient areas. Then, the gradient direction is counted, and the areas with higher direction consistency are screened. If the proportion of the horizontal or vertical gradient direction in this area exceeds the preset threshold (such as 80%), it is determined as a valid barcode area.
[0063] After detecting the presence of a barcode area in the image, identify the product information based on the barcode and set the first confidence level to 100%, indicating that the product information obtained through the barcode must be correct. If the barcode area is not recognized, use the deep recognition model to recognize the product image to obtain the product information and the corresponding first confidence level.
[0064] The preset type of this embodiment is fruits or vegetables. The present invention introduces a quality determination function in the self-checkout machine, which can automatically generate a reminder when detecting quality problems of fruits or vegetables, thereby improving user purchase satisfaction. For example, when detecting the presence of Red Fuji apples in the product image based on the deep recognition model, continue to perform quality detection on them. When it is found that some of the Red Fuji apples are damaged, a reminder is generated. In particular, fruits are usually settled in bags, and the bottom fruits are often blocked by the top fruits. In this case, when it is recognized that the top fruits are damaged, a reminder will still be generated.
[0065] Evaluating the product quality in this embodiment includes the following steps:
[0066] Perform grid segmentation on the product image to obtain multiple grid regions, obtain the original average pixel value of each grid region, set multiple interval ranges, each interval range corresponds to a mapped pixel, and map the original average pixel value to the corresponding mapped pixel based on the interval range where the original average pixel value is located.
[0067] For ease of description, the product image is divided into 6*6 grid regions, and each grid region includes 4*4 pixels. In particular, the finer the grid segmentation, the more accurate the final judgment result will be, but the judgment speed will be slower. The coarser the grid segmentation, the less accurate the final judgment result will be, but the judgment speed will be increased. The original average pixel value is obtained by calculating the average pixel value of all the same channels within the grid region. In this embodiment, 10 interval ranges are set for each of the RGB three channels. For example, for the R, G, and B channels, 0-25 is set as an interval unit, and the corresponding mapped pixel point is 13. When the original average pixel value of a certain grid region in the R channel is 19 and is within the interval range of 0-25, it is mapped to 13. The processing methods for other channels are the same. By performing mapping, the calculation complexity can be reduced and the calculation speed can be increased.
[0068] Divide the grid regions into multiple types of sub-regions, obtain the statistical features of the sub-regions. The statistical features include the mapped pixels existing in the sub-regions and the occupancy ratio. Establish a standard feature library. The standard feature library includes the standard features of each sub-region of each product type in the damaged state. Compare the statistical features of the sub-regions within the product image with the standard features, and determine the product quality based on the comparison result.
[0069] Such asFigure 3 As shown, centering on the commodity image, a plurality of grid regions are divided into one region around its circumferential direction, so that the commodity image is divided into 3 layers from the outside to the inside. Among them, the outermost region is defined as sub-region A, the middle layer region is defined as sub-region B, and the innermost region is defined as sub-region C. In other embodiments, the commodity image can also be evenly divided into a plurality of 3*3 rectangular regions. Then, the statistical features of each sub-region are calculated. For example, in the R channel of sub-region B, the mapped pixel values included are 13, 38, and 63, and the existing ratio in sub-region B is 1:1:2. Then, the occupancy ratio of the mapped pixel 38 is 25%. In this way, the mapped pixels and occupancy ratios of the G channel and B channel are continuously obtained.
[0070] The standard database includes the standard features of the commodity in a damaged state. For example, when there is rot on the surface of a Red Fuji apple, sub-region B should include the mapped pixel 38 in the R channel, and the occupancy ratio of the mapped pixel 38 should be greater than or equal to 20%. The G channel should include the mapped pixel 88, and the occupancy ratio of the mapped pixel 88 should be greater than or equal to 20%. The B channel should include the mapped pixel 138, and the occupancy ratio of the mapped pixel 138 should be greater than or equal to 20%. When the mapped pixels and ratios included in the sub-region all meet the above conditions, it is considered that the sub-region conforms to the standard features. For example, in the R channel of sub-region B above, the occupancy ratio of the mapped pixel 38 is 25%, which is greater than the ratio in the standard features, that is, it conforms to the standard features of the R channel. When the statistical features of the 3 channels all conform to the standard features, it is considered that there is commodity damage in the sub-region. In the commodity image, when at least one sub-region shows damage features, it is determined that there is a problem with the commodity quality. In other embodiments, to speed up the calculation, the commodity image can also be processed as a grayscale image before processing.
[0071] Since using a deep neural network for recognition requires a large amount of training sets for training, in the case where a large number of training sets cannot be obtained and the transfer learning method cannot achieve good recognition results, the above method proposed by the present invention can be quickly deployed.
[0072] In this embodiment, verifying the commodity information based on the picking and placing information of the selling shelf includes the following steps:
[0073] Obtain the shelf image of the sales shelf. When a commodity is removed from the shelf in the shelf image, determine the commodity information of the removed commodity based on the shelf data, perform target tracking on the user who removes the commodity to obtain the user movement image, intercept the facial image and body image in the user movement image for pre-storage, and label the commodity information of the sales shelf in the facial image and body image as the image tag. When there is a verification requirement, perform a search based on the commodity information of the unrecognized commodity and the image tag. If a facial image and body image containing the unrecognized commodity information are retrieved, compare them with the current user image. If the comparison passes, determine that the currently recognized commodity information is correct.
[0074] First, set up monitoring cameras in each shelf area. The monitoring cameras capture images of the shelf area and perform recognition based on a depth recognition model to determine whether there is a picking or placing behavior on the commodity shelves in the shelf area. The shelf data includes the commodity information sold on each shelf within the monitoring area of each monitoring camera. When it is detected that a commodity is removed from the sales shelf, first determine the position of the shelf based on the monitoring image, then locate the corresponding shelf data based on the shelf position, and obtain the commodity information of the sales shelf from it. If it is determined that the commodity sold on this sales shelf is Commodity A, continue to intercept the image of the person performing the removal operation and track this person based on the target tracking algorithm until their facial image and body image are obtained. The features that need to be recorded in the body image include the clothing color. Finally, pre-store the facial image and body image, and mark the commodity information as Commodity A.
[0075] When the commodity information of the unrecognized commodity determined by the self-service cash register is Commodity A, but the first confidence level or the second confidence level is less than the first threshold, search for facial data and body data containing the image tag of Commodity A in the pre-stored image data. If not found, the verification of the commodity information fails. If facial images and body images containing Commodity A are found, continue to obtain the facial image and body image of the current settlement user, and compare them with the found facial image and body image. If the comparison result shows that they are both the same person, the verification of the commodity information passes.
[0076] In this embodiment, if it is detected that the user moves the commodity back to the shelf in the shelf image, or the user leaves the mall area, delete the user's facial image and body image.
[0077] When the user moves the commodity back to the shelf, it indicates that the user has given up the purchase, so delete it to avoid subsequent verification errors. After the user leaves the mall, remove the corresponding facial image and body image to protect the user's privacy.
[0078] In this embodiment, the joint inference based on all commodity images and weight information includes the following steps:
[0079] Establish a weight database, where the weight database includes the standard weight and floating range of each commodity, set the error distribution of the weight sensor, obtain the first weight of the identified commodity based on the weight database, calculate the second weight of the unidentified commodity based on the total weight of the commodities within a predetermined area and the first weight, and correct the second weight based on the floating range and error distribution to obtain its weight distribution range.
[0080] The weight database includes the standard weight of each commodity, such as 200g for potato chips A, and the floating range of each commodity. For example, the floating range of potato chips A can be ±5g. For the weight sensor, the error distribution is set to 5g. The following is an example to introduce the reasoning process. Suppose there are currently two identified commodities and two unidentified commodities, and both identified commodities are potato chips A, then the first weight is 200g, and the total weight of the current predetermined area is 1200g, then the second weight is 1200 - 2 * 200 = 800g. Since there are two potato chips A, plus the error of the sensor, the weight distribution of the second weight after correction is [800 - 3 * 5g, 800 + 3 * 5g], that is, [785g, 815g].
[0081] Obtain candidate information. The candidate information is the commodity information with the first confidence greater than the second threshold in the recognition result of a single commodity image. Obtain the standard weight and floating range of the candidate information based on the weight database, generate candidate combinations that meet the weight distribution range based on the standard weight and floating range, and calculate the combination score of each candidate combination based on the weight distribution range of the candidate combination.
[0082] Identify the image of the unidentified commodity. The first confidence that it is commodity A is 0.6, the first confidence that it is commodity B is 0.55, and the first confidence that it is commodity C is 0.4. The second threshold is set to 0.5, then commodity A and commodity B are taken as two pieces of candidate information. In particular, if at least one of commodity A and commodity B is a fresh fruit, the reasoning is abandoned and the user is reminded to settle separately, because due to the interference of other commodities, it is impossible to accurately obtain its weight during settlement. For commodity A, its standard weight is 400g and the floating range is ±20g. For commodity B, its standard weight is 395g and the floating range is ±25g. Then the following candidate combinations are generated: Candidate combination A: 2 commodity A, with the weight between [760g, 840g]; Candidate combination B: 2 commodity B, with the weight between [765g, 815g]. Other candidate combinations are not listed one by one.
[0083] In this embodiment, it is first defined that the weights of product A and product B follow a normal distribution. For example, product A follows a normal distribution with μ = 400 and σ = 20. Based on the probability density function of the normal distribution, the probability density value of product A at the standard weight of 400g and the probability density value of product B at the standard weight of 395g are calculated respectively. The probability density function of the normal distribution is common knowledge and will not be introduced here. After calculation, the probability density value P(400) of product A at the standard weight of 400g is 0.0199, and the probability density value P(400) of product B at the standard weight of 395g is 0.0158.
[0084] Calculate the combined score of each candidate combination based on the following first formula: , where is the combined score, is the number of candidate information included in the candidate combination, is the probability density value of the i-th candidate information in the candidate combination at the standard weight, is the natural constant, is the penalty factor, and its value is 0.1, is the total weight of the candidate combination, which is specifically obtained by accumulating the standard weights of the candidate information included in the candidate combination. For example, the standard weight of candidate combination A is 400 + 400 = 800g, is the second weight. By substituting the above values into the first formula, the combined score of candidate combination A can be obtained, which is specifically . Similarly, the combined score of candidate combination B can be calculated as 0.000226. In the first formula, the larger the probability density value or the closer the total weight of the candidate combination is to the second weight, the larger the calculated combined score.
[0085] Calculate the candidate score of each candidate combination based on the first confidence level and the combined score, and select the candidate combination corresponding to the maximum candidate score as the target combination. The candidate information included therein is used as the inference result of the unrecognized product, and the corresponding candidate score is used as the second confidence level.
[0086] The steps for calculating the candidate score of the candidate combination in this embodiment are as follows:
[0087] Obtain the purchase probability of each candidate information based on the historical purchase records, perform a weighted sum of the first confidence level, the combined score, and the purchase probability based on the preset weights to obtain the joint score of the candidate combination, normalize the joint score to obtain the normalized score, generate a penalty value based on the current number of existing candidate combinations, and correct all the normalized scores based on the penalty value to obtain the candidate score.
[0088] In this embodiment, before settlement, the user needs to log in to the account for settlement. Although this method sacrifices some convenience, it can store the user's purchase records in the user account, facilitating the user to view at any time and perform subsequent after-sales processing. After the user logs in to the account, the purchase probability of the candidate information is determined based on the user's purchase records. For example, the ratio of the historical purchase times of product A to the total historical shopping times is used as the purchase probability of product A.
[0089] The weighted sum of the purchase probability, the first confidence level, and the combined score is obtained according to the preset weights built into the system to obtain the combined score. For example, the preset weights of the purchase probability, the first confidence level, and the combined score are 0.1, 0.6, and 0.3 respectively. In this embodiment, the weighted sum is performed based on the second formula, and the second formula is: , where S is the combined score of the candidate combination, is the average purchase probability of the candidate information (products) included in the candidate combination, is the average first confidence level of the products included in the candidate combination, R is the combined score of the candidate combination, 、 and are the preset weights of the purchase probability, the first confidence level, and the combined score respectively. For candidate combination A, the average purchase probability of the candidate information it includes is 0.16, the average first confidence level of the candidate information it includes is 0.85, and the combined score is 0.000396. Substituting into the second formula, the combined score of candidate combination A is obtained as -2.02. Similarly, the combined score of candidate combination B is calculated as -2.39.
[0090] In this embodiment, the Softmax function is used for normalization. The normalized score of candidate combination A is 0.592, and the normalized score of candidate combination B is 0.408. The more candidate combinations are generated, the higher the penalty value, indicating that the number of combinations that meet the weight distribution range is too large, and the calculation accuracy will be greatly reduced. For example, if the penalty value here is 0.95, then the candidate score of candidate combination A after correction is 0.562, and the candidate score of candidate combination B is 0.3876.
[0091] After the above calculations, it is determined that the candidate score of candidate combination A is the largest, so candidate combination A is the target combination, which includes two product As as the inference result of unrecognized products. The candidate score of candidate combination A is the second confidence level.
[0092] As Figure 4 shown, the present invention also provides a self-service cash register system based on multimodal interaction. This system is used to implement the above-mentioned self-service cash register method based on multimodal interaction. This system includes:
[0093] The recognition module, after receiving the start instruction, captures the initial image of the predetermined area, splits the initial image into multiple product images, recognizes each product image based on the depth recognition model to obtain the product information and the first confidence level included therein, and generates the first product list based on the product information with the first confidence level greater than the first threshold.
[0094] The first verification module counts the first quantity and the second quantity of the recognized products and unrecognized products in the initial image. If the first quantity and the second quantity meet the preset conditions, it obtains the weight information of the predetermined area, and performs joint reasoning based on all product images and the weight information to obtain the product information and the second confidence level of the unrecognized products.
[0095] The second verification module, if the second confidence level is less than the first threshold, locates the corresponding sales shelf based on the product information of the unrecognized products, verifies the product information based on the picking and placing information of the sales shelf. If the verification is successful, it adds the product information of the unrecognized products to the first product list to obtain the second product list. If the verification fails, it generates a manual entry reminder.
[0096] The settlement module, after receiving the completion information, generates the amount information based on the second product list and displays the payment interface until it receives the completion settlement information or the cancel payment information.
[0097] The present invention also discloses a computer storage medium storing program instructions, wherein when the program instructions run, they control the device where the computer storage medium is located to execute the method described above.
[0098] It can be understood that the terms "first", "second", etc. used in the present application may be used herein to describe various elements, but unless otherwise specified, these elements are not limited by these terms. These terms are only used to distinguish the first element from another element. For example, without departing from the scope of the present application, the first xx script may be referred to as the second xx script, and similarly, the second xx script may be referred to as the first xx script.
[0099] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0100] The above embodiments only represent several implementation manners of the present invention. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the patent of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several deformations and improvements can be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the appended claims.
[0101] The above is only the preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A self-service checkout method based on multimodal interaction, characterized in that: include: After receiving the start instruction, an initial image of a predetermined area is captured, the initial image is split into a plurality of product images, and each of the product images is recognized based on a deep recognition model to obtain product information and a first confidence level included therein; generating a first commodity list based on the commodity information whose first confidence level is greater than a first threshold, and counting a first number and a second number of recognized commodities and unrecognized commodities in the initial image; If the first quantity and the second quantity meet a preset condition, obtaining weight information of the predetermined area, performing joint reasoning based on all the commodity images and the weight information, and obtaining the commodity information and a second confidence level of the unrecognized commodity; If the second confidence level is less than the first threshold, locating a corresponding sales shelf based on the product information of the unidentified product, and verifying the product information based on the pick-and-place information of the sales shelf; If the verification is successful, the product information of the unidentified product is added to the first product list to obtain the second product list; if the verification fails, a manual entry reminder is generated; After receiving the completion information, generating amount information based on the second commodity list and displaying the payment interface until receiving the completion information or the payment cancellation information; Performing joint reasoning based on all the product images and the weight information includes the following steps: Establishing a weight database, the weight database including a standard weight and a floating range of each commodity, setting an error distribution of a weight sensor, obtaining a first weight of an identified commodity based on the weight database, calculating a second weight of an unidentified commodity based on a total weight of commodities in the predetermined area and the first weight, and correcting the second weight based on the floating range and the error distribution to obtain a weight distribution range thereof; Acquire candidate information, where the candidate information is the commodity information whose first confidence is greater than a second threshold in a single commodity image recognition result, acquire the standard weight and the floating range of the candidate information based on the weight database, generate a candidate combination that satisfies the weight distribution range based on the standard weight and the floating range, and calculate a combination score for each candidate combination based on the weight distribution range of the candidate combination; The candidate score of each candidate combination is calculated based on the first confidence and the combination score, and the candidate combination corresponding to the maximum candidate score is screened as the target combination, wherein the candidate information included therein is used as the inference result of the unidentified commodity, and the corresponding candidate score is used as the second confidence.
2. The method according to claim 1, characterized in that Identifying the product image includes the following steps: If a barcode area is detected in the product image, the product information is generated based on the barcode, and the corresponding first confidence is set to 100%; if the barcode area is not identified, the product image is identified based on the deep recognition model to obtain the product information and the first confidence; the product type is determined based on the product information; if the product type belongs to a preset type and the first confidence is greater than the first threshold, the product quality continues to be identified; if damaged products are identified, a product quality reminder is generated.
3. The method according to claim 2, characterized in that Identifying the quality of the product includes the following steps: The product image is divided into grids to obtain a plurality of grid areas, an original average pixel value of each of the grid areas is obtained, a plurality of interval ranges are set, each of the interval ranges corresponds to a mapping pixel, and based on the interval range where the original average pixel value is located, it is mapped to the corresponding mapping pixel; The grid area is divided into multiple types of sub-areas, and statistical features of the sub-areas are obtained, wherein the statistical features include the mapped pixels existing in the sub-areas and their occupation ratios; a standard feature library is established, wherein the standard feature library includes standard features of each sub-area of each product type in a damaged state; the statistical features of the sub-areas in the product image are compared with the standard features, and the quality of the product is determined based on the comparison results.
4. The method according to claim 1, characterized in that Verifying the commodity information based on the pick-and-place information of the selling shelf includes the following steps: Acquire a shelf image of the selling shelf, and when a product is removed from the shelf in the shelf image, determine the product information of the removed product based on the shelf data, track the user who removed the product to obtain a user movement image, capture a facial image and a body image in the user movement image for pre-storage, and mark the product information of the selling shelf in the facial image and the body image as an image tag; when a verification requirement arises, perform a search based on the product information of the unrecognized product and the image tag; if the facial image and the body image containing the product information of the unrecognized product are retrieved, compare them with the current user image; if the comparison is successful, determine that the currently recognized product information is correct.
5. The method according to claim 1, characterized in that: Calculating the candidate score of the candidate combination comprises the following steps: The purchase probability of each type of candidate information is obtained based on historical purchase records, and the first confidence, the combination score and the purchase probability are weighted and summed based on preset weights to obtain a joint score of the candidate combination, and the joint score is normalized to obtain a normalized score, a penalty value is generated based on the number of candidate combinations currently existing, and all the normalized scores are corrected based on the penalty value to obtain the candidate score.
6. The method according to claim 5, characterized in that The preset condition includes that both the first number and the second number are smaller than a third threshold.
7. The method according to claim 4, characterized in that If it is detected in the shelf image that the user moves the product back to the shelf, or the user leaves the shopping mall area, the facial image and the body image of the user are deleted.
8. A self-service cashier system based on multimodal interaction, used to implement the method according to any one of claims 1 to 7, characterized in that: include, The recognition module, after receiving the start instruction, captures an initial image of a predetermined area, splits the initial image into a plurality of product images, recognizes each of the product images based on a deep recognition model to obtain product information and a first confidence level included therein, and generates a first product list based on the product information whose first confidence level is greater than a first threshold. a first verification module, which counts the first number and the second number of identified commodities and unidentified commodities in the initial image, obtains the weight information of the predetermined area if the first number and the second number meet the preset conditions, performs joint reasoning based on all the commodity images and the weight information, and obtains the commodity information and the second confidence of the unidentified commodities, wherein when performing joint reasoning based on all the commodity images and the weight information, a weight database is established, the weight database includes the standard weight and floating range of each commodity, the error distribution of the weight sensor is set, the first weight of the identified commodity is obtained based on the weight database, the second weight of the unidentified commodity is calculated based on the total weight of the commodities in the predetermined area and the first weight, and the second weight is corrected based on the floating range and the error distribution to obtain its weight distribution range; Acquire candidate information, where the candidate information is the commodity information whose first confidence is greater than a second threshold in a single commodity image recognition result, acquire the standard weight and the floating range of the candidate information based on the weight database, generate a candidate combination that satisfies the weight distribution range based on the standard weight and the floating range, and calculate a combination score for each candidate combination based on the weight distribution range of the candidate combination; Calculating a candidate score for each candidate combination based on the first confidence and the combination score, screening the candidate combination corresponding to the maximum candidate score as the target combination, wherein the candidate information included therein is used as the inference result of the unidentified commodity, and the corresponding candidate score is used as the second confidence; a second verification module, if the second confidence level is less than the first threshold, locating a corresponding sales shelf based on the product information of the unidentified product, verifying the product information based on the pick-and-place information of the sales shelf, and if the verification is successful, adding the product information of the unidentified product to the first product list to obtain a second product list, and generating a manual entry reminder if the verification fails; The settlement module, after receiving the completion information, generates amount information based on the second commodity list and displays a payment collection interface until receiving the completion information or the payment cancellation information.
9. A computer storage medium, characterized in that The computer storage medium stores program instructions, wherein when the program instructions are executed, the device where the computer storage medium is located is controlled to execute the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Self-service settlement method, device and storage medium
CN109214806A
Method and device for product checkout in unmanned store
WO2024232665A1