A product settlement data determination method, device, equipment and storage medium
By using a fine-grained recognition model to identify the types of image blocks on the smart electronic scale, the problem of low accuracy in product type identification is solved, enabling efficient and accurate calculation of settlement data and improving the user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BOC FINANCIAL TECH (SUZHOU) CO LTD
- Filing Date
- 2023-10-11
- Publication Date
- 2026-07-21
AI Technical Summary
Existing smart electronic scales have low accuracy in identifying product types, leading to errors in settlement data calculation and reducing user experience.
A fine-grained recognition model is used to identify the type of multiple image blocks of the product to be weighed. The transformer module and the TransFGC module are used for position encoding and downsampling operations. Convolutional layers are combined to improve recognition accuracy and efficiency. Data augmentation techniques are used to expand the training dataset.
It improved the accuracy of product type identification and settlement data calculation, thus enhancing the user experience.
Smart Images

Figure CN117152740B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing, and more specifically, to a method, apparatus, device, and storage medium for determining product settlement data. Background Technology
[0002] Currently, the application of smart electronic scales is becoming more and more widespread. When weighing products, such as agricultural and sideline products, people usually identify the types of agricultural and sideline products by eye. They then select the corresponding type button on the smart electronic scale, and the smart electronic scale automatically weighs the product and calculates the product price to obtain the product settlement data.
[0003] When using the above method to determine product settlement data, it is necessary to manually remember the types of vegetables. However, in practical applications, it is easy to encounter situations where human error in identifying product types leads to the wrong selection of the type button on the smart electronic scale, resulting in incorrect calculation of the final product settlement data. In other words, human identification of product types has a low accuracy problem, which leads to low accuracy in product settlement data calculation and reduces user experience. Summary of the Invention
[0004] In view of this, the present invention discloses a method, apparatus, device and storage medium for determining product settlement data, in order to solve the problem that low accuracy in product type identification leads to low accuracy in product settlement data calculation and reduces user experience.
[0005] To solve the above-mentioned technical problems, the present invention adopts the following technical solution:
[0006] A method for determining product settlement data includes:
[0007] Obtain the initial image and weight of the product to be weighed;
[0008] The initial image is segmented to obtain multiple image blocks;
[0009] A fine-grained recognition model is invoked to identify the types of the multiple image patches to determine the types of the products to be weighed. The fine-grained recognition model includes a predetermined number of sequentially connected transformer modules. The first transformer module includes a sequentially connected position encoding module and a TransFGC module. The second to last transformer modules each include a sequentially connected downsampling module and a TransFGC module. In the TransFGC module, a convolutional layer is configured after the last normalization layer. The fine-grained recognition model is trained based on training data, which includes product image samples and type labels.
[0010] Based on the type and weight of the product to be weighed, calculate the product settlement data for the product to be weighed.
[0011] Optionally, the initial image is segmented to obtain multiple image blocks, including:
[0012] Obtain the segmentation parameters used during image segmentation;
[0013] The initial image is segmented according to the segmentation parameters to obtain multiple image blocks.
[0014] Optionally, a fine-grained recognition model is invoked to perform category identification on the plurality of image blocks to obtain the category of the product to be weighed, including:
[0015] The first transformer module in the fine-grained recognition model performs position encoding and image filtering operations on the multiple image blocks, and outputs the image block filtering results.
[0016] The image block filtering results are subjected to downsampling and image filtering operations sequentially through the second to last transformer modules in the fine-grained recognition model to obtain the type of the product to be weighed.
[0017] Optionally, a downsampling operation is performed on the image patch filtering results, including:
[0018] The image block filtering results are divided into multiple sub-images;
[0019] The pixel blocks located at the same position in the multiple sub-images are combined to obtain a feature map;
[0020] The feature maps are stitched together along the depth direction to obtain the stitching result.
[0021] The splicing result is normalized and pooled to obtain the sampling result.
[0022] Optionally, a feature extraction module is configured between the last transformer module and the penultimate transformer module;
[0023] The image patch selection results are sequentially processed through the second to last transformer modules in the fine-grained recognition model to perform downsampling and image selection operations, thereby obtaining the types of the products to be weighed, including:
[0024] The image patch selection results are then subjected to downsampling and image selection operations sequentially through the second to second-to-last transformer modules in the fine-grained recognition model to obtain new image patch selection results.
[0025] The new image patch filtering results are processed by the feature extraction module to extract features, resulting in feature extraction results.
[0026] The last transformer module performs downsampling and image filtering operations on the feature extraction results to obtain the type of the product to be weighed.
[0027] Optionally, the generation process of the fine-grained recognition model includes:
[0028] Acquire training data; the training data includes product image samples and category labels;
[0029] The training data is augmented to obtain training samples;
[0030] The fine-grained recognition model is trained using the training samples until a preset training stopping condition is met.
[0031] Optionally, the formula for calculating the loss function used during the training of the fine-grained recognition model is as follows:
[0032]
[0033] in, For loss function, For hyperparameters, For weighted cross-entropy loss, To predict product categories, For category labels, To compare the feature learning parameters, where,
[0034]
[0035] in, and It is a constant. , To identify different image blocks, For image blocks Location information, For image blocks Location information, For image blocks Predicted product categories For image blocks Predicted product categories.
[0036] A device for determining product settlement data, comprising:
[0037] The data acquisition module is used to acquire the initial image and weight of the product to be weighed;
[0038] The image segmentation module is used to segment the initial image to obtain multiple image blocks;
[0039] A category identification module is used to call a fine-grained identification model to identify the category of the multiple image blocks in order to obtain the category of the product to be weighed. The fine-grained identification model includes a preset number of sequentially connected transformer modules. The first transformer module includes a position encoding module and a TransFGC module connected in sequence. The second to the last transformer modules each include a downsampling module and a TransFGC module connected in sequence. In the TransFGC module, a convolutional layer is configured after the last normalization layer. The fine-grained identification model is trained based on training data. The training data includes product image samples and category labels.
[0040] The data calculation module is used to calculate the product settlement data of the product to be weighed based on the type and weight of the product to be weighed.
[0041] An electronic device, the electronic device including a memory and a processor;
[0042] The memory is used to store at least one instruction;
[0043] The processor is used to execute the at least one instruction to implement the above-described method for determining product settlement data.
[0044] A computer-readable storage medium storing at least one instruction that, when executed by a processor, implements the above-described method for determining product settlement data.
[0045] Compared with the prior art, the present invention has the following beneficial effects:
[0046] As can be seen from the above technical solution, the present invention provides a method, apparatus, device, and storage medium for determining product settlement data. In this invention, a fine-grained recognition model is invoked to identify the type of multiple image blocks of the product to be weighed, thereby obtaining the type of the product and enabling the calculation of product settlement data based on the type and weight of the product. The fine-grained recognition model in this invention is trained on training data. Therefore, after training, compared to manual product type identification, the fine-grained recognition model can improve the accuracy of product type identification, thereby improving the accuracy of product settlement data calculation and enhancing the user experience. Furthermore, a convolutional layer is configured after the last normalization layer in the TransFGC module; that is, the present invention improves the accuracy of product type identification through convolutional layers, which in turn also improves the accuracy of product settlement data calculation. Attached Figure Description
[0047] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the published drawings without creative effort.
[0048] Figure 1 This is a flowchart of a method for determining product settlement data disclosed in an embodiment of the present invention;
[0049] Figure 2 This is a schematic diagram of a fine-grained recognition model disclosed in an embodiment of the present invention;
[0050] Figure 3 This is a schematic diagram of a location encoding disclosed in an embodiment of the present invention;
[0051] Figure 4 This is a schematic diagram of a TransFGC module disclosed in an embodiment of the present invention;
[0052] Figure 5 This is a schematic diagram of a downsampling method disclosed in an embodiment of the present invention;
[0053] Figure 6 This is a flowchart of a downsampling method disclosed in an embodiment of the present invention;
[0054] Figure 7 This is a schematic diagram of a feature extraction module disclosed in an embodiment of the present invention;
[0055] Figure 8 This is a flowchart of a method for generating a fine-grained recognition model disclosed in an embodiment of the present invention;
[0056] Figure 9 This is a schematic diagram illustrating a scenario of a method for determining product settlement data disclosed in an embodiment of the present invention;
[0057] Figure 10 This is a schematic diagram of the structure of a product settlement data determination device disclosed in an embodiment of the present invention;
[0058] Figure 11 This is a schematic diagram of the structure of an electronic device disclosed in an embodiment of the present invention. Detailed Implementation
[0059] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0060] Currently, the application of smart electronic scales is becoming more and more widespread. When weighing products, such as agricultural and sideline products, people usually identify the types of agricultural and sideline products by eye. They then select the corresponding type button on the smart electronic scale, and the smart electronic scale automatically weighs the product and calculates the product price to obtain the product settlement data.
[0061] When using the above method to determine product settlement data, it is necessary to manually remember the types of vegetables. However, in practical applications, it is easy to make mistakes in human identification of product types, resulting in the wrong type button being selected on the smart electronic scale, which leads to errors in the final product settlement data calculation.
[0062] Furthermore, in practical applications, such as large agricultural markets and supermarkets, products, like agricultural by-products, often have multiple varieties within the same category. For example, grapes are divided into Kyoho grapes, Xiangfeng grapes, and Red Globe grapes. In such cases, human error in identifying the variety is more likely to occur, resulting in lower accuracy. Moreover, product weighing relies heavily on human intervention, leading to lower efficiency.
[0063] To improve the accuracy of product identification, increase efficiency, and reduce manpower, smart electronic scales can be improved. However, current improvements mainly focus on relying on a precisely designed structure to achieve accurate weighing, or on the rotation of the scale's bottom tray to collect images from multiple angles. Neither of these methods can achieve the technical effects of improving the accuracy of product identification, increasing efficiency, and reducing manpower.
[0064] To address this, the inventors discovered that with the rapid development of computer vision and deep learning technologies, fine-grained recognition algorithms have made significant progress in target classification and recognition tasks. These algorithms can identify and distinguish subtle differences with similar appearance features, such as different varieties of the same vegetable or meat. However, research and practical applications of fine-grained recognition algorithms in traditional electronic scales are relatively limited. Furthermore, in practical applications, fine-grained algorithms generally employ neural networks or deep learning, resulting in low accuracy and potential misclassification or misjudgment when distinguishing similar products. Especially when appearances are similar or deformed, recognition accuracy may decrease, leading to inaccurate calculation of the total transaction price. Additionally, current fine-grained algorithms have high computational complexity, potentially requiring long processing times and resulting in slow response times, making them unsuitable for scenarios requiring rapid and accurate product identification and pricing, such as product weighing. Moreover, environmental factors affecting product placement (such as lighting and background noise), as well as the product's position and orientation on the smart electronic scale, can impact image quality and recognition accuracy.
[0065] In summary, while some progress has been made in computer vision-based intelligent electronic scales, challenges and shortcomings remain, including issues with recognition accuracy, training sample requirements, real-time performance, and environmental factors. These problems need to be addressed to further improve the performance and reliability of intelligent electronic scales.
[0066] Therefore, this invention improves the fine-grained algorithm by employing a fine-grained recognition model for product type identification. This model includes a predetermined number of sequentially connected transformer modules. The first transformer module comprises a position encoding module and a TransFGC module connected in sequence. The second to last transformer modules each comprise a downsampling module and a TransFGC module connected in sequence. In the TransFGC module, a convolutional layer is configured after the last normalization layer to improve the accuracy and efficiency of product type identification. Furthermore, the downsampling module reduces computational complexity, thereby improving product type identification efficiency. Additionally, to avoid the influence of environmental factors and the position and posture of the product on the smart scale on product type identification, this embodiment uses data augmentation techniques to expand the dataset, obtaining samples under different environmental factors and the position and posture of the product on the smart scale.
[0067] Based on the above, this invention provides a method, apparatus, device, and storage medium for determining product settlement data. In this invention, a fine-grained recognition model is invoked to identify the type of multiple image blocks of the product to be weighed, thereby obtaining the type of the product. Based on the type and weight of the product, the product settlement data can be calculated. The fine-grained recognition model in this invention is trained on training data. Therefore, after training, compared to manual product type identification, the fine-grained recognition model can improve the accuracy of product type identification, thereby improving the accuracy of product settlement data calculation and enhancing the user experience. Furthermore, a convolutional layer is configured after the last normalization layer in the TransFGC module. That is, this invention uses a convolutional layer to improve the accuracy of product type identification, which in turn also improves the accuracy of product settlement data calculation.
[0068] It should be noted that the method, apparatus, device, and storage medium for determining product settlement data provided by this invention can be used in the fields of artificial intelligence or finance. The above are merely examples and do not limit the application areas of the method, apparatus, device, and storage medium for determining product settlement data provided by this invention.
[0069] Based on the above, one embodiment of the present invention provides a method for determining product settlement data, which can be applied to smart electronic scales. (Refer to...) Figure 1 A method for determining product settlement data may include:
[0070] S11. Obtain the initial image and weight of the product to be weighed.
[0071] In practical applications, when it is necessary to weigh a product, referred to as the product to be weighed in this embodiment, the product to be weighed is placed in the smart electronic scale. The image acquisition module in the smart electronic scale, such as a high-precision camera, will acquire an image of the product to be weighed as the initial image.
[0072] In addition, the smart electronic scale will collect the weight of the product to be weighed and display it on the screen.
[0073] It should be noted that the products to be weighed in this embodiment include, but are not limited to, agricultural and sideline products.
[0074] S12. Perform a segmentation operation on the initial image to obtain multiple image blocks.
[0075] In this embodiment, the initial image is segmented according to segmentation parameters. First, the segmentation parameters used for image segmentation are obtained. For example, if the image is to be divided into 16*16 blocks, the segmentation parameters are 16*16. Then, the initial image is segmented according to these parameters to obtain multiple image blocks. For instance, if the initial image is 224*224, segmenting it according to 16*16 will result in 14*14=196 small blocks, each of which is called an image block.
[0076] In this embodiment, the initial image is segmented into image blocks according to the segmentation parameters in order to meet the requirement that the input of the transformer module is the segmented image blocks, so that the transformer module can recognize the image blocks and process them.
[0077] S13. Call the fine-grained recognition model to identify the types of the multiple image blocks in order to obtain the types of the products to be weighed.
[0078] In practical applications, we will first introduce the fine-grained recognition model. (Refer to...) Figure 2 Taking vegetables as the input to the fine-grained recognition model, the model is segmented. In this embodiment, the segmentation is divided into 9 image blocks, which are then input into the fine-grained recognition model.
[0079] Fine-grained recognition models, such as the TransFGC model, consist of sequentially connected transformer modules of a predetermined number. Specifically, refer to Stage 1 through Stage N; each Stage contains one transformer module.
[0080] It should be noted that existing fine-grained recognition models typically use 9 transformer layers. However, in this embodiment of the invention, the target types are generally vegetables, meat, and most agricultural products. The original 9 transformer layers are somewhat overly complex, and too many parameters lead to longer training time, thus slowing down the detection speed and even causing model overfitting. To improve the target detection speed of vegetables, meat, and other agricultural products, and from the perspective of maintaining existing accuracy and simplifying the network structure, a fine-grained recognition model with fewer parameters and lower computational complexity is proposed. This fine-grained recognition model generally uses any number of transformer layers from 6 to 9. Generally, using 6 transformer layers can significantly reduce computational complexity, so the TransFGC model can retain only 6 transformer layers. That is, N is any number from 6 to 9, preferably 6. Since downsampling is added to the transformer layers to reduce data computational complexity, even when using 9 transformer layers, the computational complexity can still be reduced compared to existing technologies.
[0081] Reference Figure 2 Taking N as 6 as an example, the 6 transformer layers correspond to Stage1-StageN respectively. The first layer corresponds to Stage1 in the structure diagram, and the second to sixth layers correspond to StageN in the structure diagram.
[0082] The first transformer module includes position encoding modules connected in sequence ( Figure 2 The Linear Projection of Flattened Patches and the TransFGC module ( Figure 2 The Trans FGC Block in the middle.
[0083] In practical applications, the image (224*224 pixels) is first divided into 16*16 blocks, resulting in 14*14=196 blocks. Each block consists of RGB channels, making each block a 16*16*3=768-dimensional vector. Therefore, the 196 segmented image blocks form a 196*768 two-dimensional matrix. The segmented image blocks are then input into the `Linear Projection of Flattenend Patches`, which performs positional encoding on all image blocks, recording the position of each block. Therefore, `Linear Projection of Flattenend Patches` requires a 1*768 learnable matrix to record the position of each image block. Finally, a 197*768 matrix is output from `Linear Projection of Flattenend Patches`. This series of positional encoding embedding operations is shown in the following structure. Figure 3 As shown. It should be noted that, Figure 3 The example used is nine image blocks.
[0084] After the positional encoding operation, the encoded image patch is input into the TransFGC module of the TransFGC model, such as... Figure 2 The TransFGC module is located in the TransFGC Block. Its structure is as follows: Figure 4 As shown.
[0085] In practical applications, since the recognition targets of this invention are generally agricultural products such as vegetables and meat, and different vegetables have different varieties, in order to improve the recognition accuracy and speed, TransFGC Blocks introduces Conv 1d (1d convolutional layer) compared to the TransFGC module in the prior art. That is, in the TransFGC module of this embodiment, a convolutional layer is configured after the last normalization layer to improve the recognition accuracy and speed.
[0086] In the TransFGC module, the specific connection structures of Multi-head Attention, MLP (Multi-layer Perceptron), Layer Norm (Normalization Layer), and Conv 1d (1d Convolutional Layer) are detailed below. Figure 4 .
[0087] Figure 4 In the middle, the features of the input image are used as For example, firstly, through Layer Norm and Multi-head Attention, a portion of the output features are obtained. ,Will Make two identical copies, each as... Used as input to the two modules on the right side of the graph, one part passes through Layer Norm and MLP, and the other part passes through Layer Norm and Conv 1d, respectively, to obtain Perform feature merging operation to obtain .
[0088] In embodiments of the present invention, the second to the last transformer modules each include sequentially connected downsampling modules ( Figure 2 The Patch Merging layer and the TransFGC module. For the specific structure of the TransFGC module, please refer to [link to relevant documentation]. Figure 4 The specific workflow of the downsampling module in this embodiment can be found in [reference needed]. Figure 5 The downsampling module is used to reduce the depth of the feature map from C to C / 2.
[0089] It should be noted that in this embodiment, the number of small blocks that are input into the Stage in each transformer module is:
[0090]
[0091] in, Indicates the image height. Indicates the image width. Represents a stage, such as stage1. =1, This indicates the feature map depth.
[0092] In practical applications, the fine-grained recognition model is trained based on training data, which includes product image samples and category labels.
[0093] Specifically, you can acquire a large amount of training data and use it to train a fine-grained recognition model.
[0094] The type of product to be weighed can be obtained through the fine-grained recognition model described above. When using the fine-grained recognition model, the first step is to use the first transformer module, stage1, to perform position encoding. The position encoding result is then used for image filtering through Trans FGC Blocks to output the image block filtering result.
[0095] In this embodiment, the image filtering operation involves adjusting the weights of each image block according to... Filter out the corresponding number of data blocks.
[0096] Then, through the second transformer module, stage2, the downsampling module performs downsampling operations to reduce the amount of data. The sampling results are then input into Trans FGC Blocks for image filtering operations to update the image block filtering results.
[0097] The working logic of Trans FGC Blocks in stage 2 is the same as that of Trans FGC Blocks in stage 1.
[0098] Then, by sequentially passing through stage 3, stage 4, stage 5, and stage 6, the type of product to be weighed can be obtained.
[0099] It should be noted that in this embodiment, stages 2-6 have the same structure and the same working logic.
[0100] In addition, in this embodiment, the downsampling module is only set in stage2-stage6 and not in stage1 because the amount of data output from the position encoding module is small. The amount of data will only increase after passing through Trans FGC Blocks. Therefore, the downsampling module is set in stage2-stage6 after passing through Trans FGC Blocks to reduce the amount of data processing and reduce the computational complexity.
[0101] In summary, when the fine-grained recognition model is invoked to identify the types of the multiple image blocks in order to obtain the types of the product to be weighed, the first transformer module in the fine-grained recognition model performs position encoding and image filtering operations on the multiple image blocks, outputting the image block filtering results. The second to last transformer modules in the fine-grained recognition model then perform downsampling and image filtering operations on the image block filtering results to obtain the types of the product to be weighed.
[0102] In this embodiment, the downsampling module is used to reduce the amount of data processing and the computational complexity. At the same time, by setting convolutional layers in Trans FGC Blocks, the data processing speed and accuracy are improved. Thus, the fine-grained recognition model in this embodiment of the invention can quickly and accurately identify product types and has a high adaptability to similar varieties of the same product.
[0103] In another implementation of the present invention, the specific working logic of the downsampling module is given, specifically, refer to... Figure 6 Performing a downsampling operation on the image patch filtering results may include:
[0104] S21. Divide the image block filtering results to obtain multiple sub-images.
[0105] Specifically, refer to Figure 5 Suppose the image patch selection result input into Patch Merging is a 4x4 single-channel feature map. In this process, Patch Merging divides each 2x2 adjacent pixel into a patch, thus obtaining multiple sub-images. Figure 5 Four sub-images were obtained.
[0106] S22. Combine the pixel blocks located at the same position in the multiple sub-images to obtain a feature map.
[0107] Specifically, pixels at the same location (identified by the same location) in each patch are grouped together to obtain four new feature maps. For example, the pixels in the top left corner of each patch are grouped together, and the pixels in the top right corner are grouped together.
[0108] S23. Perform a stitching operation on the feature map in the depth direction to obtain the stitching result.
[0109] Specifically, these four feature maps are concatted together along the depth direction to obtain the concatenated result.
[0110] S24. The splicing result is normalized and pooled to obtain the sampling result.
[0111] Specifically, the concatenated result is processed through a Layer Norm layer. Finally, two pooling operations are used to extract important information in the depth direction of the feature map, reducing the depth of the feature map from C to C / 2.
[0112] By using downsampling in this embodiment, the amount of data can be reduced by half, thereby speeding up data processing and improving the efficiency of category identification.
[0113] In another implementation of the present invention, in order to further extract better features and improve classification efficiency, a feature extraction module, Part Selection Module, can be configured between the last transformer module and the penultimate transformer module, as detailed below. Figure 7 .
[0114] In practical applications, the image block filtering results are sequentially downsampled and filtered through the second to penultimate transformer modules in the fine-grained recognition model to obtain new image block filtering results. The new image block filtering results are then processed by the feature extraction module to extract features, and finally, the feature extraction results are downsampled and filtered through the last transformer module to obtain the type of the product to be weighed.
[0115] Specifically, the data processed through Stage 1 to Stage 5, combined with the position weight results of the first 5 Trans FGC Blocks layers, retains only the top 12 image blocks with higher activation levels. These image blocks and their position codes are then output to Stage 6 via the Part Selection Module, ultimately outputting the category recognition results.
[0116] In this embodiment, by setting a feature extraction module between Stage 5 and Stage 6, further feature extraction can be performed to extract features that can be quickly and accurately identified, thereby accelerating the efficiency of category identification.
[0117] Finally, the fine-grained recognition model combines six TransFGC Blocks (Stage 1 to Stage 6) to identify the categories and output the results. Each input is further refined according to the value of N in Stage N.
[0118] S14. Based on the type and weight of the product to be weighed, calculate the product settlement data of the product to be weighed.
[0119] In practical applications, after obtaining the type and weight of the product to be weighed, the corresponding price for that type can be obtained. Specifically, large agricultural markets establish a unified pricing system, and a price is set in this system for each type of agricultural product entering the market. The smart scale automatically retrieves the preset price for that category from the unified pricing system based on the type identification result. The total price of the vegetables is calculated based on the weight of the agricultural product obtained by the smart scale and its associated price. This can be done through multiplication, i.e., multiplying the vegetable weight by the preset price for that category. Finally, the calculated total price is displayed on the smart scale's screen, or it can be output in other ways, such as printing a receipt, voice announcement, or data interaction with other systems. The total price of the vegetables is the product settlement data in this embodiment.
[0120] It should be noted that in practical applications, the fine-grained recognition model may incorrectly identify the product category. To avoid errors in calculating the total product price, the category identified by the fine-grained model can be displayed on the smart scale's screen for manual verification. If correct, no adjustment is needed; otherwise, manual adjustment is required. Then, the initial image of the product to be weighed and the manually corrected category are used as samples to retrain the fine-grained recognition model, thereby improving recognition accuracy.
[0121] In this embodiment, a fine-grained recognition model is invoked to identify the type of multiple image blocks of the product to be weighed, thereby determining the type of the product. Based on the type and weight of the product, the product settlement data can be calculated. The fine-grained recognition model in this invention is trained on training data. Therefore, after training, compared to manual product type identification, the fine-grained recognition model can improve the accuracy of product type identification, thereby improving the accuracy of product settlement data calculation and enhancing the user experience. Furthermore, a convolutional layer is configured after the last normalization layer in the TransFGC module. This invention uses a convolutional layer to improve the accuracy of product type identification, which in turn also improves the accuracy of product settlement data calculation.
[0122] The fine-grained recognition model in this invention can identify products with similar appearances within the same variety, improving the accuracy of category identification. Furthermore, the downsampling module in this invention reduces the amount of data computation, thereby lowering computational complexity. Additionally, by configuring a convolutional layer after the last normalization layer in the TransFGC module, recognition efficiency and accuracy are improved.
[0123] The fine-grained recognition model described above is obtained through training. This embodiment describes the generation process of the fine-grained recognition model, specifically including the following steps:
[0124] S31. Obtain training data.
[0125] The training data includes product image samples and category labels.
[0126] Product image samples can be referenced. Figure 2 The vegetable image on the far left has a category label indicating the type of product sample, such as Kyoho grapes.
[0127] In practical applications, a fine-grained recognition model can be trained for different agricultural products, such as using the same model for vegetables and fruits. Alternatively, separate models can be trained for different categories, such as one fine-grained recognition model for fruits and another for vegetables. The specific number of fine-grained recognition models trained can be determined based on actual configuration. This embodiment uses the same fine-grained recognition model for vegetables and fruits as an example.
[0128] Training data can form a dataset, which includes product image samples not only of various agricultural and sideline products in the agricultural market, but also a portion of vegetables with soil on them. First, a large amount of agricultural product image data is collected, covering different varieties and appearances of various vegetables and meats. Suitable image data can be obtained by taking photos of actual agricultural product samples or from sources such as the internet, and the collected image data is labeled (labeling object location and category). To ensure data diversity and balance, the collected data includes image samples of agricultural products with various varieties, sizes, shapes, colors, and ripeness. The collected image samples are cleaned and preprocessed, including noise removal, image resizing, and contrast adjustment, to ensure data quality and consistency.
[0129] It should be noted that, in order to improve the efficiency of data annotation, images of the same type can be placed in the same folder and only the folder can be labeled with the type. This can improve the annotation efficiency compared to labeling each image individually.
[0130] S32. Expand the training data to obtain training samples.
[0131] In practical applications, the color and texture of goods may appear differently under different lighting conditions, potentially affecting the accuracy of the recognition results. At the same time, the product's position or posture on the smart scale can also influence the recognition results; for example, if the product is partially obscured or in an abnormal posture, it may lead to the failure or error of the recognition algorithm.
[0132] Therefore, to avoid the impact of environmental factors and the position and posture of the products on the smart electronic scale on image quality and recognition accuracy, this embodiment uses data augmentation techniques to expand the dataset, increasing the diversity and quantity of samples. For example, operations such as random cropping, rotation, scaling, and flipping can be performed to generate more training samples.
[0133] In this embodiment, for newly emerging product categories, data augmentation technology can quickly obtain a large number of data samples for that product category, avoiding the problem of low identification accuracy due to a lack of samples.
[0134] S33. Use the training samples to train a fine-grained recognition model until the preset training stop condition is met.
[0135] The final dataset is divided into training, validation, and test sets. The training set is used for model training and parameter optimization, the validation set is used for model selection and hyperparameter tuning, and the test set is used to evaluate the model's performance and accuracy.
[0136] During model training, the loss function used in training the fine-grained recognition model is calculated using the following formula:
[0137]
[0138] in, For loss function, For hyperparameters, For weighted cross-entropy loss, To predict product categories, For category labels, To compare the feature learning parameters, where,
[0139]
[0140] in, and It is a constant. =0.5, This refers to the batch size, specifically 32. , To identify different image blocks, This represents the token's position within the input category, i.e., its location information. For image blocks Location information, For image blocks Location information, For image blocks Predicted product categories For image blocks Predicted product categories.
[0141] Due to the imbalanced effect of the dataset, and to accelerate model convergence and prevent overfitting, hyperparameters are introduced here. This further accelerates the model's computation speed and ensures its real-time performance.
[0142] In this embodiment, the current network training will be stopped when a preset training stop condition is met. The preset training condition may be that the number of training iterations reaches a set number, or the loss function value is less than a set value.
[0143] Once the model is trained, fine-grained identification of agricultural and sideline products can be achieved. The final output is the top three vegetable varieties with the highest confidence level, with the one with the highest confidence level being the first choice and the other two being the alternative varieties.
[0144] In this embodiment, adding hyperparameters to the model's loss function can speed up the model's calculation, ensure the model's real-time performance, and also avoid overfitting.
[0145] By expanding the training samples in this embodiment, the robustness of the model is increased, allowing it to adapt to different complex environments. This enables accurate identification of product types in various environments, or under different postures and positions on the smart electronic scale, thus improving the accuracy of product type identification.
[0146] In practical applications, in order to enable those skilled in the art to better understand the present invention, the calculation process of the entire product settlement data in the present invention will now be described.
[0147] Reference Figure 9 In steps S41-S48, the user-selected vegetables are placed on the electronic scale. The smart electronic scale uses the aforementioned fine-grained recognition model to identify the vegetable type. After identifying the vegetable type, the user manually checks its accuracy. If correct, the price of the vegetables is obtained, the total price is calculated, and the result is output. If incorrect, the user checks if the model's output of alternative varieties is correct. If correct, the product type is corrected. If the alternative varieties are also incorrect, the user manually selects a vegetable type, corrects the variety, obtains the price, calculates the total price, and outputs the result.
[0148] In this embodiment, by combining a fine-grained recognition algorithm with a smart electronic scale, the present invention can automatically identify the types of products such as vegetables and meat, obtain the uniform preset prices for vegetable and meat varieties in the pricing system of large-scale agricultural markets, and then calculate the final transaction price in real time. This smart electronic scale can improve the accuracy and efficiency of the transaction process, reduce manual calculation costs, and provide a better transaction experience for merchants and consumers, possessing broad application prospects and commercial value.
[0149] Optionally, based on the above-described embodiment of a method for determining product settlement data, another embodiment of the present invention provides a device for determining product settlement data, referring to... Figure 10 It can include:
[0150] Data acquisition module 11 is used to acquire the initial image and weight of the product to be weighed;
[0151] Image segmentation module 12 is used to segment the initial image to obtain multiple image blocks;
[0152] The category identification module 13 is used to call a fine-grained identification model to identify the category of the multiple image blocks in order to obtain the category of the product to be weighed. The fine-grained identification model includes a preset number of transformer modules connected in sequence. The first transformer module includes a position encoding module and a TransFGC module connected in sequence. The second to the last transformer modules each include a downsampling module and a TransFGC module connected in sequence. In the TransFGC module, a convolutional layer is configured after the last normalization layer. The fine-grained identification model is trained based on training data. The training data includes product image samples and category labels.
[0153] The data calculation module 14 is used to calculate the product settlement data of the product to be weighed based on the type and weight of the product to be weighed.
[0154] Furthermore, the image segmentation module 12 is specifically used for:
[0155] The segmentation parameters used in image segmentation are obtained, and the initial image is segmented according to the segmentation parameters to obtain multiple image blocks.
[0156] Furthermore, the category identification module 13 is specifically used for:
[0157] The first transformer module in the fine-grained recognition model performs position encoding and image filtering operations on the multiple image blocks, outputting the image block filtering results. The second to last transformer modules in the fine-grained recognition model then perform downsampling and image filtering operations on the image block filtering results to obtain the type of the product to be weighed.
[0158] Furthermore, a downsampling operation is performed on the image patch filtering results, including:
[0159] The image block filtering results are divided into multiple sub-images;
[0160] The pixel blocks located at the same position in the multiple sub-images are combined to obtain a feature map;
[0161] The feature maps are stitched together along the depth direction to obtain the stitching result.
[0162] The splicing result is normalized and pooled to obtain the sampling result.
[0163] Furthermore, a feature extraction module is configured between the last transformer module and the penultimate transformer module;
[0164] The image patch selection results are sequentially processed through the second to last transformer modules in the fine-grained recognition model to perform downsampling and image selection operations, thereby obtaining the types of the products to be weighed, including:
[0165] The image patch selection results are then subjected to downsampling and image selection operations sequentially through the second to second-to-last transformer modules in the fine-grained recognition model to obtain new image patch selection results.
[0166] The new image patch filtering results are processed by the feature extraction module to extract features, resulting in feature extraction results.
[0167] The last transformer module performs downsampling and image filtering operations on the feature extraction results to obtain the type of the product to be weighed.
[0168] Furthermore, it also includes a model generation module for:
[0169] Acquire training data; the training data includes product image samples and category labels;
[0170] The training data is augmented to obtain training samples;
[0171] The fine-grained recognition model is trained using the training samples until a preset training stopping condition is met.
[0172] Furthermore, the formula for calculating the loss function used during the training of the fine-grained recognition model is as follows:
[0173]
[0174] in, For loss function, For hyperparameters, For weighted cross-entropy loss, To predict product categories, For category labels, To compare the feature learning parameters, where,
[0175]
[0176] in, and It is a constant. , To identify different image blocks, For image blocks Location information, For image blocks Location information, For image blocks Predicted product categories For image blocks Predicted product categories.
[0177] In this embodiment, a fine-grained recognition model is invoked to identify the type of multiple image blocks of the product to be weighed, thereby determining the type of the product. Based on the type and weight of the product, the product settlement data can be calculated. The fine-grained recognition model in this invention is trained on training data. Therefore, after training, compared to manual product type identification, the fine-grained recognition model can improve the accuracy of product type identification, thereby improving the accuracy of product settlement data calculation and enhancing the user experience. Furthermore, a convolutional layer is configured after the last normalization layer in the TransFGC module. This invention uses a convolutional layer to improve the accuracy of product type identification, which in turn also improves the accuracy of product settlement data calculation.
[0178] It should be noted that the working process of each module in this embodiment is explained in the above description, and will not be repeated here.
[0179] Corresponding to the above embodiments, such as Figure 11 As shown, the present invention also provides an electronic device, which may include: a processor 1 and a memory 2;
[0180] The processor 1 and memory 2 communicate with each other via communication bus 3.
[0181] Processor 1, for executing at least one instruction;
[0182] Memory 2 is used to store at least one instruction; the instruction is stored in a computer program.
[0183] Processor 1 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement embodiments of the present invention.
[0184] Memory 2 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk storage device.
[0185] The processor executes at least one instruction to implement the above-described method for determining product settlement data.
[0186] In this embodiment, a fine-grained recognition model is invoked to identify the type of multiple image blocks of the product to be weighed, thereby determining the type of the product. Based on the type and weight of the product, the product settlement data can be calculated. The fine-grained recognition model in this invention is trained on training data. Therefore, after training, compared to manual product type identification, the fine-grained recognition model can improve the accuracy of product type identification, thereby improving the accuracy of product settlement data calculation and enhancing the user experience. Furthermore, a convolutional layer is configured after the last normalization layer in the TransFGC module. This invention uses a convolutional layer to improve the accuracy of product type identification, which in turn also improves the accuracy of product settlement data calculation.
[0187] Corresponding to the above embodiments, another embodiment of the present invention provides a computer-readable storage medium that stores at least one instruction, which, when executed by a processor, implements the above-described method for determining product settlement data.
[0188] In this embodiment, a fine-grained recognition model is invoked to identify the type of multiple image blocks of the product to be weighed, thereby determining the type of the product. Based on the type and weight of the product, the product settlement data can be calculated. The fine-grained recognition model in this invention is trained on training data. Therefore, after training, compared to manual product type identification, the fine-grained recognition model can improve the accuracy of product type identification, thereby improving the accuracy of product settlement data calculation and enhancing the user experience. Furthermore, a convolutional layer is configured after the last normalization layer in the TransFGC module. This invention uses a convolutional layer to improve the accuracy of product type identification, which in turn also improves the accuracy of product settlement data calculation.
[0189] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0190] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.
[0191] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for determining product settlement data, characterized in that, include: Obtain the initial image and weight of the product to be weighed; The initial image is segmented to obtain multiple image blocks; The fine-grained recognition model uses the first transformer module to perform position encoding and image filtering operations on multiple image blocks, outputting image block filtering results. The second to penultimate transformer modules in the model then perform downsampling and image filtering operations on the image block filtering results to obtain new image block filtering results. These new results are then processed by the feature extraction module to extract features. Finally, the last transformer module performs downsampling and image filtering operations on the feature extraction results to determine the type of product to be weighed. The fine-grained recognition model includes a predetermined number of sequentially connected transformer modules. The first transformer module includes a sequentially connected position encoding module and a TransFGC module. The second to last transformer modules each include a sequentially connected downsampling module and a TransFGC module. A feature extraction module is configured between the last and penultimate transformer modules. In the TransFGC module, a convolutional layer is configured after the last normalization layer. The fine-grained recognition model is trained based on training data. The training data includes product image samples and category labels; wherein, the downsampling operation on the image block selection results includes: dividing the image block selection results into multiple sub-images; combining pixel blocks located at the same position in the multiple sub-images to obtain a feature map; performing a stitching operation on the feature map in the depth direction to obtain a stitching result; and performing normalization and pooling operations on the stitching result to obtain a sampling result; Based on the type and weight of the product to be weighed, calculate the product settlement data of the product to be weighed; The TransFGC module includes: a multi-head attention layer, a multi-layer perceptron, a normalization layer, and a 1D convolutional layer; the image features input to the TransFGC module are... First, a normalization layer and a multi-head attention layer are used to obtain a portion of the output features. ,Will Make two identical copies, each as... The input is used as the input to two modules. One part passes through the normalization layer and multilayer perceptron of one module, and the other part passes through the normalization layer and 1D convolutional layer of the other module, yielding the outputs of the two modules respectively. , to output the two The output of the TransFGC module is obtained by performing a feature merging operation. ; The formula for calculating the loss function used during the training of the fine-grained recognition model is as follows: Where L is the loss function and w is the hyperparameter. The weighted cross-entropy loss is y, where y is the predicted product category. For category labels, To compare the feature learning parameters, where, in, And N are constants. To identify different image blocks, For image blocks Location information, For image blocks Location information, For image blocks Predicted product categories For image blocks Predicted product categories.
2. The determination method according to claim 1, characterized in that, The initial image is segmented to obtain multiple image patches, including: Obtain the segmentation parameters used during image segmentation; The initial image is segmented according to the segmentation parameters to obtain multiple image blocks.
3. The determination method according to claim 1, characterized in that, The generation process of the fine-grained recognition model includes: Acquire training data; the training data includes product image samples and category labels; The training data is augmented to obtain training samples; The fine-grained recognition model is trained using the training samples until a preset training stopping condition is met.
4. A device for determining product settlement data, characterized in that, include: The data acquisition module is used to acquire the initial image and weight of the product to be weighed; The image segmentation module is used to segment the initial image to obtain multiple image blocks; The category identification module is used to perform position encoding and image filtering operations on the multiple image blocks through the first transformer module in the fine-grained identification model, and output the image block filtering result; then, through the second to second-to-last transformer modules in the fine-grained identification model, downsampling and image filtering operations are performed on the image block filtering result to obtain a new image block filtering result; the new image block filtering result is then processed by the feature extraction module to extract features, and finally, through the last transformer module, downsampling and image filtering operations are performed on the feature extraction result to obtain the category of the product to be weighed; the fine-grained identification model includes a predetermined number of sequentially connected transformer modules, the first transformer module includes a sequentially connected position encoding module and a TransFGC module, and the second to last transformer modules each include a sequentially connected downsampling module and a TransFGC module; a feature extraction module is configured between the last transformer module and the second-to-last transformer module; in the TransFGC module, a convolutional layer is configured after the last normalization layer, and the fine-grained identification model is trained based on training data; The training data includes product image samples and category labels; wherein, the downsampling operation on the image block selection results includes: dividing the image block selection results into multiple sub-images; combining pixel blocks located at the same position in the multiple sub-images to obtain a feature map; performing a stitching operation on the feature map in the depth direction to obtain a stitching result; and performing normalization and pooling operations on the stitching result to obtain a sampling result; The data calculation module is used to calculate the product settlement data of the product to be weighed based on the type and weight of the product to be weighed. The TransFGC module includes: a multi-head attention layer, a multi-layer perceptron, a normalization layer, and a 1D convolutional layer; the image features input to the TransFGC module are... First, a normalization layer and a multi-head attention layer are used to obtain a portion of the output features. ,Will Make two identical copies, each as... The input is used as the input to two modules. One part passes through the normalization layer and multilayer perceptron of one module, and the other part passes through the normalization layer and 1D convolutional layer of the other module, yielding the outputs of the two modules respectively. , to output the two The output of the TransFGC module is obtained by performing a feature merging operation. ; The formula for calculating the loss function used during the training of the fine-grained recognition model is as follows: Where L is the loss function and w is the hyperparameter. The weighted cross-entropy loss is y, where y is the predicted product category. For category labels, To compare the feature learning parameters, where, in, And N are constants. To identify different image blocks, For image blocks Location information, For image blocks Location information, For image blocks Predicted product categories For image blocks Predicted product categories.
5. An electronic device, characterized in that, The electronic device includes a memory and a processor; The memory is used to store at least one instruction; The processor is used to execute the at least one instruction to implement a method for determining product settlement data as described in any one of claims 1 to 3.
6. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores at least one instruction, which, when executed by a processor, implements a method for determining product settlement data as described in any one of claims 1 to 3.