Article image acquisition method and device for weighing equipment, equipment and medium

Through the deep learning model, the item images on the weighing equipment are classified and the plastic bag feature recognition is identified, and the available images are screened out, which solves the problem of image quality and lack of data caused by the items on the weighing equipment being blocked by the plastic bag, and improves the image acquisition efficiency.

CN120236143AActive Publication Date: 2025-07-01杭州食方科技有限公司 +1
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510532602.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-07-01
Estimated Expiration
2045-04-25

AI Technical Summary

Technical Problem

The items on the weighing equipment are blocked by plastic bags, resulting in a decrease in image quality, affecting the database image quality and losing effective data. The prior art has failed to effectively screen out available images.

Method used

The deep learning model is used to classify the object images on the weighing equipment and identify the plastic bag feature. The plastic bag occlusion ratio information is used to determine whether the image meets the retention conditions and filter out the available images.

Benefits of technology

While ensuring image quality, the number of available images collected is increased, the image acquisition efficiency is improved, and the data lack of diversity caused by plastic bag occlusion is avoided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120236143A_ABST
    Figure CN120236143A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an article image acquisition method and device for weighing equipment, equipment and a medium. A specific embodiment of the method comprises the following steps: shooting an object on weighing equipment to obtain an initial object image; inputting the initial article image into a pre-trained image classification model to obtain an initial article image classification result; in response to determining that the classification result of the initial article image represents that the plastic bag is shielded, determining that the initial article image meets a temporary retention condition; in response to determining that the initial article image meets the temporary retention condition, performing plastic bag feature recognition on the initial article image, and determining whether the initial article image meets the retention condition; in response to determining that the initial article image satisfies the retention condition, the initial article image is determined as the acquired article image. According to the embodiment, while the quality of the acquired article images is ensured, the number of the effectively acquired article images is increased, and the efficiency of acquiring the available article images is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the field of computer technology, and in particular, to a method, apparatus, device, and medium for collecting item images for weighing devices. Background Art

[0002] During the establishment of an item database, the types of items are rich and the image angles are diverse, showing great variety. Therefore, a large number of real item images need to be collected. Currently, most food scales use the method of adding a camera on the scale to collect item images.

[0003] However, when collecting item images in the above manner, the following technical problems often exist: Items on the food scale are often blocked by plastic bags. If directly added to the database, it will reduce the image quality of the database and affect the corresponding model training. If directly screened out completely, a large amount of valid data will be lost, resulting in a lack of diversity in the item image data.

[0004] The above information disclosed in this background art section is only used to enhance the understanding of the background of the concept of the present disclosure, and thus, it may include information that does not form the prior art known to those of ordinary skill in the art. Summary of the Invention

[0005] This summary of the present disclosure is used to introduce concepts in a brief form, which will be described in detail in the subsequent detailed implementation section. This summary of the present disclosure is not intended to identify the key features or essential features of the claimed technical solution, nor is it intended to be used to limit the scope of the claimed technical solution.

[0006] Some embodiments of the present disclosure propose a method, apparatus, electronic device, and computer-readable medium for collecting item images for weighing devices to solve one or more of the technical problems mentioned in the above background art section.

[0007] In a first aspect, some embodiments of the present disclosure provide an article image acquisition method for a weighing device. The method includes: photographing an object on the weighing device to obtain an initial article image; inputting the initial article image into a pre-trained image classification model to obtain an initial article image classification result; in response to determining that the initial article image classification result indicates that there is a plastic bag occlusion, determining that the initial article image meets the temporary retention condition; in response to determining that the initial article image meets the temporary retention condition, performing plastic bag feature recognition on the initial article image to obtain plastic bag feature information corresponding to the initial article image, where the plastic bag feature information includes plastic bag occlusion ratio information; determining whether the initial article image meets the retention condition according to the plastic bag occlusion ratio information included in the plastic bag feature information; and in response to determining that the initial article image meets the retention condition, determining the initial article image as the acquired article image.

[0008] In a second aspect, some embodiments of the present disclosure provide an article image acquisition device for a weighing device. The device includes: a photographing unit configured to photograph an object on the weighing device to obtain an initial article image; a classification unit configured to input the initial article image into a pre-trained image classification model to obtain an initial article image classification result; a first determination unit configured to determine that the initial article image meets the temporary retention condition in response to determining that the initial article image classification result indicates that there is a plastic bag occlusion; a second determination unit configured to perform plastic bag feature recognition on the initial article image to obtain plastic bag feature information corresponding to the initial article image in response to determining that the initial article image meets the temporary retention condition, where the plastic bag feature information includes plastic bag occlusion ratio information; a third determination unit configured to determine whether the initial article image meets the retention condition according to the plastic bag occlusion ratio information included in the plastic bag feature information; and a fourth determination unit configured to determine the initial article image as the acquired article image in response to determining that the initial article image meets the retention condition.

[0009] In a third aspect, some embodiments of the present disclosure provide an electronic device, including: one or more processors; a storage device storing one or more programs thereon, and when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any implementation manner of the first aspect above.

[0010] In a fourth aspect, some embodiments of the present disclosure provide a computer-readable medium storing a computer program thereon, where the program, when executed by a processor, implements the method described in any implementation manner of the first aspect above.

[0011] The above embodiments of the present disclosure have the following beneficial effects: Through an article image acquisition method for a weighing device of the present disclosure, the plastic bag information in a picture containing a plastic bag can be predicted, and the retention conditions of the acquired article images can be set by itself according to a downstream model (such as article recognition). Thus, pictures containing plastic bags but with high usability can be flexibly screened out. Specifically, the reason why most article image acquisition methods cannot flexibly select images with high usability is that most article image acquisition methods do not further screen the article pictures blocked by plastic bags, resulting in a small number and poor quality of the acquired article images. Based on this, the article image acquisition method for a weighing device in some embodiments of the present disclosure includes: First, photograph the object on the weighing device to obtain an initial article image. Thus, the article image to be processed is obtained. Then, input the above initial article image into a pre-trained image classification model to obtain an initial article image classification result. Thus, the article can be roughly classified for the first time through the first model. Then, in response to determining that the above initial article image classification result indicates that there is a plastic bag blockage, determine that the above initial article image meets the temporary retention condition. Thus, a basis is laid for further classification later. Then, in response to determining that the above initial article image meets the above temporary retention condition, perform plastic bag feature recognition on the above initial article image to obtain plastic bag feature information corresponding to the above initial article image, wherein the above plastic bag feature information includes plastic bag occlusion ratio information. Then, determine whether the above initial article image meets the retention condition according to the plastic bag occlusion ratio information included in the above plastic bag feature information. Finally, in response to determining that the above initial article image meets the above retention condition, determine the above initial article image as the acquired article image. Thus, the data cleaning of the acquired article images is completed. The above article image acquisition method for a weighing device uses two deep learning models to perform data cleaning on the acquired article images, and filters out available article images by using preset conditions. Thus, while ensuring the quality of the acquired article images, the number of the acquired article images is increased as much as possible, and the efficiency of acquiring available article images is improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] In combination with the accompanying drawings and referring to the following specific embodiments, the above and other features, advantages and aspects of the various embodiments of the present disclosure will become more obvious. Throughout the accompanying drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic, and the elements and elements are not necessarily drawn to scale.

[0013] Figure 1 is a flowchart of some embodiments of an article image acquisition method for a weighing device according to the present disclosure; Figure 2 is a schematic structural diagram of some embodiments of an article image acquisition device for a weighing device according to the present disclosure; Figure 3 It is a schematic structural diagram of an electronic device suitable for implementing some embodiments of the present disclosure. Detailed implementation manners

[0014] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.

[0015] In addition, it should be noted that for the sake of convenience of description, only parts related to the relevant invention are shown in the drawings. Without conflict, the embodiments in the present disclosure and the features in the embodiments can be combined with each other.

[0016] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependent relationships.

[0017] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive. Those skilled in the art should understand that unless clearly specified otherwise in the context, it should be understood as "one or more".

[0018] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.

[0019] The present disclosure will be described in detail below with reference to the drawings and in combination with embodiments.

[0020] Figure 1 It is a flowchart of some embodiments of an article image acquisition method for a weighing device according to the present disclosure. The article image acquisition method for the weighing device includes the following steps: Step 101, photograph an object on the weighing device to obtain an initial article image.

[0021] In some embodiments, the execution subject of the method for acquiring an item image for a weighing device (such as a food scale with a camera) can obtain an initial item image from an integrated device through a wired connection or a wireless connection, and can connect the above initial item image to a server through a wired connection or a wireless connection for processing. The above item image can be acquired by setting a camera on the food scale. The above camera can be a fixed-focus camera or a high-definition wide-angle camera, without specific limitation. It should be noted that the above wireless connection method can include but is not limited to 3G / 4G connection, WiFi connection, Bluetooth connection, WiMAX connection, Zigbee connection, UWB (ultra wideband) connection, and other currently known or future-developed wireless connection methods. The above initial item image can include but is not limited to: food image or dish image, without specific limitation here.

[0022] Optionally, the above subject can also perform the following steps: First step, save the above initial item image locally. The above execution subject can temporarily save the above initial item image in the above execution subject. The above saving method can include but is not limited to: local flash memory or SD card, without specific limitation.

[0023] Second step, upload each initial item image saved within a preset time period to the server. The above preset time period can be set as needed. For example, it can be to upload all the above item images collected within the current month once a month, and after the upload is completed, automatically clear the uploaded above item images to prevent the local memory from being full. The above cloud server can be a private cloud server, without specific limitation. The above execution subject can receive item images uploaded by different food scales and distinguish them by food scale number. The above execution subject can distinguish the received item images by time period, and users can download them as needed.

[0024] Step 102, input the initial item image into a pre-trained image classification model to obtain an initial item image classification result.

[0025] In some embodiments, the above execution subject can input the above initial item image into a pre-trained image classification model to obtain an initial item image classification result. Among them, the above image classification model is a neural network model that takes the initial item image as input and the initial item image classification result as output.

[0026] In practice, the above execution subject can classify the above initial item image into four categories. For example, the above initial item image is classified into: normal image, image blocked by a plastic bag, unclear image, and image blocked by foreign objects.

[0027] In some alternative implementation manners of some embodiments, the above-mentioned execution subject may input the initial item image into a pre-trained image classification model through the following steps to obtain an initial item image classification result: In the first step, input the initial object image into the classification convolutional network included in the above-mentioned pre-trained image classification model to obtain a classification feature map. Among them, the above-mentioned pre-trained image classification model may further include a classification model fully connected layer and a classification model output layer. The above-mentioned classification convolutional network may include a first classification model convolutional layer set, a second classification model convolutional layer set, a third classification model convolutional layer set, a fourth classification model convolutional layer set, and a fifth classification model convolutional layer set. Each convolutional layer in the above-mentioned first classification model convolutional layer set may be provided with a first preset number of convolutional kernels. Each convolutional layer in the above-mentioned second classification model convolutional layer set may be provided with a second preset number of convolutional kernels. The above-mentioned second preset number may be greater than the above-mentioned first preset number. Each convolutional layer in the above-mentioned third classification model convolutional layer set may be provided with a third preset number of convolutional kernels. The above-mentioned third preset number may be greater than the above-mentioned second preset number. Each convolutional layer in the above-mentioned fourth classification model convolutional layer set may be provided with a fourth preset number of convolutional kernels. The above-mentioned fourth preset number may be greater than the above-mentioned third preset number. Each convolutional layer in the above-mentioned fifth classification model convolutional layer set may be provided with a fifth preset number of convolutional kernels. The above-mentioned fifth preset number may be equal to the above-mentioned fourth preset number. The above-mentioned first classification model convolutional layer set may include two convolutional layers with a convolutional kernel of 3*3 and a channel number of 64 and one pooling layer. In practice, the above-mentioned execution entity may extract local features of the above-mentioned initial object image, such as edges and textures, through the convolutional layers in the above-mentioned first classification model convolutional layer set. The pooling layer in the above-mentioned first classification model convolutional layer set may perform a downsampling operation on the above-mentioned local features to reduce the amount of data, while retaining the main features and preventing overfitting. The above-mentioned second classification model convolutional layer set may include two convolutional layers with a convolutional kernel of 3*3 and a channel number of 128 and one pooling layer. The convolutional layers in the above-mentioned second classification model convolutional layer set may be used to extract deep features, such as shapes and colors, and perform a downsampling operation through the above-mentioned pooling layer. The above-mentioned third classification model convolutional layer set may include four convolutional layers with a convolutional kernel of 3*3 and a channel number of 256 and one pooling layer. The above-mentioned third classification model convolutional layer set may be used to further increase the number of channels to extract deeper features, such as combinations of shapes, and also perform a downsampling operation through the pooling layer. Both the above-mentioned fourth classification model convolutional layer set and the fifth classification model convolutional layer set may include four convolutional layers with a convolutional kernel of 3*3 and a channel number of 512 and one pooling layer. The functions of the above-mentioned fourth classification model convolutional layer set and the above-mentioned fifth classification model convolutional layer set are both to extract deeper features, such as combination features of more shapes, using convolutional layers with further expanded channel numbers. The pooling layer of the above-mentioned fourth classification model convolutional layer set is connected to the above-mentioned fifth classification model convolutional layer set to maximize the feature extraction functions of the two convolutional layer sets.It should be noted that the number of convolutional kernels of each of the above convolutional kernels and the number of convolutional layers in each of the above convolutional layer sets can be adjusted according to requirements, and no specific limitations are made here.

[0028] In the second step, input the above classification feature map into the fully connected layer of the above classification model to obtain item anomaly detection classification information. Among them, the fully connected layer of the above classification model may include two fully connected layers with 4096 neurons and one fully connected layer with 1000 neurons.

[0029] In practice, the above execution entity can integrate and further process the features extracted by the convolutional layer and the pooling layer through the fully connected layer in the fully connected layer of the above classification model, then perform non-linear transformation and map the features to specific categories to obtain the final classification result. For example, the above item anomaly detection classification information can be judged and classified according to the following criteria: 1. Whether there is a foreign object (non-plastic bag) blocking (the value is set between 0 and 1, and the closer it is to 0, the higher the probability of no foreign object blocking) 2. Whether there is a plastic bag blocking (the value is set between 0 and 1, and the closer it is to 0, the higher the probability of no plastic bag blocking) 3. Whether it is clear (the value is set between 0 and 1, and the closer it is to 1, the higher the probability that the image is clear) 4. According to the three outputs of each picture, set a threshold to screen pictures (for example, if the clarity value is less than 0.8, it is determined to be an unclear picture; if the foreign object blocking value is greater than 0.6, it is determined to be a picture blocked by a foreign object; if the plastic bag blocking value is greater than 0.7, it is determined to be a picture blocked by a plastic bag; the rest of the pictures are normal pictures).

[0030] In the third step, input the above item anomaly detection classification information into the output layer of the above classification model to obtain the initial item image classification result. Among them, the output layer of the above classification model may be a Sigmoid output layer. The output layer of the above classification model can be used to finally classify or predict the input above item image and output the corresponding result. For example, output image 1 and attach the label "normal" to it, and automatically classify it into the "normal image collection".

[0031] In some optional implementation manners of some embodiments, the above execution entity can input the initial item image into the classification convolutional network included in the above pre-trained image classification model through the following steps to obtain a classification feature map: In the first step, input the above initial item image into the first classification model convolutional layer set to obtain the first feature map of the initial item image. Among them, the information included in the above first feature map may be local features of the above initial item image, such as edges, textures, etc.

[0032] In the second step, input the above first feature map into the convolutional layer set of the second classification model to obtain a second feature map. Among them, the information contained in the above second feature map can be relatively abstract features of the above initial object image, such as shape, color, etc.

[0033] In the third step, input the above second feature map into the convolutional layer set of the third classification model to obtain a third feature map. Among them, the information contained in the above third feature map can be the depth features of the above initial object image, such as the combined relationship of shapes, etc.

[0034] In the fourth step, input the above third feature map into the convolutional layer set of the fourth classification model to obtain a fourth feature map. Among them, the information contained in the above fourth feature map can be deeper features of the above initial object image, such as the combined relationship of shapes, etc.

[0035] In the fifth step, input the above fourth feature map into the convolutional layer set of the fifth classification model to obtain a fifth feature map as the classification feature map. Among them, the information contained in the above fifth feature map can be further deeper features of the above initial object image, such as more combined relationships of shapes, etc.

[0036] Optionally, the above pre-trained image classification model is trained through the following steps: In the first step, obtain a set of initial object image samples with pre-annotations. Among them, each initial object image sample in the above set of initial object image samples can correspond to image category information. The above image category information can include each label value of the corresponding image category set.

[0037] In practice, first, manually judge and classify each initial object image in the above set of initial object image samples, and give the category of each picture (three labels). For example, for a picture that is unclear, has a plastic bag, and no other foreign objects, the label can be (clear: 0, foreign object occlusion: 0, plastic bag occlusion: 1).

[0038] In the second step, divide the above set of initial object image samples into a training set and a test set according to a preset ratio.

[0039] In practice, after the above annotation is completed, the execution entity can calculate the number of pictures with each of the three labels being 0 and 1. For example, the initial data set has 1000 object pictures, including 200 clear object images, 800 unclear object images, 700 object images with foreign object occlusion, 300 object images without foreign object occlusion, 500 object images with plastic bags, and 500 object images without plastic bags. Flexibly supplement the data to make the proportion of each classification quantity close to balance (1:1) to obtain an expanded model data set.

[0040] In the third step, based on the above training set and the above test set, the above image classification model is trained to obtain a trained image classification model. In practice, the above execution entity can train the above image classification model with the above training set, update the parameters through forward and backward propagation, and evaluate with the above test set. Adjust the hyperparameters according to the results. If the expected result is achieved, save the above image classification model that has completed training.

[0041] Step 103, in response to determining that the initial item image classification result indicates that there is a plastic bag occlusion, determine that the initial item image meets the temporary retention condition.

[0042] In some embodiments, the above execution entity can, in response to determining that the initial item image classification result indicates that there is a plastic bag occlusion, determine that the initial item image meets the temporary retention condition. Among them, the above temporary retention condition may be that there is a plastic bag occlusion on the item in the item image and the item image is clear without afterimages.

[0043] Step 104, in response to determining that the initial item image meets the temporary retention condition, perform plastic bag feature recognition on the initial item image to obtain plastic bag feature information corresponding to the initial item image.

[0044] In some embodiments, the above execution entity can, in response to determining that the initial item image meets the temporary retention condition, perform plastic bag feature recognition on the initial item image to obtain plastic bag feature information corresponding to the initial item image. Among them, the above plastic bag feature information includes plastic bag occlusion ratio information, plastic bag color, and plastic bag transparency. The above plastic bag occlusion ratio information may be the plastic bag occlusion ratio. Among them, the above plastic bag occlusion ratio can be defined as the ratio of the number of pixels containing the item and having a plastic bag to the number of pixels containing the item.

[0045] In some alternative implementation manners of some embodiments, the above execution entity can, in response to determining that the initial item image meets the above temporary retention condition, perform plastic bag feature recognition on the initial item image through the following steps to obtain plastic bag feature information corresponding to the above initial item image: First step, input the initial object image that meets the above temporary retention conditions into the extraction convolutional network included in the pre-trained plastic bag feature extraction model to obtain a plastic bag information feature map. The above plastic bag feature extraction model is a neural network model that takes the initial object image that meets the above temporary retention conditions as input and outputs plastic bag feature information. Among them, the above pre-trained plastic bag feature extraction model further includes an extraction model average pooling layer, an extraction model fully connected layer, and an extraction model output layer. The above extraction convolutional network may include a first extraction model convolutional layer, a second extraction model convolutional layer set, a third extraction model convolutional layer set, a fourth extraction model convolutional layer set, and a fifth extraction model convolutional layer set. The above first extraction model convolutional layer is provided with a sixth preset number of convolutional kernels. The above second extraction model convolutional layer set is provided with a seventh preset number of convolutional kernels. The above seventh preset number is greater than the above sixth preset number. The above third extraction model convolutional layer set is provided with an eighth preset number of convolutional kernels. The above eighth preset number is greater than the above seventh preset number. The above fourth extraction model convolutional layer set is provided with a ninth preset number of convolutional kernels. The above ninth preset number is greater than the above eighth preset number. The above fifth extraction model convolutional layer set is provided with a tenth preset number of convolutional kernels. The above tenth preset number is greater than the above ninth preset number. The above first extraction model convolutional layer may include a convolutional layer with a convolutional kernel of 7*7, a channel number of 64, and a stride of 2, and a pooling layer with a stride of 2. In practice, the above execution entity can extract low-level features of the plastic bag in the image through the above first extraction model convolutional layer. For example, the plastic bag contour and texture, and at the same time perform downsampling with a stride of 2 to reduce the image size. Then perform downsampling through the above pooling layer to prevent overfitting. The above second extraction model convolutional layer set may include six convolutional layers with a convolutional kernel of 3*3 and a channel number of 64. The above third extraction model convolutional layer set may include eight convolutional layers with a convolutional kernel of 3*3 and a channel number of 128. The above fourth extraction model convolutional layer set may include twelve convolutional layers with a convolutional kernel of 3*3 and a channel number of 256. The above fifth extraction model convolutional layer set may include twelve convolutional layers with a convolutional kernel of 3*3 and a channel number of 512. The working principles of the above second extraction model convolutional layer set, the above third extraction model convolutional layer set, the above fourth extraction model convolutional layer set, and the above fifth extraction model convolutional layer set are the same as those of each convolutional layer set in the above classification convolutional network, that is, continuously extract higher-level and more abstract plastic bag image features. Each layer of convolution can capture features at different levels. Shallow layers may capture edges, simple shapes, etc., while deep layers can capture more complex object components and overall structures.A residual connection (Residual Connection) can be provided between the output end of the second extraction model convolutional layer set and the first two convolutional layers of the third extraction model convolutional layer set, between the output end of the third extraction model convolutional layer set and the first two convolutional layers of the fourth extraction model convolutional layer set, and between the output end of the fourth extraction model convolutional layer set and the first two convolutional layers of the fifth extraction model convolutional layer set. The above residual connection can be used to effectively update the weights of both shallow and deep convolutional layers, ensure the continuous optimization of the convolutional layer, and improve the feature extraction effect. It should be noted that the number of convolutional kernels of each of the above convolutional kernels and the number of convolutional layers in each of the above convolutional layer sets can be adjusted according to requirements, and no specific limitation is made here.

[0046] In the second step, input the above plastic bag information feature map into the average pooling layer of the above extraction model to obtain a downsampled plastic bag information feature map. The above average pooling layer can be used to convert the above plastic bag information feature map into a feature vector of a fixed length, further reduce the feature dimension, and at the same time comprehensively represent the entire feature map to prepare for the subsequent fully connected layer.

[0047] In the third step, input the above downsampled plastic bag information feature map into the fully connected layer of the above extraction model to obtain plastic bag extraction information. The fully connected layer of the above extraction model can be a fully connected layer with 1000 neurons, which can be used to convert the feature vector into the distribution of each category and is finally used for prediction in tasks such as image classification.

[0048] In the fourth step, output the above plastic bag extraction information through the output layer of the above extraction model to obtain plastic bag feature information. Among them, the above plastic bag feature information can be various information related to the plastic bag in the above initial item image that meets the above temporary retention condition, for example: "The color of the plastic bag in this image is white, the transparency is 80%, the occlusion ratio is 18%, and it can be retained."

[0049] In practice, the above execution entity can automatically store the initial item image label with the above temporary retention condition of "can be retained" and delete the initial item image with the above temporary retention condition of "cannot be retained".

[0050] In some optional implementation manners of some embodiments, the above execution entity can input the above initial item image that meets the above temporary retention condition into the extraction convolutional network included in the pre-trained plastic bag feature extraction model through the following steps to obtain a plastic bag information feature map: In the first step, input the initial item image that meets the above temporary retention conditions into the first extraction model convolutional layer of the pre-trained plastic bag feature extraction model to obtain the sixth feature map of the initial item image that meets the above temporary retention conditions. The information contained in the sixth feature map can be local features of the plastic bag in the initial item image that meets the above temporary retention conditions, such as edges, textures, etc.

[0051] In the second step, input the sixth feature map into the set of second extraction model convolutional layers to obtain the seventh feature map. The information contained in the seventh feature map can be more abstract features of the plastic bag in the initial item image that meets the above temporary retention conditions, such as shape, color, etc.

[0052] In the third step, input the seventh feature map into the set of third extraction model convolutional layers to obtain the eighth feature map. The information contained in the eighth feature map can be the depth features of the plastic bag in the initial item image that meets the above temporary retention conditions, such as the combined relationship between the plastic bag and the item.

[0053] In the fourth step, input the eighth feature map into the set of fourth extraction model convolutional layers to obtain the ninth feature map. The information contained in the ninth feature map can be deeper features of the plastic bag in the initial item image that meets the above temporary retention conditions, such as more detailed relationships between the plastic bag and the item, such as color differences and color fusions with the item.

[0054] In the fifth step, input the ninth feature map into the set of fifth extraction model convolutional layers to obtain the tenth feature map as the plastic bag information feature map. The information contained in the plastic bag information feature map can be further depth features of the plastic bag in the initial item image that meets the above temporary retention conditions, such as feature information in the case of the relationship between multiple items and multiple plastic bags.

[0055] Step 105: Determine whether the initial item image meets the retention conditions according to the plastic bag occlusion ratio information included in the plastic bag feature information.

[0056] In some embodiments, the above execution subject can determine whether the initial item image meets the retention conditions according to the plastic bag occlusion ratio information included in the plastic bag feature information. Among them, determining whether the initial item image meets the retention conditions can be achieved by comparing the plastic bag occlusion ratio information with a pre-set occlusion ratio threshold. The occlusion ratio threshold can be set according to the actual situation. For example, when the plastic bag color is white and the transparency is not less than 70%, if the occlusion ratio does not exceed 20%, it can be retained, otherwise it cannot be retained.

[0057] Step 106: In response to determining that the initial item image meets the retention conditions, determine the initial item image as the collected item image.

[0058] In some embodiments, the above-mentioned execution entity may determine the initial item image as the acquired item image in response to determining that the initial item image meets the retention condition. Among them, the above-mentioned item images that do not meet the retention condition will be automatically deleted. The above-mentioned acquired item images can be automatically stored in the server.

[0059] Optionally, the above-mentioned execution entity may also perform the following steps: In the first step, in response to determining that the above-mentioned initial item image classification result indicates occlusion and afterimage, an abnormal visualization prompt message is output. Among them, the above-mentioned abnormal prompt message can be visualized through a visualization device. The device is not specifically limited. For example, outputting the above-mentioned abnormal prompt message can be outputting the message "Error, cannot be recognized" and displaying it.

[0060] In the second step, the following steps are performed on the above-mentioned acquired item image: In the first sub-step, the above-mentioned acquired item image is preprocessed to obtain a preprocessed item image. The above-mentioned item image may include a food image or a dish image, which is not specifically limited. The above-mentioned preprocessing may include adjusting the above-mentioned acquired item image to a size adapted to subsequent input.

[0061] The second sub-step is to input the pre-processed item image into the feature embedding sub-network included in the item recognition model to obtain a feature recognition map. The item recognition model includes a feature embedding sub-network, a relationship learning sub-network, and an output layer. The item recognition model can be a neural network model that takes the pre-processed item image as input and outputs item recognition information. The feature embedding sub-network can include at least one convolutional layer with a convolutional kernel of 7*7, and two or more of the convolutional layers can form a residual block. For example, when the pre-processed item image is input into the residual block, it first passes through a convolutional layer with a convolutional kernel of 3*3, and the number of output channels is 64. Then, batch normalization and ReLU activation are performed. After that, it passes through another 3*3 convolutional layer, and the number of output channels is still 64. Then, this output is added to the input feature map element by element to achieve residual connection, which helps to solve the problems of gradient disappearance and degradation that occur as the network deepens, enabling the network to be trained deeper and extract more advanced features. The relationship learning sub-network can include a Graph Neural Network (GNN). The graph neural network can be a data structure composed of nodes and edges, which can represent many complex relational data. In practice, the execution entity can receive the pre-processed item image through the feature embedding sub-network and convert the high-dimensional and redundant image data into low-dimensional, compact, and semantic-rich feature vectors to obtain a feature recognition map. Among them, the relationship learning sub-network can be a neural network that mines the internal relationships between different features in the feature recognition map. For example, in a picture containing multiple fruits, the feature embedding sub-network extracts the shape and color features of each fruit, and the relationship learning sub-network needs to analyze logical relationships such as "red circular objects are more likely to be apples" and "yellow long objects are probably bananas", so that the model not only knows what isolated features are in the image but also understands how they are related and matched, thereby improving the recognition accuracy.

[0062] The third sub-step is to perform feature fusion on the feature recognition map to obtain a fused feature recognition map. The fusion can be the element-by-element addition of the residual module to the feature map. The feature map can include edges, textures, etc. in the item image.

[0063] The fourth sub-step is to input the fused feature recognition map into the relationship learning sub-network to obtain relationship learning information. The relationship learning information can include the relationship between the category and price of the item, etc. For example, "apple" corresponds to "1 yuan".

[0064] The fifth sub-step is to input the relationship learning information into the output layer to obtain item recognition information. For example, output "Item type: apple, price: 1 yuan".

[0065] The sixth sub-step is to perform visualization processing on the above-mentioned item identification information. The device for the above-mentioned visualization processing can be the same as the device for the above-mentioned abnormal prompt information visualized in the first step.

[0066] The above first step - second step and their related content are an inventive point of an embodiment of the present disclosure, which solves the technical problem of "lack of utilization of the data of the items after cleaning". The factors leading to the lack of utilization of the data of the items after cleaning are often as follows: Most of the screened item image data obtained by current data cleaning methods are mostly applied to subsequent model training and are not further utilized, resulting in low utilization rate. If the above factors are solved, the utilization rate of the screened item images can be improved. To achieve this effect, first, the above first step - second step and their related content provide a method for identifying the screened item image data and feed back relevant information to the user through visualization means. Therefore, the screened item images are further effectively utilized. Thus, the utilization rate of the screened item image data is improved.

[0067] Optionally, the above-mentioned execution subject can also perform the following steps: The first step is to, in response to determining that the above-mentioned initial item image classification result is characterized as unclear, perform format conversion on the initial item image characterized as unclear to obtain an item image in the form of an array. The clarity of the image can be determined by PSNR. The above PSNR is an objective index widely used to measure the quality of an image and can measure the degree of distortion of the image by calculating the mean square error (MSE) between the original image and the image to be evaluated. In practice, the above-mentioned execution subject can convert the unclear item image from a file form to a digital format recognizable and processable by a computer through conversion technology. Among them, the above-mentioned conversion technology can be to first use the OpenCV library of Python to read the image, and then use the cv2.imread() function to convert the above-mentioned unclear item image into an item image in the form of an array.

[0068] The second step is to perform grayscale conversion on the above-mentioned item image in the form of an array to obtain a grayscale-converted item image. In practice, the above-mentioned execution subject can perform grayscale conversion on the above-mentioned item image through a grayscale conversion function. Among them, the above-mentioned grayscale conversion function can be the cv2.cvtColor() function in OpenCV.

[0069] The third step is to perform the following steps on the above-mentioned grayscale-converted item image: The first sub-step is to divide the above-mentioned grayscale-converted object image into a preset number of non-overlapping regions to obtain the divided object image. Among them, the above-mentioned preset number can be set according to the image size and actual requirements. For example, the above-mentioned grayscale-converted object image is divided into multiple small blocks of 10x10 pixels. In practice, the above-mentioned execution entity can divide the above-mentioned grayscale-converted object image through a segmentation algorithm. Among them, the above-mentioned segmentation algorithm can include, but is not limited to: superpixel segmentation algorithm and adaptive division based on image content, etc.

[0070] The second sub-step is to generate the standard deviation of the pixel values of each non-overlapping region in the above-mentioned divided object image to obtain a standard deviation set. Among them, the above-mentioned standard deviation can reflect the degree of dispersion of pixel values in the corresponding region, providing a data basis for subsequent calculation of the noise level estimation value. In practice, the above-mentioned execution entity can calculate the standard deviation of the pixel values of each non-overlapping region through a standard deviation function. Among them, the above-mentioned standard deviation function can be the np.std() function in the numpy library of Python.

[0071] The third sub-step is to generate the average value of each standard deviation in the above-mentioned standard deviation set to obtain the noise level estimation value. Among them, the above-mentioned noise level estimation value can be a numerical index for measuring the degree of noise in the image. For example, for a landscape photo taken with a mobile phone in a low-light environment, the average value of the standard deviations of each pixel block is 8.6, that is, the noise level estimation value is 8.6, indicating that the photo is highly interfered by noise and there may be many noise points in the picture.

[0072] The fourth step is to generate the image gradient of each pixel point in the above-mentioned grayscale-converted object image to obtain an image gradient set in a preset direction. Among them, the above-mentioned image gradient set includes the image gradients in the preset direction of each pixel point. The above-mentioned image gradient can be the change rate and change direction of pixel values in the grayscale-converted object image. The above-mentioned preset direction can be the horizontal direction (x direction) and vertical direction (y direction) of the current pixel point. In practice, the above-mentioned execution entity can calculate the image gradient of each pixel point through a Sobel operator. For example, the above-mentioned Sobel operator can be implemented through the cv2.Sobel() function in the OpenCV library. The horizontal gradient of a pixel point in the x direction can be denoted as gradient_x, and the vertical gradient in the y direction can be denoted as gradient_y.

[0073] Step 5: Generate the gradient magnitude corresponding to each pixel point in the grayscale-converted item image based on the set of image gradients in the above-mentioned preset direction, obtaining a set of gradient magnitudes. Among them, the above-mentioned gradient magnitude can be a value reflecting the degree of change of local pixel values in the image. In practice, the above-mentioned execution entity can preset a calculation method and calculate the corresponding gradient magnitude based on the above-mentioned image gradient. Among them, the gradient magnitude can be denoted as gradient_magnitude. The above-mentioned preset calculation method can be to add the square of gradient_x and the square of gradient_y and then take the square root.

[0074] Step 6: Divide the grayscale-converted item image into different neighborhood blocks according to the above-mentioned set of gradient magnitudes and a preset gradient magnitude threshold. Among them, the above-mentioned neighborhood block can be a non-overlapping area in the grayscale-converted item image. The above-mentioned preset gradient magnitude threshold can be a value used to distinguish different image feature regions. For example, the above-mentioned preset gradient magnitude threshold can be 50. In practice, first, the above-mentioned execution entity can traverse each pixel point in the grayscale-converted item image to obtain its gradient magnitude. Then, compare the gradient magnitude of each pixel with the above-mentioned preset gradient magnitude threshold. Then, according to the comparison result, add the coordinates of each pixel to the pre-created smooth area list or edge or texture area list respectively. Finally, piece together the pixels corresponding to the coordinates of each pixel in the same area to form the above-mentioned different neighborhood blocks.

[0075] Step 7: Determine a set of similar blocks based on each neighborhood block in the above-mentioned different neighborhood blocks. In practice, the above-mentioned execution entity can judge whether they are similar by calculating the difference in pixel values within each neighborhood block through a difference algorithm. Among them, the above-mentioned difference algorithm can be a search method based on a hash table. For example, first, the hash value of each above-mentioned neighborhood block can be generated using the zlib library. Then, store the hash values of each above-mentioned neighborhood block in the pre-created hash table. Finally, find the blocks with the same or similar (such as the difference being less than 0.2) hash values in the pre-created hash table, and determine these blocks with similar hash values as similar blocks.

[0076] Step 8: Generate the structure tensor components of each pixel point in the grayscale-converted item image based on the image gradient in the above-mentioned preset direction, obtaining a set of structure tensor components. Among them, the above-mentioned structure tensor components can be three values, which are respectively 、 and . In practice, the above-mentioned execution entity can generate the above-mentioned structure tensor components through a preset gradient calculation method and the above-mentioned image gradient. Among them, the methods for calculating the above-mentioned three values by the above-mentioned preset gradient calculation method can be respectively: It can be the square of the above-mentioned gradient_x of the current pixel point, It can be the square of the above-mentioned gradient_y, It can be the product of gradient_x and gradient_y.

[0077] In the ninth step, according to the above-mentioned set of structure tensor components, generate the structure tensor matrix of each pixel point in the item image after the above-mentioned gray conversion, and obtain a set of structure tensor matrices. Among them, the above-mentioned structure tensor matrix can comprehensively reflect the local structure characteristics of the image, and it can capture the distribution law of pixels in the image and characteristics such as edges and textures. In practice, the above-mentioned execution entity can generate the structure tensor matrix of each pixel point through the following formula .

[0078]

[0079] Among them, can represent the pixel coordinates in the image. is the neighborhood of the current pixel point. The above-mentioned neighborhood can be a local area centered on the current pixel point, and the above-mentioned local area can include a preset number (such as 60) of pixels around the current pixel point. can be the Gaussian weight, can be the row, and the Gaussian weight of the pixel point in the column. The above can be the row, and the column pixel point's . can be the row, and the column pixel point's . The above can be the row, and the column pixel point's . The above-mentioned structure tensor matrix comprehensively reflects the local structure characteristics of the image, and it can capture the distribution law of pixels in the image and characteristics such as edges and textures.

[0080] In the tenth step, generate the difference degree values of the corresponding structure tensor matrices of each similar block in the above-mentioned set of similar blocks through the matrix difference algorithm, and obtain a set of similarity values. Among them, the above-mentioned matrix difference algorithm can be to calculate the Euclidean distance between the elements in the structure tensor matrix corresponding to each of the above-mentioned similar blocks. In practice, first, the above-mentioned execution entity can use the Euclidean distance formula to generate the structural similarity metric values between the structure tensor matrices corresponding to each pair of similar blocks. Then, the structural similarity metric values between each similar block can be summarized to form the above-mentioned set of similarity metrics.

[0081] The eleventh step is to assign weights to each of the above-mentioned similar blocks according to the above-mentioned similarity value set and the preset weight assignment information, so as to obtain a preliminary weight assignment information set. Among them, the above-mentioned preset weight assignment information can be linear information based on structural similarity measurement. The above-mentioned linear information based on structural similarity measurement can be that the more similar the structure of the block is considered, the greater the contribution to the current pixel repair, and a higher weight is assigned. The above-mentioned preliminary weight assignment information set can include the initial weight values corresponding to each similar block. In practice, first, the above-mentioned execution entity can set a weight assignment method. The above-mentioned weight assignment method can be a threshold-based assignment method. The above-mentioned threshold can be set as needed. For example, initialize an empty dictionary weight_dict to store the weights of each similar block. For each pair of similar blocks, determine the weight according to its structural similarity measurement value d. If d is less than a set smaller threshold (such as 0.1, indicating very similar structures), then assign a higher initial weight value, such as 0.8. If d is greater than a set larger threshold (such as 0.5, indicating a large structural difference), then assign a lower initial weight value, such as 0.2. For those in between, the weight can be determined by linear interpolation. Then, determine the weights of each similar block through the above-mentioned weight assignment method. Finally, the obtained weight values can be stored, using the position coordinates of each similar block in the image as the key and the corresponding weight value as the value, to obtain a preliminary weight assignment information set.

[0082] The twelfth step is to adjust the above-mentioned preliminary weight assignment information set according to the above-mentioned noise level estimation value to obtain a finally determined weight distribution information set. Among them, the above-mentioned finally determined weight distribution information set can include the finally assigned weight values of each similar block updated after being adjusted by the above-mentioned noise level estimation value. In practice, the above-mentioned execution entity can adjust the above-mentioned preliminary weight assignment information set by the weight threshold method. Among them, the above-mentioned weight threshold method can be to set different weight thresholds through the above-mentioned noise level estimation value, and then amplify or reduce each weight in the above-mentioned preliminary weight assignment information set. For example, raise the threshold for distinguishing high weights and low weights from 0.3 to 0.4. When the noise level is low, lower the above-mentioned threshold (such as lowering it to 0.2). Then reassign the weights of each similar block according to the adjusted threshold (such as assigning lower weights to those with high noise level estimation values and higher weights to those with low noise level estimation values), and update the preliminary weight assignment information set, so as to obtain a finally determined weight distribution information set.

[0083] Step 13: Repair the item image after the above grayscale conversion according to the above set of similar blocks and the finally determined weight distribution to obtain the repaired item image. In practice, for each pixel in the item image after the above grayscale conversion, the above execution entity can generate the repaired pixel values of the pixels in the item image after the above grayscale conversion through a weight algorithm based on each similar block and its corresponding weight, so as to obtain the above preliminarily repaired item image. Among them, the above weight algorithm can be a non-local means filtering algorithm. The above data type is numpy.ndarray, and its shape is the same as when input. The above preliminarily repaired item image can be an image that repairs the blurred or residual image parts on the initial item image characterized as unclear.

[0084] Step 14: In response to the above repaired item image meeting the above retention conditions, retain the above repaired item image. In practice, the above execution entity can also determine whether the above repaired item image meets the above retention conditions through an image sharpness determination method. If it meets the conditions, the repaired item image will be automatically stored. If it does not meet the conditions, it will be discarded. Among them, the above image sharpness determination method can also be PSNR.

[0085] The above steps 1 - 14 and their related content are an inventive point of an embodiment of the present disclosure, which solves the technical problem of "lacking a solution for further processing the unclear item images collected". The factors that lead to the lack of a solution for further processing the unclear item images collected are often as follows: Currently, most data cleaning methods will directly discard the item images collected and characterized as unclear instead of further processing them, which results in the loss of many collected item images. If the above factors are solved, the number of available item images collected can be further increased. To achieve this effect, first, the above steps 1 - 14 and their related content provide a solution for repairing unclear item images, which can traverse each pixel of the image, dynamically divide regions, and calculate and allocate corresponding weights to repair the residual image and blurred regions in the image, etc. Thus, while ensuring the image quality, the number of available item image data collected is further increased.

[0086] The above-mentioned various embodiments of the present disclosure have the following beneficial effects: Through an article image acquisition method for a weighing device of the present disclosure, the plastic bag information in a picture containing a plastic bag can be predicted, and the retention conditions of the acquired article images can be set by itself according to a downstream model (such as article recognition). Thus, pictures containing plastic bags but with high usability can be flexibly screened out. Specifically, the reason why most article image acquisition methods cannot flexibly select images with high usability is that most article image acquisition methods do not further screen the article pictures blocked by plastic bags, resulting in a small number and poor quality of the acquired article images. Based on this, the article image acquisition method for a weighing device in some embodiments of the present disclosure includes: First, photograph the object on the weighing device to obtain an initial article image. Thus, the article image to be processed is obtained. Then, input the above initial article image into a pre-trained image classification model to obtain an initial article image classification result. Thus, the article can be initially roughly classified by the first model. Then, in response to determining that the above initial article image classification result indicates that there is a plastic bag occlusion, determine that the above initial article image meets the temporary retention condition. Thus, a basis is laid for further classification later. Then, in response to determining that the above initial article image meets the above temporary retention condition, perform plastic bag feature recognition on the above initial article image to obtain plastic bag feature information corresponding to the above initial article image, where the above plastic bag feature information includes plastic bag occlusion ratio information. Then, determine whether the above initial article image meets the retention condition according to the plastic bag occlusion ratio information included in the above plastic bag feature information. Finally, in response to determining that the above initial article image meets the above retention condition, determine the above initial article image as the acquired article image. Thus, the data cleaning of the acquired article images is completed. The above article image acquisition method for a weighing device uses two deep learning models to perform data cleaning on the acquired article images, and filters out available article images by using preset conditions. Thus, while ensuring the quality of the acquired article images, the number of acquired article images is increased as much as possible, and the efficiency of acquiring available article images is improved.

[0087] Further referring to Figure 2 , as an implementation of the methods shown in the above figures, the present disclosure provides some embodiments of a device for article image acquisition for a weighing device. These device embodiments correspond to Figure 1 the method embodiments shown, and the device can be specifically applied to various electronic devices.

[0088] As Figure 2As shown in the figure, the device 200 for item image acquisition of a weighing device in some embodiments includes: a photographing unit 201, a classification unit 202, a first determination unit 203, a second determination unit 204, a third determination unit 205, and a fourth determination unit 206. Among them, the photographing unit 201 is configured to photograph an object on the weighing device to obtain an initial item image; the classification unit 202 is configured to input the above initial item image into a pre-trained image classification model to obtain an initial item image classification result; the first determination unit 203 is configured to determine that the above initial item image satisfies the temporary retention condition in response to determining that the above initial item image classification result indicates that there is a plastic bag occlusion; the second determination unit 204 is configured to perform plastic bag feature recognition on the above initial item image in response to determining that the above initial item image satisfies the above temporary retention condition to obtain plastic bag feature information corresponding to the above initial item image, where the above plastic bag feature information includes plastic bag occlusion ratio information; the third determination unit 205 is configured to determine whether the above initial item image satisfies the retention condition according to the plastic bag occlusion ratio information included in the above plastic bag feature information; the fourth determination unit 206 is configured to determine the above initial item image as the acquired item image in response to determining that the above initial item image satisfies the above retention condition.

[0089] It can be understood that the various units described in the device 200 correspond to the respective steps in the method described with reference to Figure 1 Therefore, the operations, features, and beneficial effects described above for the method also apply to the device 200 and the units included therein, and will not be repeated here.

[0090] The following refers to Figure 3 , which shows a schematic structural diagram of an electronic device (such as the computing device 101 shown in Figure 1 ) 300 suitable for implementing some embodiments of the present disclosure. Figure 3 The electronic device shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present disclosure.

[0091] As Figure 3 shown, the electronic device 300 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 301, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from a storage device 308 into a random access memory (RAM) 303. In the RAM 303, various programs and data required for the operation of the electronic device 300 are also stored. The processing device 301, the ROM 302, and the RAM 303 are connected to each other through a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.

[0092] Typically, the following devices can be connected to the I / O interface 305: input devices 306 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; output devices 307 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; storage devices 308 including, for example, magnetic tape, a hard disk, etc.; and a communication device 309. The communication device 309 can allow the electronic device 300 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 3 the electronic device 300 with various devices is shown, it should be understood that it is not required to implement or include all the shown devices. Instead, more or fewer devices can be implemented or included. Figure 3 Each block shown in can represent one device or, as required, multiple devices.

[0093] In particular, according to some embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, some embodiments of the present disclosure include a computer program product that includes a computer program carried on a computer-readable medium, and the computer program includes program codes for performing the methods shown in the flowcharts. In such some embodiments, the computer program can be downloaded and installed from a network through the communication device 309, or installed from the storage device 308, or installed from the ROM 302. When the computer program is executed by the processing device 301, the above functions defined in the methods of some embodiments of the present disclosure are performed.

[0094] It should be noted that the computer-readable media described in some embodiments of the present disclosure may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In some embodiments of the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In some embodiments of the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium may also be any computer-readable medium other than the computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0095] In some embodiments, the client and the server can communicate using any currently known or future-developed network protocol such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks ("LAN"), wide area networks ("WAN"), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed network.

[0096] The above computer-readable medium may be included in the above electronic device; or it may exist separately and not be assembled into the electronic device. The above computer-readable medium carries one or more programs, and when the one or more programs are executed by the electronic device, the electronic device is caused to: take a picture of an object on the weighing device to obtain an initial article image; input the initial article image into a pre-trained image classification model to obtain an initial article image classification result; in response to determining that the initial article image classification result indicates that there is a plastic bag occlusion, determine that the initial article image meets the temporary retention condition; in response to determining that the initial article image meets the temporary retention condition, perform plastic bag feature recognition on the initial article image to obtain plastic bag feature information corresponding to the initial article image, wherein the plastic bag feature information includes plastic bag occlusion ratio information; determine whether the initial article image meets the retention condition according to the plastic bag occlusion ratio information included in the plastic bag feature information; in response to determining that the initial article image meets the retention condition, determine the initial article image as the collected article image.

[0097] Computer program code for performing the operations of some embodiments of the present disclosure may be written in one or more programming languages or combinations thereof. The above programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).

[0098] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or by a combination of dedicated hardware and computer instructions.

[0099] The units described in some embodiments of the present disclosure can be implemented in software or in hardware. The described units can also be provided in a processor. For example, it can be described as: a processor includes a photographing unit, a classification unit, a first determination unit, a second determination unit, a third determination unit, and a fourth determination unit. Among them, the names of these units do not constitute a limitation on the unit itself in some cases. For example, the photographing unit can also be described as "a unit that photographs an object on a weighing device to obtain an initial article image".

[0100] The functions described above can be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), and so on.

[0101] The above description is only some preferred embodiments of the present disclosure and an explanation of the technical principles applied. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, technical solutions formed by mutually replacing the above features with technical features having similar functions (but not limited to) disclosed in the embodiments of the present disclosure.

Claims

1. A method for collecting an image of an object for a weighing device, comprising: Photograph the object on the weighing device to obtain an initial object image; Inputting the initial object image into a pre-trained image classification model to obtain an initial object image classification result; In response to determining that the classification result of the initial object image indicates that it is blocked by a plastic bag, determining that the initial object image satisfies a temporary retention condition; In response to determining that the initial object image satisfies the temporary retention condition, performing plastic bag feature recognition on the initial object image to obtain plastic bag feature information corresponding to the initial object image, wherein the plastic bag feature information includes plastic bag occlusion ratio information; Determining whether the initial object image meets a retention condition according to the plastic bag occlusion ratio information included in the plastic bag feature information; In response to determining that the initial object image satisfies the retention condition, the initial object image is determined as the acquired object image.

2. The method according to claim 1, wherein: The method further comprises: Saving the initial item image locally; Each initial object image saved within a preset time period is uploaded to the server.

3. The method according to claim 1, wherein: The step of inputting the initial object image into a pre-trained image classification model to obtain an initial object image classification result includes: Inputting the initial object image into the classification convolution network included in the pre-trained image classification model to obtain a classification feature map, wherein the pre-trained image classification model also includes a classification model fully connected layer and a classification model output layer, and the classification convolution network includes a first classification model convolution layer set, a second classification model convolution layer set, a third classification model convolution layer set, a fourth classification model convolution layer set, and a fifth classification model convolution layer set, each convolution layer in the first classification model convolution layer set is provided with a first preset number of convolution kernels, each convolution layer in the second classification model convolution layer set is provided with a second preset number of convolution kernels, the second preset number is greater than the first preset number, each convolution layer in the third classification model convolution layer set is provided with a third preset number of convolution kernels, the third preset number is greater than the second preset number, each convolution layer in the fourth classification model convolution layer set is provided with a fourth preset number of convolution kernels, the fourth preset number is greater than the third preset number, each convolution layer in the fifth classification model convolution layer set is provided with a fifth preset number of convolution kernels, and the fifth preset number is equal to the fourth preset number; Inputting the classification feature map into the fully connected layer of the classification model to obtain object anomaly detection classification information; The object anomaly detection classification information is input into the classification model output layer to obtain an initial object image classification result.

4. The method according to claim 3, wherein: The inputting the initial object image into the classification convolution network included in the pre-trained image classification model to obtain a classification feature map includes: Inputting the initial object image into the first classification model convolution layer set to obtain a first feature map of the initial object image; Inputting the first feature map into a set of convolutional layers of a second classification model to obtain a second feature map; Inputting the second feature map into a third classification model convolutional layer set to obtain a third feature map; Inputting the third feature map into a fourth classification model convolutional layer set to obtain a fourth feature map; The fourth feature map is input into a fifth classification model convolutional layer set to obtain a fifth feature map as a classification feature map.

5. The method according to claim 3, wherein: The pre-trained image classification model is trained by the following steps: Acquire a pre-labeled initial object image sample set, wherein each initial object image sample in the initial object image sample set corresponds to image category information, and the image category information includes each label value corresponding to the image category set; Dividing the initial object image sample set into a training set and a test set according to a preset ratio; The image classification model is trained according to the training set and the test set to obtain a trained image classification model.

6. The method according to claim 1, wherein: In response to determining that the initial object image satisfies the temporary retention condition, performing plastic bag feature recognition on the initial object image to obtain plastic bag feature information corresponding to the initial object image includes: Inputting the initial item image satisfying the temporary retention condition into the extraction convolution network included in the pre-trained plastic bag feature extraction model to obtain a plastic bag information feature map, wherein the pre-trained plastic bag feature extraction model also includes an extraction model average pooling layer, an extraction model fully connected layer and an extraction model output layer, the extraction convolution network includes a first extraction model convolution layer, a second extraction model convolution layer set, a third extraction model convolution layer set, a fourth extraction model convolution layer set and a fifth extraction model convolution layer set, the first extraction model convolution layer is provided with a sixth preset number of convolution kernels, the second extraction model convolution layer set is provided with a seventh preset number of convolution kernels, the seventh preset number is greater than the sixth preset number, the third extraction model convolution layer set is provided with an eighth preset number of convolution kernels, the eighth preset number is greater than the seventh preset number, the fourth extraction model convolution layer set is provided with a ninth preset number of convolution kernels, the ninth preset number is greater than the eighth preset number, the fifth extraction model convolution layer set is provided with a tenth preset number of convolution kernels, the tenth preset number is greater than the ninth preset number; Inputting the plastic bag information feature map into the average pooling layer of the extraction model to obtain a dimensionality-reduced plastic bag information feature map; Inputting the dimension-reduced plastic bag information feature map into the fully connected layer of the extraction model to obtain plastic bag extraction information; The plastic bag extraction information is output through the extraction model output layer to obtain plastic bag feature information.

7. The method according to claim 6, wherein: The step of inputting the initial object image satisfying the temporary retention condition into an extraction convolution network included in a pre-trained plastic bag feature extraction model to obtain a plastic bag information feature map includes: Inputting the initial object image that meets the temporary retention condition into a first extraction model convolution layer of a pre-trained plastic bag feature extraction model to obtain a sixth feature map of the initial object image that meets the temporary retention condition; Inputting the sixth feature map into the second extraction model convolutional layer set to obtain a seventh feature map; Inputting the seventh feature map into the third extraction model convolutional layer set to obtain an eighth feature map; Inputting the eighth feature map into the fourth extraction model convolutional layer set to obtain a ninth feature map; The ninth feature map is input into the fifth extraction model convolutional layer set to obtain a tenth feature map as a plastic bag information feature map.

8. An object image acquisition device for a weighing device, comprising: A photographing unit, configured to photograph an object on the weighing device to obtain an initial object image; a classification unit configured to input the initial object image into a pre-trained image classification model to obtain an initial object image classification result; A first determining unit is configured to determine that the initial object image satisfies a temporary retention condition in response to determining that the classification result of the initial object image indicates that the image is blocked by a plastic bag; A second determination unit is configured to, in response to determining that the initial object image satisfies the temporary retention condition, perform plastic bag feature recognition on the initial object image to obtain plastic bag feature information corresponding to the initial object image, wherein the plastic bag feature information includes plastic bag occlusion ratio information; A third determination unit is configured to determine whether the initial object image meets a retention condition according to the plastic bag occlusion ratio information included in the plastic bag feature information; A fourth determining unit is configured to, in response to determining that the initial object image satisfies the retention condition, determine the initial object image as the acquired object image.

9. An electronic device, comprising: one or more processors; a storage device having one or more programs stored thereon; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 7.

10. A computer readable medium having a computer program stored thereon, wherein: When the program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Commodity identification method

    CN111626150A

  • Occluded image detection method and device and medium

    CN112200040A

  • Image occlusion detection method and vehicle-mounted terminal

    CN112446246A

  • Intelligent scale interference identification method, device and system and storage medium

    CN114782812A

  • Kitchen garbage shielding object recognition and shielding relation judgment method

    CN119027675A