Model training method, package image recognition method and device thereof
By training a food classification model by parsing and replacing text information on food packaging images, the problem of large workload for manual annotation is solved, training efficiency is improved, and the model's generalization ability and robustness are enhanced.
Patent Information
- Application Number
- CN202210324021.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-29
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2042-03-29
AI Technical Summary
Existing technologies require manual annotation of character information on a large number of food packaging images to train food classification models, resulting in a large workload and low efficiency.
By acquiring the annotation information of the packaging image, parsing it into the first text, and partially replacing it with the target preset text, the second text is generated. The first and second texts are then used to train a food classification model, reducing the need for manual annotation.
This approach expands the training samples, reduces the annotation workload for staff, improves model training efficiency, and the trained model exhibits good generalization ability and robustness.
Smart Images

Figure CN114663874B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer, in particular to a model training method, a packaging image recognition method and a device thereof. BACKGROUND
[0002] Generally, various characters such as Kai font, artistic characters, etc. are printed on the packaging of food. Part or all of the text information in the characters is a description of the properties of the food. It has great value to accurately extract the text information describing the properties of the food.
[0003] However, due to the variety of food and the richness of information on the packaging, in order to recognize the information on the packaging of different food, different food classification models need to be trained in advance for different food. At present, when different food classification models are trained, the packaging images corresponding to the food packaging need to be collected, and then the characters on the packaging images are recognized by OCR (Optical Character Recognition), and then the staff need to label the characters recognized by OCR, and then obtain accurate training data for the training of the food classification model. This requires manual labeling of a large number of characters on the packaging images, thereby increasing the workload of the staff. SUMMARY
[0004] Aspects of the present application provide a model training method, a packaging image recognition method and a device thereof to reduce the workload of the staff.
[0005] The first aspect of the embodiment of the present application provides a model training method, comprising: obtaining annotation information corresponding to a first packaging image, the annotation information being actual character information on the first packaging image, and the first packaging image being a packaging image of a first sample food; analyzing the annotation information to obtain a first text, the semantic of the first text being a description of the first sample food; replacing at least part of the first text with a target preset text to obtain a second text, the target preset text having an associated relationship with the first text, and the semantic of the second text being a description of a second sample food; training a food classification model according to the first text and the second text to obtain a trained food classification model.
[0006] The second aspect of the embodiment of the present application provides a packaging image recognition method, comprising: obtaining a packaging image of a target food; recognizing character information in the packaging image by using image recognition technology; analyzing the recognized character information to obtain text information, the semantic of the text information being a description of the target food; inputting the text information into a food classification model for classification processing to obtain a classification result of the target food, the food classification model being trained by the model training method of the first aspect.
[0007] The third aspect of the embodiment of the present application provides a packaging image recognition device, comprising:
[0008] An acquisition module is configured to acquire a packaging image of a target food product;
[0009] An identification module is configured to identify, by using an image recognition technology, identification character information in the packaging image;
[0010] An analysis module is configured to analyze the identification character information to obtain text information, the semantics of the text information being a description of the target food product;
[0011] A processing module is configured to input the text information into a food classification model for classification processing to obtain a classification result of the target food product, the food classification model being obtained by training according to the model training method of the first aspect.
[0012] The fourth aspect of the embodiment of the present application provides an electronic device, comprising a processor, a memory, and a computer program stored in the memory and executable on the processor, and the processor implements the model training method of the first aspect or the packaging image recognition method of the second aspect when executing the computer program.
[0013] The fifth aspect of the embodiment of the present application provides a computer readable storage medium, which stores a computer program, and when the computer program is executed by a processor, the processor implements the model training method of the first aspect or the packaging image recognition method of the second aspect.
[0014] The embodiment of the present application is applied to a text recognition scene on food product packaging, and the provided model training method comprises the following steps: acquiring annotation information corresponding to a first packaging image, the annotation information being actual character information on the first packaging image, and the first packaging image being a packaging image of a first sample food product; analyzing the annotation information to obtain a first text, the semantics of the first text being a description of the first sample food product; replacing at least part of the first text with a target preset text to obtain a second text, the target preset text having an associated relationship with the first text, and the semantics of the second text being a description of a second sample food product; and training a food classification model according to the first text and the second text to obtain a trained food classification model. The embodiment of the present application replaces at least part of the first text on the first packaging image with the target preset text, expands the training sample, avoids the need for manual annotation of characters on packaging images of various food products to obtain sufficient training samples, and thus reduces the annotation workload of the staff and improves the efficiency of model training. BRIEF DESCRIPTION OF DRAWINGS
[0015] The accompanying drawings described herein are used to provide further understanding of the present application, form a part of the present application, and are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:
[0016] Figure 1 A determination diagram of a first text provided for an exemplary embodiment of the present application;
[0017] Figure 2 A step flow chart of a model training method provided for an exemplary embodiment of the present application;
[0018] Figure 3 A step flow chart of another model training method provided for an exemplary embodiment of the present application;
[0019] Figure 4 A diagram of a second text provided for an exemplary embodiment of the present application;
[0020] Figure 5 Another diagram of a second text provided for an exemplary embodiment of the present application;
[0021] Figure 6 A diagram of a third text provided for an exemplary embodiment of the present application;
[0022] Figure 7 A step flow chart of a packaging image recognition method provided for an exemplary embodiment of the present application;
[0023] Figure 8 A structural block diagram of a packaging image recognition device provided for an exemplary embodiment of the present application;
[0024] Figure 9 A structural diagram of an electronic device provided for an exemplary embodiment of the present application. DETAILED DESCRIPTION
[0025] In order to make the objectives, technical solutions and advantages of the present application clearer, the technical solutions of the present application will be described below in conjunction with specific embodiments of the present application and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of protection of the present application.
[0026] For the text recognition scene on the food packaging, the provided model training method comprises: obtaining label information corresponding to a first packaging image, the label information being actual character information on the first packaging image, the first packaging image being a packaging image of a first sample food; analyzing the label information to obtain a first text, the semantic of the first text being a description of the first sample food; replacing at least part of the first text with a target preset text to obtain a second text, the target preset text having an associated relationship with the first text, the semantic of the second text being a description of a second sample food; training a food classification model according to the first text and the second text to obtain a trained food classification model. The embodiment of the present application realizes the expansion of the training sample by replacing at least part of the first text on the first packaging image with the target preset text, avoids the need for manual labeling of characters on packaging images of various foods to obtain sufficient training samples, thereby reducing the labeling workload of the staff and improving the efficiency of model training.
[0027] In the present embodiment, the execution device of the model training method is not limited. Alternatively, the model training method can implement the overall model training method by means of a cloud computing system. For example, the model training method can be applied to a cloud server so as to run various neural network models by means of the advantages of resources on the cloud; relative to the application to the cloud, the model training method can also be applied to a server device such as a conventional server, a cloud server or a server array.
[0028] In addition, with reference to Figure 1A schematic diagram of obtaining a first text 14 from a packaging image 11 of a beverage is provided in the embodiments of the present application. The packaging image 11 can be captured by a camera. Then the characters on the packaging image 11 are recognized by an OCR recognition technology to obtain a recognition result 12, which includes serial numbers (such as 1 to 14) and characters corresponding to the serial numbers. Since the characters on the packaging image 11 are in multiple fonts, the recognition result 12 obtained has problems such as character recognition errors and missing characters. For example, "no" in "no sugar" in serial number 1 is recognized as "husband", "lemon" in serial number 2 is not recognized, "a" in "aspartame" in serial number 3 is recognized as "river", "honey" in serial number 4 is not recognized, "fragrance" in serial number 4 is recognized as "mysterious", "trademark" in serial number 5 is recognized as "cup", "place of origin: Nanning City" in serial number 7 is not recognized, "23" in "GH12345" in serial number 8 is not recognized, "00" in serial number 9 is recognized as "88", and "return" in serial number 13 is recognized as "furnace". In the embodiments of the present application, the recognition result 12 needs to be labeled to obtain labeled information 13. The labeled information 13 is a correction of the part of the recognition result 12 that is recognized incorrectly, so the labeled information 13 is the actual character information on the packaging image 11. Further, the first text 14 is an extraction result of the text associated with the beverage in the labeled information 13.
[0029] Referring to Figure 1 The labeled information 13 can be obtained by manual labeling from the recognition result 12. At present, since different foods have different packaging, and the characters on the packaging are also various, if manual labeling is needed for the recognition results of different packaging of different foods to obtain corresponding labeled information, a large amount of manual operation is required, and the efficiency will be very low. However, in the model training method provided in the embodiments of the present application, only the labeled information of limited foods needs to be obtained, and a food classification model with generalization ability and robustness can be trained, and manual labeling of different packaging of various foods is not required, so the training efficiency of the food classification model can be improved.
[0030] This application provides a model training method that can expand the training samples by at least partially replacing the first text on the first packaging image with a target preset text. This avoids the need for manual annotation of characters on various food packaging images to obtain sufficient training samples, thereby reducing the workload of staff and improving the efficiency of model training. The resulting food classification model has good generalization ability and robustness. From the consumer's perspective, the trained food classification model can help consumers quickly and effectively determine whether a food is suitable for them. For example, a diabetic consumer may only want to determine whether a food is high in sugar. From the manufacturer's perspective, it can also help verify whether the text information printed on the food packaging is consistent with the text preset for that food, thereby enabling efficient quality control.
[0031] The technical solutions provided by the various embodiments of this application are described in detail below with reference to the accompanying drawings.
[0032] Figure 2 A flowchart illustrating the steps of a model training method provided for an exemplary embodiment of this application. Figure 2 The model training method shown includes the following steps:
[0033] S201, Obtain the annotation information corresponding to the first packaging image.
[0034] The labeled information refers to the actual character information on the first packaging image, which is the packaging image of the first sample food.
[0035] For example, refer to Figure 1 The first sample food is the beverage corresponding to packaging image 11, and the first packaging image is as shown in packaging image 11. The annotation information corresponding to the first packaging image is as shown in annotation information 13. Among them, annotation information 13 is the actual character information on packaging image 11, which includes: text, letters, numbers, punctuation marks, etc.
[0036] S202, parse the annotation information to obtain the first text.
[0037] The semantics of the first text are a description of the first sample food. Specifically, some characters in the annotation information are irrelevant to the first sample food; these irrelevant characters need to be removed to obtain the first text describing the first sample food. For example, refer to... Figure 1 The first text 14 includes: the description of the first sample food, “Sugar-free yet refreshing”; the ingredients of the first sample food, “water, food additives (carbon dioxide, citric acid, potassium citrate, sodium benzoate, aspartame (containing phenylalanine), acesulfame potassium, sucralose), edible flavorings); the manufacturer of the first sample food, “Guangxi C Beverage Co., Ltd.”; and the place of origin of the first sample food, “Nanning City”.
[0038] In addition, the first text is an extraction of information in the labeling information, and the first text can be in any format, which is not limited.
[0039] S203, at least partially replacing the first text with the target preset text to obtain a second text.
[0040] The target preset text has an association relationship with the first text, and the semantics of the second text is a description of the second sample food. Specifically, the target preset text has an association relationship with the corresponding text to be replaced in the first text, wherein the association relationship includes the same attribute type. For example, if the attribute type of the text to be replaced is an ingredient, the attribute type of the target preset text is also an ingredient, and if the attribute type of the text to be replaced is a place of origin, the attribute type of the target preset text is also a place of origin. If the text to be replaced is also a main ingredient in the ingredient, the target preset text is also a main ingredient in the ingredient, and if the target preset text is an additive in the ingredient, the target preset text is also an additive in the ingredient.
[0041] In addition, the obtained second text is a description of the second sample food. The second sample food can be a food of the same type as the first sample food, or a food of a different type. For example, if the first sample food is beverage A, the target preset text is an additive or a place of origin in the ingredient, and after replacing at least part of the first text with the target preset text, the obtained second text is essentially a description of another beverage B, then the second sample food is beverage B which is the same type as the first sample food. If the target preset text is a main ingredient in the ingredient, for example, wheat. Then replace the main ingredient "water" in the first text with "wheat", and the second sample food can be a food such as "biscuit", then it can be understood that the second sample food is different from the first sample food.
[0042] The training of the food classification model requires a large number of training samples, and in the embodiment of the present application, the first text is at least partially replaced by the target preset text to obtain the second text, which can expand the training samples, and does not need to process a large number of food packaging images in the manner of Figure 1 to obtain the corresponding training text, that is, the embodiment of the present application only needs to artificially label a small amount of food as shown in Figure 1 to obtain the labeling information, and further obtain the first text and the expanded second text according to the labeling information, so as to train the following food classification model.
[0043] S204, training a food classification model according to the first text and the second text to obtain a trained food classification model.
[0044] In the embodiment of the present application, the first text and the second text are both training samples, and the food classification model is trained. In addition, the food classification model can be used to predict the category of food, such as beverage, biscuit, butcher shop, preserved fruit, etc. The category can also be: sugar-free, low-sugar, medium-sugar or high-sugar food, or low-fat, full-fat or high-fat food, etc. In the embodiment of the present application, the category of food can also be set to other as needed, which is not limited here.
[0045] Further, in the training process of the food classification model, the first text and the second text are trained as training samples respectively. For example, the first text is used to train the food classification model for 100 times, and the second is used to train the food classification model for 10 times. In addition, in the embodiment of the present application, the second text can include multiple, such as 100, each second text corresponds to a food, and each second text is used as a separate training sample to train the food classification model. Therefore, the trained food classification model has good generalization ability and can classify various types of food.
[0046] Illustratively, the first sample food object in the embodiment of the present application is beverage, and the second sample food object is multiple, which is biscuit, preserved fruit, butcher shop, etc. The first text corresponding to the first sample food object and the second text corresponding to the second sample food object are used as training samples to train the food classification model, and the obtained food classification model can classify beverage, biscuit, preserved fruit, butcher shop, etc. Therefore, the food classification model has good generalization ability. And only the recognition result of the packaging image of the beverage needs to be manually labeled, and the recognition result of the packaging image of the biscuit, preserved fruit, butcher shop, etc. does not need to be manually labeled, so as to reduce the workload and improve the training efficiency of the food classification model.
[0047] Referring to Figure 3 , another model training method provided by the exemplary embodiment of the present application is provided. As Figure 3 shown, the model training method specifically includes the following steps:
[0048] S301, obtaining the annotation information corresponding to the first packaging image.
[0049] S302, extracting the association text related to the first sample food in the annotation information.
[0050] Among them, referring to Figure 1 , the first text 14 includes: association text. The association text is all the text related to the first sample food in the annotation information 13.
[0051] Illustratively, in Figure 1In the example, the associated text is "Ingredients: water, food additives (carbon dioxide, citric acid, potassium citrate, sodium benzoate, aspartame (containing phenylalanine), acesulfame, sucralose), food flavoring manufacturer: Guangxi C Beverage Co., Ltd. Address: FFFF Place of production: Nanning City."
[0052] In addition, in the embodiment of the present application, the associated text can be part or all of the text in the annotation information that is related to the first sample food. This can be set according to actual needs, and is not limited herein. For example, in the example, "Shelf life: 9 months" in the annotation information 13 is also a text related to the first sample food, but it is not extracted as the associated text in the first text 14. Figure 1
[0053] S303, determining at least one attribute information according to the associated text.
[0054] In the example, the attribute information includes: an attribute type of the first sample food, at least one attribute content text corresponding to the attribute type, and position information of the attribute content text in the associated text.
[0055] Specifically, the attribute type of the first sample food includes: ingredients, manufacturer, place of production, and / or shelf life. The attribute content text is the specific text corresponding to the attribute type. Referring to the example, in the attribute information 1, when the attribute type is ingredients, the attribute content text is: water, food additives (carbon dioxide, citric acid, potassium citrate, sodium benzoate, aspartame (containing phenylalanine), acesulfame, sucralose), food flavoring. In the attribute information 2, when the attribute type is manufacturer, the attribute content text is: Guangxi C Beverage Co., Ltd. In the attribute information 3, when the attribute type is place of production, the attribute content text is: Nanning City. Further, the position information of the attribute content text in the associated text can refer to the character position of the first letter of the attribute content text in the associated text. Referring to the example, in the attribute information 1, the first letter of the attribute content text is "water", and the character position in the associated text is 12. In the attribute information 2, the first letter of the attribute content text is "Guang", and the character position in the associated text is 69. In the attribute information 3, the first letter of the attribute content text is "Nanning", and the character position in the associated text is 84. Figure 1 Figure 1
[0056] In the embodiment of the present application, the attribute information is a further extraction of the associated text. The content text accurately describing the attribute type of the first sample food is extracted.
[0057] S304, combining the associated text and the at least one attribute information to obtain the first text.
[0058] Specifically, referring to the example, the first text 14 is obtained by combining the associated text and the attribute information 1, the attribute information 2, and the attribute information 3. Figure 1 The first text 14 is a combination of associated text and at least one attribute information.
[0059] S305, in the preset knowledge base, determine the preset text that belongs to the same attribute type as at least one attribute content text as the target preset text.
[0060] The preset knowledge base includes: multiple preset texts, and the attribute types of the preset texts.
[0061] In this embodiment, a preset knowledge base can store multiple preset texts and their attribute types. The attribute types of the preset texts can include ingredients, manufacturers, or places of origin. Ingredients can include main ingredients, additives, and auxiliary ingredients. Furthermore, the preset knowledge base can also include the types of food to which different ingredients can be applied. For example, if the main ingredient is wheat, the food type can be biscuits, bread, etc. If the main ingredient is meat, the food type can be meat jerky, dried meat, etc. For example, referring to Table 1, a preset knowledge base provided in this embodiment is illustrated. In the preset knowledge base, various main ingredients, food additives, and auxiliary ingredients are provided for different food types. Ingredients and places of origin are attribute types of the preset texts, and the preset texts are specific text content.
[0062] Table 1
[0063]
[0064] In the aforementioned preset knowledge base, there are multiple preset texts belonging to the same attribute type as the attribute content text, and you can select some to replace the corresponding parts. For example, Figure 1 If all the attribute text of the ingredients in the table, such as "water, food additives (carbon dioxide, citric acid, potassium citrate, sodium benzoate, aspartame (containing phenylalanine), acesulfame potassium, sucralose), and edible flavorings), are to be replaced, then the corresponding target preset text is "wheat, ascorbic acid, citric acid, emulsifier, eggs" in Table 1, or "corn flour, sodium bicarbonate, natural carotene, vitamin C, pepper powder", or "pork, sorbic acid and its potassium salt, lactic acid bacteria, sodium diacetate, sodium dehydroacetate, sugar", or "green plum, cyclamate, tartrazine, sugar", etc.
[0065] like Figure 1 If the text "water" is the attribute of the ingredients and is the text to be replaced, then the corresponding target preset text can be "wheat", "corn flour", "pork" or "apple" in Table 1.
[0066] In one optional embodiment, if Figure 1The attribute content text "Nanning City" corresponding to the middle-class corresponds to the content text to be replaced, and the corresponding target preset text can be "Beijing", "Shanghai", "Xi'an", "Shenzhen", etc. in Table 1.
[0067] In the embodiments of the present application, different food categories can correspond to the same main ingredient or food additive or auxiliary material.
[0068] In S306, the target preset text is used to replace the corresponding attribute content text in the first text to obtain the second text.
[0069] The replacement refers to the replacement of the corresponding associated text and the attribute content text in the attribute information in the first text. For example, referring to Figure 4 The second text 41 obtained by replacing the first text 14 in Figure 1
[0070] In the embodiments of the present application, at least one attribute content text in at least one attribute information in the first text can be replaced to obtain the second text.
[0071] In an optional embodiment, if the attribute type is an ingredient, the preset text in the preset knowledge base that belongs to the same attribute type as the at least one attribute content text is determined as the target preset text, including: in the attribute information, the attribute content text belonging to the target addition category is determined as the target content text, and the target addition category includes: main ingredient and / or food additive; in the preset knowledge base, the preset text belonging to the target addition category is determined as the target preset text, and the target preset text is used to replace the target content text.
[0072] In the embodiments of the present application, the attribute text content corresponding to the ingredient can be multiple, such as water, carbon dioxide, citric acid, etc. The attribute text content of other attribute information is usually one, such as the attribute content text corresponding to the place of origin is Nanning City, and the attribute text content corresponding to the manufacturer is Guangxi C Beverage Co., Ltd. Therefore, in the embodiments of the present application, one or all attribute text contents corresponding to the ingredient can be replaced, and all attribute content texts corresponding to the manufacturer, place of origin, etc. are replaced.
[0073] Further, the ingredient includes multiple addition categories, such as main ingredient, food additive, and auxiliary material. For example, in Figure 1 In the embodiment, the main ingredient in the ingredient is "water", the food additive is "carbon dioxide, potassium citrate, sodium benzoate, aspartame (containing phenylalanine), acesulfame, sucralose", and the auxiliary material is "edible essence". Referring to Table 1, the addition category of the ingredient in the preset knowledge base also includes the main ingredient, the food additive, and the auxiliary material. One of the addition categories in the ingredient can be taken as the target addition category, for example, the main ingredient "water" is taken as the target addition category, and only "water" in the ingredient is replaced by "wheat".
[0074] In the embodiment, a plurality of second texts can be obtained, and each second text corresponds to one or one kind of food. The food classification model trained by using the second text has good generalization ability.
[0075] In an optional embodiment, the method further includes: determining a similar character similar to the at least one target character in the first text, the similar character and the target character having shape similarity; and replacing the at least one target character with the similar character to obtain a second text.
[0076] In the embodiment, the similar character is a homograph of the target character, and the purpose is to simulate the situation that the character on the packaging image is recognized incorrectly by using the OCR technology. The second text obtained by replacing the target character with the similar character can improve the robustness of the food classification model.
[0077] Specifically, in the embodiment, the similar character can be directly used to replace the target character in the first text to obtain the second text. In the training phase, the first text, the second text obtained by replacing the attribute content text, and the second text obtained by replacing the first text with the similar character are used as training samples to train the food classification model.
[0078] In another optional embodiment, the similar character can be replaced based on the second text obtained by replacing the attribute content text to obtain a third text. The first text, the second text obtained by replacing the attribute content text, the second text obtained by replacing the first text with the similar character, and the third text are used as training samples to train the food classification model.
[0079] The application can preset a homograph knowledge base, wherein the homograph knowledge base includes a plurality of groups of homographs and an identification error rate between the groups of homographs. For example, referring to Table 2.
[0080] Table 2
[0081]
[0082] In the embodiment, the similar character can be used to randomly replace the target character in the first text or the second text, for example, the target character in the first text is replaced by the similar characterFigure 1 Replace the target text in the first text 14 therein. Specifically, randomly replace the "wu" in sugar-free with "da", the "pin" in food additive with "lv", and the "a" in aspartame with "he", and the obtained second text is referred to Figure 5 . Also, for example: for Figure 4 Replace the target text in the second text 41 therein. Randomly replace the "wu" in sugar-free with "da", the "pin" in food additive with "lv", and the "hua" in emulsifier with "|匕", and refer to Figure 6 , and the obtained third text 61.
[0083] S307, respectively determine the label data corresponding to the first text and the second text.
[0084] Among them, the label data corresponding to the first text is the normalization analysis result or classification result of the first text, and the label data corresponding to the second text is the normalization analysis result or classification result of the second text.
[0085] Specifically, the normalization analysis result means that the label data can be 0 or 1. For example, 0 can represent sugar-free, and 1 represents sugar-containing; and the classification result means that the label data can be 0, 1, 2, 3, etc. For example, 0 represents beverage, 1 represents biscuit, 2 represents dried meat, and 3 represents preserved fruit.
[0086] Furthermore, the label data can be determined according to the required function of the food classification model. For example, if the food classification model needs to determine whether the food is suitable for diabetics, then determine the label data corresponding to the first text as suitable, and the label data corresponding to the second text as not suitable. If the food classification model needs to determine the food type of the food, then determine the label data corresponding to the first text as beverage, and the label data corresponding to the second text as biscuit. In the embodiments of the present application, the label data can be set according to the training needs. Among them, the first text and the second text have their respective label data.
[0087] S308, train the food classification model according to the first text, the second text, and the label data to obtain the trained food classification model.
[0088] Among them, the number of times of training the food classification model with the first text and the label data corresponding to the first text is greater than the first number threshold; the number of times of training the food classification model with the second text and the label data corresponding to the second text is less than the second number threshold, and the second number threshold is less than the first number threshold.
[0089] In the embodiment of the present application, if a first text is replaced by m times of attribute content text to obtain m second texts, and each second text is replaced by n times of target character, then m x n new training samples are added, the first text and the label data of the first text can be used for k times of training, and each of the m x n new training samples is trained for L times. When k is greater than m x n x L, the robustness of the food classification model can be ensured.
[0090] Further, in the embodiment of the present application, the first text can also be replaced by O times of similar characters, so that one first text can be expanded to (1+O+m+m x n) training samples. That is, only one recognition result of a packaging image needs to be manually labeled to obtain (1+O+m+m x n) training samples, thereby reducing the workload of manually labeling the recognition result, reducing the dependence on labeled information, improving the training efficiency of the food classification model, and obtaining a food classification model with high generalization ability and robustness.
[0091] Figure 7 A step flowchart of a packaging image recognition method provided by an exemplary embodiment of the present application is shown in FIG. 7. Figure 7 The packaging image recognition method specifically includes the following steps:
[0092] S701, obtaining a packaging image of a target food.
[0093] In the embodiment of the present application, the target food can be any kind of food, such as beverage, biscuit, preserved fruit, and meat biscuit, etc. The packaging image of the target food can be collected by a camera.
[0094] S702, recognizing the recognition character information in the packaging image by using an image recognition technology.
[0095] The image recognition technology is, for example, an OCR technology, and the recognized recognition character information can be deviated from the actual character information on the packaging image.
[0096] S703, analyzing the recognition character information to obtain text information.
[0097] The format of the text information is as shown in the format of the first text 14 in FIG. 7, and the specific analysis manner can refer to the analysis of the actual character information described above, which will not be repeated here. Figure 1
[0098] S704, inputting the text information into a food classification model for classification processing to obtain a classification result of the target food.
[0099] The food classification model is trained by the model training method in the above embodiment. The obtained classification result can be sugar or sugar-free, and can also be low-fat or high-fat, etc.
[0100] In the embodiment of the present application, the food classification model trained as described above can recognize the recognition character information recognized from the packaging images of various target foods, and obtain the classification result of the corresponding target food. The food classification model has good generalization ability and robustness.
[0101] In the embodiment of the present application, in addition to providing a packaging image recognition method, a packaging image recognition device is also provided, as shown in Figure 8 The packaging image recognition device 80 includes:
[0102] The acquisition module 81 is configured to acquire a packaging image of a target food.
[0103] The recognition module 82 is configured to recognize, by using an image recognition technology, recognition character information in the packaging image.
[0104] The analysis module 83 is configured to analyze the recognition character information to obtain text information, the semantics of the text information being a description of the target food.
[0105] The processing module 84 is configured to input the text information into a food classification model for classification processing to obtain a classification result of the target food, the food classification model being trained according to the model training method in the above embodiment.
[0106] The packaging image recognition device provided in the embodiment of the present application can recognize the recognition character information recognized from the packaging images of various target foods by using the food classification model trained as described above, and obtain the classification result of the corresponding target food.
[0107] In addition, the embodiment of the present application also provides a model training device (not shown), which includes:
[0108] The acquisition module is configured to acquire annotation information corresponding to a first packaging image, the annotation information being actual character information on the first packaging image, and the first packaging image being a packaging image of a first sample food.
[0109] The analysis module is configured to analyze the annotation information to obtain a first text, the semantics of the first text being a description of the first sample food.
[0110] The replacement module is configured to replace at least part of the first text with a target preset text to obtain a second text, the target preset text having an associated relationship with the first text, and the semantics of the second text being a description of a second sample food.
[0111] The training module is configured to train a food classification model according to the first text and the second text to obtain a trained food classification model.
[0112] In an optional embodiment, the parsing module is specifically configured to: extract association text associated with the first sample food in the annotation information; determine at least one attribute information according to the association text, the attribute information comprising: an attribute type of the first sample food, at least one attribute content text corresponding to the attribute type, and position information of the attribute content text in the association text; and combine the association text and the at least one attribute information to obtain the first text.
[0113] In an optional embodiment, the replacing module is specifically configured to: determine, in the preset knowledge base, a preset text that is of the same attribute type as the at least one attribute content text as a target preset text, the preset knowledge base comprising: a plurality of preset texts, and attribute types of the preset texts; and replace the corresponding attribute content text in the first text with the target preset text to obtain the second text.
[0114] In an optional embodiment, if the attribute type is: ingredients, when the replacing module determines, in the preset knowledge base, a preset text that is of the same attribute type as the at least one attribute content text as a target preset text, the replacing module is specifically configured to: determine, in the attribute information, an attribute content text that belongs to a target addition category as a target content text, the target addition category comprising: main ingredients and / or food additives; and determine, in the preset knowledge base, a preset text that belongs to the target addition category as the target preset text, the target preset text being used to replace the target content text.
[0115] In an optional embodiment, the replacing module is further configured to: determine a similar character that is similar to at least one target character in the first text, the similar character and the target character having a similarity in shape; and replace the at least one target character with the similar character as the target preset text to obtain the second text.
[0116] In an optional embodiment, the training module is specifically configured to: determine label data corresponding to the first text and the second text respectively, the label data corresponding to the first text being a normalization analysis result or a classification result of the first text, and the label data corresponding to the second text being a normalization analysis result or a classification result of the second text; train the food classification model according to the first text, the second text, and the label data to obtain the trained food classification model; wherein the number of times of training the food classification model by using the first text and the label data corresponding to the first text is greater than a first threshold number of times; and the number of times of training the food classification model by using the second text and the label data corresponding to the second text is less than a second threshold number of times, the second threshold number of times being less than the first threshold number of times.
[0117] The model training apparatus provided in the embodiments of the present application can reduce the dependence on annotation information, and can train a food classification model with good generalization ability and robustness.
[0118] In addition, in some of the processes described in this specification, including the processes described in the description and the drawings, multiple operations are described in a particular, sequential order. However, it should be understood that unless otherwise specifically stated in this specification, the order of execution or performance of the operations can be changed, or omitted, or combined, or performed in parallel, or performed at different times, without departing from the spirit of the application. In addition, the processes can include more or fewer operations than those described in this specification, and these operations can be performed in sequential order or in parallel, or performed at different times. It should be noted that the descriptions "first", "second", and the like in this specification are used to distinguish different messages, devices, modules, and the like, and do not represent the order of execution or the type of "first" and "second".
[0119] Figure 9 A structural diagram of an electronic device is provided for the exemplary embodiments of the present application. The electronic device is used to run the model training method and the package image recognition method described above. As shown in the figure, the electronic device includes a memory 94 and a processor 95. Figure 9
[0120] The memory 94 is used to store computer programs and can be configured to store various other data to support operations on the electronic device. The memory 94 can be an object storage service (OSS).
[0121] The memory 94 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk or optical disk.
[0122] The processor 95 is coupled to the memory 94 and is used to execute the computer programs in the memory 94, to: obtain annotation information corresponding to a first package image, the annotation information being actual character information on the first package image, the first package image being a package image of a first sample food; parse the annotation information to obtain a first text, the semantic of the first text being a description of the first sample food; replace at least part of the first text with a target preset text to obtain a second text, the target preset text having an associated relationship with the first text, the semantic of the second text being a description of a second sample food; and train a food classification model according to the first text and the second text to obtain a trained food classification model.
[0123] Further optionally, the processor 95, when parsing the annotation information to obtain the first text, is specifically configured to: extract associated text in the annotation information related to the first sample food; determine at least one attribute information according to the associated text, the attribute information including: an attribute type of the first sample food, at least one attribute content text corresponding to the attribute type, and position information of the attribute content text in the associated text; and combine the associated text and the at least one attribute information to obtain the first text.
[0124] Further optionally, the processor 95, when replacing the first text with the target preset text to obtain the second text, is specifically configured to: determine, in the preset knowledge base, a preset text that is of the same attribute type as the at least one attribute content text as the target preset text, the preset knowledge base including: a plurality of preset texts, and attribute types of the preset texts; and replace the corresponding attribute content text in the first text with the target preset text to obtain the second text.
[0125] Further optionally, the processor 95, when determining, in the preset knowledge base, a preset text that is of the same attribute type as the at least one attribute content text as the target preset text, is specifically configured to: determine, in the attribute information, an attribute content text belonging to a target addition category as a target content text, the target addition category including: main ingredients and / or food additives; and determine, in the preset knowledge base, a preset text belonging to the target addition category as the target preset text, the target preset text being used to replace the target content text.
[0126] Further optionally, the processor 95 is further configured to determine a similar character similar to at least one target character in the first text, the similar character and the target character having a similarity in shape; and replace the at least one target character with the similar character as the target preset text to obtain the second text.
[0127] Further optionally, the processor 95, when training the food classification model according to the first text and the second text to obtain the trained food classification model, is specifically configured to: determine label data corresponding to the first text and the second text respectively, the label data corresponding to the first text being a normalized analysis result or a classification result of the first text, and the label data corresponding to the second text being a normalized analysis result or a classification result of the second text; train the food classification model according to the first text, the second text, and the label data to obtain the trained food classification model; wherein the number of times of training the food classification model with the first text and the label data corresponding to the first text is greater than a first threshold number of times; and the number of times of training the food classification model with the second text and the label data corresponding to the second text is less than a second threshold number of times, the second threshold number of times being less than the first threshold number of times.
[0128] In an optional embodiment, the processor 95, coupled with the memory 94, is configured to execute a computer program in the memory 94, and is further configured to: acquire a packaging image of a target food; identify, by using an image recognition technology, identification character information in the packaging image; parse the identification character information to obtain text information, the semantics of the text information being a description of the target food; and input the text information into a food classification model for classification processing to obtain a classification result of the target food, the food classification model being obtained by the above-mentioned model training method.
[0129] Further, as shown in Figure 9 , the electronic device further includes a firewall 91, a load balancer 92, a communication component 96, a power supply component 98, and other components. Figure 9 Some components are only schematically shown in the electronic device, and it does not mean that the electronic device only includes Figure 9 the components shown.
[0130] The electronic device provided by the embodiments of the present application realizes expansion of the training sample by using the target preset text to at least partially replace the first text on the first packaging image, avoids the need for manual labeling of characters on packaging images of various foods to obtain sufficient training samples, and thus reduces the labeling workload of the staff and improves the efficiency of model training. Moreover, the obtained food classification model can accurately classify target food.
[0131] Correspondingly, the embodiments of the present application also provide a computer readable storage medium storing a computer program, when the computer program / instruction is executed by a processor, the processor is caused to implement the steps in the method shown in Figure 2 , Figure 3 or Figure 7 .
[0132] Correspondingly, the embodiments of the present application also provide a computer program product, including a computer program / instruction, when the computer program / instruction is executed by a processor, the processor is caused to implement the steps in the method shown in Figure 2 , Figure 3 or Figure 7 .
[0133] The above Figure 9The communication component in the electronic device 100 is configured to facilitate wired or wireless communication between the electronic device 100 and other devices. The electronic device 100 can access a wireless network based on a communication standard, such as WiFi, a 2G, 3G, 4G / LTE, 5G, or the like mobile communication network, or a combination thereof. In an example embodiment, the communication component receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In an example embodiment, the communication component further includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on Radio Frequency Identification (RFID) techniques, infrared data association (IrDA) techniques, ultra-wideband (UWB) techniques, Bluetooth (BT) techniques, and other techniques.
[0134] The power component in the electronic device 100 provides power to various components of the electronic device 100. The power component can include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the electronic device 100. Figure 9
[0135] Those skilled in the art will appreciate that embodiments of the present application can be readily used as a method, apparatus, or computer program product. The present application can be implemented in hardware- or software-only embodiments, or in embodiments that include both software and hardware. Furthermore, the present application can be implemented in a computer program product that can be executed on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, and the like).
[0136] The present application is described in reference to the flowchart illustrations and / or block diagrams of methods, apparatus (systems) and computer program products according to embodiments of the application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general purpose computer, special purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions specified in the flowchart illustrations and / or block diagrams. Figure 1 The means for performing each of the functions specified in the flowchart illustrations and / or block diagrams can be embodied in one or more computer-readable storage media that include computer-readable instructions that, when executed by a processor of a computer, other programmable data processing apparatus, or other means for parallel processing, generate a machine implemented process such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions specified in the flowchart illustrations and / or block diagrams. Figure 1 The means for performing each of the functions specified in the flowchart illustrations and / or block diagrams can be embodied in one or more computer-readable storage media that include computer-readable instructions that, when executed by a processor of a computer, other programmable data processing apparatus, or other means for parallel processing, generate a machine implemented process such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions specified in the flowchart illustrations and / or block diagrams.
[0137] These computer program instructions can also be stored in a computer-readable storage medium that can direct a computer, other programmable data processing apparatus, or other means to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instructions which implement the functions specified in the flowchart illustrations and / or block diagrams. Figure 1 The means for performing each of the functions specified in the flowchart illustrations and / or block diagrams can be embodied in one or more computer-readable storage media that include computer-readable instructions that, when executed by a processor of a computer, other programmable data processing apparatus, or other means for parallel processing, generate a machine implemented process such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions specified in the flowchart illustrations and / or block diagrams. Figure 1 the function(s) specified in the block or blocks.
[0138] These computer program instructions can also be loaded into computer or other programmable data processing devices to cause a series of operational steps to be performed on the computer or other programmable devices to produce a computer implemented process such that the instructions which execute on the computer or other programmable devices provide steps for implementing the flowchart block(s) or flowchart flow(s) and / or portions thereof. Figure 1 the flowchart block(s) or flowchart flow(s) and / or portions thereof. Figure 1 the function(s) specified in the block or blocks.
[0139] In one typical configuration, the computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.
[0140] The memory can include non-persistent memory and / or volatile memory, such as random access memory (RAM) about which the computer stores the information. The memory is an example of computer readable media.
[0141] Computer readable media includes permanent and non-permanent, removable and non-removable media implemented in any method or technology for storage of information. The information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read only memory (ROM), electrically erasable programmable read only memory (EEPROM), flash memory or other memory technology, compact disc read only memory (CD-ROM), digital versatile disc (DVD), or other optical storage, magnetic cassettes, magnetic tapes, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible to computing devices. According to the definition herein, computer readable media does not include transitory media, such as modulated data signals and carrier waves.
[0142] It should also be noted that the terms "comprising", "including", or any other variation thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article or apparatus that comprises a list of elements does not include only those elements recited, but can also include other elements not expressly listed or inherent to such process, method, article or apparatus. Without further limitation, an element preceded by "comprises a" does not, without more constraints, foreclose the existence of additional identical elements in the process, method, article or apparatus that includes the element.
[0143] The above merely provides an example of the present application, and is not intended to limit the present application. For those skilled in the art, the present application can have various modifications and changes. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall fall into the scope of claims of the present application.
Claims
1. A model training method, characterized in that, The method comprises the following steps: obtaining annotation information corresponding to a first packaging image, the annotation information being actual character information on the first packaging image, the first packaging image being a packaging image of a first sample food; extracting associated text in the annotation information related to the first sample food; determining at least one attribute information according to the associated text, the attribute information comprising an attribute type of the first sample food, at least one attribute content text corresponding to the attribute type, and position information of the attribute content text in the associated text; combining the associated text and the at least one attribute information to obtain a first text, the semantic of the first text being a description of the first sample food; adopting a target preset text to at least partially replace the first text to obtain a second text, the target preset text having an associated relationship with the first text, the semantic of the second text being a description of a second sample food; training a food classification model according to the first text and the second text to obtain a trained food classification model.
2. The model training method of claim 1, wherein, The method comprises the following steps: determining, in a preset knowledge base, a preset text belonging to the same attribute type as at least one attribute content text as the target preset text, the preset knowledge base comprising a plurality of preset texts and attribute types of the preset texts; adopting the target preset text to replace the corresponding attribute content text in the first text to obtain the second text.
3. The model training method of claim 2, wherein, If the attribute type is ingredients, the method comprises the following steps: determining, in the attribute information, an attribute content text belonging to a target addition category as a target content text, the target addition category comprising main ingredients and / or food additives; determining, in the preset knowledge base, a preset text belonging to the target addition category as the target preset text, the target preset text being used to replace the target content text.
4. The model training method of claim 2, wherein, The method further comprises the following steps: determining a similar character similar to at least one target character in the first text, the similar character and the target character having similarity in shape; adopting the similar character as the target preset text to replace the at least one target character to obtain the second text. 5.The model training method of any one of claims 1 to 4, wherein, The method comprises the following steps: determining label data corresponding to the first text and the second text respectively, the label data corresponding to the first text being a normalized analysis result or a classification result of the first text, the label data corresponding to the second text being a normalized analysis result or a classification result of the second text; training a food classification model according to the first text, the second text, and the label data to obtain a trained food classification model. The number of times of training the food classification model by using the first text and the label data corresponding to the first text is greater than a first number threshold, and the number of times of training the food classification model by using the second text and the label data corresponding to the second text is less than a second number threshold, the second number threshold being less than the first number threshold.
6. A method of identifying a package image, characterized by, The method comprises: obtaining a packaging image of a target food; recognizing, by using an image recognition technology, recognition character information in the packaging image; parsing the recognition character information to obtain text information, the semantics of the text information being a description of the target food; inputting the text information into the food classification model for classification processing to obtain a classification result of the target food, the food classification model being trained according to the model training method in any one of claims 1 to 5.
7. An apparatus for identifying a package image, the apparatus comprising: The method comprises: an obtaining module configured to obtain a packaging image of a target food; an identifying module configured to recognize, by using an image recognition technology, recognition character information in the packaging image; a parsing module configured to parse the recognition character information to obtain text information, the semantics of the text information being a description of the target food; a processing module configured to input the text information into the food classification model for classification processing to obtain a classification result of the target food, the food classification model being trained according to the model training method in any one of claims 1 to 5.
8. An electronic device, comprising: The method comprises: a processor, a memory, and a computer program stored on the memory and executable on the processor, the processor implementing the model training method in any one of claims 1 to 5 or the packaging image recognition method in claim 6 when executing the computer program.
9. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program, which, when executed by the processor, causes the processor to implement the model training method in any one of claims 1 to 5 or the packaging image recognition method in claim 6.
10. A computer program product, characterised in that, The computer program, when executed by the processor, causes the processor to implement the model training method in any one of claims 1 to 5 or the packaging image recognition method in claim 6.
Citation Information
Patent Citations
Systems and methods for food analysis, personalized recommendations, and health management
CN111902878A
Information extraction model training method and device and electronic equipment
CN113420533A
Cited By
Packaging image acquisition method and apparatus for eye movement experiment
CN122618407A