A picture annotation method based on Chinese herbal medicine graphic and text modality data
By extracting the graphic and text annotation pairs in the graphic and text mode data of Chinese herbal medicine, combining edge detection, maximum inter-class variance method and graphic technology, OCR and NLP technologies are used for identification and analysis, the problem of low accuracy of the graphic and text annotation data of Chinese herbal medicine is solved, and efficient and accurate automatic labeling and identification effects are achieved.
Patent Information
- Application Number
- CN202210388664.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-14
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2042-04-14
AI Technical Summary
The prior art is difficult to effectively process the graphic and text modal data of Chinese herbal medicines, resulting in the problem of decreasing the accuracy of picture labeling and loss of information.
A picture annotation method based on Chinese herbal graphic and text modal data is proposed. By extracting the picture annotation pair from the picture image data, the edge detection and maximum inter-class variance algorithm are used to extract the edge and front background thresholds of the picture, combined with graphic expansion and water flooding method to obtain the sub-picture mask, and then the text is identified and parsed through OCR and NLP technology to generate the picture annotation pair.
Automatic labeling of the original and medicinal forms of Chinese herbal medicines is realized, the labeling efficiency is improved, the error introduced by manual labeling is reduced, and the recognition effect is improved through the semantic consistency constraint model.
Smart Images

Figure CN114742161B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image annotation, and in particular to a method for annotating images based on Chinese herbal medicine graphic and text modality data. Background Art
[0002] As an important part of Chinese traditional culture, Chinese herbal medicine contains rich cultural symbols and cultural connotations. With the advent of the era of cultural big data, Chinese herbal medicine data also has rich digital resources. Integrating technology and culture to explore the cultural content contained in Chinese herbal medicine is a widely mentioned topic. Using annotation methods to analyze the categories and functions of Chinese herbal medicine is a scientific and technological method that can achieve cultural identification, cultural interpretation, and cultural inheritance.
[0003] Traditional Chinese medicine is mainly composed of botanicals, animal medicines and mineral medicines. Because botanicals account for the majority of traditional Chinese medicine, traditional Chinese medicine is also called Chinese herbal medicine. There are about 5,000 kinds of traditional Chinese medicine used in various parts of China, and the prescriptions formed by combining various medicinal materials are countless. After thousands of years of research, an independent science has been formed - herbal medicine. The significance of studying the association annotation algorithm of traditional Chinese medicine is that it can free traditional Chinese medicine plant workers from tedious and repetitive work, reduce labor costs, and allow traditional Chinese medicine plant workers to spend more time and energy on valuable traditional Chinese medicine research.
[0004] There are two modes of Chinese herbal medicine image data. The first is the original form of Chinese herbal medicine, which is generally manifested as wild plants, animals, etc. The second is the medicinal form of Chinese herbal medicine, that is, the Chinese herbal medicine that can be directly used after artificial processing and other industrial processes. We call it the finished drug form. In general annotation algorithms, due to the large differences and weak correlation between the two modes, they should be treated as different classification and annotation tasks. However, this will lose the fact that both modes have the same label attribute (semantic consistency) and have certain connections in some special categories (such as the high similarity between the two forms of honeycomb or tortoise shell data). In this case, if we use the two forms of image data separately, it will lead to information loss in the algorithm model, resulting in reduced accuracy, inaccurate object recognition and other problems; at the same time, if we confuse the two modes for recognition, it will cause more serious recognition accuracy problems. It can be seen that both methods are feasible to a certain extent, but there are also some problems in specific situations.
[0005] Currently, no effective solution has been proposed for the problems in the related technologies. Summary of the invention
[0006] In view of the problems in the related art, the present invention proposes an image annotation method based on Chinese herbal medicine graphic and text modal data to overcome the above-mentioned technical problems existing in the existing related art.
[0007] To this end, the specific technical solution adopted by the present invention is as follows:
[0008] A method for annotating images based on Chinese herbal medicine graphic and text modal data, the method comprising the following steps:
[0009] S1, extracting image-text annotation pairs from Chinese herbal medicine image data;
[0010] S2, extract the category labels and images separately to produce a Chinese herbal medicine image and text annotation dataset;
[0011] S3. Build a Chinese herbal medicine association labeling algorithm model with semantic consistency constraints and conduct training.
[0012] Furthermore, the step of extracting the image-text annotation pairs from the image-text data comprises the following steps:
[0013] S11, extracting and binarizing the images in the graphic image data;
[0014] S12, connecting adjacent connected domains, obtaining a sub-graph mask, and extracting the remaining text;
[0015] S13, recognizing the extracted text by optical character recognition technology, and outputting the recognition result;
[0016] S14. Use natural language processing technology and keyword extraction technology to perform semantic analysis and semantic splitting to obtain graphic and text results.
[0017] Furthermore, the extraction and binarization of the images in the graphic image data comprises the following steps:
[0018] S111, extracting an edge binary image of the image using an edge detection operator;
[0019] S112, extracting the foreground and background thresholds of the image using the maximum inter-class variance algorithm, and performing binarization processing on the image;
[0020] S113, superimposing the edge binary image and the foreground and background binary images to obtain a mask.
[0021] Furthermore, the connecting of adjacent connected domains, obtaining a subgraph mask, and extracting the remaining text includes the following steps:
[0022] S121. Use graphics expansion to open up adjacent connected domains;
[0023] S122, filtering the text part according to the size of the connected domain, and retaining the sub-graph part;
[0024] S123, using the flooding method to fill the sub-image connected domain to obtain the sub-image mask;
[0025] S124, extracting the sub-image part through the mask, and extracting the remaining text part.
[0026] Furthermore, the process of recognizing the extracted text by optical character recognition technology and outputting the recognition result comprises the following steps:
[0027] S131, select the convolutional recurrent neural network + text recognition network structure network model, and capture the forward and backward information simultaneously through the double-layer long short-term memory network structure in the recurrent neural network;
[0028] S132, the text recognition network transcription layer judges the results of the model output through statistical principles, selects the result with the highest probability to output, and obtains the final recognition result.
[0029] Furthermore, the use of natural language processing technology and keyword extraction technology to perform semantic analysis and semantic splitting to obtain graphic and text results includes the following steps:
[0030] S141. Use text segmentation technology to break up the paragraphs into multi-level word vectors;
[0031] S142, removing stop words through a stop word list;
[0032] S143, using an improved text sorting algorithm to extract core words in the text in the form of graph vectors, and then extracting text content corresponding to different keywords from the text to achieve parsing and extraction of semantic tags;
[0033] S144: assign labels to the separated corresponding sub-graphs to obtain the final graphic and text results.
[0034] Furthermore, the construction of a Chinese herbal medicine association annotation algorithm model with semantic consistency constraints and training includes the following steps:
[0035] S31, using the Chinese herbal medicine graphic and text annotation dataset as a training set for the algorithm model;
[0036] S32, build a model algorithm based on the extreme version of the beginning feature extraction network model;
[0037] S33, using the idea of transfer learning to train a single-form algorithm model;
[0038] S34. Train the double-layer algorithm model to improve the training speed and model effect.
[0039] Furthermore, the model algorithm is constructed based on the extreme version of the beginning feature extraction network model, including the following steps:
[0040] S321. Based on the extreme version of the initial feature extraction network model, the model structure of input and recognition results through dual streams;
[0041] S322. Connect a group of pooling layers and fully connected layers in series, use the splicing method to splice the inputs of the two input streams into the same feature dimension, and then connect them to the fully connected layer for labeling and recognition.
[0042] Furthermore, the idea of using transfer learning to train a single-form algorithm model includes the following steps:
[0043] S331, using the idea of transfer learning to load the single morphology algorithm model with the starting feature extraction network model parameters of the extreme version trained on the visualization database data set, and adding a fully connected layer;
[0044] S332, using the preprocessed training set, locking all layers except the fully connected layer for training;
[0045] S333. Unlock the last eight layers of the model and retrain them, saving the final model data.
[0046] Furthermore, the two-layer algorithm model is trained to improve the training speed and model effect, including the following steps:
[0047] S341, training each layer model, and importing the trained weight parameters into the double-layer model for further training;
[0048] S342, building a two-layer algorithm model and loading the parameters of the trained single-form algorithm model except the fully connected layer, and after importing the model, locking the layers except the last fully connected layer for training;
[0049] S343, extracting data images of different modes as training data by random sampling, and randomly adding blank images as noise;
[0050] S344. Unlock the last six layers of the model, retrain the model, and save the model results.
[0051] The beneficial effects of the present invention are as follows: extracting a picture and a label pair from the graphic modal data, extracting the edge binary image of the picture using the Sobel operator and extracting the foreground and background threshold of the picture using the OTSU algorithm, superimposing the two results, using the graphics expansion method to open up the adjacent connected domains, filtering the text part by the size of the connected domain, and filling the connected domain with the water flooding method to obtain the sub-image mask, then identifying the text part by the OCR technology, and completing the label extraction work by disassembling the semantics by the NLP technology to obtain the graphic and text annotation pairs, and obtaining two sets of data sets of the original form and the medicinal form of Chinese herbal medicine in this way. Thus, the data set can be produced by machine automatic labeling instead of manual labeling, improving efficiency while reducing errors caused by the uncertainty of manual labeling.
[0052] After obtaining and constructing a complete Chinese herbal medicine data set, it will be used as the training set of the model. Through the constructed semantic consistency constraint model, the image and text pairs are input into the model as the training set for training. After loading the Xception network model parameters pre-trained by ImageNet and adding a fully connected layer, the network training in a single form is performed twice. Two sets of parameters are trained twice, and then a semantic consistency constraint model is built. The trained parameters are pre-loaded into the model for the third training until convergence, until the construction and training of the recognition model are completed. By building the algorithm model system and using the system, the processing, handling, display, collection and other functions of Chinese herbal medicine image and text data are provided and implemented, and the accuracy of the algorithm is continuously optimized using the collected data, greatly improving the recognition effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0054] Figure 1 is a flow chart of a method for annotating an image based on Chinese herbal medicine graphic and text modality data according to an embodiment of the present invention;
[0055] Figure 2 This is a diagram of the overall operation steps of a method for annotating an image based on Chinese herbal medicine graphic and text modality data according to an embodiment of the present invention;
[0056] Figure 3 It is a schematic diagram of the construction and training process of a semantic consistency constraint model for an image annotation method based on Chinese herbal medicine graphic and text modal data according to an embodiment of the present invention. DETAILED DESCRIPTION
[0057] According to an embodiment of the present invention, a method for annotating an image based on Chinese herbal medicine graphic and text modality data is provided.
[0058] The present invention will now be further described with reference to the accompanying drawings and specific embodiments. Figure 1-3 As shown, according to an embodiment of the present invention, a method for labeling an image based on Chinese herbal medicine graphic and text modality data comprises the following steps:
[0059] S1, extracting picture-text annotation pairs from Chinese herbal medicine picture-text image data, including the following steps:
[0060] S11, extracting and binarizing the images in the graphic image data;
[0061] Wherein, step S11 includes the following steps:
[0062] S111, extracting an edge binary image of the image using an edge detection (Sobel) operator;
[0063] S112, extracting the foreground and background thresholds of the image using the OTSU algorithm, and performing binarization processing on the image;
[0064] S113, superimposing the edge binary image with the foreground and background binary images to obtain a mask, which can not only distinguish the foreground and background information, but also retain the original edge information.
[0065] S12, connecting adjacent connected domains, obtaining a sub-graph mask, and extracting the remaining text;
[0066] Wherein, step S12 includes the following steps:
[0067] S121. Use graphics expansion to open up adjacent connected domains;
[0068] S122, filtering the text part according to the size of the connected domain, and retaining the sub-graph part;
[0069] S123, using the flooding method to fill the sub-image connected domain to obtain the sub-image mask;
[0070] S124, extracting the sub-image part through the mask, and extracting the remaining text part.
[0071] S13, recognizing the extracted text by optical character recognition (OCR) technology, and outputting the recognition result;
[0072] Wherein, step S13 comprises the following steps:
[0073] S131. Select the convolutional recurrent neural network + text recognition network (CRNN+CTC) structure network model, and use the double-layer long short-term memory network (LSTM) structure in the recurrent neural network (RNN) to simultaneously capture forward and backward information, thereby improving the model's ability to capture contextual information and improving the model effect;
[0074] S132, the character recognition network (CTC) transcription layer uses statistical principles to judge the results of the model output, selects the result with the highest probability to output, and obtains the final recognition result.
[0075] S14. Use natural language processing (NLP) keyword extraction technology to perform semantic analysis and semantic splitting to obtain graphic and text results.
[0076] Wherein, step S14 comprises the following steps:
[0077] S141. Use text segmentation technology to break up the paragraph into multi-level word vectors;
[0078] S142, removing stop words through a stop word list;
[0079] S143, using an improved text ranking (TextRank) algorithm to extract core words in the text in the form of graph vectors, and then extracting text content corresponding to different keywords from the text to achieve parsing and extraction of semantic tags;
[0080] S144: assign labels to the separated corresponding sub-graphs to obtain the final graphic and text results.
[0081] S2, extract the category labels and images separately to produce a Chinese herbal medicine image and text annotation dataset;
[0082] Since each image already has a set of sub-labels and a category main label, the category labels and images are extracted separately to produce a Chinese herbal medicine graphic annotation dataset, and used as a training set for the following algorithm model. The main purpose of step S1 is to produce a dataset through automatic machine labeling instead of manual labeling, thereby improving efficiency and reducing errors caused by the uncertainty of manual labeling.
[0083] S3. Build a Chinese herbal medicine association annotation algorithm model with semantic consistency constraints and train it, including the following steps:
[0084] S31, using the Chinese herbal medicine graphic and text annotation dataset as a training set for the algorithm model;
[0085] S32, build a model algorithm based on the extreme version of the beginning feature extraction network (Xception network) model;
[0086] Wherein, step S32 comprises the following steps:
[0087] S321, based on the extreme version of the beginning feature extraction network (Xception network) model, the model structure of input and recognition results through dual streams;
[0088] S322. Connect a group of pooling layers and fully connected layers in series, use the splicing method to splice the inputs of the two input streams into the same feature dimension, and then connect them to the fully connected layer for labeling and recognition.
[0089] S33, using the idea of transfer learning to train a single-form algorithm model;
[0090] Wherein, step S33 comprises the following steps:
[0091] S331, using the idea of transfer learning to load the model parameters of the beginning feature extraction network (Xception network) in the extreme version trained on the visualization database (ImageNet) data set into the single morphology algorithm model, and adding a fully connected layer;
[0092] S332, using the preprocessed training set (preprocessing includes performing random rotation and affine transformation on the training data to expand the data set), locking all layers except the fully connected layer for training;
[0093] S333. Unlock the last eight layers of the model and retrain them, saving the final model data.
[0094] For example, the early-stop monitor is set to val-loss for training, and the final convergence results on val-accuracy are 0.482 and 0.431. Then the last eight layers of the model are unlocked (the experiment shows the best effect) and re-trained. The final convergence results are 0.772 and 0.731. It can be seen that good results have been achieved in the single-modality classification problem, and the model data is saved.
[0095] S34. Train the double-layer algorithm model to improve the training speed and model effect.
[0096] Wherein, step S34 comprises the following steps:
[0097] S341, training each layer model, and importing the trained weight parameters into the double-layer model for further training;
[0098] S342, building a two-layer algorithm model and loading the parameters of the trained single-form algorithm model except the fully connected layer, and after importing the model, locking the layers except the last fully connected layer for training;
[0099] S343, extracting data images of different modes as training data by random sampling (random rotation and affine transformation, etc. are also required), and randomly adding blank images as noise;
[0100] S344. Unlock the last six layers of the model, retrain the model, and save the model results.
[0101] In the experiment, the probability was taken as 0.2, and the final convergence result was 0.492. After that, the last 6 layers of the model were unlocked and the model was retrained. The final convergence result was val-accuracy = 0.801, val-loss = 0.23, for a total of 30 epochs.
[0102] Therefore, the principle of the present invention can be summarized as follows: first, the sub-image part and the text part in the image are disassembled by a graphics method; then, the text part is recognized by OCR and the paragraph is broken into words by NLP technology, and the keywords and key contents are extracted, and the results are assigned to the corresponding sub-images; then, the image and text data are input into the Figure 3 The trained annotation model is obtained by training in the semantic consistency constraint model shown in the figure; finally, an algorithm model system is built, and the system is used to annotate and identify Chinese herbal medicine image data, and the algorithm is continuously optimized by collecting data through the front-end system to expand the data set.
[0103] In summary, with the help of the above technical solution of the present invention, the image and label pair are extracted from the image-text modal data, the edge binary image of the image is extracted using the Sobel operator and the foreground and background threshold of the image is extracted using the OTSU algorithm, the two results are superimposed, the adjacent connected domains are opened up using the expansion method of graphics, and the text part is filtered by the size of the connected domain. At the same time, the connected domain is filled with the water flooding method to obtain the sub-image mask, and then the text part is recognized by the OCR technology, and the semantics is disassembled by the NLP technology to complete the label extraction work, and the image-text annotation pair is obtained. In this way, two sets of data sets of the original form and the medicinal form of Chinese herbal medicine are obtained. Therefore, the data set can be produced by automatic machine labeling instead of manual labeling, which improves efficiency while reducing errors caused by the uncertainty of manual labeling.
[0104] After obtaining and constructing a complete Chinese herbal medicine data set, it will be used as the training set of the model. Through the constructed semantic consistency constraint model, the image and text pairs are input into the model as the training set for training. After loading the Xception network model parameters pre-trained by ImageNet and adding a fully connected layer, the network training in a single form is performed twice. Two sets of parameters are trained twice, and then a semantic consistency constraint model is built. The trained parameters are pre-loaded into the model for the third training until convergence, until the construction and training of the recognition model are completed. By building the algorithm model system and using the system, the processing, handling, display, collection and other functions of Chinese herbal medicine image and text data are provided and implemented, and the accuracy of the algorithm is continuously optimized using the collected data, greatly improving the recognition effect.
[0105] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. A picture annotation method based on Chinese herbal medicine graphic and text modality data, It is characterized in that The method comprises the following steps: S1, extracting image-text annotation pairs from Chinese herbal medicine image data; S2, extract the category labels and images separately to produce a Chinese herbal medicine image and text annotation dataset; S3. Build a Chinese herbal medicine association annotation algorithm model with semantic consistency constraints and conduct training; The method of building a semantic consistency-constrained Chinese herbal medicine association annotation algorithm model and training it includes the following steps: S31, using the Chinese herbal medicine image and text annotation dataset as a training set for the algorithm model, wherein the Chinese herbal medicine image and text annotation dataset includes Chinese herbal medicine plant morphology pictures and Chinese herbal medicine medicinal morphology pictures; S32, build a model algorithm based on the extreme version of the beginning feature extraction network model; S33, using the idea of transfer learning to train a single-form algorithm model; S34, training the double-layer algorithm model to improve the training speed and model effect; The model algorithm based on the extreme version of the beginning feature extraction network model includes the following steps: S321. Based on the extreme version of the initial feature extraction network model, the model structure of input and recognition results through dual streams; S322, connect a group of pooling layers and fully connected layers in series, use the splicing method to splice the inputs of the two input streams into the same feature dimension, and then connect to the fully connected layer for labeling and recognition; The method of using the idea of transfer learning to train a single-form algorithm model includes the following steps: S331, using the idea of transfer learning to load the single morphology algorithm model with the starting feature extraction network model parameters of the extreme version trained on the visualization database data set, and adding a fully connected layer; S332, using the preprocessed training set, locking all layers except the fully connected layer for training; S333, unlock the last eight layers of the model and retrain, and save the final model data; The method of training the double-layer algorithm model to improve the training speed and model effect includes the following steps: S341, training each layer model, and importing the trained weight parameters into the double-layer model for further training; S342, building a two-layer algorithm model and loading the parameters of the trained single-form algorithm model except the fully connected layer, and after importing the model, locking the layers except the last fully connected layer for training; S343, extracting data images of different modes as training data by random sampling, and randomly adding blank images as noise; S344. Unlock the last six layers of the model, retrain the model, and save the model results.
2. A method for labeling images based on Chinese herbal medicine graphic and text modality data according to claim 1, It is characterized in that The method of extracting the image-text annotation pairs from the Chinese herbal medicine image data comprises the following steps: S11, extracting and binarizing the images in the graphic image data; S12, connecting adjacent connected domains, obtaining a sub-graph mask, and extracting the remaining text; S13, recognizing the extracted text by optical character recognition technology, and outputting the recognition result; S14. Use natural language processing technology and keyword extraction technology to perform semantic analysis and semantic splitting to obtain graphic and text results.
3. A method for labeling images based on Chinese herbal medicine graphic and text modality data according to claim 2, It is characterized in that The extraction and binarization of the pictures in the graphic image data comprises the following steps: S111, extracting an edge binary image of the image using an edge detection operator; S112, extracting the foreground and background thresholds of the image using the maximum inter-class variance algorithm, and performing binarization processing on the image; S113, superimposing the edge binary image and the foreground and background binary images to obtain a mask.
4. A method for labeling images based on Chinese herbal medicine graphic and text modality data according to claim 3, It is characterized in that The method of connecting adjacent connected domains, obtaining a subgraph mask, and extracting the remaining text includes the following steps: S121. Use graphics expansion to open up adjacent connected domains; S122, filtering the text part according to the size of the connected domain, and retaining the sub-graph part; S123, using the flooding method to fill the sub-image connected domain to obtain the sub-image mask; S124, extracting the sub-image part through the mask, and extracting the remaining text part.
5. A method for labeling images based on Chinese herbal medicine graphic and text modality data according to claim 4, It is characterized in that The method of recognizing the extracted text by optical character recognition technology and outputting the recognition result comprises the following steps: S131, select the convolutional recurrent neural network + text recognition network structure network model, and capture the forward and backward information simultaneously through the double-layer long short-term memory network structure in the recurrent neural network; S132, the text recognition network transcription layer judges the results of the model output through statistical principles, selects the result with the highest probability to output, and obtains the final recognition result.
6. A method for labeling images based on Chinese herbal medicine graphic and text modality data according to claim 5, It is characterized in that The method of using natural language processing technology keyword extraction technology to perform semantic analysis and semantic splitting to obtain graphic and text results includes the following steps: S141. Use text segmentation technology to break up the paragraph into multi-level word vectors; S142, removing stop words through a stop word list; S143, using an improved text sorting algorithm to extract core words in the text in the form of graph vectors, and then extracting text content corresponding to different keywords from the text to achieve parsing and extraction of semantic tags; S144: assign labels to the separated corresponding sub-graphs to obtain the final graphic and text results.
Citation Information
Patent Citations
Chinese speech decoding nursing system based on transfer learning
CN110807386A
Remote sensing image semantic segmentation method and device
CN111582104A