Formula identification method and device, equipment and storage medium
By introducing the target features of the feature extraction model as supervision information into the organic chemical formula recognition model, high-precision recognition of organic chemical formulas is achieved, solving the problem of low recognition accuracy in existing technologies.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GUANGZHOU SHIYUAN ELECTRONICS CO LTD
- Filing Date
- 2024-10-25
- Publication Date
- 2026-04-28
AI Technical Summary
In existing technologies, images based on organic chemical formulas are difficult to identify effectively, resulting in low identification accuracy.
The target features of organic chemical formula samples are extracted using a pre-trained feature extraction model, and these features are used as supervised information to train the formula recognition model. The recognition ability of the model is improved through self-supervised training.
It improves the accuracy of organic chemical formula recognition, enabling more accurate identification of the structural information of handwritten organic chemical formulas.
Smart Images

Figure CN121938009A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, such as a method, apparatus, device, and storage medium for formula recognition. Background Technology
[0002] Currently, in fields such as academic research, education, automated office work, and intelligent conferencing, it is often necessary to intelligently recognize handwritten formulas based on their images to facilitate user access. However, because images of organic chemical formulas differ significantly from those of mathematical and physical formulas, effective recognition is challenging.
[0003] In order to achieve intelligent recognition of handwritten organic chemical formulas, related technologies can train neural network models based on sample images of organic chemical formulas to obtain organic chemical formula recognition models. These models can then be used to recognize images of organic chemical formulas and output the corresponding character sequences.
[0004] However, due to the complex structure of organic chemical formulas, the organic chemical formula recognition model trained based on sample images of organic chemical formulas is difficult to accurately recognize handwritten organic chemical formulas, resulting in low recognition accuracy. Summary of the Invention
[0005] To provide a basic understanding of some aspects of the disclosed embodiments, a brief summary is given below. This summary is not intended as a general commentary, nor is it intended to identify key / important components or describe the scope of protection of these embodiments, but rather as a prelude to the detailed description that follows.
[0006] This application provides a method, apparatus, device, and storage medium for formula recognition, which can improve the recognition accuracy of handwritten organic chemical formulas.
[0007] In a first aspect, embodiments of this application provide a method for formula recognition, the method comprising:
[0008] Obtain organic chemical formula samples from a pre-set dataset;
[0009] The organic chemical formula samples are processed using a trained feature extraction model to obtain the target features of the organic chemical formula samples;
[0010] Using the target features of the organic chemical formula samples, a formula recognition model is trained to obtain a well-trained formula recognition model.
[0011] The trained formula recognition model is used to identify the trajectory image of the target organic chemical formula, and the recognition result of the target organic chemical formula is obtained.
[0012] Secondly, embodiments of this application provide a formula recognition device, which includes:
[0013] The sample acquisition module is used to acquire organic chemical formula samples from a preset dataset;
[0014] The feature extraction module is used to process the organic chemical formula sample using a trained feature extraction model to obtain the target features of the organic chemical formula sample.
[0015] The model training module is used to train the formula recognition model using the target features of the organic chemical formula sample to obtain a trained formula recognition model.
[0016] The formula recognition module is used to recognize the trajectory image of the target organic chemical formula using the trained formula recognition model, and to obtain the recognition result of the target organic chemical formula.
[0017] Thirdly, embodiments of this application provide an electronic device, including a processor and a memory storing program instructions, wherein the processor is configured to execute the formula recognition method as described in the first aspect when running the program instructions.
[0018] Fourthly, embodiments of this application provide a storage medium storing program instructions, wherein the program instructions, when executed, perform the formula recognition method as described in the first aspect.
[0019] The formula recognition method, apparatus, device, and storage medium provided in the embodiments of this application can achieve the following technical effects:
[0020] An electronic device acquires organic chemical formula samples from a preset dataset and processes these samples using a pre-trained feature extraction model to obtain target features. These target features are then used as supervisory information to train a formula recognition model, resulting in a trained model. The trained model is then used to identify the trajectory image of the target organic chemical formula, outputting the recognition result. In this embodiment, because the target features output by the feature extraction model are used as supervisory information during the formula recognition model training process, the model learns richer organic chemical formula information. This allows the trained model to better recognize organic chemical formulas, improving the accuracy when recognizing handwritten organic chemical formulas. Attached Figure Description
[0021] One or more embodiments are illustrated by way of example with reference to the accompanying drawings. These illustrations and drawings do not constitute a limitation on the embodiments. Elements having the same reference numerals in the drawings are shown as similar elements. The drawings are not to be scaled. And wherein:
[0022] Figure 1 This is an example diagram of organic chemical formula recognition;
[0023] Figure 2 This is another example diagram for recognizing organic chemical formulas;
[0024] Figure 3 This is a schematic diagram illustrating the environment in which the formula recognition method provided in this application is applicable;
[0025] Figure 4 This is an example diagram illustrating the conversion of organic chemical formula recognition results provided in an embodiment of this application;
[0026] Figure 5 This is a schematic diagram of a formula recognition method provided in an embodiment of this application;
[0027] Figure 6 This is a schematic diagram of a training method for a feature extraction model provided in an embodiment of this application;
[0028] Figure 7 This is an example diagram of a positive sample provided in an embodiment of this application;
[0029] Figure 8 This is an example diagram of a negative sample provided in an embodiment of this application;
[0030] Figure 9 This is a schematic diagram of another formula recognition method provided in an embodiment of this application;
[0031] Figure 10 This is a schematic diagram of a formula recognition device provided in an embodiment of this application;
[0032] Figure 11 This is a schematic diagram of another formula recognition device provided in an embodiment of this application;
[0033] Figure 12 This is a schematic diagram of an electronic device provided in an embodiment of this application. Detailed Implementation
[0034] The terms "first," "second," etc., used in the specification, claims, and drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate for the embodiments of this application described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion.
[0035] Unless otherwise stated, the term "multiple" means two or more. In the embodiments of this application, the character " / " indicates that the preceding and following objects are in an "or" relationship. For example, A / B means: A or B. The term "and / or" describes an association relationship between objects, indicating that three relationships can exist. For example, A and / or B means: A or B, or, A and B. The term "corresponding" can refer to an association relationship or binding relationship; A corresponding to B means that there is an association relationship or binding relationship between A and B.
[0036] To provide a more detailed understanding of the features and technical content of the embodiments of this application, the implementation of the embodiments of this application will be described in detail below with reference to the accompanying drawings. The accompanying drawings are for illustrative purposes only and are not intended to limit the embodiments of this application. In the following technical description, for ease of explanation, several details are used to provide a full understanding of the disclosed embodiments. However, one or more embodiments may still be implemented without these details. In other cases, well-known structures and devices may be simplified in their depiction to simplify the drawings.
[0037] Currently, in fields such as academic research, education, automated office work, and intelligent conferencing, it is often necessary to intelligently recognize handwritten formulas based on their images to facilitate user access. However, because images of organic chemical formulas differ significantly from those of mathematical and physical formulas, effective recognition is challenging.
[0038] To enable intelligent recognition of handwritten organic chemical formulas, related technologies can input sample images of organic chemical formulas into a neural network model with an encoder-decoder architecture for training, thereby obtaining an organic chemical formula recognition model. This model can then be used to recognize images of organic chemical formulas and output the corresponding character sequences.
[0039] However, in related technologies, when training organic chemical formula recognition models, only sample images of organic chemical formulas are input into the neural network model for training. Because the structures of organic chemical formulas are relatively complex, relying solely on sample images to train the neural network model makes it difficult to learn the structural information of the organic chemical formula. This results in organic chemical formula recognition models trained on neural networks having poor ability to identify the structure of organic chemical formulas, or even lacking the ability to identify the structure altogether, thus failing to accurately identify organic chemical formulas. The following will combine... Figure 1 and Figure 2 As shown, an example is provided illustrating the situation where organic chemical formula recognition models in related technologies cannot accurately identify organic chemical formulas.
[0040] For example, the organic chemical formula sequence "C1CCC(C)CC1Br" corresponds to the structure under normal circumstances. Figure 11 In related technologies, because organic chemical formula recognition models cannot accurately identify the structure of organic chemical formulas, they are prone to misinterpreting the structure. Figure 11 Incorrectly identified as structure Figure 12 However, the structure Figure 12 This corresponds to the organic chemical formula sequence "C1CCC(C)C(Br)C1". Therefore, the branching position of the organic chemical formula was incorrectly identified. The organic chemical formula sequence "C1CCCCC1" normally corresponds to structure diagram 21. However, in related technologies, because the organic chemical formula recognition model cannot accurately identify the structure of the organic chemical formula, it is easy to incorrectly identify structure diagram 21 as structure diagram 22. However, structure diagram 22 corresponds to the organic chemical formula sequence "C1CCCC1". Therefore, the number of atoms in the organic chemical formula was incorrectly identified.
[0041] As can be seen from the above, the organic chemical formula recognition model in the relevant technologies has poor accuracy in recognizing organic chemical formulas.
[0042] In view of this, embodiments of this application provide a method, apparatus, device, and storage medium for formula recognition. This method utilizes a trained feature extraction model to pre-extract target features from organic chemical formula samples and uses these target features as supervisory information to monitor the training process of the formula recognition model. Thus, by adjusting the training process of the formula recognition model, its ability to recognize organic chemical formulas can be improved, thereby enhancing the accuracy of organic chemical formula recognition.
[0043] Combination Figure 3 As shown, this application embodiment provides an environment in which the formula recognition method is applicable, the environment including a first electronic device 31, a second electronic device 32, a database 33, and a display device 34. Wherein:
[0044] The first electronic device 31 is the executing entity of the formula recognition scheme. It can recognize organic chemical formulas through a deployed feature extraction model and formula recognition model. The first electronic device 31 can be a laptop, desktop computer, mobile phone, tablet computer, conference tablet, or smart interactive tablet. It should be noted that when a user recognizes a handwritten organic chemical formula through the first electronic device 31, if the first electronic device 31 is a device integrating touch and display (such as a mobile phone, tablet computer, conference tablet, or smart interactive tablet), the user can directly write the organic chemical formula on the first electronic device 31, allowing the first electronic device 31 to directly recognize the image of the written organic chemical formula and store or display the recognition result. If the first electronic device 31 is a laptop, desktop computer, or other device without touch functionality, the user can write the organic chemical formula on a second electronic device 32 with touch functionality and save the image of the written organic chemical formula through the second electronic device 32. The second electronic device 32 then sends the image of the organic chemical formula to the first electronic device 31, which recognizes the image of the organic chemical formula and stores or displays the recognition result.
[0045] When storing the identification results of organic chemical formulas, the first electronic device 31 can store the identification results locally or upload them to the database 33 for storage. It is understood that storing the identification results of organic chemical formulas involves converting them into a string format (e.g., SMILES format). SMILES can be considered a syntax rule that converts organic chemical formulas into one-dimensional ASCII strings according to preset rules. Figure 4 For example, if the organic chemical formula is "cyclohexane", the recognition result in the SMILES format is "C1CCCCC1".
[0046] When displaying the recognition results of organic chemical formulas, the first electronic device 31 can directly display the recognition results on the screen, or it can send the recognition results to a display device 34 equipped with a screen for display. It is also understood that when displaying the recognition results of organic chemical formulas, the corresponding character sequence is not directly displayed; instead, the handwriting of the organic chemical formula is restored based on the recognition results and displayed.
[0047] Combination Figure 5 As shown, this application provides a formula recognition method, which can be applied to the first electronic device in the above embodiments. The method includes:
[0048] S51, Obtain organic chemical formula samples from the preset dataset.
[0049] S52, use the trained feature extraction model to process organic chemical formula samples to obtain the target features of organic chemical formula samples.
[0050] S53 utilizes the target features of organic chemical formula samples to train a formula recognition model to obtain a well-trained formula recognition model.
[0051] S54. Use the trained formula recognition model to identify the trajectory image of the target organic chemical formula and obtain the recognition result of the target organic chemical formula.
[0052] Using the formula recognition method provided in this application, a first electronic device can acquire organic chemical formula samples from a preset dataset and process the samples using a pre-trained feature extraction model to obtain target features of the organic chemical formula samples. The target features of the organic chemical formula samples are used as supervisory information to train the formula recognition model, resulting in a trained formula recognition model. The trained formula recognition model is then used to recognize the trajectory image of the target organic chemical formula, outputting the recognition result of the target organic chemical formula. In this application embodiment, during the training process of the formula recognition model, since the target features output by the feature extraction model are used as supervisory information for training, the formula recognition model can learn richer organic chemical formula information during training. Therefore, the trained formula recognition model can better recognize organic chemical formulas, thereby improving the accuracy when recognizing handwritten organic chemical formulas using the trained formula recognition model.
[0053] In step S51 above, an organic chemical formula dataset is pre-constructed to facilitate model training. When organic chemical formula samples are needed to train the formula recognition model, they can be directly obtained from the pre-constructed dataset.
[0054] In steps S52 and S53 above, during the training of the formula recognition model, the first electronic device uses a pre-trained feature extraction model to identify organic chemical formula samples, obtain the target features of the organic chemical formula samples, and use the target features of the organic chemical formula samples as supervisory information. Through this supervisory information, the training process of the formula recognition model is supervised, achieving self-supervised training of the formula recognition model. Since the target features are specifically extracted from organic chemical formula samples using a pre-trained feature extraction model, the target features more accurately reflect the attributes or characteristics of organic chemical formulas.
[0055] The target characteristics of organic chemical formulas will be explained below.
[0056] The target features of organic chemical formula samples include at least the chemical features of atoms and the topological features between atoms. The chemical features of atoms can be multi-dimensional, such as 64-dimensional features, as shown in Table 1 below:
[0057]
[0058] Table 1
[0059] In organic chemical formulas, chemical bonds between atoms can be classified into single bonds, double bonds, triple bonds, hybrid bonds, virtual ring bonds, and other types of bonds. The topological characteristics of atoms are used to characterize which of these chemical bonds connects the atoms. For example, the topological characteristics of atoms can be represented using an adjacency matrix.
[0060] Optionally, the formula recognition model is trained using the target features of the organic chemical formula samples to obtain a trained formula recognition model. This includes: inputting the organic chemical formula samples into the formula recognition model for training, so that the formula recognition model outputs the predicted features of the organic chemical formula samples; calculating the learning loss of the formula recognition model based on the predicted features and target features of the organic chemical formula samples; adjusting the parameters of the formula recognition model according to the learning loss of the formula recognition model, and retraining the formula recognition model; and completing the training of the formula recognition model when the learning loss of the formula recognition model meets a first preset condition, thus obtaining a trained formula recognition model.
[0061] In this implementation, the learning loss of the formula recognition model is used to characterize the difference between the predicted features output by the formula recognition model when recognizing organic chemical formula samples and the target features of the organic chemical formula samples during the model training process. The purpose of calculating the learning loss of the formula recognition model is to adjust the parameters of the model based on the difference between the predicted and target features during training, thereby guiding the training process and enabling the trained model to better recognize organic chemical formulas. Specifically, the learning loss of the formula recognition model can be represented by KL divergence (Kullback-Leibler divergence). Based on the target features of the organic chemical formula, the probability distribution of the target features is calculated as the true distribution. Based on the predicted features of the organic chemical formula samples, the probability distribution of the predicted features is calculated as the predicted distribution. By calculating the difference between the predicted distribution and the true distribution, the KL divergence of the formula recognition model is obtained, which is the learning loss of the formula recognition model. After obtaining the KL divergence, the parameters of the formula recognition model are adjusted based on the KL divergence. In this way, using organic chemical formula samples, the formula recognition model with adjusted parameters is retrained. The formula recognition model outputs predicted features of the organic chemical formula samples again, and based on the predicted features and target features of the organic chemical formulas, the KL divergence is calculated again. This process is repeated to train the formula recognition model until the learning loss of the formula recognition model meets the first preset condition. Meeting the first preset condition means that the KL divergence value of the formula recognition model decreases to a preset value. The smaller the KL divergence value, the higher the similarity between the distribution of the predicted features output by the formula recognition model and the distribution of the target features; that is, the predicted distribution gets closer to the true distribution. When the KL divergence value of the formula recognition model decreases to the preset value, the predicted distribution is considered to be consistent with the true distribution, indicating that the formula recognition model training is complete, and a well-trained formula recognition model is obtained.
[0062] This implementation method utilizes the probability distribution of target features as the true distribution to supervise the training of the formula classification model. This allows the model to learn detailed features of organic chemical formulas (chemical characteristics of atoms and topological features between atoms) during training. Consequently, the trained formula recognition model can better identify organic chemical formulas, thus improving its recognition performance.
[0063] The following will explain the training process of the feature extraction model used in the above formula recognition method. The feature extraction model can be a graph convolutional neural network model.
[0064] Combination Figure 6As shown, this application provides a method for training a feature extraction model. This method can be applied to the first electronic device described above, and the trained feature extraction model is applied to the formula recognition method described above. The training method includes:
[0065] S61, construct positive and negative samples based on the preset organic chemical formula sequence.
[0066] S62, input positive and negative samples into the feature extraction model for training to obtain the feature distribution of positive samples and the feature distribution of negative samples.
[0067] S63, based on the feature distributions of positive and negative samples, calculates the contrastive learning loss of the feature extraction model.
[0068] S64. Based on the contrastive learning loss of the feature extraction model, adjust the parameters of the feature extraction model and retrain the feature extraction model.
[0069] S65, when the contrastive learning loss of the feature extraction model meets the second preset condition, the feature extraction model training is completed, and a trained feature extraction model is obtained.
[0070] In this embodiment, contrastive learning is used to train the feature extraction model. Contrastive learning is an unsupervised learning method that improves model performance by learning the similarity and differences between samples. Positive and negative samples guide the model's training process. Specifically, positive samples refer to samples with high similarity in the feature space, while negative samples refer to samples with low similarity. The purpose of training the feature extraction model using contrastive learning is to enable it to distinguish between positive and negative samples, i.e., to bring positive samples closer together and push negative samples further apart, so that similar samples are closer together in the feature space, while dissimilar samples are farther apart.
[0071] In this embodiment, the contrastive learning loss of the feature extraction model is used to quantify the similarity between samples, thereby facilitating the adjustment of the parameters of the feature extraction model based on the contrastive learning loss of the feature extraction model during the training process of the feature extraction model.
[0072] Therefore, the training method of the feature extraction model provided in this application, which trains the feature extraction model through comparative learning, enables the feature extraction model to extract detailed features of organic chemical formula samples. These detailed features participate in the training process of the formula recognition model, allowing the formula recognition model to learn the detailed features of organic chemical formulas during training. This improves the recognition ability of the trained formula recognition model when identifying organic chemical formulas.
[0073] In step S61 above, constructing positive samples based on a preset organic chemical formula sequence includes: obtaining a first organic chemical formula sequence from the preset organic chemical formula sequence; restoring the first organic chemical formula sequence to a structural diagram of the first organic chemical formula; performing enhancement processing on the structural diagram of the first organic chemical formula to construct a structural diagram of a second organic chemical formula similar to the structure of the first organic chemical formula; and using the structural diagrams of the first and second organic chemical formulas as positive samples.
[0074] In this embodiment, the preset organic chemical formula sequence can be a character sequence stored in SMILES format. The preset organic chemical formula sequence may include one or more organic chemical formula sequences. The first electronic device can reconstruct the first organic chemical formula sequence into a structural diagram of the first organic chemical formula. Referring to the structure of the first organic chemical formula, the structural diagram of the first organic chemical formula is enhanced. This enhancement process can involve adjusting the structure of the first organic chemical formula to obtain a second organic chemical formula, thereby obtaining a structural diagram of the second organic chemical formula. Since the second organic chemical formula is generated with reference to the first organic chemical formula, the structures of the first and second organic chemical formulas are similar, making the structural diagrams of the first and second organic chemical formulas similar as well. After constructing a large number of second organic chemical formulas based on the first organic chemical formula, the structural diagrams of the first and second organic chemical formulas are used as positive samples.
[0075] In step S61 above, constructing negative samples based on a preset organic chemical formula sequence includes: obtaining a third organic chemical formula sequence from the preset organic chemical formula sequence; restoring the third organic chemical formula sequence to a structural diagram of the third organic chemical formula; enhancing the structural diagram of the third organic chemical formula to construct a structural diagram of a fourth organic chemical formula similar to the structure of the third organic chemical formula; and using the structural diagrams of the third and fourth organic chemical formulas as negative samples.
[0076] In this embodiment, negative samples refer to samples with a significant difference in similarity to positive samples. When constructing negative samples, the first electronic device obtains a third organic chemical formula sequence that differs from the first organic chemical formula sequence from a preset organic chemical formula sequence, and restores the third organic chemical formula sequence to its structural diagram. Referring to the structure of the third organic chemical formula, the structural diagram of the third organic chemical formula is enhanced to obtain a structural diagram of a fourth organic chemical formula similar to the third organic chemical formula. Based on the third organic chemical formula, a large number of fourth organic chemical formulas are constructed, and the structural diagrams of the third and fourth organic chemical formulas are used as negative samples.
[0077] Understandably, positive samples are constructed based on the first organic chemical formula sequence, while negative samples are constructed based on the third organic chemical formula sequence. Since the first organic chemical formula sequence differs from the third organic chemical formula sequence, the structural diagram of the first organic chemical formula obtained by reconstructing the first organic chemical formula sequence will show significant differences compared to the structural diagram of the third organic chemical formula obtained by reconstructing the third organic chemical formula sequence. Therefore, the similarity between positive and negative samples can be guaranteed to be low.
[0078] In step S61 above, a large number of positive and negative samples can be obtained through enhancement processing. With a sufficient number of samples, training the feature extraction model can avoid overfitting or underfitting, thereby improving the generalization ability of the feature extraction model.
[0079] In step S61 above, when constructing a positive sample, the structure diagram of the first organic chemical formula is enhanced to construct a structure diagram of the second organic chemical formula that is similar to the structure of the first organic chemical formula. This includes: based on the structure diagram of the first organic chemical formula, deleting a preset number of nodes from the first organic chemical formula to form a structure diagram of the second organic chemical formula.
[0080] In this embodiment, when enhancing the structure diagram of the first organic chemical formula, a predetermined number of nodes can be deleted from the first organic chemical formula, referring to its structure, to obtain the second organic chemical formula, and then a structure diagram of the second organic chemical formula can be constructed. Here, the nodes in the first organic chemical formula refer to the atoms within it.
[0081] It should be noted that deleting a large number of nodes from the first organic chemical formula can easily lead to a lower similarity between the first organic chemical formula after node deletion and the first organic chemical formula without deleted nodes. Therefore, by limiting the number of nodes that can be deleted from the first organic chemical formula, a high similarity between the first and second organic chemical formulas can be ensured, that is, a high similarity between sample data in positive samples can be guaranteed. The preset number of nodes that can be deleted is determined based on the total number of nodes in the first organic chemical formula. For example, in the first organic chemical formula, the preset number should be less than 1 / 4 of the total number of nodes.
[0082] It should also be noted that in the process of constructing the second organic chemical formula based on the first organic chemical formula, the corresponding second organic chemical formula can be obtained by deleting one node from the first organic chemical formula. Alternatively, the corresponding second organic chemical formula can be obtained by deleting multiple nodes from the first organic chemical formula. The specific configuration can be set according to requirements, and this application does not limit this.
[0083] Below, we will combine Figure 7The process of enhancing the first organic chemical formula is illustrated in the figure.
[0084] For example, based on the structural diagram 71 of the first organic chemical formula, after deleting nodes 72 and 73, a structural diagram 74 of the second organic chemical formula is obtained. And, based on the structural diagram 71 of the first organic chemical formula, after deleting node 75, another structural diagram 76 of the second organic chemical formula is obtained.
[0085] Optionally, in step S61 above, when constructing a negative sample, the structure diagram of the third organic chemical formula is enhanced to construct a structure diagram of a fourth organic chemical formula with a structure similar to the third organic chemical formula. This includes: based on the structure diagram of the third organic chemical formula, deleting a predetermined number of nodes from the fourth organic chemical formula to form a structure diagram of the fourth organic chemical formula. For example, in a positive sample as shown... Figure 7 When the structure diagram is shown, the molecular formula of the negative sample can be as follows: Figure 8 As shown.
[0086] In this embodiment, when enhancing the structural diagram of the third organic chemical formula, the process of enhancing the structural diagram of the first organic chemical formula described above can be referred to, and the implementation method is similar, so it will not be repeated here.
[0087] In this embodiment, by deleting organic chemical formula nodes, another organic chemical formula similar to the original organic chemical formula can be quickly constructed. This improves the efficiency of constructing positive and negative samples, thereby enhancing the training efficiency of the feature extraction model.
[0088] It should be noted that, in this embodiment, the structural diagram of an organic chemical formula refers to its data structure graph, typically represented as G(V, E), where V represents atoms in the organic chemical formula and E represents the bonds between atoms. The structural diagram of an organic chemical formula can reflect its structural information. Thus, by training the feature extraction model using positive and negative samples constructed based on the organic chemical formula's structural diagram, the model can output the structural diagram of the organic chemical formula. Embedding the structural diagram of the organic chemical formula into the training process of the formula recognition model enables the model to identify the structural information of the organic chemical formula when recognizing its trajectory image, thereby improving the accuracy of organic chemical formula recognition.
[0089] In steps S62 to S65 above, after constructing positive and negative samples, the positive and negative samples can be input into the feature extraction model for training, so that the feature extraction model outputs feature information of positive samples and feature information of negative samples respectively. Based on the feature information of positive and negative samples, the feature distributions of positive and negative samples can be calculated respectively. Based on the feature distributions of positive and negative samples, the contrastive learning loss of the feature extraction model is calculated. The contrastive learning loss is used to characterize the correlation between the feature distributions of positive and negative samples, and can be represented by mutual information loss. Thus, according to the contrastive learning loss, the parameters of the feature extraction model are adjusted, and the feature extraction model with adjusted parameters is retrained, so that the feature extraction model outputs feature information of positive and negative samples again. Therefore, the contrastive learning loss of the feature extraction model is calculated by outputting the feature information of positive and negative samples again. This process of repeatedly training the feature extraction model continues until the contrastive learning loss of the feature extraction model satisfies the second preset condition. Satisfying the second preset condition for the contrastive learning loss of the feature extraction model can be maximizing the contrastive learning loss between the feature distributions of positive and negative samples. The large contrastive learning loss between the feature distributions of positive and negative samples indicates that the feature extraction model can distinguish between positive and negative samples and classify them into different classes. Therefore, if the contrastive learning loss meets the second preset condition, the feature extraction model can be considered successfully trained, and a well-trained feature extraction model can be obtained.
[0090] Combination Figure 9 As shown in the embodiments of this application, another method for formula recognition is provided. This method can be applied to the first electronic device described in the above embodiments. The method includes:
[0091] S91, obtain the trajectory image of the target organic chemical formula.
[0092] S92 uses a trained formula recognition model to identify the trajectory image of the target organic chemical formula in order to obtain the recognition result of the target organic chemical formula.
[0093] The trained formula recognition model is obtained by training the model using the target features of organic chemical formula samples. The target features of the organic chemical formula samples are obtained by processing organic chemical formula samples obtained from a pre-set dataset using a trained feature extraction model.
[0094] The formula recognition method provided in this application is explained from the perspective of recognizing target organic chemical formulas using a trained formula recognition model. For details, please refer to the above-mentioned examples of training and application of the formula recognition model, which will not be repeated here.
[0095] Combination Figure 10 As shown, this application provides a formula recognition device that can be integrated into the first electronic device described above. The device includes a sample acquisition module 101, a feature extraction module 102, a model training module 103, and a formula recognition module 104.
[0096] The sample acquisition module 101 is used to acquire organic chemical formula samples from a preset dataset.
[0097] The feature extraction module 102 is used to process organic chemical formula samples using a trained feature extraction model to obtain the target features of the organic chemical formula samples.
[0098] The model training module 103 is used to train the formula recognition model using the target features of organic chemical formula samples to obtain a trained formula recognition model.
[0099] Formula recognition module 104 is used to recognize the trajectory image of the target organic chemical formula using a trained formula recognition model, and obtain the recognition result of the target organic chemical formula.
[0100] Optionally, the model training module 103 is specifically used to input organic chemical formula samples into the formula recognition model for training, so that the formula recognition model outputs the predicted features of the organic chemical formula samples. Based on the predicted features and target features of the organic chemical formula samples, the learning loss of the formula recognition model is calculated. According to the learning loss of the formula recognition model, the parameters of the formula recognition model are adjusted, and the formula recognition model is trained again. When the learning loss of the formula recognition model meets the first preset condition, the training of the formula recognition model is completed, and the trained formula recognition model is obtained.
[0101] Optionally, the model training module 103 is also used to train the feature extraction model, specifically to construct positive and negative samples based on preset organic chemical formula sequences. The positive and negative samples are input into the feature extraction model for training to obtain the feature distributions of the positive and negative samples. Based on the feature distributions of the positive and negative samples, the contrastive learning loss of the feature extraction model is calculated. The parameters of the feature extraction model are adjusted according to the contrastive learning loss, and the model is trained again. When the contrastive learning loss of the feature extraction model meets a second preset condition, the feature extraction model training is complete, and a trained feature extraction model is obtained.
[0102] Optionally, the model training module 103 is specifically used to obtain a first organic chemical formula sequence from a preset organic chemical formula sequence. The first organic chemical formula sequence is then reconstructed into a structural diagram of the first organic chemical formula. The structural diagram of the first organic chemical formula is then enhanced to construct a structural diagram of a second organic chemical formula that is structurally similar to the first organic chemical formula. The structural diagrams of the first and second organic chemical formulas are used as positive samples.
[0103] Optionally, the model training module 103 is specifically used to obtain a third organic chemical formula sequence from a preset organic chemical formula sequence. The third organic chemical formula sequence is then restored to its structural diagram. The structural diagram of the third organic chemical formula is enhanced to construct a structural diagram of a fourth organic chemical formula that is structurally similar to the third organic chemical formula. The structural diagrams of the third and fourth organic chemical formulas are used as negative samples.
[0104] Optionally, the model training module 103 is specifically used to delete a preset number of nodes from the structure diagram of the first organic chemical formula to form the structure diagram of the second organic chemical formula.
[0105] The formula recognition device provided in this application embodiment can execute the formula recognition method in the above embodiment. Its implementation principle and technical effect are similar, and will not be described again here.
[0106] Combination Figure 11 As shown, this application embodiment provides another formula recognition device, which can be integrated into the first electronic device of the above embodiment. The device includes: an acquisition module 111 and a recognition module 112.
[0107] The acquisition module 111 is used to acquire the trajectory image of the target organic chemical formula.
[0108] The recognition module 112 is used to recognize the trajectory image of the target organic chemical formula using a trained formula recognition model, so as to obtain the recognition result of the target organic chemical formula.
[0109] The trained formula recognition model is obtained by training the model using the target features of organic chemical formula samples. The target features of the organic chemical formula samples are obtained by processing organic chemical formula samples obtained from a pre-set dataset using a trained feature extraction model.
[0110] The formula recognition device provided in this application embodiment can execute the formula recognition method in the above embodiment. Its implementation principle and technical effect are similar, and will not be described again here.
[0111] Combination Figure 12As shown, this application embodiment provides an electronic device 120, including a processor 121 and a memory 122. Optionally, the electronic device 120 may further include a communication interface 123 and a bus 124. The processor 121, communication interface 123, and memory 122 can communicate with each other via the bus 124. The communication interface 123 can be used for information transmission. The processor 121 can call logical instructions in the memory 122 to execute the formula recognition method described in the above embodiment.
[0112] Furthermore, the logic instructions in the aforementioned memory 122 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium.
[0113] The memory 122, as a computer-readable storage medium, can be used to store software programs and computer-executable programs, such as program instructions / modules corresponding to the methods in the embodiments of this application. The processor 121 executes functional applications and data processing by running the program instructions / modules stored in the memory 122, that is, it implements the formula recognition method in the above embodiments.
[0114] The memory 122 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the terminal device. Furthermore, the memory 122 may include high-speed random access memory and may also include non-volatile memory.
[0115] This application provides a storage medium storing computer-executable instructions configured to perform the formula recognition method described in the above embodiments.
[0116] The aforementioned storage medium can be a transient computer-readable storage medium or a non-transitory computer-readable storage medium.
[0117] The technical solutions of this application embodiment can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes one or more instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in this application embodiment. The aforementioned storage medium can be a non-transitory storage medium, including: USB flash drive, portable hard drive, read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk, and other media capable of storing program code; it can also be a transient storage medium.
[0118] The foregoing description and accompanying drawings fully illustrate embodiments of this disclosure to enable those skilled in the art to practice them. Other embodiments may include structural, logical, electrical, procedural, and other changes. The embodiments represent only possible variations. Individual components and functions are optional unless explicitly required, and the order of operation may vary. Parts and features of some embodiments may be included in or replace parts and features of other embodiments. Moreover, the terminology used in this application is for describing embodiments only and is not intended to limit the claims. As used in the description of embodiments and claims, the singular forms “a,” “an,” and “the” are intended to equally include the plural forms unless the context clearly indicates otherwise. Similarly, the term “and / or” as used in this application means including one or more of the associated listed items and all possible combinations thereof. Additionally, when used in this application, the term "comprise" and its variations "comprises" and / or "comprising" refer to the presence of stated features, integrals, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or groups thereof. Without further limitations, an element defined by the phrase "comprises a..." does not exclude the presence of other identical elements in the process, method, or apparatus that includes said element. In this document, each embodiment may focus on the differences from other embodiments, and similar or identical parts between embodiments can be referred to mutually. For methods, products, etc., disclosed in the embodiments, if they correspond to the method section disclosed in the embodiments, the relevant parts can be referred to the description of the method section.
[0119] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the embodiments of this application. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0120] The methods and products (including but not limited to devices and equipment) disclosed in the embodiments herein can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For instance, the division of units may be merely a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces, and the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to implement this embodiment according to actual needs. In addition, the functional units in the embodiments of this application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0121] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than that shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. In the descriptions corresponding to the flowcharts and block diagrams in the accompanying drawings, the operations or steps corresponding to different blocks may also occur in a different order than disclosed in the description; sometimes there is no specific order between different operations or steps. For example, two consecutive operations or steps may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. Each block in a block diagram and / or flowchart, and combinations of blocks in a block diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
Claims
1. A method for formula recognition, characterized in that, include: Obtain organic chemical formula samples from a pre-set dataset; The organic chemical formula samples are processed using a trained feature extraction model to obtain the target features of the organic chemical formula samples; Using the target features of the organic chemical formula samples, a formula recognition model is trained to obtain a well-trained formula recognition model. The trained formula recognition model is used to identify the trajectory image of the target organic chemical formula, and the recognition result of the target organic chemical formula is obtained.
2. The method according to claim 1, characterized in that, Using the target features of the organic chemical formula samples, a formula recognition model is trained to obtain a trained formula recognition model, including: The organic chemical formula sample is input into the formula recognition model for training, so that the formula recognition model outputs the predicted features of the organic chemical formula sample; Based on the predicted features of the organic chemical formula samples and the target features of the organic chemical formula samples, the learning loss of the formula recognition model is calculated. Based on the learning loss of the formula recognition model, adjust the parameters of the formula recognition model and retrain the formula recognition model; When the learning loss of the formula recognition model meets the first preset condition, the training of the formula recognition model is completed, and the trained formula recognition model is obtained.
3. The method according to claim 1, characterized in that, Training the feature extraction model includes: Based on the preset organic chemical formula sequence, construct positive and negative samples; The positive samples and the negative samples are input into the feature extraction model for training to obtain the feature distribution of the positive samples and the feature distribution of the negative samples; Based on the feature distributions of the positive samples and the feature distributions of the negative samples, the contrastive learning loss of the feature extraction model is calculated; Based on the contrastive learning loss of the feature extraction model, adjust the parameters of the feature extraction model and retrain the feature extraction model; When the contrastive learning loss of the feature extraction model meets the second preset condition, the feature extraction model is trained and the trained feature extraction model is obtained.
4. The method according to claim 3, characterized in that, Based on a pre-defined organic chemical formula sequence, construct positive samples, including: Obtain the first organic chemical formula sequence from the preset organic chemical formula sequence; The first organic chemical formula sequence is restored to the structural diagram of the first organic chemical formula; The structural diagram of the first organic chemical formula is enhanced to construct a structural diagram of a second organic chemical formula that is similar to the structure of the first organic chemical formula; The structural diagrams of the first organic chemical formula and the second organic chemical formula are used as positive samples.
5. The method according to claim 3, characterized in that, Based on a pre-defined organic chemical formula sequence, negative samples are constructed, including: Obtain a third organic chemical formula sequence from the preset organic chemical formula sequence; The sequence of the third organic chemical formula is reduced to the structural diagram of the third organic chemical formula; The structural diagram of the third organic chemical formula is enhanced to construct a structural diagram of a fourth organic chemical formula that is similar to the structure of the third organic chemical formula; The structural diagrams of the third organic chemical formula and the fourth organic chemical formula are used as negative samples.
6. The method according to claim 4, characterized in that, Enhancement processing is performed on the structural diagram of the first organic chemical formula to construct a structural diagram of a second organic chemical formula that is similar in structure to the first organic chemical formula, including: Based on the structural diagram of the first organic chemical formula, a predetermined number of nodes are deleted from the first organic chemical formula to form the structural diagram of the second organic chemical formula.
7. A formula recognition device, characterized in that, include: The sample acquisition module is used to acquire organic chemical formula samples from a preset dataset; The feature extraction module is used to process the organic chemical formula sample using a trained feature extraction model to obtain the target features of the organic chemical formula sample. The model training module is used to train the formula recognition model using the target features of the organic chemical formula sample to obtain a trained formula recognition model. The formula recognition module is used to recognize the trajectory image of the target organic chemical formula using the trained formula recognition model, and to obtain the recognition result of the target organic chemical formula.
8. An electronic device comprising a processor and a memory storing program instructions, characterized in that, The processor is configured to perform the formula recognition method as described in any one of claims 1 to 6 when executing the program instructions.
9. A storage medium storing program instructions, characterized in that, When the program instructions are executed, they perform the formula recognition method as described in any one of claims 1 to 6.