Polymer structural formula identification method and system, electronic equipment and storage medium

Through the visual large language model, the polymer chemical structure image is aligned with the description text to generate SMILES strings, which solves the problems of low efficiency and poor accuracy of identifying polymer structural formulas in the prior art, and achieves efficient and accurate polymer structural formula recognition.

CN120047941APending Publication Date: 2025-05-27CHANGCHUN INSTITUTE OF APPLIED CHEMISTRY CHINESE ACADEMY OF SCIENCES +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510182929.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-19
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently and accurately identify polymer structural formulas from images, and existing optical character recognition and rule-based image analysis methods have limited ability to process complex chemical bonds and structural symbols.

Method used

The Visual Large Language Model (VLLM) is used, which includes a visual encoder, a text encoder, an alignment module and a fully connected neural network. Through iterative training, the polymer chemical structure image is aligned with the description text to generate a SMILES string of polymer structural formula.

Benefits of technology

It realizes efficient and accurate identification of polymer structural formulas from images, overcomes the problems of low recognition efficiency and poor accuracy in the prior art, and can automatically identify complex polymer structures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047941A_ABST
    Figure CN120047941A_ABST
Patent Text Reader

Abstract

The invention discloses a polymer structural formula identification method and system, electronic equipment and a storage medium, and belongs to the technical field of image analysis technologies. The polymer structural formula identification method comprises the following steps: acquiring a training sample from a data set; inputting the training sample into a visual large language model, and updating model parameters of the visual large language model by using a loss function so as to iteratively train the visual large language model; wherein the visual large language model comprises a visual encoder, a text encoder, an alignment module and a full-connection neural network; the trained visual large language model is deployed; and if a recognition task is received, determining an unknown polymer chemical structure image corresponding to the recognition task, and generating an SMILES character string of a polymer structural formula of the unknown polymer chemical structure image by using the visual big language model. According to the invention, the structural formula of the polymer can be efficiently and accurately identified from the image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image analysis technology, and particularly to a method, system, electronic device and storage medium for identifying polymer structural formulas. Background Art

[0002] Polymer materials play a crucial role in modern industry, technology and social development. From daily necessities to aerospace materials, the application scope of polymer materials is extremely wide. The development of new materials, the performance optimization of existing materials, and the failure analysis of materials highly depend on the accurate understanding and effective utilization of polymer material data. These data include chemical composition, molecular structure, physical properties, processing conditions, etc. Among them, molecular structure information, especially the precise structural formula of polymers, is the key to understanding and predicting material properties.

[0003] In this field, a large amount of polymer material information exists in the form of images. For example, in scientific research papers, patent documents, textbooks, and experimental records, chemical structural formulas often appear in two-dimensional or three-dimensional image forms. It is difficult to automatically extract polymer structure information in image form. Currently, the extraction and analysis of polymer structural formulas mainly rely on manual operations, which are time-consuming, laborious, inefficient, and prone to introducing human errors. In addition, when consulting literature materials to extract structural formula information, problems such as scattered information, inconsistent formats, and uneven drawing quality are also faced, seriously hindering scientific research efficiency.

[0004] Existing optical character recognition (OCR) technologies and rule-based image analysis methods are helpless in dealing with complex chemical bonds and structure symbols. Even if the chemical name or molecular formula can be recognized, the structural information of polymers cannot be completely restored. Usually, a large number of manually designed features and rules are required, and the generalization ability is poor, making it difficult to process various images in the real world. Although deep learning has made remarkable progress in the fields of object detection and image recognition, current models applied to chemical structure recognition usually only focus on single molecules or simple structures, lacking the ability to understand and model the complex chain structures of polymers.

[0005] Therefore, how to efficiently and accurately identify polymer structural formulas from images is a technical problem that needs to be solved by those skilled in the art at present. Summary of the Invention

[0006] The purpose of this application is to provide a method, system, electronic device and storage medium for identifying polymer structural formulas, which can efficiently and accurately identify polymer structural formulas from images.

[0007] To solve the above technical problems, this application provides a method for identifying polymer structural formulas, including:

[0008] Obtain training samples from the dataset; wherein, the training samples include polymer chemical structure images and descriptive texts, and the descriptive texts include SMILES strings corresponding to polymer structural formulas, chemical names, and structural feature information;

[0009] Input the training samples into the vision-language model and update the model parameters of the vision-language model using a loss function to iteratively train the vision-language model; wherein, the vision-language model includes a vision encoder, a text encoder, an alignment module, and a fully connected neural network. The vision encoder is used to extract image features from polymer chemical structure images, the text encoder is used to extract text features from descriptive texts, the alignment module is used to align the image features with the text features, and the fully connected neural network is used to predict the SMILES string of the polymer structural formula based on the aligned image features and text features;

[0010] Deploy the trained vision-language model;

[0011] If a recognition task is received, determine the unknown polymer chemical structure image corresponding to the recognition task, and use the vision-language model to generate the SMILES string of the polymer structural formula of the unknown polymer chemical structure image.

[0012] Optionally, using the vision-language model to generate the SMILES string of the polymer structural formula of the unknown polymer chemical structure image includes:

[0013] Input the unknown polymer chemical structure image into the vision encoder of the vision-language model to obtain the current image features;

[0014] Use the text encoder of the vision-language model to process the initial text to obtain the current text features; wherein, the initial text includes prompt words and / or relevant polymer information;

[0015] Use the alignment module to align the current image features with the current text features;

[0016] Concatenate or fuse the aligned current image features and the current text features to obtain text-image features;

[0017] Input the text-image features into the fully connected neural network of the vision-language model to obtain probability distribution information; wherein, the probability distribution information is used to describe the probabilities of each word in the SMILES vocabulary;

[0018] Use a decoding algorithm to transform the probability distribution information to obtain the SMILES string of the polymer structural formula corresponding to the unknown polymer chemical structure image.

[0019] Optionally, updating the model parameters of the visual large language model using a loss function includes:

[0020] Calculating the contrastive learning loss value by using a contrastive learning loss function for the aligned image features and the aligned text features;

[0021] Calculating the cross-entropy loss value by using a cross-entropy loss function for the aligned text features and the description text;

[0022] Performing weighted calculation on the contrastive learning loss value and the cross-entropy loss value to obtain a total loss value;

[0023] Updating the model parameters of the visual large language model according to the total loss value.

[0024] Optionally, during the iterative training of the visual large language model, it further includes:

[0025] Adjusting the learning rate of model training by using a cosine annealing learning rate scheduler;

[0026] Adjusting the weight values of the contrastive learning loss value and the cross-entropy loss value through grid search or random search.

[0027] Optionally, updating the model parameters of the visual large language model according to the total loss value includes:

[0028] Updating the model parameters of the visual large language model through an optimizer according to the total loss value; wherein, the optimizer is an optimizer combined with weight decay regularization.

[0029] Optionally, before obtaining training samples from the dataset, it further includes:

[0030] Performing data augmentation on the dataset through a preset operation;

[0031] Wherein, the preset operation includes:

[0032] Rendering the polymer chemical structure images in the dataset according to multiple parameters;

[0033] And / or, randomly cropping and randomly scaling the polymer chemical structure images in the dataset;

[0034] And / or, randomly adjusting the color channels of the polymer chemical structure images in the dataset;

[0035] And / or, adding Gaussian noise or salt-and-pepper noise to the polymer chemical structure images in the dataset.

[0036] Optionally, deploying the trained vision-language model includes:

[0037] Deploying the trained vision-language model on a local device and / or a cloud server.

[0038] This application also provides a polymer structural formula recognition system, including:

[0039] A sample acquisition module for acquiring training samples from a dataset; wherein, the training samples include polymer chemical structure images and description texts, and the description texts include SMILES strings corresponding to the polymer structural formulas, chemical names, and structural feature information.

[0040] A model training module for inputting the training samples into a vision-language model and updating the model parameters of the vision-language model using a loss function to iteratively train the vision-language model; wherein, the vision-language model includes a vision encoder, a text encoder, an alignment module, and a fully connected neural network. The vision encoder is used to extract image features from the polymer chemical structure images, the text encoder is used to extract text features from the description texts, the alignment module is used to align the image features with the text features, and the fully connected neural network is used to predict the SMILES string of the polymer structural formula based on the aligned image features and text features.

[0041] A model deployment module for deploying the trained vision-language model.

[0042] A model invocation module for, if a recognition task is received, determining the unknown polymer chemical structure image corresponding to the recognition task and generating the SMILES string of the polymer structural formula of the unknown polymer chemical structure image using the vision-language model.

[0043] This application also provides a storage medium with a computer program stored thereon, and when the computer program is executed, it implements the steps performed by the above polymer structural formula recognition method.

[0044] This application also provides an electronic device, including a memory and a processor. When the processor calls the computer program stored in the memory, it implements the steps performed by the above polymer structural formula recognition method.

[0045] This application realizes the recognition of polymer structural formulas based on a vision-language model architecture. The above vision-language model architecture includes a vision encoder, a text encoder, an alignment module, and a fully connected neural network. During the process of training the model, this application obtains training samples containing polymer chemical structure images and descriptive texts from the dataset, extracts the visual features of the polymer chemical structure images through the vision encoder, and simultaneously extracts the semantic features of the descriptive texts using the text encoder. The alignment module aligns the image features with the text features to ensure their consistency in the semantic space. The aligned image features and text features are fused through a fully connected neural network, and then the SMILES string of the polymer structural formula is predicted. After the model training is completed, this application deploys the vision-language model and directly calls the vision-language model to generate the SMILES string of the polymer structural formula for an unknown polymer chemical structure image when receiving a recognition task. This application uses multi-modal data for training, which enhances the understanding ability of the vision-language model for polymer structures. The feature alignment mechanism improves the robustness and accuracy of the vision-language model. This application can use the deployed vision-language model to automatically recognize the polymer structural formula of an unknown polymer chemical structure image without manual participation. Therefore, this application can efficiently and accurately recognize polymer structural formulas from images. This application also provides a polymer structural formula recognition system, a storage medium, and an electronic device, which have the above beneficial effects and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] To more clearly illustrate the embodiments of this application, the following will briefly introduce the drawings required for the embodiments. Obviously, the drawings in the following description are only some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0047] Figure 1 It is a flowchart of a method for recognizing polymer structural formulas provided by an embodiment of this application;

[0048] Figure 2 It is a schematic diagram of a polymer chemical structure provided by an embodiment of this application;

[0049] Figure 3 It is a flowchart of another method for recognizing polymer structural formulas provided by an embodiment of this application;

[0050] Figure 4 It is a schematic diagram of the recognition result of a polymer homopolymer structural formula provided by an embodiment of this application;

[0051] Figure 5 It is a schematic diagram of the recognition result of a block copolymer structural formula provided by an embodiment of this application;

[0052] Figure 6 This is a schematic structural diagram of a polymer structural formula recognition system provided by an embodiment of the present application. Detailed implementation manners

[0053] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0054] Obtaining the polymer structural formula is the key to understanding and predicting the properties of materials. Although some software can help draw or edit chemical structures, currently, it is impossible to directly extract polymer structural formula information from images, and users must manually redraw it. This manual operation is not only time-consuming and laborious but also prone to introducing errors, especially for complex polymer structures, such as copolymers containing multiple monomer units or polymers with special topological structures.

[0055] Currently, there is an urgent need for an automated and efficient polymer structural formula extraction solution to solve the above problems and accelerate the research and development and application of polymer materials. The polymer structural formula recognition solution based on the vision large language model proposed by the present invention can fill this gap and break through the bottleneck of polymer material data extraction, especially the challenge of obtaining polymer structural formula information.

[0056] Please refer to the following Figure 1 , Figure 1 This is a flowchart of a polymer structural formula recognition method provided by an embodiment of the present application.

[0057] The specific steps may include:

[0058] S101: Obtain training samples from the dataset;

[0059] Among them, the training samples include polymer chemical structure images and description texts, and the description texts include SMILES (Simplified Molecular Input Line Entry System) strings corresponding to the polymer structural formula, chemical names, and structural feature information.

[0060] The above dataset includes multiple training samples, and each training sample includes a polymer chemical structure image and its corresponding descriptive text. The above polymer chemical structure diagram can be synthetic data. In this embodiment, synthetic data is used to train and fine-tune the Visual Large Language Model (VLLM). These synthetic data contain polymers of various types and topologies. Please refer to Figure 2 , Figure 2 which is a schematic diagram of a polymer chemical structure provided by an embodiment of the present application. The polymer chemical structures of random copolymers, block copolymers, alternating copolymers, and branched copolymers are shown in the figure.

[0061] The synthetic data covers a variety of common polymer structure types, including homopolymers, copolymers (such as random copolymers, alternating copolymers, block copolymers), and branched polymers with different degrees and ways of branching, etc. These diverse structures ensure that the trained visual large language model can handle various complex polymer structures.

[0062] The polymer dataset includes virtual creations and polymer molecules that have been verified by experiments or simulations. The representation of the polymer structural formula can be SELFIES (Self-Referencing Embedded Strings) or BigSMILES (a polymer characterization method).

[0063] To increase the diversity of the data and improve the robustness of the model, in this embodiment, Indigo (an open-source chemical structure drawing tool) can be used to generate rendered images of the same chemical structure in different styles. Indigo provides various rendering options, such as different line thicknesses, colors, backgrounds, and different perspectives, so that the visual large language model can better adapt to various images in the real world.

[0064] A significant feature of the polymer structural formula is that the molecular monomers are wrapped in a pair of brackets, and there may also be cases with different subscripts. Therefore, in this embodiment, the SMILES strings corresponding to the polymer structural formula can use multiple different types of brackets, such as the following three: "[]", "{}", and "()". To further enrich the diversity of the polymer structural formula, in this embodiment, the size, color, and thickness of the brackets are varied. The subscripts of the SMILES strings corresponding to the polymer structural formula can be represented by common single letters, such as m, n, x, y, z. In addition, a large number of subscripts in the polymer structural formula use numbers. Therefore, the subscripts of the SMILES strings corresponding to the polymer structural formula in this embodiment also include different numbers (integers 0 - 100, decimals 0.0 - 1.0).

[0065] S102: Input the training samples into the vision-language model, and update the model parameters of the vision-language model using a loss function to iteratively train the vision-language model;

[0066] Among them, the vision-language model includes a vision encoder, a text encoder, an alignment module, and a fully connected neural network. The vision encoder is used to extract image features from polymer chemical structure images, the text encoder is used to extract text features from descriptive texts, the alignment module is used to align the image features with the text features, and the fully connected neural network is used to predict the SMILES string of the polymer structural formula based on the aligned image features and text features. Specifically, the fully connected neural network can fuse and further process the aligned image features and text features to generate a feature representation for prediction. The fully connected neural network can fuse the aligned image features and text features to generate a feature representation for prediction, and then the decoder generates the SMILES string of the polymer structural formula based on these feature representations.

[0067] In this embodiment, the training samples of the current batch can be selected and input into the vision-language model. The vision-language model processes the input data through its various modules (such as the vision encoder, text encoder, alignment module, and fully connected neural network, etc.), extracts features, and generates prediction results. In this embodiment, a loss function can be used to calculate the difference between the prediction result and the true label, and the model parameters are updated by backpropagation. By repeating the above process, the model performance can be gradually optimized.

[0068] In this embodiment, the iterative training can be stopped after the number of iterations reaches or the loss function meets the requirements.

[0069] S103: Deploy the trained vision-language model;

[0070] Among them, after the vision-language model is trained, the vision-language model can be saved and deployed to the actual application environment. After deployment, the vision-language model can receive the input of new polymer chemical structure images and quickly generate the corresponding SMILES string to achieve automatic recognition. Specifically, in this embodiment, the trained vision-language model can be deployed on a local device and / or a cloud server.

[0071] S104: If a recognition task is received, determine the unknown polymer chemical structure image corresponding to the recognition task, and use the vision-language model to generate the SMILES string of the polymer structural formula of the unknown polymer chemical structure image.

[0072] Among them, when receiving the recognition task of polymer structural formula, this embodiment can identify and determine the unknown polymer chemical structure image corresponding to the task, so as to use the trained vision-language large model to process the unknown polymer chemical structure image. The vision-language large model extracts image features using a vision encoder, and text features extracted by a text encoder. After being aligned by an alignment module, the image features and text features are input into a fully connected neural network to generate the SMILES string of the polymer structural formula. The above process realizes the efficient conversion from image to text description and can automatically recognize chemical structures.

[0073] This embodiment realizes the recognition of polymer structural formula based on the vision-language large model architecture. The above vision-language large model architecture includes a vision encoder, a text encoder, an alignment module, and a fully connected neural network. During the process of training the model, this embodiment obtains training samples containing polymer chemical structure images and descriptive texts from the dataset, extracts the visual features of the polymer chemical structure images through the vision encoder, and at the same time extracts the semantic features of the descriptive texts using the text encoder. The alignment module aligns the image features with the text features to ensure their consistency in the semantic space. The aligned image features and text features are fused through a fully connected neural network, and then the SMILES string of the polymer structural formula is predicted. After the model training is completed, this embodiment deploys the vision-language large model and directly calls the vision-language large model to generate the SMILES string of the polymer structural formula of the unknown polymer chemical structure image when receiving the recognition task. This embodiment uses multi-modal data for training, which enhances the vision-language large model's understanding ability of polymer structures, and the feature alignment mechanism improves the robustness and accuracy of the vision-language large model. This embodiment can use the deployed vision-language large model to automatically recognize the polymer structural formula of the unknown polymer chemical structure image without manual participation. Therefore, this embodiment can efficiently and accurately recognize the polymer structural formula from the image.

[0074] It can be seen that this application can overcome the limitations of the prior art in recognizing polymer chemical structures from images. By introducing an innovative system and method based on the vision-language large model (VLLM), it realizes the automatic, accurate, and efficient extraction of polymer structural formulas. This application utilizes the powerful visual perception ability and SMILES understanding ability of VLLM to directly parse the chemical structure information of polymers from images and output it in a structured form. This system can process polymer structure images of different qualities and styles, including real pictures and synthetic data.

[0075] As a feasible implementation manner, before obtaining training samples from the dataset, the dataset can also be enhanced through preset operations;

[0076] Among them, the preset operations include:

[0077] Rendering the polymer chemical structure images in the dataset according to multiple parameters;

[0078] And / or, randomly cropping and randomly scaling the polymer chemical structure images in the dataset;

[0079] And / or, randomly adjusting the color channels of the polymer chemical structure images in the dataset;

[0080] And / or, adding Gaussian noise or salt-and-pepper noise to the polymer chemical structure images in the dataset.

[0081] The above visual encoder can adopt a pre-trained VitDet (Vision Transformer Detector) model with 8M parameters. As a polymer structural formula image feature extractor, VitDet can effectively capture local and global features in the image, providing rich visual information for the recognition of polymer structural formulas.

[0082] The SMILES encoder can adopt the Qwen 2.5 0.5B model as the text encoder. Qwen 2.5 is a pre-trained language model with text understanding ability, which is responsible for processing text information related to chemical structures, such as chemical names, molecular formulas, and descriptions of polymer topological structures.

[0083] In order to align the polymer structural formula image features extracted by VitDet with the text features generated by Qwen 2.5, this embodiment adopts a fully connected layer as the alignment module. The alignment module is responsible for learning the mapping relationship between polymer structural formula image features and text features, projecting them into the same semantic space, so that the vision-language model can effectively understand the chemical structure information in the image.

[0084] The above visual encoder can also adopt ResNet (Residual Network), EfficientNet (efficient convolutional neural network), the SMILES encoder can also adopt other large language models, and the feature alignment module can also adopt an attention mechanism or other non-linear mapping methods.

[0085] The introduction to the visual encoder VitDet is as follows:

[0086] VitDet is an object detection model based on the Vision Transformer (ViT) architecture, which combines the global modeling ability of ViT and the specific requirements of the object detection task. The core idea of VitDet is to divide the input image into a series of patches (image blocks), and then input these patches into the Transformer encoder for feature extraction. The polymer structural formula image is first divided into patches of a fixed size of 16x16 pixels. Each patch is flattened into a vector and then mapped to the embedding space through a linear transformation to form patch embeddings.

[0087] Let the polymer structural formula image , where H, W, and C represent the height, width, and number of channels of the image respectively. The size of each patch is , then the number of patches is . Each patch can be represented as , P i represents the i-th image block patch after being segmented. After flattening the patch, we get . The embedding of the patch is , where , , D represents the embedding dimension, b represents the bias term, represents the domain.

[0088] Since the Transformer lacks the perception of positional information, positional encoding needs to be added to the patch embeddings. In this embodiment, sine and cosine functions are used to encode the positions of the patches. Let the position coordinates of the patch in the image be (x, y), then the positional encoding PE can be expressed as:

[0089] ;

[0090] ;

[0091] In the above formula, pos represents the position, i represents the embedding dimension index, and d represents the size of the embedding dimension.

[0092] In this embodiment, the sum of the patch embedding and the positional encoding can be input into the Transformer encoder. The Transformer encoder is stacked by multiple Transformer layers, and each layer contains self-attention, allowing the embedding of each patch to interact with the embeddings of other patches, so as to learn global context information. Given the input X, the self-attention module generates queries (Q), keys (K), and values (V) through three linear transformations.

[0093] ;

[0094] ;

[0095] ;

[0096] W Q 、W K 、W V represent the weight matrices of query (Q), key (K), and value (V) respectively.

[0097] The self-attention mechanism is calculated as follows:

[0098] ;

[0099] denotes attention, denotes the function that converts a vector into a probability distribution, denotes the matrix multiplication of the transpose of the query matrix and the key matrix, denotes the dimensions of Q and K.

[0100] The multi-head self-attention adopted in this embodiment is an extension of the above self-attention. It calculates multiple self-attentions in parallel and concatenates the results together.

[0101] ;

[0102] Among them, , is the output mapping matrix.

[0103] denotes multi-head self-attention, denotes concatenation, denotes the head, denotes the query weight matrix of the i-th head, denotes the key weight matrix of the i-th head, denotes the value weight matrix of the i-th head.

[0104] Feed-Forward Network (FFN). Each layer contains a two-layer multi-layer perceptron, which acts on each position of the self-attention output to increase the non-linearity of the model.

[0105] ;

[0106] represents the activation function, represents the input vector, represents the weight matrix of the first layer, represents the bias vector of the first layer, represents the weight matrix of the second layer, represents the bias vector of the second layer.

[0107] Feature Pyramid Network (FPN): VitDet usually uses FPN to construct multi-scale feature maps to better detect objects of different sizes. FPN fuses feature maps of different layers through top-down and lateral connections, thereby generating feature maps with multi-scale information.

[0108] In this embodiment, the total number of parameters of the visual encoder VitDet is 8M, and it can also run under limited computing resources. In order to utilize its polymer structural formula image feature extraction ability in this embodiment, the output feature map is used.

[0109] The introduction of the SMILES encoder is as follows:

[0110] This embodiment uses Qwen 2.5 as the encoder for SMILES strings. Based on the Transformer architecture, it extracts features of chemical structures through powerful text understanding and generation capabilities. First, the input SMILES string is converted into word embeddings, that is, each word or sub-word is mapped to a high-dimensional vector space, and its embedding matrix is represented as , where is the vocabulary size and D is the embedding dimension. To capture the sequential features of words in the sequence, the model also adds positional encoding, similar to the method in ViT, to make it have the ability to perceive the SMILES sequence order. Subsequently, the model uses the Transformer decoder to encode the text features, including the self-attention mechanism (capturing internal dependencies in the sequence) and the feed-forward network (enhancing the non-linear expression ability of the model), but does not involve the cross-attention mechanism in the decoding stage. Finally, the model outputs the encoded text features, providing an efficient feature representation for subsequent chemical information processing and analysis.

[0111] The number of parameters of the SMILES encoder used in this embodiment is 0.5B, which belongs to a medium-scale model. While ensuring performance, it reduces the computational overhead. Qwen 2.5 has been pre-trained on a large-scale text dataset, including text data in various fields, enabling it to possess extensive knowledge and language capabilities. Qwen 2.5 typically outputs the embedded representation of text and the probability distribution of text generation. In this embodiment, its text embedding is used as text features.

[0112] In this embodiment, other methods can be used to extract text features, such as using multiple layers in the SMILES encoder. Different decoding strategies can be used in this embodiment, such as beam search and nucleus sampling. Post-processing tools (such as the open-source toolkit RDKit for cheminformatics and the chemical structure type file format conversion software OpenBabel) can also be used in this embodiment to standardize the chemical structures. The recognition results in this embodiment can be saved in various formats, such as JSON, CSV, Excel, etc. The results can be provided to other systems through the application programming interface API in this embodiment.

[0113] The introduction of the alignment module is as follows:

[0114] The alignment module uses a fully connected layer to project the polymer structural formula image features output by VitDet and the text features output by Qwen 2.5 into the same semantic space.

[0115] Let the polymer structural formula image feature be , and the text feature be . Through the fully connected layer: , . Among them , , , . In the formula, the polymer structural formula image feature and the text feature are respectively mapped to the same feature space dimension D3 through the fully connected layer for alignment comparison or fusion. The image feature transformation is represented as , where is the weight matrix that maps the image feature from dimension D1 to D3, is the bias vector; the text feature transformation is represented as , where is the weight matrix that maps the text feature from dimension D2 to D3, is the bias vector.

[0116] Based on the above visual large language model architecture, synthetic data is used for model training during training. This data includes polymer chemical structure images with different topological structures (such as homopolymers, copolymers, branched polymers, etc.), as well as corresponding SMILES descriptions (such as the SMILES strings corresponding to the structural formulas, chemical names, and text descriptions of structural features).

[0117] It can be seen that this embodiment provides a polymer chemical structure recognition framework based on a visual large language model (VLLM). This embodiment utilizes an innovative framework of a visual large language model (VLLM) to directly recognize polymer chemical structures from images. This framework combines the capabilities of a polymer structural formula image encoder and a SMILES encoder, and maps polymer structural formula images and SMILES features to the same semantic space through an alignment module, thereby achieving the understanding and extraction of polymer structural formulas in images. The visual large language model architecture realizes its application in polymer structural formula recognition by virtue of its ability to directly analyze chemical structures using visual information.

[0118] To improve the robustness of the model, the following data augmentation strategies can be adopted:

[0119] Indigo-style rendering: Use Indigo to render images of the same structure in different styles, including different line thicknesses, colors, backgrounds, and different perspectives;

[0120] Random cropping / zooming: Randomly crop and zoom the images to simulate different shooting angles and distances;

[0121] Color jittering: Randomly adjust the color channels of the images to simulate different lighting conditions;

[0122] Noise addition: Add Gaussian noise or salt-and-pepper noise to simulate the situation of poor image quality.

[0123] This embodiment uses a contrastive learning loss function combined with a cross-entropy loss function for model training to achieve the alignment of polymer structure images and text features and the accurate prediction of structural information.

[0124] It can be seen that this embodiment provides a training scheme based on synthetic data and data augmentation. Among them, the synthetic data covers various common polymer structure types and topological structures, uses tools such as Indigo to generate structure images in different styles, and data augmentation includes various means such as random cropping, zooming, color jittering, noise addition, and random rotation / translation to improve the robustness and generalization ability of the model. The idea of using synthetic data (covering various polymer topological structures) rather than relying solely on real annotated data to train the model, as well as specific data augmentation strategies, including using Indigo to generate diverse rendered images and various image enhancement methods, to improve the robustness of the model.

[0125] As a feasible implementation, this embodiment can update the model parameters of the vision-language model in the following way:

[0126] Step A1: Calculate the contrastive learning loss value by using the contrastive learning loss function for the aligned image features and the aligned text features;

[0127] Step A2: Calculate the cross-entropy loss value by using the cross-entropy loss function for the aligned text features and the description text;

[0128] Step A3: Perform weighted calculation on the contrastive learning loss value and the cross-entropy loss value to obtain the total loss value;

[0129] Step A4: Update the model parameters of the vision-language model according to the total loss value.

[0130] Specifically, this embodiment can update the model parameters of the vision-language model through an optimizer according to the total loss value; wherein, the optimizer is an optimizer combined with weight decay regularization.

[0131] Furthermore, during the iterative training of the vision-language model, the learning rate of model training can also be adjusted by using a cosine annealing learning rate scheduler; the weight values of the contrastive learning loss value and the cross-entropy loss value are adjusted through grid search or random search.

[0132] In order to enable the vision-language model to learn the mapping relationship between the polymer structural formula picture features and the text features, this embodiment uses the contrastive learning loss function InfoNCE (Information Noise Contrastive Estimation) to calculate the loss. The goal of the contrastive learning loss InfoNCE is to maximize the similarity between the polymer structural formula image features and the text features corresponding to the same polymer structure, while minimizing the similarity between the features corresponding to different structures.

[0133] InfoNCE loss can be expressed as:

[0134] ;

[0135] where and are the aligned polymer structure images and text features of the i-th sample respectively. sim(·,·) is a similarity function, such as cosine similarity. is the temperature coefficient, which is used to control the sharpness of the similarity distribution, and exp represents the exponential function. j represents the index for traversing all samples within a batch.

[0136] For the given polymer structural formula picture features and the corresponding text feature t, the contrastive learning loss function can be expressed as:

[0137] ;

[0138] where is the dot product of feature vectors (which can be understood as a simplified version of cosine similarity), represents the text features of other samples within the batch except t.

[0139] To ensure that the model can accurately predict chemical structure information, in this embodiment, a cross-entropy loss is added at the SMILES output end. Specifically, in this embodiment, the text feature t' can be input into a linear layer to output a probability distribution corresponding to the SMILES vocabulary, and then the cross-entropy loss between the predicted probability and the true SMILES label is calculated.

[0140] The cross-entropy loss is calculated as:

[0141] ;

[0142] where is the true label (one-hot vector) of the i-th word. is the predicted probability distribution of the model. The above true label is the descriptive text in the training sample.

[0143] By minimizing the cross-entropy, the model will tend to generate more accurate SMILES outputs to better describe the chemical structure.

[0144] The total loss function is the weighted sum of the contrastive learning loss and the cross-entropy loss:

[0145] ;

[0146] where and are weight coefficients used to balance the effects of the two loss functions.

[0147] It can be seen that this embodiment provides a training strategy that combines contrastive learning and cross-entropy loss. Specifically, this embodiment adopts a training strategy that combines contrastive learning and cross-entropy loss to optimize the performance of the VLLM. The contrastive learning loss aims to make the visual and SMILES features of the same polymer structure similar, while the cross-entropy loss ensures that the model accurately predicts the SMILES description corresponding to the structural formula.

[0148] By means of the combination method of the loss function and the setting method of the weight coefficients of the contrast loss and the cross-entropy loss ( and ), it is used to optimize VLLM, achieve the alignment of visual and SMILES features, and accurately predict chemical structure information.

[0149] In this embodiment, the loss function can also adopt other contrastive learning losses such as SupCon, the optimizer can adopt Adam or RMSprop, mixed-precision training can be used to accelerate model training, gradient accumulation can be used to increase the batch size, and knowledge distillation can be used to improve the model performance.

[0150] In this embodiment, the AdamW optimizer is used to update the model parameters. AdamW is an improved version of the Adam optimizer, which adds weight decay regularization to help prevent model overfitting.

[0151] The parameter update method of AdamW is as follows:

[0152] ;

[0153] ;

[0154] ;

[0155] ;

[0156] ;

[0157] where is the exponential moving average of the gradient, is the exponential moving average of the squared gradient. and are the exponential decay rates. is the current gradient. is the learning rate. is the weight decay coefficient. is a small constant to avoid division by zero error. represents the bias correction of the exponential moving average of the gradient, represents the bias correction of the exponential moving average of the squared gradient, represents the model parameters at the t-th moment.

[0158] In this embodiment, the hyperparameters can be adjusted in the following ways:

[0159] Learning rate, an initial learning rate and a learning rate decay strategy are adopted. In the present invention, a cosine annealing learning rate scheduler is used to gradually reduce the learning rate as the training progresses. Batch size, the choice of batch size depends on the computing power of the hardware, and the present invention adopts 16. Weight decay coefficient, the weight decay coefficient is used to control the strength of L2 regularization, and is usually set to 0.01 or a smaller value. Temperature coefficient, the temperature coefficient ( ) is used to control the sharpness of the similarity distribution in the contrastive learning loss. It is usually adjusted between 0.01 and 0.1. and The weight coefficient used to balance the contrastive learning loss and the cross-entropy loss can be adjusted by grid search or random search. Number of training epochs, the number of training epochs is adjusted according to the performance of the model until the performance of the model on the validation set reaches the optimum. Early stopping, when the performance on the validation set no longer improves, the training is terminated early to prevent overfitting of the model.

[0160] The training process of the vision-language model is as follows:

[0161] Step B1, Data loading: Load the training data set, which includes polymer structural formula pictures and corresponding SMILES.

[0162] Step B2, Data preprocessing: Perform data augmentation on the images and tokenize the SMILES.

[0163] Step B3, Feature extraction: Input the polymer structural formula image into VitDet to obtain the polymer structural formula image features, and input the SMILES into Qwen 2.5 to obtain the text features.

[0164] Step B4, Feature alignment: Project the polymer structural formula image features and the text features into the same space through a fully connected layer.

[0165] Step B5, Calculate loss: Calculate the contrastive learning loss and the cross-entropy loss.

[0166] Step B6, Backpropagation: Backpropagate the loss into the model to update the model parameters.

[0167] Step B7, Model evaluation: Regularly evaluate the model performance on the validation set and adjust the hyperparameters. Model saving, save the trained model weights.

[0168] The instructions for deploying the vision-language model are as follows

[0169] The operating system used in this embodiment is Ubuntu 20.04 because it has better compatibility and performance for deep learning frameworks. The Python version is Python 3.11, the PyTorch version is 2.20, and the corresponding CUDA drivers and libraries. Other software dependencies are as follows: numpy for numerical calculations; opencv-python for image processing; Pillow for image reading and saving; scikit-learn for model evaluation; matplotlib for result visualization; indigo-python for rendering using indigo.

[0170] In this embodiment, the vision-language model can be deployed in the following ways: (1) Local deployment, where the model is deployed on a local machine for single-machine or small-scale testing and use; (2) Cloud deployment, where the model is deployed to a cloud server for large-scale data processing and online services, and a Python web framework such as Flask or FastAPI is used to build an API service.

[0171] As a feasible implementation, the process of using the vision-language model to generate the SMILES string of the polymer structural formula of the unknown polymer chemical structure image in this embodiment includes:

[0172] Step C1: Input the unknown polymer chemical structure image into the vision encoder of the vision-language model to obtain the current image features;

[0173] Step C2: Use the text encoder of the vision-language model to process the initial text to obtain the current text features;

[0174] Among them, the above initial text can be an empty text, or the initial text can also include prompt words and / or relevant polymer information. The prompt word can be "Extract the SMILES string of the polymer structural formula in the figure".

[0175] Step C3: Use the alignment module to align the current image features with the current text features;

[0176] Step C4: Concatenate or fuse the aligned current image features and the current text features to obtain text-image features;

[0177] Step C5: Input the text-image features into the fully connected neural network of the vision-language model to obtain probability distribution information; where the probability distribution information is used to describe the probabilities of each word in the SMILES vocabulary.

[0178] Step D6: Use a decoding algorithm to convert the probability distribution information to obtain the SMILES string of the polymer structural formula corresponding to the unknown polymer chemical structure image.

[0179] The inference process of the vision-language model is as follows: Input the polymer structural formula image: Receive the polymer structure image to be analyzed. Data preprocessing: Preprocess the image, including image resizing and normalization. Feature extraction: Input the preprocessed image into the VitDet model to obtain the polymer structural formula image feature map. Use the Qwen 2.5 model to generate the embedding representation of the SMILES. Feature alignment: Use a fully connected layer to map the polymer structure image features and text features to the same semantic space. Structure prediction: Input the aligned features into a linear classifier to predict the probability distribution of the SMILES vocabulary. Use greedy search or beam search for SMILES decoding to generate the polymer structural formula description.

[0180] The pseudocode of the inference process is as follows:

[0181] import torch

[0182] from transformers import AutoModel, AutoTokenizer

[0183] import numpy as np

[0184] import cv2

[0185] from PIL import Image

[0186] # Load the model and tokenizer

[0187] vitdet_model = AutoModel.from_pretrained("pretrained_vitdet_model") # Replace with the pre-trained VitDet model

[0188] vitdet_model.eval() # Set to inference mode

[0189] qwen_model = AutoModel.from_pretrained("pretrained_qwen_model") # Replace with the pre-trained Qwen model

[0190] qwen_model.eval() # Set to inference mode

[0191] qwen_tokenizer = AutoTokenizer.from_pretrained("pretrained_qwen_model") # Replace with the pre-trained Qwen tokenizer

[0192] def preprocess_image(image_path):

[0193] image = Image.open(image_path)

[0194] image = image.resize((224, 224)) # Resize the image

[0195] image = np.array(image) / 255.0 # Normalize the image

[0196] image = image.transpose((2, 0, 1)) # Change the channel order

[0197] image = torch.tensor(image, dtype=torch.float32).unsqueeze(0) # Convert to tensor and add batch dimension

[0198] return image

[0199] def extract_text_features(text):

[0200] inputs = qwen_tokenizer(text, return_tensors='pt')

[0201] with torch.no_grad():

[0202] outputs = qwen_model( inputs)

[0203] return outputs.last_hidden_state.mean(dim=1) # Get the text features after average pooling

[0204] def extract_visual_features(image_path):

[0205] image = preprocess_image(image_path)

[0206] with torch.no_grad():

[0207] outputs = vitdet_model(image)

[0208] return outputs.last_hidden_state.mean(dim=(2, 3)) # Take the visual features after average pooling

[0209] def align_features(visual_feature, text_feature):

[0210] # Align visual and text features with a fully connected layer

[0211] aligned_visual_feature = visual_feature @ Wv + bv # Wv and bv are the weights and biases of the fully connected layer

[0212] aligned_text_feature = text_feature @ Wt + bt # Wt and bt are the weights and biases of the fully connected layer

[0213] return aligned_visual_feature, aligned_text_feature

[0214] def predict_structure(aligned_visual_feature, aligned_text_feature):

[0215] # The linear classifier outputs a probability distribution

[0216] combined_feature = torch.cat((aligned_visual_feature, aligned_text_feature), dim=1)

[0217] probability = combined_feature @ Wc + bc # Wc, bc are the weights and biases of the linear classification layer

[0218] predicted_id = torch.argmax(probability, -1)

[0219] predicted_text = qwen_tokenizer.decode(predicted_id) # Convert the ID to text

[0220] return predicted_text

[0221] # Inference process

[0222] image_path = "path_to_image.jpg" # Replace with the image path

[0223] text_description = "This is the description of the polymer structure"# Replace with the initial text description (can also be an empty string)

[0224] visual_feature = extract_visual_features(image_path)

[0225] text_feature = extract_text_features(text_description)

[0226] aligned_visual_feature, aligned_text_feature = align_features(visual_feature, text_feature)

[0227] predicted_text = predict_structure(aligned_visual_feature, aligned_text_feature)

[0228] print("Predicted polymer structure: ", predicted_text).

[0229] It can be seen that this embodiment provides a specific deployment and inference process for polymer structural formula recognition, including hardware configuration, software dependencies, model deployment methods, data preprocessing, feature extraction, feature alignment, structure prediction, result postprocessing, and result output, etc., and provides corresponding code examples and preferred implementation methods, emphasizing the actual application ability of the model.

[0230] The inference and application of VLLM in the present invention are divided into the following steps:

[0231] Step D1: Input image reception;

[0232] The VLLM system receives an image containing a polymer structural formula. The image can be a digital image in various formats (such as JPG, PNG, TIFF), which can be a scanned literature figure, an experimental record, or a hand-drawn sketch. The system supports single-image or batch-image input.

[0233] Step D2: Image preprocessing;

[0234] Adjust the image to a predefined size, 1024 × 1024 pixels, to meet the input requirements of the VitDet model. Normalize the pixel values of the image to the range of [0, 1] to accelerate the training and inference of the model. Convert the image into a Pytorch data type that the model can accept.

[0235] Step D3: Polymer structural formula image feature extraction;

[0236] Input the preprocessed polymer structural formula image into the VitDet visual encoder to extract multi-layer polymer structural formula image feature maps. Use the last layer feature map of VitDet as the polymer structural formula image feature representation.

[0237] Step D4: Text feature extraction;

[0238] In the initial stage of VLLM training, there may be no clear requirements for the input SMILES. However, for better results, an initial text description can be provided to the SMILES encoder. This initial text can be empty or can also contain some information about the relevant polymer, such as the name of the monomer, the category of the polymer, etc. Tokenize the SMILES description and convert it into token ids. Input the token ids into the Qwen 2.5 SMILES encoder to obtain the embedding representation of the SMILES. The average pooling vector of the last layer hidden state output by Qwen 2.5 can be used as the text feature.

[0239] Step D5: Alignment of image features and text features;

[0240] Use a fully connected layer or other feature alignment modules to project the polymer structural formula image features and text features into the same semantic space to obtain the aligned feature representation.

[0241] Step D6: Structure prediction;

[0242] Concatenate the aligned polymer structural formula picture features and text features, or fuse them using methods such as weighted summation. Input the fused features into a linear classifier or a small neural network to predict the probabilities of each word in the SMILES vocabulary. Use a decoding algorithm (such as greedy decoding, beam search decoding) to convert the probability distribution into a SMILES sequence, thereby obtaining the SMILES description of the polymer structural formula.

[0243] Step D7: Post-processing of results;

[0244] Perform verification of the chemical structure, such as using RDKit to check the correctness of the SMILES string.

[0245] Step D8: Output of results;

[0246] Output the final polymer structure recognition result, including: SMILES description (e.g., polymer name, SMILES string, etc.).

[0247] This embodiment also provides a system for identifying polymer chemical structures from images, including: an image input module for receiving input images; a vision-language large model (VLLM) configured to process the input images and identify chemical structures; and a structure output module configured to output a description of the identified chemical structure. The VLLM is trained on a multi-modal dataset containing chemical structure images and corresponding SMILES descriptions. The structure output module is also configured to convert the identified chemical structure into a standardized chemical format SMILES string. The above system further includes a preprocessing module configured to improve the quality of the input images.

[0248] This embodiment also provides a method for identifying polymer chemical structures from images, including the following steps: receiving an input polymer structural formula image; using a vision-language large model (VLLM) to process the input image to identify the chemical structure; generating a SMILES description of the identified chemical structure; and outputting the SMILES and description. This embodiment can also convert the SMILES description into a standardized chemical format. The above embodiment can also include an image preprocessing step. The above VLLM is trained to be able to identify different polymer chemical structures, including monomers, chain configurations, and specific functional groups. The above embodiment also includes a post-processing module for looking up the identified structure in a database.

[0249] The above embodiments provide an innovative polymer structural formula recognition solution based on a vision-large language model, which can automatically, efficiently, and accurately extract polymer structure information from images, solving the problems of time-consuming, laborious, and error-prone traditional methods. The above embodiments can process images of different qualities and styles, have strong robustness and generalization capabilities, and are applicable to polymer structural formula recognition in various scenarios. The above embodiments utilize the powerful SMILES understanding ability of large language models, which can better understand and describe the complex structures of polymers and convert them into standardized SMILES strings for subsequent calculations and analyses. The above embodiments provide an innovative polymer natural language-based representation method, which can comprehensively describe the structures and properties of polymers and contribute to efficient polymer material design and discovery. The polymer natural language representation method can be conveniently applied end-to-end to high-throughput material screening and simulation experiments to accelerate the R & D process of polymer materials.

[0250] The following uses examples in practical applications to illustrate the processes described in the above embodiments.

[0251] Please refer to Figure 3 , Figure 3 which is a flowchart of another polymer structural formula recognition method provided by the embodiments of the present application. The process includes: polymer molecular structural formula picture synthesis, constructing a vision-large language model, model training, model deployment, and recognizing polymer molecular structural formulas.

[0252] This embodiment can be applied to specific polymer field scenarios, such as extracting structural formulas from scientific literature, verifying laboratory synthesis results, and copolymer recognition, improving the practical value and application scope of the polymer structural formula recognition solution.

[0253] The following uses the process of recognizing the structural formula of a polymer homopolymer in scientific literature to illustrate the above solution:

[0254] A large amount of polymer chemical structure information is contained in scientific journal papers, but this information often exists in the form of images and is very time-consuming to extract manually. This case aims to show how to automatically extract the structural formula information of polymer homopolymers from scientific literature using the VLLM system.

[0255] Scan or take a screenshot of a scientific document (such as a PDF file) containing a polymer structural formula image, and preprocess the scanned image, such as denoising, enhancing, cropping, and normalizing. Input the preprocessed image into the VLLM model, and use an image processor to convert it into a tensor form, while dynamically generating a query string containing image tokens. Subsequently, encode the query text using a tokenizer, and combine the image tensor as the model input to generate a text result through mixed-precision calculation. The VLLM system outputs the recognition result, including the SMILES description (SMILES string, polymer name) and the structural formula image, as Figure 4 shown. The SMILES representation of the polymer structural formula recognized by the VLLM model is: C1CCC2C(C1)C(=O)N(C2=O)C1CCC(CC1)OCCN(C1CCCCC1)CCOC1CCC(CC1)N1C(=O)C2C(C1=O)CC(CC2)C(C(F)(F)F)(C(F)(F)F) .y, the recognition is correct. Please refer to Figure 4 , Figure 4 which is a schematic diagram of the recognition result of a polymer homopolymer structural formula provided by an embodiment of this application. C represents carbon, O represents oxygen, N represents nitrogen, F represents fluorine, Ph represents phenyl, and y represents the subscript of the polymer structural formula.

[0256] The following uses the recognition process of block copolymers to illustrate the above solution:

[0257] A block copolymer is a polymer material formed by connecting homopolymer segments with two or more different chemical compositions through chemical bonds. Its unique structure endows it with special properties and has a wide range of applications in the field of materials science. However, the structural formula of block copolymers is often relatively complex, containing multiple different segments, and it is difficult to identify manually. This case aims to show how the VLLM system can be used to accurately identify and distinguish various types of block copolymer structural formulas.

[0258] The acquisition, input, preprocessing, and inference process of the polymer picture are the same as in Case 2. The polymer structural formula to be recognized is shown in Figure 5 , Figure 5 which is a schematic diagram of the recognition result of a block copolymer structural formula provided by an embodiment of this application. The final output result of the model is as follows:

[0259] CCC(CC1C(SC2C1SC(C2)C1=C2C(=O)N(C(=C2C(=O)N1CC(CCCC)CC) )CC(CCCC)CC)C1CC2C(S1)C(C1SC(C(C1)CC(CCCC)CC) )C1C(C2)SC(C1)CC(CCCC)CC) )CC.x.CCC(CC1CC(SC1C1CC(C(S1)C1CCC(S1) )C1=C2C(=O)N(C(=C2C(=O)C2C1CC(S2)C1SC2C(C1)C(F)C(C(C2F)F) )C1=O)CC(CC)C)C.y. C represents carbon, O represents oxygen, N represents nitrogen, S represents sulfur, F represents fluorine, and x and y are the subscripts of the polymer structural formula. It can be seen that in this embodiment, not only the SMILES representation of the block copolymer monomers is recognized, but also the ratio information of the two monomers, namely x and y, is accurately recognized.

[0260] Please refer to Figure 6 , Figure 6 which is a schematic structural diagram of a polymer structural formula recognition system provided by an embodiment of the present application, including:

[0261] A sample acquisition module 601, configured to obtain training samples from a dataset; wherein, the training samples include polymer chemical structure images and description texts, and the description texts include SMILES strings corresponding to polymer structural formulas, chemical names, and structural feature information;

[0262] A model training module 602, configured to input the training samples into a vision-language model and update the model parameters of the vision-language model using a loss function to iteratively train the vision-language model; wherein, the vision-language model includes a vision encoder, a text encoder, an alignment module, and a fully connected neural network, the vision encoder is configured to extract image features from polymer chemical structure images, the text encoder is configured to extract text features from description texts, the alignment module is configured to align the image features with the text features, and the fully connected neural network is configured to predict the SMILES string of the polymer structural formula according to the aligned image features and text features;

[0263] A model deployment module 603, configured to deploy the trained vision-language model;

[0264] A model invocation module 604, configured to, if a recognition task is received, determine the unknown polymer chemical structure image corresponding to the recognition task and generate the SMILES string of the polymer structural formula of the unknown polymer chemical structure image using the vision-language model.

[0265] This embodiment realizes the recognition of polymer structural formulas based on a vision-language model architecture. The above vision-language model architecture includes a vision encoder, a text encoder, an alignment module, and a fully connected neural network. During the process of training the model, this embodiment obtains training samples containing polymer chemical structure images and descriptive texts from a dataset, extracts visual features of the polymer chemical structure images through the vision encoder, and simultaneously extracts semantic features of the descriptive texts using the text encoder. The alignment module aligns the image features with the text features to ensure their consistency in the semantic space. The aligned image features and text features are fused through the fully connected neural network, and then the SMILES string of the polymer structural formula is predicted. After the model training is completed, this embodiment deploys the vision-language model and directly calls the vision-language model to generate the SMILES string of the polymer structural formula for an unknown polymer chemical structure image when receiving a recognition task. Using multi-modal data for training in this embodiment enhances the vision-language model's understanding ability of polymer structures, and the feature alignment mechanism improves the robustness and accuracy of the vision-language model. This embodiment can automatically recognize the polymer structural formula of an unknown polymer chemical structure image using the deployed vision-language model without manual participation. Therefore, this embodiment can efficiently and accurately recognize polymer structural formulas from images.

[0266] Further, the process by which the model call module 604 generates the SMILES string of the polymer structural formula for the unknown polymer chemical structure image using the vision-language model includes: inputting the unknown polymer chemical structure image into the vision encoder of the vision-language model to obtain the current image features; processing the initial text using the text encoder of the vision-language model to obtain the current text features; where the initial text includes prompt words and / or relevant polymer information; using the alignment module to align the current image features with the current text features; splicing or fusing the aligned current image features and current text features to obtain text-image features; inputting the text-image features into the fully connected neural network of the vision-language model to obtain probability distribution information; where the probability distribution information is used to describe the probabilities of each word in the SMILES vocabulary;

[0267] Using a decoding algorithm to transform the probability distribution information to obtain the SMILES string of the polymer structural formula corresponding to the unknown polymer chemical structure image.

[0268] Further, the process of the model training module 602 updating the model parameters of the vision-language model includes: calculating the aligned image features and the aligned text features using a contrastive learning loss function to obtain a contrastive learning loss value; calculating the aligned text features and the description text using a cross-entropy loss function to obtain a cross-entropy loss value; performing weighted calculation on the contrastive learning loss value and the cross-entropy loss value to obtain a total loss value; and updating the model parameters of the vision-language model according to the total loss value.

[0269] Further, it also includes:

[0270] A hyperparameter adjustment module, which is used to adjust the learning rate of model training using a cosine annealing learning rate scheduler during the iterative training of the vision-language model; and is also used to adjust the weight values of the contrastive learning loss value and the cross-entropy loss value through grid search or random search.

[0271] Further, the process of the model training module 602 updating the model parameters of the vision-language model according to the total loss value includes: updating the model parameters of the vision-language model through an optimizer according to the total loss value; wherein, the optimizer is an optimizer combined with weight decay regularization.

[0272] Further, before obtaining training samples from the dataset, it also includes:

[0273] A data augmentation module, which is used to perform data augmentation on the dataset through preset operations;

[0274] Wherein, the preset operations include:

[0275] Rendering the polymer chemical structure images in the dataset according to multiple parameters;

[0276] and / or, randomly cropping and randomly scaling the polymer chemical structure images in the dataset;

[0277] and / or, randomly adjusting the color channels of the polymer chemical structure images in the dataset;

[0278] and / or, adding Gaussian noise or salt-and-pepper noise to the polymer chemical structure images in the dataset.

[0279] Further, the process of the model deployment module 603 deploying the trained vision-language model includes: deploying the trained vision-language model on a local device and / or a cloud server.

[0280] Since the embodiments in the system part correspond to those in the method part, for the descriptions of the embodiments in the system part, please refer to the descriptions of the embodiments in the method part, which will not be elaborated here.

[0281] This application also provides a storage medium, on which a computer program is stored. When the computer program is executed, the steps provided in the above embodiments can be implemented. The storage medium may include: various media such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs that can store program codes.

[0282] This application also provides an electronic device, which may include a memory and a processor. When the processor calls the computer program stored in the memory, the steps provided in the above embodiments can be implemented. Of course, the electronic device may also include various network interfaces, power supplies and other components.

[0283] The embodiments in the specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other. For the system disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple. For the relevant parts, reference can be made to the description in the method part. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of this application, several improvements and modifications can be made to this application, and these improvements and modifications also fall within the protection scope of this application.

[0284] It should also be noted that in this specification, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article or device including the said element.

Claims

1. A method for identifying a polymer structure, characterized in that: include: Acquire training samples from a data set; wherein the training samples include polymer chemical structure images and description texts, and the description texts include SMILES strings corresponding to the polymer structural formulas, chemical names, and structural feature information; Input the training samples into a visual large language model, and use a loss function to update the model parameters of the visual large language model so as to iteratively train the visual large language model; wherein the visual large language model includes a visual encoder, a text encoder, an alignment module and a fully connected neural network, the visual encoder is used to extract image features from a polymer chemical structure image, the text encoder is used to extract text features from a description text, the alignment module is used to align image features with text features, and the fully connected neural network is used to predict a SMILES string of a polymer structure formula based on the aligned image features and text features; Deploying the trained visual large language model; If a recognition task is received, an unknown polymer chemical structure image corresponding to the recognition task is determined, and a SMILES string of a polymer structural formula of the unknown polymer chemical structure image is generated using the visual large language model.

2. The polymer structural formula identification method according to claim 1, characterized in that: Generating a SMILES string of a polymer structural formula of the unknown polymer chemical structure image using the visual large language model includes: Inputting the unknown polymer chemical structure image into the visual encoder of the visual large language model to obtain current image features; Processing the initial text using the text encoder of the visual large language model to obtain current text features; wherein the initial text includes prompt words and / or related polymer information; Utilizing the alignment module to align the current image feature with the current text feature; Splicing or fusing the aligned current image features with the current text features to obtain image-text features; Inputting the image and text features into the fully connected neural network of the visual large language model to obtain probability distribution information; wherein the probability distribution information is used to describe the probability of each word in the SMILES vocabulary; The probability distribution information is converted using a decoding algorithm to obtain a SMILES character string of a polymer structural formula corresponding to the unknown polymer chemical structure image.

3. The polymer structural formula identification method according to claim 1, characterized in that: The loss function is used to update the model parameters of the visual large language model, including: The contrastive learning loss function is used to calculate the aligned image features and the aligned text features to obtain the contrastive learning loss value; The aligned text features and the description text are calculated using a cross entropy loss function to obtain a cross entropy loss value; Performing weighted calculation on the contrastive learning loss value and the cross entropy loss value to obtain a total loss value; The model parameters of the visual large language model are updated according to the total loss value.

4. The polymer structural formula identification method according to claim 3, characterized in that: The process of iteratively training the visual large language model also includes: Use the cosine annealing learning rate scheduler to adjust the learning rate of model training; The weight values ​​of the contrastive learning loss value and the cross entropy loss value are adjusted by grid search or random search.

5. The polymer structural formula identification method according to claim 3, characterized in that: Updating the model parameters of the visual large language model according to the total loss value includes: According to the total loss value, the model parameters of the visual large language model are updated through an optimizer; wherein the optimizer is an optimizer combined with weight decay regularization.

6. The polymer structural formula identification method according to claim 1, characterized in that: Before obtaining training samples from the dataset, it also includes: Performing data enhancement on the data set through a preset operation; The preset operation includes: Rendering the polymer chemical structure images in the data set according to a plurality of parameters; and / or, randomly cropping and randomly scaling the polymer chemical structure images in the data set; and / or, randomly adjusting the color channels of the polymer chemical structure images in the data set; And / or, adding Gaussian noise or salt and pepper noise to the polymer chemical structure images in the data set.

7. The polymer structural formula identification method according to claim 1, characterized in that: Deploying the trained visual large language model includes: The trained visual large language model is deployed on a local device and / or a cloud server.

8. A polymer structure recognition system, characterized in that: include: A sample acquisition module, used to acquire training samples from a data set; wherein the training samples include polymer chemical structure images and description texts, and the description texts include SMILES strings corresponding to the polymer structural formulas, chemical names, and structural feature information; A model training module, used for inputting the training samples into a visual large language model, and updating the model parameters of the visual large language model using a loss function, so as to iteratively train the visual large language model; wherein the visual large language model comprises a visual encoder, a text encoder, an alignment module and a fully connected neural network, wherein the visual encoder is used for extracting image features from a polymer chemical structure image, the text encoder is used for extracting text features from a description text, the alignment module is used for aligning image features with text features, and the fully connected neural network is used for predicting a SMILES string of a polymer structure formula according to the aligned image features and text features; A model deployment module, used to deploy the trained visual large language model; The model calling module is used to determine the unknown polymer chemical structure image corresponding to the recognition task if a recognition task is received, and generate a SMILES string of the polymer structure formula of the unknown polymer chemical structure image by using the visual large language model.

9. An electronic device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the polymer structure formula identification method according to any one of claims 1 to 7 when calling the computer program in the memory.

10. A storage medium, characterized in that: The storage medium stores computer executable instructions, and when the computer executable instructions are loaded and executed by the processor, the steps of the polymer structure identification method according to any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Polymer performance prediction system and synthesis scheme recommendation system based on two channels

    CN121075491A