A text word vectorization method and related system in the field of environmental protection
By using dynamic weighting functions and GLOVE models in text data processing in the environmental protection field to generate initial word vectors, and through M3E model optimization training, the problem of low accuracy of word vectorization in the environmental protection field in the existing technology is solved, and more efficient and accurate text data representation is achieved.
Patent Information
- Application Number
- CN202510485995.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-04-17
AI Technical Summary
The existing technology lacks word vectorization methods specifically for the environmental protection field, and the existing models cannot be dynamically adjusted to adapt to changes in the vocabulary, resulting in a low accuracy of word vectorization.
A text word vectorization method in the field of environmental protection is adopted. By obtaining text data in the field of environmental protection, a vocabulary library is established, a co-occurrence list is constructed, and the weight value of word pairs is calculated using dynamic weighting functions, the initial word vector is generated using the GLOVE model, and the final text word vector is generated through the optimization training of the M3E model.
It improves the accuracy and efficiency of text data processing and analysis in the field of environmental protection, enhances the ability to capture the core meaning of text, solves the problem of multimodal data fusion, and improves the model's ability to generalize data without seeing.
Smart Images

Figure CN119990127B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of natural language processing, and particularly relates to a method for text word vectorization in the environmental protection field and related systems. Background Art
[0002] In the environmental protection field, data plays a crucial role. It not only faces multiple entities such as government agencies and enterprises, but also takes environmental protection as the core orientation, deeply influencing various changes in environmental quality protection. These environmental protection field data play an active role in promoting the big data networking and technological progress of environmental protection work, and at the same time help to reduce the implementation delay of environmental protection measures and greatly promote the popularization and improvement of environmental protection awareness. However, it is worth noting that the data in the environmental protection field shows a significant characteristic, that is, the proportion of data textification is relatively high. Among these text data, there are a large number of professional vocabulary and technical terms in the environmental protection field, which makes the data in the environmental protection field a typical text-based data set. In addition, there are some tabular data and a small amount of image materials interspersed in these data sets, further enriching the diversity and complexity of the data.
[0003] In order to improve the public's environmental protection awareness and facilitate the popularization of environmental protection knowledge, it is particularly important to integrate and clean the data in the environmental protection field. Since these data come from multiple channels such as administrative region documents, press conferences, research reports, etc., the data collection process is often relatively complex and requires close cooperation and coordination among different departments and professionals. In the process of data integration, not only the integrity and accuracy of the data should be ensured, but also the timeliness and availability of the data should be paid attention to in order to provide strong data support for subsequent environmental protection decision-making.
[0004] However, in the process of processing and analyzing the data in the environmental protection field, an issue that cannot be ignored is that there is currently a lack of a word vectorization method specifically for the environmental protection field, and the existing models cannot be dynamically adjusted with the change of the vocabulary library, resulting in a relatively low accuracy of word vectorization of the models. Most of the existing word vectorization methods are of general nature. Although they can process the text data in the environmental protection field to a certain extent, in the actual application process, their effects are often not ideal. This is mainly because the professional vocabulary and technical terms in the environmental protection field have their uniqueness and professionalism, and general word vectorization methods are often difficult to accurately capture the internal relationships and semantic features between these words. Therefore, developing a word vectorization method specifically for the environmental protection field is of great significance for improving the accuracy and efficiency of data processing and analysis in the environmental protection field. Summary of the Invention
[0005] The object of the present invention is to overcome the problem of insufficient generalization ability when dealing with unseen data in the process of processing and analyzing environmental protection data, and to provide a method for text word vectorization in the field of environmental protection and related systems.
[0006] To achieve the above object, the present invention adopts the following technical solutions:
[0007] In a first aspect, the present invention provides a method for text word vectorization in the field of environmental protection, including the following steps:
[0008] Obtain text data in the field of environmental protection and establish a vocabulary library;
[0009] Based on the established vocabulary library, construct a co-occurrence list, calculate the weight value of each word pair in the co-occurrence list by using a dynamic weight function, and perform word vectorization on the word pairs with high weight values through the GLOVE (Global Vectors for Word Representation) model to generate initial word vectors;
[0010] Based on the generated initial word vectors, train the M3E (Multimodal Multi-task Embedding) model, jointly optimize the M3E model through in-batch negative sampling contrast learning and cross-modal loss functions, and perform optimization training on the M3E model;
[0011] Input the text data into the trained M3E model to generate text word vectors.
[0012] In the step of obtaining text data in the field of environmental protection and establishing a vocabulary library, the specific method is as follows:
[0013] Obtain text in the field of environmental protection;
[0014] Perform data cleaning and structuring processing on the obtained text in the field of environmental protection;
[0015] Based on the text in the field of environmental protection after cleaning and structuring processing, establish a vocabulary library.
[0016] In the step of constructing a co-occurrence list based on the established vocabulary library, calculating the weight value of each word pair in the co-occurrence list by using a dynamic weight function, and performing word vectorization on the word pairs with high weight values through the GLOVE model to generate initial word vectors, the specific method is as follows:
[0017] Traverse the vocabulary library, count the co-occurrence times of word pairs, and generate a co-occurrence matrix X; the elements of the co-occurrence matrix X are , indicating the word and the word in a window.
[0018] Traverse the vocabulary library, store each word and its occurrence frequency into dictionary D, and return a dictionary D→(a,f), which maps to the ID of the professional term and the occurrence frequency of the professional term; where a represents the ID of the professional term, and f represents the occurrence frequency of the professional term;
[0019] Extract each word pair and its corresponding co-occurrence count from the co-occurrence matrix X, and establish a co-occurrence list based on the co-occurrence count of the word pair and the content of dictionary D;
[0020] Design a dynamic weight function, assign higher weights to the word pairs with higher occurrence counts in the co-occurrence list, and supplement the co-occurrence relationships of the word pairs with lower occurrence counts but related to the environmental protection field to obtain the weight value of each word;
[0021] Initialize a random word vector representation for the word pairs with high weight values in the co-occurrence list as the initial parameters of the GLOVE model, and define the loss function of the GLOVE model;
[0022] Based on the initial parameters of the GLOVE model, use the gradient descent method to minimize the loss function. When the convergence of the loss function no longer changes, obtain the trained GLOVE model;
[0023] Extract the word vector representation of the word pairs with high weight values from the trained GLOVE model as the initial word vectors;
[0024] Among them, the word pairs with high occurrence counts refer to the word pairs with occurrence counts greater than 50 times, and the word pairs with low occurrence counts refer to the word pairs with occurrence counts less than 10 times. Assign a higher weight value of 2 to the word pairs with high occurrence counts, and the word pairs with high weight values refer to the word pairs ranked in the top 10% of the weight values.
[0025] The designed dynamic weight function is as follows:
[0026]
[0027] Among them, represents the weight value, represents the first weight coefficient, , represents the word frequency amplification coefficient, represents the word 's occurrence frequency, represents the word 's occurrence frequency; represents the supplementary term, represents the second weight coefficient, , represents the environmental protection field relevance score of the word pair, Represents a threshold smoothing function, Represents the threshold for setting the co-occurrence count. When happens, , the supplementary item takes effect; when happens, , the supplementary item does not take effect, avoiding repeated enhancement of word pairs with high occurrence frequencies.
[0028] The threshold smoothing function is Sigmoid a function used to smoothly transition the weight supplement between word pairs with high occurrence frequencies and word pairs with low occurrence frequencies. The formula is as follows:
[0029]
[0030] Among them, represents any real number, used to adjust Sigmoid the steepness of the function.
[0031] In the step of training the M3E model based on the generated initial word vectors and optimizing the M3E model through in-batch negative sampling contrast learning, the specific method is as follows:
[0032] Use the gensim library (a Python library for text semantic modeling) to load the initial word vectors generated by the GLOVE model, map each initial word vector to the word embedding space of the GLOVE model, and calculate the proportion of the initial word vectors that are not mapped to the word embedding space of the GLOVE model. If the proportion is greater than 10%, retrain the GLOVE model; otherwise, obtain the trained word vectors of the GLOVE model and proceed to the next step;
[0033] Based on the trained word vectors of the GLOVE model, perform intra-modal optimization training and cross-modal optimization training on the M3E model. The specific method of intra-modal optimization training is as follows:
[0034] Select sentence pairs with similar semantics from the same batch as positive sample pairs;
[0035] Randomly select the text of other sentences from the same batch as negative samples;
[0036] Train the M3E model based on the positive sample pairs and negative samples. Use the InfoNCE loss function to maximize the similarity of the positive sample pairs and minimize the similarity of the negative samples. Continuously update the parameters of the M3E model according to the similarity results until the convergence of the InfoNCE loss function no longer changes, obtaining the trained text data.
[0037] The method of cross-modal optimization training is as follows:
[0038] Load the structured table data in the environmental protection field, and use a Bi-LSTM encoder to generate text word vectors from the trained text data ; Use a multi-layer perceptron to map the table data features into table data vectors ;
[0039] Calculate the cross-modal loss function, perform cross-modal alignment on the text word vectors and the table data vectors, and compare the losses through numerical similarity, which is specifically expressed as follows:
[0040]
[0041] Among them, represents the cross-modal loss function, is the modal balance coefficient and ∈[0.6, 0.8], is the text word vector comparison loss function, is the table data vector comparison loss function;
[0042] When the convergence of the cross-modal loss function no longer changes, the trained M3E model is obtained.
[0043] In a second aspect, the present invention provides a text word vectorization system in the environmental protection field, including
[0044] A data acquisition and establishment module: used to acquire text data in the environmental protection field and establish a vocabulary library;
[0045] An initial word vector generation module, used to construct a co-occurrence list based on the established vocabulary library, calculate the weight values of each word pair in the co-occurrence list using a dynamic weight function, and perform word vectorization on the word pairs with high weight values through the GLOVE model to generate initial word vectors;
[0046] An M3E model optimization module, used to train the M3E model based on the generated initial word vectors, jointly optimize the M3E model through in-batch negative sampling contrast learning and the cross-modal loss function, and perform optimization training on the M3E model;
[0047] A word vector generation module, used to input text data into the trained M3E model to generate text word vectors.
[0048] In a third aspect, the present invention provides an electronic device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the steps of a text word vectorization method in the environmental protection field are implemented.
[0049] In a fourth aspect, the present invention provides a storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of a text word vectorization method in the environmental protection field are implemented.
[0050] Compared with the prior art, the present invention has the following beneficial effects:
[0051] The present invention provides a method for text word vectorization in the field of environmental protection, including the following steps: obtaining text data in the field of environmental protection and establishing a vocabulary library; constructing a co-occurrence list based on the established vocabulary library, calculating the weight value of each word pair in the co-occurrence list using a dynamic weight function, and performing word vectorization on the word pairs with high weight values through the GLOVE model to generate initial word vectors; training the M3E model based on the generated initial word vectors, jointly optimizing the M3E model through in-batch negative sampling contrast learning and a cross-modal loss function, and performing optimized training on the M3E model; inputting the text data into the trained M3E model to generate text word vectors. By collecting text data in the field of environmental protection and establishing a vocabulary library, and calculating the weight value of each word in the co-occurrence list using a dynamic weight function, and performing word vectorization on the word pairs with high weight values through the GLOVE model, it can ensure that the GLOVE model has higher accuracy when processing text in the field of environmental protection. Selecting word pairs with high weight values based on the dynamic weight function for optimizing the M3E model enables the M3E model to more accurately capture the core meaning of the text, thereby improving the representation ability of the text; by jointly optimizing the M3E model through in-batch negative sampling contrast learning and a cross-modal loss function, solving the problem of multi-modal data fusion in the field of environmental protection, enhancing the representation ability of word vectors for numerical data, and enhancing the generalization ability of the M3E model for unseen text data in the field of environmental protection; using the GLOVE model and the M3E model comprehensively for word vectorization provides richer semantic information, capable of capturing the statistical relationships between words and the subtle differences of words in different contexts. By performing word vectorization processing on text data in the field of environmental protection, hidden semantic relationships and patterns in the text can be discovered, providing strong support for knowledge discovery and decision-making support in the field of environmental protection.
[0052] Furthermore, the introduction of the dynamic weight function can, according to different contexts and task requirements, be able to dynamically adjust the weights of vocabulary in real-time by perceiving the changes in vocabulary semantics, which helps the M3E model better understand the dynamic semantics of the text and improves the accuracy of the model's vectorized representation.
[0053] Furthermore, the trained M3E model can generate high-quality text word vectors, and these text word vectors can be used as feature inputs for subsequent text analysis, classification, retrieval, and other tasks, thus supporting the requirements of various text processing tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1 is a flowchart of the method of the present invention;
[0055] Figure 2 is a system module diagram of the present invention;
[0056] Figure 3 This is the system diagram of Embodiment 5 of the present invention. Detailed implementation manners
[0057] To further understand the content of the present invention, the following describes the present invention in detail with reference to the accompanying drawings and specific embodiments. It should be understood that the embodiments are only for explaining the present invention and not for limiting it.
[0058] Embodiment 1
[0059] As Figure 1 shown, a method for text word vectorization in the field of environmental protection includes the following steps:
[0060] S1: Obtain text data in the field of environmental protection and establish a vocabulary library;
[0061] S2: Construct a co-occurrence list based on the established vocabulary library, calculate the weight value of each word pair in the co-occurrence list using a dynamic weight function, and perform word vectorization on the word pairs with high weight values through the GLOVE model to generate initial word vectors;
[0062] S3: Train the M3E model based on the generated initial word vectors, jointly optimize the M3E model through in-batch negative sampling contrast learning and a cross-modal loss function, and perform optimized training on the M3E model;
[0063] S4: Input the text data into the trained M3E model to generate text word vectors.
[0064] Embodiment 2
[0065] This embodiment provides a method for text word vectorization in the field of environmental protection. The specific steps are as follows:
[0066] S1: Obtain text data in the field of environmental protection and establish a vocabulary library. The specific method is as follows:
[0067] Obtain text in the field of environmental protection, including unstructured data (policy documents, academic papers, industry reports) and structured data (records in environmental protection databases, standardized tables);
[0068] Perform data cleaning and structuring on the obtained text in the field of environmental protection: Use regular expressions to remove special symbols (such as HTML (web page) tags, garbled characters), numbers, and irrelevant punctuation; Perform word segmentation on the text in the field of environmental protection using the jieba word segmentation tool, and load a professional vocabulary library in the field of environmental protection (such as "carbon emission", "PM2.5", "biodiversity") to improve the accuracy of word segmentation; Apply a custom stop word list (including general stop words and words irrelevant to the field of environmental protection, such as "in summary", "it is reported", etc.);
[0069] Build a vocabulary library: Build a vocabulary library based on the environmentally friendly domain texts after cleaning and structuring;
[0070] S2: Build a co-occurrence list based on the established vocabulary library, calculate the weight value of each word pair in the co-occurrence list using a dynamic weight function, and perform word vectorization on the word pairs with high weight values through the GLOVE model to generate initial word vectors. The specific method is as follows:
[0071] 1) Build a co-occurrence matrix: Define the sliding window size as 5 (that is, 5 words before and after the current word are the context), traverse the vocabulary library, count the co-occurrence times of word pairs, and generate a co-occurrence matrix X. The elements of the co-occurrence matrix X are , indicating the word and the word co-occurrence times within one window;
[0072] 2) Build a dictionary D of the occurrence frequencies of professional vocabulary: Traverse the vocabulary library and count the occurrence times of each word; store each word and its occurrence frequency in the dictionary D, and return a dictionary D→(a,f), mapping to the ID of the professional vocabulary and the occurrence frequency of the professional vocabulary; where, a represents the ID of the professional vocabulary, and f represents the occurrence frequency of this professional vocabulary;
[0073] 3) Extract each word pair and its corresponding co-occurrence times from the co-occurrence matrix X, and establish a co-occurrence list based on the co-occurrence times of the word pairs and the content of the dictionary D.
[0074] 4) Design a dynamic weight function to assign higher weights to the word pairs with high occurrence times in the co-occurrence list, and supplement the co-occurrence relationships of the word pairs with low occurrence times but related to the environmentally friendly domain to obtain the weight value of each word pair.
[0075] The designed dynamic weight function is as follows:
[0076]
[0077] Among them, represents the weight value, represents the first weight coefficient, , represents the word frequency amplification coefficient, represents the word occurrence frequency, represents the word occurrence frequency; represents the supplementary term, represents the second weight coefficient, , represents the environmentally friendly domain relevance score of the word pair, represents the threshold smoothing function, Represents the threshold for setting the co-occurrence count. When occurs, , the supplementary item takes effect; when occurs, , the supplementary item does not take effect, avoiding repeated enhancement of word pairs with high occurrence counts;
[0078] Among them, is Sigmoid a function, , used for smooth transition of weight supplementation between word pairs with high occurrence counts and word pairs with low occurrence counts, represents any real number, used to adjust Sigmoid the steepness of the function.
[0079] Among them, word pairs with high occurrence counts refer to word pairs with occurrence counts greater than 50 times, word pairs with low occurrence counts refer to word pairs with occurrence counts less than 10 times, and word pairs with high weight values are given a higher weight value of 2, and word pairs with high weight values refer to word pairs ranked in the top 10% of weight values.
[0080] 5) Initialize a random word vector representation for word pairs with high weight values (words related to the environmental protection field) in the co-occurrence list as the initial parameters of the GLOVE model;
[0081] Define the loss function of the GLOVE model as follows:
[0082]
[0083] Among them, represents the loss function of the GLOVE model, represents the size of the vocabulary, is the word vector representation of word , is the word vector representation of word , and are two bias terms, is the weight function, used to adjust the weights of different co-occurrence counts, and the weight function is expressed as follows:
[0084]
[0085] Among them, is a constant, is the parameter for adjusting the weight function, with a value of 0.8;
[0086] The function of the above weight function is as follows: control the weight of the co-occurrence matrix so that it is not too large and does not increase after reaching a certain degree; if two words do not co-occur within a window, then they will not participate in the calculation of the loss function, that is, h(0)=0;
[0087] That is, the above weight function can be expressed as:
[0088] Among them, is the weight function in , taking the value of 100, which plays the role of a threshold. This threshold determines the upper limit of the co-occurrence times, and the co-occurrence times exceeding this upper limit will no longer increase the impact when calculating the weight.
[0089] 6) Based on the initial parameters of the GLOVE model, use the gradient descent method to minimize the loss function, and iteratively adjust the word vector representations and bias terms of word pairs with high weight values to minimize the loss function. When the convergence of the loss function no longer changes, obtain the trained GLOVE model;
[0090] 7) Extract the word vector representations of word pairs with high weight values from the trained GLOVE model as the initial word vectors.
[0091] S3: Train the M3E model based on the generated initial word vectors, and jointly optimize the M3E model through in-batch negative sampling contrast learning and cross-modal loss function, and optimize and train the M3E model. The specific method is as follows:
[0092] Use the gensim library to load the initial word vectors generated by the GLOVE model, map each initial word vector to the word embedding space of the GLOVE model, and assign the initial word vectors that are not mapped to the word embedding space of the GLOVE model as zero vectors; calculate the proportion of zero vectors. If the proportion is greater than 10%, then retrain the GLOVE model. Otherwise, obtain the trained word vectors of the GLOVE model and proceed to the next step;
[0093] Based on the trained word vectors of the GLOVE model, perform intra-modal optimization training and cross-modal optimization training on the M3E model. The specific method is as follows:
[0094] The intra-modal (text data) optimization training method is as follows:
[0095] Construct positive sample pairs: Select sentence pairs with similar semantics from the same batch (such as "Industrial wastewater treatment technology" and "Wastewater purification process") as positive sample pairs;
[0096] Construct negative samples: Randomly select texts of other sentences from the same batch as negative samples; these negative samples can be samples that are similar to positive samples but are incorrectly labeled, or completely unrelated samples;
[0097] When constructing negative samples, a semantic similarity threshold control is introduced. A semantic similarity matrix is constructed based on the text data in the same batch, and the similarity score is calculated. If the similarity score between the selected positive sample pair and the negative sample exceeds 0.7, the negative sample is judged as a false negative sample and the false negative sample data is removed and resampled to achieve dynamic negative sample sampling, avoid interference from false negative samples, improve the accuracy of contrastive learning, and ensure that the negative sample and positive sample pairs are truly semantically unrelated.
[0098] The M3E model is trained based on the constructed positive sample pairs and negative samples. The InfoNCE loss function is used to maximize the similarity of the positive sample pairs and minimize the similarity of the negative samples. The parameters of the M3E model are continuously updated according to the similarity results until the convergence of the InfoNCE loss function does not change, and the trained text data is obtained.
[0099] Among them, the definition of InfoNCE loss function is as follows:
[0100]
[0101] in, represents the InfoNCE loss function, Represents the index variable when summing, ranging from 1 to , represents the total number of negative samples in the same batch, represents the score of a specific category, is the similarity of the positive sample pair, is the similarity of negative samples; is the temperature parameter, which is used to adjust the scale of the similarity score;
[0102] The cross-modal optimization training method is as follows:
[0103] During the optimization training phase of the M3E model, structured table data in the field of environmental protection (such as pollutant concentration values, equipment operating parameters, etc.) are loaded simultaneously. The trained text data and structured table data are processed in the following ways:
[0104] Use Bi-LSTM encoder to generate text word vectors from trained text data ;
[0105] Use a multi-layer perceptron to map tabular data features into tabular data vectors ;
[0106] Calculate the cross-modal loss function, perform cross-modal alignment on the text word vectors and the table data vectors, and compare the loss through numerical similarity, which is specifically expressed as follows:
[0107]
[0108] Among them, represents the cross-modal loss function, is the modal balance coefficient and ∈[0.6, 0.8], is the text word vector comparison loss function, is the table data vector comparison loss function. When the convergence of the cross-modal loss function no longer changes, the trained M3E model is obtained.
[0109] S4: Use the sentence-transformers library to load the trained M3E model, input the text data into the trained M3E model, and generate text word vectors.
[0110] Example 3
[0111] Apply the method described in Example 2, and compare it with using the GLOVE model alone and using the M3E model alone for word vectorization of texts in the environmental protection field. The results are as follows:
[0112] Word vectorization is mainly applied to the retrieval system in the environmental protection field to achieve the purpose of intelligently answering users' questions. Therefore, after word vectorization, the performance indicators of the information retrieval system are selected, and then the indicators are compared.
[0113] Select Recall (recall rate), MRR (Mean Reciprocal Rank, average reciprocal rank), and NDCG (Normalized Discounted Cumulative Gain, normalized discounted cumulative gain) to measure the performance of the retrieval system after word vectorization.
[0114] Recall@q: Among the top q results, the proportion of the number of relevant documents to the total number of all relevant documents, and its formula is expressed as follows:
[0115] .
[0116] MRR@q: Among the top q results, the ranking quality of the first relevant document, and its formula is expressed as follows:
[0117]
[0118] Among them, represents the total number of relevant documents for the query, represents for the The rank of the first relevant document among the first q results for a query.
[0119] NDCG@q: The normalized discounted cumulative gain value of relevant documents among the first q results, and its formula is as follows:
[0120]
[0121] Where
[0122] DCG@q (Discounted Cumulative Gain): Represents the relevance score at the position of each relevant document among the first q results;
[0123] IDCG@q (Ideal Discounted Cumulative Gain): Represents the relevance score at the position of each relevant document among the first q results when all relevant documents are arranged in descending order of relevance.
[0124] After setting the basic fine-tuning parameters of the GLOVE model and the M3E model, the index comparison of querying using the GLOVE model alone, the M3E model alone, and the combination of the GLOVE model and the M3E model is as follows in the table:
[0125] Table 1 Comparison of Index Results of Three Models
[0126]
[0127] It can be clearly seen that through the manifestation of the above indexes, the result of using the GLOVE model alone is the lowest, and the result of using the M3E model alone is higher than that of using the GLOVE model alone.
[0128] All indexes after combining the GLOVE model and the M3E model have obvious improvements compared with using only one model. The experimental results verify the effectiveness of the present invention.
[0129] Example 4
[0130] An environmental protection field text word vectorization system, including:
[0131] Data acquisition and establishment module: Used to acquire environmental protection field text data and establish a vocabulary library;
[0132] Initial word vector generation module: Used to construct a co-occurrence list based on the established vocabulary library, calculate the weight value of each word pair in the co-occurrence list using a dynamic weight function, and perform word vectorization on the word pairs with high weight values through the GLOVE model to generate initial word vectors;
[0133] The M3E model optimization module is used to train the M3E model based on the generated initial word vectors, and jointly optimize the M3E model through in-batch negative sampling contrast learning and cross-modal loss functions to perform optimization training on the M3E model;
[0134] The word vector generation module is used to input text data into the trained M3E model to generate text word vectors.
[0135] Embodiment 5
[0136] As Figure 3 shown, the present invention also provides an electronic device 100 for a text word vectorization method in the environmental protection field; the electronic device 100 includes a memory 101, at least one processor 102, a computer program 103 stored in the memory 101 and executable on the at least one processor 102, and at least one communication bus 104.
[0137] The memory 101 can be used to store the computer program 103. The processor 102 realizes the steps of the text word vectorization method in the environmental protection field described in Embodiment 1 by running or executing the computer program stored in the memory 101 and calling the data stored in the memory 101. The memory 101 may mainly include a program storage area and a data storage area. Among them, the program storage area may store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area may store data created according to the use of the electronic device 100 (such as audio data, etc.). In addition, the memory 101 may include non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, at least one magnetic disk storage device, a flash device, or other non-volatile solid-state storage devices.
[0138] The at least one processor 102 may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor 102 may be a microprocessor or the processor 102 may also be any conventional processor, etc. The processor 102 is the control center of the electronic device 100, and connects various parts of the entire electronic device 100 through various interfaces and lines.
[0139] The memory 101 in the electronic device 100 stores multiple instructions to implement a method for text word vectorization in the field of environmental protection. The processor 102 can execute the multiple instructions to implement:
[0140] Obtain text data in the field of environmental protection and establish a vocabulary library;
[0141] Based on the established vocabulary library, construct a co-occurrence list, calculate the weight value of each word pair in the co-occurrence list using a dynamic weight function, and perform word vectorization on the word pairs with high weight values through the GLOVE model to generate initial word vectors;
[0142] Based on the generated initial word vectors, train the M3E model, and jointly optimize the M3E model through in-batch negative sampling contrast learning and a cross-modal loss function to perform optimized training on the M3E model;
[0143] Input the text data into the trained M3E model to generate text word vectors.
[0144] Embodiment 6
[0145] If the modules / units integrated in the electronic device 100 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-described embodiment methods of the present invention, it can also be completed by a computer program instructing relevant hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-described method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, and read-only memory (ROM, Read-Only Memory).
[0146] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, optical storage, etc.) containing computer-usable program code.
[0147] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of processes and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0148] These computer program instructions can also be stored in a computer-readable memory capable of guiding a computer or other programmable data processing devices to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0149] These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are performed on the computer or other programmable apparatus to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one process or a plurality of processes and / or blocks Figure 1 one process or a plurality of processes and / or blocks Figure 1 steps for implementing the functions specified in one block or a plurality of blocks.
[0150] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: the specific implementation manners of the present invention can still be modified or equivalently replaced, and any modification or equivalent replacement that does not depart from the spirit and scope of the present invention shall be covered by the protection scope of the claims of the present invention.
Claims
1. A method for word vectorization of environmental protection text, characterized in that: The steps include: Obtain text data in the field of environmental protection and establish a vocabulary database; Based on the established vocabulary library, a co-occurrence list is constructed, and the weight value of each word pair in the co-occurrence list is calculated using a dynamic weight function. The word pairs with high weight values are vectorized through the GLOVE model to generate the initial word vector; The specific method is as follows: Traverse the vocabulary, count the number of co-occurrences of word pairs, and generate a co-occurrence matrix X; the elements of the co-occurrence matrix X are , Representing words With words The number of co-occurrences within a window; Traverse the vocabulary library, store each word and its frequency of occurrence in the dictionary D, and return a dictionary D→(a,f), which maps the ID of professional vocabulary and the frequency of occurrence of professional vocabulary; where a represents the ID of professional vocabulary and f represents the frequency of occurrence of the professional vocabulary; Extract each word pair and its corresponding co-occurrence count from the co-occurrence matrix X, and build a co-occurrence list based on the co-occurrence count of the word pair and the content of the dictionary D; Design a dynamic weight function to assign higher weights to word pairs with high occurrences in the co-occurrence list, and supplement the co-occurrence relationship of word pairs with low occurrences but related to the environmental protection field to obtain the weight value of each word pair; Initialize a random word vector representation for the word pairs with high weight values in the co-occurrence list as the initial parameters of the GLOVE model and define the loss function of the GLOVE model; Based on the initial parameters of the GLOVE model, the gradient descent method is used to minimize the loss function. When the convergence of the loss function no longer changes, the trained GLOVE model is obtained. Extract the word vector representation of the word pair with high weight value from the trained GLOVE model as the initial word vector; Among them, the word pairs with high occurrence times refer to the word pairs with more than 50 occurrence times, the word pairs with low occurrence times refer to the word pairs with less than 10 occurrence times, and the word pairs with high occurrence times are given a higher weight value of 2. The word pairs with high weight values refer to the word pairs with the top 10% weight values. Based on the generated initial word vectors, the M3E model is trained and optimized by jointly optimizing the M3E model through in-batch negative sampling contrastive learning and cross-modal loss function. Specifically, the cross-modal loss function is: in, represents the cross-modal loss function, is the modal balance coefficient and ∈[0.6,0.8], is the text word vector contrast loss function, is the contrast loss function for the tabular data vector; Input text data into the trained M3E model to generate text word vectors.
2. According to claim 1, a method for word vectorization of environmental protection text, characterized in that: In the step of obtaining text data in the field of environmental protection and establishing a vocabulary library, the specific method is as follows: Get texts in the field of environmental protection; Carry out data cleaning and structural processing on the acquired environmental protection texts; A vocabulary database is established based on the environmental protection texts after cleaning and structuring.
3. According to the environmental protection field text word vectorization method according to claim 1, it is characterized in that: In the step of constructing a co-occurrence list based on the established vocabulary library, calculating the weight value of each word pair in the co-occurrence list using a dynamic weight function, and vectorizing the word pairs with high weight values using the GLOVE model to generate the initial word vector, the designed dynamic weight function is as follows: in, represents the weight value, represents the first weight coefficient, , represents the word frequency magnification factor, Representing words The frequency of occurrence, Representing words The frequency of occurrence; Indicates supplementary items. represents the second weight coefficient, , represents the environmental protection domain relevance score of the word pair, represents the threshold smoothing function, Indicates setting the threshold of co-occurrence times. hour, , the supplementary item takes effect; when hour, , the supplementary items are not effective, avoiding repeated enhancement of word pairs with high occurrence frequency.
4. A method for word vectorization of environmental protection text according to claim 3, characterized in that: The threshold smoothing function for Sigmoid The function is used to smoothly transition the weight complement between word pairs with high occurrence frequency and word pairs with low occurrence frequency. The formula is as follows: in, represents any real number, For adjustment Sigmoid The steepness of the function.
5. According to claim 1, a method for word vectorization of environmental protection text, characterized in that: The M3E model is trained based on the generated initial word vector, and the M3E model is jointly optimized by In-batch negative sampling contrastive learning and cross-modal loss function. In the step of optimizing the training of the M3E model, the specific method is as follows: Use the gensim library to load the initial word vectors generated by the GLOVE model, map each initial word vector to the word embedding space of the GLOVE model, and calculate the proportion of the initial word vectors that are not mapped to the word embedding space of the GLOVE model. If the proportion is greater than 10%, retrain the GLOVE model; otherwise, obtain the word vectors trained by the GLOVE model and proceed to the next step; Based on the word vectors trained by the GLOVE model, the M3E model is trained with intra-modal optimization and cross-modal optimization. The specific method of intra-modal optimization training is as follows: Select sentence pairs with similar semantics from the same batch as positive sample pairs; Randomly select texts of other sentences from the same batch as negative samples; The M3E model is trained based on positive sample pairs and negative samples. The InfoNCE loss function is used to maximize the similarity of positive sample pairs and minimize the similarity of negative samples. The parameters of the M3E model are continuously updated according to the similarity results until the convergence of the InfoNCE loss function no longer changes, and the trained text data is obtained.
6. A method for word vectorization of environmental protection text according to claim 5, characterized in that: The cross-modal optimization training method is as follows: Load structured table data in the field of environmental protection, and use the Bi-LSTM encoder to generate text word vectors from the trained text data ; Use a multi-layer perceptron to map tabular data features into tabular data vectors ; Calculate the cross-modal loss function, align the text word vector with the tabular data vector cross-modally, and compare the loss through numerical similarity; When the convergence of the cross-modal loss function no longer changes, the trained M3E model is obtained.
7. A text word vectorization system in the field of environmental protection, based on a text word vectorization method in the field of environmental protection as claimed in any one of claims 1 to 6, characterized in that: include: Data acquisition and establishment module: used to acquire text data in the field of environmental protection and establish a vocabulary database; The initial word vector generation module is used to build a co-occurrence list based on the established vocabulary library, calculate the weight value of each word pair in the co-occurrence list using a dynamic weight function, and use the GLOVE model to vectorize the word pairs with high weight values to generate the initial word vector; The M3E model optimization module is used to train the M3E model based on the generated initial word vectors, and optimize the M3E model by jointly optimizing the M3E model through in-batch negative sampling contrastive learning and cross-modal loss function. The word vector generation module is used to input text data into the trained M3E model to generate text word vectors.
8. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of a text word vectorization method in the environmental protection field described in any one of claims 1 to 6 are implemented.
9. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of a text word vectorization method in the environmental protection field described in any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Image-text retrieval system and method based on multi-mode consensus perception and momentum comparison
CN118051630A
Retrieval enhancement generation method based on optimized word embedding
CN118332170A