Environmental protection field text word vectorization method and related system
By obtaining text data in the field of environmental protection, establishing a vocabulary library and using GLOVE and M3E models for word vectorization, the problem of low accuracy of word vectorization in the field of environmental protection in the existing technology is solved, and more efficient data processing and analysis is achieved.
Patent Information
- Application Number
- CN202510485995.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-04-17
AI Technical Summary
The existing technology lacks word vectorization methods specifically for the environmental protection field, and the existing models are difficult to adjust dynamically, resulting in a low accuracy of word vectorization.
A method for text word vectorization in the field of environmental protection is proposed. By obtaining text data in the field of environmental protection, a vocabulary library is established, a co-occurrence list is constructed, the weight value of word pairs is calculated using dynamic weighting function, word vectorization is used using GLOVE model, and the M3E model is optimized and trained to generate high-quality text word vectors.
It improves the accuracy and efficiency of data processing and analysis in the field of environmental protection, enhances the ability to represent text in the field of environmental protection, and better captures the relationship between the core meaning of the text and the multimodal data.
Smart Images

Figure CN119990127A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of natural language processing, and specifically relates to a text word vectorization method and a related system in the field of environmental protection. Background Art
[0002] In the field of environmental protection, data plays a vital role. It is not only for multiple subjects such as government agencies and enterprises, but also takes environmental protection as its core orientation, profoundly affecting various changes in environmental quality protection. These environmental protection data play a positive role in promoting the networking of big data and technological progress in environmental protection work, while helping to reduce the delay in the implementation of environmental protection measures and greatly promoting the popularization and improvement of environmental awareness. However, it is worth noting that the data in the field of environmental protection presents a significant feature, that is, the proportion of data text is relatively high. Among these text data, there are a large number of professional vocabulary and technical terms in the field of environmental protection, which makes the data in the field of environmental protection a typical text data set. In addition, these data sets are interspersed with some tabular data and a small amount of image data, which further enriches the diversity and complexity of the data.
[0003] In order to improve the public's environmental awareness and facilitate the popularization of environmental protection knowledge, it is particularly important to integrate and clean the data in the field of environmental protection. Since these data come from multiple channels such as documents from different administrative regions, press conferences, research reports, etc., the data collection process is often complicated and requires close cooperation and coordination between different departments and professionals. In the process of data integration, it is necessary not only to ensure the integrity and accuracy of the data, but also to pay attention to the timeliness and availability of the data, so as to provide strong data support for subsequent environmental protection decisions.
[0004] However, in the process of processing and analyzing data in the field of environmental protection, a problem that cannot be ignored is that there is currently a lack of a word vectorization method specifically for the field of environmental protection, and the existing model cannot be dynamically adjusted as the vocabulary changes, resulting in a low accuracy rate of the model's word vectorization. Most of the existing word vectorization methods are general in nature. Although they can process text data in the field of environmental protection to a certain extent, their effects are often not ideal in actual application. This is mainly because the professional vocabulary and technical terms in the field of environmental protection are unique and professional, and general word vectorization methods often find it difficult to accurately capture the intrinsic connections and semantic features between these words. Therefore, developing a word vectorization method specifically for the field of environmental protection is of great significance to improving the accuracy and efficiency of data processing and analysis in the field of environmental protection. Summary of the invention
[0005] The purpose of the present invention is to overcome the problem of insufficient generalization ability when encountering unseen data during the processing and analysis of existing environmental protection data, and to provide a text word vectorization method and related system in the environmental protection field.
[0006] In order to achieve the above object, the present invention adopts the following technical solution: In a first aspect, the present invention provides a method for word vectorization of text in the field of environmental protection, comprising the following steps: Obtain text data in the field of environmental protection and establish a vocabulary database; Based on the established vocabulary library, a co-occurrence list is constructed. The weight value of each word pair in the co-occurrence list is calculated using a dynamic weight function. The word pairs with high weight values are vectorized through the GLOVE (Global Vectors for Word Representation) model to generate initial word vectors. Based on the generated initial word vectors, the M3E (Multimodal Multi-task Embedding) model is trained. The M3E model is optimized and trained by jointly optimizing the M3E model through in-batch negative sampling contrastive learning and cross-modal loss function. Input text data into the trained M3E model to generate text word vectors.
[0007] In the step of obtaining text data in the field of environmental protection and establishing a vocabulary library, the specific method is as follows: Get texts in the field of environmental protection; Carry out data cleaning and structural processing on the acquired environmental protection texts; A vocabulary database is established based on the environmental protection texts after cleaning and structuring.
[0008] In the step of building a co-occurrence list based on the established vocabulary library, calculating the weight value of each word pair in the co-occurrence list using a dynamic weight function, vectorizing the word pairs with high weight values using the GLOVE model, and generating the initial word vector, the specific method is as follows: Traverse the vocabulary, count the number of co-occurrences of word pairs, and generate a co-occurrence matrix X; the elements of the co-occurrence matrix X are , Representing words With words The number of co-occurrences within a window; Traverse the vocabulary library, store each word and its frequency of occurrence in the dictionary D, and return a dictionary D→(a,f), which maps the ID of professional vocabulary and the frequency of occurrence of professional vocabulary; where a represents the ID of professional vocabulary and f represents the frequency of occurrence of the professional vocabulary; Extract each word pair and its corresponding co-occurrence count from the co-occurrence matrix X, and build a co-occurrence list based on the co-occurrence count of the word pair and the content of the dictionary D; Design a dynamic weight function to assign higher weights to word pairs with high occurrences in the co-occurrence list, and supplement the co-occurrence relationship of word pairs with low occurrences but related to the environmental protection field to obtain the weight value of each word; Initialize a random word vector representation for the word pairs with high weight values in the co-occurrence list as the initial parameters of the GLOVE model and define the loss function of the GLOVE model; Based on the initial parameters of the GLOVE model, the gradient descent method is used to minimize the loss function. When the convergence of the loss function no longer changes, the trained GLOVE model is obtained. Extract the word vector representation of the word pair with high weight value from the trained GLOVE model as the initial word vector; Among them, word pairs with high occurrence times refer to word pairs with more than 50 occurrences, and word pairs with low occurrence times refer to word pairs with less than 10 occurrences. A higher weight value of 2 is assigned to word pairs with high occurrence times, and word pairs with high weight values refer to word pairs with weight values ranking in the top 10%.
[0009] The dynamic weight function of the design is as follows:
[0010] in, represents the weight value, represents the first weight coefficient, , represents the word frequency magnification factor, Representing words The frequency of occurrence, Representing words The frequency of occurrence; Indicates supplementary items. represents the second weight coefficient, , represents the environmental protection domain relevance score of the word pair, represents the threshold smoothing function, Indicates setting the threshold of co-occurrence times. hour, , the supplementary item takes effect; when hour, , the supplementary items are not effective, avoiding repeated enhancement of word pairs with high occurrence frequency.
[0011] The threshold smoothing function for SigmoidThe function is used to smoothly transition the weight complement between word pairs with high occurrence frequency and word pairs with low occurrence frequency. The formula is as follows:
[0012] in, represents any real number, For adjustment Sigmoid The steepness of the function.
[0013] The M3E model is trained based on the generated initial word vector, and the M3E model is optimized by In-batch negative sampling contrast learning. In the step of optimizing the training of the M3E model, the specific method is as follows: Use the gensim library (a Python library for text semantic modeling) to load the initial word vectors generated by the GLOVE model, map each initial word vector to the word embedding space of the GLOVE model, and calculate the proportion of initial word vectors that are not mapped to the word embedding space of the GLOVE model. If the proportion is greater than 10%, retrain the GLOVE model; otherwise, obtain the word vectors trained by the GLOVE model and proceed to the next step. Based on the word vectors trained by the GLOVE model, the M3E model is trained with intra-modal optimization and cross-modal optimization. The specific method of intra-modal optimization training is as follows: Select sentence pairs with similar semantics from the same batch as positive sample pairs; Randomly select texts of other sentences from the same batch as negative samples; The M3E model is trained based on positive sample pairs and negative samples. The InfoNCE loss function is used to maximize the similarity of positive sample pairs and minimize the similarity of negative samples. The parameters of the M3E model are continuously updated according to the similarity results until the convergence of the InfoNCE loss function no longer changes, and the trained text data is obtained.
[0014] The cross-modal optimization training method is as follows: Load structured table data in the field of environmental protection, and use the Bi-LSTM encoder to generate text word vectors from the trained text data ; Use a multi-layer perceptron to map tabular data features into tabular data vectors ; Calculate the cross-modal joint loss function, align the text word vector with the table data vector cross-modally, and compare the loss through numerical similarity, which is specifically expressed as follows:
[0015] in, represents the cross-modal loss function, is the modal balance coefficient and ∈[0.6,0.8], is the text word vector contrast loss function, is the contrast loss function for the tabular data vector; When the convergence of the cross-modal loss function no longer changes, the trained M3E model is obtained.
[0016] In a second aspect, the present invention provides a text word vectorization system in the field of environmental protection, comprising Data acquisition and establishment module: used to acquire text data in the field of environmental protection and establish a vocabulary database; The initial word vector generation module is used to build a co-occurrence list based on the established vocabulary library, calculate the weight value of each word pair in the co-occurrence list using a dynamic weight function, and use the GLOVE model to vectorize the word pairs with high weight values to generate the initial word vector; The M3E model optimization module is used to train the M3E model based on the generated initial word vectors, and optimize the M3E model by jointly optimizing the M3E model through in-batch negative sampling contrastive learning and cross-modal loss function. The word vector generation module is used to input text data into the trained M3E model to generate text word vectors.
[0017] In a third aspect, the present invention provides an electronic device including a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of a method for text word vectorization in the field of environmental protection when executing the computer program.
[0018] In a fourth aspect, the present invention provides a storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of a method for text word vectorization in the field of environmental protection.
[0019] Compared with the prior art, the present invention has the following beneficial effects: The invention provides a word vectorization method for text in the field of environmental protection, comprising the following steps: acquiring text data in the field of environmental protection and establishing a vocabulary library; constructing a co-occurrence list based on the established vocabulary library, calculating the weight value of each word pair in the co-occurrence list by using a dynamic weight function, performing word vectorization on the word pairs with high weight values by using a GLOVE model, and generating initial word vectors; training an M3E model based on the generated initial word vectors, jointly optimizing the M3E model by using In-batch negative sampling contrast learning and a cross-modal loss function, and optimizing the M3E model; inputting text data into the trained M3E model to generate text word vectors. By collecting text data and establishing a vocabulary library for the environmental protection field, and using a dynamic weight function to calculate the weight of each word in the co-occurrence list, the GLOVE model is used to vectorize the word pairs with high weight values, which can ensure that the GLOVE model has higher accuracy when processing environmental protection texts. Based on the dynamic weight function, word pairs with high weight values are selected for optimization of the M3E model, so that the M3E model can more accurately capture the core meaning of the text, thereby improving the representation ability of the text; the M3E model is jointly optimized by In-batch negative sampling contrast learning and cross-modal loss function to solve the problem of multimodal data fusion in the environmental protection field, enhance the representation ability of word vectors for numerical data, and enhance the generalization ability of the M3E model for unseen environmental protection text data; the GLOVE model and the M3E model are used to comprehensively vectorize words, which provides richer semantic information and can capture the statistical relationship between words and the subtle differences of words in different contexts. By vectorizing the text data in the environmental protection field, the hidden semantic relationships and patterns in the text can be discovered, providing strong support for knowledge discovery and decision support in the environmental protection field.
[0020] Furthermore, the introduction of the dynamic weight function can perceive the changes in lexical semantics in real time and dynamically adjust the weight of vocabulary according to different contexts and task requirements, which helps the M3E model better understand the dynamic semantics of the text and improves the accuracy of the model's vectorized representation.
[0021] Furthermore, the trained M3E model can generate high-quality text word vectors, which can be used as feature inputs for subsequent text analysis, classification, retrieval and other tasks, thereby supporting the needs of various text processing tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 is a flow chart of the method of the present invention; Figure 2 It is a system module diagram of the present invention; Figure 3 This is a system diagram of Example 5 of the present invention. DETAILED DESCRIPTION
[0023] In order to further understand the content of the present invention, the present invention is described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the embodiments are only for explaining the present invention and are not intended to limit it.
[0024] Example 1 like Figure 1 As shown, a method for word vectorization of environmental protection text includes the following steps: S1: Obtain text data in the field of environmental protection and establish a vocabulary database; S2: Build a co-occurrence list based on the established vocabulary library, use the dynamic weight function to calculate the weight value of each word pair in the co-occurrence list, and use the GLOVE model to vectorize the word pairs with high weight values to generate the initial word vector; S3: Based on the generated initial word vectors, the M3E model is trained. The M3E model is optimized by jointly optimizing the In-batch negative sampling contrastive learning and the cross-modal loss function. S4: Input text data into the trained M3E model to generate text word vectors.
[0025] Example 2 This embodiment provides a method for word vectorization of text in the field of environmental protection, and the specific steps are as follows: S1: Obtain text data in the field of environmental protection and establish a vocabulary database. The specific method is as follows: Acquire environmental protection texts, including unstructured data (policy documents, academic papers, industry reports) and structured data (records in environmental protection databases, standardized forms); Data cleaning and structural processing of the acquired environmental protection texts: use regular expressions to remove special symbols (such as HTML (webpage) tags, garbled characters), numbers and irrelevant punctuation; use the Jieba word segmentation tool to segment environmental protection texts, load professional dictionaries in the environmental protection field (such as "carbon emissions", "PM2.5", "biodiversity") to improve the accuracy of word segmentation; apply a custom stop word list (including common stop words and irrelevant words in the environmental protection field, such as "in summary", "it is reported", etc.); Constructing a vocabulary database: Building a vocabulary database based on the environmental protection texts after cleaning and structuring; S2: Build a co-occurrence list based on the established vocabulary library, use the dynamic weight function to calculate the weight value of each word pair in the co-occurrence list, and use the GLOVE model to vectorize the word pairs with high weight values to generate the initial word vector. The specific method is as follows: 1) Construct a co-occurrence matrix: define the sliding window size as 5 (i.e., 5 words before and after the current word are the context), traverse the vocabulary, count the number of co-occurrences of word pairs, and generate a co-occurrence matrix X. The elements of the co-occurrence matrix X are , Representing words With words The number of co-occurrences within a window; 2) Construct a dictionary D of the frequency of occurrence of professional vocabulary: traverse the vocabulary database and count the number of occurrences of each word; store each word and its frequency of occurrence in the dictionary D, and return a dictionary D→(a,f), which maps the ID of the professional vocabulary and the frequency of occurrence of the professional vocabulary; where a represents the ID of the professional vocabulary and f represents the frequency of occurrence of the professional vocabulary; 3) Extract each word pair and its corresponding co-occurrence count from the co-occurrence matrix X, and build a co-occurrence list based on the co-occurrence count of the word pair and the content of the dictionary D; 4) Design a dynamic weight function to assign higher weights to word pairs with high occurrence frequencies in the co-occurrence list, and supplement the co-occurrence relationship of word pairs with low occurrence frequencies but related to the environmental protection field to obtain the weight value of each word pair.
[0026] The designed dynamic weight function is as follows:
[0027] in, represents the weight value, represents the first weight coefficient, , represents the word frequency magnification factor, Representing words The frequency of occurrence, Representing words The frequency of occurrence; Indicates supplementary items. represents the second weight coefficient, , represents the environmental protection domain relevance score of the word pair, represents the threshold smoothing function, Indicates setting the threshold of co-occurrence times. hour, , the supplementary item takes effect; when hour, , the supplementary items are not effective, avoiding repeated enhancement of word pairs with high occurrence frequency; in, for Sigmoid function, , used to smoothly transition the weights between word pairs with high occurrences and word pairs with low occurrences, represents any real number, For adjustment Sigmoid The steepness of the function.
[0028] Among them, word pairs with high occurrence times refer to word pairs with more than 50 occurrences, and word pairs with low occurrence times refer to word pairs with less than 10 occurrences. A higher weight value of 2 is assigned to word pairs with high occurrence times, and word pairs with high weight values refer to word pairs with weight values ranking in the top 10%.
[0029] 5) Initialize a random word vector representation for the word pairs with high weight values in the co-occurrence list (words related to the environmental protection field) as the initial parameters of the GLOVE model; Define the loss function of the GLOVE model as follows:
[0030] in, represents the loss function of the GLOVE model, represents the size of the vocabulary, It's a word The word vector representation of It's a word The word vector representation of and are two bias terms, Is a weight function, which is used to adjust the weights of different co-occurrence times. It is expressed as follows:
[0031] in, is a constant, To adjust the parameters of the weight function, the value is 0.8; The role of the above weight function is to control the weight of the co-occurrence matrix not to be too large, and it will not increase after reaching a certain level; if two words do not co-occur in a window, they will not participate in the calculation of the loss function, that is, h(0)=0; That is, the above weight function can be expressed as:
[0032] in, is the weight function In , takes a value of 100, and acts as a threshold. This threshold determines the upper limit of the number of co-occurrences. The number of co-occurrences exceeding this upper limit will no longer increase the impact when calculating the weight.
[0033] 6) Based on the initial parameters of the GLOVE model, the gradient descent method is used to minimize the loss function. The word vector representation and bias term of the word pairs with high weight values are iteratively adjusted to minimize the loss function. When the convergence of the loss function no longer changes, the trained GLOVE model is obtained. 7) Extract the word vector representation of the word pairs with high weight values from the trained GLOVE model as the initial word vector.
[0034] S3: Based on the generated initial word vector training, the M3E model is optimized by in-batch negative sampling contrastive learning and cross-modal loss function. The specific method is as follows: Use the gensim library to load the initial word vectors generated by the GLOVE model, map each initial word vector to the word embedding space of the GLOVE model, and assign the initial word vectors that are not mapped to the word embedding space of the GLOVE model as zero vectors; calculate the proportion of zero vectors, if the proportion is greater than 10%, retrain the GLOVE model, otherwise, obtain the word vectors trained by the GLOVE model and proceed to the next step; Based on the word vectors trained by the GLOVE model, the M3E model is trained for intra-modality optimization and cross-modality optimization. The specific methods are as follows: The in-modality (text data) optimization training method is as follows: Construct positive sample pairs: select sentence pairs with similar semantics (such as "industrial wastewater treatment technology" and "wastewater purification process") from the same batch as positive sample pairs; Construct negative samples: Randomly select texts of other sentences from the same batch as negative samples; these negative samples can be samples that are similar to positive samples but are incorrectly labeled, or completely unrelated samples; When constructing negative samples, a semantic similarity threshold control is introduced. A semantic similarity matrix is constructed based on the text data in the same batch, and the similarity score is calculated. If the similarity score between the selected positive sample pair and the negative sample exceeds 0.7, the negative sample is judged as a false negative sample and the false negative sample data is removed and resampled to achieve dynamic negative sample sampling, avoid interference from false negative samples, improve the accuracy of contrastive learning, and ensure that the negative sample and positive sample pairs are truly semantically unrelated.
[0035] The M3E model is trained based on the constructed positive sample pairs and negative samples. The InfoNCE loss function is used to maximize the similarity of the positive sample pairs and minimize the similarity of the negative samples. The parameters of the M3E model are continuously updated according to the similarity results until the convergence of the InfoNCE loss function does not change, and the trained text data is obtained. Among them, the definition of InfoNCE loss function is as follows:
[0036] in, represents the InfoNCE loss function, Represents the index variable when summing, ranging from 1 to , represents the total number of negative samples in the same batch, represents the score of a specific category, is the similarity of the positive sample pair, is the similarity of negative samples; is the temperature parameter, which is used to adjust the scale of the similarity score; The cross-modal optimization training method is as follows: During the optimization training phase of the M3E model, structured table data in the field of environmental protection (such as pollutant concentration values, equipment operating parameters, etc.) are loaded simultaneously. The trained text data and structured table data are processed in the following ways: Use Bi-LSTM encoder to generate text word vectors from trained text data ; Use a multi-layer perceptron to map tabular data features into tabular data vectors ; Calculate the cross-modal joint loss function, align the text word vector with the table data vector cross-modally, and compare the loss through numerical similarity, which is specifically expressed as follows:
[0037] in, represents the cross-modal loss function, is the modal balance coefficient and ∈[0.6,0.8], is the text word vector contrast loss function, The loss function is compared with the tabular data vector. When the convergence of the cross-modal loss function does not change, the trained M3E model is obtained.
[0038] S4: Use the sentence-transformers library to load the trained M3E model, input the text data into the trained M3E model, and generate text word vectors.
[0039] Example 3 The method described in Example 2 is applied to compare the word vectorization of environmental protection texts using the GLOVE model alone and the M3E model alone. The results are as follows: Word vectorization is mainly used in retrieval systems in the field of environmental protection to intelligently answer user questions. Therefore, word vectorization is selected to perform information retrieval system performance indicators, and then the indicators are compared.
[0040] Recall, MRR (Mean Reciprocal Rank), and NDCG (Normalized Discounted Cumulative Gain) are selected to measure the performance of the retrieval system after word vectorization.
[0041] Recall@q: The ratio of the number of relevant documents to the total number of relevant documents in the first q results. The formula is as follows: .
[0042] MRR@q: The ranking quality of the first relevant document among the first q results, and its formula is as follows:
[0043] in, Represents the total number of relevant documents for the query, Indicates that for For a query, the rank of the first relevant document among the first q results.
[0044] NDCG@q: The normalized discounted cumulative gain value of relevant documents in the first q results, which is expressed as follows:
[0045] in, DCG@q (Discounted Cumulative Gain): represents the relevance score of each relevant document position in the first q results; IDCG@q (Ideal Discounted Cumulative Gain): represents the relevance score of each relevant document position in the first q results when all relevant documents are arranged in descending order of relevance.
[0046] After setting the basic fine-tuning parameters of the GLOVE model and the M3E model, we query the comparison of the indicators using the GLOVE model alone, the M3E model alone, and the combination of the GLOVE model and the M3E model. The comparison results are shown in the following table: Table 1 Comparison of the three model indicators
[0047] It can be clearly seen from the above indicators that the result of using the GLOVE model alone is the lowest, and the result of using the M3E model alone is higher than that of using the GLOVE model alone.
[0048] The various indicators after combining the GLOVE model and the M3E model are significantly improved compared with using only one model. The experimental results verify the effectiveness of the present invention.
[0049] Example 4 A text word vectorization system in the field of environmental protection, comprising: Data acquisition and establishment module: used to acquire text data in the field of environmental protection and establish a vocabulary database; The initial word vector generation module is used to build a co-occurrence list based on the established vocabulary library, calculate the weight value of each word pair in the co-occurrence list using a dynamic weight function, and use the GLOVE model to vectorize the word pairs with high weight values to generate the initial word vector; The M3E model optimization module is used to train the M3E model based on the generated initial word vectors, and optimize the M3E model by jointly optimizing the M3E model through in-batch negative sampling contrastive learning and cross-modal loss function. The word vector generation module is used to input text data into the trained M3E model to generate text word vectors.
[0050] Example 5 like Figure 3 As shown, the present invention also provides an electronic device 100 for a text word vectorization method in the field of environmental protection; the electronic device 100 includes a memory 101, at least one processor 102, a computer program 103 stored in the memory 101 and executable on the at least one processor 102, and at least one communication bus 104.
[0051] The memory 101 can be used to store the computer program 103, and the processor 102 implements the steps of the environmental protection field text word vectorization method described in Example 1 by running or executing the computer program stored in the memory 101 and calling the data stored in the memory 101. The memory 101 may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area may store data (such as audio data) created according to the use of the electronic device 100, etc. In addition, the memory 101 may include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), at least one disk storage device, a flash memory device, or other non-volatile solid-state storage devices.
[0052] The at least one processor 102 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor 102 may be a microprocessor or any conventional processor, etc. The processor 102 is the control center of the electronic device 100, and uses various interfaces and lines to connect various parts of the entire electronic device 100.
[0053] The memory 101 in the electronic device 100 stores a plurality of instructions to implement a text word vectorization method in the field of environmental protection, and the processor 102 can execute the plurality of instructions to implement: Obtain text data in the field of environmental protection and establish a vocabulary database; Based on the established vocabulary library, a co-occurrence list is constructed, and the weight value of each word pair in the co-occurrence list is calculated using a dynamic weight function. The word pairs with high weight values are vectorized through the GLOVE model to generate the initial word vector; Based on the generated initial word vectors, the M3E model is trained and optimized by jointly optimizing the M3E model through in-batch negative sampling contrastive learning and cross-modal loss function. Input text data into the trained M3E model to generate text word vectors.
[0054] Example 6 If the module / unit integrated in the electronic device 100 is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present invention implements all or part of the process in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and the computer program can implement the steps of the above-mentioned method embodiments when executed by the processor. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include: any entity or device that can carry the computer program code, recording medium, U disk, mobile hard disk, disk, optical disk, computer memory and read-only memory (ROM, Read-Only Memory).
[0055] It will be appreciated by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, optical storage, etc.) containing computer-usable program code.
[0056] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0057] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0058] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0059] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, ordinary technicians in the relevant field should understand that the specific implementation methods of the present invention can still be modified or replaced by equivalents. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.
Claims
1. A method for word vectorization of environmental protection text, characterized in that: The steps include: Obtain text data in the field of environmental protection and establish a vocabulary database; Based on the established vocabulary library, a co-occurrence list is constructed, and the weight value of each word pair in the co-occurrence list is calculated using a dynamic weight function. The word pairs with high weight values are vectorized through the GLOVE model to generate the initial word vector; Based on the generated initial word vectors, the M3E model is trained and optimized by jointly optimizing the M3E model through in-batch negative sampling contrastive learning and cross-modal loss function. Input text data into the trained M3E model to generate text word vectors.
2. According to claim 1, a method for word vectorization of environmental protection text, characterized in that: In the step of obtaining text data in the field of environmental protection and establishing a vocabulary library, the specific method is as follows: Get texts in the field of environmental protection; Carry out data cleaning and structural processing on the acquired environmental protection texts; A vocabulary database is established based on the environmental protection texts after cleaning and structuring.
3. According to the environmental protection field text word vectorization method according to claim 1, it is characterized in that: In the step of building a co-occurrence list based on the established vocabulary library, calculating the weight value of each word pair in the co-occurrence list using a dynamic weight function, vectorizing the word pairs with high weight values using the GLOVE model, and generating the initial word vector, the specific method is as follows: Traverse the vocabulary, count the number of co-occurrences of word pairs, and generate a co-occurrence matrix X; the elements of the co-occurrence matrix X are , Representing words With words The number of co-occurrences within a window; Traverse the vocabulary library, store each word and its frequency of occurrence in the dictionary D, and return a dictionary D→(a,f), which maps the ID of professional vocabulary and the frequency of occurrence of professional vocabulary; where a represents the ID of professional vocabulary and f represents the frequency of occurrence of the professional vocabulary; Extract each word pair and its corresponding co-occurrence count from the co-occurrence matrix X, and build a co-occurrence list based on the co-occurrence count of the word pair and the content of the dictionary D; Design a dynamic weight function to assign higher weights to word pairs with high occurrences in the co-occurrence list, and supplement the co-occurrence relationship of word pairs with low occurrences but related to the environmental protection field to obtain the weight value of each word pair; Initialize a random word vector representation for the word pairs with high weight values in the co-occurrence list as the initial parameters of the GLOVE model and define the loss function of the GLOVE model; Based on the initial parameters of the GLOVE model, the gradient descent method is used to minimize the loss function. When the convergence of the loss function no longer changes, the trained GLOVE model is obtained. Extract the word vector representation of the word pair with high weight value from the trained GLOVE model as the initial word vector; Among them, word pairs with high occurrence times refer to word pairs with more than 50 occurrences, and word pairs with low occurrence times refer to word pairs with less than 10 occurrences. A higher weight value of 2 is assigned to word pairs with high occurrence times, and word pairs with high weight values refer to word pairs with weight values ranking in the top 10%.
4. A method for word vectorization of environmental protection text according to claim 3, characterized in that: The dynamic weight function of the design is as follows: in, represents the weight value, represents the first weight coefficient, , represents the word frequency magnification factor, Representing words The frequency of occurrence, Representing words The frequency of occurrence; Indicates supplementary items. represents the second weight coefficient, , represents the environmental protection domain relevance score of the word pair, represents the threshold smoothing function, Indicates setting the threshold of co-occurrence times. hour, , the supplementary item takes effect; when hour, , the supplementary items are not effective, avoiding repeated enhancement of word pairs with high occurrence frequency.
5. According to claim 4, a method for word vectorization of environmental protection text is characterized in that: The threshold smoothing function for Sigmoid The function is used to smoothly transition the weight complement between word pairs with high occurrence frequency and word pairs with low occurrence frequency. The formula is as follows: in, represents any real number, For adjustment Sigmoid The steepness of the function.
6. The method for word vectorization of environmental protection text according to claim 1 is characterized in that: The M3E model is trained based on the generated initial word vector, and the M3E model is jointly optimized by In-batch negative sampling contrastive learning and cross-modal loss function. In the step of optimizing the training of the M3E model, the specific method is as follows: Use the gensim library to load the initial word vectors generated by the GLOVE model, map each initial word vector to the word embedding space of the GLOVE model, and calculate the proportion of the initial word vectors that are not mapped to the word embedding space of the GLOVE model. If the proportion is greater than 10%, retrain the GLOVE model; otherwise, obtain the word vectors trained by the GLOVE model and proceed to the next step; Based on the word vectors trained by the GLOVE model, the M3E model is trained with intra-modal optimization and cross-modal optimization. The specific method of intra-modal optimization training is as follows: Select sentence pairs with similar semantics from the same batch as positive sample pairs; Randomly select texts of other sentences from the same batch as negative samples; The M3E model is trained based on positive sample pairs and negative samples. The InfoNCE loss function is used to maximize the similarity of positive sample pairs and minimize the similarity of negative samples. The parameters of the M3E model are continuously updated according to the similarity results until the convergence of the InfoNCE loss function no longer changes, and the trained text data is obtained.
7. A method for word vectorization of environmental protection text according to claim 6, characterized in that: The cross-modal optimization training method is as follows: Load structured table data in the field of environmental protection, and use the Bi-LSTM encoder to generate text word vectors from the trained text data ; Use a multi-layer perceptron to map tabular data features into tabular data vectors ; Calculate the cross-modal joint loss function, align the text word vector with the table data vector cross-modally, and compare the loss through numerical similarity, which is specifically expressed as follows: in, represents the cross-modal loss function, is the modal balance coefficient and ∈[0.6,0.8], is the text word vector contrast loss function, is the contrast loss function for the tabular data vector; When the convergence of the cross-modal loss function no longer changes, the trained M3E model is obtained.
8. A text word vectorization system in the field of environmental protection, characterized in that: include: Data acquisition and establishment module: used to acquire text data in the field of environmental protection and establish a vocabulary database; The initial word vector generation module is used to build a co-occurrence list based on the established vocabulary library, calculate the weight value of each word pair in the co-occurrence list using a dynamic weight function, and use the GLOVE model to vectorize the word pairs with high weight values to generate the initial word vector; The M3E model optimization module is used to train the M3E model based on the generated initial word vectors, and optimize the M3E model by jointly optimizing the M3E model through in-batch negative sampling contrastive learning and cross-modal loss function. The word vector generation module is used to input text data into the trained M3E model to generate text word vectors.
9. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of a text word vectorization method in the environmental protection field described in any one of claims 1 to 7 are implemented.
10. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of a text word vectorization method in the environmental protection field described in any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Medical text word vectorization method
CN114004225A
Scoring dictionary construction method and system based on English vocabulary linguistic attribute prediction
CN117874242A
Image-text retrieval system and method based on multi-mode consensus perception and momentum comparison
CN118051630A
Retrieval enhancement generation method based on optimized word embedding
CN118332170A
Case analysis method and device, equipment, storage medium and program product
CN119314695A