ESG evaluation method based on AI language model and computer equipment
Through the ESG evaluation method based on AI language model, combining quantitative and qualitative data, and using BERT model for text analysis, the problem of ignoring text qualitative data in traditional ESG evaluation is solved, and a more accurate and comprehensive enterprise ESG evaluation is achieved.
Patent Information
- Application Number
- CN202510011533.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-04
- Publication Date
- 2025-06-06
AI Technical Summary
Traditional ESG evaluation methods rely on quantitative analysis and ignore the rich value and in-depth information of text qualitative data, resulting in inaccurate and comprehensive evaluation results.
The ESG evaluation method based on AI language model is adopted to obtain training text for preprocessing, assign text tags, and use the BERT model for training and testing, and comprehensive evaluation is carried out in combination with quantitative data.
It provides a more accurate and comprehensive assessment of corporate environmental, social and governance performance, helps companies identify and manage risks, and provides a more comprehensive and accurate ESG evaluation tool.
Smart Images

Figure CN120106631A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of deep learning and data processing, and in particular to an ESG evaluation method and computer device based on an AI language model. Background Art
[0002] For the key areas of ESG evaluation system, namely the construction of environmental, social and governance evaluation system, the traditional mainstream rating method has long relied on quantitative analysis of data to set and evaluate various indicators, including energy consumption, carbon emissions and waste recycling rate. However, although this method can provide quantifiable evaluation results, it ignores the rich value and in-depth information of textual qualitative data. Textual qualitative data, such as company annual reports, social responsibility reports and news reports, contain a large amount of unstructured information related to the company's environmental performance, social responsibility and governance practices.
[0003] When judging whether the text content is green and environmentally friendly, we face the following practical problems: First, what is the style of the corpus collected by the company. For example, the corpus style of the business scope and the corpus style of the annual report are completely different in terms of sentence length and expression characteristics; second, how the company interprets the text expression of this type of collection object. For different industries and different corpus contexts, the same expression may produce different understandings. For example, in the packaging industry, the use of biodegradable materials is considered a green and environmentally friendly practice, but in the medical industry, the use of biodegradable materials may not be green and environmentally friendly. If the degradable materials are used improperly and classification and disinfection measures are not taken, they may still cause pollution to the environment.
[0004] Therefore, there is a need for a method that can combine quantitative data with textual information to comprehensively evaluate a company's environmental, social and governance performance and provide more accurate and objective ESG evaluation results. Summary of the invention
[0005] In order to overcome the problems existing in the related technology, the present invention provides an ESG evaluation method and computer device based on an AI language model, wherein the method can combine quantitative data with text information to comprehensively evaluate the environmental, social and governance performance of the enterprise, and provide more accurate and objective ESG evaluation results.
[0006] An ESG evaluation method based on an AI language model, comprising:
[0007] Acquire a training text, and preprocess the training text to obtain a first preprocessed text;
[0008] assigning a text label to the first preprocessed text, wherein the text label corresponds to a text score;
[0009] Inputting the first preprocessed text and the text label into a to-be-trained model for training to obtain a trained model;
[0010] A test text is obtained, and the test text is input into the trained model for text analysis to obtain an environment score.
[0011] In a preferred technical solution of the present invention, the step of inputting the first preprocessed text and the text label into a model to be trained to obtain a trained model comprises:
[0012] Inputting the first preprocessed text and the text label into the model to be trained, and calculating the loss function;
[0013] Check whether the loss function is less than the loss function threshold. If so, stop training to obtain a trained model.
[0014] In a preferred technical solution of the present invention, after detecting whether the loss function is less than the loss function threshold, the method further includes:
[0015] If the loss function is greater than or equal to the loss function threshold, the model parameters of the model to be trained are adjusted in a supervised learning manner according to the loss function; wherein the model parameters include a learning rate.
[0016] In a preferred technical solution of the present invention, adjusting the model parameters of the model to be trained by supervised learning according to the loss function includes:
[0017] Design the optimizer and set the initial learning rate;
[0018] The optimizer is used to dynamically adjust the initial learning rate of the model to be trained to obtain a final learning rate.
[0019] In a preferred technical solution of the present invention, the use of the optimizer to dynamically adjust the initial learning rate of the model to be trained to obtain the final learning rate includes:
[0020] Adopting the AdamW algorithm to locally adaptively adjust the initial learning rate of each model parameter of the model to be trained to obtain the corresponding preliminary adjusted parameters;
[0021] The StepLR algorithm is used to perform global learning rate adjustment on all the initially adjusted parameters of the model to be trained to obtain a final learning rate.
[0022] In a preferred technical solution of the present invention, the step of assigning a text label to the first preprocessed text, wherein the text label corresponds to a text score, includes:
[0023] Randomly extracting a portion of text from the first preprocessed text;
[0024] Scoring the part of the text according to the sentence length, sentence expression characteristics and corpus context to obtain a text score; wherein the text score is 1 or 0;
[0025] A text label is assigned to the portion of text according to the text score.
[0026] In a preferred technical solution of the present invention, the preprocessing of the training text to obtain a first preprocessed text includes:
[0027] Performing data cleaning on the training text to obtain a cleaned text;
[0028] Segmenting the cleaned text to obtain text words;
[0029] Vectorization is performed on the text words to obtain a first preprocessed text.
[0030] In a preferred technical solution of the present invention, the data cleaning of the training text to obtain the cleaned text includes:
[0031] Redundant data in the training text is removed to obtain a cleaned text; wherein the redundant data includes HTML tags, special characters and stop words.
[0032] In a preferred technical solution of the present invention, the model to be trained is a BERT model, the BERT model includes an encoder and a decoder, and the encoder includes a bidirectional Transformer model; the BERT model adopts a cross-entropy loss function, and the training process of the BERT model includes a pre-training stage and a fine-tuning stage.
[0033] The present invention also provides a computer device, comprising a memory and a processor, wherein the memory stores a computer program, and wherein the processor implements the steps of any one of the above-mentioned ESG evaluation methods based on the AI language model when executing the computer program.
[0034] The beneficial effects of the present invention are:
[0035] The ESG evaluation method based on the AI language model provided by the present invention includes obtaining a training text, preprocessing the training text, and obtaining a first preprocessed text. A text label is assigned to the first preprocessed text, and the text label corresponds to a text score. The first preprocessed text and the text label are input into the model to be trained to obtain a trained model. A test text is obtained, and the test text is input into the trained model for text analysis to obtain an environmental score. The present invention introduces qualitative data in the training text, inputs the first preprocessed text and the text label containing the qualitative data into the model to be trained for deep text analysis, and utilizes the language understanding ability and language processing ability of the model to be trained to extract key information from a large amount of training text, providing a more comprehensive and in-depth perspective for ESG evaluation. The present invention combines qualitative data and a trained model with a deep text analysis function to solve the problem that only quantitative analysis of text data in the past evaluation system leads to the neglect of unstructured information, making the environmental evaluation results more accurate and comprehensive, thereby more comprehensively evaluating the performance of enterprises in the environment, society and governance, helping enterprises to more accurately identify and manage risks, and providing enterprises and regulators with a more comprehensive and accurate ESG evaluation tool, and providing enterprises with a more accurate decision-making basis. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 is a flow chart of the ESG evaluation method based on the AI language model of the present invention;
[0037] Figure 2 It is a schematic diagram of the structure of the BERT model adopted in the present invention. DETAILED DESCRIPTION
[0038] The preferred embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although the preferred embodiments of the present invention are shown in the accompanying drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments described herein. On the contrary, these embodiments are provided to make the present invention more thorough and complete, and to fully convey the scope of the present invention to those skilled in the art.
[0039] Example 1
[0040] like Figure 1 As shown, this embodiment provides an ESG evaluation method based on an AI language model, including:
[0041] S1: Acquire a training text, and preprocess the training text to obtain a first preprocessed text.
[0042] S2: Assign a text label to the first preprocessed text, where the text label corresponds to a text score.
[0043] S3: Input the first preprocessed text and the text label into the model to be trained to obtain a trained model.
[0044] S4: Obtain a test text, input the test text into the trained model for text analysis, and obtain an environment score.
[0045] The step of inputting the first preprocessed text and the text label into the to-be-trained model for training to obtain a trained model includes:
[0046] S31: Input the first preprocessed text and the text label into the model to be trained, and calculate the loss function.
[0047] S32: Detect whether the loss function is less than a loss function threshold.
[0048] S33: If the loss function is less than the loss function threshold, the training is stopped to obtain the trained model.
[0049] S34: If the loss function is greater than or equal to the loss function threshold, adjusting the model parameters of the model to be trained by supervised learning according to the loss function; wherein the model parameters include a learning rate.
[0050] Step S33 and step S34 are parallel steps. Figure 2 As shown, the model to be trained is a BERT model, the BERT model includes an encoder and a decoder, and the encoder includes a bidirectional Transformer model; the BERT model adopts a cross entropy loss function, and the training process of the BERT model includes a pre-training stage and a fine-tuning stage.
[0051] The BERT model is based on natural language processing technology and can deeply understand the contextual information in the text. By inputting a training set containing a large amount of first preprocessed text and text labels into the BERT model and training the BERT model, rich semantic knowledge can be obtained, so that the trained model, i.e. the trained BERT model, can efficiently and accurately process the text information in the ESG report and extract key information related to the company's environmental performance.
[0052] The assigning of a text label to the first preprocessed text, wherein the text label corresponds to a text score, includes:
[0053] S21: Randomly extracting a portion of text from the first preprocessed text.
[0054] S22: Score the portion of text according to sentence length, sentence expression characteristics and corpus context to obtain a text score; wherein the text score is 1 or 0.
[0055] S23: Assigning a text label to the portion of text according to the text score.
[0056] 1 / 50 of the first preprocessed texts were randomly selected and scored. The score for those that meet the green environmental protection requirements was 1, and the score for those that do not meet the green environmental protection requirements was 0. When judging whether the text content is "green and environmentally friendly", the following practical problems are faced: First, what is the style of the corpus collected by the enterprise? For example, the corpus style of the business scope and the corpus style of the annual report are completely different in sentence length and expression characteristics; second, how does the enterprise interpret the text expression of this type of collected objects? For different industries and different corpus contexts, the same expression may have different understandings. For example, the use of biodegradable materials in the packaging industry is considered a green and environmentally friendly practice, but the use of biodegradable materials in the medical industry may not be green and environmentally friendly. If the biodegradable materials are used improperly and classification and disinfection measures are not taken, it may still cause pollution to the environment. Based on the above two points, when judging whether the text expression of a specific object meets the requirements of green environmental protection, the judgment criteria will change with the corpus style and industry characteristics.
[0057] The advantage of the BERT model is that, once it has received specific training, it can effectively and accurately score tens of thousands of corpora according to specific standards. When dealing with tasks in different situations, only a small number of sample objects need to be manually classified to create a specific BERT model, which is a reflection of the generalization ability of this intelligent AI system. The scored text labels are used in the training set, and these text labels identify whether the test text represents environmentally friendly behavior or practice.
[0058] The BERT model of this embodiment adopts the version bert-base-Chinese for Chinese text. This version of the BERT model is pre-trained for Chinese text, which enhances the BERT model's ability to understand Chinese text. The BERT model is trained based on multiple first preprocessed texts and multiple text labels, and the model parameters of the BERT model are adjusted by supervised learning. The loss function is used to guide the update of the model parameters, and the learning rate is gradually reduced to improve the generalization performance of the BERT model, avoid overfitting, and ensure that the loss function value continues to decrease until the loss function value converges. After the loss function value of the BERT model converges, the trained model, that is, the trained BERT model, is used to analyze the test text. The test text includes different types of corporate texts and outputs an environmental score.
[0059] Preferably, after obtaining the environmental score, in addition to processing the text information using the BERT model, the evaluation system of the present invention also combines quantitative data obtained from the enterprise level, such as energy consumption, carbon emissions, and waste recycling rates. By combining quantitative data with qualitative text information, the evaluation system can more comprehensively evaluate the environmental, social, and governance performance of the enterprise, namely, ESG performance, and provide more accurate and objective evaluation results. This fusion strategy enables the evaluation system to have a certain degree of objectivity and verifiability while maintaining flexibility.
[0060] The ESG evaluation method based on the AI language model provided in this embodiment includes obtaining a training text, preprocessing the training text, and obtaining a first preprocessed text. A text label is assigned to the first preprocessed text, and the text label corresponds to a text score. The first preprocessed text and the text label are input into the model to be trained to obtain a trained model. A test text is obtained, and the test text is input into the trained model for text analysis to obtain an environmental score. The present invention introduces qualitative data in the training text, inputs the first preprocessed text and the text label containing the qualitative data into the model to be trained for deep text analysis, and utilizes the language understanding ability and language processing ability of the model to be trained to extract key information from a large amount of training text, providing a more comprehensive and in-depth perspective for ESG evaluation. The present invention combines qualitative data and a trained model with a deep text analysis function to solve the problem of only quantitative analysis of text data in the past evaluation system, resulting in the neglect of unstructured information, making the environmental evaluation results more accurate and comprehensive, thereby more comprehensively evaluating the performance of enterprises in the environment, society and governance, helping enterprises to more accurately identify and manage risks, and providing enterprises and regulators with a more comprehensive and accurate ESG evaluation tool, and providing enterprises with a more accurate decision-making basis.
[0061] Example 2
[0062] This embodiment provides an ESG evaluation method based on an AI language model. This embodiment only describes the differences from Embodiment 1. The method of adjusting the model parameters of the model to be trained by supervised learning according to the loss function includes:
[0063] S341: Design optimizer and set initial learning rate.
[0064] S342: Using the optimizer to dynamically adjust the initial learning rate of the model to be trained to obtain a final learning rate.
[0065] The step of dynamically adjusting the initial learning rate of the model to be trained by the optimizer to obtain a final learning rate includes:
[0066] S3421: Adopting the AdamW algorithm to locally adaptively adjust the initial learning rate of each model parameter of the model to be trained, and obtain corresponding preliminary adjusted parameters.
[0067] S3422: Use the StepLR algorithm to perform global learning rate adjustment on all the initially adjusted parameters of the model to be trained to obtain a final learning rate.
[0068] When training deep learning models, the choice of optimizer is crucial. The optimizer is responsible for updating the weight parameters of the model to minimize the loss function. Common optimizers include SGD, Adam, and AdamW. SGD is one of the most basic optimization algorithms in deep learning, but it faces the problem of difficult selection and adjustment of learning rate. SGD usually uses a fixed learning rate. If the learning rate is set too high, the optimization process may oscillate and it will be difficult to converge to the optimal solution. There may also be a gradient explosion. If the learning rate is set too low, the optimization process will be very slow and may fall into a local optimum.
[0069] AdamW is an optimization algorithm based on Adam, which has the function of weight decay. AdamW processes weight decay, i.e. L2 regularization, and learning rate update separately, thus solving some problems of the traditional Adam optimizer in regularization, especially making it more stable when dealing with larger neural networks.
[0070] StepLR is a fixed-step learning rate scheduler. StepLR reduces the learning rate by multiplying the learning rate by a scaling factor gamma every fixed epoch. Preferably, the scaling factor gamma in this embodiment is 0.1.
[0071] In the actual training process, AdamW can locally and adaptively adjust the learning rate of each parameter of the model to be trained, automatically handle the gradient changes of different parameters, and make different network layers or different parameters in the same network layer learn with different step sizes. StepLR is a global learning rate scheduler, which periodically reduces the learning rate of the entire model. In the early stages of training, the model to be trained will converge quickly with a relatively large learning rate, and as StepLR gradually decays, the learning rate decreases, and the model enters a more refined adjustment stage. In this way, in the early stages of training, AdamW allows the model to be trained to quickly find a suitable solution for the parameters, and StepLR gradually reduces the learning rate to avoid parameter oscillation or missing the optimal solution due to excessive learning rates in the later stages.
[0072] AdamW can be used together with SterpLR, the learning rate scheduler, to dynamically adjust the learning rate and further optimize the training process of the model to be trained. In the code, the learning rate is adjusted every 50 epochs, and the learning rate is multiplied by 0.1 each time. This ensures that the model to be trained can converge quickly during the training process and can be fine-tuned in the later stage, ultimately achieving higher performance.
[0073] The initial learning rate was set to 1*e-4, 2*e-4, 1*e-5, 2*e-5, 1*e-6 and 2*e-6 respectively. The learning rate scheduler tried to adjust the learning rate every 50 or 100 epochs respectively. The learning rate iteration was multiplied by 0.1 or 0.01 each time. After multiple experiments, the results were analyzed to obtain the optimal parameter combination.
[0074] The experimental results of the ESG evaluation method based on the language model provided in this embodiment are shown in Table 1. Compared with the previous semantic recognition method, the evaluation system of this embodiment is adjusted from keyword recognition to semantic recognition of the entire sentence, which improves the model's comprehensive understanding and accurate judgment of text semantics.
[0075] Table 1 Experimental results of ESG evaluation method based on language model
[0076]
[0077]
[0078] This embodiment uses the AdamW algorithm to perform local adaptive adjustment on the initial learning rate of each model parameter of the model to be trained to obtain the corresponding preliminary adjusted parameters. The StepLR algorithm is used to perform global learning rate adjustment on all the preliminary adjusted parameters of the model to be trained to obtain the final learning rate. In the actual training process, AdamW can perform local adaptive adjustment on the learning rate of each parameter of the model to be trained, automatically handle the gradient changes of different parameters, so that different network layers or different parameters in the same network layer are learned with different step sizes. StepLR is a global learning rate scheduler, which periodically reduces the learning rate of the entire model.
[0079] Example 3
[0080] This embodiment provides an ESG evaluation method based on an AI language model. This embodiment only describes the differences from Embodiment 1. The preprocessing of the training text to obtain a first preprocessed text includes:
[0081] S12: performing data cleaning on the training text to obtain a cleaned text.
[0082] S13: Segment the cleaned text to obtain text words.
[0083] S14: performing vectorization processing on the text words to obtain a first preprocessed text.
[0084] Before step S12, the method further includes step S11: acquiring training text.
[0085] The step of performing data cleaning on the training text to obtain a cleaned text includes:
[0086] Redundant data in the training text is removed to obtain a cleaned text; wherein the redundant data includes HTML tags, special characters and stop words.
[0087] Collect text data related to the enterprise, that is, the description of the enterprise's business scope. The collected data may contain noise and irrelevant information, so data cleaning is required. Data cleaning includes removing HTML tags, special characters and stop words. Special characters include extra spaces and line breaks. Stop words include: (1) of, (2) is.
[0088] The second is word segmentation. An important step in Chinese text processing is word segmentation. Because there are no obvious separators between words in Chinese text, Chinese text needs to be segmented. Commonly used word segmentation tools include Jieba and THULAC. The purpose of word segmentation is to split the text into multiple meaningful words. Finally, input vectorization is performed. All preprocessed data will be converted into vector form as the input of the BERT model. Through the above preprocessing steps, the text data is systematically converted into a format suitable for BERT model processing. This not only ensures the consistency and standardization of the model input, but also retains the semantic information of the text to the maximum extent, thereby improving the performance and accuracy of the model.
[0089] This embodiment preprocesses the training text to obtain a first preprocessed text, including performing data cleaning on the training text to obtain a cleaned text. The cleaned text is segmented to obtain text words. The text words are vectorized to obtain a first preprocessed text. Through the above preprocessing steps, the text data is systematically converted into a format suitable for BERT model processing. This not only ensures the consistency and standardization of the model input, but also retains the semantic information of the text to the maximum extent, thereby improving the performance and accuracy of the model.
[0090] Example 4
[0091] This embodiment provides an ESG evaluation device based on an AI language model, including:
[0092] (1) Text processing module: AI models such as the BERT model are used as the core technology for text processing. Through a large amount of pre-training and fine-tuning of sample texts, the model can deeply understand the text information provided by companies around ESG and extract key information related to the company's environmental, social and governance performance.
[0093] (2) Data fusion module: This module integrates the key information extracted by the text processing module with the quantitative data obtained from the enterprise level to form a comprehensive and integrated ESG data set.
[0094] (3) Evaluation module: Based on the ESG data set, an evaluation model is constructed to achieve a comprehensive evaluation of the company's ESG performance by setting evaluation indicators and weights.
[0095] (4) Result output module: The evaluation results are displayed in a visual form, including scoring, ranking and reporting, so that users can intuitively understand the company's ESG performance.
[0096] This technical solution, by introducing advanced technologies such as AI models such as the BERT model, achieves efficient processing of text information and comprehensive analysis of quantitative data in the ESG evaluation system, improving the efficiency and accuracy of the evaluation. At the same time, through data fusion and the construction of an evaluation model, a comprehensive evaluation of the company's ESG performance is achieved, making the evaluation results more objective, comprehensive and accurate. In addition, this technical solution is also scalable and customizable, and can be customized and optimized according to user needs to meet different evaluation needs and scenarios.
[0097] This embodiment also provides a computer device, which may be a server, wherein the computer device includes a processor, a memory, a network interface, and a database connected via a system bus. The processor of the computer design is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal via a network connection.
[0098] This embodiment also provides a computer-readable storage medium on which a computer program is stored, and when the computer program is executed by a processor, an ESG evaluation method based on an AI language model is implemented. It can be understood that the computer-readable storage medium in this embodiment can be a volatile readable storage medium or a non-volatile readable storage medium.
[0099] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, device, article or method including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, device, article or method. In the absence of further restrictions, an element defined by the sentence "includes a ..." does not exclude the presence of other identical elements in the process, device, article or method including the element.
[0100] The above description is only a preferred embodiment of the present application, and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.
Claims
1. An ESG evaluation method based on an AI language model, characterized in that: include: Acquire a training text, and preprocess the training text to obtain a first preprocessed text; assigning a text label to the first preprocessed text, wherein the text label corresponds to a text score; Inputting the first preprocessed text and the text label into a to-be-trained model for training to obtain a trained model; A test text is obtained, and the test text is input into the trained model for text analysis to obtain an environment score.
2. The ESG evaluation method based on the AI language model according to claim 1, characterized in that: The step of inputting the first preprocessed text and the text label into the to-be-trained model for training to obtain a trained model includes: Inputting the first preprocessed text and the text label into the model to be trained, and calculating the loss function; Check whether the loss function is less than the loss function threshold. If so, stop training to obtain a trained model.
3. The ESG evaluation method based on the AI language model according to claim 2 is characterized in that: After detecting whether the loss function is less than the loss function threshold, the method further includes: If the loss function is greater than or equal to the loss function threshold, the model parameters of the model to be trained are adjusted in a supervised learning manner according to the loss function; wherein the model parameters include a learning rate.
4. The ESG evaluation method based on the AI language model according to claim 3 is characterized in that: The step of adjusting the model parameters of the model to be trained by supervised learning according to the loss function includes: Design the optimizer and set the initial learning rate; The optimizer is used to dynamically adjust the initial learning rate of the model to be trained to obtain a final learning rate.
5. The ESG evaluation method based on the AI language model according to claim 4 is characterized in that: The step of dynamically adjusting the initial learning rate of the model to be trained by the optimizer to obtain a final learning rate includes: Adopting the AdamW algorithm to locally adaptively adjust the initial learning rate of each model parameter of the model to be trained to obtain the corresponding preliminary adjusted parameters; The StepLR algorithm is used to perform global learning rate adjustment on all the initially adjusted parameters of the model to be trained to obtain a final learning rate.
6. The ESG evaluation method based on the AI language model according to claim 1, characterized in that: The assigning of a text label to the first preprocessed text, wherein the text label corresponds to a text score, includes: Randomly extracting a portion of text from the first preprocessed text; Scoring the part of the text according to the sentence length, sentence expression characteristics and corpus context to obtain a text score; wherein the text score is 1 or 0; A text label is assigned to the portion of text according to the text score.
7. The ESG evaluation method based on the AI language model according to claim 1, characterized in that: The preprocessing of the training text to obtain a first preprocessed text includes: Performing data cleaning on the training text to obtain a cleaned text; Segmenting the cleaned text to obtain text words; Vectorization is performed on the text words to obtain a first preprocessed text.
8. The ESG evaluation method based on the AI language model according to claim 7 is characterized in that: The step of performing data cleaning on the training text to obtain a cleaned text includes: Redundant data in the training text is removed to obtain a cleaned text; wherein the redundant data includes HTML tags, special characters and stop words.
9. The ESG evaluation method based on the AI language model according to claim 1, characterized in that: The model to be trained is a BERT model, which includes an encoder and a decoder, and the encoder includes a bidirectional Transformer model; the BERT model adopts a cross-entropy loss function, and the training process of the BERT model includes a pre-training stage and a fine-tuning stage.
10. A computer device comprising a memory and a processor, wherein a computer program is stored in the memory, wherein: When the processor executes the computer program, the steps of the ESG evaluation method based on the AI language model as described in any one of claims 1 to 9 are implemented.