System and method for compliance management of photovoltaic project
Through the photovoltaic project compliance management system, natural language processing and rule matching technology are used to automatically process the photovoltaic project policy text and the project text to be evaluated, solving the problem of inefficient interpretation and compliance evaluation of photovoltaic project policy, achieving the effect of rapid response to policy requirements and accurate judgment of compliance.
Patent Information
- Application Number
- CN202510199761.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-06-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The text data of the policy documents involved in photovoltaic projects are large and complex, which leads to inefficiency in interpreting and following policies, prone to omissions and misjudgments, and increases compliance risks. At the same time, it is difficult to achieve efficient and accurate assessment of the compliance audit of a large number of photovoltaic projects to be evaluated.
It provides a compliance management system for photovoltaic projects, including acquisition module, text preprocessing module, text extraction module, rule definition module, rule matching module and evaluation module. Through natural language processing and rule matching technology, it automatically processes the photovoltaic project policy text and the project text to be evaluated, accurately divide words, build a rule database, and conduct multi-dimensional matching evaluation.
Significantly speed up the compliance evaluation process, improve overall evaluation efficiency, reduce manual interpretation and review workload, ensure that the project responds quickly to policy requirements, reduce the risk of misjudgment, and accurately judge whether the project is compliant.
Smart Images

Figure CN120146456A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data management, and particularly relates to a compliance management system and method for photovoltaic projects. Background Art
[0002] As an important form of renewable energy utilization, photovoltaic projects are booming at an unprecedented speed. Governments around the world have introduced a large number of relevant policies, and detailed and complex regulatory requirements have been formulated for each link from project establishment, construction, operation to subsidy acquisition, aiming to ensure the quality, safety of photovoltaic projects and the healthy and orderly development of the industry.
[0003] On the one hand, the policies of photovoltaic projects have the characteristic of dynamic update. New technological breakthroughs, changes in the market environment, and upgrades in environmental protection standards have prompted frequent revisions and supplements to policies and regulations, resulting in a huge and complex amount of text data involved in photovoltaic projects. Policy documents at different regions and different levels have different formats and expression styles, which pose great challenges for enterprises and regulatory authorities to interpret and comply with. The traditional manual method of sorting out policies is inefficient, not only consuming a large amount of human and time costs, but also extremely prone to omissions and misjudgments, leading to compliance risks during the project implementation process, and may result in penalties and loss of subsidy qualifications, seriously affecting the economic and social benefits of the project.
[0004] On the other hand, in the face of a large number of photovoltaic projects to be evaluated, it is an even more arduous task to review their compliance one by one. Each project to be evaluated carries a large amount of text information, covering project plans, technical solutions, operation reports, etc. To accurately extract key information and compare it with the current policies, it is almost impossible to achieve efficient and accurate evaluation simply relying on manpower. There is an urgent need in the current market for an intelligent and automated system to solve the connection problem between the interpretation of photovoltaic project policies and project evaluation, and ensure that the entire life cycle of the project is in a compliant state.
[0005] In view of this, there is an urgent need to propose a compliance management system and method for photovoltaic projects. Summary of the Invention
[0006] To this end, the present invention provides a compliance management system for photovoltaic projects to improve the overall evaluation efficiency and ensure that the project can quickly respond to policy requirements and compliance.
[0007] In the first aspect of the present invention, there is provided a compliance management system for photovoltaic projects, including:
[0008] A first acquisition module configured to acquire first text data of photovoltaic project policies from at least one data source;
[0009] A second acquisition module configured to receive a project to be evaluated and acquire second text data from the project to be evaluated;
[0010] A text preprocessing module, configured to perform word segmentation on the first text data and the second text data by using a preset natural language processing algorithm to obtain first word segmentation data and second word segmentation data;
[0011] A text extraction module, configured to determine a first reference text and a first description text from the first word segmentation data, and at the same time configured to determine a second reference text and a second description text of the second word segmentation data; the reference text is a text including a first preset number of characters classified according to part of speech, and the description text is a descriptive vocabulary of the reference text, and each description text includes at least two groups of texts within a range from the first preset number of characters to the second preset number of characters;
[0012] A rule definition module, configured to generate multiple groups of rule libraries established based on the part-of-speech types of the first reference text, and each group of rule libraries includes multiple single rules with the same part of speech;
[0013] A rule matching module, configured to match the second reference text with the single rules in the rule library to obtain a matching degree;
[0014] An evaluation module, configured to obtain the similarity data in the rule matching module, generate an evaluation value based on the weights of the single rules, and determine whether the item to be evaluated is compliant according to the evaluation value;
[0015] A server, configured to store the changes of the first reference text and the second description text in the photovoltaic project policy according to a time period, and at the same time configured to store the stage data logs of the item to be evaluated after the corresponding configured items are executed by the second acquisition module, the text preprocessing module, and the rule matching module.
[0016] As a preferred method, when the rule definition module executes its configured item, it also executes the following configuration:
[0017] Establish an order sequence and an association index in the rule library,
[0018] Each of the single rules is composed of a first reference text and a first description text, and a position code is generated through the text positions of the first reference texts in the photovoltaic project policy between adjacent single rules, and the entire position codes of the rule library constitute the order sequence;
[0019] Locate the bytes at preset adjacent positions in the photovoltaic project policy data according to the position code, and establish an association index if there is a logical relationship between the adjacent first reference texts in the photovoltaic project policy data;
[0020] Write the position codes of the order sequence into the association index and generate an association index set.
[0021] As a preferred manner, the matching degrees set in the rule matching module include:
[0022] The first matching degree is set as the matching ratio of the second reference text matching to the first reference text;
[0023] The second matching degree is set as the ratio of the matched second reference text and its associated second description text matching to the corresponding first description text;
[0024] The third matching degree is set as the matching degree between the association index set generated between the second reference texts and the association index set generated between the first reference texts.
[0025] As a preferred manner, before the text preprocessing module and the text processing module execute the processing, they are further configured to train the natural language processing algorithm, specifically including:
[0026] Obtain the public project policy text before the occurrence of the project to be evaluated, and extract the public project policy text as the first training set;
[0027] Obtain the public evaluation text before the occurrence of the project to be evaluated, and extract the public evaluation text as the second training set;
[0028] Set the minimum truncation unit and the maximum truncation unit;
[0029] Perform word segmentation starting from the minimum truncation unit to obtain the initial word segmentation training set;
[0030] Expand the minimum truncation unit for word segmentation according to a fixed learning rate until the maximum truncation unit to obtain the intermediate word segmentation training sets at the times of each expansion;
[0031] Extract the first feature regarding the first matching degree, and the first feature is based on the part-of-speech distribution and represents the similarity in the grammatical structure between the policy text and the project to be evaluated;
[0032] Extract the second feature regarding the second matching degree, and the second feature is based on the semantic vector and represents the similarity of the description text;
[0033] Extract the third feature regarding the third matching degree, and the third feature is based on the association index and represents the similarity in logic between the policy text and the project to be evaluated;
[0034] Obtain the truncation unit and the description text logic unit with the maximum average similarity regarding the first feature, the second feature, and the third feature as the parameters for configuring the text preprocessing module and the text extraction module.
[0035] As a preferred manner, it further includes a correction module, which is configured to obtain and store the corrected text after the evaluation of the item to be evaluated, and perform a preset adjustment on the first reference text and the first description text according to the corrected text until the evaluation similarity between the evaluation result of the evaluation module and the corrected text is within a preset threshold range.
[0036] As a preferred manner, the second acquisition module includes:
[0037] A text acquisition sub-module, configured to extract the second text data from the text of the item to be evaluated;
[0038] An audio acquisition sub-module, configured to convert the audio into text and supplement it to the second text data;
[0039] A graphic acquisition sub-module, configured to supplement the text in the graphic to the second text data through ocr recognition, and also configured to supplement the image to the second text data after matching it with the text library.
[0040] As a preferred manner, it further includes an output module, configured to receive the evaluation value of the evaluation module, and output the evaluation text and risk text of the item to be evaluated according to the historical evaluation situation.
[0041] In the second aspect of the present invention, a compliance management method for a photovoltaic project is provided, including the following steps:
[0042] S1. Obtain the first text data corresponding to the photovoltaic project policy from at least one data source;
[0043] S2. Receive the photovoltaic project to be evaluated, and extract the second text data therefrom. The extraction process includes:
[0044] Extract text information from the text part of the item to be evaluated to form preliminary text data;
[0045] Convert the audio content into text form and incorporate it into the preliminary text data;
[0046] Use optical character recognition technology to recognize the text contained in the graphics in the project, and after matching with the preset text library, supplement the corresponding text to the preliminary text data to finally obtain the complete second text data;
[0047] S3. Perform word segmentation processing on the obtained first text data and second text data respectively to obtain first word segmentation data and second word segmentation data;
[0048] Determine the first reference text and the first description text from the first word segmentation data, and determine the second reference text and the second description text from the second word segmentation data, where:
[0049] The first reference text is a text unit with a dynamically adjusted value of 50 - 200 characters, which is automatically determined according to the paragraph structure of the policy text;
[0050] The first description text is a text unit used to describe the first reference text, and each first description text covers at least two semantically complete units within the range of 10 - 50 characters;
[0051] S4. Construct multiple groups of rule bases according to the part-of-speech types of the first reference text, and each group of rule bases accommodates multiple single rules with consistent parts of speech;
[0052] Generate an order sequence and an associated index, where:
[0053] Each single rule is composed of the corresponding first reference text and the first description text;
[0054] Generate position codes based on the positions of the first reference texts in adjacent single rules in the photovoltaic project policy text, and these position codes form an order sequence;
[0055] Locate the adjacent position bytes of the photovoltaic project policy data according to the position codes. If there is a logical relationship, establish an associated index, and write the position codes of the order sequence into the associated index to generate an associated index set;
[0056] S5. Match the second reference text with each single rule in the rule base one by one to obtain the matching degrees including the following:
[0057] The first matching degree of the second reference text matching the first reference text;
[0058] The second matching degree of the matched second reference text and its associated second description text matching the corresponding first description text;
[0059] The third matching degree of the associated index set generated among the second reference texts and the associated index set generated among the first reference texts;
[0060] S6. Calculate an evaluation value based on the weights of each single rule, and determine whether the photovoltaic project to be evaluated is compliant according to this evaluation value;
[0061] When the similarity between the evaluation value and the calibration text exceeds the preset threshold, adjust the weight coefficient of the first reference text.
[0062] As a preferred method, calculating the evaluation value based on the weights of each single rule specifically includes:
[0063] Extract the part-of-speech distribution of the second reference text as the first feature, calculate its first matching degree with the first reference text, and the weight proportion is A;
[0064] Extract the semantic vector of the second description text as the second feature, calculate its second matching degree with the first description text, and the weight ratio is B;
[0065] Extract the logical relationship of the associated index set between the second reference texts as the third feature, calculate its third matching degree with the associated index set between the first reference texts, and the weight ratio is C;
[0066] Calculate the evaluation value according to the following formula:
[0067] Evaluation value = A × First matching degree + B × Second matching degree + C × Third matching degree
[0068] Determine whether the photovoltaic project to be evaluated is compliant based on the evaluation value, where:
[0069] If the evaluation value ≥ 0.8, it is determined to be compliant;
[0070] If 0.6 ≤ evaluation value < 0.8, it is determined to be partially compliant and supplementary materials are required;
[0071] If the evaluation value < 0.6, it is determined to be non-compliant;
[0072] When the similarity between the evaluation value and the corrected text exceeds the preset threshold, adjust the weight coefficient of the first reference text.
[0073] The above technical solution of the present invention has the following advantages compared with the prior art:
[0074] With the help of a series of modules such as text preprocessing, extraction, and rule matching, the system can automatically process a large amount of complex photovoltaic project policy texts and texts of projects to be evaluated. It performs accurate word segmentation, scientifically constructs a rule library, and conducts evaluations based on multi-dimensional matching degrees, greatly reducing the workload of manual policy interpretation and project review, significantly speeding up the originally cumbersome and time-consuming compliance evaluation process, improving the overall evaluation efficiency, and ensuring that projects can quickly respond to policy requirements.
[0075] The present invention determines the reference text and the description text according to part-of-speech classification, and is equipped with a well-designed set of rule libraries. The rule matching module calculates the matching degree from multiple angles, comprehensively considering the matching situations at the text reference, description content, and logical association levels. Each single rule also has a reasonable weight, making the evaluation value more scientific, effectively reducing the risk of misjudgment, and accurately judging whether the project is compliant.
[0076] Through the correction module, the present invention uses the corrected text after evaluation to reversely optimize the reference text and the description text of the system, making the evaluation results gradually approach accuracy, the system performance gradually improve with actual applications, and the evaluation system become more mature and perfect, meeting the requirements of real business scenarios.
[0077] The present invention also trains natural language processing algorithms before text preprocessing to find the optimal word segmentation parameters, making the algorithms fit the characteristics of photovoltaic project texts, so that the subsequent text processing and extraction processes are smoother and more accurate, strengthening the overall intelligent processing ability of the system and continuously empowering efficient and accurate evaluation. BRIEF DESCRIPTION OF THE DRAWINGS
[0078] Figure 1 is a structural block diagram of a compliance management system for a photovoltaic project provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0079] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0080] In the first aspect of the embodiments of the present disclosure, a compliance management system for a photovoltaic project is provided, as Figure 1 shown, including:
[0081] A first acquisition module configured to acquire first text data of photovoltaic project policies from at least one data source;
[0082] A second acquisition module configured to receive a project to be evaluated and acquire second text data from the project to be evaluated;
[0083] A text preprocessing module configured to perform word segmentation on the first text data and the second text data using a preset natural language processing algorithm to obtain first segmented data and second segmented data;
[0084] A text extraction module configured to determine a first reference text and a first description text from the first segmented data, and at the same time configured to determine a second reference text and a second description text of the second segmented data; the reference text is a text including a first preset number of characters classified according to part of speech, and the description text is a descriptive vocabulary of the reference text, and each description text includes at least two groups of texts within the range from the first preset number of characters to the second preset number of characters;
[0085] A rule definition module configured to generate multiple groups of rule libraries established based on the part-of-speech types of the first reference text, and each group of rule libraries includes multiple single rules with the same part of speech;
[0086] A rule matching module configured to match the second reference text with the single rules in the rule library to obtain a matching degree;
[0087] An evaluation module, configured to obtain the similarity data in the rule matching module, generate an evaluation value based on the weights of each single rule, and determine whether the item to be evaluated complies with the regulations according to the evaluation value;
[0088] A server, configured to store the changes of the first reference text and the second description text in the photovoltaic project policy in a time period, and is also configured to store the stage data logs of the item to be evaluated after the corresponding configured items are executed by the second acquisition module, the text preprocessing module, and the rule matching module.
[0089] As a preferred method, when the rule definition module executes its configured item, the following configurations are also executed:
[0090] Establish an order sequence and an association index in the rule library,
[0091] Each single rule is composed of a first reference text and a first description text. A position code is generated by the text positions of the first reference texts in the photovoltaic project policy between adjacent single rules, and the entire position codes of the rule library constitute the order sequence;
[0092] Locate the bytes at preset adjacent positions in the photovoltaic project policy data according to the position code. If there is a logical relationship between adjacent first reference texts in the photovoltaic project policy data, establish their association index;
[0093] Write the position codes of the order sequence into the association index and generate an association index set.
[0094] As a preferred method, the matching degrees set in the rule matching module include:
[0095] The first matching degree, set as the matching ratio of the second reference text matching to the first reference text;
[0096] The second matching degree, set as the ratio of the matched second reference text and its associated second description text matching to the corresponding first description text;
[0097] The third matching degree, set as the matching degree between the association index set generated between the second reference texts and the association index set generated between the first reference texts.
[0098] As a preferred method, before the text preprocessing module and the text processing module execute the processing, they are also configured to train the natural language processing algorithm, specifically including:
[0099] Obtain the public project policy text before the item to be evaluated occurs, and extract the public project policy text as the first training set;
[0100] Obtain the publicly available evaluation text before the project to be evaluated occurs, and extract the publicly available evaluation text as the second training set;
[0101] Set the minimum truncation unit and the maximum truncation unit;
[0102] Perform word segmentation starting from the minimum truncation unit to obtain the initial word segmentation training set;
[0103] Expand the minimum truncation unit for word segmentation at a fixed learning rate until the maximum truncation unit to obtain the intermediate word segmentation training sets when each is expanded;
[0104] Extract the first feature regarding the first matching degree, where the first feature is based on the part-of-speech distribution and represents the similarity in grammatical structure between the policy text and the project to be evaluated;
[0105] Extract the second feature regarding the second matching degree, where the second feature is based on the semantic vector and represents the similarity of the descriptive text;
[0106] Extract the third feature regarding the third matching degree, where the third feature is based on the association index and represents the logical similarity between the policy text and the project to be evaluated;
[0107] Obtain the truncation unit with the maximum average similarity regarding the first feature, the second feature, and the third feature, and the descriptive text logical unit as the parameters for configuring the text preprocessing module and the text extraction module.
[0108] As a preferred method, it further includes a correction module, which is configured to obtain and store the corrected text after the evaluation of the project to be evaluated, and perform a preset adjustment to the first reference text and the first descriptive text according to the corrected text until the evaluation result of the evaluation module and the evaluation similarity of the corrected text are within a preset threshold range. The preset adjustment provided in the embodiments of the present disclosure is to modify the character size of the first reference text and the first descriptive text, and the weight coefficients when the first feature, the second feature, and the third feature are executed.
[0109] As a preferred method, the second acquisition module includes:
[0110] A text acquisition sub-module, configured to extract the second text data from the text of the project to be evaluated;
[0111] An audio acquisition sub-module, configured to convert the audio into text and supplement it to the second text data;
[0112] A graphic acquisition sub-module, configured to supplement the text in the graphic to the second text data through ocr recognition, and is also configured to supplement the image after matching with the text library to the second text data.
[0113] As a preferred method, it further includes an output module, which is configured to receive the evaluation value of the evaluation module and output an evaluation text and a risk text of the item to be evaluated according to the historical evaluation situation.
[0114] In the second aspect of the embodiments of the present disclosure, a compliance management method for a photovoltaic project is provided, including the following steps:
[0115] S1. Obtain first text data corresponding to the photovoltaic project policy from at least one data source;
[0116] S2. Receive the photovoltaic project to be evaluated and extract second text data therefrom. The extraction process includes:
[0117] Extract text information from the text part of the item to be evaluated to form preliminary text data;
[0118] Convert the audio content into text form and incorporate it into the preliminary text data;
[0119] Use optical character recognition technology to recognize the text contained in the graphics in the project, match it with a preset text library, and supplement the corresponding text to the preliminary text data to finally obtain complete second text data;
[0120] S3. Perform word segmentation processing on the obtained first text data and second text data respectively to obtain first word segmentation data and second word segmentation data;
[0121] Determine a first reference text and a first description text from the first word segmentation data, and determine a second reference text and a second description text from the second word segmentation data, where:
[0122] The first reference text is a text unit with a dynamically adjustable value of 50-200 characters, which is automatically determined according to the paragraph structure of the policy text;
[0123] The first description text is a text unit used to describe the first reference text, and each first description text covers at least two semantically complete units within the range of 10-50 characters;
[0124] S4. Construct multiple groups of rule libraries according to the part-of-speech types of the first reference text, and each group of rule libraries contains multiple single rules with the same part of speech;
[0125] Generate a sequence sequence and an association index, where:
[0126] Each single rule is composed of the corresponding first reference text and first description text;
[0127] Generate a position code according to the position of the first reference text in the adjacent single rules in the photovoltaic project policy text, and these position codes form a sequence sequence;
[0128] Locate the bytes at adjacent positions of the photovoltaic project policy data according to the position encoding. If there is a logical relationship, establish an associated index, and write the position encoding of the sequential sequence into the associated index to generate an associated index set;
[0129] S5. Match the second reference text with each single rule in the rule library to obtain the matching degrees including the following contents:
[0130] The first matching degree of the second reference text matching the first reference text;
[0131] The second matching degree of the matched second reference text and its associated second description text matching the corresponding first description text;
[0132] The third matching degree of the associated index set generated between the second reference texts and the associated index set generated between the first reference texts;
[0133] S6. Calculate the evaluation value based on the weights of each single rule, and determine whether the photovoltaic project to be evaluated complies with the regulations according to the evaluation value;
[0134] When the similarity between the evaluation value and the calibration text exceeds the preset threshold, perform the adjustment of the weight coefficient and the character count of the first reference text and the first description text.
[0135] As a preferred method, calculating the evaluation value based on the weights of each single rule specifically includes:
[0136] Extract the part-of-speech distribution of the second reference text as the first feature, calculate its first matching degree with the first reference text, and the weight ratio is A;
[0137] Extract the semantic vector of the second description text as the second feature, calculate its second matching degree with the first description text, and the weight ratio is B;
[0138] Extract the logical relationship of the associated index set between the second reference texts as the third feature, calculate its third matching degree with the associated index set between the first reference texts, and the weight ratio is C;
[0139] Calculate the evaluation value according to the following formula:
[0140] Evaluation value = A × First matching degree + B × Second matching degree + C × Third matching degree
[0141] Determine whether the photovoltaic project to be evaluated complies with the regulations according to the evaluation value, where:
[0142] Evaluation value ≥ 0.8, it is determined to be compliant;
[0143] 0.6 ≤ Evaluation value < 0.8, it is determined to be partially compliant and supplementary materials are required;
[0144] The evaluation value < 0.6 is determined as non - compliant;
[0145] When the similarity between the evaluation value and the corrected text exceeds the preset threshold, perform the adjustment of the weight coefficients and the number of characters of the first reference text and the first description text.
[0146] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of the apparatus, method, and computer program product according to embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code that contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions marked in the block may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. In the descriptions corresponding to the flowcharts and block diagrams in the accompanying drawings, the operations or steps corresponding to different blocks may also occur in a different order than disclosed in the description, and sometimes there is no specific order between different operations or steps. For example, two consecutive operations or steps may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. Each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware device that performs the specified function or action, or can be implemented by a combination of dedicated hardware and computer instructions.
Claims
1. A compliance management system for photovoltaic projects, characterized in that: include: A first acquisition module is configured to acquire first text data of a photovoltaic project policy from at least one data source; A second acquisition module is configured to receive the item to be evaluated and acquire second text data from the item to be evaluated; A text preprocessing module is configured to use a preset natural language processing algorithm to perform word segmentation processing on the first text data and the second text data to obtain first word segmentation data and second word segmentation data; a text extraction module configured to determine a first reference text and a first description text from the first word segmentation data, and configured to determine a second reference text and a second description text of the second word segmentation data; the reference text is a text including a first preset number of characters classified according to part of speech, the description text is a description vocabulary of the reference text, and each description text includes at least two groups of texts ranging from the first preset number of characters to the second preset number of characters; A rule definition module is configured to generate a plurality of rule bases established based on the part-of-speech type of the first reference text, each rule base comprising a plurality of single rules with the same part-of-speech; A rule matching module is configured to match the second reference text with a single rule in the rule base to obtain a matching degree; An evaluation module is configured to obtain the similarity data in the rule matching module, generate an evaluation value based on the weight of each single rule, and determine whether the item to be evaluated is compliant according to the evaluation value; The server is configured to store changes in the first benchmark text and the second description text in the photovoltaic project policy according to a time period, and is also configured to store the stage data log of the project to be evaluated after the second acquisition module, the text preprocessing module, and the rule matching module execute the corresponding configured items.
2. A photovoltaic project compliance management system according to claim 1, characterized in that: When the rule definition module executes its configuration items, it also executes the following configurations: Establishing sequential sequences and associative indexes in the rule base, Each of the single rules is composed of a first reference text and a first description text, and a position code is generated between adjacent single rules by the text position of each of the first reference texts in the photovoltaic project policy, and all the position codes of the rule base constitute the sequential sequence; Locate bytes at preset adjacent positions in the photovoltaic project policy data according to the position code, and establish an associated index if there is a logical relationship between the adjacent first reference texts of the photovoltaic project policy data; The position codes of the sequential sequences are written into the associated indexes, and an associated index set is generated.
3. A photovoltaic project compliance management system according to claim 2, characterized in that: The matching degree set in the rule matching module includes: A first matching degree, which is set as a matching ratio of the second reference text to the first reference text; A second matching degree is set as a ratio of the matched second reference text and its associated second description text to the first description text; The third matching degree is set as the matching degree between the associated index set generated between the second reference texts and the associated index set generated between the first reference texts.
4. A photovoltaic project compliance management system according to claim 3, characterized in that: Before the text preprocessing module and the text processing module perform processing, they are also configured to train the natural language processing algorithm, specifically including: Obtaining a public project policy text before the project to be evaluated occurs, and extracting the public project policy text as a first training set; Obtaining a public evaluation text before the project to be evaluated occurs, and extracting the public evaluation text as a second training set; Set the minimum word truncation unit and the maximum word truncation unit; Segment words from the smallest word segmentation unit to obtain the initial word segmentation training set; Expand the minimum word segmentation unit according to a fixed learning rate to perform word segmentation until the maximum word segmentation unit is reached, and obtain the intermediate word segmentation training set for each expanded word segmentation; Extracting a first feature about the first matching degree, wherein the first feature is based on part-of-speech distribution and represents the similarity between the policy text and the item to be evaluated in grammatical structure; Extracting a second feature related to the second matching degree, wherein the second feature is based on a semantic vector and represents the similarity of the description text; Extracting a third feature about the third matching degree, wherein the third feature is based on the association index and represents the logical similarity between the policy text and the item to be evaluated; The word truncation unit and the description text logic unit with the largest average similarity with respect to the first feature, the second feature and the third feature are obtained as parameters of configuration items for executing the text preprocessing module and the text extraction module.
5. A photovoltaic project compliance management system according to claim 1, characterized in that: It also includes a correction module, which is configured to obtain and store a corrected text after the evaluation of the item to be evaluated is completed, and to perform preset adjustments on the first benchmark text and the first description text according to the corrected text until the evaluation result of the evaluation module and the evaluation similarity of the corrected text are within a preset threshold range.
6. A photovoltaic project compliance management system according to claim 4, characterized in that: The second acquisition module includes: A text acquisition submodule, configured to extract the second text data from the text of the item to be evaluated; An audio acquisition submodule, configured to convert the audio into text and then add it to the second text data; The graphic acquisition submodule is configured to supplement the text in the graphic into the second text data through OCR recognition, and is also configured to supplement the image into the second text data after matching with the text library.
7. A photovoltaic project compliance management system according to claim 1, characterized in that: It also includes an output module, which is configured to receive the evaluation value of the evaluation module and output the evaluation text and risk text of the project to be evaluated according to the historical evaluation situation.
8. A compliance management method for a photovoltaic project, characterized in that: The steps include: S1. Acquire first text data corresponding to a photovoltaic project policy from at least one data source; S2. Receive a photovoltaic project to be evaluated, and extract second text data therefrom, wherein the extraction process includes: Extracting text information from the text portion of the item to be evaluated to form preliminary text data; Converting audio content into text form and incorporating it into preliminary text data; Using optical character recognition technology to identify the text contained in the graphics in the project, and after matching with the preset text library, the corresponding text is added to the preliminary text data to finally obtain complete second text data; S3, performing word segmentation processing on the acquired first text data and the acquired second text data respectively to obtain first word segmentation data and second word segmentation data; A first reference text and a first description text are determined from the first word segmentation data, and a second reference text and a second description text are determined from the second word segmentation data, wherein: The first reference text is a text unit having a dynamically adjusted value of 50-200 characters, which is automatically determined according to the paragraph structure of the policy text; The first description text is a text unit used to describe the first reference text, and each first description text includes at least two groups of semantically complete units within a range of 10-50 characters; S4, constructing multiple groups of rule bases according to the part-of-speech type of the first reference text, each group of rule bases containing multiple single rules with the same part-of-speech; Generates a sequential sequence and associated index where: Each single rule is composed of a corresponding first reference text and a first description text; Generate position codes according to the position of the first benchmark text in the adjacent single rule in the photovoltaic project policy text, and these position codes form a sequential sequence; Locate the adjacent bytes of the photovoltaic project policy data according to the position code, establish an associated index if there is a logical relationship, and write the position code of the sequential sequence into the associated index to generate an associated index set; S5. Match the second reference text with the single rules in the rule base one by one to obtain the matching degree including the following contents: The second reference text matches the first reference text to a first degree of matching; The matched second reference text and its associated second description text match the corresponding first description text to a second matching degree; a third matching degree between the associated index set generated between the second reference texts and the associated index set generated between the first reference texts; S6. Calculate the evaluation value based on the weight of each single rule, and determine whether the photovoltaic project to be evaluated is compliant based on the evaluation value; When the similarity between the evaluation value and the corrected text exceeds a preset threshold, the weight coefficient of the first reference text is adjusted.
9. The compliance management method for photovoltaic projects according to claim 8, characterized in that: The evaluation value is calculated based on the weight of each single rule, including: Extract the part-of-speech distribution of the second benchmark text as the first feature, and calculate the first matching degree between the second benchmark text and the first benchmark text, with a weight of A. Extract the semantic vector of the second description text as the second feature, and calculate the second matching degree between the semantic vector and the first description text, with a weight ratio of B; Extract the logical relationship of the second reference text association index set as the third feature, and calculate the third matching degree between the second reference text association index set and the first reference text association index set, with a weight ratio of C; The evaluation value is calculated according to the following formula: Evaluation value = A × first matching degree + B × second matching degree + C × third matching degree The compliance of the photovoltaic project to be assessed is determined based on the assessment value, including: The evaluation value is ≥0.8, which is considered compliant; 0.6≤assessment value<0.8, it is judged as partially compliant and requires additional materials; If the evaluation value is <0.6, it is considered non-compliant; When the similarity between the evaluation value and the correction text exceeds a preset threshold, the weight coefficient adjustment and the character number adjustment of the first reference text and the first description text are performed.