A method and apparatus for assessing social stability risks of major engineering projects
By generating semantic vectors for engineering projects and using the autoregressive processing of the large language model GLM4, the inefficiency of traditional risk assessment methods is solved, enabling rapid and accurate social stability risk assessment for major engineering projects, and adapting to the ever-changing engineering project environment.
Patent Information
- Application Number
- CN202510897824.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-01
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2045-07-01
AI Technical Summary
Traditional risk assessment methods in major engineering projects suffer from high reliance on experts, long assessment cycles, low efficiency, lack of real-time and dynamic adjustment capabilities, resulting in insufficient accuracy and timeliness of assessments.
A social stability risk assessment method is adopted. By acquiring project information, risk indicator weights, and scoring questions, query vectors are generated. These vectors are then converted into semantic vectors using the BERT text embedding model. Finally, the GLM4 large language model is used for knowledge retrieval and autoregressive generation to output risk indicator weights and scoring results, thus constructing an automated assessment closed loop.
It automates risk assessment, shortens the assessment cycle, improves assessment efficiency, reduces human interference and subjective bias, provides accurate assessment based on deep semantic understanding, and is adaptable to multi-dimensional risk assessment of complex engineering projects.
Smart Images

Figure CN120410229B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of risk assessment technology, and in particular to a method for assessing social stability risks of major engineering projects, a device for assessing social stability risks of major engineering projects, an electronic device, and a computer-readable medium. Background Technology
[0002] Traditional risk assessment methods are widely used in large-scale infrastructure and major engineering projects. These methods primarily rely on expert subjective judgment to comprehensively analyze factors such as the project's social impact, environmental impact, and economic benefits. While traditional methods have certain advantages, their limitations are particularly evident in complex and ever-changing engineering projects: on the one hand, they suffer from strong reliance on experts, long assessment cycles, and low efficiency; on the other hand, they lack real-time and dynamic adjustment capabilities. These problems severely restrict the accuracy and timeliness of risk assessments. Summary of the Invention
[0003] In view of the above problems, the present invention is proposed to provide a method for assessing the social stability risks of major engineering projects, a corresponding device for assessing the social stability risks of major engineering projects, an electronic device, and a computer-readable medium to overcome or at least partially solve the above problems.
[0004] This invention discloses a method for assessing social stability risks of major engineering projects, the method comprising:
[0005] Access to information on major engineering projects, questions regarding risk indicator weights, and questions regarding risk scoring;
[0006] A weighted query vector is generated based on information about major engineering projects and risk indicator weights, and a scoring query vector is generated based on information about major engineering projects and risk scores.
[0007] Based on the weighted query vector and the scoring query vector, knowledge retrieval is performed in the pre-set social stability risk vector database of major engineering projects to obtain the semantic vector of social stability risk of major engineering projects that matches the weighted query vector and the semantic vector of social stability risk of major engineering projects that matches the scoring query vector.
[0008] Based on the weighted query vector and the semantic vector of social stability risk of major engineering projects that matches the weighted query vector, a weighted query hint is generated, and the risk indicator weight result of major engineering projects is generated based on the weighted query hint.
[0009] Based on the scoring query vector, the semantic vector of social stability risk of major engineering projects that matches the scoring query vector, and the risk indicator weight results, a scoring query prompt is generated, and the risk scoring result of the major engineering project is generated based on the scoring query prompt.
[0010] Optionally, a weighted query vector is generated based on information about major engineering projects and risk indicator weights, and a scoring query vector is generated based on information about major engineering projects and risk scores, including:
[0011] The BERT text embedding model is used to convert questions about major engineering projects and risk indicator weights into weighted query vectors, and questions about major engineering projects and risk scores into score query vectors.
[0012] Optionally, based on the weighted query vector and the scoring query vector, a knowledge retrieval is performed in a pre-defined database of social stability risks for major engineering projects to obtain semantic vectors of social stability risks for major engineering projects that match the weighted query vectors and semantic vectors of social stability risks for major engineering projects that match the scoring query vectors, including:
[0013] Based on the weighted query vector, the candidate semantic vectors of social stability risk of major engineering projects are selected from the pre-set database of social stability risk vectors of major engineering projects by means of maximum inner product search. Then, by means of normalized inner product score and maximum likelihood estimation, a conditional probability distribution is constructed for the candidate semantic vectors of social stability risk of major engineering projects of the weighted query vector. The first pre-set number of candidate semantic vectors of social stability risk of major engineering projects corresponding to the maximum probability are selected as the semantic vectors of social stability risk of major engineering projects that match the weighted query vector.
[0014] Based on the scoring query vector, candidate major engineering project social stability risk semantic vectors are selected from the pre-defined database of major engineering project social stability risk vectors through maximum inner product search. Then, a conditional probability distribution is constructed for the candidate major engineering project social stability risk semantic vectors of the scoring query vector through normalized inner product score and maximum likelihood estimation. The top pre-defined number of candidate major engineering project social stability risk semantic vectors corresponding to the highest probability are selected as the major engineering project social stability risk semantic vectors that match the scoring query vector.
[0015] Optionally, a weight query hint is generated based on the weight query vector and the semantic vector of social stability risk of major engineering projects that matches the weight query vector. Based on the weight query hint, the risk indicator weight results for major engineering projects are generated, including:
[0016] The semantic vector of social stability risk of major engineering projects that matches the weight query vector is merged with the weight query vector to obtain the weight query hint. Then, an autoregressive generation is performed based on the weight query hint using a large language model. When generating each word, the probability distribution is calculated by combining the semantic vector of social stability risk of major engineering projects that matches the weight query vector and the generated context information. The risk index weight results of all major engineering projects are obtained through multiple iterations.
[0017] Optionally, a scoring query prompt is generated based on the scoring query vector, the semantic vector of social stability risk of major engineering projects matching the scoring query vector, and the risk indicator weight results. Based on the scoring query prompt, a risk scoring result for the major engineering projects is generated, including:
[0018] The semantic vector of social stability risk of major engineering projects that matches the scoring query vector is merged with the scoring query vector and the risk indicator weight results to obtain the scoring query prompt. Then, an autoregressive generation is performed based on the scoring query prompt using a large language model. When generating each word, the probability distribution is calculated by combining the semantic vector of social stability risk of major engineering projects that matches the scoring query vector and the generated context information. The risk scoring results of all major engineering projects are obtained through multiple iterations.
[0019] Optionally, the method further includes:
[0020] Collect multi-source textual data related to social stability risks of major engineering projects, including social stability risk assessment reports, historical major engineering project emergencies, policy and regulatory data, and news and social data;
[0021] Data clarity and preprocessing are performed on multi-source text data related to social stability risks of major engineering projects to obtain a dataset of social stability risks of major engineering projects.
[0022] For each major historical engineering project emergency, a multi-dimensional scenario space decomposition is performed to obtain multiple scenarios for each event and information for each scenario. The multiple scenarios for each event and information for each scenario are then converted into semantic vectors and stored in the major engineering project social stability risk vector database. The information for each scenario includes scenario objects and scenario elements.
[0023] Social stability risk assessment reports, policy and regulatory data, news and social data are converted into semantic vectors and stored in the social stability risk vector database for major engineering projects.
[0024] Optionally, data cleaning and preprocessing are performed on multi-source text data related to social stability risks of major engineering projects to obtain a dataset on social stability risks of major engineering projects, including:
[0025] Based on the screening criteria, entries that are irrelevant to or incomplete in multi-source text data related to the social stability risks of major engineering projects were removed; the screening criteria were based on the social impact of the project, the scale of mass incidents, and its relevance to the risk assessment objectives.
[0026] The selected multi-source texts related to social stability risks of major engineering projects are deduplicated and standardized to obtain standard multi-source texts related to social stability risks of major engineering projects.
[0027] The multi-source texts related to the social stability risks of major engineering projects are classified according to preset labels, and relevant metadata is attached to each data entry to obtain the social stability risk dataset of major engineering projects.
[0028] This invention also discloses a device for assessing social stability risks of major engineering projects, the device comprising:
[0029] The project information and question acquisition module is used to obtain information on major engineering projects, ask questions about risk indicator weights, and ask questions about risk scoring.
[0030] The project information and question vectorization module is used to generate weighted query vectors based on major engineering project information and risk indicator weights, and to generate score query vectors based on major engineering project information and risk score questions.
[0031] The knowledge retrieval module is used to perform knowledge retrieval in a preset database of social stability risk vectors for major engineering projects based on weighted query vectors and scoring query vectors, and to obtain semantic vectors of social stability risks for major engineering projects that match the weighted query vectors and semantic vectors of social stability risks for major engineering projects that match the scoring query vectors.
[0032] The risk indicator weight result generation module is used to generate weight query prompts based on the weight query vector and the semantic vector of social stability risk of major engineering projects that matches the weight query vector, and to generate risk indicator weight results of major engineering projects based on the weight query prompts.
[0033] The risk scoring result generation module is used to generate scoring query prompts based on the scoring query vector, the semantic vector of social stability risk of major engineering projects that matches the scoring query vector, and the risk indicator weight results, and to generate risk scoring results for major engineering projects based on the scoring query prompts.
[0034] The present invention also discloses an electronic device, including a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus;
[0035] The memory is used to store computer programs;
[0036] When the processor executes the program stored in the memory, it implements the social stability risk assessment method for major engineering projects as described in this invention.
[0037] The present invention also discloses one or more computer-readable media having instructions stored thereon that, when executed by one or more processors, cause the processors to perform the social stability risk assessment method for major engineering projects as described in the present invention.
[0038] This invention has the following advantages:
[0039] The present invention provides a method for assessing social stability risks of major engineering projects. First, it acquires project information, weights, and scoring questions, generating corresponding weight and scoring query vectors. Through semantic matching in a pre-set risk vector database, it retrieves similar risk vectors. Then, it uses the weight query vectors and matching results to generate prompts and outputs risk indicator weights. Finally, it combines the scoring query vectors, matching results, and weight values to generate scoring prompts and outputs a risk score. This method significantly shortens the assessment cycle and improves efficiency through an automated risk assessment process. Furthermore, by combining historical data, contextual information, and the large language model GLM4, it achieves accurate assessment based on deep semantic understanding, reducing human interference and subjective bias. Attached Figure Description
[0040] Figure 1 This is a flowchart illustrating the steps of a method for assessing social stability risks in major engineering projects, as provided in an embodiment of the present invention.
[0041] Figure 2 This is a flowchart for assessing the social stability risks of major engineering projects provided in an embodiment of the present invention;
[0042] Figure 3 This is a structural block diagram of a social stability risk assessment device for major engineering projects provided in an embodiment of the present invention;
[0043] Figure 4 This is a block diagram of an electronic device provided in an embodiment of the present invention;
[0044] Figure 5 This is a schematic diagram of a computer-readable medium provided in an embodiment of the present invention. Detailed Implementation
[0045] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0046] Reference Figure 1The diagram illustrates a flowchart of a method for assessing the social stability risks of major engineering projects according to an embodiment of the present invention, which may specifically include the following steps:
[0047] Step 101: Obtain information on major engineering projects, ask questions about risk indicator weights, and ask questions about risk scoring;
[0048] Step 102: Generate a weight query vector based on information about major engineering projects and risk indicator weights, and generate a score query vector based on information about major engineering projects and risk scores.
[0049] Step 103: Based on the weight query vector and the score query vector, perform knowledge retrieval in the preset major engineering project social stability risk vector database to obtain the major engineering project social stability risk semantic vector that matches the weight query vector and the major engineering project social stability risk semantic vector that matches the score query vector.
[0050] Step 104: Generate weight query hints based on the weight query vector and the semantic vector of social stability risk of major engineering projects that matches the weight query vector, and generate risk indicator weight results for major engineering projects based on the weight query hints;
[0051] Step 105: Generate a scoring query prompt based on the scoring query vector, the semantic vector of social stability risk of major engineering projects that matches the scoring query vector, and the risk indicator weight results, and generate the risk scoring results of major engineering projects based on the scoring query prompt.
[0052] In this invention, the system first acquires basic project information, risk indicator weights, and scoring questions, generating weight query vectors and scoring query vectors respectively. Then, it performs similarity matching in a pre-set risk vector database to retrieve the semantic vector closest to the query vector. Next, the system uses the weight query vector and its matching results to construct prompts and automatically generates the weight allocation for each risk indicator. Finally, it combines the scoring query vector, matching results, and weight data to construct scoring prompts, ultimately outputting the project's comprehensive risk score. The entire assessment process employs vectorized retrieval technology for knowledge matching and uses an intelligent prompt generation mechanism to complete weight calculation and risk quantification, forming a complete closed loop for automatic risk assessment.
[0053] Reference Figure 2 The specific steps are as follows:
[0054] Input: Social stability risk dataset for major engineering projects User input Among them, project case P, project case weight question Project case evaluation questions
[0055] Output: Project case weight results Project Case Evaluation Results
[0056] 1: Dataset Data and through embedded models Transform text data into semantic vectors
[0057] 2: The generated semantic vector Stored in a vector database middle
[0058] 3: User Input
[0059] 4: Through embedding models Generate the corresponding weighted query vector And rating query vector
[0060] 5: According to and exist Perform knowledge retrieval to obtain relevant knowledge. and
[0061] 6: = , Weighted query hints
[0062] 7: {Weight query hints provided to GLM4 to generate project case weight results}
[0063] 8: Generate rating query prompts
[0064] 9: {Rating query prompts are provided for GLM4 to generate project case rating results}
[0065] As an example, a question about risk indicator weights could be, "Based on the social stability risk assessment indicator system and indicator definitions for major engineering projects, assess the social stability risk assessment system and indicator weights for this major engineering project case." Similarly, as an example, a question about risk scoring could be, "Based on the social stability risk assessment indicator system and indicator definitions for major engineering projects, assess the social stability risk score for this major engineering project case." Specifically:
[0066] Instruction: Summarize the following text to help answer the question, rather than directly answering the question itself. If the text is irrelevant, reply "Not applicable". At the end of your answer, provide a relevance score from 0 to 1 for each indicator weight, the sum of all indicator weight scores must be 1, and explain the scores.
[0067] Query: Based on the following major engineering project case study, please use the social stability risk assessment indicator system for major engineering projects and the corresponding indicator definitions to evaluate the social stability risk assessment system and the weights of each indicator in this case study. The weight and description of each indicator should be clear and specific.
[0068] Context: The relevant risk assessment is as follows: {}
[0069] The social stability risk assessment indicator system and indicator definitions are as follows: { Index_Context}
[0070] Input: Major engineering projects: {}
[0071] Answer: Event: {Event_Name}
[0072] Risk type and weight score:
[0073] Project decision-making risks
[0074] Project approval process: {weight_1}
[0075] Site selection planning: {weight_2}
[0076] …
[0077] Case Analysis and Scoring:
[0078] Project decision-making risks
[0079] - Project approval process: {score_1} ({description_1})
[0080] - Site planning and selection: {score_2} ({description_2})
[0081] …
[0082] Template 1: The template section for weight issues includes the template's Instruction, Query, Context, Input, and Answer, which respectively represent the template's instruction (prompting the model to perform specific tasks), query (prompting the model to query questions), input (specific information that the model needs to process), and answer (defining the expected type or format of the model's output to generate inference text). {*} are placeholders for the corresponding descriptive content.
[0083] In an optional embodiment of the present invention, a weight query vector is generated based on information about major engineering projects and risk indicator weights, and a score query vector is generated based on information about major engineering projects and risk scoring, including:
[0084] The BERT text embedding model is used to convert questions about major engineering projects and risk indicator weights into weighted query vectors, and questions about major engineering projects and risk scores into score query vectors.
[0085] In an optional embodiment of the present invention, knowledge retrieval is performed in a preset major engineering project social stability risk vector database based on the weight query vector and the scoring query vector to obtain a major engineering project social stability risk semantic vector matching the weight query vector and a major engineering project social stability risk semantic vector matching the scoring query vector, including:
[0086] Based on the weighted query vector, the candidate semantic vectors of social stability risk of major engineering projects are selected from the pre-set database of social stability risk vectors of major engineering projects by means of maximum inner product search. Then, by means of normalized inner product score and maximum likelihood estimation, a conditional probability distribution is constructed for the candidate semantic vectors of social stability risk of major engineering projects of the weighted query vector. The first pre-set number of candidate semantic vectors of social stability risk of major engineering projects corresponding to the maximum probability are selected as the semantic vectors of social stability risk of major engineering projects that match the weighted query vector.
[0087] Based on the scoring query vector, candidate major engineering project social stability risk semantic vectors are selected from the pre-defined database of major engineering project social stability risk vectors through maximum inner product search. Then, a conditional probability distribution is constructed for the candidate major engineering project social stability risk semantic vectors of the scoring query vector through normalized inner product score and maximum likelihood estimation. The top pre-defined number of candidate major engineering project social stability risk semantic vectors corresponding to the highest probability are selected as the major engineering project social stability risk semantic vectors that match the scoring query vector.
[0088] In an optional embodiment of the present invention, a weight query prompt is generated based on the weight query vector and the semantic vector of social stability risk of major engineering projects that matches the weight query vector, and a risk indicator weight result of the major engineering projects is generated based on the weight query prompt, including:
[0089] The semantic vector of social stability risk of major engineering projects that matches the weight query vector is merged with the weight query vector to obtain the weight query hint. Then, an autoregressive generation is performed based on the weight query hint using a large language model. When generating each word, the probability distribution is calculated by combining the semantic vector of social stability risk of major engineering projects that matches the weight query vector and the generated context information. The risk index weight results of all major engineering projects are obtained through multiple iterations.
[0090] In this invention, during the indicator weighting stage, a weight-related question template, "Based on the social stability risk assessment indicator system and indicator weights for a major engineering project case, evaluate the social stability risk assessment system and indicator weights for that major engineering project case," is input to generate a prompt. This method first inputs this question template into the nonparametric memory part, as shown in Template 1. The nonparametric memory expands GLM4's capabilities by dynamically retrieving relevant information from an external knowledge base, enhancing the quality of the answers and reducing "illusions." The nonparametric memory part consists of two sub-parts: a query encoder and a document vector index. Let the term sequence of the weighting question be... After the non-parametric memory part of the problem input is processed, it is first transformed into a fixed-length weighted problem vector representation by the query encoder. The question vector is then matched against an index containing a large number of document vectors, which have been pre-processed by a query encoder and stored in an external vector database. In this process, this paper uses... The text embedding model, acting as a query encoder, ensures a semantic match between the input question and the document vector. The document vector index is applied to a given weighted question vector. The maximum inner product search is used to compute the set of document vectors in the vector database. Maximize the inner product with the problem vector, i.e.
[0091]
[0092] For the weight problem vector and document vector set Maximize the inner product The inner product represents the weight problem vector in the vector space. and each document vector The document vector index only needs to return the first few documents in the document set to ensure that the documents that best match the input question are extracted from the large document set and to speed up the search. Maximize the inner product As document vectors This includes knowledge related to weighting issues. This process can be described using conditional probability distributions. This represents the set of document vectors. Each document vector in Scoring is performed, and the results are calculated for a given weighted problem vector. Select document under the condition The probability of this conditional probability. It reflects the degree of matching between the retrieved document vector and the input vector, and its calculation is based on the weighted question vector. and document vectors The inner product similarity, by normalizing these similarity scores, can be used to select each document vector. The probability of:
[0093]
[0094] in, Represents the weight problem vector and document vectors The similarity score is calculated, with the denominator being the sum of the similarity scores of all candidate documents. The document vector that best fits the problem is selected through normalization and maximum likelihood estimation. Furthermore, due to the normalization and return process during retrieval... The operation of maximizing the inner product as a document set combines multiple highly similar document vectors. This allows for correction of discrepancies even within a single document by integrating information from other relevant documents, ensuring the scientific validity and rationality of the retrieved knowledge. Furthermore, the horizontal comparison between multiple information sources reduces errors from a single source. This multi-source fusion approach enhances the reliability of the automatic scoring method, particularly when facing the complex backgrounds and multi-dimensional risk factors of major engineering projects, providing comprehensive and referable risk assessment knowledge. Finally, this step will also find the previous... A collection of document vectors And weight problem vector Merging into indicator weights: a prompt .
[0095] Tips on using indicator weights As input to the parameter memory section, the parameter memory section uses GLM4 as the generator and employs an autoregressive approach, providing hints regarding weight issues. Weight problem vector in The former -1 word unit; the generator not only combines contextual information and relies on the generated text, but also on the retrieved document vectors. External relevant documents can be used as references to enhance the generated text, making it more logical and consistent, and reducing subjective bias. Specifically, for each word in the text generated for this lexical unit... At that time, the generator calculates the probability distribution of generating the word from each document vector, and then sums them to obtain the set of all document vectors. probability distribution under certain conditions :
[0096]
[0097] in, This represents all words that were generated before this word was generated. It is a collection of document vectors A document vector from the document. In this way, the generator processes the previous... Adding -1 lexical unit can generate more accurate and context-sensitive text, and predict the first lexical unit. Analyze each word until a complete case weight result is generated. .
[0098] In an optional embodiment of the present invention, a scoring query prompt is generated based on a scoring query vector, a semantic vector of social stability risk of a major engineering project matching the scoring query vector, and risk indicator weight results. A risk scoring result for the major engineering project is then generated based on the scoring query prompt, including:
[0099] The semantic vector of social stability risk of major engineering projects that matches the scoring query vector is merged with the scoring query vector and the risk indicator weight results to obtain the scoring query prompt. Then, an autoregressive generation is performed based on the scoring query prompt using a large language model. When generating each word, the probability distribution is calculated by combining the semantic vector of social stability risk of major engineering projects that matches the scoring query vector and the generated context information. The risk scoring results of all major engineering projects are obtained through multiple iterations.
[0100] After generating the social stability risk assessment indicator system and its weights for major engineering projects in the indicator weighting stage, the system will move into the probability scoring stage. The core objective of this stage is to automatically score the social stability risk of specific projects based on the indicator weights generated in the previous stage and the scoring criteria within the social stability risk assessment indicator system. This requires inputting a scoring question template: "Based on the social stability risk assessment indicator system and indicator definitions for major engineering projects, assess the social stability risk score of this major engineering project case." Similar to the indicator weighting stage, the scoring question template will also undergo non-parametric memory processing. Specifically, the system first transforms the scoring question into a vector representation using a query encoder. This process ensures a structured representation of the question's semantics, thus providing a foundation for subsequent retrieval and generation. Next, The documents will be fed into the document vector index and matched against a set of document vectors in an external knowledge base. The system uses a maximum inner product search mechanism to find the documents most relevant to the scoring question, forming a set of document vectors. Document vector collection It contains knowledge and information related to social stability risk scoring, which may include historical assessment cases and risk scoring standards in relevant fields. By dynamically retrieving information from different sources, the system not only ensures the multi-dimensionality of the scoring process but also reduces biases introduced by a single information source. Integrating information from different documents enhances the comprehensiveness and accuracy of the scoring. At this stage, the scoring question vector... With document vector set They are not processed in isolation. Instead, they are combined with the case weights generated during the indicator weighting phase. Together, you will receive the final rating query prompt. .
[0101] Use rating query tips As input to the parameter memory section, the process is similar to that in the indicator weighting stage. The input information will undergo autoregressive processing. Autoregression utilizes the current term, relies on the generated context, and incorporates the document vector set. External reference information is used to deduce subsequent lexical units in order to generate a complete case scoring result. Unlike traditional expert scoring methods that rely on subjective judgment, this method, through dynamic retrieval of large amounts of data and deep learning of the model, greatly reduces the errors caused by subjectivity in traditional manual scoring during the automated scoring process, making it more scientifically based.
[0102] In an optional embodiment of the present invention, the method further includes:
[0103] Collect multi-source textual data related to social stability risks of major engineering projects, including social stability risk assessment reports, historical major engineering project emergencies, policy and regulatory data, and news and social data;
[0104] Data clarity and preprocessing are performed on multi-source text data related to social stability risks of major engineering projects to obtain a dataset of social stability risks of major engineering projects.
[0105] For each major historical engineering project emergency, a multi-dimensional scenario space decomposition is performed to obtain multiple scenarios for each event and information for each scenario. The multiple scenarios for each event and information for each scenario are then converted into semantic vectors and stored in the major engineering project social stability risk vector database. The information for each scenario includes scenario objects and scenario elements.
[0106] Social stability risk assessment reports, policy and regulatory data, news and social data are converted into semantic vectors and stored in the social stability risk vector database for major engineering projects.
[0107] In an optional embodiment of the present invention, data cleaning and preprocessing are performed on multi-source text data related to social stability risks of major engineering projects to obtain a dataset on social stability risks of major engineering projects, including:
[0108] Based on the screening criteria, entries that are irrelevant to or incomplete in multi-source text data related to the social stability risks of major engineering projects were removed; the screening criteria were based on the social impact of the project, the scale of mass incidents, and its relevance to the risk assessment objectives.
[0109] The selected multi-source texts related to social stability risks of major engineering projects are deduplicated and standardized to obtain standard multi-source texts related to social stability risks of major engineering projects.
[0110] The multi-source texts related to the social stability risks of major engineering projects are classified according to preset labels, and relevant metadata is attached to each data entry to obtain the social stability risk dataset of major engineering projects.
[0111] This invention constructs a comprehensive, effective, and dynamic dataset of social stability risk scenarios as the foundation for model training and evaluation. To achieve this goal, this invention collects a large amount of textual data related to major engineering projects through various methods. The sources of this data mainly include the following categories:
[0112] 1. Social Stability Risk Assessment Reports: These reports are compiled by governments at all levels and third-party organizations based on the social impact of actual engineering projects. They encompass a comprehensive analysis of the social stability risks that the project may trigger. The reports typically describe in detail the project's social, economic, and environmental impacts, as well as its potential impact on local communities and public opinion. To ensure data representativeness, this invention focuses on collecting data covering engineering projects of different types and scales, especially large-scale infrastructure construction projects, urban planning projects, and projects with significant social impact.
[0113] 2. Historical Accident Cases: This includes historical accidents or mass incidents involving major engineering projects from 2007 to 2024. These incidents are typically social instability factors triggered by the implementation of engineering projects, and the data on the background, development process, and participant reactions are of significant reference value. By collecting and organizing these historical cases, valuable lessons can be extracted from actual events, helping the model to better conduct risk assessments. Case sources include news reports, government-issued accident reports, and academic literature.
[0114] 3. Policy and Regulatory Documents: This section covers national and local policies and regulations related to major engineering projects, including social stability risk assessment standards and construction project management regulations. These documents not only provide a legal framework for project management but also provide a basis for the legality and compliance of subsequent risk assessments. Policy documents guide the social stability risk assessment of projects, clarifying relevant regulations and procedures for risk management and ensuring the legal compliance of the assessment.
[0115] 4. News Reports and Academic Papers: To ensure the dataset is as comprehensive as possible, in addition to collecting information from official reports and accident cases, this invention also extensively collected news reports and academic papers on the social stability impacts of major engineering projects. This data provides a diverse perspective and rich information for risk assessment, particularly in reflecting public opinion and policy dynamics, helping assessment models better capture changes in the social context.
[0116] All collected data must undergo rigorous screening, noise reduction, and standardization to ensure that the data meets quality requirements before being used for model training.
[0117] 1. Data Screening: The collected data first undergoes preliminary screening to remove entries that are irrelevant to the project or lack complete information. For example, cases that do not involve social stability risks or lack key information will be eliminated. Screening criteria are based on the project's social impact, the scale of any potential mass incidents, and its relevance to the risk assessment objectives.
[0118] 2. Data Deduplication: Since data from different sources may contain duplicate content, one of the important tasks in the data cleaning process is to remove duplicate information. Deduplication algorithms are used to remove duplicate text records, ensuring the simplicity and effectiveness of the dataset.
[0119] 3. Text Normalization: To ensure the model can effectively process text data, all text data needs to be standardized. This includes converting different encoding formats in the text, eliminating irrelevant characters, standardizing terminology and abbreviations, etc., so that data from different sources can be uniformly entered into the subsequent processing flow.
[0120] 4. Tagging and Classification: To better organize the data, the collected text data will be classified according to certain tags, such as grouping by time, event type, project scale, etc. Each data entry will be attached with relevant metadata, such as the time of the event, the scope of impact, and the number of participants, which will help with subsequent scenario analysis and model training.
[0121] Finally, a multi-dimensional, dynamic dataset of social stability risk scenarios was constructed, containing textual information from multiple sources. This data, after rigorous screening and organization, will be categorized into several classes to facilitate subsequent model learning and training. The dataset will include the following core components:
[0122] 1. Scenario Case Data: This section includes all historical event cases that passed the screening, categorized by time period, project type, and group incident. Each case records the detailed process and impact of the event and is labeled with its corresponding social stability risk level. This data provides core historical event information for model training.
[0123] 2. Policy and Regulatory Data: This section includes all policies and legal documents related to the project, covering national and local legal frameworks for construction, management, and risk assessment. These documents provide legal compliance for risk assessment and background information for the model.
[0124] 3. Public Opinion and Social Data: In addition to hard data, public opinion and social reaction data are also an important component of this invention. These data sources include news reports, social media articles, public opinions, etc., which can reflect the public's attitude and reaction to the project, providing dynamic contextual information for risk assessment.
[0125] After completing the data collection and preprocessing, the data in the dataset is vectorized to construct a vector database of social stability risks for major engineering projects.
[0126] By decomposing the scenario case data in the dataset using a multidimensional scenario space approach, the model can gain a more comprehensive understanding of the dynamic process of events and provide more refined scenario information for risk assessment. The steps are as follows:
[0127] (1) Segmenting the case: The case is objectively divided into multiple scenarios in chronological order. Each scenario needs to reflect the time and space information of the case, and the scenarios are continuous. Use the letter S to represent the scenario, and use the serial numbers 1, 2, 3... to represent the order of scenario evolution.
[0128] (2) Sub-situations: Sub-situations are also obtained from the objective analysis of the case and are a further expansion of the content of the situation. They are represented by the letter S'.
[0129] (3) Define the case subjects: The subjects refer to the real entities involved in the case. It is necessary to identify the bearers of events and the decision-makers in the event response in the primary and secondary scenarios. Use the letter O to represent the subjects.
[0130] (4) Extract case elements: Extract case elements based on the typical characteristics and frequency of information appearance in each scenario, and use the letter F to represent the elements.
[0131] The BERT embedding model was employed to transform scenario data obtained from multidimensional scenario space decomposition into semantic vectors. These vectors not only accurately express the deep semantics of the text but also capture subtle differences. In this process, each scenario is transformed into a fixed-length vector representing the core information within that scenario. This approach makes the scenario data more structured, enabling the model to effectively process and analyze it. The generated vectors can be used to quickly retrieve similar cases by calculating their similarity, providing effective supporting data. Finally, to effectively store and manage the generated high-dimensional semantic vectors, the vectorized scenario data is stored in a dedicated vector database.
[0132] For the processing of other parts of the dataset, the present invention uses the same embedding technique to directly convert the remaining text data into semantic vectors, and the generated semantic vectors will be stored in a vector database.
[0133] The present invention has the following advantages:
[0134] 1. The assessment cycle has been significantly shortened through an automated risk assessment process;
[0135] 2. By combining historical data, contextual information, and the large language model GLM4, accurate evaluation based on deep semantic understanding was achieved, reducing human interference and subjective bias.
[0136] 3. As new project data and historical events increase, this invention can continuously update the knowledge base to achieve real-time assessment of the latest risks;
[0137] 4. This solution is particularly suitable for complex large-scale engineering projects, and can adapt to different scenarios and changing environments, enabling more accurate and flexible risk assessment.
[0138] It should be noted that, for the sake of simplicity, the method embodiments are all described as a series of actions. However, those skilled in the art should understand that the embodiments of the present invention are not limited to the described order of actions, because according to the embodiments of the present invention, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions involved are not necessarily essential to the embodiments of the present invention.
[0139] Reference Figure 3 The diagram illustrates a structural block diagram of a social stability risk assessment device for major engineering projects provided in an embodiment of the present invention, which may specifically include the following modules:
[0140] The project information and problem acquisition module 301 is used to acquire information on major engineering projects, ask questions about risk indicator weights, and ask questions about risk scoring.
[0141] The project information and question vectorization module 302 is used to generate weight query vectors based on major engineering project information and risk indicator weights, and to generate score query vectors based on major engineering project information and risk score questions.
[0142] The knowledge retrieval module 303 is used to perform knowledge retrieval in a preset major engineering project social stability risk vector database based on the weight query vector and the score query vector, and obtain the major engineering project social stability risk semantic vector that matches the weight query vector and the major engineering project social stability risk semantic vector that matches the score query vector.
[0143] The risk indicator weight result generation module 304 is used to generate weight query prompts based on the weight query vector and the semantic vector of social stability risk of major engineering projects that matches the weight query vector, and to generate risk indicator weight results of major engineering projects based on the weight query prompts.
[0144] The risk scoring result generation module 305 is used to generate scoring query prompts based on the scoring query vector, the semantic vector of social stability risk of major engineering projects that matches the scoring query vector, and the risk indicator weight results, and to generate risk scoring results for major engineering projects based on the scoring query prompts.
[0145] As the device embodiment is basically similar to the method embodiment, the description is relatively simple, and relevant parts can be found in the description of the method embodiment.
[0146] In addition, embodiments of the present invention also provide an electronic device, such as... Figure 4 As shown, it includes a processor 401, a communication interface 402, a memory 403, and a communication bus 404, wherein the processor 401, the communication interface 402, and the memory 403 communicate with each other through the communication bus 404.
[0147] Memory 403 is used to store computer programs;
[0148] When the processor 401 executes the program stored in the memory 403, it implements the social stability risk assessment method for major engineering projects as described in the above embodiments.
[0149] The communication bus mentioned above can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used to represent it in the diagram, but this does not mean that there is only one bus or one type of bus.
[0150] The communication interface is used for communication between the aforementioned terminal and other devices.
[0151] The memory may include random access memory (RAM) or non-volatile memory, such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.
[0152] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0153] like Figure 5 As shown, in another embodiment of the present invention, a computer-readable storage medium 501 is also provided, which stores instructions that, when executed on a computer, cause the computer to perform the social stability risk assessment method for major engineering projects described in the above embodiments.
[0154] In another embodiment of the present invention, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to execute the social stability risk assessment method for major engineering projects described in the above embodiments.
[0155] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state disk (SSD)).
[0156] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0157] The various embodiments in this specification are described in a related manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the apparatus embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.
[0158] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention are included within the scope of protection of the present invention.
Claims
1. A method for assessing social stability risks of major engineering projects, characterized in that, The method includes: Access to information on major engineering projects, questions regarding risk indicator weights, and questions regarding risk scoring; A weighted query vector is generated based on information about major engineering projects and risk indicator weights, and a scoring query vector is generated based on information about major engineering projects and risk scores. Based on the weighted query vector and the scoring query vector, knowledge retrieval is performed in the pre-set social stability risk vector database of major engineering projects to obtain the semantic vector of social stability risk of major engineering projects that matches the weighted query vector and the semantic vector of social stability risk of major engineering projects that matches the scoring query vector. Based on the weighted query vector and the semantic vector of social stability risk of major engineering projects that matches the weighted query vector, a weighted query hint is generated, and the risk indicator weight result of major engineering projects is generated based on the weighted query hint. Based on the scoring query vector, the semantic vector of social stability risk of major engineering projects that matches the scoring query vector, and the risk indicator weight results, a scoring query prompt is generated, and the risk scoring result of the major engineering project is generated based on the scoring query prompt.
2. The method according to claim 1, characterized in that, Based on information about major engineering projects and risk indicator weights, a weighted query vector is generated. Similarly, based on information about major engineering projects and risk scoring, a scoring query vector is generated, including: The BERT text embedding model is used to convert questions about major engineering projects and risk indicator weights into weighted query vectors, and questions about major engineering projects and risk scores into score query vectors.
3. The method according to claim 1, characterized in that, Based on the weighted query vector and the scoring query vector, knowledge retrieval is performed in a pre-defined database of social stability risks for major engineering projects to obtain semantic vectors of social stability risks for major engineering projects that match the weighted query vectors and semantic vectors of social stability risks for major engineering projects that match the scoring query vectors, including: Based on the weighted query vector, the candidate semantic vectors of social stability risk of major engineering projects are selected from the pre-set database of social stability risk vectors of major engineering projects by means of maximum inner product search. Then, by means of normalized inner product score and maximum likelihood estimation, a conditional probability distribution is constructed for the candidate semantic vectors of social stability risk of major engineering projects of the weighted query vector. The first pre-set number of candidate semantic vectors of social stability risk of major engineering projects corresponding to the maximum probability are selected as the semantic vectors of social stability risk of major engineering projects that match the weighted query vector. Based on the scoring query vector, candidate major engineering project social stability risk semantic vectors are selected from the pre-defined database of major engineering project social stability risk vectors through maximum inner product search. Then, a conditional probability distribution is constructed for the candidate major engineering project social stability risk semantic vectors of the scoring query vector through normalized inner product score and maximum likelihood estimation. The top pre-defined number of candidate major engineering project social stability risk semantic vectors corresponding to the highest probability are selected as the major engineering project social stability risk semantic vectors that match the scoring query vector.
4. The method according to claim 1, characterized in that, Based on the weighted query vector and the semantic vector of social stability risk of major engineering projects that matches the weighted query vector, a weighted query hint is generated. Then, based on the weighted query hint, the risk indicator weights of major engineering projects are generated, including: The semantic vector of social stability risk of major engineering projects that matches the weight query vector is merged with the weight query vector to obtain the weight query hint. Then, an autoregressive generation is performed based on the weight query hint using a large language model. When generating each word, the probability distribution is calculated by combining the semantic vector of social stability risk of major engineering projects that matches the weight query vector and the generated context information. The risk index weight results of all major engineering projects are obtained through multiple iterations.
5. The method according to claim 1, characterized in that, Based on the scoring query vector, the semantic vector of social stability risk of major engineering projects matching the scoring query vector, and the risk indicator weight results, a scoring query prompt is generated. Then, based on the scoring query prompt, a risk scoring result for the major engineering project is generated, including: The semantic vector of social stability risk of major engineering projects that matches the scoring query vector is merged with the scoring query vector and the risk indicator weight results to obtain the scoring query prompt. Then, an autoregressive generation is performed based on the scoring query prompt using a large language model. When generating each word, the probability distribution is calculated by combining the semantic vector of social stability risk of major engineering projects that matches the scoring query vector and the generated context information. The risk scoring results of all major engineering projects are obtained through multiple iterations.
6. The method according to claim 1, characterized in that, The method further includes: Collect multi-source textual data related to social stability risks of major engineering projects, including social stability risk assessment reports, historical major engineering project emergencies, policy and regulatory data, and news and social data; Data cleaning and preprocessing were performed on multi-source text data related to social stability risks of major engineering projects to obtain a dataset of social stability risks of major engineering projects. For each major historical engineering project emergency, a multi-dimensional scenario space decomposition is performed to obtain multiple scenarios for each event and information for each scenario. The multiple scenarios for each event and information for each scenario are then converted into semantic vectors and stored in the major engineering project social stability risk vector database. The information for each scenario includes scenario objects and scenario elements. Social stability risk assessment reports, policy and regulatory data, news and social data are converted into semantic vectors and stored in the social stability risk vector database for major engineering projects.
7. The method according to claim 6, characterized in that, Data cleaning and preprocessing were performed on multi-source text data related to social stability risks of major engineering projects to obtain a dataset on social stability risks of major engineering projects, including: Based on the screening criteria, entries that are irrelevant to or incomplete in multi-source text data related to the social stability risks of major engineering projects were removed; the screening criteria were based on the social impact of the project, the scale of mass incidents, and its relevance to the risk assessment objectives. The selected multi-source texts related to social stability risks of major engineering projects are deduplicated and standardized to obtain standard multi-source texts related to social stability risks of major engineering projects. The multi-source texts related to the social stability risks of major engineering projects are classified according to preset labels, and relevant metadata is attached to each data entry to obtain the social stability risk dataset of major engineering projects.
8. A device for assessing social stability risks of major engineering projects, characterized in that, The device includes: The project information and question acquisition module is used to obtain information on major engineering projects, ask questions about risk indicator weights, and ask questions about risk scoring. The project information and question vectorization module is used to generate weighted query vectors based on major engineering project information and risk indicator weights, and to generate score query vectors based on major engineering project information and risk score questions. The knowledge retrieval module is used to perform knowledge retrieval in a preset database of social stability risk vectors for major engineering projects based on weighted query vectors and scoring query vectors, and to obtain semantic vectors of social stability risks for major engineering projects that match the weighted query vectors and semantic vectors of social stability risks for major engineering projects that match the scoring query vectors. The risk indicator weight result generation module is used to generate weight query prompts based on the weight query vector and the semantic vector of social stability risk of major engineering projects that matches the weight query vector, and to generate risk indicator weight results of major engineering projects based on the weight query prompts. The risk scoring result generation module is used to generate scoring query prompts based on the scoring query vector, the semantic vector of social stability risk of major engineering projects that matches the scoring query vector, and the risk indicator weight results, and to generate risk scoring results for major engineering projects based on the scoring query prompts.
9. An electronic device, characterized in that, It includes a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; The memory is used to store computer programs; When the processor executes the program stored in the memory, it implements the social stability risk assessment method for major engineering projects as described in any one of claims 1-7.
10. One or more computer-readable media, characterized in that, It stores instructions that, when executed by one or more processors, cause the processors to perform the social stability risk assessment method for major engineering projects as described in any one of claims 1-7.
Citation Information
Patent Citations
Method for evaluating early-stage social risks of construction project
CN112132375A
Social risk early warning method and system based on multi-source data
CN119719813A