Method and system for automatic generation of secure code based on full information theory and mechanism theory

Through the automatic generation method of security code with full information theory and mechanismism, the problems of insufficient security, interpretability and adaptability in the existing technology are solved, and the high security and autonomous defense capabilities of code generation are achieved, which are suitable for high security demand scenarios such as industrial control systems and smart contract development.

CN120447881BActive Publication Date: 2025-09-02SHENZHEN CESTBON TECH CO
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510965201.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-14
Publication Date
2025-09-02
Estimated Expiration
2045-07-14

AI Technical Summary

Technical Problem

Existing code generation methods have limitations in terms of security, interpretability and adaptability, and cannot effectively identify semantic vulnerabilities or pragmatic risks in the deep logic of the code. Moreover, security strategies lack adaptability, making it difficult to build a closed-loop repair mechanism.

Method used

Using a secure code automatic generation method based on full information theory and mechanismism, a high-security code generation framework is built through the deep collaboration mechanism of the grammar verification layer, semantic analysis layer and pragmatic optimization layer to realize the endogenous fusion of grammar, semantic and pragmatic three-dimensional information and dynamic strategy generation.

Benefits of technology

It improves the security, interpretability and adaptability of code generation, realizes the endogenous fusion and autonomous evolution of security attributes, and can prevent vulnerabilities in the code generation stage and adapt to dynamic changes in complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120447881B_ABST
    Figure CN120447881B_ABST
Patent Text Reader

Abstract

The present invention provides a method and system for automatically generating secure code based on full information theory and mechanism theory, comprising: step S1, extracting features of the syntax verification layer of the input requirements, extracting the syntax features of the code through the topological features of the abstract syntax tree, and performing formal rule verification on the extracted syntax features; step S2, constructing a semantic association model across code fragments through the semantic parsing layer to form a propagation path for tracking data flow and control flow; step S3, in the pragmatic optimization layer, generating a security policy through a dynamic utility function in combination with the pragmatic constraints of the business scenario; step S4, optimizing the parameters of the dynamic utility function and the security policy through the knowledge federation evolution and closed-loop feedback mechanism. The present invention reconstructs the deep collaborative mechanism and dynamic policy generation model of the three-dimensional information of syntax, semantics and pragmatics, transforms the security logic from an external attribute to an endogenous cognitive ability, and improves the security, explainability and adaptability of automatic code generation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a code generation method, in particular to a method for automatically generating secure codes based on full information theory and mechanismism, and further to a system adopting the method for automatically generating secure codes based on full information theory and mechanismism. Background Art

[0002] Currently, code generation methods based on large language models are widely used in the field of artificial intelligence. However, these existing code generation methods face severe security challenges. Existing approaches to AI often treat security as an external attribute of intelligent systems, rather than an inherent element of the system. Under the dominance of the material science paradigm, security attributes have been isolated and formalized, while semantic and pragmatic information, which should be the primary focus, have become secondary dimensions.

[0003] This traditional code generation method is limited by the divide-and-conquer mentality under the material science paradigm, separating formal verification, semantic analysis, and security constraints into independent links, resulting in the following problems in the generated code: First, there are limitations in security. Formal verification only focuses on syntactic compliance and cannot identify semantic vulnerabilities or pragmatic risks in the deep logic of the code; second, there are limitations in applicability. The static security rule base is difficult to cope with the dynamic evolution of complex attack scenarios and lacks the ability to make adaptive decisions based on value rationality; third, there are limitations in interpretability. Existing methods rely on black box models to generate code logic. Its unexplainability makes it difficult to trace security defects. In addition, existing technologies usually adopt a fixed defense strategy solution and cannot build a closed-loop repair mechanism. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to provide a method for automatically generating secure code based on full information theory and mechanismism. By constructing a high-security automatic code generation framework based on a unified intelligent model of full information theory and mechanismism, a deep collaborative mechanism and dynamic strategy generation model of grammatical, semantic, and pragmatic three-dimensional information are reconstructed to achieve the endogenous integration and autonomous evolution of security attributes, thereby improving the security, explainability, and adaptability of automatic code generation. On this basis, a system that adopts this method for automatically generating secure code based on full information theory and mechanismism is further provided.

[0005] To this end, the present invention provides a method for automatically generating secure codes based on full information theory and mechanism theory, comprising the following steps:

[0006] Step S1, performing feature extraction of the syntax verification layer on the input requirements, in which the syntax features of the code are extracted through the topological features of the abstract syntax tree, and the extracted syntax features are verified by formal rules to establish a standardized representation corresponding to the syntax features;

[0007] Step S2: constructing a semantic association model across code snippets through the semantic parsing layer to form a propagation path for tracing data flow and control flow;

[0008] Step S3: In the pragmatic optimization layer, a security policy is generated through a dynamic utility function in combination with pragmatic constraints of the business scenario.

[0009] Step S4: Optimize the parameters of the dynamic utility function and the security policy through knowledge federation evolution and closed-loop feedback mechanism.

[0010] A further improvement of the present invention is that step S1 includes the following sub-steps:

[0011] In step S101, the parser corresponding to the programming language is first called to convert the input source code text into an abstract syntax tree. The abstract syntax tree is then traversed based on a predefined security sensitivity rule base, and the node weights are calculated using a preset node type weight table. Nodes with weights above a preset threshold are then marked as key syntax nodes. A multi-dimensional feature analysis is then performed on the key syntax nodes to construct a structured feature set. Finally, the structured features in the structured feature set are vectorized and encoded to form a fixed-length syntax feature vector.

[0012] Step S102: statically verify the grammatical feature vector based on preset formal verification rules;

[0013] In step S103, a graph node entity is first created for each key syntax node based on the static verification results. Then, the topological dependencies between the key syntax nodes are analyzed to construct three types of core edge structures, including control flow edges, data flow edges, and call relationship edges. Finally, a graph convolutional network is used to perform third-order neighborhood feature aggregation to collect risk features of surrounding nodes. The diffusion probability distribution of vulnerability impact is calculated using the Softmax activation function algorithm, and the PageRank algorithm is used to locate key syntax nodes. The core nodes with the greatest risk radiation capability are identified to complete the mapping of static verification results to syntax feature graphs.

[0014] Step S104: convert the key grammatical nodes corresponding to the grammatical feature graph into predefined semantic labels, and then convert the grammatical dependency into an association path with security semantics.

[0015] A further improvement of the present invention is that step S2 includes the following sub-steps:

[0016] Step S201: constructing a semantic association model across code snippets based on the knowledge graph;

[0017] Step S202: Analyze the constructed semantic association model using a graph neural network to identify potential attack intentions.

[0018] Step S203: Generate a semantic risk vector based on the identified potential attack intentions to quantify the logical hazard level of the security vulnerability.

[0019] A further improvement of the present invention is that in step S201, based on the knowledge graph, semantic association is performed on user input nodes and database operation interfaces across code fragments through a semantic parsing layer to construct a semantic association model across code fragments.

[0020] A further improvement of the present invention is that step S3 includes the following sub-steps:

[0021] Step S301: construct a dynamic strategy generation model based on pragmatic constraints of business scenarios;

[0022] Step S302: Call the risk entropy model and calculate the utility value of each strategy through the dynamic utility function , the calculation process of the dynamic utility function is: ,in, 、 and Represent the coefficients under different business scenarios respectively; represents the security level parameter, and the dynamic utility function is expressed by Indicates the principle of safety first; Represents the performance loss parameter, which is used to reflect the computing resource overhead introduced by the candidate strategy set during the deployment process. The dynamic utility function is Represents computing resource constraints; Represents the user experience parameter, which is used to reflect the impact of the strategy on the user experience. The dynamic utility function is Represents the law of diminishing marginal utility of user experience; through the formula Calculate the safety level parameters , Indicates the dynamic adjustment coefficient of the business scenario, Indicates the security level parameter The baseline value, It represents the statically configured security baseline; through the formula Calculate benchmark value , represents the risk entropy model, , It represents the triggering probability of the kth vulnerability. It represents the pragmatic value weight of the associated asset, n represents the number of vulnerabilities, and k represents the sequence number of the vulnerability; Indicates the maximum theoretical entropy value of the current scene;

[0023] Step S303: Select utility value The highest strategy.

[0024] A further improvement of the present invention is that in step S302, when the attack risk increases, the security level parameter is increased by increasing the authentication strength. Approaching 1, ensuring utility value Positive growth, and the coefficient Dynamic adjustment based on asset value; performance loss parameters Normalized calculation is performed based on real-time monitoring. When the edge node load exceeds the preset threshold, the coefficient Automatically adjust upwards; dynamically adjust downwards during peak business hours .

[0025] A further improvement of the present invention is that in step S303, in the current business scenario, a candidate strategy set is generated according to the preset strategy response, and the utility values ​​of all feasible strategies are continuously calculated through the dynamic utility function. , select the utility value The largest strategy is used as the global optimal path. The preset strategy responses include only enhanced log auditing, mandatory MFA verification, and rate limiting and downgrading. MFA verification refers to multi-factor authentication.

[0026] A further improvement of the present invention is that step S4 includes the following sub-steps:

[0027] Step S401: The edge node extracts attack features and generates a knowledge vector based on the extracted attack features;

[0028] Step S402: Upload the generated knowledge vector to a global knowledge base in the cloud through a privacy protection protocol. The global knowledge base includes a semantic knowledge base and a grammar verification rule base.

[0029] Step S403: The semantic knowledge base updates the attack pattern graph and reversely injects the updated content into the syntax verification rule base of the edge node;

[0030] Step S404: Optimizing the parameters of the dynamic utility function and the security policy through a closed-loop feedback mechanism.

[0031] A further improvement of the present invention is that, by formula Calculating the attack missed detection rate when automatically generating security codes ,in, represents the missed detection rate of the black box model, represents the full information synergistic enhancement coefficient, 、 and Represent the security weights of grammatical, semantic and pragmatic dimensions respectively.

[0032] The present invention also provides a security code automatic generation system based on full information theory and mechanismism, which adopts the security code automatic generation method based on full information theory and mechanismism as described above, and includes:

[0033] The feature extraction module of the syntax verification layer is used to extract the features of the syntax verification layer based on the input requirements. In the syntax verification layer, the syntax features of the code are extracted through the topological features of the abstract syntax tree, and the extracted syntax features are formalized and checked to establish a standardized representation corresponding to the syntax features.

[0034] The semantic association module of the semantic parsing layer builds a semantic association model across code snippets through the semantic parsing layer, forming a propagation path for tracing data flow and control flow;

[0035] The dynamic policy generation module of the pragmatic optimization layer generates security policies through dynamic utility functions in combination with pragmatic constraints of business scenarios.

[0036] The knowledge federation evolution and strategy closed-loop optimization module is used to optimize the parameters of the dynamic utility function and the security strategy through the knowledge federation evolution and closed-loop feedback mechanism.

[0037] Compared with the prior art, the beneficial effects of the present invention are: first, feature extraction is performed on the input requirements at the syntax verification layer, and then a semantic association model across code fragments is constructed through the semantic parsing layer to form a propagation path for tracking data flow and control flow. Then, at the pragmatic optimization layer, pragmatic constraints of the business scenario are combined to generate security policies through dynamic utility functions. Finally, the parameters of the dynamic utility function and the security policy are optimized through the knowledge federation evolution and closed-loop feedback mechanism, thereby constructing a high-security code automatic generation framework based on a unified intelligent model of full information theory and mechanismism, and reconstructing a deep collaborative mechanism and dynamic policy generation model of grammatical, semantic and pragmatic three-dimensional information. The present invention jointly models the grammatical features, semantic logic and pragmatic constraints of the code, transforming the security logic from an external attached attribute to an endogenous cognitive ability of the system, and realizing vulnerability prevention in the code generation stage. Therefore, it can effectively realize the endogenous fusion and autonomous evolution capability of security attributes, and improve the security, explainability and adaptability of code automatic generation. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 This is a schematic diagram of a high-security code automatic generation framework according to an embodiment of the present invention;

[0039] Figure 2 This is a diagram of the existing security issues of the three major schools of artificial intelligence;

[0040] Figure 3 It is a three-dimensional perspective diagram of full information theory;

[0041] Figure 4 This is a schematic diagram of the workflow of an embodiment of the present invention;

[0042] Figure 5 It is a schematic diagram of a closed-loop feedback mechanism optimization process according to an embodiment of the present invention. DETAILED DESCRIPTION

[0043] Before describing the specific embodiments of the present invention in detail, the key terms and related technologies of the present invention are first explained.

[0044] PH stands for Package Hallucination; COT stands for Chain of Thought; AST stands for Abstract Syntax Tree; m-GTAI stands for mechanism-based theory of AI.

[0045] The revolution in the artificial intelligence (AI) paradigm has brought new technological potential to cyberspace security. Mechanism-based AI theory, by reconstructing the three-dimensional information synergy of "syntax, semantics, and pragmatics," has significantly improved the interpretability and inherent security of intelligent systems. Traditional AI research has long been constrained by the traditional material science paradigm, leaving security defense systems in a difficult, reactive state. With the widespread application of AI technology in code generation, traditional code security assurance systems are facing the daunting challenge of a paradigm shift. Currently, the core contradiction in code generation security lies in establishing an interpretable, verifiable, and traceable synergy between intelligent generation efficiency and inherent security.

[0046] First, the dilemma of artificial intelligence security research is introduced.

[0047] Because current AI research adheres to the traditional disciplinary paradigm of divide-and-conquer and pure formalism, the overall field of AI research has been fragmented, divided into three major schools: structuralism, functionalism, and behaviorism. These three schools all have significant limitations in the areas of cybersecurity and code security. More importantly, these limitations are not due to accidental errors in AI technology, but rather to systemic flaws in the underlying research paradigm.

[0048] Structuralism simplified intelligence to a simple imitation of the human brain's organizational structure. Over time, researchers from this school developed black-box models in the security field that rely heavily on statistical learning. These models rely on vast amounts of data to train their ability to identify attack patterns. However, during training, these models overly focus on the data's syntactical features and ignore its semantic and pragmatic information, resulting in a loss of interpretability. In practical cybersecurity, the black-box nature of these models makes vulnerability tracing extremely difficult. Once-clear attack paths are obscured by the black-box model's uninterpretable weight parameters, rendering all defenses largely passive. Furthermore, in the face of zero-day attacks or adversarial attacks, these models' lack of prior knowledge makes their ability to identify unknown attack patterns even less than ideal. Relying solely on statistical patterns in historical data, black-box models often fail to understand the semantic intent behind malicious behavior. The fundamental reason for these phenomena is that the divide-and-conquer approach to training neural networks ignores the inherent connections within the code. This approach focuses solely on local features such as byte sequences or traffic statistics, disrupting the integrated understanding of program semantics, pragmatics, and grammar.

[0049] Functionalism attempts to construct formal security rules using expert systems, but many scholars often fall into the cognitive trap of physical symbolic assumptions. In the process of constructing formal security rules, security logic must be abstracted into static symbolic reasoning, and much dynamic semantic information is rigidly fixed into various conditional statements. In real-world attack scenarios, especially when attackers employ polymorphic code obfuscation or lateral movement, the rule bases pre-set by security personnel often exhibit significant lags. Even more unacceptable is that some legitimate user operations are mistakenly labeled as anomalies by the system because they reach the threshold of the formal rules, while some malicious actions carefully disguised by the attacker are successfully executed because they conform to the syntax of the pre-set rule base. The root cause of these phenomena stems from the material science paradigm's emphasis on form. Its security policies are often divorced from real-world security scenarios and rely solely on superficial reasoning based on static symbols.

[0050] Behaviorism focuses primarily on the immediate feedback between perception and action, viewing human thought as a process of conditioned reflexes and stimulus responses. This process has gradually evolved in the security field into a threat response mechanism, centered around behavioral feature matching. This mechanism blocks malicious behavior by monitoring superficial metrics like API call frequency and network traffic peaks. However, due to a lack of in-depth understanding of malicious behavior, normal, high-frequency user access may be mechanically classified as an attack, leading to the blocking of critical services. Meanwhile, some low-frequency malicious data infiltrations can evade detection due to their "mild" behavior patterns. These phenomena indicate that behaviorism can lead to global imbalances when faced with local optimization problems and cannot effectively balance security and utility, often making binary decisions when privacy and efficiency conflict.

[0051] The current status of security issues of these three schools is as follows Figure 2 Although the three schools of thought have different research paths, they share similar issues in the security field: they all view security as an external, attached attribute of intelligent systems, rather than an inherent element of the system. Structuralism's statistical black-box model, functionalism's symbolic assumptions and authentication traps, and behaviorism's mechanical reflexive behavior all demonstrate a lack of inherent security. When attackers leverage semantic information to construct logic bombs or rely on pragmatic information to infiltrate permission systems, the limitations of traditional AI become the weakest point in security defenses.

[0052] Then, the breakthrough of total information theory (also known as total information theory) is introduced.

[0053] During the paradigm shift in artificial intelligence, the emergence of full information theory marked a fundamental overturn in traditional views of information. Traditional intelligence research, long constrained by the material science paradigm, simplified information into symbolic manipulation at the "syntax" level, ignoring the deeper connections between "semantics" and "pragmatics" in human intelligence. This disconnect directly contributed to AI's inability to understand complex scenarios. The key point of full information theory is to reconstruct the information environment that integrates syntax, semantics, and pragmatics, providing a foundation for the explainability of intelligent systems.

[0054] Syntactic information is the most superficial dimension of information, presented as a data structure, such as the character sequence of code or the message structure of a network protocol. Neuroscience research shows that the human perceptual system can only directly grasp the formal characteristics of external stimuli: visual neurons record the wavelength and intensity of light waves, while auditory neurons respond to the frequency and amplitude of sound waves. This mechanism is related to the field of artificial intelligence, specifically the ability of sensors and statistical models to extract syntactic information. If this were the case, intelligent systems would become "blind signal processors." Firewalls could only detect unusual byte distributions in packets, but could not distinguish between malicious payloads. Code review tools could identify syntactic errors but could not understand the attack intent behind logical vulnerabilities.

[0055] The introduction of pragmatic information imbues grammatical data with a value dimension and serves as the core basis for intelligent systems' decision-making. Cognitive science research shows that the human attention mechanism is essentially a form of pragmatic screening: the brain does not passively receive all sensory signals, but rather prioritizes information based on its relevance to the target. When operations personnel begin monitoring network traffic, they focus on anomalous conversations rather than regular requests. This choice is not based on grammatical differences in the data, but rather on pragmatic judgments based on experience. In the context of artificial intelligence, pragmatic information is reflected in the generation logic of security policies: when an intrusion detection system identifies signs of high-frequency port scanning, it needs to determine the response level based on the value weight of the network assets, rather than simply making mechanical interception actions based on grammatical features.

[0056] Semantic information can connect form and meaning. Regarding the content of information, relevant research in cognitive linguistics shows that human understanding of language does not rely on the statistics of symbols, but rather relies on semantic networks to combine concepts into knowledge systems. The semantics of the word encryption not only includes the grammatical rules of character combinations, but is also related to deep-level content such as key management and algorithm strength. In the field of security, the lack of semantics will lead to serious logical breaks.

[0057] The major breakthrough of full information theory is that it reveals the dynamic transformation mechanism of grammar, pragmatics and semantics. The human brain uses the "information-knowledge-intelligence" transformation chain to integrate the grammatical information captured by the senses with prior knowledge, realize the semantic analysis of the current scene, and make corresponding decisions based on the target value. This process is presented in network security as a multi-layer defense collaborative system. The grammatical layer in this collaborative system can quickly filter out noise, the semantic analysis layer identifies the attack logic, and the pragmatic evaluation dynamically adjusts the response measures. Because the core defect of the three major schools of traditional artificial intelligence is to tear apart the integrity of full information, this invention uses full information theory as a link to provide a foundation for building an intelligent security architecture that is "formally verifiable, content explainable, and value traceable."

[0058] Full information theory provides a new "three-dimensional perspective" for the paradigm change of artificial intelligence, such as Figure 3 As shown, grammatical information enables machines to interact precisely with the physical world, semantic information gives them a deep understanding of behavioral logic, and pragmatic information plays a key role in their value rationality in complex environments. For a long time, academics viewed semantics, grammar, and pragmatics as three parallel concepts at the same level. However, grammar and pragmatics are intuitive concepts, while semantics is not. Semantic information cannot be obtained through direct observation or testing; it can only be obtained through abstraction from the grammatical and pragmatic community. Therefore, grammar and pragmatics are concepts at the same level, while semantics is an abstraction from their community, a higher-level concept than grammar and pragmatics.

[0059] Next, the core advantages of mechanism-based artificial intelligence are introduced.

[0060] Traditional AI research has long been constrained by the "divide and conquer" and "pure formalism" paradigms of traditional material science. This has led to the intelligent generation mechanism being trapped in an unexplainable black box state and a fragmented cognitive dilemma. Mechanistic AI theory, on the other hand, forms a universal intelligent generation mechanism by constructing a "full information-knowledge-intelligence" transformation chain. Its theoretical breakthroughs are concentrated in the two dimensions of explainability and universality.

[0061] Interpretability stems from the transparent transformation of semantic information. Mechanism-based AI theory emphasizes that the essence of intelligence lies not in the mechanical permutation and combination of symbols or the statistical correlation of data, but rather in the co-evolution of the three-dimensional information ecosystem of "syntax, semantics, and pragmatics." Taking secure code generation as an example, traditional neural network models focus only on surface-level formal features such as variable naming conventions and function call patterns. Mechanism-based AI theory, however, employs the principle of first-class information conversion, transforming perceptual information into full information, mapping the grammatical structure of the code into the pragmatic dimensions of semantic logic and compliance constraints that express functional intent. For example, the grammatical feature of "file read and write operations" appearing in the code is transformed into a semantic representation of "data leakage risk," and vulnerability prevention is achieved through a dynamic permission constraint mechanism. This parseable path from formal features to semantic connotations frees security policy construction from a reliance on empirical rules and instead builds on a deeper semantic understanding of code behavior.

[0062] Universality stems from the inclusiveness of the theoretical framework of the intelligent generation mechanism. The four types of information conversion principles proposed by mechanismism constitute a closed-loop feedback system that can flexibly adapt to security needs in multiple fields. In network behavior authentication scenarios, traditional behaviorist models rely on shallow indicators such as traffic peaks and API call frequencies to trigger responses. In contrast, mechanismist artificial intelligence models transform network behavior data into semantic representations of user behavior intentions and generate dynamic authentication strategies based on business value weights. Mechanismist agents can not only recognize shallow syntactic information but also generate multi-path solutions through semantic mapping of utility functions and pragmatic trade-offs of energy optimization, thus breaking through the theoretical limitations of the rigid rules of traditional expert systems.

[0063] Finally, the current status of the code generation security field is introduced.

[0064] Currently, code generation technology based on large language models faces severe security challenges. Research has shown that code generated using large language models is prone to widespread security vulnerabilities. The root causes of these vulnerabilities can be categorized into three main categories: contamination of training data, structural flaws in the generation mechanism, and fragmentation of the security verification system.

[0065] Empirical analysis of mainstream open-source codebases revealed security or maintainability flaws in a small percentage of training samples, directly leading to a 5.85% defect rate in the code generated by fine-tuned models. This contamination of training data not only creates explicit vulnerabilities but also potentially introduces the risk of package hallucination. Research has found that large language models are highly likely to fabricate unverified dependency packages when generating Python code, providing a covert channel for attackers to inject malicious code. Attackers can successfully trick models like CodeBERT into generating vulnerable code by contaminating only 3% of the training data, without requiring explicit trigger activation. This type of targeted poisoning attack renders traditional defense mechanisms ineffective.

[0066] At the generation mechanism level, large language models prioritize grammatical correctness and functional integrity, leading to a systemic lack of security constraints. The EXACT evaluation framework shows that while the SHA1 encrypted code generated by large language models compiles, it has a high crash rate, significantly inferior to human-written code. In formal verification scenarios, the semantic error rate of assertions generated by commercial LLMs reached 68%, exposing vulnerabilities in core security logic. This mechanism flaw presents multi-dimensional risks. Research has shown that while using thought chaining prompts can improve C / C++ vulnerability detection accuracy by 21.6%, it reduces Java detection accuracy by 4.6%, demonstrating the fragmentation of cross-language security cognition. Other research has found that GPT-4's success rate in repairing its own generated code is significantly lower than that of repairing code from other models, demonstrating a structural flaw in its self-healing capabilities. In Linux kernel vulnerability detection, research has confirmed that traditional static analysis tools have a 42% false positive rate for UBI vulnerabilities, while large language models, without guidance strategies, have a detection success rate of less than 15%. These data reflect systemic failures in complex scenarios.

[0067] Existing security enhancement solutions exhibit a fragmented "local patching" nature, making it difficult to build a systematic defense system. While the code security prefix technology developed by researchers has reduced the vulnerability rate of GPT-4o, it has also led to a decrease in functional integrity, exposing a fundamental contradiction between security and performance. While the VSP strategy achieves high detection accuracy, it relies on a manually designed vulnerability semantic mapping chain, limiting its adaptability. The exploration of collaborative verification mechanisms also faces bottlenecks. While the cross-model comparison strategy reduces the probability of malicious code retention to below 5%, the computational cost surges. The AutoSafeCoder framework reduces vulnerabilities by 13% through three-agent collaboration, but struggles to cope with the dynamic evolution of polymorphic code.

[0068] A more profound security risk stems from the cognitive limitations of technological approaches. Currently, security enhancement for large language models is split into two approaches: fine-tuning general models and training specialized models. The former suffers from high misjudgment rates due to domain knowledge bias, while the latter is limited by data scarcity and training costs. Multi-dimensional risk analysis from some studies warns that large language models can be affected by training data, with 38% of generated code exhibiting intellectual property ambiguity. In multi-agent scenarios, package hallucinations in a single model can contaminate the entire system through API calls, creating cascading security risks.

[0069] Currently, the field of code generation security is transitioning from vulnerability patching to cognitive immunity. Although local breakthroughs have been achieved through data cleaning, prompt engineering, and evaluation framework optimization, the hidden contamination of training data, structural flaws in generation mechanisms, and the fragmentation of security verification systems still pose systemic security threats. Future breakthroughs will require going beyond the boundaries of existing tools. Therefore, this invention embeds security attributes into the model architecture, training paradigm, and verification mechanism, achieving a three-dimensional collaborative defense guided by full information theory.

[0070] Vulnerability semantics guided prompting (VSP), a prior art related to the present invention, is a vulnerability analysis method based on a large language model. Its core goal is to improve the ability to identify, discover, and remediate software vulnerabilities by guiding the model to focus on the essential semantic logic of the vulnerability. This technology breaks down complex vulnerability analysis tasks into interpretable, chained reasoning steps, enabling the model to simulate the analysis process of security experts.

[0071] In terms of technical implementation, VSP first defines vulnerability semantics as the core data and control flow paths within a program that lead to vulnerable behavior. This includes the code statements that directly trigger the vulnerability, such as unvalidated buffer operations, and the context they depend on. When the large language model discovers missing bounds checks or uninitialized pointers, it uses VSP strategies to map the vulnerability semantics into thought nodes for chained reasoning. The model can then gradually verify whether there are logical paths in the code that violate security rules.

[0072] VSP's reasoning framework design combines static code analysis with dynamic behavior simulation. In vulnerability discovery tasks, the model analyzes potentially dangerous operations in code snippets, such as dynamic memory allocation or file reading and writing, combined with contextual dependencies, such as unvalidated user input paths, to infer the likelihood of vulnerabilities and ultimately classify them into standard vulnerability types. For vulnerability remediation tasks, VSP not only locates the vulnerability points, but also traces back to the root causes of the vulnerabilities, such as conditional branch logic errors, and generates minimal code changes. This process is achieved through a small number of example prompts, which clearly show the reasoning path of the vulnerability semantics, thereby guiding the model to learn the patterned logic of vulnerability analysis.

[0073] Furthermore, VSP adapts to the needs of different scenarios through task-customized prompting strategies. In vulnerability identification, prompt examples clearly specify the definition of the target CWE and typical code patterns; in vulnerability discovery, examples cover the semantic characteristics of multiple vulnerability types, requiring the model to autonomously reason and classify; and in vulnerability remediation, examples emphasize the correlation between root cause analysis and patch generation, ensuring that remediation solutions not only eliminate superficial symptoms but also maintain the functional integrity of the code. This layered design enables VSP to flexibly respond to changes in code size, vulnerability types, and context complexity. However, its effectiveness is highly dependent on the model's depth of understanding of code semantics and the quality of the prompting project.

[0074] Although VSP technology significantly improves the efficiency and accuracy of vulnerability analysis through vulnerability semantics guidance, its practical application still faces multiple limitations and has the following shortcomings: First, VSP technology is highly dependent on the reasoning capabilities of large language models, and existing large language models still lack a deep understanding of complex code semantics. For example, the model may miss vulnerabilities due to ignoring cross-function or cross-module data flows, or ignore the dynamic changes of loop boundary conditions in multi-layer nested control structures, resulting in false positives or missed detections. Secondly, the context input of VSP technology is limited by the token length constraints of the model, making it difficult to cover large-scale code bases or cross-file dependency scenarios. In real projects, vulnerabilities often involve global variables, external interfaces, or the interaction logic of distributed systems, and the input of local code snippets will result in the loss of key contextual information, which in turn affects the integrity of vulnerability analysis.

[0075] Furthermore, VSP's static semantic analysis framework struggles to capture dynamic runtime behavior, such as vulnerabilities caused by race conditions, memory fragmentation, or differences in environment configuration. Accurately locating these issues requires program execution traces or state monitoring, and VSP, which relies solely on static code snippets, may fail to identify these potential risks. Furthermore, while model-generated patch solutions can address surface vulnerability symptoms, they may not address fundamental design flaws or even introduce new vulnerabilities due to incomplete logic coverage.

[0076] A deeper challenge lies in the complexity of prompt engineering and the cost of model adaptation. Constructing high-quality vulnerability semantic examples requires the participation of security experts in the design to ensure the rigor and coverage of the reasoning chain. This is different from the responses of large language models to the same prompts, and the level of refinement of the reasoning logic needs to be adjusted according to the model architecture. In summary, VSP technology often faces problems such as contextual input limitations, forgotten prompt words, misjudgment of data flow paths, and semantic understanding deviations in specific tasks. These technical issues make the large-scale deployment of VSP technology face a high engineering threshold, especially when dealing with heterogeneous code bases or real-time analysis needs. Its efficiency and robustness still need to be further optimized.

[0077] Another prior art related to the present invention is the AutoSafeCoder multi-agent framework, a large language model-based multi-agent framework designed to improve the security of generated code through an iterative verification mechanism combining static analysis and dynamic testing. The framework's core design concept is to decompose the code generation process into three steps: function implementation, vulnerability detection, and dynamic verification, each performed by a coding agent (Coding Agent), a static analyzer agent (Static AnalyzerAgent), and a fuzzing agent (Fuzzing Agent).

[0078] The coding agent implements basic code generation based on the GPT-4 model. It generates initial code snippets after receiving natural language requirements. During the generation process, the agent uses few-shot learning techniques to infer the target functional logic by combining the code requirement description with partial source code. The generated code is then passed to the static analysis agent, which uses prompt engineering techniques to inject the MITRE CWE vulnerability knowledge base to conduct a security audit of the code.

[0079] The fuzz testing agent is responsible for vulnerability discovery during the dynamic verification phase. Its workflow consists of two phases: initial seed generation and mutation enhancement. The initial seed generator uses LLM to parse the code's semantic requirements and synthesize input samples that meet functional expectations. This seed is then dynamically enhanced using a type-aware mutation strategy. The generated mutated inputs are executed in a sandbox environment, detecting runtime vulnerabilities by monitoring program crash signals. The input that triggered the exception and its error context are recorded and fed back to the coding agent, driving targeted remediation.

[0080] The framework adopts a phased iterative strategy for collaborative mechanism design: During the static analysis phase, a maximum of four rounds of interaction are performed to balance efficiency and security, while during the dynamic testing phase, 150 rounds of mutation cycles are used to explore the boundary conditions of code execution paths. Interactions between agents are implemented using structured prompt templates.

[0081] This existing technology has the following shortcomings: Although the AutoSafeCoder multi-agent framework has achieved a significant reduction in vulnerability rates through a multi-agent collaboration mechanism, its technical paradigm still has systemic flaws. The serial collaboration mode between agents leads to a significant decrease in code generation efficiency. In particular, when dealing with complex scenarios such as cross-module dependencies or distributed system interactions, the alternating iterative process of static analysis and dynamic testing results in significant resource consumption and time delays. The static analysis agent is highly dependent on the preset CWE vulnerability rule library and exhibits significant lag when faced with new logical vulnerabilities. Its mechanized rule matching mechanism is unable to capture the hidden attack chains constructed by attackers through semantic ambiguity.

[0082] Secondly, the dynamic verification capabilities of fuzz testing agents are limited by the quality of the initial seed and the limitations of the mutation strategy. While type-aware mutation can generate diverse inputs, it lacks coverage of the code's deep semantic logic, resulting in blind spots in the detection of critical execution paths. The sandbox environment's execution monitoring mechanism can only capture explicit crash signals but cannot identify security vulnerabilities that do not cause program exceptions. More seriously, the vulnerability repair process exhibits "patching"-style local optimization characteristics. The coding agent tends to adopt pattern replacement rather than root-cause logic reconstruction, which may introduce new potential risks to the repair operation.

[0083] The deep-seated flaws stem from the cognitive limitations of the technical architecture: the framework separates syntax verification, semantic parsing, and pragmatic optimization into independent links, and fails to build a dynamic coordination mechanism for three-dimensional information. The static analysis agent only focuses on surface grammatical features such as dangerous function calls, and lacks semantic correlation analysis between code intent and business scenarios; the fuzz testing agent over-relies on input mutation strategies and ignores the pragmatic value assessment of runtime behavior. In addition, the static characteristics of the rule base and the lack of a knowledge update mechanism make it impossible for defense strategies to adapt to rapidly evolving new attack patterns. Compared with the present invention, the core problem of the AutoSafeCoder multi-agent framework is that security attributes are still imposed on the generation process as external constraints, rather than achieving "generation is immunity" through the endogenous integration of full information theory. Its mechanical collaboration mechanism and fragmented verification system are also difficult to support dynamic defense needs in high-security scenarios.

[0084] In order to overcome the above-mentioned defects of the prior art, the present invention proposes a method and system for automatic generation of security codes based on the universal intelligent generation mechanism model of full information theory and mechanism-based artificial intelligence, aiming to achieve the endogenous fusion of security attributes and dynamic strategy generation by reconstructing the deep collaborative mechanism of grammatical, semantic and pragmatic three-dimensional information.

[0085] This invention, through a fully information-theoretic verification framework, transforms the code generation process into a co-evolving ecosystem of "syntax, semantics, and pragmatics." It jointly models the code's grammatical features, semantic logic, and pragmatic constraints, addressing technical issues inherent in traditional approaches, such as externalized security attributes, fixed decision-making basis, and lack of interpretability. For example, when generating code involving sensitive operations, this invention not only verifies the syntactical compliance of function calls but also traces the propagation path of input parameters through a semantic risk map (composed of semantic risk vectors). It also dynamically assesses permission requirements and risk entropy values ​​in conjunction with a pragmatic optimization layer, thereby generating code logic that both conforms to formal specifications and embeds security constraints. Furthermore, based on the "information-knowledge-intelligence" transformation chain, it enables autonomous optimization of security policies. When new attack patterns are detected, it automatically extracts semantic features and pragmatic value indicators, updates the knowledge base, and dynamically adjusts the generated security policy, thus overcoming the lag bottleneck of traditional rule bases.

[0086] This invention can be applied to high-security demand scenarios such as industrial control systems and smart contract development. Through the integration of full information theory and mechanismism, it has the advantage of explainability, moving security attributes from the end link of code review to the generation stage, achieving the inherent security goal of "generation is immunity", and transforming code generation technology from passive defense to active cognition.

[0087] The semantic risk graph is explained and illustrated as follows: The semantic risk graph is the input to the pragmatic optimization layer. Its data structure is a weighted directed graph composed of semantic risk vectors. Semantic risk vectors refer to the node risk values ​​and edge weights in the graph. The semantic risk graph is generated each time code is parsed and is the instantiation result of the current code.

[0088] The relationship between the semantic risk graph and the knowledge graph is as follows: the knowledge graph is equivalent to a global knowledge base and is static; the semantic risk graph is generated based on each real-time parsed code and is dynamic. For example, the knowledge graph is used to define "what is potentially dangerous," while the semantic risk graph is used to quantify "how dangerous it is currently."

[0089] The preferred embodiments of the present invention will be described in further detail below with reference to the accompanying drawings.

[0090] like Figure 1 As shown, this embodiment provides a high-security code automatic generation framework, the core of which is to reconstruct the entire information conversion ecosystem. The syntax verification layer of the code breaks through the limitations of traditional formal tools. By enhancing the syntax model, the formal fragments of the code are mapped into semantic logic networks, which facilitates the construction of semantic logic networks across code fragments through the semantic parsing layer. The semantic logic network is the output of the syntax verification layer and the input of the semantic parsing layer. The semantic parsing layer of the code realizes the cognitive dimensionality upgrade of program intent and evaluates the command injection risk in combination with the context. For example, a semantic association model is used to form a semantic risk vector. The semantic association model is an intermediate quantity of the semantic parsing layer and is used to generate a semantic risk vector. The pragmatic optimization layer of the code directly faces the multi-dimensional contradictions in security decision-making, constructs a dynamic utility function, and realizes dynamic trade-offs between dimensions such as vulnerability exploitation complexity and asset value weight, making the security policy a continuous function with goals, environment and cost as variables, rather than a discrete rule.

[0091] Specifically, such as Figure 4 and Figure 5 As shown, this embodiment provides a method for automatically generating secure codes based on full information theory and mechanism theory, comprising the following steps:

[0092] Step S1: extracting features of the syntax verification layer for the input requirements. In the syntax verification layer, the syntax features of the code are extracted through the topological features of the abstract syntax tree, and the extracted syntax features are verified by formal rules to establish a standardized representation corresponding to the syntax features for subsequent processing;

[0093] Step S2: constructing semantic risk vectors across code snippets through the semantic parsing layer to form propagation paths for tracing data flows and control flows;

[0094] Step S3: In the pragmatic optimization layer, a security policy is generated through a dynamic utility function in combination with pragmatic constraints of the business scenario.

[0095] Step S4: Optimize the parameters of the dynamic utility function and the security strategy through knowledge federation evolution and closed-loop feedback mechanism.

[0096] The step S1 of this embodiment is used to realize feature extraction of the syntax verification layer. The step S1 of this embodiment preferably includes steps S101 to S104.

[0097] Step S101, first call the programming language specific parser, convert the input source code text into an abstract syntax tree, and traverse the abstract syntax tree based on the predefined security sensitivity rule base, and calculate the node weight through the preset node type weight table; then, mark the nodes with weights higher than the preset threshold as key syntax nodes; then, perform multi-dimensional feature analysis on the key syntax nodes to construct a structured feature set; finally, vectorize the extracted structured features to form a fixed-length syntax feature vector. Among them, the structured features of the structured feature set include data-type features, category-type features and descriptive features, the data-type features include control flow depth and data dependency, the category-type features include permission changes and variable pollution, the descriptive features include semantic labels, etc., all of the above structured features can be used as features corresponding to multi-dimensional feature analysis. In the process of vectorizing the extracted structured features, since the category-type features and the descriptive features are not numerical values, in order to facilitate vectorization coding, the category-type features are preferably one-hot encoded, and the descriptive features are preferably converted into vectors through pre-trained word embedding, and finally the syntax feature vector corresponding to the key syntax node is constructed. The security sensitivity rule base is a rule base that is predefined and configured as needed. The preset threshold is a pre-set weight threshold for determining whether a grammatical node is a key grammatical node and can be adjusted according to actual needs. A fixed-length grammatical feature vector refers to a pre-set grammatical feature vector of a fixed length.

[0098] In step S102, based on the preset formal verification rules, the constructed grammatical feature vector is statically verified to ensure that the code complies with the basic specifications of the programming language and generate a static verification result of the grammatical feature vector; taking the SQL injection defense scenario as an example, the grammatical verification layer uses regular expressions to screen the input spliced ​​keywords and generate a grammatical risk weight.

[0099] In step S103, a graph node entity is first created for each key syntax node based on the static verification results. These graph node entities not only inherit the original information of the abstract syntax tree but also integrate the risk attributes generated during the verification phase. The topological dependencies between key syntax nodes are then analyzed, and three types of core edge structures are constructed and their weights are calculated: control flow edges, data flow edges, and call relationship edges. The weights of control flow edges are dynamically calculated based on the complexity of the execution path, ensuring that high-frequency execution paths receive higher attention. The weights of data flow edges focus on risk propagation characteristics and are calculated using the risk propagation coefficient and contamination probability. The weights of call relationship edges are calculated based on function call frequency and parameter risk. To optimize the risk representation capabilities of the graph, this embodiment introduces a graph convolutional network for topological enhancement. This graph convolutional network performs third-order neighborhood feature aggregation to collect risk features of surrounding nodes. It then calculates the diffusion probability distribution of vulnerability impact using the Softmax activation function algorithm. Finally, the PageRank algorithm is used to locate key hub nodes, identify core nodes with the greatest risk radiation capacity, and complete the mapping of the syntax feature graph.

[0100] Step S104 converts the key grammatical nodes corresponding to the grammatical feature graph into predefined semantic labels, and then converts the grammatical dependencies into association paths with security semantics. Specifically, based on the categorical and descriptive features contained in the grammatical feature vectors of the key grammatical nodes, a pre-set semantic label mapping table is consulted to classify the key grammatical nodes into corresponding semantic types. For example, input variable nodes can be labeled as "external input," permission change nodes can be labeled as "permission-sensitive points," and function call nodes can be labeled as "potential execution entry points," thereby completing the grammatical-to-semantic label conversion. Subsequently, based on the three types of edge structures established in the graph (control flow edges, data flow edges, and call relationship edges), semantic reinterpretation is performed in combination with edge type and weight information. For example, the path connecting "external input" and "database query" in the data flow edge is labeled as an "input contamination propagation path," the path connecting "user operation" and "permission-sensitive points" in the call relationship edge is labeled as an "unauthorized call path," and the path representing an abnormal branch jump in the control flow edge is labeled as a "logical bypass path." During the semantic reinterpretation process, all grammatical dependencies are embedded with security semantics through label annotations and path naming without changing the original structure, so as to be converted into associated paths with security semantics, thereby constructing a semantic logic network and providing explanatory input for the subsequent semantic parsing layer.

[0101] The preset formal verification rules of this embodiment refer to pre-set rules or standards used in the formal verification process, which can be customized and adjusted according to actual conditions and needs. SQL refers to Structured Query Language.

[0102] In this embodiment, step S2 is used to implement logical association at the semantic parsing layer. Step S2 preferably includes steps S201 to S203.

[0103] Step S201: construct a semantic association model across code snippets based on the knowledge graph.

[0104] The semantic parsing layer receives the semantic logic network output by the syntax validation layer as input and constructs a semantic association model across code snippets based on the knowledge graph. This semantic logic network, generated by the syntax validation layer through abstract syntax tree parsing, static verification, and graph enhancement, already includes semantic labels for key syntax nodes and multiple risk paths, including control flow, data flow, and call relationships. Based on this, the semantic parsing layer further leverages global knowledge graph resources to semantically complete and structurally optimize the semantic labels and edge attributes of the nodes in the semantic logic network. Specifically, the construction process is as follows: Graph node entity alignment is first used to map the graph node entities in the network to structured knowledge in the knowledge graph, such as function semantics, attack patterns, and dependency templates. Specifically, each graph node entity is matched to different categories of structured knowledge items, such as function semantics, attack patterns, and dependency templates. Semantic consistency is then verified based on the context, resulting in a fused semantic graph structure. Subsequently, a graph neural network is used to embed model the fused semantic graph structure, resulting in a semantic association model that can be used for high-level semantic reasoning, providing a foundation for subsequent steps.

[0105] Step S202: Calling a graph neural network to analyze the constructed semantic association model to identify potential attack intentions.

[0106] This embodiment uses an existing graph convolutional network (GCN) model to analyze the semantic association model and identify potential attack intent. Each node in the semantic association model corresponds to a graph node entity of a key syntax node. Its node features include its syntax risk weight, semantic label, and contextual structure information. Edges represent control flow, data flow, and call relationships, and are assigned corresponding semantic weights. The GCN model captures high-order relationships in multi-hop dependency paths by aggregating and propagating features of node neighborhoods in the semantic graph structure, thereby identifying potential attack paths and behavior patterns. Finally, based on the aggregation results, it outputs attack intent labels such as information leakage, command injection, and access control bypass.

[0107] Step S203: Generate a semantic risk vector based on the identified potential attack intentions to quantify the logical hazard level of the security vulnerability.

[0108] This embodiment generates a semantic risk vector based on the potential attack intention labels and their associated paths output by the graph convolutional network in step S202, and uses it to quantify the logical hazard level of the security vulnerability. Specifically, the key grammatical nodes and their adjacent paths marked as attack-related are first extracted to construct a structural representation of the semantic attack chain. Subsequently, based on the attack intention category, the preset semantic risk mapping table is called to map each attack intention into risk indicators of multiple dimensions, including propagation range, dependency depth, data sensitivity, execution frequency, etc. Finally, the risk indicator values ​​of each dimension are combined into a fixed-length vector, which is a semantic risk vector, which is used to represent the comprehensive risk representation of the current potential vulnerability at the semantic level. This semantic risk vector can not only be used for multi-objective utility evaluation by the pragmatic optimization layer, but also has good traceability and interpretability, providing a structural basis for subsequent strategy generation and result auditing.

[0109] In step S201 of this embodiment, a semantic association model across code snippets is constructed. For example, user input nodes in SQL query statements are associated with database operation interfaces to form a data flow propagation path. When user input is detected as being directly embedded in an HTML document, the semantic parsing layer traces the propagation link from the input to the executable context through DOM tree analysis.

[0110] Therefore, in step S201 described in this embodiment, based on the knowledge graph, the user input nodes across code fragments and the database operation interface that calls the global knowledge graph resources are semantically associated through the semantic parsing layer to construct a semantic risk vector across code fragments.

[0111] In this embodiment, step S3 is used to implement dynamic policy generation of the pragmatic optimization layer, which generates security policies through dynamic utility functions. Step S3 preferably includes steps S301 to S303.

[0112] Step S301, a dynamic policy generation model is constructed in combination with the pragmatic constraints of the business scenario. The dynamic policy generation model preferably consists of three parts: pragmatic constraint modeling, utility function system, and optimization mechanism. First, the dynamic parameters in the business scenario feature library are parsed, and boundary conditions such as response delay and asset value threshold are extracted as rigid constraints, while combining elastic constraints such as resource consumption tolerance and user experience baseline. The pragmatic constraints described in this embodiment include preset rigid constraints and elastic constraints; the constraint parameters of these rigid constraints and elastic constraints are continuously updated through real-time monitors and integrated into the pragmatic constraint library. Then, the security goal is quantified as a risk attenuation rate function, the defense gain is dynamically calculated through the vulnerability probability of the semantic risk vector, the performance goal is converted into a resource loss function, and a normalized model is established based on the CPU / memory overhead of the policy simulation execution; the user experience goal is mapped to a logarithmic benefit function to capture the marginal diminishing characteristics of experience optimization. Each logarithmic benefit function dynamically adjusts the target priority through a weighting coefficient, and the weight value automatically switches with the scenario mode. Finally, a multi-objective evolutionary algorithm is used for iterative optimization. The Pareto front solution set is constructed through non-dominated sorting, and the optimal strategy is converged based on the utility function value. The strategy execution effect is fed back to the parameter optimizer to form a closed-loop tuning loop of the coefficient weights, thereby obtaining a dynamic strategy generation model.

[0113] Step S302: Call the risk entropy model to dynamically calculate the security level parameters , the formula of the risk entropy model is ,in It represents the triggering probability of the kth vulnerability, which comes from the semantic risk vector. It represents the pragmatic value weight of the associated asset, which comes from the pragmatic constraint library. Then calculate the security level parameter Baseline value: ,in The maximum entropy theoretical value of the current scene is added, and the scene coefficient is finally superimposed to generate the safety level parameter , , It represents the statically configured security baseline, and λ represents the dynamic adjustment coefficient of the scenario. The utility value of each strategy is calculated through the dynamic utility function. , the calculation process of the dynamic utility function is: ,in, 、 and Represent the coefficients under different business scenarios respectively; is the security level parameter obtained through dynamic calculation, and the dynamic utility function is obtained through Indicates the principle of safety first; represents the performance loss parameter, and the dynamic utility function is expressed by Represents computing resource constraints; represents the user experience parameter, and the dynamic utility function is expressed by Indicates the law of diminishing marginal utility of user experience. Performance loss parameter Refers to the computing resource overhead parameter, which is used to reflect the computing resource overhead introduced by the candidate strategy set during the deployment process. It can be customized and adjusted according to actual conditions, or it can be obtained through monitoring and calculation. The process of obtaining through monitoring and calculation is as follows: monitor the computing resource overhead during runtime and normalize the computing resource overhead to obtain the performance loss parameter The computing resource overhead includes performance indicators such as CPU usage, memory usage, and response delay. Preferably, the normalized range defaults to 0-1, where 0 represents no overhead and 1 represents maximum overhead. When performing normalized calculations, the weights of each performance indicator can be customized according to actual needs, such as through expert experience or historical data. User experience parameters Refers to the parameters used to reflect the impact of the strategy on user experience. Normalized calculations are performed based on user experience indicators such as the verification waiting time, the degree of increase in operation complexity, and the misjudgment rate introduced by the policy. The normalization range here also defaults to 0-1. In this case, 0 indicates no impact and 1 indicates the maximum impact. Preferably, the weights of various user experience indicators can be customized and adjusted according to actual needs, or can be set based on expert experience or historical data to fully reflect changes in user experience. Specifically, the verification waiting time refers to the time the user needs to wait during the verification process; the degree of increase in operation complexity refers to the proportion of the increase in the complexity of the user operation after the implementation of the policy relative to the original operation. For example, it can be measured by the number of steps required for the user to complete the task; the misjudgment rate refers to the probability of the system making an incorrect judgment after the implementation of the policy.

[0114] Step S303: Based on the result of the aforementioned dynamic utility function, the strategy with the largest utility value is selected from the candidate strategy set. 、 and They are weight coefficients associated with business scenarios, which are used to dynamically adjust the balance between security, performance and experience in different scenarios.

[0115] In this embodiment, the coefficients are not fixed configurations, but are dynamically set based on business environment characteristics, autonomous identification results, and knowledge base feedback. Table 1 shows examples of typical parameter configurations and default policy responses in different business scenarios.

[0116] Table 1 Coefficient settings for different business scenarios

[0117] Scene characteristics Parameter configuration Default policy response Low risk: Regular user access (0.6, 0.3, 0.5) IP address filtering + request frequency monitoring High risk: transfer payment operations (0.9, 0.2, 0.3) Mandatory MFA+semantic parameter verification Peak hours: System overload (0.7, 0.8, 0.2) Downgrade to token fast verification + asynchronous audit

[0118] In Table 1, scenario characteristics refer to the scenario characteristics corresponding to different business scenarios. For example, in low-risk business scenarios, the corresponding scenario characteristics are user routine access; in high-risk business scenarios, the corresponding scenario characteristics are transfer and payment operations; in peak-hour business scenarios, the corresponding scenario characteristics are system overload. The example lists the default corresponding parameter configurations for different business scenarios. The parameter configurations in the second column of Table 1 refer to 、 and The configuration of these three parameters, for example, in a low-risk business scenario, the parameter configuration is (0.6, 0.3, 0.5), which means , , , these parameters 、 and The configuration can be adjusted according to the actual situation and needs. In addition, the corresponding default policy responses are given, which refer to: In low-risk business scenarios, the default policy responses are IP address filtering and request frequency monitoring. IP address filtering refers to checking the source IP address of the network data packet to determine whether to allow the data packet to pass. Request frequency monitoring refers to monitoring and analyzing the user's request frequency. The frequency threshold can be set in advance. If the request frequency exceeds the frequency threshold, it will be automatically triggered to increase the coefficient. In high-risk business scenarios, the default policy response is mandatory MFA and semantic-level parameter verification. Mandatory MFA refers to mandatory multi-factor authentication, requiring users to provide two or more types of verification information to prove their identity. Semantic-level parameter verification refers to verifying the meaning and context of input data. In business scenarios during peak hours, the default policy response is downgrading to token fast verification and asynchronous auditing. Downgrading to token fast verification means temporarily lowering the level of security verification, that is, reducing the coefficient. ,Asynchronous auditing refers to analyzing and recording security events and ,detecting potential security threats.

[0119] In Table 1, scenario characteristics refer to typical usage scenarios in different business environments. For example, in low-risk business scenarios, scenario characteristics include routine user access operations; in high-risk business scenarios, scenario characteristics include transfer and payment operations; and in peak-hour business scenarios, scenario characteristics include system overload. For these different scenarios, the system has preset corresponding parameter configuration combinations and default policy response plans. These parameter configurations can be dynamically adjusted based on actual conditions and business needs to achieve an optimal balance between security, resource utilization, and user experience.

[0120] In low-risk business scenarios, the corresponding scenario characteristics are regular user access, etc. The system's default policy response is IP address filtering and request frequency monitoring. Among them, IP address filtering refers to checking the source IP address of the network data packet to determine its access legitimacy, thereby achieving static access control; request frequency monitoring refers to statistics and analysis of the frequency of user access requests. When the request rate exceeds the preset threshold, the system will automatically trigger the risk control mechanism and increase the safety factor accordingly. Since the system defaults to ensuring user experience and the response efficiency of basic resources in this scenario, the parameters are set to (0.6, 0.3, 0.5), which means To ensure the user experience, moderately control resource consumption and security strength.

[0121] In high-risk business scenarios, the corresponding scenario characteristics are transfer and payment operations, etc. The system puts security first, so the default policy response is mandatory multi-factor authentication (MFA) and semantic-level parameter verification. Mandatory MFA means requiring users to provide two or more authentication factors to complete login or transaction operations to prevent the risk of credential theft; semantic-level parameter verification refers to verifying user input content in the semantic dimension, such as analyzing whether the input is consistent with historical behavior, whether there is command injection intent, etc. The corresponding parameter configuration is (0.9, 0.2, 0.3), which significantly improves the security factor. weight, while moderately lowering , tolerate a certain degree of performance and experience sacrifice in exchange for higher security protection strength.

[0122] During peak hours, when the system faces resource constraints or concurrent pressure, the policy response focuses on ensuring the availability of the entire system. In this case, the default policy is to downgrade to token fast verification and asynchronous audit. The former means simplifying the authentication mechanism to a one-time fast authentication process based on tokens, temporarily reducing the verification intensity to alleviate the system computing pressure, and correspondingly reducing The latter uses asynchronous logging and post-event behavior analysis to conduct delayed audits and assessments of security incidents, ensuring uninterrupted system operation. The corresponding parameter configuration is (0.7, 0.8, 0.2), which significantly increases the resource weight. , and compress the proportion of experience and security to achieve overall stable operation of the system.

[0123] In step S302 of this embodiment, the dynamic utility function Item is used to directly reflect the principle of security priority. When the risk of attack increases, the system will increase the authentication strength to make the parameter Approaching 1, making the utility value Positive growth, at this time, the coefficient Dynamic adjustment based on asset value. If the asset value increases, the parameter becomes larger; if the asset value decreases, the parameter For example, in financial trading scenarios, the parameter The value of may be 0.9, while in internal office systems the parameter A possible value of is 0.6.

[0124] In the dynamic utility function The term reflects the computing resource constraints and the performance loss parameters Normalized calculation is performed based on the real-time monitoring system indicators. When the edge node load exceeds the preset threshold, the coefficient Automatically adjust the coefficient upwards to force the policy engine to prioritize lightweight authentication solutions. Dynamically adjust the coefficient downwards during peak business hours. The preset threshold refers to a preset edge node load threshold, which can be set and adjusted according to actual conditions.

[0125] In the dynamic utility function The item adopts the logarithmic function form, focusing on showing the law of diminishing marginal utility of user experience, and initially improving the user experience parameters Can significantly increase utility value, but when user experience parameters When it is close to 1, the benefits of further optimization slow down. For example, when the authentication delay is compressed from 2 seconds to 1 second, the user experience is significantly enhanced. When the delay is less than a certain delay threshold, that is, when the user experience parameter When it is close to 1, the user experience is not significantly enhanced. In the peak hours of the business scenario, the coefficient is dynamically lowered. , to prioritize system stability.

[0126] In step S303 of this embodiment, in the current business scenario, a candidate strategy set is generated according to the preset strategy response, and the utility values ​​of all feasible strategies are continuously calculated through the dynamic utility function. , select the utility value The largest strategy is used as the global optimal path. The preset strategy responses include only enhanced log auditing, mandatory MFA verification, and rate limiting and downgrading. MFA verification refers to multi-factor authentication.

[0127] For example, assuming that in the current business scenario, the coefficients are: , , When the model detects abnormal characteristics in a user behavior chain, the system generates a candidate strategy set based on the preset strategy response:

[0128] Strategy A: enforce MFA verification and improve security level parameters , correspondingly adjust the performance loss parameters , user experience parameters ;

[0129] Strategy B: only enhance log auditing and reduce security level parameters , correspondingly adjust the performance loss parameters , user experience parameters ;

[0130] Strategy C: perform current limiting and downgrade, lowering the security level parameters , correspondingly adjust the performance loss parameters , user experience parameters .

[0131] Substituting the coefficients of the business scenario and calculating through the dynamic utility function, we can get: the utility value corresponding to strategy A for ; The utility value corresponding to strategy B for ; The utility value corresponding to strategy C for .

[0132] Since the utility value corresponding to strategy A is Therefore, the model selects strategy A as the global optimal path, that is, the optimal solution. At this time, although its user experience is low, the security benefit dominates the decision.

[0133] Step S4 in this embodiment is used to realize knowledge federation evolution and strategy closed-loop optimization. Figure 5 As shown, this embodiment achieves continuous evolution of defense capabilities through distributed federated learning. After the generated code is deployed, continuous iteration of defense capabilities is achieved through the following paths: Figure 5 What is shown is the process of optimizing a closed-loop feedback mechanism.

[0134] Specifically, step S4 in this embodiment preferably includes steps S401 to S404.

[0135] In step S401 , the edge node extracts attack features and generates a knowledge vector based on the extracted attack features.

[0136] Edge nodes monitor attack signatures such as runtime logs, system calls, and abnormal behavior in real time. Upon identifying suspicious events, they initiate the attack signature extraction and knowledge vector generation process. This embodiment first uses an abstract syntax tree (AST) to extract structured grammatical features, such as dangerous functions and unusual splicing patterns. It then uses a program dependency graph (PDG) to track the control flow and data flow relationships between variables and sensitive operations, obtaining semantic dependency paths across functions. Furthermore, the pragmatic optimization layer performs a value-weighted analysis based on the business scenario, resource access sensitivity, and operation cost of the current operation, thereby determining the true impact of the attack behavior within the overall system context. During the feature fusion phase, the three types of features are mapped into a unified semantic graph structure, where key grammatical nodes serve as graph entities (i.e., graph node entities), semantic dependencies serve as edge connections, and pragmatic information is encoded as node weights or labels. A graph convolutional neural network (GCN) is used to perform high-level semantic encoding of this structure, combining the attack behavior label information with contextual features to generate a dense, fixed-dimensional knowledge vector. This knowledge vector refers to the attack knowledge vector.

[0137] In step S402 , the generated knowledge vector is uploaded to a global knowledge base in the cloud through a privacy protection protocol. The global knowledge base includes a semantic knowledge base and a grammar verification rule base.

[0138] Step S403: The semantic knowledge base updates the attack pattern graph and reversely injects the updated content into the syntax verification rule base of the edge node.

[0139] Step S404: Optimizing the parameters of the dynamic utility function and the security policy through a closed-loop feedback mechanism.

[0140] Through all the above steps, this embodiment can effectively overcome the limitations of the existing technology's separation of syntax, semantics, and pragmatics, forming an integrated "generate-verify-defense" architecture, ultimately achieving the automatic generation of highly secure code. Among them, the attack pattern graph is one of the core substructures in the semantic knowledge base. The attack pattern graph uses a semantic graph structure (nodes + edges) to express the semantic logic chain of attack actions, such as the nodes and their connections in the information collection, utilization, privilege escalation, and data leakage stages of the attack chain. This attack pattern graph is similar to the graphical representation of the ATT&CK matrix and is a semantic network of attack knowledge that supports model evolution and strategic reasoning.

[0141] The difference between the attack pattern graph and the knowledge graph described in this embodiment is as follows: The "knowledge graph" in step S2 is a local semantic graph constructed locally during the code generation phase, which is used to analyze semantic relationships such as code structure and variable dependencies, and to extract semantics from the code. The "semantic knowledge base" and "reverse injection" in step S4, on the other hand, are based on attack knowledge and reversely update the grammar verification rule base. That is, in step S4, the semantic knowledge base first identifies new attack path patterns based on the attack knowledge vectors aggregated from multiple edge nodes, such as a certain splicing structure that is often used to bypass verification, and then compresses these attack path patterns into concise grammar-level rules, such as adding a matching pattern for a certain function call combination, and sends it to the grammar verification rule base of the edge node to improve its initial interception capability.

[0142] To summarize, this embodiment first extracts features from the input requirements at the syntax verification layer, then constructs semantic risk vectors across code fragments through the semantic parsing layer to form a propagation path for tracking data flow and control flow. Next, at the pragmatic optimization layer, pragmatic constraints of the business scenario are combined to generate security policies through dynamic utility functions. Finally, through the knowledge federation evolution and closed-loop feedback mechanism, the parameters of the dynamic utility function and the security policy are optimized, thereby constructing a unified intelligent model based on full information theory and mechanismism, and realizing a high-security code automatic generation framework.

[0143] This embodiment reconstructs the deep collaborative mechanism and dynamic strategy generation model of grammatical, semantic and pragmatic three-dimensional information, jointly models the grammatical features, semantic logic and pragmatic constraints of the code, builds a full-information verification framework, transforms security logic from an external attached attribute to an endogenous cognitive ability of the system, and realizes vulnerability prevention in the code generation stage. Therefore, it can effectively realize the endogenous integration and autonomous evolution capability of security attributes, and improve the security, explainability and adaptability of automatic code generation.

[0144] Compared with the defects of mechanical segmentation of formal verification, semantic analysis and security constraints in traditional code generation technology, this embodiment transforms security logic from an external attribute to an endogenous cognitive ability of the system, realizing vulnerability prevention in the code generation stage, significantly improving the security, interpretability and adaptability of code generated by large language models.

[0145] Existing technologies rely on static rule bases or black-box models, making them difficult to cope with the dynamic evolution of complex attack scenarios and unable to trace the underlying logical links of security flaws, leaving defense systems in a very passive position. Unlike existing technologies, this embodiment, through the organic combination of a full information verification framework and a dynamic policy generation model, overcomes the irreconcilable contradictions between security, performance, and user experience in traditional technologies, providing a feasible theoretical foundation for high-security scenarios such as industrial control systems and smart contract development. Its beneficial effects are mainly reflected in the following aspects.

[0146] The first is the essential improvement of security. Traditional code generation technology is limited by the fragmented thinking of the material discipline paradigm, and its security verification focuses on compliance checks at the syntax layer, while ignoring logical vulnerabilities at the semantic layer and dynamic risks at the pragmatic layer. This embodiment achieves in-depth mining of hidden vulnerabilities through multi-dimensional feature extraction at the syntax verification layer in step S1 and intention recognition at the semantic parsing layer in step S2. Taking SQL injection attacks as an example, the regular matching filtering of traditional methods can only intercept explicit attack features, while this embodiment associates user input nodes with database operation interfaces through the semantic parsing layer, combined with the dynamic risk assessment of the pragmatic optimization layer, generates parameterized query statements and embeds them into the runtime sandbox isolation mechanism, thereby blocking the possibility of injection attacks from the root.

[0147] The mathematical model analysis shows that under the same attack load, this embodiment uses the formula Calculating the attack missed detection rate when automatically generating security codes , can effectively reduce the missed detection rate; among them, represents the missed detection rate of the traditional black box model, represents the full information synergistic enhancement coefficient, 、 and The security weights of the grammatical, semantic, and pragmatic dimensions are represented respectively. Therefore, this embodiment adopts the collaborative verification of three-dimensional information. The attacker needs to break through the triple constraints of formal compliance, logical legitimacy, and value rationality at the same time, and the attack cost increases exponentially.

[0148] The second is a breakthrough in explainability and adaptability. Existing code generation models generally have the problem of black box models being unexplainable, and the unexplainability of their security logic leads to vulnerability repair relying on manual reverse analysis. Unlike existing technologies, this embodiment constructs a white-box strategy generation path through the "information-knowledge-intelligence" transformation chain of mechanism-based artificial intelligence. The grammatical feature graph of the grammatical verification layer, the semantic risk vector of the semantic parsing layer, and the dynamic utility function parameters of the pragmatic optimization layer are all presented in a visual form, making the system's security decision-making basis traceable and auditable. This transparent mechanism can significantly reduce the difficulty of locating security defects and shorten the vulnerability repair cycle. In addition, this embodiment implements distributed knowledge evolution based on the federated learning protocol. The edge nodes refine local defense experience into knowledge vectors and synchronize them to the global knowledge base, thereby promoting the continuous evolution and adaptive adjustment of the grammatical verification rule base and the semantic risk vector, and enhancing the system's ability to respond quickly to new attacks.

[0149] The third is the dynamic balance between performance and utility. Traditional security solutions often face the dilemma of over-defense and under-defense: while strict formal verification can improve security, it also leads to performance degradation and a poor user experience. This embodiment achieves Pareto optimality of security policies through the dynamic utility function of the pragmatic optimization layer.

[0150] Fourth, the versatility of the method and the construction of an ecosystem. The core theoretical breakthrough of this embodiment is to define security attributes as intrinsic constraints of intelligent systems, rather than as externally imposed rule bases. This paradigm innovation enables the technical solution to have strong generalization capabilities. In the code generation scenario of industrial control systems, the separation of traditional functional safety and information security has led to the fragmentation of defense strategies. This embodiment generates composite code that integrates functional safety logic and information security mechanisms through the traceability of device control instructions at the semantic parsing layer and the production risk assessment at the pragmatic optimization layer. In the code generation scenario of IoT device firmware development, the system dynamically selects the encryption algorithm strength based on device resource constraints.

[0151] Even more profoundly, this embodiment lays the foundation for building an intelligent security ecosystem. Through human-machine cognitive collaboration, security experts can define defense objectives using natural language. The system automatically decomposes these objectives into grammatical rules, semantic constraints, and pragmatic weights, generating interpretable defense code. Simultaneously, the characteristics of new attack patterns are fed back to the experts, forming a closed loop of "local perception and global immunity." This ecological capability provides a strong foundation for the trusted evolution of the digital society.

[0152] This embodiment further provides a system for automatically generating secure codes based on full information theory and mechanismism, which adopts the above-mentioned method for automatically generating secure codes based on full information theory and mechanismism, and includes:

[0153] The feature extraction module of the syntax verification layer is used to extract the features of the syntax verification layer based on the input requirements. In the syntax verification layer, the syntax features of the code are extracted through the topological features of the abstract syntax tree, and the extracted syntax features are formalized and checked to establish a standardized representation corresponding to the syntax features.

[0154] The semantic association module of the semantic parsing layer builds a semantic association model across code snippets through the semantic parsing layer, forming a propagation path for tracing data flow and control flow;

[0155] The dynamic policy generation module of the pragmatic optimization layer generates security policies through dynamic utility functions in combination with pragmatic constraints of business scenarios.

[0156] The knowledge federation evolution and strategy closed-loop optimization module is used to optimize the parameters of the dynamic utility function and the security strategy through the knowledge federation evolution and closed-loop feedback mechanism.

[0157] This embodiment realizes the three-dimensional collaborative mechanism of the full information verification framework. By constructing a full information verification framework based on grammatical, semantic and pragmatic information, the grammatical verification layer extracts grammatical features through the topological features of the abstract syntax tree, and performs formal rule verification to establish a standardized representation of the code structure. The semantic parsing layer constructs a cross-module semantic association model based on the knowledge graph to track the propagation path of data flow and control flow; cross-module includes cross-code snippets, cross-functions, cross-classes or cross-code files, etc., collectively referred to as cross-modules. The pragmatic optimization layer combines the dynamic constraints of the business scenario to construct a multi-objective optimization utility function model (using a dynamic utility function) to quantify the comprehensive benefits of the security strategy.

[0158] This embodiment possesses the autonomous decision-making capabilities of a dynamic policy generation model. Based on the "information-knowledge-intelligence" transformation chain of mechanism-based artificial intelligence, a dynamic policy generation model is constructed using a dynamic utility function. This model implements distributed knowledge evolution through a federated learning protocol. The dynamic policy generation module uses a real-time knowledge base to perform multi-dimensional evaluations of the dynamic utility function and generate dynamic security policies tailored to the current attack scenario.

[0159] This embodiment implements a closed-loop feedback system for endogenous security attributes. This embodiment establishes a closed-loop feedback and self-evolution mechanism for security attributes. In existing technologies, remediation of security vulnerabilities relies on external manual intervention. However, this embodiment, through an integrated "generate-verify-defense" architecture, embeds security constraints during the code generation phase. This process achieves closed-loop optimization through an "information-knowledge-intelligence" transformation chain, and the execution results are fed back to the global knowledge base, driving the adaptive optimization of the parameters of the dynamic utility function, resulting in a continuous enhancement of defense capabilities.

[0160] In summary, this embodiment redefines the endogenous paradigm of code security through the innovative integration of full-information collaborative verification, dynamic strategy generation, and a closed-loop feedback system. It breaks through the paradigm shift of security logic from external rules to endogenous cognition, and provides a foundation for the automatic generation of high-security code and the improvement of defense capabilities.

[0161] The above is a further detailed description of the present invention in conjunction with specific preferred embodiments, and the specific implementation of the present invention should not be considered to be limited to these descriptions. For those skilled in the art of the present invention, without departing from the concept of the present invention, several simple deductions or substitutions can be made, which should be considered to fall within the scope of protection of the present invention.

Claims

1. A method for automatically generating secure codes based on full information theory and mechanism theory, characterized in that: The following steps are involved: Step S1, performing feature extraction of the syntax verification layer on the input requirements, in which the syntax features of the code are extracted through the topological features of the abstract syntax tree, and the extracted syntax features are verified by formal rules to establish a standardized representation corresponding to the syntax features; Step S2: constructing a semantic association model across code snippets through the semantic parsing layer to form a propagation path for tracing data flow and control flow; Step S3: In the pragmatic optimization layer, a security policy is generated through a dynamic utility function in combination with pragmatic constraints of the business scenario. Step S4: Optimizing the parameters of the dynamic utility function and the security policy through knowledge federation evolution and closed-loop feedback mechanism; The step S1 includes the following sub-steps: In step S101, the parser corresponding to the programming language is first called to convert the input source code text into an abstract syntax tree. The abstract syntax tree is then traversed based on a predefined security sensitivity rule base, and the weights of the nodes are calculated using a preset node type weight table. Nodes with weights above a preset threshold are then marked as key syntax nodes. A multi-dimensional feature analysis is then performed on the key syntax nodes to construct a structured feature set. Finally, the structured features in the structured feature set are vectorized and encoded to form a fixed-length grammatical feature vector; Step S102: statically verify the grammatical feature vector based on preset formal verification rules; In step S103, a graph node entity is first created for each key syntax node based on the static verification results. Then, the topological dependencies between the key syntax nodes are analyzed to construct three types of core edge structures, including control flow edges, data flow edges, and call relationship edges. Finally, a graph convolutional network is used to perform third-order neighborhood feature aggregation to collect risk features of surrounding nodes. The diffusion probability distribution of vulnerability impact is calculated using the Softmax activation function algorithm, and the PageRank algorithm is used to locate key syntax nodes. The core nodes with the greatest risk radiation capability are identified to complete the mapping of static verification results to syntax feature graphs. Step S104: convert the key grammatical nodes corresponding to the grammatical feature graph into predefined semantic labels, and then convert the grammatical dependency into an association path with security semantics; The step S2 includes the following sub-steps: Step S201: constructing a semantic association model across code snippets based on the knowledge graph; Step S202: Analyze the constructed semantic association model using a graph neural network to identify potential attack intentions. Step S203: Generate a semantic risk vector based on the identified potential attack intentions to quantify the logical hazard level of the security vulnerability.

2. The method for automatically generating secure codes based on full information theory and mechanism theory according to claim 1 is characterized in that: In step S201, based on the knowledge graph, semantic association is performed on user input nodes and database operation interfaces across code snippets through a semantic parsing layer to construct a semantic association model across code snippets.

3. The method for automatically generating secure codes based on full information theory and mechanism theory according to claim 1 or 2, characterized in that: The step S3 includes the following sub-steps: Step S301: construct a dynamic strategy generation model based on pragmatic constraints of business scenarios; Step S302: Call the risk entropy model and calculate the utility value of each strategy through the dynamic utility function , the calculation process of the dynamic utility function is: ,in, 、 and Represent the coefficients under different business scenarios respectively; represents the security level parameter, and the dynamic utility function is expressed by Indicates the principle of safety first; represents the performance loss parameter, and the dynamic utility function is expressed by Represents computing resource constraints; represents the user experience parameter, and the dynamic utility function is expressed by Represents the law of diminishing marginal utility of user experience; through the formula Calculate the safety level parameters , Indicates the dynamic adjustment coefficient of the business scenario, Indicates the security level parameter The baseline value, It represents the statically configured security baseline; Step S303: Select utility value The highest strategy.

4. The method for automatically generating secure codes based on full information theory and mechanism theory according to claim 3 is characterized in that: In step S302, when the attack risk increases, the security level parameter is increased by increasing the authentication strength. Approaching 1, ensuring utility value Positive growth, and the coefficient Dynamic adjustment based on asset value; performance loss parameters Normalized calculation is performed based on real-time monitoring. When the edge node load exceeds the preset threshold, the coefficient Automatically adjust upwards; dynamically adjust downwards during peak business hours .

5. The method for automatically generating secure codes based on full information theory and mechanism theory according to claim 3 is characterized in that: In step S303, in the current business scenario, a candidate strategy set is generated according to the preset strategy response, and the utility values ​​of all feasible strategies are continuously calculated through the dynamic utility function. , select the utility value The largest strategy is used as the global optimal path. The preset strategy responses include only enhanced log auditing, mandatory MFA verification, and rate limiting and downgrading. MFA verification refers to multi-factor authentication.

6. The method for automatically generating secure codes based on full information theory and mechanism theory according to claim 1 or 2, characterized in that: The step S4 includes the following sub-steps: Step S401: The edge node extracts attack features and generates a knowledge vector based on the extracted attack features; Step S402: Upload the generated knowledge vector to a global knowledge base in the cloud through a privacy protection protocol. The global knowledge base includes a semantic knowledge base and a grammar verification rule base. Step S403: The semantic knowledge base updates the attack pattern graph and reversely injects the updated content into the syntax verification rule base of the edge node; Step S404: Optimizing the parameters of the dynamic utility function and the security policy through a closed-loop feedback mechanism.

7. The method for automatically generating secure codes based on full information theory and mechanism theory according to claim 1 or 2, characterized in that: By formula Calculating the attack missed detection rate when automatically generating security codes ,in, represents the missed detection rate of the black box model, represents the full information synergistic enhancement coefficient, 、 and Represent the security weights of grammatical, semantic and pragmatic dimensions respectively.

8. A secure code automatic generation system based on full information theory and mechanism theory, characterized by: The method for automatically generating secure codes based on full information theory and mechanism theory as claimed in any one of claims 1 to 7 is adopted, and includes: The feature extraction module of the syntax verification layer is used to extract the features of the syntax verification layer based on the input requirements. In the syntax verification layer, the syntax features of the code are extracted through the topological features of the abstract syntax tree, and the extracted syntax features are formalized and checked to establish a standardized representation corresponding to the syntax features. The semantic association module of the semantic parsing layer builds a semantic association model across code snippets through the semantic parsing layer, forming a propagation path for tracing data flow and control flow; The dynamic policy generation module of the pragmatic optimization layer generates security policies through dynamic utility functions in combination with pragmatic constraints of business scenarios. The knowledge federation evolution and strategy closed-loop optimization module is used to optimize the parameters of the dynamic utility function and the security strategy through the knowledge federation evolution and closed-loop feedback mechanism.

Citation Information

Patent Citations

  • Method for constructing effectiveness verification architecture of capability evaluation index system based on grammar and semantic pragmatic

    CN118940741A

  • Intelligent retrieval method and system for genuine medicinal materials based on atlas

    CN120216741A