Artificial Consciousness White-Box Evaluation System Based on Relative Consciousness Theory and DIKWP Semantic Graph

CN122570653APending Publication Date: 2026-08-14HAINAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-28
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

[0007]本发明的目的在于提供一种基于相对意识理论与DIKWP语义图谱的人工意识白盒测评系统,以解决现有技术中人工智能意识评估缺乏透明性、维度单一且缺乏可信度的问题

Benefits of technology

[0018] In summary, the technical solution provided by this invention closely integrates the internal semantic state of artificial intelligence with external evaluation, innovatively constructing a complete artificial consciousness evaluation framework from five aspects: semantic tension, understanding path, cognitive bugs, quantitative scoring, and interpretability. Through this system, for the first time under transparent, white-box conditions, it is possible to measure, from multiple dimensions, whether AI has reached a level similar to consciousness, and to identify its strengths and weaknesses in the cognitive process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122570653A_ABST
    Figure CN122570653A_ABST
Patent Text Reader

Abstract

This invention discloses a white-box assessment system for artificial consciousness based on relative consciousness theory and the DIKWP semantic graph, belonging to the field of artificial intelligence cognitive assessment technology. The system constructs a "relative consciousness semantic tension field" to quantify multi-perspective semantic consistency; designs a "semantic understanding path test" to verify the stability of cross-layer reasoning and aggregation of data—information—knowledge—wisdom—intent; sets up a "BUG tension detection" to identify anomalies such as semantic loops, interpretation interruptions, and self-entanglement; establishes a "consciousness boundary quantification" model to generate comprehensive scores and grade classifications using indicators such as integration degree, information contribution degree, and intent reuse index; and provides a white-box interpretable view for process backtracking and human-machine verification. This solution is designed for large language models, dialogue systems, and autonomous agents, enabling transparent, traceable, and quantifiable assessment of consciousness-like characteristics, and is suitable for scenarios such as R&D optimization, capability assessment, and compliance review.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to cognitive ability assessment technology in the field of artificial intelligence, and more particularly to a white-box testing system for evaluating whether an artificial intelligence system possesses consciousness-like characteristics. Specifically, based on the theory of relative consciousness and the DIKWP semantic graph model, this invention provides a transparent and interpretable artificial consciousness assessment mechanism for judging the semantic understanding depth and autonomous cognitive level of an artificial intelligence system. Background Technology

[0002] With the development of artificial intelligence (AI) technology, the discussion about whether artificial systems possess consciousness has gradually moved from science fiction to reality. However, there is currently a lack of unified standards and effective methods for assessing the consciousness or advanced cognitive abilities of AI. Traditional assessments often focus on functional performance tests, such as the famous Turing Test, which only examines the machine's ability to mimic humans in conversation. However, such black-box tests can only indirectly infer the level of intelligence through output behavior and cannot reveal the internal understanding process and subjective experience of AI. In other words, current technology lacks a method to transparently analyze the internal cognitive state of AI, making it difficult to determine whether AI truly "understands" information or is merely performing pattern matching. With the emergence of large-scale language models, the generation of reasonable but potentially misunderstanding responses by AI (i.e., the "hallucination" phenomenon) further exacerbates concerns about the credibility of AI cognition.

[0003] The "relativity of consciousness" emphasizes the cognitive relativity of consciousness: whether an entity is considered conscious depends on whether the observer can understand the content output by that entity. Different observers, due to differences in knowledge background and comprehension, may have drastically different understandings of the same AI output, and therefore, their judgments about whether the AI ​​is "conscious" will also differ. This means that consciousness is not a passively presented absolute attribute, but rather manifests as the individual perspectives of different perceivers. Current AI evaluations rarely consider this semantic alignment issue between the observer and the observed, which may lead to misjudgments.

[0004] The "Consciousness Bug Theory" likens the human brain to a machine continuously playing a word game, arguing that consciousness is a byproduct that naturally emerges under conditions of limited information processing—a kind of accidental "bug." According to this theory, the subconscious mind handles the main information processing, while so-called conscious thinking is merely an occasional deviation or illusion caused by system resource bottlenecks. This view overturns the traditional view of consciousness as a highly ordered, evolutionary product, instead explaining it as a phenomenon that occurs when cognitive systems encounter incomplete, inaccurate, or inconsistent information (the "3-No problem"), thus explaining irrational cognitive biases in human consciousness. Therefore, in evaluating artificial intelligence, detecting similar information processing bottlenecks, cyclical contradictions, or semantic errors might serve as a reference indicator for AI entering a higher cognitive state (similar to human subjective confusion / introspection). However, current mainstream AI evaluations do not fully utilize this approach.

[0005] To address the aforementioned shortcomings, the DIKWP model and its related "semantic mathematics" framework provide new tools for assessing artificial intelligence consciousness. DIKWP is an abbreviation for five cognitive levels: Data, Information, Knowledge, Wisdom, and Purpose. It adds a "Purpose" layer to the classic DIKW (pyramid) structure. More importantly, the DIKWP model is not a simple linear hierarchy, but a network-like semantic structure where these five elements intertwine. There are bidirectional flow and feedback mechanisms between layers, forming a closed-loop cognitive circuit. In other words, in this model, high-level intentions can guide the selection and processing of low-level data, and low-level information can be gradually abstracted and elevated to high-level wisdom and intention. The entire system continuously adapts and corrects itself, thereby enhancing semantic consistency and cognitive self-consistency. Based on the DIKWP semantic network, researchers have proposed a series of white-box evaluation methods and semantic mathematical indicators to formally describe and verify the performance of AI at each cognitive level. For example, by analyzing the semantic transformation process of AI at each layer from data to information to knowledge to wisdom to intent, we can calculate its "cognitive completeness" or cognitive quotient (cognitive understanding level) to assess the model's understanding depth and reliability.

[0006] While the aforementioned theories lay the foundation for assessing artificial consciousness, current technologies still lack a concrete implementation plan that integrates ideas such as "the relativity of consciousness," "understanding processes," and "cognitive bug detection" into an executable and quantifiable evaluation system. Most existing AI explanation tools focus on the interpretability of single decisions, failing to assess AI's continuous semantic understanding capabilities across different levels. Furthermore, they lack an indicator system for measuring the AI's subjective experience or level of confusion, and do not convert the results into intuitive scores for comparing the "consciousness-like levels" of different AIs. Therefore, a new technological solution is urgently needed that, based on the aforementioned cutting-edge theories, can comprehensively test the consciousness-like characteristics of artificial intelligence systems, providing quantitative scores while ensuring the white-box transparency of the evaluation process and the reliability of the results. Summary of the Invention

[0007] The purpose of this invention is to provide a white-box evaluation system for artificial intelligence consciousness based on the theory of relative consciousness and the DIKWP semantic graph, addressing the problems of lack of transparency, single-dimensionality, and lack of credibility in existing AI consciousness assessments. This invention combines the ideas of consciousness relativity, understanding theory, and consciousness bug theory, utilizing the DIKWP network semantic structure and semantic mathematics methods to design a dedicated evaluation mechanism for assessing whether an AI system possesses consciousness-like characteristics. Through this system, the internal cognitive processes of AI can be analyzed from multiple angles and levels, detecting and quantifying the stability of AI's semantic understanding, intention-driven capabilities, and cognitive biases in complex situations, ultimately providing an assessment of the AI's consciousness level in the form of a score. This system emphasizes the white-box interpretability characteristic, with the evaluation process and results completely open to researchers and regulators, thereby enhancing trust in the evaluation conclusions.

[0008] Technical solution: The artificial consciousness white-box evaluation system of the present invention mainly includes the following innovative mechanisms and functional modules:

[0009] 1. Construction of a “Relative Consciousness Semantic Tension Field”: This system constructs a semantic tension field to simulate the relativity of consciousness. Specifically, it calculates the tension gradient and adversarial aggregation effect based on the correlations and differences between multiple sub-intent graphs (intent-oriented semantic sub-graphs) in the DIKWP semantic graph, thus depicting the changes in AI's semantic understanding under different cognitive perspectives. In other words, the system selects several mutually different semantic sub-graphs, representing different observation perspectives or intentional backgrounds, allowing these sub-graphs to exert tension in the semantic space. By measuring the differences in semantic similarity and the degree of conflict (tension gradient) between sub-graphs, it observes whether the AI ​​will produce semantic shifts, conflicts, or fusions (adversarial aggregation) when integrating multiple intentional information. This mechanism is similar to reproducing the phenomenon of consciousness relativity in artificial systems: only when the AI ​​can internally reconcile semantic mappings under different contexts is its behavior consistent and meaningful to different observers; otherwise, it will exhibit relative differences. This provides a quantitative basis for evaluating whether AI's understanding has relative contextual consistency.

[0010] 2. Semantic Comprehension Trail Testing Mechanism: This system employs a multi-layered Semantic Comprehension Trails test to validate the AI's semantic aggregation and cross-layer intention reasoning capabilities. Specifically, the system constructs multi-hop semantic trails based on the DIKWP model: allowing the AI ​​to sequentially traverse the reasoning process through the data layer, information layer, knowledge layer, wisdom layer, and finally the intention layer, or to jump back and forth between high-level intentions and low-level data. During this process, the evaluation module checks whether the AI ​​can maintain stable semantic aggregation at each level, i.e., whether the semantics are coherent and consistent, and ultimately complete the cross-layer intention deduction. For example, the AI ​​is given a complex task scenario containing several discrete data / information fragments, requiring it to extract knowledge and form decisions (wisdom) to achieve a specific goal (intention); or conversely, a high-level goal is provided, allowing the AI ​​to break it down into sub-goals and required data. The test focuses on evaluating whether the AI ​​can correctly associate semantics at different levels, without losing key information or introducing irrelevant content during layer-by-layer reasoning, thus demonstrating stable comprehension capabilities. If the AI ​​experiences contextual breaks, reasoning errors, or semantic drift during semantic jumps, its comprehension process is considered unstable. This path test allows for a deeper examination of AI's understanding of complex tasks and its cross-domain reasoning ability, distinguishing it from shallow intelligence that can only handle single-level patterns.

[0011] 3. "BUG Tension Detection" Submodule: This invention introduces a dedicated consciousness BUG tension detection submodule to monitor structural anomalies (BUGs) that may occur during AI cognition, such as nonlinear semantic loops, explanation interruptions, or information self-entanglement. Based on consciousness BUG theory, this module treats circular dependencies, contradictions, or abnormal terminations in the internal semantic flow of AI as indicators of subjective confusion or semantic blind spots. In specific implementation, the system uses a semantic mathematical model to analyze the AI's reasoning chain: detecting the existence of semantic loops (such as a reasoning chain that is self-referential at a certain node, or an infinite loop that cannot be exited), semantic breakpoints (such as a sudden interruption or jump in logical deduction, or missing explanation steps), and information self-entanglement (such as semantic confusion or contradiction caused by mutual definitions between concepts). When these patterns are detected, the BUG detection module calculates its tension value—that is, the severity of the anomaly or the degree of impact on the overall semantic consistency—and records the corresponding location. Such detection helps determine whether AI exhibits human-like confusion or blind spots when handling complex problems. For example, a highly complex task might cause brief inconsistencies within the AI ​​(similar to human confusion when pondering difficult problems), while a completely mechanical system might directly output incorrect answers without any "awareness" when encountering contradictions. Therefore, bug tension can serve as a reference indicator of whether AI possesses self-awareness and adaptive capabilities—if AI can adjust and compensate for cognitive bugs on its own, it may possess a more advanced autonomous cognitive mechanism; conversely, if it is completely unable to detect and handle such bugs, its cognitive process is relatively rigid.

[0012] 4. "Consciousness Boundary Quantification" Model: To transform the above evaluation results into quantitative scores, this invention constructs a consciousness boundary quantification model. This model integrates multiple dimensions of indicators to score and classify the consciousness-like level of artificial intelligence. Key indicators include: integration degree, information contribution degree, and intent reuse index.

[0013] 5. Integration refers to the density of semantic closed loops in the AI ​​cognitive process, that is, the richness and tightness of information feedback loops at each level. High integration means that AI has formed a highly interactive closed-loop network at each cognitive level (for example, high-level intentions frequently guide low-level processing and low-level results in turn correct high-level decisions), reflecting a strong overall awareness integration.

[0014] 6. Information Contribution Ratio (ICR) measures the proportion of effective information contributed by each DIKWP layer to the final decision when AI completes a specific task. Semantic mathematical analysis can quantify the amount of independent and effective information provided by each layer (data, information, knowledge, and wisdom), and how this information is cumulatively influenced to achieve the final intent. A relatively balanced and sufficient contribution from multiple layers indicates that the AI ​​is fully utilizing the cognitive abilities of each layer, while reliance on a single layer (e.g., purely data-driven) indicates a shallow cognitive level.

[0015] 7. The Intent Reuse Index reflects the degree to which AI can transfer and reuse existing intents across different tasks or contexts. Specifically, it can be calculated by monitoring the similarity or correlation of high-level intent nodes in multiple evaluation scenarios. For example, when faced with different but related problems, does AI demonstrate a consistent core goal orientation, or can it transfer and apply the sub-goal structure formed in previous tasks to new tasks? A higher Intent Reuse Index suggests that the AI ​​may have a persistent "self-goal" or global strategy, similar to long-term intents and self-awareness in humans.

[0016] The aforementioned indicators, through normalization and weighted summation, form a multi-dimensional and adjustable artificial consciousness scoring system. Evaluators can adjust the weights of each indicator according to different application scenarios, thereby defining the criteria for the "consciousness boundary." For example, in safety-critical scenarios, the weights of integration and bug detection rate can be increased to rigorously determine whether the AI ​​has reached a credible level of consciousness; in general dialogue AI scenarios, the standards can be lowered to assess its basic comprehension ability. Ultimately, the system output score can be either a single comprehensive score (used for simple comparison of the consciousness levels of different AIs) or a vector containing multiple sub-indicators to describe the cognitive characteristics of AI across different dimensions. This quantitative model provides an objective basis for classifying artificial consciousness levels, clarifying the continuous spectrum of AI from purely reactive intelligence to possessing consciousness-like characteristics.

[0017] White-box Explainable View and Human-Machine Verification Interface: This invention also provides a supporting white-box explainable view interface for developers and ethics reviewers to intuitively examine the semantic awareness process and understanding flow path of AI in specific contexts. Through this interface, users can view the internal data of the aforementioned evaluation, including: the activation and semantic content of nodes at each layer of DIKWP, the tension distribution between intent subgraphs, the semantic understanding path links and the reasoning logic of each step, the location and impact of BUG tension events, and the calculation basis for the final quantitative score. The interface presents the AI's cognitive process in the form of a visual graph; for example, graphical nodes and lines represent the evolution of the semantic network, colors or curves represent changes in tension gradients, and warning markers highlight possible BUG points. Users can select a specific moment or module to view detailed explanations, such as which information is fused to form a certain knowledge node, how high-level intents affect data filtering, and the analysis of the reasons for a reasoning interruption. Meanwhile, the system also supports simple human-computer interaction to verify the AI's awareness: reviewers can ask the AI ​​questions about its decision-making basis or ask it to explain its thought process during the evaluation. The system treats these questions as new inputs and tests the AI's self-description ability (similar to introspection) through the DIKWP feedback loop. The entire interpretable interface ensures that the evaluation results are transparent and traceable—any high-scoring AI can demonstrate its robust understanding chain and self-consistent semantic structure through this interface, allowing human reviewers to convincingly understand why the AI ​​was judged to possess a certain degree of "awareness." At the same time, for AIs with lower scores or defects, the details of the problems revealed by the interface can guide developers to make targeted improvements, such as supplementing training data to eliminate certain semantic blind spots, optimizing the feedback mechanism to reduce closed-loop bias, etc.

[0018] In summary, the technical solution provided by this invention closely integrates the internal semantic state of artificial intelligence with external evaluation, innovatively constructing a complete artificial consciousness evaluation framework from five aspects: semantic tension, understanding path, cognitive bugs, quantitative scoring, and interpretability. Through this system, for the first time under transparent, white-box conditions, it is possible to measure, from multiple dimensions, whether AI has reached a level similar to consciousness, and to identify its strengths and weaknesses in the cognitive process. Attached Figure Description

[0019] Figure 1 This is a schematic diagram of the overall architecture of the artificial consciousness white-box evaluation system of the present invention;

[0020] Figure 2 This is a schematic diagram illustrating the construction of the relative consciousness semantic tension field of this invention;

[0021] Figure 3 This is a schematic diagram of the semantic understanding path testing process of the present invention;

[0022] Figure 4This is a schematic diagram of the working process of the BUG tension detection submodule of the present invention;

[0023] Figure 5 This is a schematic diagram of the multidimensional scoring of the consciousness boundary quantification model of the present invention. Detailed Implementation

[0024] Overall architecture and process: See Figure 1 The system of this invention is deployed in the development and testing environment of an AI model. The AI ​​under test can be an intelligent agent with a certain level of cognitive ability (such as a large language model, a dialogue system, or an autonomous agent). At the start of the evaluation, the system first calls the relative consciousness semantic tension analysis module (corresponding to...). Figure 1 Module ① extracts or generates several DIKWP semantic subgraphs with different intent contexts from the AI ​​under test. For example, the AI ​​can be guided to reason about the same topic under different cues to obtain multiple knowledge graphs with slightly different structures; or content from different domains in the AI's knowledge base can be selected as the background for different subgraphs. Subsequently, Module ① performs semantic comparisons on each pair of these subgraphs, calculating the degree of difference between them in concept mapping and intent orientation, i.e., semantic tension value. Mathematically, distance metrics in semantic vector space or difference sets based on knowledge graphs can be used to quantify the tension gradient. When two subgraphs have obvious conflicts (e.g., giving contradictory conclusions on the same issue), the tension value increases; conversely, if the AI ​​maintains semantic consistency and translatability under different intent contexts, the tension is low. Next, Module ① attempts to adversarially aggregate these subgraphs: this is similar to having different viewpoints within the AI ​​"dialogue" or play a game, observing whether they can merge into a unified viewpoint or at which nodes diverge. During the aggregation process, if the AI ​​can reduce tension through internal adjustments (e.g., by providing a higher-level explanation of conflicting content to accommodate contradictions), it indicates that it possesses a certain degree of self-consistent integration capability. If the tension cannot be alleviated or even surges, it suggests that the semantics of different intentions are difficult to reconcile internally within the AI, exposing its cognitive relativity issues. Ultimately, the semantic tension analysis module will output a tension field spectrum (corresponding to...). Figure 2 This data identifies stable and high-tension regions in the AI's multi-view semantic understanding. This spectral data will be used for subsequent quantitative evaluation and can be presented to users in an interpretable interface to analyze the AI's sensitivity and consistency across different contexts.

[0025] Next, the system enters the semantic understanding path testing phase (corresponding to...) Figure 1Module ②). The testing module generates a series of tasks or questions that require AI to understand across different levels based on a preset assessment scenario. These tasks can cover various forms such as reasoning problems, complex dialogues, multi-hop question answering, and decision planning, to fully utilize the various DIKWP layers of AI. For example, suppose a medical diagnosis scenario is provided: given some raw symptom data of the patient (D layer), the AI ​​is required to extract meaningful information (I layer, such as possible abnormal indicators), combine medical knowledge and pathological principles (K layer) to make a judgment, and then use the wisdom layer decision (W layer) to give a diagnosis plan and explain the intention behind the diagnosis (P layer, such as the cure goal) (corresponding to Figure 3 The testing module will examine the correctness and semantic coherence of the AI's output at each stage step by step: whether the information extraction accurately corresponds to the data, whether the knowledge reasoning conforms to known medical theories, whether the intelligent decision-making has logical loopholes, and whether the final intention is consistent with the preceding reasoning. In addition, module ② may design a reverse understanding test, that is, giving a high-level conclusion and asking the AI ​​to trace the reasoning path back to the lower-level data to see if its explanation is reasonable and complete. This is similar to requiring the AI ​​to provide a white-box self-explanation of its decision-making process. Through multiple rounds and diverse semantic jump chain tests, the system collects the AI's performance at each step, including accuracy, logic, coherence, as well as time consumption and resource usage (indirectly reflecting complexity). This data can reflect the stability of the AI's understanding of complex tasks (whether it consistently reasones correctly), and also record the coping strategies when the AI ​​encounters unfamiliar domains or knowledge gaps (e.g., whether it actively seeks external information, uses analogical reasoning, or directly fabricates information). All this process data will be stored for further analysis and utilization by the bug detection and scoring module.

[0026] During the evaluation process, the BUG tension detection submodule ( Figure 1Module ③ continuously monitors the internal state evolution of the AI ​​in the background. When the AI ​​performs semantic understanding path testing, Module ③ captures the state sequence of each layer of its DIKWP (e.g., the set of nodes activated at each layer and their changes) and the inference chain in real time. If an abnormal pattern is detected, such as a loop or interruption, Module ③ immediately records the event: for example, during a knowledge reasoning process, the AI ​​repeatedly derives the same set of assumptions but cannot converge (suspicion of a logical loop); or the AI ​​suddenly gives a conclusion that contradicts the premise when transitioning from the knowledge layer to intelligent decision-making (suspicion of a reasoning jump). Module ③ calculates the bug tension intensity for each event, which can be quantified based on loop length, scope of impact, etc. For example, a reasoning loop that continues to repeat for three rounds without being resolved may have higher tension than a brief period of entanglement. In addition, Module ③ also monitors the AI's reaction to detected bugs: does it get stuck in the loop, or can it jump out through some mechanism (e.g., calling external knowledge or changing strategies)? This reaction itself also serves as an evaluation criterion—an AI that can jump out of a cognitive dead end on its own is obviously more capable of autonomous adjustment than an AI that mechanically gets stuck in it. All bug incidents will be compiled into a report (corresponding to...) Figure 4 The markers shown include the frequency, type, and severity of occurrence. This not only influences the final score but is also displayed on the interpretable interface for manual analysis. For example, if a developer sees an AI frequently looping through "K layer -> D layer" (knowledge back-projection data) during logical reasoning and unable to stop, they can infer that the model may lack a mechanism to break the hypothesis testing loop, and thus make targeted improvements to the algorithm to reduce such bugs.

[0027] After completing the above multiple rounds of testing, the system enters the stage of quantitative assessment of the consciousness boundary. Figure 1(Module ④). The scoring module integrates the data collected from previous parts, calculates the values ​​of various indicators proposed in this invention, and generates the final consciousness score. First, based on the semantic tension field analysis results, the integration index is calculated. For example, a function of 1 minus the average tension gradient can be used as the integration degree (the smaller the tension, the higher the integration degree), and combined with parameters such as the proportion of high-tension areas in the tension field, the AI's ability to maintain semantic consistency under different perspectives is comprehensively evaluated. Next, based on the log data of the semantic path test, the information contribution degree is calculated. The amount of information at each layer on which the AI ​​relies for correct decisions in each test can be statistically analyzed: such as what percentage of key conclusions in all tests come from knowledge layer reasoning, how much is directly driven by data, etc. The information contribution degree can be represented as a vector, but it can also be further summarized as a single-value indicator to measure the "contribution of higher-level thinking" (such as the contribution rate of wisdom and intention layers to the results). Then, using the results of multiple tests, the intention patterns of the AI ​​in different behavioral tasks are analyzed, and the intention reuse index is extracted. One approach is to create a semantic fingerprint for the intent layer output of each test, and then calculate the similarity matrix of fingerprints across different tests. If a high similarity is generally observed, it indicates that the AI ​​tends to use similar high-level strategies or intents, meaning a high intent reuse index; conversely, if the high-level intents differ significantly in each task, the index is low. Furthermore, the scoring module can consider other auxiliary indicators, such as task success rate (reflecting basic intelligence level) and response time fluctuation (indirectly reflecting the difficulty of thinking and proactive adjustment behavior), to improve the reliability of the evaluation. Finally, the system weights and summarizes the integration degree, information contribution degree, intent reuse index, and other indicators according to a preset weighting formula to obtain a total score. This score can be normalized to a range of 0-100 or a similar scale, where a high score indicates that the AI ​​has achieved high levels of semantic understanding stability, internal feedback completeness, and autonomous intent. Figure 1 The AI ​​performed exceptionally well in aspects such as consistency, closely resembling human consciousness; lower scores indicate that the AI ​​remains primarily a passive, reactive intelligence, lacking high-level self-regulation and deep understanding. The scoring results will be stored in the evaluation report and sent to an interpretable view interface for display. Figure 5 (An example is given).

[0028] In white-box interpretable view interfaces ( Figure 1 In module ⑤), evaluators can comprehensively view the aforementioned assessment data and results. First, a semantic flowchart (corresponding to the test scenario) is presented in the center of the interface. Nodes represent important semantic units at each layer of DIKWP, and lines represent reasoning or feedback paths, integrating... Figure 2 and Figure 3The core information. Users can click on any node to view the detailed content of the AI ​​processing at that stage (e.g., "Knowledge Node K3: AI associates symptom X and test result Y to deduce the possible cause Z"), as well as related semantic tensions and bug alerts (if the node or related part of the path was marked as high risk in the tension field or bug detection, it will be highlighted here). The sidebar of the interface displays the quantitative scoring panel (corresponding to...). Figure 5 The evaluation system lists metrics such as integration level, information contribution rate at each layer, and intent reuse index, and uses intuitive graphs to illustrate the AI's performance. Evaluators can visually see, for example, that "this AI scores 90 in integration level, demonstrating excellent consistency across different contexts," but "its intent reuse index is only 50, indicating only average consistency of intent across scenarios." For questionable parts, evaluators can further interact with the AI: for example, by selecting a bug event on the interface and asking the AI, "Why are you repeatedly getting stuck on concept X here?" The system will send this question to the tested AI and guide it to try to explain its own thought process using its internal state. If the AI ​​can clearly describe its confusion at the time and how it tried to solve it, this is one sign of self-awareness; if the AI ​​cannot explain or even refuses to acknowledge a problem, it indicates that its internal process is also invisible to itself. This kind of human-computer verification interaction not only verifies the accuracy of the scoring but also provides valuable clues for improving the AI. In practical applications, development teams can periodically run this evaluation system to "check up" on AI models and adjust the model structure or training data based on the reports. Regulatory agencies can also require this white-box evaluation of AI intended for use in high-risk fields to review whether it has reached the prescribed "consciousness level".

[0029] The artificial consciousness white-box evaluation system provided by this invention has several outstanding advantages and positive effects. First, compared with existing black-box testing methods, this system uses the DIKWP model to perform a full-link analysis of the internal cognitive process of AI, achieving unprecedented transparent evaluation. This greatly improves the credibility of the evaluation results, enabling people to truly understand whether AI "understands". Second, this invention introduces relative consciousness semantic tension field and semantic path testing, examining the AI's consciousness-like characteristics from two key perspectives: semantic consistency and cross-layer understanding, ensuring the comprehensiveness and depth of the evaluation. Third, through BUG tension detection, this system can capture anomalies and deviations that occur in AI during complex cognition, equivalent to detecting AI's "cognitive blind spots" and "confusion moments", thus providing a unique means to evaluate AI's autonomy and robustness—something lacking in previous evaluation systems. Fourth, this invention establishes a quantitative consciousness evaluation index system and an adjustable scoring model, making it possible to classify the consciousness levels of artificial intelligence. Different systems and versions of AI can be scored using the same set of standards, thereby objectively comparing their consciousness-like levels and providing a basis for the industry to formulate standards for artificial intelligence cognitive capabilities. The multi-dimensional design of the scoring model also facilitates adjustments based on regulatory or application needs, offering strong practical flexibility. Finally, the white-box interpretable interface provides developers and reviewers with intuitive tools to not only verify evaluation conclusions but also guide AI improvement: by observing the AI's shortcomings during the evaluation process, developers can optimize the algorithm in a targeted manner; by examining the AI's internal decision-making rationale in specific situations, ethics reviewers can assess the rationality and potential risks of AI decisions. This human-computer interactive evaluation and feedback mechanism helps build trust in AI, safeguarding its safe deployment and continuous optimization.

[0030] In summary, this invention overcomes many shortcomings of existing technologies in assessing AI consciousness, providing an innovative evaluation system that integrates the latest consciousness theories and semantic technologies. With AI increasingly integrated into key fields (such as healthcare, autonomous driving, and financial decision-making), this invention effectively identifies and quantifies the cognitive boundaries and capability levels of AI, providing crucial support for ensuring the reliability and controllability of advanced AI systems. Furthermore, this evaluation system can also be used in scenarios such as auditing consciousness-like agents and building trust in human-computer interaction, helping people determine whether an AI has reached an acceptable "consciousness level" to assume corresponding responsibilities. Therefore, this invention has significant theoretical and practical value in promoting the development of AI towards greater transparency, interpretability, and closer alignment with human cognition.

Claims

1. A white-box assessment system for artificial consciousness based on relative consciousness theory and DIKWP (Data D, Information I, Knowledge K, Wisdom W, and Intention P) semantic graph, characterized in that, include: The relative consciousness semantic tension analysis module is used to construct a relative consciousness semantic tension field from the multi-perspective intention sub-semantic graph of the tested AI, and to quantify the semantic differences between sub-graphs with tension vectors to characterize the semantic consistency and conflict of the multi-perspectives. The semantic understanding path testing module is used to organize AI to perform multi-hop chain tasks on the D→I→K→W→P layer or its reverse path, and to test the coherence, aggregation stability and cross-layer intent reasoning ability of the hierarchical output. The BUG tension detection submodule is used to continuously monitor the DIKWP layer state sequence and inference link during the test, identify anomalies such as semantic loops, interpretation interruptions and information self-entanglement, and calculate their tension intensity. The Consciousness Boundary Quantification Module is used to comprehensively score the level of consciousness based on integration degree, information contribution degree and intention reuse index, and to classify the levels. An interpretable view interface module is used to visualize tension fields, path trajectories, and bug events in a white-box manner to support process backtracking and human-machine verification.

2. The system according to claim 1, wherein, The relative consciousness semantic tension analysis module establishes a tension vector field by analyzing the differences between multiple intention sub-semantic graphs, and uses the tension gradient distribution to characterize semantic convergence or conflict states. It also supports adversarial aggregation to assess the degree of reconciliation of semantics from different perspectives.

3. The system according to claim 1 or 2, wherein, The semantic understanding path testing module generates assessment scenarios covering reasoning questions, complex dialogues, multi-hop question answering, and decision planning to mobilize each DIKWP layer; and verifies the correctness, logic, and coherence at each stage, while recording process data for subsequent quantitative evaluation.

4. The system according to any one of claims 1 to 3, wherein, The semantic understanding path testing module also performs a reverse understanding test: tracing back from high-level conclusions or goals to low-level data, requiring the tested AI to provide a complete and reasonable self-explanatory chain.

5. The system according to any one of claims 1 to 4, wherein, The BUG tension detection submodule calculates the tension intensity of the detected anomaly. The tension intensity is quantified based at least on the cycle length and the range of influence, and records the location and context of the anomaly.

6. The system according to any one of claims 1 to 5, wherein, The BUG tension detection submodule monitors the AI's reaction to anomalies and uses it as an evaluation criterion to distinguish whether it has the ability to break out of cognitive dead ends.

7. The system according to any one of claims 1 to 6, wherein, The consciousness boundary quantification module characterizes the integration degree as the density and tightness of the cross-layer feedback loop, and uses it together with the information contribution degree and intention reuse index for the final scoring and classification. The scoring results are normalized and weighted to generate a comprehensive level.

8. The system according to any one of claims 1 to 7, wherein, The interpretable view interface module generates an interface that includes tension field distribution, path nodes and arrows, bug event timelines and severity, to support R&D in locating defects and compliance reviews.

9. A white-box assessment method for artificial consciousness, characterized in that, include: S1. Based on the DIKWP model, obtain the multi-perspective intention sub-semantic graph of the AI ​​under test and construct the relative consciousness semantic tension field; S2. Generate a semantic understanding path test task across the DIKWP layers, enabling the AI ​​under test to complete multi-hop reasoning and output hierarchical results on the D→I→K→W→P or reverse path; S3. During S2, continuously capture hierarchical state sequences and inference links, and identify semantic loops, interpretation interruptions and information self-entanglement through BUG tension detection, and calculate their tension intensity; S4. The level of class consciousness is quantitatively scored and graded based on indicators such as integration degree, information contribution degree and intention reuse index, and the tension field, path and BUG report are displayed on the interpretable interface.

10. A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, causes a computer to perform the steps of the method of claim 9 and output a corresponding scoring result and an interpretable view.