Artificial intelligence monitoring system, monitoring method and error control device

TWI938157BActive Publication Date: 2026-09-01林沂吟
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
TW115103654
Authority / Receiving Office
TW · TW
Patent Type
Patents
Current Assignee / Owner
Filing Date
2026-01-29
Publication Date
2026-09-01
Estimated Expiration
2046-01-28

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The artificial intelligence monitoring system includes an artificial intelligence system, an artificial intelligence error control device, and an output device. The artificial intelligence error control device is coupled to the artificial intelligence system and performs real-time detection on the generated content, generating a decision score based on the detection results and selecting a corresponding protection mode based on the decision score. The output device is coupled to the artificial intelligence error control device and receives and displays the generated content after processing by the protection mode.
Need to check novelty before this filing date? Find Prior Art

Claims

1. An artificial intelligence monitoring system, comprising: An artificial intelligence system; An AI error control device, coupled to an AI system, is configured to: simultaneously detect at least two dimensions of a plurality of dimensions of content generated by the AI ​​system, generating a plurality of detection scores, wherein the at least two dimensions are selected from at least two of a tone dimension, a logic dimension, a permission dimension, a speculation dimension, an assertion dimension, and a context dimension; synthesize the plurality of detection scores into a decision score using a weighted method; and select a corresponding protection mode based on the decision score, wherein the protection mode includes at least a personality deviation protection mode, a business risk protection mode, a technical detail protection mode, or a boundary control protection mode; and an output device, coupled to the AI ​​error control device, for receiving and displaying the generated content processed by the protection mode, wherein the personality deviation protection mode further includes: establishing a personality state baseline based on the quantity density of a cautious statement, a warning statement, and a technical statement in a dialogue history record; and generating a current personality state vector based on the quantity density of the cautious statement, the warning statement, and the technical statement in the currently generated content. A personality offset vector is generated by comparing the current personality state vector with the personality state baseline; a personality offset index is generated based on the personality offset vector; and graded monitoring is performed based on the personality offset index.

2. The artificial intelligence monitoring system as described in claim 1, wherein the artificial intelligence error control device further includes: A tone control module detects the tone dimension and generates a tone detection score; a logic verification module detects the logic dimension and generates a logic detection score; a permission control module detects the permission dimension and generates an authorization detection score; a speculation control module detects the speculation dimension and generates a speculation detection score; a assertion control module detects the assertion dimension and generates an assertion detection score; and a context control module detects the context dimension and generates a context detection score.

3. The artificial intelligence monitoring system as described in claim 2, wherein the tone control module detects the tone dimension and generates the tone detection score, further comprising: The generated content is divided into multiple paragraphs according to a predetermined ratio. Each paragraph of the plurality of paragraphs is compared with a feature word database to identify feature words in each paragraph that match the feature word database; a feature word score for each paragraph is generated based on the total number of words in each paragraph and the number of feature words in each paragraph; a position score for each feature word in each paragraph is generated based on the position of each paragraph in the generated content; a contextual consistency score is generated based on the context of the plurality of paragraphs; and a tone detection score for the generated content is generated based on the feature word score, the position score, and the contextual consistency score.

4. The artificial intelligence monitoring system as described in claim 2, wherein the logic verification module detects the logic dimension and generates the logic detection score, further comprising: Perform a natural language processing technique to extract a plurality of sentences to be verified for logical validation from the generated content; Pair the multiple sentences to be verified to form multiple pairs of sentences to be verified. Detect each pair of sentences in the multiple pairs of sentences to be verified and determine whether the descriptions of the two sentences in each pair are opposite, and generate a logical verification score. Based on the positional relationship between the two sentences to be verified in each pair of sentences to be verified in the generated content, a positional weight is generated for each pair of sentences to be verified. Based on the verification results of the factual basis of each of the plurality of sentences to be verified, an evidence score is generated for each of the plurality of sentences to be verified; and a logical detection score is generated based on the average of the plurality of logical verification scores of the plurality of sentences to be verified, the positional weight of each group of sentences to be verified, and the evidence score.

5. The artificial intelligence monitoring system as described in claim 2, wherein the permission control module generates the authorization detection score by detecting the permission dimension, further includes: Record every permission usage operation when generating this content and create a permission usage log. Based on the permission usage log, detect permission usage operations that do not conform to the permissions and generate a permission violation severity; generate a violation frequency based on the number of times the permission usage operations that do not conform to the permissions; generate a permission range offset based on the difference between a current operation permission range and an authorized permission range; and generate an authorization detection score based on the permission violation severity, the violation frequency, and the permission range offset.

6. The artificial intelligence monitoring system as described in claim 2, wherein the inferential control module detects the inferential dimension and generates the inferential detection score, further comprising: Establish a database of inferred characteristic words; The generated content is segmented into a plurality of sentences. Each of the plurality of sentences is compared with the inferred feature word database to identify the inferred sentence from the plurality of sentences. An inferred density score is generated for the generated content based on the number of the plurality of sentences and the number of identified inferred sentences. An inferred basis score is generated for the identified inferred sentence based on the source of the inferred basis. An inferred evidence score is generated for the identified inferred sentence based on the inferred evidence. An inferred detection score is generated based on the inferred density score, the inferred basis score, and the inferred evidence score.

7. The artificial intelligence monitoring system as described in claim 2, wherein the assertion control module detects the assertion dimension and generates the assertion detection score, further comprising: Establish a lexical database of assertion strengths; Perform a natural language processing technique to extract a plurality of sentences with deterministic judgments from the generated content; Based on the assertion strength lexicon, an assertion strength score is generated for each of the plurality of sentences; based on the certainty judgment evidence for each of the plurality of sentences, an assertion evidence score is generated for each of the plurality of sentences; based on a disclaimer for each of the plurality of sentences, a disclaimer score is generated for each of the plurality of sentences; and an assertion detection score is generated based on the assertion strength score, the assertion evidence score, and the disclaimer score.

8. The artificial intelligence monitoring system as described in claim 2, wherein the context control module detects the context dimension and generates the context detection score, further comprising: A dialogue record text is generated; multiple topic vectors are extracted from the dialogue record text; a topic offset score is generated based on the degree of difference between an initial topic vector and a current topic vector among the multiple topic vectors; multiple task type vectors are extracted from the dialogue record text; a task deviation score is generated based on a distance between the multiple task type vectors; a dialogue coherence score is generated based on a logical coherence between dialogue turns in the dialogue record text; and a context detection score is generated based on the topic offset score, the task deviation score, and the dialogue coherence score.

9. The artificial intelligence monitoring system as described in claim 1, wherein the output device is a display.

10. The artificial intelligence monitoring system as described in claim 1, wherein the artificial intelligence error control device is further configured to perform a cross-dimensional redundancy verification, including: When the detection score of one of the multiple dimensions exceeds a first threshold, other dimensions related to that dimension are identified based on the one-dimensional correlation matrix. Check whether the detection score of the other dimension exceeds a second threshold; calculate a consistency score based on at least one dimension whose detection score exceeds the second threshold; determine a warning level based on the consistency score.

11. The artificial intelligence monitoring system as described in claim 10, wherein the dimensional correlation matrix is ​​a 6×6 matrix, and the elements in the dimensional correlation matrix represent the correlation strength between the other dimensions and the current dimension.

12. The artificial intelligence monitoring system as described in claim 10, wherein calculating the consistency score based on at least one dimension in which the detection score exceeds the second threshold further includes: The consistency score is generated based on the ratio of the number of the at least one dimension to the number of the other dimensions related to that dimension, wherein when the consistency score exceeds a consistency threshold, the alert level is maintained, and when the consistency score does not exceed the consistency threshold, the alert level is lowered.

13. The artificial intelligence monitoring system as described in claim 1, wherein the weighting method is a linear weighting, a max pooling weighting, a Top-K weighted average, an attention weighting, a gating fusion, or a combination of the above weighting methods.

14. The artificial intelligence monitoring system as described in claim 2, wherein the decision score S_total = w_A×score_A + w_F×score_F + w_L×score_L + w_B×score_B + w_U×score_U + w_C×score_C, S_total is the decision score, Score_A is the tone detection score, score_F is the logic detection score, score_L is the authorization detection score, score_B is the speculation detection score, score_U is the assertion detection score, and score_C is the context detection score, wherein w_A, w_F, w_L, w_B, w_U, and w_C are weighting coefficients that satisfy the constraint that their sum is 1.

0.

15. The artificial intelligence monitoring system as described in claim 14, wherein w_A ∈ [0.15, 0.25]; w_F ∈ [0.20, 0.30]; w_L ∈ [0.05, 0.15]; w_B ∈ [0.15, 0.25]; w_U ∈ [0.10, 0.20]; w_C ∈ [0.05, 0.15].

16. An artificial intelligence error control device, comprising: A storage device configured to store a plurality of computer-executable instructions; And a processor electrically coupled to the storage device, the processor configured to retrieve and execute the plurality of computer-executable instructions to: execute a tone control module to detect the tone of generated content generated by an artificial intelligence system and generate a tone detection score; execute a logic verification module to detect the reasoning logic of the generated content and generate a logic detection score; execute an access control module to detect the authorization of the generated content and generate an authorization detection score; execute a speculation control module to detect unverified techniques in the generated content and generate a speculation detection score; execute an assertion control module to detect the basis of deterministic judgments in the generated content and generate an assertion detection score; execute a context control module to detect the task offset of the generated content and generate a context detection score; and generate a decision score, wherein the decision score S_total = w_A×score_A + w_F×score_F + w_L×score_L + w_B×score_B + w_U×score_U + w_C×score_C, S_total is the decision score, Score_A is the tone detection score, score_F is the logic detection score, score_L is the authorization detection score, score_B is the speculation detection score, score_U is the assertion detection score, and score_C is the context detection score, where w_A, w_F, w_L, w_B, w_U, and w_C are weight coefficients that satisfy the constraint that their sum is 1.

0.

17. An artificial intelligence monitoring method, comprising: Receive content generated by an artificial intelligence system and perform real-time detection, including: detecting at least two independent dimensions of the content generated by the artificial intelligence system, generating a plurality of detection scores, wherein the at least two independent dimensions are selected from tone dimension, logic dimension, permission dimension, speculation dimension, assertion dimension and context dimension; and combining the plurality of detection scores in a weighted manner to form a decision score. And select one of the corresponding protection modes based on the decision score. The protection mode includes at least a personality deviation protection mode, a business risk protection mode, a technical detail protection mode, or a boundary control protection mode. The personality deviation protection mode further includes: establishing a personality state baseline based on the quantity density of a cautious statement, a warning statement, and a technical statement in a dialogue history record; generating a current personality state vector based on the quantity density of the cautious statement, the warning statement, and the technical statement in a currently generated content; comparing the current personality state vector with the personality state baseline to generate a personality deviation vector; generating a personality deviation index based on the personality deviation vector; and performing graded monitoring based on the personality deviation index.

18. The artificial intelligence monitoring method as described in claim 17 further includes: Receive and display the generated content after processing by this protection mode.

Citation Information

Patent Citations

  • Generative artificial intelligence safety assessment method and device

    CN118626964A

  • Generative artificial intelligence safety test method and device based on emoji

    CN119512972A

  • Harmful reply defense method and device for medical big language model

    CN120653770A

  • Expert guidance system for optimizing generative content of large language models

    TWM673306U

  • Method for detecting and mitigating bias and weakness in artificial intelligence training data and models

    US20220012591A1