AI model performance evaluation method and system based on big data analysis
Through the improvement of the performance evaluation method of AI model and the use of technologies such as big data analysis and graph neural networks, the problem that existing methods cannot fully cover the performance of the model is solved, and the comprehensive evaluation and dynamic optimization of the performance of AI model is achieved, and the adaptability and execution quality of the model in different task scenarios is improved.
Patent Information
- Application Number
- CN202510614972.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-14
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-05-14
AI Technical Summary
The existing AI model performance evaluation methods rely on manually set evaluation criteria and limited data sets, and cannot fully cover the performance of the model in different tasks and application scenarios, and fail to effectively consider the diversity and complexity of the tasks, ignore the adaptability of the model when facing unknown tasks, and it is difficult to achieve dynamic optimization based on user interaction.
By obtaining task text data, using embedding models for word segmentation processing and vectorization, using UMAP algorithm for dimensionality reduction and clustering, building a task map and using a graph neural network to extract task risk scores, generating optimization prompt templates and optimizing through adversarial networks, combining user feedback and data backup mechanisms to perform performance evaluation and optimization.
It realizes a comprehensive and accurate evaluation of the performance of AI models, and can provide optimization prompt templates in different tasks and application scenarios, improve the quality of model execution, ensure that the model is adaptable when facing unknown tasks, and can perform dynamic optimization in real time.
Smart Images

Figure CN120123714A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of big data analysis, and in particular to an AI model performance evaluation method and system based on big data analysis. Background Art
[0002] AI models, especially deep learning models, learn potential patterns by processing large amounts of data and are used to perform various tasks. However, how to effectively evaluate the performance of these AI models and ensure their continuous optimization has always been an important issue in research and practice. With the rapid development of big data technology, data-driven AI performance evaluation methods have gradually become mainstream. By means of big data analysis, various factors can be comprehensively considered to evaluate indicators such as the generalization ability, accuracy, and stability of AI models, thus providing support for model improvement and optimization. Existing AI performance evaluation methods usually rely on standard evaluation datasets or task-specific evaluation metrics; However, traditional methods usually rely on manually set evaluation criteria and limited datasets, and cannot comprehensively cover the performance of models in different tasks and different application scenarios. In addition, many existing evaluation methods fail to effectively consider the diversity and complexity of tasks, ignore the adaptability of models when facing unknown tasks, and current AI model evaluation methods also fail to fully combine user feedback and problems encountered in actual applications, making it difficult to achieve dynamic optimization based on user interaction. Summary of the Invention
[0003] In view of the above existing problems, the present invention is proposed.
[0004] Therefore, the present invention provides an AI model performance evaluation method based on big data analysis to solve the problems that traditional methods usually rely on manually set evaluation criteria and limited datasets, cannot comprehensively cover the performance of models in different tasks and different application scenarios. In addition, many existing evaluation methods fail to effectively consider the diversity and complexity of tasks, ignore the adaptability of models when facing unknown tasks, and current AI model evaluation methods also fail to fully combine user feedback and problems encountered in actual applications, making it difficult to achieve dynamic optimization based on user interaction.
[0005] To solve the above technical problems, the present invention provides the following technical solutions: In the first aspect, the present invention provides an AI model performance evaluation method based on big data analysis, which includes: Obtain task text data for the target AI model, perform word segmentation processing through an embedding model to output text vectors, use the UMAP algorithm to reduce the dimension of the task vectors and perform clustering calculations to obtain task clusters, calculate the cosine similarity between task text vectors to construct a task graph, and use a graph neural network to extract task risk scores, create a prompt template for learning and optimization, and generate different optimized prompt templates through an adversarial network; Calculate the Cosine similarity and Jaccard similarity between the output of the target task model and the standard output according to the optimized prompt template, screen the optimized prompt template, perform BERT encoding on the text output by the target task model and the standard output text respectively and calculate the semantic matching degree for quantitative scoring, perform LCS comparison on the model output text and the standard text, generate a quality data set of the text, calculate the clustering divergence of different time windows based on BERT encoding, and perform divergence early warning; Obtain user feedback records, calculate the n-gram overlap degree between the output content of the error text and the standard text, and calculate the sliding deviation of the vocabulary distribution according to the occurrence frequency of the label words, and perform score early warning and deviation early warning respectively; Generate log records for regular analysis, encrypt and store encrypted data items, and perform data backup through secure transmission.
[0006] As a preferred solution of the AI model performance evaluation method based on big data analysis described in the present invention, wherein: the generation of different optimized prompt templates includes, Based on the query task text data set of the target AI model, perform data cleaning and perform word segmentation processing based on the NLTK library; Use the pre-trained embedding model BERT to perform embedding vector dimension on the text of each task, convert each task text into a text vector, and use the UMAP algorithm to reduce the dimension of the task vector; Use the K-means clustering algorithm to cluster the dimension-reduced vectors to obtain task clusters. Each clustering cluster contains a set of semantically similar task text vectors, and calculate the cosine similarity between task text vectors as a measure of the similarity between tasks; According to the similarity between tasks and task clusters, use the graph database Neo4j to construct a task graph. Each task belonging to the cluster is used as a node according to the task cluster, and the edge weight between nodes is the similarity of the task. Use the graph neural network GNN to update the feature representation of each node by passing information between nodes through graph convolution operations; Update the node representation layer by layer according to multiple layers of convolution, output the risk score of each task through a fully connected layer, and perform a descending order sorting of the tasks based on the risk score to clarify the task objectives of the task cluster, and create a prompt template, including task objectives, input examples, and output formats; Using the reinforcement learning RL algorithm, initialize the prompt of the task cluster with a standard prompt template. The intelligent agent generates a model output based on the initial prompt and interacts with the input and output requirements of the task cluster as the environment. Calculate the accuracy and completeness of the output text and the standard output. Adjust the prompt template according to the reward function. The intelligent agent iterates by interacting with the environment multiple times until the changes in accuracy and completeness are no longer obvious, then stop the iteration and generate an optimized prompt template; Using the generative adversarial network GANs algorithm, conduct adversarial training through the generator and discriminator to generate multiple different versions of the optimized prompt template.
[0007] As a preferred solution of the AI model performance evaluation method based on big data analysis described in the present invention, wherein: for the quality dataset of the generated text, perform divergence warning, including, According to the optimized prompt template for the target task model input, calculate the text similarity value between the output text and the standard output through the Jaccard similarity algorithm, calculate the completeness score between the output text and the standard output using the Cosine similarity, adaptively adjust the weights of the text similarity value and the completeness score through historical task data, and calculate the quality score by weighted summation. Calculate the quality score threshold based on the sum of the historical mean and twice the standard deviation of the quality score. If the quality score is less than or equal to the quality score threshold, then eliminate the corresponding optimized prompt template used; Using the pre-trained BERTScore model, perform BERT encoding on the text output by the target task model and the standard output text respectively, calculate the similarity between the encoded text vectors, and compare each word to obtain the per-word similarity. Weight-average the per-word similarity to obtain the similarity score of the entire output, and quantitatively score the semantic matching degree between the output text and the standard text; Using the ROUGE-L tool, perform LCS comparison on the model output text and the standard text, find the longest common subsequence, and calculate the recall rate, precision, and F1 score based on the length and matching degree of the LCS to obtain a quality dataset for comprehensively evaluating the generated text; Mark the target task model according to the quality dataset and the quantitative score to generate an evaluation signature vector; Based on BERT encoding, cluster the input BERT encodings within a time window to obtain the probability distribution of each input belonging to each cluster center. For the cluster distributions of two time windows, calculate the change using the KL divergence; Set a trigger threshold based on historical experience. If the divergence between two time windows is greater than or equal to the trigger threshold, then conduct re-verification and update the evaluation signature vector.
[0008] As a preferred solution of the AI model performance evaluation method based on big data analysis according to the present invention, wherein: the score warning and deviation warning are carried out, including, Implement real-time feedback scoring of users of the monitored target task model, identify user feedback scores based on preset criteria, identify the output content of the error text marked by users, and calculate the n-gram overlap degree between the output content of the error text and the standard text; Perform group label recognition based on user attributes, calculate the occurrence frequency in the output text of the target person model for different label words, and calculate the deviation change of the continuous window as the sliding deviation of the vocabulary distribution; Respectively, use the sum of the historical mean and standard deviation of the BLEU score and the sliding deviation as the score threshold and the deviation threshold. If the BLEU score is greater than or equal to the score threshold and the sliding deviation is greater than or equal to the deviation threshold, then trigger a deviation warning.
[0009] As a preferred solution of the AI model performance evaluation method based on big data analysis according to the present invention, wherein: the generated log records are analyzed regularly, including, Determine the user question text data of the deviation warning and the text output data of the target model, and at the same time record the timestamp, model version, and alarm information. Use a log generation tool to generate a log record file, use the ELK Stack for log storage, and perform regular analysis on the log record file.
[0010] As a preferred solution of the AI model performance evaluation method based on big data analysis according to the present invention, wherein: the encrypted data items are encrypted and stored, including, Determine the encrypted data items, including log files, model input and output, and any relevant user information, and use the AES-256 encryption algorithm to securely encrypt the encrypted data items; Based on the role-based access control RBAC mechanism, set up audit logs to record the user, time, and operation type of each operation, ensure that the data access process is traceable, and perform data desensitization and anonymization processing on the user personal data in the input and output.
[0011] As a preferred solution of the AI model performance evaluation method based on big data analysis according to the present invention, wherein: the data backup is carried out through secure transmission, including, Regularly back up the encrypted data items, audit logs, and log record files, use the TLS / SSL protocol for secure transmission, and perform off-line environment backup.
[0012] In a second aspect, the present invention provides a system for an AI model performance evaluation method based on big data analysis, including, The data word segmentation processing module performs word segmentation on the text data of the target task, converts the text into a vector representation through an embedding model, uses the UMAP algorithm to reduce the dimension of the task vectors, clusters the data, and calculates the task clusters. The task graph construction module calculates the cosine similarity between task text vectors, constructs a task graph, and further extracts the risk scores of the tasks. The prompt template optimization module generates different optimized prompt templates through an adversarial network and optimizes the learning of the target task model according to the optimized prompt templates. The model output comparison module calculates the similarity between the output of the target task model and the standard output and scores it. The text quality evaluation module compares the quality data set of the generated text. The clustering divergence warning module calculates the clustering divergence for different time windows based on BERT encoding and issues divergence warnings. The user feedback module calculates the n-gram overlap between the output content of the error text and the standard text, analyzes the sliding deviation of the vocabulary distribution, and issues score warnings and deviation warnings. The log backup module generates log records for regular analysis, encrypts and stores encrypted data items, and performs data backup through secure transmission.
[0013] In a third aspect, the present invention provides a computer device, including a memory and a processor, where the memory stores a computer program, and: when the computer program is executed by the processor, any step of the AI model performance evaluation method based on big data analysis as described in the first aspect of the present invention is implemented.
[0014] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and: when the computer program is executed by the processor, any step of the AI model performance evaluation method based on big data analysis as described in the first aspect of the present invention is implemented.
[0015] The beneficial effects of the present invention are as follows: By leveraging the hierarchy of graph convolution to strengthen the risk assessment of each task, the risk scores become more accurate and reliable, providing a priority reference for subsequent task management and optimization. The RL algorithm continuously optimizes and generates strategies based on the reward mechanism, enabling the prompt templates to better conform to the task objectives and improving the task execution quality. Multiple high-quality optimized prompt templates are generated through the generative adversarial network to ensure the execution quality in different task scenarios. By using methods such as Jaccard similarity, Cosine similarity, BERTScore, and ROUGE-L to conduct multi-level quality assessments on the model output text, it can not only accurately measure the similarity, integrity, and semantic matching degree between the model output and the standard output, but also ensure that the optimized prompt templates meet the requirements of accuracy and content integrity during the generation process. By implementing the monitoring of the real-time feedback scores of the target task model users and calculating the n-gram overlap degree between the error text and the standard text in combination with the BLEU score, a quantitative quality assessment can be provided for the output generated by each model. By calculating the frequency of occurrence of words related to the group labels in the output text of the target task model and detecting whether there are deviations in text generation based on the deviation changes in consecutive windows, it is possible to effectively identify whether there are output biases or unfair problems in the model for certain groups. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0017] Figure 1 FIG. is a schematic flow chart of the AI model performance evaluation method based on big data analysis in Embodiment 1; Figure 2 FIG. is a schematic structural diagram of the AI model performance evaluation system based on big data analysis in Embodiment 1. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0018] To make the above objects, features, and advantages of the present invention more obvious and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings of the specification.
[0019] Many specific details are set forth in the following description to facilitate a thorough understanding of the present invention. However, the present invention can also be implemented in other ways different from those described herein. Those skilled in the art can make similar extensions without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.
[0020] Second, the "one embodiment" or "embodiment" referred to herein means a specific feature, structure, or characteristic that may be included in at least one implementation manner of the present invention. The "in one embodiment" that appears in different places in this specification does not all refer to the same embodiment, nor is it an embodiment that is separate or selectively mutually exclusive with other embodiments.
[0021] Embodiment 1, referring to From Figure 1 to Figure 2 , is the first embodiment of the present invention. This embodiment provides an AI model performance evaluation method based on big data analysis, including the following steps: S1. Obtain task text data for the target AI model, perform word segmentation processing through an embedding model to output text vectors, use the UMAP algorithm to reduce the dimension of the task vectors and perform clustering calculations to obtain task clusters, calculate the cosine similarity between task text vectors to construct a task graph, and use a graph neural network to extract task risk scores, create a prompt template for learning and optimization, and generate different optimized prompt templates through an adversarial network; Preferably, generating different optimized prompt templates includes, Based on the inquiry task text data set of the target AI model, perform data cleaning and perform word segmentation processing based on the NLTK library; Use the pre-trained embedding model BERT to embed the dimension of the text of each task, and convert each task text into a text vector; Use the UMAP algorithm to reduce the dimension of the task vectors (using local retention and global retention strategies for high-dimensional data), (retaining the relative distance and similarity between task vectors, ensuring that in the low-dimensional space, semantically similar task vectors are still close, while semantically different task vectors are separated); Use the K-means clustering algorithm to cluster the reduced-dimensional vectors to obtain task clusters, and each clustering cluster contains a group of semantically similar task text vectors; Calculate the cosine similarity between task text vectors as a measure of the similarity between tasks; According to the similarity between tasks and task clusters, use the graph database Neo4j to construct a task graph. Each task belonging to the cluster is used as a node according to the task cluster, and the edge weight between nodes is the similarity of the task. Use the graph neural network GNN to update the feature representation of each node by passing information between nodes through graph convolution operations, where each layer of graph convolution operation is expressed as: ; Where represents the convolutional representation of task i at the L-th layer, represents the activation function, and respectively represent the neighbor nodes of task i and j, Denote the convolutional representation of task j at the L-th layer. Denote the bias term of the L-th layer convolution. Denote the similarity between task i and task j. Update the node representation layer by layer according to multi-layer convolution, output the risk score of each task through the fully connected layer, and sort the tasks in descending order of priority based on the risk score. Clarify the task objectives of the task cluster (for example, the objective of the "text classification" cluster is to assign the text to the most appropriate category, and the objective of the "text generation" cluster is to generate relevant text based on the given prompt), create a prompt template, including the task objective (based on the task objective of the task cluster, a clear task description), input examples (uniform input examples designed based on the task cluster), and output format (provide a standard output format requirement for each task to ensure that the output structures of all tasks are consistent). Use the reinforcement learning RL algorithm to initialize the prompts of the task cluster with the standard prompt template. The intelligent agent generates model outputs based on the initial prompts and interacts with the input-output requirements of the task cluster as the environment. Calculate the accuracy Accuracy and completeness Completeness of the output text and the standard output, adjust the prompt template according to the reward function, and the intelligent agent iterates through multiple interactions with the environment until the changes in accuracy and completeness are no longer obvious, then stop the iteration and generate an optimized prompt template. Use the generative adversarial network GANs algorithm. (Based on the optimized prompt template generated by the reinforcement learning through the generator, generate multiple diverse versions of the prompt template. The discriminator judges whether the generated template is accurate, complete, and operable.) Conduct adversarial training through the generator and the discriminator. (The generator continuously adjusts the generated prompt template, and the discriminator continuously improves the ability to identify the generated template to ensure that the generated template can meet a certain quality standard.) Generate multiple different versions of the optimized prompt template.
[0022] Through word segmentation, each text data becomes more structured, laying a good foundation for subsequent embedding and model training, ensuring that each individual word and its grammatical role in the sentence can be processed in subsequent steps, thus making semantic understanding more accurate. By using the pre-trained BERT model to embed the text and convert each task text into a vector, powerful semantic information representation can be provided for each task text. The BERT model captures context relationships, enabling each text vector to accurately reflect the semantic features of the text, providing rich semantic dimension support for subsequent dimensionality reduction and clustering analysis. UMAP ensures that after dimensionality reduction through local and global retention strategies, task texts with similar semantics are still close to each other in the low-dimensional space, while task texts with different semantics are separated, guaranteeing the relative relationship between tasks and the separability between tasks. By constructing a task graph and updating node information through a graph neural network (GNN), and updating the features of task nodes through graph convolution operations, the feature representation of each task becomes more accurate. By fully considering the mutual relationships between tasks in the form of a graph structure, the GNN can propagate information between nodes and learn the similarity between tasks, thus providing a more informative feature representation for the risk scoring of tasks. By leveraging the hierarchy of graph convolution to strengthen the risk assessment of each task, the risk scoring becomes more accurate and reliable, providing a priority reference for subsequent task management and optimization; By clarifying the goals for each task cluster and creating standardized prompt templates, the unity and operability of tasks are ensured. By continuously adjusting the prompts through interaction with the environment, it can be ensured that the generated prompt templates reach the optimal level in terms of accuracy and integrity. The RL algorithm continuously optimizes the generation strategy based on the reward mechanism, enabling the prompt templates to better conform to the task goals and improving the task execution quality. By using a generative adversarial network to generate multiple high-quality optimized prompt templates, these templates can provide more diverse inputs for the target task and ensure the execution quality in different task scenarios.
[0023] S2. Calculate the Cosine similarity and Jaccard similarity between the output of the target task model and the standard output according to the optimized prompt template, screen the optimized prompt template, perform BERT encoding on the text output by the target task model and the standard output text respectively, calculate the semantic matching degree for quantitative scoring, perform LCS comparison on the model output text and the standard text, generate a quality dataset for the text, calculate the clustering divergence for different time windows based on BERT encoding, and issue a divergence warning; Preferably, generating a quality dataset for the text and issuing a divergence warning includes Optimize the prompt template according to the input of the target task model, calculate the text similarity value between the output text and the standard output through the Jaccard similarity algorithm, calculate the integrity score of the output text and the standard output using the Cosine similarity, (adaptively adjust the weights of the text similarity value and the integrity score based on historical task data, and calculate the quality score by weighted summation. Calculate the quality score threshold based on the sum of the historical mean and twice the standard deviation of the quality score. If the quality score is less than or equal to the quality score threshold, then eliminate the corresponding optimized prompt template; Use the pre-trained BERTScore model to perform BERT encoding on the text output by the target task model and the standard output text respectively, calculate the similarity between the encoded text vectors, and compare each word to obtain the word-by-word similarity. Weight and average the word-by-word similarity to obtain the similarity score of the entire output, and quantitatively score the semantic matching degree between the output text and the standard text; Use the ROUGE-L tool to perform LCS comparison on the model output text and the standard text, find the longest common subsequence, and calculate the recall rate, precision, and F1 score based on the length and matching degree of the LCS to obtain a quality dataset for comprehensively evaluating the generated text; Generate an evaluation signature vector based on the quality dataset and the quantitative score to label the target task model; Based on BERT encoding, cluster the input BERT encoding within a time window to obtain the probability distribution of each input belonging to each cluster center. For the cluster distributions of two time windows, use KL divergence to calculate the change, expressed as: ; where represents the divergence between two time windows, represents the number of cluster categories, and represent the probability distributions of the k-th category in time windows t and t + 1 respectively; Set a trigger threshold based on historical experience. If the divergence between two time windows is greater than or equal to the trigger threshold, then perform re-verification and update the evaluation signature vector.
[0024] By comparing the optimized prompt template input to the task model with the standard output, using Jaccard similarity and Cosine similarity to calculate the similarity and integrity of the output text, it is possible to ensure that the model's output remains consistent in terms of semantic matching with the standard output. Through the adaptive adjustment of task data, it is possible to dynamically adjust the evaluation weights for different task types, thereby improving the accuracy of the quality score. By using the BERTScore model to encode the output and standard output of the target task model, the text can be transformed into deep semantic embeddings, and the quality score of the output can be obtained by comparing the similarity between word vectors. By combining the ROUGE-L tool to perform the longest common subsequence (LCS) comparison between the generated text and the standard text, it is possible to further quantitatively evaluate the content similarity of the text. By clustering the input text within two time windows based on BERT encoding and KL divergence calculation, and analyzing the changes in the clustering distribution, it is possible to effectively detect changes in the model's input over consecutive time periods; by quantifying the changes in the data to capture potential drifts or pattern changes; By using methods such as Jaccard similarity, Cosine similarity, BERTScore, and ROUGE-L to conduct multi-level quality evaluations on the model output text, it is not only possible to accurately measure the similarity, integrity, and semantic matching between the model output and the standard output, but also to ensure that the optimized prompt template can meet the requirements of accuracy and content integrity during the generation process. Combining dimensionality reduction clustering and data drift monitoring based on BERT encoding and KL divergence calculation enables the model to continuously track changes in the distribution of input data and identify potential drifts based on the different characteristics of task clusters. This process provides strong support for the continuous verification of the task model: when the data distribution changes, it triggers re-verification in a timely manner to ensure that the model can adapt to changes in the input patterns of different time windows and prevent the model performance from degrading due to data bias.
[0025] S3, obtain the user feedback records, calculate the n-gram overlap between the content of the incorrect text output and the standard text, and calculate the sliding deviation of the lexical distribution based on the occurrence frequency of the labeled words, and issue score warnings and deviation warnings respectively; Preferably, issuing score warnings and deviation warnings includes, Implement real-time feedback scoring for users monitoring the target task model, identify user feedback scores based on preset criteria, identify the content of incorrect text output marked by users, and calculate the n-gram overlap between the content of the incorrect text output and the standard text, expressed as: ; where represents the BLEU score, represents the n-gram precision value between the generated text and the standard text, Indicates the text length of the error text output; Perform group label recognition based on user attributes, calculate the occurrence frequency in the output text of the target person model for different label words, and calculate the deviation change of the continuous window as the sliding deviation of the vocabulary distribution; Use the sum of the historical mean and standard deviation of the BLEU score and the sliding deviation as the score threshold and deviation threshold respectively. If the BLEU score is greater than or equal to the score threshold and the sliding deviation is greater than or equal to the deviation threshold, then trigger a deviation warning.
[0026] By implementing the monitoring of the real-time feedback score of the target task model for users and calculating the n-gram overlap degree between the error text and the standard text in combination with the BLEU score, a quantitative quality assessment can be provided for the output generated by each model. As an important tool for measuring the similarity between the generated text and the standard text, the BLEU score can accurately capture the degree of vocabulary overlap in the model output. Through the analysis of the n-gram overlap degree of the error text output content, it can help us clarify in which aspects the model has deviations and which specific words or phrases have caused the model errors. Using the attributes of users for group label recognition enables the task model to optimize the output content according to the needs of different groups. Especially for the task execution of diverse user groups, by calculating the occurrence frequency of words related to group labels in the output text of the target task model and detecting whether there are deviations in text generation based on the deviation change of the continuous window, it can effectively identify whether the model has output biases or unfair problems for certain groups.
[0027] S4. Generate log records for regular analysis, encrypt and store encrypted data items, and perform data backup through secure transmission; Preferably, generate log records for regular analysis, including, Determine the user question text data for deviation warning and the text output data of the target model, and at the same time record the timestamp, model version, and alarm information. Use a log generation tool to generate a log record file, use the ELK Stack for log storage, and perform regular analysis on the log record file.
[0028] By implementing a deviation warning mechanism and recording the user question text data and the output data of the target model, the generation time and related information of each warning can be accurately tracked, and detailed logs such as timestamps, model versions, and alarm information can be recorded. This can not only ensure the accurate positioning of each deviation warning but also provide important reference data for model performance monitoring and subsequent problem troubleshooting.
[0029] Furthermore, encrypt and store encrypted data items, including, Identify encrypted data items, including log files, model input and output, and any relevant user information, and securely encrypt the encrypted data items using the AES-256 encryption algorithm; Based on the role-based access control (RBAC) mechanism, set up audit logs to record the user, time, and operation type of each operation, ensure that the data access process is traceable, and perform data desensitization and anonymization on the user's personal data in the input and output.
[0030] By encrypting log files, model input and output, and relevant user information using AES-256, the security and confidentiality of these sensitive data can be effectively ensured, avoiding unauthorized access or leakage. The AES-256 algorithm provides strong encryption protection to ensure the security of data during storage and transmission. Using the role-based access control (RBAC) mechanism, it can ensure that only authorized users can access sensitive data, allocate appropriate permissions according to roles, thereby implementing the principle of minimizing data access and reducing potential security vulnerabilities.
[0031] Furthermore, perform data backup through secure transmission, including, Regularly back up encrypted data items, audit logs, and log record files, use the TLS / SSL protocol for secure transmission, and perform offline environment backups.
[0032] By regularly backing up encrypted data items, audit logs, and log record files, it can effectively ensure that important data can be quickly restored in case of system failures or data loss, ensuring the high availability of the system. By backing up encrypted data, it ensures that even if the backup data is stolen, the data is still under encryption protection, avoiding the risk of leaking sensitive information.
[0033] This embodiment also provides a system for evaluating the performance of an AI model based on big data analysis, including, A data tokenization processing module that tokenizes the text data of the target task, converts the text into a vector representation through an embedding model, uses the UMAP algorithm to reduce the dimension of the task vector, and clusters the data to calculate the task clusters A task graph construction module that calculates the cosine similarity between task text vectors, constructs a task graph, and further extracts the risk score of the task; A prompt template optimization module that generates different optimized prompt templates through an adversarial network and learns and optimizes the target task model according to the optimized prompt templates; A model output comparison module that calculates the similarity between the output of the target task model and the standard output and scores it; A text quality evaluation module that compares the quality data set of the generated text; The clustering divergence warning module calculates the clustering divergence for different time windows based on BERT encoding and issues divergence warnings. The user feedback module calculates the n-gram overlap between the output content of the error text and the standard text, analyzes the sliding deviation of the vocabulary distribution, and issues score warnings and deviation warnings. The log backup module generates log records for regular analysis, encrypts and stores encrypted data items, and performs data backup through secure transmission.
[0034] This embodiment also provides a computer device applicable to the case of the AI model performance evaluation method based on big data analysis, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the AI model performance evaluation method based on big data analysis as proposed in the above embodiment.
[0035] This computer device can be a terminal. The computer device includes a processor, a memory, a communication interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner. The wireless manner can be achieved through WIFI, a carrier network, NFC (Near Field Communication), or other technologies. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, touchpad, or mouse, etc.
[0036] This embodiment also provides a storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the method for evaluating the performance of an AI model based on big data analysis as proposed in the above embodiment; the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (Static Random Access Memory, abbreviated as SRAM), electrically erasable programmable read-only memory (Electrically Erasable Programmable Read-Only Memory, abbreviated as EEPROM), erasable programmable read-only memory (Erasable Programmable Read Only Memory, abbreviated as EPROM), programmable read-only memory (Programmable Red-Only Memory, abbreviated as PROM), read-only memory (Read-Only Memory, abbreviated as ROM), magnetic memory, flash memory, a magnetic disk or an optical disc.
[0037] In summary, the present invention strengthens the risk assessment of each task by utilizing the hierarchy of graph convolution, making the risk score more accurate and reliable, providing a priority reference for subsequent task management and optimization. The RL algorithm continuously optimizes the generation strategy based on the reward mechanism, enabling the prompt template to better conform to the task objective and improving the task execution quality. Multiple high-quality optimized prompt templates are generated through a generative adversarial network to ensure the execution quality in different task scenarios. By using methods such as Jaccard similarity, Cosine similarity, BERTScore, and ROUGE-L to perform multi-level quality assessment on the model output text, it can not only accurately measure the similarity, integrity, and semantic matching degree between the model output and the standard output, but also ensure that the optimized prompt template can meet the requirements of accuracy and content integrity during the generation process. By implementing the monitoring of the real-time feedback score of the target task model users and calculating the n-gram overlap degree between the error text and the standard text in combination with the BLEU score, a quantitative quality assessment can be provided for the output generated by each model. By calculating the frequency of occurrence of words related to the group label in the output text of the target task model and detecting whether there is a deviation in text generation based on the deviation change of the continuous window, it can effectively identify whether there is an output bias or unfairness problem of the model in certain groups.
[0038] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered by the scope of the claims of the present invention.
Claims
1. A method for evaluating AI model performance based on big data analysis, characterized in that: include: Obtain task text data for the target AI model, perform word segmentation processing through the embedding model to output text vectors, use the UMAP algorithm to reduce the dimension of the task vectors and perform clustering calculations to obtain task clusters, calculate the cosine similarity between task text vectors to build a task graph, and use a graph neural network to extract task risk scores, create prompt templates for learning optimization, and generate different optimized prompt templates through adversarial networks; Calculate the Cosine similarity and Jaccard similarity between the target task model output and the standard output according to the optimization prompt template, select the optimization prompt template, perform BERT encoding on the target task model output text and the standard output text respectively, calculate the semantic matching degree for quantitative scoring, perform LCS comparison on the model output text and the standard text, generate a text quality data set, calculate the clustering divergence of different time windows based on BERT encoding, and issue divergence warning; Obtain user feedback records, calculate the n-gram overlap between the erroneous text output and the standard text, and calculate the sliding deviation of the vocabulary distribution based on the frequency of occurrence of the label words, and perform score warning and deviation warning respectively; Generate log records for regular analysis, encrypt data items for storage, and back up data via secure transmission.
2. The AI model performance evaluation method based on big data analysis according to claim 1, characterized in that: The generation of different optimization prompt templates includes: Based on the query task text dataset of the target AI model, perform data cleaning and perform word segmentation processing based on the NLTK library; Use the pre-trained embedding model BERT to process the embedding vector dimension of each task text, convert each task text into a text vector, and use the UMAP algorithm to reduce the dimension of the task vector; Use the K-means clustering algorithm to cluster the reduced-dimensional vectors to obtain task clusters. Each cluster contains a group of semantically similar task text vectors. Calculate the cosine similarity between task text vectors as a measure of similarity between tasks. Based on the similarity between tasks and task clusters, we use the graph database Neo4j to build a task graph. According to the task cluster, each task in the cluster is regarded as a node. The edge weight between nodes is the similarity of the task. We use the graph neural network GNN to transfer information between nodes through graph convolution operations to update the feature representation of each node. Update the node representation layer by layer according to the multi-layer convolution, output the risk score of each task through the fully connected layer, sort the tasks in descending order of priority based on the risk score, clarify the task objectives of the task cluster, and create a prompt template, including the task objectives, input examples, and output formats; Using the reinforcement learning RL algorithm, the prompts of the task cluster are initialized with the standard prompt template. The intelligent agent generates the model output according to the initial prompt, and interacts with the input and output requirements of the task cluster as the environment, calculates the accuracy and completeness of the output text and the standard output, and adjusts the prompt template according to the reward function. The intelligent agent iterates through multiple interactions with the environment until the changes in accuracy and completeness are no longer obvious, then stops iterating and generates an optimized prompt template. Using the generative adversarial network (GANs) algorithm, adversarial training is performed between the generator and the discriminator to generate multiple different versions of optimized prompt templates.
3. The AI model performance evaluation method based on big data analysis according to claim 2, characterized in that: The quality data set of the generated text is used for divergence warning, including: According to the target task model, the optimization prompt template is input, the text similarity value between the output text and the standard output is calculated by the Jaccard similarity algorithm, the integrity score between the output text and the standard output is calculated by using the Cosine similarity, the weights of the text similarity value and the integrity score are adaptively adjusted through historical task data, and the quality score is calculated by weighted summation, and the quality score threshold is calculated based on the sum of the historical mean and double standard deviation of the quality score. If the quality score is less than or equal to the quality score threshold, the corresponding optimization prompt template is eliminated; Using the pre-trained BERTScore model, the text output by the target task model and the standard output text are BERT encoded respectively, the similarity between the encoded text vectors is calculated, and each word is compared to obtain the word-by-word similarity. The word-by-word similarity is weighted averaged to obtain the similarity score of the entire output, and the semantic match between the output text and the standard text is quantitatively scored; Use the ROUGE-L tool to perform LCS comparison between the model output text and the standard text to find the longest common subsequence. Based on the length and matching degree of the LCS, calculate the recall rate, precision, and F1 score to obtain a comprehensive quality dataset for evaluating the generated text. Generate an evaluation signature vector based on the quality dataset and quantitative score to label the target task model; Based on BERT encoding, cluster the input BERT encoding within a time window to obtain the probability distribution of each input belonging to each cluster center. For the cluster distribution of two time windows, use KL divergence to calculate the change; The trigger threshold is set based on historical experience. If the divergence of the two time windows is greater than or equal to the trigger threshold, re-verification is performed and the evaluation signature vector is updated.
4. The AI model performance evaluation method based on big data analysis according to claim 3, characterized in that: The score warning and deviation warning include: Implement real-time feedback scoring of users of the monitoring target task model, identify user feedback scores based on preset standards, identify erroneous text output content marked by users, and calculate the n-gram overlap between the erroneous text output content and the standard text; Identify group labels based on user attributes, calculate the frequency of occurrence of different label words in the output text of the target persona model, and calculate the deviation change of continuous windows as the sliding deviation of vocabulary distribution; The score threshold and deviation threshold are respectively based on the sum of the historical mean and standard deviation of the BLEU score and the sliding deviation. If the BLEU score is greater than or equal to the score threshold, and the sliding deviation is greater than or equal to the deviation threshold, a deviation warning is triggered.
5. The AI model performance evaluation method based on big data analysis according to claim 4, characterized in that: The generated log records are regularly analyzed, including, Determine the user question text data for deviation warning and the text output data of the target model, record the timestamp, model version and alarm information, use the log generation tool to generate log files, use ELK Stack for log storage and regularly analyze the log files.
6. The AI model performance evaluation method based on big data analysis according to claim 5, characterized in that: The encrypted data item is encrypted and stored, including: Identify encrypted data items, including log files, model input and output, and any relevant user information, and securely encrypt them using the AES-256 encryption algorithm; The role-based access control (RBAC) mechanism sets up audit logs to record the user, time, and operation type of each operation, ensures that the data access process is traceable, and performs data desensitization and anonymization on user personal data in input and output.
7. The AI model performance evaluation method based on big data analysis according to claim 6, characterized in that: The data backup is performed through secure transmission. include, Encrypted data items, audit logs, and log record files are backed up regularly, using TLS / SSL protocols for secure transmission and offline environment backup.
8. A system for an AI model performance evaluation method based on big data analysis, based on the AI model performance evaluation method based on big data analysis according to any one of claims 1 to 7, characterized in that: include, The data segmentation processing module performs word segmentation on the text data of the target task, converts the text into a vector representation through an embedding model, uses the UMAP algorithm to reduce the dimension of the task vector, clusters the data, and calculates the task clusters. The task graph construction module calculates the cosine similarity between task text vectors, constructs the task graph, and further extracts the risk score of the task; The prompt template optimization module generates different optimized prompt templates through the adversarial network, and learns and optimizes the target task model based on the optimized prompt templates; Model output comparison module, which calculates the similarity between the target task model output and the standard output and scores them; The text quality assessment module compares the quality dataset of the generated text; Cluster divergence warning module, which calculates cluster divergence of different time windows based on BERT encoding and issues divergence warning; User feedback module, which calculates the n-gram overlap between the erroneous text output and the standard text, analyzes the sliding deviation of vocabulary distribution, and issues score warning and deviation warning; The log backup module generates log records for regular analysis, encrypts and stores encrypted data items, and backs up data through secure transmission.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the AI model performance evaluation method based on big data analysis described in any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the AI model performance evaluation method based on big data analysis described in any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Standard think tank knowledge recommendation method and system based on big data
CN118193845A
Question generation method and device based on reinforcement learning and storage medium
CN118536585A
Task-oriented large language model type selection and construction method and system
CN118760740A
Decision-making method and model for offline reinforcement learning and continuous online fine tuning
CN119249360A
Large model text generation method and system based on adaptive cue words
CN119623475A
Cited By
Artificial intelligence model evaluation optimization method and equipment based on task network
CN122309316A
Methods and Equipment for Evaluating and Optimizing Artificial Intelligence Models Based on Task Networks
CN122309316B