AI model decision-making method and device based on multi-modal data and electronic equipment
By acquiring multimodal data and user review data to generate user portraits and network graphs, the learning samples of the AI model are purified and expanded, which solves the problem of excessive dirty data in the AI model training samples and improves the training effect of the model.
Patent Information
- Application Number
- CN202510985030.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-17
- Publication Date
- 2025-10-14
- Estimated Expiration
- 2045-07-17
AI Technical Summary
The large amount of dirty data in the AI model training samples leads to poor training results.
By acquiring multimodal data that integrates text, time series, and images, as well as real user review data, we generate user portraits and real user network graphs, purify the learning samples of the initial AI model, and use the intermediate AI model to interact with users for credibility evaluation and data expansion, ultimately optimizing the target AI model.
It improves the quality of AI model training samples, reduces the impact of dirty data, enhances the robustness and generalization ability of the model, and provides higher accuracy and reliability.
Smart Images

Figure CN120781082A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence, in particular to an AI model decision method and device based on multi-modal data and electronic equipment. BACKGROUND
[0002] At present, model training samples are the core basis for constructing efficient models, and their quality and scale directly affect model performance. Models learn language rules, visual features or decision logic through massive samples, for example, in the pre-training stage, using academic papers, blogs and other text data to learn word association and grammar structure to form basic language ability. However, there are many AI model training sample data that do not have certain credibility, resulting in more dirty data in AI model training samples, which leads to poor AI model training effect. SUMMARY
[0003] The purpose of the present application is to provide an AI model decision method and device based on multi-modal data and electronic equipment to solve the technical problem of poor AI model training effect caused by too much dirty data in AI model training samples.
[0004] In a first aspect, the present application provides an AI model decision method based on multi-modal data, which comprises: Obtaining multi-modal data fused with text, time sequence and image, and real user comment data related to the multi-modal data; Generating a user portrait of a target user corresponding to the real user comment data according to the real user comment data and the multi-modal data corresponding to the comment; Generating a user real network graph of a plurality of target users based on the relationship between the multi-modal data and the real user comment data according to a plurality of user portraits corresponding to the target users; Adding the user real network graph to the learning samples of an initial AI model to obtain first learning samples, and using the first learning samples to train the initial AI model to improve the real credibility of the learning samples of the initial AI model through the user real network graph and purify the dirty data learned by the initial AI model to obtain a trained intermediate AI model; Through the intermediate AI model and the first user, obtaining interaction data, and using the intermediate AI model to evaluate the credibility of the first user according to the interaction data to obtain a confidence value evaluation result of the first user; In response to the confidence value evaluation result of the first user being greater than a specified confidence value, determining the user comment data of the first user corresponding to a plurality of fields as the real user comment data to expand the user range of the user real network graph using a plurality of first users; Add the user real network graph extended in the user range to the learning sample of the intermediate AI model to obtain a second learning sample, and perform model optimization on the intermediate AI model by using the second learning sample to obtain an optimized target AI model.
[0005] In one possible implementation, after the model optimization on the intermediate AI model by using the second learning sample to obtain the optimized target AI model, the method further includes: In response to a learning request of a user for a target stream media, simulate a real user's stream media query process by increasing the speed of browsing by using the target AI model, and obtain first-level multi-modal data in the target stream media by simulating the stream media query process; Diverge peripheral related data of related user request learning data through the user real network graph with the first-level multi-modal data as the center to obtain second-level multi-modal data of related user request learning; Preferentially perform main user request trend training on the target AI model by using the first-level multi-modal data, and perform secondary related request trend training on the target AI model by using the second-level multi-modal data after the release of the computing power of the main user request trend training is completed to obtain a final AI model.
[0006] In one possible implementation, the confidence value evaluation of the first user by using the intermediate AI model according to the interaction data includes: The confidence value evaluation of the first user by using the intermediate AI model according to the similarity of the interaction data to the interaction data of other real users except the first user, the number of target real users whose similarity of the interaction data is greater than a specified similarity threshold, the confidence value evaluation result corresponding to the target real user, and the activity of the first user.
[0007] In one possible implementation, the confidence value evaluation of the first user by using the intermediate AI model according to the similarity of the interaction data to the interaction data of other real users except the first user, the number of target real users whose similarity of the interaction data is greater than a specified similarity threshold, the confidence value evaluation result corresponding to the target real user, and the activity of the first user includes: The intermediate AI model is used to perform a credibility evaluation on the first user using the following formula based on the similarity between the interaction data and the interaction data of other real users other than the first user, the number of target real users whose interaction data similarity is greater than a specified similarity threshold, the confidence value evaluation result corresponding to the target real user, and the activity of the first user, to obtain the confidence value evaluation result of the first user: = + α × + β × ; in, express; Indicates the similarity of the interaction data between the i-th target real user and the first user; Represents the confidence value evaluation result of the i-th target real user; Indicates the activity level of the first user; The number of users representing the target real users whose interaction data similarity is greater than a specified similarity threshold; α Indicates the degree of influence of adjusting behavioral deviation on confidence value evaluation; β Indicates the impact of activity on confidence evaluation.
[0008] In one possible implementation, after using the intermediate AI model to perform a credibility evaluation on the first user based on the interaction data to obtain a confidence value evaluation result of the first user, the method further includes: In response to a confidence value evaluation result of the first user being less than or equal to the specified confidence value, determining the user comment data of the first user corresponding to the multiple fields as dirty data samples; The intermediate AI model is reversely trained using the dirty data samples to purify the dirty data learned by the intermediate AI model.
[0009] In one possible implementation, it also includes: Using heterogeneous data adapters, real user behavior logs, sensor time series signals, and detection images are encoded into feature vectors, and the feature vectors are aligned using cross-modal attention. The risk features generated by the AI model are input into the dual-channel model of the XGBoost algorithm model and the LightGBM algorithm model for preliminary risk assessment; The real-time feedback module based on the preliminary risk assessment and reinforcement learning dynamically adjusts the risk threshold of the risk feature by interacting with the user environment.
[0010] In one possible implementation, it also includes: a contrastive attribution analysis mode is used to contrast feature differences between specified high-risk features and specified low-risk features, and a multi-granularity explanation engine is constructed using the feature differences; natural language analysis data is generated based on the multi-granularity explanation engine, and an AI model decision is converted into a compliance explanation result by a large model based on the natural language analysis data.
[0011] In a second aspect, the present application provides an AI model decision device based on multi-modal data, comprising: An acquisition module is configured to acquire multi-modal data fused with text, time sequence, and images, and real user comment data related to the multi-modal data. A first generation module is configured to generate a user portrait of a target user corresponding to the real user comment data according to the real user comment data and the multi-modal data corresponding to the comment. A second generation module is configured to generate a user real network graph of a plurality of target users based on relationships between the multi-modal data and the real user comment data according to the user portraits of the target users. A training module is configured to add the user real network graph to learning samples of an initial AI model to obtain first learning samples, and perform model training on the initial AI model using the first learning samples to improve real credibility of the learning samples of the initial AI model and purify dirty data learned by the initial AI model through the user real network graph, thereby obtaining a trained intermediate AI model. An evaluation module is configured to interact with a first user through the intermediate AI model to obtain interaction data, and perform credibility evaluation on the first user according to the interaction data using the intermediate AI model to obtain a confidence value evaluation result of the first user. A determination module is configured to determine user comment data of a plurality of fields corresponding to the first user as the real user comment data in response to the confidence value evaluation result of the first user being greater than a specified confidence value, so as to expand a user range of the user real network graph using a plurality of first users. An optimization module is configured to add the user real network graph with the expanded user range to learning samples of the intermediate AI model to obtain second learning samples, and perform model optimization on the intermediate AI model using the second learning samples to obtain an optimized target AI model.
[0012] In a third aspect, the present application further provides an electronic device comprising a memory and a processor, wherein the memory stores a computer program executable on the processor, and the processor implements the method of the first aspect when executing the computer program.
[0013] In a fourth aspect, the present application provides a computer readable storage medium storing computer executable instructions, which when executed by a processor, cause the processor to perform the method of the first aspect.
[0014] The application brings the following beneficial effects: the AI model decision method, device and electronic equipment based on multi-modal data provided by the application can obtain multi-modal data fused with text, time sequence and image and real user comment data related to the multi-modal data; generate a user portrait of a target user corresponding to the real user comment data according to the real user comment data and the multi-modal data corresponding to the comment; generate a user real network graph of a plurality of target users based on the relationship between the multi-modal data and the real user comment data according to the user portraits corresponding to the target users; add the user real network graph to the learning sample of an initial AI model to obtain a first learning sample, and perform model training on the initial AI model by using the first learning sample, so as to improve the real credibility of the learning sample of the initial AI model through the user real network graph and purify the dirty data learned by the initial AI model, and obtain a trained intermediate AI model; interact with a first user through the intermediate AI model to obtain interaction data, and perform credibility evaluation on the first user according to the interaction data by using the intermediate AI model to obtain a confidence value evaluation result of the first user; in response to the confidence value evaluation result of the first user being greater than a specified confidence value, determine the user comment data of the first user corresponding to a plurality of fields as the real user comment data, so as to expand the user range of the user real network graph by using a plurality of first users; add the user real network graph with the expanded user range to the learning sample of the intermediate AI model to obtain a second learning sample, and perform model optimization on the intermediate AI model by using the second learning sample to obtain an optimized target AI model.In the scheme, by fusing multi-modal data of text, time sequence and image and real user comment data, the system first obtains rich information sources, which are used to generate user portraits of target users, ensuring accurate understanding of user behavior and preferences. Using multiple user portraits and their related multi-modal data and comment data, the system further constructs a user real network graph, which not only captures the behavior characteristics of individual users, but also reveals the mutual relationship and influence mode between users, thereby providing a more three-dimensional data structure for subsequent analysis. Moreover, by adding the user real network graph to the learning samples of the initial AI model, the model can learn in a situation closer to the real world, thereby effectively improving the authenticity and representativeness of the learning samples and reducing the problem of dirty data caused by sample bias. Furthermore, through the interaction between the intermediate AI model and the first user, the credibility of the user is evaluated based on the interaction data, which helps to filter out unreliable or malicious data input, further purifying the training sample set. When the confidence value of the first user reaches the specified standard, its comment data in multiple fields will be included in the real user comment data pool, expanding the coverage of the user real network graph. This expansion not only increases the data volume, but also enhances the data diversity, which is conducive to improving the robustness and generalization ability of the model. Finally, using the purified and expanded user real network graph as the second learning sample, the intermediate AI model is optimized to obtain the target AI model. This model has higher accuracy and reliability because it has higher quality learning materials, greatly reducing the impact of dirty data. Therefore, through the above data processing and model training strategies, the quality of the AI model training sample is gradually improved, thereby achieving the technical effect of significantly reducing dirty data, and solving the technical problem of poor AI model training effect caused by too much dirty data in the AI model training sample.
[0015] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the following preferred embodiments are described in detail below, and the accompanying drawings are described as follows. BRIEF DESCRIPTION OF DRAWINGS
[0016] In order to more clearly illustrate the specific embodiments of the present application or the technical solutions in the prior art, the following will briefly introduce the drawings needed to be used in the specific embodiments or the prior art description. Obviously, the drawings in the following description are some embodiments of the present application, and those skilled in the art can obtain other drawings according to these drawings without creative labor.
[0017] Figure 1 Flowchart of the AI model decision method based on multi-modal data provided by the embodiments of the present application; Figure 2Another flowchart of the AI model decision method based on multi-modal data provided by the embodiment of the present application is shown. Figure 3 A structural diagram of an AI model decision device based on multi-modal data provided by the embodiment of the present application is shown. Figure 4 A structural diagram of an electronic device provided by the embodiment of the present application is shown. DETAILED DESCRIPTION
[0018] To make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions of the present application will be described below in detail with reference to the accompanying drawings. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of protection of the present application.
[0019] The terms "comprise" and "have" and any variations thereof mentioned in the embodiments of the present application are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units is not limited to the listed steps or units, but can optionally further comprise other steps or units not listed, or can optionally further comprise other steps or units inherent to the process, method, product or device.
[0020] At present, many AI model training sample data do not have certain credibility, resulting in more dirty data in the AI model training sample, which leads to poor AI model training effect. Based on this, the embodiments of the present application provide an AI model decision method, device and electronic equipment based on multi-modal data, which can solve the technical problem that more dirty data in the AI model training sample leads to poor AI model training effect.
[0021] The embodiments of the present application will be further described below with reference to the accompanying drawings.
[0022] Figure 1 A flowchart of an AI model decision method based on multi-modal data provided by the embodiment of the present application is shown. As shown in Figure 1 The method comprises: Step S110, obtaining multi-modal data fused with text, time sequence and image and real user comment data related to the multi-modal data.
[0023] As an optional implementation, natural language processing (NLP) techniques are used to extract text information from selected data sources, including product reviews, service feedback, etc. Time series of user behavior are recorded, such as login time, browsing history, time points of purchase behavior, etc. Image recognition techniques are used to automatically capture and analyze visual elements in photos or videos uploaded by users. Irrelevant or duplicate data is cleaned up to ensure the quality of the data set. Data from different sources is converted to a unified format for subsequent processing. The internal relationships between text, time series, and image data are analyzed. For example, user comments in a certain time period may be related to a certain picture or a series of operation behaviors. Based on existing user behavior data and interaction records, the authenticity of the comments is labeled, and false or misleading information is excluded. The cleaned and correlated text, time series, and image data are fused together to form a comprehensive data set. Indexes are created for the fused multi-modal data to facilitate fast retrieval and query. Based on data volume and access requirements, a suitable database system is selected to store the multi-modal data. Ensure data security to prevent data leakage and unauthorized access.
[0024] Through the above steps, the system can effectively obtain and process multi-modal data that integrates text, time series, and images, and accurately associate with real user comments related to these data, laying a foundation for further user portrait construction, network graph analysis, etc.
[0025] Step S120, generating a user portrait of a target user corresponding to the real user comment data according to the real user comment data and the multi-modal data corresponding to the comment.
[0026] Among them, the user portrait contains user attributes, user behavior and user expectation data of the target user.
[0027] In an alternative embodiment, natural language processing techniques are used to extract keywords, topics, sentiment trends, and other information from user reviews. Time series analysis of user behavior, such as login frequency, browsing habits, purchase cycles, and other factors, is used to capture user behavior patterns. Computer vision techniques are used to extract information from user-uploaded photos or videos, such as scene recognition, object detection, and other factors. Based on the extracted features, models are established to describe user behavior. This may include, but is not limited to, user preference models, user demand models, and other models. Machine learning algorithms are used to predict future user behavior trends, providing the basis for personalized recommendations. Based on existing data, user age range, gender, geographic location, and other basic information are inferred. By analyzing user behavior and preferences in different situations, user interests, hobbies, and consumption habits are described. If the data allows, the user's social network can also be analyzed to understand their influence and degree of influence. Data from different channels is integrated to form a comprehensive and detailed user portrait. As new data continues to flow in, the user portrait is regularly updated and optimized to maintain its accuracy and timeliness. Based on the constructed user portrait, more personalized services and product recommendations are provided to users. By collecting user feedback on recommended content, the accuracy and effectiveness of the user portrait are evaluated, and the algorithms and strategies are adjusted accordingly.
[0028] Step S130, based on the user portraits of the plurality of target users, the relationship between the multi-modal data and the real user comment data is used to generate a user real network graph of the plurality of target users.
[0029] Exemplarily, the collected data is cleaned, including removing noise, standardizing format, filling in missing values, etc. Based on the existing multi-modal data and user comment data, a detailed user portrait is constructed for each target user, covering basic information, interests and hobbies, social behavior, etc. With the inflow of new data, the user portrait is constantly updated and optimized to ensure its accuracy and timeliness. Useful feature information is extracted from multi-modal data, such as sentiment orientation in text, behavior patterns in time series, scene recognition in images, etc. The internal relationship between different modal data is analyzed to explore how they jointly act on user behavior and preferences. By analyzing the interactions between users (such as comment replies, likes, shares, etc.), direct or indirect relationships between users are identified. According to the interaction frequency, content similarity, etc. between users, the strength or weight of each relationship is calculated. All target users and their relationships are drawn into a network graph, where nodes represent users and edges represent relationships between users. The thickness or color of the edge can represent the strength or type of the relationship. In the generated network graph, a community discovery algorithm is applied to find out the closely connected user groups or communities. Use appropriate graphical tools and techniques to visualize the network graph to better understand the relationship structure between users. Based on the generated user network graph, more accurate personalized recommendation services are provided for users. By collecting user feedback on recommended content, the effectiveness of the network graph is evaluated, and the algorithm and strategy are adjusted accordingly to achieve continuous improvement of the system.
[0030] In step S140, the user real network graph is added to the learning sample of the initial AI model to obtain a first learning sample, and the initial AI model is trained using the first learning sample to improve the real and reliable learning sample of the initial AI model through the user real network graph and purify the dirty data learned by the initial AI model, to obtain a trained intermediate AI model.
[0031] The information in the user's real network map (such as node attributes, edge relationships, etc.) is used to expand the original learning sample set to form the first learning sample set. This step aims to increase the authenticity and diversity of the data. According to the newly added network map data, update or re-label the label information in the learning sample to ensure that it accurately reflects the reality. Load the pre-set initial AI model architecture and prepare to start the training process. Use the first learning sample set to train the initial AI model. In this process, the focus is on how to effectively use the newly added network map data to improve the performance of the model. Evaluate the model performance through cross-validation and other methods, and adjust the parameters or improve the algorithm to optimize the model performance. Identify and label the outliers or dirty data in the first learning sample set using statistical methods or machine learning techniques. Based on the results of the previous step, use appropriate technical means (such as data filtering, repair or deletion) to process the identified dirty data, thereby improving the quality of the entire data set. Use the first learning sample set after cleaning to further train the model and obtain the trained intermediate AI model. Perform a comprehensive performance evaluation on the final intermediate AI model, including but not limited to accuracy, recall rate, F1 score and other indicators, to ensure that it has significantly improved compared to the initial model. Collect feedback on the actual operation of the model after deployment and analyze the performance of the model in real-world scenarios. Based on the feedback, adjust and optimize the model as needed to achieve a closed-loop development process and continuously improve the applicability and accuracy of the model.
[0032] In step S150, the intermediate AI model interacts with the first user to obtain interaction data, and the intermediate AI model evaluates the credibility of the first user based on the interaction data to obtain a confidence value evaluation result of the first user.
[0033] In some embodiments, the above-mentioned evaluation of the credibility of the first user based on the interaction data using the intermediate AI model can include the following steps: The intermediate AI model evaluates the credibility of the first user based on the similarity of the interaction data of the first user and other real users, the number of target real users whose interaction data similarity is greater than a specified similarity threshold, the confidence value evaluation result of the target real user, and the activity level of the first user, to obtain a confidence value evaluation result of the first user.
[0034] By comparing the interaction data of the first user with the interaction data of other real users except for the first user, similarities in behavior patterns, preferences or other characteristics are identified. This similarity analysis helps to discover potential group behavior rules and preliminarily classify users based on these rules. Furthermore, for target real users whose interaction data similarity is greater than a specified threshold, the number of target real users is counted, and a weighted calculation is performed in combination with the confidence value evaluation result corresponding to each target real user. In this way, not only the behavior patterns of similar users are considered, but also the trust scores of existing users are introduced, thereby improving the reliability and accuracy of the evaluation result.
[0035] Further, the above-mentioned use of the intermediate AI model to evaluate the trustworthiness of the first user according to the similarity of the interaction data to the interaction data of other real users except for the first user, the number of target real users whose interaction data similarity is greater than a specified similarity threshold, the confidence value evaluation result corresponding to the target real user, and the activity level of the first user, to obtain the confidence value evaluation result of the first user, can specifically include the following steps: The intermediate AI model is used to evaluate the trustworthiness of the first user according to the similarity of the interaction data to the interaction data of other real users except for the first user, the number of target real users whose interaction data similarity is greater than a specified similarity threshold, the confidence value evaluation result corresponding to the target real user, and the activity level of the first user, to obtain the confidence value evaluation result of the first user, by the following formula: = + α × + β × ; wherein, represents; represents the interaction data similarity of the i-th target real user and the first user; represents the confidence value evaluation result of the i-th target real user; represents the activity level of the first user; represents the number of target real users whose interaction data similarity is greater than a specified similarity threshold; α represents the influence degree of the adjustment behavior bias on the confidence value evaluation; β represents the influence degree of the activity level on the confidence value evaluation.
[0036] Through the above-mentioned data processing method of the calculation formula, the data of the confidence value evaluation result of the first user can be more accurate.
[0037] Step S160, in response to the confidence value evaluation result of the first user being greater than the specified confidence value, determining the user comment data of the first user corresponding to the multiple fields as real user comment data, to expand the user range of the user real network graph by using the multiple first users.
[0038] For example, the confidence value of the first user is compared with the specified confidence value. If the confidence value of a certain first user is greater than the specified value, the user is identified as a trusted user. For the first user identified as trusted, the system automatically extracts all the comment data published by the user in multiple fields and marks it as real user comment data. Based on the extracted real user comment data, the system starts to build or update the user real network graph. This includes but is not limited to the interaction between users, the interest distribution of users on different topics, the influence of users, etc. For each newly confirmed real user comment data, the system adds it to the graph as a new node, and creates corresponding connection edges according to the comment content, reply object and other factors to reflect the relationship between users and the association between users and topics. In order to ensure the accuracy and reliability of the user real network graph, the system should regularly check the quality of the newly added data to ensure that there is no misjudgment or introduction of false information. Based on the feedback of actual application effect, the confidence value calculation method is continuously adjusted, the data source range is expanded or the graph construction strategy is improved to improve the performance and user experience of the entire system. Using the updated user real network graph, more accurate services can be provided in multiple scenarios such as recommendation system, public opinion analysis, market research, etc. Since user behavior and network environment are constantly changing, it is necessary to continuously monitor and update the user real network graph in a timely manner to ensure its timeliness and effectiveness.
[0039] Through such a process, the system not only can effectively identify real user comment data, but also can build a more rich and accurate user real network graph on this basis, thereby providing strong support for various application scenarios.
[0040] Step S170, adding the user real network graph after expanding the user range to the learning sample of the intermediate AI model to obtain a second learning sample, and using the second learning sample to optimize the intermediate AI model to obtain an optimized target AI model.
[0041] In the embodiment of the present application, by fusing multimodal data of text, time series and images and real user comment data, the system first obtains a rich source of information. These data are used to generate user portraits of target users, ensuring an accurate understanding of user behavior and preferences. Using multiple user portraits and their related multimodal data and comment data, the system further constructs a real network graph of users. This graph not only captures the behavioral characteristics of individual users, but also reveals the relationships and influence patterns between users, thereby providing a more three-dimensional data structure for subsequent analysis. Moreover, by adding the real network graph of users to the learning samples of the initial AI model, the model can learn in a context closer to the real world, thereby effectively improving the authenticity and representativeness of the learning samples and reducing the dirty data problem caused by sample bias. Furthermore, through the interaction between the intermediate AI model and the first user, the user's credibility is evaluated based on the interaction data. Estimation helps to filter out unreliable or malicious data inputs and further purifies the training sample set. When the confidence value of the first user reaches the specified standard, his comment data in multiple fields will be included in the real user comment data pool, expanding the coverage of the user's real network graph. This expansion not only increases the amount of data, but also enhances data diversity, which is conducive to improving the robustness and generalization ability of the model. Finally, the purified and expanded user real network graph is used as the second learning sample to optimize the intermediate AI model to obtain the target AI model. This model has higher accuracy and reliability because it has higher quality learning materials, which greatly reduces the impact of dirty data. Therefore, through the above data processing and model training strategies, the quality of AI model training samples is gradually improved, thereby achieving the technical effect of significantly reducing dirty data, and solving the technical problem that the AI model training effect is poor due to the large amount of dirty data in the AI model training samples.
[0042] In some embodiments, as Figure 2 As shown, after optimizing the intermediate AI model using the second learning sample in step S170 to obtain the optimized target AI model, the method may further include the following steps: Step S210: In response to a user's learning request for a target streaming media, the target AI model is used to simulate a real user's streaming media query process by increasing the browsing speed, and the first-level multimodal data in the target streaming media is obtained through the simulated streaming media query process; Step S220: Using the first-level multimodal data as the center, diverge the peripheral related data of the relevant user request learning data through the user's real network graph to obtain the second-level multimodal data of multiple relevant user request learning data; Step S230, preferentially using the first-level multi-modal data to train the target AI model for the main user request trend, and using the second-level multi-modal data to train the target AI model for the secondary related request trend after the release of the computing power of the main user request trend training is completed, to obtain a final AI model.
[0043] By simulating the real user's query process through fast-forward browsing of the target streaming media, first-hand information (i.e., first-level multi-modal data) about the streaming media can be quickly obtained, including but not limited to video content, audio, subtitles, and other information. Moreover, based on the first-level multi-modal data, related user request learning data is further mined through a real user network graph to obtain second-level multi-modal data. In this way, the data sources can be expanded, and the diversity and accuracy of the recommendation system can be increased. Furthermore, by preferentially using the first-level multi-modal data for main user request trend training of the AI model, it is ensured that the model first understands the core content and directly related information, improving the accuracy of preliminary training. After completing the main trend training, the second-level multi-modal data is used for more extensive secondary related request trend training, so that the model can recognize and respond to more diverse user needs, improving overall performance.
[0044] In some embodiments, after obtaining the confidence value evaluation result of the first user by using the intermediate AI model to evaluate the credibility of the first user according to the interaction data in step S150 described above, the method can further include the following steps: In response to the confidence value evaluation result of the first user being less than or equal to a specified confidence value, determining the user comment data of the first user corresponding to multiple fields as dirty data samples; and using the dirty data samples to perform reverse training on the intermediate AI model to purify the dirty data learned by the intermediate AI model.
[0045] In the embodiments of the present application, the dirty data samples are determined according to the confidence value evaluation result of the first user, which can effectively identify data that may contain bias, false information or malicious content, preventing these data from polluting the model. By using these dirty data samples to perform reverse training on the intermediate AI model, the model can learn how to correct or ignore the influence of these bad data, thereby improving the accuracy and reliability of the model in real application scenarios.
[0046] In this way, the AI model can better cope with the inevitable dirty data problem in actual application, enhance its performance when facing imperfect input, and make the model more stable and reliable. Furthermore, the purified model can provide higher quality decision support, whether for product recommendation, content filtering or other applications requiring data analysis, to bring better user experience and service quality.
[0047] In some embodiments, the method can further include the steps of: encoding the real user's behavior logs, sensor time series signals, and detection images into feature vectors using a heterogeneous data adapter, and aligning the feature vectors through a cross-modal attention method; inputting the risk features generated by the AI model into a dual-channel model of XGBoost algorithm model and LightGBM algorithm model for preliminary risk assessment; and dynamically adjusting the risk threshold of the risk features based on the preliminary risk assessment and real-time feedback module of reinforcement learning through interaction with the user environment.
[0048] By using a heterogeneous data adapter to unify data from different sources (such as user behavior logs, sensor time series signals, and detection images) into feature vectors and aligning them through a cross-modal attention mechanism, the effective fusion of multi-source heterogeneous data is ensured. This step greatly enhances the system's understanding of complex environments, making risk assessment more comprehensive and accurate. Then, these feature vectors are input into a dual-channel model composed of XGBoost algorithm model and LightGBM algorithm model for preliminary risk assessment. These two gradient boosting decision tree-based algorithms are known for their efficiency and accuracy in the field of machine learning, and their combined use not only improves the speed of risk identification but also enhances the reliability of the results.
[0049] In the embodiments of the present application, based on the results of preliminary risk assessment and the real-time feedback module of reinforcement learning, the system can interact with the user environment and dynamically adjust the risk threshold of the risk features. This approach allows the system to update its decision-making strategy in real time based on the latest data and environmental changes, greatly improving the adaptability and flexibility of the risk assessment system.
[0050] In some embodiments, the method can further include the steps of: comparing the feature differences between specified high-risk features and specified low-risk features using a comparative attribution analysis method, and using the feature differences to build a multi-granularity explanation engine; generating natural language analysis data based on the multi-granularity explanation engine, and converting the AI model decision into compliance explanation results through a large model based on the natural language analysis data.
[0051] In the embodiments of the present application, the comparative attribution analysis method is used to compare the specified high-risk features and low-risk features, and identify the significant differences between them. This process helps to clarify which features are crucial for risk assessment, thereby improving the accuracy and reliability of the model. Based on the above feature differences, a multi-level (multi-granularity) explanation engine is constructed. This engine can analyze the model decisions from different levels (such as individual instances, group behaviors, etc.), providing a more comprehensive and detailed understanding. It not only focuses on macro-level risk patterns, but also values micro-level case analysis. Subsequently, based on the analysis data generated by the multi-granularity explanation engine, natural language processing technology is used to convert complex AI decisions into easy-to-understand language descriptions. This step is particularly important for non-technical personnel, as it makes the AI model's decision-making process transparent and easy to understand. Further, through the large model, these natural language descriptions are converted into compliance explanation results that meet regulatory requirements, making AI decisions not only effective but also compliant.
[0052] Figure 3 A structural diagram of an AI model decision device based on multi-modal data is provided. As shown in Figure 3 The AI model decision device based on multi-modal data 300 includes: An acquisition module 301 is configured to acquire multi-modal data fused with text, time sequence, and images, and real user comment data related to the multi-modal data. A first generation module 302 is configured to generate a user portrait of a target user corresponding to the real user comment data according to the real user comment data and the multi-modal data corresponding to the comment. A second generation module 303 is configured to generate a user real network graph of a plurality of target users based on the relationship between the multi-modal data and the real user comment data according to the user portraits of the target users. A training module 304 is configured to add the user real network graph to a learning sample of an initial AI model to obtain a first learning sample, and perform model training on the initial AI model using the first learning sample to improve the real credibility of the learning sample of the initial AI model through the user real network graph and purify the dirty data learned by the initial AI model, thereby obtaining a trained intermediate AI model. An evaluation module 305 is configured to interact with a first user through the intermediate AI model to obtain interaction data, and perform credibility evaluation on the first user according to the interaction data using the intermediate AI model to obtain a confidence value evaluation result of the first user. The determining module 306 is configured to determine the user comment data of the first user in a plurality of fields as the real user comment data to expand the user range of the user real network graph by using a plurality of the first users, in response to the confidence value evaluation result of the first user being greater than the specified confidence value. The optimization module 307 is configured to add the user real network graph after the user range is expanded to the learning sample of the intermediate AI model to obtain a second learning sample, and perform model optimization on the intermediate AI model by using the second learning sample to obtain the optimized target AI model.
[0053] The AI model decision device based on multi-modal data provided by the embodiments of the present application has the same technical features as the AI model decision method based on multi-modal data provided by the above embodiments, and can solve the same technical problems and achieve the same technical effects.
[0054] The electronic device provided by the embodiments of the present application, as shown in Figure 4 The electronic device 400 includes a processor 402, a memory 401, and the memory stores a computer program executable on the processor, and the processor implements the steps of the method provided by the above embodiments when executing the computer program.
[0055] Referring to Figure 4 , the electronic device further includes a bus 403 and a communication interface 404, and the processor 402, the communication interface 404 and the memory 401 are connected through the bus 403; the processor 402 is configured to execute the executable modules stored in the memory 401, such as computer programs.
[0056] The memory 401 can include a high-speed random access memory (RAM) and can also include a non-volatile memory such as at least one disk memory. The communication between the system network element and at least one other network element is realized through at least one communication interface 404 (which can be wired or wireless), and the Internet, a wide area network, a local area network, a metropolitan area network, etc. can be used.
[0057] The bus 403 can be an ISA bus, a PCI bus, or an EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the convenience of representation, Figure 4 Only one bidirectional arrow is used in the figure, but it does not mean that there is only one bus or one type of bus.
[0058] The memory 401 is configured to store a program, and the processor 402 is configured to execute the program after receiving an execution instruction. The method performed by the device defined by the process disclosed in any embodiment of the present application can be applied to the processor 402, or implemented by the processor 402.
[0059] The processor 402 can be an integrated circuit chip having a processing capability. In the implementation process, each step of the above method can be completed by the integrated logic circuit of hardware in the processor 402 or the instruction in the form of software. The processor 402 mentioned above can be a general processor, including a central processing unit (CPU), a network processor (NP), etc.; can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. Each method, step and logic block disclosed in the embodiments of the present application can be implemented or executed. The general processor can be a microprocessor or the processor can also be any conventional processor. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as a hardware code processor for execution, or a combination of hardware and software modules in the code processor for execution. The software module can be located in a random access memory, a flash memory, a read only memory, a programmable read only memory or an electrically erasable programmable memory, a register, etc. The storage medium in the art. The storage medium is located in the memory 401, and the processor 402 reads the information in the memory 401, and combines the hardware to complete the steps of the above method.
[0060] Corresponding to the above AI model decision method based on multi-modal data, the embodiment of the present application also provides a computer readable storage medium, the computer readable storage medium stores computer executable instructions, when the processor calls and runs the computer executable instructions, the computer executable instructions make the processor run the steps of the above AI model decision method based on multi-modal data.
[0061] The AI model decision device based on multi-modal data provided in the embodiments of the present application can be specific hardware on a device or software or firmware installed on the device, etc. The device provided in the embodiments of the present application has the same implementation principle and technical effects as the foregoing method embodiments, and for brief description, the part not mentioned in the device embodiment part can refer to the corresponding content in the foregoing method embodiments. Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the system, device and unit described above can refer to the corresponding process in the foregoing method embodiments, which will not be repeated here.
[0062] In the embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are only schematic. For example, the division of the units is only a logical function division, and there can be another division manner in actual implementation. For example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some communication interfaces, devices or units, and can be electrical, mechanical or other forms.
[0063] For another example, the flowcharts and block diagrams in the drawings show the possible implementation architecture, function and operation of the devices, methods and computer program products according to the embodiments of the present application. In this regard, each block in the flowcharts or block diagrams can represent a module, a program segment or a part of code containing one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur in different orders from that shown in the drawings. For example, two consecutive blocks can actually be executed substantially in parallel, and sometimes they can be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and the combination of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0064] The units described as separate components can or can not be physically separated, and the components shown as units can or can not be physical units, i.e. they can be located in one place or distributed on multiple network units. Some or all of the units can be selected to achieve the purpose of the present embodiment according to actual needs.
[0065] In addition, each of the functional units in the embodiments of the present application can be integrated in one processing unit, or each unit can exist alone physically, or two or more units can be integrated in one unit.
[0066] The functions described can be implemented in hardware, software, firmware or any combination thereof. If implemented in software, the functions can be stored or transmitted over as one or more instructions or code on a computer-readable medium. Computer-readable media include both computer storage media and communication media including any medium that facilitates transfer of a computer program from one place to another. A storage medium can be any available media that can be accessed by a computer. By way of example, and not limitation, such computer-readable media can comprise RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and that can be accessed by a computer. Also, any connection is properly termed a computer-readable medium. For example, if the software is transmitted from a website, server, or other remote source using a coaxial cable, fiber optic cable, or twisted pair, then the coaxial cable, fiber optic cable, or twisted pair are included in the definition of medium. Disk and disc, as used herein, include compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), and Blu-Ray® disc where disks usually reproduce data magnetically, while discs reproduce data optically with lasers. Combinations of the above should also be included within the scope of computer-readable media.
[0067] It should be noted that: similar reference numbers and letters represent similar items in the following drawings, therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings, in addition, the terms "first", "second", "third" and the like are only used to distinguish description, and cannot be understood as indicating or implying relative importance.
[0068] Finally, it should be noted that: the above-described embodiments are merely specific embodiments of the present application, used to illustrate the technical solutions of the present application, and not to limit the same, the protection scope of the present application is not limited thereto, although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand: any skilled person familiar with the technical field can still modify or easily think of changes to the technical solutions recorded in the foregoing embodiments within the technical scope disclosed by the present application, or make equivalent replacement to part of the technical features; and these modifications, changes or replacements do not make the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application. All should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. An AI model decision-making method based on multimodal data, characterized in that: The method comprises: Acquire multimodal data that integrates text, time series, and images, as well as real user review data related to the multimodal data; Generating a user profile of a target user corresponding to the real user comment data according to the real user comment data and the multimodal data of the corresponding comments; Generating a plurality of real network graphs of the target users based on the user portraits corresponding to the target users and utilizing the relationship between the multimodal data and the real user comment data; Adding the user's real network graph to the learning sample of the initial AI model to obtain a first learning sample, and using the first learning sample to train the initial AI model, so as to improve the authenticity and credibility of the learning sample of the initial AI model through the user's real network graph and purify the dirty data learned by the initial AI model, thereby obtaining a trained intermediate AI model; Interacting with the first user through the intermediate AI model to obtain interaction data, and using the intermediate AI model to perform a credibility evaluation on the first user based on the interaction data to obtain a confidence value evaluation result of the first user; In response to the confidence value evaluation result of the first user being greater than a specified confidence value, determining the user review data corresponding to the first user in multiple fields as the real user review data, so as to expand the user range of the user real network graph by utilizing the multiple first users; The real network map of users after the user range is expanded is added to the learning sample of the intermediate AI model to obtain a second learning sample, and the intermediate AI model is optimized using the second learning sample to obtain an optimized target AI model.
2. The method according to claim 1, characterized in that After optimizing the intermediate AI model using the second learning sample to obtain an optimized target AI model, the method further includes: In response to a user's learning request for a target streaming media, the target AI model is used to simulate a real user's streaming media query process by increasing the browsing speed, and the first-level multimodal data in the target streaming media is obtained by simulating the streaming media query process; Taking the first-level multimodal data as the center, diverging peripheral related data of the related user request learning data through the user real network graph to obtain second-level multimodal data of multiple related user request learning data; The first-level multimodal data is preferentially used to train the target AI model for the main user request trend, and after the computing power release of the main user request trend training is completed, the second-level multimodal data is used to train the target AI model for the secondary related request trend to obtain the final AI model.
3. The method according to claim 1, characterized in that The using the intermediate AI model to perform a credibility evaluation on the first user based on the interaction data to obtain a confidence value evaluation result of the first user includes: The intermediate AI model is used to perform a credibility evaluation on the first user based on the similarity between the interaction data and the interaction data of other real users except the first user, the number of target real users whose interaction data similarity is greater than a specified similarity threshold, the confidence value evaluation result corresponding to the target real user, and the activity of the first user to obtain the confidence value evaluation result of the first user.
4. The method according to claim 3, characterized in that The intermediate AI model is used to perform a credibility evaluation on the first user based on the similarity between the interaction data and the interaction data of other real users other than the first user, the number of target real users whose interaction data similarity is greater than a specified similarity threshold, the confidence value evaluation result corresponding to the target real user, and the activity of the first user, to obtain the confidence value evaluation result of the first user, including: The intermediate AI model is used to perform a credibility evaluation on the first user using the following formula based on the similarity between the interaction data and the interaction data of other real users other than the first user, the number of target real users whose interaction data similarity is greater than a specified similarity threshold, the confidence value evaluation result corresponding to the target real user, and the activity of the first user, to obtain the confidence value evaluation result of the first user: = + α × + β × ; in, express; Indicates the similarity of the interaction data between the i-th target real user and the first user; Represents the confidence value evaluation result of the i-th target real user; Indicates the activity level of the first user; The number of users representing the target real users whose interaction data similarity is greater than a specified similarity threshold; α Indicates the degree of influence of adjusting behavioral deviation on confidence value evaluation; β Indicates the impact of activity on confidence evaluation.
5. The method according to claim 1, wherein After using the intermediate AI model to perform a credibility evaluation on the first user based on the interaction data to obtain a confidence value evaluation result of the first user, the method further includes: In response to a confidence value evaluation result of the first user being less than or equal to the specified confidence value, determining the user comment data of the first user corresponding to the multiple fields as dirty data samples; The intermediate AI model is reversely trained using the dirty data samples to purify the dirty data learned by the intermediate AI model.
6. The method according to claim 1, characterized in that Also includes: Using heterogeneous data adapters, real user behavior logs, sensor time series signals, and detection images are encoded into feature vectors, and the feature vectors are aligned using cross-modal attention. The risk features generated by the AI model are input into the dual-channel model of the XGBoost algorithm model and the LightGBM algorithm model for preliminary risk assessment; The real-time feedback module based on the preliminary risk assessment and reinforcement learning dynamically adjusts the risk threshold of the risk feature by interacting with the user environment.
7. The method according to claim 6, characterized in that Also includes: Comparing the feature differences between designated high-risk features and designated low-risk features using a comparative attribution analysis method, and building a multi-granularity explanation engine using the feature differences; Natural language analysis data is generated based on the multi-granularity interpretation engine, and AI model decisions are converted into compliance interpretation results through a large model based on the natural language analysis data.
8. An AI model decision-making device based on multimodal data, characterized in that: include: An acquisition module is used to acquire multimodal data that integrates text, time series, and images, as well as real user review data related to the multimodal data; A first generating module, configured to generate a user profile of a target user corresponding to the real user comment data based on the real user comment data and the multimodal data of the corresponding comments; A second generating module is configured to generate a plurality of real network graphs of the target users based on the user portraits corresponding to the target users and utilizing the relationship between the multimodal data and the real user comment data; A training module, configured to add the user's real network graph to the learning samples of the initial AI model to obtain a first learning sample, and use the first learning sample to train the initial AI model, so as to improve the authenticity and credibility of the learning samples of the initial AI model through the user's real network graph and purify dirty data learned by the initial AI model, thereby obtaining a trained intermediate AI model; an evaluation module, configured to interact with a first user through the intermediate AI model to obtain interaction data, and use the intermediate AI model to perform a credibility evaluation on the first user based on the interaction data to obtain a confidence value evaluation result of the first user; a determination module configured to, in response to a confidence value evaluation result of the first user being greater than a specified confidence value, determine user review data corresponding to multiple fields of the first user as the real user review data, so as to expand the user scope of the user real network graph by utilizing the multiple first users; An optimization module is used to add the user's real network map after the user range is expanded to the learning sample of the intermediate AI model to obtain a second learning sample, and use the second learning sample to optimize the intermediate AI model to obtain an optimized target AI model.
9. An electronic device comprising a memory and a processor, wherein the memory stores a computer program that can be run on the processor, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-executable instructions. When the computer-executable instructions are called and executed by a processor, the computer-executable instructions prompt the processor to execute the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Method and device for identifying enterprise risk
CN113361963A
User portrait generation method and device, electronic equipment and storage medium
CN117743848A
Information processing method, device and equipment and computer readable medium
CN119441489A
Software-defined science and education resource dynamic collaborative recommendation method and system based on knowledge graph
CN119622115A
Adaptive learning question-answering system and method based on multi-modal interaction
CN120256575A