Multi-modal large model auxiliary interpretable auditing and self-adaptive evaluation method and device

Through the multimodal large-model auxiliary interpretable audit method, the problems of high manual review pressure and indefinite guidance in the existing video audit mechanism are solved, efficient and transparent video auditing is achieved, and user autonomy and responsible content management of the platform are enhanced.

CN120339911APending Publication Date: 2025-07-18SHANGHAI JIAOTONG UNIV
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510404236.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing video review mechanism relies on manual and machine learning models, with high false positive rates, high pressure on manual review, low enthusiasm for creators, insufficient review guidance, and lack of autonomy and self-drivenity of users, making it difficult to effectively improve the social responsibility level of videos.

Method used

The multimodal large-modal model assisted interpretable audit method is adopted to obtain visual and audio elements by disassembling video data, generate keyframe information and text information, and conduct multimodal analysis in combination with social responsibility standards, quantify it into a comprehensive rating, and provide detailed audit suggestions.

Benefits of technology

It improves the efficiency of video review and the initiative of creators, reduces human resources and time costs, enhances users' autonomy and platform reputation in content management, and promotes a healthy platform ecosystem.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339911A_ABST
    Figure CN120339911A_ABST
Patent Text Reader

Abstract

The invention provides a multi-modal large model auxiliary interpretable auditing and self-adaptive evaluation method and device, and relates to the technical field of artificial intelligence. The method comprises the following steps: acquiring video data created by a user, disassembling the video data to obtain a visual element and an audio element, and obtaining key frame information and text information according to the visual element and the audio element; generating cue words according to the text information; analyzing the key frame information and the text information according to the cue word and a preset social responsibility standard to obtain a multi-modal analysis result; the multi-modal analysis result is quantified into comprehensive rating through an evaluation function; and obtaining suggestions of the video data according to the comprehensive rating, the key frame information, the text information and the multi-modal analysis result. The invention provides a video content responsibility evaluation and suggestion framework system based on a multi-modal large model, and creates an auditing framework which takes a user as a center and has high responsiveness and transparency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a multimodal large model-assisted interpretable review and adaptive evaluation method and device. Background Art

[0002] In the current rapid development of technologies such as streaming media and short videos, a vast amount of video information is published on the public Internet every day. The quality of these video contents is uneven. Therefore, it is particularly important to explore a more convenient and efficient review mechanism.

[0003] Currently, video review mainly relies on manual review, adopting a review mechanism of "user upload - platform review - video release". Although an automated review plus manual review process is also adopted in some steps, on the one hand, this review mechanism requires a large amount of human and material resources, and on the other hand, it will also affect the enthusiasm and initiative of creators during this process, and has limited effect on improving the social responsibility level of the entire video.

[0004] Currently, some researchers also use models such as machine learning to detect hate speech, etc., but these are all for clear violations, with limited attention to morality and social responsibility, and there are also problems such as lag and poor interpretability.

[0005] Currently, the review mechanism of UGC (User Generated Content) platforms is mainly carried out in two stages: after the user completes and uploads the content, the platform will first preliminarily detect possible problems of the content according to internal detection algorithms such as machine learning, and then manual intervention is carried out for further review and marking. This review mechanism has the following problems: 1. The platform's automated detection is often inaccurate, resulting in a high false positive rate, increasing the pressure and cost of manual review, and misjudgment will also dampen the enthusiasm of creators. 2. For the current review mechanism of the platform, the reasons and improvement measures given are too general, with limited guiding significance for creators to modify, and the modification is more likely to be superficial and does not improve the social responsibility quality of the overall content. 3. The current review framework makes the review process concentrated on the platform, and users are passive recipients of the results, restricting the self-driving nature of user content evaluation. Summary of the Invention

[0006] In order to solve the problems existing in the existing review framework, embodiments of the present invention provide a multimodal large model-assisted interpretable review and adaptive evaluation method and device. The technical solutions are as follows:

[0007] On the one hand, a multimodal large model-assisted interpretable review and adaptive evaluation method is provided. This method is implemented by an assisted interpretable review and adaptive evaluation device, and the method includes:

[0008] S1. Obtain the video data created by the user, disassemble the video data to obtain visual elements and audio elements, and obtain key frame information and text information based on the visual elements and audio elements.

[0009] S2. Generate prompts according to the text information.

[0010] S3. Analyze the key frame information and text information according to the prompts and the preset social responsibility standards to obtain multi-modal analysis results.

[0011] S4. Quantify the multi-modal analysis results into a comprehensive rating through an evaluation function.

[0012] S5. Obtain suggestions for the video data based on the comprehensive rating, key frame information, text information, and multi-modal analysis results.

[0013] Optionally, obtaining the key frame information and text information based on the visual elements and audio elements in S1 includes:

[0014] S11. Preprocess the video data;

[0015] S12. Calculate the inter-frame feature difference:

[0016] Input the preprocessed video data into a convolutional feature extraction neural network model to extract feature vectors, and calculate the inter-frame feature difference based on the extracted feature vectors. The calculation formula is as follows:

[0017] Δ f (t) = ||f(t) - f(t - 1)||

[0018] Where f(t) is the feature vector extracted from the t-th frame, f(t - 1) is the feature vector extracted from the (t - 1)-th frame, and ||·|| represents the L2 norm of the vector;

[0019] S13. Calculate the optical flow difference:

[0020] Use the Farneback algorithm to calculate the optical flow between two consecutive frames. After obtaining the two-dimensional optical flow field V(t), convert it to polar coordinate form and calculate the average value of the optical flow amplitude as the optical flow difference. The calculation formula is as follows:

[0021]

[0022] Where u i (t) and v i (t) are the horizontal and vertical components of the i-th pixel in the t-th frame optical flow field respectively, and N is the total number of pixels;

[0023] S14. Calculate the audio difference:

[0024] Calculate the pitch trajectory of the audio using Pratt and calculate the audio difference. The calculation formula is as follows:

[0025] Δ pitch (t) = pitch(t) - pitch(t - 3)

[0026] Among them, pitch(t) is the pitch at time t. Taking a 3 - second interval is to avoid having key frames at the start of each sentence due to speaking tone;

[0027] S15. Text sentiment polarity calculation:

[0028] Process the audio element using the Whisper model to obtain text information. Then divide the text into sentences and calculate the sentiment polarity of each sentence using the Flair natural language framework. When the sentiment polarity polarity_score > 0.7, locate the time period of the sentence in the video, and take the frame with the largest optical flow difference within the time period as the key frame;

[0029] S16. Set the threshold automatic adjustment mechanism:

[0030] For the inter - frame feature difference, optical flow difference, and audio difference, calculate the 75th percentile of all their difference values as the dynamic threshold. Denote the feature difference set as {Δ f (t)}, the optical flow difference set as {Δ flow (t)}, and the audio difference set as {Δ pitch (t)}. Then the automatically adjusted thresholds are respectively:

[0031] T f = percentile 75 ({Δ f (t)})

[0032] T flow = percentile 75 ({Δ flow (t)})

[0033] T pitch = percentile 75 ({Δ pitch (t)})

[0034] When any difference value exceeds the corresponding threshold and meets the minimum time interval condition between adjacent key frames, it is determined as a key frame;

[0035] S17. Key frame determination:

[0036] When traversing the video frames, only when the calculated Δ f (t) ≥ T f or Δflow (t) ≥ T flow or Δ pitch (t) ≥ T pitch or when the polarity_score > 0.7 and the time interval Δt between the current frame and the previous key frame is ≥ the minimum time interval, the video frame is saved as a key frame;

[0037] S18. Result visualization for auxiliary analysis:

[0038] To intuitively reflect the changing trends of the inter-frame feature differences and optical flow differences, a composite line chart is drawn, the positions of the key frames are marked, and it is saved as an image file for subsequent parameter optimization and quality assessment.

[0039] Optionally, the convolutional feature extraction neural network model includes five stages for extracting feature vectors:

[0040] Stage 1 includes four parts, namely the convolutional layer CONV, the batch normalization BN layer, the RELU activation function, and the max pooling MaxPool layer;

[0041] Stages 2 - 5 include two parts, Conv Block and ID Block, which are two different residual blocks. The ConvBlock is used to change the dimension of the feature map; the ID Block is used to keep the dimension of the feature map unchanged and focuses on enhancing the feature representation. There are differences in the parameter settings for Stages 2 - 5;

[0042] After the five stages, it includes three parts: the global average pooling layer Avg Pool, the Dropout layer, and the Flattening layer, which are used to optimize the output and improve the generalization ability of the model;

[0043] The final output layer outputs the extracted feature vectors.

[0044] Optionally, generating prompt words according to the text information in S2 includes:

[0045] S21. Input the text information into a large language model to generate a high - level summary, and capture the type and content of the video according to the high - level summary.

[0046] S22. According to the high - level summary, determine the key monitoring points through the large language model and the preset social responsibility standards; among them, the preset social responsibility standards include: compliance, credibility, fairness and inclusiveness, privacy protection and data security, social impact and responsibility, and the appropriateness of emotional expression.

[0047] S23. According to the key monitoring points, generate prompt words through the large language model; among them, the types of prompt words include: summary generation prompt words, monitoring point prompt words, interpretability reason prompt words, comprehensive evaluation prompt words, and suggestion generation prompt words.

[0048] Optionally, analyze the key frame information and text information according to the prompt words and the preset social responsibility criteria in S3 to obtain the multimodal analysis results, including:

[0049] S31. For the key frame information, identify and analyze the core elements of each frame through a large language model to obtain the analysis results of each frame.

[0050] S32. Evaluate the analysis results of each frame based on the preset social responsibility criteria to obtain the evaluation results of each frame, and synthesize the evaluation results of all frames to obtain the key frame modal analysis result.

[0051] S33. Analyze the text information through a large language model and prompt words to obtain the text modal analysis result.

[0052] S34. Aggregate the key frame modal analysis result and the text modal analysis result to obtain the multimodal analysis result.

[0053] Optionally, quantify the multimodal analysis result into a comprehensive rating through an evaluation function in S4, including:

[0054] S41. Integrate the content, type, key frame modal analysis result, text modal analysis result, and multimodal analysis result of the video to obtain the input of the large language model.

[0055] S42. Define the dimension grading information according to the preset social responsibility criteria to establish the rating standard information; among them, the dimension grading information includes: evaluating whether the video complies with relevant laws, regulations, and platform rules to avoid violations or irregularities; evaluating whether the content is accurate and based on reliable sources to avoid misleading statements or false information; evaluating whether to avoid unequal remarks and reflect respect for different backgrounds, cultures, and viewpoints; evaluating whether to respect and protect the privacy of others to avoid unauthorized exposure of personal information; evaluating the potential impact of the video on society and whether to avoid promoting negative issues or misleading the public; evaluating whether the emotional expression is consistent with the content and whether to avoid excessive manipulation of emotions or triggering unnecessary emotional reactions.

[0056] S43. According to the input of the large language model and the rating standard information, use the chain of thought method to obtain the rating results of six dimensions, and convert the rating results of the six dimensions into scores of the six dimensions according to the mapping function.

[0057] S44. Calculate the comprehensive score according to the scores of the six dimensions, and convert the comprehensive score into a comprehensive rating according to the inverse mapping function; among them, the comprehensive rating includes grade A, grade B, grade C, and grade D.

[0058] On the other hand, a multimodal large model-assisted interpretable auditing and adaptive evaluation device is provided. This device is applied to the multimodal large model-assisted interpretable auditing and adaptive evaluation method, and the device includes:

[0059] An acquisition module, configured to acquire video data created by a user, disassemble the video data to obtain visual elements and audio elements, and obtain key frame information and text information based on the visual elements and audio elements.

[0060] A prompt word generation module, configured to generate prompt words according to the text information.

[0061] An analysis module, configured to analyze the key frame information and text information according to the prompt words and a preset social responsibility standard to obtain a multimodal analysis result.

[0062] A comprehensive rating module, configured to quantify the multimodal analysis result into a comprehensive rating through an evaluation function.

[0063] An output module, configured to obtain suggestions for the video data according to the comprehensive rating, key frame information, text information, and multimodal analysis result.

[0064] Optionally, the acquisition module is further configured to:

[0065] S11. Process the visual elements using a convolutional neural network ResNet-50 model and optical flow analysis to obtain key frame information.

[0066] S12. Process the audio elements using a Whisper model to obtain text information.

[0067] Optionally, the prompt word generation module is further configured to:

[0068] S21. Input the text information into a large language model to generate a high-level summary, and capture the type and content of the video according to the high-level summary.

[0069] S22. Determine key monitoring points according to the high-level summary through the large language model and a preset social responsibility standard; wherein, the preset social responsibility standards include: compliance, credibility, fairness and inclusiveness, privacy protection and data security, social impact and responsibility, and moderation of emotional expression.

[0070] S23. Generate prompt words through the large language model according to the key monitoring points; wherein, the types of prompt words include: summary generation prompt words, monitoring point prompt words, interpretability reason prompt words, comprehensive judgment prompt words, and suggestion generation prompt words.

[0071] Optionally, the analysis module is further configured to:

[0072] S31. For the key frame information, use a large language model to identify and analyze the core elements of each frame, and obtain the analysis results of each frame.

[0073] S32. For the analysis results of each frame, evaluate them based on the preset social responsibility standards to obtain the evaluation results of each frame, and integrate the evaluation results of all frames to obtain the key frame modal analysis results.

[0074] S33. For the text information, analyze it through a large language model and prompt words to obtain the text modal analysis results.

[0075] S34. Aggregate the key frame modal analysis results and the text modal analysis results to obtain the multi-modal analysis results.

[0076] Optionally, the comprehensive rating module is further used for:

[0077] S41. Integrate the content, type, key frame modal analysis results, text modal analysis results, and multi-modal analysis results of the video to obtain the input of the large language model.

[0078] S42. According to the preset social responsibility standards, define the dimension grading information and establish the rating standard information; wherein, the dimension grading information includes: evaluating whether the video complies with relevant laws, regulations, and platform rules to avoid violations or irregularities; evaluating whether the content is accurate and based on reliable sources to avoid misleading statements or false information; evaluating whether to avoid unequal remarks and reflect respect for different backgrounds, cultures, and viewpoints; evaluating whether to respect and protect the privacy of others to avoid unauthorized exposure of personal information; evaluating the potential impact of the video on society and whether to avoid promoting negative issues or misleading the public; evaluating whether the emotional expression is consistent with the content and whether to avoid excessive manipulation of emotions or triggering unnecessary emotional reactions.

[0079] S43. According to the input of the large language model and the rating standard information, use the chain of thought method to obtain the rating results of six dimensions, and convert the rating results of the six dimensions into scores of the six dimensions according to the mapping function.

[0080] S44. Calculate the comprehensive score according to the scores of the six dimensions, and convert the comprehensive score into a comprehensive rating according to the inverse mapping function; wherein, the comprehensive rating includes grade A, grade B, grade C, and grade D.

[0081] On the other hand, provide an auxiliary interpretable review and adaptive evaluation device, where the auxiliary interpretable review and adaptive evaluation device includes: a processor; a memory, and computer-readable instructions are stored on the memory, and when the computer-readable instructions are executed by the processor, any one of the methods in the above multi-modal large model assisted interpretable review and adaptive evaluation method is implemented.

[0082] On the other hand, a computer-readable storage medium is provided, in which at least one instruction is stored, and the at least one instruction is loaded and executed by a processor to implement any one of the above-mentioned multi-modal large model-assisted interpretable review and adaptive evaluation methods.

[0083] The beneficial effects brought by the technical solutions provided in the embodiments of the present invention at least include:

[0084] In the present invention, not only the autonomy of users in content refinement is enhanced, but also the content approval process is accelerated, the long "review - feedback - revision" cycle is reduced, and content review is transformed from reactive supervision to proactive guidance. Embedding this system into a social media platform can improve the efficiency of video review and the initiative of video creators, saving human resources and time costs; at the same time, by reducing harmful content, it can encourage users to actively participate in self-regulation, improve the platform reputation of users in responsible content management, and thus support a healthier platform ecosystem. BRIEF DESCRIPTION OF THE DRAWINGS

[0085] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for description in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.

[0086] Figure 1 is a flowchart of a multi-modal large model-assisted interpretable review and adaptive evaluation method provided by an embodiment of the present invention;

[0087] Figure 2 is a video publishing flowchart of an embedded intelligent detection machine provided by an embodiment of the present invention;

[0088] Figure 3 is an interpretability evaluation framework diagram provided by an embodiment of the present invention;

[0089] Figure 4 is a block diagram of the structure of a convolutional feature extraction neural network model provided by an embodiment of the present invention;

[0090] Figure 5 is an adaptation suggestion framework diagram provided by an embodiment of the present invention;

[0091] Figure 6 is a block diagram of a multi-modal large model-assisted interpretable review and adaptive evaluation device provided by an embodiment of the present invention;

[0092] Figure 7 is a structural schematic diagram of an auxiliary interpretable review and adaptive evaluation device provided by an embodiment of the present invention. Detailed implementation manners

[0093] The technical solutions in the present invention will be described below with reference to the accompanying drawings.

[0094] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as an "example" in the present invention should not be construed as being more preferred or more advantageous than other embodiments or design solutions. Rather, the use of the word "example" is intended to present concepts in a specific manner. In addition, in the embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one of the two can be selected.

[0095] In the embodiments of the present invention, "image" and "picture" can sometimes be used interchangeably. It should be noted that when the difference is not emphasized, the meanings they express are the same. "(of)", "corresponding" and "corresponding" can sometimes be used interchangeably. It should be noted that when the difference is not emphasized, the meanings they express are the same.

[0096] In the embodiments of the present invention, sometimes a subscript such as W1 may be written in a non-subscript form such as W1. When the difference is not emphasized, the meanings they express are the same.

[0097] To make the technical problems, technical solutions and advantages to be solved by the present invention clearer, the following will be described in detail with reference to the accompanying drawings and specific embodiments.

[0098] The embodiments of the present invention provide a multimodal large model-assisted interpretable review and adaptive evaluation method, which can be implemented by an assisted interpretable review and adaptive evaluation device, and the assisted interpretable review and adaptive evaluation device can be a terminal or a server. As Figure 1 shown in the flowchart of the multimodal large model-assisted interpretable review and adaptive evaluation method, the processing flow of the method can include the following steps:

[0099] S1. Obtain the video data created by the user, disassemble the video data to obtain visual elements and audio elements, and obtain key frame information and text information according to the visual elements and audio elements.

[0100] In a feasible implementation manner, the assisted interpretable review and adaptive evaluation system based on the multimodal large model of the present invention includes an interpretability evaluation model and an adaptive recommendation model.

[0101] An interpretability evaluation model for obtaining a scoring result of video data based on the video data; the interpretability evaluation model includes: a video parsing module, a prompt word generation module, a detection module based on a multimodal large language model, and an evaluation module.

[0102] An adaptive suggestion model for obtaining modification suggestions for the video data according to the scoring result.

[0103] The present invention aims to provide a self-audit system for content responsibility before a user uploads video content to a platform, providing the user with an audit result, reasons, and improvement measures, and the functional framework is as Figure 2 shown.

[0104] Furthermore, the interpretability evaluation model: analyzes the user's video content in combination with social responsibility standards and gives a specific score, and the framework structure is as Figure 3 shown.

[0105] Optionally, obtaining the key frame information and text information according to the visual elements and audio elements in S1 includes:

[0106] S11. Preprocess the video data;

[0107] S12. Calculate the inter-frame feature difference:

[0108] Input the preprocessed video data into a convolutional feature extraction neural network model to extract feature vectors, and calculate the inter-frame feature difference according to the extracted feature vectors. The calculation formula is as follows:

[0109] Δ f (t) = ||f(t) - f(t - 1)||

[0110] where f(t) is the feature vector extracted from the t-th frame, f(t - 1) is the feature vector extracted from the (t - 1)-th frame, and ||·|| represents the L2 norm of the vector;

[0111] S13. Calculate the optical flow difference:

[0112] Use the Farneback algorithm to calculate the optical flow between two consecutive frames, obtain the two-dimensional optical flow field V(t), and then convert it into polar coordinate form, and calculate the average value of the optical flow amplitude as the optical flow difference. The calculation formula is as follows:

[0113]

[0114] where u i (t) and v i (t) are the horizontal and vertical components of the i-th pixel in the t-th frame optical flow field respectively, and N is the total number of pixels;

[0115] S14. Calculate the audio difference:

[0116] Use Pratt to calculate the pitch trajectory of the audio and calculate the audio difference. The calculation formula is as follows:

[0117] Δ pitch (t) = pitch(t) - pitch(t - 3)

[0118] Where pitch(t) is the pitch at time t. Taking a 3 - second interval is to avoid having key frames at the start of each sentence due to speaking tone;

[0119] S15. Text sentiment polarity calculation:

[0120] Process the audio element using the Whisper model to obtain text information. Then divide the text into sentences and use the Flair natural language framework to calculate the sentiment polarity of each sentence. When the sentiment polarity polarity_score > 0.7, locate the time period of the sentence in the video, and take the frame with the largest optical flow difference within the time period as the key frame;

[0121] S16. Set the threshold automatic adjustment mechanism:

[0122] For the inter - frame feature difference, optical flow difference, and audio difference, calculate the 75th percentile of all their difference values as the dynamic threshold. Denote the feature difference set as {Δ f (t)}, the optical flow difference set as {Δ flow (t)}, and the audio difference set as {Δ pitch (t)}. Then the automatically adjusted thresholds are respectively:

[0123] T f = percentile 75 ({Δ f (t)})

[0124] T flow = percentile 75 ({Δ flow (t)})

[0125] T pitch = percentile 75 ({Δ pitch (t)})

[0126] When any difference value exceeds the corresponding threshold and meets the condition of the minimum time interval (such as 2 seconds) between adjacent key frames, it is determined as a key frame;

[0127] S17. Key frame determination:

[0128] When traversing the video frames, only when the calculated Δf (t) ≥ T f or Δ flow (t) ≥ T flow or Δ pitch (t) ≥ T pitch or when polarity_score > 0.7 and the time interval Δt between the current frame and the previous key frame ≥ the minimum time interval (such as 2 seconds), the video frame is saved as a key frame;

[0129] S18. Result visualization for auxiliary analysis:

[0130] To intuitively reflect the changing trends of inter-frame feature differences and optical flow differences, a composite line chart is drawn, the positions of key frames are marked, and it is saved as an image file for subsequent parameter optimization and quality assessment.

[0131] Optionally, as Figure 4 shown, the convolutional feature extraction neural network model includes five stages for extracting feature vectors:

[0132] Stage 1 includes four parts, namely the convolutional layer CONV, the batch normalization BN layer, the RELU activation function, and the max pooling MaxPool layer;

[0133] Stages 2 - 5 include two parts, Conv Block and ID Block, which are two different residual blocks. ConvBlock is used to change the dimension of the feature map; ID Block is used to keep the dimension of the feature map unchanged and focus on enhancing the feature representation. There are differences in the parameter settings for Stages 2 - 5;

[0134] After the five stages, it includes three parts: the global average pooling layer Avg Pool, the Dropout layer, and the Flattening layer, which are used to optimize the output and improve the generalization ability of the model;

[0135] The final output layer outputs the extracted feature vectors.

[0136] S2. Generate prompt words according to the text information.

[0137] Optionally, the above step S2 may include the following steps S21 - S23:

[0138] S21. Input the text information into the large language model to generate a high - level summary, and capture the type and content of the video according to the high - level summary.

[0139] S22. According to the high - level summary, determine the key monitoring points through the large language model and the six dimensions of responsibility; among them, the six dimensions of responsibility include: compliance, credibility, fairness and inclusiveness, privacy protection and data security, social impact and responsibility, and the appropriateness of emotional expression.

[0140] S23. Generate prompting words through a large language model based on key monitoring points; among them, the types of prompting words include: abstract generation prompting words, monitoring point prompting words, interpretability reason prompting words, comprehensive evaluation prompting words, and suggestion generation prompting words.

[0141] In a feasible implementation manner, a prompting word generation module is used to generate targeted prompting words by the large language model according to the audio text content for subsequent detection.

[0142] Specifically, this task is divided into three subtasks: First, input the text into an LLM (Large Language Model) to generate a high-level abstract to capture the theme and key content of the video, so that the model can background the type and theme of the video; Second, let the large model refer to the six dimensions of responsibility (see Table 1 for the video content responsibility evaluation dimensions and rating criteria) to determine the key monitoring points and explain why there may be risks in certain areas; Third, based on these key monitoring points, let the large model generate an adaptive prompting word to guide the subsequent model for detailed content analysis. With the complete type and theme background and clear monitoring dimensions and directions, the large model can generate targeted and clear prompting words (see Table 2 for the system prompting word architecture).

[0143] Table 1

[0144]

[0145]

[0146] Table 2

[0147]

[0148]

[0149]

[0150] S3. Analyze the key frame information and text information according to the prompting words and the preset social responsibility standards to obtain the multimodal analysis results.

[0151] Optionally, the above step S3 may include the following steps S31 - S34:

[0152] S31. For the key frame information, identify and analyze the core elements of each frame through the large language model to obtain the analysis result of each frame.

[0153] S32. Analyze the results of each frame, evaluate them based on specific responsibility dimensions to obtain the evaluation results of each frame, and synthesize the evaluation results of all frames to obtain the key-frame modal analysis results.

[0154] S33. Analyze the text information through a large language model and prompting words to obtain the text modal analysis results.

[0155] S34. Aggregate the key-frame modal analysis results and the text modal analysis results to obtain the multi-modal analysis results.

[0156] In a feasible implementation, a detection module based on MLLM (Multimodal Large Language Model) is used to detect key-frame information and audio-text information according to the prompting words generated in the previous step.

[0157] Specifically, first, let the large model process all key frames, identify and analyze the core elements within each frame, such as prominent visual aspects like the environment, people, emotions, body language, interactions, and screen text. Then evaluate potential sensitivity or risk-related aspects based on specific responsibility dimensions. Finally, synthesize the results of all key frames to form a comprehensive evaluation of the entire video content and style; second, let the large model analyze the audio text according to the prompting words, focusing on areas such as conversations, scenes, actions, and narrative progress that may raise responsibility issues.

[0158] S4. Quantify the multi-modal analysis results into a comprehensive rating through an evaluation function.

[0159] Optionally, the above step S4 may include the following steps S41 - S44:

[0160] S41. Integrate the content, type, key-frame modal analysis results, text modal analysis results, and multi-modal analysis results of the video to obtain the input of the large language model.

[0161] S42. Define the dimension grading information according to the six dimensions of responsibility to establish the rating standard information; among them, the dimension grading information includes: evaluating whether the video complies with relevant laws, regulations, and platform rules to avoid violations or irregularities; evaluating whether the content is accurate and based on reliable sources to avoid misleading statements or false information; evaluating whether to avoid unequal remarks and demonstrate respect for different backgrounds, cultures, and viewpoints; evaluating whether to respect and protect the privacy of others to avoid unauthorized exposure of personal information; evaluating the potential impact of the video on society and whether to avoid promoting negative issues or misleading the public; evaluating whether the emotional expression is consistent with the content and whether to avoid excessive manipulation of emotions or triggering unnecessary emotional reactions.

[0162] S43. Based on the input of the large language model and the rating standard information, use the chain of thought method to obtain the rating results of six dimensions, and convert the rating results of the six dimensions into scores of the six dimensions according to the mapping function.

[0163] S44. Calculate the comprehensive score according to the scores of the six dimensions, and convert the comprehensive score into a comprehensive rating according to the inverse mapping function to obtain the rating result of the video data; among them, the comprehensive rating includes grade A, grade B, grade C, and grade D.

[0164] In a feasible implementation manner, the evaluation module is used to map the analysis result into a quantitative index for intuitive presentation and quantitative comparison.

[0165] Specifically, according to the chain of prompt method, the evaluation process is divided into four subtasks: First, integrate the video content, type, unimodal recognition result, and multimodal analysis output into the unified input of the large model to ensure the integrity of the materials and establish a complete and detailed background information for rating; Second, the present invention designs six responsibility dimensions and defines the dimension grading information in detail to establish complete rating standard information (see Table 1 for details); Third, use the chain of thought method to obtain the rating results of the six dimensions, and convert the dimension ratings into specific scores according to the mapping function; Fourth, calculate the comprehensive score according to the average scores of the six dimensions, and convert the comprehensive score into a comprehensive rating according to the inverse mapping function.

[0166] S5. Obtain suggestions for the video data according to the comprehensive rating, key frame information, text information, and multimodal analysis results.

[0167] In a feasible implementation manner, for the video data with a comprehensive rating of grade A, the modification suggestions given include the improvement of visibility and popularity.

[0168] For the video data with a comprehensive rating of grade B or grade C, the modification suggestions given include the improvement of responsible content.

[0169] For the video data with a comprehensive rating of grade D, the modification suggestions given include substantial revisions.

[0170] The adaptive suggestion model also includes a visual component suggestion module and a text component suggestion module.

[0171] Adaptive suggestion framework: Combine video image information, audio text information, rating information, etc., and give appropriate and feasible modification suggestions through the suggestion module based on the large language model. The framework structure is as Figure 5 shown.

[0172] Specifically, the model generates suggestions based on the rating level of the video: for highly rated videos (Level A), it mainly focuses on improving visibility and popularity, such as refining the title, description, and keywords to increase coverage. For videos rated B or C, the suggestions emphasize responsible content improvement, including improving factual accuracy, enhancing transparency, and adding clarifications when necessary. For videos with lower ratings (Level D), substantial revisions are recommended to address core issues and provide clear guidance for correcting misleading or inaccurate elements; in addition to content suggestions, the model also provides specific suggestions for visual and text components. Image-based suggestions include adjusting visual elements, updating scene settings, or replacing inappropriate images, and text-based guidance may focus on enriching content depth, adjusting the mood or tone, or improving logical accuracy, etc. Each suggestion includes specific examples to facilitate implementation, ensuring that creators can make precise and effective improvements according to the platform's expectations.

[0173] Through the present invention: 1. Users can conduct self-assessment and timely adjust the created content, breaking out of the traditional "review - feedback - modification" cycle and shortening the user's creation cycle; 2. This system generates clear and actionable feedback based on the user's video content, improving the transparency and practicality of the review, and at the same time deepening the creator's understanding of laws, regulations, and responsibility standards during this process, cultivating the creator's sense of responsibility; 3. Centered around the user, enhancing the user's subjective awareness, and at the same time continuously improving the accuracy and professionalism of the review through the user feedback system, achieving good human-machine interaction.

[0174] In the embodiment of the present invention, it not only enhances the user's autonomy in content refinement but also accelerates the content approval process, reduces the long "review - feedback - revision" cycle, and transforms content review from reactive supervision to proactive guidance. Embedding this system into a social media platform can improve the efficiency of video review and the initiative of video creators, saving human resources and time costs; at the same time, it can encourage users to actively participate in self-regulation by reducing harmful content, improving the platform reputation of users in responsible content management, thus supporting a healthier platform ecosystem.

[0175] Figure 6 It is a block diagram of a multimodal large model-assisted interpretable review and adaptive evaluation device shown according to an exemplary embodiment. This device is used for the multimodal large model-assisted interpretable review and adaptive evaluation method. Referring to Figure 6 , this device includes an acquisition module 610, a prompt word generation module 620, an analysis module 630, a comprehensive rating module 640, and an output module 650. Among them:

[0176] An acquisition module 610 is configured to acquire video data created by a user, disassemble the video data to obtain visual elements and audio elements, and obtain key frame information and text information based on the visual elements and the audio elements.

[0177] A prompt generation module 620 is configured to generate prompts according to the text information.

[0178] An analysis module 630 is configured to analyze the key frame information and the text information according to the prompts and a preset social responsibility standard to obtain a multimodal analysis result.

[0179] A comprehensive rating module 640 is configured to quantify the multimodal analysis result into a comprehensive rating through an evaluation function.

[0180] An output module 650 is configured to obtain suggestions for the video data according to the comprehensive rating, the key frame information, the text information, and the multimodal analysis result.

[0181] In an embodiment of the present invention, not only the autonomy of the user in content refinement is enhanced, but also the content approval process is accelerated, the long "review - feedback - revision" cycle is reduced, and the content review is changed from reactive supervision to proactive guidance. Embedding this system into a social media platform can improve the efficiency of video review and the initiative of video creators, save human resources and time costs; at the same time, it can encourage users to actively participate in self - regulation by reducing harmful content, improve the platform reputation of users in responsible content management, thereby supporting a healthier platform ecosystem.

[0182] Figure 7 It is a schematic structural diagram of an auxiliary interpretable review and adaptive evaluation device provided by an embodiment of the present invention. As Figure 7 shown, the auxiliary interpretable review and adaptive evaluation device may include the above - mentioned Figure 6 multimodal large - model - assisted interpretable review and adaptive evaluation device shown. Optionally, the auxiliary interpretable review and adaptive evaluation device 700 may include a first processor 7001.

[0183] Optionally, the auxiliary interpretable review and adaptive evaluation device 700 may further include a memory 7002 and a transceiver 7003.

[0184] Wherein, the first processor 7001 is connected to the memory 7002 and the transceiver 7003, such as through a communication bus.

[0185] Next, in combination with Figure 6 each component of the auxiliary interpretable review and adaptive evaluation device 700 will be specifically introduced:

[0186] Among them, the first processor 7001 is the control center of the auxiliary interpretable audit and adaptive evaluation device 700, which can be a single processor or a collective term for multiple processing elements. For example, the first processor 7001 is one or more central processing units (CPUs), or can be an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention. For example: one or more digital signal processors (DSPs), or one or more field programmable gate arrays (FPGAs).

[0187] Optionally, the first processor 7001 can execute various functions of the auxiliary interpretable audit and adaptive evaluation device 700 by running or executing software programs stored in the memory 7002 and calling data stored in the memory 7002.

[0188] In a specific implementation, as an embodiment, the first processor 7001 may include one or more CPUs, such as Figure 7 CPU0 and CPU1 shown in

[0189] In a specific implementation, as an embodiment, the auxiliary interpretable audit and adaptive evaluation device 700 may also include multiple processors, such as Figure 6 the first processor 7001 and the second processor 7004 shown in

[0190] Each of these processors can be a single-core processor (single-CPU), a multi-core processor (multi-CPU), or a graphics processing unit (GPU). Here, the processor can refer to one or more devices, circuits, and / or processing cores for processing data (such as computer program instructions).

[0191] Optionally, the memory 7002 may be a read-only memory (ROM) or other type of static storage device that can store static information and instructions, a random access memory (RAM) or other type of dynamic storage device that can store information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 7002 may be integrated with the first processor 7001 or may exist independently and be coupled to the first processor 7001 through the interface circuit of the assistive interpretable auditing and adaptive evaluation device 700 ( Figure 6 not shown in the figure), and the embodiments of the present invention do not make specific limitations on this.

[0192] The transceiver 7003 is used to communicate with a network device or with a terminal device.

[0193] Optionally, the transceiver 7003 may include a receiver and a transmitter ( Figure 7 not shown separately in the figure). Among them, the receiver is used to implement the receiving function, and the transmitter is used to implement the sending function.

[0194] Optionally, the transceiver 7003 may be integrated with the first processor 7001 or may exist independently and be coupled to the first processor 7001 through the interface circuit of the assistive interpretable auditing and adaptive evaluation device 700 ( Figure 7 not shown in the figure), and the embodiments of the present invention do not make specific limitations on this.

[0195] It should be noted that Figure 7 the structure of the assistive interpretable auditing and adaptive evaluation device 700 shown in the figure does not constitute a limitation on the router. The actual knowledge structure recognition device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.

[0196] In addition, the technical effects of the assistive interpretable auditing and adaptive evaluation device 700 may refer to the technical effects of the multi-modal large model assisted interpretable auditing and adaptive evaluation method described in the above method embodiments, and will not be elaborated here.

[0197] It should be understood that the first processor 7001 in the embodiments of the present invention may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0198] It should also be understood that the memory in the embodiments of the present invention may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM) or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of random access memory (RAM) are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM) and direct rambus RAM (DR RAM).

[0199] The above embodiments can be implemented in whole or in part by software, hardware (such as circuits), firmware, or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that contains one or more collections of available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, or magnetic tape), an optical medium (such as a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.

[0200] It should be understood that the term "and / or" in this document is merely a description of the association relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B can be singular or plural. In addition, the character " / " in this document generally represents an "or" relationship between the associated objects before and after, but it may also represent an "and / or" relationship, which can be understood specifically with reference to the context before and after.

[0201] In the present invention, "at least one" means one or more, and "a plurality" means two or more. "At least one of the following" or its similar expressions refer to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c can be single or multiple.

[0202] It should be understood that in various embodiments of the present invention, the magnitudes of the sequence numbers of the above processes do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.

[0203] Those of ordinary skill in the art will realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0204] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the devices, apparatuses, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0205] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.

[0206] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0207] In addition, the functional units in each embodiment of the present invention can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0208] When the above-mentioned functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or a part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0209] As described above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A multimodal large model-assisted interpretable auditing and adaptive evaluation method, characterized in that, The method includes: S1. Obtain the video data created by the user, disassemble the video data to obtain visual elements and audio elements, and obtain key frame information and text information based on the visual elements and audio elements; S2. Generate a prompt word according to the text information; S3. Analyze the key frame information and text information according to the prompt word and the preset social responsibility standard to obtain a multi-modal analysis result; S4. Quantify the multi-modal analysis result into a comprehensive rating through an evaluation function; S5. Obtain suggestions for the video data according to the comprehensive rating, key frame information, text information, and multi-modal analysis result.

2. The multimodal large model-assisted interpretable auditing and adaptive evaluation method according to claim 1, wherein, The obtaining of the key frame information and text information based on the visual elements and audio elements in S1 includes: S11. Preprocess the video data; S12. Frame-interval feature difference calculation: Input the preprocessed video data into a convolutional feature extraction neural network model to extract feature vectors, and calculate the frame-interval feature difference according to the extracted feature vectors. The calculation formula is as follows: Δ f (t) = ||f(t) - f(t - 1)|| where f(t) is the feature vector extracted from the t-th frame, f(t - 1) is the feature vector extracted from the (t - 1)-th frame, and ||·|| represents the L2 norm of the vector; S13. Optical flow difference calculation: Use the Farneback algorithm to calculate the optical flow between two consecutive frames. After obtaining the two-dimensional optical flow field V(t), convert it to polar coordinate form, and calculate the average value of the optical flow amplitude as the optical flow difference. The calculation formula is as follows: where, u i (t) and v i (t) are the horizontal and vertical components of the i-th pixel in the t-th frame optical flow field respectively, and N is the total number of pixels; S14. Audio difference calculation: Use Pratt to calculate the pitch trajectory of the audio and calculate the audio difference. The calculation formula is as follows: Δ pitch (t) = pitch(t) - pitch(t - 3) where pitch(t) is the pitch at time t. Taking 3 seconds as the interval is to avoid that every sentence starts as a key frame due to the speaking tone; S15. Text sentiment polarity calculation: Use the Whisper model to process the audio elements to obtain text information, then divide the text into sentences, and use the Flair natural language framework to calculate the sentiment polarity of each sentence. When the sentiment polarity polarity_score > 0.7, locate the time period of the sentence in the video, and take the frame with the largest optical flow difference within the time period as the key frame; S16. Set a threshold automatic adjustment mechanism: For the inter-frame feature differences, optical flow differences, and audio differences, calculate the 75th percentile of all their difference values as the dynamic thresholds. Denote the feature difference set as {Δ f (t)}, the optical flow difference set as {Δ flow (t)}, and the audio difference set as {Δ pitch (t)}. Then the automatically adjusted thresholds are respectively: T f = percentile 75 ({Δ f (t)}) T flow = percentile 75 ({Δ flow (t)}) T pitch = percentile 75 ({Δ pitch (t)}) When any difference value exceeds the corresponding threshold and meets the minimum time interval condition between adjacent key frames, it is determined as a key frame; S17. Key frame determination: When traversing video frames, only when the calculated Δ f (t) ≥ T f or Δ flow (t) ≥ T flow or Δ pitch (t) ≥ T pitch or polarity_score > 0.7, and the time interval Δt between the current frame and the previous key frame is ≥ the minimum time interval, will the video frame be saved as a key frame; S18. Result visualization-assisted analysis: To intuitively reflect the change trends of the frame-interval feature difference and the optical flow difference, draw a compound line chart, mark the positions of the key frames, and save them as image files for subsequent parameter optimization and quality evaluation.

3. The multimodal large model-assisted interpretable auditing and adaptive evaluation method according to claim 2, wherein The convolutional feature extraction neural network model includes five stages of extracting feature vectors: Stage 1 includes four parts, namely the convolutional layer CONV, the batch normalization BN layer, the RELU activation function, and the max pooling MaxPool layer; Stages 2 - 5 include two parts, the Conv Block and the ID Block. These are two different residual blocks, and the Conv Block is used to change the dimension of the feature map; The ID Block is used to keep the dimension of the feature map unchanged and focuses on enhancing the feature representation. There are differences in the parameter settings for stages 2 - 5; After five stages, it includes three parts: the global average pooling layer Avg Pool, the Dropout layer, and the Flattening layer, which are used to optimize the output and improve the generalization ability of the model; The final output layer outputs the extracted feature vector.

4. The multimodal large model-assisted interpretable auditing and adaptive evaluation method according to claim 1, characterized in that Generating prompting words according to the text information in S2 includes: S21. Input the text information into a large language model to generate a high - level summary, and capture the type and content of the video according to the high - level summary; S22. According to the high - level summary, determine the key monitoring points through the large language model and the preset social responsibility criteria; among them, the preset social responsibility criteria include: compliance, credibility, fairness and inclusiveness, privacy protection and data security, social impact and responsibility, and appropriateness of emotional expression; S23. Generate prompting words through the large language model according to the key monitoring points; among them, the types of the prompting words include: summary generation prompting words, monitoring point prompting words, interpretability reason prompting words, comprehensive evaluation prompting words, and suggestion generation prompting words.

5. The multimodal large model-assisted interpretable auditing and adaptive evaluation method according to claim 1, characterized in that Analyzing the key frame information and text information according to the prompting words and the preset social responsibility criteria in S3 to obtain the multi - modal analysis results, including: S31. For the key frame information, identify and analyze the core elements of each frame through the large language model to obtain the analysis results of each frame; S32. Evaluate the analysis results of each frame based on the preset social responsibility criteria to obtain the evaluation results of each frame, and synthesize the evaluation results of all frames to obtain the key frame modal analysis result; S33. Analyze the text information through the large language model and the prompting words to obtain the text modal analysis result; S34. Aggregate the key frame modal analysis result and the text modal analysis result to obtain the multi - modal analysis result.

6. The multimodal large model-assisted interpretable auditing and adaptive evaluation method according to claim 1, wherein Quantifying the multi - modal analysis result into a comprehensive rating in S4 includes: S41. Integrate the content, type, key frame modal analysis result, text modal analysis result, and multi - modal analysis result of the video to obtain the input of the large language model; S42. Define the dimension grading information according to the preset social responsibility criteria to establish the rating standard information; among them, the dimension grading information includes: evaluating whether the video complies with relevant laws, regulations, and platform rules to avoid violations or irregularities; evaluating whether the content is accurate and based on reliable sources to avoid misleading statements or false information; evaluating whether to avoid unequal remarks and reflect respect for different backgrounds, cultures, and viewpoints; evaluating whether to respect and protect the privacy of others to avoid unauthorized exposure of personal information; evaluating the potential social impact of the video and whether to avoid promoting negative issues or misleading the public; evaluating whether the emotional expression is consistent with the content and whether to avoid excessive manipulation of emotions or triggering unnecessary emotional reactions; S43. Based on the input of the large language model and the rating standard information, use the chain of thought method to obtain the rating results in six dimensions, and convert the rating results in the six dimensions into scores in the six dimensions according to the mapping function; S44. Calculate the comprehensive score according to the scores in the six dimensions, and convert the comprehensive score into a comprehensive rating according to the inverse mapping function; wherein, the comprehensive rating includes grade A, grade B, grade C, and grade D.

7. A multimodal large model-assisted interpretable auditing and adaptive evaluation device, which is used to implement the multimodal large model-assisted interpretable auditing and adaptive evaluation method according to any one of claims 1-6, characterized in that, The device includes: An acquisition module, configured to acquire video data created by a user, disassemble the video data to obtain visual elements and audio elements, and obtain key frame information and text information according to the visual elements and audio elements; A prompt word generation module, configured to generate prompt words according to the text information; An analysis module, configured to analyze the key frame information and text information according to the prompt words and a preset social responsibility standard to obtain a multi-modal analysis result; A comprehensive rating module, configured to quantify the multi-modal analysis result into a comprehensive rating through an evaluation function; An output module, configured to obtain suggestions for the video data according to the comprehensive rating, key frame information, text information, and multi-modal analysis result.

8. An auxiliary interpretable audit and adaptive evaluation device, characterized in that, The auxiliary interpretable audit and adaptive evaluation device includes: A processor; A memory, on which computer-readable instructions are stored, and when the computer-readable instructions are executed by the processor, the method described in any one of claims 1 to 6 is implemented.

9. A computer-readable storage medium, characterized in that, Program code is stored in the computer-readable storage medium, and the program code can be called by the processor to execute the method described in any one of claims 1 to 6.

Citation Information

Cited By

  • Law enforcement supervision scene-oriented method and device for realizing multi-modal large model key frame extraction, processor and readable storage medium thereof

    CN121236432A

  • Key frame determination method, electronic equipment, storage medium and computer program product

    CN121415320A

  • Video propagation effect analysis method and device, electronic equipment and storage medium

    CN121418623A