Video content quality analysis and knowledge recommendation method and system based on large model
By performing multi-dimensional quality scoring and similarity scoring on video content, combining it with user input questions, and dynamically adjusting the weights, we solved the problem that existing video recommendation algorithms are unable to identify professionalism and personalized recommendations, and achieved efficient and personalized influenza A knowledge recommendations.
Patent Information
- Application Number
- CN202510892447.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-09-30
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing video recommendation algorithms cannot effectively identify the professionalism and reliability of video content, cannot deeply understand medical knowledge, and lack personalized recommendation capabilities, making it difficult for users to obtain high-quality influenza A knowledge.
Through the video content quality analysis method based on large models, multi-dimensional quality scoring and similarity scoring are performed. Combined with user input questions and video content, the weights are dynamically adjusted to achieve personalized and accurate recommendations.
It improves the efficiency and experience of users in obtaining high-quality medical knowledge, ensures the professionalism and reliability of recommended video content, and meets personalized needs.
Smart Images

Figure CN120723976A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of biomedicine technology, and specifically relates to a video content quality analysis and knowledge recommendation method and system based on a large model. Background Art
[0002] With the widespread use of the internet and the rapid development of short video platforms, a vast amount of video content related to influenza A has emerged. The quality of this video content varies widely, ranging from popular science videos published by professional organizations to experiences or rumors shared by individual users. In this vast ocean of information, users struggle to quickly and accurately access high-quality, personalized information about influenza A. Traditional video recommendation algorithms primarily rely on user behavioral data or simple content features. These algorithms are inadequate in identifying the professionalism and reliability of video content, and they are unable to deeply understand the medical knowledge contained in the videos, hindering the efficiency and experience of users in acquiring high-quality medical knowledge. Furthermore, existing video recommendation algorithms often only process single-modal video data, neglecting the integration and complementarity of multimodal data. This results in incomplete and inaccurate information extraction, a lack of personalized recommendation capabilities, and an inability to provide precise knowledge recommendation services based on users' specific needs and preferences. Summary of the Invention
[0003] The present invention provides a video content quality analysis and knowledge recommendation method and system based on a large model. By performing multi-dimensional quality scoring on the acquired video content, combining the user input questions and the video content with similarity scores, and dynamically adjusting the relevant weights based on the user's personal preferences, it is possible to achieve a deep understanding of the user input questions and related video content and make personalized and accurate recommendations, thereby effectively improving the efficiency and experience of users in obtaining high-quality medical knowledge.
[0004] A video content quality analysis and knowledge recommendation method based on a large model, comprising: Based on the short video platform, the video is acquired by inputting keywords to obtain the video ID, video content text and video metadata; Based on the video content text, multi-dimensional text information extraction and problem diagnosis are performed to generate structured video text; Perform multi-dimensional scoring on structured video text, calculate the video quality score, and construct an initial high-quality video candidate set; Based on the initial high-quality video candidate set and the user input question, the similarity score between the user input question and the initial high-quality video candidate set is calculated, and the high-quality video candidate set is constructed; Combining the video quality score and similarity score, the weighted comprehensive recommendation score of each video in the high-quality video candidate set is calculated, and the top three recommended videos and the reasons for the recommendation are output to achieve precise recommendation.
[0005] By performing multi-dimensional quality scoring on the acquired video content, combining similarity scores between user-input questions and video content, and dynamically adjusting the relevant weights based on user personal preferences, we can achieve an in-depth understanding of user-input questions and related video content and make personalized and accurate recommendations, thereby effectively improving the efficiency and experience of users in obtaining high-quality medical knowledge.
[0006] Furthermore, the short video platform is based on the use of input keywords to obtain videos, obtain video ID, video content text and video metadata, including: Based on the short video platform, keywords are used for batch crawling to obtain the video ID and first video metadata corresponding to the keywords; the first video metadata includes the video title, video tags, video release time, and video source platform; Based on the obtained video ID and the first video metadata, blacklist label filtering and deduplication processing are performed to generate structured video ID and the first video metadata; Download and parse the video content corresponding to the video ID, and convert the video content audio into video content text to generate structured video content text and video secondary metadata; the video secondary metadata includes the video playback volume and the video like number; The structured video ID, video content text, video first video metadata, and video second video metadata are structured and saved.
[0007] By adopting multimodal data processing technology, video image and audio information is obtained and converted into video content text with a unified format, ensuring the integrity and integration of the data.
[0008] Furthermore, the multi-dimensional text information extraction and problem diagnosis based on the video content text to generate structured video text includes: Based on the video content, entity recognition technology in natural language processing is used to perform multi-dimensional information extraction and problem diagnosis, obtaining multi-dimensional entities associated with keywords; the multi-dimensional entities include symptom description, comparison with the common cold, medical advice, preventive measures, and treatment suggestions; Relationship extraction technology is used to identify the relationships between multi-dimensional entities, and a video knowledge graph is constructed based on the entities and the relationships between them to generate structured video text.
[0009] By identifying entities associated with keywords in video content and the relationships between entities, a video knowledge graph is constructed to assist manual review and quality improvement.
[0010] Furthermore, the structured video text is scored in multiple dimensions to calculate the video quality score and construct an initial high-quality video candidate set, including: Use a large language model to perform multi-dimensional scoring on structured video texts and construct an evaluation matrix for video texts; Based on the multi-dimensional scores of the video text and the weights of different large language models, the final scores of the video text in different dimensions are calculated; Based on the final scores of the video text in different dimensions, the video quality score is weighted and calculated; The expression of the video quality score is: ; Where, Indicates the The video quality score of each video text corresponding to the video; , Indicates the total number of dimensions; Indicates the The weights of the dimensions, and ; Indicates the Dimensions, Final rating of the video text; All videos whose video quality scores are not lower than the preset quality scores are selected to construct an initial high-quality video candidate set.
[0011] Through the use of multiple large language models, the video content text is evaluated in multiple dimensions to ensure the professionalism and reliability of the final recommended video content; and through the construction of an initial high-quality video candidate set, it is possible to preliminarily screen out videos that do not meet the video quality requirements.
[0012] Furthermore, the method of calculating a similarity score between the user input question and the initial high-quality video candidate set based on the initial high-quality video candidate set and the user input question, and constructing the high-quality video candidate set, includes: Preprocessing is performed based on the user input question and the video text in the initial high-quality video candidate set; the preprocessing includes word segmentation, stop word removal, and medical terminology standardization; Based on the user input question and the video text, the weights of each word in the user input question and the video text after word segmentation processing are calculated to obtain the user input question weight and the video text weight, and the first cosine similarity score of the user input question weight and the video text weight is calculated; The expression of the first cosine similarity score is: ; Where, The first cosine similarity score representing the video text and the user input question; Represents the weight of the term in the user input question; Represents the weight of the term in the video text; To indicate a term; Based on the user input question and the video text in the initial high-quality video candidate set, a pre-trained semantic large model is used to encode and generate a high-dimensional semantic vector, thereby obtaining a high-dimensional user input question vector and video text vector, and calculating the second cosine similarity score between the user input question vector and the video text vector; The expression of the second cosine similarity score is: ; Where, The second cosine similarity score representing the video text and the user input question; Represents a high-dimensional user input question vector, i.e. , represents the encoding function; Represents the video text vector, that is ; represents the norm; Combining the first cosine similarity score and the second similarity score, weightedly calculating the similarity score between the user input question and the initial high-quality video candidate set; The similarity score between the user input question and the initial high-quality video candidate set is expressed as: ; Where, Represents the similarity score between the user input question and the initial high-quality video candidate set; represents the first cosine similarity score weight; represents the second cosine similarity score weight; All videos with similarity scores not lower than the preset similarity score are selected to construct a high-quality video candidate set.
[0013] By performing dual-dimensional calculations of keyword matching and semantic understanding, a dual filtering mechanism with deep semantic matching can be quickly screened.
[0014] Furthermore, the method of calculating the weights of each word in the user input question and the video text after word segmentation processing based on the user input question and the video text to obtain the user input question weight and the video text weight, and calculating the first cosine similarity score of the user input question weight and the video text weight includes: Based on the user input question and video text, word segmentation is performed to obtain multiple terms; Calculate the frequency of each term in the user input question and video text respectively; Count the number of video texts containing the term and calculate the inverse document frequency of each term; Combine the word frequency and inverse document frequency corresponding to each term to calculate the weight of each term in the user input question and video text, and obtain the user input question weight and video text weight; Combine the user input question weight and the video text weight to calculate the first cosine similarity score of the user input question weight and the video text weight.
[0015] Furthermore, the video quality score and similarity score are combined to calculate the weighted comprehensive recommendation score of each video in the high-quality video candidate set, and the top three recommended videos and the reasons for the recommendation are output to achieve precise recommendation, including: Combining the video quality score and similarity score, a weighted comprehensive recommendation score is calculated for each video in the high-quality video candidate set. The expression of the comprehensive recommendation score of each video is: ; Where, Indicates the The comprehensive recommendation score of the video corresponding to the video text; Represents the similarity score weight; Indicates the video quality score weight; Arrange the videos in descending order according to their comprehensive recommendation scores, output the three videos with the highest comprehensive recommendation scores and the reasons for the recommendations, and achieve precise recommendations.
[0016] Furthermore, it also includes customizing the weights of each dimension according to the user's personal preferences.
[0017] A system for video content quality analysis and knowledge recommendation based on a large model, comprising: The video data acquisition module is used to acquire videos based on short video platforms by inputting keywords, obtaining video IDs, video content text, and video metadata; and to perform multi-dimensional text information extraction and problem diagnosis based on the video content text to generate structured video text; A high-quality video candidate set construction module is used to perform multi-dimensional scoring on structured video text, calculate video quality scores, and construct an initial high-quality video candidate set; and based on the initial high-quality video candidate set and the user input question, calculate the similarity score between the user input question and the initial high-quality video candidate set, and construct a high-quality video candidate set; The recommended video output module is used to combine the video quality score and similarity score, weightedly calculate the comprehensive recommendation score of each video in the high-quality video candidate set, and output the top three recommended videos and the reasons for the recommendation to achieve precise recommendation.
[0018] The beneficial effects of the present invention are: By performing multi-dimensional quality scoring on the acquired video content, combining similarity scores between user-input questions and video content, and dynamically adjusting the relevant weights based on user personal preferences, we can achieve an in-depth understanding of user-input questions and related video content and make personalized and accurate recommendations, thereby effectively improving the efficiency and experience of users in obtaining high-quality medical knowledge. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 is a flow chart of the present invention; Figure 2 Schematic diagram of the structure of the system in the present invention. DETAILED DESCRIPTION
[0020] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0021] It should be noted that various aspects of the embodiments within the scope of the appended claims are described below. It should be apparent that the aspects described herein can be embodied in a wide variety of forms, and any specific structure and / or function described herein is merely illustrative. Based on this disclosure, it should be understood by those skilled in the art that an aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number of aspects described herein can be used to implement an apparatus and / or practice a method. In addition, other structures and / or functionalities other than one or more of the aspects described herein can be used to implement this apparatus and / or practice this method.
[0022] In addition, in the following description, specific details are provided to facilitate a thorough understanding of the examples. For ordinary technicians in this field, the specific meanings of the above terms in this application can be understood in specific situations.
[0023] Example 1 Figure 1 The method is a large-scale model-based video content quality analysis and knowledge recommendation method. By performing multi-dimensional quality scoring on the acquired video content, combining the similarity score between the user-input question and the video content, and dynamically adjusting the weights based on the user's personal preferences, it can achieve a deep understanding of the user's input question and related video content and make personalized and accurate recommendations, thereby effectively improving the user's efficiency and experience in obtaining high-quality medical knowledge. The specific steps include the following: S1: Based on the short video platform, use the input keywords to obtain videos and obtain the video ID, video content text and video metadata; S11: Based on the short video platform, use keywords to perform batch crawling to obtain the video ID and primary video metadata corresponding to the keyword; the primary video metadata includes the video title, video tags, video release time, and video source platform; In this embodiment, based on the open API interfaces of short video platforms such as Douyin and Kuaishou, the core keywords and extended keywords of "Influenza A (H1N1)", "Influenza A symptoms", and "H1N1 treatment" are captured, and the captured video IDs are saved in Coze's Lark document.
[0024] S12: Based on the obtained video ID and first video metadata, perform blacklist labeling filtering and deduplication processing to generate structured video ID and first video metadata; In this embodiment, regular expressions are used to perform blacklist labeling to filter out irrelevant or low-quality content videos such as advertisements and entertainment, and deduplication is performed based on the uniqueness of the video content. A structured video ID list is output and saved in a JON format file to provide basic data support for subsequent analysis.
[0025] The video ID set is represented as , then the function for blacklist label filtering is defined as: ; Where, Represents a filtering function based on video ID; For video All video tags of ; Indicates that the video tag does not belong to the blacklist set; Indicates the video ID; Indicates filtering out the output of the video; S13: Download and parse the video content corresponding to the video ID, and convert the video content audio into video content text to generate structured video content text and video secondary metadata; the video secondary metadata includes the video playback volume and the video like number; In this embodiment, the Coze plug-in download-video is called to parse and download the video content corresponding to the video ID, and the SpeechToxt plug-in in the Alibaba Cloud Bailian platform is called to obtain the audio content of the downloaded video. The video content is converted into video content text by integrating speech recognition (ASR), and OCR technology is used to extract graphic information such as cover graphics. The video content text with a unified format is generated through multimodal data fusion technology, and the video second metadata in JSON format is output for storage to provide basic data support for the integrity of subsequent data analysis and multimodal information fusion.
[0026] S14: Structurally save the structured video ID, video content text, video first video metadata, and video second video metadata; S2: Based on the video content text, perform multi-dimensional text information extraction and problem diagnosis to generate structured video text; S21: Based on the video content text, use entity recognition technology in natural language processing to perform multi-dimensional information extraction and problem diagnosis, and obtain multi-dimensional entities associated with keywords; In this embodiment, when multi-dimensional information is extracted from the input ASR video content text, the multi-dimensional information includes: Symptom description (Symptom), which includes systemic symptoms, respiratory system symptoms, nervous system symptoms, and digestive system symptoms; Comparison with the common cold (Comparison), which includes the differences in symptoms, course of disease, and risk of complications compared with the common cold; Medical advice, including critical symptoms such as abnormal fever, respiratory system, impaired consciousness, and risk of dehydration; Precautionary measures, which include personal protection, environmental management, and health management measures; Treatment recommendations (Treatment), which include identifying the type of medication, marking applicable symptoms, and mentioning the specific name of the medication; It should be noted that when performing the above-mentioned multi-dimensional information extraction, if the information of a certain dimension cannot be extracted, "None" is output, and the information that can be extracted under the dimension is structured and output in JSON format.
[0027] In this embodiment, a text classification model is used to identify three typical problems, including incorrect information, ambiguous expressions, and missing information, and a diagnostic report is generated that includes the problem type, location fragment, and correction suggestions.
[0028] S22: Use relationship extraction technology to identify the relationships between multi-dimensional entities, build a video knowledge graph based on the entities and the relationships between them, and generate structured video text; In this embodiment, a visualization tool is used to output a heat map showing the distribution of high-frequency terms and a word cloud to highlight the knowledge graph of core knowledge points. The structured video text includes entity lists, relationship pairs, question tags, etc.
[0029] S3: Perform multi-dimensional scoring on the structured video text, calculate the video quality score, and construct an initial high-quality video candidate set; S31: Use a large language model to perform multi-dimensional scoring on structured video texts and construct an evaluation matrix for video texts; Among them, the large language model is set to ,Right now , the video text is ,Right now , and define the multi-dimensional evaluation index as ,Right now In this embodiment, there are five multi-dimensional evaluation indicators, namely Description of symptoms, Indicates comparison with the common cold, Indicate medical advice, Indicates precautionary measures, Indicates treatment recommendations.
[0030] Among them, after multi-dimensional scoring of the video text, the evaluation matrix of the constructed video text is: ; Where, Indicates the Evaluation matrix for video texts; Indicates the A large language model is In the dimension The rating of each video text ranges from 0 to 100 points; Represents the evaluation matrix dimension, where ; S32: Based on the scores of the video text in multiple dimensions and combined with the weights of different large language models, the final scores of the video text in different dimensions are calculated; Among them, the definition of major language models is in The weights of the dimensions are: ; ; ; ; Where, Indicates the A large language model is In the dimension The weight of the video text; Indicates the A large language model is In the dimension The recognition degree of a video text is in the range of , that is, the larger the recognition value is, the higher the consistency is; Indicates the A large language model and A large model in In the dimension The distance between the video and text; Indicates the A large language model is In the dimension Rating of video texts; Indicates the A large language model is In the dimension Rating of video texts; Indicates the A large language model and A large language model is In the dimension The absolute value of the difference in ratings of the video texts; Among them, video text In the The expression of the final score on each dimension is: ; Where, Indicates in In the dimension Final rating of the video text; S33: Based on the final scores of the video text in different dimensions, the video quality score is weighted and calculated; The expression of video quality score is: ; Where, Indicates the Each video text corresponds to a video quality score, which ranges from 0 to 100 points; , Indicates the total number of dimensions; Indicates the The weights of the dimensions, and ; Indicates the Dimensions, Final rating of the video text; S34: Select all videos whose video quality scores are not lower than a preset quality score to construct an initial high-quality video candidate set; In this embodiment, the preset quality score is set to 80 points, that is, video texts with a video quality score of not less than 80 points are selected to be included in the initial high-quality video candidate set. ,Right now And by constructing an initial high-quality video candidate set, videos that do not meet the video quality requirements are preliminarily screened.
[0031] In this embodiment, when using a large language model to perform multi-dimensional scoring on structured video text, each dimension is divided into four levels according to the extracted video text information, and the first level score range is 90~100 points, the second level score range is 80~90 points, the third level score range is 60~80 points, and the fourth level score range is below 60 points.
[0032] Specifically, based on the symptom description entity, typical symptoms, symptom severity and symptom duration are compared, and videos that mention all typical symptoms, symptom severity and symptom duration are classified as the first level, videos that mention major symptoms are classified as the second level, videos that mention some symptoms are classified as the third level, and videos with incomplete or incorrect symptom descriptions are classified as the fourth level. For example, comparing the typical symptoms of influenza A such as high fever, headache, and muscle aches, the severity of symptoms with a high fever above 39°C, and the duration of symptoms that peak in 3 to 5 days and last for 1 to 2 weeks, videos that mention sudden high fever accompanied by obvious headache and muscle aches are classified as the first level, videos that mention high fever, headache, muscle aches and fatigue that peak within a few days are classified as the second level, videos that have fever and general discomfort and symptoms that last for a week are classified as the third level, and videos that mention fever but do not mention other symptoms are classified as the fourth level. Specifically, based on the comparison entities with the common cold, the course of the disease, and the differences in complications, videos that mention all typical symptoms, symptom severity, and symptom duration are classified as the first level, videos that mention the main symptoms are classified as the second level, videos that mention some symptom differences are classified as the third level, and videos with incorrect or no symptom comparisons are classified as the fourth level. For example, comparing the course of influenza A, which is 3-5 days, and the common cold, which is 1-2 days, as well as the differences in the risks of influenza A pneumonia and myocarditis, videos that mention that the incidence of influenza A pneumonia is 6 times higher than that of the common cold, that the course of influenza A is usually 3-5 days, and that the course of the common cold is 1-2 days in Tongchuan are classified as the first level. Videos that mention that influenza A usually presents with high fever and severe systemic symptoms, and that the common cold presents with nasal congestion, runny nose, and milder systemic symptoms are classified as the second level. Videos that mention that influenza A is more serious than the common cold and that its symptoms last longer are classified as the third level. Videos that claim that the common cold is more dangerous than influenza A are classified as the fourth level. Specifically, based on the medical advice entity, by comparing symptoms that require medical treatment and precautions for high-risk groups, videos that mention multiple symptoms that require medical treatment and precautions for high-risk groups will be classified as the first level, videos that mention major medical symptoms will be classified as the second level, videos that mention simple medical tips will be classified as the third level, and videos that do not involve medical advice will be classified as the fourth level. For example, when comparing symptoms that require medical treatment such as persistent high fever and difficulty breathing, as well as high-risk groups such as pregnant women, the elderly, and children, videos that mention symptoms that require immediate medical treatment such as persistent high fever, difficulty breathing, and confusion will be classified as the first level, videos that mention the need for medical treatment for persistent high fever and difficulty breathing and the need for special attention for high-risk groups will be classified as the second level, videos that mention serious symptoms that require medical treatment but do not mention specific symptoms will be classified as the third level, and videos that do not mention medical advice will be classified as the fourth level; Specifically, based on the preventive measures entity, videos that mention frequent hand washing and mask wearing are classified as Level 1, videos that mention major preventive measures are classified as Level 2, videos that mention some preventive measures are classified as Level 3, and videos that do not mention preventive measures are classified as Level 4. For example, when comparing the recommendation to wash hands frequently with soap and running water for at least 20 seconds each time and the recommendation to wear a medical surgical mask or N95 mask, videos that mention preventive measures including frequent hand washing are classified as Level 1, videos that mention frequent hand washing and wearing masks as ways to prevent influenza A are classified as Level 2, videos that mention paying attention to personal hygiene and avoiding contact with patients are classified as Level 3, and videos that do not mention preventive measures are classified as Level 4. Specifically, based on the treatment recommendation entity, compare the medication advice and early medication reminder. Videos that mention antiviral medication advice, antipyretic and analgesic medication advice, and early medication reminder are classified into the first level. Videos that mention antipyretic and analgesic medication advice are classified into the second level. Videos that mention taking medicine but do not involve specific drugs are classified into the third level. Videos that do not involve medication advice are classified into the fourth level. For example, compare the medication advice of antiviral drugs such as oseltamivir, antipyretic analgesics such as ibuprofen and acetaminophen, and early medication reminder. Videos that mention antiviral drugs such as oseltamivir and zanamivir, antipyretic analgesics such as ibuprofen and acetaminophen, and the importance of early medication are classified into the first level. Videos that mention antipyretic analgesics such as ibuprofen or acetaminophen are classified into the second level. Videos that mention symptomatic treatment but do not specifically specify the drugs are classified into the third level. Videos that do not mention treatment advice are classified into the fourth level.
[0033] S4: Based on the initial high-quality video candidate set and the user input question, calculate the similarity score between the user input question and the initial high-quality video candidate set, and construct a high-quality video candidate set; S41: Based on the user input question and the video text in the initial high-quality video candidate set, perform preprocessing; In this embodiment, the preprocessing includes word segmentation, stop word removal, and medical term standardization. That is, use a word segmentation tool to segment the user input question, then perform stop word removal processing for words such as "of", "right", "excuse me" according to the cleaning rules, and perform medical term standardization processing. It should be noted that in practical applications, jieba can be used for Chinese questions and spaCy can be used for English questions for word segmentation.
[0034] S42: Based on the user input question and the video text, calculate the weights of each term after word segmentation in the user input question and the video text, obtain the user input question weight and the video text weight, and calculate the first cosine similarity score between the user input question weight and the video text weight; S421: Based on the user input question and the video text, perform word segmentation to obtain multiple terms; Among them, define the user input question as , based on the video text , after word segmentation, multiple terms are obtained ; S422: Calculate the term frequency of each term in the user input question and the video text respectively; Among them, the term in the user input question is: ; In the formula, Representing terms In user input problem word frequency in ; Among them, the term Text in video The word frequencies in are: ; Where, Representing terms Text in video word frequency in ; S423: Counting the number of video texts containing the term and calculating the inverse document frequency of each term; The definition includes the terms The number of video texts is , that is, its inverse document frequency is: ; Where, Representing terms The inverse document frequency of Indicates the total number of video texts; The operation avoids zero denominator and smoothes extreme values; S424: Calculate the weight of each term in the user input question and the video text by combining the term frequency and inverse document frequency corresponding to each term, and obtain the user input question weight and the video text weight; Among them, the term In user input problem The weights in are: ; Where, Represents the weight of the term in the user input question; Among them, the term Text in video The weights in are: ; Where, Represents the weight of the term in the video text; S425: Calculating a first cosine similarity score of the user input question weight and the video text weight by combining the user input question weight and the video text weight; Among them, combining the user input question and video text to construct the TF-IDF vector, the expression of the first cosine similarity score is obtained: ; Where, The first cosine similarity score represents the video text and the user input question. Its value range is [0,1]. That is, the larger the first cosine similarity score, the more problematic the user input. With video text The higher the similarity between them; S43: Based on the user input question and the video text in the initial high-quality video candidate set, a pre-trained semantic large model is used to encode and generate a high-dimensional semantic vector, thereby obtaining a high-dimensional user input question vector and video text vector, and calculating the second cosine similarity score between the user input question vector and the video text vector; In this embodiment, the pre-trained semantic large model is the coROM-Chinese-Medical model, which is used to encode user input questions and video text into high-dimensional vectors, that is, to generate high-dimensional user input question vectors and video text vectors.
[0035] Among them, the encoding process of the user input question vector is: ; Where, Represents a high-dimensional user input question vector; represents the encoding function; Among them, the encoding process of video text vector is: ; Where, Represents the video text vector; Among them, combining the user input question vector and the video text vector, the expression of the second cosine similarity score is obtained as follows: ; Where, The second cosine similarity score representing the video text and the user input question; represents the norm; S44: combining the first cosine similarity score and the second similarity score, and weightedly calculating a similarity score between the user input question and the initial high-quality video candidate set; The similarity score between the user input question and the initial high-quality video candidate set is expressed as: ; Where, Represents the similarity score between the user input question and the initial high-quality video candidate set; represents the first cosine similarity score weight; represents the second cosine similarity score weight; in this embodiment, , ; S45: Select all videos whose similarity scores are not less than a preset similarity score to construct a high-quality video candidate set; In this embodiment, the preset similarity score is set to 0.7, that is, video texts with a similarity score not less than 0.7 are selected to be included in the high-quality video candidate set. ,Right now And by constructing a high-quality video candidate set, we can filter out videos whose similarity scores do not meet the requirements.
[0036] By performing dual-dimensional calculations of keyword matching and semantic understanding, a dual filtering mechanism with deep semantic matching can be quickly screened.
[0037] S5: Combine the video quality score and similarity score to calculate the weighted comprehensive recommendation score of each video in the high-quality video candidate set, and output the top three recommended videos and the reasons for the recommendation to achieve precise recommendation; S51: Combine the video quality score and similarity score to calculate the weighted comprehensive recommendation score of each video in the high-quality video candidate set; Among them, the expression of the comprehensive recommendation score of each video is: ; Where, Indicates the The comprehensive recommendation score of the video corresponding to the video text; Represents the similarity score weight; Indicates the video quality score weight; S52: Arrange the videos in descending order according to their comprehensive recommendation scores, output the three videos with the highest comprehensive recommendation scores and the reasons for the recommendations, and achieve precise recommendations.
[0038] In actual application, the weights of each dimension can be customized according to the user's personal preferences. The system can prioritize the display of video texts that focus on dimension details according to the user's weight settings, allowing different types of users to flexibly optimize recommendation results based on actual needs and achieve accurate adaptation of medical knowledge services.
[0039] Example 2 In this embodiment, a video content quality analysis and knowledge recommendation method based on a large model is provided to capture and process videos related to the core keyword and extended keyword of "Jiaxing Influenza (H1N1)", analyze and evaluate the video quality, and recommend related videos.
[0040] T1: Based on the short video platform, use the input keywords to obtain videos and obtain the video ID, video content text and video metadata; T11: Based on the short video platform, use keywords to perform batch crawling to obtain the video ID and primary video metadata corresponding to the keyword; the primary video metadata includes the video title, video tags, video release time, and video source platform; In this embodiment, based on core keywords, the influenza A-related video IDs and video metadata released between 2020 and 2025 are obtained, and multi-source medical information such as Wanfang Medical Network document abstracts and electronic medical records are simultaneously integrated to construct an influenza A knowledge base containing 1,200 structured video texts and save it in Coze's Feishu document.
[0041] T12: Based on the obtained video ID and first video metadata, perform blacklist labeling filtering and deduplication processing to generate structured video ID and first video metadata; In this embodiment, advertisements and entertainment content are filtered out through regular expressions, and valid videos with a duration of more than 60 seconds are retained. Medical entities are then identified through the BERT model, and three types of professional terms, symptoms and drug examination items, are retained. Then, the TF-IDF threshold method is combined with Sentence-BERT semantic deduplication to remove duplicate videos with a cosine similarity score greater than 0.95. Finally, the quality of video recommendations is improved through video metadata enhancement, including video clarity scores, professional scores annotated by medical experts, authoritative scores of cited literature, and timeliness scores based on publication time. Multimodal medical texts are integrated, and sparse matrix calculation and normalization processing are used to ensure the consistency of basic data.
[0042] T13: Download and parse the video content corresponding to the video ID, convert the video content audio into video content text, and generate structured video content text and video secondary metadata; video secondary metadata includes video playback volume and video like count; In this embodiment, the Coze plug-in download-video is called to parse and download the video content corresponding to the video ID, and the SpeechToxt plug-in in the Alibaba Cloud Bailian platform is called to obtain the audio content of the downloaded video; the audio content is converted into video text by integrating speech recognition (ASR), and the deep neural network model is used for fine-tuning to achieve an audio content recognition accuracy of 95.2%; OCR technology is used to extract video subtitles and graphic information; a dedicated recognition library is established through special symbols in medical charts, so that the extraction completeness reaches 98.7%; video content text with a unified format is generated through multimodal data fusion technology, and the Transformer architecture is used for semantic fusion. The weights of the three types of information, audio, subtitles, and images, are dynamically allocated in combination with the attention mechanism, and the second video metadata in JSON format is output for storage to provide basic data support for the integrity of subsequent data analysis and multimodal information fusion; Table 1 shows the video ID is Video content analysis and video metadata diagram.
[0043] Table 1 Video content analysis and video metadata diagram
[0044] T14: Save the structured video ID, video content text, video first video metadata, and video second video metadata in a structured manner; T2: Based on the video content text, perform multi-dimensional text information extraction and problem diagnosis to generate structured video text; T21: Based on the video content text, use entity recognition technology in natural language processing to perform multi-dimensional information extraction and problem diagnosis, and obtain multi-dimensional entities associated with keywords; In this embodiment, when extracting multi-dimensional information from the input ASR video content text, the multi-dimensional information includes symptom description , compared with the common cold , medical advice , preventive measures , treatment recommendations Table 2 shows a schematic diagram of multi-dimensional information category recognition.
[0045] Table 2 Schematic diagram of multi-dimensional information category identification
[0046] In this embodiment, a text classification model is used to identify three typical problems, including incorrect information, ambiguous expressions, and missing information, and a diagnostic report is generated that includes the problem type, location fragment, and correction suggestions.
[0047] T22: Use relationship extraction technology to identify the relationships between multi-dimensional entities, build a video knowledge graph based on the entities and the relationships between them, and generate structured video text; In this example, the Pyecharts visualization tool was used to output a heatmap to display the distribution of high-frequency terms and a word cloud to highlight the core knowledge points. The structured video text includes entity lists, relationship pairs, question tags, and TF-IDF weights.
[0048] T3: Perform multi-dimensional scoring on the structured video text, calculate the video quality score, and construct an initial high-quality video candidate set; T31: Use a large language model to perform multi-dimensional scoring on structured video texts and construct an evaluation matrix for video texts; In this embodiment, the video text Input to multiple large language models In the example, each large language model has five dimensions Perform independent scoring to obtain the In the dimension Rating of video texts , integrated to form an evaluation matrix .
[0049] T32: Based on the multi-dimensional scores of the video text and the weights of different large language models, the final scores of the video text in different dimensions are calculated; In this embodiment, the distance between the score differences between large language models is calculated. , normalize it to eliminate the dimension effect and convert it into the first A large language model is In the dimension Recognition of video text , and then calculate the major language models based on the recognition The weight of the dimension , and finally combine each large language model in the In the dimension Rating of video texts , weighted calculation of video text In the Final rating on each dimension ; T33: Calculate the video quality score based on the final scores of the video text in different dimensions; The video quality score is calculated based on the final score of the video text in different dimensions, combined with user-defined weights or default weights. .
[0050] In this embodiment, the symptom description is set Compared with the common cold , medical advice , preventive measures , treatment recommendations The weights of are all 0.2. Table 3 shows the video text Ratings in different dimensions and video quality ratings.
[0051] Table 3 Video text Ratings in different dimensions and video quality ratings
[0052] T34: Select all videos with a video quality score not lower than a preset quality score to construct an initial high-quality video candidate set; In this embodiment, the preset quality score is set to 80 points, that is, video texts with a video quality score of not less than 80 points are selected to be included in the initial high-quality video candidate set. ,Right now By constructing an initial set of high-quality video candidates, we can preliminarily screen out videos that do not meet the quality requirements. In practical applications, we can require that the average quality scores of multiple large language models be greater than a preset quality score for video screening.
[0053] T4: Based on the initial high-quality video candidate set and the user input question, calculate the similarity score between the user input question and the initial high-quality video candidate set, and construct a high-quality video candidate set; T41: Preprocessing based on user input questions and video text in the initial high-quality video candidate set; In this embodiment, the preprocessing includes word segmentation, stop word removal, and medical term standardization.
[0054] Specifically, for the Chinese user input question "I suspect I have been infected with the influenza A virus and currently have symptoms of fever and nasal congestion. What medicines should I take to effectively relieve the symptoms?", the word segmentation results are obtained through word segmentation processing, such as "suspected", "infection", "influenza A virus", "fever", "nasal congestion", "symptoms", "taking", "medication", "relieve", and "symptoms". Then, according to the cleaning rules, stop words such as "please ask" are removed, so that the final pre-processed keyword items are "influenza A", "fever", "nasal congestion", "medication", and "relieve".
[0055] T42: Based on the user input question and the video text, calculate the weight of each word in the user input question and the video text after word segmentation processing, obtain the user input question weight and the video text weight, and calculate the first cosine similarity score of the user input question weight and the video text weight; T421: Perform word segmentation based on the user input question and video text to obtain multiple terms; T422: Calculate the frequency of each term in the user input question and video text respectively; T423: Count the number of video texts containing the term and calculate the inverse document frequency of each term; T424: Calculate the weight of each term in the user input question and video text by combining the term frequency and inverse document frequency corresponding to each term, and obtain the user input question weight and video text weight; T425: Calculate the first cosine similarity score of the user input question weight and the video text weight by combining the user input question weight and the video text weight; Combining user input issues and video text , construct the TF-IDF vector and get the first cosine similarity score .
[0056] T43: Based on the user input question and the video text in the initial high-quality video candidate set, a pre-trained semantic large model is used to encode and generate a high-dimensional semantic vector. This high-dimensional user input question vector and video text vector are obtained, and the second cosine similarity score of the user input question vector and video text vector is calculated. In this embodiment, the pre-trained semantic model is the coROM-Chinese-Medical model, which is used to encode the user input question and video text into a high-dimensional vector, that is, to generate a high-dimensional user input question vector and video text vector. Combining the user input question vector and the video text vector, the second cosine similarity score is obtained. .
[0057] T44: Combine the first cosine similarity score and the second similarity score to weightedly calculate the similarity score between the user input question and the initial high-quality video candidate set; Among them, the first cosine similarity score and the second similarity score are combined, and the weight of the first cosine similarity score is set to 0.4 and the second similarity score is set to 0.6 to obtain the similarity score between the user input question and the initial high-quality video candidate set. In this embodiment, Table 4 shows a similarity score table based on the initial high-quality video candidate set.
[0058] Table 4. Similarity scores based on the initial high-quality video candidate set
[0059] T45: Select all videos with similarity scores not less than the preset similarity score to build a high-quality video candidate set; In this embodiment, the preset similarity score is set to 0.7, that is, video texts with a similarity score not less than 0.7 are selected to be included in the high-quality video candidate set. ,Right now And by constructing a high-quality video candidate set, we can filter out videos whose similarity scores do not meet the requirements.
[0060] T5: Combines the video quality score and similarity score to calculate the weighted comprehensive recommendation score for each video in the high-quality video candidate set, and outputs the top three recommended videos and the reasons for the recommendation to achieve precise recommendation; T51: Combine the video quality score and similarity score to calculate the weighted comprehensive recommendation score of each video in the high-quality video candidate set; In this embodiment, it is set that the user is more concerned about the symptom description and treatment suggestions, so the symptom description is set Compared with the common cold , medical advice , preventive measures , treatment recommendations The personalized weights are 0.3, 0.1, 0.1, 0.1, and 0.4 respectively, and the video quality score weight is set to 0.2 and the similarity score weight is set to 0.8. The video text is obtained by weighted average operator aggregation Comprehensive recommendation score And the Top-3 ranking results. Table 5 shows the video text Comprehensive recommendation score , video quality rating , similarity score and sorting result table; Table 5 Video text Scores and ranking results table
[0061] T52: Arrange the videos in descending order based on their comprehensive recommendation scores, output the three videos with the highest comprehensive recommendation scores and the reasons for the recommendations, and achieve precise recommendations.
[0062] In this embodiment, the third place in the comprehensive recommendation score is video text , which is specifically "The flu has been quite strong recently. Data released by the China Centers for Disease Control and Prevention show that the positive rate of influenza virus is still rising, of which more than 99% are influenza A. I and other colleagues have also been infected once. How did we quickly relieve the symptoms? Today I will share with you. The first is oseltamivir, which is a commonly used antiviral drug for the treatment of influenza. It is generally more effective if taken within 48 hours after infection. It should be noted that oseltamivir is a prescription drug. Different dosage forms and specifications have different usage and dosage. It is best to use it under the guidance of a doctor according to your own situation. You can also choose Chinese patent medicines recently recommended by the Health and Health Commissions of Beijing, Anhui, Tianjin and other places, such as Lianhua Qingwen. I often keep it at home. It has good efficacy and can effectively inhibit the influenza A virus. As early as during the 2009 influenza A period, the Xunzheng Research has confirmed that annualized mild has a significant effect in reducing the severity of the disease in influenza A patients, and can significantly shorten the duration of influenza symptoms such as fever, cough, muscle aches, fatigue, and headache. In the same year, it was included in the list of human infection with influenza A (H1N) by the Ministry of Health. 1. Influenza diagnosis and treatment plan. Among the recommended drugs for treating influenza A (H1N1), I usually prepare some Lianhua Qingwen when the seasons change. It can also deal with the common cold and other respiratory diseases, improve the body's resistance, and help recover from the disease. And Chinese patent medicines have fewer side effects, and can be used by both the elderly and children. If the body temperature exceeds 38.5 degrees during the infection period of influenza A, oral antipyretics should also be used in time, such as hemaminophen, pineapple phenol, etc. Some people say what to do if the cough never gets better after influenza A? You can take cough suppressants in time and treat the symptoms. The whole influenza A symptom lasts for about a week. If the symptoms have not improved after this time, go to the hospital as soon as possible. Do you remember?" Output the video ratings for each dimension and the comprehensive recommendation score , video quality rating , similarity score The reasons for recommendation are shown in Table 6.
[0063] Table 6 Video text Ratings and reasons for recommendation
[0064] In this embodiment, the second place in comprehensive recommendation score is video text , which is specifically "What should I do if I am diagnosed with influenza A? In fact, after being infected with influenza A, we have a 48-hour golden control period. Timely medication can effectively relieve symptoms, shorten the course of the disease, and reduce suffering. If it is delayed, the condition may worsen and recovery will be very troublesome. What medicine should I take after being infected with influenza A? Then anti-influenza virus drugs are the key. Oseltamivir is commonly used in clinical practice. It can inhibit the replication and spread of the virus. It is best taken indoors within 48 hours. Its usage is suitable for 1-3 For people over the age of 12, it should be taken in the way prescribed by the doctor. Another drug is called Mabaloxavir, which is also an anti-influenza virus drug. Its advantage is that it only needs to be taken once, which is more convenient to take. It is suitable for people over the age of 5, and it also needs to be taken according to the doctor's orders. As for relieving symptoms, if fever, headache, muscle aches occur, then these serious manifestations can be treated with acetaminophen, pineapple phenol and other antipyretic and analgesic drugs. For children under 12 years old, use some aminophenol, and for children over 12 years old and adults, acetaminophen and ibuprofen are applicable. If the cough and sputum are obvious, you can use Anxiusuo, Methaphan and other cough and expectorant drugs. The specific dosage can be Refer to the method in the instructions. For some friends with underlying diseases, especially those with digestive system diseases, including gastric ulcers, etc., they must be cautious when taking these antipyretic and analgesic drugs. So everyone should pay attention to the drugs, not abuse them, and must choose according to the symptoms and the doctor's advice. During the medication period, if the symptoms are not relieved, or there is difficulty breathing, worsening symptoms, chest pain, etc., then you must seek medical attention immediately. Everyone must pay attention to this influenza A. Only by using the right medicine at the critical moment can you recover as soon as possible! Do you think this video is useful? Don't forget to forward it to relatives and friends, follow me, and pay attention to your physical health! "Output the video ratings in each dimension and the comprehensive recommendation score , video quality rating , similarity score The reasons for recommendation are shown in Table 7.
[0065] Table 7 Video text Ratings and reasons for recommendation
[0066] In this embodiment, the video text with the highest comprehensive recommendation score is , which is specifically "Recently, many people have been infected with viruses, especially influenza A, and many issues about medication have been discussed. Here are a few important points to talk about, and everyone should remember to tell their families. First, if you have flu symptoms, such as sudden high fever, accompanied by body muscle aches, sore throat, headache, etc., go for a pathogen test. If it is methylsulfide or ethylsulfide, you can take oseltamivir for 5 days or mabaloxavir once in the early stage. Second, amoxicillin, various cephalosporins, and azithromycin are not antipyretics, but antibiotics for treating bacterial infections. Axinomycin is also used to treat mycoplasma and antibacterial infections. Third, antibiotics do not treat viral infections. Fourth, ibuprofen and acetaminophen are antipyretic and analgesic drugs that can help reduce fever. Fifth, it is generally recommended to use antipyretics when the body temperature is above 38.5 degrees Celsius. For some elderly people with underlying diseases, or those with obvious symptoms but a body temperature below 38.5 Patients with fever who are mentally depressed or have other systemic symptoms should take antipyretics with caution. Sixth, pay attention to the dosage of antipyretic and analgesic drugs and take them according to the instructions. Do not use them continuously or overdose, otherwise there is a risk of liver and kidney damage. Seventh, tablets should be taken as a whole tablet, and do not crush or dissolve them before taking. Eighth, do not drink alcohol while taking the medicine. Ninth, people with underlying diseases, pregnant women and breastfeeding women should consult a doctor before taking the medicine. 10. Do not give adult medicines to infants and children. Eleventh, many people may take complex medicines when they have uncomfortable symptoms. For cold medicines prepared with prescriptions, you must make sure to read the various ingredients and contents clearly, such as acetaminophen and ephedrine magnesium tablets, which contain para-aminophenol and ephedrine, acetaminophen and ephedrine magnesium tablets, and acetaminophen yellow nanoparticles for children, as well as artificial bezoar green nanoparticles and acetaminophen yellow nanoparticles for redemption points. Basically, cold medicines with the words ammonia or acetaminophen have redemption points. Before taking the medicine, you must read the dosage in the instructions carefully to protect your body. You don’t need to take medicine if you are not sick. You must take the right medicine and don’t take too much. Please everyone! "Output the video ratings of each dimension and the comprehensive recommendation score , video quality rating , similarity score The reasons for recommendation are shown in Table 8.
[0067] Table 8 Video text Ratings and reasons for recommendation
[0068] Example 3 In this embodiment, a video content quality analysis and knowledge recommendation system based on a large model is provided, which includes a video data acquisition module, a high-quality video candidate set construction module, and a recommended video output module.
[0069] Specifically, the video data acquisition module is used to acquire videos based on short video platforms by inputting keywords, obtaining video IDs, video content texts, and video metadata; and based on the video content texts, it performs multi-dimensional text information extraction and problem diagnosis to generate structured video texts; Specifically, the high-quality video candidate set construction module is used to perform multi-dimensional scoring on structured video text, calculate the video quality score, and construct an initial high-quality video candidate set; and based on the initial high-quality video candidate set and the user input question, calculate the similarity score between the user input question and the initial high-quality video candidate set, and construct a high-quality video candidate set; Specifically, the recommended video output module is used to combine the video quality score and similarity score, weightedly calculate the comprehensive recommendation score of each video in the high-quality video candidate set, and output the top three recommended videos and the reasons for the recommendation to achieve precise recommendation.
[0070] Throughout this specification, reference to terms such as "one embodiment," "some embodiments," "examples," "specific examples," or "some examples" means that a specific feature, structure, material, or characteristic described in conjunction with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, schematic representations of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.
[0071] Although the embodiments of the present invention have been shown and described above, it will be understood that the above embodiments are illustrative and are not to be construed as limitations on the present invention. A person skilled in the art may change, modify, replace and modify the above embodiments within the scope of the present invention.
Claims
1. A video content quality analysis and knowledge recommendation method based on a large model, characterized in that: include: Based on the short video platform, the video is acquired by inputting keywords to obtain the video ID, video content text and video metadata; Based on the video content text, multi-dimensional text information extraction and problem diagnosis are performed to generate structured video text; Perform multi-dimensional scoring on structured video text, calculate the video quality score, and construct an initial high-quality video candidate set; Based on the initial high-quality video candidate set and the user input question, the similarity score between the user input question and the initial high-quality video candidate set is calculated, and the high-quality video candidate set is constructed; Combining the video quality score and similarity score, the weighted comprehensive recommendation score of each video in the high-quality video candidate set is calculated, and the top three recommended videos and the reasons for the recommendation are output to achieve precise recommendation.
2. The method for video content quality analysis and knowledge recommendation based on a large model according to claim 1, characterized in that: The short video platform uses the input keywords to obtain videos, obtain video ID, video content text and video metadata, including: Based on the short video platform, keywords are used for batch crawling to obtain the video ID and first video metadata corresponding to the keywords; the first video metadata includes the video title, video tags, video release time, and video source platform; Based on the obtained video ID and the first video metadata, blacklist label filtering and deduplication processing are performed to generate structured video ID and the first video metadata; Download and parse the video content corresponding to the video ID, and convert the video content audio into video content text to generate structured video content text and video secondary metadata; the video secondary metadata includes the video playback volume and the video like number; The structured video ID, video content text, video first video metadata, and video second video metadata are structured and saved.
3. The method for video content quality analysis and knowledge recommendation based on a large model according to claim 1, characterized in that: The multi-dimensional text information extraction and problem diagnosis based on the video content text to generate structured video text includes: Based on the video content, entity recognition technology in natural language processing is used to perform multi-dimensional information extraction and problem diagnosis, obtaining multi-dimensional entities associated with keywords; the multi-dimensional entities include symptom description, comparison with the common cold, medical advice, preventive measures, and treatment suggestions; Relationship extraction technology is used to identify the relationships between multi-dimensional entities, and a video knowledge graph is constructed based on the entities and the relationships between them to generate structured video text.
4. The method for video content quality analysis and knowledge recommendation based on a large model according to claim 1, characterized in that: The multi-dimensional scoring of the structured video text is performed to calculate the video quality score and construct an initial high-quality video candidate set, including: Use a large language model to perform multi-dimensional scoring on structured video texts and construct an evaluation matrix for video texts; Based on the multi-dimensional scores of the video text and the weights of different large language models, the final scores of the video text in different dimensions are calculated; Based on the final scores of the video text in different dimensions, the video quality score is weighted and calculated; The expression of the video quality score is: ; Where, Indicates the The video quality score of each video text corresponding to the video; , Indicates the total number of dimensions; Indicates the The weights of the dimensions, and ; Indicates the Dimensions, Final rating of the video text; All videos whose video quality scores are not lower than the preset quality scores are selected to construct an initial high-quality video candidate set.
5. The method for video content quality analysis and knowledge recommendation based on a large model according to claim 4 is characterized in that: The method of calculating the similarity score between the user input question and the initial high-quality video candidate set based on the initial high-quality video candidate set and the user input question, and constructing the high-quality video candidate set, includes: Preprocessing is performed based on the user input question and the video text in the initial high-quality video candidate set; the preprocessing includes word segmentation, stop word removal, and medical terminology standardization; Based on the user input question and the video text, the weights of each word in the user input question and the video text after word segmentation processing are calculated to obtain the user input question weight and the video text weight, and the first cosine similarity score of the user input question weight and the video text weight is calculated; The expression of the first cosine similarity score is: ; Where, The first cosine similarity score representing the video text and the user input question; Represents the weight of the term in the user input question; Represents the weight of the term in the video text; To indicate a term; Based on the user input question and the video text in the initial high-quality video candidate set, a pre-trained semantic large model is used to encode and generate a high-dimensional semantic vector, thereby obtaining a high-dimensional user input question vector and video text vector, and calculating the second cosine similarity score between the user input question vector and the video text vector; The expression of the second cosine similarity score is: ; Where, The second cosine similarity score representing the video text and the user input question; Represents a high-dimensional user input question vector, i.e. , represents the encoding function; Represents the video text vector, that is ; represents the norm; Combining the first cosine similarity score and the second similarity score, weightedly calculating the similarity score between the user input question and the initial high-quality video candidate set; The similarity score between the user input question and the initial high-quality video candidate set is expressed as: ; Where, Represents the similarity score between the user input question and the initial high-quality video candidate set; represents the first cosine similarity score weight; represents the second cosine similarity score weight; All videos with similarity scores not lower than the preset similarity score are selected to construct a high-quality video candidate set.
6. The method for video content quality analysis and knowledge recommendation based on a large model according to claim 5, characterized in that: The method of calculating the weights of each word in the user input question and the video text after word segmentation processing based on the user input question and the video text to obtain the user input question weight and the video text weight, and calculating the first cosine similarity score of the user input question weight and the video text weight includes: Based on the user input question and video text, word segmentation is performed to obtain multiple terms; Calculate the frequency of each term in the user input question and video text respectively; Count the number of video texts containing the term and calculate the inverse document frequency of each term; Combine the word frequency and inverse document frequency corresponding to each term to calculate the weight of each term in the user input question and video text, and obtain the user input question weight and video text weight; Combine the user input question weight and the video text weight to calculate the first cosine similarity score of the user input question weight and the video text weight.
7. The method for video content quality analysis and knowledge recommendation based on a large model according to claim 6, characterized in that: The video quality score and similarity score are combined to calculate the weighted comprehensive recommendation score of each video in the high-quality video candidate set, and the top three recommended videos and the reasons for the recommendation are output to achieve precise recommendation, including: Combining the video quality score and similarity score, a weighted comprehensive recommendation score is calculated for each video in the high-quality video candidate set. The expression of the comprehensive recommendation score of each video is: ; Where, Indicates the The comprehensive recommendation score of the video corresponding to the video text; Represents the similarity score weight; Indicates the video quality score weight; Arrange the videos in descending order according to their comprehensive recommendation scores, output the three videos with the highest comprehensive recommendation scores and the reasons for the recommendations, and achieve precise recommendations.
8. The method for video content quality analysis and knowledge recommendation based on a large model according to claim 1, characterized in that: It also includes customizing the weights of each dimension according to the user's personal preferences.
9. A system for the large model-based video content quality analysis and knowledge recommendation method according to claim 1, characterized in that: include: The video data acquisition module is used to acquire videos based on the short video platform by inputting keywords, and obtain the video ID, video content text and video metadata; And based on the video content text, perform multi-dimensional text information extraction and problem diagnosis to generate structured video text; A high-quality video candidate set construction module is used to perform multi-dimensional scoring on structured video text, calculate the video quality score, and construct an initial high-quality video candidate set; and based on the initial high-quality video candidate set and the user input question, calculating the similarity score between the user input question and the initial high-quality video candidate set, and constructing the high-quality video candidate set; The recommended video output module is used to combine the video quality score and similarity score, weightedly calculate the comprehensive recommendation score of each video in the high-quality video candidate set, and output the top three recommended videos and the reasons for the recommendation to achieve precise recommendation.
Citation Information
Cited By
Short video recommendation method, apparatus and device, and computer readable storage medium
CN121051266A
Short video recommendation method, device and computer readable storage medium
CN121051266B
Hypertension group knowledge recommendation method and system based on large language model multi-agent
CN121075697A
High blood pressure population knowledge recommendation method and system based on large language model multi-agent
CN121075697B
Shot quality automatic scoring and optimal selection method for short episode materials
CN121644912A