Methods and Evaluation Systems for Generating Customized Quality Scales for Health Science Popularization Short Videos

By generating a customized scale for health science popularization short videos through principal component analysis and maximum variance rotation algorithm, the problem that existing tools cannot effectively assess multimodal features is solved, achieving efficient and accurate short video quality assessment and supporting real-time review by the platform.

CN122489797APending Publication Date: 2026-07-31SHANXI BETHUNE HOSPITAL (SHANXI ACAD OF MEDICAL SCI SHANXI HOSPITAL OF TONGJI HOSPITAL AFFILIATED TO TONGJI MEDICAL COLLEGE OF HUAZHONG UNIV OF SCI & TECH SHANXI MEDICAL UNIV THIRD HOSPITAL SHANXI MEDICAL UNIV THIRD CLINICAL COLLEGE OF MEDICINE)
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANXI BETHUNE HOSPITAL (SHANXI ACAD OF MEDICAL SCI SHANXI HOSPITAL OF TONGJI HOSPITAL AFFILIATED TO TONGJI MEDICAL COLLEGE OF HUAZHONG UNIV OF SCI & TECH SHANXI MEDICAL UNIV THIRD HOSPITAL SHANXI MEDICAL UNIV THIRD CLINICAL COLLEGE OF MEDICINE)
Filing Date
2026-05-07
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing health science popularization short video quality assessment tools lack a systematic dimensional design for multimodal features, cannot effectively identify the unique safety risks and content vividness differences of short videos, and the assessment process relies on manual operation, making it difficult to achieve standardization and automation.

Method used

By setting clear statistical thresholds and automatic elimination rules, and using principal component analysis and maximum variance rotation algorithms, a customized evaluation scale adapted to the characteristics of short video media is generated, including core dimensions such as safety, professionalism, practicality, emotionality, quality, image, and interactivity. A structural equation model is then constructed for verification.

Benefits of technology

The generated scale can objectively and reproducibly extract multimodal features, the assessment results are close to the real user perception experience, support real-time quality screening, and improve the efficiency and accuracy of the assessment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122489797A_ABST
    Figure CN122489797A_ABST
Patent Text Reader

Abstract

This invention relates to the field of short video quality assessment, specifically to a method and system for generating a customized quality scale for health science popularization short videos. The method includes: selecting initial indicators based on expert scoring thresholds to generate a first-level database; calculating consistency coefficients and content validity indices based on small sample test data, and removing indicators below the threshold to generate a second-level database; using large-scale user rating data as input, performing data suitability tests, and then invoking principal component analysis and maximum variance rotation algorithms, with the processor automatically executing dynamic elimination rules to generate a third-level database; constructing a structural equation model and performing confirmatory factor analysis, outputting a customized scale architecture when the root mean square error of approximation is <0.05 and the comparison fit index is >0.95. This invention achieves automated and standardized scale development, fills the gap in multimodal assessment of short videos, and can be directly embedded into intelligent review systems.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of short video quality assessment, and in particular to a method and assessment system for generating a customized quality scale for health science popularization short videos. Background Technology

[0002] Short health education videos have become an important medium for the public to acquire medical knowledge, garnering massive daily views across the internet. However, these videos, typically 15 to 60 seconds long, are characterized by fragmented narratives, high-density audiovisual cues, and algorithm-driven dissemination, resulting in inconsistent content quality. Currently, internationally used health information quality assessment tools include the DISCERN scale, JAMA Benchmark, and HONcode. These tools are primarily designed for standardized plain text health information or traditional long-form video content, focusing on textual elements such as information accuracy, citation sources, and treatment choices. They lack consideration for the multimodal characteristics of mobile short videos, including visual presentation, emotional expression, creator image, and user interaction feedback. Some scholars have attempted to directly apply these traditional tools to short video assessment, but verification has revealed a serious mismatch between their items and the media attributes of short videos, failing to effectively identify the unique safety risks and content vividness differences inherent in short videos.

[0003] Existing assessment methods and scale development processes suffer from the following technical deficiencies: First, the assessment perspective is singular and highly subjective. The selection of existing indicators largely relies on the personal experience of medical experts or researchers, employing the Delphi method or expert panel method for qualitative judgment. This neglects the cognitive load and perceptual experience of ordinary users as the core audience, making it difficult for the assessment results to accurately reflect the influence of short videos in actual dissemination scenarios. Second, there is a lack of systematic dimensional design tailored to the characteristics of multimodal short videos, particularly a lack of objective quantitative indicators for short video-specific quality dimensions such as "image presentation," "emotional resonance," and "interactive appeal." Third, the scale development process heavily relies on manual operation. Researchers must manually use statistical software such as SPSS and Mplus to complete reliability and validity tests and factor analyses step by step. The steps are cumbersome, time-consuming, and require repeated human judgment (such as deciding which indicator to remove and which dimension to retain). This makes it impossible to form a standardized, repeatable, automated data processing workflow, and it is also difficult to embed into the real-time review and distribution systems of short video platforms as an underlying measurement tool. Summary of the Invention

[0004] The purpose of this invention is to provide a method and evaluation system for generating a customized quality scale for health science popularization short videos. By setting clear statistical thresholds and automatic elimination rules, it aims to achieve automated dimensionality reduction and verification of the quality dimensions of multimodal health science popularization short videos, thereby generating a customized evaluation scale that is adapted to the characteristics of short video media and can be embedded in an intelligent review system.

[0005] To achieve the above objectives, the present invention adopts the following technical solution: In a first aspect, this invention provides a method for generating a customized quality scale for short health science videos, comprising: S1: screening initial evaluation indicators for short health science videos based on expert scoring thresholds to generate a first-level indicator database; S2: calculating the consistency coefficient and content validity index of each indicator in the first-level indicator database based on small sample test data, removing indicators below a preset consistency threshold, and generating a second-level indicator database; S3: using user ratings of each indicator in the second-level indicator database as input, performing a data applicability test, and after passing the test, calling principal component analysis and maximum variance rotation algorithms, with the computer processor automatically executing the following dynamic elimination rules to extract multi-dimensional features in parallel: in, For the first Commonality of the indicators; For the first The first indicator in the Factor loadings in dimensionality For the first The first indicator in the Factor loadings in dimensions; For the first The first indicator in the Factor loadings in dimensions; For the first The total correlation of the correction terms of each indicator is calculated. After iterative iteration, multiple core dimensions and related indicators with feature values ​​greater than 1 are extracted from the data to generate a third-level indicator database. S4: Based on the third-level indicator database, a structural equation model is constructed and confirmatory factor analysis is performed. When the global fit parameter meets the preset threshold, a customized scale architecture is output.

[0006] In step S3, the data suitability test includes calculating the Kaiser-Mayer-Holkin measure and performing the Bartlett test for sphericity.

[0007] In step S4, the global fit parameters include the root mean square of the approximation error and the comparison fit index. The preset threshold is that the root mean square of the approximation error is less than 0.05 and the comparison fit index is greater than 0.95.

[0008] Step S4 is followed by step S5: using multiple core dimensions as input variables and user feedback satisfaction and willingness to continue using as output variables, a multiple linear regression model is constructed to determine the predictive power of the customized scale architecture.

[0009] Secondly, this invention also provides a customized quality evaluation system for short health science videos, comprising: a data receiving and preprocessing module, configured to receive first quantitative scores from experts on initial evaluation indicators, filter the initial evaluation indicators according to preset expert scoring thresholds, and generate a first-level indicator database; receive small sample test data from target user mobile terminals, calculate the consistency coefficient and content validity index of each indicator in the first-level indicator database, and remove indicators with a consistency coefficient lower than 0.80 and a content validity index lower than a preset value, generating a second-level indicator database; and a feature dimensionality reduction and indicator removal engine, connected to the data receiving and preprocessing module, embedding a principal component analyzer and a logical rule judge, configured to... To receive the second-level indicator database and the second quantitative scores of each indicator in the second-level indicator database from a large number of users, a data suitability test is performed. After the test passes, the principal component analysis algorithm and the maximum variance rotation algorithm are called. The built-in processor automatically executes the dynamic elimination rules. After iterative iteration, multiple core dimensions and related indicators with eigenvalues ​​greater than 1 are extracted to generate the third-level indicator database. The model validation and architecture output module is connected to the feature dimensionality reduction and indicator elimination engine. It is configured to receive the third-level indicator database, construct a structural equation model and perform confirmatory factor analysis. When the root mean square error of the global fit parameter is less than 0.05 and the comparison fit index is greater than 0.95, a customized scale architecture is output.

[0010] The data receiving and preprocessing module includes an expert review terminal communication unit and a user test terminal communication unit. The expert review terminal communication unit is configured to receive the first quantitative scores from experts based on relevance, clarity, and comprehensiveness of the initial evaluation indicators. The user test terminal communication unit is configured to distribute the preliminary scale to the target user's mobile terminal and collect small sample test data.

[0011] The logical rule judge in the feature dimensionality reduction and indicator elimination engine is configured to execute the following judgment conditions: automatically eliminate an indicator when the communality of an indicator is less than 0.5; automatically eliminate an indicator when the absolute value of the factor loading of an indicator is less than 0.5; automatically eliminate an indicator when the absolute value of the difference in factor loadings of an indicator on two or more dimensions is greater than 0.4; and automatically eliminate an indicator when the total correlation of the correction terms of an indicator is less than 0.4.

[0012] When executing dynamic elimination rules, the feature dimensionality reduction and indicator elimination engine iterates in the following order: first, it checks and eliminates indicators with a communality of less than 0.5; then it checks and eliminates indicators with an absolute value of factor loading of less than 0.5; then it checks and eliminates indicators with an absolute value of cross-dimensional factor loading difference of greater than 0.4; and finally it checks and eliminates indicators with a total correlation of correction terms of less than 0.4.

[0013] The model validation and architecture output module is also configured to construct a multiple linear regression model with multiple core dimensions as independent variables and user feedback satisfaction and willingness to continue using as dependent variables before outputting the customized scale architecture, to determine the predictive power of the customized scale architecture, and output the customized scale architecture only when the determination coefficient of the predictive power is greater than 0.50 and there is no multicollinearity.

[0014] The data applicability test in the feature reduction and indicator elimination engine includes: calculating the Kaiser-Mayer-Holkin measure value and comparing it with a preset threshold, and performing the Bartlett test for sphericity. The test is considered passed when the Kaiser-Mayer-Holkin measure value is greater than 0.70 and the significance level of the Bartlett test for sphericity is less than 0.05.

[0015] Compared with the prior art, this application has the following advantages: 1. This invention automatically extracts seven core dimensions—safety, professionalism, practicality, emotional appeal, quality, image, and interactivity—from massive user rating data through principal component analysis and maximum variance rotation algorithm. Among these, the dimensions of "image," "interactivity," and "emotional appeal" are quality assessment elements constructed for the first time specifically for the characteristics of short video media. They can effectively capture non-textual information such as visual presentation, creator's personal charm, user interactions (likes, comments, shares, etc.), and emotional resonance, which cannot be quantified by existing tools (such as DISCERN and JAMA Benchmark). Compared to the practice of directly transplanting text-based assessment tools in existing technologies, the scale generated by this invention is highly compatible with the multimodal information ecosystem of 15-60 second short videos, and the assessment results are closer to the real user's perception and experience.

[0016] 2. Existing table compilation methods (including the EFA / CFA process described in Comparative Document 1) heavily rely on researchers manually using statistical software such as SPSS and Mplus for step-by-step operations. Each elimination decision (such as which low-loading indicator to delete or which cross-loading item to retain) requires manual judgment, and different researchers may reach different conclusions. This invention, for the first time, encodes four statistical thresholds—communicality less than 0.5, absolute factor loading less than 0.5, cross-dimensional factor loading difference greater than 0.4, and total correlation of correction terms less than 0.4—into a "dynamic elimination rule" that is automatically executed by a computer processor, and incorporates iterative loop logic. This technique transforms the originally vague and subjective qualitative sociological screening process into an objective and reproducible algorithmic process, eliminating biases caused by differences in researchers' experience, and resulting in a final extracted dimensional structure and indicator set with higher internal consistency and structural stability.

[0017] 3. This invention not only performs global fit verification using structural equation modeling after dimensionality reduction (forcing the root mean square error of the approximation to be less than 0.05 and the comparison fit index to be greater than 0.95), but also further incorporates a multiple linear regression prediction module. Using seven core dimensions as independent variables and user satisfaction and willingness to continue using the product as dependent variables, it measures a determination coefficient greater than 0.50 and eliminates multicollinearity. Unlike existing technologies that only report the fit index, this invention's dual verification mechanism ensures that the generated scale not only perfectly fits the data using a theoretical model, but also possesses a high predictive ability for real-world user adoption behavior and dissemination effects.

[0018] 4. This invention encapsulates preprocessing, feature reduction, model validation, and prediction output into a continuous data processing workflow, and correspondingly designs a complete system architecture including a data receiving and preprocessing module, a feature reduction and indicator elimination engine, and a model validation and architecture output module. This system supports real-time communication with expert review terminals and user mobile terminals, and can automatically receive scoring data, perform threshold filtering and iterative iterations, and output customized scales. Compared to the traditional scale development process in existing technologies, which takes weeks to months, this invention compresses the entire process to be completed automatically within hours. Furthermore, the generated application programming interface can be seamlessly deployed on the backend review servers of short video platforms such as Douyin and Kuaishou, enabling batch and real-time quality screening of massive amounts of health science popularization short videos, demonstrating significant technical efficiency and commercial conversion value. Attached Figure Description

[0019] Figure 1 This is a flowchart illustrating a method for generating a customized quality scale for short health science videos, as provided in an embodiment of this application. Detailed Implementation

[0020] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.

[0021] In the description of the invention, it should be understood that the terms "upper," "lower," "left," "right," "front," "rear," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or relative positional relationship shown in the accompanying drawings. They are used only for the convenience of describing the invention and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the invention. Unless otherwise specified, the above-mentioned orientational descriptions can be flexibly set in practical applications, provided that the relative positional relationships shown in the accompanying drawings are satisfied.

[0022] The terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this invention, unless otherwise stated, "a plurality of" means two or more.

[0023] In the description of this invention, it should be noted that, unless otherwise explicitly specified and limited, the terms "installation," "connection," "linking," and "communication" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection. They can refer to a direct connection or an indirect connection through an intermediate medium, or a communication between the internal components of two elements. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.

[0024] In embodiments of the invention, the terms "comprising," "including," or any other variations thereof are intended to cover a non-exclusive inclusion, such that a process, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, article, or apparatus that includes that element.

[0025] In embodiments of the present invention, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or design described as "exemplary" or "for example" in embodiments of the present invention should not be construed as being more preferred or advantageous than other embodiments or designs. Rather, the use of the terms "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.

[0026] This application provides a customized quality evaluation system for short health science videos, including: The data receiving and preprocessing module is configured to receive the first quantitative scores of experts on the initial evaluation indicators, filter the initial evaluation indicators according to the preset expert scoring threshold, and generate a first-level indicator database; receive small sample test data from the target user's mobile terminal, calculate the consistency coefficient and content validity index of each indicator in the first-level indicator database, remove indicators with a consistency coefficient lower than 0.80 and a content validity index lower than the preset value, and generate a second-level indicator database.

[0027] The data receiving and preprocessing module includes an expert review terminal communication unit and a user test terminal communication unit. The expert review terminal communication unit is configured to receive the first quantitative scores from experts based on relevance, clarity, and comprehensiveness of the initial evaluation indicators. The user test terminal communication unit is configured to distribute the preliminary scale to the target user's mobile terminal and collect small sample test data.

[0028] This module comprises two core communication units: an expert review terminal communication unit and a user testing terminal communication unit. The expert review terminal communication unit connects to at least three computers or tablets used by experts via a wireless network (such as WiFi, 4G / 5G). It receives expert scores of 1-5 points for each initial evaluation indicator across the dimensions of "relevance," "clarity," and "comprehensiveness," and stores these scores in local memory. Internally, this unit sets an expert scoring threshold (a preset value is configurable, for example, 4.0). The processor automatically calculates the average score for each indicator across the three dimensions, retaining only indicators with an average score not lower than the threshold, thus generating the first-level indicator database. The user testing terminal communication unit connects to a WeChat mini-program or dedicated app on the target users' smartphones via an application interface. It pushes the preliminary scale (electronic questionnaire) generated based on the first-level indicator database to a small sample of users (e.g., 30-100 people) and collects their scoring data. Upon receiving the small sample test data, this unit calls a built-in reliability and validity calculation function to calculate the CITC value for each indicator, the Cronbach's α for the overall scale, and the IOC value for each indicator. Set consistency thresholds (e.g., Cronbach's α not lower than 0.80, CITC not lower than 0.40, IOC not lower than 0.78) to automatically remove unqualified indicators and generate a second-level indicator database.

[0029] The feature reduction and indicator removal engine, connected to the data receiving and preprocessing module, is configured to receive the second-level indicator database and the second quantitative scores of each indicator in the second-level indicator database from a large number of users. It performs data applicability testing, and after the test is passed, it calls the principal component analysis algorithm and the maximum variance rotation algorithm. The built-in processor automatically executes the dynamic removal rules, and after iterative iteration, it extracts multiple core dimensions and related indicators with feature values ​​greater than 1, generating the third-level indicator database.

[0030] The model validation and architecture output module, connected to the feature dimensionality reduction and index elimination engine, is configured to receive a third-level index database, construct a structural equation model and perform confirmatory factor analysis. When the root mean square error of the global fit parameter is less than 0.05 and the comparison fit index is greater than 0.95, it outputs a customized scale architecture.

[0031] The feature reduction and indicator removal engine is connected to the output of the data receiving and preprocessing module via a high-speed data bus. This engine reads the second-level indicator database and a large-scale user rating dataset (typically with a sample size of at least 300, each user giving a rating of 1-7 for each indicator). The engine first runs a data suitability test subroutine, calculating the KMO value and the p-value of Bartlett's test of sphericity. Subsequent steps are only executed if KMO > 0.70 and p < 0.05. Then, the engine calls the principal component analyzer to extract principal components with eigenvalues ​​greater than 1 based on the covariance matrix or correlation matrix. During extraction, the logical rule judge checks the following four rules round by round for each indicator. If any one of these rules is met, the indicator is marked for removal and deleted from the dataset of the current iteration: Commonality of Indicators The absolute value of factor loadings of the indicator on its principal dimension The difference between the maximum factor loadings of the indicator on other dimensions and the factor loadings of the main dimension. , Same sign. Total correlation of the indicator's correction term. .

[0032] After each rejection, the engine re-executes PCA and rotation, recalculates the corresponding statistics of the remaining indicators, and repeats the iteration until all retained indicators satisfy the four rules. After the iteration converges, the engine outputs the core dimensions with eigenvalues ​​greater than 1 and the indicators with absolute factor loadings greater than 0.5 (or a higher threshold such as 0.55) under each dimension, as the third-level indicator database.

[0033] When executing dynamic elimination rules, the feature dimensionality reduction and indicator elimination engine iterates in the following order: first, it checks and eliminates indicators with a communality of less than 0.5; then it checks and eliminates indicators with an absolute value of factor loading of less than 0.5; then it checks and eliminates indicators with an absolute value of cross-dimensional factor loading difference of greater than 0.4; and finally it checks and eliminates indicators with a total correlation of correction terms of less than 0.4.

[0034] In some embodiments, the engine can also execute the iterations in a preset order: first, examine and remove indicators with a communality of less than 0.5; then, examine and remove indicators with an absolute factor loading of less than 0.5; next, examine and remove indicators with an absolute difference in cross-dimensional factor loadings greater than 0.4; and finally, examine and remove indicators with a total correlation of correction terms of less than 0.4. This sequential iterative approach can reduce unnecessary calculations and accelerate convergence.

[0035] The model validation and architecture output module is also configured to construct a multiple linear regression model with multiple core dimensions as independent variables and user feedback satisfaction and willingness to continue using as dependent variables before outputting the customized scale architecture, to determine the predictive power of the customized scale architecture, and output the customized scale architecture only when the determination coefficient of the predictive power is greater than 0.50 and there is no multicollinearity.

[0036] The model validation and architecture output module connects to the output port of the feature dimensionality reduction engine, receiving a third-level indicator database (containing dimension-indicator correspondences and user rating data). Internally, the module integrates a SEM algorithm library (such as a secondary development library based on lavaan or OpenMX) to automatically construct a measurement model: creating a corresponding observation variable path for each latent variable (dimension). The module performs maximum likelihood estimation or generalized least squares estimation, calculating the parameter estimates and standard errors for each path. It extracts global fit indices: root mean square approximation error (RMSEA), comparison fit index (CFI), and standardized residual root mean square (SRMR), etc. Based on preset conditions, if RMSEA < 0.05 and CFI > 0.95, the model is considered to have an excellent fit, and the module will output a complete customized scale, including the definition of each dimension, the description of each indicator, the scoring method (e.g., a 7-point Likert scale), and a reliability and validity statistical report.

[0037] The data applicability test in the feature reduction and indicator elimination engine includes: calculating the Kaiser-Mayer-Holkin measure value and comparing it with a preset threshold, and performing the Bartlett test for sphericity. The test is considered passed when the Kaiser-Mayer-Holkin measure value is greater than 0.70 and the significance level of the Bartlett test for sphericity is less than 0.05.

[0038] This application provides a method for generating a customized quality scale for short health science videos, including: S1: Based on expert scoring thresholds, the initial evaluation indicators for health science popularization short videos are screened to generate a first-level indicator database.

[0039] Step S1 involves collecting multi-source evaluation data from short health science videos and using expert scoring thresholds to quantitatively screen initial indicators, forming a structured first-level indicator database that provides a data foundation for subsequent automated processing.

[0040] Specifically, the system first extracts preliminary evaluation indicators related to the quality of health science popularization short videos from literature databases, existing evaluation tools (such as the DISCERN scale and JAMABenchmark scale), and user comments on short video platforms. These indicators cover multiple aspects, including content accuracy, information sources, expression methods, visual presentation, and interactive features. Then, the system sends indicator evaluation requests to connected expert review terminals, receiving quantitative scores from at least three experts based on three dimensions: "relevance" (whether the indicator is closely related to the short video quality), "clarity" (whether the indicator description is unambiguous), and "comprehensiveness" (whether the indicator system covers the main aspects of short video quality). For example, a five-point Likert scale is used, where 1 point indicates strong disagreement and 5 points indicate strong agreement. The system processor automatically calculates the average score of each indicator across the three components and determines whether it exceeds a preset expert scoring threshold. For example, a threshold of 4.0 is set, meaning indicators with an average score of at least 4.0 are retained. The remaining indicators after this threshold filtering constitute the first-level indicator database.

[0041] For example, suppose 50 evaluation indicators are initially collected from existing literature, including "medical information in the video has clear sources," "the video is clear and shaky," "the doctor's white coat is properly presented," "the video ends with encouragement for users to like it," and "the diagnostic suggestions in the video include risk warnings." The system sends evaluation forms to five health education experts from top-tier hospitals and three communication experts, with each expert scoring each indicator between 1 and 5. The system calculates that the average expert score for the indicator "diagnostic suggestions in the video include risk warnings" is 4.8, exceeding the threshold of 4.0, and is therefore retained; while the average expert score for the indicator "video length meets 15 seconds" is 2.5, below 4.0, and is automatically removed. Ultimately, 30 indicators are retained to form the first-level indicator database.

[0042] S2: Calculate the consistency coefficient and content validity index of each indicator in the first-level indicator database based on small sample test data, remove indicators that are below the preset consistency threshold, and generate the second-level indicator database.

[0043] The system transforms the first-level indicator database into a preliminary scale (e.g., each indicator corresponds to a statement, and users rate their feelings after watching health education short videos, typically using a 1-5 or 1-7 Likert scale). This preliminary scale is then distributed to 30 to 100 representative target users' mobile terminals via user testing terminal communication units. After watching the designated health education short videos, users rate the statements corresponding to each indicator. After collecting all small sample test data, the system calls the built-in reliability and validity calculation module to calculate the total correlation coefficient (CITC) of corrected items for each indicator and the Cronbach's α coefficient for the overall scale. Simultaneously, it calculates the content validity index (IOC) for each indicator. This index is calculated by experts rating the consistency between the indicator and the target construct, typically using an expert scoring method. The IOC value for each indicator is equal to the number of experts who believe "this indicator effectively measures the target dimension" divided by the total number of experts. The system presets consistency thresholds, such as removing indicators whose Cronbach's α is below 0.80, or removing an indicator whose CITC is less than 0.40 and whose α coefficient increases significantly after deleting the indicator. The remaining indicators after step S2 constitute the second-level indicator database.

[0044] For example, the system distributed a preliminary questionnaire containing 30 indicators to 50 ordinary users (aged 20-60, covering different educational backgrounds) who frequently watch health short videos. After watching a compilation of three health science short videos (each approximately 45 seconds long, themed "Daily Dietary Precautions for Hypertension"), users rated each indicator on a scale of 1-5. After collecting 50 valid questionnaires, the system calculated the overall scale's Cronbach's α value to be 0.85, which is within the good range. However, it was found that the CITC value for the indicator "Medical information in the video has a clear source" was 0.32 (below 0.40), and removing this indicator would increase the overall α value to 0.87; the CITC value for the indicator "The video ends by encouraging users to like it" was 0.28, which was also poor. The system automatically removed these two indicators. In addition, the system calculates the IOC value for each indicator. For the indicator "professional image of doctors in videos," four out of five experts believed that this indicator could effectively measure professionalism, with an IOC of 0.8, which is greater than the preset value of 0.78, and it is retained. However, for the indicator "whether the background music used in the video is pleasant to listen to," only one expert believed that it could measure quality, with an IOC of 0.2, which is lower than the threshold, and it is automatically removed. The remaining 25 indicators constitute the second-level indicator database.

[0045] S3: Using user ratings of each indicator in the second-level indicator database as input, perform a data suitability test. If the test passes, invoke principal component analysis and maximum variance rotation algorithms. The computer processor will then automatically execute the following dynamic elimination rules to extract multi-dimensional features in parallel: in, For the first Commonality of the indicators; For the first The first indicator in the Factor loadings in dimensionality For the first The first indicator in the Factor loadings in dimensions; For the first The first indicator in the Factor loadings in dimensions; For the first The total correlation of the correction items of each indicator is calculated; after iterative iteration, multiple core dimensions and related indicators with feature values ​​greater than 1 are extracted from the data to generate a third-level indicator database.

[0046] Step S3 uses large-scale real user rating data to perform data applicability tests, then calls principal component analysis (PCA) and maximum variance rotation algorithm, and the computer processor automatically executes a set of hard-coded "dynamic elimination rules". After iterative iteration, the core dimensions with feature values ​​greater than 1 and the corresponding strongly correlated indicators are extracted from the data to generate a third-level indicator database.

[0047] First, the system distributes the formal scale based on the second-level indicator database to a larger group of ordinary users (typically with a sample size of over 300 people) through the user terminal communication unit, collecting large-scale rating data (e.g., users rating each indicator from 1 to 7). The data receiving and preprocessing module then inputs this dataset into the feature dimensionality reduction and indicator removal engine.

[0048] As one possible implementation, in step S3, the data suitability test includes calculating the Kaiser-Mayer-Holkin measure and performing the Bartlett test for sphericity.

[0049] The engine first performs a data suitability check, which includes two steps: (1) Calculate the Kaiser-Mayer-Orgin (KMO) measure. This statistic is used to determine whether factor analysis is suitable by comparing the magnitude of the simple correlation coefficient and the partial correlation coefficient between variables. Generally, KMO > 0.70 is considered suitable.

[0050] (2) Perform Bartlett's test of sphericity to test whether the variables are independent. The null hypothesis is that the correlation matrix is ​​an identity matrix. If the significance level p < 0.05, the null hypothesis is rejected, indicating that there is a significant correlation between the variables, which is suitable for factor analysis.

[0051] After the validation passes, the engine invokes Principal Component Analysis (PCA) to extract principal components with initial eigenvalues ​​greater than 1 as candidate dimensions. Simultaneously, a Varimax rotation algorithm is employed to concentrate high-loading indicators and disperse low-loading indicators on each factor, thereby enhancing the interpretability of the factors. During the PCA iteration process, the computer processor automatically determines the retention of each indicator according to pre-stored dynamic elimination rules in memory. For the first The communality of an indicator is the proportion of its variance that can be explained by all the extracted common factors in factor analysis. It ranges from 0 to 1. The lower the communality, the weaker the correlation between the indicator and all common factors, and the smaller its contribution to the scale. For the first The first indicator in the Factor loadings on a dimension are the correlation coefficients between indicators and factors. Positive values ​​indicate positive influences, negative values ​​indicate negative influences, and the larger the absolute value, the closer the relationship between the indicator and the factor. For the first The first indicator in the Factor loadings in dimensions; For the first The first indicator in the Factor loadings in dimensions; For the first The total correlation of the correction terms for each indicator is calculated by performing a Pearson correlation analysis between the score of that indicator and the total score of the remaining indicators. This indicates that the indicator has poor coordination with the overall scale.

[0052] The processor automatically evaluates each of the above rules. For any indicator that meets any elimination condition, the system automatically removes it from the current indicator pool, then re-runs the PCA and rotation algorithms to recalculate the communality, factor loadings, and CITC of the remaining indicators. This process is repeated until all remaining indicators no longer meet any elimination conditions, or the preset maximum number of iterations is reached. After each iteration, the system records factors with eigenvalues ​​greater than 1 (core dimensions) and outputs indicators whose absolute factor loadings under each factor are greater than a certain threshold (e.g., 0.5) as associated indicators for that dimension. Finally, the system generates a third-level indicator database containing multiple core dimensions and associated indicators.

[0053] For example, the system collected ratings from 800 ordinary movie platform users on 25 indicators (from a secondary database) through an online survey platform. First, a KMO test was performed, yielding a KMO value of 0.87, greater than 0.70, and a Bartlett's test of sphericity with a p-value less than 0.001, indicating the data was suitable for factor analysis. The system then performed its first PCA run, extracting 7 factors with eigenvalues ​​greater than 1, explaining 68.5% of the cumulative variance, preliminarily determining that 7 dimensions could be extracted.

[0054] In the first iteration, the system calculated the communality of each indicator. The indicator "Hospital level of doctors in the video" had a communality of 0.42, less than 0.5, and was automatically removed. Simultaneously, the factor loading matrix was calculated. The indicator "Background music style in the video" had a loading of 0.63 on the first factor and 0.51 on the second factor. The difference was calculated as 0.63 - 0.51 = 0.12, which is less than 0.4, so it was not removed. However, the indicator "Branded health products appeared in the video" had a loading of 0.45 on the third factor (safety) (less than 0.5) and 0.52 on the fourth factor (interactivity). The maximum value was calculated as max(0.45, 0.52) - 0.45 = 0.07, which did not meet the removal criteria. However, since its absolute loading value was less than 0.5, it was removed according to the second rule. Additionally, the indicator "Number of likes in the user comment section" had a CITC value of 0.38, less than 0.4, and was automatically removed. After the first round of removal, 22 indicators remained.

[0055] The system enters its second iteration, rerunning PCA. All 22 indicators meet their respective thresholds (communality ≥ 0.5, absolute factor loadings ≥ 0.5, cross-factor loading differences ≤ 0.4, CITC ≥ 0.4), and the number of dimensions with eigenvalues ​​greater than 1 remains at 7. The system stops iterating, outputting 7 core dimensions and their corresponding 22 strongly correlated indicators. These 7 dimensions are named according to their content: Safety, Professionalism, Usability, Emotionality, Quality, Image, and Interactivity. For example, the "Safety" dimension includes indicators such as "No unproven treatments recommended in the video," "Risk warnings provided in the video," and "No absolute medical promises made in the video"; the "Image" dimension includes indicators such as "Doctors dressed professionally and neatly," "Aesthetically pleasing video composition," and "Natural and believable presenter expressions." At this point, the third-level indicator database is generated.

[0056] S4: Construct a structural equation model based on the third-level indicator database and perform confirmatory factor analysis. When the global fit parameter meets the preset threshold, output a customized scale architecture.

[0057] As one possible implementation, in step S4, the global fit parameters include the root mean square of the approximation error and the comparison fit index, with a preset threshold of the root mean square of the approximation error being less than 0.05 and the comparison fit index being greater than 0.95.

[0058] The system model validation and architecture output module receives a third-level indicator database from the feature reduction engine, containing identified latent variables (core dimensions, such as security and professionalism) and observed variables (corresponding to 22 indicators). The module calls the SEM library to construct a structural equation model, specifically including: setting the measurement path for each latent variable and defining the correlations between latent variables. Then, confirmatory factor analysis (CFA) is performed to estimate the standardized factor loadings, error variance, and other parameters for each path. After CFA completion, the module extracts key global fit parameters, mainly including the root mean square error of approximation (RMSEA) and the comparison fit index (CFI). RMSEA reflects the gap between the hypothetical model and the perfectly fitted saturated model; a smaller value is better. Generally, RMSEA < 0.05 indicates a "good fit," 0.05–0.08 indicates a "reasonable fit," and greater than 0.10 indicates a "poor fit." The Comparative Fidelity Index (CFI) assesses model quality by comparing the fit between the hypothetical model and the independent model (assuming all variables are uncorrelated). The CFI ranges from 0 to 1, with a CFI > 0.90 indicating acceptable performance and a CFI > 0.95 indicating excellent performance. In this embodiment, the preset thresholds are RMSEA < 0.05 and CFI > 0.95. Only when both conditions are met will the system consider the constructed seven-dimensional scale architecture to have excellent construct validity and output this architecture as the final customized scale. If the fit parameters do not meet the thresholds, the system will automatically output a diagnostic report, indicating the possible existence of model correction indices for user reference and adjustment (e.g., allowing correlation in certain indicator error terms).

[0059] For example, based on 800 user data points and the aforementioned 7-dimensional, 22-indicator structure, the system constructs a SEM roadmap: Security (latent variable) – its corresponding 4 indicators (observed variables), Professionalism – its corresponding 3 indicators, Usability – its corresponding 3 indicators, Emotionality – its corresponding 3 indicators, Quality – its corresponding 3 indicators, Image – its corresponding 3 indicators, and Interactivity – its corresponding 3 indicators. After running maximum likelihood estimation, the CFA results show: RMSEA = 0.042 (less than 0.05), CFI = 0.972 (greater than 0.95), and SRMR (standardized root mean square residuals) is 0.035, verifying the excellent model fit. The system outputs this customized scale architecture, which includes 7 assessment dimensions and 22 specific assessment items, each item accompanied by a standard 1-7 point rating scale and rating instructions. For example, the first item under the "Safety" dimension is: "Does the health education information in this video provide sufficient risk warnings and safety boundaries (such as not recommending stopping medication on one's own, avoiding misdiagnosis, etc.)?"

[0060] For example, after step S4, step S5 is also included: using multiple core dimensions as input independent variables and user feedback satisfaction and willingness to continue using as output dependent variables, a multiple linear regression model is constructed to determine the predictive power of the customized scale architecture.

[0061] The system validity prediction module uses multiple core dimensions (such as safety, professionalism, etc., seven dimensions) identified in the third-level indicator database as independent variables (X1, X2, ..., X7), and the "overall satisfaction" and "willingness to continue using" filled in by users after watching health science popularization short videos as dependent variables (Y), constructing a multiple linear regression model: Y = β0 + β1X1 + β2X2 + ... + β7X7 + ε. The system uses the least squares method to estimate each regression coefficient β and calculates the coefficient of determination (R²). 2 This value represents the proportion of the dependent variable variance that the model can explain. Simultaneously, the system calculates the variance inflation factor (VIF) among the independent variables to diagnose multicollinearity; a VIF > 5 is generally considered to indicate severe multicollinearity. Only when R... 2 The system outputs a validation signal only when the VIF value is >0.50 (indicating that the model has moderate or higher explanatory power) and all VIF values ​​are less than 5 (indicating no serious multicollinearity). This indicates that the customized scale has reliable predictive power for user satisfaction and willingness to continue using the product, and can serve as a basis for content recommendation and risk control on short video platforms.

[0062] For example, in a regression analysis of 800 users, the average scores of seven dimensions—security, professionalism, practicality, emotional appeal, quality, image, and interactivity—were used as independent variables, and user satisfaction (1-7 points) as the dependent variable. The regression results showed that the adjusted R²... 2 = 0.63, indicating that the seven dimensions collectively explained 63% of the variance in user satisfaction. The VIF values ​​for each dimension were 1.2, 1.5, 1.8, 1.4, 1.6, 1.3, and 1.7, all less than 5, indicating no multicollinearity. Among them, the image dimension (β=0.28, p<0.01) and the safety dimension (β=0.25, p<0.01) contributed the most to satisfaction. Meanwhile, the regression with continued use intention as the dependent variable also showed R... 2 = 0.57, and VIF values ​​are all less than 5. Therefore, the system determines that the scale has good predictive power and outputs it as the final deployable version.

[0063] In the description of this specification, specific features, structures, materials, or characteristics may be combined in any suitable manner in one or more embodiments or examples.

[0064] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A method for generating a health popular science short video quality customization scale, characterized in that, include: S1: Based on expert scoring thresholds, the initial evaluation indicators for health science popularization short videos are screened to generate a first-level indicator database; S2: Based on small sample test data, calculate the consistency coefficient and content validity index of each indicator in the first-level indicator database, remove indicators below the preset consistency threshold, and generate the second-level indicator database; S3: Using user ratings of each indicator in the second-level indicator database as input, perform a data suitability test. After the test passes, call the principal component analysis algorithm and the maximum variance rotation algorithm, and the computer processor automatically executes the following dynamic elimination rules to extract multi-dimensional features in parallel: in, For the first Commonality of the indicators; For the first The first indicator in the Factor loadings in dimensionality For the first The first indicator in the Factor loadings in dimensions; For the first The first indicator in the Factor loadings in dimensions; For the first The total correlation of the correction terms of each indicator is obtained. After iterative iteration, multiple core dimensions and related indicators with feature values ​​greater than 1 are extracted from the data to generate a third-level indicator database; S4: Based on the third-level indicator database, a structural equation model is constructed and confirmatory factor analysis is performed. When the global fit parameter meets the preset threshold, a customized scale architecture is output.

2. The method for generating a customized quality scale for health science popularization short videos according to claim 1, characterized in that, In step S3, the data suitability test includes calculating the Kaiser-Mayer-Holkin measure and performing the Bartlett test for sphericity.

3. The method for generating a customized quality scale for health science popularization short videos according to claim 1, characterized in that, In step S4, the global fit parameter includes the root mean square of the approximation error and the comparison fit index. The preset threshold is that the root mean square of the approximation error is less than 0.05 and the comparison fit index is greater than 0.

95.

4. The method for generating a customized quality scale for health science popularization short videos according to claim 1, characterized in that, Following step S4, step S5 is also included: using the multiple core dimensions as input independent variables and user feedback satisfaction and willingness to continue using as output dependent variables, a multiple linear regression model is constructed to determine the predictive power of the customized scale architecture.

5. A customized quality evaluation system for short health science videos, characterized in that, include: The data receiving and preprocessing module is configured to receive the first quantitative scores from experts on the initial evaluation indicators, filter the initial evaluation indicators according to a preset expert scoring threshold, and generate a first-level indicator database; receive small sample test data from the target user's mobile terminal, calculate the consistency coefficient and content validity index of each indicator in the first-level indicator database, and remove indicators with a consistency coefficient lower than 0.80 and a content validity index lower than a preset value, generating a second-level indicator database; the feature dimensionality reduction and indicator removal engine, connected to the data receiving and preprocessing module, embeds a principal component analyzer and a logical rule judge, and is configured to receive the second-level indicator database and large-scale... Users assign second quantitative scores to each indicator in the second-level indicator database. A data suitability test is then performed. If the test passes, principal component analysis and maximum variance rotation algorithms are invoked. The built-in processor automatically executes dynamic elimination rules, extracting multiple core dimensions and related indicators with eigenvalues ​​greater than 1 through iterative iteration, generating the third-level indicator database. The model validation and architecture output module, connected to the feature reduction and indicator elimination engine, is configured to receive the third-level indicator database, construct a structural equation model, and perform confirmatory factor analysis. When the root mean square error of the global fit parameter is less than 0.05 and the comparison fit index is greater than 0.95, a customized scale architecture is output.

6. The customized quality evaluation system for health science popularization short videos according to claim 5, characterized in that, The data receiving and preprocessing module includes an expert review terminal communication unit and a user test terminal communication unit. The expert review terminal communication unit is configured to receive the first quantitative scores of the initial evaluation indicators based on relevance, clarity and comprehensiveness by experts. The user test terminal communication unit is configured to distribute the preliminary scale to the target user's mobile terminal and collect small sample test data.

7. The customized quality evaluation system for health science popularization short videos according to claim 5, characterized in that, The logical rule judge in the feature dimensionality reduction and indicator elimination engine is configured to execute the following judgment conditions: automatically eliminate an indicator when the communality of an indicator is less than 0.5; automatically eliminate an indicator when the absolute value of the factor loading of an indicator is less than 0.5; automatically eliminate an indicator when the absolute value of the difference in factor loadings of an indicator on two or more dimensions is greater than 0.4; and automatically eliminate an indicator when the total correlation of the correction terms of an indicator is less than 0.

4.

8. The customized quality evaluation system for health science popularization short videos according to claim 5, characterized in that, When executing the dynamic elimination rules, the feature dimensionality reduction and indicator elimination engine iterates in the following order: first, it checks and eliminates indicators with a communality of less than 0.5; then it checks and eliminates indicators with an absolute value of factor loading of less than 0.5; then it checks and eliminates indicators with an absolute value of cross-dimensional factor loading difference of greater than 0.4; and finally it checks and eliminates indicators with a total correlation of correction terms of less than 0.

4.

9. A customized quality evaluation system for health science popularization short videos according to claim 5, characterized in that, The model validation and architecture output module is further configured to construct a multiple linear regression model with the multiple core dimensions as independent variables and user feedback satisfaction and willingness to continue using as dependent variables before outputting the customized scale architecture, to determine the predictive power of the customized scale architecture, and to output the customized scale architecture only when the determination coefficient of the predictive power is greater than 0.50 and there is no multicollinearity.

10. A customized quality evaluation system for health science popularization short videos according to claim 5, characterized in that, The data applicability test in the feature reduction and index elimination engine includes: calculating the Kaiser-Mayer-Holkin measure value and comparing it with a preset threshold, and performing a Bartlett test of sphericity. The test is considered passed when the Kaiser-Mayer-Holkin measure value is greater than 0.70 and the significance level of the Bartlett test of sphericity is less than 0.05.