An old-age health management AI video generation system based on multi-modal interaction

The AI ​​video generation system for elderly health management, based on a front-end and back-end separation architecture, combines a multimodal interaction module and algorithms specifically designed for the elderly. This system addresses the issues of misunderstandings about the needs of the elderly and inappropriate content in the generation of health management videos, ensuring that the video content is professional, compliant, and age-friendly, thereby improving the targeting and effectiveness of health information dissemination for the elderly.

CN122340334APending Publication Date: 2026-07-03MACAO POLYTECHNIC INST
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
MACAO POLYTECHNIC INST
Filing Date
2026-04-17
Publication Date
2026-07-03

AI Technical Summary

Technical Problem

Existing video generation technologies related to health management do not fully take into account the physiological and cognitive characteristics of the elderly, resulting in misunderstandings of needs, inappropriate content generation, poor viewing experience, and a lack of data feedback and optimization mechanisms, making it difficult to meet the actual needs of elderly health management.

Method used

The AI ​​video generation system for elderly health management, which adopts a front-end and back-end separation architecture, achieves demand normalization, script compliance, material adaptation, and closed-loop management through a multimodal interaction module. Combined with elderly-specific algorithms, it performs intelligent processing throughout the entire process, including accurate demand matching, video parameter adaptation, and content optimization.

Benefits of technology

This has improved the targeting and effectiveness of health information dissemination for the elderly. The video content is professional and compliant, and the presentation format is adapted to the needs of the elderly, thus enhancing the viewing experience and the system's continuous optimization capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122340334A_ABST
    Figure CN122340334A_ABST
Patent Text Reader

Abstract

This invention discloses an AI video generation system for elderly health management based on multimodal interaction, relating to the fields of artificial intelligence and health management technology. The system includes: a demand normalization module that receives multimodal demands from elderly users, generates standardized instructions, and transmits them to a script compliance module; the script compliance module generates and optimizes scripts, which are then transmitted to a material adaptation module; the material adaptation module generates three types of materials, which are transmitted to an audio-visual synthesis module; and the synthesis module processes the materials and transmits them to a closed-loop control module, which collects data feedback to optimize the algorithm and also provides save and share functions. This invention, by building a multimodal interactive AI video generation system for elderly health management, accurately captures demands and complies with script regulations, generates scripts and adapts materials based on a large model, constructs an algorithm closed-loop optimization mechanism, adopts a front-end and back-end separation architecture to ensure efficiency and standardization, provides save and share functions, and is age-friendly throughout the entire process, improving the effectiveness of information dissemination.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of artificial intelligence and health management technology, specifically to an AI video generation system for elderly health management based on multimodal interaction. Background Technology

[0002] As the population ages, the health management needs of the elderly are becoming increasingly diversified and personalized, making it a key focus in public services. With the development of artificial intelligence and multimodal interaction technologies, AI video generation technology is gradually being applied to the field of health information dissemination, providing a new digital and visual path for elderly health management. The elderly have unique cognitive and physiological characteristics in acquiring health knowledge, and have clear age-appropriate needs for the presentation and expression of health information. Traditional forms of health information dissemination are no longer suitable for the actual user experience of elderly users. AI video generation technology based on multimodal interaction can integrate multiple input forms such as voice and text, and combine algorithms to achieve personalized content generation, becoming an important technical means to adapt to the dissemination of health management information for the elderly, and also promoting the upgrading of elderly health management services towards intelligence and age-appropriateness.

[0003] Existing health management video generation technologies are mostly designed for the general public, failing to adequately consider the physiological and cognitive characteristics of the elderly. They exhibit significant shortcomings in age-friendly design. In the needs assessment phase, these technologies struggle to accurately identify the multimodal health needs of elderly users, lacking professional age-friendly semantic matching capabilities and prone to misunderstandings. In the content generation phase, they fail to adequately control the professionalism of health knowledge and the safety of its expression, neglecting age-friendly expression requirements. This results in content with excessive technical jargon and unreasonable information density. Furthermore, video material parameter settings lack age-friendly adaptation, with insufficient control over the synchronization of visuals, audio, and subtitles, leading to a poor viewing experience. Most technologies also lack a closed-loop optimization mechanism for data feedback, making it impossible to continuously optimize the generated results based on user feedback, thus failing to truly meet the practical application needs of elderly health management. Summary of the Invention

[0004] The purpose of this invention is to overcome the shortcomings of existing technologies and provide an AI video generation system for elderly health management based on multimodal interaction. It adopts a front-end and back-end separation architecture, building five core modules: demand normalization, script compliance, material adaptation, audio-visual synthesis, and closed-loop management. This enables intelligent processing of the entire process, from inputting multimodal health needs from elderly users to outputting finished AI videos. The system utilizes multiple age-appropriate algorithms to achieve precise demand matching, script compliance optimization, video parameter adaptation, and audio-visual synchronization calibration. Simultaneously, it constructs a closed-loop optimization mechanism for data feedback, using anonymized user data for algorithm parameter iteration. The entire process is designed around the physiological and cognitive characteristics of the elderly, ensuring professional and compliant video content and a presentation format adapted to their needs. This provides an intelligent and personalized video generation solution for elderly health management, significantly improving the targeting and effectiveness of elderly health information dissemination.

[0005] To address the aforementioned technical problems, this invention provides the following technical solution: an AI video generation system for elderly health management based on multimodal interaction. This system adopts a front-end / back-end separation architecture. The front-end uses the React framework to construct the user interface and handle operation responses, while the back-end uses Python to process business logic and connect with AI services. The system includes:

[0006] Demand normalization module: Receives multimodal health demand input from elderly users, uses an age-friendly demand semantic matching algorithm to calculate the semantic similarity between user demands and the standard demand library, generates standardized health demand structured instructions, and passes them to the script compliance module.

[0007] Script Compliance Module: Receives the standardized health needs structured instructions, calls the large language model to generate the initial video script, uses the health script compliance adaptation algorithm to calculate the compliance score of the script in three dimensions: accuracy of health knowledge, safety of expression, and age-appropriate expression, iteratively optimizes the initial video script based on the compliance score until the compliance score reaches the preset threshold, generates a time-series storyboard script and passes it to the material adaptation module.

[0008] When calculating compliance scores across three dimensions, the script compliance module assesses the accuracy of health knowledge based on standardized content from a knowledge graph specific to elderly health; the safety of expression based on the requirements for medical information dissemination; and the age-appropriate expression based on sentence length, vocabulary fluency, and information density. The module optimizes the script based on the overall compliance score. If the overall compliance score is below a preset threshold, the script content for the corresponding low-scoring dimension is modified and adjusted. If the overall compliance score reaches the preset threshold, the initial video script is divided into 6-8 sequential storyboards based on a 3-5 minute video length. Each sequential storyboard is labeled with three types of information: scene description, narration, and subtitle text.

[0009] Material adaptation module: Receives the time-series storyboard script, generates three types of video materials: storyboard image sequence, voice narration, and synchronized subtitles, uses a multimodal aging-adaptation parameter adaptation algorithm to calculate the aging-adaptation parameters for video frame rate, subtitle font size, and voice playback speed, and transmits the three types of video materials to the audio-visual synthesis module;

[0010] When using a multimodal age-appropriate parameter adaptation algorithm to calculate age-appropriate parameters for video frame rate, subtitle font size, and voice playback speed, three types of indicators are first extracted from the time-series storyboard: content rhythm, information density, and screen description complexity. Combined with the visual and auditory physiological characteristics of the elderly, fixed parameter value ranges are set for video frame rate, subtitle font size, and voice playback speed. The algorithm calculates the appropriate value for video frame rate based on screen description complexity, calculates the appropriate value for subtitle font size based on the screen display area and the number of subtitle characters, and calculates the appropriate value for voice playback speed based on the difficulty and sentence length of the narration content. All calculated age-appropriate parameters fall within the corresponding preset parameter value range.

[0011] Audio and video synthesis module: Receives three types of video materials, uses a multimodal aging-adaptive synchronization algorithm to calculate the time axis alignment deviation and synchronization quality score of video frame, audio and subtitle, performs video rendering and storage based on the calculation results, and transmits the finished video and the time axis alignment deviation and synchronization quality score to the closed-loop control module;

[0012] When using a multimodal aging-adaptive synchronization algorithm to calculate the timeline alignment deviation and synchronization quality score of video images, audio, and subtitles, the video is first divided into several storyboard units according to the division criteria of the time-series storyboard. Based on the preset display time of the time-series storyboard, the difference between the actual display time and the preset display time of video images, audio, and subtitles in each storyboard unit is calculated to obtain the timeline alignment deviation of a single storyboard unit. Then, the average value of the timeline alignment deviation of all storyboard units is calculated as the overall timeline alignment deviation of the video. The algorithm quantifies and converts the overall timeline alignment deviation of the video based on the magnitude of the deviation and the preset scoring criteria to obtain the synchronization quality score of video images, audio, and subtitles.

[0013] Closed-loop management module: Receives the finished video, the timeline alignment deviation, and the synchronization quality score; collects user operation data and feedback, and performs anonymization processing; feeds the anonymized data back to the requirement normalization module and the script compliance module for iterative optimization of algorithm parameters; and provides video saving and sharing functions.

[0014] Furthermore, the demand normalization module receives multimodal health demands in four input formats: voice, text, button click, and touch swipe. For voice-based health demands, preprocessing operations such as voice-to-text conversion and filtering of invalid interjections are performed sequentially. For text-based health demands, preprocessing operations such as typo correction and colloquial content standardization are performed sequentially. For button click and touch swipe-based health demands, information extraction of demand options is performed. The standard demand library contains health demand classification information for elderly chronic disease management, daily health care for the elderly, and first aid knowledge for the elderly. Before calculating semantic similarity, the demand normalization module extracts three types of information—age, health problem, and demand orientation—from the preprocessed user demands and matches them with health entities from the standard demand library.

[0015] Furthermore, in the demand normalization module, the mathematical expression for the aging-friendly demand semantic matching algorithm is:

[0016] ;

[0017] in The semantic similarity between user requirements and the standard requirements library. For multimodal fusion coefficients, The number of healthy entities identified in the user's requirements. For the first The weight value of each healthy entity, The colloquial redundancy penalty coefficient, The variance of semantic recognition error for user needs description. For the natural constant An exponential function with base π is used to perform a non-linear mapping on the weighted similarity sum, enhancing the discriminative power of the matching results. The cosine similarity function is used to calculate the first similarity among user requirements. A healthy entity Corresponding health entity in the standard requirements library The cosine value of the angle between the vector spaces is [0, 1];

[0018] When using the age-friendly demand semantic matching algorithm to calculate the semantic similarity between user needs and the standard demand library, three core elements are first extracted from the preprocessed user needs: health entities, demand orientation, and user features. Health entity labeling and demand classification are then performed on all demand information in the standard demand library. First, user needs are matched with similar health needs in the standard demand library. Then, the semantic similarity between each health entity in the user needs and its corresponding health entity in the standard demand library is calculated one by one. Preset weight values ​​are assigned to different types of health entities, and the comprehensive similarity of each health entity is obtained through weighted calculation. Finally, the features from multimodal input are combined to complete the fusion calculation, resulting in the overall semantic similarity between the user needs and the standard demand library.

[0019] Furthermore, the large language model called by the script compliance module adopts a hierarchical structure combining pre-training and domain fine-tuning. The bottom layer of the model is the natural language understanding layer, which performs semantic parsing, health entity recognition, and precise extraction of demand intent for standardized health demand structured instructions. The middle layer of the model is the elderly health knowledge fusion layer, which continuously accesses the elderly health-specific knowledge graph to complete the real-time embedding of health knowledge and knowledge constraints of generated content. The upper layer of the model is the age-friendly expression generation layer, which is equipped with an elderly health-specific expression corpus to complete the adaptation of generation style and expression habits. The calling process of the large language model adopts the input form of instruction fine-tuning. The standardized health demand structured instructions and the structured prompt words of the elderly health theme are concatenated and input into the model. The model outputs the initial video script according to the preset generation length and fixed expression rules. The generation length is strictly limited to the text volume corresponding to a 3-5 minute video. The expression rules are executed according to the age-friendly setting requirements of single sentence character count and single paragraph sentence count.

[0020] When the script compliance module calls the large language model to generate the initial video script, it first concatenates the standardized health needs structured instructions with the structured prompts on the topic of elderly health to form the model input content. Fixed generation parameters are set for the large language model. The generation parameters include the text generation length matching the 3-5 minute video length, the expression style that fits the elderly group, and the wording standard that avoids professional terminology. After receiving the input content, the large language model completes the semantic understanding, elderly health knowledge integration, and content logic construction processes in sequence. The generated initial video script includes the corresponding information of the screen description, narration content, and subtitle text, and the script content is segmented according to the narrative rhythm of the video.

[0021] Furthermore, in the script compliance module, the mathematical expression for the health script compliance adaptation algorithm is:

[0022] ;

[0023] in The script is evaluated based on a comprehensive compliance score across three dimensions: accuracy of health knowledge, safety of expression, and age-appropriate language. The weighting value for the accuracy dimension of health knowledge. The score for the accuracy dimension of health knowledge. To represent the weight values ​​of the security dimension, To express the score of the security dimension, The weight values ​​for the aging-appropriate expression dimension. The score is for the aging-appropriate expression dimension. The penalty coefficient for violations. This is the identifier value for script violations;

[0024] When using a health script compliance adaptation algorithm to calculate the compliance score of a script across three dimensions—accuracy of health knowledge, safety of expression, and age-appropriate expression—the algorithm compares the content of the initial video script sentence by sentence with the standardized content of the elderly health-specific knowledge graph, and quantifies the score based on the degree of fit. For the safety of expression dimension, the algorithm conducts a full review of the initial video script content, determining and quantifying the presence of any non-compliant medical expressions according to the regulations for medical information publication. For the age-appropriate expression dimension, the algorithm statistically analyzes the sentence length, colloquialism of vocabulary, and information density of the initial video script, and scores it according to preset quantitative indicators. The algorithm assigns preset weight values ​​to the scores of the three dimensions, and the overall compliance score of the script is obtained through weighted calculation.

[0025] Furthermore, the material adaptation module generates a storyboard image sequence based on the scene descriptions in the time-series storyboard, generates voice narration based on the narration content in the time-series storyboard, and generates synchronized subtitles based on the time-series nodes of the voice narration. In the material adaptation module, the video frame rate calculated by the multimodal aging-adaptation parameter adaptation algorithm ranges from 20fps to 24fps, the subtitle font size ranges from 16pt to 48pt, and the voice playback speed ranges from 100 to 140 words per minute. The material adaptation module associates the calculated aging-adaptation parameters of video frame rate, subtitle font size, and voice playback speed with the three types of video materials, and then transmits them to the audio-visual synthesis module.

[0026] Furthermore, in the multimodal aging-adaptive synchronization algorithm, the mathematical expression of the multimodal aging-adaptive synchronization algorithm is:

[0027] ;

[0028] in The calculation results are for the multimodal aging-adaptive synchronization algorithm. This represents the total number of time-series storyboards. For the first Timeline alignment discrepancies in the visuals, audio, and subtitles within a single scene. The video frame rate adaptation factor is expressed as follows: ; This represents the video frame rate adaptation factor, with a value range of [0, 1], and is positively correlated with video smoothness. Represents the actual frame rate of the video, with a value range of [0, 24]. The closer the value is to 24fps, the closer the coefficient is to 1, indicating a higher level of smoothness in the video. The subtitle font size adaptation factor is expressed as follows: , This represents the font size adaptation factor for subtitles, with a value range of [0, 1], and is positively correlated with font readability. This represents the actual font size of the subtitles, and its value range is [value range missing]. When the font size reaches 32 points or above, the coefficient is 1, indicating that it is completely suitable for the elderly. The speech playback speed adaptation coefficient is expressed as follows: , The speech rate adaptation coefficient represents the speech rate adaptation factor, with a value range of [0, 1], and is positively correlated with speech rate suitability. This represents the actual speech rate, with a value range of [0, 240]. When the speech rate is close to 120 words per minute, this coefficient is at its maximum of 1, indicating that the speech rate is most suitable for elderly users.

[0029] Furthermore, the audio-visual synthesis module calculates the time axis alignment deviation as the difference between the actual display time of the image, audio, and subtitles and the preset display time. The synchronization quality score is calculated on a 100-point scale. When the synchronization quality score reaches the preset value, video rendering is performed directly. When the synchronization quality score does not reach the preset value, the timing of the image, audio, and subtitles is calibrated based on the time axis alignment deviation before video rendering. The video rendering is output in MP4 format with a resolution of 1080P. When storing the video, the finished video is archived along with user requirements information and a time-series storyboard. The naming rule for video storage is a unique identifier consisting of a combination of numbers and letters, along with the video format suffix.

[0030] Furthermore, the user operation data collected by the closed-loop management module includes the user's selected health need type, the adjusted values ​​of video playback parameters, the modified content of the script, and the duration and segments of video playback. The collected feedback includes adjustments to the video content, voice narration, and synchronized subtitles. The closed-loop management module performs personal information removal and anonymization processing on the user operation data and semantic extraction and anonymization processing on the user feedback. The anonymized data is classified according to the algorithm parameter categories of the demand normalization module and the script compliance module, and then fed back to the corresponding modules.

[0031] Furthermore, after the closed-loop management module feeds back the anonymized data to the requirement normalization module, the requirement normalization module adjusts the multimodal fusion coefficient, health entity weight value, and colloquial redundancy penalty coefficient in the age-friendly requirement semantic matching algorithm based on the anonymized data. After the anonymized data feeds back to the script compliance module, the script compliance module adjusts the three-dimensional weight values ​​and violation content penalty coefficient in the health script compliance adaptation algorithm based on the anonymized data. The video saving function provided by the closed-loop management module supports storing the finished video to local storage media, and the video sharing function supports generating and outputting the storage link of the finished video. The closed-loop management module also associates the timeline alignment deviation and synchronization quality score with the anonymized data, and participates in the algorithm parameter iteration optimization together.

[0032] Compared with existing technologies, this AI video generation system for elderly health management based on multimodal interaction has the following beneficial effects:

[0033] I. This invention establishes a multimodal interactive AI video generation system for elderly health management, enabling precise capture and standardized transformation of elderly health needs. It combines a dedicated knowledge graph with medical publishing standards to complete multi-dimensional compliance verification of video scripts, ensuring that generated content meets both the professional requirements of elderly health knowledge and the age-friendly expression needs. The system relies on a layered, large language model to generate scripts and uses algorithms to adapt video materials to age-friendly parameters, taking into account the visual and auditory physiological characteristics of the elderly. This ensures that the video's frame rate, subtitles, and voice narration are all adapted to the elderly user's reception capabilities. Simultaneously, through refined synchronous quality assessment of the storyboard unit, it guarantees the temporal consistency of video footage, audio, and subtitles, improving the overall viewing experience and allowing elderly users to clearly and smoothly obtain health management-related information.

[0034] Second, this invention constructs a closed-loop algorithm optimization mechanism throughout the entire process. User operation data and feedback are anonymized and then fed back to the core algorithm module, enabling dynamic iteration of algorithm parameters for demand matching and script compliance. This allows the system to continuously adapt to the actual usage needs and expression habits of elderly users. The system adopts a front-end and back-end separation architecture, achieving efficient collaboration between the interactive interface and business logic. Each module completes data transmission and processing according to standardized processes, ensuring the efficiency and standardization of video generation. At the same time, it provides video saving and sharing functions to meet the usage and dissemination needs of elderly users. The entire process is designed with age-friendliness as the core principle. From demand input to video output, each link is tailored to the operation and cognitive characteristics of the elderly, significantly improving the targeting and effectiveness of elderly health management information dissemination.

[0035] Other advantages, objectives and features of the invention will be set forth in part in the description which follows, and in part will be apparent to those skilled in the art from the following examination or study, or may be learned from the practice of the invention. Attached Figure Description

[0036] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are merely some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without any creative effort.

[0037] Figure 1 A flowchart of an AI video generation system for elderly health management based on multimodal interaction;

[0038] Figure 2 This is a framework diagram of an AI video generation system for elderly health management based on multimodal interaction. Detailed Implementation

[0039] To further illustrate the technical means and effects of the present invention in achieving its intended purpose, the following detailed description of the specific implementation methods, structures, features, and effects of the present invention, in conjunction with the accompanying drawings and preferred embodiments, is provided below.

[0040] Example 1:

[0041] This invention relates to an AI video generation system for elderly health management based on multimodal interaction. It adopts a front-end / back-end separation architecture. The front-end, based on the React framework, completes the visual construction of the user interface and real-time response to various operation commands. The back-end, based on Python, handles the entire process of business logic processing and seamless integration with various AI services. The system consists of a requirement normalization module, a script compliance module, a material adaptation module, an audio / video synthesis module, and a closed-loop control module, which establish signal connections sequentially. Each module operates collaboratively according to the order of data transmission and processing, achieving intelligent processing from the input of elderly users' health needs to the output of age-appropriate AI video products. Simultaneously, it collects and feeds back user data and continuously iterates and optimizes system algorithm parameters to adapt to the physiological characteristics, operating habits, and health information acquisition needs of the elderly population. Figure 1 As shown.

[0042] Requirements normalization module:

[0043] The demand normalization module, serving as the system's demand input and standardization processing unit, is responsible for receiving, preprocessing, entity matching, and semantic similarity calculation of multimodal health needs of elderly users. It ultimately generates standardized structured health demand instructions, which are then passed to subsequent modules. This is a fundamental step in achieving accurate transformation of health needs. The specific implementation method is as follows:

[0044] Multimodal health needs reception: The module supports four forms of health needs input: voice, text, button click, and touch swipe. It is fully adapted to the operation and expression habits of elderly users or their family members. It can receive a single form of needs input alone, or it can receive a combination of multiple forms of needs input simultaneously, realizing the synchronous collection and integrated storage of multi-source needs data, and ensuring the convenience and comprehensiveness of needs input.

[0045] Differentiated Preprocessing of Needs: For health needs in different input formats, customized preprocessing operations are performed to remove invalid information and standardize data format. Specifically: For health needs in voice format, speech recognition technology is first used to convert speech to text, and then natural language processing technology is used to filter out invalid interjections, removing words without actual semantic meaning such as "um," "ah," and "ya," while retaining the core need expression; For health needs in text format, Chinese error correction algorithms are first used to correct typos, and then colloquial and fragmented expressions are standardized and converted into logically clear and standardized written content; For health needs in button click or touch swipe format, information extraction of need options is directly performed to clarify the selected health need category, specific problem, and other key information from the user's touch operation, forming structured raw need data.

[0046] Health Entity Matching: The system's pre-built standard requirement library contains three main categories of health requirement information: chronic disease management for the elderly, daily health care for the elderly, and first aid knowledge for the elderly. Each category contains multiple subcategories and specific standard requirement items. After completing requirement preprocessing, the module first accurately extracts three core information categories—age, health problem, and requirement focus—from the preprocessed user requirements. This extracted information is then precisely matched with health entities in the standard requirement library to initially determine the health requirement category to which the user's requirement belongs, thus defining the matching range for subsequent semantic similarity calculations.

[0047] Semantic similarity calculation for age-friendly requirements: The semantic similarity between user requirements and similar requirements in the standard requirement library is calculated using an age-friendly requirement semantic matching algorithm. The mathematical expression for the age-friendly requirement semantic matching algorithm is as follows:

[0048] ;

[0049] in The semantic similarity between user requirements and the standard requirements library. For multimodal fusion coefficients, The number of healthy entities identified in the user's requirements. For the first The weight value of each healthy entity, The colloquial redundancy penalty coefficient, The variance of semantic recognition error for user needs description. For the natural constant An exponential function with base π is used to perform a non-linear mapping on the weighted similarity sum, enhancing the discriminative power of the matching results. The cosine similarity function is used to calculate the first similarity among user requirements. A healthy entity Corresponding health entity in the standard requirements library The cosine of the angle between the vector spaces is [0, 1]. During the calculation, three core elements—health entities, demand orientation, and user features—are first extracted from the preprocessed user requirements. Health entity labeling and demand classification are then performed on similar demand information within the standard demand library. Next, the semantic similarity between each health entity in the user requirements and its corresponding counterpart in the standard demand library is calculated. A weighted calculation is then performed to obtain the comprehensive similarity of each health entity. Finally, the features from the multimodal input are combined to complete the fusion calculation, yielding the overall semantic similarity between the user requirements and the standard demand library.

[0050] Standardized Health Request Structured Instruction Generation: Based on semantic similarity calculations, users' non-standardized and unstructured health requests are converted into standardized structured health request instructions containing key information such as health request classification, specific request content, user characteristics, and request priority. These instructions are encapsulated in a fixed data format to ensure accurate parsing by subsequent modules. After generation, the instruction is transmitted to the script compliance module in real time. Figure 2 As shown.

[0051] Script compliance module:

[0052] The script compliance module receives standardized, structured health requirement instructions from the requirement normalization module. Its core functions include initial video script generation, multi-dimensional compliance scoring, script optimization, and sequential storyboard splitting. This is a crucial step in ensuring the health, compliance, and age-appropriateness of video scripts. The specific implementation method is as follows:

[0053] The large language model, used in both the initial video script generation module and the invocation of the large language model, employs a hierarchical structure combining pre-training and domain fine-tuning. The bottom layer is the natural language understanding layer, responsible for deep semantic parsing of standardized health-related structured instructions, accurate identification of health entities, and extraction of intent, ensuring a precise understanding of the user's core health needs. The middle layer is the elderly health knowledge fusion layer, continuously accessing an elderly-specific health knowledge graph in real-time to embed health knowledge and constrain the generated content, preventing errors in health knowledge. The top layer is the age-friendly expression generation layer, equipped with a specially constructed corpus of elderly health-related expressions, adapting the generation style to the expression habits of the elderly, ensuring the script's expression aligns with the comprehension abilities of elderly users. The large language model invocation process uses an instruction fine-tuning input format. Standardized health-related structured instructions are first concatenated with structured prompts on the topic of elderly health to form the model's standardized input content. The structured prompts clearly define the script's age-friendliness requirements, content scope, and expression standards. Simultaneously, fixed generation parameters are set for the large language model. These parameters include text generation length matching the 3-5 minute video duration, a colloquial style of expression suitable for the elderly, standardized word choice avoiding professional medical terminology, and restrictions on sentence length and information density. After receiving the input content, the large language model sequentially completes the processes of semantic understanding, elderly health knowledge integration, and content logic construction, ultimately generating an initial video script. The script contains one-to-one correspondence information between visual descriptions, narration, and subtitle text, and the script content is strictly segmented according to the video's narrative rhythm to ensure ease of subsequent scene breakdown.

[0054] The health script compliance adaptation score calculation uses a health script compliance adaptation algorithm to calculate the compliance score of the initial video script in three dimensions: accuracy of health knowledge, safety of expression, and age-appropriate expression. The mathematical expression of the health script compliance adaptation algorithm is:

[0055] ;

[0056] in The script is evaluated based on a comprehensive compliance score across three dimensions: accuracy of health knowledge, safety of expression, and age-appropriate language. The weighting value for the accuracy dimension of health knowledge. The score for the accuracy dimension of health knowledge. To represent the weight values ​​of the security dimension, To express the score of the security dimension, The weight values ​​for the aging-appropriate expression dimension. The score is for the aging-appropriate expression dimension. The penalty coefficient for violations. This is a value used to identify script violations. If violations are found, the value is assigned based on the number and severity of the violation. The scoring for all three dimensions is quantified on a 100-point scale. The specific judgment rules are as follows: For the accuracy of health knowledge, the content of the initial video script is precisely compared sentence by sentence with the standardized content of the elderly-specific health knowledge graph. Scoring is based on the degree of fit, with perfect fit receiving full marks. Minor deviations result in proportional deductions, and factual errors result in a score of 0 for this dimension. For the safety of expression, the initial video script is fully reviewed, and in accordance with relevant national regulations on medical information release, it is judged and quantified for the presence of illegal medical expressions, exaggerated claims, and absolute statements. No illegal expressions result in full marks, minor illegal expressions result in proportional deductions, and serious illegal expressions result in a score of 0 for this dimension. For the age-appropriate expression dimension, the initial video script's sentence length, vocabulary accessibility, and information density are statistically analyzed and scored according to preset quantitative indicators. Sentence length within a preset reasonable range, completely accessible vocabulary, and information density suitable for the elderly's comprehension result in full marks; any indicator failing to meet the requirements results in proportional deductions. After scoring each of the three dimensions individually, the script's overall compliance score is calculated based on the algorithm formula and the preset weight values ​​for each dimension.

[0057] The script optimization and sequential storyboard splitting system has a preset comprehensive compliance score threshold. If the initial video script's comprehensive compliance score is lower than this threshold, the module will identify the dimension corresponding to the low score and make targeted modifications and adjustments to the script content for that dimension. If the health knowledge accuracy dimension is incorrect, it will be replaced with standardized content from the knowledge graph; if the expression safety dimension is in violation, the inappropriate expression will be deleted or corrected; if the age-appropriate expression dimension is not up to standard, the sentences will be simplified, common words will be replaced, and information density will be reduced. After modification, the comprehensive compliance score will be recalculated until the score reaches the preset threshold. If the initial video script's comprehensive compliance score reaches the preset threshold, the module will evenly split the initial video script into 6-8 sequential storyboards according to the total video length of 3-5 minutes. Each sequential storyboard corresponds to an independent segment of the video, and each storyboard clearly labels three types of information: scene description, narration content, and subtitle text. The scene description clearly defines the scene, scene elements, and camera movements. The narration content and subtitle text are completely corresponding and adapted to the auditory and visual reception habits of the elderly. After splitting, all sequential storyboards are passed to the material adaptation module.

[0058] Material adaptation module:

[0059] The material adaptation module receives the time-series storyboard script from the script compliance module, generates three types of video materials: storyboard image sequences, voice-over narration, and synchronized subtitles. Simultaneously, it calculates age-appropriate parameters using algorithms and associates them with the materials, providing standardized, age-appropriate video materials for audio-visual synthesis. The specific implementation method is as follows:

[0060] Multi-type video material generation: Based on the scene descriptions in the time-series storyboard, the module generates a storyboard image sequence that highly matches the scene descriptions using AI image generation technology. The image sequence has a simple style and soft colors, which is suitable for the visual characteristics of the elderly. Based on the narration content in the time-series storyboard, the module generates voice narration using speech synthesis technology. The voice tone is selected to be gentle and clear, making it age-friendly and ensuring that elderly users can clearly recognize it. Based on the time sequence nodes and text content of the voice narration, synchronized subtitles are generated. The subtitle text is completely consistent with the narration content and is displayed 1-2 frames in advance, which is adapted to the auditory and visual rhythm of elderly users.

[0061] Multimodal aging-friendly parameter calculation: A multimodal aging-friendly parameter adaptation algorithm is used to calculate three core aging-friendly parameters: video frame rate, subtitle font size, and voice playback speed. In the calculation process, three key indicators, namely content rhythm, information density, and screen description complexity, are first accurately extracted from the time-series storyboard. Then, combined with the visual and auditory physiological characteristics of the elderly, fixed parameter value ranges are set for the three types of aging-friendly parameters. The video frame rate ranges from 20fps to 24fps, the subtitle font size ranges from 16 to 48, and the voice playback speed ranges from 100 to 140 words per minute. The algorithm calculates the appropriate video frame rate based on the complexity of the scene description: 20fps-22fps for low scene description complexity and high proportion of static scenes, and 22fps-24fps for high scene description complexity and high proportion of dynamic scenes. It also calculates the appropriate subtitle font size based on the size of the screen display area and the number of subtitle characters: a larger font size for a large screen display area and fewer subtitle characters, and a moderate font size for a small screen display area and more subtitle characters, with all values ​​falling within the range of 16-48 points. Finally, it calculates the appropriate voice playback speed based on the difficulty and length of the narration: 120-140 words per minute for simple narration and short sentences, and 100-120 words per minute for complex narration and longer sentences. All calculated aging-appropriate parameters strictly fall within the corresponding preset parameter value ranges to ensure adaptability and rationality.

[0062] Material and Aging-Friendly Parameter Association Transfer: The module associates the calculated three types of aging-friendly parameters—video frame rate, subtitle font size, and voice playback speed—with the generated three types of video materials—storyboard image sequence, voice narration, and synchronized subtitles—one by one, and marks the corresponding aging-friendly parameter requirements for each type of material to ensure accurate application in the subsequent audio and video synthesis process. After the association is completed, the three types of video materials and their corresponding aging-friendly parameters are transferred to the audio and video synthesis module.

[0063] Audio and video synthesis module:

[0064] The audio-visual synthesis module receives three types of video materials with associated aging-appropriate parameters from the material adaptation module, completes the synchronization quality assessment, timing calibration, video rendering, and storage archiving of video images, audio, and subtitles, and finally outputs an AI video product that meets aging-appropriate requirements. The specific implementation method is as follows:

[0065] Multimodal synchronization quality assessment employs a multimodal aging-adaptive synchronization algorithm to calculate the timeline alignment deviation and synchronization quality score of video footage, audio, and subtitles. The calculation process consists of two steps: First, according to the division criteria of the time-series storyboard, the video to be synthesized is completely divided into storyboard units, the same number as the storyboard. Using the preset display time of each storyboard unit in the time-series storyboard as a benchmark, the difference between the actual display time and the preset display time of video footage, audio, and subtitles within each storyboard unit is calculated. This difference is the timeline alignment deviation of a single storyboard unit. Second, the arithmetic mean of the timeline alignment deviations of all storyboard units is calculated. This average is the overall timeline alignment deviation of the video. After completing the timeline alignment deviation calculation, based on the magnitude of the overall timeline alignment deviation of the video, a quantification conversion is performed using the system's preset percentage scoring standard to obtain the synchronization quality score of video footage, audio, and subtitles. A timeline alignment deviation of 0 is a perfect score of 100 points; the larger the deviation value, the lower the score; a deviation value exceeding the preset range results in a score of 0. Meanwhile, the module further optimizes the synchronization effect of multimodal elements in the video through a multimodal aging-adaptive synchronization algorithm. The mathematical expression of the multimodal aging-adaptive synchronization algorithm is as follows:

[0066] ;

[0067] in The calculation results are for the multimodal aging-adaptive synchronization algorithm. This represents the total number of time-series storyboards. For the first Timeline alignment discrepancies in the visuals, audio, and subtitles within a single scene. This is a function for calculating the aging parameters for video frame rate. For the video frame rate, This is a function for calculating the aging-appropriate parameters for subtitle font size. For the font size of the subtitles, This is a function for calculating the aging-appropriate parameters for voice broadcast speed. Set a value for the voice broadcast speed.

[0068] Video Timing Calibration and Rendering: The system presets a passing score for synchronization quality. If the synchronization quality score reaches this passing score, the module directly performs audio-visual synthesis rendering on the storyboard image sequence, voice narration, and synchronized subtitles based on the associated age-appropriate parameters. If the synchronization quality score does not reach the preset passing score, the module will precisely calibrate the timing of the video, audio, and subtitles based on the calculated timeline alignment deviation, adjusting the display time of each element to ensure a high degree of timeline consistency. After calibration, the audio-visual synthesis rendering operation is performed. Video rendering uniformly uses the MP4 universal video format for output, with a fixed resolution of 1080P. This ensures both video clarity and compatibility with the playback requirements of various terminal devices. Simultaneously, the rendering process strictly adheres to the associated age-appropriate parameters, ensuring that the video frame rate, subtitle font size, and voice playback speed all meet the reception characteristics of elderly users.

[0069] Video Finished Product Storage and Archiving: This module associates and archives the rendered video finished product with the user's original requirement information, standardized health requirement structured instructions, and time-sequential storyboards, enabling traceability of the video finished product and all process data. The naming convention for video storage uses a unique identifier combining numbers and letters with an MP4 video format suffix. The unique identifier contains implicit information such as user requirement category and generation time, facilitating subsequent retrieval, access, and management. Simultaneously, the module transmits the video finished product, the overall timeline alignment deviation of the video, and the synchronization quality score to the closed-loop management module in real time.

[0070] Closed-loop control module:

[0071] The closed-loop management module, serving as the system's feedback optimization and function expansion unit, completes the collection, anonymization, and algorithm parameter iterative optimization of user operation data and feedback. It also provides users with practical functions for video saving and sharing, achieving closed-loop system upgrades. The specific implementation method is as follows:

[0072] Comprehensive User Data Collection: The module continuously collects user operation data and subjective feedback during system use. The collected user operation data includes the user's selected health needs type, manual adjustments to video playback parameters, manual modifications to script content, video playback duration, and key video segments watched. The collected user feedback includes suggestions for adjusting video content, evaluations of the tone and speed of voice narration, and modifications to the font and display rhythm of synchronized subtitles, achieving comprehensive data collection of user behavior and subjective experience.

[0073] User data anonymization: To ensure user information security, the module performs strict anonymization on all collected user data. Specifically, user operation data undergoes personal information removal anonymization, completely deleting all sensitive personal information such as name, age, contact information, and home address, retaining only non-sensitive data related to operational behavior. User feedback undergoes semantic extraction anonymization, using natural language processing technology to extract core adjustment requests and evaluation opinions from the feedback, removing irrelevant statements and potentially sensitive information. The anonymized data retains only the effective information needed for system optimization, ensuring no risk of user privacy leakage during data return.

[0074] Data anonymization feedback and algorithm parameter iteration optimization: The module accurately categorizes the anonymized user data according to the algorithm parameter categories of the requirement normalization module and the script compliance module. After categorization, the data is fed back to the corresponding modules for iterative optimization of algorithm parameters. When the anonymized data is fed back to the requirement normalization module, the module dynamically adjusts the multimodal fusion coefficient, the weight values ​​of each health entity, and the colloquial redundancy penalty coefficient in the age-appropriate requirement semantic matching algorithm, making the semantic similarity calculation results more aligned with users' actual needs and expression habits. When the anonymized data is fed back to the script compliance module, the module dynamically adjusts the weight values ​​of the three dimensions of health knowledge accuracy, expression security, and age-appropriate expression in the health script compliance adaptation algorithm, as well as the penalty coefficient for violations, making the script compliance scoring standards more aligned with users' actual usage needs and health information acquisition requirements. Simultaneously, the module correlates the video's timeline alignment deviation and synchronization quality score with the anonymized user data, participating in the iterative optimization of algorithm parameters in both modules. This allows the system to continuously improve the accuracy of requirement processing and script generation based on the actual effect of video synthesis and user feedback.

[0075] Video saving and sharing functionality: This module provides users with convenient video saving and sharing capabilities, meeting their needs for offline viewing and knowledge dissemination. The video saving function supports direct storage of finished videos to local computers, tablets, mobile phones, and other storage media, while also supporting multiple resolution options to adapt to the storage requirements of different devices. The video sharing function supports automatically generating a unique storage link for the finished video, with a reasonable validity period set. Users can use this link to share the finished video to social media platforms, friends' and family's mobile devices, etc., achieving efficient dissemination of elderly health management knowledge. Furthermore, the module records users' video saving and sharing operations. This recorded data, after being anonymized, will be fed back into the system, providing data support for future function optimization.

[0076] This invention relates to an AI video generation system for elderly health management based on multimodal interaction. Through multimodal interaction, it accurately captures and standardizes the health needs of elderly users. Combining a hierarchical large language model with an elderly-specific health knowledge graph, it ensures the health, compliance, and age-friendliness of video scripts. Furthermore, through material adaptation and audio-visual synthesis, it achieves accurate synthesis and rendering of age-friendly video materials, ultimately outputting AI video products that fit the physiological characteristics of the elderly. Simultaneously, the system uses a closed-loop management module to collect and feedback user data and continuously iterate and optimize algorithm parameters. This ensures that the system's functions and performance continuously adapt to the changing needs of elderly users, providing an efficient and intelligent solution for the visualization and popularization of elderly health management knowledge, effectively improving the convenience and effectiveness of health information access for the elderly.

[0077] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any simple modifications, equivalent changes and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention.

Claims

1. An AI video generation system for elderly health management based on multimodal interaction, characterized in that, The system includes: Demand normalization module: Receives multimodal health demand input from elderly users, uses an age-friendly demand semantic matching algorithm to calculate the semantic similarity between user demands and the standard demand library, generates standardized health demand structured instructions, and passes them to the script compliance module. Script Compliance Module: Receives the standardized health needs structured instructions, calls the large language model to generate the initial video script, uses the health script compliance adaptation algorithm to calculate the compliance score of the script in three dimensions: accuracy of health knowledge, safety of expression, and age-appropriate expression, iteratively optimizes the initial video script based on the compliance score until the compliance score reaches the preset threshold, generates a time-series storyboard script and passes it to the material adaptation module. Material adaptation module: Receives the time-series storyboard script, generates three types of video materials: storyboard image sequence, voice narration, and synchronized subtitles, uses a multimodal aging-adaptation parameter adaptation algorithm to calculate the aging-adaptation parameters for video frame rate, subtitle font size, and voice playback speed, and transmits the three types of video materials to the audio-visual synthesis module; Audio and video synthesis module: Receives three types of video materials, uses a multimodal aging-adaptive synchronization algorithm to calculate the time axis alignment deviation and synchronization quality score of video frame, audio and subtitle, performs video rendering and storage based on the calculation results, and transmits the finished video and the time axis alignment deviation and synchronization quality score to the closed-loop control module; Closed-loop management module: Receives the finished video, the timeline alignment deviation, and the synchronization quality score; collects user operation data and feedback, and performs anonymization processing; feeds the anonymized data back to the requirement normalization module and the script compliance module for iterative optimization of algorithm parameters; and provides video saving and sharing functions.

2. The AI ​​video generation system for elderly health management based on multimodal interaction according to claim 1, characterized in that, The demand normalization module receives multimodal health demands including four input forms: voice, text, button click, and touch swipe. For voice-based health demands, it performs preprocessing operations such as voice-to-text conversion and filtering of invalid interjections. For text-based health demands, it performs preprocessing operations such as typo correction and colloquial content standardization. For button click and touch swipe-based health demands, it performs information extraction operations for demand options. The standard demand library contains health demand classification information for elderly chronic disease management, elderly daily health care, and elderly first aid knowledge. Before calculating semantic similarity, the demand normalization module extracts three types of information from the preprocessed user demands: age, health problems, and demand orientation, and matches them with the standard demand library for health entity matching.

3. The AI ​​video generation system for elderly health management based on multimodal interaction according to claim 1, characterized in that, In the demand normalization module, the mathematical expression for the aging-friendly demand semantic matching algorithm is: ; in The semantic similarity between user requirements and the standard requirements library. For multimodal fusion coefficients, The number of healthy entities identified in the user's requirements. For the first The weight value of each healthy entity. The colloquial redundancy penalty coefficient, The variance of semantic recognition error for user needs description. For the natural constant An exponential function with base π is used to perform a non-linear mapping on the weighted similarity sum, enhancing the discriminative power of the matching results. The cosine similarity function is used to calculate the first similarity among user requirements. A healthy entity Corresponding health entity in the standard requirements library The cosine of the angle between the vectors in the vector space ranges from [0, 1].

4. The AI ​​video generation system for elderly health management based on multimodal interaction according to claim 1, characterized in that, The large language model called by the script compliance module adopts a hierarchical structure combining pre-training and domain fine-tuning. The bottom layer of the model is the natural language understanding layer, which performs semantic parsing, health entity recognition, and accurate extraction of demand intent for standardized health demand structured instructions. The middle layer of the model is the elderly health knowledge fusion layer, which continuously connects to the elderly health-specific knowledge graph to complete the real-time embedding of health knowledge and knowledge constraints of generated content. The top layer of the model is the age-friendly expression generation layer, which is equipped with an elderly health-specific expression corpus to complete the adaptation of generation style and expression habits. The large language model is called in the input form of instruction fine-tuning. The standardized health demand structured instructions and the structured prompt words of the elderly health theme are concatenated and input into the model. The model outputs the initial video script according to the preset generation length and fixed expression rules. The generation length is strictly limited to the text volume corresponding to a 3-5 minute video. The expression rules are executed according to the age-friendly setting requirements of single sentence character count and single paragraph sentence count.

5. The AI ​​video generation system for elderly health management based on multimodal interaction according to claim 1, characterized in that, In the script compliance module, the mathematical expression for the health script compliance adaptation algorithm is: ; in The script is evaluated based on a comprehensive compliance score across three dimensions: accuracy of health knowledge, safety of expression, and age-appropriate language. The weighting value for the accuracy dimension of health knowledge. The score for the accuracy dimension of health knowledge. To represent the weight values ​​of the security dimension, To express the score of the security dimension, The weight values ​​for the aging-appropriate expression dimension. The score is for the aging-appropriate expression dimension. The penalty coefficient for violations. This is the identifier value for script violations.

6. The AI ​​video generation system for elderly health management based on multimodal interaction according to claim 1, characterized in that, The material adaptation module generates a storyboard image sequence based on the scene descriptions in the time-series storyboard, generates voice narration based on the narration content in the time-series storyboard, and generates synchronized subtitles based on the time sequence nodes of the voice narration. In the material adaptation module, the multimodal aging-adaptive parameter adaptation algorithm calculates the video frame rate in the range of 20fps-24fps, the subtitle font size in the range of 16-48pt, and the voice playback speed in the range of 100-140 words per minute. The material adaptation module associates the calculated aging-adaptive parameters of video frame rate, subtitle font size, and voice playback speed with the three types of video materials, and then transmits them to the audio-visual synthesis module.

7. The AI ​​video generation system for elderly health management based on multimodal interaction according to claim 1, characterized in that, In the multimodal aging-adaptive synchronization algorithm, the mathematical expression of the multimodal aging-adaptive synchronization algorithm is: ; in The calculation results are for the multimodal aging-adaptive synchronization algorithm. This represents the total number of time-series storyboards. For the first Timeline alignment discrepancies in the visuals, audio, and subtitles within a single scene. This is a function for calculating the aging parameters for video frame rate. For the video frame rate, This is a function for calculating the aging-appropriate parameters for subtitle font size. For the font size of the subtitles, This is a function for calculating the aging-appropriate parameters for voice broadcast speed. Set a value for the voice broadcast speed.

8. The AI ​​video generation system for elderly health management based on multimodal interaction according to claim 1, characterized in that, The audio-visual synthesis module calculates the time axis alignment deviation as the difference between the actual display time of the image, audio, and subtitles and the preset display time. The synchronization quality score is calculated on a 100-point scale. When the synchronization quality score reaches the preset value, video rendering is performed directly. When the synchronization quality score does not reach the preset value, the timing of the image, audio, and subtitles is calibrated according to the time axis alignment deviation before video rendering. The video rendering is output in MP4 format with a resolution of 1080P. When storing the video, the finished video is archived along with user requirements information and a time-series storyboard. The naming rule for video storage is a unique identifier consisting of a combination of numbers and letters, along with the video format suffix.

9. The AI ​​video generation system for elderly health management based on multimodal interaction according to claim 1, characterized in that, The closed-loop management module collects user operation data including the user's selected health need type, adjusted values ​​of video playback parameters, modified script content, and video playback duration and segments. The collected feedback includes adjustments to video content, voice narration, and synchronized subtitles. The closed-loop management module performs personal information removal and anonymization processing on user operation data and semantic extraction and anonymization processing on user feedback. The anonymized data is classified according to the algorithm parameter categories of the requirement normalization module and the script compliance module, and then fed back to the corresponding modules.

10. The AI ​​video generation system for elderly health management based on multimodal interaction according to claim 1, characterized in that, The closed-loop management module feeds back the anonymized data to the requirement normalization module. The requirement normalization module then adjusts the multimodal fusion coefficient, health entity weight value, and colloquial redundancy penalty coefficient in the age-friendly requirement semantic matching algorithm based on the anonymized data. After the anonymized data is fed back to the script compliance module, the script compliance module adjusts the three-dimensional weight values ​​and violation content penalty coefficient in the health script compliance adaptation algorithm based on the anonymized data. The video saving function provided by the closed-loop management module supports storing the finished video to local storage media, and the video sharing function supports generating and outputting the storage link of the finished video. The closed-loop management module also correlates the timeline alignment deviation and synchronization quality score with the anonymized data, participating in the algorithm parameter iteration optimization.