Exercise training recommendation method and system based on multi-modal data fusion

By constructing the role portrait of elderly disabled patients and using a collaborative filtering model of multimodal data fusion, a personalized exercise training scheme was generated, which solved the problem of failing to effectively integrate speech and non-verbal system information in the existing technology, and achieved individualized exercise training for people with disability risk, improving the training effect.

CN120108640AActive Publication Date: 2025-06-06BEIJING REHABILITATION HOSPITAL CAPITAL MEDICAL UNIVERSITY(BEIJING WORKERS SANATORIUM)

Patent Information

Application Number
CN202510120592.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-25
Publication Date
2025-06-06
Estimated Expiration
2045-01-25

AI Technical Summary

Technical Problem

The prior art has failed to effectively integrate the information of the speech system and non-verbal system, and lacks individualized exercise training programs for people with disability risk, making it difficult to improve or improve their cognitive, emotional, exercise, speech and other abilities.

Method used

A motion training recommendation method based on multimodal data fusion is adopted to build a role portrait by obtaining the disease type, clinical characteristics, demographic information, sports research information and motor behavior characteristics of elderly disabled patients, and using collaborative filtering models and deep learning technology to integrate structured data, text data and image data to generate a personalized motion training plan.

Benefits of technology

The recommendation of personalized exercise training plans based on the patient's multi-dimensional exercise data is realized, which improves the pertinence and effectiveness of the training, and enhances the patient's cognitive, emotional, exercise and other abilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120108640A_ABST
    Figure CN120108640A_ABST
Patent Text Reader

Abstract

The invention discloses an exercise training recommendation method and system based on multi-modal data fusion. The method comprises the following steps: obtaining disease types and clinical features of the elderly disabled patients, fusing and constructing role portraits, and obtaining static features; generating and pushing a first exercise training scheme based on the role portrait; collecting structured data, text data and image data during exercise training; and inputting the data into the collaborative filtering model, and generating and pushing a next exercise training scheme in combination with the role portrait. By utilizing the method, an individualized targeted exercise training scheme can be formulated, and the capabilities of cognition, emotion, exercise, speech and the like of the elderly disabled patient can be remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to a sports training recommendation method based on multimodal data fusion, and also relates to a corresponding sports training recommendation system, belonging to the technical field of medical care informatics. Background Art

[0002] According to the dual coding theory in cognitive theory, it is generally believed that humans have two main information processing systems: the verbal system and the non-verbal system. These two systems differ in the way they are encoded. Studies have shown that the brain's memory effect and speed for image materials are better than semantic memory, which suggests that the non-verbal system may have advantages in processing image and spatial information. Although the two systems are functionally independent, they can also activate and influence each other.

[0003] In the cognitive field, there is a theory called constructivism, the core of which is that knowledge is constructed through the interaction between individuals and the environment. Through communication and cooperation with others, individuals can construct and revise their own understanding; when individuals encounter new information that conflicts with existing knowledge, cognitive conflict will arise. This conflict prompts individuals to reassess and adjust their cognitive structures to adapt to the new information. Moreover, knowledge construction is an ongoing process. As individuals gain experience and develop cognitive abilities, their knowledge structures will continue to evolve. Knowledge is constructed in a specific context, so learning in real contexts can better promote the understanding and application of knowledge.

[0004] However, there is currently no exercise training method that integrates information from the speech system and the non-verbal system in the treatment of elderly patients with disabilities. Moreover, for people at risk of disability, it is necessary to develop individualized and targeted exercise training programs to enhance or improve their cognitive, emotional, motor, and speech abilities. Summary of the invention

[0005] The primary technical problem to be solved by the present invention is to provide a sports training recommendation method based on multimodal data fusion.

[0006] Another technical problem to be solved by the present invention is to provide a sports training recommendation system based on multimodal data fusion.

[0007] In order to achieve the above technical objectives, the present invention adopts the following technical solutions:

[0008] According to a first aspect of an embodiment of the present invention, a sports training recommendation method based on multimodal data fusion is provided, comprising the following steps:

[0009] Step 1: For elderly disabled patients, obtain the patient's disease type and clinical characteristics, demographic information, exercise survey information and exercise behavior characteristics, so as to integrate and construct a role portrait and obtain static characteristics;

[0010] Step 2: Generate and push the first sports training plan based on the role portrait;

[0011] Step 3: Collect structured data, text data, and image data during sports training;

[0012] Step 4: Input the data in step 3 into the collaborative filtering model and combine it with the role portrait to generate and push the next sports training plan;

[0013] The collaborative filtering model adopts multiple autoencoders and collaborative filtering networks, and uses graph embedding technology to map entities and relationships in the graph structure into a low-dimensional vector space. It is implemented through a Bayesian TransR embedding model, a Bayesian stack denoising autoencoder, and a Bayesian stack convolutional autoencoder to convert structured data, text data, and image data into structure vectors, text vectors, and image vectors, respectively, thereby obtaining item latent vectors and user latent vectors for collaborative filtering learning to recommend the next sports training plan.

[0014] Preferably, the step 4 includes the following sub-steps:

[0015] 1) Determine the key variables as one-dimensional data, two-dimensional data, and three-dimensional data;

[0016] 2) Create a data matrix based on key variables;

[0017] 3) Add hysteresis features as four-dimensional data;

[0018] 4) Convert it into feature vector using embedding technology;

[0019] 5) Based on the feature vector, the item latent vector and the user latent vector are generated for collaborative filtering learning to recommend sports training plans.

[0020] Preferably, in the sub-step 4, the entities and relationships are converted into vectors using the Bayesian TransR knowledge graph embedding method; the text data is converted into text vectors using the Bayesian sparse autoencoder embedding method; and the Bayesian sparse convolutional autoencoder is used to extract image features and convert them into image vectors.

[0021] Preferably, the multimodal features and role-portrait features extracted from the patient's condition are used as independent variable data, and the exercise prescription recommendations clinically prescribed by the doctor are used as dependent variables, which are input into the collaborative filtering model;

[0022] Among them, the loss function of the collaborative filtering model is cosine contrast loss, which is used to maximize the similarity between positive sample pairs and minimize the similarity of negative sample pairs under margin constraints.

[0023] Preferably, in the sub-step 1, the individual data space is constructed with one-dimensional data as the W axis, two-dimensional data as the Y axis, three-dimensional data as the X axis, and four-dimensional data as the Z axis.

[0024] Preferably, the one-dimensional data includes a time variable, the two-dimensional data includes multiple variables describing exercise preferences; the three-dimensional data includes variables describing physical and psychological characteristics; and the four-dimensional data includes post-training status data.

[0025] Preferably, when creating a data matrix based on key variables, different diseases correspond to different key variables; the character portrait includes multiple attributes, including cognitive cortical damage, physical function damage, and emotional control damage attributes.

[0026] Preferably, the exercise training recommendation method further comprises the following steps:

[0027] Step 5: Collect structured data, text data, and image data during the sports training multiple times in the manner of step 3 to obtain sports training cycle data;

[0028] Step 6: Determine whether the exercise training program is the Nth one, if yes, proceed to step 7; if not, return to step 4, and generate the next exercise training program according to the features in step 5; wherein N is a positive integer;

[0029] Step 7: Based on the patient's data collected in steps 3 and 5, the data are jointly input into the RRN model to iteratively optimize the exercise training program.

[0030] The RRN model is optimized using the following function:

[0031]

[0032] Among them, θ represents the parameters to be learned; is the set of tuples observed in the training set, which includes patients, exercise training programs and time series; r ij |t represents patient i’s score for movement j at time t; For r ij |The predicted value of t; R represents the regularization function.

[0033] Preferably, the RRN model adopts an alternating subspace descent strategy, assumes that the motion state is fixed, does not propagate gradients to these motion training scheme sequences, and simultaneously back-propagates the gradients of all patient scores to update patient sequence parameters, and then switches between updating the user sequence and updating the motion training scheme sequence.

[0034] According to a second aspect of an embodiment of the present invention, a sports training recommendation system based on multimodal data fusion is provided, comprising a processor and a memory, wherein the memory is coupled to the processor and is used to store one or more programs, and when the program is executed by the processor, the processor implements the sports training recommendation method based on multimodal data fusion.

[0035] Compared with the prior art, the present invention has the following beneficial technical effects:

[0036] 1) Use the fusion analysis system of multimodal data in collaborative filtering and autoencoding learning to ensure that multi-source information enters the recommendation model, increase the analysis source, and realize the continuous iteration of exercise push content based on the patient's own multi-dimensional exercise data, so that the training can best meet the individual patient's current status and improve the training effect;

[0037] 2) Construct a training cycle data system, divide the cycle of the exercise training process into annual cycle, large cycle, medium cycle and small cycle, and create different time series data sets. When introducing the model for exercise recommendation calculation, the time series dimension feature is added, so as to more fully and regularly apply historical exercise data to assist in the generation of periodic plans for exercise recommendation;

[0038] 3) Construct a multi-level patient portrait system to gradually and deeply deconstruct and evaluate the patient's attribute characteristics, behavioral characteristics, and comprehensive characteristics (including but not limited to cognitive, emotional, motor, and verbal characteristics) to ensure that the individual label system has a fine granularity, which provides a large amount of effective information on the patient sequence for the subsequent input model to make accurate recommendations on exercise;

[0039] 4) Combined with the RRN model, the recurrent network is used to effectively calculate the time series data, and the data of patients' exercise at different time periods can be effectively traced back, so as to ensure the consistency and sustainability of exercise recommendations and improve patient compliance. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 This is a flow chart of a sports training recommendation method based on multimodal data fusion in the first embodiment of the present invention;

[0041] Figure 2 This is a schematic diagram of multimodal data fusion in the first embodiment of the present invention;

[0042] Figure 3A This is a schematic diagram of a process of generating a sports training program based on a collaborative filtering model in the first embodiment of the present invention;

[0043] Figure 3B This is a schematic diagram of an individual data space in the first embodiment of the present invention;

[0044] Figure 3C This is a schematic diagram of the structure of a collaborative filtering model in the first embodiment of the present invention;

[0045] Figure 4 This is a flow chart of a sports training recommendation method based on multimodal data fusion in the second embodiment of the present invention;

[0046] Figure 5 This is a schematic diagram of the structure of an RRN (recurrent neural network) model in the second embodiment of the present invention;

[0047] Figure 6 Schematic diagram of the structure of a sports training recommendation system based on multimodal data fusion in the third embodiment of the present invention. DETAILED DESCRIPTION

[0048] The technical content of the present invention is described in detail below with reference to the accompanying drawings and specific embodiments.

[0049] The technical concept of the embodiment of the present invention is to combine the content-based recommendation method with the collaborative filtering model, so as to integrate the feature learning of patients or projects and the project recommendation process into a unified framework. First, various deep learning models are used to learn the latent features of patients or projects, and a unified optimization function is constructed in combination with the collaborative filtering method for parameter training. Then, the trained model is used to obtain the final latent vectors of patients and projects, thereby realizing the recommendation of exercise training plans for patients (mainly elderly disabled patients).

[0050] First embodiment

[0051] like Figure 1 As shown in FIG3 , the sports training recommendation method based on multimodal data fusion provided by the first embodiment of the present invention at least includes the following steps.

[0052] Step 1: For elderly disabled patients, obtain the patient's disease type and clinical characteristics, demographic information, exercise survey information and exercise behavior characteristics, so as to integrate and construct a role portrait and obtain static characteristics.

[0053] Using the basic model in the sports training recommendation based on multimodal data fusion, the core profile can be constructed using the patient's disease type and clinical characteristics; the basic profile can be constructed using the patient's demographic information, sports survey information and sports behavior characteristics; and the core profile and basic profile are then integrated to construct the patient's role profile. In this process, static features are obtained.

[0054] It should be noted that the construction method of the above basic model can be found in the prior patent application with application number 202411764001.3 and titled “Profile construction method for people at risk of disability, exercise training recommendation method and system”, which is briefly described as follows:

[0055] The base model was constructed using data from multiple patients using the following steps:

[0056] S1: Determine the key elements of the portrait mapping according to the pathological mechanism;

[0057] S2: Build a core profile based on key elements;

[0058] S3: Build a basic profile based on demographic information, sports survey information and sports behavior characteristics;

[0059] S4: Construct character portraits based on the core portrait set and basic portrait set.

[0060] Among them, key elements include cognitive cortical damage, physical function damage, and emotional control damage; the attributes of the core portrait of disabling diseases include cognitive cortical damage, physical function damage, and emotional control damage. Different attributes include different labels, and the label data comes from the patient's diagnostic information or survey information, etc. Sports survey information includes sports preferences, etc.

[0061] like Figure 2 As shown, damage to the cognitive cortex can lead to cognitive impairment, which is manifested in complex attention, executive function, learning and memory, language, etc. Damage to physical function can lead to movement disorders, which are manifested in muscle strength, muscle tension, coordination, balance, etc. Damage to emotional control can lead to mental and behavioral disorders, which are manifested in hallucinations / delusions or depression / anxiety, etc., and in behavior, aggression / hoarding or depression / wandering, etc. By evaluating the health status (disorders or diseases) of these patients, the label of the patient can be mined.

[0062] The label for a Category A cognitive cortical injury includes at least:

[0063] A1. Frontal lobe damage: Patients may have symptoms of mental disorders such as memory loss and apathy, as well as epilepsy, monoplegia, motor aphasia, high fever, sweating, abnormal vision and smell, etc.

[0064] A2. Parietal lobe injury: The patient may have complex sensory disorders in the contralateral limbs, body image disorders, calculation and other functional disorders, etc.

[0065] A3. Temporal lobe damage: Patients may have sensory aphasia, anomia, olfactory hallucinations, epilepsy, mental abnormalities, memory impairment, visual field changes, etc.

[0066] A4. Occipital lobe damage: The patient may have visual impairment.

[0067] A5. Insular lobe damage: Patients may experience visceral movement and sensory disorders such as increased salivation, nausea, and a feeling of fullness.

[0068] A6. Limbic lobe damage: Patients may have mental abnormalities such as abnormal emotions and memory, hallucinations, and visceral activity disorders.

[0069] A7……

[0070] The label for a Class B impairment includes at least:

[0071] B1. Headache

[0072] B2. Nausea or vomiting

[0073] B3. Fatigue or sleepiness

[0074] B4. Tremor, stiffness and bradykinesia

[0075] B5. Dizziness or loss of balance

[0076] B6. Sensory problems, such as blurred vision, ringing in the ears, bad taste in the mouth, or changes in sense of smell

[0077] B7. Sensitivity to light or sound

[0078] B8……

[0079] The label for Category C emotional control impairment includes at least:

[0080] C1. Mood changes or mood swings

[0081] C2. Depression or anxiety

[0082] C3. Difficulty falling asleep

[0083] C4. Sleepier than usual

[0084] C5……

[0085] Quantify and calibrate the attribute characteristics of the core portrait to construct the core portrait. The attributes of the core portrait can correspond to the core key elements one by one, or there can be more than the core key elements. Multiple labels for each attribute constitute a label vocabulary, and different label vocabulary represent different disease courses, injury types, and disability characteristics. In this step, the core portrait label of the patient is determined based on the patient's confirmed disease type and clinical characteristics.

[0086] Obtain demographic information, exercise survey information, exercise behavior characteristics and social participation information through surveys or medical records, from which the patient's label is obtained and a basic portrait is constructed.

[0087] Demographic information includes gender, age, education level, etc.; sports survey information includes sports preferences and activity ability, etc. Sports preferences include sports types (such as skipping, running, etc.), sports venues (such as gyms, parks, etc.), sports atmosphere (quiet, dynamic), etc. Activity ability includes BADL (basic daily living activities) and IADL (instrumental daily living activities).

[0088] Movement behavior characteristics include movement frequency, movement cycle, etc.

[0089] Social participation information includes participation in family activities and participation in social activities. Family participation activities include family decision-making, child-rearing, etc.; social participation activities include work, study, and entertainment activities.

[0090] The basic portrait is not a core factor that affects the occurrence of disability, but as a basic feature, it has a synergistic effect on the patient's disease course, injury type, disability characteristics and other core portrait information. Therefore, it is included in the portrait system as a potential influencing variable and participates in the analysis.

[0091] The attributes of the basic portrait include at least:

[0092] Natural attributes: Gender attributes are widely used labels. People of different genders have different preferences for different content. Moreover, it is easier to analyze the basic proportion of the patient group through natural attribute labels such as age, region, education, occupation, marital status, and children's status.

[0093] Vertical attributes: reflect the patient's exercise needs, such as different types of exercise preferences;

[0094] Training attributes: Training attributes are also a relatively important attribute category, which helps to determine the patient's exercise intention, exercise cycle, and exercise frequency;

[0095] Label attributes: When a patient starts using the system and generates the first piece of data, the system can assign him the first label - newcomer; afterwards, as the number of patients accumulates, low-frequency patients, active patients, and high-frequency patients can be gradually divided.

[0096] The basic attribute features are quantified and calibrated to form a basic attribute portrait set.

[0097] Through the previous steps, a core portrait set and a basic portrait set are constructed for each patient. Next, a role portrait needs to be constructed based on the core portrait set and the basic portrait set.

[0098] The core portrait and the attribute portrait are in a parallel relationship. The core portrait is the core feature of the individual, and the attribute portrait is the basic feature of the individual. The role portrait uses the basic attribute portrait as a one-dimensional structure and the core portrait as a two-dimensional structure. As time goes by, the characteristics of the core portrait at different time points (different stages of disease development) are a three-dimensional structure. In this step, the portrait set is superimposed (union) with the weight difference added to form the individual's role portrait.

[0099] The portrait sets constructed based on different features are merged according to predetermined rules. For example, different weights are given to different features according to their importance, and then weighted superposition is performed to merge the features in the two portrait sets. Here, the weight of the features in the core portrait is greater than the weight of the features in the basic portrait. Specifically, factors that need to be considered in weight allocation include the feature coverage of each portrait set, the consistency of the data, the rationality of the weight allocation, etc., to ensure that the superimposed portraits can accurately reflect the comprehensive characteristics of different patient groups and will not lose balance due to over-emphasis on certain features.

[0100] The basic portrait set is used as the coarse-grained feature of the individual, and the core portrait is used as the fine-grained feature of the individual, and they are merged and superimposed. By merging features of different granularities, the characteristic information of the user can be captured more comprehensively.

[0101] Coarse-grained features refer to broad, general features that provide general information about an individual or object. These features include some basic attributes, such as age, gender, occupation, etc., which are helpful for understanding the basic profile of an individual or patient. This feature is mainly used in scenarios where the overall profile of an individual is quickly located, and as a basic data classification feature label. These features help with preliminary screening and classification, and provide direction for subsequent in-depth analysis.

[0102] Fine-grained features refer to more specific and detailed features that provide deeper information about individuals or objects. These features include some specific behaviors, interests, habits, etc., which are very important for understanding the uniqueness and personalized needs of individuals or objects. Fine-grained features are suitable for fields that require a deep understanding of the uniqueness of individuals or objects, such as personalized recommendation systems, precision training, and behavioral feedback analysis. These features can capture subtle differences and provide more accurate calculation and push support.

[0103] Therefore, coarse-grained features provide general information about the user, while fine-grained features provide more specific details. This combination of coarse and fine feature representation can improve the accuracy of portrait mapping; the fine-grained feature selection module can extract key local and fine-grained discriminative features in the object, while the coarse-grained feature selection module can obtain coarse-grained diversity features that provide contextual information and improve the model's discriminative power; by combining features of different granularities, the model can better generalize to different scenarios and conditions.

[0104] Specifically, the basic image set (coarse-grained features) and the core image set (fine-grained features) are combined and superimposed. Different weights need to be assigned to different features in this process. When different weights are given according to the importance of different features, the following conditions are met: 1) The feature weight in the core image is greater than the feature weight in the basic image; 2) The feature coverage, data consistency, and rationality of weight allocation of each feature in the core image; 3) The feature coverage, data consistency, and rationality of weight allocation of each feature in the basic image.

[0105] Quantify and calibrate the core attribute features one by one to form a core attribute portrait set.

[0106] The character portrait is a three-dimensional mapping system that integrates the basic portrait features and the core portrait features. The attributes of the core portrait include at least cognitive cortical damage, physical function damage, and emotional control damage. When obtaining the character portrait, it is necessary to construct a three-dimensional time slice based on the basic portrait and the core portrait. The three-dimensional time slice is obtained by the following steps:

[0107] First, the stages are divided. According to the progression of different types of elderly disability diseases, the course of the disease is divided into several stages, including early, middle and late stages. Each major stage can also be further divided into more specific sub-stages to understand the development of the disease in more detail.

[0108] The next step is feature extraction. In each stage, key features of the patient are extracted, including cognitive function, emotional state, and motor ability. These features can be measured by standardized assessment tools, such as the Mini Mental State Examination (MMSE), the Montreal Cognitive Assessment (MoCA), and the Berg Balance Scale.

[0109] The last step is to construct a timeline. A timeline starting from the discovery of the disease (e.g., every month, every 2 months, every year, etc.) is used to represent the development of the disease over time. On the timeline (as a time dimension), the starting and ending points of each stage, as well as the key features of each stage, can be marked. The timeline can be linear or nonlinear to reflect the complexity and uncertainty of the disease course. Such a timeline helps to understand the progression of the disease more intuitively and provide guidance for treatment and intervention.

[0110] According to industry consensus, quantify the time series dimension features in the role portrait and perform portrait fusion. Incorporate knowledge graph information such as clinician experience, industry diagnostic standards, and expert consensus in the field, use the stage of disease progression as the time dimension to divide nodes, and associate different disease progression stages, quantified clinical characteristics, and other dimensional information in the role portrait. Taking AD-induced dementia as an example, according to industry consensus, the disease progression can be divided into early (asymptomatic cerebral amyloid stage: SCD, MCI stage), middle (amyloid positive + synaptic dysfunction and (or) neurodegeneration stage), and late (amyloid positive + evidence of neurodegeneration + severe cognitive decline). Different stages of disease progression have different corresponding exercise training plans. The content of the exercise training plan is as follows: Figure 2 As shown, it includes cognitive function rehabilitation, motor function rehabilitation, mental and behavioral symptom rehabilitation, activity and participation rehabilitation, comprehensive rehabilitation, etc.

[0111] The different stages of disease development, quantified clinical characteristics, and other dimensional information in the role portrait are associated as multi-dimensional portrait feature input for the pushed sports training plan, preparing for the push of precise sports training.

[0112] Step 2: Generate and push the first sports training plan based on the role portrait.

[0113] Based on the role profile, the first exercise training plan is automatically generated. The first exercise training plan can be generated based on a pre-trained model (such as the collaborative filtering model mentioned later) or designed based on the doctor's experience.

[0114] Step 3: Collect structured data, text data, and image data during sports training.

[0115] Collect the structured data, text data, and image data of the patient exercising according to the first exercise training program. Use the structured data, text data, and image data (video data or image data) as independent variable data for the input collaborative filtering model. Specifically, it includes but is not limited to:

[0116] ■Structured data: sports diagnosis, life ability assessment scales, demographic characteristics data, etc.;

[0117] Text data: clinical diagnosis and case report data from the past two years (including patient movement diagnosis data obtained through medical records, test reports, etc., such as upper limb motor function impairment);

[0118] ■ Image data: MRI data, CT imaging data, etc.

[0119] Step 4: Input the data in step 3 into the collaborative filtering model and combine it with the role portrait to generate and push the next sports training plan.

[0120] In this embodiment, step 4 can be repeated until the desired effect is achieved or the patient stops training.

[0121] On the one hand, multiple autoencoders and collaborative filtering networks are used in the collaborative filtering model, and graph embedding technology is used to map entities (such as patients, exercise types, physiological indicators, etc.) and relationships in the graph structure into a low-dimensional vector space. This step is achieved through deep learning embedding technologies such as the Bayesian TransR embedding model, Bayesian stacked denoising autoencoder (SDAE), and Bayesian stacked convolutional autoencoder (SCAE), which convert structured data, text data, and image data into structure vectors, text vectors, and image vectors, respectively.

[0122] Specifically, the autoencoder is a layer-by-layer unsupervised learning model that mainly includes two processes, decoding and encoding, and is used to process high-dimensional data. As a tool for feature extraction, it enables the collaborative filtering model to learn low-dimensional representations (latent vectors) of users and items. These latent vectors can capture the intrinsic connection between users and items, thereby improving the performance of the recommendation system. Here, the autoencoder fuses three types of information (structured data, text data, and image data) to achieve fusion analysis of multimodal information and obtain item latent vectors. Moreover, based on the role portrait data, the autoencoder uses feature extraction to obtain user latent vectors.

[0123] On the other hand, in order to fully learn the auxiliary information features, the collaborative filtering model uses the Bayesian stacked denoising autoencoder (SDAE) to learn the vector representation of text information, the Bayesian stacked convolutional autoencoder (SCAE) to learn the vector representation of image information, and the Bayesian TransR embedding model to learn the vector representation of structural information. SDAE is a deep learning model that extracts high-level features of data through multi-layer unsupervised learning. It can capture the semantic and structural information in the text, which is particularly useful for processing the noise and uncertainty of text data. SCAE is used to learn the vector representation of image information because the convolutional neural network (CNN) has advantages in processing image data and can capture the spatial hierarchy and local features of the image. By stacking multiple convolutional layers and pooling layers, SCAE can learn more abstract and advanced image features, which helps the model make more robust predictions when faced with the diversity and complexity of image data.

[0124] TransR is an embedding model designed specifically for knowledge graphs that can effectively handle complex relationships in knowledge graphs. It achieves more accurate knowledge graph completion by mapping entities and relationships to different semantic spaces respectively, and using relationship-specific transformation matrices to project entities from entity space to relationship space, thereby performing vector operations in the projected space. This design not only captures the multidimensional relationships between entities, but also effectively handles incomplete information and noise in knowledge graphs. In addition, the flexibility and scalability of TransR enable it to adapt to different structured data, improving the accuracy and robustness when processing structured data. By converting text, images, and structured data into a unified vector representation, TransR ensures that data from different sources are comparable in the feature space, thereby providing strong support for the efficient representation and application of knowledge graphs.

[0125] After feature extraction, graph embedding technology is conducive to constructing individual data space, integrating multi-dimensional data such as time, sports preference, physiological and psychological characteristics, and post-training status to form a four-dimensional data structure. Furthermore, it integrates structure vectors, text vectors, and image vectors to form item latent vectors. Item latent vectors integrate the features of multiple data types and can more comprehensively represent the features of items.

[0126] Specifically, if Figure 3A As shown, the collaborative filtering model is used to generate and push the next sports training plan, including the following sub-steps:

[0127] 1) Determine the key variables as one-dimensional data, two-dimensional data, and three-dimensional data

[0128] First, it is necessary to determine which variables are important for assessing and predicting the patient's recovery process. By combining cognitive tasks (such as memory and attention tasks) and motor tasks (such as walking and balance training), the patient's motor function and balance ability can be improved. Through exercise, the patient's brain blood flow and blood oxygen level are improved, dormant neurons are activated, neuronal growth and connection are promoted, specific nerve conduction patterns are formed, and neurotransmitter levels are regulated; while the improvement of cognitive function helps to improve the patient's motor skills and efficiency, enhance motivation and self-control (the ability to control emotions, speech, etc.), and optimize exercise strategies.

[0129] In summary, different key independent variables need to be determined for patients with different diseases (different damaged brain areas).

[0130] Key independent variables that were important and common to all patients with cognitive disorders included:

[0131] Time (T)

[0132] Type of Exercise

[0133] Intensity

[0134] Duration

[0135] Frequency

[0136] Patient Status

[0137] Physiological Indicators (such as heart rate, blood pressure, etc.)

[0138] Psychological Indicators (such as anxiety level, depression level, etc.)

[0139] Recovery Progress

[0140] Different key variables are determined for patients with different diseases, including:

[0141] AD patients:

[0142] The MRI ROI indicators (Region of Interest) for brain injury were damage to the hippocampus, medial temporal lobe, parietal lobe, and frontal lobe;

[0143] Cognitive Ability Scores include memory loss, spatial disorientation, inattention, and executive dysfunction.

[0144] Patients with PD:

[0145] The MRI ROI indicators (Region of Interest) for brain damage were damage to the substantia nigra, striatum, prefrontal cortex, and parietal lobe;

[0146] Activity of Daily Living Indicators: movement disorders, cognitive decline, and sensory abnormalities;

[0147] Hormone Level Score, dopamine level, etc.

[0148] 2) Create a data matrix based on key variables

[0149] Based on the key variables of each patient, create a data matrix where each row represents a data record at a time point and each column represents a variable. For example:

[0150] One-dimensional data: [time (T)]

[0151] Two-dimensional data (exercise preference): [Type of Exercise

[0152] Intensity

[0153] Duration

[0154] Frequency

[0155] Three-dimensional data (physiological and psychological characteristics): [Heart Rate

[0156] Blood Pressure

[0157] Anxiety Level

[0158] Depression Level]

[0159]

[0160] 3) Add hysteresis features as four-dimensional data

[0161] To capture dependencies in time series, lagged features can be added.

[0162] Four dimensions: [Patient Status after training, Recovery Progress], for example:

[0163] time Patient status after training Rehabilitation Progress T1 (e.g. 10:00 am on October 3, 2024) good 5% T2 (e.g. October 3, 2024 at 3:00 p.m.) good 5% T3 (e.g. October 4, 2024 at 10:00 AM) good 5% T4 (e.g. October 4, 2024 at 3:00 p.m.) improve 10% ... ... ...

[0164] Based on the aforementioned one-dimensional data, two-dimensional data, three-dimensional data, and four-dimensional data, a four-dimensional data structure is formed. That is, the individual data space is constructed with one-dimensional data (time) as the W axis, two-dimensional data as the Y axis, three-dimensional data as the X axis, and four-dimensional data as the Z axis.

[0165] 4) Use embedding technology to convert it into a feature vector.

[0166] The hidden layer of the deep learning model is used for internal processing and fusion, and the four-dimensional data is converted into feature vector representation through embedding technology.

[0167] like FIG. 3A to FIG. 3C As shown in Figure 1, the structured data in the aforementioned individual data space of each patient is converted into a structured vector through structured embedding. For example, using the Bayesian TransR knowledge graph embedding method, entities and relations can be converted into vectors, which can capture the complex relationships between entities.

[0168] Through text embedding, text data is converted into text vectors. For example, using the Bayesian Stacked Denoising Autoencoder (SDAE), we can learn the low-dimensional representation of text data and capture the semantic information of the text.

[0169] Through image embedding, image-like data is converted into image vectors. For example, the Bayesian stacked convolutional autoencoder (SCAE) can be used to extract image features and convert them into vectors.

[0170] Using the methods provided by TransR, SDAE, and SCAE, the structure vector, text vector, and image vector are integrated into a project latent vector. This project latent vector is a vector that integrates the features of multiple data types and can more comprehensively represent the features of different types of patients. It is a vector that can represent project features.

[0171] 5) Based on the feature vector, the item latent vector and the user latent vector are generated for collaborative filtering learning to recommend sports training plans.

[0172] The attribute fusion graph convolutional network AF-GCN is used to fuse feature vectors from different modal data, and the attention mechanism is used to weight the importance of different modal data to gradually quantize the feature vectors layer by layer.

[0173] Thus, this embodiment integrates data from different sources (such as text, images, etc.), with one-dimensional data as the W axis; two-dimensional data as the Y axis, three-dimensional data as the X axis, and four-dimensional data as the Z axis to construct an individual data space; the one-dimensional data includes a time variable, the two-dimensional data includes multiple variables describing exercise preferences; the three-dimensional data includes variables describing physical and psychological characteristics; the four-dimensional data includes post-training state data. Moreover, the deep learning model is used to extract the features of each data type, so that these features are combined with knowledge representation to form a unified feature space, so as to better understand the interests and needs of patients and realize multimodal information feature-level fusion (fusion after feature extraction and before decision-making). Moreover, the deep learning model (such as SDAE) is used to convert the information in the role portrait of the patient (user) (behavior, preference, rating of different training programs, etc. in the portrait) into a user latent vector, which is a vector that can represent user characteristics, and is input into the collaborative filtering network together with the project latent vector to calculate the user's preference for different projects, and generate the final recommendation plan (exercise training plan) based on this.

[0174] This embodiment adopts a collaborative filtering model that combines graph structure and deep learning technology, and uses graph structure to represent the knowledge base, which includes entities (such as sports categories, user attributes, user behaviors) and their relationships. Therefore, the graph structure can be used to capture the interaction data between users and projects, and learn user preferences and project characteristics.

[0175] The multimodal features (item latent vectors) and role portrait features (user latent vectors) extracted for the patient's condition are used as independent variable data, and the first exercise training plan in the previous step (for example, the exercise prescription prescribed by the doctor in the clinic) is used as the dependent variable, which is input into the collaborative filtering model to generate the next exercise training plan. The loss function of the collaborative filtering model is the cosine contrast loss, which is used to maximize the similarity between positive sample pairs and minimize the similarity of negative sample pairs under margin constraints.

[0176] As mentioned above, through graph embedding technology, each entity and relationship in the graph structure of the individual data space is mapped into a low-dimensional vector space, which allows them to be calculated and compared, so that the collaborative filtering model can analyze the interaction data between users and items, capture the complex nonlinear relationship between patients and items, learn user preferences and item characteristics, and thus predict items that users may be interested in.

[0177] The collaborative filtering model itself has a cold start problem. In other words, for new patients, due to the lack of their historical behavior data, it is impossible to calculate the similarity between patients according to the conventional collaborative filtering model, and it is difficult to make personalized recommendations for them. This results in that when the collaborative filtering model just enters the hospital and starts to be used, the model is difficult to provide users with accurate personalized recommendations. However, for structured data, image data, text data, etc., the collaborative filtering model can predict the preferences of new users or new projects by analyzing the behaviors of other similar patients or similar projects. Only a small amount of exercise training program data with clinical authoritative diagnosis is needed to achieve the recommendation of rapid exercise training programs for people with similar characteristics (cold start function), without a large amount of training. This is because in the collaborative filtering model provided by the embodiment of the present invention, based on the role portrait data and multimodal data of new patients, it is easy to find similar patients and similar projects.

[0178] That is, the collaborative filtering model provided by the embodiment of the present invention takes advantage of the fusion of role portraits and multimodal data. For example, for patients with cognitive cortical damage, abnormal signals in specific brain areas will be shown in their MRI images, and similar cognitive dysfunction manifestations will also be mentioned in clinical diagnosis. In the multimodal data corresponding to cognitive cortical damage and physical function damage, the feature similarity is high, and the model can be used to obtain the commonalities and regularities of data from different patients in key features, thereby improving the accuracy of feature similarity judgment. For another example, although two patients differ in emotional control damage, if they are highly similar in cognitive cortical damage and physical function damage, then when recommending exercise training programs, these two patients are more likely to be classified as similar patients, and then recommend suitable exercise training programs for new patients. By using fine-grained core portraits, the model can accurately grasp the key needs of new patients and find similar patients and projects.

[0179] Based on the rich information in the role portrait, matching and screening are carried out from multiple levels to find patient groups that are similar to new patients in multiple dimensions, and then recommend exercise training programs that are more in line with their comprehensive characteristics to new patients, thereby improving the personalization of recommendations. For example, for patients with more severe cognitive cortical damage, more emphasis is placed on cognitive training and exercise training programs that activate brain functions; while for patients with prominent physical function damage, more emphasis is placed on training to restore and enhance limb motor functions. Compared with traditional collaborative filtering models that rely only on historical behavioral data, this personalized recommendation method can better solve the cold start problem and provide new patients with accurate and effective exercise training program recommendations to help them improve their cognitive, emotional, motor, and speech abilities.

[0180] In addition, based on basic portraits such as medical history and demographic characteristics and core portraits such as examination results and doctor's advice, the patient status sequence data is generated, and in the process of fusing different modal data such as imaging data (image data) and scale data (text data), the use of collaborative filtering models can more fully grasp the relationship between different media modalities through multi-layer attribute fusion, thereby improving the accuracy and generalization ability of exercise training program recommendations.

[0181] Second embodiment

[0182] On the basis of the first embodiment, in this embodiment, the recurrent neural network model (ie, the RRN model) is combined with the collaborative filtering model in the first embodiment to further improve the consistency and sustainability of the exercise training program recommendations and enhance patient compliance.

[0183] like Figure 4 The sports training recommendation method based on multimodal data fusion provided in the second embodiment of the present invention at least comprises the following steps:

[0184] Steps 1 to 4 are the same as those in the first embodiment (but step 4 does not need to be repeated multiple times), and are not described in detail here.

[0185] Step 5: Collect structured data, text data, and image data during the sports training multiple times in the manner of step 3 to obtain sports training cycle data.

[0186] The patient exercises multiple times according to the exercise training plan, and according to the method of step three, the data (structured data, text data, image data) from the multiple exercises are accumulated to obtain the exercise training cycle data.

[0187] The training cycle data in the sports training plan is generally divided into: multi-year cycle data, large cycle data, medium cycle data and small cycle data.

[0188] Multi-year cycle: Due to the large individual differences among the elderly, specific issues need to be analyzed specifically. For example, (1) the physical condition and recovery ability of each elderly person are different, so the exercise rehabilitation cycle will vary from person to person. Generally speaking, elderly people with better physical condition and less disability may recover faster, with a cycle of 2 years. (2) The more severe the disability, the longer it usually takes to recover. The cycle for elderly people who are completely bedridden is 8 years, and the cycle for elderly people who can partially take care of themselves is 4 years. During the rehabilitation process, it is necessary to continuously monitor the physical condition and rehabilitation progress of the elderly and make adjustments as needed. This helps to ensure the effectiveness of the rehabilitation plan and shorten the rehabilitation cycle as much as possible.

[0189] The macrocycle is usually based on the current year or one year as the time limit, and a half-year or full-year training plan is formulated based on it. Sometimes a year is divided into three macrocycles, which is generally related to the classification of sports.

[0190] In training practice, the medium cycle is called stage and monthly training, which usually lasts for 4-8 weeks, and is used to formulate stage or monthly training plans.

[0191] In training practice, micro-cycles are called weekly training, usually with a calendar week as the deadline to formulate a weekly training plan. You can also arrange a micro-cycle of 4 to 10 days. If you want to make a small recovery adjustment, you can complete the adjustment task in 4 days, so you can arrange a 4-day recovery micro-cycle.

[0192] Step 6: Determine whether the exercise training program is the Nth one. If yes, proceed to step 7; if not, return to step 4 and generate the next exercise training program according to the features in step 5.

[0193] N is a preset number of times, for example, 3 times. Through this step, the dynamic characteristics of the patient when training according to multiple exercise training programs can be obtained, and the characteristics are of a temporal dimension. Such a design is conducive to the RRN model to learn more complex dynamic evolution representations.

[0194] Step 7: Based on the data collected from the patient in steps 3 and 5, the data are jointly input into the RRN model to iteratively optimize the exercise training program.

[0195] In this step, use Figure 5 The RRN model shown can better learn more complex dynamic evolution representations. In the RRN model, the historical interaction information between patients and movements is the key data that drives changes in patient preferences and movement states, so the use of a co-evolution model can capture the evolutionary implicit representations of patients and movements. Here, using the training cycle vector input to the RRN model, the RRN model uses a recurrent neural network to learn the dynamic feature representations of patients and movements, and obtains a dynamic evolution representation with a temporal relationship. The training cycle vector refers to the vector obtained by converting each attribute (such as time, frequency, intensity, etc.) in the training cycle data into numerical features and then converting it through embedding or encoding technology.

[0196] like Figure 5 As shown, the RRN model uses two long short-term memory networks (LSTM) based on the traditional recurrent neural network (RNN) to learn the temporal changes of patients' exercise preferences and the long-term (e.g., seasonal) evolution of exercise respectively. Figure 5 In the example, y is learned through an LSTM network. i , which represents the dynamic feature vector of patient i in the time series (time t, t+1, etc.); y is obtained by learning through another LSTM network j, which represents the dynamic feature vector of motion j in time series. Moreover, y i and j Both are affected by the patient's static characteristics u i and the static characteristics of the motion m j This is to take into account the patient's long-term motion preference and the static properties of the motion, so that the RRN model can simultaneously learn the patient's static latent representation and the static latent representation of the motion.

[0197] In other words, at each time step, the RRN model uses an LSTM network to combine the patient's previous time step dynamic feature vector and the current static feature vector to update the dynamic feature vector; another LSTM network is used to combine the dynamic feature vector of the previous time step of the movement and the current static feature vector to update the dynamic feature vector. This RRN model that combines static and dynamic features can simultaneously consider the patient's long-term exercise preferences and short-term changes, as well as the long-term characteristics and short-term changes of the exercise training program.

[0198] Specifically, the RRN model uses a LSTM-based recurrent neural network to model the dynamic changes of patients and exercise training programs. For patient i and exercise j (a sports training program includes multiple exercises), assuming u i represents the static feature vector of patient i, m j represents the static eigenvector of motion j. At time t, assuming u it is the dynamic feature vector of patient i (i.e., y i,t ), m jt is the dynamic eigenvector of motion j (i.e., y j,t ). The dynamic characteristics u at time t+1 i,t+1 and m j,t+1 The solutions can be serialized through an LSTM network as follows:

[0199] u i,t+1 =g(u it ,{r ij|t})

[0200] m j,t+1 =h(m jt ,{r ij|t})

[0201] Among them, g and h are functions that need to be learned. g represents the function used in the LSTM network to update the dynamic characteristics of the patient; h represents the function used in the LSTM network to update the dynamic characteristics of the exercise training program; r ij|t represents the score of patient i for motion j at time t+1. The dynamic feature u of the patient at time t+1 i,t+1 is the dynamic feature u of the patient at time t through the LSTM networki, t and response r ij|t Calculated. The dynamic characteristics of the motion at time t+1 m j,t+1 is the dynamic feature m of the movement at time t through the LSTM network j,t and the response r ij|t This shows that the dynamic features of the patient and the exercise training regimen are updated by the LSTM network based on their state at the previous time point and the current response. Therefore, the RRN model is able to capture dynamic features that change over time, thereby better simulating and predicting the patient's response to different exercise training regimens.

[0202] Input patient i’s dynamic and static features (u i,t+1 , m j,t+1 ,u it , m jt ), the model can predict the patient’s current interest (score), r ij The predicted value of Among them, f is also a function that needs to be learned.

[0203] The loss function of the RRN model is the sum of the squared error loss function and the regularization term. Adjusting the parameter θ to minimize the loss function can achieve overall optimization of the RRN model. That is, the following function is used for optimization:

[0204]

[0205] It can be seen that the goal of RRN model optimization is to make the exercise training program pushed by the model more and more consistent with the patient's current exercise preference, exercise level, and exercise status, that is, to make the prediction close to the actual through the parameters generated by training. Among them, θ represents the parameters to be learned, is the set of (patients, exercise training regimens, time series) tuples observed in the training set, and R represents the regularization function.

[0206] While the objective function and building blocks in the RRN model are quite standard, a simple application of backpropagation cannot easily solve this optimization problem. The key challenge is that each patient score depends on both the patient state and the motor training regimen. Backpropagation through the 2 sequences is computationally prohibitive. This problem is mitigated by backpropagating gradients from the patient's motor feedback, however, each score still depends on the patient state, which in turn acts on the complete sequence pushed by the motor training regimen.

[0207] Therefore, in the embodiment of the present invention, an alternating subspace descent strategy is adopted. That is, it is assumed that the motion state is fixed, so there is no need to propagate the gradient to these motion training program sequences, and the gradients of all patient scores are back-propagated to update the patient sequence parameters, and then switch between updating the user sequence and updating the motion training program sequence. In this way, only one standard feedforward and back-propagation can be performed for each patient's motion training program push, and finally the model optimization is achieved.

[0208] In summary, the embodiment of the present invention uses a hybrid recommendation model (collaborative filtering model and RRN model) as the core tool for sports training recommendation. Compared with the traditional content recommendation model, the collaborative filtering model in the hybrid recommendation model can simultaneously learn the representation of multi-source data (structured data, text data, image data, training cycle data), using structural information, content information, and image information. Multi-dimensional information. At the same time, the recurrent neural network in the hybrid recommendation model is used to capture the patient's long-term preference evolution representation and short-term preference representation. Therefore, by integrating different models, it is possible to achieve effective evaluation of individual exercise dosage, exercise status, and exercise mode of elderly disabled patients, and ultimately achieve the push of the best exercise training program.

[0209] Each time the patient trains, the data collected during the previous exercise training is used to output a new exercise training program in the manner of step 7. This is repeated until the predetermined goal is achieved, such as the recovery or maintenance of cognitive, emotional, motor, and speech abilities.

[0210] As an optional step, step 8 can be added after step 7: evaluate the training effect based on the data in the last exercise training program, and if the effect is improved, make and output the exercise training report or directly end it; if the effect is not improved, return to step 7 to regenerate the exercise training program. The method for evaluating the training effect can adopt conventional evaluation methods, including evaluation by a doctor, or compare and evaluate the data in the last exercise training program with the data in the previous exercise training program.

[0211] Compared with the prior art, the present invention has the following technical advantages:

[0212] 1) Construct a multi-level patient portrait system to gradually and deeply deconstruct and evaluate the patient's attribute characteristics, behavioral characteristics, and comprehensive characteristics, ensuring that the individual label system has a detailed granularity, providing a large amount of effective information on the patient sequence for the subsequent input model to make accurate recommendations for exercise;

[0213] 2) An exercise recommendation framework that integrates the algorithmic advantages of multiple models. On the one hand, it uses the fusion analysis system of multimodal data in collaborative filtering and autoencoding learning to ensure that multi-source information enters the recommendation model and increase the source of analysis; on the other hand, it uses the effective calculation model of recurrent networks for time series data to effectively trace back the patient's exercise data at different time periods, thereby ensuring the consistency and sustainability of exercise recommendations and improving patient compliance;

[0214] 3) Construct a training cycle data system, divide the cycle of the exercise training process into annual cycle, large cycle, medium cycle and small cycle, and create different time series data sets. When introducing the model for exercise recommendation calculation, the time series dimension feature is added, so as to more fully and regularly apply historical exercise data to assist in the generation of periodic plans for exercise recommendation;

[0215] 4) Build an exercise push system based on deep learning, and build the multi-dimensional exercise training push logic into a system process, which can realize the continuous iteration of exercise push content based on the patient's own multi-dimensional exercise data, so that the training can best suit the individual patient and the current state, thereby improving the training effect.

[0216] Third embodiment

[0217] Based on the above-mentioned sports training recommendation method based on multimodal data fusion, the third embodiment of the present invention further provides a sports training recommendation system based on multimodal data fusion. Figure 6 As shown, the sports training recommendation system includes a processor and a memory. The memory is coupled to the processor and is used to store one or more programs. When the programs are executed by the processor, the processor implements the sports training recommendation method based on multimodal data fusion as in the above embodiment.

[0218] Wherein, the processor is used to control the overall operation of the sports training recommendation system to complete all or part of the steps of the sports training recommendation method based on multimodal data fusion. The processor can be a central processing unit (CPU), a graphics processing unit (GPU), a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a digital signal processing (DSP) chip, etc. The memory is used to store various types of data to support the operation of the sports training recommendation system, and these data may include instructions for any application or method for operating on the sports training recommendation system, and application-related data. The memory can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as a static random access memory (SRAM), an electrically erasable programmable read-only memory (EEPROM), an erasable programmable read-only memory (EPROM), a programmable read-only memory (PROM), a read-only memory (ROM), a magnetic memory, a flash memory, etc.

[0219] In another exemplary embodiment, the present invention further provides a computer-readable storage medium including program instructions, which, when executed by a processor, implements the steps of the sports training recommendation method based on multimodal data fusion in any of the above embodiments. For example, the computer-readable storage medium may be the above-mentioned memory including program instructions, and the above-mentioned program instructions may be executed by a processor of the system to complete the above-mentioned sports training recommendation method based on multimodal data fusion, and achieve the same technical effect as the above-mentioned method.

[0220] It should be noted that the above embodiments are only examples, and the technical solutions of the various embodiments can be combined, and the order of the steps can be changed, all within the protection scope of the present invention.

[0221] The above is a detailed description of the sports training recommendation method and system based on multimodal data fusion provided by the present invention. For those skilled in the art, any obvious changes made to it without departing from the essence of the present invention will constitute an infringement of the patent right of the present invention and will bear corresponding legal responsibilities.

Claims

1. A sports training recommendation method based on multimodal data fusion, characterized in that The steps include: Step 1: For elderly disabled patients, obtain the disease type and clinical characteristics, demographic information, exercise survey information and exercise behavior characteristics of the patient, so as to integrate and construct a role portrait and obtain static characteristics; Step 2: Generate and push the first sports training plan based on the role portrait; Step 3: Collect structured data, text data, and image data during sports training; Step 4: Input the data in step 3 into the collaborative filtering model, combined with the role portrait, to generate and push the next sports training plan; wherein the collaborative filtering model uses multiple autoencoders and collaborative filtering networks, and uses graph embedding technology to map entities and relationships in the graph structure into a low-dimensional vector space, and is implemented through a Bayesian TransR embedding model, a Bayesian stack denoising autoencoder, and a Bayesian stack convolutional autoencoder to convert structured data, text data, and image data into structure vectors, text vectors, and image vectors, respectively, thereby obtaining project latent vectors and user latent vectors for collaborative filtering learning to recommend the next sports training plan.

2. The sports training recommendation method based on multimodal data fusion according to claim 1, characterized in that The step 4 includes the following sub-steps: 1) Determine the key variables as one-dimensional data, two-dimensional data, and three-dimensional data; 2) Create a data matrix based on key variables; 3) Add hysteresis features as four-dimensional data; 4) Convert it into feature vector using embedding technology; 5) Based on the feature vector, the item latent vector and the user latent vector are generated for collaborative filtering learning to recommend sports training plans.

3. The sports training recommendation method based on multimodal data fusion according to claim 2, characterized in that: In the sub-step 4, the entities and relationships are converted into vectors using the Bayesian TransR knowledge graph embedding method; the text data is converted into text vectors using the Bayesian sparse autoencoder embedding method; and the Bayesian sparse convolutional autoencoder is used to extract image features and convert them into image vectors.

4. The sports training recommendation method based on multimodal data fusion according to claim 3, characterized in that: The multimodal features and role-portrait features extracted from the patient's condition are used as independent variable data, and the exercise prescription recommendations prescribed by the clinic are used as dependent variables, which are input into the collaborative filtering model; The loss function of the collaborative filtering model is cosine contrast loss, which is used to maximize the similarity between positive sample pairs and minimize the similarity of negative sample pairs under margin constraints.

5. The sports training recommendation method based on multimodal data fusion according to claim 4, characterized in that: In the sub-step 1, an individual data space is constructed with one-dimensional data as the W axis, two-dimensional data as the Y axis, three-dimensional data as the X axis, and four-dimensional data as the Z axis.

6. The sports training recommendation method based on multimodal data fusion according to claim 5, characterized in that: The one-dimensional data includes a time variable, the two-dimensional data includes a plurality of variables describing exercise preferences; the three-dimensional data includes variables describing physical and psychological characteristics; and the four-dimensional data includes post-training status data.

7. The sports training recommendation method based on multimodal data fusion according to claim 6, characterized in that: When creating a data matrix based on key variables, different diseases correspond to different key variables; the role portrait includes multiple attributes, including cognitive cortical damage, physical function damage, and emotional control damage attributes.

8. The sports training recommendation method based on multimodal data fusion as claimed in claim 6, characterized in that The following steps are also included: Step 5: Collect structured data, text data, and image data during the sports training multiple times in the manner of step 3 to obtain sports training cycle data; Step 6: Determine whether the exercise training program is the Nth one, if yes, proceed to step 7; if not, return to step 4, and generate the next exercise training program according to the features in step 5; wherein N is a positive integer; Step 7: Based on the patient's data collected in steps 3 and 5, the data are jointly input into the RRN model to iteratively optimize the exercise training program. The RRN model is optimized using the following function: Among them, θ represents the parameters to be learned; is the set of tuples observed in the training set, which includes patients, exercise training programs and time series; rij|t represents the score of patient i on exercise j at time t; is the predicted value of rij|t; R represents the regularization function.

9. The sports training recommendation method based on multimodal data fusion according to claim 8, characterized in that: The RRN model adopts an alternating subspace descent strategy, assuming that the motion state is fixed, does not propagate gradients to these motion training scheme sequences, and simultaneously back-propagates the gradients of all patient scores to update the patient sequence parameters, and then switches between updating the user sequence and updating the motion training scheme sequence.

10. A sports training recommendation system based on multimodal data fusion, characterized in that It includes a processor and a memory, wherein the memory is coupled to the processor and is used to store one or more programs. When the program is executed by the processor, the processor implements the sports training recommendation method based on multimodal data fusion as described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Portrayal construction method, exercise training recommendation method and exercise training recommendation system for disability risk population

    CN119864123A

  • Hyperspherical collaborative metric recommendation device and method based on pre-trained semantic model

    CN111651558A

  • Recommendation method based on self-attention mechanism

    CN113822742A

  • Amblyopia training scheme recommendation method fusing GMF and CDAE

    CN115019933A

  • Recommendation method based on knowledge graph and attention mechanism

    CN115374288A

Cited By

  • Multi-modal medical information intelligent integration and decision support system for acupuncture rehabilitation

    CN120766879A

  • Cross-modal fusion intelligent joint training method and system

    CN121237323A

  • Intelligent joint training method and system for cross-modal fusion

    CN121237323B