Multimodal behavior prediction model training system and method
By training a multimodal behavior prediction model system using multiple sensors and neural processors, the lack of reminder mechanisms in existing technologies has been solved, enabling highly accurate prediction and reminders of users' daily behaviors and improving memory support.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- REALTEK SEMICON CORP
- Filing Date
- 2024-12-04
- Publication Date
- 2026-06-05
AI Technical Summary
The lack of effective reminder mechanisms in existing technologies to remind users of important events that are easily forgotten in daily life, especially for the elderly, may lead to memory decline and an increased risk of dementia.
By training a multimodal behavior prediction model system, various sensor data are acquired using multiple sensors. Combined with a neural processor and a large language model, prediction models for both trusted and untrusted sensors are trained to establish a multimodal behavior prediction model that can identify and alert on critical events.
It achieves highly accurate prediction and reminders of users' daily behaviors, reduces the possibility of forgetting important events, and improves memory support, especially for the elderly.
Smart Images

Figure CN122153425A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a method for training a model by learning user behavior, and more particularly to a multimodal behavior prediction model training system and method that uses multiple models and multiple sensors to learn user behavior. Background Technology
[0002] Everyone records important events in their calendars or memos, but often forgets seemingly insignificant things, such as turning off the gas before leaving the house, locking the car, drinking water, and watering plants. Especially as society ages, it may be necessary to frequently remind the elderly to take their medication or engage in regular exercise. While these things may not cause significant inconvenience, over time they can lead to a lack of confidence in one's memory and may even increase the risk of dementia.
[0003] Taking exercise reminders as an example, existing technologies can achieve the purpose of reminders through technology. For instance, users can wear fitness trackers or smartwatches. These smart devices are equipped with various motion sensors, which, together with algorithms, detect the user's movements. This allows them to detect the number of times a certain posture is repeated and determine whether a set threshold has been exceeded. However, for daily tasks that are easily forgotten, existing technologies lack effective reminder mechanisms. Summary of the Invention
[0004] In order to provide effective solutions for reminding users of things to pay attention to in daily life, this invention proposes a multimodal behavior prediction model training system and method.
[0005] According to an embodiment of a multimodal behavior prediction model training system, the computing device includes a processor and a neural processor. The processor can acquire various types of sensing data generated by multiple sensors, including one or more untrusted sensors and at least one trusted sensor. The neural processor can apply corresponding multiple models to predict the user's behavior based on the various types of sensing data and identify key events. Thus, sensing data generated by at least one trusted sensor can be acquired for key events, thereby training one or more prediction models based on sensing data generated by one or more untrusted sensors at the same time for key events.
[0006] The trusted model running in the trusted sensor, which has a probability of accurately predicting the key event exceeding a threshold, can be used to train the prediction model running in the untrusted sensor until the probability of the prediction model predicting the key event exceeds the threshold.
[0007] Thus, according to an embodiment, the neural processor uses one or more trained prediction models and at least one trusted model to jointly build a multimodal behavior prediction model to predict the user's behavior based on any one or more of a variety of sensing data.
[0008] Furthermore, the various sensing data generated by the plurality of sensors include image sensing data obtained by at least one image acquisition device, sound sensing data obtained by at least one audio receiving device, and sensing data generated by at least one user device.
[0009] Furthermore, the image acquisition device and the audio receiving device are set in a scene to acquire images and sounds in the scene, and the user device is a wearable sensing device or mobile device worn on the user's body; wherein, the user's position is obtained by using a positioning circuit in the user device, and the user's movement and actions in the scene are obtained by using a motion sensing circuit in the user device.
[0010] Furthermore, the various sensor data are mainly used to detect the user's location, movement, and actions. In addition, a trained prediction model and a trusted model are used to predict the user's behavior based on the various sensor data and identify key events that are repetitive or periodic, thereby establishing a reminder calendar.
[0011] Wherein, after obtaining the one or more prediction models, the computing device can deploy the one or more prediction models, or the multimodal behavior prediction model, to at least one user device running edge computing.
[0012] Subsequently, a multimodal behavior prediction model can be used to predict the user's behavior based on various sensing data generated by multiple sensors, and to determine key events. By comparing the key event with the reminder calendar, when the key event matches one of the reminder items set in the reminder calendar, a reminder in the form of text, images, or sound is generated through at least one user device.
[0013] Furthermore, the trusted model running in the trusted sensor can be a large language model, while the prediction model running in the untrusted sensor can be a trained language model. After the trusted sensor generates sensing data, the sensing data can be converted into data that the large language model can recognize through a coding program. By predicting the user's behavior and marking key events in the recognizable data, the trained language model can be trained based on the marked key events.
[0014] The trained language model can be a model formed by limiting the operation of a large language model through prompts, or an acquisition-enhanced generative model formed by limiting the prediction of specific user behaviors by a large language model.
[0015] To further understand the features and technical content of the present invention, please refer to the following detailed description and drawings of the present invention. However, the drawings provided are for reference and illustration only and are not intended to limit the present invention. Attached Figure Description
[0016] Figure 1 A scenario embodiment diagram of a multimodal behavior prediction model training system is shown;
[0017] Figure 2 A schematic diagram of a functional module embodiment of a multimodal behavior prediction model training system is shown;
[0018] Figure 3 A flowchart illustrating an embodiment of a multimodal behavior prediction model training method is shown.
[0019] Figure 4 The figure shows an embodiment of the training system and operation method for a multimodal behavior prediction model;
[0020] Figure 5 A schematic diagram illustrating an embodiment of an alert calendar built using a multimodal behavior prediction model is shown; and
[0021] Figure 6 An example diagram is shown illustrating the use of a multimodal behavior prediction model to execute alerts.
[0022] Explanation of reference numerals in the attached figures:
[0023] 10: Scene 101: Image Acquisition Device; 103: Audio Receiving Device
[0024] 11: First User; 12: Second User; 112: Wearable Sensing Device
[0025] 111: Mobile device; 201: Image acquisition unit; 202: Audio receiving unit
[0026] 203: Localization Unit; 205: Neural Processor; 207: Trained Language Model
[0027] 204: Sensing Unit; 206: Large Language Model; 208: Behavior Detection Unit
[0028] 401: Processor; 405: Large Language Model; 407: Untrusted Behavior Prediction
[0029] 403: Neural Processor; 411: First Sensor; 408: Trustworthy Behavior Prediction
[0030] 410: Comparator; 412: Second sensor; 409: Trained language model
[0031] 60: Processor; 50: Reminder Calendar; 61: Multimodal Behavior Prediction Model
[0032] 63: Reminder Unit; 65: Behavior Detection Unit; 67: Reminder Calendar
[0033] Steps S301-S319: Training process for multimodal behavior prediction model Detailed Implementation
[0034] The following specific embodiments illustrate the implementation of the present invention. Those skilled in the art can understand the advantages and effects of the present invention from the content disclosed in this specification. The present invention can be implemented or applied through other different specific embodiments, and various details in this specification can also be modified and changed based on different viewpoints and applications without departing from the concept of the present invention. Furthermore, the accompanying drawings of the present invention are for simple illustrative purposes only and are not depictions of actual dimensions; this is stated beforehand. The following embodiments will further describe the relevant technical content of the present invention in detail, but the disclosed content is not intended to limit the scope of protection of the present invention.
[0035] It should be understood that while terms such as "first," "second," and "third" may be used in this document to describe various components or signals, these components or signals should not be limited by these terms. These terms are primarily used to distinguish one component from another, or one signal from another. Furthermore, the term "or" as used herein should, as appropriate, include any combination of one or more of the associated listed items.
[0036] This invention proposes a multimodal behavior prediction model training system and method. The method uses a neural processor, also known as a neural network processing unit (NPU), to process various types of sensing data generated by multiple sensors using artificial intelligence and machine learning algorithms, and establishes a multimodal model by learning the features of the sensing data.
[0037] First refer to Figure 1 The diagram illustrates a scenario of a multimodal behavior prediction model training system. The diagram shows a scene 10 in which several people are present. This example demonstrates that scene 10 includes an image acquisition device 101 and an audio receiving device 103, and may also include other environmental sensors. The image acquisition device 101 captures continuous images of all persons (such as the first user 11 and the second user 12) within scene 10, while the audio receiving device 103, such as a microphone, records sounds generated within scene 10.
[0038] Furthermore, the first user 11 holds a mobile device 111, which can sense the user's location through its positioning circuit (such as GPS-related circuit), sense the first user 11's movements through its motion sensing circuit, and obtain sensing data such as images through a camera lens and sound through a microphone, thereby obtaining multimodal sensing data; while the second user 12 can wear a wearable sensing device 112, which can similarly sense the second user 12's location and movements, and can include other sensing data.
[0039] Thus, in the scenario 10 shown, diverse sensing data are obtained through at least one image acquisition device 101, at least one audio receiving device 103, at least one user device (such as mobile device 111, wearable sensing device 112), and various other sensors (such as location information obtained by positioning circuit, and user movement and action obtained by motion sensing circuit) to learn user behavior using a multimodal artificial intelligence model.
[0040] According to an embodiment of the multimodal behavior prediction model training system proposed in this invention, a large language model (LLM) runs in the neural processor. Multimodal sensing data generated by various sensors, such as images, sounds, location information, and time data derived for specific behaviors, is tokenized into data that the large language model can recognize. After manually labeling events within the data, the system learns user behavior, identifies repetitive or periodic behaviors, establishes corresponding multimodal behavior prediction models, and builds corresponding models and schedules. Thus, in addition to allowing users to check at any time whether an event has been completed via text, images, or sound, the system can also remind users of tasks to be performed at predetermined times.
[0041] The aforementioned periodic behaviors are those that users repeat over time, which can be used to establish periodic reminders, such as sleep reminders, wake-up reminders, reminders to take medication after meals, and scheduled exercise reminders. Repetitive behaviors, on the other hand, refer to user actions that do not have a periodic time pattern but are performed repeatedly. These also become the behaviors used in machine learning training models. For example, by using surveillance cameras installed at home to obtain sensor data such as images of users entering and leaving the house, the sound of door locks, and time, a multimodal behavior prediction model can be built. This model can then predict when a user is about to leave based on various sensor data, generating corresponding reminders through various mechanisms. For example, it can use a mobile phone to remind the user via voice, text, vibration, or sound to turn off the gas, lock the door, and remember to take their keys before leaving.
[0042] Furthermore, existing technologies have provided trained prediction models that can correctly identify objects, such as using a specific large-scale language model to correctly identify the door and lock in an image taken at home. However, they cannot effectively determine related behaviors based on other sensor data (such as sound or others). To address this, the multimodal behavior prediction model training system and method proposed in this invention provides a solution for training another trainable model using a model with high accuracy in recognizing specific behaviors. This achieves the purpose of multimodal behavior prediction and subsequent application of providing alerts.
[0043] like Figure 2 The diagram illustrates a functional module embodiment of a multimodal behavior prediction model training system. It describes various circuits and functional components implemented by the collaboration of software and hardware (processing circuits, memory, and storage devices, etc.) in a computer system. In this system embodiment, a large language model communication neural processor with high-precision recognition capability is used to train another trained language model. A multimodal behavior prediction model can be established by training the trained language model to derive one or more prediction models and at least one trustworthy model.
[0044] According to an embodiment, the multimodal behavior prediction model training system is implemented by a computing device, which includes a neural processor 205 that runs neural network models and machine learning calculations. The various types of sensors shown in the figure include an image acquisition unit 201, an audio receiving unit 202, a positioning unit 203, and other sensing units 204, including at least one trusted sensor and untrusted sensors. In a scene, these sensors generate sensing data, which is processed by the neural processor 205. The trusted sensors can run trusted models with a probability exceeding a threshold of accurate prediction for specific key events to train the sensing data generated by one or more untrusted sensors, establishing one or more prediction models running therein, until the probability of one or more prediction models predicting the key event exceeds the threshold. In this way, trusted sensors and untrusted sensors deploying trained untrusted prediction models can have the same or similar predictive capabilities.
[0045] In this example, the image acquisition unit 201 acquires images of a scene, and the generated image data is processed by the neural processor 205. The image data is then coded and converted into data that can be processed by a large language model 206, enabling accurate identification of user behavior within the scene and the determination of key events, such as specific actions performed by the user. Simultaneously, the audio receiving unit 202 generates sound data, and the positioning unit 203 generates positioning information, which can be represented by spatial coordinates (e.g., x, y, z and time). Additionally, other sensing units 204 can be used to generate sensing data within the same timeframe, such as mobile devices held by the user or wearable sensors.
[0046] Because the large language model 206 has a high confidence level in correctly identifying user behavior when processing image data, the sensing data generated by the image acquisition unit 201 is considered trustworthy sensing data. After processing by the neural processor 205, it can be used to train sensing data generated by other untrustworthy sensors. According to the illustrated embodiment, the behavior detection unit 208 labels key events in user behavior to form trustworthy sensing data. The neural processor 205 then trains the untrustworthy sensing data, resulting in a trained language model 207 operating on untrustworthy sensors. Continuous training ensures that the trained language model 207 can correctly identify user behavior and judge key events with the same confidence level as the large language model 206.
[0047] Subsequently, in the multimodal behavior prediction model training system, trained language models 207 with the same or similar accuracy can assist the large language model 206 in operation, establish a multimodal behavior prediction model with high discriminativeness, effectively identify user behavior, and generate reminders after judging key events.
[0048] It is worth mentioning that the trained language model 207 can be a new large-scale language model, or a language model augmented on top of a trusted large-scale language model. For example, it can be formed by using prompts built for a specific large-scale language model to create a large-scale language model with specific purposes and functions restricted by the prompts; or it can be a retrievalaugmented generation (RAG) model formed by limiting the prediction of specific user behaviors by a large-scale language model; in addition, a domain adaptation method, such as LoRA (Low-Rank Adaptation), can be used to train a local part of the large-scale language model to train a small-scale language model with additional weight values, which can be used to identify specific user behaviors.
[0049] Figure 3A flowchart illustrating an embodiment of a multimodal behavior prediction model training method is shown.
[0050] Initially, the system acquires various sensing data from a multimodal sensor consisting of multiple sensors located within a scene. This data may include environmental image and sound sensing data sensed simultaneously (step S301), as well as sensing data generated by various user devices simultaneously (step S303). The various sensing data generated by the multiple sensors include image sensing data acquired by at least one image acquisition device, sound sensing data acquired by at least one audio receiving device, and sensing data generated by at least one user device.
[0051] The environmental images and sounds can be obtained from Figure 1 (or Figure 2 The image acquisition device and audio receiving device in the system described herein, and the user device such as Figure 1 The wearable sensors shown are either worn on the user's body or mobile devices held by the user. The various types of sensors, which operate independently or are installed in a specific device, can generate corresponding sensing data at the same time.
[0052] Next, the various types of sensing data are tokenized and converted into codes (tokens) that can be recognized by the corresponding models (step S305). The system can then use multiple corresponding models to predict user behavior for each type of sensing data (step S307). For example, an image-based behavior recognition model can identify one or more of the user's position, movement, and actions based on image sensing data; an audio-based behavior recognition model can identify one or more of the user's position, movement, and actions based on sound sensing data; and a prediction model can identify one or more of the user's position, movement, and actions based on sensing data generated by the user's device.
[0053] In particular, the system's multimodal behavior prediction model can predict user behavior based on various sensor data generated by multiple sensors, including one or more untrusted sensors and at least one trusted sensor. Furthermore, the system has a corresponding intelligent model for the at least one trusted sensor, such as a Large Language Model (LLM). The trusted sensor is implemented as an edge computing device, which can run a trusted model with a probability exceeding a threshold of accurate prediction for key events. The trusted model can perform trusted behavior prediction based on the sensor data generated by the trusted sensor, pre-setting the user's behavior.
[0054] On the other hand, the system also provides one or more corresponding trained prediction models for one or more untrusted sensors. These prediction models can run on the untrusted sensors and predict user behavior based on the sensing data generated by the untrusted sensors. Thus, at least one key behavior can be determined from the user behavior predicted by various sensing data (step S309). According to an embodiment, the key behavior is determined from the behavior predicted by trusted sensing data generated by trusted sensors.
[0055] Subsequently, the system obtains trusted sensing data from the at least one trusted sensor (step S311). This trusted sensing data generated for the key behavior can be used to train other untrusted sensors to generate untrusted sensing data from the object at the same time, so as to train one or more prediction models running in one or more untrusted sensors, and continue training until the probability of one or more prediction models predicting the key event exceeds the threshold (step S313).
[0056] Through the above process, one or more prediction models that have been trained are obtained. Together with the at least one trustworthy model, a multimodal behavior prediction model can be established (step S315), which can predict the user's behavior based on any one or more types of sensing data generated by various types of sensors.
[0057] Next, user behavior can be detected through a multimodal behavior prediction model or any of the multiple models, and reminders can be created based on the detected repeatable or periodic behaviors (step S317), such as creating a reminder calendar.
[0058] According to one embodiment, once a multimodal behavior prediction model is derived, the system's computing device can deploy the multimodal behavior prediction model to at least one user device running edge computing. Subsequently, in various scenarios, the multimodal behavior prediction model, or any of several models, can be used to predict user behavior based on multiple types of sensor data and detect key events. These events can be compared with a reminder calendar. When a key event matches one of the reminder items set in the reminder calendar, a reminder in the form of text, images, or sound is generated through at least one user device. For example, a reminder can be generated through the user device in the form of text, voice, or other means (step S319).
[0059] It is worth mentioning that a trained language model, after training, can produce the same prediction results as a large language model trained using sensor data generated by reliable sensing devices, and if the probability of correct prediction exceeds a set threshold, it can be deployed to user devices to determine user behavior. In particular, the trained language model can be deployed to edge computing devices that consume less computing power, have low power consumption, and / or have small data volumes, such as one of the sensors.
[0060] As can be seen from the above embodiments, in a multimodal behavior prediction model training system, the trusted model running in a trusted sensor can be a large language model, while the one running in an untrusted sensor is a trained language model. The method of operation can be found in [reference needed]. Figure 4 The diagram shows an embodiment of the multimodal behavior prediction model training system and its operation method.
[0061] The computing device for running the multimodal behavior prediction model training system may include a processor 401 that performs general system operation and data processing, and a neural processor 403 that runs the neural network model. The processor 401 first obtains the sensing data generated by the trusted first sensor 411, and pre-processes it, for example, by converting the sensing data into data that can be recognized by the large language model 405 through a coding program. The neural processor 403 then runs the large language model 405 to perform trusted behavior prediction 408.
[0062] On the other hand, the processor 401 obtains untrusted sensing data generated by the untrusted second sensor 412. During the training process, the processor 401 first performs preprocessing and converts the sensing data into data that can be recognized by the large language model 405. The neural processor 403 uses the sensing data that has been labeled with user behavior and key events after trusted behavior prediction 408 to perform untrusted behavior prediction 407 to train the untrusted sensing data, and then trains the trained language model 409.
[0063] During the training of the trained language model 409, the prediction results of the untrusted behavior prediction 407 and the prediction results of the trustworthy behavior prediction 408 are continuously compared through the comparator 410. When the difference between the two reaches the preset threshold set by the system, it indicates that the trained language model 409 and the large language model 405 have similar confidence in predicting user behavior and judging whether a key event has occurred. The trained language model 409 forms a trustworthy prediction model. At the same time, the untrustworthy sensing data originally generated by the second sensor 412 can obtain trustworthy prediction results through the trained prediction model.
[0064] In this way, the prediction model trained by the neural processor 403 and the trusted model can jointly build a multimodal behavior prediction model to predict the user's behavior based on any one or more of the various sensing data.
[0065] For example, a trusted model, namely the large language model 405 shown in the figure, runs in the trusted first sensor 411 and has a higher probability of accurately predicting key events than a threshold. The image data generated by the first sensor 411 can accurately identify user behavior through the large language model. However, the large language model 405 cannot accurately identify user behavior based on the sound data generated by the second sensor 412 (which is coded into data that the large language model 405 can recognize). Therefore, the neural processor 403 trains the untrusted sound data generated by the second sensor 412 by labeling the trusted image data until the probability of the trained language model 409 predicting the key event exceeds the set threshold. This results in a trusted prediction model that can accurately identify user behavior.
[0066] Once one or more predictive models trained using the above-described embodiments, along with at least one trustworthy model, are obtained, they can accurately predict user behavior based on various sensor data and identify repetitive or periodic key events. Furthermore, a reminder calendar can be created for these repetitive or periodic key events, as can be referred to... Figure 5 The diagram shows an example of an reminder calendar 50 built using a multimodal behavior prediction model.
[0067] The reminder calendar 50 shown in the figure is merely an example and is not intended to limit the implementation. The reminders recorded therein are mainly for events with repetitive or periodic characteristics. Thus, the predictive model trained through the above process can be deployed by a computing device to sensors running edge computing or specific user devices to create the reminder calendar 50 and set reminders.
[0068] Therefore, when a predictive model running on a user's device, or a multimodal behavior prediction model deployed in a specific scenario, predicts a user's behavior based on various sensor data generated by multiple sensors, and identifies key events, a reminder can be generated by comparing it with the reminder calendar. For example, a reminder can be generated through the user's device in the form of text, images, or sound.
[0069] Figure 6 An example diagram is shown illustrating an implementation of a multimodal behavior prediction model that uses hardware and software collaboration within a computer system to execute alerts.
[0070] The figure illustrates a computing device running a multimodal behavior prediction model 61 via processor 60, which can also be a microprocessor in an edge computing device. Sensing data within a scene is acquired by a multimodal sensor in the behavior detection unit 65 of the computing device. This multimodal behavior prediction model 61 processes one or more types of sensing data, such as image, sound, location, and time data, to predict user behavior. Notably, this can involve using a large language model to directly predict user behavior, or using a trained model to assist in predicting user behavior; alternatively, the sensing device itself can perform edge computing, using a trained model that has achieved a correct prediction probability reaching a system-defined threshold to predict user behavior.
[0071] The reminder calendar 67 contains one or more reminders, which are set by the multimodal behavior prediction model 61 based on repetitive or periodic behaviors detected by the behavior detection unit 65. The processor 60 compares the predicted user behavior with the reminders set in the reminder calendar 67. If a reminder matches the user behavior, the reminder unit 63 can issue a reminder via voice, text, or vibration.
[0072] In summary, the main technical concept of the multimodal behavior prediction model training system and method described in the above embodiments is that multiple types of sensors generate multiple sensing data for the same event, and the trusted data can be used to train the data generated by the untrusted sensing device to train a model that is sufficient to predict user behavior. This enables the implementation of a multimodal behavior prediction model and the execution of reminders for key events.
[0073] The content disclosed above is only a preferred and feasible embodiment of the present invention, and is not intended to limit the scope of patent protection of the present invention. Therefore, all equivalent technical changes made based on the content of the present invention specification and drawings shall fall within the scope of patent protection of the present invention.
Claims
1. A method for training a multimodal behavior prediction model, running in a computing device, comprising: For a user to obtain multiple sensing data generated by multiple sensors, wherein the multiple sensors include one or more untrusted sensors and at least one trusted sensor; The user's behavior is predicted using various corresponding models based on the multiple types of sensor data, and a key event is identified from them. For the critical event, obtain sensing data generated by the at least one trusted sensor, wherein the at least one trusted sensor runs at least one trusted model that has a probability of accurately predicting the critical event exceeding a threshold. as well as The sensing data generated by the at least one trusted sensor is used to train one or more prediction models running in the one or more untrusted sensors for the critical event generated by the same or other untrusted sensors at the same time, until the probability of the one or more prediction models predicting the critical event exceeds the threshold.
2. The multimodal behavior prediction model training method according to claim 1, characterized in that, The various sensing data generated by the plurality of sensors include image sensing data obtained by at least one image acquisition device, sound sensing data obtained by at least one audio receiving device, and sensing data generated by at least one user device. The various sensing data are used to detect the position, movement, and actions of the at least one user.
3. The multimodal behavior prediction model training method according to claim 2, characterized in that, The at least one image acquisition device and the at least one audio receiving device are disposed in a scene to acquire images and sounds in the scene, and the user device is a wearable sensing device or a motion device worn on the user; wherein, the user's position is obtained through a positioning circuit in the user device, and the user's movement and actions in the scene are obtained through a motion sensing circuit in the user device.
4. The multimodal behavior prediction model training method according to claim 2, characterized in that, Using the one or more prediction models that have been trained and the at least one trusted model, the system predicts the user's behavior based on the various sensor data, identifies key events that are repetitive or periodic, and establishes a reminder calendar for the key events that are repetitive or periodic.
5. The multimodal behavior prediction model training method according to claim 4, characterized in that, The user's behavior is predicted using one or more prediction models based on the various sensing data generated by the multiple sensors, the key event is determined, and a reminder is generated by comparing it with the reminder calendar.
6. The multimodal behavior prediction model training method according to claim 5, characterized in that, When the key event matches one of the reminders set in the reminder calendar, the reminder is generated via the at least one user device in the form of text, images, or sound.
7. The multimodal behavior prediction model training method according to any one of claims 1 to 6, characterized in that, A multimodal behavior prediction model is established by combining the trained one or more prediction models and the at least one trusted model to predict the user's behavior based on any one or more of the various sensing data.
8. The multimodal behavior prediction model training method according to claim 7, characterized in that, The trusted model running in the trusted sensor is a large language model, and the prediction model running in the untrusted sensor is a trained language model; wherein, after the trusted sensor generates the sensing data, the sensing data is converted into data recognizable by the large language model through a coding program, the user's behavior is predicted, the key events are marked in the recognizable data, and then the trained language model is trained based on the marked key events.
9. The multimodal behavior prediction model training method according to claim 8, characterized in that, The trained language model is either a model formed by limiting the operation of the large language model through a prompt, or an acquisition-enhanced generative model formed by limiting the large language model to predict specific user behaviors.
10. A multimodal behavior prediction model training system, comprising: A computing device, comprising a processor and a neural processor; The processor acquires various sensing data generated by multiple sensors, including one or more untrusted sensors and at least one trusted sensor. Specifically, the neural processor uses various models corresponding to different types of sensor data to predict a user's behavior and identify a key event; for the key event, it obtains sensor data generated by at least one trusted sensor and uses it to train one or more prediction models for sensor data generated by one or more untrusted sensors simultaneously for the key event; and The at least one trusted model is run in the at least one trusted sensor, and the probability of accurately predicting the key event exceeds a threshold. The one or more prediction models running in the one or more untrusted sensors are trained until the probability of the one or more prediction models predicting the key event exceeds the threshold.