Data processing method and device, electronic equipment and computer readable storage medium
By using multimodal fusion model to process multimodal physiological data collected by physiological perception devices, the problem of ineffective supervision in traditional education is solved and efficient monitoring of students' health and learning efficiency is achieved.
Patent Information
- Application Number
- CN202411795291.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-06
- Publication Date
- 2025-05-06
AI Technical Summary
In traditional education, teachers or parents are unable to continuously and efficiently supervise due to limited energy and time, which makes it difficult to effectively monitor students' health and learning efficiency.
The state data of the evaluation object is obtained by obtaining the multimodal physiological data collected by the physiological perception device worn by the evaluation object and inputting it into the pre-trained multimodal fusion model.
It realizes accurate identification of the status of the evaluation object, improves the accuracy of monitoring students' health and learning efficiency, and reduces dependence on teachers or parents.
Smart Images

Figure CN119939495A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of data processing technology, and in particular to the technical fields of large models, intelligent devices, artificial intelligence, etc. Specifically, the present disclosure relates to a data processing method and device, an electronic device, and a computer-readable storage medium. Background Art
[0002] Traditional education mainly relies on manual supervision by teachers or parents, but teachers and parents have limited energy and time and are unable to provide continuous and efficient supervision.
[0003] With the development of technology, especially the development of smart device technology, more and more smart devices are being used to help teachers or parents monitor students' health or learning efficiency. Summary of the invention
[0004] The present disclosure provides a data processing method and device, an electronic device, and a computer-readable storage medium.
[0005] According to a first aspect of the present disclosure, a data processing method is provided, the method comprising:
[0006] Acquiring physiological data collected by at least one physiological sensing device worn by the evaluation subject; the at least one physiological sensing device is used to collect physiological data of different modalities;
[0007] Physiological data of different modalities are input into a pre-trained multimodal fusion model to obtain status data of the evaluation object.
[0008] According to a second aspect of the present disclosure, there is provided a data processing device, the device comprising:
[0009] A data acquisition module, used to acquire physiological data collected by at least one physiological sensing device worn by the evaluation subject; the at least one physiological sensing device is used to collect physiological data of different modalities;
[0010] The model prediction module is used to input physiological data of different modalities into a pre-trained multimodal fusion model to obtain the state data of the evaluation object.
[0011] According to a third aspect of the present disclosure, an electronic device is provided, the electronic device comprising:
[0012] at least one processor; and
[0013] A memory in communication with the at least one processor; wherein,
[0014] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the data processing method.
[0015] According to a fourth aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable a computer to execute the above-mentioned data processing method.
[0016] According to a fifth aspect of the present disclosure, a computer program product is provided, comprising a computer program, which implements the above data processing method when executed by a processor.
[0017] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The accompanying drawings are used to better understand the present solution and do not constitute a limitation of the present disclosure.
[0019] Figure 1 is a flowchart of a data processing method provided by an embodiment of the present disclosure;
[0020] Figure 2 It is a flowchart of some steps of a data processing method provided by an embodiment of the present disclosure;
[0021] Figure 3 It is a flowchart of some steps of a data processing method provided by an embodiment of the present disclosure;
[0022] Figure 4 It is a flowchart of some steps of a data processing method provided by an embodiment of the present disclosure;
[0023] Figure 5 It is a flowchart of some steps of a data processing method provided by an embodiment of the present disclosure;
[0024] Figure 6 It is a flowchart of some steps of a data processing method provided by an embodiment of the present disclosure;
[0025] Figure 7 It is a flowchart of some steps of a data processing method provided by an embodiment of the present disclosure;
[0026] Figure 8 It is a flowchart of a complete process of a data processing method according to an embodiment of the present disclosure;
[0027] Fig. 9 is a structural schematic diagram of a data processing device provided by an embodiment of the present disclosure;
[0028] Fig.10It is a block diagram of an electronic device used to implement the data processing method of the embodiment of the present disclosure. DETAILED DESCRIPTION
[0029] The following is a description of exemplary embodiments of the present disclosure in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0030] In some related technologies, smart devices usually focus on a single data type (such as heart rate, EEG, or blood oxygen), but lack a complete system that truly integrates multimodal signals. For example, heart rate may indicate physical fatigue, but it cannot be associated with mental concentration.
[0031] The data processing method and device, electronic device, and computer-readable storage medium provided by the embodiments of the present disclosure are intended to solve at least one of the above technical problems in the prior art.
[0032] The data processing method provided in the embodiments of the present disclosure may be executed by an electronic device such as a terminal device or a server, and the terminal device may be a vehicle-mounted device, a user equipment (UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (PDA), a handheld device, a computing device, a vehicle-mounted device, a wearable device, etc. The method may be implemented by a processor calling a computer-readable program instruction stored in a memory. Alternatively, the method may be executed by a server.
[0033] Figure 1 FIG. 1 is a flow chart showing a data processing method provided by an embodiment of the present disclosure. Figure 1 As shown in , the data processing method provided by the embodiment of the present disclosure may include step S110 and step S120.
[0034] In step S110, physiological data collected by at least one physiological sensing device worn by the evaluation subject is obtained; the at least one physiological sensing device is used to collect physiological data of different modalities;
[0035] In step S120, physiological data of different modalities are input into a pre-trained multimodal fusion model to obtain status data of the evaluation object.
[0036] For example, the evaluation object may be a student who is currently studying and other persons whose status needs to be evaluated.
[0037] A physiological sensing device can be any device that collects physiological data of the human body, and can be a device worn at any position on the human body.
[0038] In some possible implementations, the same physiological sensing device can be used to collect physiological data of different modalities, different physiological sensing devices can be used to collect physiological data of different modalities, and different physiological sensing devices can also collect different types of physiological data of the same modality.
[0039] In the data processing method provided in the embodiment of the present disclosure, no matter how many physiological sensing devices the evaluation subject wears, the physiological data finally collected must include at least two physiological data of different modalities.
[0040] Among them, modality refers to the form of expression of information or sensory channel. Therefore, in the data processing method provided in the embodiment of the present disclosure, physiological data of different sensory channels (such as vision, hearing, touch, etc.) can be obtained through physiological perception equipment.
[0041] In some possible implementations, the physiological sensing device may be a smart watch, a smart neck ring, etc., and the physiological data collected may include physiological data of body sensory modalities such as heart rate and skin conductivity (GSR, Galvanic Skin Response, also known as skin electrical response).
[0042] In some possible implementations, the physiological sensing device may be a visual acquisition device that can acquire physiological data of a visual modality, such as the wearer's visual content, visual trajectory, etc.
[0043] In some possible implementations, the physiological sensing device may be a voice acquisition device, which may be used to acquire physiological data of an auditory modality, such as voice data emitted by a wearer.
[0044] In some possible implementations, the physiological sensing device may be a real-time acquisition device, that is, it can acquire the physiological data of the evaluation object in real time. The physiological sensing device may also be a periodic acquisition device, that is, it can acquire the physiological data of the evaluation object according to a certain period.
[0045] In some possible implementations, in step S110, obtaining the physiological data collected by the physiological sensing device may be establishing a connection with the physiological sensing device through a wireless method (such as Bluetooth, WiFi, etc.) to obtain the physiological data collected by the physiological sensing device.
[0046] In some possible implementations, in step S120, the pre-trained multimodal fusion model may be a large model trained using a large amount of physiological data of multiple modalities, which may determine the correspondence between physiological data and physiological state by analyzing data of multiple modalities.
[0047] In some possible implementations, the acquired physiological data of the evaluation object is input into a pre-trained multimodal fusion model to determine the physiological state of the evaluation object, thereby identifying other states of the evaluation object and generating state data of the evaluation object.
[0048] In some possible implementations, the status data of the evaluation subject may include the evaluation subject's concentration, the evaluation subject's emotional state, and the evaluation subject's fatigue state, which are closely related to the evaluation subject's physiological state and can be inferred from the physiological state.
[0049] The embodiments of the present disclosure do not limit the specific model types of the multimodal fusion model, and any model that can achieve the corresponding functions is within the protection scope of the embodiments of the present disclosure.
[0050] In the data processing method provided in the embodiment of the present disclosure, physiological data of different modalities of the evaluation object are obtained, and the physiological data of different modalities are uniformly processed through a large model, and then the physiological data of different modalities are associated, providing a complete system for integrating multimodal data, thereby improving the accuracy of identifying the status data of the evaluation object.
[0051] The data processing method provided by the embodiment of the present disclosure is introduced in detail below.
[0052] As described above, physiological data is used to identify status data of the evaluation subject.
[0053] When the subject is in a state of high emotional arousal, such as tension, anxiety, excitement, or stress, the activity of the sympathetic nervous system increases, causing the heart rate to increase, the microsweat glands of the skin to secrete, and the skin conductivity to increase.
[0054] When the subject is in a state of high concentration on a task, or encountering pressure while solving a difficult problem, which leads to increased sympathetic nervous activity, their skin conductivity will also increase.
[0055] When the subject moves from a high-pressure task to a relaxed state or gradually enters a state of fatigue, their sympathetic nerve activity weakens and their skin conductivity also decreases.
[0056] In other words, the heart rate and skin conductivity of the subject are closely related to the state of the subject. In particular, skin conductivity can sensitively reflect short-term fluctuations in emotional changes. Whether it is the increased pressure of encountering a difficult task or the high tension when emotions are out of control, it will cause a large fluctuation in short-term skin conductivity.
[0057] In view of this, in some possible implementations, the physiological sensing device may include a wearable device, and the wearable device is used to collect behavioral sequence data of the evaluation object.
[0058] Among them, the wearable device can use PPG (photoplethysmography) to measure the heart rate of the evaluation subject; use GSR (skin conductance sensor) to monitor the sweat gland activity of the evaluation subject and measure the skin conductivity of the evaluation subject. The behavior sequence data of the evaluation subject collected by the wearable device may include the heart rate data and skin conductivity data of the evaluation subject.
[0059] In some possible implementations, the collected behavior sequence data may contain noise, and the collected behavior sequence data may be missing due to the insensitivity of the sensor. Therefore, before inputting the behavior sequence data into the multimodal fusion model to obtain the status data of the evaluation object, the behavior sequence data needs to be preprocessed to improve the data quality of the behavior sequence data, thereby improving the recognition accuracy of the multimodal fusion model.
[0060] Figure 2 A flowchart of an implementation method for preprocessing behavior sequence data is shown in FIG. Figure 2 As shown, preprocessing the behavior sequence data may include step S210, step S220, and step S230.
[0061] In step S210, the behavior sequence data is subjected to denoising by filtering to obtain the denoised behavior sequence data;
[0062] In step S220, time synchronization processing is performed on the denoised behavior sequence data to obtain time-synchronized behavior sequence data;
[0063] In step S230, the time-synchronized behavior sequence data is input into a feature extractor to extract behavior data features of the time-synchronized behavior sequence data.
[0064] In some possible implementations, in step S210, a filtering operation is used to eliminate data in the heart rate data that is too large or too small and exceeds the average change range of the heart rate, so as to eliminate abnormal data (i.e., noise data) that occurs during the collection of the heart rate data due to external interference factors such as the tightness of the wearable device and the body surface temperature.
[0065] In some specific implementations, filtering algorithms such as average filtering and Kalman filtering may be used to implement denoising of the heart rate data to obtain denoised heart rate data.
[0066] In some possible implementations, during the collection of skin conductivity data, skin conductivity may produce unreasonable change curves due to emotions and other irrelevant behaviors (such as rubbing hands, etc.). Therefore, it is necessary to make a rationality judgment on the skin conductivity change curve and eliminate obviously unreasonable abnormal data (i.e., noise data). For example, the data with a sharp drop in the steadily rising skin conductivity change curve can be eliminated to achieve denoising of the skin conductivity data and obtain the denoised skin conductivity data.
[0067] Since both the heart rate data and the skin conductivity data are sequence data, it is necessary to ensure the time synchronization of the heart rate data and the skin conductivity data, that is, the heart rate data and skin conductivity data included in the behavior sequence data are data within the same period of time, and the heart rate data and skin conductivity data at the same time point are corresponding.
[0068] In some possible implementations, in step S220, precise timestamps are used to perform time synchronization processing on the denoised behavior sequence data to ensure the time consistency of the heart rate data and the skin conductivity data, as well as the time consistency of the heart rate data and the skin conductivity data with other physiological data, to obtain time-synchronized behavior sequence data.
[0069] In some possible implementations, in step S230, the feature extractor may be generated through pre-training.
[0070] The feature extractor is used to calculate behavioral data features such as heart rate variability based on heart rate data, and to calculate behavioral data features such as skin conductance change based on skin conductivity.
[0071] In some possible implementations, after using a feature extractor to extract the behavior sequence data, i.e., the data features of heart rate data and skin conductivity, the acquired behavior data features can be screened, and the data features with greater relevance to the evaluation object state data can be input into the multimodal fusion model to reduce the impact of irrelevant data features on the recognition of the multimodal fusion model and improve the recognition accuracy of the multimodal fusion model.
[0072] In some possible implementations, the time-synchronized behavior sequence data may be directly input into the multimodal fusion model, and the multimodal fusion model may be directly used to process the behavior sequence data, thereby avoiding the waste of resources and increase in recognition time caused by the use of a feature extractor.
[0073] In some possible implementations, the behavior sequence data and the behavior data features of the behavior sequence data may also be input into the multimodal fusion model together, so as to improve the recognition accuracy of the multimodal fusion model by increasing the amount of information entering and exiting the multimodal fusion model.
[0074] When the subject is highly focused on a task, his eyes should be in a state of quickly scanning to understand the task content or focusing on a part of the task; when the subject is fatigued, his concentration is not there and his eyes may be in a state of wandering.
[0075] In view of this, in some possible implementations, the physiological perception device may include an eye tracking device for collecting visual data of the evaluation object.
[0076] Among them, visual data can include the eye's activity path, the length of time the eyes stay on different targets, and the eye's movement speed and acceleration.
[0077] Eye tracking devices can use infrared tracking technology or camera tracking technology. The advantage of infrared technology is that even in poor lighting conditions, eye tracking devices can still accurately track eye movements, while camera tracking technology has high requirements for ambient light, but it is low cost to implement.
[0078] In some possible implementations, some scattered, too-fast gaze movements and invalid tracking caused by blinking or external interference may bring more "empty data", which may interfere with recognition. Therefore, before inputting the visual data into the multimodal fusion model to identify and evaluate the state data of the object, the visual data needs to be preprocessed to improve the data quality of the visual data, thereby improving the recognition accuracy of the multimodal fusion model.
[0079] Figure 3 A schematic diagram of a process for implementing a method of preprocessing visual data is shown in FIG. Figure 3 As shown, preprocessing the visual data may include step S310 and step S320.
[0080] In step S310, abnormal data in the visual data is removed according to the fluctuation standard deviation of the visual data;
[0081] In step S320, the visual data is input into a feature extractor to extract visual data features of the visual data.
[0082] In some possible implementations, in step S310, the spacing between eye trajectory points and the standard deviation of trajectory fluctuations can be calculated to eliminate "empty data", i.e., abnormal data, caused by scattered, overly fast gaze movements and invalid tracking due to blinking or external interference.
[0083] In some possible implementations, interference such as rapid look back may also be handled by statistical mean point regression.
[0084] In some possible implementations, the time synchronization of visual data and behavioral sequence data is achieved through precise timestamps, that is, the visual data and behavioral sequence data of the same period of time are obtained.
[0085] In some possible implementations, in step S320, the feature extractor may be pre-trained and generated. The feature extractor is used to obtain visual data features related to the state data of the evaluation object, such as the duration of eye fixation, eye fixation position, eye movement direction, and eye movement speed of the evaluation object according to the visual data.
[0086] In some possible implementations, after using a feature extractor to extract visual data features of visual data, the acquired visual data features can be screened, and data features with greater relevance to the evaluation object state data can be input into the multimodal fusion model to reduce the impact of irrelevant data features on the recognition of the multimodal fusion model and improve the recognition accuracy of the multimodal fusion model.
[0087] In some possible implementations, the visual data that is time-synchronized and has eliminated abnormal data can be directly input into the multimodal fusion model, and the multimodal fusion model can be used directly to process the visual data, thereby avoiding the waste of resources and increase in recognition time caused by the use of feature extractors.
[0088] In some possible implementations, the visual data and the visual data features of the visual data may also be input into the multimodal fusion model together, so as to improve the recognition accuracy of the multimodal fusion model by increasing the amount of information entering and exiting the multimodal fusion model.
[0089] In some possible implementations, the wearable device may also include other sensors to collect other physiological data of the evaluation object. The processing method of other physiological data can be consistent with the behavior sequence data and visual data, removing noise data, using precise timestamps for data synchronization, and using feature extractors for feature extraction.
[0090] The state of the subject may also be affected by the surrounding environment. The more suitable the surrounding environment is for the task, the more likely the subject is to focus. The noisier and less suitable the surrounding environment is for the task, the more likely the subject is to be affected by the environment.
[0091] In view of this, in some possible implementations, environmental data of the environment in which the evaluation object is located can be collected, and the environmental data and physiological data can be input into the multimodal fusion model together to help the multimodal fusion model to identify and improve the recognition accuracy of the multimodal fusion model.
[0092] Figure 4A flowchart of an implementation method of collecting environmental data of the environment where the evaluation object is located and inputting the environmental data and physiological data into a multimodal fusion model is shown, such as Figure 4 As shown, it may include step S410, step S420, and step S430.
[0093] In step S410, environmental data collected by an environmental sensing device of the environment where the evaluation object is located is obtained;
[0094] In step S420, the environmental data is subjected to denoising using a smoothing filtering technique to obtain denoised environmental data;
[0095] In step S430, the environmental data is input into a feature extractor to extract environmental data features of the environmental data; the environmental data features are input into a multimodal fusion model to obtain state data of the evaluation object.
[0096] In some possible implementations, in step S410, the environment sensing device may include a light sensor, a noise detector, and a temperature sensor.
[0097] The light sensor is used to obtain the light conditions of the environment where the evaluation object is located; the temperature sensor is used to obtain the temperature conditions of the environment where the evaluation object is located; and the noise detector is used to obtain whether there is noise in the environment where the evaluation object is located.
[0098] Therefore, the environmental data collected by the environmental sensing device may include light data, temperature data, and noise data.
[0099] Since the accuracy, robustness and corresponding transmission methods of environmental perception devices will greatly affect the timeliness of environmental data, environmental perception devices need to communicate in real time through optimized transmission protocols and faster transmission speeds (such as 5G or Wi-Fi6).
[0100] In some possible implementations, in step S420, since environmental data may fluctuate frequently due to rapid changes in external conditions, smoothing filtering technology may be used to denoise the environmental data, uniformly process short-term large fluctuations into an acceptable buffer zone, and generate denoised environmental data.
[0101] In some specific implementations, the smoothing filtering technique used may be sliding average.
[0102] In some possible implementations, time synchronization between environmental data and physiological data is achieved through precise time stamps, that is, physiological data and environmental data of the same period of time are acquired.
[0103] In some possible implementations, in step S430, the feature extractor may be generated by pre-training. The feature extractor is used to obtain the environmental change rate of the environment where the evaluation object is located according to the environmental data.
[0104] In some possible implementations, after using a feature extractor to extract environmental data features of environmental data, the acquired environmental data features can be screened, and data features with greater relevance to the evaluation object state data can be input into the multimodal fusion model to reduce the impact of irrelevant data features on the recognition of the multimodal fusion model and improve the recognition accuracy of the multimodal fusion model.
[0105] In some possible implementations, the environmental data that is time synchronized and has abnormal data eliminated can be directly input into the multimodal fusion model, and the multimodal fusion model can be directly used to process the environmental data, thereby avoiding the waste of resources and increase in recognition time caused by the use of feature extractors.
[0106] In some possible implementations, the environmental data and the environmental data features of the environmental data may also be input into the multimodal fusion model together, so as to improve the recognition accuracy of the multimodal fusion model by increasing the amount of information entering and exiting the multimodal fusion model.
[0107] During the task of the evaluation object, the state data of the evaluation object can also be judged by the quality of the evaluation object's writing data. When the evaluation object is in a highly concentrated state data, the accuracy of its writing internal skills is higher than that in a non-concentrated state, and the handwriting may also be relatively neat.
[0108] In view of this, in some possible implementations, the writing data of the evaluation object can be collected and input into the multimodal fusion model to assist the multimodal fusion model in recognition so as to improve the recognition accuracy of the multimodal fusion model.
[0109] Figure 5 FIG. 1 shows a flow chart of an implementation method of collecting the writing data of the evaluation object and inputting the writing data into the multimodal fusion model. Figure 5 As shown, collecting the writing data of the evaluation object and inputting the writing data into the multimodal fusion model may include step S510, step S520, and step S530.
[0110] In step S510, the writing data of the evaluation object is obtained, and the text data corresponding to the writing data is obtained by optical character recognition;
[0111] In step S520, the text data is input into a feature extractor to extract text data features of the text data;
[0112] In step S530, the text data features are input into the multimodal fusion model to obtain the status data of the evaluation object.
[0113] In some possible implementations, in step S510, the writing data of the evaluation object may be acquired through pressure sensing or optical scanning by a smart pen with an embedded sensor. The data acquired by the smart pen may be transmitted in real time through communication technology.
[0114] In some possible implementations, OCR (Optical Character Recognition) technology is used to obtain text data corresponding to the written data. The text data may specifically include text content data and handwriting data.
[0115] In some possible implementations, since the pen pressure of the evaluation subject when writing may be very high when the evaluation subject is nervous and highly focused, the pen pressure data may also be recorded as part of the writing data.
[0116] Similarly, the time synchronization between written data and physiological data is achieved through precise timestamps, that is, the physiological data and written data of the same period of time are obtained.
[0117] In some possible implementations, in step S520, the feature extractor may be pre-trained and generated. The feature extractor is used to extract text data features of text data. Specifically, the text content data may be semantically detected to determine the accuracy of the text content; the handwriting data may be analyzed to determine the neatness of the handwriting. The accuracy of the text content and the neatness of the handwriting are both text data features.
[0118] In some possible implementations, the feature extractor may also be used to obtain the pen pressure change rate based on the pen pressure data, and use the pen pressure change rate as a component of the text data feature.
[0119] In some possible implementations, in step S530, after obtaining the text data features, they can be screened to select data features that are more relevant to the state data of the evaluation object and input into the multimodal fusion model to reduce the impact of irrelevant data features on the recognition of the multimodal fusion model and improve the recognition accuracy of the multimodal fusion model.
[0120] In some possible implementations, the written data that is time-synchronized and has had abnormal data removed can be directly input into the multimodal fusion model, and the multimodal fusion model can be used directly to process the written data, thereby avoiding the waste of resources and increase in recognition time caused by the use of a feature extractor.
[0121] In some possible implementations, the writing data and text data features may be input into the multimodal fusion model together, thereby improving the recognition accuracy of the multimodal fusion model by increasing the amount of information entering and exiting the multimodal fusion model.
[0122] In some possible implementations, the acquired writing data of the assessment subject may also be used to provide guidance on the assessment subject's tasks.
[0123] Figure 6 A flowchart of an implementation method of tutoring an assessment subject's task based on writing data is shown. Figure 6 As shown, coaching the task of the assessment subject according to the writing data may include step S610 and step S620.
[0124] In step S610, semantic detection is performed on the text data, and errors in the text data are detected and determined based on the semantic detection result;
[0125] In step S620, the erroneous content is analyzed, and modification suggestions for the erroneous content are generated and fed back to the evaluation object so that the evaluation object can modify the erroneous content according to the modification suggestions.
[0126] In some possible implementations, in step S610, the text data is obtained in the same manner as in step S510, which will not be described in detail here. The text data is semantically detected using a pre-trained artificial intelligence model, and the artificial intelligence model performs semantic detection on the text content of each line of text data identified, and determines the correctness of the text content based on the semantic detection result, and further determines the erroneous content in the text data.
[0127] In some possible implementations, in step S620, when erroneous content is detected, a search is performed in the database according to the question corresponding to the erroneous content, the correct answer of the corresponding question is searched, and a right or wrong question analysis and modification suggestions are generated based on the correct answer and the erroneous text content, and the wrong question analysis and corresponding modification suggestions are fed back to the evaluation object so that the evaluation object can modify the erroneous content after receiving the error analysis and modification opinions.
[0128] The above method can effectively improve the real-time performance of the inspection during the task, thereby greatly reducing the time burden of other personnel in assisting the task.
[0129] As mentioned above, since the data features of the input multimodal fusion model are data features of different modal data, considering that physiological data, visual data, environmental data, and writing data are time series data, visual modal data, speech modal data (such as noise data), and text modal data, respectively, the multimodal fusion model based on the deep learning framework should be a visual-speech-text multimodal fusion model, and it has a time series module, such as LSTM, which can process time series information.
[0130] At the same time, since the data features of different modalities are not data features of the same feature space, the multimodal fusion model should include a feature fusion module that can fuse the data features of different modalities into the same feature space, and a prediction module that predicts the state data of the evaluation object based on the fused features fused into the same feature space.
[0131] That is to say, the multimodal fusion model is used to input the data features corresponding to different modal data into the multimodal fusion model and fuse them into the same feature space, and predict the state data of the evaluation object based on the fusion results.
[0132] Among them, the data features corresponding to different modal data are two or more of the above-mentioned behavioral data features, visual data features, environmental data features, and text data features.
[0133] Since the state data of the evaluation object is a composite state, it includes the concentration of the evaluation object, the emotional state of the evaluation object, and the fatigue state of the evaluation object.
[0134] Therefore, the multimodal fusion model is also a multi-task model, which can perform multi-task learning based on the fusion features. Its prediction module includes a submodule for identifying the concentration of the evaluation object, a submodule for identifying the emotional state of the evaluation object, and a submodule for identifying the fatigue state of the evaluation object. The parameters of different submodules are different and are obtained based on the training of different tasks in the multi-task.
[0135] That is, the multimodal fusion model is used to input physiological data of different modalities into the multimodal fusion model to identify the concentration of the evaluation object, the emotional state of the evaluation object, and the fatigue state of the evaluation object.
[0136] In some possible implementations, after acquiring the status data of the evaluation object according to the multimodal fusion model, the environmental perception device of the environment where the evaluation object is located can be controlled according to the status data of the evaluation object to adjust the environment where the evaluation object is located.
[0137] Specifically, the status data of the evaluation object is evaluated, feedback suggestions are generated based on the evaluation results, and the environmental perception device regulation actions are controlled. For example, when the evaluation object is in a state of fatigue, a rest reminder is issued through the audio device; when the evaluation object's concentration does not meet the requirements, the learning environment parameters (light, volume, etc.) are adjusted to create an environment more suitable for the evaluation object's concentration. At the same time, when regulating the environmental perception device, the current concentration of the evaluation object will be taken into account to avoid sudden environmental changes interfering with the evaluation object's concentration state.
[0138] Feedback suggestions are to automatically provide adjustment plans to the subject or other people related to the subject based on the status data of the subject. For example, if the subject's concentration decreases or fatigue increases, the subject will be informed through a message reminder tool in a timely manner that he needs to take a break. In addition, external factors such as indoor lighting and music can be adjusted to help the subject stay focused and reduce external interference.
[0139] The basis of this feedback mechanism is that it not only focuses on the physiological or behavioral signals of individuals, but also combines the state of the environment in which the evaluation object is located to achieve comprehensive regulation and improve the overall task efficiency. For example, if the system detects that the evaluation object has emotional fluctuations, it can play soothing light music and provide necessary stress relief strategies. At the same time, when the large model recognizes that the concentration has dropped sharply, it will also recommend that the evaluation object temporarily stop and relax for 5 to 10 minutes.
[0140] In some possible implementations, after controlling the environmental perception device of the environment in which the evaluation object is located to adjust the environment in which the evaluation object is located based on the status data of the evaluation object, it is possible to determine whether the judgment on the status data of the evaluation object is correct based on the status change of the evaluation object, and then perform an "adaptive" update of the multimodal fusion model based on the result of whether the judgment on the status data of the evaluation object is correct, so as to generate a dedicated model that conforms to the personal habits of the evaluation object.
[0141] Figure 7 A flowchart showing an implementation method of "adaptive" updating of a multimodal fusion model is shown, Figure 7 As shown, the “adaptive” update of the multimodal fusion model may include step S710, step S720, and step S730.
[0142] In step S710, after adjusting the environment where the evaluation object is located, physiological data collected by at least one physiological sensing device is obtained;
[0143] In step S720, the physiological data of different modalities are input into the multimodal fusion model to identify the state data of the evaluation object;
[0144] In step S730, the parameters of the multimodal fusion model are adjusted through back propagation according to the change between the state data of the evaluation object before the environment where the evaluation object is located is adjusted and the state data of the evaluation object after the environment where the evaluation object is located is adjusted.
[0145] In some possible implementations, the method of obtaining the status data of the evaluation object after adjusting the learning environment in step S710 and step S720 is as described above and will not be repeated here.
[0146] In some possible implementations, in step S730, the change between the state data of the evaluation object before adjusting the environment where the evaluation object is located and the state data of the evaluation object after adjusting the environment where the evaluation object is located can be used as the loss value of the loss function of the model, and back propagation is performed based on the loss value to adjust the parameters of the multimodal fusion model, so as to generate an adaptive multimodal fusion model.
[0147] When the evaluation object's concentration decreases or fatigue occurs during the learning process, the system will issue a reminder or suggestion. After the reminder or suggestion is issued, the evaluation object's response (whether to adopt the rest suggestion, adjust behavior, etc.) is used as a feedback input and then fed back into the multimodal fusion model to improve the parameters of the multimodal fusion model so that the multimodal fusion model is more in line with the evaluation object's habits.
[0148] Specifically, if the system detects fatigue or mood swings and prompts the subject, if the subject chooses to follow the suggested behavior, the multimodal fusion model will identify this feedback as effective regulation.
[0149] If the subject ignores the suggestions and the learning effect of the subsequent performance continues to decline, the multimodal fusion model will re-analyze whether the current recognition of the subject's status data is correct, whether the generated control plan is effective, and try to improve the content of the suggestions and reminder methods, such as adjusting the frequency of suggestions or switching to other strategies, such as environmental control (for example, light adjustment).
[0150] This feedback loop is constantly closed, allowing the system to better adapt to the individual needs of each evaluation object, not only providing solutions for the current state, but also laying the foundation for future state evaluation. Ultimately, the multimodal fusion model gradually becomes an adaptive prediction system with the individual data of the evaluation object as the core, providing more accurate and personalized learning for the evaluation object.
[0151] Based on the above statement, Figure 8 A flow chart showing a complete process of the data processing method provided by the embodiment of the present disclosure is shown as follows: Figure 8As shown, the data of different modalities of the evaluation object, namely physiological data, visual data, environmental data, and text data, are collected by the above-mentioned smart device, and features are extracted from these data to obtain data features. The state data of the evaluation object is obtained using a multimodal fusion model based on the extracted data features. If the state data of the evaluation object is normal, the data of different modalities, environmental data, and text data of the evaluation object continue to be collected by the above-mentioned smart device; if the state data of the evaluation object is abnormal, feedback suggestions for adjusting the state data of the evaluation object are generated, and whether the state data of the evaluation object changes after the feedback suggestions are executed is used to determine whether the feedback suggestions are useful. If the feedback suggestions are useful, the state data of the evaluation object will change and become better. At this time, the data of different modalities, environmental data, and text data of the evaluation object can continue to be collected by the above-mentioned smart device; if the feedback suggestions are useless, the state data of the evaluation object will not change and may even continue to decline. At this time, the multimodal fusion model is "adaptively" adjusted using the method described above to make the multimodal fusion model more in line with the evaluation object's wishes.
[0152] The data processing method of the embodiment of the present disclosure can be specifically applied to the following application scenarios.
[0153] Scenario 1: A wears a smart wearable device and an eye tracker to review for a math test. During the learning process, the system continuously records his physiological parameters (such as heart rate, skin conductivity, etc.) and eye movement changes and body posture through the wearable device. After a period of review, the data collected by the device is processed by the model, and it is concluded that A has developed a more obvious state of fatigue, and the randomness of the eye movement trajectory and the shortening of the gaze time indicate that his concentration has begun to decline. The system immediately issues a prompt sound in the learning environment and reminds A to pause the review for a few minutes through the parent's mobile phone App (application). If the mood fluctuates significantly, a soothing music will be recommended for him to listen to. In addition, the system automatically adjusts the indoor light to a gentler mode to avoid eye fatigue caused by excessive light.
[0154] Scenario 2: B is completing his Chinese homework at home. The homework content is fed back to the big model at the back end in real time through the system smart pen for analysis. When Xiaohua's writing content has incoherent sentences or format errors, the system identifies the relevant wrong questions and generates detailed correction suggestions and sentence structure instructions for it based on the built-in big language model. These correction information is automatically fed back to the parents' mobile phones through the App, helping them to check their children's wrong questions more accurately and efficiently. When the child completes the homework, the system will suggest parents to provide relevant tutoring, and provide detailed supplementary explanations to help parents understand the difficulties encountered by their children.
[0155] Scenario 3: C has a fixed review time arranged by the system every day, but because parents are sometimes unable to supervise their learning progress in the first place, the system will generate a "Parent Supervision Assistant" based on the length of review time, continuous concentration, and emotional fluctuations collected. With the help of this assistant, parents can see their children's learning progress in real time, including long-term concentration mode or whether there are too many interruptions in a short period of time. Feedback mechanism: The assistant will generate a specific review time plan. When it is found that the child has been reviewing effectively for 45 minutes, it will send a reminder to the parent to take a 5-minute break. At the same time, it can also recommend parents to continue to provide supplementary tutoring suggestions through online courses based on the subject of study.
[0156] Based on Fig. 9 The same principle as shown in the method, Fig. 9 A schematic diagram of the structure of a data processing device provided by an embodiment of the present disclosure is shown. Fig. 9 As shown, the data processing device 90 may include:
[0157] The data acquisition module 910 is used to acquire physiological data collected by at least one physiological sensing device worn by the evaluation subject; the at least one physiological sensing device is used to collect physiological data of different modalities;
[0158] The model prediction module 920 is used to input physiological data of different modalities into a pre-trained multimodal fusion model to obtain status data of the evaluation object.
[0159] In the data processing device provided in the embodiment of the present disclosure, physiological data of different modalities of the evaluation object are obtained, and the physiological data of different modalities are uniformly processed through a large model, and then the physiological data of different modalities are associated, providing a complete system for integrating multimodal data, thereby improving the accuracy of identifying the status data of the evaluation object.
[0160] In some possible implementations, the data processing device further includes: a feedback adjustment module, configured to control an environmental perception device of an environment where the evaluation object is located to adjust the environment where the evaluation object is located according to the state data of the evaluation object.
[0161] In some possible implementations, the data processing device also includes: an adaptive module, which is used to obtain physiological data collected by at least one physiological sensing device after adjusting the environment where the evaluation object is located; input physiological data of different modalities into a multimodal fusion model to obtain state data of the evaluation object; and adjust the parameters of the multimodal fusion model through back propagation based on changes in the state data of the evaluation object before adjusting the environment where the evaluation object is located and the state data of the evaluation object after adjusting the environment where the evaluation object is located.
[0162] In some possible implementations, the physiological sensing device includes a wearable device, which is used to collect behavior sequence data of the evaluation object; the model prediction module includes: a behavior sequence unit, which is used to denoise the behavior sequence data by filtering to obtain the denoised behavior sequence data; perform time synchronization on the denoised behavior sequence data to obtain the time-synchronized behavior sequence data; input the time-synchronized behavior sequence data into a feature extractor to extract behavior data features of the time-synchronized behavior sequence data; input the behavior data features into a multimodal fusion model to obtain state data of the evaluation object.
[0163] In some possible implementations, the physiological perception device includes an eye tracking device, which is used to collect visual data of the evaluation object; the model prediction module includes: a visual unit, which is used to eliminate abnormal data in the visual data according to the fluctuation standard deviation of the visual data; inputting the visual data into a feature extractor to extract visual data features of the visual data; inputting the visual data features into a multimodal fusion model to obtain status data of the evaluation object.
[0164] In some possible implementations, the data processing device also includes: an environmental data module, which is used to obtain environmental data collected by an environmental perception device in the environment where the evaluation object is located; use smoothing filtering technology to denoise the environmental data to obtain the denoised environmental data; input the environmental data into a feature extractor to extract environmental data features of the environmental data; input the environmental data features into a multimodal fusion model to obtain status data of the evaluation object.
[0165] In some possible implementations, the data processing device also includes: a writing data module, used to obtain writing data corresponding to the evaluation object, and obtain text data of the writing data through optical character recognition; input the text data into a feature extractor to extract text data features of the text data; input the text data features into a multimodal fusion model to obtain status data of the evaluation object.
[0166] In some possible implementations, the writing data module is also used to: perform semantic detection on the text data, and determine the erroneous content in the text data based on the semantic detection results; analyze the erroneous content, generate modification suggestions for the erroneous content and feed back to the evaluation object so that the evaluation object can modify the erroneous content according to the modification suggestions.
[0167] In some possible implementations, the model prediction module is also used to: input data features corresponding to different modal data into a multimodal fusion model to fuse them into the same feature space, and predict the state data of the evaluation object based on the fusion results.
[0168] In some possible implementations, the multimodal fusion model is a multi-task model, and the state data of the evaluation object includes the concentration of the evaluation object, the emotional state of the evaluation object, and the fatigue state of the evaluation object; the model prediction module is also used to: input physiological data of different modalities into the multimodal fusion model to identify the concentration of the evaluation object, the emotional state of the evaluation object, and the fatigue state of the evaluation object.
[0169] It can be understood that the above modules of the data processing device in the embodiment of the present disclosure have the function of implementing Figure 1 The functions of the corresponding steps of the data processing method in the embodiment shown in . The function can be implemented by hardware, or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions. The above modules can be software and / or hardware, and the above modules can be implemented separately or integrated with multiple modules. For the functional description of each module of the above data processing device, please refer to Figure 1 The corresponding description of the data processing method in the embodiment shown in will not be repeated here.
[0170] In the technical solution disclosed herein, the acquisition, storage and application of user personal information involved are in compliance with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0171] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium and a computer program product.
[0172] The electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the data processing method provided in the embodiment of the present disclosure.
[0173] Compared with the existing technology, this electronic device obtains physiological data of different modalities of the evaluation object, uniformly processes the physiological data of different modalities through a large model, and then associates the physiological data of different modalities, providing a complete system for integrating multimodal data and improving the accuracy of identifying the status data of the evaluation object.
[0174] The readable storage medium is a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to enable a computer to execute a data processing method as provided in an embodiment of the present disclosure.
[0175] Compared with the existing technology, this readable storage medium obtains physiological data of different modalities of the evaluation object, uniformly processes the physiological data of different modalities through a large model, and then associates the physiological data of different modalities, providing a complete system for integrating multimodal data and improving the accuracy of identifying the status data of the evaluation object.
[0176] The computer program product includes a computer program, and when the computer program is executed by a processor, the computer program implements the data processing method provided in the embodiment of the present disclosure.
[0177] Compared with the existing technology, this computer program product obtains physiological data of different modalities of the evaluation object, uniformly processes the physiological data of different modalities through a large model, and then associates the physiological data of different modalities, providing a complete system for integrating multimodal data and improving the accuracy of identifying the status data of the evaluation object.
[0178] Fig.10 A schematic block diagram of an example electronic device 1000 that can be used to implement an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.
[0179] like Fig.10 As shown, the device 1000 includes a computing unit 1001, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1002 or a computer program loaded from a storage unit 1008 into a random access memory (RAM) 1003. In the RAM 1003, various programs and data required for the operation of the device 1000 can also be stored. The computing unit 1001, the ROM 1002, and the RAM 1003 are connected to each other via a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.
[0180] A number of components in the device 1000 are connected to the I / O interface 1005, including: an input unit 1006, such as a keyboard, a mouse, etc.; an output unit 1007, such as various types of displays, speakers, etc.; a storage unit 1008, such as a disk, an optical disk, etc.; and a communication unit 1009, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 1009 allows the device 1000 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0181] The computing unit 1001 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 1001 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 1001 performs the various methods and processes described above, such as data processing methods. For example, in some embodiments, the data processing method may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 1008. In some embodiments, part or all of the computer program may be loaded and / or installed on the device 1000 via ROM 1002 and / or communication unit 1009. When the computer program is loaded into RAM 1003 and executed by the computing unit 1001, one or more steps of the data processing method described above may be performed. Alternatively, in other embodiments, the computing unit 1001 may be configured to perform the data processing method in any other appropriate manner (e.g., by means of firmware).
[0182] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0183] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.
[0184] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0185] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0186] The systems and techniques described herein may be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0187] A computer system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship of client and server is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, a server of a distributed system, or a server combined with a blockchain.
[0188] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this document does not limit this.
[0189] The above specific implementations do not constitute a limitation on the protection scope of the present disclosure. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.
Claims
1. A data processing method, comprising: Acquiring physiological data collected by at least one physiological sensing device worn by the evaluation subject; The at least one physiological sensing device is used to collect physiological data of different modalities; Physiological data of different modalities are input into a pre-trained multimodal fusion model to obtain status data of the evaluation object.
2. The method according to claim 1, further comprising: According to the status data of the evaluation object, an environment perception device of the environment where the evaluation object is located is controlled to adjust the environment where the evaluation object is located.
3. The method according to claim 2, wherein: After controlling the environment perception device of the environment where the evaluation object is located to adjust the environment where the evaluation object is located according to the state data of the evaluation object, the method further includes: After adjusting the environment where the evaluation object is located, obtaining physiological data collected by the at least one physiological sensing device; inputting the physiological data of different modalities into the multimodal fusion model to obtain the state data of the evaluation object; According to the change of the state data of the evaluation object before the environment where the evaluation object is located is adjusted and the state data of the evaluation object after the environment where the evaluation object is located is adjusted, the parameters of the multimodal fusion model are adjusted through back propagation.
4. The method according to claim 1, wherein: The physiological sensing device includes a wearable device, which is used to collect the behavior sequence data of the evaluation object; the physiological data of different modalities are input into a pre-trained multimodal fusion model to obtain the state data of the evaluation object, including: Performing denoising processing on the behavior sequence data by filtering to obtain denoised behavior sequence data; Performing time synchronization processing on the denoised behavior sequence data to obtain time synchronized behavior sequence data; Inputting the time-synchronized behavior sequence data into a feature extractor to extract behavior data features of the time-synchronized behavior sequence data; The behavior data features are input into the multimodal fusion model to obtain the status data of the evaluation object.
5. The method according to claim 1, wherein: The physiological perception device includes an eye tracking device, which is used to collect visual data of the evaluation object; The step of inputting physiological data of different modalities into a pre-trained multimodal fusion model to obtain status data of the evaluation object includes: Eliminating abnormal data in the visual data according to the standard deviation of fluctuations of the visual data; Inputting the visual data into a feature extractor to extract visual data features of the visual data; The visual data features are input into the multimodal fusion model to obtain the status data of the evaluation object.
6. The method according to claim 1, further comprising: Acquire environmental data collected by an environmental sensing device of the environment where the evaluation object is located; Using a smoothing filter technique to perform denoising on the environmental data, and obtaining denoised environmental data; Inputting the environmental data into a feature extractor to extract environmental data features of the environmental data; The environmental data features are input into the multimodal fusion model to obtain the status data of the evaluation object.
7. The method according to claim 1, further comprising: Acquire the written data of the evaluation object, and acquire text data corresponding to the written data through optical character recognition; Inputting the text data into a feature extractor to extract text data features of the text data; The text data features are input into the multimodal fusion model to obtain the status data of the evaluation object.
8. The method according to claim 7, wherein: After obtaining the written data of the evaluation object and obtaining text data corresponding to the written data through optical character recognition, the method further includes: Performing semantic detection on the text data, and detecting and determining erroneous content in the text data according to the semantic detection result; The erroneous content is analyzed, and modification suggestions for the erroneous content are generated and fed back to the evaluation object so that the evaluation object can modify the erroneous content according to the modification suggestions.
9. The method according to any one of claims 4 to 8, wherein: The step of inputting physiological data of different modalities into a pre-trained multimodal fusion model to obtain status data of the evaluation object includes: The data features corresponding to the different modal data are input into the multimodal fusion model and fused into the same feature space, and the state data of the evaluation object is predicted according to the fusion result.
10. The method according to claim 1, wherein: The multimodal fusion model is a multi-task model, and the state data of the evaluation object includes the concentration of the evaluation object, the emotional state of the evaluation object, and the fatigue state of the evaluation object; The step of inputting physiological data of different modalities into a pre-trained multimodal fusion model to obtain status data of the evaluation object includes: Physiological data of different modalities are input into the multimodal fusion model to identify the concentration of the evaluation object, the emotional state of the evaluation object, and the fatigue state of the evaluation object.
11. A data processing device, comprising: A data acquisition module, used to acquire physiological data collected by at least one physiological sensing device worn by the assessment subject; The at least one physiological sensing device is used to collect physiological data of different modalities; The model prediction module is used to input physiological data of different modalities into a pre-trained multimodal fusion model to obtain the state data of the evaluation object.
12. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 10.
13. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-10.
14. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 10.