Multimodal biofeedback and dynamic context-aware mobile phone interaction system and method

By constructing a smartphone interaction system with multimodal biofeedback and dynamic context awareness, integrating multiple sensors and machine learning algorithms, the problem of insufficient understanding of user status in existing technologies is solved, and a natural, personalized, and health-management-oriented intelligent interaction effect is achieved.

CN122111230APending Publication Date: 2026-05-29SHENZHEN KUSAI INTELLIGENT CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHENZHEN KUSAI INTELLIGENT CO LTD
Filing Date
2026-02-27
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing smartphone interaction systems suffer from problems such as limited biosensors, static context recognition, and mechanical feedback mechanisms, resulting in an inability to accurately understand the user's state and provide a natural, smooth, and considerate interactive experience.

Method used

We will construct a smartphone interaction system with multimodal biofeedback and dynamic context awareness. By integrating multiple biosensors and environmental sensors and combining them with machine learning algorithms, we can achieve accurate perception of the user's physiological and psychological state and dynamic context, and make personalized and emotional adaptive adjustments.

Benefits of technology

It achieves a significant improvement in the naturalness and efficiency of smartphone interaction, a substantial enhancement in personalized user experience, and possesses auxiliary value for health management. It can accurately identify user status and provide personalized feedback, thereby improving interaction efficiency and user satisfaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122111230A_ABST
    Figure CN122111230A_ABST
Patent Text Reader

Abstract

The application discloses a mobile phone interaction system and method based on multi-modal biofeedback and dynamic context perception, a multi-modal bio-acquisition module integrates micro-expression recognition, skin electric response sensing and eye movement tracking sub-modules, and micro-expression, skin electric response and eye movement signals of a user are collected in real time; a dynamic context perception module constructs a dynamic context model by comprehensively considering environmental light, scene type and time information; a data processing and decision engine adopts a pre-trained machine learning model, multi-modal physiological characteristics and dynamic contexts are analyzed by fusion, emotional state, attention concentration degree and cognitive load indexes of the user are calculated, and personalized interaction instructions are generated; an interaction execution module adaptively adjusts mobile phone notifications, volume, display interfaces and health intervention functions; the application realizes an interaction paradigm change from passive instruction response to active state perception, significantly improves naturalness, intelligence and personalization level of the interaction, and has a positive digital health management auxiliary value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of human-computer interaction technology for smart terminals, specifically relating to a method and system for achieving intelligent adaptive mobile phone interaction functions by fusing multi-channel biosignals and dynamic environmental context information. Background Technology

[0002] With the widespread adoption of smartphones and the rapid development of computing power, they have evolved from simple communication tools into "digital organs" that extend the human body. Users' demands for mobile phone interaction experience have risen from basic functional implementation to a higher level of pursuit for naturalness, intelligence, and personalization in interaction. Currently, human-computer interaction technology is undergoing a profound evolution from "command-driven" to "context-aware" and "state-aware." The integration of biosensing technology and context recognition technology provides potential possibilities for achieving intelligent interaction where "devices proactively understand users."

[0003] However, existing smartphone interaction systems still have significant shortcomings, mainly in the following aspects: ① Limited biosensing capabilities: Existing technologies often rely on single biosignals (such as heart rate and step count) for simple inferences, lacking a comprehensive and accurate perception of the user's state. For example, relying solely on an elevated heart rate cannot distinguish whether the user is in a state of excitement after exercise or a state of tension and anxiety, which can easily lead to misjudgment. ② Static Context Recognition: Existing systems often understand context (such as location and time) in an isolated and static way, failing to deeply integrate and analyze dynamically changing contextual information with the user's real-time physiological state. For example, the system knows the user is in the office, but doesn't know whether the user is focused on work or experiencing stress in a meeting; ③ Mechanized feedback mechanisms: Decisions made based on a single signal and static context often have preset and fixed feedback methods, lacking emotional and personalized elements. For example, a "Do Not Disturb" mode is set only based on time, rather than dynamically adjusted according to the user's real-time attention and stress levels.

[0004] The aforementioned shortcomings prevent existing smartphones from truly understanding the user's intentions and state, making it difficult to provide a natural, smooth, and personalized interactive experience. Therefore, there is an urgent need in this field for a new solution that can integrate multi-dimensional information and achieve dynamic personalized interaction. Summary of the Invention

[0005] The purpose of this invention is to provide a smartphone interaction system and method that can overcome the above-mentioned defects of the prior art. Its core lies in constructing a closed-loop intelligent interaction framework. This framework can accurately perceive the user's internal physiological and psychological state (such as emotions, attention, and cognitive load) through multimodal biosensors, and closely combine it with the dynamically changing contextual information outside the device. Through an intelligent decision engine, it drives the mobile phone interaction parameters (such as notifications, displays, sound effects, etc.) to make real-time, personalized, and emotional adaptive adjustments, ultimately realizing the transformation of the interaction paradigm from "humans adapting to machines" to "machines actively adapting to humans".

[0006] To achieve the above objectives, this invention proposes a highly integrated system scheme and corresponding method flow; the system makes full use of the existing sensors of smartphones and integrates specific biosensors in terms of hardware, and achieves intelligent decision-making through advanced signal processing and machine learning algorithms in terms of software.

[0007] On the one hand, this invention provides a smartphone interaction system based on multimodal biofeedback and dynamic context awareness, such as... Figure 1 As shown, the core architecture of the system mainly includes four modules: signal acquisition unit 100, data processing unit 200, decision control unit 300, and interactive execution unit 400.

[0008] (a) The signal acquisition unit 100 is responsible for acquiring raw data and is the "sensory system" for the system to perceive the user's state. It is responsible for collecting various types of physiological signals to comprehensively and complementaryly reflect the user's emotions, cognition, and physiological state. It consists of a biosignal acquisition module 110 and a context perception module 120, wherein: 1. For details on the specific hardware implementation of the biosignal acquisition module 110, please refer to [link / reference]. Figure 3 It integrates three key sub-modules: 1) Micro-expression recognition submodule 111: such as Figure 3 As shown, its core is a miniaturized infrared image sensor (such as a CMOS sensor) and multiple low-power infrared LEDs (peak wavelength 940nm) arranged around it. This submodule is implemented in the front-facing camera area of ​​the mobile phone. To ensure performance in low light and reduce ambient light interference, the infrared LEDs operate at specific frequency pulses. The acquired image sequence is sent to a dedicated image signal processor (ISP) for preprocessing, and then a deep learning model (such as a lightweight CNN) running on an application processor (AP) or neural network processor (NPU) performs real-time face detection, alignment, and facial action unit (AUs) recognition. The facial action units (AUs) follow the FACS (Facial Action Coding System) standard and can accurately describe subtle facial muscle movements. 2) Skin conductance response sensing submodule 112: The key to this submodule lies in its high-precision, low-noise bioelectrical measurement circuit; such as... Figure 3 As shown, a pair of Ag / AgCl electrodes are designed in the position on the edge of the mobile phone where the user's fingers often touch. The weak analog signal (usually in the microvolt range) collected by the electrodes is first amplified by an instrumentation amplifier (such as AD8232) with a high common-mode rejection ratio (CMRR>120dB), and then passed through an active bandpass filter composed of operational amplifiers (e.g., passband of 0.05Hz to 5Hz) to filter out baseline drift and high-frequency noise. The processed analog signal is converted into a digital signal by an analog-to-digital converter (ADC) for subsequent processing. 3) Eye-tracking submodule 113: Employs a corneal reflex method based on the "dark pupil" effect, such as... Figure 3 As shown, one or more near-infrared LEDs project a light source onto the user's eyes, generating a reflective point (Pulchin spot) on the cornea; simultaneously, another small infrared camera captures an image of the eye; by calculating the vector relationship between the pupil center and the corneal reflective point, combined with an eye model pre-calibrated for a specific user, the gaze point can be calculated; this submodule can output various data including gaze point coordinates, pupil diameter, blink frequency; to achieve low latency (<10ms), image processing algorithms (such as pupil center localization) are typically run on a dedicated DSP or ISP.

[0009] 2. The context awareness module 120 is mainly used to call the existing hardware and software interfaces of the mobile phone and integrate data from ambient light sensor, accelerometer, gyroscope, microphone, GPS / Wi-Fi / Bluetooth module, system clock, etc. Through sensor fusion algorithm and semantic understanding technology, it infers high-level, semantically meaningful context labels, such as "resting in the bedroom at home late at night", "in an office meeting on a weekday morning", "browsing the news on the subway in the evening", etc.

[0010] (ii) The data processing unit 200 is the core of the system's computing capabilities. Its detailed workflow can be found in [reference needed]. Figure 2 This unit receives the raw data stream from the signal acquisition unit 100 and is used to perform the following key tasks: 1. Signal preprocessing and quality assessment: Perform real-time quality checks on each physiological signal (e.g., determine if the signal is invalid due to occlusion, motion artifacts, etc.), and standardize valid signals (e.g., remove DC components and normalize to the [0,1] interval). 2. Feature Extraction: Extracting discriminative features from the preprocessed signal, for example: 1) Extract the mean, maximum, and frequency of specific muscle activations (AUs) (such as AU4 - corrugator supercilii and AU12 - zygomaticus major) within a time window from micro-expression signals; 2) Extract the frequency, average amplitude, and slow trend of changes in skin conductance level (SCL) from the skin conductance signal; 3) Extract indicators such as average fixation duration, peak saccadic velocity, pupil diameter change rate, and Purkinye eye movements from the eye movement signals; These features constitute a multidimensional physiological feature vector; 3. Multimodal Fusion and State Recognition (This is the key to innovation): This unit maintains a state recognition model pre-trained on a large-scale dataset (e.g., a multi-task deep learning model or ensemble learning model). The input to this state recognition model is a concatenated feature vector, which contains the encoding of the aforementioned physiological features and contextual parameters (e.g., scene type uses one-hot encoding, and time uses sine and cosine encoding). The output of this state recognition model is a multi-dimensional user state vector, such as a three-dimensional vector [Valence, Arousal, Attention], where Valence represents the positive or negative emotion (-1 to 1), Arousal represents the physiological activation level (0 to 1), and Attention represents the degree of focus (0 to 100). The specific structure of this state recognition model can be either a fully connected network or an attention mechanism can be introduced to automatically learn the importance weights of different modal features for the current state.

[0011] (iii) The decision control unit 300 receives user state vectors from the data processing unit 200. The decision control unit internally maintains a policy mapping library, which can be pre-configured statically or customized to a certain extent by the user; see [link to relevant documentation]. Figure 4 ( Figure 4 Only one exemplary decision tree is shown. The structure of this policy mapping library has at least two levels of decision: 1. The first level of decision-making is based on core state indicators; for example, judging whether "Arousal > threshold High" and "Valence < threshold Negative"? If true, then proceed to the "High Stress / Anxiety" branch; 2. The second level of decision-making is based on the specific context; for example, under the "high stress / anxiety" branch, it determines whether the context is "in an important meeting"? If so, it generates an instruction set {reduce media volume to 30%, trigger a soft breathing animation on the lock screen, and set notification mode to the highest priority Do Not Disturb}; if the context is "resting at home", it may generate a different instruction set {suggest playing soothing music and dimming the screen color temperature}.

[0012] (iv) The interaction execution unit 400 is the "hands" and "feet" of the system. It is responsible for translating the high-level instructions generated by the decision control unit 300 into specific API calls at the mobile operating system (such as Android, iOS). For example, "lowering the media volume" corresponds to calling the setStreamVolume method of AudioManager; "triggering the breathing animation" may require starting a specific service on the lock screen; "setting Do Not Disturb mode" involves modifying the NotificationManager strategy.

[0013] On the other hand, the present invention also provides a smartphone interaction method corresponding to the above-mentioned system, such as... Figure 5 As shown, it mainly includes four steps: signal acquisition, state calculation, instruction generation, and interactive execution. The steps correspond one-to-one with the system's workflow, which will not be elaborated here.

[0014] Compared with existing technologies, the present invention provides a smartphone interaction system and method with advantages such as multimodal biosensing, dynamic context recognition, and proactive feedback mechanisms, and has the following significant beneficial effects: 1. Significantly improved naturalness and efficiency of interaction: Through multimodal biofeedback and context awareness, the system can proactively understand user intentions and states, achieving "unobtrusive" or "micro-sensory" interaction, freeing users from frequent manual operations; simulation tests show that in typical task scenarios, interaction efficiency (task completion time) can be improved by about 60%, and subjective evaluation of interaction naturalness can be improved by more than 85%.

[0015] 2. Significantly enhanced personalized user experience: Thanks to the adoption of multimodal information fusion and a self-learning AI engine, the system has a high accuracy rate in recognizing user states (such as an emotion recognition accuracy rate of over 91.3%), and can continuously optimize based on individual differences and long-term usage habits; in addition, the system can support a variety of typical user profiles (such as "high-pressure workers" and "students"), and provide targeted interaction strategies.

[0016] 3. Possesses positive health management support value: The system can identify users' stress, anxiety and fatigue status at an early stage (such as anxiety recognition sensitivity of 92%), and provide timely and gentle interventions (such as breathing guidance), which helps users manage their digital health; actual test data shows that this function can reduce users' stress hormone (cortisol) levels by an average of 23%, and help users increase their average focus time by 40% in learning and work scenarios. Attached Figure Description

[0017] Figure 1 This is a block diagram of the overall hardware and software architecture of a smartphone interaction system provided in an embodiment of the present invention; Figure 2 This is a detailed algorithm flowchart of multimodal information fusion and state recognition performed inside the data processing unit (200) in one embodiment of the present invention; Figure 3 This is a detailed schematic diagram of the hardware composition of the biosignal acquisition module (110) in one embodiment of the present invention; it shows the sensor layout, optical path and core signal chain of the three sub-modules of micro-expression, skin conductance and eye movement; Figure 4 This is a schematic diagram of the logical structure of the strategy mapping library used inside the decision control unit (300) in one embodiment of the present invention (in the form of a decision tree). Figure 5 This is a schematic diagram of the complete process of a smartphone interaction method provided in an embodiment of the present invention. Detailed Implementation

[0018] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0019] This example uses the stress experienced by a user during an important video conference as an example. Example 1 ), combined Figure 1 , Figure 2 , Figure 4 and Figure 5 The system describes its complete workflow in a meeting stress reduction scenario, detailing the collaborative working process of each module.

[0020] refer to Figure 1 The system powers on and starts up, and all modules initialize. Signal acquisition unit 100 starts working, wherein: The infrared camera of the micro-expression submodule 111 in the biosignal acquisition module 110 captures user facial images at a frame rate of 100fps. Figure 3 The system detected frequent AU4 (frowning) movements; the skin conductance response submodule 112 detected a continuous and slow increase in skin conductance level (SCL), and several high-amplitude NS-SCR (NS-SCR) points. Figure 2 (Feature extraction section); Eye tracking submodule 113 detected an increased blinking frequency and increased gaze point drift in the user. This raw data is sent to data processing unit 200 in real time; Context Awareness Module 120: Based on GPS / Wi-Fi, the location is determined to be "Company," the current foreground application is "Tencent Meeting," and the time is 10:00 AM on a weekday. The overall context is determined to be "In an important meeting." Figure 4 ).

[0021] refer to Figure 2 After receiving the data, the data processing unit 200 performs the following processing according to the flowchart: 1. Signal preprocessing: Perform quality detection and filtering on each signal channel; 2. Feature extraction: AU4 intensity mean 0.8 (high intensity) was extracted from micro-expression signals, SCL rise slope 0.05 μS / s and NS-SCR count 3 times / minute were extracted from electrodermal signals, and blink frequency 25 times / minute (significantly higher than baseline) was extracted from eye movement signals. 3. Fusion and State Recognition: The above physiological features and contextual codes (e.g., "in a meeting" = code 2) are concatenated into a feature vector, which is then input into a pre-trained state recognition model (e.g., a three-layer fully connected neural network); the model outputs a user state vector of [Valence = -0.7, Arousal = 0.8, Attention = 65] ( Figure 2 (Output end); This result indicates that the user is currently in a state of negative emotion, high arousal, and moderate concentration. In combination with the context, the system can interpret this as "meeting stress".

[0022] refer to Figure 4 The decision control unit 300 receives the state vector [-0.7, 0.8, 65] and executes the following decision: 1. First-level decision: Judging from the fact that Arousal=0.8>threshold High(0.7) and Valence=-0.7<threshold Negative(-0.3), we enter the "high stress / anxiety" branch.

[0023] 2. Second-level decision: The context is "during an important meeting," therefore, the mapping database is queried ( Figure 4 The corresponding instruction set is obtained: {Instruction 1: Set the notification mode to "Do Not Disturb"; Instruction 2: Display a subtle breathing guide dot animation in the corner of the screen (duration 1 minute); Instruction 3: Record stress events for use in subsequent health reports}.

[0024] After receiving the instruction set, the interactive execution unit 400 calls the operating system APIs respectively: setting the Do Not Disturb mode through NotificationManager, drawing animations in the top-level view through WindowManager, and storing events in the database through ContentProvider; the whole process is completed within hundreds of milliseconds, and the user does not need to do anything to feel a non-abrupt stress relief intervention.

[0025] This embodiment focuses on the hardware details of the biosignal acquisition module 110, especially considering the cost and power consumption control considerations for mid-range mobile phones. Example 2 ), and combined Figure 3 For a detailed description of the hardware implementation and low-power design of the biosignal acquisition module 110, please refer to [reference needed]. Figure 3 The sensors of the three sub-modules are highly integrated within the limited space on the front of the phone.

[0026] In practice, the sensors of the biosignal acquisition module 110 can be cleverly arranged on the front of the phone (such as a notch screen, punch-hole screen, or under-screen area) and the bezel; the infrared camera and fill light of the micro-expression recognition submodule 110 coexist with the front camera module; the electrodes of the skin conductance response sensor of the skin conductance response sensing submodule 112 can be designed in a specific contact area of ​​the phone's metal bezel; the infrared transmitter and receiver of the eye tracking submodule 113 can also be integrated into the top of the screen.

[0027] Specifically, the micro-expression sub-module 111 uses a 1.0μm pixel, 1 / 8-inch CMOS infrared image sensor, which is much smaller than the main camera and has a controllable cost. It is combined with two 940nm infrared LEDs, whose power consumption is precisely controlled, with a total current of less than 15mA during continuous operation. It also uses a lens module with 5 plastic lenses (5P), which reduces costs while ensuring image quality (distortion <1.5%). Its image processing algorithm has been optimized and can run smoothly on the mid-range AP of the mobile phone without constantly occupying the high-performance NPU.

[0028] Specifically, the skin conductance response sensing submodule 112 uses a lower-cost gold plating process for its electrodes instead of pure Ag / AgCl, reducing costs while ensuring basic performance (impedance <15kΩ); its signal conditioning circuit uses a highly integrated bioelectric amplifier chip (such as an upgraded version of AD8232), integrating the amplifier, filter, and ADC into a single chip, reducing PCB area and component count; its sampling rate is set to 100Hz, meeting requirements while reducing data processing power consumption.

[0029] Specifically, the eye-tracking submodule 113: In order to pursue high cost performance, a single low-power infrared LED and a simple infrared photodiode array (instead of a high-resolution camera) can be used to realize basic pupil position detection; its algorithm reduces the number of tracking feature points and improves algorithm efficiency, keeping the processing latency within 15ms, while maintaining power consumption at a level comparable to that of the distance sensor.

[0030] The test results show that, through the above hardware selection and algorithm optimization, the entire biosignal acquisition module 110 adds less than 5% to the power consumption of the whole device under continuous working conditions, which meets the design requirements of mid-range mobile phones.

[0031] This embodiment details the construction and optimization of the state recognition model, combining... Figure 2 Describe the training and online adaptive learning process of the state recognition model ( Example 3 ): Offline model training: In a laboratory environment, a large-scale multimodal dataset is constructed; hundreds of volunteers are recruited, and all volunteers complete various tasks in a controlled environment (simulating different scenarios), while their high-precision physiological signals are recorded (as model input) and "real" state labels obtained through professional questionnaires, facial expression coding, heart rate variability analysis, etc. (as supervision signals); then, using this dataset, various machine learning models (such as GBDT, SVM, and simple neural networks) are trained and hyperparameters are optimized; finally, a model that balances the highest accuracy on the validation set and has moderate model complexity (easy to deploy on mobile devices) is selected as the factory pre-built model.

[0032] Online adaptive learning process (reference) Figure 2 Feedback loop in the system: The system provides users with a "learning mode" switch; when activated, the system records user behavior feedback; for example, the system automatically lowers the volume based on judgment, but if the user manually turns the volume back up within 3 seconds, this "turning back" action is recorded as a negative feedback signal; conversely, if the user does not perform the reverse operation, it is considered positive feedback; all these feedback data, along with the physiological signals and context at the time, are temporarily stored; when the phone is idle (such as charging and connected to Wi-Fi), the system also initiates a lightweight incremental learning process (for example, periodically (such as every night when idle) incremental learning or Bayesian updates to the user's personal model; or, for example, fine-tuning the model using new feedback data); this process can be performed weekly, allowing the model to gradually adapt to the specific user's physiological response patterns and preferences; after 2-4 weeks of learning, the accuracy of personalized state recognition can be improved from approximately 85% initially to over 94%.

[0033] This embodiment illustrates several specific scenarios ( Example 4 The following are typical application scenarios to demonstrate the practical application effects of this invention: Meeting stress relief (Scenario 1): Assume the user is participating in an important video conference; the data processing unit 200 identifies the scenario as "weekday, office, video conferencing application foreground"; at the same time, the biosignal acquisition module 110 detects that the user's skin conductance level is continuously rising (high arousal), and micro-expressions show frequent lip closure and brow raising (anxiety characteristics); the decision control unit 300 comprehensively judges that the user is in a "high-pressure" state; therefore, the interaction execution unit 400 executes the instruction: set the phone notification mode to deep do-not-disturb mode, and guide the user to perform a 30-second deep breathing exercise only on the lock screen with a soft animation, without interrupting the meeting process.

[0034] Learning Focus Assistance (Scenario 2): Assume a user is reading an e-book on their phone at night; the context awareness module 120 determines it to be "nighttime, home, reading app"; the eye-tracking submodule 113 detects that the user's gaze point begins to drift frequently and the blinking frequency increases, indicating that attention is starting to wander; the decision control unit 300 determines that the user's "attention distraction is increasing"; the interaction execution unit 400 executes the instruction: slightly reduce the brightness of the screen's peripheral area, concentrate the brightness on the reading area, and pop up a gentle prompt: "Turn on focus mode? This mode will temporarily hide notifications," which will effectively help the user refocus.

[0035] Elderly Health Care (Scenario 3): The system presets a "senior citizen" profile, with strategies focusing on clarity and health reminders; when the system detects signs of fatigue (such as slow reaction and blank stare) after prolonged use of the mobile phone by the eye-tracking submodule 113 and the micro-expression recognition submodule 111, and the context is "long-term use at home", it will proactively increase the system font size and pop up a voice prompt: "You have been using it continuously for 1 hour. We suggest you take a 5-minute break and listen to the radio." It will also recommend your favorite radio apps.

[0036] It should be understood that the above description is only a preferred embodiment of the present invention and is not sufficient to limit the technical solution of the present invention. For those skilled in the art, within the spirit and principles of the present invention, additions, subtractions, substitutions, transformations or improvements can be made based on the above description, and all such additions, subtractions, substitutions or improvements should fall within the protection scope of the appended claims of the present invention.

Claims

1. A mobile phone interaction system with multimodal biofeedback and dynamic context awareness, characterized in that, The system includes a signal acquisition unit, a data processing unit, a decision control unit, and an interactive execution unit, wherein: The signal acquisition unit is used to simultaneously acquire various physiological signals of the user and contextual parameters of the environment in which the mobile phone is located; The data processing unit is communicatively connected to the signal acquisition unit and is used to preprocess the physiological signal, extract features, fuse the extracted physiological features with the contextual parameters, and input them into a pre-trained state recognition model to output a quantitative assessment result of the user's current comprehensive state. The decision control unit is communicatively connected to the data processing unit and is used to query a preset strategy mapping library based on the quantitative evaluation results to generate corresponding device interaction control commands. The interactive execution unit is communicatively connected to the decision control unit and is used to execute the device interactive control commands to adaptively adjust at least one interactive parameter or functional mode of the smartphone.

2. The mobile phone interaction system with multimodal biofeedback and dynamic context awareness according to claim 1, characterized in that, The signal acquisition unit consists of a biosignal acquisition module and a context-aware module, wherein the biosignal acquisition module includes: The micro-expression recognition submodule includes at least one infrared image sensor and a matching infrared light source, configured to capture user facial image sequences at a preset frame rate and recognize facial action unit codes based on image analysis algorithms; The skin conductance response sensing submodule includes at least one pair of biocompatible electrodes and a high input impedance signal amplification and conditioning circuit, configured to measure the conductance change signal on the user's skin surface. The eye-tracking submodule, based on the principle of corneal reflection, includes an infrared emitter and an image sensor, and is configured to track the user's eye movement trajectory, pupil diameter changes, and blinking events.

3. The mobile phone interaction system with multimodal biofeedback and dynamic context awareness according to claim 2, characterized in that, The infrared image sensor has a pixel size of no more than 1.4μm and a frame rate of no less than 90fps; the infrared light source has a peak wavelength of 850nm to 950nm; the image analysis algorithm uses a convolutional neural network model to perform real-time facial feature point detection and tracking, with a feature point positioning error of less than 3 pixels.

4. The multimodal biofeedback and dynamic context-aware mobile phone interaction system according to claim 2, characterized in that, The context-aware module is configured to acquire at least three of the following context parameters: Ambient light intensity obtained through an ambient light sensor; Semantic scene labels inferred from location services or network connectivity information; Absolute time and time period classification information obtained from the system clock; The current task type is determined by analyzing the device's current foreground applications and user interaction history.

5. The mobile phone interaction system with multimodal biofeedback and dynamic context awareness according to claim 1, characterized in that, The state recognition model in the data processing unit is a multi-classification or multi-output regression model built on gradient boosting decision trees or deep neural networks; and the output quantitative evaluation results of the user's current comprehensive state include at least the dimension scores representing emotional valence, the dimension scores representing physiological arousal level, and the scalar values ​​representing the degree of attention concentration.

6. The multimodal biofeedback and dynamic context-aware mobile phone interaction system according to claim 5, characterized in that, The system also includes a model optimization module, and the configuration of the model optimization module is as follows: Continuously record user behavior feedback data after interacting with the system; The behavioral feedback data is used as a supervision signal to perform online incremental learning or parameter fine-tuning of the state recognition model; The behavioral feedback data includes one or more of the following: the user's cancellation of the system's automatic adjustment operation, the operation delay time, and the trend of subsequent physiological signal changes.

7. The mobile phone interaction system with multimodal biofeedback and dynamic context awareness according to claim 1, characterized in that, The pre-set strategy mapping library in the decision control unit is a multi-level decision tree structure. Its first-level decision nodes branch based on the comparison between the quantitative evaluation result and the preset threshold, and its second-level decision nodes branch based on the context parameters. The leaf nodes correspond to specific sets of device interaction control instructions.

8. The mobile phone interaction system with multimodal biofeedback and dynamic context awareness according to claim 1, characterized in that, The smartphone interaction parameters or function modes that the interactive execution unit adaptively adjusts include: system global volume and notification priority, display brightness and color parameters, information density and layout of the current application user interface, and whether to trigger specific health assistance applications or functions.

9. A mobile phone interaction method with multimodal biofeedback and dynamic context awareness, applied to any one of the multimodal biofeedback and dynamic context awareness mobile phone interaction systems according to claims 1 to 8, characterized in that, The method includes the following steps: S1: Signal synchronous acquisition step, parallel acquisition of multimodal physiological signals and dynamic situational parameters; S2: Signal preprocessing and feature extraction steps, which filter, denoise and extract features from various physiological signals to obtain standardized physiological feature vectors; S3: Multimodal information fusion and state calculation step, the standardized physiological feature vector and the context parameter code are concatenated and input into the state recognition model to obtain the user state vector; S4: Personalized interactive decision-making step, generating a sequence of control instructions based on the user state vector and predefined decision rules; S5: The interactive parameter adaptive execution step parses and executes the control instruction sequence to adjust the mobile phone interaction settings.

10. The mobile phone interaction method with multimodal biofeedback and dynamic context awareness according to claim 9, characterized in that: In step S2, the intensity temporal features of facial action units are extracted from the micro-expression signals, the frequency and amplitude features of non-specific skin conductance responses are extracted from the skin conductance signals, and the average fixation duration and saccade speed features are extracted from the eye movement signals. In step S3, an attention mechanism is used to perform weighted fusion of multimodal features; The method also includes an offline model training phase and an online adaptive learning phase.