Input device interaction method and apparatus

By hashing input device operation events, content and action information are separated, and irreversible hash values ​​are generated for emotion perception, thus solving the risk of information leakage and achieving secure emotion perception and interaction.

CN121349331BActive Publication Date: 2026-03-10BEIJING QIMIAO KINGDOM TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-18
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

Existing technologies pose a risk of information leakage when collecting input device information and transmitting it to the cloud for analysis during human-computer interaction, leading to threats to user information security.

Method used

By parsing the operation events of the input device, the content information and action information are separated, and a hash operation is performed on the content information to generate an irreversible hash value, which is used for emotion perception without revealing the user's input content. Emotion perception and interaction are performed by combining preset content and action information.

Benefits of technology

It enables accurate perception of user emotions and interaction without disclosing user input, thereby improving user information security and the naturalness of interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121349331B_ABST
    Figure CN121349331B_ABST
Patent Text Reader

Abstract

This application provides an input device interaction method and apparatus. The method includes: parsing operation events targeting the input device to obtain content information and first action information; performing a hash operation on the content information to obtain multiple first hash values ​​corresponding to characters; determining preset content included in the content information in response to the multiple first hash values, including a second hash value; performing emotion perception based on the preset content, the second action information, and the first action information to obtain an emotion perception result; and controlling the input device to perform interaction based on a preset interaction mode corresponding to the emotion perception result. This application enables the prediction of a user's emotions while ensuring that the user's input content information is not leaked.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to interactive technologies, and more particularly to an input device interaction method and apparatus. Background Technology

[0002] With the rapid integration of smart terminals and IoT technology, human-computer interaction is evolving from function-driven to emotion-driven. This means that users not only need devices to complete basic operations, but also expect them to be able to sense emotional states and adaptively adjust the environment to alleviate stress, anxiety and other emotional problems.

[0003] In related technologies, the content input by input devices is collected and sent to the cloud for analysis. Although this can detect the user's emotions, it also poses a risk of information leakage. Summary of the Invention

[0004] This application provides an input device interaction method, apparatus, computer-readable storage medium, and computer program product, which can perceive the user's emotions while protecting user information and interact based on the perceived emotions.

[0005] The technical solution of this application embodiment is implemented as follows:

[0006] This application provides an input device interaction method, the method comprising:

[0007] The operation event for the input device is parsed to obtain content information and first action information. The content information is used to characterize the characters included in the operation event, and the first action information is used to characterize the start time and end time of the operation in the operation event. The content information and the action information have a first correspondence relationship.

[0008] A hash operation is performed on the content information to obtain multiple first hash values ​​corresponding to the character, wherein the first hash value has a second correspondence with the first action information, and the second correspondence is determined based on the first correspondence.

[0009] In response to multiple first hash values ​​including second hash values, preset content included in the content information is determined, wherein the second hash value is a hash value obtained by performing a hash operation on the preset content, the second hash value has a third correspondence with the second action information, and the second action information is information in the first action information;

[0010] Emotion perception is performed based on the preset content, the second action information, and the first action information included in the content information to obtain an emotion perception result;

[0011] Based on the preset interaction mode corresponding to the emotion perception result, the input device is controlled to perform interaction.

[0012] This application provides an input device interaction apparatus, including:

[0013] The acquisition module is used to parse operation events for the input device to obtain content information and first action information, wherein the content information is used to characterize the characters included in the operation event, and the first action information is used to characterize the start time and end time of the operation in the operation event, and the content information and the action information have a first correspondence relationship;

[0014] A conversion module is used to perform a hash operation on the content information to obtain multiple first hash values ​​corresponding to the character, wherein the first hash value has a second correspondence with the first action information, and the second correspondence is determined based on the first correspondence.

[0015] The detection module is configured to, in response to a plurality of first hash values ​​including a second hash value, determine the preset content included in the content information, wherein the second hash value is a hash value obtained by performing a hash operation on the preset content, the second hash value has a third correspondence with the second action information, and the second action information is information in the first action information;

[0016] The perception module is used to predict emotions based on the preset content, the second action information, and the first action information included in the content information, and to obtain an emotion perception result.

[0017] The interaction module is used to control the input device to interact based on the preset interaction mode corresponding to the emotion perception result.

[0018] In the above scheme, the perception module is also used to parse the preset content, the second action information and the first action information to obtain feature vectors with multiple set dimensions;

[0019] The feature vectors of the multiple defined dimensions are combined to obtain the fused vector;

[0020] Regression is performed based on the fusion vector to obtain the emotion perception result.

[0021] The perception module is further configured to analyze the preset content, the second action information, and the first action information to obtain set statistical features, wherein each preset content corresponds to a set weight;

[0022] Calculate the deviation of the statistical feature from a pre-set reference statistical feature;

[0023] Based on the second action information, the number of times the preset content appears within the first time window is determined, and the content score of the operation event is calculated based on the weight corresponding to the preset content and the number of times it appears;

[0024] The deviation value and the content score are used as feature vectors with multiple defined dimensions.

[0025] In the above scheme, the sensing module is further configured to obtain the reference statistical features by performing the following processing:

[0026] In response to a detection command, the operation event within the second time window is detected, and a reference statistical feature with the same dimension as the statistical feature is parsed based on the operation event within the second time window, wherein the length of the second time window is greater than or equal to the length of the first time window.

[0027] In the above scheme, the sensing module is further configured to determine the duration of the first time window by performing the following processing:

[0028] The character input speed of the input device is determined based on the first action information;

[0029] In response to the character input speed of the input device being lower than a set first speed threshold, the duration of the first time window is increased;

[0030] In response to the character input speed of the input device being higher than or equal to a set second speed threshold, the duration of the first time window is reduced.

[0031] In the above scheme, the deviation value in the feature vector of the perception module includes error score, deletion score, and rhythm deviation score. The error score is determined based on the number of erroneous characters in the input and the number of erroneous characters in the reference input. The deletion score is determined based on the length of the deleted characters and the length of the reference deleted characters. The rhythm deviation score is determined based on the character input speed and the reference character input speed. The number of erroneous characters in the input and the length of the deleted characters are determined based on parsing the preset content, the second action information, and the content information. The character input speed is determined based on the first action information.

[0032] The regression based on the fusion vector to obtain the emotion perception result includes:

[0033] The error score, deletion score, and rhythm deviation score are summed to obtain a first score vector. The sum of the first score vector and the content score is used as a second score vector. The emotion confidence score is obtained by regression based on the second score vector using a normalized exponential function. The emotion confidence score is used as the emotion perception result.

[0034] In the above scheme, the perception module is further configured to set the emotion confidence level to a preset value in response to the content score being 0 and the rhythm deviation score being lower than a set rhythm deviation score threshold.

[0035] In the above scheme, the sensing module is further configured to determine the weight by performing the following process:

[0036] The weight corresponding to each of the preset contents is preset;

[0037] In response to the detection that the deviation between the occurrence count of the first word and the reference occurrence count of the first word in the reference statistical features is greater than a set deviation threshold within multiple consecutive first time windows, the weight corresponding to the first word is increased, wherein the first word is a word in the preset content.

[0038] In the above scheme, the acquisition module is also used to record the current time of the input device and the host time in response to the detection of an operation event;

[0039] The difference between the host time and the current time of the input device is embedded into the first frame of the operation event to obtain a new operation event;

[0040] The new operation event is parsed to obtain the content information and the first action information.

[0041] In the above scheme, the interaction module is further used to select the current situation from a plurality of preset situations in response to the operation events of the input device, wherein each situation is configured with a plurality of preset interaction modes, and each preset interaction mode has a third correspondence with the emotion perception result;

[0042] From the multiple preset interaction modes corresponding to the current situation, determine the preset interaction mode corresponding to the emotion perception result.

[0043] This application provides an electronic device, the electronic device comprising:

[0044] Memory is used to store executable instructions or computer programs.

[0045] The processor, when executing computer-executable instructions or computer programs stored in the memory, implements the input device interaction method provided in the embodiments of this application.

[0046] This application provides a computer-readable storage medium storing a computer program or computer-executable instructions, which, when executed by a processor, implements the input device interaction method provided in this application.

[0047] This application provides a computer program product, including a computer program or computer executable instructions. When the computer program or computer executable instructions are executed by a processor, they implement the input device interaction method provided in this application.

[0048] The embodiments of this application have the following beneficial effects:

[0049] By parsing the operation events of the input device, the information contained in the operation events is divided into content information and action information. A hash operation is performed on the content information to desensitize it. Then, by comparing the desensitized first hash value with a pre-stored second hash value, it is possible to verify whether the user has entered the set content without revealing the user's input content information. Furthermore, based on the user's input content and the action information when entering the set content, the user's emotion is predicted, thus achieving the prediction of the user's emotion while ensuring that the user's input content information is not disclosed. Attached Figure Description

[0050] Figure 1 This is a schematic diagram of the input device interaction system architecture provided in the embodiments of this application;

[0051] Figure 2 This is a schematic diagram of the structure of the device provided in the embodiments of this application;

[0052] Figure 3 This is a flowchart illustrating the input device interaction method provided in the embodiments of this application. Figure 1 ;

[0053] Figure 4 This is a flowchart illustrating the input device interaction method provided in the embodiments of this application. Figure 2 ;

[0054] Figure 5 This is a flowchart illustrating the input device interaction method provided in the embodiments of this application. Figure 3 ;

[0055] Figure 6 This is a flowchart illustrating the input device interaction method provided in the embodiments of this application. Figure 4 ;

[0056] Figure 7 This is a flowchart illustrating the input device interaction method provided in the embodiments of this application. Figure 5 ;

[0057] Figure 8 This is a flowchart illustrating the input device interaction method provided in the embodiments of this application. Figure 6 ;

[0058] Figure 9This is a flowchart illustrating an input device interaction method in an application scenario provided in an embodiment of this application.

[0059] It should be noted that the terms "first" and "second" mentioned above are only used to distinguish between different options and do not represent the degree of superiority or inferiority of the options or their priority in the implementation process. Detailed Implementation

[0060] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0061] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0062] In the following description, the terms "first, second, third" are used merely to distinguish similar objects and do not represent a specific ordering of objects. It is understood that "first, second, third" may be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.

[0063] In the embodiments of this application, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.

[0064] Unless otherwise defined, all technical and scientific terms used in the embodiments of this application have the same meaning as commonly understood by one of ordinary skill in the art. The terminology used in the embodiments of this application is for the purpose of describing the embodiments of this application only and is not intended to limit this application.

[0065] In the implementation of this application, the collection and processing of relevant data should strictly comply with the requirements of relevant laws and regulations, obtain the informed consent or separate consent of the personal information subject, and carry out subsequent data use and processing within the scope of laws and regulations and the authorization of the personal information subject.

[0066] Before providing a further detailed description of the embodiments of this application, the nouns and terms involved in the embodiments of this application will be explained, and the nouns and terms involved in the embodiments of this application shall be interpreted as follows.

[0067] 1) Hash operation: This refers to the process of converting input data into a fixed-length hash value using a hash function. This operation is deterministic, meaning that the same input will always produce the same output. It is typically designed for efficient computation and collision resistance, making the probability of different inputs producing the same output extremely low. Hash operations can be used in scenarios such as data integrity verification, cryptographic security, and hash table retrieval. For example, common algorithms such as SHA-256 or MD5 are widely used in file verification, digital signatures, and distributed systems.

[0068] 2) Operation events are events that mark a single user interaction or a series of user interactions triggered by an input device within a specific time period. These events are generated when the user interacts with the device and are the basic unit for the system to perceive and process user intent. For example, continuous keystrokes on a keyboard or rapid swipes on a touchscreen will generate corresponding operation event sequences. By parsing such event streams, the system can obtain the user's interaction content and action patterns within a specific time period.

[0069] 3) Content information refers to the actual input content contained in the operation event; this information directly corresponds to the final target data of the interaction. For example, in keyboard operations, content information is the multiple characters or symbols that the user intends to input, such as the letter "K", the number "9" or a space.

[0070] 4) Action information, used to describe the state and attributes related to the interactive behavior itself in an operation event; this information reflects the way and process by which the user performs the operation. For example, in keyboard operation, action information can be reflected in the pressing and releasing of keys, the speed of continuous input, the combination relationship of specific keys, or the duration of a long press.

[0071] 5) Emotion perception, used to infer a user's current or future emotional state based on their behavior or contextual data; this perception can be based on single-modal or multi-modal information. Its implementation includes rule-based analysis and machine learning model-based inference, and can be applied to scenarios such as emotion adaptation in human-computer interaction systems, user experience evaluation, mental health monitoring, or personalized optimization of recommendation systems.

[0072] 6) Bloom filtering is used to efficiently determine whether an element does not belong to a specific set. This structure is implemented using a binary vector and a set of hash functions, enabling fast queries with minimal storage space. Its characteristic is that it will never falsely identify existing elements (i.e., no false negatives), but it may falsely identify non-existent elements as existing (i.e., false positives are allowed). It is suitable for scenarios requiring the rapid elimination of large amounts of irrelevant data, such as spam filtering or initial database query optimization.

[0073] In related technologies, the system detects the information input by a user using an input device such as a keyboard and the pressure intensity of the keystrokes. The detected information and pressure intensity are recorded and sent to a system server. The system server then analyzes the detected information and pressure intensity records to determine the user's emotions.

[0074] However, while the above methods can predict users' emotions, there is a risk that the detected information and the stress level of the case will be intercepted or leaked by unauthorized third parties during the transmission of the data to the server, since the data usually needs to be transmitted over the network. Such information leakage may be maliciously used and pose a potential security threat to users.

[0075] Based on this, embodiments of this application provide an input device interaction method. This method is a way to recognize the user's emotions and interact through the input device while ensuring that the user's input information is not disclosed. The following describes exemplary applications of the input device provided in this application embodiment. This input device can be implemented as a laptop keyboard, desktop computer keyboard, mobile phone keyboard, touchscreen, virtual reality device, integrated MCU keyboard, standard HID keyboard, folding keyboard, touchpad, etc. Exemplary applications when the device is implemented as a keyboard will be described below.

[0076] See Figure 1 , Figure 1This is a schematic diagram of the architecture of the input device interaction system provided in this application embodiment. To perform input device interaction operations, an input device interaction application can be provided. For example, this input device interaction application can be an application specifically for handling input device interaction, or it can be a functional module in other applications (such as an input device interaction module in a keyboard environment application). The input device interaction system 100 in this application embodiment includes at least an input device 400, a network 300, and an execution terminal 200, wherein the execution terminal 200 is a device for processing or executing the content input by the input device 400. In this application embodiment, the input device 400 can constitute the input device interaction device of this application embodiment, that is, the input device interaction method of this application embodiment is implemented through the input device 400. The input device 400 is connected to the execution terminal 200 through the network 300. The network 300 can be a wide area network (WAN), a local area network (LAN), or a combination of both; it can be a wired network or a wireless network.

[0077] Users can perform interactive operations through input device 400. Upon receiving a user's interactive operation, input device 400 parses the operation event for the input device to obtain content information and first action information. The content information represents the characters included in the operation event, and the first action information represents the start and end times of the operation. A first correspondence exists between the content information and the action information. Input device 400 performs a hash operation on the content information to obtain multiple first hash values ​​corresponding to the characters. A second correspondence exists between the first hash values ​​and the first action information, determined based on the first correspondence. In response to the multiple first hash values, including the second hash value, input device 400 determines the preset content included in the content information. The second hash value is the hash value obtained by hashing the preset content. A third correspondence exists between the second hash value and the second action information, which is the information in the first action information. Input device 400 performs emotion perception based on the preset content, the second action information, and the first action information to obtain an emotion perception result. Input device 400 controls the input device to interact based on the preset interaction mode corresponding to the emotion perception result.

[0078] In some embodiments, see Figure 1The input device interaction method of this application embodiment can also be executed collaboratively by the execution terminal 200 and the input device 400. That is, the user can input an operation event through the input device 400. After the input device 400 receives the operation event, it parses the operation event for the input device to obtain content information and first action information. The content information is used to represent the characters included in the operation event, and the first action information is used to represent the start time and end time of the operation in the operation event. There is a first correspondence between the content information and the action information. The input device 400 performs a hash operation on the content information to obtain multiple first hash values ​​corresponding to the characters. There is a second correspondence between the first hash values ​​and the first action information. The second correspondence is based on the first correspondence. Based on the established relationship, the input device 400 sends a first hash value and first action information to the execution terminal 200. The execution terminal 200 detects the first hash value and the first action information. In response to multiple first hash values, including a second hash value, the execution terminal 200 determines the preset content included in the content information. The second hash value is a hash value obtained by hashing the preset content. The second hash value and the second action information have a third correspondence, and the second action information is the information in the first action information. The execution terminal 200 performs emotion perception based on the preset content, the second action information, and the first action information to obtain an emotion perception result. Based on the preset interaction mode corresponding to the emotion perception result, the execution terminal 200 controls the input device to perform interaction.

[0079] As an example, in keyboard interaction scenarios, it is often necessary to perceive and adaptively respond to the user's emotional state during typing. For instance, when a user inputs using a keyboard, the dynamics of keystrokes (such as keystroke speed, interval, and pressure) may imply emotional signals (such as tension, relaxation, or excitement). This application's embodiment parses the user's operation events on the input device into content information and first action information, and desensitizes the content information (i.e., performs a hash operation). Emotion perception is then performed based on the desensitized content information and first action information. Subsequently, the system adaptively adjusts ambient lighting (e.g., using cool-toned blue light to correspond to a relaxed state, and warm-toned red light to correspond to a tense state) and background sounds (e.g., playing soft, natural sounds to promote focus, or dynamic sound effects to enhance alertness) based on the emotion perception results, thereby improving the naturalness of the interaction and user comfort. This is suitable for businesses such as emotion-assisted office work and intelligent health management.

[0080] As an example, in touchpad interaction scenarios, such as creative design or online learning environments, it is often necessary to non-intrusively perceive and provide guided feedback on the user's focus and emotional engagement during operation. For instance, when a user uses a pressure-sensitive touchpad for drawing or interface manipulation, the smoothness of their touch trajectory, changes in pressure intensity, and frequency of pauses may reflect their level of focus (e.g., smooth and precise touch indicates immersion, while hesitation or trembling indicates distraction or frustration). This application's embodiment parses user operation events on the input device into content information and first action information, and desensitizes the content information (i.e., performs a hash operation). Emotional perception is then performed based on the desensitized content information and first action information. Subsequently, the system adaptively adjusts the ambient lighting based on the emotional perception results. For example, when signs of distraction or agitation are detected, the light tone can be gently adjusted (e.g., switched to a calming light blue) and a short, soothing ambient sound (e.g., a soft water sound) can be played to non-intrusively assist the user in refocusing, thereby improving the flow experience and efficiency during digital creation or deep learning.

[0081] In some embodiments, the electronic device may be Figure 1 Input device 400, see Figure 2 , Figure 2 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Figure 2 The illustrated electronic device includes at least one processor 410, a memory 450, at least one network interface 420, and a user interface 430. The various components of the electronic device are coupled together via a bus system 440. It is understood that the bus system 440 is used to implement communication between these components. In addition to a data bus, the bus system 440 also includes a power bus, a control bus, and a status signal bus. However, for clarity, in… Figure 2 The general labeled all buses as Bus System 440.

[0082] Processor 410 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or any conventional processor, etc.

[0083] User interface 430 includes one or more output devices 431 that enable the presentation of media content, including one or more speakers and / or one or more visual displays. User interface 430 also includes one or more input devices 432, including user interface components that facilitate user input, such as a keyboard, mouse, microphone, touch screen display, camera, other input buttons and controls.

[0084] The memory 450 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state storage, hard disk drives, optical disk drives, etc. The memory 450 may optionally include one or more storage devices physically located away from the processor 410.

[0085] The memory 450 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), and the volatile memory may be random access memory (RAM). The memory 450 described in this application embodiment is intended to include any suitable type of memory.

[0086] In some embodiments, memory 450 is capable of storing data to support various operations, examples of which include programs, modules, and data structures or subsets or supersets thereof, as illustrated below.

[0087] Operating system 451 includes system programs for handling various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, driver layer, etc., for implementing various basic business functions and handling hardware-based tasks;

[0088] The network communication module 452 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 420, exemplary network interfaces 420 including: Bluetooth, WiFi, and Universal Serial Bus (USB), etc.

[0089] The presentation module 453 enables the presentation of information (e.g., a user interface for operating peripheral devices and displaying content and information) via one or more output devices 431 associated with the user interface 430 (e.g., a display screen, a speaker, etc.).

[0090] The input processing module 454 is used to detect and translate one or more user inputs or interactions from one or more input devices 432.

[0091] In some embodiments, the apparatus provided in this application can be implemented in software. Figure 2An input device interaction device 455 stored in memory 450 is shown. This device can be software in the form of programs and plug-ins, and includes the following software modules: an acquisition module 4551, a conversion module 4552, a detection module 4553, a sensing module 4554, and an interaction module 4555. These modules are logically connected and can therefore be arbitrarily combined or further separated according to the functions they implement. The functions of each module will be described below.

[0092] In other embodiments, the apparatus provided in this application can be implemented in hardware. As an example, the apparatus provided in this application can be a processor in the form of a hardware decoding processor, which is programmed to execute the data processing method provided in this application. For example, the processor in the form of a hardware decoding processor can be one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components.

[0093] See Figure 3 , Figure 3 This is a flowchart illustrating the input device interaction method provided in the embodiments of this application. Figure 1 , will combine Figure 3 The steps shown are explained as follows: Figure 3 As shown, the method of interaction between input devices is illustrated by taking the input terminal and the execution terminal working together as the main body of the interaction as an example. The method includes the following steps 101 to 105.

[0094] In step 101, the operation events for the input device are parsed to obtain content information and first action information.

[0095] The content information is used to characterize the characters included in the operation event, and the first action information is used to characterize the start and end times of the operation in the operation event. There is a first correspondence between the content information and the action information.

[0096] Here, an input device is a physical component that provides users with raw operating instructions to a computer system. In this embodiment, it specifically refers to two types: keyboards and touchscreens. The core interactive unit of an input device is a functional area within the device that directly receives user operations and typically has a clear interactive purpose, such as the physical key array of a keyboard or the capacitive sensing layer of a touchscreen.

[0097] It's important to note that operation events are standardized data structures derived from hardware signals through system abstraction, carrying specific semantic information about user operations. Operation events can include event types and event parameters. The event type defines the essential characteristics of the interactive behavior, such as "key press" or "touch movement"; the event parameters record the specific attributes of the event, such as key encoding, touch coordinates, and timestamps. Event processing is implemented through the following mechanism: the device driver converts hardware signals into system events, which are then distributed to the application's event detector via an event queue. Common processing methods include direct callback processing and event bubbling mechanisms. Taking keyboard events as an example, when a user presses the "Ctrl+C" key combination, the system sequentially generates a "keydown" event sequence. The event parameters include the control key status identifier and the letter key value, ultimately triggering the application's copy operation.

[0098] As an example, an operation event could be generated by periodically scanning the row and column circuits of an input device and detecting changes in current or voltage in the circuits. For instance, a keyboard MCU scans the row and column circuits of a keyboard at a period of less than or equal to 1 ms, generating an event packet {scanCode, t_down, t_up} whenever a press or release is detected, where scanCode represents the keyboard scan code.

[0099] Here, action information refers to low-level behavioral metrics extracted from the raw interaction stream between the user and the input device, without carrying any semantic content. It strips away the "what was done" (such as which word was typed or which button was clicked) from the action event, focusing instead on the dynamic process of "how to perform" the action itself. Examples include the mean keystroke interval, the variance of the keystroke interval, and words per minute (local WPM).

[0100] In this embodiment of the application, parsing operation events refers to the process of separating content information with explicit interactive semantics from operation events and first action information without substantial semantics. For example, specific content information points to the user's explicit intention operation, such as which key was pressed, the order in which the keys were pressed, etc. The first action information without substantial semantics refers to physical signals generated during the operation but without explicit semantics, such as the pressure applied to the key or the frequency at which multiple keys are pressed.

[0101] In some embodiments, see Figure 4 The parsing in step 101 targets the operation events of the input device to obtain content information and first action information, which can be achieved through the following steps 1011 to 1013.

[0102] In step 1011, in response to the detection of an operation event, the current time of the input device and the host time are recorded.

[0103] Here, the host refers to the aforementioned processing device, and the host time is the processing device's time. Specifically, host time refers to the system timestamp marked by the computer's operating system kernel after the data packet of the operation event generated by the keyboard is transmitted to the motherboard of the processing device (such as a computer) through an interface (such as USB). It reflects the moment when the operating system kernel received this operation event.

[0104] Here, the current time of the input device (keyboard) refers to the timestamp marked by the keyboard's own microcontroller according to its internal clock when it detects a change in key state (such as pressing or releasing). This timestamp originates from a separate, typically low-precision clock source on the keyboard hardware, reflecting the moment when the key press occurred at the keyboard hardware level.

[0105] It should be noted that there is a small offset between the host time and the current time of the input device. This offset consists of two main parts: one is the fixed delay introduced by the internal processing of the keyboard and the scanning of the key matrix; the other is the variable delay caused by the signal waiting in the bus transmission and the operating system scheduling queue.

[0106] In step 1012, the difference between the host time and the current time of the input device is embedded into the first frame of the operation event to obtain a new operation event.

[0107] Here, since the first frame is the starting data for the handshake between the input device and the processing device, it does not contain user data, but is used to transmit metadata (such as offset). By embedding the difference between the host time and the current time of the input device in the first frame, i.e. the aforementioned offset, it is possible to avoid sending the offset in every data packet. On this basis, each data packet can be corrected for the offset, thus realizing real-time compensation for operation events.

[0108] As an example, offset The MCU can reserve bytes in the first HID packet. The embedding is accomplished in a way that allows subsequent timestamps to be corrected after the data packet is read.

[0109] In step 1013, the new operation event is parsed to obtain content information and first action information.

[0110] This application embodiment elevates low-level hardware signals into high-fidelity behavioral timing data by inserting the offset between the current time of the input device and the host time into the first frame of the HID message. It unifies all operation events to a unique and precise host timeline, enabling the subsequent perception of operation events to reliably capture these behavioral rhythms and timing characteristics. This effectively avoids "false rhythms" introduced by the deviation between the current time of the input device and the host time, as well as false anomalies caused by host load such as USB bus and OS scheduling jitter. This ensures that the extracted behavioral features (such as input speed, pressure change spectrum, and error rate timing) truly reflect the user's physiological state and emotional fluctuations.

[0111] In step 102, a hash operation is performed on the content information to obtain multiple first hash values ​​corresponding to the characters.

[0112] Among them, there is a second correspondence between the first hash value and the first action information, and the second correspondence is determined based on the first correspondence.

[0113] Here, hashing is a mathematical process that transforms input data of arbitrary length (such as a piece of text, a file, or a sequence of operation events) into a fixed-length, unique, and irreversible digital digest (i.e., a hash value) using a specific algorithm. Hash operations can include deterministic computation and collision resistance. Deterministic computation ensures that the same input always produces the same unique hash value under any circumstances, providing a reliable foundation for data identification and verification. Collision resistance means that hash algorithms are designed to make it extremely difficult to find two different inputs that produce the same output hash value, thus guaranteeing the uniqueness and security of the digest. The core function of hashing is to generate a compact representation of the original data that serves as its unique "digital fingerprint."

[0114] As an example, hash operations can be performed using the FNV-1a hash method. For instance, segments of a specific length from the content information can be used as input for the hash operation, or specific identifiers can be set in the content information. Whenever this specific identifier appears in the content information, a segment of the content information is extracted, and each segment is used as input for the hash operation. For each input, each byte is XORed with the initial hash value, and the XOR result is multiplied by a prime number base until all bytes have been processed. The resulting integer value is the first hash value obtained by the FNV-1a hash method. Alternatively, the FNV-1a hash method can be replaced with the SHA-256 truncation algorithm or threshold homomorphic encryption to obtain the fingerprint of the content information, and this fingerprint can replace the first hash value in this embodiment to determine the preset content included in the content information.

[0115] It should be noted that since the first hash value is obtained by performing a hash operation on the content information that has a first correspondence with the first action information, the first hash value and the first action information have a second correspondence.

[0116] In step 103, in response to multiple first hash values ​​including second hash values, preset content included in the content information is determined.

[0117] The second hash value is the hash value obtained by performing a hash operation on the preset content. The second hash value has a third correspondence with the second action information, and the second action information is the information in the first action information.

[0118] Here, the preset content used to determine the second hash value can be words indicating emotional intensity, such as words indicating a user is under high pressure, like "annoyed" or "working overtime." Performing a hash operation on the preset content yields a list of second hash values. A Bloom filter can be built for this list, with the first hash value as input. The Bloom filter outputs whether the first hash value includes the second hash value. If the Bloom filter outputs that the first hash value includes the second hash value, it means that the content information obtained from the operation event includes keywords indicating the user's emotional intensity (e.g., under high pressure). Keywords indicating high pressure can be obtained through cross-filtering of negative words with the highest TF-IDF from publicly available work stress vocabulary, mental health questionnaires, and industry work order corpora. In this embodiment, the preset content can also be cases indicating user input errors, such as the backspace and delete keys.

[0119] In this embodiment of the application, since the first hash value corresponds to the action information, after confirming that multiple first hash values ​​include the second hash value, it is possible to confirm the user's action information when inputting the content corresponding to the second hash value.

[0120] As an example, in a customer service consultation scenario, a user inputs "The XX function of product X is really garbage" through an input device. A hash operation is performed on the words in the input information to obtain first hash values ​​representing "product X," "of," "XX function," "really," and "garbage." Here, "garbage" is a word in the preset content, and the host has pre-determined a second hash value representing "garbage." Using Bloom filtering, the second hash value included among multiple first hash values ​​can be quickly determined. That is, by detecting the second hash value among the first hash values, it is possible to determine that the user's input information includes high-pressure words expressing emotions, i.e., the aforementioned preset content, without ensuring that the host does not interpret the user's input information. Since the input device only transmits the first hash values ​​representing "product X," "of," "XX function," and "really" to the host, due to the irreversible nature of hash operations, the host cannot obtain the user's input information by parsing the first hash value. Therefore, it is possible to perceive the user's emotions by statistically analyzing the frequency of the preset content input while protecting the user's input information from leakage.

[0121] In step 104, emotion perception is performed based on the preset content, second action information, and first action information included in the content information to obtain the emotion perception result.

[0122] In this embodiment of the application, the user's input state under normal conditions can be pre-statistically counted based on content information and first action information, i.e., relaxation baseline calibration can be performed. The relaxation baseline can be determined by requiring the user to perform free input for a set duration (such as 2-3 minutes) before the user officially performs input.

[0123] Here, based on the preset content and second action information included in the content information, it is possible to parse the frequency of user input of set content and the state when inputting specific content without transmitting the content information of the input device. Thus, by analyzing the state when the user inputs set content and the state when the user inputs other content, the user's emotional perception result can be obtained.

[0124] In some embodiments, see 5. Figure 5 The emotion perception based on preset content, second action information and first action information in step 104 is shown to be achieved by performing steps 1041 to 1043.

[0125] In step 1041, the preset content, second action information and first action information included in the content information are parsed to obtain feature vectors with multiple set dimensions.

[0126] Here, the feature vector with defined dimensions is a numerical vector formed by extracting multiple dispersed feature values ​​from the preset content, the second action information, and the first action information, and then arranging and combining these feature values ​​in an ordered manner according to a predefined, fixed-length structure. This numerical vector may include feature dimensions and numerical elements, where the feature dimension represents the specific physical or statistical meaning of each position in the vector.

[0127] As an example, the dimensions of the feature vector can be set as content score, rhythm deviation score, error rate score, and deletion score. The content score is a score determined based on the preset content information and the second action information included in the content information. The rhythm deviation score can be determined by comparing the input rhythm extracted from the first action information with the rhythm of a pre-collected user relaxation baseline. The error rate score can be determined based on the preset content and the second action information, and is obtained by counting the number of times the user discretely presses the delete or backspace key. The deletion score can be calculated by counting the length of consecutive delete key presses.

[0128] In some embodiments, see Figure 6 , Figure 6 The step 1041, which involves parsing the preset content, the second action information, the content information, and the first action information to obtain feature vectors with multiple set dimensions, can be achieved through the following steps 10411 to 10414.

[0129] 10411, parse the preset content, second action information and first action information to obtain the set statistical features.

[0130] Each preset content item has a corresponding set weight.

[0131] Here, statistical features are numerical indicators that quantify behavioral patterns, extracted from the user's interaction with input devices through mathematical statistical methods. They do not focus on the specific semantic content of the interaction events, but rather extract high-level information representing the overall trend, volatility, and rhythmic patterns of behavior from massive amounts of underlying data through aggregation and computation. Statistical features can include central tendency features, dispersion features, and distribution pattern features. Central tendency features describe the average level of behavioral data, such as the mean or median of keystroke intervals; dispersion features quantify the instability or fluctuation range of behavior, such as the standard deviation or variance of keystroke intervals; and distribution pattern features capture the rhythmic patterns of behavior, such as the operation rate per unit time (e.g., local WPM) or its variation.

[0132] As an example, a local WPM can be obtained by counting the number of characters successfully entered within a time window in the first action information and standardizing it to a rate per minute. Alternatively, the keystroke interval variance can be determined by calculating the standard deviation of the time differences between all consecutive keystrokes in a first action information. The frequency of a preset content occurrence can also be determined based on the ratio of preset content to the first action information in the second action information.

[0133] Here, the weighted index assigns a quantitative coefficient representing the emotional discrimination and intensity of each preset content. It is based on a core understanding: not all keywords in preset content are equally important; certain short phrases have a stronger indicative role and higher credibility than other words when determining a specific emotional state. For example, when preset content includes both "bad review" and "not good," the word "bad review" expresses a higher level of negative emotion; therefore, it is assigned a higher weight.

[0134] In some embodiments, the weights set in step 10411 can be determined by the following steps.

[0135] First, pre-set the weight corresponding to each preset content.

[0136] As an example, weights can be pre-set as follows. First, a sentiment lexicon is constructed, where each keyword is pre-assigned one or more weight values ​​based on its relevance to the target emotion (such as joy, anger, sadness, surprise, etc.). This process can be accomplished through expert experience, statistical models based on large-scale corpora, or machine learning model training. During actual analysis, the system retrieves keywords from the user input and obtains their corresponding weights, ultimately calculating the overall sentiment tendency score through weighted aggregation (rather than simple counting). For example, the keyword "garbage" has extremely high sentiment polarity, directly expressing a strong negative emotion, and is assigned a weight of 0.9; the keyword "disappointment" has relatively high sentiment polarity, representing a clear negative emotion, and is assigned a weight of 0.7; and the keyword "bad" has relatively weak sentiment polarity, representing a general negative evaluation, and is assigned a weight of 0.4.

[0137] Then, in response to the detection that the deviation between the number of occurrences of the first word and the reference number of occurrences of the first word in the reference statistical features is greater than a set deviation threshold within multiple consecutive first time windows, the weight corresponding to the first word is increased, where the first word is a word in the preset content.

[0138] This application embodiment obtains the deviation between the frequency of keywords (i.e., the first vocabulary) in the user's real-time input and their personal reference frequency by finding the frequency of the second hash value in the first hash value. It then dynamically adjusts the weight coefficient of each keyword in emotion recognition, effectively eliminating biases caused by different users' inherent language habits and transforming a general emotion dictionary into a personalized emotion metric. When a user's frequency of using a keyword in a specific context significantly deviates from their personal relaxation baseline, the system assigns that keyword a higher weight. This allows the system to keenly capture signals that truly reflect emotional fluctuations while filtering out background noise from personal expression styles. Furthermore, this personalized calibration provides a standardized basis for comparing emotional states across users, ensuring fair and effective emotional intensity assessments among users with different language habits. For example, in a customer service scenario, even if two customer service representatives use the same keyword at different frequencies, the system can accurately identify who experiences greater emotional stress based on their respective deviations from their personal baselines, thereby significantly improving the accuracy and practical value of emotion perception.

[0139] 10412, calculates the deviation of the statistical feature from the pre-set reference statistical feature.

[0140] Here, the deviation value is a standardized measure of the difference between statistical features and reference statistical features established in a relaxed state, used to accurately quantify the degree of deviation of the current behavioral pattern from the individual's normal state.

[0141] In some embodiments, before calculating the deviation of the statistical feature from the preset reference statistical feature in step 10412, the following steps may be performed to obtain the reference statistical feature.

[0142] First, in response to the detection command, detect the operation events within the second time window.

[0143] Here, the detection command is a clear guiding signal issued by the system to trigger and standardize the user's relaxed state data collection process. After receiving the detection command, the user can input freely or be required to input specific text.

[0144] Then, based on the operation events within the second time window, reference statistical features with the same dimensions as the statistical features are parsed.

[0145] The length of the second time window is greater than or equal to the length of the first time window.

[0146] In this embodiment, the length of the first time window is the length of the time window during which the frequency of user input of set content is statistically analyzed. Within the first time window, it is necessary to detect short-term behavioral fluctuations influenced by the user's initial emotions. Within the second time window, it is necessary to capture the user's stable and inherent long-term behavioral norms. Therefore, the length of the second time window is greater than or equal to the length of the first time window to ensure that the established baseline has sufficient statistical robustness and state representativeness, thereby providing a reliable reference anchor for short-term real-time comparisons.

[0147] This application embodiment collects operation events within a second time window and constructs a highly robust personal relaxation baseline, i.e., reference statistical features, based on the collected operation events. This effectively eliminates interference from inherent individual fluctuations. The obtained reference statistical features can provide a high signal-to-noise ratio reference benchmark for short-term detection of user operation events, thereby ensuring that the emotion perception system maintains high response speed while possessing higher judgment accuracy and personalized adaptation capabilities.

[0148] 10413. Based on the second action information, determine the number of times the preset content appears in the first time window, and calculate the content score of the operation event based on the weight and number of times the preset content appears.

[0149] It should be noted that the difference between a user's emotional state and a relaxed state is quantified by statistically analyzing the characteristics of user input, the frequency of inputting preset content, and the weight of the input preset content. Furthermore, by assigning different weights to different preset content, the system avoids confusing high-frequency but emotionally weak words with low-frequency but emotionally strong words, thus ensuring the accuracy of emotion perception.

[0150] In some embodiments, the duration of the first window can also be determined by performing the following steps.

[0151] First, the character input speed of the input device is determined based on the first action information.

[0152] Then, in response to the character input speed of the input device being lower than a set first speed threshold, the duration of the first time window is increased;

[0153] Alternatively, in response to the character input speed of the input device being higher than or equal to a set second speed threshold, the duration of the first time window is reduced.

[0154] Here, the method of adjusting the duration of the first window according to the character input speed is an optimization method that adaptively adjusts the sampling duration based on the user's real-time input speed. Its core purpose is to solve the problems of response delay and data sparsity in emotion perception with a fixed time window. By matching the analysis window with the user's information output rate, the optimal balance between timeliness and reliability in emotion perception can be achieved.

[0155] As an example, when a user types rapidly due to dissatisfaction (at a speed of 80 WPM), the system automatically shortens the detection window to 30 seconds. Once multiple high-weight keywords such as "disappointed," "terrible," and "complaint" are detected within the window, a "high dissatisfaction" alert is immediately triggered, allowing administrators to intervene promptly. When a user carefully organizes their feedback (at a speed of only 20 WPM), the system extends the window to 90 seconds to ensure sufficient keyword samples are collected before conducting an emotion assessment, avoiding misjudgments due to insufficient data.

[0156] In addition, the input speed can be divided into multiple intervals. For example, the first interval can be set to [0WPM, 30WPM], the second interval to (30WPM, 60WPM], and the third interval to (60WPM, 100WPM]. Each interval corresponds to a time window duration, and the length of the time string corresponding to the interval in which the input speed is located is selected as the duration of the first time window.

[0157] This application's embodiments adjust the time window for statistical analysis of preset user input based on the user's character input speed. This automatically adapts to different users' input habits and the same user's varying states, providing the emotion perception system with superior time-scale adaptability. It also ensures faster response when the user's emotions are strong and guarantees sufficient data for each emotion perception, thus improving the accuracy of emotion perception.

[0158] 10414 uses deviation value and content rating as feature vectors for multiple defined dimensions.

[0159] This application embodiment constructs a feature vector from multiple deviation values ​​and content scores, enabling the artificial intelligence model to automatically learn the relationship between emotional state and the statistical features and content input by the user. This allows the model to obtain a more refined emotional state by recognizing the feature vector.

[0160] In step 1042, feature vectors of multiple defined dimensions are combined to obtain a fused vector.

[0161] Here, feature vector combination refers to the process of integrating and associating feature vectors from different sources and with different properties (such as behavioral statistical deviation vectors and weighted keyword frequency vectors) to form a unified, higher-order representation that can more comprehensively and stably describe the user's emotional state. Its core purpose is to address the limitations and one-sidedness of single-dimensional features by constructing a comprehensive descriptor with higher fidelity and stronger resistance to interference through the complementarity and enhancement of multi-source information. For example, before inputting the data into the model, the values ​​of different feature vectors can be concatenated, weighted, or have their dimensionality reduced at the original data level to form a new, higher-dimensional composite feature vector.

[0162] In step 1043, regression is performed based on the fusion vector to obtain the emotion perception result.

[0163] Here, regression of the fused vector can be achieved by inputting it into a decision tree. For example, the fused vector is fed into the decision tree. The decision tree learns rules that "when high stability deviations and high-weighted confusing keywords appear simultaneously, it is a strong indicator of 'anxiety'." Therefore, the system accurately determines that the student is in an anxious state. If only behavioral features are relied upon, it may be misjudged as ordinary lack of proficiency; if only keywords are relied upon, it may be missed due to low word frequency. The fusion mechanism ensures the accuracy and robustness of perception. In this embodiment, the decision tree can be a Classification and Regression Tree (CART).

[0164] This application's embodiments combine feature vectors from different dimensions and use the resulting fused vector for emotion perception. Through the complementarity and synergy of multiple dimensions of information, it improves perception accuracy, robustness, and interpretability while protecting the specific content of user input from being leaked. This constructs a three-dimensional description of the user's emotional state, enabling the model to undergo cross-validation and effectively avoiding noise interference and misjudgments from a single data source, thus maintaining high reliability in complex real-world scenarios. Simultaneously, feature fusion allows the system to capture finer-grained emotions. For example, by recognizing different combinations of "high behavioral volatility + high-weight negative words" and "high behavioral volatility + high-weight positive words," it accurately distinguishes easily confused states such as "anger" and "excitement."

[0165] In some embodiments, see Figure 7 The deviation values ​​in the feature vector include error score, deletion score, and rhythm deviation score. The error score is determined based on the number of incorrectly input characters and the number of incorrectly input characters in the reference input. The deletion score is determined based on the length of the deleted characters and the length of the deleted characters in the reference input. The rhythm deviation score is determined based on the character input speed and the reference character input speed. The number of incorrectly input characters and the length of the deleted characters are determined based on parsing the preset content, the second action information, and the content information. The character input speed is determined based on the first action information. Based on this, step 1043 can be achieved by executing steps 10431 to 10433.

[0166] In step 10431, the error score, deletion score, and rhythm deviation score are summed to obtain the first score vector.

[0167] For example, errors are divided into ,in, The number of incorrect characters. This is the average number of incorrect characters. The standard deviation of the number of erroneous characters; deletion is divided into... ,in, The length of the characters to be deleted. This is the average length of the deleted characters. The standard deviation of the length of deleted characters; rhythm deviation is divided into ,in, The input interval is obtained from the first action information. The average value of the input interval. is the standard deviation of the input interval. Error scores, deletion scores, and rhythm deviation scores represent the deviations of the statistical characteristics of the user's current input state from the reference statistical characteristics in different dimensions. The first score vector is... ,in, , and As weight.

[0168] In step 10432, the sum of the first score vector and the content score is used as the second score vector.

[0169] As an example, the content rating is ,in, This represents the total number of characters within the first time window. The word frequency of the i-th preset content within the first time window. The weight of the i-th preset content, ,in, This represents the inverse document frequency of the i-th preset content. The second scoring vector is the user identifier factor. Inverse document frequency (IVF) indicates that the lower the frequency of a word's appearance in documents within a time window, the stronger its ability to distinguish different words, and therefore it should be given higher weight. The user identifier factor is used to characterize a user's unique emotional expression patterns. It can be obtained through historical data on a user's use of a word. For example, the stronger the emotional fluctuations accompanying a user's use of a word in historical data, the higher the intensity of the user's emotional expression expressed by that word, and therefore it should be given higher weight. Furthermore, the user identifier factor can also be manually set by the user; for example, a user could set the preset user identifier factor representing high stress to 0.7. ,in, and As weight.

[0170] In step 10433, the emotion confidence score is obtained by regression based on the second score vector using a normalized exponential function, and the emotion confidence score is used as the emotion perception result.

[0171] Here, the normalization exponential function is a mathematical transformation function that "compresses" an arbitrary real-valued K-dimensional vector into another K-dimensional probability distribution vector, such as the Sigmoid function. The Sigmoid function is a mathematical function that maps any real value to an S-shaped curve in the (0,1) interval. Its core purpose is to transform the single raw output value obtained from the model calculation of a multi-dimensional feature vector into an independent scalar representing the "probability of occurrence" or "activation level." Processed with the Sigmoid function, regardless of the input value, the output is smoothly limited to between 0 and 1, providing a perfect probability scale; and the processed value is most sensitive to input changes near zero, while saturating towards both ends, allowing it to better capture the complex relationship between features and results. In this embodiment, an independent classifier can be pre-built for each emotion to be detected (such as "joy," "anger," or "sadness"). Each classifier outputs an emotion score based on the second score vector mentioned above. This emotion score is fed into the Sigmoid function, which converts it into a confidence level for the occurrence of that emotion. The final emotion confidence level is... .

[0172] This application's embodiments normalize the combined vector into emotion types using a normalization function, modeling the determination of each emotion as an independent classification task. This enables the input terminal and execution system to identify and quantify complex psychological states where multiple emotions coexist, thus more realistically reflecting the mixed and non-mutually exclusive nature of human emotions. The Sigmoid function generates independent probability scores for each emotion, effectively capturing subtle emotional combinations such as "both joyful and moved" or "slightly anxious anticipation," breaking through the limitations of traditional mutually exclusive classification that simply categorizes emotions into a single class. Simultaneously, the probability output for each emotion provides a clear confidence reference for decision-making, allowing the system to not only know "what emotions are present" but also "the intensity of each emotion," providing richer dimensional information for subsequent interactions.

[0173] In some embodiments, before using the emotion confidence score as the emotion perception result in step 10433, the emotion confidence score is further set by performing the following steps.

[0174] In response to a content score of 0 and a pacing deviation score below the set pacing deviation score threshold, the emotional confidence level is set to the preset value.

[0175] Here, a content score of 0 indicates that no user input preset content was detected. Based on this, if the user's input rhythm deviates from the input rhythm determined under the relaxation baseline by less than a threshold, it indicates that the user's emotions are stable. There is no need to count the user's input error score and deletion score. Therefore, the emotion confidence can be set to a preset value, for example, setting the confidence to 0 indicates that the user's emotions are in a relaxed state.

[0176] This application embodiment sets the emotion confidence level to a set value when the rhythm deviation score is below a threshold and no preset content indicating negative emotions is detected from user input. This avoids running high-cost analysis in a low-risk state and improves overall computational efficiency. It effectively enhances judgment accuracy, avoids misjudgment interference caused by non-emotional factors (such as equipment failure or ordinary errors), and improves system robustness.

[0177] In step 105, the input device is controlled to interact based on the preset interaction mode corresponding to the emotion perception result.

[0178] Here, different preset interaction modes have different lighting and sound effects.

[0179] This application embodiment parses operation events on the input device into content information and first action information, and performs a hash operation on the content information to determine whether the user has entered preset content based on the hash value. This enables the perception of the user's emotions while ensuring that the specific content entered by the user is not disclosed, thereby improving the user experience.

[0180] In some embodiments, see Figure 8 The preset interaction mode corresponding to the emotion perception result can be determined by performing the following steps 106 to 107.

[0181] In step 106, for an operation event of the input device, the current scenario is selected from a plurality of preset scenarios.

[0182] Each scenario has multiple preset interaction modes, and each preset interaction mode has a third correspondence with the emotion perception result.

[0183] Here, multiple preset scenarios include situations with various emotional states, such as focused writing, public presentations, or late-night work. Users can select the current scenario from these preset scenarios by setting a shortcut key. Preset interaction modes could be, for example, that in the initial scenario, no prompt is triggered when the emotional confidence level is below 30%; a gentle prompt is triggered when the emotional confidence level is greater than or equal to 30% but less than 60%; a moderate prompt is triggered when the emotional confidence level is greater than or equal to 60% but less than 80%; and a strong prompt is triggered when the emotional confidence level is greater than or equal to 80%. Prompts can include audio and RGB lighting effects with a specific rate of change.

[0184] In step 107, based on the third correspondence, the preset interaction mode corresponding to the emotion perception result is determined from multiple preset interaction modes corresponding to the current situation.

[0185] Here, each scenario has a different interaction mode set according to different emotional confidence levels. For example, in a focused writing scenario, the emotional confidence threshold for issuing prompts is lowered. For example, a gentle prompt is triggered when the emotional confidence level is greater than or equal to 20% and less than 50%.

[0186] This application embodiment sets multiple scenarios for the light and sound modes of the input device and dynamically adjusts its interactive performance in each scenario based on the emotional perception results. It can give concrete expression to emotions through the adaptation of light color and sound changes (such as triggering soothing blue light fluctuations when anxious), thereby improving the user's self-awareness and regulation ability. By predefining differentiated feedback strategies for different scenarios (such as games and office work), it achieves a deep fit between the interactive experience and the mood of the scene, significantly enhancing the sense of immersion and personalization. In addition, it provides emotional support without interruption through implicit interaction (such as conveying soothing through the breathing rhythm of light), thus building a more empathetic input device and improving the user experience.

[0187] The following will describe an exemplary application of the embodiments of this application in a practical application scenario.

[0188] In related technologies, the system detects the information input by a user using an input device such as a keyboard and the pressure intensity of the keystrokes. The detected information and pressure intensity are recorded and sent to a system server. The system server then analyzes the detected information and pressure intensity records to determine the user's emotions.

[0189] See Figure 9 , Figure 9 This is a flowchart illustrating an input device interaction method in an application scenario provided by an embodiment of this application, including the following steps 201 to 205.

[0190] It should be noted that users can start executing the input device interaction method provided in this application embodiment by clicking the "Enable Emotion Guardian" button on the desktop interactive interface or input device.

[0191] In this embodiment, the input device is a keyboard, and the processing device is a host. The mobile device can embed a soft keyboard SDK and execute the same model using the SoC NPU; the server can also run in real-time on the KVM virtual keyboard layer for monitoring customer service seats on a large screen.

[0192] In step 201, it is determined whether it is necessary to re-establish the relaxation baseline. If it is necessary to re-establish the relaxation baseline, proceed to step 202; if the relaxation baseline has been established, proceed to step 203.

[0193] In step 202, a relaxation baseline is established. Return to step 201.

[0194] When a user first launches an application that implements the input device interaction method of this application, or through a pre-set recalibration button, they enter a mode for establishing a relaxed baseline. The recalibration button can be a button in the application's interface, or a setting button or key combination on the keyboard. Furthermore, if a change is detected in the keyboard's hardware ID, firmware version, or OS timezone, the recalibration button will pop up, prompting the user whether recalibration is required.

[0195] After the user clicks the recalibrate button, they are taken to the relaxation baseline calibration page. On this page, the user is prompted to engage in 2-3 minutes of free input, such as entering a diary entry or a specific test document. During this free input, the user's keyboard input events are collected. By analyzing these events, content and action information from the calibration phase are obtained. Analysis and feature extraction are then performed based on this information to obtain data such as error rate, pause duration, mean keystroke interval, and keystroke interval variance. A personal relaxation baseline is then established based on this information. Since keystroke interval is significantly correlated with tension levels, subsequent steps in this embodiment use keystroke interval as the primary analysis dimension. In actual testing, data such as the user's backspace key frequency, consecutive deletion length, local WPM, long pause count, burst high-speed segments, keystroke duration, and force can also be collected for recalibrating the relaxation baseline.

[0196] In this embodiment, to ensure that baseline calibration can stably perceive the user's emotions, a relaxation baseline needs to be established from at least two dimensions. The collected data can be divided into two dimensions: 1. Rhythm-related dimensions, such as keystroke interval, pause duration, WPM, and error rate; 2. Lexical dimension: the density of high-pressure words (i.e., the preset content in the above embodiment). It should be noted that even when establishing a relaxation baseline using only rhythm-related data, the input device can still provide emotion perception prompts, but the reliability is insufficient.

[0197] In step 203, user keyboard operation events are collected.

[0198] The keyboard MCU scans the voltage of the row and column circuits in the keyboard at a period of less than or equal to 1ms. Whenever a key in the row or column circuit is detected to be pressed or released, an event packet {scanCode,t_down,t_up} is generated, and an HID message is generated.

[0199] In step 204, the operation event is parsed to obtain the vocabulary dimension information and action information. The vocabulary dimension information is hashed to obtain the hash fingerprint of the input vocabulary.

[0200] It should be noted that if the input device interaction method in this embodiment is implemented through the collaboration of the keyboard and the host, due to the time difference between the keyboard and the host, when the host receives the HID message from the keyboard, it is necessary to compensate for the time difference of the message, for example, by the MCU calculating the time offset. And embed the time offset into the first frame of the HID message, where Indicates the keyboard time. This indicates the host's time so that subsequent hosts can read the corrected timestamp when reading HID packets. Furthermore, compensating for the time offset also compensates for jitter in the USB bus and OS scheduling, ensuring the duration of a single key press. The error is less than 1ms, where, Indicates the duration the button was released. This indicates the time the button was pressed, thereby improving the accuracy of rhythm judgment and avoiding false anomalies caused by host load affecting the subsequent judgment of rhythm deviation value.

[0201] In this embodiment, the lexical fingerprint of the lexical dimension information can be calculated by the FNV-1a algorithm, and a 64-bit hash value is obtained as the hash fingerprint of the lexical (i.e., the first hash value in the above embodiment).

[0202] In step 205, the high-voltage words input by the user are determined based on the hash fingerprint of the input words, and features are extracted from the action information.

[0203] In this embodiment, by comparing the hash fingerprint of the input word with the hash fingerprint of high-pressure words in the high-pressure word database, it is possible to determine whether the user has entered a high-pressure word while ensuring that the user's input words are not leaked. Furthermore, Bloom filtering can quickly determine whether the hash fingerprint of the user's input word includes the hash fingerprint of a high-pressure word in the high-pressure word database.

[0204] It should be noted that the high-pressure terminology database can be a dataset that is cross-filtered from publicly available work stress vocabulary, mental health questionnaires, and IT / customer service industry work order corpora containing negative words with the highest TF-IDF. The database file version is fixed and is uniformly replaced or expanded with software upgrades. The high-pressure terminology database interface can be located on the left side of the main interface. When clicking on the high-pressure terminology database, a general high-pressure terminology base database is loaded by default. Users can enter custom words in the text box and press the "+" key to add them, or select words in the list and click "-" to delete them.

[0205] Each high-pressure word in the high-pressure term library has a corresponding weight. Weight ,in, This represents the inverse document frequency of the i-th preset content. This is a user-defined flag factor. Users can manually set a flag factor for high-pressure words, for example, setting it to 0.7, with a weight of [missing information]. The value range is [0.3, 1.5]. The weights of high-pressure words in the high-pressure word library can be periodically re-regularized according to their usage frequency.

[0206] In step 206, the high-pressure words input by the user and the features extracted from the action information are analyzed to determine the fusion vector.

[0207] The extracted features include: mean keystroke interval, keystroke interval variance, local WPM, backspace frequency, and long pause count. The local WPM can be an estimate within the current time window, for example... .

[0208] In this embodiment of the application, the fusion vector is [vocabulary score, rhythm deviation score, error score, deletion score];

[0209] Vocabulary points: ,in, This represents the total number of characters within the first time window. The word frequency of the i-th preset content within the first time window. The weight of the i-th preset content.

[0210] Error score: ,in, The number of incorrect characters. This is the average number of incorrect characters. This represents the standard deviation of the number of incorrect characters.

[0211] Delete points: ,in, The length of the characters to be deleted. This is the average length of the deleted characters. This represents the standard deviation of the length of characters deleted.

[0212] Rhythm deviation score: ,in, The input interval is obtained from the first action information. The average value of the input interval. is the standard deviation of the input interval.

[0213] In addition, dimensions of behavioral signals such as mouse movement acceleration and window switching frequency can be added to the fusion vector.

[0214] In step 207, emotion perception is performed based on the fusion vector.

[0215] In this embodiment of the application, the fusion vector can be input into the decision tree model, so that the decision tree model outputs the emotion perception result. The emotion perception result is expressed as emotion confidence. The higher the emotion confidence, the higher the intensity of the user's emotion under the set emotion type. For example, the higher the emotion confidence, the higher the tension.

[0216] In addition, the decision tree model can be adaptively set with a depth limit of 4 and leaf nodes ≥ 25 samples in the following way;

[0217] Training is performed using desensitized feature vectors (without plaintext);

[0218] Add “vocabulary density = 0 and <1” Direct routing to low-confidence leaves reduces false alarms.

[0219] After N > 2000 inputs, Naïve Bayes incremental learning is enabled to fine-tune the leaf node probabilities, forming a personalized model.

[0220] Here, we have made adaptive settings for the current scenario, specifically as follows:

[0221] Only fine-tune the conditional probabilities of leaf nodes, without changing the tree structure;

[0222] The circular buffer has 512 rows and is updated using a recursive β distribution to avoid sample sparsity.

[0223] If the observed KL divergence is >0.4, freeze and update for 24 hours to prevent drift.

[0224] In step 208, the interaction mode is executed based on the emotion perception result, and the process returns to step 203.

[0225] The model returns an emotion confidence score p∈[0,1]; the keyboard provides four levels of emotion cues: "silent / mild / moderate / enhanced" based on thresholds {0.3,0.6,0.8}. Alternatively, an exponential moving average can be used to track the confidence score distribution over the most recent 50 windows, resetting the threshold in real time to reduce the need for manual calibration when the environment changes.

[0226] In this embodiment of the application, by continuously collecting user operation events on the keyboard, the user's emotions are periodically sensed. When it is detected that the user's emotional confidence level changes from the range indicating high tension (e.g., 0.8 to 0.9) or the range indicating medium tension (e.g., 0.6 to 0.8) to the range indicating low tension (e.g., 0.3 to 0.6) within a set time period (e.g., 30 seconds), the input device can trigger a positive light effect, such as a short green flash, and display the words "Relaxation successful" on the input device's interactive interface.

[0227] Emotional cues can be generated by reading the LUT (color, brightness, period) from the LED controller to produce a PWM (pulse control signal). Additionally, if sound effects are enabled, the driver calls the WASAPI to play a ≤500mswav clip and implements a throttling strategy (interval between cues ≥8s) to prevent noise interference. Besides RGB / sound effects, it can drive a vibration motor, a miniature thermoelectric element, or overlay a semi-transparent breathing animation on the screen; in open spaces, the cues can be output to a wristband or mobile phone notification to avoid disturbing others.

[0228] The entire process, from button press to light illumination, takes an average of 18–25 ms, with the inference phase accounting for only 2–4 ms. In MCU integration mode, the time can be further reduced to 12 ms.

[0229] Based on user feedback on the interaction patterns, the decision tree surface model can be incrementally learned online. For example, when a user clicks the feedback button for false alarms or confirmed issues, the feature information of the current time window and the label based on user feedback are written into a circular buffer. The prior is updated through recursive Naive Bayes, adapting to individual differences without full retraining. For example: the buffer collects (v, label), where label ∈ {0, 1} comes from the user's "false alarm / confirmation" clicks; each feature dimension is binary discretized (≤ threshold / >\threshold); the conditional probability is updated: P_k(new) = (n·P_k + x) / (n + 1) (x = 1 if the event occurs); the prior π(new) is updated = (m·π + label) / (m + 1); n and m are each limited to 1500, and exponential decay is applied after exceeding the limit. In addition, a statistics page can be set up on the desktop client with input device interaction applications. The statistics page can display the frequency of user tension, average duration, and a hot list of high-pressure words (determined based on the number and frequency of high-pressure words being triggered) by date, week, and month. The statistics data can be de-identified and can be saved locally for users to export as CSV.

[0230] In some embodiments, the decision tree can be replaced with a 1-Hidden-Layer MLP or LightGBM; hash matching can be replaced with an LSH index or a pure Bloom filter, or the decision tree can be switched to LightGBM, Tiny-BERT, or even a model-free rule: it can be triggered when "a high-pressure word appears and Z-Score > 1.5", which is suitable for IoT keyboards with more limited computing power.

[0231] Through the above embodiments, this solution embeds high-pressure word matching and behavioral rhythm analysis entirely within the keyboard MCU or local driver. Input content is immediately desensitized via hashing or Bloom Filtering, and no reversible strings or complete time-series data are transmitted externally. Compared to existing technologies that rely on cloud-based inference, it completely eliminates plaintext leakage and compliance risks, making it applicable even in high-security scenarios such as enterprise intranets and government confidentiality. Combined with a lightweight dual-channel model of high-pressure word density and rhythm deviation, confidence levels are obtained in an average of 2–4ms; the entire closed-loop latency is <25ms, triggering gentle RGB / sound effect guidance at the moment of user emotional peak, avoiding the lag of traditional daily report-style feedback. The online incremental learning mechanism automatically updates the baseline based on user annotations, reducing the false positive rate by more than 31% compared to static models. Model weights are <50KB, the inference cycle and keyboard scanning interruption are reused, and the overall power consumption is equivalent to only a single backlight LED; the driver version's CPU usage is <1%, having almost no impact on the battery life of laptops and mobile office devices. Compared to multimodal solutions requiring cameras and heart rate monitors, deployment costs and user psychological resistance are significantly reduced. Lighting, sound effects, and thresholds are all user-adjustable, with preset scenario templates such as "Focused Writing" and "Public Presentation." When a decrease in tension is detected, positive lighting effects reinforce the positive behavior. The algorithm-hardware decoupling design allows for rapid portability to soft keyboards, touchpads, virtual keyboards, or KVM input layers. It also allows for the addition of various feedback devices such as vibration, temperature, and breathing light strips to meet the needs of multiple scenarios including office work, gaming, and assistive devices for people with disabilities. The open API also allows third-party health or productivity applications to read tension confidence levels, achieving a cross-application emotion-friendly ecosystem. In this embodiment, an abnormal peak in rhythm can also be selected (e.g., any feature Z-Score > 1.5 and lasting ≥ 3 sliding windows is considered an abnormal peak; example: baseline interval of 120 ms, real-time measurement of 60 ms, σ = 20 ms). Z = 3>1.5 → marked as abnormal), only then is a high-pressure word searched in the backtracking interval, thereby reducing false alarms in regular fast-paced scenarios. The embodiments of this application can be applied to game input, customer service ticket systems, and the console of programming IDEs; media data (real-time speech transcription, stylus handwriting trajectory) can also be converted to text locally and then use the same algorithm to complete emotion monitoring.

[0232] The following description continues to illustrate that the input device interaction device 455 provided in the embodiments of this application is an exemplary structure of a software module. In some embodiments, such as Figure 2 The software module stored in the input device interaction device 455 of the memory 450 may include:

[0233] The acquisition module 4551 is used to parse the operation event for the input device to obtain content information and first action information, wherein the content information is used to characterize the characters included in the operation event, and the first action information is used to characterize the start time and end time of the operation in the operation event, and the content information and the action information have a first correspondence relationship;

[0234] The conversion module 4552 is used to perform a hash operation on the content information to obtain multiple first hash values ​​corresponding to the character, wherein the first hash value has a second correspondence with the first action information, and the second correspondence is determined based on the first correspondence.

[0235] Detection module 4553 is used to determine preset content included in the content information in response to a plurality of first hash values ​​including second hash values, wherein the second hash value is a hash value obtained by performing a hash operation on the preset content, the second hash value has a third correspondence with the second action information, and the second action information is information in the first action information;

[0236] The perception module 4554 is used to perform emotion prediction based on the preset content, the second action information and the first action information included in the content information, and to obtain an emotion perception result.

[0237] The interaction module 4555 is used to control the input device to interact based on the preset interaction mode corresponding to the emotion perception result.

[0238] In some embodiments, the perception module 4554 is further configured to parse the preset content, the second action information, and the first action information to obtain feature vectors of multiple set dimensions;

[0239] The feature vectors of the multiple defined dimensions are combined to obtain the fused vector;

[0240] Regression is performed based on the fusion vector to obtain the emotion perception result.

[0241] The perception module 4554 is further configured to parse the preset content, the second action information, and the first action information to obtain set statistical features, wherein each preset content corresponds to a set weight.

[0242] Calculate the deviation of the statistical feature from a pre-set reference statistical feature;

[0243] Based on the second action information, the number of times the preset content appears within the first time window is determined, and the content score of the operation event is calculated based on the weight corresponding to the preset content and the number of times it appears;

[0244] The deviation value and the content score are used as feature vectors with multiple defined dimensions.

[0245] In some embodiments, the sensing module 4554 is further configured to obtain the reference statistical features by performing the following processing:

[0246] In response to a detection command, the operation event within the second time window is detected, and a reference statistical feature with the same dimension as the statistical feature is parsed based on the operation event within the second time window, wherein the length of the second time window is greater than or equal to the length of the first time window.

[0247] In some embodiments, the sensing module 4554 is further configured to determine the duration of the first time window by performing the following processes:

[0248] The character input speed of the input device is determined based on the first action information;

[0249] In response to the character input speed of the input device being lower than a set first speed threshold, the duration of the first time window is increased;

[0250] In response to the character input speed of the input device being higher than or equal to a set second speed threshold, the duration of the first time window is reduced.

[0251] In some embodiments, the deviation values ​​in the feature vector of the perception module 4554 include error score, deletion score, and rhythm deviation score. The error score is determined based on the number of incorrectly input characters and the number of incorrectly input characters in a reference input. The deletion score is determined based on the length of the deleted characters and the length of the deleted characters in a reference input. The rhythm deviation score is determined based on the character input speed and the speed of the character input in a reference input. The number of incorrectly input characters and the length of the deleted characters are determined based on parsing the preset content, the second action information, and the content information. The character input speed is determined based on the first action information.

[0252] The regression based on the fusion vector to obtain the emotion perception result includes:

[0253] The error score, deletion score, and rhythm deviation score are summed to obtain a first score vector. The sum of the first score vector and the content score is used as a second score vector. The emotion confidence score is obtained by regression based on the second score vector using a normalized exponential function. The emotion confidence score is used as the emotion perception result.

[0254] In some embodiments, the perception module 4554 is further configured to set the emotion confidence level to a preset value in response to the content score being 0 and the rhythm deviation score being lower than a set rhythm deviation score threshold.

[0255] In some embodiments, the sensing module 4554 is further configured to determine the weight by performing the following process:

[0256] The weight corresponding to each of the preset contents is preset;

[0257] In response to the detection that the deviation between the occurrence count of the first word and the reference occurrence count of the first word in the reference statistical features is greater than a set deviation threshold within multiple consecutive first time windows, the weight corresponding to the first word is increased, wherein the first word is a word in the preset content.

[0258] In some embodiments, the acquisition module 4551 is further configured to record the current time of the input device and the host time in response to detecting an operation event;

[0259] The difference between the host time and the current time of the input device is embedded into the first frame of the operation event to obtain a new operation event;

[0260] The new operation event is parsed to obtain the content information and the first action information.

[0261] In some embodiments, the interaction module 4555 is further configured to select the current context from a plurality of preset contexts in response to the operation events of the input device, wherein each context is configured with a plurality of preset interaction modes, and each preset interaction mode has a third correspondence with the emotion perception result;

[0262] From the multiple preset interaction modes corresponding to the current situation, determine the preset interaction mode corresponding to the emotion perception result.

[0263] This application provides a computer program product, which includes a computer program or computer-executable instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer-executable instructions from the computer-readable storage medium and executes the computer-executable instructions, causing the electronic device to perform the input device interaction method described above in this application.

[0264] This application provides a computer-readable storage medium storing computer-executable instructions or a computer program. When the computer-executable instructions or the computer program are executed by a processor, the processor will execute the input device interaction method provided in this application, for example, such as... Figure 3 The input device interaction method is shown.

[0265] In some embodiments, the computer-readable storage medium may be a memory such as RAM, ROM, flash memory, magnetic surface memory, optical disk, or CD-ROM; or it may be a variety of devices including one or any combination of the above-mentioned memories.

[0266] In some embodiments, computer-executable instructions may take the form of programs, software, software modules, scripts, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as stand-alone programs or as modules, components, subroutines, or other units suitable for use in a computing environment.

[0267] As an example, computer-executable instructions may, but do not necessarily, correspond to files in a file system. They may be stored as part of a file that holds other programs or data, for example, in one or more scripts in a Hyper Text Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple co-located files (e.g., files that store one or more modules, subroutines, or code sections).

[0268] As an example, computer-executable instructions can be deployed to execute on a single electronic device, or on multiple electronic devices located at one location, or on multiple electronic devices distributed across multiple locations and interconnected via a communication network.

[0269] In summary, this application's embodiments, by parsing the operation events of the input device, divide the information contained in the operation events into content information and action information, perform a hash operation on the content information to desensitize it, and then compare the desensitized first hash value with a pre-stored second hash value, thereby verifying whether the user has entered the set content without disclosing the user's input content information. Furthermore, based on the user's input of the set content and the action information when inputting the set content, the user's emotion is predicted, achieving the prediction of the user's emotion while ensuring that the user's input content information is not disclosed.

[0270] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, and improvements made within the spirit and scope of this application are included within the scope of protection of this application.

Claims

1. An input device interaction method, characterized by, The method comprises: parsing an operation event for the input device to obtain content information and first action information, wherein the content information is used to represent characters included in the operation event, the first action information is used to represent a start time and an end time of the operation in the operation event, and the content information and the action information have a first correspondence relationship; performing a hash operation on the content information to obtain a plurality of first hash values corresponding to the characters, wherein the first hash values and the first action information have a second correspondence relationship, and the second correspondence relationship is determined based on the first correspondence relationship; in response to the plurality of first hash values including a second hash value, determining that preset content is included in the content information, wherein the second hash value is a hash value obtained by performing a hash operation on the preset content, the second hash value and second action information have a third correspondence relationship, and the second action information is information in the first action information; performing emotion perception based on the preset content included in the content information, the second action information and the first action information to obtain an emotion perception result; controlling the input device to interact based on a preset interaction mode corresponding to the emotion perception result.

2. The method of claim 1, wherein, The emotion perception based on the preset content, the second action information and the first action information comprises: parsing the preset content, the second action information and the first action information to obtain a plurality of feature vectors of a set dimension; combining the plurality of feature vectors of the set dimension to obtain a fusion vector; performing regression based on the fusion vector to obtain an emotion perception result.

3. The method of claim 2, wherein, The parsing of the preset content, the second action information and the first action information to obtain a plurality of feature vectors of a set dimension comprises: parsing the preset content, the second action information and the first action information to obtain a set of statistical features, wherein each preset content corresponds to a set of weights; calculating a deviation value of the statistical features relative to a pre-set reference statistical feature; based on the second action information, determining the number of occurrences of the preset content within a first time window, and based on the weight corresponding to the preset content and the number of occurrences, calculating a content score of the operation event; the deviation value and the content score are used as the feature vectors of a plurality of set dimensions.

4. The method of claim 3, wherein, Before the calculation of the deviation value of the statistical features relative to the pre-set reference statistical feature, the method further comprises: obtaining the reference statistical features by performing the following processing: in response to a detection instruction, detecting the operation event within a second time window; based on the operation event within the second time window, parsing to obtain a reference statistical feature with the same dimension as the statistical feature, wherein the length of the second time window is greater than or equal to the length of the first time window.

5. The method of claim 3, wherein, The method further comprises: determining the length of the first time window by performing the following processing: determining the character input speed of the input device according to the first action information; in response to the character input speed of the input device being lower than a set first speed threshold, increasing the length of the first time window; in response to the character input speed of the input device being higher than or equal to a set second speed threshold, reducing the length of the first time window.

6. The method of claim 3, wherein, The deviation values in the feature vector include an error score, a deletion score, and a rhythm deviation score, wherein the error score is determined based on a number of input errors and a reference number of input errors, the deletion score is determined based on a length of deleted characters and a reference length of deleted characters, and the rhythm deviation score is determined based on a character input speed and a reference character input speed, the number of input errors and the length of deleted characters are determined based on the parsing of the preset content, the second action information, and the content information, and the character input speed is determined based on the first action information; The regression based on the fusion vector includes: The error score, the deletion score, and the rhythm deviation score are summed to obtain a first score vector; The sum of the first score vector and the content score is taken as a second score vector; The emotion confidence is obtained by regression based on the second score vector through a normalized exponential function, and the emotion confidence is taken as the emotion perception result.

7. The method of claim 6, wherein, Before the emotion confidence is taken as the emotion perception result, the method further includes: In response to the content score being 0 and the rhythm deviation score being lower than a set rhythm deviation score threshold, the emotion confidence is set to a preset value.

8. The method of claim 3, wherein, The weight is determined by performing the following processing: The weight corresponding to each preset content is set in advance; In response to detecting, in a plurality of consecutive first time windows, that a deviation value of a number of occurrences of a first vocabulary from a reference number of occurrences of the first vocabulary in the reference statistical features is greater than a set deviation threshold, the weight corresponding to the first vocabulary is increased, wherein the first vocabulary is a vocabulary in the preset content.

9. The method of claim 1, wherein, The parsing of the operation event for the input device includes: In response to detecting an operation event, recording a current time of the input device and a host time; Embedding a difference between the host time and the current time of the input device into a first frame of the operation event to obtain a new operation event; Parsing the new operation event to obtain the content information and the first action information.

10. An input device interaction apparatus, characterized by The device includes: An acquisition module configured to parse an operation event for the input device to obtain content information and first action information, wherein the content information is used to represent characters included in the operation event, the first action information is used to represent start and end times of an operation in the operation event, and the content information and the action information have a first correspondence relationship; A conversion module configured to perform a hash operation on the content information to obtain a plurality of first hash values corresponding to the characters, wherein the first hash values and the first action information have a second correspondence relationship, and the second correspondence relationship is determined based on the first correspondence relationship. The detection module is configured to determine preset content included in the content information in response to the plurality of first hash values including a second hash value, the second hash value being a hash value obtained by performing a hash operation on the preset content, the second hash value having a third correspondence relationship with second action information, the second action information being information in the first action information; The perception module is configured to perform emotion prediction based on the preset content included in the content information, the second action information, and the first action information, to obtain an emotion perception result; The interaction module is configured to control the input device to interact based on a preset interaction mode corresponding to the emotion perception result.

Citation Information

Patent Citations

  • Emotion processing method and device and medium

    CN114816036A

  • Text recognition method and device, electronic equipment and storage medium

    CN119760141A