A feedback control method based on sensor fusion and operation intention recognition

By fusing inertial measurement, pressure sensing, and external visual posture data, multidimensional features are extracted and operational intentions are identified to generate interest weight coefficients. This solves the problem of lag in the feedback process in existing technologies and realizes dynamic feedback control of intelligent interactive devices.

CN122450291APending Publication Date: 2026-07-24HUBEI ZELIN EDUCATION TECHNOLOGY GROUP CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HUBEI ZELIN EDUCATION TECHNOLOGY GROUP CO LTD
Filing Date
2026-04-17
Publication Date
2026-07-24

Smart Images

  • Figure CN122450291A_ABST
    Figure CN122450291A_ABST
Patent Text Reader

Abstract

The application provides a feedback control method based on sensor fusion and operation intention recognition, and relates to the technical field of intelligent sensors, which comprises the following steps: acquiring inertial measurement data, pressure sensing data and external visual pose data arranged in a target interactive carrier, and performing preprocessing to form a multi-modal operation data sequence; performing sliding window segmentation processing based on the multi-modal operation data sequence, and extracting multi-dimensional feature parameters reflecting operation intensity, change trend and periodic characteristics to construct a feature vector; performing multi-dimensional decision mapping on the feature vector to identify the corresponding operation intention category; generating an interest weight coefficient according to the operation intention category and its continuous distribution characteristics in the time dimension; constructing a feedback control strategy based on the interest weight coefficient and the operation intention category, and driving an execution unit to output a response signal. The application can fuse multi-source data of an interactive device, and improve the timeliness of dynamic adjustment of a feedback strategy in combination with operation intention recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the technical field of intelligent sensors, and specifically to a feedback control method based on sensor fusion and operation intention recognition. Background Technology

[0002] Existing intelligent interactive early childhood education tools or toys typically integrate only a single type of sensing device, such as a touch sensor or a simple inertial sensor, and output fixed forms of sound, light, or mechanical feedback signals based on preset trigger logic to achieve basic human-computer interaction functions. Although some solutions introduce multiple sensors for data collection, they mostly remain at the level of simple superposition or independent processing, lacking the ability to perform temporal alignment and joint modeling of multi-source heterogeneous data. This results in the inability to effectively extract high-dimensional features that reflect the intensity, trend, and periodicity of operational behaviors. At the same time, in terms of understanding operational behaviors, they mostly rely on rule matching or low-complexity classification models, making it difficult to accurately model and recognize the complex and ever-changing operational behaviors of young children. The overall interaction process is still mainly passive response, lacking the ability to dynamically adjust to the user's real-time operational status.

[0003] However, existing technologies have a key technical deficiency in practical applications: they cannot accurately identify operational intentions based on multimodal operational data and further construct feedback control mechanisms associated with interest tendencies. Due to the lack of continuous distribution feature analysis of multimodal data over time and a unified multidimensional feature mapping mechanism, existing solutions struggle to distinguish the true intentions behind different operational behaviors. Consequently, they cannot generate weight parameters that reflect changes in user interests, nor can they dynamically adjust feedback strategies based on such parameters. This results in feedback content exhibiting significant lag and mechanicalness, making it difficult to achieve continuous guidance and closed-loop control of user exploration behavior. Summary of the Invention

[0004] This invention provides a feedback control method based on sensor fusion and operation intention recognition, which can integrate multi-source data from interactive devices and combine operation intention recognition to improve the timeliness of dynamic adjustment of feedback strategies.

[0005] In a first aspect, the present invention provides a feedback control method based on sensor fusion and operation intention recognition, the method comprising: Acquire inertial measurement data, pressure sensing data, and external visual attitude data set in the target interactive carrier, and perform preprocessing to form a multimodal operation data sequence; Based on the multimodal operation data sequence, a sliding window segmentation process is performed, and multidimensional feature parameters reflecting operation intensity, change trend and periodic characteristics are extracted to construct a feature vector; Perform a multidimensional decision mapping on the feature vector to identify the corresponding operation intent category; Based on the category of operational intent and its continuous distribution characteristics over time, interest weight coefficients are generated; A feedback control strategy is constructed based on the interest weight coefficient and the operation intention category, and the execution unit is driven to output a response signal.

[0006] In a second aspect, the present invention provides a feedback control device based on sensor fusion and operation intent recognition, the device being used to execute a feedback control method based on sensor fusion and operation intent recognition as described in any of the above embodiments, the device comprising an acquisition module, a processing module, and an output module, wherein: The acquisition module is used to acquire inertial measurement data, pressure sensing data and external visual attitude data set in the target interactive carrier, and perform preprocessing to form a multimodal operation data sequence. The processing module is used to perform sliding window segmentation processing based on the multimodal operation data sequence, and extract multidimensional feature parameters that reflect operation intensity, change trend and periodic characteristics to construct feature vectors; The processing module is used to perform multidimensional decision mapping on the feature vector to identify the corresponding operation intention category; The processing module is used to generate interest weight coefficients based on the category of the operation intention and its continuous distribution characteristics in the time dimension. The output module is used to construct a feedback control strategy based on the interest weight coefficient and the operation intention category, and drive the execution unit to output a response signal.

[0007] A third aspect of the invention provides an electronic device including a processor, a memory, a user interface, and a network interface, wherein the memory is used to store instructions, the user interface and the network interface are both used to communicate with other devices, and the processor is used to execute the instructions stored in the memory to cause the electronic device to perform the method as described in any of the preceding embodiments.

[0008] In a fourth aspect of the invention, a non-transitory computer-readable storage medium is provided, the computer-readable storage medium storing instructions that, when executed, perform the method as described in any of the preceding claims.

[0009] In summary, one or more technical solutions provided in the embodiments of the present invention have at least the following technical effects or advantages: This invention aligns and fuses inertial measurement data, pressure sensing data, and external visual attitude data using a unified time reference, enabling data from different sources to form a structurally consistent multimodal operation data sequence within the same temporal framework. This allows for the extraction of multidimensional feature parameters that simultaneously characterize operation intensity, change trends, and periodicity based on sliding window segmentation. Furthermore, multidimensional decision mapping enables refined identification of complex operational intentions. Building upon this, the invention further combines the continuous distribution characteristics of operational intentions over time to generate interest weight coefficients, allowing the system to depict the dynamic evolution of user interests in real time. Based on this, a feedback control strategy matching the current operational state is constructed, transforming the feedback process from passive triggering to proactive adjustment based on interest changes. This significantly improves the timeliness of feedback strategy updates and the continuity of interactive responses. Attached Figure Description

[0010] Figure 1 This is a flowchart illustrating a feedback control method based on sensor fusion and operation intent recognition disclosed in an embodiment of the present invention. Figure 2 This is a schematic diagram of a target interaction carrier disclosed in an embodiment of the present invention; Figure 3 This is a schematic diagram of a feedback control device based on sensor fusion and operation intention recognition disclosed in an embodiment of the present invention; Figure 4 This is a schematic diagram of the structure of an electronic device disclosed in an embodiment of the present invention.

[0011] Explanation of reference numerals in the attached drawings: 301, acquisition module; 302, processing module; 303, output module; 401, processor; 402, communication bus; 403, user interface; 404, network interface; 405, memory. Detailed Implementation

[0012] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.

[0013] In the description of the embodiments of the present invention, words such as "for example" or "for instance" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as "for example" or "for instance" in the embodiments of the present invention should not be construed as being more preferred or advantageous than other embodiments or designs. Rather, the use of words such as "for example" or "for instance" is intended to present the relevant concepts in a specific manner.

[0014] In the description of the embodiments of the present invention, the term "multiple" means two or more. For example, multiple systems means two or more systems, and multiple screen terminals means two or more screen terminals. Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the indicated technical features. Thus, a feature defined with "first" or "second" may explicitly or implicitly include one or more of that feature. The terms "comprising," "including," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized.

[0015] Existing intelligent interactive early childhood education tools generally rely on single or simply superimposed sensor data processing methods, lacking the ability to align and jointly model multi-source heterogeneous data over time. They are unable to extract high-dimensional features that can characterize the intensity of operations, trends of change, and periodic characteristics. At the same time, in terms of operation intention recognition, they mostly use rule matching or low-complexity models, which cannot accurately depict children's complex behaviors and their true intentions. Furthermore, due to the lack of interest weight modeling based on the continuous distribution characteristics of time dimension and dynamic adjustment mechanism of feedback strategy, the overall feedback process is lagging and mechanical, making it difficult to achieve continuous guidance and closed-loop control for user exploration behavior.

[0016] This invention discloses a feedback control method based on sensor fusion and operation intent recognition, which is applied to a server. The server includes, but is not limited to, electronic devices such as mobile phones, tablets, wearable devices, and PCs (Personal Computers), and can also be a backend server running a feedback control method based on sensor fusion and operation intent recognition. The server can be implemented using a standalone server or a server cluster composed of multiple servers.

[0017] This embodiment discloses a feedback control method based on sensor fusion and operation intention recognition, referring to... Figure 1 It includes the following steps: S110: Acquire inertial measurement data, pressure sensing data, and external visual attitude data set in the target interactive carrier, and perform preprocessing to form a multimodal operation data sequence.

[0018] S120 performs sliding window segmentation based on multimodal operation data sequences and extracts multidimensional feature parameters that reflect operation intensity, change trend and periodic characteristics to construct feature vectors.

[0019] S130, Perform a multidimensional decision mapping on the feature vector to identify the corresponding operational intent category.

[0020] S140 generates interest weight coefficients based on the category of operational intent and its continuous distribution characteristics over time.

[0021] S150 constructs a feedback control strategy based on interest weight coefficients and operation intention categories, and drives the execution unit to output response signals.

[0022] Reference Figure 2 This is a schematic diagram of a target interactive carrier disclosed in an embodiment of the present invention. The carrier has an overall desktop-style shell shape, integrating multiple sensing and feedback units to support multimodal operation data acquisition and response output functions. Specifically, a camera module and an infrared sensor together constitute an external visual posture data acquisition unit, used to sense the position, movement trajectory, and spatial posture changes of a child's hand. The camera module acquires continuous frame sequences through image acquisition, and the infrared sensor enhances contour recognition capabilities in low-light environments, thereby ensuring posture recognition stability. A microphone and speaker unit are used for voice input and voice feedback, respectively. The microphone collects ambient audio to assist in behavior judgment or voice interaction, and the speaker unit outputs feedback voice signals. A rotating dial and a pressure-sensing pad constitute the main tactile interaction components. The rotating dial generates angular displacement operations, and the pressure-sensing pad detects pressure intensity and contact state. A motion sensor is located on the side of the shell to detect overall movement, tilting, or vibration. This motion sensor typically includes an inertial measurement unit composed of an accelerometer and a gyroscope, thus forming an inertial measurement data source. Lights and a display panel are used to output visual feedback information, achieving interactive guidance through light changes, pattern displays, or symbol prompts. All the above components are structurally arranged around the same target interactive carrier, and data is aggregated and uniformly scheduled through an internal control unit, so that the target interactive carrier has both multimodal perception capability and multi-channel feedback capability.

[0023] In practice, when acquiring inertial measurement data, pressure sensing data, and external visual posture data around the target interactive carrier, firstly, triaxial acceleration data and triaxial angular velocity data are collected in real time within the motion sensor using an inertial measurement unit. Acceleration data characterizes the linear motion state of the target interactive carrier in space, while angular velocity data characterizes its rotational state; together, they constitute the inertial measurement data. At the pressure sensing pad, a flexible pressure sensing array acquires contact pressure values ​​and pressure distribution information. Pressure sensing data reflects the magnitude of the force applied by the child's hand, the duration of contact, and the process of force change. External visual posture data is acquired through a camera module and an infrared sensor, specifically including the location of key points on the child's hand, posture angle, movement trajectory, and the occupancy of the target interactive area. This external visual posture data describes the structural characteristics of the child's operational behavior in space. During the acquisition process, all three types of data are output at a uniform sampling frequency or a mappable sampling frequency and transmitted to the control processing unit via an internal bus to form an original multi-source data set. Among them, inertial measurement data emphasizes the overall motion state, pressure sensing data emphasizes local contact behavior, and external visual attitude data emphasizes spatial interaction relationships, thus forming a complementary information expression system.

[0024] When preprocessing inertial measurement data, pressure sensing data, and external visual attitude data to form a multimodal operational data sequence, a unified time reference alignment process is first performed on the three types of data to establish a one-to-one correspondence between data from different sources on the time axis. This is achieved by using timestamp interpolation or resampling methods to address the issue of inconsistent sampling frequencies among different sensors. Subsequently, denoising filtering and gravity component separation are performed on the inertial measurement data to eliminate the influence of environmental noise and static bias on motion characteristics. Baseline drift correction and abnormal peak removal are performed on the pressure sensing data to ensure the stability of the pressure change curve. Keypoint smoothing and occlusion compensation are performed on the external visual attitude data to improve the continuity of the attitude trajectory. After cleaning the data from each channel, a unified scale normalization process is performed on the three types of data to express data of different physical quantities within the same numerical range, thereby avoiding weight bias in subsequent fusion processes for data from one channel. Furthermore, the aligned inertial measurement data, pressure sensing data, and external visual attitude data are stitched together time-by-time according to chronological order to form a multidimensional data vector sequence. This multidimensional data vector sequence is the multimodal operation data sequence, where each time step contains information on the motion state, force application state, and attitude state at the corresponding moment. Through the above preprocessing, the originally heterogeneous, scaled, and noisy data is transformed into a data representation form with a unified structure, continuous time, and consistent semantics, thus providing a stable input foundation for subsequent sliding window segmentation processing and feature vector construction.

[0025] In one possible implementation, a sliding window segmentation process is performed on the multimodal operation data sequence, and multidimensional feature parameters reflecting operation intensity, change trend, and periodic characteristics are extracted to construct a feature vector. Specifically, this includes: performing unified time base alignment, missing data completion, and anomaly removal on the multimodal operation data sequence to form a standardized multimodal operation data sequence; configuring window parameters and performing sliding segmentation based on the standardized multimodal operation data sequence, and adaptively adjusting the window boundaries when the action changes abruptly to generate a target operation segment; extracting operation intensity features, change trend features, and periodic characteristic features from the target operation segment, and performing fusion encoding and correlation filtering processing to construct a feature vector.

[0026] Specifically, when performing unified time reference alignment, missing data completion, and anomaly removal on multimodal operation data sequences, a unified time axis is first established around the inertial measurement data, pressure sensing data, and external visual posture data within the target interactive carrier. The original local sampling times of different sensor channels are then mapped to the same reference time set. Unified time reference alignment refers to correcting data sequences with different sampling frequencies, transmission delays, and triggering points to the same time reference frame, so that changes in gripping, motion, and posture at the same moment can correspond to the same operational behavior segment. Missing data completion refers to restoring the estimated value within a time interval based on the changes in adjacent sampling points when a sensor channel does not output a valid sampling value within a certain time interval, thus maintaining the continuity of the multimodal operation data sequence. Anomaly removal refers to identifying and removing unreasonable sampling points caused by instantaneous electrical noise, false detection due to occlusion, impact jitter, or communication packet loss, avoiding the amplification effect of local anomalies on subsequent feature extraction. In specific processing, a unified reference sampling time sequence is first established, and timestamp matching is performed on inertial measurement data, pressure sensing data, and external visual attitude data respectively. When the sampling time of a certain sensor channel does not completely coincide with the reference sampling time, an interpolation completion method is used to obtain the alignment value. The linear interpolation expression is as follows:

[0027]

[0028] In the formula, This indicates that sensor channel c is at the reference time. The completion value at the location; and These represent the known sampled values ​​located on either side of the reference time; and Let each represent an adjacent valid sampling time, and satisfy the following conditions: The principle behind this formula is that it assumes the changes in the same sensor channel over a short period of time have local continuity. Therefore, the value at an intermediate moment can be estimated using the slope of the change between two consecutive points. After completing the missing data completion, a sliding statistics method is used to identify outliers. The outlier criterion can be expressed as:

[0029]

[0030]

[0031] In the formula, This indicates that sensor channel c is at time... The degree of standardization deviation; Indicated by time The mean value of this sensor channel within a local time window centered on it; This represents the standard deviation within the same local time window; This represents the abnormal threshold for the corresponding sensor channel, which is usually preset based on the channel's historical fluctuation range. When the standardization deviation exceeds the abnormal threshold, the sampling point is determined to be an abnormal point and is removed or replaced. After the above processing, a standardized multimodal operation data sequence that is continuous in time, consistent in sampling, and has controlled noise can be obtained, providing a unified data foundation for subsequent window segmentation.

[0032] Based on standardized multimodal operation data sequences, window parameter configuration and sliding segmentation are performed. When adaptively adjusting the window boundaries during action mutations, the window length, window step size, and window overlap rate are first set according to the typical operation cycle of the target interactive carrier, the sensor sampling frequency, and the desired recognition sensitivity. Window parameter configuration refers to determining the data length covered by a single analysis window in the time dimension and the moving distance between adjacent windows. Sliding segmentation refers to continuously moving the preset window along the standardized multimodal operation data sequence and capturing multiple local time segments. The target operation segment refers to a data segment that can relatively completely represent the evolution process of a local operation behavior. Action mutations refer to behavioral changes such as sudden increases in pressure, sudden increases in angular velocity, sudden changes in acceleration direction, or rapid switching of visual posture within a short period of time. These changes typically correspond to a child transitioning from stable contact to significant operational states such as pressing, swinging, rotating, or grabbing. Therefore, if fixed window boundaries are still used, a complete action may be easily truncated into two different windows, leading to subsequent feature distortion. In specific implementation, the standardized multimodal operation data sequence is first initially slid segmented according to the preset window length L and window step size S. The sample index range covered by each window can be represented as:

[0033]

[0034] In the formula, denotes the set of sample indices corresponding to the m-th initial window; L represents the number of sampling points included in a single window, and its value is determined according to the average duration of a single operation behavior; S represents the number of sampling points for each window slide. When S < L, there is an overlapping interval between adjacent windows, which is beneficial to retaining the continuity of the action. Subsequently, to identify the action mutation position, a comprehensive change intensity function is constructed around the pressure sensing data, inertial measurement data, and external visual pose data, and the expression is:

[0035]

[0036] In the formula, represents the comprehensive change intensity at time ; represents the difference value of the pressure sensing data between adjacent times, which is used to characterize the force application change speed; represents the difference of the three-axis acceleration vector, which is used to characterize the degree of linear motion mutation; represents the difference of the three-axis angular velocity vector, which is used to characterize the degree of rotation state mutation; represents the difference of the visual pose parameter vector, which is used to characterize the degree of change in the hand or limb pose observed externally; , , and respectively represent the weight coefficients of the change items of each sensor channel, and their values are determined according to the importance of each channel for the recognition of the operation intention. When exceeds the preset action mutation threshold, this moment is identified as a candidate boundary adjustment point, and the start boundary or end boundary of the current window is shifted towards the candidate boundary adjustment point, so that the main energy interval of the action尽可能 falls within the same target operation segment. The principle of this processing is to capture the turning position of the behavior structure by comprehensively considering the change amplitudes of multiple channels, so that each target operation segment尽量 corresponds to a local operation behavior with a complete start, development, and decay process.

[0037] When extracting operational intensity features, trend features, and periodicity features from target operation segments and performing fusion encoding and correlation screening to construct feature vectors, the representation results of multimodal operational data in three dimensions—amplitude, rate of change, and repetition rhythm—are first calculated sequentially within each target operation segment. Operational intensity features describe the strength of force exerted, swing amplitude, rotation intensity, and postural involvement of the child within the current target operation segment, reflecting the absolute level of behavioral activity. Trend features describe the direction of increase or decrease, slope of change, and frequency of abrupt changes in parameters such as pressure, speed, and posture over time, reflecting how behavior evolves over time. Periodicity features describe the presence of repetitive patterns such as repeated pressing, continuous rotation, or rhythmic swinging within the current target operation segment, reflecting the rhythmic structure of the behavior. Specifically, the pressure sequence within the target operation segment is analyzed. , combined acceleration sequence and resultant angular velocity sequence Extracting operation intensity features, such as the root mean square value, can be represented as:

[0038]

[0039] In the formula, This represents the root mean square intensity characteristic of a certain sensor channel within the current target operation segment; This represents the value of the current sensor channel at the nth sampling point, which can correspond to pressure, resultant acceleration, or resultant angular velocity; N represents the number of sampling points contained in the current target operation segment; this formula, by averaging the energy within the segment and then taking the square root, can comprehensively reflect the overall activity level within the operation segment. The trend characteristics can be represented by first-order slope statistics as follows:

[0040]

[0041] In the formula, This represents the average slope of change of the current sensor channel within the target operation segment; This represents the time interval between adjacent sampling points; the formula reflects whether the overall data within a segment tends to rise, fall, or remain relatively stable. Periodic characteristics can be obtained by extracting the dominant frequency of the discrete spectrum, expressed as:

[0042]

[0043]

[0044] In the formula, Let x(n) represent the spectral coefficients at the k-th discrete frequency point; x(n) represent the time-domain sequence within the current target operation segment. The index representing the dominant frequency at which the spectral amplitude reaches its maximum value; Indicates the sampling frequency; This represents the dominant frequency feature corresponding to the current target operation segment, used to characterize the speed of repetitive operation rhythm. Its principle lies in converting the time-domain repetitive pattern into frequency clustering peaks through frequency domain expansion, thereby identifying periodic behaviors such as repeated pressing and regular rotation. After obtaining various features, features from different channels and dimensions are concatenated according to a fixed field order to form candidate feature groups, and then fusion encoding is performed to ensure that each target operation segment corresponds to a structurally unified multidimensional expression result. Correlation filtering is used to eliminate feature terms with excessive redundancy or weak discriminative ability, which can be achieved through correlation coefficient criteria, expressed as:

[0045]

[0046] In the formula, This represents the correlation coefficient between the i-th feature term and the j-th feature term across all training samples; This represents the value of the m-th sample in the i-th feature term; The i-th feature term represents the average value of the i-th feature term across all samples; M represents the total number of samples; when When the threshold is exceeded, it indicates that the two feature terms highly overlap in their content, and the one with stronger discriminative ability can be retained. The result obtained after fusion encoding and correlation filtering is the feature vector, which can simultaneously characterize the operational intensity, trend of change, and periodicity of the target operation segment within the same expression space.

[0047] In one possible implementation, a multi-dimensional decision mapping is performed on the feature vector to identify the corresponding operation intention category. Specifically, this includes: constructing a set of operation intention categories and establishing a non-linear mapping relationship between the feature vector and the operation intention category based on an ensemble classification model; inputting the feature vector into multiple basic decision units in the ensemble classification model to output candidate operation intention categories respectively; performing ensemble voting on multiple candidate operation intention categories to determine the target operation intention category; when there is a conflict among candidate operation intention categories, performing conflict resolution processing based on the changing trend features and periodic characteristics in the feature vector to correct the target operation intention category; generating an intention confidence parameter based on the support of the basic decision units, and performing delayed joint discrimination processing when the intention confidence parameter is lower than a preset threshold; and performing temporal smoothing processing on the target operation intention category corresponding to continuous target operation segments to output the operation intention category.

[0048] Specifically, when constructing a set of operational intention categories and establishing a nonlinear mapping relationship between feature vectors and operational intention categories based on an ensemble classification model, the typical behavioral patterns of the target interactive carrier during actual use are first combined to semantically categorize the possible operational behaviors of young children, forming a set of operational intention categories that are distinguishable and have clear behavioral orientations. The operational intention category set refers to the set of categories used to represent different operational purposes or behavioral states, such as observing details, repeated pressing, rapid rotation, detailed disassembly, vigorous play, synchronized tapping, or competition and interference. The specific number of categories can be set according to the interaction structure of the target interactive carrier and the distribution of training samples. The nonlinear mapping relationship means that the relationship between feature vectors and operational intention categories is not a simple one-to-one linear correspondence, but a mapping structure determined by the coupling relationship, threshold splitting relationship, and combination discrimination relationship among multiple feature terms. The structure of the ensemble classification model can adopt a parallel combination architecture of multiple basic decision units, where each basic decision unit is an independently trained decision tree, and multiple basic decision units together form a random forest-style ensemble discrimination structure. This structure comprises a feature input layer, a parallel discriminant layer, and an ensemble output layer. The feature input layer receives the feature vector corresponding to the current target operation segment and inputs the operation intensity feature, change trend feature, and periodic characteristic feature into each basic decision unit according to a fixed field order. The parallel discriminant layer consists of multiple basic decision units, each learning local discrimination rules based on different training subsets and different candidate feature subspaces, allowing the same feature vector to be analyzed in parallel from multiple different discriminative perspectives. The ensemble output layer aggregates the candidate operation intention categories output by each basic decision unit and performs a voting decision to obtain the final category result. The principle behind this structure is that by allowing multiple basic decision units to learn different local nonlinear boundaries, it avoids the oversensitivity of a single classifier to local sample noise or local feature bias, thereby improving its adaptability to complex child operations.

[0049] When inputting feature vectors into multiple basic decision units in an ensemble classification model to output candidate operation intent categories, the feature vector corresponding to the current target operation fragment is first broadcast to each basic decision unit, allowing each basic decision unit to independently perform tree-structured discrimination path search. A basic decision unit is the smallest structural unit within the ensemble classification model used to complete a single local classification judgment. Essentially, it is a feature splitting structure organized in a tree hierarchy, containing a root node, intermediate nodes, and leaf nodes. The root node receives the complete feature vector, the intermediate nodes perform conditional judgments around a specific feature term and its threshold, and the leaf nodes output candidate operation intent categories. Specifically, in a basic decision unit, after the current feature vector enters from the root node, the system first reads the splitting feature index and splitting threshold corresponding to that node, and determines whether the value of the current feature vector on that splitting feature term meets the preset conditions. If it does, it enters the left child node; otherwise, it enters the right child node, and this process is repeated until a leaf node is reached. The leaf node pre-stores the category distribution results corresponding to the path, so each basic decision unit will output a candidate operation intent category. Taking the r-th basic decision unit as an example, its discrimination process can be represented as follows:

[0050]

[0051] In the formula, Let f represent the discriminant function of the r-th basic decision unit for the feature vector f; f represents the feature vector corresponding to the current target operation segment. Let represent the candidate operation intent category output by the r-th basic decision unit. The principle behind this formula is that each basic decision unit is equivalent to an independent local classifier, which maps its current feature vector to the candidate operation intent category corresponding to a certain leaf node based on the splitting rules it has trained. Since the training samples and feature selection paths of multiple basic decision units are not completely identical, the same feature vector can yield multiple candidate operation intent categories from different discriminative perspectives, providing a basis for subsequent ensemble voting.

[0052] To determine the target operation intent category, an integrated voting process is performed on multiple candidate operation intent categories. First, the distribution of candidate operation intent categories output by all basic decision-making units is statistically analyzed. Then, the dominant category of the current target operation segment is determined based on the number of support votes or the support ratio. Integrated voting refers to the process of uniformly summarizing the judgment results of multiple basic decision-making units and forming the final category output through majority support principles or weighted support principles. The target operation intent category refers to the category result determined as the true category of the current target operation segment after integrated voting. Specifically, a simple majority voting method can be used. First, the occurrence frequency of each candidate operation intent category is counted to obtain the number of support votes for each category. Then, the category with the largest number of support votes is selected as the target operation intent category. The expression is as follows:

[0053]

[0054]

[0055] In the formula, V(c) represents the number of votes supporting operation intention category c; R represents the total number of basic decision-making units; The indicator function represents the category of candidate operational intent output by the r-th basic decision unit. The value is 1 if it matches the current statistical category c, otherwise the value is 0; Represents a set of operational intent categories; This represents the final determined category of the target operation intent. The principle behind this formula is that by discretely counting the judgment results of multiple basic decision-making units, the category with the highest frequency is selected as the final category, effectively reducing the impact of misjudgments by individual basic decision-making units on the overall result. In practice, different weights can be assigned to different basic decision-making units to form a weighted voting mechanism. However, regardless of the form used, the essence is to improve classification stability through consensus among multiple judgment units.

[0056] When conflicting candidate operation intent categories exist, conflict resolution processing is performed based on the trend and periodic characteristics in the feature vectors to correct the target operation intent category. First, it is determined whether the integrated voting results show similar support votes, small differences in the support ratios of the first two categories, or multiple categories simultaneously reaching a preset support threshold. Conflict resolution processing refers to the process of using more discriminative key features to perform a secondary discrimination when multiple candidate operation intent categories are difficult to clearly distinguish based solely on voting results. Trend features reflect the enhancement, weakening, abrupt changes, and switching processes of the operation behavior over time, while periodic characteristics reflect the repetitive rhythm and frequency structure of the operation behavior; both are often important bases for distinguishing similar categories. For example, both rapid spinning and intense play may exhibit high resultant acceleration and resultant angular velocity, but rapid spinning usually has a more stable peak frequency and a more sustained rotation trend, while intense play often shows frequent switching of direction and a more discrete distribution of the peak frequency. In specific processing, the conflict category set can be targeted. Construct a conflict scoring function:

[0057]

[0058]

[0059] In the formula, T(c) represents the conflict resolution score for conflict category c; T(c) represents the trend matching score calculated around the trend characteristics, used to measure the degree of matching between the current feature vector and category c in terms of slope, frequency of mutation, consistency of change direction, etc.; P(c) represents the period matching score calculated around the periodic characteristics, used to measure the degree of matching between the current feature vector and category c in terms of dominant frequency position, spectral concentration, periodic stability, etc.; V(c) represents the number of support votes or the support ratio obtained by the original integrated voting for this category. , and These represent the weighting coefficients of the trend score, cycle score, and voting support item, respectively. This represents the target operational intent category obtained after conflict resolution. The principle behind this formula is to jointly re-discriminate the original voting results with more discriminative trend and cyclical characteristics, making the category boundaries closer to the actual behavioral evolution structure, thereby improving the ability to distinguish between similar categories.

[0060] Intent confidence parameters are generated based on the support of basic decision-making units. When delayed joint discrimination processing is performed when the intent confidence parameter is lower than a preset threshold, the credibility level of the current discrimination result is first calculated based on the degree of concentration, dominance, and conflict of support among multiple basic decision-making units for the target operation intent category. The intent confidence parameter is a quantitative parameter characterizing the reliability of the current target operation intent category discrimination result; a higher value indicates more concentrated and consistent support from multiple basic decision-making units for the current category. Delayed joint discrimination processing means that when the individual discrimination result of the current target operation segment is not reliable enough, the final category is not immediately output. Instead, it waits for subsequent adjacent time segments to generate new feature vectors before combining the information from multiple adjacent target operation segments for re-discrimination. Specifically, the intent confidence parameter can be defined as a combination function of the difference between the main category support ratio and the secondary main category support ratio, and the absolute support ratio of the main category, for example:

[0061]

[0062] In the formula, Q represents the intention confidence parameter; This indicates the number of votes supporting the category of the target operational intent; R represents the number of votes for the candidate operational intent category that received the second-highest number of votes; R represents the total number of basic decision-making units. and and represent the weight coefficients of the absolute support term and the relative dominance term, respectively. The principle behind this formula is to simultaneously consider the total support received by the main class and the dominance of the main class relative to the secondary main class, thus more accurately reflecting the stability of the classification results. When Q is lower than a preset threshold, instead of directly using the single discrimination result of the current target operation segment, the current target operation segment is combined with its preceding and following adjacent target operation segments to form a joint discrimination window. Multiple feature vectors within the joint window are then re-aggregated and discriminated to reduce misclassification caused by short-term sporadic jitter, single-frame errors, or local ambiguity. During joint discrimination, the feature vectors of multiple adjacent target operation segments can first be time-weighted and fused before being re-inputted into the ensemble classification model to complete a new category output, thereby improving the robustness of recognition in low-confidence scenarios.

[0063] When performing temporal smoothing on the target operation intent category corresponding to consecutive target operation segments to output the operation intent category, a category sequence of consecutive target operation segments is first established according to time order, and isolated category jumps with excessively short durations or significant inconsistencies with the preceding and following dominant categories are identified. Temporal smoothing refers to the process of correcting the initial category output of a single target operation segment for temporal consistency by utilizing the behavioral continuity constraints between adjacent time points. Its purpose is to avoid frequent jittering of the operation intent category between adjacent time segments due to local noise, momentary malfunctions, or short-term sensor fluctuations. In specific implementation, a sequence of categories with a length of [length missing] can be used. Within the local temporal neighborhood, majority neighborhood smoothing is performed on the category at the current time, expressed as:

[0064]

[0065] In the formula, This indicates the output operation intent category after time-series smoothing at time t; Represents the local time neighborhood of the i-th The initial target operation intent category for each target operation segment; This represents the neighborhood weight corresponding to the time position, and generally, segments closer to the current time have a larger weight; K represents the number of segments covered forward and backward by the local time neighborhood; This indicates an indicator function used to calculate class consistency. This represents the set of operational intention categories. The principle behind this formula is that real child operational behaviors usually have a certain degree of continuity. Therefore, the actual operational intention category at the current moment is likely to remain consistent with that at adjacent moments. By using the weighted majority principle within the local temporal neighborhood, transient abnormal outputs can be corrected to category results that better conform to continuous behavioral trajectories. The operational intention categories output after temporal smoothing retain sensitivity to real behavior switching while avoiding overly fragmented category sequences, thus making them more suitable for subsequent interest weight coefficient generation and feedback control strategy construction.

[0066] Furthermore, the ensemble classification model can be structurally organized into training and inference phases. In the training phase, multiple training subsets are first constructed based on historical samples. Then, a basic decision unit is generated for each training subset. Each basic decision unit randomly selects a subset of candidate features from all features during node splitting to enhance the differentiation between different basic decision units. In the inference phase, the current feature vector is simultaneously fed into all basic decision units for parallel discrimination. The candidate operation intent categories output by each basic decision unit undergo voting convergence, conflict resolution, confidence assessment, and temporal smoothing in the ensemble output layer. Thus, the ensemble classification model is not simply a stack of multiple classifiers, but rather a multi-perspective parallel discrimination system constructed through sample differences, feature differences, and discrimination path differences. This allows it to be sensitive to local features while maintaining strong overall noise resistance and generalization ability, making it particularly suitable for recognition scenarios involving rapidly changing, highly fluctuating, and ambiguous category boundaries, such as children's operational behaviors.

[0067] In one possible implementation, before performing multidimensional decision mapping on the feature vectors to identify the corresponding operation intent category, the method further includes: acquiring a multimodal operation sample data sequence covering multiple operation intent categories to construct an original sample set; performing time alignment, missing data completion, and sliding window segmentation on the original sample set to form a target operation fragment set; extracting sample feature vectors from each target operation fragment in the target operation fragment set to form a sample feature vector set; constructing an ensemble classification model containing multiple basic decision units based on the sample feature vector set; performing recursive split training on each basic decision unit to establish a nonlinear mapping relationship between the sample feature vectors and the operation intent category; and performing recognition performance verification and parameter tuning on the ensemble classification model based on the verification feature vector set to form a target ensemble classification model.

[0068] Specifically, when acquiring multimodal operation sample data sequences covering multiple operation intention categories to construct the original sample set, multiple rounds of sample collection are first organized around the actual use process of the target interactive carrier. This allows different children to complete various behaviors such as pressing, rotating, grasping, shaking, observing, continuous touching, rapid switching, and combined operations in a natural interactive state, thereby ensuring that the collection results cover multiple operation intention categories. The multimodal operation sample data sequence refers to the data sequence formed by the continuous change of inertial measurement data, pressure perception data, and external visual posture data recorded simultaneously during the same operation process over time. The original sample set refers to the untrained data population composed of multimodal operation sample data sequences corresponding to multiple operation intention categories. Operation intention categories are category labels used to characterize different operation behavior goals or behavioral states. In practice, to ensure the original sample set adequately covers real-world application scenarios, multimodal operation sample data sequences need to be repeatedly collected under different age groups, grip methods, operation speeds, and ambient lighting conditions. Each collection result is then labeled with a corresponding category, collection duration, and sample source. This ensures the original sample set contains not only category information but also contextual information to support subsequent training and validation. If the number of samples for a particular operation intent category in the original sample set is significantly insufficient, additional samples of that category are collected to reduce the risk of bias caused by class imbalance during subsequent training. To measure the coverage balance of each operation intent category in the original sample set, a category distribution ratio can be constructed:

[0069]

[0070] In the formula, This represents the proportion of the c-th type of operational intent in the original sample set; U represents the number of sample sequences corresponding to the c-th type of operation intent; U represents the total number of operation intent categories. This represents the sum of the number of sample sequences in all categories. This formula is used to characterize the distribution of each operational intent category in the original sample set. When the proportion of a certain category is significantly lower than the preset lower limit, samples of that category can be further supplemented, thereby improving the learning completeness of the subsequent integrated classification model for various operational behaviors.

[0071] When performing time alignment, missing data completion, and sliding window segmentation on the original sample set to form the target operation fragment set, a unified time reference frame is first established for each multimodal operation sample data sequence in the original sample set. Inertial measurement data, pressure sensing data, and external visual attitude data are then mapped to a unified sampling time, thereby eliminating time misalignment caused by inconsistent sampling frequencies of different sensors and differences in communication latency. Time alignment refers to aligning the data corresponding to the same physical moment of multiple sensor channels to a unified time reference. Missing data completion refers to restoring missing sample values ​​caused by temporary occlusion, instantaneous frame loss, or unstable transmission. Sliding window segmentation refers to continuously sliding an analysis window of a preset length along the time axis to extract multiple local continuous segments from the long sequence. The target operation fragment set is a set composed of all local continuous segments that meet the requirements of effective length and effective information density. In specific processing, resampling and interpolation are first performed on each multimodal operation sample data sequence. Then, the window length, window step size, and window overlap rate are set around each sequence to divide the entire sequence into multiple interrelated target operation segments. When the overall change intensity corresponding to a location with drastic data change exceeds a preset threshold, the window boundary can be offset to ensure that a complete action is located within the same target operation segment as much as possible, thereby reducing the phenomenon of action truncation across windows. The overall change intensity can be expressed as:

[0072]

[0073] In the formula, Indicates time The overall intensity of change at the location; This represents the difference between pressure sensing data at adjacent sampling times, used to reflect changes in the applied force state; The difference value of the three-axis acceleration vector is used to reflect the degree of abrupt change in linear motion; The difference value represents the three-axis angular velocity vector, used to reflect the degree of abrupt change in rotational state; The difference value represents the visual pose parameter vector, used to reflect the rate of change of external pose. , , and These represent the weight coefficients of the corresponding change terms; this formula identifies the turning points of the action structure by combining multi-channel change information, thereby improving the fit between the target operation segment set and the real behavior unit.

[0074] When extracting sample feature vectors from each target operation segment in the target operation segment set to form a sample feature vector set, firstly, multidimensional feature parameters representing the operation intensity, change trend, and periodic characteristics of each target operation segment are calculated, and then concatenated and encoded according to a fixed field order, thereby generating a sample feature vector with a unified structure for each target operation segment. Here, the sample feature vector refers to a multidimensional numerical expression used to represent the statistical attributes, temporal evolution attributes, and rhythmic attributes of a target operation segment, and the sample feature vector set is the set composed of the sample feature vectors corresponding to all target operation segments. Specifically, based on pressure sensing data, the mean pressure, peak pressure, root mean square pressure, and pressure change slope can be extracted; based on inertial measurement data, the peak resultant acceleration, mean resultant angular velocity, angular velocity fluctuation, and acceleration mutation frequency can be extracted; based on external visual attitude data, attitude deflection amplitude, attitude change rate, keypoint displacement statistics, and attitude maintenance duration can be extracted. Simultaneously, to enhance the ability of the sample feature vectors to represent repetitive behaviors, frequency domain expansion can be performed on the target operation segments to extract the dominant frequency characteristics, frequency band energy proportion, and spectral concentration. The resultant acceleration can be expressed as:

[0075]

[0076] In the formula, This represents the resultant acceleration at the nth sampling point; , and These represent the triaxial acceleration components at the same sampling point; this formula synthesizes the triaxial linear motion state into a single scalar, reflecting the overall motion intensity corresponding to the current sampling point. The pressure change slope can be expressed as:

[0077]

[0078] In the formula, This represents the slope of the average pressure change within the current target operation segment; This represents the pressure value at the nth sampling point; N represents the number of sampling points within the current target operation segment. This represents the time interval between adjacent sampling points; this formula describes whether the overall force application behavior in the current target operation segment is enhanced, weakened, or remains stable. After all features are extracted, each feature parameter is sequentially written into a unified location to form a sample feature vector. Then, all sample feature vectors are aggregated to form a sample feature vector set for subsequent ensemble classification model construction and training.

[0079] When constructing an ensemble classification model containing multiple basic decision units based on a set of sample feature vectors, the overall size of the ensemble classification model is first determined according to the number of samples, category distribution, feature dimensions, and target recognition complexity in the sample feature vector set. Then, multiple basic decision units are generated through random sampling and parallel construction. The ensemble classification model refers to the overall classification model composed of multiple basic decision units with differentiated discrimination paths connected in parallel. A basic decision unit is an independent discrimination unit capable of performing tree-like conditional splitting around several feature terms in the sample feature vector and outputting the category result. The structure of the ensemble classification model can be divided into an input layer, a parallel discrimination layer, and an ensemble output layer. The input layer receives sample feature vectors from the sample feature vector set and distributes each feature term to all basic decision units according to a unified field order. The parallel discrimination layer consists of multiple basic decision units, each learning from different training subsets and different candidate feature subspaces to form distinct local discrimination structures. The ensemble output layer receives the discrimination results of each basic decision unit on the same sample feature vector and forms the final category output accordingly. In practical implementation, multiple training subsets can be constructed from the sample feature vector set using sampling with replacement. Each training subset corresponds to a basic decision unit, ensuring that different basic decision units have natural differences in sample distribution. Simultaneously, during node splitting, each basic decision unit is only allowed to randomly select a portion of candidate feature terms from all feature terms to participate in the optimal split determination, thereby further enhancing the structural differences between basic decision units. If the first... The training subsets corresponding to each basic decision unit are denoted as follows: Then it can be expressed as:

[0080]

[0081] In the formula, This represents the training subset corresponding to the r-th basic decision unit; This represents the feature vector of the m-th training sample; This indicates the category label representing the operational intent of the sample. The formula represents the number of samples contained in the r-th training subset. This formula shows that each basic decision unit is a local discriminative structure built around a set of labeled sample feature vectors, so that multiple basic decision units together form an ensemble classification model with multi-view discriminative capabilities.

[0082] When performing recursive splitting training on each basic decision unit to establish a nonlinear mapping relationship between sample feature vectors and operational intent categories, each basic decision unit first starts with all samples in its corresponding training subset. At the root node, the proportion of each operational intent category in the current sample distribution is statistically analyzed. Then, from randomly selected candidate features, the splitting feature and splitting threshold that minimize category confounding are searched, and the current node is divided into two child nodes accordingly. The same splitting process is then repeated on each child node until the splitting stop condition is met. Here, recursive splitting training refers to the process of the tree-like discriminative structure continuously expanding from the parent node to the child node and progressively refining the sample partitioning boundary during training. The nonlinear mapping relationship refers to the complex mapping structure between sample feature vectors and operational intent categories formed through multi-layered condition combinations, asymmetric splitting, and the superposition of local rules, rather than a simple linear hyperplane segmentation. In specific implementation, Gini impurity can be used as the node splitting evaluation index. In this context, its impurity can be expressed as:

[0083]

[0084] In the formula, Represents the current node's sample set Gini impurity; U represents the total number of operational intent categories; This represents the proportion of samples of the c-th operation intent category in the current node's sample set; this formula describes the degree of category mixing within the current node's samples, with a smaller value indicating a more concentrated category. When a candidate feature term... and candidate threshold When used to split the current node's sample set, the split gain can be expressed as:

[0085]

[0086] In the formula, This indicates the amount of purity increase resulting from the current splitting operation; and They represent the characteristics respectively. and threshold The resulting set of left child node samples and the set of right child node samples; , and , respectively, represent the number of samples in the corresponding sample set; this formula measures the effectiveness of the current split by comparing the change in the degree of class mixing before and after the split. The system continuously repeats the above recursive splitting process for each basic decision unit, eventually forming a multi-layered discriminative path from the root node to the leaf node, thereby establishing a non-linear mapping relationship between the sample feature vector and the operation intention category.

[0087] When performing recognition performance verification and parameter tuning on the ensemble classification model based on the validation feature vector set to form the target ensemble classification model, the following steps are taken: First, a validation portion that does not participate in training is separated from all labeled samples. The same feature extraction process as in the training phase is then performed on this portion of samples to obtain the validation feature vector set. This set is then input into the ensemble classification model one by one. The model's classification results for each validation sample are statistically analyzed. Based on the statistical results, the model's recognition performance is quantitatively evaluated, and parameters such as the number of basic decision units, maximum node depth, minimum leaf node sample count, candidate feature sampling ratio, and voting threshold are adjusted accordingly. Recognition performance verification refers to testing the generalization ability of the current ensemble classification model on unseen samples using independent validation samples. Parameter tuning refers to reversing the model's structural and training parameters based on the performance evaluation results to achieve a better balance between recognition accuracy, stability, and real-time performance. The target ensemble classification model is the final model determined for online discrimination after performance verification and parameter tuning. In specific implementation, the overall recognition accuracy can be calculated.

[0088]

[0089] In the formula, Indicates the overall recognition accuracy; This represents the number of samples in the set of validation feature vectors that were correctly classified. This represents the total number of samples in the validation feature vector set; this formula measures the overall classification ability of the ensemble classification model for all validation samples. Additionally, recall can be calculated for each operational intent category.

[0090]

[0091] In the formula, This represents the recall rate for the c-th type of operational intent. This represents the number of samples that truly belong to class c and are correctly identified as class c; This represents the number of samples that truly belong to class c but are misidentified as other classes; this formula is used to measure the completeness of the model's recognition of a specific class. If the recall rate of certain operational intent classes is found to be low, or there is serious confusion between adjacent classes, the number of basic decision units is adjusted, the sampling ratio of candidate features is increased or decreased, or the node depth is restricted again, and the training and validation process is repeated until the overall recognition accuracy, recall rate of each class, and class stability meet the preset requirements; after the above recognition performance verification and parameter tuning, the target ensemble classification model suitable for online operational intent recognition of target interaction carriers is finally obtained.

[0092] In one possible implementation, interest weight coefficients are generated based on the category of operational intent and its persistent distribution characteristics over time. Specifically, this includes: constructing an operational intent time series based on the operational intent categories and time information corresponding to multiple consecutive target operational segments; performing category aggregation processing on the operational intent time series to form a set of intent activity segments; extracting persistent distribution features for each operational intent category based on the set of intent activity segments; performing confidence weighting processing on the persistent distribution features in conjunction with intent confidence parameters to form a persistent distribution feature expression vector; calculating an initial interest score based on the persistent distribution feature expression vector; performing time decay correction processing and sudden behavior suppression processing on the initial interest score to generate a corrected interest score; and performing normalization mapping processing on the corrected interest score to generate interest weight coefficients.

[0093] Specifically, when constructing an operational intention time series based on the operational intention categories and time information corresponding to multiple consecutive target operational segments, the operational intention category, start time, end time, segment duration, and segment number corresponding to each target operational segment are first uniformly registered according to the order of appearance of the target operational segments in the overall interaction process, and mapped onto the same continuous time axis, thereby forming an operational intention time series that reflects the evolution of children's operational behavior. Here, the operational intention time series refers to the temporal expression structure formed by arranging multiple operational intention category results obtained from discrete identification continuously according to their actual occurrence time. It not only retains the category result of each target operational segment but also retains the order of appearance, duration, and sequential relationship of the categories in the time dimension. Time information refers to the set of time attributes corresponding to each target operational segment, including at least the start time, end time, and duration. In specific implementation, the i-th target operational segment can be represented as a time-stamped temporal unit:

[0094]

[0095] In the formula, This represents the timing unit corresponding to the i-th target operation segment; This represents the operation intent category corresponding to the i-th target operation fragment; Indicates the start time of the i-th target operation segment; This indicates the end time of the i-th target operation segment; This represents the duration of the i-th target operation segment, typically determined by... The principle behind this formula is that it couples the category results with the time attribute into a unified temporal unit, enabling subsequent processing to move beyond static category labels and instead analyze the distribution patterns of categories over time within a continuous interactive process. After completing the temporal registration of all target operation segments, proceed as follows: By sorting from smallest to largest, a time series of operational intentions can be formed, thus providing a foundation for subsequent category aggregation and interest intensity assessment.

[0096] When performing category aggregation processing on the time series of operational intentions to form a set of intention activity segments, the first step is to determine the merging of adjacent time series units based on category consistency and temporal continuity. Multiple adjacent time series units with the same category and a time interval less than a preset threshold are merged into the same intention activity segment. Category aggregation processing refers to merging multiple originally discrete and local operational intention category results according to the principles of category consistency and temporal continuity, thereby extracting a higher-level continuous behavioral structure. The set of intention activity segments refers to a collection composed of several intention activity segments, each representing a stable activity state of a certain operational intention category within a continuous time interval. Specifically, if the i-th time series unit and the (i+1)-th time series unit satisfy... And the start and end intervals of the two Less than the preset aggregation threshold If the behavior is of the same category, it is considered a continuous continuation and merged; if the categories are different or the time interval exceeds a preset aggregation threshold, it is divided into different intent activity segments. The merged m-th intent activity segment can be represented as:

[0097]

[0098] In the formula, This represents the m-th intentional activity segment; This indicates the category of the operation intent corresponding to the intent activity segment; Indicates the overall start time of this intended activity segment; Indicates the end time of the entire intended activity segment; Indicates the overall duration of the intended activity segment; This indicates the number of target operation segments aggregated into the intended activity segment. The principle behind this formula is that by merging multiple local segments that are temporally adjacent and categorically unified into a whole activity segment, it can more realistically reflect whether children are continuously exploring around the same function or the same operation method, rather than being fragmented by the discrete recognition results of a single short segment.

[0099] When extracting persistent distribution features based on the set of intention activity segments for each category of operational intent, the set of intention activity segments is first grouped according to the operational intent category. Then, for each category of operational intent, persistent distribution features such as cumulative duration, continuous holding duration, repetition frequency, adjacent occurrence interval, time coverage ratio, category switching density, and revisit intensity are calculated within the current observation window. The persistent distribution features refer to the feature set used to characterize the distribution state of a certain operational intent category in the time dimension. They reflect not whether a single occurrence is valid, but rather the persistence, clustering, repetition, and regression of the category throughout the continuous interaction process. In specific implementation, the cumulative duration is used to characterize the total length of time a certain operational intent category is observed by children within the current observation window, which can be expressed as:

[0100]

[0101] In the formula, This represents the cumulative duration characteristic of operation intent category c; This represents the set of indices of intent activity segments belonging to intent category c; This represents the duration of the m-th intentional activity segment; this formula reflects the overall attention level of that category during the overall interaction process by summing the durations of all activity segments of the same category. The frequency of recurrence is used to characterize whether the child repeatedly returns to the same operational intention category, and can be expressed as:

[0102]

[0103] In the formula, This indicates the frequency of recurrence of operation intent category c; This indicates the number of intent activity segments belonging to this category; this formula shows that if a category appears multiple times within the observation window, its corresponding interest tendency has a repeated revisit characteristic. The time coverage ratio is used to characterize the proportion of a certain operational intent category within the overall observation window, and can be expressed as:

[0104]

[0105] In the formula, This indicates the time coverage percentage for operation intent category c; This represents the total duration of the current observation window; its principle lies in measuring the dominance of this category in the overall interaction time by normalizing the total duration. The interval between adjacent occurrences can be used to reflect the temporal density between two occurrences of the same category; the shorter the interval, the more obvious the temporal clustering of this category. Through the joint extraction of the above persistent distribution features, the original category results at the single-segment level can be expanded into a temporal distribution expression for the overall behavioral trajectory.

[0106] When performing confidence-weighted processing on persistent distribution features combined with intent confidence parameters to form a persistent distribution feature expression vector, the persistent distribution features corresponding to each operation intent category are first fused with the discriminative credibility of that category on each target operation segment. This ensures that the persistent distribution features not only reflect the temporal distribution state but also the credibility level of the recognition result. The intent confidence parameter refers to the category reliability quantification value generated for each target operation segment in the preceding multi-dimensional decision mapping stage. A higher value indicates a more stable and reliable category judgment for that target operation segment. Confidence-weighted processing involves using the intent confidence parameter to weight and correct various statistics in the persistent distribution features, making the time contribution of high-confidence segments greater and appropriately suppressing the time contribution of low-confidence segments. Specifically, the confidence-weighted cumulative duration of operation intent category c can be calculated first.

[0107]

[0108] In the formula, This represents the confidence-weighted cumulative duration of operation intent category c; This represents the average intent confidence parameter corresponding to the m-th intent activity segment; This represents the duration of the m-th intentional activity segment. The principle behind this formula is that if an intentional activity segment, although long in duration, has generally low confidence levels in the category judgments of its multiple target operation segments, its contribution to interest modeling should not be completely equivalent to that of high-confidence activity segments. Furthermore, features such as the confidence-weighted cumulative duration, confidence-weighted repetition frequency, confidence-weighted time coverage ratio, and continuous duration can be concatenated in a unified order to form a persistent distribution feature expression vector.

[0109]

[0110] In the formula, The vector representing the persistent distribution features of the operational intent category c; This represents the frequency of recurrence after confidence weighting; This indicates the confidence-weighted time coverage ratio characteristic; This represents the continuous dwell time characteristic after confidence weighting; This indicates the confidence-weighted intensity of follow-up visits; This represents the category switching density related features after confidence weighting; this formula incorporates multiple weighted persistent distribution indicators into the subsequent interest score calculation process through vectorization.

[0111] When calculating the initial interest score based on the persistent distribution feature representation vector, a comprehensive scoring rule is first set for each operation intention category, including persistent enhancement, repetition enhancement, coverage enhancement, and suppression. This ensures that the contributions of different persistent distribution features to the strength of interest are output in a unified numerical form. The initial interest score refers to the original interest intensity value directly mapped from the persistent distribution feature representation vector before time decay correction and sudden behavior suppression correction. It reflects the basic level of attention presented by a certain operation intention category within the current observation window. Specifically, a linearly weighted interest scoring function coupled with a suppression term can be constructed.

[0112]

[0113] In the formula, The initial interest score represents the operational intent category c; Indicates the confidence-weighted cumulative duration; Indicates the confidence-weighted frequency of repetition; Indicates the confidence-weighted time coverage ratio; Indicates the duration of confidence-weighted continuous holding; Indicates the confidence-weighted follow-up intensity; This indicates a confidence-weighted adjacent interval feature; This represents the confidence-weighted class switching density feature; to These represent the weight coefficients of each feature, and their values ​​are determined based on empirical rules, historical sample statistics, or offline calibration results. The principle behind this formula is that features that are conducive to representing stable interests, such as continuous, repetitive, and covering features, are used as positive contribution terms, while features that reflect distracted attention or unstable interests, such as interval and switching features, are used as negative suppression terms. This allows the initial interest score to simultaneously take into account both interest intensity and interest stability.

[0114] The initial interest score undergoes time decay correction and sudden behavior suppression processing. When generating the corrected interest score, the initial interest score is first modified based on the most recent occurrence time of the operational intent category, the historical time distribution of occurrence, and the structural characteristics of local high-frequency short-duration action clusters. Time decay correction reduces the influence of historical activity segments distant from the current time on the current interest judgment, giving more weight to recently occurring operational behaviors in interest modeling. Sudden behavior suppression identifies operational patterns that occur frequently within a short period but are extremely short-lived and lack continuous exploration characteristics, and suppresses their excessive amplification of the interest score. In practice, a time decay factor can be introduced for each intentional activity segment.

[0115]

[0116] In the formula, This represents the time decay factor corresponding to the m-th intentional activity segment; This represents the time decay coefficient; the larger the value, the faster the impact of historical behavior decreases. Indicates the current moment of interest assessment; This represents the end time of the m-th intentional activity segment. The principle behind this formula is that activity segments closer to the current time are more representative of the child's current true interest, and the influence of earlier activity segments should gradually decrease. Furthermore, burst behavior inhibitory factors can be constructed around short-duration, high-frequency activities.

[0117]

[0118] In the formula, Indicates the sudden behavioral inhibition factor for operational intent category c; This indicates the burst density of the category within a preset short-term observation interval, which is usually determined by the number of short-term repetitions and the average activity segment length. This represents the intensity coefficient of sudden behavior inhibition. The principle behind this formula is that if a certain category appears multiple times instantaneously within a very short period of time, but each occurrence is extremely brief, it is more likely to correspond to accidental touches, random tapping, or sporadic actions rather than stable interests. Therefore, its score should be compressed proportionally. After combining time decay and sudden inhibition, the corrected interest score can be expressed as:

[0119]

[0120] In the formula, This indicates the corrected interest score for category c of the operational intent; Indicates the initial interest rating; This represents the overall time decay factor obtained by weighted summation of all intentional activity segments in this category; This indicates a sudden behavior inhibition factor; by simultaneously introducing timeliness correction and abnormal activity inhibition, the corrected interest score is closer to the child's current real, stable and continuous interest state.

[0121] When performing normalization mapping on corrected interest scores to generate interest weight coefficients, the corrected interest scores corresponding to all operational intention categories are first aggregated into the same score set, and then normalized to map them into comparable relative weight values. The interest weight coefficient is a numerical coefficient representing the relative proportion of interest for each operational intention category within the current observation window. A larger value indicates a higher level of attention the child pays to the corresponding operational intention category, and a higher priority for guidance in subsequent feedback control strategies. In practice, proportional normalization can be used to generate the interest weight coefficients.

[0122]

[0123] In the formula, This represents the interest weight coefficient corresponding to the category c of operational intent; This indicates the corrected interest score for that category; Indicates the total number of operation intent categories; This represents the sum of all operational intent category-corrected interest scores; the principle behind this formula is to transform the absolute scores of each category into a relative proportion within the overall score, providing a unified scale for direct comparison between different categories. If it is necessary to enhance the separation of high-interest categories, an exponential normalization form can also be used.

[0124]

[0125] In the formula, The value represents the mapping temperature coefficient; the larger the value, the more significant the weight difference between high-rated and low-rated categories. This formula enhances the distinctiveness of the dominant interest category through an exponential amplification mechanism. After completing the normalization mapping, the set of interest weight coefficients corresponding to all operational intent categories is obtained. The category corresponding to the maximum value is taken as the current dominant interest intent category, thus providing a direct basis for constructing a feedback control strategy based on the interest weight coefficients and operational intent categories.

[0126] Furthermore, in collaborative interactive teaching aid scenarios, multiple children no longer interact with the teaching aids solely through their own independent actions. Instead, they form interconnected group behavior patterns around the target interactive carrier within the same time frame. Collaborative operation intentions typically manifest as multiple children performing rhythmically consistent, sequentially linked, or goal-oriented joint operations around the same functional area or interactive mechanism. For example, multiple children simultaneously tap the same teaching aid surface or alternately rotate around a rotating structure, thus jointly triggering a certain interactive effect. Disruptive behavior patterns, on the other hand, manifest as conflicts among multiple children regarding space occupancy, action rhythm, or functional competition. For instance, they may compete for the same knob, pressing area, or operable component, leading to overlapping operation sequences, conflicting action directions, or chaotic feedback triggering. Therefore, behavior recognition in such scenarios requires not only determining what operation a single child is performing but also further analyzing the synchronous, alternating, and conflicting relationships among multiple children in the temporal dimension to distinguish whether group interaction belongs to collaborative exploration, alternating cooperation, or disruptive competition.

[0127] In one possible implementation, interest weight coefficients are generated based on the operation intent category and its persistent distribution characteristics over time. Specifically, this includes: constructing a group operation time series based on the operation intent categories and time information corresponding to multiple target interaction carriers; performing cross-device category aggregation processing and group behavior segmentation processing on the group operation time series to form a set of group behavior activity segments; extracting persistent distribution features of the group based on the set of group behavior activity segments for each group operation intent category; performing joint confidence weighting processing on the persistent distribution features combined with intent confidence parameters, device activity parameters, and participation stability parameters to form a group distribution feature expression vector; calculating an initial group interest score based on the group distribution feature expression vector; performing group time decay correction processing, instantaneous aggregation suppression processing, and conflict enhancement identification processing based on the initial group interest score to generate a corrected group interest score; and generating interest weight coefficients based on the corrected group interest score.

[0128] Specifically, when constructing a group operation time sequence based on the operation intent categories and time information corresponding to multiple target interaction carriers, firstly, for each target interaction carrier in the group collaborative interactive teaching aid scenario, obtain its corresponding target operation segment sequence, operation intent category sequence, and segment time information. Then, map the operation intent category results on different target interaction carriers to the same global time axis to form a group operation time sequence that can simultaneously represent the group behavior change process of multiple devices, multiple children, and multiple moments. Here, the target interaction carrier refers to the teaching aid body or teaching aid functional unit in the group collaborative interactive teaching aid scenario that allows different children to touch, press, rotate, tap, or manipulate. The time information includes at least the start time, end time, duration, and time index position of the target operation segment. The group operation time sequence refers to the temporal expression structure formed by jointly arranging the operation intent categories identified on multiple target interaction carriers according to their occurrence order and concurrency relationship within a unified time framework. In specific implementation, the first... The temporal unit of a target interaction carrier on the i-th target operation segment is represented as:

[0129]

[0130] In the formula, This represents the i-th timing unit corresponding to the u-th target interaction carrier; This indicates the type of operation intent corresponding to the timing unit; Indicates the start time of this timing unit; Indicates the end time of this timing unit; The duration of this timing unit is typically represented by... The principle behind this formula is to uniformly convert the originally independent local category results on each target interaction carrier into group temporal units with time attributes, thereby enabling the identification in subsequent processing whether different target interaction carriers are in a state of synchronous activity, alternating activity, or conflicting activity at the same time. After completing the temporal registration of all target interaction carriers, they are sorted and organized in parallel according to the global time order to form a group operation time sequence.

[0131] When performing cross-device category aggregation and group behavior segmentation on the time sequence of group operations to form a set of group behavior activity segments, the process first searches for multiple time-series units that overlap in time, are related in category, and whose behavioral relationships meet preset conditions on a unified global timeline. These time-series units are then aggregated into the same group behavior activity segment. Cross-device category aggregation means that instead of simply merging based on the category continuity within a single target interaction carrier, the process further compares the degree of overlap in time and the degree of semantic association between multiple target interaction carriers to identify whether multiple children engage in synchronized tapping, alternating rotation, or competing behaviors around the same interactive theme. Group behavior segmentation means that the entire group interaction process is divided into multiple distinguishable group behavior activity segments based on the overlap, sequential connection, and conflict relationships between the time-series units of multiple devices. The set of group behavior activity segments is the total set of all group behavior activity segments. In specific processing, if time-series units on multiple target interaction carriers appear simultaneously in time and have the same operational intent category or satisfy a preset collaborative relationship, they are grouped into a collaborative behavior activity segment; if time-series units on multiple target interaction carriers overlap in time, exhibit action conflict, or class exclusion in the same functional area or the same type of component, they are grouped into a disruptive behavior activity segment. The time overlap ratio can be expressed as:

[0132]

[0133] In the formula, This represents the time overlap ratio between the i-th timing unit of the u-th target interaction carrier and the j-th timing unit of the v-th target interaction carrier; and These represent the start and end times of the preceding time unit, respectively; and These represent the start and end times of the next time series unit, respectively. This formula characterizes the degree of overlap between two time series units on the time axis. Aggregation is triggered when this ratio exceeds a preset threshold and the category relationship meets the conditions for cooperation or conflict. The m-th group behavior activity segment after aggregation can be represented as:

[0134]

[0135] In the formula, This represents the m-th group behavior activity segment; This indicates the category of group operational intent corresponding to the group's behavioral activity segment; Indicates the start time of the group's behavioral activity segment; Indicates the end time of the group's behavioral activity segment; Indicates the duration of the group's behavioral activity segment; This represents the set of target interaction carriers participating in the group's behavioral activities. This indicates the number of target interaction carriers participating in the group's behavioral activity segment. The principle behind this formula is to abstract the group interaction state formed by multiple target interaction carriers within the same time interval into a unified activity segment structure, thereby enabling subsequent group interest modeling to be built at the group level rather than the individual level.

[0136] When extracting persistent distribution features of groups based on the set of group behavior activity segments for each group's operational intention category, the set of group behavior activity segments is first classified and grouped according to the group's operational intention category. Then, persistent distribution features of the group are extracted based on the persistence, repetition, participation, and synchronization states of each group's operational intention category within the current observation window. Here, the group operational intention category refers to the category identifier defined for a scenario of joint interaction between multiple children and multiple target interactive devices, used to characterize the semantic direction of group behavior's coordination or interference, such as synchronized tapping, alternating rotation, sequential pressing, competition for possession, or conflict intervention. Persistent distribution features refer to the feature set used to characterize the distribution state of a certain group's operational intention category in the time and device dimensions. It reflects not only how long the category lasted and how many times it appeared, but also how many target interactive devices participated, whether the participation was stable, whether the group rhythm was consistent, and whether conflicts were concentrated. In specific implementation, at least the cumulative duration of the group, the continuous duration of the group, the frequency of repeated occurrences of the group, the device participation coverage ratio, the synchronization maintenance ratio, the alternation connection density, the conflict occurrence density, and the intensity of follow-up aggregation can be extracted. The cumulative duration of the group can be expressed as:

[0137]

[0138] In the formula, The cumulative duration of the group's operational intent category c; This represents the set of indexes for group behavior activity segments belonging to group operation intent category c; This represents the duration of the m-th group behavior activity segment; this formula describes the total duration for which a certain group's operational intention category is maintained within the observation window. The device participation coverage ratio can be expressed as:

[0139]

[0140] In the formula, The percentage of devices participating in coverage, indicating the category c of group operational intent. This represents the union of all target interaction carriers that have participated in group behavior activities belonging to the group's operational intent category; U represents the total number of target interaction carriers; this formula reflects the coverage of a certain group's operational intent category in the device dimension, and the larger the value, the more target interaction carriers are involved in the group behavior. The synchronization maintenance ratio can be expressed as:

[0141]

[0142] In the formula, Indicates the synchronization maintenance ratio of group operation intention category c; This represents the number of target interaction carriers participating in the m-th group behavior activity segment; its principle lies in measuring what proportion of target interaction carriers participate simultaneously during the duration of this type of group behavior. Through the above feature extraction, the set of group behavior activity segments can be transformed into a multidimensional temporal distribution expression for group interest modeling.

[0143] When performing joint confidence weighting processing on the persistent distribution characteristics of a group, combining intent confidence parameters, device activity parameters, and participation stability parameters to form a group distribution characteristic expression vector, the intent confidence parameters corresponding to all target interaction carriers within each group's behavioral activity segment are first aggregated. Then, the extracted persistent distribution characteristics of the group are weighted and corrected by combining the activity level and participation stability of each target interaction carrier within the observation window. Among them, the intent confidence parameter refers to the category reliability quantification parameter output by the preceding multidimensional decision mapping stage for a single target operation segment. The larger the value, the more stable the corresponding category judgment. The device activity parameter refers to the quantification parameter used to characterize the degree of operation participation of a target interaction carrier within the current observation window, such as the total duration of the target interaction carrier being triggered, the total trigger frequency, or the cumulative value of local action energy. The participation stability parameter refers to the quantification parameter used to characterize whether the continuous participation of a target interaction carrier in a certain group behavioral activity segment is stable, such as the continuous participation ratio, the reciprocal of the disconnection rate, or the time overlap stability. The purpose of joint confidence weighting is to avoid excessive interference from individual low-confidence, low-activity, or transiently participating target interaction carriers in the judgment of group interests, so that the final group distribution feature expression vector focuses more on those group interaction units that are reliably identified, fully participated, and whose behavior is consistently stable. In specific implementation, the joint weighting factor for the m-th group behavior activity segment can be defined first:

[0144]

[0145] In the formula, This represents the joint weighting factor for the m-th group behavior activity segment; This indicates the number of target interactive carriers participating in the group's behavioral activities. The average intent confidence parameter of the target interaction carrier u during the group's behavioral activity segment; Represents the target interaction carrier The active parameters of the equipment during the group's behavioral activity segment; Represents the target interaction carrier Participation stability parameters during the group's behavioral activity segment; , and These represent the weighting coefficients for the three types of parameters. The principle behind this formula is to first integrate the reliability, activity, and stability of different target interaction carriers within an activity segment, and then correct the group's persistence distribution characteristics on a unit of activity segment. The weighted cumulative group persistence duration can be expressed as:

[0146]

[0147] In the formula, This represents the joint confidence-weighted cumulative duration of group operational intent category c; the principle is to give greater contribution to the total duration of high-confidence, high-activity, and high-stability group behavior segments. Subsequently, all weighted group persistence distribution features can be concatenated in a fixed field order to form a group distribution feature representation vector.

[0148]

[0149] In the formula, A vector representing the group distribution characteristics of group operation intention category c; This indicates the frequency of group repetitions after joint confidence weighting; This indicates the coverage ratio of devices after joint confidence weighting; This indicates the proportion of simultaneous maintenance after joint confidence weighting; This represents the alternation connection density after joint confidence weighting; This represents the conflict occurrence density after joint confidence weighting; this formula provides a unified input for subsequent group interest score calculation through vectorization encoding.

[0150] When calculating the initial group interest score based on the group distribution feature expression vector, a group-level scoring rule is first established for each category of group operational intent. Features that reflect the group's synergy strength, coverage, and behavioral persistence are used as positive contribution terms, while features that reflect conflict concentration, dispersed participation, or unstable switching are used as suppression terms, thus obtaining the corresponding initial group interest score. The initial group interest score refers to the original group interest intensity value directly obtained by mapping the group distribution feature expression vector before introducing group time decay, instantaneous aggregation suppression, and conflict intensification correction. It is used to characterize the basic dominance of a certain group's operational intent category within the current observation window. Specifically, the following scoring expression can be constructed:

[0151]

[0152] In the formula, The initial group interest score represents the category c of the group's operational intention. Indicates the weighted cumulative duration of joint confidence; Indicates the frequency of repeated occurrences based on joint confidence weighting; This indicates the coverage ratio of jointly trusted weighted devices; This indicates the joint confidence weighted synchronous maintenance ratio; Indicates the joint confidence weighted alternation connection density; Indicates the joint confidence weighted conflict density; to These represent the scoring weight coefficients for the corresponding features. The principle behind this formula is that if a group's operational intention category persists for a long time within the observation window, appears repeatedly, covers multiple target interaction carriers, has a high degree of synchronization, or exhibits obvious alternating cooperation, then the group's behavior is more likely to correspond to genuine shared group interests and should be given a higher score. Conversely, if the density of conflict occurrences is too high, then the group's behavior is more likely to correspond to competition or interference, and its scoring level as a positive interest should be suppressed to leave room for subsequent conflict reinforcement identification.

[0153] Based on the initial group interest score, group time decay correction, instantaneous aggregation suppression, and conflict enhancement identification are performed. When generating the corrected group interest score, the initial group interest score is first modified from three dimensions: timeliness, short-term abnormal aggregation, and group conflict risk. Specifically, group time decay correction reduces the contribution of historical group behavior segments that are far removed from the current time to the current group interest judgment, making the most recent group interaction more representative of the current group's focus. Instantaneous aggregation suppression compresses instantaneous aggregation events that suddenly trigger multiple target interaction carriers simultaneously within a very short time but lack stable follow-up behavior, preventing accidental gatherings, random collisions, or short-term imitation actions from being misjudged as stable group interests. Conflict enhancement identification increases the intervention weight for group behaviors with high conflict density, large category dispersion, and severe temporal overlap, enabling subsequent feedback control strategies to prioritize and handle group interference behaviors. Specifically, the group time decay factor for the m-th group behavior segment can be expressed as:

[0154]

[0155] In the formula, This represents the group time decay factor corresponding to the m-th group behavior activity segment; Indicates the group time decay coefficient; Indicates the current moment of group interest assessment; This represents the end time of the m-th group behavior segment; the principle behind this formula is that group behavior closer to the current moment better reflects the current true group state. The instantaneous aggregation inhibition factor can be expressed as:

[0156]

[0157] In the formula, The instantaneous aggregation inhibition factor representing the group's operational intent category c; This indicates the instantaneous cluster density of the group's operational intent category within a preset short-term observation interval; This represents the group aggregation suppression coefficient; this formula is used to compress the amplification effect of short-term anomalous aggregation on group interest scores. For conflict reinforcement identification, a conflict reinforcement factor can be constructed:

[0158]

[0159] In the formula, The conflict intensification factor representing category c of group operational intent; Indicates the conflict intensification coefficient; This represents the conflict occurrence density after joint confidence weighting; the principle behind this formula is that when group conflict behavior continues to intensify, the salience of the corresponding category in intervention priority should be increased. After combining the above three types of corrections, the corrected group interest score can be expressed as:

[0160]

[0161] In the formula, The corrected group interest score represents category c of the group's operational intent. This indicates the initial group interest score for this category; This represents the overall population time decay factor corresponding to this category; Indicates the transient aggregation inhibition factor; This represents the conflict intensification factor; by simultaneously considering the timeliness, abnormal aggregation, and conflict risk of group behavior, this formula enables the corrected group interest score to both characterize positive group collaborative interests and highlight group conflict states that require priority intervention.

[0162] When generating interest weight coefficients based on corrected group interest scores, the corrected group interest scores corresponding to all group operational intention categories are first aggregated into the same score set. Then, a uniform scale normalization mapping is performed on these scores, allowing comparisons between different group operational intention categories in the form of relative weights. In the context of group collaborative interactive teaching aids, the interest weight coefficient characterizes the relative dominance or relative intervention priority of each group operational intention category within the current observation window. A higher value indicates that the group operational intention category is more likely to represent the interaction direction that multiple children are currently focusing on, or that it should be prioritized by the system as a conflict state requiring focused attention. In specific implementation, a proportional normalization method can be used to generate the interest weight coefficients.

[0163]

[0164] In the formula, This represents the interest weight coefficient corresponding to category c of the group's operational intent; V represents the corrected group interest score for this category; V represents the total number of categories of group operational intent. This represents the sum of the corrected group interest scores for all group operation intention categories. The principle behind this formula is to transform the absolute scores of each group operation intention category into a relative proportion within the overall score, allowing the system to directly compare the dominance of different group behavior patterns, such as synchronized tapping, alternating rotation, and competing behavior, in the current group interaction process. If it is necessary to further enhance the distinction between dominant and non-dominant group operation intention categories, an exponential normalization mapping method can also be used.

[0165]

[0166] In the formula, The group mapping temperature coefficient represents the group's behavior. A larger value indicates a more significant difference in interest weights between the high-scoring and low-scoring groups' operational intention categories. This formula enhances the recognizability of the main group's behavioral patterns through an exponential amplification mechanism. After normalization mapping, a set of interest weight coefficients for group collaborative interaction teaching aid scenarios is obtained. Furthermore, the group operational intention category corresponding to the largest interest weight coefficient can be identified as the current main group's interest category, and the group operational intention category whose interest weight coefficient significantly increases after conflict intensification can be identified as the current priority intervention category. This provides a direct basis for constructing group feedback control strategies based on interest weight coefficients and operational intention categories.

[0167] In one possible implementation, a feedback control strategy is constructed based on interest weight coefficients and operation intent categories, and the execution unit is driven to output a response signal. Specifically, this includes: constructing a set of group feedback decision parameters based on interest weight coefficients, operation intent categories, intent confidence parameters, number of participating devices, device participation coverage ratio, synchronization maintenance ratio, alternating connection density, and conflict occurrence density; performing group state discrimination processing on the group feedback decision parameter set to generate a group interaction state; establishing a set of strategy mapping rules and determining the control objective based on the group interaction state and operation intent categories; calculating strategy priority parameters based on the operation intent category, group interaction state, and control objective; constructing a collaborative feedback strategy set based on the operation intent category, group interaction state, and strategy priority parameters; performing execution resource mapping processing on the collaborative feedback strategy set to form a response signal generation scheme; and performing timing orchestration and intensity adaptive processing on the response signal generation scheme to generate a target response signal sequence, thereby driving the execution unit to perform cross-device collaborative output according to the target response signal sequence to generate a response signal.

[0168] Specifically, when constructing the group feedback decision parameter set based on interest weight coefficients, operation intention categories, intention confidence parameters, number of participating devices, device participation coverage ratio, synchronization maintenance ratio, alternation connection density, and conflict occurrence density, the core discriminant quantities of multiple target interaction carriers in the group interaction process are first uniformly aggregated based on the group interest assessment results and group behavior identification results already obtained within the current observation window. These are then organized into a group feedback decision parameter set for subsequent control decisions according to a fixed field order. Among them, the interest weight coefficient is used to characterize the relative dominance of each group's operation intention category in the current time period; the operation intention category is used to characterize the semantic type of the current group behavior; the intention confidence parameter is used to characterize the reliability of the preceding identification results; the number of participating devices is used to characterize the number of target interaction carriers that actually participate in the interaction in the current group behavior; the device participation coverage ratio is used to characterize the coverage range of all target interaction carriers involved in the current group behavior; the synchronization maintenance ratio is used to characterize the degree to which multiple target interaction carriers maintain consistent behavior in the same time interval; the alternation connection density is used to characterize whether there is a clear rotational succession relationship between multiple target interaction carriers; and the conflict occurrence density is used to characterize the concentration of interference behaviors such as competition, overlapping occupation, and adversarial operation in the time dimension. In practice, a corresponding group feedback decision parameter vector can be constructed for each group's operational intent category:

[0169]

[0170] In the formula, This represents the group feedback decision parameter vector corresponding to the group operation intention category c; This represents the interest weight coefficient corresponding to this category; This represents the comprehensive intent confidence parameter corresponding to this category; This indicates the number of participating devices corresponding to this category; This indicates the coverage percentage of devices corresponding to this category; This indicates the synchronization maintenance ratio corresponding to this category; This indicates the alternation density corresponding to this category; This indicates the conflict density corresponding to the category. The principle behind this formula is to encapsulate multiple core quantities that characterize the group's interest level, participation scale, coordination status, and conflict status into the same parameter space, so that subsequent group state discrimination processing can be calculated based on structurally consistent data, thereby ensuring that the generation of feedback control strategies is based on a unified, complete, and comparable group discrimination foundation.

[0171] When performing group state discrimination processing on the set of group feedback decision parameters to generate group interaction states, a comprehensive analysis is first performed on the group feedback decision parameter vectors corresponding to the operational intention categories of each group within the current time period. Then, the interaction state of the current group behavior is discriminated based on the combination relationship between the parameters. Among them, group state discrimination processing refers to the process of mapping the group's interest level, group synchronization level, group rotation level, group participation scope, and group conflict level into a finite number of executable control states. The group interaction state is a state label defined for the group collaborative interaction teaching aid scenario, used to characterize whether the current multiple children are more likely to explore collaboratively, take turns to cooperate, have scattered attention, or compete for attention. In practical implementation, at least four types of group interaction states can be defined: enhanced collaboration state, alternating collaboration state, distracted attention state, and conflict intervention state. When a group's operational intent category has a high interest weight coefficient, and both the synchronization maintenance ratio and the device participation coverage ratio are high, it can be determined as an enhanced collaboration state. When a group's operational intent category has a high interest weight coefficient, and the alternation density is significantly higher than the synchronization maintenance ratio, it can be determined as an alternating collaboration state. When the interest weight coefficients of multiple group operational intent categories are similar, and the device participation coverage ratio is wide but the category consistency is insufficient, it can be determined as a distracted attention state. When the conflict occurrence density continues to increase and the overall intent confidence parameter remains at a high level, it can be determined as a conflict intervention state. To achieve the above state discrimination, a group state scoring function can be constructed:

[0172]

[0173] In the formula, This represents the state matching score of group operation intention category c under group interaction state s; to These represent the weight coefficients corresponding to the group interaction state s. The values ​​of these weight coefficients differ across different group interaction states to reflect that the enhanced collaboration state emphasizes the synchronization maintenance ratio and the equipment participation coverage ratio, the collaborative rotation state emphasizes the alternation density, and the conflict intervention state emphasizes the difference in conflict occurrence density. The principle behind this formula is that by assigning different state-specific weights to the same set of group feedback decision parameters, it is possible to determine which control state the current group behavior is closer to. Subsequently, the state with the highest score is selected as the current group interaction state, thus providing a state basis for determining subsequent control objectives.

[0174] When establishing a set of strategy mapping rules and determining control objectives based on group interaction states and operational intent categories, the first step is to construct a set of strategy mapping rules under the dual constraints of state and category, using group interaction states as the state axis and operational intent categories as the behavioral semantic axis. This ensures that the same operational intent category can be assigned to different control objectives under different group interaction states. The set of strategy mapping rules refers to a set of mapping relationships that are pre-defined or obtained offline, used to map group interaction states and operational intent categories to a clear control objective. The control objective is the guiding or intervention direction that the feedback control strategy attempts to achieve in the current time period, such as strengthening joint exploration, maintaining rotation order, focusing attention, or alleviating competition and conflict. In specific handling, when the group interaction state is one of enhanced collaboration and the operational intent is synchronous tapping or pressing, the control objective can be set as strengthening the consistency of the group's rhythm and extending the duration of joint exploration. When the group interaction state is one of collaborative rotation and the operational intent is alternating rotation or sequential triggering, the control objective can be set as maintaining rotation order and enhancing the continuity of subsequent interactions. When the group interaction state is one of distracted attention, regardless of the neutral interaction category of the operational intent, the control objective can be prioritized as converging the group's attention center and highlighting the main interaction direction. When the group interaction state is one of conflict intervention and the operational intent is one of competing for possession or conflict-overlapping behavior, the control objective can be set as reducing the intensity of competition, guiding the diversion of behavior, and restoring orderly interaction relationships. The principle behind this approach is that simply relying on the operational intent category cannot determine whether the feedback should focus on enhancement or inhibition, and simply relying on the group interaction state cannot clarify what interactive theme the feedback should revolve around. Therefore, it is necessary to establish a two-dimensional mapping by combining the group interaction state and the operational intent category to ensure that the control objective conforms to both the current semantics of the group behavior and the current risk state of the group interaction.

[0175] When calculating the strategy priority parameter based on the operational intent category, group interaction state, and control objective, the priority order for entering the feedback control process is first calculated for all candidate group operational intent categories within the current time period. This determines which behavior mode the system should prioritize for feedback output when multiple group behaviors coexist. The strategy priority parameter is a quantitative parameter characterizing the degree to which a particular group's operational intent category is prioritized within the current control cycle. A higher value indicates that the category should be prioritized in the collaborative feedback strategy set construction process. In specific implementation, the interest weight coefficient, comprehensive intent confidence parameter, equipment participation coverage ratio, synchronization maintenance ratio, and alternating connection density can be considered positive promoting factors, while conflict occurrence density can be considered an intervention priority promoting factor. The contribution weights of each factor are dynamically adjusted based on the group interaction state and control objective. For example, in a collaborative enhancement state, the weights of the interest weight coefficient and synchronization maintenance ratio should be increased; in a collaborative alternation state, the weight of the alternating connection density should be increased; and in a conflict intervention state, the weights of the conflict occurrence density and equipment participation coverage ratio should be increased. Accordingly, the strategy priority parameter expression can be constructed as follows:

[0176]

[0177] In the formula, Indicates the category of group operation intention The corresponding strategy priority parameter; Indicates the current group interaction state; Indicates the current control objective; to These represent the interaction states of a specific group. and control objectives Priority weight coefficients for each input item; , , , , and These represent the interest weight coefficient, comprehensive intent confidence parameter, device participation coverage ratio, synchronization maintenance ratio, alternation connection density, and conflict occurrence density for the corresponding categories, respectively. The principle of this formula is that, through dynamic weight design based on state correlation and goal correlation, the priority ranking is no longer fixed, but can adaptively change according to the nature of the current group interaction and the control direction that the system wants to achieve, thereby ensuring that feedback control resources are preferentially applied to the most critical group behavior patterns.

[0178] When constructing a collaborative feedback strategy set based on operational intent categories, group interaction states, and strategy priority parameters, candidate group operational intent categories are first sorted from high to low according to the strategy priority parameters. Then, for the top-ranked group operational intent categories, corresponding collaborative feedback strategy items are generated one by one, taking into account the current group interaction state and control objectives. These strategy items are then organized into a collaborative feedback strategy set. The collaborative feedback strategy set refers to a collection of executable feedback strategies established for multiple potentially coexisting group behavior patterns within the current control cycle. It can contain various types such as collaborative guidance strategies, rotation coordination strategies, attention focus strategies, and conflict mitigation strategies. In practical implementation, when the group interaction state is in a collaborative enhancement state and the high-priority group operation intention category is synchronous tapping, rhythm-enhanced voice prompt strategy, synchronous light resonance strategy, and linked image feedback strategy can be added to the collaborative feedback strategy set; when the group interaction state is in a collaborative rotation state and the high-priority group operation intention category is alternating rotation, a round prompt strategy, a sequential light flow strategy, and a successive completion reward strategy can be added; when the group interaction state is in a distracted state, a focus target highlighting strategy, a non-dominant target weakening strategy, and a main interaction direction prompting strategy can be added; when the group interaction state is in a conflict intervention state and the high-priority group operation intention category is competitive behavior, a conflict area desensitization strategy, an alternative path guidance strategy, and a rhythm slowing down suppression strategy can be added. The principle behind constructing this collaborative feedback strategy set is that different group behavior patterns require different feedback semantics and different feedback functional structures, and the strategy priority parameter can determine which strategy items are prioritized when resources are limited or time is tight, thus making the collaborative feedback strategy set both targeted to the current group state and practically executable.

[0179] When performing resource mapping processing on the collaborative feedback strategy set to form a response signal generation scheme, the control intent corresponding to each strategy item in the collaborative feedback strategy set is first mapped to a combination of feedback resources that can be invoked within the target interactive carrier. Then, the resource invocation objects are allocated according to the number of target interactive carriers, spatial distribution relationships, and the scope of strategy application. Among them, the resource mapping processing refers to the process of converting abstract strategy items into concrete execution resource configuration relationships such as sound, light, display, and motion. Execution resources refer to hardware feedback resources or feedback content resources in the target interactive carrier that can be directly scheduled by the control unit, such as the voice resources of the speaker unit, the lighting and display resources of the light and display panel, the motion resources of mechanical execution components, and vibration or rhythm output resources. In specific processing, a resource mapping template can be first established for each type of strategy item. For example, the collaborative guidance strategy is preferentially mapped to synchronous voice prompt resources, globally consistent lighting resources, and positive reward display resources; the rotation coordination strategy is preferentially mapped to sequential broadcast voice resources, flowing light resources, and staged action resources; the attention focus strategy is preferentially mapped to main area highlight resources, non-main area dimming resources, and focus prompt voice resources; and the conflict mitigation strategy is preferentially mapped to diversion prompt voice resources, conflict area desensitization display resources, and mechanical action deceleration resources. Then, a response signal generation scheme is generated based on the current set of controlled objects. If the k-th strategy item is assigned to the... The resource mapping results on each target interaction carrier are denoted as The overall response signal generation scheme can then be expressed as:

[0180]

[0181] In the formula, This indicates the response signal generation scheme; This represents the resource mapping unit of the k-th strategy item on the u-th target interaction carrier; This represents the set of indexes of the selected strategy items; Let represent the set of target interaction carriers that the k-th strategy item actually affects. The principle behind this formula is that by establishing a correspondence between strategy items and specific resource instances, the collaborative feedback strategy set can be transformed into a response execution plan that can be directly scheduled in the future.

[0182] When performing timing orchestration and intensity adaptive processing on the response signal generation scheme to generate a target response signal sequence, and driving the execution unit to perform cross-device collaborative output according to the target response signal sequence to generate response signals, firstly, all resource mapping units in the response signal generation scheme are uniformly scheduled in time to determine the trigger time, duration, synchronous output relationship, alternating output relationship, and execution order of each resource mapping unit. Then, the output intensity of each resource is adaptively adjusted according to the current group interest weight coefficient, strategy priority parameter, and group interaction state, ultimately forming a target response signal sequence that can directly drive the execution unit. Timing orchestration refers to the process of coordinating the output behavior of multiple target interaction carriers in the time dimension, aiming to keep the group feedback consistent with the rhythm of the children's group behavior and avoid resource output conflicts. Intensity adaptive processing refers to the process of dynamically adjusting output intensity parameters such as voice volume, light brightness, display salience, and mechanical movement speed according to the current group behavior intensity, the degree of group interest concentration, and the degree of conflict risk. The target response signal sequence refers to the feedback signal sequence arranged in chronological order that can be directly executed by the execution unit. In specific implementation, the response parameter of the z-th output unit can be defined as follows:

[0183]

[0184] In the formula, This represents the z-th response signal unit; This indicates the resource type and output content identifier corresponding to the response signal unit; This indicates the start time of the response signal unit's output; This indicates the duration of the response signal unit; This represents the output intensity of the response signal unit. The principle behind this formula is that each feedback action to be output is uniformly encoded into a signal unit determined by both time and intensity parameters. Then, by sorting all the signal units, a target response signal sequence can be formed. The output intensity can be further defined as:

[0185]

[0186] In the formula, Indicates the output strength of the current response signal unit; This represents the interest weight coefficient for the corresponding group's operational intention category; This indicates the strategy priority parameter for the corresponding category; Indicates the synchronization maintenance ratio; Indicates alternating connection density; Indicates the density of conflict occurrences; to These represent the contribution coefficients of each input to the intensity adjustment. The principle behind this formula is that when the group interest weight coefficient is high, the strategy priority is high, the synchronization maintenance ratio or alternation density is high, the system should enhance the significance of the corresponding positive feedback signal. Conversely, when the conflict density increases, the stimulating output should be reduced and the guiding output increased to avoid feedback exacerbating the conflict. After completing the timing arrangement and intensity adaptive processing, a target response signal sequence can be generated. The execution unit then performs cross-device collaborative output among multiple target interaction carriers according to the target response signal sequence. For example, in the collaborative enhancement state, multiple target interaction carriers synchronously output lights and encouraging voices with consistent rhythms; in the collaborative rotation state, multiple target interaction carriers sequentially output round prompt signals; and in the conflict intervention state, the target interaction carriers directly related to the conflict output inhibitory prompt signals instead of the path-related target interaction carriers outputting transfer guidance signals. After this processing, the feedback control strategy can start from the interest weight coefficient and the category of operation intention, and through parameter integration, state discrimination, strategy generation, resource mapping and time-series scheduling, finally form a response signal that matches the current group's interest tendency, and achieve organized, rhythmic and guiding collaborative output among multiple target interaction carriers.

[0187] This embodiment also discloses a feedback control device based on sensor fusion and operation intention recognition, referring to... Figure 3 The device includes an acquisition module 301, a processing module 302, and an output module 303. It is used to execute any of the above-described feedback control methods based on sensor fusion and operation intention recognition, wherein: The acquisition module 301 is used to acquire inertial measurement data, pressure sensing data and external visual attitude data set in the target interactive carrier, and perform preprocessing to form a multimodal operation data sequence. The processing module 302 is used to perform sliding window segmentation processing based on the multimodal operation data sequence, and extract multidimensional feature parameters that reflect the operation intensity, change trend and periodic characteristics to construct feature vectors; Processing module 302 is used to perform multidimensional decision mapping on feature vectors to identify the corresponding operation intent category; Processing module 302 is used to generate interest weight coefficients based on the category of operation intention and its continuous distribution characteristics in the time dimension; The output module 303 is used to construct a feedback control strategy based on the interest weight coefficient and the operation intention category, and drive the execution unit to output a response signal.

[0188] It should be noted that the above embodiments of the apparatus are only illustrated by the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the apparatus and method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.

[0189] This embodiment also discloses an electronic device, as shown in the reference. Figure 4 The electronic device may include: at least one processor 401, at least one communication bus 402, user interface 403, network interface 404, and at least one memory 405.

[0190] The communication bus 402 is used to enable communication between these components.

[0191] The user interface 403 may include a display screen and a camera. Optionally, the user interface 403 may also include a standard wired interface and a wireless interface.

[0192] The network interface 404 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface).

[0193] The processor 401 may include one or more processing cores. The processor 401 connects to various parts of the server using various interfaces and lines, and performs various server functions and processes data by running or executing instructions, programs, code sets, or instruction sets stored in memory 405, and by calling data stored in memory 405. Optionally, the processor 401 may be implemented using at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). The processor 401 may integrate one or a combination of several of the following: Central Processing Unit (CPU), Graphics Processing Unit (GPU), and modem. The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the content required for display; and the modem handles wireless communication. It is understood that the modem may also not be integrated into the processor 401 and may be implemented as a separate chip.

[0194] The memory 405 may include random access memory (RAM) or read-only memory. Optionally, the memory may include a non-transitory computer-readable storage medium. The memory 405 may be used to store instructions, programs, code, code sets, or instruction sets. The memory 405 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the above-described method embodiments, etc.; the data storage area may store data involved in the above-described method embodiments, etc. Optionally, the memory 405 may also be at least one storage device located remotely from the aforementioned processor 401. As a computer storage medium, the memory 405 may include an operating system, a network communication module, a user interface 403 module, and an application program of a feedback control method based on sensor fusion and operation intention recognition.

[0195] exist Figure 4 In the electronic device shown, the user interface 403 is mainly used to provide an input interface for the user and to obtain the user input data; while the processor 401 can be used to call an application program stored in the memory 405 that is a feedback control method based on sensor fusion and operation intention recognition. When executed by one or more processors 401, the electronic device executes one or more methods as described in the above embodiments.

[0196] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the present invention is not limited to the described order of actions, as some steps can be performed in other orders or simultaneously according to the present invention. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to the present invention.

[0197] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0198] In the several embodiments provided by this invention, it should be understood that the disclosed apparatus can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some service interface; the indirect coupling or communication connection between apparatuses or units may be electrical or other forms.

[0199] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0200] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0201] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage device (CMD). Based on this understanding, the technical solution of this invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory 405 and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this invention. The aforementioned memory 405 includes various media capable of storing program code, such as a USB flash drive, external hard drive, magnetic disk, or optical disk.

[0202] The present invention also discloses a non-transitory computer-readable storage medium storing instructions. When executed by one or more processors 401, these instructions cause an electronic device to perform one or more methods as described in the above embodiments.

[0203] The above are merely exemplary embodiments of this disclosure and should not be construed as limiting the scope of this disclosure. Any equivalent changes and modifications made in accordance with the teachings of this disclosure shall still fall within the scope of this disclosure. Those skilled in the art will readily conceive of other embodiments of this disclosure upon considering the specification and the disclosure of practical truths. This invention is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not described in this disclosure. The specification and embodiments are to be considered exemplary only, and the scope and spirit of this disclosure are defined by the claims.

Claims

1. A feedback control method based on sensor fusion and operation intention recognition, characterized in that, The method includes: Acquire inertial measurement data, pressure sensing data, and external visual attitude data set in the target interactive carrier, and perform preprocessing to form a multimodal operation data sequence; Based on the multimodal operation data sequence, a sliding window segmentation process is performed, and multidimensional feature parameters reflecting operation intensity, change trend and periodic characteristics are extracted to construct a feature vector; Perform a multidimensional decision mapping on the feature vector to identify the corresponding operation intent category; Based on the category of operational intent and its continuous distribution characteristics over time, interest weight coefficients are generated; A feedback control strategy is constructed based on the interest weight coefficient and the operation intention category, and the execution unit is driven to output a response signal.

2. The feedback control method based on sensor fusion and operation intention recognition according to claim 1, characterized in that, The process of performing sliding window segmentation based on the multimodal operation data sequence and extracting multidimensional feature parameters reflecting operation intensity, trend of change, and periodic characteristics to construct a feature vector specifically includes: The multimodal operation data sequence is aligned to a unified time base, missing data is filled in, and anomaly is removed to form a standardized multimodal operation data sequence. Based on the standardized multimodal operation data sequence, window parameter configuration and sliding segmentation are performed, and the window boundary is adaptively adjusted when the action changes abruptly to generate the target operation segment; The operation intensity features, change trend features, and periodic characteristics are extracted from the target operation segment, and fusion encoding and correlation filtering are performed to construct the feature vector.

3. The feedback control method based on sensor fusion and operation intention recognition according to claim 1, characterized in that, The step of performing multidimensional decision mapping on the feature vector to identify the corresponding operation intention category specifically includes: Construct a set of operation intent categories, and establish a non-linear mapping relationship between the feature vectors and the operation intent categories based on an ensemble classification model; The feature vector is input into multiple basic decision units in the ensemble classification model to output candidate operation intent categories respectively; Perform integrated voting on multiple candidate operation intent categories to determine the target operation intent category; When there is a conflict between candidate operation intention categories, conflict resolution processing is performed based on the change trend features and periodic characteristics in the feature vector to correct the target operation intention category. Based on the support status of the basic decision-making unit, an intent confidence parameter is generated, and when the intent confidence parameter is lower than a preset threshold, a delayed joint discrimination process is performed. Perform temporal smoothing on the target operation intent category corresponding to the continuous target operation segments to output the operation intent category.

4. The feedback control method based on sensor fusion and operation intention recognition according to claim 3, characterized in that, Before performing multidimensional decision mapping on the feature vector to identify the corresponding operational intent category, the method further includes: Obtain multimodal operation sample data sequences covering multiple operation intent categories to construct an original sample set; The original sample set is processed by time alignment, missing data completion, and sliding window segmentation to form the target operation fragment set; Extract sample feature vectors from each target operation fragment in the target operation fragment set to form a sample feature vector set; An integrated classification model containing multiple basic decision units is constructed based on the set of sample feature vectors. Recursive splitting training is performed on each of the basic decision units to establish a nonlinear mapping relationship between the sample feature vector and the operation intention category; Based on the set of verification feature vectors, the ensemble classification model is subjected to recognition performance verification and parameter tuning to form the target ensemble classification model.

5. The feedback control method based on sensor fusion and operation intention recognition according to claim 1, characterized in that, The step of generating interest weight coefficients based on the operation intention category and its continuous distribution characteristics over time specifically includes: Construct an operation intent time series based on the operation intent category and time information corresponding to multiple consecutive target operation segments; Perform category aggregation processing on the time series of the operation intentions to form a set of intention activity segments; For each operational intent category, persistent distribution features are extracted based on the set of intent activity segments; The persistent distribution features are combined with the intent confidence parameter to perform confidence-weighted processing to form a persistent distribution feature expression vector; Calculate the initial interest score based on the persistent distribution feature expression vector; The initial interest score is subjected to time decay correction and sudden behavior suppression processing to generate a corrected interest score; The corrected interest score is subjected to a normalization mapping process to generate the interest weight coefficients.

6. The feedback control method based on sensor fusion and operation intention recognition according to claim 1, characterized in that, The step of generating interest weight coefficients based on the operation intention category and its continuous distribution characteristics over time specifically includes: A group operation time sequence is constructed based on the operation intent categories and time information corresponding to multiple target interaction carriers; The time series of group operations is subjected to cross-device category aggregation processing and group behavior segmentation processing to form a set of group behavior activity segments; Based on the set of group behavior activity segments, extract persistent distribution features of each group's operational intent category. The group's persistent distribution characteristics are combined with the intention confidence parameter, device activity parameter, and participation stability parameter to perform joint confidence weighting processing to form a group distribution characteristic expression vector; Calculate the initial group interest score based on the group distribution feature expression vector; Based on the initial group interest score, perform group time decay correction processing, instantaneous aggregation suppression processing, and conflict enhancement identification processing to generate a corrected group interest score; The interest weight coefficient is generated based on the corrected group interest score.

7. The feedback control method based on sensor fusion and operation intention recognition according to claim 1, characterized in that, The step of constructing a feedback control strategy based on the interest weight coefficient and the operation intention category, and driving the execution unit to output a response signal, specifically includes: A set of group feedback decision parameters is constructed based on the interest weight coefficient, the operation intention category, the intention confidence parameter, the number of participating devices, the device participation coverage ratio, the synchronization maintenance ratio, the alternation connection density, and the conflict occurrence density. Perform group state discrimination processing on the group feedback decision parameter set to generate group interaction state; A set of strategy mapping rules is established based on the group interaction state and the operation intention category, and the control target is determined. Calculate the strategy priority parameter based on the operation intention category, the group interaction state, and the control objective; Construct a collaborative feedback strategy set based on the operation intention category, the group interaction state, and the strategy priority parameter; The collaborative feedback strategy set is subjected to resource mapping processing to form a response signal generation scheme; The response signal generation scheme is subjected to timing orchestration and intensity adaptive processing to generate a target response signal sequence, thereby driving the execution unit to perform cross-device collaborative output according to the target response signal sequence to generate the response signal.

8. A feedback control device based on sensor fusion and operation intention recognition, characterized in that, The device is used to execute a feedback control method based on sensor fusion and operation intention recognition as described in any one of claims 1-7. The device includes an acquisition module, a processing module, and an output module, wherein: The acquisition module is used to acquire inertial measurement data, pressure sensing data and external visual attitude data set in the target interactive carrier, and perform preprocessing to form a multimodal operation data sequence. The processing module is used to perform sliding window segmentation processing based on the multimodal operation data sequence, and extract multidimensional feature parameters that reflect operation intensity, change trend and periodic characteristics to construct feature vectors; The processing module is used to perform multidimensional decision mapping on the feature vector to identify the corresponding operation intention category; The processing module is used to generate interest weight coefficients based on the category of the operation intention and its continuous distribution characteristics in the time dimension. The output module is used to construct a feedback control strategy based on the interest weight coefficient and the operation intention category, and drive the execution unit to output a response signal.

9. An electronic device, characterized in that, The device includes a processor, a communication bus, a user interface, a network interface, and a memory. The memory is used to store instructions. The user interface and the network interface are both used to communicate with other devices. The communication bus is used to enable communication between the various components within the electronic device. The processor is used to execute the instructions stored in the memory to cause the electronic device to perform the method described in any one of claims 1-7.

10. A non-transitory computer-readable storage medium, characterized in that, The computer-readable storage medium stores instructions that, when executed, perform the method as described in any one of claims 1-7.