A disabled employment auxiliary intelligent agent management system and method based on a multi-modal reflection architecture

CN122550339APending Publication Date: 2026-08-11BEIJING GUOHUA ZHONGLIAN TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-30
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0005]本申请实施例通过提供一种基于多模态反思架构的残疾人就业辅助智能体管理系统及方法,解决了现有技术中标准化岗位指令与残疾劳动者多元化认知/生理障碍不匹配、智能辅助系统单向输出缺乏实时监测与自我纠偏机制、边缘终端算力受限无法支撑多模态数据实时处理与高频反思三大核心技术痛点

Benefits of technology

1、通过获取包含行为影像、生理体征、音频交互及岗位标准化指令的多模态数据,并进行特征提取与基于最大互信息和动态时间规整的跨模态语义对齐,从而建立统一语义维度的多模态语义子图,进而实现了对残疾劳动者作业状态与情境的精准、全面感知。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122550339A_ABST
    Figure CN122550339A_ABST
Patent Text Reader

Abstract

This invention discloses a management system and method for assisting employment of people with disabilities based on a multimodal reflective architecture, relating to the field of computer technology. The system collects multi-source data from disabled workers, including behavioral images, physiological signs, audio interactions, and standardized job instructions, through a multimodal perception fusion module. It then performs cross-modal semantic alignment and fusion to construct a multimodal semantic subgraph based on task IDs. A three-level reflective architecture module leverages this subgraph to perform millisecond-level result reflection, second-level process reflection, and periodic strategy reflection, generating multi-level feedback data. An instruction adaptive evolution module drives prompt word optimization, automatically transforming original job instructions into personalized interactive prompts tailored to cognitive, hearing, or visual impairments. A heterogeneous hardware scheduling module allocates computationally intensive and real-time tasks to the central GPU or edge devices based on resource consumption and task type, ensuring end-to-end collaboration and low-latency response.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer technology, and in particular to a management system and method for assisting employment of persons with disabilities based on a multimodal reflective architecture. Background Technology

[0002] Employment is a crucial pathway for people with disabilities to ensure their basic livelihood and achieve social integration. my country has over 85 million people with disabilities, with a large number of those of working age and holding relevant certificates. However, current employment for people with disabilities still faces significant digital divides and skills mismatches. Traditional supported employment models heavily rely on human mentors. While this model can provide necessary humanistic care, its large-scale implementation has revealed pain points such as high mentoring costs, limited coverage, and difficulty in providing 24 / 7 real-time support.

[0003] At the technical level, existing standardized job operation instructions are mostly designed for able-bodied individuals, and their information density and expression often fail to consider the diverse cognitive and physiological impairments of disabled workers. For example, for workers with cognitive impairments, complex long sentences may lead to comprehension difficulties; for visually impaired workers, voice prompts lacking spatial orientation are difficult to guide them in delicate assembly operations. In recent years, multimodal artificial intelligence has made significant progress in areas such as speech recognition and limb detection, but most existing intelligent assistance systems are limited to a one-way output mode, lacking real-time monitoring and self-correction mechanisms for the worker's work status. Especially when faced with complex task chains, the system cannot accurately identify whether the worker is in a state of high cognitive load, leading to a disconnect between assistive information and actual needs, thereby causing operational errors or even safety risks.

[0004] Furthermore, the practical deployment of assistive devices for people with disabilities is limited by the computing power of the terminal. When real-time video streams, physiological sign streams, and high-frequency reflective logic need to be processed simultaneously, achieving millisecond-level response on resource-constrained edge devices while ensuring deep semantic fusion of multimodal data remains a technical challenge that the industry has not yet fully overcome. Therefore, there is an urgent need to develop an intelligent agent management system with perception, reflection, and self-evolutionary closed-loop capabilities to reduce the cost of assisting people with disabilities and improve the quality of employment for them, which has significant social and technological value. Summary of the Invention

[0005] This application provides a multimodal reflective architecture-based intelligent agent management system and method for assisting employment of people with disabilities. It addresses three core technical pain points in the prior art: the mismatch between standardized job instructions and the diverse cognitive / physiological impairments of disabled workers; the lack of real-time monitoring and self-correction mechanisms in the one-way output of intelligent assistance systems; and the inability of edge terminals to support real-time processing of multimodal data and high-frequency reflection due to limited computing power.

[0006] This application provides a multimodal reflection architecture-based intelligent agent management system for assisting employment of persons with disabilities, comprising: a multimodal perception and fusion module, used to acquire multimodal data including behavioral images, physiological signs, audio interactions, and standardized job instructions of disabled workers; to perform feature extraction and cross-modal semantic fusion on the multimodal data; and to establish a multimodal semantic subgraph based on task ID; a three-level reflection architecture module, used to perform three-level reflection processing based on the multimodal semantic subgraph; to generate feedback data at corresponding levels; the three-level reflection processing includes result reflection for millisecond-level response, process reflection for second-level response, and strategy reflection for periodic response; an instruction adaptive evolution module, used to automatically transform the original standardized job instructions into personalized interactive prompts that conform to specific disability types based on the feedback data generated by the three-level reflection processing and using a prompt word optimization algorithm; and a heterogeneous hardware scheduling module, used to perform heterogeneous hardware collaborative scheduling based on resource consumption data and task type.

[0007] This application provides a management method for intelligent agents assisting employment of persons with disabilities based on a multimodal reflection architecture. The method acquires multimodal data including behavioral images, physiological signs, audio interactions, and standardized job instructions of disabled workers. Feature extraction and cross-modal semantic fusion are performed on the multimodal data to establish a multimodal semantic subgraph based on task IDs. A three-level reflection process is executed based on the multimodal semantic subgraph to generate corresponding feedback data. This three-level reflection process includes result reflection for millisecond-level responses, process reflection for second-level responses, and strategy reflection for periodic responses. Based on the feedback data generated by the three-level reflection process, a prompt word optimization algorithm is used to automatically transform the original standardized job instructions into personalized interactive prompts that conform to specific disability types. Based on resource consumption data and task types, heterogeneous hardware collaborative scheduling is performed.

[0008] This application also provides an electronic device for assisting employment of persons with disabilities based on a multimodal reflective architecture, including one or more processors and one or more memories. The one or more memories store at least one computer program, which is loaded and executed by the one or more processors to implement an intelligent agent management system for assisting employment of persons with disabilities based on a multimodal reflective architecture.

[0009] One or more technical solutions provided in the embodiments of this application have at least the following technical effects or advantages: 1. By acquiring multimodal data including behavioral images, physiological signs, audio interactions, and standardized job instructions, and performing feature extraction and cross-modal semantic alignment based on maximum mutual information and dynamic time warping, a multimodal semantic subgraph with a unified semantic dimension is established, thereby achieving accurate and comprehensive perception of the work status and context of disabled workers.

[0010] 2. By reflecting on the results in milliseconds and comparing the physical state with the standard digital twin in real time to calculate the operational error, combined with the process reflection at the second level using path decision trees and heart rate variability indicators to diagnose operational bottlenecks and cognitive load, and the periodic strategy reflection based on Bayesian network models to analyze the distribution of cognitive biases, a three-level progressive feedback loop from immediate error correction and process optimization to long-term strategy adjustment is constructed, thereby realizing the dynamic self-adaptation of auxiliary strategies and the continuous enhancement of workers' own operational capabilities.

[0011] 3. By simplifying text for cognitive impairments, augmented reality visual mapping for hearing impairments, and spatial audio conversion for visual impairments, standardized job instructions can be adaptively transformed into personalized interactive prompts that conform to the cognitive characteristics of specific impairment types. This ensures that the information delivery method is precisely matched with the worker's remaining dominant senses, thereby effectively eliminating cognitive, auditory, and visual barriers and ensuring that workers accurately understand the work intentions. Attached Figure Description

[0012] Figure 1 A schematic diagram of the structure of the intelligent agent management system for assisting employment of persons with disabilities based on a multimodal reflective architecture provided in this application embodiment; Figure 2 A flowchart illustrating the management method for assistive intelligent agents in employment of persons with disabilities based on a multimodal reflective architecture, as provided in this application embodiment. Detailed Implementation

[0013] This application provides a multimodal reflective architecture-based intelligent agent management system and method for assisting employment of people with disabilities, addressing three core challenges in existing technologies: First, the lack of instruction adaptability. Traditional standardized instructions do not consider the cognitive and physiological characteristics of workers with different types of disabilities. This system captures individual behavior, physical signs, and interaction characteristics through multimodal perception fusion, and utilizes instruction adaptive evolution to transform standard instructions into simplified language, visual projections, or spatial audio, thus eliminating comprehension barriers. Second, the lack of real-time reflection and closed-loop correction. Existing systems are mostly unidirectional outputs, unable to perceive high cognitive loads and operational states. This system introduces millisecond-level results. The system employs a progressive self-correction and adaptation mechanism, encompassing reflection, second-level process reflection, and periodic strategy reflection. This mechanism ranges from immediate warnings of operational errors and simplified guidance for high-frequency, inefficient nodes to cognitive bias analysis and knowledge base reconstruction. Thirdly, it addresses the contradiction between edge computing power and multimodal deep processing. On resource-constrained terminal devices, it is difficult to simultaneously process images, vital sign streams, and high-frequency reflection logic. This system addresses this by using heterogeneous hardware collaborative scheduling to offload computationally intensive tasks to GPU-accelerated nodes or the edge cloud, while retaining real-time instructions and sensor acquisition locally. It can also dynamically bind visual inference to independent NPU units, thereby ensuring millisecond-level response while achieving deep cross-modal semantic fusion.

[0014] To better understand the above technical solutions, the following will provide a detailed explanation of the technical solutions in conjunction with the accompanying drawings and specific implementation methods.

[0015] like Figure 1 The diagram shows the structure of the intelligent agent management system for assisting employment of persons with disabilities based on a multimodal reflection architecture provided in this application embodiment. This system includes: a multimodal perception fusion module, a three-level reflection architecture module, an instruction adaptive evolution module, and a heterogeneous hardware scheduling module. The multimodal perception fusion module acquires multimodal data including behavioral images, physiological signs, audio interactions, and standardized job instructions of disabled workers. It performs feature extraction and cross-modal semantic fusion on the multimodal data to establish a multimodal semantic subgraph based on task IDs. The three-level reflection architecture module performs three-level reflection processing based on the multimodal semantic subgraph, generating corresponding levels of feedback data. The three-level reflection processing includes result reflection for millisecond-level responses, process reflection for second-level responses, and strategy reflection for periodic responses. The instruction adaptive evolution module automatically transforms the original standardized job instructions into personalized interactive prompts that conform to specific disability types using a prompt word optimization algorithm based on the feedback data. This algorithm adopts a template-based hierarchical instruction transformation strategy: first, the original... The initial instruction is parsed into a four-tuple of {operation object, operation action, operation condition, operation target}. Then, based on the obstacle type identifier in the feedback data, a predefined instruction conversion rule base is queried—for cognitive impairments, simplified rules are invoked to perform redundant information removal, long sentence splitting, and terminology replacement; for hearing impairments, visualization mapping rules are invoked to map the four-tuple into a graphical prompt template for an AR overlay; for visual impairments, spatial audio rules are invoked to convert the spatial coordinates of the operation object into HRTF parameters—finally, personalized interactive prompts are generated through template filling. A heterogeneous hardware scheduling module is used to perform heterogeneous hardware collaborative scheduling based on resource consumption data and task type, allocating tasks with different computing power requirements to corresponding hardware nodes to ensure that the end-to-end response latency meets real-time requirements.

[0016] In an application example at an industrial assembly station, the system uses a head-mounted camera and physiological sensors to perceive workers' behavioral images and heart rate variability in real time. When the process reflection layer detects a cognitive bottleneck in the worker's cognitive process, such as an abnormally high LF / HF ratio and frequent reference to the standard operating procedure (SOP) during the installation of the sealing ring, the strategy reflection layer dynamically evolves the standardized work process into a cyclically projected 3D disassembly animation based on the worker's inherent visual cognitive preference. Finally, the result reflection layer confirms that the installation accuracy meets the standards through visual comparison and provides positive light effect feedback, thus forming a complete closed loop of perception, reflection, and instruction evolution from pressure detection and cognitive assistance to quality verification.

[0017] In summary, this system integrates four core modules, constructing a collaborative processing chain that starts with multimodal perception fusion to generate semantic subgraphs, then produces layered feedback through a three-level reflective architecture, followed by personalized prompts through adaptive instruction evolution, and finally relies on heterogeneous hardware scheduling to ensure real-time response. This architecture can significantly increase the number of people a single human counselor can care for and greatly improve the accuracy of independent operations in complex positions, effectively overcoming the technical bottlenecks of traditional disability employment assistance systems, such as slow response and limited adaptability.

[0018] Furthermore, feature extraction and cross-modal semantic fusion are performed on the multimodal data to establish a multimodal semantic subgraph based on task ID. The multimodal semantic graph is a keyword semantic association graph constructed by extracting core keywords from the multimodal semantic subgraph according to the task objective. Specifically, this includes: collecting behavioral video streams of workers through an RGB-D camera, extracting eye movement trajectory features using an eye-tracking algorithm, and calculating eye movement trajectory feature indicators based on the eye movement trajectory features. The eye movement trajectory feature indicators include fixation duration and saccade path. Calculating fixation duration specifically includes: cleaning the time-stamped eye spatial coordinate sampling sequence of eye movement trajectory features, removing blink frames, missing sampling frames, and frames with abnormal pupil diameters; using linear interpolation to complete discontinuous data with a duration less than a preset missing threshold to obtain a calibrated continuous eye movement trajectory sequence; and using the velocity thresholding method (I-VT) to calculate eye movements between adjacent sampling points. Angular velocity is the ratio of the visual offset of adjacent sampling points to the sampling time interval. Continuous sampling points with angular velocities below a preset threshold are identified as candidate gaze segments, while those above the threshold are identified as candidate saccade segments. Adjacent candidate gaze segments with time intervals less than a preset minimum gaze interval are merged to generate independent gaze events. The difference between the start and end timestamps of a single gaze event is used as the single gaze duration for that event. Simultaneously, the mean of the coordinates of all sampling points within the gaze event is calculated as the centroid coordinates of the gaze point. The centroid coordinates of the gaze point for each gaze event are mapped to a preset set of regions of interest (ROIs) on the work surface. These ROIs correspond to the material area, tool area, operation assembly area, and instruction display area. The cumulative gaze duration within a preset statistical window is grouped by ROI, and the gaze duration for each area is output. The calculation of the saccade path specifically includes: extracting all candidate saccade segments from the calibrated eye movement trajectory sequence, recording the start coordinates, end coordinates, start time, and duration of each saccade segment; sequentially concatenating the centroid coordinates of the fixation points of all fixation events in chronological order to form an ordered coordinate sequence of alternating "fixation point, saccade segment, fixation point," which is the basic saccade path; the line connecting two adjacent fixation points corresponds to the motion trajectory of this saccade, and the spatial distance between the lines is the saccade amplitude; matching the centroid coordinates of each fixation point in the basic saccade path with preset regions of interest and work process nodes one by one, converting the coordinate sequence into a node sequence with work process semantic labels, generating a semantic saccade path corresponding to the work process, used for subsequent analysis of operational logic coherence and cognitive pauses; collecting the physiological signal flow of workers through photoplethysmography (PPG) using wearable sensors, extracting heart rate variability indicators and skin conductance indicators; and using the maximum mutual information criterion to perform cross-modal alignment of each modal feature.Because the movement frequencies of disabled workers vary significantly (e.g., limb disabilities may lead to slow movements), dynamic time warping technology is used to align physiological sign streams and behavioral image streams. The aligned data is then converted into semantic vectors of a unified dimension and entered into a multi-source knowledge base retrieval process. The multi-source knowledge base includes three sub-bases: a standard operation knowledge base, storing standard operation steps and their semantic vectors corresponding to each task ID; an auxiliary strategy knowledge base, storing auxiliary strategies and their semantic vectors corresponding to each cognitive dimension; and a historical case knowledge base, storing historically successfully handled anomalies and their semantic vectors. During retrieval, the cosine similarity of the currently aligned unified semantic vector is calculated with the semantic vectors stored in the three sub-bases. The top three records with a similarity ≥ 0.75 are taken as retrieval results. The retrieval threshold is set to 0.75 to ensure that the retrieved operation guidelines are highly relevant to the current work status. The aligned unified semantic vector is then associated with the above retrieval results to jointly construct a multimodal semantic subgraph based on the task ID.

[0019] In this embodiment, the multimodal semantic subgraph is a cross-modal semantic fusion basic representation built based on task IDs. Internally, it stores a unified-dimensional semantic vector transformed by maximum mutual information and dynamic time warping alignment, as well as complete feature information such as relevant standard operations, auxiliary strategies, and historical case semantic nodes retrieved from multi-source knowledge bases. The multimodal semantic graph, on the other hand, is a simplified semantic structure dynamically drawn by extracting the core keywords of the target task from this subgraph during result reflection. It highlights the semantic items of key operations that have been achieved and those that have not, and is used to calculate keyword coverage and semantic matching degree in real time, supporting millisecond-level operation error determination. This solution achieves deep correlation modeling of multi-dimensional data, effectively improving the accuracy of speech recognition for people with speech impairments and effectively eliminating the feature misalignment problem caused by differences in the frequency of actions of disabled workers.

[0020] Furthermore, the millisecond-level response result reflection focuses on the immediate verification of outputs, specifically including: extracting core keywords of the target task from the multimodal semantic subgraph, drawing a multimodal semantic graph, which is a keyword semantic association graph constructed from the core keywords extracted from the multimodal semantic subgraph according to the task target, specifically including: locating the corresponding job standard knowledge node subset in the multimodal semantic subgraph according to the current task ID, taking the process target node in this subset as the root node, extracting upstream and downstream process nodes, operation constraint nodes, and quality acceptance nodes that have process dependencies with the root node, and generating an initial candidate keyword set; using a pre-trained cross-modal semantic mapping model, respectively integrating limb movement features, hand operation texture features, and eye movement tracks. The semantic vectors corresponding to trace features are mapped to candidate keywords for action description, the semantic vectors corresponding to audio features are mapped to candidate keywords for interaction commands, and the semantic vectors corresponding to physiological signs are mapped to candidate keywords for work status. These candidate keywords are then added to the initial candidate keyword set. A degree-centrality weighted mechanism is used to calculate the overall importance of each candidate keyword: the number of cross-modal connection edges between the node corresponding to the candidate keyword in the multimodal semantic subgraph is counted as the degree-centrality value. This value is then weighted and summed with the industry terminology weight coefficient and the task objective relevance coefficient to obtain the overall importance score for each candidate keyword. A predetermined number of keywords are selected as the core keywords for the target task based on their scores, while other core keywords are retained. The semantic dependencies and cross-modal mutual information weights between keywords are analyzed. Using the selected core keywords as graph nodes, each node is labeled with its modal attribute and overall importance weight. A directed keyword semantic association graph, i.e., a multimodal semantic graph, is constructed with the process target keywords as the central nodes, using the process dependencies and semantic associations between core keywords as directed edges and the corresponding cross-modal mutual information weights as edge weights. Based on the multimodal semantic graph, the physical state of the worker after their current operation is compared in real time with the standard digital twin to generate keyword coverage and semantic matching degree. The operation error value is calculated based on the keyword coverage and semantic matching degree. Keyword coverage is the ratio of achieved keywords to the total number of keywords in the multimodal semantic graph. Achieved keywords are those obtained through visual... Semantic analysis identifies completed operations from physical states that semantically match keywords in the multimodal semantic graph. The semantic matching degree is the cosine similarity between the feature vector of the physical state operation and the standard feature vector of the digital twin. The operation error value is calculated as follows: Error value = 0.4 × (1 - Keyword coverage) + 0.6 × (1 - Semantic matching degree). The weight coefficients 0.4 and 0.6 are determined as follows: 100 sets of labeled samples are collected, each set containing physical state data with known qualified / unqualified labels. The keyword coverage and semantic matching degree are extracted separately. A logistic regression model is used to fit the weight coefficients with the defect judgment accuracy as the target, resulting in a keyword coverage weight of 0.4 and a semantic matching degree weight of 0.The accuracy rate is highest at time 6, therefore this weighting is adopted. When the operation error value exceeds the preset operation error threshold, a millisecond-level error warning is triggered, and language control parameters are extracted to adjust the structural output template of subsequent instructions. Language control parameters include speech rate coefficient, simplification level, and repetition count. The structural output template is the syntactic framework of the instruction. Defective products caused by improper operation are prevented from flowing to the next process. The preset operation error threshold is the maximum permissible error value pre-set based on product tolerance requirements and acceptance standards.

[0021] In this embodiment, the standard digital twin is constructed based on a 3D model of the product assembly (extracted from an enterprise management system PDM / PLM). Specifically, this involves decomposing the 3D model into intermediate state models corresponding to each assembly step. Each intermediate state model contains the geometric constraints and appearance feature parameters that should be achieved upon completion of that step. The digital twin is constructed offline during the system initialization phase and is synchronously updated through the enterprise management system interface when the product process changes. The result reflection layer extracts the core keywords of the target task from the multimodal semantic subgraph and draws a multimodal semantic graph. It then performs real-time registration and comparison between the physical state after the worker's operation and the standard digital twin. When the error value exceeds a preset threshold, a millisecond-level error warning is immediately triggered, and language control parameters are extracted to dynamically adjust the structural output template of subsequent instructions. Through the real-time error detection and warning of the result reflection layer, the assembly quality can be determined within 100ms after the operation is completed. This allows the defect interception node to be moved from post-inspection to the work-in-process stage, thereby reducing the risk of batch rework.

[0022] Furthermore, the second-level response process reflection focuses on the diagnosis of work paths and cognitive load, specifically including: obtaining the worker's operation step frequency and logical pauses based on multimodal semantic subgraphs, and analyzing them using a path decision tree. The path decision tree uses work steps as nodes and the transition relationships between steps as edges. Each node stores the historical average processing time of that step, and analyzes the worker's operation step frequency and logical pauses; for each work step node in the path decision tree, the actual processing time corresponding to that node within a preset statistical period is statistically analyzed, and the current average processing time is calculated; if the current average processing time of any work step node exceeds a preset multiple of the historical average processing time corresponding to that node, then that node is marked as a high-frequency and inefficient node; when the frequency of change of the heart rate variability index of any high-frequency and inefficient node exceeds a preset variability change frequency threshold, the simplified guidance strategy of the cache is invoked to skip redundant instructions, directly provide core operation guidance, and configure an LRU caching strategy for that node. When the same bottleneck is encountered later, the optimized fast guidance scheme is directly read. The simplified guidance strategy is a pre-set simplified instruction sequence for high-frequency, low-efficiency nodes. Each strategy contains three fields: (1) core operation description (a verb-object phrase of no more than 10 characters, such as "align buckle"); (2) key visual markers (coordinates of the target area highlighted on the AR interface); and (3) haptic feedback mode (a short prompt tone encoding of the bone conduction headphones). This strategy is pre-configured by the job instructor based on the complexity of each work step during system initialization, and the instruction length is automatically fine-tuned based on the worker's response time after each trigger. The historical benchmark value is the historical average processing time of the worker performing this step under stable working conditions, and the preset multiple of the historical benchmark value is a pre-set multiple threshold used to determine abnormally long processing times. The frequency of change of the heart rate variability index refers to the frequency of fluctuation of the LF / HF ratio per unit time. The preset variability frequency threshold is a critical value used to identify abnormal fluctuations in cognitive load, which is set based on the baseline frequency of change of the worker under low cognitive load.

[0023] In this embodiment, the path decision tree is constructed as follows: the main path is the sequence of work steps defined in the job standardization SOP, arranged sequentially from the root node (first step) to the leaf node (last step); each step node stores the standard operation description, historical average processing time (initial value is standard working time), qualification judgment conditions, and simplified guidance strategy pointer for that step. For operations with branch selection (such as "if there are burrs, perform grinding; otherwise, perform cleaning"), a branch child node is created at the corresponding node. The historical baseline value is automatically updated to the moving average of the most recent 20 times after each work shift. The process reflection layer obtains the worker's operation step frequency and logical pause data based on the multimodal semantic subgraph, constructs a path decision tree for each node corresponding to a work step for analysis, and marks nodes whose average processing time exceeds a preset multiple of the historical baseline value as high-frequency and inefficient nodes; when a high-frequency and inefficient node is detected to be accompanied by abnormal heart rate variability indicators indicating high cognitive load, it is determined to be a cognitive bottleneck, and the cached simplified guidance strategy is immediately invoked to skip redundant instructions, directly provide core operation guidance, and configure a 1GB LRU cache strategy for the node to store the optimized fast guidance scheme. This layer can accurately locate cognitive difficulties in the operation within seconds, shorten the auxiliary response time of a single step, effectively reduce the cognitive load of workers, and avoid operational errors and emotional anxiety caused by difficulty in understanding.

[0024] Furthermore, the periodic response strategy reflection specifically includes: obtaining historical decision records within the most recent preset period based on the multimodal semantic subgraph, and analyzing the cognitive bias distribution of workers using a Bayesian network model; calculating the bias probability of each cognitive dimension based on the cognitive bias distribution, where the bias probability is the ratio of the number of erroneous operations to the total number of operations under that dimension; if the bias probability of any cognitive dimension exceeds the preset bias probability threshold, then generating a knowledge structure update requirement and reconstructing the accessibility task knowledge base corresponding to the multimodal semantic subgraph: (1) locking the cognitive dimensions where the cognitive bias probability exceeds the preset threshold; (2) retrieving the knowledge base entries corresponding to that dimension and extracting the common patterns that lead to erroneous operations; (3) correcting the rules of the original knowledge base entries (such as adding preconditions, reducing operational complexity, and changing the auxiliary strategy type); (4) replacing the original entries with the corrected entries, incrementing the version number, and verifying the correction effect in the next strategy reflection cycle. The accessibility task knowledge base is a structured knowledge base indexed by template category identifiers, storing operational norms and assistance strategies for each cognitive dimension. A pre-trained Bayesian network model is used to analyze the distribution of workers' cognitive biases. The Bayesian network model uses cognitive dimensions as nodes and the influence relationships between cognitive dimensions as directed edges; the conditional probability table is trained from historical labeled data. Based on the reconstructed knowledge base, the tutoring intensity is dynamically adjusted, and personalized job suitability suggestions are generated. The preset period is a time window for reflecting on and evaluating the execution strategy set by the system; the preset deviation probability threshold is an upper limit for the cognitive dimension error rate set according to process tolerance requirements.

[0025] In this embodiment, the strategy reflection layer obtains historical decision records within the most recent preset period based on the multimodal semantic subgraph, analyzes the cognitive bias distribution of workers using a Bayesian network model, and calculates the bias probability of each cognitive dimension (the ratio of the number of incorrect operations in that dimension to the total number of operations). When the bias probability of any cognitive dimension exceeds a preset threshold, a knowledge structure update requirement is automatically generated, and the accessibility task knowledge base corresponding to the multimodal semantic subgraph is reconstructed. Based on the reconstructed knowledge base, the tutoring intensity is dynamically adjusted, and personalized job matching suggestions are generated. The Bayesian network model defines the following cognitive dimension nodes: spatial perception ability, temporal memory ability, instruction comprehension ability, fine motor skills ability, and attention span ability, totaling 5 nodes. The causal relationship between nodes is defined as: attention span ability - instruction comprehension ability - temporal memory ability - fine motor skills ability, spatial perception ability - fine motor skills ability. The conditional probability table is initialized with the historical work data of 50 disabled workers through maximum likelihood estimation, and is updated using new historical decision records after each preset period (e.g., 40 working hours). This mechanism enables long-term dynamic adaptation of individual capabilities, automatically compensates for workers' perceptual biases (such as sound field positioning bias for visually impaired individuals), provides enterprises with a scientific basis for job allocation, and provides data support for job adaptation decisions.

[0026] Furthermore, the original standardized job instructions are automatically transformed into personalized interactive prompts that conform to specific obstacle types. Specifically, this includes: acquiring feedback data generated by the three-level reflection architecture module; this feedback data includes millisecond-level result reflection feedback operation error values ​​and language control parameters, second-level process reflection feedback high-frequency and inefficient node markers, and periodic strategy reflection feedback cognitive bias distribution; and determining the complexity adjustment requirements for the current instruction based on the feedback data. The operation error value is compared with the preset error tolerance threshold. If the operation error value exceeds the error tolerance threshold, the difference between the operation error value and the error tolerance threshold is calculated as the error excess. The error excess is input to the preset correction mapping function. The correction mapping function is a piecewise linear mapping and outputs a correction coefficient in the range of [0, 1]. The correction coefficient is multiplied by the preset maximum simplification level offset and maximum repetition number offset to obtain the simplification level increment and repetition number increment, which are used as the first complexity correction amount.

[0027] The system acquires high-frequency and inefficient node markers from the second-level process reflection feedback, extracts the identifier of the marked operation node and its marking frequency or proportion within the current task cycle; compares the marking frequency of the marked node with a preset inefficient frequency threshold; when the marking frequency is greater than the inefficient frequency threshold, calculates the difference between the marking frequency and the inefficient frequency threshold, and the ratio of the difference to the inefficient frequency threshold is the frequency excess ratio. The frequency excess ratio is input into a preset discrete mapping rule, and the discrete mapping rule outputs the corresponding instruction splitting strategy level based on the value range of the frequency excess ratio β: when β∈(0,1], it outputs "single node splitting"; when β∈(1,2], it outputs "single node splitting and sub-step simplification"; when β>2, it outputs "single node splitting, sub-step simplification, and redundancy review prompt"; the instruction splitting strategy level is converted into the corresponding second complexity correction quantity, which includes the target node identifier and the splitting and simplification strategy of that node.

[0028] Extract the cognitive bias distribution from the periodic strategy reflection feedback. The cognitive bias distribution includes the bias intensity value on at least one preset cognitive bias dimension. For each preset cognitive bias dimension, compare its bias intensity value with the corresponding bias threshold. When the bias intensity value of at least one preset cognitive bias dimension is greater than its corresponding bias threshold, extract the bias intensity values ​​of all dimensions exceeding the threshold, and calculate the multimodal redundancy demand index. ;in The preset weights for the i-th cognitive bias dimension are: Let i be the deviation intensity value of the i-th dimension. Let be the deviation threshold for the i-th dimension; match the multimodal redundancy requirement index with the preset redundancy level mapping rule to obtain the third complexity correction amount. The third complexity correction amount includes a Boolean flag indicating whether multimodal redundancy prompts need to be enabled, the modal combination type of the redundancy prompts, and the enhancement level of the key information identifiers. The preset redundancy level mapping rule is as follows: When γ∈[0, T1), the output Boolean flag is "No", multimodal redundancy prompts are not enabled, and the enhancement level is 0; When γ∈[T1,T2), the output Boolean flag is "Yes", the modality combination type is "Visual Identity Enhancement", and the enhancement level is 1; When γ∈[T2, T3), the output Boolean flag is "Yes", the modal combination type is "Visual Identification Reinforcement + Vibration Feedback", and the reinforcement level is 2; When γ≥T3, the output Boolean flag is "Yes", the modal combination type is "Visual Identification Enhancement + Vibration Feedback + Spatial Audio Verification", and the enhancement level is 3; T1, T2, and T3 are preset redundancy level division thresholds, and satisfy 0 <T1<T2<T3。

[0029] Based on the simplification level and repetition count in the language control parameters, the complexity adjustment requirements for the current instruction are obtained by comprehensively superimposing the first, second, and third complexity corrections. Specifically, the simplification level increment in the first correction is added to the baseline simplification level, and splitting and simplification processing is applied to the specified node according to the second correction; the repetition count increment in the first correction is added to the baseline repetition count, and the number of repetitions of redundant prompts is additionally increased according to the enhancement level in the third correction; finally, the complexity adjustment requirements are obtained, including the overall instruction simplification level adjustment target, the number of instruction repetitions, and the enhancement prompt strategy for specific operation steps. The system calls upon the corresponding prompt word optimization template based on the worker's disability type; it then adjusts the requirements and prompt word optimization template according to complexity to transform the original standardized job instructions. Specifically: for cognitive or intellectual disabilities, it parses the original standardized job instruction text for the current task and constructs a syntactic dependency tree; based on the simplification level in the language control parameters, it performs pruning operations on the syntactic dependency tree. When the simplification level is level one, only adjectives and adverbial modifiers are removed; when the simplification level is level two, prepositional phrases and clauses are further deleted; the pruned subject-verb-object core structure is retained to generate simplified instruction text; based on the repetition frequency in the language control parameters, the simplified instruction text is converted into voice broadcast instructions, and a repetition cycle is set. The loop playback count is the number of repetitions. For hearing impairments, when high-frequency, inefficient node markers indicate a delay in auditory command response, the voice command is mapped to an augmented reality projection, overlaid on the actual work surface to display operation instructions, and the visual identifier of the current step is highlighted based on the operation error value. For visual impairments, the azimuth weight of spatial audio is adjusted according to the operation error value and cognitive bias distribution. The spatial positioning error pattern of the worker is extracted from the cognitive bias distribution, and the deviation confidence level of each azimuth interval is calculated. The operation error value and the deviation confidence level of each azimuth interval are weighted and superimposed to generate an azimuth weight adjustment vector. The azimuth weight adjustment vector is superimposed on the azimuth increment of the HRTF filter parameters. The gain coefficient enhances the spatial cues of sound in high-error azimuth directions and weakens the audio output in low-correlation azimuth directions. It acquires the three-dimensional coordinates of the material box in physical space and calculates the HRTF filter parameters relative to the worker's ear in real time. The HRTF filter parameters include filter coefficients corresponding to the azimuth, elevation, and distance of the material box, specifically the left and right ear frequency domain amplitude spectrum, phase spectrum, or time domain impulse response coefficients of the head-related impulse response pair. The real-time calculation of the HRTF filter parameters relative to the worker's ear specifically includes: acquiring the worker's real-time head orientation data, converting the three-dimensional coordinates of the material box in the world coordinate system to spherical coordinates with the center of the worker's head as the origin, and obtaining the azimuth angle θ and elevation angle θ. and distance r; based on the preset personalized HRTF database, with θ, Using r as an index, the corresponding left and right ear filter coefficients are retrieved, and short voice commands are played through bone conduction headphones.

[0030] Material boxes are standardized material containers deployed within the work area. Their three-dimensional coordinates in physical space are coordinates in a world coordinate system referenced to a pre-set origin in the workshop. These three-dimensional coordinates can be calculated in real-time using a visual positioning system to capture reference markers on the material box and determine its pose in the world coordinate system; or by using ultra-wideband positioning tags embedded in the material box and fixed base stations for real-time ranging to calculate the three-dimensional spatial coordinates. A personalized HRTF database refers to a pre-constructed dataset of head-related transfer functions containing full-space orientation and distance information for each individual worker. It records the subject's bilateral responses to different azimuth (θ) and elevation (θ) angles using specialized measurement equipment (such as probe microphones picking up sound within the ear canal, or near-field scanning calculations). The impulse response of the sound source at distance (r) is extracted into the amplitude and phase spectra of the left and right ears via Fourier transform, or stored as impulse response coefficients in the time domain. This database uses the center of the worker's head as the origin, discretizing the entire space into several (θ, ...) components. The index grid, r), with each grid point corresponding to a set of HRTF filter coefficients, can accurately reflect the diffraction and scattering characteristics of sound waves by an individual's head, auricle, and torso, eliminating spatial positioning ambiguity, front-back confusion, and vertical misalignment caused by morphological differences in the general HRTF. In this scheme, the role of the personalized HRTF database is: when the system obtains the world coordinates of the material box and converts them into spherical coordinates (θ, r) with the worker's head as the origin. After r), the exclusive filter coefficients are directly retrieved based on these coordinate indices, and then superimposed with the orientation weight adjustment vector generated by the operation error value and cognitive bias distribution. Finally, high-precision spatial audio commands are rendered and output in real time through bone conduction headphones, enabling visually impaired workers to accurately perceive the true three-dimensional orientation of the material box by sound cues, and achieve safe and undisturbed work guidance.

[0031] In this embodiment, the solution can adaptively adjust the simplification level of instructions, the number of repetitions, the breakdown of operation steps, and the redundancy prompt strategy based on the worker's real-time operational errors, process efficiency bottlenecks, and periodic cognitive biases. It also employs differentiated interaction channels for different obstacle types: for those with cognitive impairments, redundant information is removed and core actions are reinforced; for those with hearing impairments, auditory instructions are transformed into augmented reality visual guidance superimposed on the actual work surface, with high-risk steps highlighted; for those with visual impairments, individualized HRTF spatial audio and orientation weighting based on operational errors and positioning deviations are used to accurately render the three-dimensional spatial orientation of the material box through bone conduction headphones, enabling workers to accurately perceive the target location solely by sound. This achieves safe, efficient, and interference-free work guidance, significantly improving the relevance and reliability of auxiliary instructions.

[0032] Furthermore, the system also includes real-time adaptive intervention and control based on the extracted heart rate variability (HRV) indicators, skin conductance (SCA) indicators, and eye movement trajectory feature indicators. Specifically, this includes: extracting the LF / HF ratio of the HRV indicators; when the LF / HF ratio exceeds a corresponding preset threshold, it indicates high cognitive load, triggering the process reflection layer to invoke a simplified guidance strategy, reducing the command issuance frequency according to a preset frequency adjustment ratio, and activating a simplified speech mode; LF (Low Frequency) and HF (High Frequency) are the most important components in HRV frequency domain analysis. The two core power spectrum bands, with their power ratio being the LF / HF ratio, are a standard indicator for assessing autonomic nervous system balance. The low-frequency component (LF) typically ranges from 0.04 to 0.15 Hz and is jointly regulated by the sympathetic and parasympathetic nervous systems, with the sympathetic nervous system playing a dominant role. It reflects the body's stress arousal level and vascular tone regulation. The high-frequency component (HF) typically ranges from 0.15 to 0.4 Hz and is highly synchronized with the body's respiratory rhythm. It is primarily regulated by the parasympathetic nervous system (vagus nerve) and directly reflects the parasympathetic nervous system's tone level and the degree of physical and mental relaxation. An elevated LF / HF ratio indicates that sympathetic nerve excitation is dominant, corresponding to a state of stress, anxiety, excessive cognitive load, or accumulated fatigue. A decreased LF / HF ratio indicates that the parasympathetic nervous system is dominant, corresponding to a relaxed, stable state with sufficient cognitive resources. The phase response parameters (Phasic / Tonic response parameters) of the electrodermal activity index are extracted. When the frequency of change of the Phasic / Tonic response parameters exceeds a preset frequency threshold, emotional fluctuations are identified, encouraging voice prompts are played, and a short rest is suggested. The fixation duration and saccade paths of eye movement characteristics are extracted. When the fixation duration of any saccade path exceeds a preset fixation duration threshold, it is determined that the worker's information acquisition is obstructed, and the automatic highlighting of the visual guidance spot is triggered. The preset ratio threshold is a critical value for sympathetic nerve activity based on the worker's LF / HF ratio baseline at rest; the preset frequency adjustment ratio refers to the percentage reduction in the frequency of pre-set commands. In the Phasic / Tonic response parameters of the electrodermal activity index, Phasic represents the phase component, reflecting the short-term electrodermal response induced by a specific stimulus, and Tonic represents the tonic component, reflecting the baseline level of skin conductance. The preset frequency threshold is a critical value used to determine emotional fluctuations, set according to the fluctuation frequency of the Phasic / Tonic parameters in the worker's calm state. The preset gaze duration threshold is a time limit set based on the statistical data of the gaze duration of a worker in a single area when acquiring normal information. Exceeding this threshold indicates that the acquisition of visual information is obstructed.

[0033] In this embodiment, the physiological-behavioral linkage intervention mechanism of the present invention can identify workers' cognitive overload and emotional fluctuations in advance, proactively intervene to alleviate anxiety, reduce the error rate caused by fatigue or stress during the work process, and improve work safety and workers' work experience.

[0034] Furthermore, the steps for implementing heterogeneous hardware collaborative scheduling include: identifying task types based on resource consumption data, including compute-intensive tasks and I / O-intensive tasks. Compute-intensive tasks include real-time detection of multiple image streams, while I / O-intensive tasks include historical data strategy reflection retrieval. Using a priority queue scheduler, compute-intensive tasks with computational demands exceeding a preset threshold are allocated to a central node equipped with GPU accelerators, while real-time command broadcasting and sensor acquisition tasks are retained at the edge. The priority queue scheduler employs a three-level priority preemptive scheduling strategy: Priority level 0 (highest) is for real-time command broadcasting and sensor acquisition tasks, with CPU affinity bound to a dedicated edge core and cannot be preempted; Priority level 1 is for visual recognition and result reflection tasks, executed by default on the edge NPU, automatically unloaded to the central GPU node when the waiting queue length exceeds 3; Priority level 2 (lowest) is for historical data strategy reflection and knowledge base reconstruction tasks, executed by default in batch processing on the central GPU node. The scheduler polls each task queue every 100ms, allocating computing resources in descending priority order. The formula for calculating the computing power requirement of IO-intensive tasks is: Computing power requirement = CPU utilization × Task priority coefficient. The task priority coefficient is divided into three levels according to the task's real-time requirements: 1.0 for hard real-time tasks, 0.6 for soft real-time tasks, and 0.3 for batch processing tasks. When the thread waiting time of a computing-intensive task exceeds the preset waiting time threshold, thread binding adjustment is automatically performed, binding the visual recognition thread to an independent NPU unit at the edge, while offloading the reflection logic to the campus edge cloud server. The preset requirement threshold is a boundary value for distinguishing computing power requirements between computing-intensive tasks and ordinary tasks, set according to the system's computing power resource allocation strategy; the preset waiting time threshold is an upper limit for thread waiting time preset according to real-time requirements.

[0035] In this embodiment, the heterogeneous hardware scheduling module first categorizes tasks into two types based on resource consumption data: computationally intensive (e.g., real-time detection of multi-channel image streams) and I / O intensive (e.g., historical data strategy reflection retrieval). A priority queue scheduler allocates computationally intensive tasks with computational demands exceeding a preset threshold to a central node equipped with a GPU accelerator, while real-time command broadcasting and sensor acquisition tasks remain at the edge. When the thread waiting time for a computationally intensive task exceeds a preset threshold, thread binding adjustment is automatically performed, binding the visual recognition thread to an independent NPU unit at the edge and offloading the reflection logic to the campus edge cloud server. This scheduling scheme ensures that the system's end-to-end response latency is stably controlled within extremely low latency, significantly extending the battery life of front-end wearable devices, improving the utilization of expensive computing resources, and meeting the requirements of industrial-grade real-time deployment.

[0036] like Figure 2 The flowchart shown is a flowchart of the intelligent agent management method for assisting employment of persons with disabilities based on a multimodal reflection architecture provided in this application embodiment. It involves acquiring multimodal data including behavioral images, physiological signs, audio interactions, and standardized job instructions of disabled workers; performing feature extraction and cross-modal semantic fusion on the multimodal data to establish a multimodal semantic subgraph based on task IDs; performing three-level reflection processing based on the multimodal semantic subgraph to generate corresponding levels of feedback data. The three-level reflection processing includes result reflection for millisecond-level responses, process reflection for second-level responses, and strategy reflection for periodic responses; using the feedback data generated by the three-level reflection processing, an algorithm for optimizing prompts is used to automatically transform the original standardized job instructions into personalized interactive prompts that conform to specific disability types; and performing heterogeneous hardware collaborative scheduling based on resource consumption data and task types.

[0037] Furthermore, an electronic device for assisting employment of persons with disabilities based on a multimodal reflective architecture includes one or more processors and one or more memories, wherein at least one computer program is stored in the one or more memories, and the at least one computer program is loaded and executed by the one or more processors to implement an intelligent management system for assisting employment of persons with disabilities based on a multimodal reflective architecture.

[0038] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.

Claims

1. A disability employment assistance intelligent agent management system based on a multimodal reflective architecture, characterized in that, include: The multimodal perception and fusion module is used to acquire multimodal data including behavioral images, physiological signs, audio interactions and standardized job instructions of disabled workers, perform feature extraction and cross-modal semantic fusion on the multimodal data, and establish a multimodal semantic subgraph based on task ID; The three-level reflection architecture module is used to perform three-level reflection processing based on multimodal semantic subgraphs to generate feedback data at the corresponding level. The three-level reflection processing includes result reflection for millisecond-level response, process reflection for second-level response, and strategy reflection for periodic response. The instruction adaptive evolution module is used to automatically transform the original standardized job instructions into personalized interactive prompts that conform to specific obstacle types based on the feedback data generated by the three-level reflection processing and the prompt word optimization algorithm. The heterogeneous hardware scheduling module is used to perform collaborative scheduling of heterogeneous hardware based on resource consumption data and task type.

2. The disabled employment assistance intelligent agent management system based on a multi-modal reflection architecture of claim 1, wherein, The step of performing feature extraction and cross-modal semantic fusion on multimodal data to establish a multimodal semantic subgraph based on task ID specifically includes: The process involves collecting video streams of workers' behavior, extracting eye movement trajectory features using an eye-tracking algorithm, and calculating eye movement trajectory feature indices based on these features. These indices include fixation duration and saccade path. Collect physiological data from workers and extract heart rate variability and skin electrophysiological activity indicators; The maximum mutual information criterion is used to perform cross-modal alignment of features of each modality. Combined with dynamic time warping technology, physiological sign streams and behavioral image streams are aligned. The aligned data is converted into semantic vectors of a unified dimension and a multimodal semantic subgraph based on task ID is established.

3. The disabled employment assistance intelligent agent management system based on a multi-modal reflection architecture of claim 1, wherein, The reflection on the results of the millisecond-level response specifically includes: Extract the core keywords of the target task from the multimodal semantic subgraph and draw the multimodal semantic graph; Based on the multimodal semantic graph, the physical state of the worker after the current operation is compared with the standard digital twin in real time to generate keyword coverage and semantic matching degree. The operation error value is calculated based on the keyword coverage and semantic matching degree. The keyword coverage is the ratio of the achieved keywords to the total number of keywords in the multimodal semantic graph. The semantic matching degree is the cosine similarity between the feature vector of the operation physical state and the standard feature vector of the digital twin. When the operation error value exceeds the preset operation error value threshold, a millisecond-level error warning is triggered, and language control parameters are extracted to adjust the structure output template of subsequent instructions. The language control parameters include speech rate coefficient, simplification level and repetition number, and the structure output template is the syntactic framework of the instruction.

4. The disability employment assistance intelligent agent management system based on a multimodal reflective architecture as described in claim 1, characterized in that, The process of achieving a second-level response specifically includes: The operation step frequency and logical pauses of workers are obtained based on multimodal semantic subgraphs, and analyzed using path decision trees. The path decision trees are based on operation steps as nodes and the transition relationship between steps as edges. Each node stores the historical average processing time of the step. For each task step node in the path decision tree, the actual processing time of the node within a preset statistical period is calculated and the current average processing time is calculated. If the current average processing time of any task step node exceeds a preset multiple of the historical average processing time of the node, the node is marked as a high-frequency and low-efficiency node. When the frequency of heart rate variability indices of any high-frequency, inefficient node exceeds the preset variability frequency threshold, the simplified boot strategy is invoked to skip redundant instructions, directly provide core operation guidance, and configure an LRU caching strategy for that node so that when the same bottleneck is encountered in the future, the optimized fast boot solution can be directly read.

5. The disabled employment assistance intelligent agent management system based on a multi-modal reflection architecture of claim 1, wherein, The strategic reflection on the periodic response specifically includes: Based on multimodal semantic subgraphs, historical decision records within the most recent preset period are obtained, and Bayesian network models are used to analyze the distribution of workers' cognitive biases. The deviation probability of each cognitive dimension is calculated based on the cognitive deviation distribution. The deviation probability is the ratio of the number of erroneous operations to the total number of operations in that dimension. If the deviation probability of any cognitive dimension exceeds the preset deviation probability threshold, a knowledge structure update requirement is generated to reconstruct the accessibility task knowledge base corresponding to the multimodal semantic subgraph. The tutoring intensity is dynamically adjusted based on the reconstructed accessibility task knowledge base, and personalized job matching suggestions are generated.

6. The disabled employment assistance intelligent agent management system based on a multi-modal reflection architecture of claim 1, wherein, The process of automatically converting standardized job instructions into personalized interactive prompts that conform to specific obstacle types specifically includes: Obtain the feedback data generated by the three-level reflection architecture module. The feedback data includes the operation error value and language control parameters of the millisecond-level result reflection feedback, the high-frequency and inefficient node markers of the second-level process reflection feedback, and the cognitive bias distribution of the periodic strategy reflection feedback. Based on the feedback data, the complexity adjustment requirement for the current instruction is determined, and the prompt word optimization template corresponding to the worker's obstacle type is invoked; Using the aforementioned prompt word optimization algorithm, the original standardized job instructions are transformed based on the complexity adjustment requirements and prompt word optimization template, wherein: For cognitive or intellectual disabilities, the original instructions are reduced in text according to the simplification level and repetition frequency in the language control parameters. Modifiers are removed while the subject-verb-object core structure is retained to generate simplified instructions and control the number of times they are broadcast. For hearing impairment, when the high-frequency inefficient node marker indicates a delay in auditory command response, the voice command is mapped to an augmented reality projection, and the operation guidance is superimposed and displayed on the real work surface. The visual identifier of the current step is highlighted and rendered according to the operation error value. For visual impairments, the spatial audio orientation weights are adjusted based on the operational error value and cognitive bias distribution to obtain the three-dimensional coordinates of the material box in physical space, calculate the HRTF filter parameters relative to the worker's ear in real time, and play short voice commands through bone conduction headphones.

7. The disabled employment assistance intelligent agent management system based on a multi-modal reflection architecture of claim 2, wherein, The system also includes real-time adaptive interventional regulation based on extracted heart rate variability, skin conductance, and eye movement characteristics, specifically including: Extract heart rate variability indicators. When the heart rate variability indicator is higher than the preset heart rate variability indicator threshold, trigger process reflection, reduce the command issuance frequency according to the preset frequency adjustment ratio, and start the simplified mode. The phase response parameters of the skin conductance activity index are extracted. When the frequency of change of the phase response parameters exceeds a preset frequency threshold, it is determined that there is an encouraging voice prompt and a short rest is required. The fixation duration and saccade path of eye movement trajectory feature indicators are extracted. When the fixation duration of any saccade path exceeds the preset fixation duration threshold, the visual guide spot is automatically highlighted.

8. The disabled employment assistance intelligent agent management system based on a multi-modal reflection architecture of claim 1, wherein, The steps for performing heterogeneous hardware collaborative scheduling include: Task types are identified based on resource consumption data, and these task types include compute-intensive tasks and I / O-intensive tasks. Using a priority queue scheduler, computationally intensive tasks with computing power requirements exceeding a preset threshold are assigned to central nodes equipped with GPU accelerators, while real-time instruction broadcasting and sensor acquisition tasks are kept at the edge. When the thread waiting time of a computationally intensive task is detected to exceed the preset waiting time threshold, the thread binding adjustment is automatically performed, binding the visual recognition thread to an independent NPU unit at the edge, while the reflection logic is offloaded to the campus edge cloud server.

9. A method for managing intelligent agents assisting employment of persons with disabilities based on a multimodal reflective architecture, applied in the intelligent agent management system for assisting employment of persons with disabilities based on a multimodal reflective architecture as described in any one of claims 1 to 8, characterized in that, The method includes: Acquire multimodal data containing behavioral images, physiological signs, audio interactions, and standardized job instructions of disabled workers; perform feature extraction and cross-modal semantic fusion on the multimodal data; and establish a multimodal semantic subgraph based on task ID. Three-level reflection processing is performed based on multimodal semantic subgraphs to generate corresponding level feedback data. The three-level reflection processing includes result reflection for millisecond-level response, process reflection for second-level response, and strategy reflection for periodic response. Based on the feedback data generated by the three-level reflection process, the original standardized job instructions are automatically transformed into personalized interactive prompts that conform to specific obstacle types using a prompt word optimization algorithm. Based on resource consumption data and task type, perform heterogeneous hardware collaborative scheduling.

10. An electronic device for assisting employment of persons with disabilities based on a multimodal reflective architecture, characterized in that, The electronic device includes one or more processors and one or more memories, wherein at least one computer program is stored in the one or more memories, and the at least one computer program is loaded and executed by the one or more processors to implement the multimodal reflective architecture-based intelligent agent management system for assisting employment of persons with disabilities as described in any one of claims 1 to 8.