Dynamic sign language time sequence modeling method and system based on attention mechanism

By collecting and fusing multi-source features and combining attention mechanisms to perform dynamic sign language temporal modeling, the problems of insufficient accuracy and poor adaptability in existing technologies have been solved. This has enabled efficient and accurate multi-user and multi-scenario adaptation, meeting the needs of barrier-free communication.

CN122024324APending Publication Date: 2026-05-12SHANDONG VOCATIONAL COLLEGE OF SPECIAL EDUCATION
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANDONG VOCATIONAL COLLEGE OF SPECIAL EDUCATION
Filing Date
2026-02-03
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing dynamic sign language temporal modeling technology suffers from insufficient accuracy, poor adaptability to multiple users and scenarios, and weak hardware and algorithm synergy, making it difficult to meet the actual needs of barrier-free communication.

Method used

By synchronously acquiring image data and inertial measurement data to generate a dual-domain temporal tensor, fusing three-dimensional rhythm features and scene prior features, performing three-level attention weighting processing of four-dimensional fused features at the frame, segment, and sentence levels, and loading three-level nested meta-parameters of major categories, subcategories, and semantic units, the system achieves accurate fusion and personalized adaptation of multi-source features.

Benefits of technology

It improves the accuracy and adaptability of dynamic sign language temporal modeling, realizes multi-user personalized adaptation and smooth switching of multiple scenarios, and ensures real-time and efficient execution of modeling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122024324A_ABST
    Figure CN122024324A_ABST
Patent Text Reader

Abstract

The invention provides a dynamic sign language time sequence modeling method and system based on an attention mechanism, belongs to the technical field of sign language recognition and man-machine interaction, and is used for solving the problems of poor sign language semantic recognition scene adaptability, insufficient personalization and inaccurate semantic feature capture in the related technology. According to the method, hand space coordinates and inertial data are synchronously collected and fused to generate a double-domain time sequence tensor, a four-dimensional fusion feature is constructed in combination with rhythm features and scene prior, semantic expression is enhanced through three-level attention weighting, scene adaptation is achieved based on hierarchical meta-parameters, and the accuracy of scene matching is improved. Personalized optimization is achieved through hand feature fingerprints and full-link parameter feedback, and precise sign language semantic modeling, scene self-adaption and individual adaptation are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of sign language recognition and temporal modeling, and in particular to a dynamic sign language temporal modeling method and system based on an attention mechanism. Background Technology

[0002] With the development of accessible communication technologies, dynamic sign language temporal modeling, as a core component of sign language recognition, directly impacts the accuracy of sign language semantic understanding through its modeling accuracy and adaptability. It holds significant application value in accessible interaction across various scenarios, including schools, communities, and government service halls. Currently, dynamic sign language temporal modeling technology largely relies on computer vision and temporal analysis techniques, extracting image features of sign language movements for temporal correlation modeling, and gradually evolving towards multi-feature fusion and attention modeling.

[0003] In existing technologies, dynamic sign language temporal modeling mostly adopts a single action feature-based temporal modeling approach or directly transfers general temporal attention mechanisms. Some solutions attempt to introduce simple scene features to assist modeling. However, these existing technologies still have many shortcomings: First, they ignore the unique characteristics of sign language, such as "strong binding between action and rhythm, and synchronous semantic and rhythmic abrupt changes," and relying solely on action feature modeling easily misses core semantic information; second, general attention mechanisms cannot adapt to the heterogeneity of multi-source features (action, rhythm, scene) in sign language, resulting in poor feature fusion effects; third, they lack personalized adaptation mechanisms for the differences in action habits among multiple users, and there are parameter gaps when adapting across scenes; fourth, semantic matching only focuses on temporal features, ignoring the spatiotemporal nature of sign language actions, resulting in low accuracy in multi-user adaptation; fifth, existing hardware systems are mostly general-purpose computer vision hardware combinations, which cannot adapt to the real-time and parallel computing requirements of sign language modeling, resulting in poor hardware-algorithm synergy. These shortcomings result in insufficient accuracy and limited adaptability of existing modeling methods, making it difficult to meet the needs of actual barrier-free communication. Therefore, there is an urgent need for a dynamic temporal modeling solution that adapts to the unique characteristics of sign language and takes into account personalization and multiple scenarios. Summary of the Invention

[0004] This application provides a dynamic sign language temporal modeling method and system based on an attention mechanism, which can solve the problems of insufficient accuracy, poor adaptability to multiple users and multiple scenarios, and weak hardware and algorithm synergy in existing dynamic sign language temporal modeling, and achieve accurate, efficient and personalized dynamic sign language temporal modeling.

[0005] Firstly, this application provides a dynamic sign language temporal modeling method based on an attention mechanism. Image data and inertial measurement data of dynamic sign language are simultaneously acquired to generate a dual-domain temporal tensor adapted to the amplitude and speed characteristics of sign language movements. This dual-domain temporal tensor, three-dimensional rhythmic features representing semantic transitions in sign language, and scene prior features adapted to the semantic constraints of the sign language scene are fused to form a four-dimensional fusion feature. Based on the four-dimensional fusion feature, three-level attention weighting processing (frame, segment, and sentence) adapted to the semantic hierarchy of sign language is performed to generate a semantically enhanced feature map. Three-level nested meta-parameters (category, sub-category, and semantic unit) adapted to the hierarchical characteristics of the sign language scene are loaded to perform scene-specific adaptation on the semantically enhanced feature map.

[0006] By adopting the above technical solutions, the multi-source features of sign language actions, rhythms, and scenes are systematically integrated into temporal modeling. Combined with hierarchical attention to adapt to the semantic hierarchy characteristics of sign language, it effectively adapts to the semantic expression rules of sign language, improves the accuracy of dynamic sign language temporal modeling, and provides a reliable foundation for subsequent semantic understanding.

[0007] Furthermore, the process of fusing the dual-domain temporal tensor, three-dimensional rhythm features, and scene prior features to form a four-dimensional fusion feature includes: projecting the dual-domain temporal tensor, three-dimensional rhythm features, and scene prior features onto a preset high-dimensional space; mining the intrinsic correlation between sign language actions, rhythm, and scene through attention association calculation to obtain the correlation weights between multi-source features; and outputting dynamic coefficients through a dynamic gating adjustment mechanism to increase the contribution of rhythm features for sign language semantic transition frames and increase the contribution of scene features for cross-scene switching, thereby adjusting the contribution of the three types of features in different frames or different scenes to complete the construction of the four-dimensional fusion feature.

[0008] By adopting the above technical solutions, the accurate fusion of multi-source heterogeneous features of sign language can be achieved, the feature contribution of different frames and scenes can be dynamically adapted, the redundancy or lack of rhythm or scene information can be avoided, and the effectiveness of fused features can be further improved.

[0009] Furthermore, the three-level attention weighting processing based on the four-dimensional fusion features includes: calculating the rhythmic abrupt change degree representing the semantic transition in sign language based on the abrupt change degree parameter in the three-dimensional rhythmic features; identifying semantic transition frames such as sign language negation words and sentence pauses based on the rhythmic abrupt change degree; and dynamically increasing the attention weight of the semantic transition frames by using the rhythmic abrupt change degree as an adjustment factor for the frame-level attention weight.

[0010] By adopting the above technical solutions, we can accurately capture sign language semantic transition frames, avoid missing core semantic information, and improve the targeting and accuracy of frame-level attention modeling.

[0011] Furthermore, the three-level attention weighting processing based on the four-dimensional fusion features includes: calculating the rhythm matching coefficient between the current action segment and the corresponding scene standard sign language rhythm to ensure the standardization of sign language actions; calculating the emotional rhythm coefficient based on the statistical parameters of the action segment rhythm trend to capture the emotional semantics of sign language; calculating the scene adaptation coefficient between the action segment features and the scene-specific semantic unit features to constrain the semantics of sign language context; and outputting dynamic weights through a gating dynamic fusion mechanism to adjust the contribution of the rhythm matching coefficient, emotional rhythm coefficient, and scene adaptation coefficient, thereby generating segment-level attention weights.

[0012] By adopting the above technical solutions, multi-dimensional collaborative weighting of segment-level attention is achieved, taking into account the standardization of sign language movements, emotional semantics, and contextual constraints, thereby improving the completeness of segment-level semantic modeling.

[0013] Furthermore, before performing the three-level attention weighting processing of frames, segments, and sentences, the method further includes: calculating the local rhythm density of each frame based on the three-dimensional rhythm features to characterize the clustering characteristics of sign language movement rhythms; setting an adaptive density threshold that adapts to the change pattern of sign language rhythms and selecting density peak points as the cluster centers of movement segments; and dividing movement segments by taking the midpoint of adjacent cluster centers as the boundary of movement segments and combining the continuity constraints of sign language movements.

[0014] By adopting the above technical solution, we can achieve adaptive and accurate division of action segments, adapt to the rhythm differences of different users, and avoid semantic unit splitting or merging errors caused by rigid division.

[0015] Furthermore, the loading of three-level nested meta-parameters (category, sub-category, and semantic unit) to adapt the semantic enhancement feature map to specific scenarios includes: constructing a four-level nested meta-parameter system (global, category, sub-category, and semantic unit) adapted to the hierarchical characteristics of sign language scenarios, using a storage structure combining basic and incremental parameters; achieving online fine-tuning of meta-parameters through a parameter evolution mechanism based on a small set of sign language action features; and adapting meta-parameters for different scenarios by weighted fusion based on the similarity of sign language semantic units when crossing scenarios.

[0016] By adopting the above technical solutions, we can achieve rapid adaptation to multiple scenarios and smooth transition across scenarios, reduce the adaptation cost for new scenarios and new users, and improve the scenario adaptability of modeling.

[0017] Furthermore, it also includes: acquiring static images of the user's hand, extracting hand physiological features directly related to sign language actions, and generating hand feature fingerprints; constructing a five-level storage architecture based on user fingerprints, subdivided scenarios, semantic units, and parameter types to store personalized adaptation-related data; and using the hand feature fingerprints as an index to achieve personalized parameter matching through a fast indexing mechanism.

[0018] By adopting the above technical solutions, we can achieve accurate differentiation of multiple users and rapid matching of personalized parameters, thereby improving the modeling adaptability in multi-user scenarios.

[0019] Furthermore, it also includes: using an exponential moving average algorithm to update the rhythm baseline that adapts to the slow changes in user action habits; introducing historical loss constraints to update attention weights and meta-parameters based on incremental step size; and feeding back the updated personalized parameters to the feature layer, attention layer, and meta-parameter layer respectively to achieve end-to-end parameter feedback.

[0020] By adopting the above technical solutions, the dynamic evolution of personalized parameters can be achieved, adapting to long-term changes in user behavior habits and ensuring modeling accuracy during long-term use.

[0021] Furthermore, after generating the semantic enhancement feature map, the process further includes: constructing a dynamic sign language spatiotemporal map by concatenating consecutive frames with key hand points as nodes and physical connections between key points as edges; extracting sign language spatiotemporal features through spatial graph convolution branches and temporal convolution branches; and calculating personalized semantic matching similarity by fusing the semantic enhancement feature map, spatiotemporal features, and rhythm alignment results.

[0022] By adopting the above technical solutions, the spatiotemporal essence of sign language movements can be accurately captured, adapting to the spatiotemporal differences of multiple users' movements and improving the accuracy of semantic matching.

[0023] Secondly, this application provides a system for implementing the attention-based dynamic sign language temporal modeling method as described in any of the first aspects above. It includes a data acquisition module, a collaborative computing module, a distributed intelligent storage module, and a control module, with each module interconnected in real time via a data bus. The data acquisition module includes a depth camera, an inertial measurement sensor, and an FPGA synchronous control unit. The FPGA synchronous control unit integrates parallel units for 3D rhythm calculation, hand fingerprint extraction, and spatiotemporal graph node preprocessing adapted to sign language data preprocessing. The collaborative computing module includes a GPU, a CPU, and an integrated parameter scheduling center. The GPU deploys a parallel computing core for spatiotemporal graph convolution and attention fusion adapted to sign language spatiotemporal modeling. The distributed intelligent storage module includes high-speed SSD cache and NAS storage, constructing an integrated index of user fingerprints, subdivided scenes, semantic units, and parameter types adapted to the personalized and scenario-based needs of sign language. The control module performs global collaborative loss function calculation, linking each module to complete the closed-loop control of the entire data acquisition, modeling, matching, and updating process.

[0024] By adopting the above technical solutions, the hardware modules and sign language temporal modeling algorithms are precisely matched, ensuring the synchronization of data acquisition, the efficiency of computation, and the speed of storage, and supporting the real-time closed-loop execution of the entire modeling chain.

[0025] In summary, this application has at least the following beneficial effects: A dynamic sign language temporal modeling method and system that balances accuracy, personalization, and multi-scenario adaptability is provided to achieve efficient and accurate understanding of sign language semantics; By using multi-feature fusion and hierarchical attention modeling, the semantic capture accuracy of sign language temporal modeling is improved; multi-user personalized adaptation and dynamic parameter evolution are achieved to meet long-term usage needs. By co-designing hardware and algorithms, we ensure real-time and efficient execution of the entire modeling process.

[0026] It should be understood that the description in the Summary Section is not intended to limit the key or essential features of the embodiments of this application, nor is it intended to restrict the scope of this application. Other features of this application will become readily apparent from the following description. Attached Figure Description

[0027] The above and other features, advantages, and aspects of the embodiments of this application will become more apparent from the accompanying drawings and the following detailed description. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein: Figure 1 The diagram illustrates the principle of a dynamic sign language timing modeling system based on an attention mechanism, according to an embodiment of this application.

[0028] Figure 2 A flowchart of a dynamic sign language temporal modeling method based on an attention mechanism is shown in an embodiment of this application. Detailed Implementation

[0029] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0030] Furthermore, the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.

[0031] This application provides a dynamic sign language temporal modeling method and system based on an attention mechanism. It adapts to the unique characteristics of sign language, achieves accurate fusion of multiple features such as movement, rhythm, and scene, takes into account personalized adaptation for multiple users and smooth switching between multiple scenes, and provides accurate and efficient modeling with hardware and algorithm co-adaptation, which helps to achieve efficient semantic understanding in barrier-free communication scenarios.

[0032] In a first aspect, embodiments of this application disclose a dynamic sign language temporal modeling system based on an attention mechanism.

[0033] Figure 1 The diagram illustrates the principle of a dynamic sign language timing modeling system based on an attention mechanism, according to an embodiment of this application.

[0034] Reference Figure 1 The system includes a macro-level hardware operating environment that supports the implementation of the dynamic sign language temporal modeling method based on the attention mechanism. This system is based on a multi-module collaborative architecture. Through the orderly connection and functional cooperation of various hardware support systems, it provides real-time, efficient and stable hardware support for the entire process of dynamic sign language temporal modeling, ensuring the orderly implementation of core links such as data acquisition, feature fusion, attention modeling, parameter adaptation and semantic matching in the modeling method.

[0035] The hardware support system upon which this macroscopic equipment environment relies consists of four core hardware modules. These modules are interconnected in real-time via a high-speed data bus, forming a closed-loop hardware support chain to ensure efficient collaboration in data flow, computation execution, storage retrieval, and global control. The data acquisition module, serving as the foundational hardware support, comprises a depth camera, an inertial measurement sensor, and an FPGA synchronization control unit. The depth camera accurately captures spatial posture image data of hands in dynamic sign language, while the inertial measurement sensor synchronously acquires inertial motion data of hand movements. The FPGA synchronization control unit not only achieves temporal synchronization scheduling of the two types of acquired data but also integrates parallel units for 3D rhythm calculation, hand fingerprint extraction, and spatiotemporal graph node preprocessing adapted for sign language data preprocessing, completing real-time preprocessing of data before modeling. The collaborative computing module, serving as the core computing power support, includes a GPU, a CPU, and an integrated parameter scheduling center. The GPU is specifically deployed with a spatiotemporal graph convolution and attention fusion parallel computing core adapted for sign language spatiotemporal modeling, undertaking high-intensity parallel computing tasks. U is responsible for lightweight data processing and module collaborative scheduling. The integrated parameter scheduling center realizes the linkage scheduling and updating of meta-parameters and personalized parameters during the modeling process. The distributed intelligent storage module serves as data support, including high-speed SSD cache and NAS storage. The high-speed SSD cache is used to cache frequently accessed personalized parameters and real-time modeling intermediate data to ensure low-latency reading. The NAS storage is used for long-term storage of massive amounts of user fingerprints, subdivided scenarios, semantic units, parameter types and other related data, and builds corresponding integrated indexes to support fast data retrieval and access. The control module serves as the core of global regulation, specifically executing global collaborative loss function calculation, linking the data acquisition module, collaborative computing module and distributed intelligent storage module, and coordinating the closed-loop control of the entire process of data acquisition, modeling, matching and updating.

[0036] The various hardware support systems cooperate and coordinate with each other. The preprocessed data output by the data acquisition module is transmitted to the collaborative computing module in real time through the data bus. The collaborative computing module calls the relevant parameters in the distributed intelligent storage module to complete the modeling calculation. The intermediate data and final parameters during the calculation process are transmitted back to the distributed intelligent storage module for storage and updating through the data bus. The control module monitors the running status and data flow rhythm of each module throughout the process, forming a complete hardware support closed loop. This fully meets the hardware requirements of the attention mechanism-based dynamic sign language temporal modeling method for real-time performance, accuracy, and adaptability, ensuring the efficient implementation of the modeling method.

[0037] Secondly, embodiments of this application disclose a dynamic sign language temporal modeling method based on an attention mechanism.

[0038] Figure 2 A flowchart of a dynamic sign language temporal modeling method based on an attention mechanism is shown in an embodiment of this application.

[0039] Reference Figure 2 The method specifically includes the following steps: S1: Synchronously acquire image data and inertial measurement data of dynamic sign language, and generate a dual-domain temporal tensor that adapts to the amplitude and speed characteristics of sign language movements.

[0040] The technical solution for this step is described in detail below: First, synchronous acquisition of dynamic sign language data is performed. This acquisition relies on the depth camera and inertial measurement sensor in the system's data acquisition module. The FPGA synchronization control unit ensures time synchronization between the two devices, guaranteeing the consistency of the acquired data's timestamps. Specifically, the depth camera acquires continuous frame image data of the dynamic sign language, with a preset sampling frequency of [missing information]. (This parameter is pre-set based on the dynamic characteristics of sign language movements and can be adjusted according to the actual scene.) After preprocessing, each frame of image data outputs the three-dimensional coordinate data of 21 key nodes of the hand, which is recorded as the hand spatial coordinate sequence. ,in For time step, The total number of frames captured for a single sign language gesture (based on the gesture duration and sampling frequency) Sure, (The duration of the action is in seconds, obtained in real time by the data acquisition system.) Indicates the first Time of the first The three-dimensional coordinates of key points on the hand ( The axes correspond to the horizontal, vertical, and depth directions of the world coordinate system, and the coordinate values ​​are directly acquired and output by the depth camera.

[0041] An inertial measurement sensor is worn on the user's hand and synchronously acquires three-dimensional linear acceleration and three-dimensional angular velocity data of the hand with a depth camera, outputting an inertial data sequence. ,in The first time Linear acceleration in the direction (unit: ), The first time Angular velocity in direction (unit: This data is directly acquired and output by the inertial measurement sensor, and after being aligned with the timestamp by the FPGA synchronization control unit, it is synchronized with the hand's spatial coordinate sequence. Related.

[0042] To reduce the impact of noise on subsequent modeling, the collected data... and Preprocessing is performed. A moving average filtering algorithm is used to denoise the data, with the filter window size... The preset value is 5 (based on the continuity of sign language movements, ensuring effective filtering of high-frequency noise while preserving movement details), and the filtering formula is as follows: , in This is the filtered sequence of hand spatial coordinates. This is the filtered inertial data sequence. This step uses the time step index within the sliding window as an example. The output of this step is the denoised synchronization data pair. .

[0043] Subsequently, a dual-domain time series tensor is generated based on the denoised synchronization data. The dual domains refer to the time domain and the frequency domain. The time-domain time series feature tensor and the frequency-domain time series feature tensor are constructed respectively, and then fused to form the final dual-domain time series tensor.

[0044] Construction of temporal series feature tensors: The denoised... and Concatenate the vectors according to the time step to form a single-time-step temporal feature vector. The dimension of this vector is (The 3D coordinates of 21 key hand points total 63 dimensions, and the inertial data is 6 dimensions, for a total of 69 dimensions). Each time step Arranged in chronological order, forming a temporal series feature tensor. The first dimension of the tensor is the time step. The second dimension is the time-domain feature dimension. This tensor directly represents the dynamic evolution of sign language movements over time.

[0045] Construction of frequency domain time-series feature tensors: for time-domain feature vectors The Fast Fourier Transform (FFT) is used to extract its frequency domain features to capture the frequency characteristics of sign language movements (corresponding to the variation in movement speed). For each dimension of the time domain feature sequence... ( , representing the first eigenvector in the time domain Perform an FFT transformation on each dimension, as shown in the following formula: ; in For the first Frequency domain features of time-domain features For frequency index, The number of effective frequency points (based on the conjugate symmetry of the FFT transform, only the first half of the effective frequency components are retained). The value is the imaginary unit. To avoid the influence of imaginary features on subsequent modeling, the amplitude of the frequency domain feature is taken as the effective feature, i.e. (Amplitude values ​​are real numbers, representing the intensity of the corresponding frequency component.) The frequency domain amplitude features of all dimensions are arranged by frequency index and feature dimension to form a frequency domain time-series feature tensor. The first dimension is the number of effective frequency points. The second dimension is the frequency domain feature dimension (which is consistent with the time domain feature dimension, both being...). This tensor represents the speed variation characteristics of sign language movements (the higher the frequency, the faster the movement speed).

[0046] Finally, the fusion of the two-domain temporal tensors is performed, and the temporal feature tensors are combined using a feature concatenation method. and frequency domain time series feature tensor Merged into a dual-domain temporal tensor Considering the consistency of the dimensions of the time-domain and frequency-domain features, Expand to time step using zero-filling. (like Then, padding is added to the end of the frequency domain tensor. (a row of all zeros), to obtain the expanded frequency domain tensor. The fusion formula is as follows: ; in The final generated bi-domain temporal tensor has its first dimension being the time step. The second dimension is the total dimension of the dual-domain features. (69 dimensions in the time domain + 69 dimensions in the frequency domain, totaling 138 dimensions). This tensor simultaneously contains the temporal dynamic evolution information of sign language movements (corresponding to changes in movement amplitude) and frequency domain velocity characteristics information, achieving precise adaptation of "movement amplitude and velocity characteristics," and providing basic feature support for subsequent multi-source feature fusion and attention-weighted processing.

[0047] S2: The dual-domain temporal tensor, the three-dimensional rhythmic features representing semantic transitions in sign language, and the scene prior features adapted to the semantic constraints of sign language scenarios are fused to form a four-dimensional fused feature.

[0048] The specific steps of this method include: projecting the dual-domain temporal tensor, three-dimensional rhythm features, and scene prior features onto a preset high-dimensional space; mining the intrinsic relationships between sign language actions, rhythms, and scenes through attention association calculations to obtain the association weights between multi-source features; outputting dynamic coefficients through a dynamic gating adjustment mechanism to increase the contribution of rhythm features for sign language semantic transition frames and increase the contribution of scene features for cross-scene switching, adjusting the contribution of the three types of features in different frames or different scenes, and completing the construction of four-dimensional fusion features.

[0049] First, the core features involved in this step are clearly defined and constructed, including the bi-domain temporal tensor. The output from step S1 fully includes the temporal amplitude evolution and frequency velocity characteristics of sign language movements; the three-dimensional rhythm features and scene prior features are key supplementary features for subsequent fusion, and their construction process and parameter definitions are as follows: The construction of three-dimensional rhythm features is based on the hand spatial coordinate sequence after S1 step denoising. Based on this, the core representation of sign language movements' rhythmic variation patterns (including speed, acceleration, and rhythmic abrupt changes) corresponds to the core requirement of "representing semantic transitions in sign language." The specific construction process is as follows: Calculate motion speed characteristics The speed of movement is characterized by the change in the average Euclidean distance between key hand points in adjacent time steps. The formula is as follows: ; in for The Middle The three-dimensional coordinates of the key points ( hour (due to the lack of a preceding time step reference). Units are It directly reflects the average movement amplitude of the action within a single time step; Calculate motion acceleration characteristics The rate of change of rhythm is characterized by the first-order difference of the velocity characteristics, and the formula is as follows: , ( hour The larger this parameter is, the more drastic the change in the rhythm of the movement; Calculate rhythm abrupt change characteristics The semantic transition potential is characterized by the ratio of extreme values ​​in a sliding window based on acceleration features, as shown in the formula: ; in Set the preset sliding window size (based on the continuous characteristics of sign language gestures). To avoid the minimum value where the denominator is zero, ( A frame is identified as a potential semantic transition frame when a preset mutation threshold (obtained through statistical analysis of a large number of sign language samples) is set.

[0050] By concatenating the above three dimensions according to time steps, a three-dimensional rhythmic feature is formed. The overall dimension is .

[0051] The construction of scenario-prior features is adapted to the requirement of "sign language scenario semantic constraints." These features originate from a pre-set sign language interaction scenario library (including typical accessibility scenarios such as government affairs, campuses, and communities). Each scenario corresponds to a unique set of semantic units (e.g., government affairs scenarios include semantics such as "registration" and "payment," while campus scenarios include semantics such as "enrollment" and "consultation"). Scenario-prior features are presented as scenario semantic embedding vectors, denoted as... ,in The embedding dimension is preset (determined in advance based on word vector training of scene semantic units). The time step is (because the scene remains fixed within a single sign language movement, therefore...) The embedding vectors are the same at all time steps, i.e. , This is a preset embedding vector corresponding to the current scene, which is directly called by the system based on the user's selected interaction scenario.

[0052] After constructing the three types of features, high-dimensional projection processing of the multi-source features is performed. The core purpose is to unify the dimensional space of the three types of features and eliminate the impact of dimensional differences between heterogeneous features on the fusion effect. The preset dimension of the high-dimensional space is [value missing]. (Based on a pre-set balance between feature representation capability and computational efficiency), linear projection matrices are designed for the three types of features respectively. , , The projection matrix parameters are obtained through training on a pre-defined sign language sample dataset (the training objective is to minimize the semantic discrimination loss of the projected features). The projection formulas are as follows: ; in Three types of features are respectively in The high-dimensional projection results at time t. This is a preset bias term (obtained through synchronous training with the projection matrix) used to compensate for feature offset after projection.

[0053] Subsequently, the intrinsic relationships among the three types of features are mined through attention-based association calculations to obtain association weights. A feature association model is constructed using a cross-attention mechanism, projecting the high-dimensional results of the dual-domain temporal tensor. As a query, the three-dimensional rhythmic feature projection result and scene prior feature projection results The concatenation serves as a key and value, focusing on the relationship between action features and rhythm and scene features. The specific calculation steps are as follows: Constructing a query matrix Key matrix Value matrix ; Calculate the attention score: The strength of the feature association is represented by the dot product of the query and the key, using the following formula: ; in Indicates the first Moment action characteristics and the first The correlation between the temporal rhythm and scene joint features. This is a normalization term to avoid excessively large scores; Calculate the association weights: Normalize the attention scores using the Softmax function to obtain the weight matrix. The formula is ; Characterizing the first Moment rhythm - scene features on the first The correlation contribution weight of action features at different times.

[0054] Based on attention-related weights, a dynamic gating mechanism is used to output dynamic coefficients, achieving adaptive adjustment of the contribution of the three types of features. The gating mechanism employs a dual-gating structure (semantic transition gate and scene switching gate) to adapt to the special requirements of semantic transition frames and cross-scene switching. First, two judgment flags are defined: 1) Semantic transition flag. :when ( When the mutation threshold is consistent with that in the three-dimensional rhythm features, (Indicates that the current frame is a semantic transition frame), otherwise it is 0; 2) Scene switching flag : ( )hour, (Indicates that the current frame is a cross-scene transition frame), otherwise it is 0 ( ).

[0055] The formula for calculating the dynamic coefficient is as follows: ; in These are the dynamic adjustment coefficients for the two-domain temporal tensor, the three-dimensional rhythmic features, and the scene prior features, respectively. Use the Sigmoid activation function (ensuring that the coefficients are within the range of (0,1)). This is the gating weight matrix (obtained through training with sign language samples). This is the gated bias term (obtained during synchronous training); in the formula and To enhance the contribution, the requirement is to "enhance the contribution of rhythm features by semantic transition frames and enhance the contribution of scene features by cross-scene switching". The 0.3 decay term of the dual-domain temporal tensor coefficients reserves weight space for the enhancement term to ensure the balance of the total contribution.

[0056] Finally, the three types of features are fused to construct a four-dimensional fused feature. First, the adjusted three types of features are weighted element-wise, using the following formula: ; The weighted features are then concatenated along the feature dimensions to obtain the fused feature vector at a single time step. .Will Each time step Arranged chronologically, and combining the time step dimension, feature dimension, feature type dimension, and dynamic adjustment coefficient dimension, a four-dimensional fused feature is formed. The four dimensions are: "time step" →High-dimensional feature dimensions →Feature type (action / rhythm / scene) →Dynamic coefficient weight”, this feature fully preserves the intrinsic correlation and adaptive adjustment information of multi-source features, providing accurate feature input for subsequent three-level attention weighting processing.

[0057] S3: Based on the four-dimensional fusion features, perform three-level attention weighting processing of frames, segments, and sentences to adapt to the semantic level of sign language, and generate semantically enhanced feature maps.

[0058] The specific steps of this method include: before performing the three-level attention weighting processing of frames, segments, and sentences, calculating the local rhythm density of each frame based on the three-dimensional rhythm features to characterize the clustering characteristics of sign language action rhythms; setting an adaptive density threshold that adapts to the rhythm change patterns of sign language, and selecting density peak points as action segment cluster centers; using the midpoint of adjacent cluster centers as the boundary of the action segment, and dividing the action segment in combination with the continuity constraints of sign language actions; after dividing the action segment, calculating the rhythm abruptness that characterizes the semantic transition of sign language based on the abruptness parameter in the three-dimensional rhythm features; identifying semantic transition frames such as sign language negation words and sentence pauses based on the rhythm abruptness; using the rhythm abruptness as an adjustment factor for frame-level attention weights, dynamically increasing the attention weight of the semantic transition frames, and simultaneously calculating the rhythm matching coefficient between the current action segment and the standard sign language rhythm of the corresponding scene to ensure... The system aims to standardize sign language movements; calculate emotional rhythm coefficients based on statistical parameters of movement segment rhythm trends to capture sign language emotional semantics; calculate scene adaptation coefficients of movement segment features and scene-specific semantic unit features to constrain sign language contextual semantics; output dynamic weights through a gating dynamic fusion mechanism to adjust the contribution of the rhythm matching coefficient, emotional rhythm coefficient, and scene adaptation coefficient, generating segment-level attention weights, and then completing frame, segment, and sentence-level attention weighting processing to generate a semantic enhancement feature map; after generating the semantic enhancement feature map, construct a dynamic sign language spatiotemporal graph by connecting consecutive frames with hand key points as nodes and physical connections between key points as edges; extract sign language spatiotemporal features through spatial graph convolution branches and temporal convolution branches; and calculate personalized semantic matching similarity by fusing the semantic enhancement feature map, spatiotemporal features, and rhythm alignment results.

[0059] First, clarify the source and definition of the core input parameters for this step: four-dimensional fusion features. The output from step S2 fully preserves the association and dynamic adjustment information of multi-source features; three-dimensional rhythm features Also from S2, its mutation degree parameter This is the core basis for semantic transition recognition; the subsequent standard sign language rhythms and scene-specific semantic unit features all come from the pre-set sign language scene semantic database (and the S2 scene prior features). (The system calls the parameters based on the current interaction scenario, and the personalized parameters will be associated with the output of the S5 steps to ensure personalized adaptation of semantic matching.)

[0060] Before performing frame, segment, and sentence-level attention weighting processing, action segment division is completed first. The core principle is to achieve natural segmentation of actions based on rhythmic clustering characteristics, providing the basic units for segment-level attention processing. The calculation of local rhythm density is based on velocity in the three-dimensional rhythmic features. As the core indicator (the degree of speed concentration directly reflects the intensity of rhythm), the Gaussian kernel density estimation method is used, and the formula is: ; in For the first Local rhythm density of a frame. for The temporal neighborhood window of a frame (preset window size) That is, including (Frames, set based on the continuity characteristics of sign language gestures). The preset Gaussian kernel bandwidth was determined through statistical analysis of a large number of sign language samples to ensure effective differentiation between rhythmically dense and sparse regions. Using the L2 norm, this formula characterizes the degree of clustering of tempo in the current frame by calculating the weighted sum of similarity of velocities within the neighborhood.

[0061] The adaptive density threshold setting needs to adapt to the rhythmic variations of different sign language movements to avoid segmentation bias caused by a fixed threshold. The calculation method combines global density statistics with local density correction. ; in For the first The adaptive density threshold corresponding to the frame, This represents the global rhythm density average. The global density standard deviation, This is a local density correction factor that ensures the threshold is adaptively increased in regions of high rhythm density, avoiding misidentification of cluster centers. The criteria for selecting density peak points are: and For its neighborhood window The maximum value within which the condition is satisfied. That is, the cluster center of the action segment, denoted as ( (This refers to the number of cluster centers, i.e., the number of action segments).

[0062] Action segment boundary delineation is based on the midpoint of adjacent cluster centers, while incorporating sign language action continuity constraints (to avoid boundaries crossing key action frames). Specific boundary locations... The formula for determining it is: ; in For adjacent cluster centers ( ), For floor operations, Preset continuity threshold (unit: (To ensure smooth changes in motion at the boundary), if the continuity constraint is not met, the boundary will be adjusted to... (That is, the frame with the lowest velocity between adjacent cluster centers, ensuring the integrity of the action segment). The final action segment is... , record The action segment is The included frame index is ( The first frame of the segment. (This is the end frame of the segment).

[0063] After the action segments are divided, frame-level attention weighting processing is performed. First, the abrupt change parameter of the three-dimensional rhythmic features in S2 is reused. (already defined) ), and based on this, identify semantic transition frames: when When a threshold for a sudden change is set (obtained statistically from sign language semantic transition samples), it is identified as a semantic transition frame (such as the action frame corresponding to the negation word "no" or a sentence pause), and the set of transition frames is recorded as follows. .

[0064] The calculation of frame-level attention weights is based on four-dimensional fusion features, combined with abrupt change parameters for dynamic adjustment, and a self-attention mechanism is used to calculate the basic attention weights. The formula is as follows: ; in The result of flattening the four-dimensional fusion features (two-dimensional feature dimensions after flattening). For attention projection matrix ( (The parameters for the preset attention dimension are obtained through training with sign language samples). Based on attention score, The normalized base weights represent the first... Frame to the first The correlation contribution of frames.

[0065] Using the rhythmic abruptness as a modulating factor, the attention weight of semantic transition frames is dynamically increased. The final frame-level attention weight formula is as follows: ; in The preset boosting factor (determined through experiments to balance the weights of transition frames and non-transition frames) is used. For indicator functions ( (1 if true, 0 otherwise). This weight is used to weight the flattened four-dimensional fusion features to obtain frame-level enhanced features. .

[0066] Simultaneously with frame-level processing, segment-level attention-weighted processing is performed. The core of this process involves calculating three key coefficients and generating segment-level weights through gating fusion. First, the rhythm matching coefficients are calculated. This is used to ensure the standardization of action segments, and its calculation is based on the current action segment. The similarity between the rhythmic characteristics and the standard sign language rhythm template for the scene. Standard sign language rhythm template for the scene. From the preset scene semantic library ( The frame count of the standard action segment is determined by the Dynamic Time Warping (DTW) algorithm. Align to The rhythm matching coefficient is calculated using cosine similarity. ; in This is the result of DTW alignment of the rhythmic features of the current action segment. The closer the value is to 1, the better the rhythm of the current action segment matches the standard rhythm.

[0067] Emotional rhythm coefficient It is used to capture the emotional semantics in sign language movements (e.g., fast movements correspond to excitement, slow movements correspond to calmness), and is calculated based on statistical parameters of the rhythmic trend of movement segments. The statistical parameters include the mean speed within the segment. ,variance and maximum value (All derived from the velocity components of three-dimensional rhythmic features) The formula is: ; in Use the Sigmoid activation function (ensure the coefficients are in the range of (0,1)). This is the global average sign language movement speed (preset statistical value). The weights are preset (obtained through training with sign language samples labeled with emotions), and the larger the coefficient, the stronger the emotion of the action.

[0068] Scene adaptability coefficient Used to constrain the contextual semantics of an action segment, calculating the similarity between the features of the current action segment and the features of scene-specific semantic units. Scene-specific semantic unit features From the preset scene semantic library ( (Number of semantic units in the current scene), features of the current action segment. (Mean of frame-level enhancement features within a segment), scene adaptation coefficient is: ; That is, the maximum cosine similarity between the current action segment features and the features of all semantic units in the scene. The closer the value is to 1, the better the action segment is suited to the current scene context.

[0069] A gating dynamic fusion mechanism is used to adjust the contribution of the three coefficients, generating segment-level attention weights. The input to the gating mechanism is the concatenated vector of the three coefficients, and the output is dynamic weights. (satisfy The formula is: ; in, As the input for the gating mechanism, It is the feature coefficient corresponding to the current action segment (the m-th action segment) (such as the length coefficient, complexity coefficient, overlap coefficient, etc. of the action segment, which needs to be combined with the action feature extraction logic of the method). It is a three-dimensional column vector formed by combining the three coefficients in order (the superscript T is the "transpose" operation, which converts the row vector into a column vector to meet the dimensional requirements of subsequent matrix operations). It is an activation function (usually the sigmoid function, which maps the result of the operation to the (0,1) interval to ensure that the dynamic weights calculated in subsequent calculations are positive). For the gated weight matrix, These are bias terms (all obtained through training with sign language samples). Ultimately, it is a three-dimensional vector with 3 elements ( () is the subsequent calculation of dynamic weights The basis is (through normalization operations, the sum of the three dynamic weights is made to 1). The first gating output vector There are 10 elements. The segment-level attention weight is 10. By weighting the action segment features using this weight, segment-level enhanced features are obtained. .

[0070] Sentence-level attention-weighted processing takes segment-level enhanced features as input, captures the semantic relationships between action segments, and forms a complete sentence-level semantic representation. A global self-attention mechanism is used to calculate the inter-segment association weights, using the following formula: ; in For the set of all segment-level enhanced features, For sentence-level attention projection matrix ( (training) This is a sentence-level enhancement feature.

[0071] After completing the three-level attention weighting process (frame, segment, and sentence), a semantically enhanced feature map is generated. The frame-level enhanced features are then... Segment-level enhancement features (Align to time step via zero padding) Sentence-level augmentation features (Also aligned to time step) By concatenating along the feature dimensions, a semantically enhanced feature map is obtained. Each element of this feature map corresponds to a multi-scale semantic enhancement feature at a certain time, providing a foundation for subsequent spatiotemporal feature extraction and semantic matching.

[0072] After generating semantically enhanced feature maps, a dynamic sign language spatiotemporal map is constructed. The coordinates of the key hand points in the core sign language movements are derived from the denoised hand spatial coordinate sequence in step S1. (21 key points on the hand, including 5 on the thumb, 4 on the index finger, 4 on the middle finger, 4 on the ring finger, and 4 on the little finger), each key point is used as a node in the spacetime diagram, denoted as the [node name missing]. The set of nodes at time point is The edges between nodes are constructed based on the physical connections of the hand's physiological structure, with a pre-defined set of edges. (Such as 19 fixed edges, such as thumb tip-to-thumb proximal phalanx, index finger tip-to-index finger proximal phalanx, covering hand bone connections), while introducing dynamic edge constraints: when the Euclidean distance between two key points... (When a preset distance threshold is used, based on hand size statistics) a dynamic edge is added. Finally the The set of edges at time is . will continue The nodes and edges of the frames are concatenated in chronological order to form a dynamic sign language spatiotemporal graph. The diagram contains both the spatial structure and temporal evolution information of hand movements.

[0073] The spatial-graphical convolutional branch extracts the spatiotemporal features of sign language. The spatial-graphical convolutional branch employs a Graph Convolutional Network (GCN) to extract spatial structure features. Its core principle is the aggregation of node features based on the adjacency matrix. The construction rules are as follows: (like ),otherwise and through normalization processing ( The formula for spatial graph convolution is: ; in This is the node feature matrix (3D coordinates of key points). The spatial convolution weight matrix ( (training) For bias terms, For the first The spatial features of a frame are obtained by global average pooling. ( For the first Spatial characteristics of each node.

[0074] The temporal convolution branch uses 1D convolution to extract temporal evolution features, with continuous input. Spatial feature sequence of frames The convolution formula is: ; in This represents a 1D convolution operation. For temporal convolution kernels ( The kernel size is [size]. (The number of output channels is obtained during training). Step size, To fill (ensure the output time step matches the input). It is a time-related feature.

[0075] By fusing the outputs of the spatial graph convolution branch and the temporal convolution branch, the spatiotemporal features of sign language are obtained. The fusion method is feature splicing: .

[0076] Finally, the semantic enhancement feature map, spatiotemporal features, and rhythm alignment results are fused to calculate personalized semantic matching similarity. The rhythm alignment result comes from the alignment processing of S2 three-dimensional rhythm features (the current action rhythm is aligned with the standard rhythm using the DTW algorithm to obtain the alignment matrix). , (Standard rhythm frame count), personalized parameters The personalized memory (containing adaptation parameters corresponding to the user's hand fingerprint features, which is called in advance here) comes from the subsequent S5 step. First, the semantically enhanced feature map... and spatiotemporal characteristics Unify dimensions using linear projection. : ; in Let be the projection matrix. These are bias terms (all obtained through training).

[0077] Introducing personalized parameters ( For personalized weight matrices, This is a personalized bias term (obtained by matching the user's hand fingerprint index) used to personalize the fused features: ; Combine the rhythm alignment results to calculate the standard semantic features of the scene. The formula for personalized semantic matching similarity is: ; in For the elements of the rhythm alignment matrix (representing the first...) Frame and Standard (frame alignment weights) The final personalized semantic matching similarity value is 1. The closer the value is to 1, the higher the matching degree between the current sign language action and the standard semantics of the scene, providing a core basis for subsequent semantic recognition.

[0078] The specific rules for applying personalized semantic matching results are as follows: Set a semantic matching threshold. (Determined through statistical analysis of a large number of sign language samples, balancing recognition accuracy and recall), if If the current sign language action matches the standard semantics of the scene, the corresponding semantic unit label (such as "payment" or "query") is output; if If this occurs, a secondary matching process is triggered, recalculating the matching similarity using the meta-parameters fused from the S4 cross-scene model; if If the match fails, the user is prompted to re-execute the sign language action, and the failed sample is added to the S5 support set for subsequent parameter fine-tuning. Simultaneously, the matching result is associated with the user's fingerprint ID and scene information and stored in a five-level storage architecture to provide data support for historical loss calculation (S5).

[0079] S4: Load the three-level nested meta-parameters of major categories, subcategories, and semantic units that adapt to the hierarchical characteristics of sign language scenarios, and perform scenario-based adaptation on the semantic enhancement feature map.

[0080] The specific methods in this step include: constructing a four-level nested meta-parameter system of global, major category, sub-category, and semantic unit levels adapted to the hierarchical characteristics of sign language scenarios, and adopting a storage structure that combines basic parameters and incremental parameters; achieving online fine-tuning of meta-parameters through a parameter evolution mechanism based on a small set of sign language action features; and, in cross-scenario scenarios, using weighted fusion of sign language semantic unit similarity to adapt meta-parameters to different scenarios.

[0081] First, let's clarify the source of the core input and basic dependency parameters for this step: semantic enhancement feature maps. The output from step S3 fully carries multi-scale semantic enhancement information of sign language actions; the scene-related hierarchical classification is based on a pre-defined sign language scene classification system (and the scene prior features from step S2). The system is based on the current interaction scenario and is called by the system. It is divided into four levels: "global-major category-sub-semantic unit" and covers typical barrier-free interaction scenarios such as government affairs, campus, and community. The basic data of the meta-parameters comes from the pre-training results of a large-scale sign language corpus. Subsequent online fine-tuning and cross-scenario fusion are all based on this foundation to ensure a balance between the universality and adaptability of the parameters.

[0082] When constructing a four-level nested meta-parameter system (global, major category, sub-category, and semantic unit) adapted to the hierarchical characteristics of sign language scenarios, it is necessary to first clarify the specific definition of the four levels and the functional positioning of the corresponding meta-parameters. The four-level division rules are as follows: 1) Global layer: A general parameter layer covering all sign language interaction scenarios, adapting to the common rules of sign language actions (such as spatiotemporal modeling parameters of basic hand movements); 2) Major category layer: A primary category divided according to the attributes of the interaction scenario (such as government services, campus communication, and community life), adapting to the common semantic expression rules of the same type of scenario; 3) Sub-category layer: Specific sub-scenarios under the major category scenario (such as "social security processing" and "government consultation" under the government services category, and "course consultation" and "registration" under the campus communication category), adapting to the exclusive semantic unit distribution of the sub-scenarios; 4) Semantic unit layer: The smallest semantic unit under each sub-scenarios (such as the action units corresponding to specific sign language semantics such as "payment," "inquiry," and "loss reporting" under the "social security processing" scenario), adapting to the precise semantic matching requirements.

[0083] Corresponding to the four-level hierarchy, the nested meta-parameter system is denoted as... The definitions and dimensions of parameters at each level are as follows: 1) Global meta-parameters ,in The feature dimensions of the semantically enhanced feature map (from the S3 feature concatenation result). The global parameter output dimension is obtained by training on a large-scale general sign language corpus (covering 10+ major scene categories and 100+ sub-scenes). Its core function is to provide basic semantic feature mapping capabilities; 2) Major category meta-parameters ( Preset the number of major scene categories ), each ( ), obtained through training on corpora corresponding to major scene categories, and adapted to the semantic preferences of major scene categories; 3) Subdivided meta-parameters ( To further refine the scene index, each major category contains... (each sub-scenarios) ( ), obtained through fine-tuning of segmented scene corpora; 4) Semantic unit meta-parameters ( For semantic unit indexing, each sub-scenario contains (semantic units) The threshold parameter is matched to the features of each semantic unit and trained using the semantic unit labeled samples.

[0084] A storage structure combining basic and incremental parameters is adopted to balance parameter universality and scenario adaptability, thereby reducing storage and retrieval costs. Among these, the basic parameters... These are common parameters shared across all scenarios, stored in read-only mode in the NAS storage of the distributed intelligent storage module (from the distributed intelligent storage module in the system claim); incremental parameters For specific adaptation parameters tailored to different scenarios and semantic units, a read-write mode is used to store them in a high-speed SSD cache, supporting online updates and fast retrieval. During storage, an index is built based on "Global ID - Category ID - Specific Scenario ID - Semantic Unit ID" (from the same source as the system's unified index), ensuring that during loading, the corresponding level of meta-parameters can be quickly located and loaded based on scenario information. The loading formula is as follows: ,in This serves as a hierarchical identifier for the current scene (input by the system based on the user interaction scenario). This is the set of three nested meta-parameters ("major category + sub-category + semantic unit") that need to be loaded for the current scenario.

[0085] Based on a small support set of sign language action features, online fine-tuning of meta-parameters is achieved through a parameter evolution mechanism. The core is to optimize incremental parameters using a small number of scene-specific samples, thereby improving the scene adaptation accuracy of the semantically enhanced feature map. First, the definition and source of the support set are clarified: the support set of sign language action features. ,in To support the number of samples in the set (preset) (This falls under the category of "small sample size"). For the first Semantic enhancement feature maps of support set samples (from S3's processing results of sign language actions on the support set). The semantic unit labels corresponding to the samples (obtained through manual annotation, and...) (corresponding to the semantic unit index), the supported set is provided by the user in a small amount of data collected in the current scenario or by the system's preset scenario-specific sample library.

[0086] The parameter evolution mechanism employs an incremental gradient descent algorithm, aiming to minimize the matching loss between semantically enhanced features and meta-parameters, and only applies incremental parameters. Fine-tuning (basic parameters) (Fixed to ensure generality). First, define the matching loss function as cross-entropy loss, with the formula: ; in The Sigmoid activation function maps the matching results to the (0,1) interval, representing the matching probability between features and semantic units; The matching calculation between semantic enhancement features, after being mapped by subdivided meta-parameters, and semantic unit meta-parameters is performed, and the output is the matching score.

[0087] The formula for incremental parameter update based on the loss function is: ; in For the first Incremental parameters for wheel fine-tuning For the updated parameters, The preset learning rate is determined experimentally, balancing fine-tuning speed and parameter stability. This represents the gradient of the loss function with respect to the incremental parameters. A gradient clipping mechanism is introduced during fine-tuning to limit the maximum norm of the gradient. To avoid gradient explosion leading to parameter distortion, after fine-tuning, the updated parameters will be... Write back to the high-speed SSD cache to overwrite the original incremental parameters, enabling online updates and iterations of meta-parameters.

[0088] When working across different scenarios, the core of adapting meta-parameters for different scenarios is to solve the problem of "meta-parameter discontinuity during scenario switching" and ensure a smooth transition in semantic adaptation. First, let's define "cross-scenario": Let the current scenario be... The target switching scene is The two scenarios belong to different sub-scenarios within the same broad category (such as switching from "Social Security Processing" to "Government Affairs Consultation" under the category of government services) or different broad categories (such as switching from "Government Services" to "Campus Communication"), and therefore need to be integrated. Fine-tuned meta-parameters and The original meta-parameters (Base incremental parameters that have not been fine-tuned by the current user).

[0089] The calculation of sign language semantic unit similarity is based on the semantic unit features of two types of scenarios. ( The set of semantic unit features for (number of semantic units) and ( The semantic unit feature set (the set of features) all come from a preset scene semantic library, and are composed of semantic unit meta-parameters. The mapping is obtained. Cosine similarity is used to calculate the similarity between pairwise semantic units, with the formula: ; in for The A semantic unit, for The A semantic unit, The closer the value is to 1, the stronger the semantic correlation between the two semantic units.

[0090] The fusion weights of cross-scene meta-parameters are calculated based on semantic unit similarity. First, the fusion weights are calculated for... Each semantic unit ,turn up The semantic unit with the highest similarity The corresponding similarity is used as the basic weight, and then normalization is used to obtain the final fusion weight. and (satisfy The formula is: ; in To avoid the minimum value where the denominator is zero, The maximum similarity of semantic units in two types of scenarios represents the strength of semantic association between scenarios.

[0091] The weighted fusion formula for cross-scene meta-parameters is: ; in For the fused cross-scene meta-parameters, if the semantic correlation between scenes is high ( ),but Larger values ​​(prioritize retaining fine-tuning experience for the current scenario); if the correlation strength is low ( ),but Larger values ​​are preferred (the original parameters of the target scenario are used first to ensure scenario adaptability).

[0092] Finally, the semantically enhanced feature map is adapted to the scene using the fused meta-parameters (either online fine-tuned or cross-scene fused). The adaptation process involves feature mapping and matching filtering. By inputting the fused meta-parameters, the semantic unit matching probability at each time step can be obtained. ( (The number of semantic units in the current / target scene), with a matching probability greater than a preset threshold. Semantic units (obtained through scene sample statistics) are used to form a semantic feature sequence after scene adaptation. ,in For element-wise multiplication, For diagonal matrix construction operations, Ensure that only highly matching semantic features are retained. This adapted feature sequence. It fully adapts to the semantic expression needs of the current scenario, providing a precise feature foundation for subsequent personalized semantic matching (S3 follow-up steps) and parameter feedback (S5).

[0093] S5: Collect static images of the user's hand, extract hand physiological features directly related to sign language actions, and generate hand feature fingerprints; construct a five-level storage architecture based on user fingerprints, subdivided scenarios, semantic units, and parameter types to store personalized adaptation-related data; use the hand feature fingerprints as indexes to achieve personalized parameter matching through a fast indexing mechanism.

[0094] The specific methods in this step include: updating the rhythm baseline to adapt to the slow changes in user action habits using the exponential moving average algorithm; introducing historical loss constraints and updating attention weights and meta-parameters based on incremental step size; and feeding back the updated personalized parameters to the feature layer, attention layer, and meta-parameter layer respectively to achieve end-to-end parameter feedback.

[0095] First, clarify the core hardware modules and the sources of prerequisite parameters for this step: Hand static image acquisition relies on the data acquisition module (depth camera) in the system claims, and the storage architecture depends on a distributed intelligent storage module (high-speed SSD cache + NAS storage); prerequisite parameters include the semantically enhanced feature map output by S3. S4 fine-tuned meta-parameters Three-dimensional rhythmic features of S2 And a preset template for extracting hand physiological features (constructed based on human hand anatomical features, covering the physiological structures associated with the core of sign language movements).

[0096] The acquisition of static images of the user's hand needs to ensure the stability of feature extraction. A static shooting mode using a depth camera is employed, with the following parameters set: resolution. (Preset, balancing acquisition accuracy and storage cost), Shooting distance (Adjusted by the user as prompted by the system) Light intensity (To avoid shadow interference with feature extraction). During the acquisition process, users should be guided to maintain a natural hand extension posture (palm facing forward, fingers spread, covering the most essential basic posture in sign language movements). Three consecutive still images should be acquired, and the images should be filtered based on sharpness (sharpness score). ,in The image height and width, For pixels Gradient values, filtering (Images), select the frame with the highest clarity as the valid image. .

[0097] When extracting hand physiological features directly related to sign language movements, effective images are used. Based on the 21 hand key points detected (consistent with the hand key point definitions in S1 and S3, derived from the hand pose detection algorithm output by the depth camera), 8 core physiological features (highly stable and directly related to the amplitude / flexibility of sign language movements) were selected. The specific extraction and calculation methods are as follows: 1. The hand key points in S3 are defined consistently, with index 0 representing the palm, 1 representing the wrist, and 4 / 8 / 12 / 16 / 20 representing the fingertips, etc. Based on this, eight core features with strong stability and significant impact on the amplitude / flexibility of sign language movements are selected. The extraction logic and calculation formula for each feature are as follows: 1) Index finger - middle finger length ratio The ratio of the Euclidean distance from the tip of the index finger (index 12) to the metacarpophalangeal joint of the index finger (index 9) to the distance from the tip of the middle finger (index 16) to the metacarpophalangeal joint of the middle finger (index 13). ; in For the first 1) Three-dimensional coordinates of key points (directly from depth camera data), this ratio determines the feature boundary of finger extension sign language movements; 2) Palm width-to-palm length ratio: distance between the left and right boundaries of the palm (leftmost key point) Key point on the far right The ratio of the distance from the palm (index 0) to the distance from the wrist (index 1) to the distance from the center of the palm (index 0). ; 3) The ratio of thumb metacarpophalangeal joints to the distance between the thumb and index finger joints. The ratio of the distance from the tip of the thumb (index 4) to the metacarpophalangeal joint of the thumb (index 1) to the distance from that joint to the center of the palm. ; 4) The ratio of the length of the ring finger to the little finger affects the accuracy of thumb-dominant sign language recognition; 5) Index finger - middle finger extension angle : Calculate the angle between two fingers using the vector dot product. ; The unit is radians, reflecting the basic characteristics of finger opening and closing movements; 6) Palm thickness index The difference between the palm depth coordinate and the wrist depth coordinate. ( (For depth direction coordinates), adapting motion feature modeling to the depth dimension; 7) Middle finger joint spacing ratio (Ratio of distance from fingertip to proximal joint to proximal joint to middle joint); 8) Thumb-index finger opening angle This is the core feature of thumb-related sign language movements.

[0098] Among the above features, The length ratio is dimensionless. The angular features are all calculated directly from the coordinates of key points acquired by the depth camera, without the need for additional hardware input.

[0099] When generating hand fingerprint features, the eight physiological features are first standardized (to eliminate individual scale differences and dimensional effects) using a min-max normalization algorithm, with the following formula: ; in The sample set of hand physiological features is a pre-set set (containing statistical data on hand features of 1000+ people of different ages and genders, sourced from a public physiological database of sign language users). The first The global minimum and maximum values ​​of each feature, after standardization To highlight the weights of core features, a feature weight vector is introduced. (Thumb-related features were determined through a sensitivity experiment on sign language movement characteristics) The thumb has the highest weight (because it is the most dominant in sign language gestures), resulting in the final hand feature fingerprint vector. ( (for element-wise multiplication), where After generation, uniqueness verification needs to be performed, and the cosine similarity with existing fingerprints in the distributed storage module needs to be calculated. ; like (If a preset threshold is used to ensure fingerprint uniqueness), it is determined to be a new user and a unique 16-bit string user fingerprint ID is assigned (e.g., FP2024050100000001); otherwise, the existing user fingerprint ID is reused to avoid duplicate storage.

[0100] When constructing a five-level storage architecture based on user fingerprint, sub-scenario, semantic unit, and parameter type, a "major scenario" level is added to form a complete five-level link (adapting to S4's hierarchical meta-parameter system). The final hierarchy is "User Fingerprint ID - Major Scenario - Sub-scenario - Semantic Unit - Parameter Type". The storage positioning and content of each level are as follows: 1) User Fingerprint ID layer: Using a unique ID as an index, it associates with the hand feature fingerprint vector. 1) **Storage Layer:** Stored in NAS storage (long-term fixed storage, supporting unique user identification); 2) **Category Scenario Layer:** Corresponding to the major scenario categories of S4 (e.g., government services, campus communication), storing a list of commonly used major scenario categories in the user's history (generated from system-recorded interaction logs), located in a high-speed SSD cache (high-frequency access); 3) **Sub-Scenario Layer:** Specific sub-scenarios under each major category (e.g., "social security processing" under government services), storing the user's historical modeling data in that scenario, located in an SSD cache; 4) **Semantic Unit Layer:** The smallest semantic unit under each sub-scenario (e.g., "payment" "query"), associated with the matching threshold and historical features of the user's corresponding semantic unit, located in an SSD cache; 5) **Parameter Type Layer:** Core storage of personalized adaptation data, specifically divided into three categories—attention weight parameters (frame, segment, and sentence-level weights of S3), meta-parameters (user-specific incremental meta-parameters of S4). The rhythm baseline parameters (user action habit benchmark) adopt a storage strategy of "SSD caching + NAS backup" (highly updated parameters are stored on SSD and backed up to NAS regularly).

[0101] During storage, a five-level index structure is constructed: "User Fingerprint ID - Category ID - Sub-category ID - Semantic Unit ID - Parameter Type ID" (from the same source as the system's integrated index), with the index key being the MD5 hash value of the user fingerprint ID. This ensures quick location and querying.

[0102] The fast indexing mechanism based on hand fingerprint features is implemented using a B+ tree (to meet the high-efficiency query requirements of hierarchical storage). The core process consists of two steps: 1) Fingerprint index matching: matching the current fingerprint index... 1) Generate an index key using a hash function, traverse from the root node to the leaf node in the B+ tree to locate the corresponding user fingerprint ID; if it is a new user (no matching index), insert a new index key and ID into the leaf node to complete the registration; 2) Parameter location: Based on the user fingerprint ID, traverse the major category, sub-scenario, and semantic unit levels in sequence to finally locate the target personalized parameter at the parameter type layer. .

[0103] To improve efficiency, a cache preheating mechanism is introduced—personalized parameters from the user's three most recent interactions are preloaded into the SSD cache, controlling the time spent on a single query to within [a certain threshold]. If the target parameter is not in the cache, it is loaded from NAS storage and cached. The validity of the parameter match is verified twice using fingerprint similarity. ; ( (for the stored fingerprint vector), when When a match is successful, output Otherwise, trigger the re-collection process.

[0104] The tempo baseline is updated using an exponential moving average algorithm. Defined as the average rhythmic characteristics of users in a specific segmented scenario (based on S2's three-dimensional rhythmic characteristics). The time mean (mean value) represents the user's basic action habits (e.g., users with slower actions have a lower baseline). Initial baseline This serves as the standard rhythm baseline for the current scene (from the S4 scene semantic library), and subsequent rhythm features are modeled based on user feedback each time. Incremental update, the formula is: ; in To update the iteration count, A preset smoothing coefficient is used to balance the stability of historical baselines with the adaptability of current actions. This represents the average tempo of the current action. (Updated) It will serve as the basis for adjusting the attention weight in S3 and fine-tuning the meta-parameters in S4, adapting to the gradual changes in user behavior habits.

[0105] Introducing historical loss constraints to update attention weights and meta-parameters is primarily aimed at preventing parameters from deviating excessively from historical adaptation experience. The total loss function is defined as a weighted sum of the current loss and historical losses: ; in For the current personalized semantic matching loss ( Personalized semantic matching similarity from S3). For the most recent The historical loss mean of the second model ( (For preset history window) Historical loss weights (balancing current adaptation with historical experience).

[0106] based on The parameters are updated using the incremental gradient descent algorithm, as shown in the formula: ; in, The weight matrix corresponding to the user attention calculation module. The bias parameter is used for incremental adjustment of user-side action features. The superscript (k) indicates the parameter state during the k-th iteration. The preset incremental step size (learning rate, determined through parameter sensitivity experiments) is used. The gradient of the loss function with respect to the parameters is used; to avoid gradient explosion, a gradient clipping mechanism is introduced. This ensures stable parameter updates.

[0107] The end-to-end parameter feedback pushes the updated personalized parameters to the feature layer, attention layer, and meta-parameter layer, forming a closed-loop optimization: 1) Feedback to the feature layer: The feature projection matrix (such as S2's) in the personalized meta-parameters is used to... Replacing the universal projection matrix makes the dual-domain temporal tensor of S1 and the four-dimensional fused feature of S2 more closely resemble the user's hand features. The feedback formula is as follows: 2) Feedback to the attention layer: The updated... 3) Replace the attention weight matrix of S3 to improve targeting of user-specific semantic transition frames and action segments; 4) Feedback to the meta-parameter layer: The system writes back to the five-level storage architecture and simultaneously updates the incremental meta-parameter system of S4 to support subsequent cross-scenario adaptation and reuse. After this feedback, the system performs the next round of modeling based on the updated parameters, ensuring continuous optimization of personalized adaptation accuracy. (Average increase of 5%~8%).

[0108] The closed-loop verification process of end-to-end feedback is as follows: After the parameter feedback is completed, the system automatically calls the standard sign language action samples of the current scene (from the scene semantic library), re-executes the modeling process of S1-S3, and calculates the personalized semantic matching similarity after feedback. Compare the increase in similarity before and after the feedback. ( (for the matching similarity before feedback), if (If the preset minimum improvement threshold is met, the feedback is deemed effective, and the updated parameters are retained; if...) If this happens, the parameter rollback mechanism is triggered, restoring the parameter state to its state before the feedback, and adjusting the incremental step size of S5. (Adjusted to the original) (0.8 times that of the previous method), and re-execute the parameter update process. The verification results are synchronously written to the parameter type layer of the five-level storage architecture to optimize subsequent parameter update strategies.

[0109] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, because according to the embodiments of this application, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily essential to this application.

[0110] By employing FPGA-synchronized dual-domain data acquisition and fusion, the temporal consistency and feature integrity of sign language action data are first ensured, laying a high-quality data foundation for subsequent semantic modeling. Then, through multi-source feature high-dimensional projection and dynamic gating fusion technology, the intrinsic relationship between actions, rhythm, and scene is established, achieving multi-dimensional complementarity and precise adaptation of semantic information, solving the problems of dimensional differences and semantic bias in heterogeneous feature fusion. Based on rhythm density-based action segment division and three-level attention weighting technology, semantic transition frames and inter-segment relationships can be accurately captured, strengthening the expression of key semantic information and improving the hierarchical distinguishability of semantic features. This is further enhanced by relying on four-level nesting. The meta-parameter system and cross-scenario weighted fusion technology enable hierarchical adaptation of scene semantics, taking into account both general semantic rules and specific scene requirements, and breaking down parameter gaps caused by scene switching. By constructing hand physiological feature fingerprints and a full-link parameter feedback mechanism, a personalized adaptation closed loop is established, enabling the model to dynamically adapt to differences in user action habits. At the same time, historical loss constraints ensure the stability of parameter updates. Ultimately, the entire link of "precise data collection - multi-dimensional semantic fusion - hierarchical feature enhancement - scene adaptation - personalized optimization" is achieved, resulting in the technical effects of accurate modeling of sign language action semantics, scene adaptation, and individual adaptability.

[0111] The above description is merely a preferred embodiment of this application and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of disclosure in this application is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the foregoing disclosed concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features with similar functions disclosed in this application.

Claims

1. A dynamic sign language temporal modeling method based on an attention mechanism, characterized in that, Simultaneously acquire image data and inertial measurement data of dynamic sign language, and generate a dual-domain temporal tensor that adapts to the amplitude and speed characteristics of sign language movements; The dual-domain temporal tensor, the three-dimensional rhythmic features representing semantic transitions in sign language, and the scene prior features adapted to the semantic constraints of sign language scenarios are integrated to form a four-dimensional fusion feature. Based on the aforementioned four-dimensional fusion features, a three-level attention weighting process of frames, segments, and sentences is performed to adapt to the semantic level of sign language, generating a semantically enhanced feature map; Load the three-level nested meta-parameters of major categories, subcategories, and semantic units that adapt to the hierarchical characteristics of sign language scenarios, and perform scenario-based adaptation on the semantic enhancement feature map.

2. The dynamic sign language temporal modeling method based on attention mechanism according to claim 1, characterized in that, The fusion of the dual-domain temporal tensor, three-dimensional rhythmic features, and scene prior features to form a four-dimensional fusion feature includes: The dual-domain temporal tensor, three-dimensional rhythm features, and scene prior features are projected onto a preset high-dimensional space, respectively. By using attention-based association computation, we can uncover the intrinsic connections between sign language movements, rhythms, and scenes, and obtain the association weights between multi-source features. By outputting dynamic coefficients through a dynamic gating adjustment mechanism, the contribution of rhythm features is increased for sign language semantic transition frames, and the contribution of scene features is increased for cross-scene switching. The contribution of the three types of features in different frames or different scenes is adjusted to complete the construction of four-dimensional fusion features.

3. The dynamic sign language temporal modeling method based on attention mechanism according to claim 1, characterized in that, The three-level attention weighting processing based on the four-dimensional fusion features includes: The rhythmic abruptness, representing the semantic shift in sign language, is calculated based on the abruptness parameter in the three-dimensional rhythmic features. Based on the rhythmic abruptness, semantic transition frames such as sign language negation words and sentence pauses are identified. The rhythmic abruptness is used as an adjustment factor for frame-level attention weights to dynamically increase the attention weights of semantic transition frames.

4. The dynamic sign language temporal modeling method based on attention mechanism according to claim 1, characterized in that, The three-level attention weighting processing based on the four-dimensional fusion features includes: Calculate the rhythm matching coefficient between the current action segment and the standard sign language rhythm of the corresponding scene to ensure the standardization of sign language actions; Emotional rhythm coefficients are calculated based on statistical parameters of the rhythmic trend of movement segments to capture the emotional semantics of sign language. Calculate the scene adaptation coefficient between action segment features and scene-specific semantic unit features to constrain sign language contextual semantics; Dynamic weights are output through a gating dynamic fusion mechanism to adjust the contribution of the rhythm matching coefficient, emotional rhythm coefficient, and scene adaptation coefficient, thereby generating segment-level attention weights.

5. The dynamic sign language temporal modeling method based on attention mechanism according to claim 1, characterized in that, Before performing the frame, segment, and sentence-level attention weighting processing, the following is also included: The local rhythm density of each frame is calculated based on the three-dimensional rhythm features to characterize the clustering characteristics of sign language movement rhythms. An adaptive density threshold is set to match the rhythmic variation of sign language, and density peak points are selected as cluster centers for action segments. The action segments are divided by taking the midpoint of the adjacent cluster centers as the boundary of the action segment and combining the continuity constraints of sign language actions.

6. The dynamic sign language temporal modeling method based on attention mechanism according to claim 1, characterized in that, The loading of three levels of nested meta-parameters (category, subcategory, and semantic unit) for contextual adaptation of the semantic enhancement feature map includes: A four-level nested meta-parameter system of global, major category, sub-category, and semantic unit is constructed to adapt to the hierarchical characteristics of sign language scenarios, and a storage structure combining basic parameters and incremental parameters is adopted. Based on a small set of sign language action features, online fine-tuning of meta-parameters is achieved through a parameter evolution mechanism; When crossing scenarios, meta-parameters are adapted to different scenarios by weighted fusion based on the similarity of sign language semantic units.

7. The dynamic sign language temporal modeling method based on attention mechanism according to claim 1, characterized in that, Also includes: Collect static images of the user's hand, extract the physiological features of the hand that are directly related to sign language movements, and generate hand feature fingerprints; A five-level storage architecture is built based on user fingerprints, segmented scenarios, semantic units, and parameter types to store personalized adaptation-related data. Using the hand fingerprint as an index, a fast indexing mechanism is used to match personalized parameters.

8. The dynamic sign language temporal modeling method based on attention mechanism according to claim 7, characterized in that, Also includes: The rhythm baseline is updated using an exponential moving average algorithm to adapt to the slow changes in user behavior habits; Introduce historical loss constraints and update attention weights and meta-parameters based on incremental step size; The updated personalized parameters are fed back to the feature layer, attention layer, and meta-parameter layer respectively, achieving end-to-end parameter feedback.

9. The dynamic sign language temporal modeling method based on attention mechanism according to claim 1, characterized in that, After generating the semantically enhanced feature map, the process also includes: Using key hand gestures as nodes and physical connections between key gestures as edges, a dynamic sign language spatiotemporal graph is constructed by connecting consecutive frames. Spatiotemporal features of sign language are extracted through spatial graph convolution branches and temporal convolution branches; By integrating the semantic enhancement feature map, spatiotemporal features, and rhythm alignment results, personalized semantic matching similarity is calculated.

10. A system for implementing the dynamic sign language temporal modeling method based on an attention mechanism as described in any one of claims 1-9, characterized in that, It includes a data acquisition module, a collaborative computing module, a distributed intelligent storage module, and a control module, with each module achieving real-time linkage through a data bus; The data acquisition module includes a depth camera, an inertial measurement sensor, and an FPGA synchronization control unit. The FPGA synchronization control unit integrates parallel units for three-dimensional rhythm calculation, hand fingerprint extraction, and spatiotemporal graph node preprocessing adapted for sign language data preprocessing. The collaborative computing module includes a GPU, a CPU, and an integrated parameter scheduling center. The GPU is deployed with a spatiotemporal graph convolution and attention fusion parallel computing core adapted for sign language spatiotemporal modeling. The distributed intelligent storage module includes high-speed SSD cache and NAS storage, and constructs an integrated index of user fingerprints, subdivided scenarios, semantic units and parameter types that are adapted to the personalized and scenario-based needs of sign language. The control module performs global collaborative loss function calculation and coordinates with other modules to complete the entire closed-loop control of data acquisition, modeling, matching, and updating.