Text generation method, system, terminal device and storage medium
By collecting gesture data through a 3D depth camera and reconstructing gesture trajectories and analyzing features, and combining it with emotion processing to generate intelligent description text, it solves the problems of insufficient flexibility and personalization in gesture text generation in traditional methods, and achieves more accurate and vivid text generation.
Patent Information
- Application Number
- CN202510084544.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2045-01-20
AI Technical Summary
Traditional gesture-to-text generation methods are difficult to dynamically adjust according to the user's real-time intentions and interaction scenarios, and cannot intuitively establish a mapping relationship between the user's body language and text content, resulting in insufficient flexibility and personalization in expression.
A 3D depth camera is used to capture real-time hand depth images. Through gesture trajectory reconstruction, feature analysis, and intention recognition, combined with emotional adjective processing, intelligent gesture description text is generated and visually rendered.
It can generate more expected text descriptions based on the user's real-time intentions, accurately identify the motion trajectory, action intentions and emotional colors in gestures, generate more vivid and information-rich texts, and support dynamic adjustment and visual feedback.
Smart Images

Figure CN120031012B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of text generation, and in particular to a text generation method, system, terminal device and storage medium. Background Art
[0002] Human-computer interaction methods are becoming increasingly diverse, and traditional input methods such as keyboards and mice are gradually unable to meet people's needs in specific scenarios. In recent years, gesture recognition, as a natural and intuitive interaction method, has received widespread attention and has developed rapidly, showing great application potential in smart homes, virtual reality, augmented reality, medical assistance, and other fields. As a natural and intuitive form of expression, gestures carry rich information and can convey user intentions and emotions. Using computer vision technology to recognize and understand gestures and convert them into corresponding text information can achieve more natural and convenient human-computer interaction. However, traditional gesture text generation methods mostly focus on classifying and recognizing gestures and converting them into simple command text. These methods still have significant limitations in terms of flexibility, dynamism, and personalization of expression. They are difficult to dynamically adjust according to the user's real-time intentions and interaction scenarios, and cannot intuitively map the user's body language to text content. Summary of the Invention
[0003] Based on this, the present invention provides a text generation method, system, terminal device and storage medium to solve at least one of the above technical problems.
[0004] To achieve the above object, a text generation method includes the following steps:
[0005] Step S1: using a 3D depth camera to collect real-time hand depth images of the user to obtain depth hand image sequence data; reconstructing gesture trajectories based on the depth hand image sequence data to obtain spatial gesture motion trajectory data;
[0006] Step S2: performing gesture motion feature analysis on the spatial gesture motion trajectory data to generate gesture motion feature vector data; performing basic gesture action unit segmentation based on the gesture motion feature vector data to generate basic gesture action unit data; annotating the gesture motion feature vector data with gesture motion semantic labels based on the basic gesture action unit data to generate gesture motion text description data; constructing a gesture motion intention recognition model; using the gesture motion intention recognition model to perform action intention recognition on the gesture motion feature vector data to obtain gesture action intention recognition data;
[0007] Step S3: Processing the action emotion adjectives based on the gesture action intention recognition data, and performing gesture semantic fusion based on the gesture motion text description data to generate key gesture semantic fusion data; performing intelligent text generation based on the key gesture semantic fusion data, and performing text template filling to obtain intelligent description gesture text data;
[0008] Step S4: Perform visual text rendering processing based on the intelligent description gesture text data to achieve the generation of intelligent description text for the gesture action.
[0009] This invention utilizes a 3D depth camera to capture deep hand image sequences and reconstruct gesture trajectories, accurately capturing hand motion trajectories in three-dimensional space. This overcomes the shortcomings of traditional gesture recognition methods based on two-dimensional images, which are sensitive to gesture posture, viewing angle, and lighting variations, resulting in more stable and accurate gesture recognition. By performing feature analysis on spatial gesture motion trajectories and segmenting basic gesture action units, more fine-grained gesture motion features can be extracted. Based on gesture motion feature vector data, the system identifies the user's action intent, such as "point," "grab," or "put down," thereby understanding the core meaning the user intends to convey. This enables the system to dynamically adjust based on the user's real-time intent, generating text descriptions that better align with user expectations, rather than simply providing instructions. Furthermore, the system combines action intent recognition results with emotional adjective processing to identify the emotional connotations embedded in the user's gestures, such as "quickly," "slowly," and "forcefully," making the generated text descriptions more vivid and expressive. By semantically fusing gesture motion text description data with gesture action intent recognition data, key gesture semantic fusion data is generated, achieving a comprehensive understanding and description of gesture actions. This is not just about converting gestures into simple text labels, but about integrating the gesture's motion trajectory, action intention, and emotional expression to generate a more informative and expressive text description. Intelligent text generation and text template filling based on key gesture semantic fusion data can generate more complete, fluent, and natural text descriptions, such as "He quickly pointed to a red object in the distance" rather than simply "pointing to red." By visually rendering the intelligent description gesture text data, the text can be associated with the gesture action, such as displaying the text description of the corresponding gesture on the screen in real time, and dynamically adjusting the display position and method of the text according to the gesture's motion trajectory. This visual feedback mechanism enables users to understand the system's recognition results more clearly and facilitates more effective interaction between users and the system.
[0010] Preferably, step S1 includes the following steps:
[0011] Step S11: using a 3D depth camera to collect real-time hand depth images of the user, and perform initial image correction to obtain depth hand image sequence data;
[0012] Step S12: performing background removal on the depth hand image sequence data and performing hand region segmentation to obtain user hand region data;
[0013] Step S13: performing hand posture estimation on the user's hand area data to obtain the user's hand posture data;
[0014] Step S14: Detecting skeleton key points based on the user's hand posture data, and performing hand skeleton key point coordinate analysis to obtain skeleton key point coordinate data;
[0015] Step S15: extracting pixel coordinates of gesture key points from the depth hand image sequence data using the skeleton key point coordinate data to obtain pixel coordinate data of gesture image key points;
[0016] Step S16: performing timestamp synchronization processing based on the pixel coordinate data of the key points of the gesture image, and reconstructing the gesture trajectory based on the user's hand posture data to obtain spatial gesture motion trajectory data.
[0017] This invention uses a 3D depth camera to capture real-time hand depth images and performs initial image correction, ensuring the accuracy and reliability of the raw data. Compared to traditional two-dimensional images, depth information better reflects the true shape of the hand in three-dimensional space, avoiding errors caused by factors such as perspective changes and lighting conditions. Initial image correction further eliminates the effects of camera distortion and other factors, ensuring the accuracy of subsequent processing. Hand posture estimation, skeletal keypoint detection, and hand skeletal keypoint coordinate analysis provide key information for accurate reconstruction of gesture trajectories. Hand posture estimation determines the hand's orientation and posture in space, providing a reference for subsequent keypoint detection. Skeletal keypoint detection further locates the positions of key hand parts, such as fingertips and knuckles. By analyzing the coordinate data of these keypoints, the hand's shape and motion trajectory can be accurately described. Timestamp synchronization and gesture trajectory reconstruction based on the user's hand posture data ensure the temporal and spatial consistency and accuracy of the gesture trajectory. Timestamp synchronization ensures the temporal consistency of data between each step, allowing the reconstructed gesture trajectory to accurately reflect the motion process of the gesture. By combining hand posture data for trajectory reconstruction, the rotation and displacement of the hand in space are further considered, so that the reconstructed spatial gesture motion trajectory data more accurately describes the actual movement of the gesture in three-dimensional space.
[0018] Preferably, step S2 includes the following steps:
[0019] Step S21: analyzing the spatial gesture motion trajectory data for the length of the gesture spatial trajectory and performing time regularization processing on the sampling trajectory to obtain pre-processed gesture motion trajectory data;
[0020] Step S22: performing gesture motion feature analysis based on the pre-processed gesture motion trajectory data to generate gesture motion feature vector data;
[0021] Step S23: Segmenting the basic gesture action units according to the gesture motion feature vector data to generate basic gesture action unit data;
[0022] Step S24: annotating the gesture motion feature vector data with gesture motion semantic labels using a preset gesture-semantic mapping database based on the basic gesture action unit data to generate gesture motion text description data;
[0023] Step S25: performing gesture movement intention sample transfer learning according to a preset convolutional neural network model, thereby constructing a gesture movement intention recognition model;
[0024] Step S26: using the gesture motion intention recognition model to perform motion intention recognition on the basic gesture action unit data and the gesture motion feature vector data to obtain gesture action intention recognition data.
[0025] The present invention effectively preprocesses the original trajectory data by analyzing the length of gesture spatial trajectories and time-regularizing the sampled trajectories, thereby improving the standardization and usability of the data. When different users perform the same gesture, their speeds and trajectory lengths vary. Through length analysis and time regularization, trajectory data of different lengths and speeds can be unified to the same standard. Gesture motion feature analysis converts the preprocessed trajectory data into gesture motion feature vector data, realizing the transformation from specific trajectories to abstract features, laying the foundation for subsequent semantic understanding and intention recognition. These feature vectors can effectively capture key information of gesture motion, such as motion direction, speed, amplitude, etc. Basic gesture action unit segmentation divides the continuous gesture motion trajectory into a series of semantic action units, realizing the transformation from continuous motion to discrete semantic units. This segmentation enables the system to identify the basic action elements contained in complex gestures. For example, "waving" can be divided into two action units: "lifting" and "swinging", providing more fine-grained support for understanding the meaning of gestures.
[0026] Preferably, step S22 includes the following steps:
[0027] Step S221: extracting the coordinates of the starting point and ending point of the gesture trajectory based on the pre-processed gesture motion trajectory data to obtain the trajectory starting and ending point data;
[0028] Step S222: Calculating the displacement vectors of the hand key points in adjacent frames based on the pre-processed gesture motion trajectory data to generate inter-frame displacement vector data of the hand;
[0029] Step S223: performing inter-frame instantaneous velocity vector calculation on the inter-frame displacement vector data of the hand to generate instantaneous velocity vector data;
[0030] Step S224: performing low-pass filtering on the instantaneous velocity vector data and calculating the gesture motion acceleration to obtain frame-by-frame acceleration vector data;
[0031] Step S225: performing mean processing on the instantaneous velocity vector data and the frame-by-frame acceleration vector data through a preset time window to obtain velocity and acceleration feature data;
[0032] Step S226: performing trajectory curve fitting on the pre-processed gesture motion trajectory data using the trajectory start and end point data, and performing gesture trajectory curvature and direction change analysis to obtain gesture trajectory geometric feature data;
[0033] Step S227: performing gesture posture change frequency analysis based on the pre-processed gesture motion trajectory data to obtain gesture posture change frequency data;
[0034] Step S228: performing gesture feature combination on the velocity acceleration feature data, the gesture trajectory geometric feature data, and the gesture posture change frequency data to obtain gesture motion feature vector data.
[0035] The present invention describes the dynamic characteristics of gesture motion in detail by calculating the displacement vectors of hand key points in adjacent frames, the instantaneous velocity vectors between frames, and the frame-by-frame acceleration vectors. The displacement vector reflects the direction and distance of the gesture's movement in a short period of time, the instantaneous velocity vector further reveals the speed change of the movement, and the acceleration vector reflects the rate of speed change. These dynamic features are crucial for distinguishing different types of gestures. For example, fast waving and slow movement have significantly different speed and acceleration characteristics. The geometric features of the gesture trajectory are extracted through trajectory curve fitting, gesture trajectory curvature, and direction change analysis. These geometric features reflect the shape and direction change of gesture motion. For example, circular gestures and straight gestures have different curvature characteristics, and waving to the left and waving to the right have different direction change characteristics.
[0036] Preferably, step S24 includes the following steps:
[0037] Step S231: performing key frame detection on the gesture motion feature vector data to obtain key frame data;
[0038] Step S232: segmenting the pre-processed gesture motion trajectory data into key frame time periods using the key frame data to obtain time period segmented trajectory data;
[0039] Step S233: performing hand gesture cluster analysis on the time segmented trajectory data to obtain hand gesture cluster data;
[0040] Step S234: dividing the time segmentation trajectory data into basic gesture action units based on the gesture posture clustering data to generate basic gesture action unit division data;
[0041] Step S235: performing conversion relationship processing between gesture action units on the basic gesture action unit division data to obtain unit conversion relationship data;
[0042] Step S236: Integrate the basic gesture action unit division data and the unit conversion relationship data to generate basic gesture action unit data.
[0043] The present invention clusters and analyzes hand gestures, classifies similar hand gestures into one category, and divides the time period segmentation trajectory data into basic hand gesture action units based on the hand gesture clustering data, so that the segmented action units have clear semantic boundaries. For example, "lift", "put down", "wave to the left", etc. can all be regarded as a basic action unit. This division method based on hand gestures makes the division of action units more consistent with human perception and understanding of gestures. The conversion relationship between gesture action units is processed on the basic hand gesture action unit division data, and the temporal relationship and logical relationship between action units are clarified, providing key information for understanding the overall meaning of complex gestures. For example, "lift" is usually followed by "put down" or "wave". By analyzing the conversion relationship between action units, the user's intention can be inferred. For example, continuous "click" actions mean "confirmation" or "selection".
[0044] Preferably, step S24 includes the following steps:
[0045] Step S241: determining the gesture motion intensity level based on the velocity acceleration feature data to generate motion intensity level data;
[0046] Step S242: mapping the exercise intensity level data to exercise semantic text labels using a preset gesture-semantic mapping database to obtain intense exercise semantic vocabulary data;
[0047] Step S243: performing gesture complexity level analysis on the gesture trajectory geometric feature data to obtain complexity level data;
[0048] Step S244: using a preset gesture-semantic mapping database to perform motion semantic text label mapping on the complexity level data to obtain complex motion semantic vocabulary data;
[0049] Step S245: classifying the gesture rhythm speed level of the gesture posture change frequency data to obtain gesture rhythm level data;
[0050] Step S246: mapping the gesture rhythm level data to gesture rhythm text labels using a preset gesture-semantic mapping database to obtain gesture rhythm text vocabulary data;
[0051] Step S247: performing gesture semantic description label fusion on the violent motion semantic vocabulary data, the complex motion semantic vocabulary data and the gesture rhythm text vocabulary data based on the basic gesture action unit data to obtain gesture motion text description data.
[0052] The present invention maps gesture motion feature data to different semantic dimensions by making level judgments on the intensity, complexity and rhythm speed of gesture velocity acceleration feature data, gesture trajectory geometric feature data and gesture posture change frequency data. This multi-dimensional analysis makes the description of gestures more comprehensive and detailed, such as "fast", "slow", "complex", "simple", "rhythmic", etc. Using a preset gesture-semantic mapping database, these level data are mapped to corresponding semantic text labels. This database is the key to achieving text description from feature data. It associates different feature levels with corresponding semantic vocabulary, such as "violent", "mild", "complex", "simple", "fast", "soothing", etc. Through mapping, the system can convert abstract feature data into specific semantic vocabulary.
[0053] Preferably, step S3 includes the following steps:
[0054] Step S31: performing confidence evaluation on the gesture action intention recognition data and screening high-confidence intentions to obtain high-confidence action intention data;
[0055] Step S32: performing action emotion adjective processing based on the high-confidence action intention data, and performing gesture semantic fusion based on the gesture motion text description data to generate key gesture semantic fusion data;
[0056] Step S33: performing intelligent text generation based on the key gesture semantic fusion data to generate initial gesture text content data;
[0057] Step S34: performing gesture category analysis based on the basic gesture action unit data, and performing text template keyword processing to generate text template keyword data;
[0058] Step S35: using the text template keyword data to perform text template matching on a preset text template library to obtain selected text template data;
[0059] Step S36: performing placeholder semantic filling on the initial gesture text content data by selecting text template data, thereby obtaining intelligent description gesture text data.
[0060] The present invention performs a confidence assessment on gesture action intention recognition data and selects high-confidence intention data, ensuring that subsequent text generation is based on reliable intent understanding. This step, by setting a confidence threshold, eliminates erroneous recognition results, improves the accuracy of intent understanding, and avoids generating text descriptions that do not align with user intent. High-confidence action intention data is processed using action emotion adjectives and semantically fused with the gesture motion text description data to generate key gesture semantic fusion data. This step further enriches the semantic information of the text description, particularly by incorporating emotional overtones. This ensures that the generated text is not only an objective description of the action but also reflects the user's emotional state, such as "waving joyfully" or "putting down hands in frustration." This inclusion of emotional information makes the text description more expressive and humane. Gesture category analysis is performed on basic gesture action unit data, and text template keyword processing is performed to generate text template keyword data, providing a basis for text template selection. Different gesture categories, such as interaction, instruction, and expression, typically correspond to different text templates. Keyword matching allows for rapid identification of appropriate text templates. Combining the generated text content with the template ensures that the final text output has both personalized content and good readability and fluency.
[0061] Preferably, the present invention further provides a text generation system for executing the above-mentioned text generation method, wherein the text generation system comprises:
[0062] The gesture trajectory reconstruction module is used to use a 3D depth camera to collect real-time hand depth images of the user to obtain depth hand image sequence data; based on the depth hand image sequence data, the gesture trajectory is reconstructed to obtain spatial gesture motion trajectory data;
[0063] The gesture semantic annotation module is used to analyze the gesture motion characteristics of spatial gesture motion trajectory data to generate gesture motion feature vector data; segment the gesture motion feature vector data into basic gesture action units to generate basic gesture action unit data; annotate the gesture motion feature vector data with gesture motion semantic labels based on the basic gesture action unit data to generate gesture motion text description data; construct a gesture motion intention recognition model; and use the gesture motion intention recognition model to perform action intention recognition on the gesture motion feature vector data to obtain gesture action intention recognition data.
[0064] The intelligent gesture description text generation module is used to process action emotion adjectives based on gesture action intention recognition data, and perform gesture semantic fusion based on gesture motion text description data to generate key gesture semantic fusion data; based on the key gesture semantic fusion data, intelligent text generation is performed and text template semantic filling is performed to obtain intelligent gesture description text data;
[0065] The description text visualization module is used to perform visual text rendering processing based on the intelligent description gesture text data to realize the generation of intelligent description text for gesture actions.
[0066] Preferably, the present invention further provides a terminal device, comprising:
[0067] processor;
[0068] a memory for storing processor-executable instructions;
[0069] The processor is configured to implement any one of the above text generation methods.
[0070] Preferably, the present invention further provides a computer-readable storage medium storing a computer program, wherein the computer program implements any one of the above text generation methods when executed. BRIEF DESCRIPTION OF THE DRAWINGS
[0071] Figure 1 Schematic diagram of the steps of the text generation method of the present invention;
[0072] Figure 2 for Figure 1 Detailed implementation steps of step S1 in FIG.
[0073] Figure 3 for Figure 1 Detailed implementation steps of step S3 in FIG.
[0074] The purpose, features and advantages of the present invention will be further described with reference to the accompanying drawings and in conjunction with the embodiments. DETAILED DESCRIPTION
[0075] The following is a clear and complete description of the technical method of the present invention in conjunction with the accompanying drawings. It is obvious that the embodiments described are part of the embodiments of the present invention, but not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making any creative efforts are within the scope of protection of the present invention.
[0076] In addition, the accompanying drawings are merely schematic illustrations of the present invention and are not necessarily drawn to scale. Identical reference numerals in the figures denote identical or similar parts, and thus repetitive descriptions thereof will be omitted. Some of the block diagrams shown in the accompanying drawings are functional entities that do not necessarily correspond to physically or logically separate entities. These functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor and / or microcontroller approaches.
[0077] It should be understood that although the terms "first," "second," and the like may be used herein to describe various elements, these elements should not be limited by these terms. These terms are used solely to distinguish one element from another. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element, without departing from the scope of the exemplary embodiments. The term "and / or" as used herein includes any and all combinations of one or more of the listed associated items.
[0078] To achieve this, please refer to Figures 1 to 3 The present invention provides a text generation method, comprising the following steps:
[0079] Step S1: using a 3D depth camera to collect real-time hand depth images of the user to obtain depth hand image sequence data; reconstructing gesture trajectories based on the depth hand image sequence data to obtain spatial gesture motion trajectory data;
[0080] Step S2: performing gesture motion feature analysis on the spatial gesture motion trajectory data to generate gesture motion feature vector data; performing basic gesture action unit segmentation based on the gesture motion feature vector data to generate basic gesture action unit data; annotating the gesture motion feature vector data with gesture motion semantic labels based on the basic gesture action unit data to generate gesture motion text description data; constructing a gesture motion intention recognition model; using the gesture motion intention recognition model to perform action intention recognition on the gesture motion feature vector data to obtain gesture action intention recognition data;
[0081] Step S3: Processing the action emotion adjectives based on the gesture action intention recognition data, and performing gesture semantic fusion based on the gesture motion text description data to generate key gesture semantic fusion data; performing intelligent text generation based on the key gesture semantic fusion data, and performing text template filling to obtain intelligent description gesture text data;
[0082] Step S4: Perform visual text rendering processing based on the intelligent description gesture text data to achieve the generation of intelligent description text for the gesture action.
[0083] In an embodiment of the present invention, the text generation method includes the following steps:
[0084] Step S1: using a 3D depth camera to collect real-time hand depth images of the user to obtain depth hand image sequence data; reconstructing gesture trajectories based on the depth hand image sequence data to obtain spatial gesture motion trajectory data;
[0085] In an embodiment of the present invention, a 3D depth camera, such as one using a structured light or time-of-flight device, is activated to scan the user's hand in real time. The camera captures hand images containing depth information at a rate of 30 frames per second, generating a continuous sequence of depth images. Each depth image is a two-dimensional matrix, with each element representing the distance from the corresponding pixel to the camera. For example, a pixel value of 50 indicates that the pixel is 50 centimeters away from the camera. The hand region in the depth image is first identified using image processing techniques. Common methods include threshold segmentation and morphological processing. For example, a depth threshold is set to initially identify regions within a distance of less than 80 centimeters and greater than 20 centimeters from the camera as the hand region. Morphological operations such as erosion and dilation are then used to remove noise and holes in the hand region. Next, a hand feature point needs to be determined, such as the palm center or fingertip position. In the first frame, the initial palm center can be determined by calculating the geometric center of the hand region. In each subsequent frame, tracking algorithms such as optical flow or Kalman filtering are used to track the positional changes of this feature point between frames. For example, the Lucas-Kanade optical flow algorithm is used to calculate the displacement of feature points between adjacent frames, thereby updating the coordinates of the feature points in the new frame. By tracking the positions of feature points frame by frame, a sequence of points in three-dimensional space is obtained. This point sequence is the gesture motion trajectory data, represented as a set of three-dimensional coordinate points. For example: [(x1, y1, z1), (x2, y2, z2), ..., (xn, yn, zn)], where (xi, yi, zi) represents the three-dimensional coordinates of the hand feature point in the i-th frame.
[0086] Step S2: performing gesture motion feature analysis on the spatial gesture motion trajectory data to generate gesture motion feature vector data; performing basic gesture action unit segmentation based on the gesture motion feature vector data to generate basic gesture action unit data; annotating the gesture motion feature vector data with gesture motion semantic labels based on the basic gesture action unit data to generate gesture motion text description data; constructing a gesture motion intention recognition model; using the gesture motion intention recognition model to perform action intention recognition on the gesture motion feature vector data to obtain gesture action intention recognition data;
[0087] In embodiments of the present invention, multiple features are extracted from trajectory data, including velocity, direction, acceleration, and curvature. For example, for each point in the trajectory data, the distance from the previous point is calculated and divided by the time interval between two frames to obtain the instantaneous velocity of that point. The angle between the velocity vectors of three adjacent points is then calculated to obtain the curvature of the trajectory at that point. Performing these operations for each point on the trajectory yields a velocity sequence and a curvature sequence. Similarly, features such as acceleration and directional change can also be calculated. These feature sequences are then statistically analyzed, for example, by calculating the mean, variance, maximum, minimum, and rate of change of these values. These statistical values are combined into a vector, which represents the gesture motion feature vector data. Continuous gesture motion is decomposed into a series of discrete, semantically meaningful basic action units. For example, a complete "circle" gesture can be segmented into four basic action units: "upward motion," "rightward motion," "downward motion," and "leftward motion." Segmentation can be based on the changing trends of the individual features in the feature vector, such as a sudden increase or decrease in velocity or a significant change in direction. Rules or thresholds can be set. When feature values exceed or fall below these thresholds, one basic action unit is considered to end and another basic action unit begins. For example, segmentation occurs when the rate of change of velocity exceeds a certain threshold or when the angle of change in direction is greater than 45 degrees. Each segmented basic action unit corresponds to a feature vector segment, and the collection of these segments constitutes the basic gesture action unit data. Based on this basic gesture action unit data, the original gesture motion feature vector data is annotated, using text labels to describe the semantic meaning of each basic action unit. For example, "upward movement," "rightward movement," "circle," "wave," etc. These text labels are associated with the corresponding feature vector segments to form gesture motion text description data. A gesture motion intent recognition model is constructed. A classification model is trained using deep learning methods, such as recurrent neural networks (RNNs) or long short-term memory networks (LSTMs). The annotated gesture motion text description data serves as the training set. For example, hundreds of sample data of different types of gestures are prepared, and each sample is annotated with its corresponding intent, such as "select," "confirm," "cancel," etc. Using the trained gesture motion intent recognition model, we perform intent recognition on new, unlabeled gesture motion feature vector data. The feature vector of the new data is input into the model, which outputs an intent label, such as "confirm," "zoom in," or "zoom out." This output is the gesture action intent recognition data.
[0088] Step S3: Processing the action emotion adjectives based on the gesture action intention recognition data, and performing gesture semantic fusion based on the gesture motion text description data to generate key gesture semantic fusion data; performing intelligent text generation based on the key gesture semantic fusion data, and performing text template filling to obtain intelligent description gesture text data;
[0089] In an embodiment of the present invention, a library of emotional adjectives is predefined, such as "quickly", "slowly", "forcefully", "gently", "smoothly", "augmentally", etc. According to the speed, acceleration and other features in the gesture motion feature vector data, appropriate emotional adjectives are selected. For example, if the average speed of a "waving" action is high and the acceleration is large, adjectives such as "quickly" and "forcefully" are selected to modify it. If the curvature of the trajectory of a "drawing a circle" action changes smoothly, the adjective "smoothly" is selected. The selected emotional adjectives are combined with the gesture action intention recognition data, for example, "select" is combined with "quickly" to obtain "quickly select". For example, for a gesture that is divided into two basic action units of "upward movement" and "rightward movement", the corresponding text description data is "up" and "right". If the intention recognition data is "draw" and the emotional adjective processing result is "slowly", then the result of semantic fusion is "slowly draw upwards, then draw right". Natural language generation techniques, such as rule-based template filling or neural network-based sequence-to-sequence models, take key semantic descriptions as input and generate more complete sentences. For example, for the key semantic description "slowly drawing upwards, then drawing to the right," the generated sentence could be "The user is slowly moving their hand upwards to draw, then moving their hand to the right to continue drawing." If template filling is used, predefined sentence templates can be used, such as "The user is [emotional adjective] [action intention] [basic action unit description]." The corresponding parts of the key semantic fusion data are then filled in the corresponding positions in the template. For example, connecting words and background information can be added to make the generated text smoother and easier to understand. For example, the text generated in the previous step can be supplemented with "This operation completes the drawing of the shape," resulting in the final intelligent gesture description text data: "The user is slowly moving their hand upwards to draw, then moving their hand to the right to continue drawing, and this operation completes the drawing of the shape."
[0090] Step S4: Perform visual text rendering processing based on the intelligent description gesture text data to achieve the generation of intelligent description text for the gesture action.
[0091] In an embodiment of the present invention, a suitable font and font size are selected. For example, the clear and easy-to-read Microsoft YaHei font is selected and the font size is set to 18. Then, the display position of the text on the screen is determined. For example, the text can be displayed in the lower center of the screen, 5 cm from the bottom of the screen, and horizontally centered. Next, the color and background color of the text are set. For example, white text and a black semi-transparent background are used to ensure that the text is clearly visible against various backgrounds. Automatic line wrapping is performed based on the length of the text and the width of the screen. For example, a maximum of 30 characters can be displayed per line, and if there are more than 30 characters, the line automatically wraps to the next line. Then, some animation effects can be added, such as fade in and fade out, word-by-word display, etc., to make the text presentation more vivid. For example, the text can be set to display word by word at a speed of 5 words per second and remain displayed after the display is completed. The processed text is rendered to the screen. For example, a graphics library such as OpenGL or DirectX can be used to render the text into a texture, then the texture can be applied to a rectangular area, and finally the rectangular area can be drawn to a specified position on the screen. Through the above steps, the intelligent description text data of gestures is presented in a user-friendly manner, and the generation of intelligent description text of gesture actions is completed.
[0092] Preferably, step S1 includes the following steps:
[0093] Step S11: using a 3D depth camera to collect real-time hand depth images of the user, and perform initial image correction to obtain depth hand image sequence data;
[0094] Step S12: performing background removal on the depth hand image sequence data and performing hand region segmentation to obtain user hand region data;
[0095] Step S13: performing hand posture estimation on the user's hand area data to obtain the user's hand posture data;
[0096] Step S14: Detecting skeleton key points based on the user's hand posture data, and performing hand skeleton key point coordinate analysis to obtain skeleton key point coordinate data;
[0097] Step S15: extracting pixel coordinates of gesture key points from the depth hand image sequence data using the skeleton key point coordinate data to obtain pixel coordinate data of gesture image key points;
[0098] Step S16: performing timestamp synchronization processing based on the pixel coordinate data of the key points of the gesture image, and reconstructing the gesture trajectory based on the user's hand posture data to obtain spatial gesture motion trajectory data.
[0099] As an example of the present invention, refer to Figure 2 As shown, Figure 1Detailed implementation steps of step S1 are shown in the flowchart. In this example, step S1 includes:
[0100] Step S11: using a 3D depth camera to collect real-time hand depth images of the user, and perform initial image correction to obtain depth hand image sequence data;
[0101] In an embodiment of the present invention, a 3D depth camera, such as a depth sensor using the time-of-flight principle, is activated. Its transmitter emits infrared light pulses of a specific wavelength toward the scene, and its receiver measures the arrival time of the light pulses reflected at each pixel. By calculating the round-trip time of the light pulses and combining it with the speed of light, the distance from each pixel to the camera can be determined, thereby constructing a depth image. The camera continuously captures depth images at a rate of 60 frames per second. Each frame is a matrix with a resolution of 640x480 pixels, where each element in the matrix represents the distance of that pixel from the camera in millimeters. For example, a pixel value of 1500 indicates that the pixel is 1.5 meters away from the camera. The captured data constitutes a depth image sequence. For example, the coordinates of each pixel are corrected using pre-calibrated camera intrinsic parameters and distortion coefficients. The calibration process can be performed by photographing a calibration plate of known shape and size, such as a checkerboard. Assume that the camera's focal length, as determined through calibration, is 500 pixels, and the distortion coefficients are k1 = 0.1 and k2 = -0.2. For a pixel with coordinates (u, v) in the depth image, first convert it to normalized coordinates (x, y): x = (u-cx) / fx, y = (v-cy) / fy, where cx and cy are the image center coordinates, and fx and fy are the focal lengths. Then, use the distortion model x_corrected = x(1+k1×r^2+k2×r^4) and y_corrected = y(1+k1×r^2+k2×r^4) to calculate the corrected normalized coordinates, where r^2 = x^2+y^2. Finally, convert the corrected normalized coordinates back to pixel coordinates, obtaining the corrected pixel coordinates (u_corrected, v_corrected).
[0102] Step S12: performing background removal on the depth hand image sequence data and performing hand region segmentation to obtain user hand region data;
[0103] In an embodiment of the present invention, for example, assuming that the user typically performs gestures within a range of 0.5 to 1.5 meters from the camera, pixels with depth values less than 500 mm or greater than 1500 mm can be considered background pixels. In each depth image frame, all pixels are traversed and the depth values of pixels whose depth values are not within this range are set to 0, indicating that these pixels are background pixels. This results in a depth image after removing the background. The depth image after removing the background is binarized. For example, the values of pixels with depth values greater than 0 are set to 255, indicating that these pixels are foreground pixels, i.e., the hand area, and the values of the remaining pixels remain at 0, indicating the background. Then, a connected region is searched in the binary image. For example, using the eight-connected algorithm, starting from any foreground pixel, all foreground pixels in the eight surrounding neighborhoods are marked as the same category, and then the same operation is recursively performed on these newly marked pixels until no new pixels can be marked. In this way, a connected region is found. The above steps are repeated until all foreground pixels are marked. Generally, the connected region containing the hand is the largest. Therefore, the connected region with the largest area is selected as the hand region.
[0104] Step S13: performing hand posture estimation on the user's hand area data to obtain the user's hand posture data;
[0105] In an embodiment of the present invention, a hand model is defined. For example, a simplified hand model can be used, dividing the hand into a palm and five fingers, with each finger connected by three joints. This model can be represented by a set of parameters, such as the three-dimensional coordinates of the palm center, the palm normal vector, the length of each finger, and the joint angles. Based on the shape and depth information of the hand region, the initial values of the hand model parameters are preliminarily estimated. For example, the center of mass of the hand region can be calculated as the initial position of the palm center, and the direction of the principal axis of the minimum bounding rectangle of the hand region can be calculated as the initial direction of the palm normal vector. The hand model is then rendered into a depth image, referred to as a model depth map, based on the current parameters. The distance between the model depth map and the depth map corresponding to the user's hand region data is calculated. For example, for each pixel in the model depth map, the nearest point in the user's hand region depth map is searched and the distance between these two points is calculated. The average of these distances is taken as the matching error under the current parameters. Next, the hand model parameters are adjusted to minimize the matching error. For example, the gradient descent method can be used to update the parameters along the negative gradient of the error function. Repeat the above steps until the matching error converges to a smaller value or reaches the preset maximum number of iterations. The final hand model parameters are the estimated user hand posture data.
[0106] Step S14: Detecting skeleton key points based on the user's hand posture data, and performing hand skeleton key point coordinate analysis to obtain skeleton key point coordinate data;
[0107] In an embodiment of the present invention, for example, the hand model used in the previous steps is used. This model contains 16 joint points, corresponding to the wrist, palm, four fingertips, and three joints of each finger. Based on the hand posture data, the coordinates of each joint point in three-dimensional space can be calculated. For example, given the coordinates of the palm center, the palm normal vector, and the length and joint angle of each finger, the position of each joint point relative to the palm center can be determined through simple geometric calculations. These relative positions are then added to the coordinates of the palm center to obtain the three-dimensional coordinates of each joint point in the coordinate system. The joint angle is obtained by calculating the vector between two adjacent joint points and then calculating the angle between the two vectors. For example, if the coordinates of the three joint points of a finger are A(150, 200, 900), B(160, 210, 850), and C(170, 230, 800), then vector AB = BA = (10, 10, -50), and vector BC = CB = (10, 20, -50). The angle between these two vectors can be calculated using the inverse cosine function, for example, arccos((AB·BC) / (|AB|×|BC|)), where · represents the dot product of the vectors and |AB| represents the modulus of vector AB. This method can be used to calculate the joint angles of all fingers. It is also possible to calculate the distance between the fingertips, such as the Euclidean distance between the index and middle fingertips, to form the coordinate data of the skeletal key points.
[0108] Step S15: extracting pixel coordinates of gesture key points from the depth hand image sequence data using the skeleton key point coordinate data to obtain pixel coordinate data of gesture image key points;
[0109] In an embodiment of the present invention, the three-dimensional coordinates of the skeleton key points are converted to the camera coordinate system. Since the previously obtained coordinates of the skeleton key points are expressed in the world coordinate system, and the depth image is expressed in the camera coordinate system, a coordinate transformation is required. For example, the world coordinates P_world of the skeleton key points can be converted to the coordinates P_camera in the camera coordinate system by using the pre-calibrated camera external parameters, namely the rotation matrix R and translation vector T of the camera in the world coordinate system: P_camera = R × P_world + T. Projection is performed using a pinhole camera model. Assume that the intrinsic parameters of the camera are known, including the focal length fx, fy and the principal point coordinates cx, cy. For a three-dimensional point P_camera = (X, Y, Z) in the camera coordinate system, its projection coordinates (u, v) on the image plane can be calculated by the following formula: u = fx × X / Z + cx, v = fy × Y / Z + cy. For each skeleton key point, the above-mentioned coordinate transformation and projection operations are performed in each frame of the depth image to obtain the pixel coordinates of each skeleton key point in each frame of the image. For example, for a certain fingertip key point, its projection coordinates in the first frame image are (300, 200), and its projection coordinates in the second frame image are (305, 205), and so on. These pixel coordinates constitute the pixel coordinate data of the gesture image key point.
[0110] Step S16: performing timestamp synchronization processing based on the pixel coordinate data of the key points of the gesture image, and reconstructing the gesture trajectory based on the user's hand posture data to obtain spatial gesture motion trajectory data.
[0111] In an embodiment of the present invention, the timestamps of all data frames are unified to a reference time, and interpolation or extrapolation is performed based on the time difference. Then, the gesture trajectory is reconstructed based on the pixel coordinate data of the key points of the gesture image after timestamp synchronization and the user hand posture data obtained in step S13. The specific method is to convert the pixel coordinates of the hand key points in each frame image into three-dimensional space coordinates, and connect the key point coordinates of adjacent frames to form a gesture motion trajectory. For example, by connecting the three-dimensional coordinates of the index fingertip in several consecutive frames of images, the motion trajectory of the index fingertip can be obtained. Finally, the spatial gesture motion trajectory data is obtained, which contains the coordinate sequence of the hand key points in three-dimensional space and is stored with time as the index. For example, a gesture trajectory data containing 10 frames can be represented as a 10×21×3 tensor, where 10 represents the number of frames, 21 represents the number of key points, and 3 represents the three-dimensional coordinates.
[0112] Preferably, step S2 includes the following steps:
[0113] Step S21: analyzing the spatial gesture motion trajectory data for the length of the gesture spatial trajectory and performing time regularization processing on the sampling trajectory to obtain pre-processed gesture motion trajectory data;
[0114] Step S22: performing gesture motion feature analysis based on the pre-processed gesture motion trajectory data to generate gesture motion feature vector data;
[0115] Step S23: Segmenting the basic gesture action units according to the gesture motion feature vector data to generate basic gesture action unit data;
[0116] Step S24: annotating the gesture motion feature vector data with gesture motion semantic labels using a preset gesture-semantic mapping database based on the basic gesture action unit data to generate gesture motion text description data;
[0117] Step S25: performing gesture movement intention sample transfer learning according to a preset convolutional neural network model, thereby constructing a gesture movement intention recognition model;
[0118] Step S26: using the gesture motion intention recognition model to perform motion intention recognition on the basic gesture action unit data and the gesture motion feature vector data to obtain gesture action intention recognition data.
[0119] In an embodiment of the present invention, the spatial gesture motion trajectory data obtained in step S1 is subjected to trajectory length analysis. The length of each gesture trajectory is calculated, i.e., the sum of the distances between all adjacent points in the trajectory. For example, a gesture trajectory containing 10 three-dimensional coordinate points has a trajectory length equal to the sum of the lengths of 9 line segments. Then, to eliminate the effect of varying trajectory lengths caused by speed differences when different users or the same user perform the same gesture at different times, time warping of the sampled trajectories is performed, also known as dynamic time warping (DTW). The DTW algorithm can align trajectories of different lengths and find the best matching path between them. Specifically, a distance matrix is constructed, whose elements are the Euclidean distances between corresponding points in the two trajectories. Then, a dynamic programming algorithm is used to find a path from the upper left corner of the matrix to the lower right corner that minimizes the cumulative distance and satisfies the monotonicity and continuity constraints. Using this path, the two trajectories can be aligned to the same length. For example, a 10-frame trajectory and a 12-frame trajectory are aligned to 11 frames, and the deformation between the corresponding frames is calculated. The motion speed of each trajectory point is calculated. For example, for preprocessed trajectory data P = {(x1, y1, z1), (x2, y2, z2), ..., (xN, yN, zN)}, each point corresponds to a timestamp t. Assuming the timestamps are uniform, that is, the time interval between adjacent points is the same, the velocity of each point can be calculated: v(i) = sqrt((xi-xi-1)^2+(yi-yi-1)^2+(zi-zi-1)^2) / (ti-ti-1), where i ranges from 2 to N. For example, in the trajectory data of a "waving" action, the speed of some points reaches 10 centimeters per second, while in the trajectory data of a "slowly drawing a circle" action, the speed of all points does not exceed 2 centimeters per second. The direction of motion of each trajectory point is calculated. For example, for each point i, the displacement vector between it and the previous point i-1 can be calculated: ΔP(i) = (xi-xi-1,yi-yi-1,zi-zi-1). The displacement vector is then normalized to obtain a unit direction vector: u(i) = ΔP(i) / |ΔP(i)|, where |ΔP(i)| represents the modulus of the vector ΔP(i). This unit direction vector u(i) represents the direction of motion of the trajectory at point i. The extracted features are quantized and combined into a fixed-length feature vector. For example, each feature can be quantized into 10 levels, and all features are combined into a 50-dimensional feature vector. Each dimension represents a specific feature, and its value indicates the strength of that feature. For example, a gesture representing "rapid upward movement" has a high velocity feature value and a direction feature value close to 90 degrees. Finally, gesture motion feature vector data is generated. A set of basic gesture action units is defined, such as "open palm," "clenched fist," and "pointing finger." Then, an HMM model is trained using the labeled gesture data to learn the feature vector distribution corresponding to each basic gesture action unit.The trained HMM model can identify the most hidden state sequence, that is, the basic gesture action unit sequence, based on the input feature vector sequence. For example, a gesture sequence containing two actions, "open palm" and "clench fist", will have its feature vector sequence divided into two parts by the HMM model, corresponding to the two basic gesture action units, "open palm" and "clench fist". The basic gesture action unit data contains a series of gesture action segments, each of which corresponds to a feature vector data segment and a time range. The gesture-semantic mapping database is a pre-established database that stores the correspondence between the feature vector data and semantic labels of various gesture action units. For example, the database stores semantic labels such as "upward movement", "left movement", "drawing a circle", and "waving", as well as the feature vector data corresponding to these labels. The labeling process is to match the feature vector data of each basic gesture action unit obtained in step S23 in the gesture-semantic mapping database, find the feature vector data that is most similar to it, and assign the corresponding semantic label to the basic gesture action unit. For example, for a basic gesture action unit, whose feature vector data is V, the database is searched for the feature vector data V' that is closest to V. Assuming that the semantic label corresponding to V' is "upward motion," the label "upward motion" is assigned to the basic gesture action unit. The semantic labels of all action units are arranged in chronological order to generate gesture motion textual description data. ResNet-50 is pre-trained using a publicly available large-scale image dataset (such as ImageNet) to learn rich image feature representations. The convolutional layer parameters of the pre-trained model are then frozen, and only the fully connected layer parameters are trained. The model is fine-tuned using a labeled gesture action intention dataset to adapt it to the task of gesture action intention recognition. For example, a gesture dataset containing action categories such as "grab," "drop," and "move" is used for training. During training, a cross-entropy loss function is used as the objective function, and the Adam optimizer is used for parameter updates. After training, a gesture action intention recognition model is obtained that maps gesture action feature vectors to action intention categories. The feature vectors of the basic gesture action units are input into the model to obtain the intention probability distribution for each action unit. For example, the intent probability distribution for an action unit "opening the palm and moving toward an object" is: grab (80%), drop (10%), and move (10%). The intent with the highest probability is selected as the recognition result for that action unit. The intent recognition results of all action units are arranged in chronological order to obtain gesture action intention recognition data. For example, a gesture sequence containing three action units has the intent recognition results of "grab," "move," and "drop." The final result is gesture action intention recognition data.
[0120] Preferably, step S22 includes the following steps:
[0121] Step S221: extracting the coordinates of the starting point and ending point of the gesture trajectory based on the pre-processed gesture motion trajectory data to obtain the trajectory starting and ending point data;
[0122] Step S222: Calculating the displacement vectors of the hand key points in adjacent frames based on the pre-processed gesture motion trajectory data to generate inter-frame displacement vector data of the hand;
[0123] Step S223: performing inter-frame instantaneous velocity vector calculation on the inter-frame displacement vector data of the hand to generate instantaneous velocity vector data;
[0124] Step S224: performing low-pass filtering on the instantaneous velocity vector data and calculating the gesture motion acceleration to obtain frame-by-frame acceleration vector data;
[0125] Step S225: performing mean processing on the instantaneous velocity vector data and the frame-by-frame acceleration vector data through a preset time window to obtain velocity and acceleration feature data;
[0126] Step S226: performing trajectory curve fitting on the pre-processed gesture motion trajectory data using the trajectory start and end point data, and performing gesture trajectory curvature and direction change analysis to obtain gesture trajectory geometric feature data;
[0127] Step S227: performing gesture posture change frequency analysis based on the pre-processed gesture motion trajectory data to obtain gesture posture change frequency data;
[0128] Step S228: performing gesture feature combination on the velocity acceleration feature data, the gesture trajectory geometric feature data, and the gesture posture change frequency data to obtain gesture motion feature vector data.
[0129] In an embodiment of the present invention, the pre-processed gesture motion trajectory data is a time series, in which each time point corresponds to a three-dimensional coordinate. The starting point coordinate of the trajectory is the first coordinate of the time series, and the ending point coordinate is the last coordinate of the time series. For example, a gesture trajectory data containing 10 frames has a starting point coordinate of the three-dimensional coordinate of the first frame and an ending point coordinate of the tenth frame. For each frame, the three-dimensional coordinate difference of each key point between the current frame and the previous frame is calculated to obtain a three-dimensional displacement vector. For example, the displacement vector of a key point in the second frame is the difference between the three-dimensional coordinates of the key point in the second frame and the first frame. The displacement vectors of all key points are combined into a matrix to generate the inter-frame displacement vector data of the hand. For example, a gesture containing 21 key points has an inter-frame displacement vector data of a 21x3 matrix, in which each row represents the displacement vector of a key point. Since the time interval between adjacent frames is fixed (for example, 1 / 30 second), the displacement vector can be divided by the time interval to obtain the instantaneous velocity vector. For example, if the time interval between adjacent frames is 0.033 seconds, the displacement vector obtained in step S222 is divided by 0.033 to obtain the instantaneous velocity vector. The instantaneous velocity vectors of each frame are combined into a matrix to generate instantaneous velocity vector data. A Butterworth filter or other type of low-pass filter can be used. The filtered data can more accurately reflect the movement trend of the gesture. Then, the change in the instantaneous velocity vector between adjacent frames is calculated to obtain a frame-by-frame acceleration vector. The calculation method is similar to that of step S223, that is, the difference between the instantaneous velocity vectors of adjacent frames is divided by the time interval. For example, the difference between the instantaneous velocity vectors of the third frame and the second frame is divided by the time interval to obtain the acceleration vector of the third frame. For example, the size of the time window can be set to 5 frames, that is, the data of the current moment and the 2 frames before and after it are considered. For each time window, the average value of the instantaneous velocity vector and acceleration vector of all frames in the window is calculated as the velocity and acceleration characteristics of the time window. For example, if the time window is 0.1 seconds and the frame rate is 30 frames per second, each time window contains 3 frames of data, and the average velocity and acceleration of these 3 frames are calculated. Cubic spline interpolation or other curve fitting methods can be used. The fitted curve can more smoothly represent the gesture trajectory. The fitted curve is then analyzed for curvature and direction change. Curvature represents the degree of curvature of the curve and can be calculated using the curvature formula. Direction change represents the rate of change of the tangent direction angle of the curve. The curvature and direction change are quantified and combined into a feature vector to obtain the geometric feature data of the gesture trajectory. For example, several key joints can be selected, such as the fingertip joints of the index finger, middle finger, and thumb, and the distance or angle between them can be calculated. Then, the change in these distances or angles between adjacent frames is calculated.For example, for the distance d between the tips of the index finger and the middle finger, calculate the absolute value of the difference between the two adjacent frames: Δd = |d(i)-d(i-1)|, where d(i) represents the distance between the tips of the index finger and the middle finger in the i-th frame. For the angle, the absolute value of the angle difference between the two adjacent frames can be calculated. Count the frequency of these changes over a period of time. For example, a time window can be set, such as 10 frames. For each time window, count the number of times the change exceeds a certain threshold. For example, set the threshold for distance change to 2 cm and the threshold for angle change to 10 degrees. In a 10-frame time window, if the distance change Δd exceeds 2 cm 6 times, then the distance change frequency in the time window is 6 / 10=0.6. Similarly, the angle change frequency can be calculated. The change frequencies in all time windows are calculated to form the gesture posture change frequency data. The velocity and acceleration feature data obtained in step S225, the gesture trajectory geometric feature data obtained in step S226, and the gesture posture change frequency data obtained in step S227 are combined to obtain gesture motion feature vector data. The velocity and acceleration feature data includes information about the velocity and acceleration of the gesture motion, the gesture trajectory geometric feature data includes information about the shape of the gesture trajectory, and the gesture posture change frequency data includes information about changes in the gesture posture.
[0130] Preferably, step S24 includes the following steps:
[0131] Step S231: performing key frame detection on the gesture motion feature vector data to obtain key frame data;
[0132] Step S232: segmenting the pre-processed gesture motion trajectory data into key frame time periods using the key frame data to obtain time period segmented trajectory data;
[0133] Step S233: performing hand gesture cluster analysis on the time segmented trajectory data to obtain hand gesture cluster data;
[0134] Step S234: dividing the time segmentation trajectory data into basic gesture action units based on the gesture posture clustering data to generate basic gesture action unit division data;
[0135] Step S235: performing conversion relationship processing between gesture action units on the basic gesture action unit division data to obtain unit conversion relationship data;
[0136] Step S236: Integrate the basic gesture action unit division data and the unit conversion relationship data to generate basic gesture action unit data.
[0137] In an embodiment of the present invention, key frame detection is performed based on the rate of change of feature vectors. For example, the Euclidean distance between feature vectors of adjacent frames can be calculated as the rate of change of features. For example, for a gesture motion feature vector sequence F = {F1, F2, ..., FN}, where Fi represents the feature vector of the i-th frame. Calculate the distance between the feature vectors of two adjacent frames: d(i, i+1) = ||Fi+1-Fi||, where ||.|| represents the Euclidean norm of the vector. Compare these distances with a preset threshold. If the distance is greater than the threshold, the corresponding frame is considered to be a key frame. For example, the threshold can be set to twice the average value of all distances. For the calculated distance sequence d = {d(1, 2), d(2, 3), ..., d(N-1, N)}, calculate its average value d_avg. If d(i, i+1)>2×d_avg, the i+1-th frame is marked as a key frame. The time period between two adjacent key frames is regarded as a segmentation unit. For example, if three keyframes are detected, located at frames 5, 10, and 15, the gesture trajectory is segmented into three time segments: frames 1 to 5, frames 6 to 10, and frames 11 to 15. Each time segment represents a relatively stable gesture pose or a complete action unit. The segmented trajectory data contains the start and end frame numbers of each time segment, as well as the trajectory data for that time segment. Feature extraction is performed on the trajectory data within each time segment, such as calculating average velocity, average acceleration, and trajectory length. Then, using a clustering algorithm such as K-Means or DBSCAN, time segments with similar features are clustered together. Each cluster represents a specific gesture pose. For example, all time segments representing "open palm" are clustered into one category, and all time segments representing "clenched fist" are clustered into another category. The clustering results generate gesture pose cluster data, which contains the cluster category to which each time segment belongs. Each cluster category is considered a basic gesture action unit. For example, all time periods corresponding to the "open palm" cluster are divided into one basic gesture action unit, and all time periods corresponding to the "clenched fist" cluster are divided into another basic gesture action unit. The division result generates basic gesture action unit division data, which contains all time periods corresponding to each basic gesture action unit. The basic gesture action units belonging to adjacent time periods are analyzed to determine the conversion relationship between units. For example, if one time period belongs to the "open palm" unit and the next time period belongs to the "clenched fist" unit, it is considered that there is a conversion relationship from "open palm" to "clenched fist". All conversion relationships are recorded to generate unit conversion relationship data. For example, a directed graph can be used to represent the conversion relationship between units. The nodes in the graph represent basic gesture action units, and the edges represent the conversion relationship between units.The basic gesture action unit division data generated in step S234 and the unit conversion relationship data generated in step S235 are integrated to generate the final basic gesture action unit data. This data contains information about each basic gesture action unit, such as the unit's start time, end time, corresponding trajectory data, and conversion relationships with other units.
[0138] Preferably, step S24 includes the following steps:
[0139] Step S241: determining the gesture motion intensity level based on the velocity acceleration feature data to generate motion intensity level data;
[0140] Step S242: mapping the exercise intensity level data to exercise semantic text labels using a preset gesture-semantic mapping database to obtain intense exercise semantic vocabulary data;
[0141] Step S243: performing gesture complexity level analysis on the gesture trajectory geometric feature data to obtain complexity level data;
[0142] Step S244: using a preset gesture-semantic mapping database to perform motion semantic text label mapping on the complexity level data to obtain complex motion semantic vocabulary data;
[0143] Step S245: classifying the gesture rhythm speed level of the gesture posture change frequency data to obtain gesture rhythm level data;
[0144] Step S246: mapping the gesture rhythm level data to gesture rhythm text labels using a preset gesture-semantic mapping database to obtain gesture rhythm text vocabulary data;
[0145] Step S247: performing gesture semantic description label fusion on the violent motion semantic vocabulary data, the complex motion semantic vocabulary data and the gesture rhythm text vocabulary data based on the basic gesture action unit data to obtain gesture motion text description data.
[0146] In an embodiment of the present invention, the intensity level of exercise is determined based on the average values of velocity and acceleration. For example, several thresholds can be pre-set to divide the average velocity and acceleration values into several intervals, each corresponding to an intensity level. For example, the thresholds for the average velocity can be set to 5 cm / s and 15 cm / s, dividing the average velocity into three intervals: less than 5 cm / s, 5-15 cm / s, and greater than 15 cm / s, corresponding to the three intensity levels of "mild," "moderate," and "violent," respectively. Similarly, a threshold for the average acceleration can be set to divide it into several intervals, each corresponding to different intensity levels. For each gesture, the intensity level is determined based on the average velocity and average acceleration values in its velocity and acceleration feature data. For example, if a gesture has an average velocity of 10 cm / s and an average acceleration of 20 cm / s^2, then based on the above thresholds, its velocity intensity level is "moderate" and its acceleration intensity level is also "moderate." The velocity intensity level and acceleration intensity level are combined to form the final exercise intensity level. For example, if the speed intensity level and the acceleration intensity level are consistent, then this level is used as the final motion intensity level. If they are inconsistent, the final motion intensity level can be determined according to a pre-set rule. For example, the acceleration intensity level can be given priority, or the higher of the two levels can be used as the final motion intensity level. For example, in the above example, since the speed and acceleration intensity levels are both "medium", the motion intensity level of the gesture action is also "medium". The motion intensity level data generated in step S241 is mapped to motion semantic text labels using a preset gesture-semantic mapping database. The gesture-semantic mapping database pre-defines semantic text labels corresponding to different motion intensity levels. For example, "mild" corresponds to "slowly", "medium" corresponds to "steadily", and "violent" corresponds to "quickly". Based on the motion intensity level data, the corresponding semantic text labels are searched from the database to obtain violent motion semantic vocabulary data. For example, a sequence containing three basic gesture action units, whose motion intensity level data is "mild, violent, medium", then the corresponding violent motion semantic vocabulary data is "slowly, quickly, steadily". Based on features such as the curvature and directional changes of the trajectory, gesture complexity is classified into different levels, such as "simple," "medium," and "complex." This classification can be performed using predefined thresholds or machine learning-based models. For example, complexity can be determined based on the trajectory's total curvature and directional changes, with higher values indicating greater complexity. The analysis results generate complexity level data, with each element corresponding to the complexity level of a basic gesture action unit. For this complexity level data, the gesture-semantic mapping database stores semantic labels corresponding to each complexity level.For example, the semantic label corresponding to the complexity level of "simple" can be "simply", "directly", etc., the semantic label corresponding to the complexity level of "medium" can be "moderately complex", "skillfully", etc., and the semantic label corresponding to the complexity level of "complex" can be "complexly", "finely", etc. The mapping process is to search for the corresponding semantic label in the gesture-semantic mapping database according to the complexity level of each gesture action. For example, if the complexity level of a gesture action is "complex", then the semantic label corresponding to "complex" is searched in the database, and labels such as "complexly" and "finely" are found. The gesture rhythm speed level is divided into gesture posture change frequency data (step S227). According to the frequency of gesture posture change, the gesture rhythm is divided into different levels, such as "slow", "medium speed" and "fast". Predefined thresholds can be used for division, for example, the frequency below threshold A is "slow", the frequency between threshold A and threshold B is "medium speed", and the frequency above threshold B is "fast". The division result generates gesture rhythm level data. The gesture rhythm level data is mapped to gesture rhythm text labels using the preset gesture-semantic mapping database. The database pre-defines text labels corresponding to different rhythm levels. For example, "slow" corresponds to "slow rhythm", "medium speed" corresponds to "moderate rhythm", and "fast" corresponds to "fast rhythm". According to the gesture rhythm level data, the corresponding text label is searched from the database to obtain the gesture rhythm text vocabulary data. These semantic information are associated with the basic gesture action units. For each basic gesture action unit, its corresponding violent motion semantic vocabulary, complex motion semantic vocabulary and gesture rhythm text vocabulary can be combined to form a semantic label that describes the action unit.
[0147] Preferably, step S3 includes the following steps:
[0148] Step S31: performing confidence evaluation on the gesture action intention recognition data and screening high-confidence intentions to obtain high-confidence action intention data;
[0149] Step S32: performing action emotion adjective processing based on the high-confidence action intention data, and performing gesture semantic fusion based on the gesture motion text description data to generate key gesture semantic fusion data;
[0150] Step S33: performing intelligent text generation based on the key gesture semantic fusion data to generate initial gesture text content data;
[0151] Step S34: performing gesture category analysis based on the basic gesture action unit data, and performing text template keyword processing to generate text template keyword data;
[0152] Step S35: using the text template keyword data to perform text template matching on a preset text template library to obtain selected text template data;
[0153] Step S36: performing placeholder semantic filling on the initial gesture text content data by selecting text template data, thereby obtaining intelligent description gesture text data.
[0154] As an example of the present invention, refer to Figure 3 As shown, Figure 1 Detailed implementation steps of step S3 are shown in the flowchart. In this example, step S3 includes:
[0155] Step S31: performing confidence evaluation on the gesture action intention recognition data and screening high-confidence intentions to obtain high-confidence action intention data;
[0156] In an embodiment of the present invention, the confidence level is evaluated based on the probability value output by the gesture motion intention recognition model. For example, in step S26, the gesture motion intention recognition model outputs a vector, each element of which represents the probability that the gesture belongs to each intent category. The category with the highest probability can be selected as the intent recognition result, and the probability value can be used as the confidence level. For example, if the vector output by the model is [0.1, 0.8, 0.05, 0.05, 0], it means that the probability that the gesture belongs to the first intent category is 0.1, and the probability that it belongs to the second intent category is 0.8, and so on. Then, the second category can be selected as the intent recognition result, and 0.8 can be used as the confidence level. A confidence threshold value can be set in advance, such as 0.7. If the confidence level of an intent recognition result is higher than the threshold value, the result is considered to be of high confidence and is retained; otherwise, the result is considered to be of low confidence and is discarded. For example, for the above example, since 0.8 is greater than 0.7, the intent recognition result is retained. If the confidence of another intent recognition result is 0.6, then since 0.6 is less than 0.7, this result is discarded. All high-confidence intent recognition results are filtered out to form high-confidence action intention data.
[0157] Step S32: performing action emotion adjective processing based on the high-confidence action intention data, and performing gesture semantic fusion based on the gesture motion text description data to generate key gesture semantic fusion data;
[0158] In an embodiment of the present invention, an emotional adjective library is pre-established, which contains various emotional adjectives, such as "quickly", "slowly", "carefully", "boldly", etc. Then, according to the characteristics of the gesture action, such as speed, acceleration, trajectory shape, etc., appropriate emotional adjectives are selected. For example, according to the speed and acceleration feature data obtained in step S225, if the speed and acceleration of a gesture action are relatively large, emotional adjectives such as "quickly" and "swiftly" can be selected; if the speed and acceleration are relatively small, emotional adjectives such as "slowly" and "gently" can be selected. The selection of adjectives can be based on predefined rules or machine learning models. The selected emotional adjectives are fused with the gesture motion text description data generated in step S247 to generate key gesture semantic fusion data. The fusion method can be to insert the adjective into the appropriate position of the text description, for example, inserting "quickly" before "opening the palm" to form "opening the palm quickly".
[0159] Step S33: performing intelligent text generation based on the key gesture semantic fusion data to generate initial gesture text content data;
[0160] In an embodiment of the present invention, text generation is performed using a rule-based approach or a deep learning-based approach. The rule-based approach can use predefined grammatical rules and templates to combine the various elements in the key gesture semantic fusion data into sentences. For example, a template can be defined: "user [emotional adjective] [intention] [semantic description label]", and then the corresponding elements in the key gesture semantic fusion data can be filled into the template. For example, for the fusion result of "quickly select", "user quickly selects" can be generated.
[0161] Step S34: performing gesture category analysis based on the basic gesture action unit data, and performing text template keyword processing to generate text template keyword data;
[0162] In this embodiment of the present invention, gesture category analysis is performed based on the basic gesture action unit data (step S236). For example, gestures can be divided into categories such as "grabbing," "dropping," and "pointing." Then, corresponding keywords are extracted based on the gesture category. For example, for the "grabbing" category, keywords such as "grab," "hold," and "take" can be extracted; for the "pointing" category, keywords such as "point," "point to," and "indicate" can be extracted. The extracted keywords are used for subsequent text template matching to generate text template keyword data.
[0163] Step S35: using the text template keyword data to perform text template matching on a preset text template library to obtain selected text template data;
[0164] In an embodiment of the present invention, the text template keyword data generated in step S34 is used to perform text template matching on a preset text template library. The text template library contains various text templates that describe gesture actions, and each template contains some placeholders for filling in specific semantic information. For example, a text template can be "[action][object]", where "[action]" and "[object]" are placeholders. The matching process selects the most appropriate text template based on the text template keyword data. For example, if the keyword is "grab", a text template containing "grab" or its synonyms can be selected. The matching result obtains the selected text template data.
[0165] Step S36: performing placeholder semantic filling on the initial gesture text content data by selecting text template data, thereby obtaining intelligent description gesture text data.
[0166] In an embodiment of the present invention, the selected text template data obtained in step S35 performs semantic filling of the placeholders of the initial gesture text content data generated in step S33. The semantic information extracted from the initial gesture text content data is filled into the placeholders of the text template, thereby generating complete intelligent description gesture text data. For example, if the selected text template is "[action][object]" and the initial gesture text content data is "quickly grab an apple", the filled text is "grab the apple". If the initial text content cannot completely fill the placeholders of the template, a default value can be used or inferred based on the context.
[0167] Preferably, the present invention further provides a text generation system for executing the above-mentioned text generation method, wherein the text generation system comprises:
[0168] The gesture trajectory reconstruction module is used to use a 3D depth camera to collect real-time hand depth images of the user to obtain depth hand image sequence data; based on the depth hand image sequence data, the gesture trajectory is reconstructed to obtain spatial gesture motion trajectory data;
[0169] The gesture semantic annotation module is used to analyze the gesture motion characteristics of spatial gesture motion trajectory data to generate gesture motion feature vector data; segment the gesture motion feature vector data into basic gesture action units to generate basic gesture action unit data; annotate the gesture motion feature vector data with gesture motion semantic labels based on the basic gesture action unit data to generate gesture motion text description data; construct a gesture motion intention recognition model; and use the gesture motion intention recognition model to perform action intention recognition on the gesture motion feature vector data to obtain gesture action intention recognition data.
[0170] The intelligent gesture description text generation module is used to process action emotion adjectives based on gesture action intention recognition data, and perform gesture semantic fusion based on gesture motion text description data to generate key gesture semantic fusion data; based on the key gesture semantic fusion data, intelligent text generation is performed and text template semantic filling is performed to obtain intelligent gesture description text data;
[0171] The description text visualization module is used to perform visual text rendering processing based on the intelligent description gesture text data to realize the generation of intelligent description text for gesture actions.
[0172] Preferably, the present invention further provides a terminal device, comprising:
[0173] processor;
[0174] a memory for storing processor-executable instructions;
[0175] The processor is configured to implement any one of the above text generation methods.
[0176] Preferably, the present invention further provides a computer-readable storage medium storing a computer program, wherein the computer program implements any one of the above text generation methods when executed.
[0177] This application aims to accurately capture the hand's motion trajectory in three-dimensional space by using a 3D depth camera to capture real-time depth images of the user's hand. This overcomes the shortcomings of traditional gesture recognition methods based on two-dimensional images, which are sensitive to gesture posture, viewing angle, and lighting changes, making gesture recognition more stable and accurate. This invention effectively preprocesses raw trajectory data by analyzing the length of gesture spatial trajectories and time-warping the sampled trajectories, improving the standardization and usability of the data. When different users perform the same gesture, their speed and trajectory length vary. Length analysis and time warping unify trajectory data of varying lengths and speeds to the same standard. Gesture motion feature analysis converts the preprocessed trajectory data into gesture motion feature vector data, achieving the transformation from concrete trajectories to abstract features. Geometric features of gesture trajectories are extracted through trajectory curve fitting, gesture trajectory curvature, and directional change analysis. These geometric features reflect the shape and directional changes of gesture motion. For example, circular gestures and linear gestures have different curvature characteristics, and leftward and rightward swipes have different directional change characteristics. Gesture motion feature data is mapped onto different semantic dimensions. This multi-dimensional analysis enables more comprehensive and detailed descriptions of gestures, such as "fast," "slow," "complex," "simple," and "rhythmic." Using a pre-set gesture-semantic mapping database, these levels of data are mapped to corresponding semantic text labels. This database is key to transforming feature data into textual descriptions, associating different feature levels with corresponding semantic terms, such as "violent," "mild," "complex," "simple," "fast," and "soothing." Through mapping, the system can transform abstract feature data into concrete semantic terms. High-confidence action intention data is processed into action emotion adjectives. Different categories of gestures, such as interactive, instructional, and expressive, typically correspond to different text templates. Keyword matching allows for rapid identification of appropriate text templates. Combining the generated text content with the template ensures that the final text output is both personalized and readable and fluent.
[0178] The present invention is therefore intended to be illustrative and non-restrictive in all respects, with the scope of the invention being defined by the appended claims rather than the foregoing description, and all changes that come within the meaning and range of equivalents of the application documents are intended to be embraced therein.
[0179] The foregoing description is intended only to provide specific embodiments of the present invention, which will enable those skilled in the art to understand and implement the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not intended to be limited to the embodiments shown herein, but is to be construed in the widest possible manner consistent with the principles and novel features disclosed herein.
Claims
1. A text generation method, characterized in that: The following steps are involved: Step S1: using a 3D depth camera to collect real-time hand depth images of the user to obtain depth hand image sequence data; reconstructing gesture trajectories based on the depth hand image sequence data to obtain spatial gesture motion trajectory data; Step S2: performing gesture motion feature analysis on the spatial gesture motion trajectory data to generate gesture motion feature vector data; Segment the basic gesture action units based on the gesture motion feature vector data to generate basic gesture action unit data; annotate the gesture motion semantic labels on the gesture motion feature vector data based on the basic gesture action unit data to generate gesture motion text description data; construct a gesture motion intention recognition model; use the gesture motion intention recognition model to perform action intention recognition on the gesture motion feature vector data to obtain gesture action intention recognition data; Step S3: Processing action emotion adjectives based on the gesture action intention recognition data, and performing gesture semantic fusion based on the gesture movement text description data to generate key gesture semantic fusion data; Based on the key gesture semantic fusion data, intelligent text generation is performed and text template filling is performed to obtain text data that intelligently describes gestures; Step S4: Perform visual text rendering processing based on the intelligent description gesture text data to achieve the generation of intelligent description text for the gesture action.
2. The text generation method according to claim 1, characterized in that Step S1 includes the following steps: Step S11: using a 3D depth camera to collect real-time hand depth images of the user, and perform initial image correction to obtain depth hand image sequence data; Step S12: performing background removal on the depth hand image sequence data and performing hand region segmentation to obtain user hand region data; Step S13: performing hand posture estimation on the user's hand area data to obtain the user's hand posture data; Step S14: Detecting skeleton key points based on the user's hand posture data, and performing hand skeleton key point coordinate analysis to obtain skeleton key point coordinate data; Step S15: extracting pixel coordinates of gesture key points from the depth hand image sequence data using the skeleton key point coordinate data to obtain pixel coordinate data of gesture image key points; Step S16: performing timestamp synchronization processing based on the pixel coordinate data of the key points of the gesture image, and reconstructing the gesture trajectory based on the user's hand posture data to obtain spatial gesture motion trajectory data.
3. The text generation method according to claim 1, characterized in that Step S2 includes the following steps: Step S21: analyzing the spatial gesture motion trajectory data for the length of the gesture spatial trajectory and performing time regularization processing on the sampling trajectory to obtain pre-processed gesture motion trajectory data; Step S22: performing gesture motion feature analysis based on the pre-processed gesture motion trajectory data to generate gesture motion feature vector data; Step S23: Segmenting the basic gesture action units according to the gesture motion feature vector data to generate basic gesture action unit data; Step S24: annotating the gesture motion feature vector data with gesture motion semantic labels using a preset gesture-semantic mapping database based on the basic gesture action unit data to generate gesture motion text description data; Step S25: performing gesture movement intention sample transfer learning according to a preset convolutional neural network model, thereby constructing a gesture movement intention recognition model; Step S26: using the gesture motion intention recognition model to perform motion intention recognition on the basic gesture action unit data and the gesture motion feature vector data to obtain gesture action intention recognition data.
4. The text generation method according to claim 3, characterized in that Step S22 includes the following steps: Step S221: extracting the coordinates of the starting point and ending point of the gesture trajectory based on the pre-processed gesture motion trajectory data to obtain the trajectory starting and ending point data; Step S222: Calculating the displacement vectors of the hand key points in adjacent frames based on the pre-processed gesture motion trajectory data to generate inter-frame displacement vector data of the hand; Step S223: performing inter-frame instantaneous velocity vector calculation on the inter-frame displacement vector data of the hand to generate instantaneous velocity vector data; Step S224: performing low-pass filtering on the instantaneous velocity vector data and calculating the gesture motion acceleration to obtain frame-by-frame acceleration vector data; Step S225: performing mean processing on the instantaneous velocity vector data and the frame-by-frame acceleration vector data through a preset time window to obtain velocity and acceleration feature data; Step S226: performing trajectory curve fitting on the pre-processed gesture motion trajectory data using the trajectory start and end point data, and performing gesture trajectory curvature and direction change analysis to obtain gesture trajectory geometric feature data; Step S227: performing gesture posture change frequency analysis based on the pre-processed gesture motion trajectory data to obtain gesture posture change frequency data; Step S228: performing gesture feature combination on the velocity acceleration feature data, the gesture trajectory geometric feature data, and the gesture posture change frequency data to obtain gesture motion feature vector data.
5. The text generation method according to claim 3, characterized in that Step S24 includes the following steps: Step S231: performing key frame detection on the gesture motion feature vector data to obtain key frame data; Step S232: segmenting the pre-processed gesture motion trajectory data into key frame time periods using the key frame data to obtain time period segmented trajectory data; Step S233: performing hand gesture cluster analysis on the time segmented trajectory data to obtain hand gesture cluster data; Step S234: dividing the time segmentation trajectory data into basic gesture action units based on the gesture posture clustering data to generate basic gesture action unit division data; Step S235: performing conversion relationship processing between gesture action units on the basic gesture action unit division data to obtain unit conversion relationship data; Step S236: Integrate the basic gesture action unit division data and the unit conversion relationship data to generate basic gesture action unit data.
6. The text generation method according to claim 4, characterized in that Step S24 includes the following steps: Step S241: determining the gesture motion intensity level based on the velocity acceleration feature data to generate motion intensity level data; Step S242: mapping the exercise intensity level data to exercise semantic text labels using a preset gesture-semantic mapping database to obtain intense exercise semantic vocabulary data; Step S243: performing gesture complexity level analysis on the gesture trajectory geometric feature data to obtain complexity level data; Step S244: using a preset gesture-semantic mapping database to perform motion semantic text label mapping on the complexity level data to obtain complex motion semantic vocabulary data; Step S245: classifying the gesture rhythm speed level of the gesture posture change frequency data to obtain gesture rhythm level data; Step S246: mapping the gesture rhythm level data to gesture rhythm text labels using a preset gesture-semantic mapping database to obtain gesture rhythm text vocabulary data; Step S247: performing gesture semantic description label fusion on the violent motion semantic vocabulary data, the complex motion semantic vocabulary data and the gesture rhythm text vocabulary data based on the basic gesture action unit data to obtain gesture motion text description data.
7. The text generation method according to claim 1, characterized in that Step S3 includes the following steps: Step S31: performing confidence evaluation on the gesture action intention recognition data and screening high-confidence intentions to obtain high-confidence action intention data; Step S32: performing action emotion adjective processing based on the high-confidence action intention data, and performing gesture semantic fusion based on the gesture motion text description data to generate key gesture semantic fusion data; Step S33: performing intelligent text generation based on the key gesture semantic fusion data to generate initial gesture text content data; Step S34: performing gesture category analysis based on the basic gesture action unit data, and performing text template keyword processing to generate text template keyword data; Step S35: using the text template keyword data to perform text template matching on a preset text template library to obtain selected text template data; Step S36: performing placeholder semantic filling on the initial gesture text content data by selecting text template data, thereby obtaining intelligent description gesture text data.
8. A text generation system, characterized in that: For executing the text generation method according to claim 1, the text generation system comprises: The gesture trajectory reconstruction module is used to use a 3D depth camera to collect real-time hand depth images of the user to obtain depth hand image sequence data; based on the depth hand image sequence data, the gesture trajectory is reconstructed to obtain spatial gesture motion trajectory data; The gesture semantic annotation module is used to analyze the gesture motion characteristics of spatial gesture motion trajectory data to generate gesture motion feature vector data; segment the basic gesture action units based on the gesture motion feature vector data to generate basic gesture action unit data; annotate the gesture motion feature vector data with gesture motion semantic labels based on the basic gesture action unit data to generate gesture motion text description data; construct a gesture motion intention recognition model; use the gesture motion intention recognition model to perform action intention recognition on the gesture motion feature vector data to obtain gesture action intention recognition data; The intelligent gesture description text generation module is used to process action emotion adjectives based on gesture action intention recognition data, and perform gesture semantic fusion based on gesture motion text description data to generate key gesture semantic fusion data; based on the key gesture semantic fusion data, intelligent text generation is performed and text template semantic filling is performed to obtain intelligent gesture description text data; The description text visualization module is used to perform visual text rendering processing based on the intelligent description gesture text data to realize the generation of intelligent description text for gesture actions.
9. A terminal device, characterized in that: The terminal device includes: processor; a memory for storing processor-executable instructions; The processor is configured to implement the text generation method according to any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed, the text generation method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Text recognition method and device applied to air handwriting, equipment and storage medium
CN115826762A
Air handwriting interaction method based on three-dimensional gesture reconstruction, storage medium and device
CN117058691A