Text generation method and system, terminal equipment and storage medium

Through a 3D depth camera, the gesture trajectory is reconstructed and feature analysis is performed. Combined with intention recognition and emotional processing, intelligent descriptive gesture text data is generated, solving the limitations of traditional methods in terms of flexibility and personalization, and achieving more accurate and vivid text generation.

CN120031012AActive Publication Date: 2025-05-23SHENZHEN WRITER INTELLIGENT TECHNOLOGY CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510084544.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-05-23
Estimated Expiration
2045-01-20

AI Technical Summary

Technical Problem

Traditional gesture text generation methods have limitations in expression flexibility, dynamicity and personalization, and it is difficult to dynamically adjust according to the user's real-time intentions and interactive scenarios, and it is impossible to intuitively establish a mapping relationship between the user's body language and text content.

Method used

By using a 3D depth camera to collect user's hand depth image sequence data, reconstruct gesture trajectory, perform gesture motion feature analysis and basic gesture action unit segmentation, build a gesture motion intention recognition model, combine action intention and emotional adjectives, intelligent text generation and template filling, and generate intelligent description gesture text data.

Benefits of technology

It realizes more accurate and stable gesture recognition, can dynamically adjust text generation to conform to the user's real-time intentions, generate more vivid and well-readable text descriptions, and intuitively associate gestures with text content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120031012A_ABST
    Figure CN120031012A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of text generation, in particular to a text generation method and system, terminal equipment and a storage medium. The method comprises the following steps: performing real-time gesture track reconstruction on a user by using a 3D depth camera to obtain space gesture motion track data; carrying out basic gesture action unit segmentation on the spatial gesture movement track data to generate basic gesture action unit data; performing gesture motion semantic label labeling based on the basic gesture action unit data, and performing action intention recognition to obtain gesture action intention recognition data; performing action emotion adjective processing according to the gesture action intention recognition data, and performing intelligent text generation to obtain intelligent description gesture text data; therefore, the gesture action intelligent description text is generated. According to the method, gesture motion features are analyzed in real time to recognize action intentions, and dynamic text description of gesture actions is achieved in combination with emotional adjectives.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of text generation, and in particular to a text generation method, system, terminal equipment and storage medium. Background Art

[0002] As human-computer interaction becomes increasingly diversified, traditional input methods such as keyboards and mice are gradually unable to meet people's needs in specific scenarios. In recent years, gesture recognition, as a natural and intuitive way of interaction, has received widespread attention and has developed rapidly, and has shown great application potential in smart home, virtual reality, augmented reality, medical assistance and other fields. As a natural and intuitive way of expression, gestures carry rich information and can convey users' intentions and emotions. Using computer vision technology to recognize and understand gestures and convert them into corresponding text information can achieve more natural and convenient human-computer interaction. However, traditional gesture text generation methods mostly focus on the classification and recognition of gestures and convert them into simple command texts. There are still great limitations in terms of flexibility, dynamism and personalization of expression. It is difficult to dynamically adjust according to the user's real-time intentions and interaction scenarios, and it is also impossible to intuitively map the user's body language with the text content. Summary of the invention

[0003] Based on this, the present invention provides a text generation method, system, terminal device and storage medium to solve at least one of the above technical problems.

[0004] To achieve the above object, a text generation method comprises the following steps:

[0005] Step S1: using a 3D depth camera to collect real-time hand depth images of the user to obtain depth hand image sequence data; reconstructing gesture trajectories according to the depth hand image sequence data to obtain spatial gesture motion trajectory data;

[0006] Step S2: performing gesture motion feature analysis on the spatial gesture motion trajectory data to generate gesture motion feature vector data; performing basic gesture action unit segmentation according to the gesture motion feature vector data to generate basic gesture action unit data; annotating the gesture motion semantic label on the gesture motion feature vector data based on the basic gesture action unit data to generate gesture motion text description data; constructing a gesture motion intention recognition model; using the gesture motion intention recognition model to perform action intention recognition on the gesture motion feature vector data to obtain gesture action intention recognition data;

[0007] Step S3: Processing the action emotion adjectives according to the gesture action intention recognition data, and performing gesture semantic fusion according to the gesture motion text description data to generate key gesture semantic fusion data; performing intelligent text generation based on the key gesture semantic fusion data, and performing text template filling to obtain intelligent description gesture text data;

[0008] Step S4: Perform visual text rendering processing according to the intelligent description gesture text data to realize the generation of intelligent description text for gesture actions.

[0009] The present invention uses a 3D depth camera to obtain depth hand image sequence data and reconstructs the gesture trajectory, which can accurately capture the motion trajectory of the hand in three-dimensional space, overcome the shortcomings of the traditional gesture recognition method based on two-dimensional images that is sensitive to gesture posture, viewing angle and illumination changes, and make gesture recognition more stable and accurate. By performing feature analysis on the spatial gesture motion trajectory and segmenting the basic gesture action unit, more fine-grained gesture motion features can be extracted. According to the gesture motion feature vector data, the user's action intention is identified, such as "instruction", "grab", "put down", etc., so as to understand the core meaning that the user wants to express. This enables the system to dynamically adjust according to the user's real-time intention and generate a text description that is more in line with the user's expectations, rather than a simple instruction text. Furthermore, the system combines the action intention recognition results with the emotional adjective processing, and can identify the emotional color contained in the user's gesture, such as "quickly", "slowly", "forcefully", etc., so that the generated text description is more vivid and expressive. By semantically fusing the gesture motion text description data and the gesture action intention recognition data, key gesture semantic fusion data is generated, and a comprehensive understanding and description of the gesture action is achieved. This is not just about converting gestures into simple text labels, but about integrating the motion trajectory, action intention, and emotional expression of gestures to generate a more informative and expressive text description. Intelligent text generation and text template filling based on key gesture semantic fusion data can generate more complete, fluent, and natural text descriptions, such as "He quickly pointed to the red object in the distance" instead of simply "pointing to red." By visually rendering the intelligent description gesture text data, the text can be associated with the gesture action, such as displaying the text description of the corresponding gesture on the screen in real time, and dynamically adjusting the display position and method of the text according to the motion trajectory of the gesture. This visual feedback mechanism enables users to understand the system's recognition results more clearly, and also facilitates users to interact more effectively with the system.

[0010] Preferably, step S1 comprises the following steps:

[0011] Step S11: using a 3D depth camera to collect real-time hand depth images of the user, and perform initial image correction to obtain depth hand image sequence data;

[0012] Step S12: removing the background of the depth hand image sequence data and performing hand region segmentation to obtain user hand region data;

[0013] Step S13: performing hand posture estimation on the user's hand area data to obtain the user's hand posture data;

[0014] Step S14: Detect skeleton key points according to the user's hand posture data, and perform hand skeleton key point coordinate analysis to obtain skeleton key point coordinate data;

[0015] Step S15: extracting pixel coordinates of gesture key points from the depth hand image sequence data using the skeleton key point coordinate data to obtain pixel coordinate data of gesture image key points;

[0016] Step S16: performing time stamp synchronization processing according to the pixel coordinate data of the key points of the gesture image, and reconstructing the gesture trajectory according to the user's hand posture data to obtain spatial gesture motion trajectory data.

[0017] The present invention uses a 3D depth camera to collect real-time hand depth images and performs initial image correction to ensure the accuracy and reliability of the original data. Compared with traditional two-dimensional images, depth information can better reflect the true shape of the hand in three-dimensional space, avoiding errors caused by factors such as changes in viewing angles and lighting conditions. Initial image correction further eliminates the influence of factors such as camera distortion itself, ensuring the accuracy of subsequent processing. Hand posture estimation, skeleton key point detection and hand skeleton key point coordinate analysis provide key information for accurate reconstruction of gesture trajectory. Hand posture estimation can determine the direction and posture of the hand in space, providing a reference for subsequent key point detection. Skeleton key point detection further locates the position of key parts of the hand, such as fingertips, knuckles, etc. By analyzing the coordinate data of these key points, the shape and motion trajectory of the hand can be accurately described. Timestamp synchronization processing and gesture trajectory reconstruction based on user hand posture data ensure the temporal and spatial consistency and accuracy of the gesture trajectory. Timestamp synchronization ensures the temporal consistency of data between each step, so that the reconstructed gesture trajectory can accurately reflect the motion process of the gesture. The trajectory reconstruction is performed in combination with the hand posture data, and the rotation and displacement of the hand in space are further considered, so that the reconstructed spatial gesture motion trajectory data more accurately describes the actual movement of the gesture in three-dimensional space.

[0018] Preferably, step S2 comprises the following steps:

[0019] Step S21: analyzing the spatial gesture motion trajectory data for the length of the spatial gesture trajectory, and performing time regularization processing on the sampling trajectory to obtain pre-processed gesture motion trajectory data;

[0020] Step S22: performing gesture motion feature analysis based on the pre-processed gesture motion trajectory data to generate gesture motion feature vector data;

[0021] Step S23: Segment the basic gesture action unit according to the gesture motion feature vector data to generate basic gesture action unit data;

[0022] Step S24: annotating the gesture motion feature vector data with gesture motion semantic labels based on the basic gesture action unit data using a preset gesture-semantic mapping database to generate gesture motion text description data;

[0023] Step S25: performing gesture movement intention sample transfer learning according to a preset convolutional neural network model, thereby constructing a gesture movement intention recognition model;

[0024] Step S26: using the gesture motion intention recognition model to perform motion intention recognition on the basic gesture action unit data and the gesture motion feature vector data to obtain gesture action intention recognition data.

[0025] The present invention effectively preprocesses the original trajectory data by analyzing the length of gesture space trajectory and time-regularizing the sampling trajectory, thereby improving the standardization and availability of the data. When different users perform the same gesture, their speed and trajectory length are different. Through length analysis and time regularization, trajectory data of different lengths and speeds can be unified to the same standard. Gesture motion feature analysis converts the preprocessed trajectory data into gesture motion feature vector data, realizes the conversion from specific trajectory to abstract features, and lays the foundation for subsequent semantic understanding and intention recognition. These feature vectors can effectively capture the key information of gesture motion, such as motion direction, speed, amplitude, etc. Basic gesture action unit segmentation divides the continuous gesture motion trajectory into a series of action units with semantics, realizing the conversion from continuous motion to discrete semantic units. This segmentation enables the system to identify the basic action elements contained in complex gestures. For example, "waving" can be divided into two action units, "lifting" and "swinging", providing more fine-grained support for understanding the meaning of gestures.

[0026] Preferably, step S22 comprises the following steps:

[0027] Step S221: extracting the coordinates of the starting point / ending point of the gesture trajectory according to the pre-processed gesture motion trajectory data to obtain the trajectory starting and ending point data;

[0028] Step S222: Calculating the displacement vectors of the hand key points in adjacent frames according to the pre-processed gesture motion trajectory data to generate hand inter-frame displacement vector data;

[0029] Step S223: performing inter-frame instantaneous velocity vector calculation on the inter-frame displacement vector data of the hand to generate instantaneous velocity vector data;

[0030] Step S224: performing low-pass filtering according to the instantaneous velocity vector data, and performing gesture motion acceleration calculation to obtain frame-by-frame acceleration vector data;

[0031] Step S225: performing mean processing on the instantaneous velocity vector data and the frame-by-frame acceleration vector data through a preset time window to obtain velocity acceleration characteristic data;

[0032] Step S226: performing trajectory curve fitting on the pre-processed gesture motion trajectory data through the trajectory start and end point data, and performing gesture trajectory curvature and direction change analysis to obtain gesture trajectory geometric feature data;

[0033] Step S227: performing gesture posture change frequency analysis based on the pre-processed gesture motion trajectory data to obtain gesture posture change frequency data;

[0034] Step S228: performing gesture feature combination on the velocity acceleration feature data, the gesture trajectory geometric feature data and the gesture posture change frequency data to obtain gesture motion feature vector data.

[0035] The present invention describes the dynamic characteristics of gesture motion in detail by calculating the displacement vectors of the key points of the hand in adjacent frames, the instantaneous velocity vectors between frames, and the acceleration vectors frame by frame. The displacement vector reflects the direction and distance of the gesture in a short period of time, the instantaneous velocity vector further reveals the speed change of the motion, and the acceleration vector reflects the rate of speed change. These dynamic features are crucial for distinguishing different types of gestures. For example, fast waving and slow movement have significantly different speed and acceleration characteristics. The geometric features of the gesture trajectory are extracted through trajectory curve fitting, gesture trajectory curvature, and direction change analysis. These geometric features reflect the shape and direction change of gesture motion. For example, circular gestures and straight line gestures have different curvature characteristics, and waving to the left and waving to the right have different direction change characteristics.

[0036] Preferably, step S24 includes the following steps:

[0037] Step S231: performing key frame detection on the gesture motion feature vector data to obtain key frame data;

[0038] Step S232: segmenting the pre-processed gesture motion trajectory data into key frame time periods using the key frame data to obtain time period segmented trajectory data;

[0039] Step S233: performing hand gesture cluster analysis on the time segmentation trajectory data to obtain hand gesture cluster data;

[0040] Step S234: dividing the time segmentation trajectory data into basic gesture action units based on the gesture posture clustering data to generate basic gesture action unit division data;

[0041] Step S235: performing conversion relationship processing between gesture action units on the basic gesture action unit division data to obtain unit conversion relationship data;

[0042] Step S236: Integrate the basic gesture action unit division data and the unit conversion relationship data to generate basic gesture action unit data.

[0043] The present invention clusters and analyzes hand gestures, classifies similar hand gestures into one category, and divides the time segmentation trajectory data into basic gesture action units based on the hand gesture clustering data, so that the segmented action units have clear semantic boundaries. For example, "lift", "put down", "wave to the left", etc. can all be regarded as a basic action unit. This division method based on hand gestures makes the division of action units more in line with human perception and understanding of gestures. The conversion relationship between gesture action units is processed on the data divided into basic gesture action units, and the temporal relationship and logical relationship between action units are clarified, providing key information for understanding the overall meaning of complex gestures. For example, "lift" is usually followed by "put down" or "wave". By analyzing the conversion relationship between action units, the user's intention can be inferred, for example, continuous "click" actions represent "confirmation" or "selection".

[0044] Preferably, step S24 includes the following steps:

[0045] Step S241: determining the gesture motion intensity level of the speed acceleration feature data to generate motion intensity level data;

[0046] Step S242: mapping the exercise intensity level data to exercise semantic text labels using a preset gesture-semantic mapping database to obtain intense exercise semantic vocabulary data;

[0047] Step S243: performing gesture complexity level analysis on the gesture trajectory geometric feature data to obtain complexity level data;

[0048] Step S244: using a preset gesture-semantic mapping database to map the complexity level data to motion semantic text labels, and obtaining complex motion semantic vocabulary data;

[0049] Step S245: classifying the gesture rhythm speed level of the gesture posture change frequency data to obtain gesture rhythm level data;

[0050] Step S246: mapping the gesture rhythm level data to gesture rhythm text labels using a preset gesture-semantic mapping database to obtain gesture rhythm text vocabulary data;

[0051] Step S247: Based on the basic gesture action unit data, the violent motion semantic vocabulary data, the complex motion semantic vocabulary data and the gesture rhythm text vocabulary data are fused with gesture semantic description labels to obtain gesture motion text description data.

[0052] The present invention maps gesture motion feature data to different semantic dimensions by making level judgments on the intensity, complexity and rhythm speed of gesture velocity acceleration feature data, gesture trajectory geometric feature data and gesture posture change frequency data. This multi-dimensional analysis makes the description of gestures more comprehensive and detailed, such as "fast", "slow", "complex", "simple", "rhythmic", etc. Using a preset gesture-semantic mapping database, these level data are mapped to corresponding semantic text labels. This database is the key to achieving text description from feature data. It associates different feature levels with corresponding semantic vocabulary, such as "violent", "mild", "complex", "simple", "fast", "soothing", etc. Through mapping, the system can convert abstract feature data into specific semantic vocabulary.

[0053] Preferably, step S3 comprises the following steps:

[0054] Step S31: performing confidence evaluation on the gesture action intention recognition data, and screening high-confidence intentions to obtain high-confidence action intention data;

[0055] Step S32: performing action emotion adjective processing according to the high-confidence action intention data, and performing gesture semantic fusion according to the gesture motion text description data to generate key gesture semantic fusion data;

[0056] Step S33: performing intelligent text generation based on key gesture semantic fusion data to generate initial gesture text content data;

[0057] Step S34: performing gesture category analysis according to the basic gesture action unit data, and performing text template keyword processing to generate text template keyword data;

[0058] Step S35: using the text template keyword data to perform text template matching on a preset text template library to obtain selected text template data;

[0059] Step S36: The initial gesture text content data is semantically filled with placeholders by selecting text template data, thereby obtaining intelligent description gesture text data.

[0060] The present invention performs confidence evaluation on gesture action intention recognition data and selects high-confidence intention data, thereby ensuring that subsequent text generation is based on reliable intention understanding. This step eliminates erroneous recognition results by setting a confidence threshold, improves the accuracy of intention understanding, and avoids generating text descriptions that do not match user intentions. The high-confidence action intention data is processed with action emotional adjectives, and semantically fused with gesture movement text description data to generate key gesture semantic fusion data. This step further enriches the semantic information of the text description, especially incorporates emotional color, so that the generated text is not only an objective description of the action, but also can reflect the user's emotional state, such as "waving happily" or "putting down hands in frustration". The addition of such emotional information makes the text description more expressive and humane. Gesture category analysis is performed through basic gesture action unit data, and text template keyword processing is performed to generate text template keyword data, which provides a basis for the selection of text templates. Different categories of gestures, such as interaction, instruction, expression, etc., usually correspond to different text templates. Through keyword matching, a suitable text template can be quickly found. Combining the generated text content with the template ensures that the final text output has both personalized content and good readability and fluency.

[0061] Preferably, the present invention further provides a text generation system for executing the above-mentioned text generation method, the text generation system comprising:

[0062] The gesture trajectory reconstruction module is used to collect the user's hand depth image in real time using a 3D depth camera to obtain depth hand image sequence data; the gesture trajectory is reconstructed according to the depth hand image sequence data to obtain spatial gesture motion trajectory data;

[0063] The gesture semantic annotation module is used to analyze the gesture motion characteristics of the spatial gesture motion trajectory data and generate gesture motion characteristic vector data; segment the basic gesture action units according to the gesture motion characteristic vector data and generate basic gesture action unit data; annotate the gesture motion semantic labels of the gesture motion characteristic vector data based on the basic gesture action unit data and generate gesture motion text description data; construct a gesture motion intention recognition model; use the gesture motion intention recognition model to perform action intention recognition on the gesture motion characteristic vector data and obtain gesture action intention recognition data;

[0064] The intelligent gesture description text generation module is used to process the action emotion adjectives according to the gesture action intention recognition data, and to perform gesture semantic fusion according to the gesture movement text description data to generate key gesture semantic fusion data; based on the key gesture semantic fusion data, intelligent text generation is performed, and text template semantic filling is performed to obtain intelligent gesture description text data;

[0065] The description text visualization module is used to perform visual text rendering processing based on the intelligent description gesture text data to achieve the generation of intelligent description text for gesture actions.

[0066] Preferably, the present invention further provides a terminal device, the terminal device comprising:

[0067] processor;

[0068] a memory for storing processor-executable instructions;

[0069] Wherein, the processor is configured to implement any of the text generation methods described above.

[0070] Preferably, the present invention further provides a computer-readable storage medium storing a computer program, wherein the computer program implements any one of the above text generation methods when executed. BRIEF DESCRIPTION OF THE DRAWINGS

[0071] Figure 1 A schematic diagram of the steps of the text generation method of the present invention;

[0072] Figure 2 for Figure 1 Detailed implementation steps of step S1 in FIG.

[0073] Figure 3 for Figure 1 Detailed implementation steps of step S3 in FIG.

[0074] The realization of the purpose, functional features and advantages of the present invention will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0075] The technical method of the present invention is described clearly and completely below in conjunction with the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by technicians in this field without creative work are within the scope of protection of the present invention.

[0076] In addition, the accompanying drawings are only schematic illustrations of the present invention and are not necessarily drawn to scale. The same reference numerals in the figures represent the same or similar parts, and their repeated description will be omitted. Some of the block diagrams shown in the accompanying drawings are functional entities and do not necessarily correspond to physically or logically independent entities. The functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor methods and / or microcontroller methods.

[0077] It should be understood that, although the terms "first", "second", etc. may be used herein to describe various units, these units should not be limited by these terms. These terms are used only to distinguish one unit from another unit. For example, without departing from the scope of the exemplary embodiments, the first unit may be referred to as the second unit, and similarly the second unit may be referred to as the first unit. The term "and / or" used herein includes any and all combinations of one or more of the listed associated items.

[0078] To achieve this, please refer to Figures 1 to 3 The present invention provides a text generation method, comprising the following steps:

[0079] Step S1: using a 3D depth camera to collect real-time hand depth images of the user to obtain depth hand image sequence data; reconstructing gesture trajectories according to the depth hand image sequence data to obtain spatial gesture motion trajectory data;

[0080] Step S2: performing gesture motion feature analysis on the spatial gesture motion trajectory data to generate gesture motion feature vector data; performing basic gesture action unit segmentation according to the gesture motion feature vector data to generate basic gesture action unit data; annotating the gesture motion semantic label on the gesture motion feature vector data based on the basic gesture action unit data to generate gesture motion text description data; constructing a gesture motion intention recognition model; using the gesture motion intention recognition model to perform action intention recognition on the gesture motion feature vector data to obtain gesture action intention recognition data;

[0081] Step S3: Processing the action emotion adjectives according to the gesture action intention recognition data, and performing gesture semantic fusion according to the gesture motion text description data to generate key gesture semantic fusion data; performing intelligent text generation based on the key gesture semantic fusion data, and performing text template filling to obtain intelligent description gesture text data;

[0082] Step S4: Perform visual text rendering processing according to the intelligent description gesture text data to realize the generation of intelligent description text for gesture actions.

[0083] In an embodiment of the present invention, the text generation method comprises the following steps:

[0084] Step S1: using a 3D depth camera to collect real-time hand depth images of the user to obtain depth hand image sequence data; reconstructing gesture trajectories according to the depth hand image sequence data to obtain spatial gesture motion trajectory data;

[0085] In an embodiment of the present invention, a 3D depth camera is started, for example, using a device similar to structured light or time-of-flight principle, to scan the user's hand in real time. The camera captures a hand image containing depth information at a rate of 30 frames per second to generate a continuous sequence of depth images. Each depth image is a two-dimensional matrix, and each element in the matrix represents the distance from the corresponding pixel to the camera. For example, a pixel value of 50 means that the point is 50 centimeters away from the camera. First, the hand area in the depth image is identified by image processing technology, and commonly used methods include threshold segmentation and morphological processing. For example, a depth threshold is set, and the area less than 80 centimeters and greater than 20 centimeters from the camera is preliminarily identified as the hand area. After that, corrosion and expansion in morphological operations are used to remove noise and holes in the hand area. Next, a hand feature point needs to be determined, such as the center of the palm or the position of the fingertip. In the first frame of the image, the initial palm center can be determined by calculating the geometric center of the hand area. In each subsequent frame of the image, a tracking algorithm such as optical flow or Kalman filtering is used to track the position change of the feature point between different frames. For example, the Lucas-Kanade optical flow algorithm is used to calculate the displacement of feature points between adjacent frames, thereby updating the coordinates of feature points in the new frame image. In this way, by tracking the position of feature points frame by frame, a point sequence in three-dimensional space is obtained. This point sequence is the gesture motion trajectory data, which is represented in the form of a set of three-dimensional coordinate points. For example: [(x1, y1, z1), (x2, y2, z2), ..., (xn, yn, zn)], where (xi, yi, zi) represents the three-dimensional coordinates of the hand feature points in the i-th frame image.

[0086] Step S2: performing gesture motion feature analysis on the spatial gesture motion trajectory data to generate gesture motion feature vector data; performing basic gesture action unit segmentation according to the gesture motion feature vector data to generate basic gesture action unit data; annotating the gesture motion semantic label on the gesture motion feature vector data based on the basic gesture action unit data to generate gesture motion text description data; constructing a gesture motion intention recognition model; using the gesture motion intention recognition model to perform action intention recognition on the gesture motion feature vector data to obtain gesture action intention recognition data;

[0087] In an embodiment of the present invention, features of multiple dimensions are extracted from the trajectory data, including motion speed, direction, acceleration and curvature. For example, for each point in the trajectory data, the distance between it and the previous point is calculated and divided by the time interval between two frames to obtain the instantaneous speed of the point. Then, the angle between the velocity vectors of three adjacent points is calculated to obtain the curvature of the trajectory at the point. The above operations are performed on each point on the trajectory to obtain a velocity sequence and a curvature sequence. Similarly, features such as acceleration and direction change can also be calculated. Then, these feature sequences are statistically analyzed, such as calculating the mean, variance, maximum, minimum, and the rate of change of these values. These statistical values ​​are combined into a vector, which is the gesture motion feature vector data. The continuous gesture motion is decomposed into a series of discrete basic action units with semantic meaning. For example, a complete "circle drawing" gesture can be divided into four basic action units of "upward movement", "rightward movement", "downward movement" and "leftward movement". The basis for segmentation can be the change trend of each feature in the feature vector, such as a sudden increase or decrease in speed, a significant change in direction, etc. Some rules or thresholds can be set. When the feature value exceeds or falls below these thresholds, it is considered that one basic action unit ends and another basic action unit begins. For example, when the speed change rate exceeds a certain threshold, or the direction change angle is greater than 45 degrees, segmentation is performed. Each basic action unit after segmentation corresponds to a feature vector segment, and the collection of these segments is the basic gesture action unit data. Based on these basic gesture action unit data, the original gesture motion feature vector data is annotated, and the semantic meaning of each basic action unit is described with text labels. For example, "upward movement", "right movement", "circle", "wave", etc. These text labels are associated with the corresponding feature vector segments to form gesture motion text description data. Construct a gesture motion intention recognition model. Use deep learning methods, such as recurrent neural networks (RNN) or long short-term memory networks (LSTM), to train a classification model. Use the annotated gesture motion text description data as a training set. For example, prepare hundreds of sample data of different types of gesture actions, and annotate each sample with its corresponding intention, such as "select", "confirm", "cancel", etc. Use the trained gesture motion intention recognition model to perform intent recognition on new, unlabeled gesture motion feature vector data. Input the feature vector of the new data into the model, and the model will output an intent label, such as "confirm", "zoom in", "zoom out", etc. This output is the gesture action intention recognition data.

[0088] Step S3: Processing the action emotion adjectives according to the gesture action intention recognition data, and performing gesture semantic fusion according to the gesture motion text description data to generate key gesture semantic fusion data; performing intelligent text generation based on the key gesture semantic fusion data, and performing text template filling to obtain intelligent description gesture text data;

[0089] In an embodiment of the present invention, a library of emotional adjectives is predefined, such as "quickly", "slowly", "forcefully", "gently", "fluently", "austerely", etc. According to the speed, acceleration and other features in the gesture motion feature vector data, appropriate emotional adjectives are selected. For example, if the average speed of a "waving" action is high and the acceleration is large, adjectives such as "quickly" and "forcefully" are selected to modify it. If the curvature of the trajectory of a "drawing a circle" action changes smoothly, the adjective "fluently" is selected. The selected emotional adjective is combined with the gesture action intention recognition data, for example, "select" is combined with "quickly" to obtain "quickly select". For example, for a gesture that is divided into two basic action units of "upward movement" and "rightward movement", the corresponding text description data is "upward" and "rightward". If the intention recognition data is "draw" and the emotional adjective processing result is "slowly", then the result of semantic fusion is "slowly draw upwards, then draw to the right". Using natural language generation technology, such as rule-based template filling or neural network-based sequence-to-sequence model, key semantic descriptions are used as input to generate more complete sentences. For example, for the key semantic description of "slowly drawing upwards, then drawing to the right", "the user is slowly moving his hand upwards to draw, and then moving his hand to the right to continue drawing" can be generated. If the template filling method is adopted, some sentence templates can be pre-defined, such as "the user is [emotional adjective][action intention][basic action unit description]". Then fill each part of the key semantic fusion data into the corresponding position in the template. For example, some connecting words can be added, some background information can be supplemented, etc., to make the generated text smoother and easier to understand. For example, "the drawing of the figure is completed through this operation" can be added to the text generated in the previous step to obtain the final intelligent description gesture text data: "the user is slowly moving his hand upwards to draw, and then moving his hand to the right to continue drawing, and the drawing of the figure is completed through this operation".

[0090] Step S4: Perform visual text rendering processing according to the intelligent description gesture text data to realize the generation of intelligent description text for gesture actions.

[0091] In an embodiment of the present invention, a suitable font and font size are selected. For example, a clear and easy-to-read Microsoft Yahei font is selected, and the font size is set to 18. Then, the display position of the text on the screen is determined. For example, the text can be displayed in the lower center of the screen, 5 cm from the bottom of the screen, and aligned horizontally in the center. Next, the color and background color of the text are set. For example, white text and a black semi-transparent background are used to ensure that the text can be clearly seen under various backgrounds. Automatic line wrapping is performed according to the length of the text and the width of the screen. For example, a maximum of 30 characters are displayed per line, and more than 30 characters are automatically switched to the next line. Then, some animation effects can be added, such as fade in and fade out, word by word display, etc., to make the presentation of the text more vivid. For example, the text can be set to be displayed word by word at a speed of 5 words per second, and the display state is maintained after the display is completed. The processed text is rendered on the screen. For example, a graphics library such as OpenGL or DirectX can be used to render the text into a texture, and then the texture is pasted on a rectangular area, and finally the rectangular area is drawn to a specified position on the screen. Through the above steps, the intelligent description text data of gestures is presented in a user-friendly manner, and the generation of intelligent description text of gesture actions is completed.

[0092] Preferably, step S1 comprises the following steps:

[0093] Step S11: using a 3D depth camera to collect real-time hand depth images of the user, and perform initial image correction to obtain depth hand image sequence data;

[0094] Step S12: removing the background of the depth hand image sequence data and performing hand region segmentation to obtain user hand region data;

[0095] Step S13: performing hand posture estimation on the user's hand area data to obtain the user's hand posture data;

[0096] Step S14: Detect skeleton key points according to the user's hand posture data, and perform hand skeleton key point coordinate analysis to obtain skeleton key point coordinate data;

[0097] Step S15: extracting pixel coordinates of gesture key points from the depth hand image sequence data using the skeleton key point coordinate data to obtain pixel coordinate data of gesture image key points;

[0098] Step S16: performing time stamp synchronization processing according to the pixel coordinate data of the key points of the gesture image, and reconstructing the gesture trajectory according to the user's hand posture data to obtain spatial gesture motion trajectory data.

[0099] As an example of the present invention, refer to Figure 2 As shown, Figure 1Detailed implementation steps of step S1 in the flowchart, in this example, step S1 includes:

[0100] Step S11: using a 3D depth camera to collect real-time hand depth images of the user, and perform initial image correction to obtain depth hand image sequence data;

[0101] In an embodiment of the present invention, a 3D depth camera is started, such as a depth sensor using the time-of-flight principle, whose transmitter emits infrared light pulses of a specific wavelength to the scene, and the receiver measures the arrival time of the light pulse reflected by each pixel. By calculating the round-trip time of the light pulse and combining it with the speed of light, the distance from each pixel to the camera can be obtained, thereby constructing a depth image. The camera continuously collects depth images at a rate of 60 frames per second, and each frame of the image is a matrix with a resolution of 640x480. Each element in the matrix represents the distance of the pixel from the camera in millimeters. For example, if the value of a pixel is 1500, it means that the point is 1.5 meters away from the camera. The collected data constitutes a depth image sequence. For example, the coordinates of each pixel are corrected using the pre-calibrated camera internal parameters and distortion coefficients. The calibration process can be completed by shooting a calibration plate of known shape and size, such as a chessboard. Assume that the focal length of the camera obtained by calibration is 500 pixels, and the distortion coefficients are k1=0.1, k2=-0.2. For the pixel point with coordinates (u, v) in the depth image, first convert it to normalized coordinates (x, y), that is, x = (u-cx) / fx, y = (v-cy) / fy, where cx and cy are the image center coordinates, and fx and fy are the focal lengths. Then, use the distortion model x_corrected = x(1+k1×r^2+k2×r^4), y_corrected = y(1+k1×r^2+k2×r^4) to calculate the corrected normalized coordinates, where r^2 = x^2+y^2. Finally, convert the corrected normalized coordinates back to pixel coordinates to obtain the corrected pixel coordinates (u_corrected, v_corrected).

[0102] Step S12: removing the background of the depth hand image sequence data and performing hand region segmentation to obtain user hand region data;

[0103] In an embodiment of the present invention, for example, assuming that the user usually performs gesture operations within a range of 0.5 meters to 1.5 meters from the camera, pixels with depth values ​​less than 500 mm or greater than 1500 mm can be regarded as background pixels. In each frame of the depth image, all pixels are traversed, and the depth values ​​of pixels whose depth values ​​are not within the range are set to 0, indicating that these pixels are background. In this way, a depth image after removing the background is obtained. The depth image after removing the background is binarized. For example, the values ​​of pixels with depth values ​​greater than 0 are set to 255, indicating that these pixels are foreground, that is, the hand area, and the values ​​of the remaining pixels remain at 0, indicating the background. Then, a connected area is found in the binary image. For example, using the eight-connected algorithm, starting from any foreground pixel, all foreground pixels in the eight surrounding neighborhoods are marked as the same category, and then the same operation is recursively performed on these newly marked pixels until no new pixels can be marked. In this way, a connected area is found. Repeat the above steps until all foreground pixels are marked. Usually, the connected area containing the hand is the largest. Therefore, the connected region with the largest area is selected as the hand region.

[0104] Step S13: performing hand posture estimation on the user's hand area data to obtain the user's hand posture data;

[0105] In an embodiment of the present invention, a hand model is defined. For example, a simplified hand model can be used to divide the hand into a palm and five fingers, each finger being connected by three joint points. The model can be represented by a set of parameters, such as the three-dimensional coordinates of the center of the palm, the normal vector of the palm, the length of each finger and the joint angle, etc. According to the shape and depth information of the hand area, the initial values ​​of the hand model parameters are preliminarily estimated. For example, the centroid of the hand area can be calculated as the initial position of the center of the palm, and the main axis direction of the minimum circumscribed rectangle of the hand area can be calculated as the initial direction of the palm normal vector. Then, the hand model is rendered into a depth image according to the current parameters, which is called a model depth map. The distance between the model depth map and the depth map corresponding to the user hand area data is calculated. For example, for each pixel in the model depth map, the nearest point is found in the user hand area depth map, and the distance between the two points is calculated. The average value of all these distances is taken as the matching error under the current parameters. Next, the parameters of the hand model are adjusted to minimize the matching error. For example, the gradient descent method can be used to update the parameters along the negative gradient direction of the error function. Repeat the above steps until the matching error converges to a smaller value or reaches the preset maximum number of iterations. The final hand model parameters are the estimated user hand posture data.

[0106] Step S14: Detect skeleton key points according to the user's hand posture data, and perform hand skeleton key point coordinate analysis to obtain skeleton key point coordinate data;

[0107] In an embodiment of the present invention, for example, the hand model used in the previous step is used, and the model includes 16 joint points, which correspond to the wrist, palm, four fingertips and three joints of each finger. According to the hand posture data, the coordinates of each joint point in three-dimensional space can be calculated. For example, given the palm center coordinates, palm normal vector, and the length and joint angle of each finger, the position of each joint point relative to the palm center can be obtained by simple geometric calculation. Then, these relative positions are added to the coordinates of the palm center to obtain the three-dimensional coordinates of each joint point in the coordinate system. The joint angle is obtained by calculating the vector between two adjacent joint points and then calculating the angle between the two vectors. For example, the coordinates of the three joint points of a finger are A(150,200,900), B(160,210,850), and C(170,230,800), respectively, then vector AB=BA=(10,10,-50), and vector BC=CB=(10,20,-50). The angle between these two vectors can be calculated using the inverse cosine function, such as arccos((AB·BC) / (|AB|×|BC|)), where · represents the dot product of the vectors and |AB| represents the modulus of vector AB. In this way, the joint angles of all fingers can be calculated. The distances between the fingertips of different fingers can also be calculated, such as the Euclidean distance between the index fingertip and the middle fingertip, forming the coordinate data of the skeleton key points.

[0108] Step S15: extracting pixel coordinates of gesture key points from the depth hand image sequence data using the skeleton key point coordinate data to obtain pixel coordinate data of gesture image key points;

[0109] In an embodiment of the present invention, the three-dimensional coordinates of the skeleton key points are converted to the camera coordinate system. Since the previously obtained coordinates of the skeleton key points are expressed in the world coordinate system, and the depth image is expressed in the camera coordinate system, a coordinate transformation is required. For example, the world coordinates P_world of the skeleton key points can be converted to the coordinates P_camera in the camera coordinate system by pre-calibrated camera external parameters, that is, the rotation matrix R and translation vector T of the camera in the world coordinate system: P_camera = R × P_world + T. Projection is performed using a pinhole camera model. Assume that the intrinsic parameters of the camera are known, including the focal length fx, fy and the principal point coordinates cx, cy. For a three-dimensional point P_camera = (X, Y, Z) in the camera coordinate system, its projection coordinates (u, v) on the image plane can be calculated by the following formula: u = fx × X / Z + cx, v = fy × Y / Z + cy. For each skeleton key point, the above-mentioned coordinate transformation and projection operations are performed in each frame of the depth image to obtain the pixel coordinates of each skeleton key point in each frame of the image. For example, for a certain fingertip key point, its projection coordinates in the first frame image are (300, 200), and its projection coordinates in the second frame image are (305, 205), and so on. These pixel coordinates constitute the pixel coordinate data of the key point of the gesture image.

[0110] Step S16: performing time stamp synchronization processing according to the pixel coordinate data of the key points of the gesture image, and reconstructing the gesture trajectory according to the user's hand posture data to obtain spatial gesture motion trajectory data.

[0111] In an embodiment of the present invention, the timestamps of all data frames are unified into a reference time, and interpolation or extrapolation is performed according to the time difference. Then, the gesture trajectory is reconstructed according to the pixel coordinate data of the key points of the gesture image after the timestamp synchronization and the user hand posture data obtained in step S13. The specific method is to convert the pixel coordinates of the key points of the hand in each frame image into three-dimensional space coordinates, and connect the key point coordinates of adjacent frames to form a gesture motion trajectory. For example, by connecting the three-dimensional coordinates of the fingertip of the index finger in several consecutive frames of images, the motion trajectory of the fingertip of the index finger can be obtained. Finally, the spatial gesture motion trajectory data is obtained, which contains the coordinate sequence of the key points of the hand in three-dimensional space and is stored with time as the index. For example, a gesture trajectory data containing 10 frames can be represented as a 10×21×3 tensor, where 10 represents the number of frames, 21 represents the number of key points, and 3 represents the three-dimensional coordinates.

[0112] Preferably, step S2 comprises the following steps:

[0113] Step S21: analyzing the spatial gesture motion trajectory data for the length of the spatial gesture trajectory, and performing time regularization processing on the sampling trajectory to obtain pre-processed gesture motion trajectory data;

[0114] Step S22: performing gesture motion feature analysis based on the pre-processed gesture motion trajectory data to generate gesture motion feature vector data;

[0115] Step S23: Segment the basic gesture action unit according to the gesture motion feature vector data to generate basic gesture action unit data;

[0116] Step S24: annotating the gesture motion feature vector data with gesture motion semantic labels based on the basic gesture action unit data using a preset gesture-semantic mapping database to generate gesture motion text description data;

[0117] Step S25: performing gesture movement intention sample transfer learning according to a preset convolutional neural network model, thereby constructing a gesture movement intention recognition model;

[0118] Step S26: using the gesture motion intention recognition model to perform motion intention recognition on the basic gesture action unit data and the gesture motion feature vector data to obtain gesture action intention recognition data.

[0119] In an embodiment of the present invention, the spatial gesture motion trajectory data obtained in step S1 is subjected to trajectory length analysis. The length of each gesture trajectory is calculated, that is, the sum of the distances between all adjacent points in the trajectory. For example, for a gesture trajectory containing 10 three-dimensional coordinate points, the trajectory length is the sum of the lengths of 9 line segments. Then, in order to eliminate the influence of different trajectory lengths caused by speed differences when different users or the same user perform the same gesture at different times, it is necessary to perform sampling trajectory time warping processing, also known as dynamic time warping (DTW). The DTW algorithm can align trajectories of different lengths and find the best matching path between them. Specifically, a distance matrix is ​​constructed, and the matrix elements are the Euclidean distances between corresponding points in two trajectories. Then, a dynamic programming algorithm is used to find a path from the upper left corner of the matrix to the lower right corner, on which the cumulative distance is the smallest and satisfies the monotonicity and continuity constraints. Through this path, the two trajectories can be aligned to the same length. For example, a 10-frame trajectory and a 12-frame trajectory are aligned to 11 frames, and the deformation between the corresponding frames is calculated. The motion speed of each trajectory point is calculated. For example, for the preprocessed trajectory data P = {(x1, y1, z1), (x2, y2, z2), ..., (xN, yN, zN)}, each point corresponds to a timestamp t. Assuming that the timestamps are uniform, that is, the time intervals between adjacent points are the same, the speed of each point can be calculated: v(i) = sqrt((xi-xi-1)^2+(yi-yi-1)^2+(zi-zi-1)^2) / (ti-ti-1), where i ranges from 2 to N. For example, in the trajectory data of a "waving" action, the speed of some points reaches 10 cm per second, while in the trajectory data of a "slowly drawing a circle" action, the speed of all points does not exceed 2 cm per second. Calculate the movement direction of each trajectory point. For example, for each point i, the displacement vector between it and the previous point i-1 can be calculated: ΔP(i) = (xi-xi-1, yi-yi-1, zi-zi-1). Then the displacement vector is normalized to obtain the unit direction vector: u(i) = ΔP(i) / |ΔP(i)|, where |ΔP(i)| represents the modulus of the vector ΔP(i). This unit direction vector u(i) represents the direction of motion of the trajectory at point i. The extracted features are quantized and combined into a fixed-length feature vector. For example, each feature can be quantized into 10 levels, and all features are combined into a 50-dimensional feature vector. Each dimension represents a specific feature, and its value represents the strength of the feature. For example, a gesture representing "moving upward quickly" has a high speed feature value and a direction feature value close to 90 degrees. Finally, the gesture motion feature vector data is generated. Define a set of basic gesture action units, such as "open palm", "clenched fist", "pointing finger", etc. Then, use the labeled gesture data to train the HMM model to learn the feature vector distribution corresponding to each basic gesture action unit.The trained HMM model can identify the most hidden state sequence, that is, the basic gesture action unit sequence, based on the input feature vector sequence. For example, a gesture sequence containing two actions, "open palm" and "clench fist", its feature vector sequence will be divided into two parts by the HMM model, corresponding to the two basic gesture action units of "open palm" and "clench fist". The basic gesture action unit data contains a series of gesture action fragments, each of which corresponds to a feature vector data and a time range. The gesture-semantic mapping database is a pre-established database that stores the correspondence between the feature vector data and semantic labels of various gesture action units. For example, the database stores semantic labels such as "upward movement", "left movement", "circle drawing", "waving", and the feature vector data corresponding to these labels. The labeling process is to match the feature vector data of each basic gesture action unit obtained in step S23 in the gesture-semantic mapping database, find the most similar feature vector data, and assign the corresponding semantic label to the basic gesture action unit. For example, for a basic gesture action unit, its feature vector data is V. Find the feature vector data V' that is closest to V in the database. Assuming that the semantic label corresponding to V' is "upward movement", then assign the label "upward movement" to the basic gesture action unit. Arrange the semantic labels of all action units in chronological order to generate gesture motion text description data. Use a public large-scale image dataset (such as ImageNet) to pre-train ResNet-50 so that it can learn rich image feature representations. Then, freeze the convolution layer parameters of the pre-trained model and only train the fully connected layer parameters. Use the labeled gesture action intention dataset to fine-tune the model to adapt it to the task of gesture action intention recognition. For example, use a gesture dataset containing action categories such as "grab", "put down", and "move" for training. During the training process, use the cross entropy loss function as the objective function and use the Adam optimizer to update the parameters. After the training is completed, a gesture motion intention recognition model that can map gesture motion feature vectors to action intention categories is obtained. Input the feature vector of the basic gesture action unit into the model to obtain the intention probability distribution of each action unit. For example, the intention probability distribution of an action unit of "opening the palm and moving toward the object" is: grab (80%), put down (10%), move (10%). Select the intention with the highest probability as the recognition result of the action unit. Arrange the intention recognition results of all action units in chronological order to obtain gesture action intention recognition data. For example, a gesture sequence containing three action units has the intention recognition results of "grab", "move", and "put down". Finally, the gesture action intention recognition data is obtained.

[0120] Preferably, step S22 includes the following steps:

[0121] Step S221: extracting the coordinates of the starting point / ending point of the gesture trajectory according to the pre-processed gesture motion trajectory data to obtain the trajectory starting and ending point data;

[0122] Step S222: Calculating the displacement vectors of the hand key points in adjacent frames according to the pre-processed gesture motion trajectory data to generate hand inter-frame displacement vector data;

[0123] Step S223: performing inter-frame instantaneous velocity vector calculation on the inter-frame displacement vector data of the hand to generate instantaneous velocity vector data;

[0124] Step S224: performing low-pass filtering according to the instantaneous velocity vector data, and performing gesture motion acceleration calculation to obtain frame-by-frame acceleration vector data;

[0125] Step S225: performing mean processing on the instantaneous velocity vector data and the frame-by-frame acceleration vector data through a preset time window to obtain velocity acceleration characteristic data;

[0126] Step S226: performing trajectory curve fitting on the pre-processed gesture motion trajectory data through the trajectory start and end point data, and performing gesture trajectory curvature and direction change analysis to obtain gesture trajectory geometric feature data;

[0127] Step S227: performing gesture posture change frequency analysis based on the pre-processed gesture motion trajectory data to obtain gesture posture change frequency data;

[0128] Step S228: performing gesture feature combination on the velocity acceleration feature data, the gesture trajectory geometric feature data and the gesture posture change frequency data to obtain gesture motion feature vector data.

[0129] In an embodiment of the present invention, the preprocessed gesture motion trajectory data is a time series, in which each time point corresponds to a three-dimensional coordinate. The coordinate of the starting point of the trajectory is the first coordinate of the time series, and the coordinate of the ending point is the last coordinate of the time series. For example, a gesture trajectory data containing 10 frames, the coordinate of the starting point is the three-dimensional coordinate of the first frame, and the coordinate of the ending point is the three-dimensional coordinate of the tenth frame. For each frame, the three-dimensional coordinate difference of each key point between the current frame and the previous frame is calculated to obtain a three-dimensional displacement vector. For example, the displacement vector of a key point in the second frame is the difference between the three-dimensional coordinates of the key point in the second frame and the first frame. The displacement vectors of all key points are combined into a matrix to generate the inter-frame displacement vector data of the hand. For example, a gesture containing 21 key points, the inter-frame displacement vector data is a 21x3 matrix, in which each row represents the displacement vector of a key point. Since the time interval between adjacent frames is fixed (for example, 1 / 30 seconds), the displacement vector can be divided by the time interval to obtain the instantaneous velocity vector. For example, if the time interval between adjacent frames is 0.033 seconds, the displacement vector obtained in step S222 is divided by 0.033 to obtain the instantaneous velocity vector. The instantaneous velocity vectors of each frame are combined into a matrix to generate instantaneous velocity vector data. A Butterworth filter or other types of low-pass filters can be used. The filtered data can more accurately reflect the movement trend of the gesture. Then, the change in the instantaneous velocity vector between adjacent frames is calculated to obtain a frame-by-frame acceleration vector. The calculation method is similar to step S223, that is, the difference between the instantaneous velocity vectors of adjacent frames is divided by the time interval. For example, the difference between the instantaneous velocity vectors of the third frame and the second frame is divided by the time interval to obtain the acceleration vector of the third frame. For example, the size of the time window can be set to 5 frames, that is, the data of the current moment and the 2 frames before and after it are considered. For each time window, the average value of the instantaneous velocity vector and the acceleration vector of all frames in the window is calculated as the velocity and acceleration characteristics of the time window. For example, if the time window is 0.1 seconds and the frame rate is 30 frames per second, each time window contains 3 frames of data, and the average of the velocity and acceleration of these 3 frames is calculated. Cubic spline interpolation or other curve fitting methods can be used. The fitted curve can more smoothly represent the gesture trajectory. Then, the curvature and direction change analysis of the fitted curve is performed. Curvature indicates the degree of curvature of the curve, which can be calculated using the curvature formula. Direction change indicates the rate of change of the tangent direction angle of the curve. The curvature and direction changes are quantified and combined into a feature vector to obtain the geometric feature data of the gesture trajectory. For example, several key joints, such as the fingertip joints of the index finger, middle finger, and thumb, can be selected to calculate the distance or angle between them. Then, the change in these distances or angles between adjacent frames is calculated.For example, for the distance d between the tips of the index finger and the middle finger, calculate the absolute value of the difference between the two adjacent frames: Δd = |d(i)-d(i-1)|, where d(i) represents the distance between the tips of the index finger and the middle finger in the i-th frame. For the angle, the absolute value of the angle difference between two adjacent frames can be calculated. Count the frequency of changes in these changes over a period of time. For example, a time window can be set, such as 10 frames. For each time window, count the number of times the change exceeds a certain threshold. For example, set the threshold for distance change to 2 cm and the threshold for angle change to 10 degrees. In a 10-frame time window, if the number of times the distance change Δd exceeds 2 cm is 6 times, then the distance change frequency in the time window is 6 / 10 = 0.6. Similarly, the angle change frequency can be calculated. The change frequencies in all time windows are calculated to form gesture posture change frequency data. The velocity acceleration feature data obtained in step S225, the gesture trajectory geometric feature data obtained in step S226, and the gesture posture change frequency data obtained in step S227 are combined to obtain gesture motion feature vector data. The velocity acceleration feature data includes the velocity and acceleration information of the gesture motion, the gesture trajectory geometric feature data includes the shape information of the gesture trajectory, and the gesture posture change frequency data includes the change information of the gesture posture.

[0130] Preferably, step S24 includes the following steps:

[0131] Step S231: performing key frame detection on the gesture motion feature vector data to obtain key frame data;

[0132] Step S232: segmenting the pre-processed gesture motion trajectory data into key frame time periods using the key frame data to obtain time period segmented trajectory data;

[0133] Step S233: performing hand gesture cluster analysis on the time segmentation trajectory data to obtain hand gesture cluster data;

[0134] Step S234: dividing the time segmentation trajectory data into basic gesture action units based on the gesture posture clustering data to generate basic gesture action unit division data;

[0135] Step S235: performing conversion relationship processing between gesture action units on the basic gesture action unit division data to obtain unit conversion relationship data;

[0136] Step S236: Integrate the basic gesture action unit division data and the unit conversion relationship data to generate basic gesture action unit data.

[0137] In an embodiment of the present invention, key frame detection is performed based on the rate of change of feature vectors. For example, the Euclidean distance between feature vectors of adjacent frames can be calculated as the rate of change of features. For example, for a gesture motion feature vector sequence F = {F1, F2, ..., FN}, Fi represents the feature vector of the i-th frame. Calculate the distance between the feature vectors of two adjacent frames: d(i, i+1) = ||Fi+1-Fi||, where ||.|| represents the Euclidean norm of the vector. Compare these distances with a preset threshold. If the distance is greater than the threshold, the corresponding frame is considered to be a key frame. For example, the threshold can be set to twice the average value of all distances. For the calculated distance sequence d = {d(1,2), d(2,3), ..., d(N-1, N)}, calculate its average value d_avg. If d(i, i+1)>2×d_avg, the i+1th frame is marked as a key frame. The time period between two adjacent key frames is regarded as a segmentation unit. For example, if three key frames are detected, located at the 5th, 10th and 15th frames respectively, the gesture trajectory is segmented into three time periods: from the 1st to the 5th frame, from the 6th to the 10th frame, and from the 11th to the 15th frame. Each time period represents a relatively stable gesture posture or a complete action unit. The segmented time period segmentation trajectory data contains the start frame number, the end frame number and the trajectory data in each time period. Feature extraction is performed on the trajectory data in each time period, such as calculating the average speed, the average acceleration, the trajectory length, etc. Then, a clustering algorithm, such as the K-Means algorithm or the DBSCAN algorithm, is used to cluster time periods with similar features together. Each cluster represents a specific gesture posture. For example, all time periods representing "open palm" are clustered into one category, and all time periods representing "clenched fist" are clustered into another category. The clustering result generates gesture posture clustering data, which contains the cluster category to which each time period belongs. Each cluster category is regarded as a basic gesture action unit. For example, all time periods corresponding to the "open palm" cluster are divided into a basic gesture action unit, and all time periods corresponding to the "clenched fist" cluster are divided into another basic gesture action unit. The division result generates basic gesture action unit division data, which includes all time periods corresponding to each basic gesture action unit. The basic gesture action units to which adjacent time periods belong are analyzed to determine the conversion relationship between the units. For example, if one time period belongs to the "open palm" unit and the next time period belongs to the "clenched fist" unit, it is considered that there is a conversion relationship from "open palm" to "clenched fist". All conversion relationships are recorded to generate unit conversion relationship data. For example, a directed graph can be used to represent the conversion relationship between units. The nodes in the graph represent the basic gesture action units, and the edges represent the conversion relationship between units.The basic gesture action unit division data generated in step S234 and the unit conversion relationship data generated in step S235 are integrated to generate the final basic gesture action unit data. The data includes information about each basic gesture action unit, such as the unit's start time, end time, corresponding trajectory data, and conversion relationships with other units.

[0138] Preferably, step S24 includes the following steps:

[0139] Step S241: determining the gesture motion intensity level of the speed acceleration feature data to generate motion intensity level data;

[0140] Step S242: mapping the exercise intensity level data to exercise semantic text labels using a preset gesture-semantic mapping database to obtain intense exercise semantic vocabulary data;

[0141] Step S243: performing gesture complexity level analysis on the gesture trajectory geometric feature data to obtain complexity level data;

[0142] Step S244: using a preset gesture-semantic mapping database to map the complexity level data to motion semantic text labels, and obtaining complex motion semantic vocabulary data;

[0143] Step S245: classifying the gesture rhythm speed level of the gesture posture change frequency data to obtain gesture rhythm level data;

[0144] Step S246: mapping the gesture rhythm level data to gesture rhythm text labels using a preset gesture-semantic mapping database to obtain gesture rhythm text vocabulary data;

[0145] Step S247: Based on the basic gesture action unit data, the violent motion semantic vocabulary data, the complex motion semantic vocabulary data and the gesture rhythm text vocabulary data are fused with gesture semantic description labels to obtain gesture motion text description data.

[0146] In an embodiment of the present invention, the intensity level of the movement is determined according to the average values ​​of the speed and acceleration. For example, several thresholds can be pre-set, and the average speed and the average acceleration are divided into several intervals, each of which corresponds to an intensity level. For example, the thresholds of the average speed can be set to 5 cm / s and 15 cm / s, and the average speed is divided into three intervals: less than 5 cm / s, 5-15 cm / s and greater than 15 cm / s, corresponding to the three intensity levels of "mild", "medium" and "severe", respectively. Similarly, the threshold of the average acceleration can be set, divided into several intervals, and corresponded to different intensity levels. For each gesture action, the intensity level to which it belongs is determined according to the average speed and the average acceleration in its speed acceleration feature data. For example, if the average speed of a gesture action is 10 cm / s and the average acceleration is 20 cm / s^2, then according to the above thresholds, its speed intensity level is "medium" and the acceleration intensity level is also "medium". The speed intensity level and the acceleration intensity level are combined as the final exercise intensity level. For example, if the speed intensity level and the acceleration intensity level are consistent, then this level is used as the final motion intensity level. If they are inconsistent, the final motion intensity level can be determined according to a pre-set rule. For example, the acceleration intensity level can be given priority, or the higher of the two levels can be used as the final motion intensity level. For example, in the above example, since the intensity levels of speed and acceleration are both "medium", the motion intensity level of the gesture action is also "medium". The motion intensity level data generated in step S241 is mapped to motion semantic text labels using a preset gesture-semantic mapping database. The semantic text labels corresponding to different motion intensity levels are pre-defined in the gesture-semantic mapping database. For example, "slight" corresponds to "slowly", "medium" corresponds to "steadily", and "violent" corresponds to "quickly". According to the motion intensity level data, the corresponding semantic text labels are searched from the database to obtain the violent motion semantic vocabulary data. For example, a sequence containing three basic gesture action units, whose motion intensity level data is "slight, violent, medium", then the corresponding violent motion semantic vocabulary data is "slowly, quickly, steadily". According to the curvature of the trajectory, direction change and other characteristics, the complexity of gestures is divided into different levels, such as "simple", "medium" and "complex". The division can be performed using a predefined threshold or a machine learning-based model. For example, the complexity can be judged based on the total curvature and direction change of the trajectory. The higher the value, the higher the complexity. The analysis results generate complexity level data, and each element corresponds to the complexity level of a basic gesture action unit. For the complexity level data, the gesture-semantic mapping database stores the semantic labels corresponding to each complexity level.For example, the semantic label corresponding to the complexity level of "simple" can be "simply", "directly", etc., the semantic label corresponding to the complexity level of "medium" can be "moderately complex", "skillfully", etc., and the semantic label corresponding to the complexity level of "complex" can be "complex", "finely", etc. The mapping process is to search for the corresponding semantic label in the gesture-semantic mapping database according to the complexity level of each gesture action. For example, if the complexity level of a gesture action is "complex", then the semantic label corresponding to "complex" is searched in the database, and labels such as "complex" and "finely" are found. The gesture rhythm speed level is divided into gesture posture change frequency data (step S227). According to the frequency of gesture posture change, the gesture rhythm is divided into different levels, such as "slow", "medium speed" and "fast". Predefined thresholds can be used for division, for example, the frequency below threshold A is "slow", the frequency between threshold A and threshold B is "medium speed", and the frequency above threshold B is "fast". The division result generates gesture rhythm level data. The gesture rhythm text label mapping is performed on the gesture rhythm level data using the preset gesture-semantic mapping database. The database pre-defines text labels corresponding to different rhythm levels. For example, "slow" corresponds to "slow rhythm", "medium speed" corresponds to "moderate rhythm", and "fast" corresponds to "fast rhythm". According to the gesture rhythm level data, the corresponding text label is searched from the database to obtain the gesture rhythm text vocabulary data. These semantic information are associated with the basic gesture action units. For each basic gesture action unit, its corresponding violent motion semantic vocabulary, complex motion semantic vocabulary and gesture rhythm text vocabulary can be combined to form a semantic label that describes the action unit.

[0147] Preferably, step S3 comprises the following steps:

[0148] Step S31: performing confidence evaluation on the gesture action intention recognition data, and screening high-confidence intentions to obtain high-confidence action intention data;

[0149] Step S32: performing action emotion adjective processing according to the high-confidence action intention data, and performing gesture semantic fusion according to the gesture motion text description data to generate key gesture semantic fusion data;

[0150] Step S33: performing intelligent text generation based on key gesture semantic fusion data to generate initial gesture text content data;

[0151] Step S34: performing gesture category analysis according to the basic gesture action unit data, and performing text template keyword processing to generate text template keyword data;

[0152] Step S35: using the text template keyword data to perform text template matching on a preset text template library to obtain selected text template data;

[0153] Step S36: The initial gesture text content data is semantically filled with placeholders by selecting text template data, thereby obtaining intelligent description gesture text data.

[0154] As an example of the present invention, refer to Figure 3 As shown, Figure 1 Detailed implementation steps of step S3 in the flowchart, in this example, step S3 includes:

[0155] Step S31: performing confidence evaluation on the gesture action intention recognition data, and screening high-confidence intentions to obtain high-confidence action intention data;

[0156] In an embodiment of the present invention, the confidence is evaluated according to the probability value output by the gesture motion intention recognition model. For example, in step S26, the gesture motion intention recognition model outputs a vector, each element of which represents the probability that the gesture belongs to each intent category. The category with the highest probability can be selected as the intent recognition result, and the probability value can be used as the confidence. For example, if the vector output by the model is [0.1, 0.8, 0.05, 0.05, 0], it means that the probability that the gesture belongs to the first intent category is 0.1, and the probability that it belongs to the second intent category is 0.8, and so on. Then, the second category can be selected as the intent recognition result, and 0.8 can be used as the confidence. A confidence threshold value can be set in advance, such as 0.7. If the confidence of an intent recognition result is higher than the threshold, the result is considered to be of high confidence and the result is retained; otherwise, the result is considered to be of low confidence and the result is discarded. For example, for the above example, since 0.8 is greater than 0.7, the intent recognition result is retained. If the confidence of another intention recognition result is 0.6, then since 0.6 is less than 0.7, the result is discarded. All high-confidence intention recognition results are screened out to form high-confidence action intention data.

[0157] Step S32: performing action emotion adjective processing according to the high-confidence action intention data, and performing gesture semantic fusion according to the gesture motion text description data to generate key gesture semantic fusion data;

[0158] In an embodiment of the present invention, an emotional adjective library is pre-established, which contains various emotional adjectives, such as "quickly", "slowly", "carefully", "boldly", etc. Then, according to the characteristics of the gesture action, such as speed, acceleration, trajectory shape, etc., appropriate emotional adjectives are selected. For example, according to the speed acceleration feature data obtained in step S225, if the speed and acceleration of a gesture action are relatively large, emotional adjectives such as "quickly" and "swiftly" can be selected; if the speed and acceleration are relatively small, emotional adjectives such as "slowly" and "gently" can be selected. The selection of adjectives can be based on predefined rules or machine learning models. The selected emotional adjectives are fused with the gesture motion text description data generated in step S247 to generate key gesture semantic fusion data. The fusion method can be to insert the adjective into the appropriate position of the text description, for example, inserting "quickly" before "opening the palm" to form "opening the palm quickly".

[0159] Step S33: performing intelligent text generation based on key gesture semantic fusion data to generate initial gesture text content data;

[0160] In an embodiment of the present invention, a rule-based method or a deep learning-based method is used to generate text. The rule-based method can use pre-defined grammatical rules and templates to combine the elements in the key gesture semantic fusion data into sentences. For example, a template can be defined: "user [emotional adjective] [intention] [semantic description label]", and then the corresponding elements in the key gesture semantic fusion data are filled into the template. For example, for the fusion result of "quickly select", "user quickly selects" can be generated.

[0161] Step S34: performing gesture category analysis according to the basic gesture action unit data, and performing text template keyword processing to generate text template keyword data;

[0162] In the embodiment of the present invention, gesture category analysis is performed based on the basic gesture action unit data (step S236). For example, gestures can be divided into categories such as "grabbing", "dropping", and "pointing". Then, corresponding keywords are extracted according to the gesture category. For example, for the "grabbing" category, keywords such as "grabbing", "holding", and "taking" can be extracted; for the "pointing" category, keywords such as "finger", "pointing", and "indicating" can be extracted. The extracted keywords will be used for subsequent text template matching to generate text template keyword data.

[0163] Step S35: using the text template keyword data to perform text template matching on a preset text template library to obtain selected text template data;

[0164] In an embodiment of the present invention, the text template keyword data generated in step S34 is used to perform text template matching on a preset text template library. The text template library contains various text templates that describe gesture actions, and each template contains some placeholders for filling in specific semantic information. For example, a text template may be "[action][object]", where "[action]" and "[object]" are placeholders. The matching process selects the most appropriate text template based on the text template keyword data. For example, if the keyword is "grab", a text template containing "grab" or its synonyms can be selected. The matching result obtains the selected text template data.

[0165] Step S36: The initial gesture text content data is semantically filled with placeholders by selecting text template data, thereby obtaining intelligent description gesture text data.

[0166] In an embodiment of the present invention, the selected text template data obtained in step S35 performs placeholder semantic filling on the initial gesture text content data generated in step S33. The semantic information extracted from the initial gesture text content data is filled into the placeholder of the text template, thereby generating complete intelligent description gesture text data. For example, if the selected text template is "[action][object]", and the initial gesture text content data is "quickly grab an apple", the filled text is "grab the apple". If the initial text content cannot completely fill the placeholder of the template, a default value can be used or inferred based on the context.

[0167] Preferably, the present invention further provides a text generation system for executing the above-mentioned text generation method, the text generation system comprising:

[0168] The gesture trajectory reconstruction module is used to collect the user's hand depth image in real time using a 3D depth camera to obtain depth hand image sequence data; the gesture trajectory is reconstructed according to the depth hand image sequence data to obtain spatial gesture motion trajectory data;

[0169] The gesture semantic annotation module is used to analyze the gesture motion characteristics of the spatial gesture motion trajectory data and generate gesture motion characteristic vector data; segment the basic gesture action units according to the gesture motion characteristic vector data and generate basic gesture action unit data; annotate the gesture motion semantic labels of the gesture motion characteristic vector data based on the basic gesture action unit data and generate gesture motion text description data; construct a gesture motion intention recognition model; use the gesture motion intention recognition model to perform action intention recognition on the gesture motion characteristic vector data and obtain gesture action intention recognition data;

[0170] The intelligent gesture description text generation module is used to process the action emotion adjectives according to the gesture action intention recognition data, and to perform gesture semantic fusion according to the gesture movement text description data to generate key gesture semantic fusion data; based on the key gesture semantic fusion data, intelligent text generation is performed, and text template semantic filling is performed to obtain intelligent gesture description text data;

[0171] The description text visualization module is used to perform visual text rendering processing based on the intelligent description gesture text data to achieve the generation of intelligent description text for gesture actions.

[0172] Preferably, the present invention further provides a terminal device, the terminal device comprising:

[0173] processor;

[0174] a memory for storing processor-executable instructions;

[0175] Wherein, the processor is configured to implement any of the text generation methods described above.

[0176] Preferably, the present invention further provides a computer-readable storage medium storing a computer program, wherein the computer program implements any one of the above text generation methods when executed.

[0177] The present application is to accurately capture the motion trajectory of the hand in three-dimensional space by collecting the user's hand depth image in real time through a 3D depth camera, overcome the shortcomings of the traditional gesture recognition method based on two-dimensional images that is sensitive to gesture posture, viewing angle and illumination changes, and make gesture recognition more stable and accurate. The present invention effectively preprocesses the original trajectory data by analyzing the length of the gesture space trajectory and regularizing the sampling trajectory time, thereby improving the standardization and availability of the data. When different users perform the same gesture, their speed and trajectory length are different. Through length analysis and time regularization, trajectory data of different lengths and speeds can be unified to the same standard. Gesture motion feature analysis converts the preprocessed trajectory data into gesture motion feature vector data, realizing the conversion from specific trajectory to abstract features. Through trajectory curve fitting, gesture trajectory curvature and direction change analysis, the geometric features of the gesture trajectory are extracted. These geometric features reflect the shape and direction change of gesture movement. For example, circular gestures and linear gestures have different curvature features, and waving to the left and waving to the right have different direction change features. The gesture motion feature data is mapped to different semantic dimensions. This multi-dimensional analysis makes the description of gestures more comprehensive and detailed, such as "fast", "slow", "complex", "simple", "rhythmic", etc. Using the preset gesture-semantic mapping database, these level data are mapped to the corresponding semantic text labels. This database is the key to achieving the transformation from feature data to text description. It associates different feature levels with corresponding semantic words, such as "violent", "slight", "complex", "simple", "fast", "soothing", etc. Through mapping, the system can transform abstract feature data into specific semantic words. The high-confidence action intention data is processed into action emotion adjectives. Different categories of gestures, such as interaction, instruction, expression, etc., usually correspond to different text templates. Through keyword matching, the appropriate text template can be quickly found. The generated text content is combined with the template, so that the final text output has both personalized content and good readability and fluency.

[0178] Therefore, the embodiments should be regarded as illustrative and non-restrictive from all points, and the scope of the present invention is limited by the appended claims rather than the above description, and it is therefore intended that all changes falling within the meaning and range of equivalent elements of the application documents are included in the present invention.

[0179] The above description is only a specific embodiment of the present invention, so that those skilled in the art can understand or implement the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but should conform to the widest scope consistent with the principles and novel features invented herein.

Claims

1. A text generation method, characterized in that: The following steps are involved: Step S1: using a 3D depth camera to collect real-time hand depth images of the user to obtain depth hand image sequence data; reconstructing gesture trajectories according to the depth hand image sequence data to obtain spatial gesture motion trajectory data; Step S2: performing gesture motion feature analysis on the spatial gesture motion trajectory data to generate gesture motion feature vector data; Segment the basic gesture action units according to the gesture motion feature vector data to generate basic gesture action unit data; annotate the gesture motion semantic labels on the gesture motion feature vector data based on the basic gesture action unit data to generate gesture motion text description data; construct a gesture motion intention recognition model; use the gesture motion intention recognition model to perform action intention recognition on the gesture motion feature vector data to obtain gesture action intention recognition data; Step S3: performing action emotion adjective processing according to the gesture action intention recognition data, and performing gesture semantic fusion according to the gesture movement text description data to generate key gesture semantic fusion data; Based on the key gesture semantic fusion data, intelligent text generation is performed and text template filling is performed to obtain text data that intelligently describes gestures; Step S4: Perform visual text rendering processing according to the intelligent description gesture text data to realize the generation of intelligent description text for gesture actions.

2. The text generation method according to claim 1, characterized in that: Step S1 includes the following steps: Step S11: using a 3D depth camera to collect real-time hand depth images of the user, and perform initial image correction to obtain depth hand image sequence data; Step S12: removing the background of the depth hand image sequence data and performing hand region segmentation to obtain user hand region data; Step S13: performing hand posture estimation on the user's hand area data to obtain the user's hand posture data; Step S14: Detect skeleton key points according to the user's hand posture data, and perform hand skeleton key point coordinate analysis to obtain skeleton key point coordinate data; Step S15: extracting pixel coordinates of gesture key points from the depth hand image sequence data using the skeleton key point coordinate data to obtain pixel coordinate data of gesture image key points; Step S16: performing time stamp synchronization processing according to the pixel coordinate data of the key points of the gesture image, and reconstructing the gesture trajectory according to the user's hand posture data to obtain spatial gesture motion trajectory data.

3. The text generation method according to claim 1, characterized in that: Step S2 includes the following steps: Step S21: analyzing the spatial gesture motion trajectory data for the length of the spatial gesture trajectory, and performing time regularization processing on the sampling trajectory to obtain pre-processed gesture motion trajectory data; Step S22: performing gesture motion feature analysis based on the pre-processed gesture motion trajectory data to generate gesture motion feature vector data; Step S23: Segment the basic gesture action unit according to the gesture motion feature vector data to generate basic gesture action unit data; Step S24: annotating the gesture motion feature vector data with gesture motion semantic labels based on the basic gesture action unit data using a preset gesture-semantic mapping database to generate gesture motion text description data; Step S25: performing gesture movement intention sample transfer learning according to a preset convolutional neural network model, thereby constructing a gesture movement intention recognition model; Step S26: using the gesture motion intention recognition model to perform motion intention recognition on the basic gesture action unit data and the gesture motion feature vector data to obtain gesture action intention recognition data.

4. The text generation method according to claim 3, characterized in that: Step S22 includes the following steps: Step S221: extracting the coordinates of the starting point / ending point of the gesture trajectory according to the pre-processed gesture motion trajectory data to obtain the trajectory starting and ending point data; Step S222: Calculating the displacement vectors of the hand key points in adjacent frames according to the pre-processed gesture motion trajectory data to generate hand inter-frame displacement vector data; Step S223: performing inter-frame instantaneous velocity vector calculation on the inter-frame displacement vector data of the hand to generate instantaneous velocity vector data; Step S224: performing low-pass filtering according to the instantaneous velocity vector data, and performing gesture motion acceleration calculation to obtain frame-by-frame acceleration vector data; Step S225: performing mean processing on the instantaneous velocity vector data and the frame-by-frame acceleration vector data through a preset time window to obtain velocity acceleration characteristic data; Step S226: performing trajectory curve fitting on the pre-processed gesture motion trajectory data through the trajectory start and end point data, and performing gesture trajectory curvature and direction change analysis to obtain gesture trajectory geometric feature data; Step S227: performing gesture posture change frequency analysis based on the pre-processed gesture motion trajectory data to obtain gesture posture change frequency data; Step S228: performing gesture feature combination on the velocity acceleration feature data, the gesture trajectory geometric feature data and the gesture posture change frequency data to obtain gesture motion feature vector data.

5. The text generation method according to claim 3, characterized in that: Step S24 includes the following steps: Step S231: performing key frame detection on the gesture motion feature vector data to obtain key frame data; Step S232: segmenting the pre-processed gesture motion trajectory data into key frame time periods using the key frame data to obtain time period segmented trajectory data; Step S233: performing hand gesture cluster analysis on the time segmentation trajectory data to obtain hand gesture cluster data; Step S234: dividing the time segmentation trajectory data into basic gesture action units based on the gesture posture clustering data to generate basic gesture action unit division data; Step S235: performing conversion relationship processing between gesture action units on the basic gesture action unit division data to obtain unit conversion relationship data; Step S236: Integrate the basic gesture action unit division data and the unit conversion relationship data to generate basic gesture action unit data.

6. The text generation method according to claim 4, characterized in that: Step S24 includes the following steps: Step S241: determining the gesture motion intensity level of the speed acceleration feature data to generate motion intensity level data; Step S242: mapping the exercise intensity level data to exercise semantic text labels using a preset gesture-semantic mapping database to obtain intense exercise semantic vocabulary data; Step S243: performing gesture complexity level analysis on the gesture trajectory geometric feature data to obtain complexity level data; Step S244: using a preset gesture-semantic mapping database to map the complexity level data to motion semantic text labels, and obtaining complex motion semantic vocabulary data; Step S245: classifying the gesture rhythm speed level of the gesture posture change frequency data to obtain gesture rhythm level data; Step S246: mapping the gesture rhythm level data to gesture rhythm text labels using a preset gesture-semantic mapping database to obtain gesture rhythm text vocabulary data; Step S247: Based on the basic gesture action unit data, the violent motion semantic vocabulary data, the complex motion semantic vocabulary data and the gesture rhythm text vocabulary data are fused with gesture semantic description labels to obtain gesture motion text description data.

7. The text generation method according to claim 1, characterized in that: Step S3 includes the following steps: Step S31: performing confidence evaluation on the gesture action intention recognition data, and screening high-confidence intentions to obtain high-confidence action intention data; Step S32: performing action emotion adjective processing according to the high-confidence action intention data, and performing gesture semantic fusion according to the gesture motion text description data to generate key gesture semantic fusion data; Step S33: performing intelligent text generation based on key gesture semantic fusion data to generate initial gesture text content data; Step S34: performing gesture category analysis according to the basic gesture action unit data, and performing text template keyword processing to generate text template keyword data; Step S35: using the text template keyword data to perform text template matching on a preset text template library to obtain selected text template data; Step S36: The initial gesture text content data is semantically filled with placeholders by selecting text template data, thereby obtaining intelligent description gesture text data.

8. A text generation system, characterized in that: For executing the text generation method according to claim 1, the text generation system comprises: The gesture trajectory reconstruction module is used to collect the user's hand depth image in real time using a 3D depth camera to obtain depth hand image sequence data; the gesture trajectory is reconstructed according to the depth hand image sequence data to obtain spatial gesture motion trajectory data; The gesture semantic annotation module is used to analyze the gesture motion characteristics of the spatial gesture motion trajectory data and generate gesture motion characteristic vector data; segment the basic gesture action units according to the gesture motion characteristic vector data and generate basic gesture action unit data; annotate the gesture motion semantic labels of the gesture motion characteristic vector data based on the basic gesture action unit data and generate gesture motion text description data; construct a gesture motion intention recognition model; use the gesture motion intention recognition model to perform action intention recognition on the gesture motion characteristic vector data and obtain gesture action intention recognition data; The intelligent gesture description text generation module is used to process the action emotion adjectives according to the gesture action intention recognition data, and to perform gesture semantic fusion according to the gesture movement text description data to generate key gesture semantic fusion data; based on the key gesture semantic fusion data, intelligent text generation is performed, and text template semantic filling is performed to obtain intelligent gesture description text data; The description text visualization module is used to perform visual text rendering processing based on the intelligent description gesture text data to achieve the generation of intelligent description text for gesture actions.

9. A terminal device, characterized in that: The terminal device comprises: processor; a memory for storing processor-executable instructions; Wherein, the processor is configured to implement the text generation method described in any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed, the text generation method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Text recognition method and device applied to air handwriting, equipment and storage medium

    CN115826762A

  • Air handwriting interaction method based on three-dimensional gesture reconstruction, storage medium and device

    CN117058691A

  • Oral talent training method, device and equipment and storage medium

    CN117522643A

  • Man-machine interaction method and apparatus, and man-machine interaction terminal

    WO2018219198A1