Video sign language question and answer method, system and device and storage medium
By employing a self-supervised pre-training and task-fine-tuning sign language recognition and translation model and a multi-turn interactive generative question-answering technique based on a large language model, the accuracy problem of sign language recognition, translation, and generation in existing technologies has been solved. This enables accurate recognition of sign language and knowledge answers in professional fields in complex scenarios, and generates sign language videos that conform to the sign language word order.
Patent Information
- Application Number
- CN202511537027.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-27
- Publication Date
- 2025-11-28
- Estimated Expiration
- 2045-10-27
AI Technical Summary
Existing sign language recognition and translation technologies cannot effectively handle complex scenarios. Existing technologies struggle to accurately recognize and understand the complexity of gestures. They also cannot effectively provide knowledge-based answers in complex scenarios. Sign language generation technologies cannot generate results that conform to sign language word order. Natural language question answering systems cannot accurately understand user needs.
By using a self-supervised pre-trained and task-fine-tuned sign language recognition and translation model, combined with a large language model and knowledge retrieval technology, and multi-round interactive generative question answering, sign language videos that conform to the word order of sign language are generated.
It enables accurate recognition and translation of sign language in complex scenarios, as well as the provision of professional knowledge answers, and generates sign language videos that conform to the word order of sign language, thereby improving the accuracy and efficiency of sign language communication.
Smart Images

Figure CN121033944A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of video sign language question and answer, and in particular to a video sign language question and answer method, system, device and storage medium. BACKGROUND
[0002] In modern society, the deaf community is severely affected in the realization of basic rights and social participation due to limited information access and communication ability. With the development of technology, sign language recognition and translation, natural language question and answer systems, and sign language generation technology have made some progress, but still face some challenges. These challenges mainly come from the complexity of sign language and the diversity of application scenarios. For example, the system of sign language linguistics is not strong, the numerous sign language dialects make data collection and translation difficult, and it is not possible to provide knowledge answers in professional fields in complex scenarios.
[0003] Sign language is a complex form of communication that involves not only hand movements but also facial expressions, body postures, and other modalities of information. The system needs to process and drive these multi-modal data simultaneously to achieve comprehensive and accurate recognition of sign language. By introducing prior knowledge, the semantic and contextual information of sign language can be better understood, thereby improving the accuracy and efficiency of question and answer. For example, the system can use linguistic knowledge, sign language grammar rules, and other prior knowledge in the conversation to optimize the accuracy of each module of the system.
[0004] The video sign language question and answer system is divided into three key modules: a sign language recognition and translation subsystem that translates sign language videos into text sentences to understand the sign language expressions of the hearing-impaired population; a natural language question and answer subsystem that matches the text recognition results with an established sign language question and answer knowledge base to output a text answer; and a sign language synthesis subsystem that converts the natural language text answer into a digital human sign language video based on sign language grammar rules.
[0005] (1) Sign language recognition and translation subsystem.
[0006] Video sign language recognition and translation technology can be mainly divided into two categories: deep learning and machine vision. Among them, deep learning technology is to train a large number of sign language video data, and the model can learn the characteristics and rules of sign language gestures, so as to realize the recognition of sign language in video. Recurrent neural network (RNN) and convolutional neural network (CNN) are two commonly used models. RNN is suitable for processing sign language with time sequence characteristics, which can capture the continuity and variability of sign language gestures in time, and more accurately recognize sign language. CNN mainly processes image data, extracts key features of sign language gestures such as hand shape, contour and motion trajectory, and identifies sign language through feature extraction. Machine vision technology captures gesture actions in video, and extracts information such as hand position, shape and motion trajectory through image processing technology, and inputs the extracted information into a deep learning model for sign language recognition and classification.
[0007] (2) Natural language question and answer subsystem.
[0008] Question and answer system is divided into two technical modes of natural language generation and template matching. Natural language generation (NLG) usually uses seq2seq (sequence to sequence) model or Transformer (transform neural network) model to automatically construct complete answers. NLG is responsible for converting the system action selected by the dialogue strategy module into natural language for feedback. Template matching technology matches questions with templates by defining a series of templates and rules, finds the most suitable template, and generates answers according to the template.
[0009] In general, template generation is simple and direct, and is directly applicable to a series of different practical application task scenarios. However, such technical mode needs to establish a large template library to improve the accuracy to meet the various ways of asking users. The quality of data set is required to be higher in the process of model training for NLG method, and the accuracy of answer generation is slightly inferior to that of template generation.
[0010] (3) Sign language synthesis subsystem.
[0011] Sign language synthesis consists of two parts: sign language word conversion, which constructs a large sign language word library, uses rule matching to divide the text sentence of natural language, and maps the word to sign language word; the second is sign language digital human driving, which transmits the sequence parameters of sign language word to digital human for driving display.
[0012] However, the defects of the existing scheme are: (1) The existing sign language recognition and translation technology needs a large amount of image processing and data analysis. Hand movement has high flexibility and rapidity, often accompanied by occlusion and self-occlusion. The traditional two-dimensional image processing method is difficult to accurately capture the three-dimensional information of hand movement, and the traditional machine learning and deep learning model cannot fully learn the characteristics and rules of gestures, so it is difficult to control the accuracy of recognition.
[0013] (2) The existing generative natural language question and answer technology cannot control the accuracy of question and answer due to the lack of knowledge base constraint, and is prone to "random answer". Template matching needs to build all possible question and answer libraries, which cannot control the accuracy of the answer while the construction cost is huge, and cannot achieve multi-round question and answer.
[0014] (3) The existing sign language generation technology is mainly based on rule matching of spoken language to sign language word conversion library. The sign language generation result is oral translation and does not contain sign language sequence, which makes it difficult for deaf people to read the result after fitting the sign language word of the digital person.
[0015] Therefore, the present application is proposed. SUMMARY
[0016] The purpose of the present application is to provide a video sign language question and answer method, system, device and storage medium, which not only can complete the smooth conversion between natural language and sign language, but also can efficiently and accurately understand user demand and provide actual answer to the pain point problem of deaf people group information acquisition with fault, and bring substantial convenience to deaf people group.
[0017] The purpose of the present application is realized by the following technical solutions: A video sign language question and answer method, comprising: extracting a skeleton point sequence from an input sign language video, and using a sign language recognition and translation model based on self-supervised pre-training and task fine-tuning to perform sign language recognition and translation, and obtaining a question text; generating a question and answer under the constraint of a knowledge base by a large language model combined with the question text; in the process of generative question and answer, through multi-round interaction with the large language model, finally generating an answer text corresponding to the question text; converting the answer text into a sign language word text conforming to the sign language sequence, forming a sign language word text sequence, then searching and matching a sign language sequence word library to obtain a corresponding action sequence, and processing by combining action smoothing and transition generation technology, driving a digital person to generate a corresponding sign language video.
[0018] A video sign language question and answer system for realizing the foregoing method, comprising: The sign language recognition and translation subsystem is used to extract skeleton point sequences from input sign language videos and use a sign language recognition and translation model obtained based on self-supervised pre-training and task fine-tuning to perform sign language recognition and translation to obtain the question text; The multi-turn knowledge retrieval question answering subsystem is used to perform generative question answering by combining a large language model with the question text under the constraints of a knowledge base. During the generative question answering process, the system interacts with the large language model in multiple turns to finally generate the answer text corresponding to the question text. The sign language synthesis subsystem is used to convert the answer text into sign language word text that conforms to the sign language word order, form a sign language word text sequence, and then obtain the corresponding action sequence by searching and matching the sign language sequence word library. After processing with action smoothing and transition generation technology, it drives the digital human to generate the corresponding sign language video.
[0019] A processing device includes: one or more processors; and a memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the aforementioned method.
[0020] A readable storage medium storing a computer program that, when executed by a processor, implements the aforementioned method.
[0021] As can be seen from the technical solutions provided by the present invention, the video sign language question answering scheme based on multimodal data-driven and prior knowledge modeling can construct a complete and efficient information processing closed loop. From the accurate input recognition of sign language, to the deep processing and accurate answering based on strong knowledge reserves, and then to the output in a natural and fluent sign language form, it provides a comprehensive and one-stop technical architecture for sign language communication and information transmission. In general, the video sign language question answering scheme provided by the present invention: (1) The sign language recognition and translation technology uses self-supervised pre-training technology to enhance the model's representation ability, realize sign language video translation and complete the full training data. (1) Improve recognition accuracy by utilizing large language models and knowledge retrieval technology, and perform generative question answering under the constraints of the knowledge base. At the same time, understand user intent and maintain context through multiple interactions (multi-turn dialogues) to complete intelligent dialogue under complex tasks. (2) In the process of sign language generation, the first step is to convert the text into gloss (sign language word text) that conforms to the sign language word order. The second step is to convert the gloss into an action sequence and drive the digital human. In this step, the action smoothing and transition generation technology is used to process the action sequence of sign language words to reduce the jitter problem of the action and obtain a smoother action sequence to drive the digital human. Attached Figure Description
[0022] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0023] Figure 1 A flowchart of a video sign language question-and-answer method provided in an embodiment of the present invention.
[0024] Figure 2 This is a schematic diagram of the overall architecture of a video sign language question-and-answer method provided in an embodiment of the present invention.
[0025] Figure 3 This is a schematic diagram of the sign language recognition and translation technology architecture provided in an embodiment of the present invention.
[0026] Figure 4 This is a schematic diagram of the natural language question answering technology roadmap provided in the embodiments of the present invention.
[0027] Figure 5 This is a schematic diagram of the sign language synthesis technology route provided in an embodiment of the present invention.
[0028] Figure 6 This is a schematic diagram of the training technology route for the text2gloss translation model provided in an embodiment of the present invention.
[0029] Figure 7 This is a schematic diagram of a processing device provided in an embodiment of the present invention. Detailed Implementation
[0030] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the protection scope of the present invention.
[0031] First, the following explanations are provided for the terms that may be used in this article: The terms "comprising," "including," "containing," "having," or other similar semantic descriptions should be interpreted as non-exclusive inclusion. For example, including a technical feature element (such as raw material, component, ingredient, carrier, dosage form, material, size, part, component, mechanism, device, step, process, method, reaction conditions, processing conditions, parameter, algorithm, signal, data, product or article of manufacture, etc.) should be interpreted as including not only the expressly listed technical feature element, but also other technical feature elements known in the art that are not expressly listed.
[0032] The term "composed of" excludes any technical features not expressly listed. When used in a claim, it closes the claim to exclude all technical features other than those expressly listed, except for associated conventional impurities. If the term appears only in a clause of a claim, it limits the claim to the elements expressly listed in that clause; elements recited in other clauses are not excluded from the overall claim.
[0033] Unless otherwise explicitly specified or limited, the terms "installation," "connection," "linking," and "fixing," etc., should be interpreted broadly. For example, they can refer to fixed connections, detachable connections, or integral connections; they can refer to mechanical connections or electrical connections; they can refer to direct connections or indirect connections through an intermediate medium; and they can refer to the internal connection between two components. Those skilled in the art can understand the specific meaning of the above terms in this document according to the specific circumstances.
[0034] The following is a detailed description of a video sign language question-and-answer method, system, device, and storage medium provided by the present invention. Contents not described in detail in the embodiments of the present invention are prior art known to those skilled in the art. Where specific conditions are not specified in the embodiments of the present invention, they are performed according to conventional conditions in the art or conditions recommended by the manufacturer. Where the manufacturers of the instruments used in the embodiments of the present invention are not specified, they are all conventional products that can be purchased commercially.
[0035] Example 1 This invention provides a video sign language question-and-answer method, such as... Figure 1 As shown, it mainly includes the following steps: Step 1: Extract the skeleton point sequence from the input sign language video, and use the sign language recognition and translation model obtained based on self-supervised pre-training and task fine-tuning to perform sign language recognition and translation to obtain the question text.
[0036] Preferably, the steps for obtaining a sign language recognition and translation model based on self-supervised pre-training and task fine-tuning are as follows: (1) Collect unlabeled sign language video data and labeled sign language video data, and select a sign language pre-training model; (2) Perform self-supervised pre-training on the sign language pre-training model using unlabeled sign language video data, including: extracting skeleton point sequences from unlabeled sign language video data using a human pose extractor; performing masking modeling on the skeleton point sequences and inputting them into the sign language pre-training model, and performing self-supervised pre-training by combining the skeleton sequence vector output by the sign language pre-training model to obtain a self-supervised pre-training model; (3) Perform task fine-tuning on the self-supervised pre-training model using labeled sign language video data, including: extracting skeleton point sequences from labeled sign language video data using a human pose extractor, The data is then input into a self-supervised pre-trained model. Through a skeleton sequence vector and language text alignment mechanism, and combined with a large language model for sign language recognition and translation, fine-tuning training is performed to obtain a fine-tuned self-supervised pre-trained model and a fine-tuned large language model for sign language recognition and translation. The two together constitute a sign language recognition and translation model. The language text involved in the skeleton sequence vector and language text alignment mechanism is the labeled text in the labeled sign language video data.
[0037] In this embodiment of the invention, the human posture extractor can be implemented using conventional techniques, which will not be elaborated here.
[0038] Step 2: Generative question answering is performed by combining the large language model with the question text under the constraints of the knowledge base. During the generative question answering process, multiple rounds of interaction with the large language model are conducted to finally generate the answer text corresponding to the question text.
[0039] Preferably, the large language model includes: a knowledge retrieval large language model and a question-answering large language model; the question text is input into the knowledge retrieval large language model, which retrieves the N knowledge items with the highest matching degree to the question text from the knowledge base, and inputs the N knowledge items into the question-answering large language model to obtain the corresponding N answers; based on the context learning method of embedding technology, the question text and the corresponding N answers are embedded into the same vector space and input into the knowledge retrieval large language model to continue a new round of interaction; furthermore, the response process of the new round of interaction is controlled by reinforcement learning in combination with the output of the question-answering large language model; finally, the answer text corresponding to the question text is generated.
[0040] Preferably, the knowledge retrieval large language model and the question-answering large language model are each trained using pre-collected datasets. These datasets include: a question-and-knowledge pair dataset, a question-and-answer pair dataset, and a triplet dataset consisting of a question, a high-quality answer, and a second-best answer. The question-and-knowledge pair dataset is vectorized to construct a vectorized knowledge base. The vectorized knowledge base is used to train the knowledge retrieval large language model. The question-and-answer pair dataset is used to fine-tune the question-answering large language model, enabling it to answer questions based on input questions and output corresponding answers. Simultaneously, reinforcement learning is used to optimize the fine-tuned question-answering large language model using the triplet dataset, ensuring that the optimized model's output answers align with the answers in the triplet dataset.
[0041] Step 3: Convert the answer text into sign language word text that conforms to the sign language word order, form a sign language word text sequence, then obtain the corresponding action sequence by searching and matching the sign language sequence word library, and after processing with action smoothing and transition generation technology, drive the digital human to generate the corresponding sign language video.
[0042] Preferably, the conversion of the answer text into sign language word text conforming to sign language word order is achieved through the text2gloss translation model. The text2gloss translation model converts natural language text into sign language word text and adopts an encoder-decoder architecture. The encoder is responsible for encoding the answer text into an intermediate representation vector, and the decoder is responsible for decoding the intermediate representation vector into sign language word text conforming to sign language word order.
[0043] Preferably, the text2gloss translation model is based on spoken text-sign language word sequence text pairs and is trained for the task of translating natural text into sign language word text through transfer learning and large model techniques.
[0044] Preferably, the step of obtaining the corresponding action sequence by searching and matching the sign language sequence lexicon, and then processing it with action smoothing and transition generation techniques to drive the digital human to generate the corresponding sign language video includes: (1) querying the sign language sequence lexicon for the corresponding action sequence for each sign language word in the sign language word text sequence, and splicing the action sequences corresponding to all sign language word texts to obtain the action sequence corresponding to the sign language word text sequence; (2) performing action smoothing and transition generation processing on the action sequence corresponding to the sign language word text sequence to obtain the final action sequence; (3) importing the final action sequence into the rendering engine, driving the digital human in the rendering engine to render the corresponding sign language video.
[0045] To more clearly demonstrate the technical solution and its effects provided by the present invention, the method provided by the embodiments of the present invention will be described in detail below with reference to specific examples.
[0046] I. Overall Overview of the Plan
[0047] Considering that sign language recognition and translation, sign language question answering, and sign language synthesis technologies are important technological innovations in the field of artificial intelligence aimed at addressing communication barriers for the hearing impaired, these technologies aim to break down language barriers and promote barrier-free communication. To address the problems existing in existing solutions, this invention provides a video sign language question answering method based on multimodal data-driven and prior knowledge modeling. This method integrates the entire technological chain of sign language recognition and translation, knowledge-based question answering with a large language model, and the generation of sign language sequences to drive digital human development, achieving seamless integration and collaborative operation of these three key stages.
[0048] In the field of video sign language recognition and translation, this approach first employs a large amount of unlabeled video data for self-supervised learning based on a pre-trained sign language model (e.g., the SignBERT model). Within this self-supervised pre-training framework of skeleton point sequences, prior sign language information is utilized to implicitly learn the contextual relationships within the skeleton point sequences by designing different masking strategies and reconstructing the masked content. Then, a large amount of labeled video data is used to fine-tune the large-scale language model for sign language recognition and translation based on the skeleton sequence vectors output from the self-supervised pre-trained model. This vector is aligned with the sign language video text, aiming to extract features from both visual and textual modalities through multi-modal learning, aligning and mapping them to a shared latent space. This allows the model to understand sign language and generate relevant text translations or perform contextual understanding. This method combines self-supervised pre-training techniques with the powerful semantic understanding capabilities of the large-scale language model for sign language recognition and translation, making it suitable for processing complex sign language video data.
[0049] Multi-turn knowledge retrieval and question answering based on a large language model: Large language models can better understand and handle complex dialogue scenarios, including but not limited to long-term memory retention and sentiment analysis. Utilizing advanced retrieval algorithms and matching models, it improves the accuracy and efficiency of knowledge retrieval. It can not only answer directly relevant questions but also provide more coherent and tailored answers based on previous conversations, demonstrating a strong knowledge base and deep understanding. Whether it's professional questions related to sign language or various inquiries across a wide range of fields, it can quickly grasp the core of the problem and provide accurate, detailed, and practical answers.
[0050] Text-driven digital human sign language generation technology deeply integrates natural language processing with sophisticated modeling and efficient generation algorithms for sign language movements. Based on a large language translation model, it employs an encoder-decoder architecture. The encoder converts the question-answering model's response text into an intermediate representation (vector sequence), while the decoder generates the sign language text based on this representation. This architecture allows the model to flexibly generate the target language while understanding the source language, achieving an efficient translation process. This enables the fluent, natural, and accurate conversion of input text information into vivid sign language expressions for digital humans.
[0051] like Figure 2 The diagram illustrates the overall architecture of the method described above, realizing a complete process from sign language recognition to text generation and then to sign language video synthesis. All related processes are implemented using corresponding models. Figure 2 The training and inference process is illustrated in the diagram, and specific details will be introduced later.
[0052] II. Detailed introduction of the plan.
[0053] 1. Sign language recognition and translation technology architecture (sign language recognition and translation subsystem).
[0054] This invention proposes a sign language video translation method based on a pre-trained sign language model (e.g., the SignBERT sign language pre-trained model), achieving end-to-end sign language video-to-text translation through a two-stage training process. In the first stage, a self-supervised pre-training approach is used to train the sign language pre-trained model using unlabeled sign language video data. By designing diverse skeleton point sequence masking strategies and constructing a reconstruction task, the model can implicitly learn the contextual dependencies of sign language actions. In the second stage, skeleton sequence vectors from labeled data are extracted based on the model trained in the first stage. Fine-tuning is then performed using a skeleton sequence vector-to-text alignment mechanism, combined with a large-scale sign language recognition and translation model. This framework achieves efficient translation from sign language video to natural language.
[0055] (1) Generalization processing of sign language recognition data.
[0056] The quality of sign language videos is a crucial factor affecting the quality of pre-trained models. To ensure the quality of sign language data, high-resolution camera equipment should be selected, a uniform and simple background should be used, and repeated recordings by multiple people are necessary to increase data robustness. Video editing should remove redundant parts, and the format, resolution, and frame rate should be standardized to facilitate subsequent processing and analysis. After obtaining the sign language video data, data filtering and quality control are performed. This includes invalid video removal: removing invalid video segments such as blurry ones; image quality checks: using automated or semi-automated tools to analyze the video's sharpness, brightness, and contrast to ensure all videos meet visual standards; and duration and frame rate checks: using automated tools to help check whether the duration and frame rate of each video are consistent, and standardizing the dataset format.
[0057] (2) Training of sign language video translation model based on self-supervised pre-training model.
[0058] In sign language recognition and translation, sign language videos of deaf individuals are acquired through video capture or uploading and translated into text content that the deaf community can understand. With the rapid development of deep neural network algorithms and the collection of large-scale data, self-supervised pre-trained large models exhibit powerful representation capabilities, facilitating the application of deep models in real-world scenarios. Therefore, this invention utilizes self-supervised pre-training technology to enhance the model's representation capabilities in the sign language recognition part, enabling sign language video translation.
[0059] like Figure 3 As shown, unlabeled video data is modeled using a human pose extractor skeleton point sequence for masking. This sequence is then input into a pre-trained sign language model (e.g., SignBERT) to learn a general representation of sign language, and the output is a skeleton sequence vector. This pre-training method can significantly improve the model's performance on limited labeled data. Figure 3 The right side illustrates the principle of self-supervised pre-training, which involves masking the skeleton point sequence, reconstructing it using a sign language pre-trained model, and then performing self-supervised training based on the reconstruction results. T represents the length of the skeleton point sequence. Figure 3 The gesture feature extraction, Transformer encoder, and hand model perception decoder shown on the right are all pre-trained sign language models.
[0060] The self-supervised pre-trained model described above, trained on an unlabeled dataset, outputs a skeleton sequence vector containing contextual dependencies of sign language actions. This skeleton sequence vector is then fine-tuned using a large-scale sign language recognition and translation model by aligning it with the corresponding text semantics. During training, a cross-entropy loss function is employed to enhance the model's semantic understanding of sign language. In labeled sign language video data, each video has corresponding labeled text, allowing for fine-tuning so that the model learns to map sign language videos to natural language text. The fine-tuning process integrates cross-modal alignment techniques, primarily using a human pose extractor, a pre-trained sign language model, and a large-scale sign language video recognition and translation model. The aim is to enable the model to understand sign language and generate relevant text translations or perform contextual understanding through multimodal learning. This method combines pose extraction technology with the powerful semantic understanding capabilities of a large-scale sign language recognition and translation model, making it suitable for processing complex sign language video data. Considering that the method provided in this invention can be combined with other related schemes for training, further details are omitted.
[0061] After fine-tuning, the resulting sign language recognition and translation model employs a multimodal learning architecture, fusing video and text information to understand sign language and generate semantic output. The input sign language video passes through a human pose extractor, converting the frame sequence (or pose information) into a fixed-length skeleton point sequence representing the dynamic information in the video. These skeleton point sequences contain information such as gestures and facial expressions, serving as an abstract representation of sign language. This sequence is then input into the sign language recognition and translation model, where the fine-tuned sign language model outputs a skeleton sequence vector. Finally, the fine-tuned sign language recognition and translation model outputs the video text content, serving as the question text.
[0062] In the application scenario (inference stage), deaf people record or upload sign language videos through a camera, which are then transmitted to the backend server via the network. The video data is first preprocessed, and the preprocessed video is then processed by a human pose extractor to extract the corresponding skeleton point sequence, which is then simultaneously fed into the sign language recognition and translation model trained above to predict the question text corresponding to the sign language video. All model parameters are fine-tuned by importing and freezing the weights, and the generated question text will be used as the input for the next step of the large language model.
[0063] 2. Knowledge retrieval and question answering architecture based on a large language model (multi-turn knowledge retrieval and question answering subsystem).
[0064] In knowledge retrieval and question-answering large language models, a reinforcement learning-based retrieval enhancement generation method is incorporated. Within this method, a reinforcement learning algorithm is designed, defining a reward for trustworthy alignment. This allows the question-answering large language model to proactively guide topics and achieve multi-turn dialogues based on the retrieved information. For example... Figure 4The diagram illustrates the knowledge retrieval question-answering architecture and its training and inference processes based on a large language model. The core of this approach involves training a vectorized model (called the knowledge retrieval large language model) by combining a retrieval mechanism with a knowledge vector large language model. This includes efficient use of context windows and end-to-end joint optimization of retrieval and generation. The knowledge retrieval large language model retrieves relevant fragments from external knowledge bases (such as databases and document sets) and concatenates these retrieved context fragments with the input prompts from the question-answering large language model, providing enhanced generation criteria and rich domain knowledge support for the model. Based on the retrieved credible reference information, the question-answering large language model generates accurate and traceable answers, avoiding the illusion problem.
[0065] (1) Dataset construction.
[0066] The question-answering system constructs three types of datasets: the retrieval subsystem dataset, the question-answering subsystem dataset, and the human preference dataset.
[0067] First, the dataset of the retrieval subsystem (knowledge retrieval large language model) focuses on question-knowledge pairs in order to train the model's retrieval capabilities, that is, to quickly and accurately find information related to the user's question from a large amount of data.
[0068] The dataset for the question-answering subsystem (question-answering large language model) is constructed in the form of question-and-answer pairs. The purpose of this dataset is to enable the model to directly answer specific user questions, rather than simply providing relevant information. This requires the model to understand not only the literal meaning of the question but also the deeper intent behind it.
[0069] Human preference datasets focus more on learning user preferences and satisfaction. Through triples of questions, high-quality answers, and near-high-quality answers, the model learns how to select the best option from multiple possibilities. The dataset is used to build and optimize the model based on feedback data, allowing the model to self-optimize and adjust according to actual feedback, thereby continuously improving the quality of question and answer.
[0070] These diverse datasets collectively form the database of a high-efficiency natural language question answering system. This enables the system to not only provide fast and accurate information retrieval but also to directly offer precise answers and optimize based on user preferences. Through these carefully designed and trained models, the natural language question answering system can provide professional and personalized services across various fields, significantly improving service efficiency and quality.
[0071] (2) Construction of a knowledge base of problem-knowledge pairs (problem and knowledge pairs).
[0072] This system utilizes AI (Artificial Intelligence)-driven knowledge decomposition tools to intelligently process unstructured document knowledge from diverse channels and in various formats. It automatically performs semantic understanding and knowledge decomposition using external large language model tools (such as the GPT model), accurately extracting multiple core information from the documents and generating structured question-knowledge tuples for knowledge representation. The decomposed knowledge is then manually verified and labeled to achieve standardization and unification of knowledge format. The GPT model is a generative pre-trained language model.
[0073] (3) Knowledge retrieval big language model.
[0074] Knowledge retrieval large language models are the core technology for achieving efficient semantic search. Through the contrastive prediction (CPC) method of self-supervised learning, they predict the positional relationship of text fragments in the context, capture the semantic patterns behind the data, map the text into dense vectors, and calculate semantic similarity in the vector space.
[0075] By combining the retrieval mechanism with knowledge vectors, a knowledge retrieval big language model is formed. The retrieval stage is introduced into the process of generating answers and guiding questions in multiple rounds. The knowledge retrieval big language model can effectively find and select the best answer or option based on contextual information. It can retrieve relevant knowledge information from outside its own knowledge reserves, enabling the question-answering big language model to dynamically retrieve external knowledge.
[0076] (4) Knowledge recall.
[0077] When a user asks a question, the knowledge retrieval big language model performs preliminary retrieval and knowledge recall (symmetric recall and asymmetric recall). The candidate set of the preliminary recall is coarsely ranked, and then finely ranked. The candidate set is sorted according to the matching degree of the binary pairs, and the top N knowledge (that is, the top N with the highest matching degree) are selected. Through multi-way recall using a combination of various recall algorithms, the top N retrieval results and the scores of each retrieval item are obtained, thereby accurately retrieving the relevant knowledge answers to the question.
[0078] (5) Multi-round question and answer process.
[0079] The question-answering language model is pre-trained using high-quality question-answer pairs to fine-tune its parameters, altering their weights to capture rich semantic information. Furthermore, a context-based learning approach using embedding techniques is employed to enhance performance by embedding questions and answers into the same vector space, enabling the model to capture their semantic relationships. Based on this, the question-answering language model can understand the context of questions, reasoning and making judgments accordingly, thus accurately understanding the intentions and needs of deaf users. It can then combine this understanding with knowledge retrieval results to provide the correct response.
[0080] (6) Intercom management.
[0081] A fine-tuned question-and-answer language model tracks dialogue history and contextual information to dynamically guide and adjust dialogue strategies. Based on dialogue state management technology, a state machine is used to maintain the dialogue context. Historical dialogues are recorded, including question input, the output of the question-and-answer language model, and the current state of the dialogue. Based on the dialogue history and the context of the analyzed dialogue content, the next action is determined. Simultaneously, appropriate dialogue strategies are selected based on the purpose of the dialogue and the needs of the deaf user, and these strategies are adjusted based on user feedback. By defining dialogue strategies (e.g., stopping historical dialogue analysis after a set number of rounds of question-and-answer, ending the dialogue if the user does not ask further questions after a set time), response flows for different situations are designed, such as continuing to provide more specific information or ending the dialogue.
[0082] (7) Human feedback reinforcement learning.
[0083] By leveraging the evaluation results of human experts on the knowledge pairs, we employ human feedback reinforcement learning to align the question-answering language model with human preferences. We use human preferences to learn a reward function and use reinforcement learning to optimize the learned reward to align the model, continuously improving the accuracy of question-answer pair generation and achieving continuous iteration of the question-answering language model.
[0084] The training process of human feedback reinforcement learning consists of three stages. First, supervised fine-tuning allows the pre-trained question-answering language model to initially adapt to the task. Then, data on human preferences for the model's output (such as satisfaction ratings) is collected to train a reward model (RM) that learns human evaluation criteria. Finally, reinforcement learning algorithms (such as PPO) optimize the question-answering language model to maximize the reward for its output. The entire process iteratively optimizes through human feedback, ultimately making the question-answering language model's output more aligned with human values and preferences. This transforms subjective human preferences into quantifiable reward signals, aligning the question-answering language model with human values.
[0085] The knowledge retrieval large language model and the question-answering large language model are collectively referred to as the knowledge retrieval augmented generative large language model (hereinafter simply referred to as the large language model). Its training process is divided into two stages: knowledge retrieval and answer generation. In the knowledge retrieval stage, a large amount of knowledge document data is used to obtain text vectors through comparative learning, and a vectorized knowledge base is constructed. By constraining the knowledge retrieval large language model, semantically related question and answer pairs are mapped to similar vector spaces, thereby training the knowledge retrieval large language model. In the answer generation stage, the question-answering large language model is fine-tuned based on question-answer pair data. A conditional generation task is adopted to enable the model to learn to integrate retrieved knowledge to generate accurate answers. Combined with end-to-end iterative augmentation strategies (such as iterative retrieval-generation), the knowledge retrieval and generation processes are dynamically optimized to improve the model's knowledge utilization efficiency.
[0086] The reasoning process of the knowledge retrieval-enhanced generative language model is a dynamic collaborative process of retrieval and generation. When a user inputs a question, the knowledge retrieval language model first retrieves the most relevant knowledge fragments from the knowledge base. Then, these retrieval results are concatenated with the question text as contextual hints, which are then input into the question-answering language model for answer synthesis. The entire process adopts a cascaded "retrieval-generation" architecture, where retrieval ensures the factual accuracy of the answer (providing up-to-date or proprietary knowledge), while the question-answering language model is responsible for understanding the context and outputting a fluent and natural answer. At the same time, an attention mechanism is used to integrate the retrieved content with pre-trained knowledge, ultimately achieving intelligent question answering that is both accurate and conforms to language habits.
[0087] 3. Sign language video synthesis technology architecture (sign language synthesis subsystem).
[0088] Through an encoder-decoder structure, the encoder converts spoken text into an intermediate vector representation, and the decoder generates sign language text based on this representation. During the inference process, multiple factors are considered, including the semantics, syntax, and contextual information of the spoken text, as well as the grammatical rules of the sign language, to transcribe the sign language text into a sign language text sequence for driving the digital human.
[0089] like Figure 5 The diagram illustrates the workflow of sign language video synthesis, which primarily converts text information into sign language videos for deaf individuals to understand. Its foundation lies in spoken text-to-sign language translation, 3D motion capture, and parametric modeling techniques. To generate coherent and natural sign language videos, a translation model first generates isolated sign language word sequences with sign language semantics. Then, 3D motion capture technology collects a large amount of sign language action data, establishing a database containing various sign language words and actions. Using this sign language word and action library, rule-based action sequence synthesis is employed. Isolated words are converted into action sequences through an action sequence dictionary, quickly matching words in the text to corresponding sign language actions. SmoothNet (a lightweight neural network designed to address the jitter problem in human pose estimation during video) is then used to stitch the action sequences together. During action sequence stitching, motion smoothing and transition generation techniques are also involved. These techniques help simulate smooth transitions in natural human movements, avoiding abrupt changes between actions. Finally, through real-time rendering technology, the system can convert the synthesized sign language actions into animated digital human outputs in real time, allowing deaf individuals to visually see the sign language communication content on a screen.
[0090] (1) text2gloss translation model.
[0091] Figure 6The technical roadmap for training the text2gloss translation model first involves fine-tuning the training using a large amount of well-annotated text2gloss data (spoken text and sign language word sequence text pairs) to obtain a spoken text to sign language text translation model (text2gloss).
[0092] like Figure 6 As shown, the text2gloss translation model employs an encoder-decoder architecture. The encoder converts the answer text output by the question-answering subsystem into an intermediate representation (vector sequence), while the decoder generates sign language word text conforming to sign language word order based on this representation. This architecture allows the model to flexibly generate sign language words while understanding the answer text, achieving an efficient translation process.
[0093] In addition, the text2gloss translation model is equipped with an attention mechanism, which enables the model to focus on different parts of the text during the translation process and dynamically adjust the weights according to the context information, making the translation results more accurate when dealing with long sentences and complex sentence structures.
[0094] The text2gloss translation model takes the answer text from the question-answering model as input and outputs sign language word text that conforms to the sign language word order, ultimately forming a sign language word text sequence (gloss sequence).
[0095] (2) gloss2pose (sign language text to action) sign language word to action sequence.
[0096] First, a sign language word-action sequence lexicon is constructed. The input gloss sequence is converted into an action sequence by matching the lexicon with rules. This sequence is then combined with action smoothing and transition generation techniques for subsequent digital human driving. In sign language, each gloss is a separate word, independent of the others; therefore, it is also called an isolated word sequence. Actions corresponding to multiple glosses can be directly concatenated to form the action sequence of a sign language sentence.
[0097] To achieve the above goals, the first step is to obtain the corresponding action for each gloss. Motion capture equipment is used to record motion capture data for commonly used sign language words. Each motion capture data piece includes precise body, hand, and facial movements to ensure accuracy. For the input gloss sequence, the corresponding action sequence for each gloss is looked up in the vocabulary, and multiple action sequences are concatenated to obtain the action sequence for the entire sentence. The chosen concatenation method is spherical linear interpolation, a method for smooth rotational transitions in three-dimensional space, widely used in computer graphics, game development, and robot control, capable of generating very robust and smooth action sequences.
[0098] To address the jitter issue, a neural network-based motion smoothing method, SmoothNet, is employed to de-jitter motion sequences. SmoothNet introduces an independent, time-dependent, refined network that learns the long-term temporal relationships of each joint without considering noise correlations between joints. This allows the network to efficiently capture the natural smoothness of body motion, significantly improving the temporal smoothness of existing estimators without requiring complex spatiotemporal models. Finally, the generated motion sequences are fed into a rendering engine (e.g., Unity) to drive the motion of the digital human.
[0099] (3) Synthesize sign language videos.
[0100] Import the motion sequences into rendering engines such as Unity and bind and debug them with the digital human model. Adjust animation playback parameters and binding relationships using the animation controller in the rendering engine to ensure correct animation display on the digital human model. Utilize the real-time rendering capabilities of the rendering engine to generate high-quality sign language animations. Apply real-time lighting, shadows, materials, and post-processing effects to enhance visual effects and user experience. By adjusting rendering settings and applying material optimizations, ensure the digital human's image is realistic and its movements are delicate, achieving the best visual presentation.
[0101] By applying the solutions provided in the embodiments of this invention, a complete and efficient information processing closed loop can be constructed. From accurate input recognition of sign language, to in-depth processing and accurate responses based on a strong knowledge base, and finally to output in a natural and fluent sign language form, a comprehensive, one-stop technical architecture is provided for sign language communication and information transmission. Through continuous iteration of the models corresponding to the above three parts, the development of intelligent sign language technology at this stage has been promoted, breaking through technical limitations. It mainly addresses the following pain points in various scenarios: 1) Communication barriers: It breaks down the communication barriers between hearing-impaired and hearing individuals, enabling them to communicate without barriers. 2) Difficulty in information acquisition: Through sign language synthesis technology, hearing-impaired individuals can more easily obtain information from audio-visual content, expanding their information access channels. 3) Professional service needs: In complex scenarios such as medical treatment and legal consultation, it can provide professional translation services to meet the special needs of hearing-impaired individuals. In the long term, this invention can reduce the cost of relying on human translation, improve service efficiency, reduce economic losses caused by communication barriers, significantly improve the social participation and quality of life of the deaf population, promote information equality, and reduce feelings of social isolation. At the same time, the application of video sign language Q&A technology is also a positive practice of technological ethics and social responsibility, promoting the construction of a more inclusive and harmonious social environment.
[0102] Through the above description of the embodiments, those skilled in the art can clearly understand that the above embodiments can be implemented by software, or by using software plus necessary general-purpose hardware platforms. Based on this understanding, the technical solutions of the above embodiments can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, mobile hard drive, etc.), including several instructions to cause a computer device (such as a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present invention.
[0103] Example 2 The present invention also provides a video sign language question-and-answer system, which is mainly used to implement the method provided in the foregoing embodiments, and mainly includes: The sign language recognition and translation subsystem is used to extract skeleton point sequences from input sign language videos and use a sign language recognition and translation model obtained based on self-supervised pre-training and task fine-tuning to perform sign language recognition and translation to obtain the question text; A multi-turn knowledge retrieval and question answering subsystem is used to perform generative question answering based on the question text, using a large language model and under the constraints of a knowledge base; during the generative question answering process, through multiple rounds of interaction with the large language model, the answer text corresponding to the question text is finally generated. The sign language synthesis subsystem is used to convert the answer text into sign language word text that conforms to the sign language word order, form a sign language word text sequence, and then obtain the corresponding action sequence by searching and matching the sign language sequence word library. After processing with action smoothing and transition generation technology, it drives the digital human to generate the corresponding sign language video.
[0104] The overall architecture of the system can also be found in the aforementioned appendix. Figure 2 Furthermore, the relevant technical details have been described in the previous embodiments, and therefore will not be repeated here.
[0105] Those skilled in the art will understand that, for the sake of convenience and brevity, the above-described division of functional modules is used as an example. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the system can be divided into different functional modules to complete all or part of the functions described above.
[0106] Example 3 The present invention also provides a processing device, such as Figure 7 As shown, it mainly includes: one or more processors; a memory for storing one or more programs; wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the method provided in the foregoing embodiments.
[0107] Furthermore, the processing device also includes at least one input device and at least one output device; in the processing device, the processor, memory, input device, and output device are connected via a bus.
[0108] In this embodiment of the invention, the specific types of the memory, input device, and output device are not limited; for example: Input devices can be touchscreens, image acquisition devices, physical buttons, or mice, etc. The output device can be a display terminal; The memory can be random access memory (RAM) or non-volatile memory, such as disk storage.
[0109] Example 4 The present invention also provides a readable storage medium storing a computer program that, when executed by a processor, implements the method provided in the foregoing embodiments.
[0110] In this embodiment of the invention, the readable storage medium is a computer-readable storage medium and can be disposed in the aforementioned processing device, for example, as a memory in the processing device. Furthermore, the readable storage medium can also be any medium capable of storing program code, such as a USB flash drive, portable hard drive, read-only memory (ROM), magnetic disk, or optical disk.
[0111] The above description is merely a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims. The information disclosed in the background section is intended only to enhance the understanding of the overall background technology of the present invention and should not be construed as an admission or implication in any way that such information constitutes prior art known to those skilled in the art.
Claims
1. A video sign language question-and-answer method, characterized in that, include: The skeleton point sequence is extracted from the input sign language video, and the sign language recognition and translation model obtained by self-supervised pre-training and task fine-tuning is used to perform sign language recognition and translation to obtain the question text; Generative question answering is performed by combining a large language model with the question text under the constraints of a knowledge base. During the generative question answering process, the large language model conducts multiple rounds of interaction to finally generate the answer text corresponding to the question text. The answer text is converted into sign language word text that conforms to the sign language word order, forming a sign language word text sequence. Then, the corresponding action sequence is obtained by searching and matching the sign language sequence word library. After processing with action smoothing and transition generation technology, the digital human is driven to generate the corresponding sign language video.
2. The video sign language question-and-answer method according to claim 1, characterized in that, The steps to obtain a sign language recognition and translation model using self-supervised pre-training and task fine-tuning are as follows: Collect unlabeled sign language video data and labeled sign language video data, and select a sign language pre-trained model; Self-supervised pre-training of a sign language pre-training model using unlabeled sign language video data includes: extracting skeleton point sequences from unlabeled sign language video data using a human pose extractor; performing masking modeling on the skeleton point sequences and inputting them into the sign language pre-training model; and performing self-supervised pre-training by combining the skeleton point sequences output by the sign language pre-training model to obtain a self-supervised pre-training model. The self-supervised pre-trained model is fine-tuned using labeled sign language video data, including: extracting skeleton point sequences from the labeled sign language video data using a human pose extractor and inputting them into the self-supervised pre-trained model; fine-tuning the model by aligning the skeleton point sequences with the language text and combining it with a large language model for sign language recognition and translation, thereby obtaining the fine-tuned self-supervised pre-trained model and the fine-tuned large language model for sign language recognition and translation, which together constitute the sign language recognition and translation model; wherein, the language text involved in the alignment mechanism between the skeleton point sequences and the language text is the labeled text in the labeled sign language video data.
3. The video sign language question-and-answer method according to claim 1, characterized in that, The generative question-answering method, which combines a large language model with question text and is constrained by a knowledge base, includes: The large language model includes: a knowledge retrieval large language model and a question-answering large language model; The question text is input into the knowledge retrieval language model, which retrieves the N knowledge items that best match the question text from the knowledge base, and inputs the N knowledge items into the question answering language model to obtain the corresponding N answers; The context learning method based on embedding technology embeds the question text and the corresponding N knowledge points into the same vector space and inputs them into the knowledge retrieval large language model to continue a new round of interaction; and, combined with the output of the question answering large language model, it uses reinforcement learning to control the answering process of the new round of interaction. Finally, the answer text corresponding to the question text is generated.
4. The video sign language question-and-answer method according to claim 3, characterized in that, The knowledge retrieval large language model and the question answering large language model are each trained using pre-collected datasets; The collected datasets include: question and knowledge pair datasets, question and answer pair datasets, and triple datasets of questions, top-quality answers, and second-best-quality answers; Specifically, the problems and knowledge are vectorized into the dataset to construct a vectorized knowledge base; the vectorized knowledge base is then used to train a large language model for knowledge retrieval. We fine-tuned the question-answering language model using a question-and-answer pair dataset, enabling the model to answer questions based on input questions and output corresponding answers. Simultaneously, we optimized the fine-tuned model using reinforcement learning in conjunction with a triplet dataset, ensuring that the optimized model's output answers align with those in the triplet dataset.
5. A video sign language question-and-answer method according to claim 1, characterized in that, Also includes: The answer text is converted into sign language word text that conforms to the sign language word order through the text2gloss translation model. The text2gloss translation model converts natural language text into sign language word text and adopts an encoder-decoder architecture. The encoder is responsible for encoding the answer text into an intermediate representation vector, and the decoder is responsible for decoding the intermediate representation vector into sign language word text that conforms to the sign language word order.
6. A video sign language question-and-answer method according to claim 5, characterized in that, The text2gloss translation model is based on spoken text and sign language word sequence text pairs. It is trained for the task of translating natural text into sign language word text through transfer learning and large model technology.
7. A video sign language question-and-answer method according to claim 1, characterized in that, The process of obtaining corresponding action sequences by searching and matching a sign language sequence lexicon, and then processing them using motion smoothing and transition generation techniques to drive the digital human to generate corresponding sign language videos includes: For each sign language word in the sign language word text sequence, the corresponding action sequence is queried in the sign language sequence word database. The action sequences corresponding to all sign language word texts are concatenated to obtain the action sequence corresponding to the sign language word text sequence. The action sequence corresponding to the sign language word text sequence is processed by action smoothing and transition generation to obtain the final action sequence; The final action sequence is imported into the rendering engine, which drives the digital human in the rendering engine to render the corresponding sign language video.
8. A video sign language question-and-answer system, characterized in that, To implement the method according to any one of claims 1 to 7, comprising: The sign language recognition and translation subsystem is used to extract skeleton point sequences from input sign language videos and use a sign language recognition and translation model obtained based on self-supervised pre-training and task fine-tuning to perform sign language recognition and translation to obtain the question text; The multi-turn knowledge retrieval question answering subsystem is used to perform generative question answering by combining a large language model with the question text under the constraints of a knowledge base. During the generative question answering process, the system interacts with the large language model in multiple turns to finally generate the answer text corresponding to the question text. The sign language synthesis subsystem is used to convert the answer text into sign language word text that conforms to the sign language word order, form a sign language word text sequence, and then obtain the corresponding action sequence by searching and matching the sign language sequence word library. After processing with action smoothing and transition generation technology, it drives the digital human to generate the corresponding sign language video.
9. A processing device, characterized in that, include: One or more processors; Memory, used to store one or more programs; Wherein, when the one or more programs are executed by the one or more processors, the one or more processors cause the one or more processors to implement the method as described in any one of claims 1 to 7.
10. A readable storage medium storing a computer program, characterized in that, When a computer program is executed by a processor, it implements the method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Question and answer method and system based on large language model
CN119166767A
Sign language translation method and device based on vision and word feature pre-training alignment
CN119785439A
Robot sign language communication method, related device and storage medium
CN120578294A
Video question-answer method, device and system, and storage medium
WO2024046038A1