Information processing system

CN122799485APending Publication Date: 2026-09-22SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202610329187.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2025-03-19
Filing Date
2026-03-18
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

[0003]现有技术中,针对聋哑人或听障用户的手语交流,多数系统仅依赖固定规则或传统模式识别算法,将手语动作直接映射为有限的预置短语,难以灵活地生成自然、连贯的文本内容

Benefits of technology

1. 精度提升

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122799485A_ABST
    Figure CN122799485A_ABST
Patent Text Reader

Abstract

The application provides an information processing system. An information processing system, characterized by comprising: a processor; wherein the processor is configured to: capture a sign language through a camera, and analyze image data obtained by the capturing to identify content of the sign language; the processor is configured to: in order to convert the identified sign language content into text data, generate prompt information for instructing a generative artificial intelligence model to generate text data, and generate the text data based on the prompt information; and the processor is configured to: convert the generated text data into voice data by using a voice synthesis technology, and output the voice data through a voice output device.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology disclosed herein relates to an information processing system. Background Technology

[0002] Japanese Patent Application Publication No. 2022-180282 discloses a method for controlling a role-based chatbot executed by at least one processor. The method includes the following steps: receiving a user's speech; adding the user's speech to a prompt word, the prompt word containing instruction statements associated with an explanation of the chatbot's role; encoding the prompt word; and inputting the encoded prompt word into a language model to generate a chatbot response to the user's speech.

[0003] In existing technologies, most sign language communication systems for deaf or hearing-impaired users rely solely on fixed rules or traditional pattern recognition algorithms to directly map sign language gestures into a limited set of pre-defined phrases, making it difficult to flexibly generate natural and coherent text content. Therefore, when sign language expressions are complex, contextually varied, or contain rich semantic information, existing systems often fail to accurately reproduce the meaning of the sign language, making it difficult for hearing users to understand the true intentions of hearing-impaired users in a timely and accurate manner.

[0004] In addition, traditional sign language conversion systems often do not fully consider the user's facial expressions, tone of voice, and other emotional information, resulting in the generated speech output lacking emotional color and tone variation, failing to truly reflect the user's emotional state, and hindering the improvement of the naturalness and affinity of communication.

[0005] On the other hand, some sign language recognition systems do not perform detailed frame-by-frame analysis of video data during video processing. The temporal details and subtle changes of sign language movements cannot be fully captured and utilized, thus affecting recognition accuracy. This is especially true in scenarios with continuous movements, fast speeds, or complex backgrounds, where recognition errors are more likely to occur.

[0006] Therefore, it is necessary to provide a new system that can: on the one hand, utilize generative artificial intelligence models to more flexibly convert sign language content into natural language text, thereby improving the richness and accuracy of sign language expression; on the other hand, combine information such as the user's facial expressions and voice tone to identify emotions and output them with voice tones corresponding to the emotions; and improve the high-precision recognition capability of sign language actions by analyzing video image data frame by frame, so as to solve the above-mentioned problems existing in the prior art. Summary of the Invention

[0007] To address the aforementioned issues, this invention provides an information processing system comprising a processor. The processor captures sign language using a camera and analyzes the captured image data to identify the content of the sign language. By incorporating an image analysis function into the processor, motion features can be extracted from continuous sign language videos to obtain semantic information corresponding to the user's sign language actions.

[0008] In order to convert the identified sign language content into text data, the processor generates prompting information to instruct the generative artificial intelligence model to generate text data, and generates the text data based on the prompting information. By introducing a generative artificial intelligence model into the processor and constructing prompting information for the model that includes sign language semantics and contextual information, the generative artificial intelligence model can dynamically generate natural, coherent, and semantically rich text data according to the meaning of sign language, thereby breaking through the limitations of traditional rule-based mapping methods in terms of expressive power.

[0009] The processor uses speech synthesis technology to convert the generated text data into speech data, and then outputs the speech data through a speech output device. In this way, the content expressed by hearing-impaired users through sign language can be automatically converted into a speech signal and played by the speech output device, enabling hearing users to directly understand the sign language content through hearing, thereby achieving natural communication between hearing-impaired and hearing users.

[0010] Furthermore, to make the output speech more natural and emotionally resonant, in the system of this invention, the processor identifies the user's emotions by analyzing the user's facial expressions and / or the tone of the user's voice, and generates the speech data with a corresponding tone based on the identified emotions. By adding an emotion recognition function to the processor, the user's current emotional state, such as happiness, anger, or sadness, can be obtained, and this emotional information can be mapped to parameters such as tone, speech rate, and volume of the speech, making the synthesized speech not only semantically accurate but also emotionally consistent with the user.

[0011] Furthermore, to improve the accuracy of sign language recognition, in the system of this invention, the processor captures sign language images using a camera and analyzes the captured image data frame by frame, thereby recognizing sign language movements with high precision. By analyzing each frame, the processor can precisely capture the shape and positional changes of the gestures in each frame, as well as their relative relationships with other parts of the body. Based on time-series information, it can model the sign language movements, significantly improving the recognition accuracy of continuous sign language movements and effectively handling scenarios with rapidly changing movements or complex backgrounds.

[0012] Through the above-mentioned technical means, the system of the present invention uses a camera to collect sign language image data, adopts frame-by-frame analysis and high-precision motion recognition technology to obtain sign language semantics, and on this basis, converts the sign language content into natural language text through a generative artificial intelligence model. Then, combined with emotion recognition and speech synthesis technology, the system outputs the text content with emotion matching speech, thereby solving the problems of unnatural expression of sign language content, lack of emotional information and low recognition accuracy in the existing technology.

[0013] "System" refers to an overall device or combination of devices including at least one processor and hardware and / or software components that work together with it to realize functions such as sign language acquisition, parsing, text generation, speech synthesis and output.

[0014] A "processor" is an electronic processing unit that can execute program instructions, perform calculations, control, and logical processing on input data, thereby enabling functions such as sign language recognition, prompt information generation, calling generative artificial intelligence models, and speech synthesis. It can be a single chip, multiple chips, or a combination of processing units.

[0015] "Camera" refers to an image acquisition device used to collect sign language video data, including but not limited to built-in camera modules, external camera devices, or other imaging devices capable of acquiring continuous image frames.

[0016] "Sign language" refers to a symbolic system that expresses information through non-verbal means such as hand movements, hand shape changes, body postures, and facial expressions, including but not limited to natural sign languages ​​of various countries or regions and conventional signals based on gestures.

[0017] "Image data" refers to image information captured by a camera to represent sign language actions, including single-frame images, sequences of consecutive multi-frame images, or encoded video data.

[0018] "Analysis" refers to the process of processing, analyzing, and extracting features from image data in order to identify movement information, spatial location, and temporal changes related to sign language.

[0019] "The content of sign language" refers to the semantic information conveyed through sign language gestures, including the meaning of words, phrases, sentences, or longer texts.

[0020] “Text data” refers to natural language information represented in the form of characters or strings, used to express the semantic content contained in sign language in written form.

[0021] "Generative artificial intelligence models" refer to models built based on machine learning or deep learning techniques that can automatically generate text data based on input prompts, including but not limited to large language models or other text generation models.

[0022] "Prompt information" refers to instructional or guiding text or structured data constructed by a processor and input into a generative artificial intelligence model, used to instruct the generative artificial intelligence model to generate corresponding text data according to the expected semantics and format.

[0023] "Speech synthesis technology" refers to the technical process of converting text data into audible speech signals, including steps such as text analysis, acoustic feature generation, and speech waveform synthesis.

[0024] "Speech data" refers to digital audio information generated by speech synthesis technology that can be played back, used to express text content in sound form.

[0025] "Speech output device" refers to a device used to convert speech data into sound wave signals that can be heard by the human ear, including but not limited to speakers, headphones or other audio playback devices.

[0026] "User's facial expressions" refer to the visual facial expressions displayed by users through changes in facial muscles during sign language expression or interaction, which are used to reflect the user's emotional state or attitude.

[0027] "User voice pitch" refers to the acoustic characteristics of a user's voice, such as pitch and intonation, which are used to reflect the user's emotional inclination and tone.

[0028] "Emotion" refers to the psychological and emotional state reflected by a user's facial expressions and / or tone of voice, including but not limited to emotions such as happiness, anger, sadness, surprise, and calmness.

[0029] "Pitch" refers to the combination of parameters such as pitch, stress, speech rate, and tone changes reflected in speech data, which are used to express different emotional tones or tone styles.

[0030] "Frame-by-frame analysis" refers to an analysis method that divides continuous image data into basic time units of frames and performs image processing and feature extraction on each frame separately.

[0031] "High-precision sign language recognition" refers to the process of accurately reconstructing the content expressed by sign language by comprehensively analyzing frame-by-frame image data and its temporal relationships, thereby identifying the specific type, sequence, and combination of sign language movements with high accuracy and robustness. Attached Figure Description

[0032] Figure 1 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the first embodiment.

[0033] Figure 2This is a conceptual diagram illustrating an example of the main functions of the data processing apparatus and smart device according to the first embodiment.

[0034] Figure 3 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the second embodiment.

[0035] Figure 4 This is a conceptual diagram illustrating an example of the main functions of the data processing device and smart glasses according to the second embodiment.

[0036] Figure 5 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the third embodiment.

[0037] Figure 6 This is a conceptual diagram illustrating an example of the main functions of the data processing apparatus and head-mounted terminal according to the third embodiment.

[0038] Figure 7 This is a conceptual diagram illustrating an example of the configuration of the data processing system according to the fourth embodiment.

[0039] Figure 8 This is a conceptual diagram illustrating an example of the main functions of the data processing device and robot according to the fourth embodiment.

[0040] Figure 9 This represents an emotion map that maps multiple emotions.

[0041] Figure 10 This represents an emotion map that maps multiple emotions.

[0042] Figure 11 This is a sequence diagram illustrating the processing flow of the data processing system of the first embodiment.

[0043] Figure 12 This is a sequence diagram illustrating the processing flow of the data processing system in Application Example 1.

[0044] Figure 13 This is a sequence diagram illustrating the processing flow of the data processing system of the second embodiment.

[0045] Figure 14 This is a sequence diagram illustrating the processing flow of the data processing system in Application Example 2. Detailed Implementation

[0046] Hereinafter, an example of an implementation of the system according to the present disclosure will be described with reference to the accompanying drawings.

[0047] First, let me explain the terminology used in the following instructions.

[0048] In the following embodiments, the processor (hereinafter referred to as "processor") with reference numerals may be a single computing device or a combination of multiple computing devices. Furthermore, the processor may be a single computing device or a combination of multiple computing devices. Examples of computing devices include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), etc.

[0049] In the following embodiments, RAM (Random Access Memory), as indicated in the figures, is a memory that temporarily stores information and is used as working memory by the processor.

[0050] In the following embodiments, the memory, as indicated by the reference numerals, is one or more non-volatile storage devices that store various programs and parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), disks (e.g., hard disks), or magnetic tapes.

[0051] In the following embodiments, the communication I / F (Interface) with reference numerals is an interface that includes a communication processor and an antenna, etc. The communication I / F is responsible for communication between multiple computers. As an example of a communication specification applicable to the communication I / F, wireless communication specifications such as 5G (5th Generation Mobile Communication System), Wi-Fi (wireless fidelity) (registered trademark), or Bluetooth (registered trademark) can be listed.

[0052] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B". That is, "A and / or B" means that it can be only A, only B, or a combination of A and B. Furthermore, in this specification, when "and / or" connects to express more than three items, the same interpretation as "A and / or B" applies.

[0053] First Implementation Method Figure 1 An example of the configuration of the data processing system 10 according to the first embodiment is shown.

[0054] like Figure 1 As shown, the data processing system 10 includes a data processing device 12 and an intelligent device 14. A server can be cited as an example of the data processing device 12.

[0055] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0056] The smart device 14 includes a computer 36, a receiving device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. In addition, the receiving device 38, output device 40, camera 42, and communication I / F 44 are also connected to the bus 52.

[0057] The receiving device 38 includes a touchscreen 38A and a microphone 38B, and receives user input. The touchscreen 38A receives user input via touch by detecting contact with an indicator (e.g., a pen or finger). The microphone 38B receives user input via sound by detecting the user's voice. The control unit 46A in the processor 46 sends data representing the user input received by the touchscreen 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data representing the user input.

[0058] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting data in a form perceptible to the user 20 (e.g., sound and / or text). The display 40A displays visual information such as text and images according to instructions from the processor 46. The speaker 40B outputs sound according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0059] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for sending and receiving various information between processor 46 and processor 28 via network 54.

[0060] Figure 2 The diagram shows an example of the main functions of the data processing device 12 and the smart device 14.

[0061] like Figure 2 As shown, in the data processing apparatus 12, specific processing is performed by the processor 28. A specific processing program 56 is stored in the memory 32. The specific processing program 56 is an example of a "program" as understood in this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0062] The memory 32 stores a data generation model 58 and an emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290. The specific processing unit 290 can use the emotion-specific model 59 to infer the user's emotions and perform specific processing based on the user's emotions. In the emotion inference function (emotion-specific function) using the emotion-specific model 59, various inferences and predictions related to the user's emotions are performed, including inferences and predictions of the user's emotions, but this is not limited to this example. Furthermore, emotion inference and prediction may also include, for example, emotion analysis (parsing).

[0063] In the smart device 14, the processor 46 performs the acceptance output processing. The memory 50 stores the acceptance output program 60. The acceptance output program 60 is used in conjunction with the data processing system 10 and the specific processing program 56. The processor 46 reads the acceptance output program 60 from the memory 50 and executes the read acceptance output program 60 on the RAM 48. The acceptance output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance output program 60 executed on the RAM 48. Furthermore, the smart device 14 has the same data generation model and emotion-specific model as the data generation model 58 and the emotion-specific model 59, and these models can also be used to perform the same processing as the specific processing unit 290. The acceptance output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance output program 60 executed on the RAM 48.

[0064] Alternatively, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device (e.g., a generation server) may have the data generation model 58. In this case, the data processing device 12 obtains the processing results (prediction results, etc.) using the data generation model 58 by communicating with the server device that has the data generation model 58. Furthermore, the data processing device 12 may be a server device or a user-held terminal device (e.g., a mobile phone, robot, home appliance, etc.). Next, an example of the processing of the data processing system 10 of the first embodiment will be described.

[0065] Example 1 The flow of a specific process in Example 1 will be described. Each part of the system described below is implemented by the data processing device 12 and the smart device 14. Furthermore, the data processing device 12 is referred to as the "server," and the smart device 14 is referred to as the "terminal."

[0066] In recent years, with the development of computer vision technology, speech synthesis technology, and generative artificial intelligence models, human-computer interaction methods have become increasingly diversified. However, when it comes to real-time communication between deaf and other hearing-impaired individuals and hearing individuals, existing technologies still have the following shortcomings in terms of computer processing flow and system architecture.

[0067] First, existing sign language recognition systems mostly use a fixed end-to-end model, directly inputting the image data obtained from the camera into the recognition model to output text results. This lacks structured modeling of intermediate features such as temporal body part position information and shape information, resulting in poor robustness of the model to complex postures, continuous movements and environmental changes, insufficient recognition accuracy and interpretability, and difficulty in working stably in low-latency real-time scenarios.

[0068] Second, existing systems typically embed text generation logic within the recognition model, lacking a unified interface and prompt design mechanism for flexible natural language generation using generative artificial intelligence models. This prevents them from fully utilizing the advantages of generative artificial intelligence models in contextual understanding, word order adjustment, and semantic refinement, resulting in output text that is often stiff and lacks contextual coherence, thus affecting the quality of human-computer dialogue.

[0069] Third, existing text-to-speech conversion processes mostly involve simply calling a speech synthesis engine to directly convert the recognition result into audio output. They lack system-level structured management and associated storage of the entire data chain, from "image features—intermediate recognition information—prompt statements—text data—audio data." Therefore, it is difficult to systematically backtrack and optimize the parameters of image processing algorithms and generative artificial intelligence models based on user corrections at the front end, making it difficult for the system to continuously learn and improve its performance through online or offline methods.

[0070] Fourth, from a computer system architecture perspective, existing technologies do not fully reveal how servers hierarchically divide tasks among internal modules: for example, how image feature extraction is performed at the lower level, intermediate recognition information is generated, and then generative artificial intelligence models and speech synthesis algorithms work together at the upper level; nor do they explain how text information and audio data are associated and stored as historical information in the form of interaction with the terminal to support subsequent model retraining and algorithm optimization. This results in deficiencies in the system's scalability, maintainability, and resource scheduling.

[0071] Therefore, a technical solution is needed to improve the computer system and program processing flow, enabling the server to: (1) Efficiently and structurally extract temporal position information and shape information of body parts from video images and generate intermediate recognition information; (2) Construct prompt statements adapted to the generative artificial intelligence model based on intermediate recognition information, generate text data in natural language form by the generative artificial intelligence model, and optimize it through language processing algorithms; (3) Unify the management of the entire process of data from image information to text information and then to audio data, and store it in association with user correction information to provide a foundation for subsequent model learning and continuous improvement of system performance; This improves the conversion process from sign language and other physical expressions to speech output at the computer technology level, thereby enhancing the overall system's recognition accuracy, real-time performance, scalability, and maintainability.

[0072] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Embodiment 1 is achieved by the following means.

[0073] In this invention, the server includes: a device for acquiring image information containing body movements using an image acquisition device; a device for performing feature extraction processing on the image information, calculating temporal body part position information and shape information using an image processing algorithm, and generating intermediate recognition information representing the semantic content of posture expression based on the calculation results; a device for generating instruction information consisting of prompt statements to instruct a generative artificial intelligence model to generate text information based on the intermediate recognition information, and inputting the instruction information into the generative artificial intelligence model to generate text data in natural language form; a device for performing language processing algorithms on the generated text data, performing word order correction and expression unification, and outputting text information organized into sentences; a device for using the organized text information as input to execute a speech synthesis algorithm to generate audio data; a device for sending the audio data to a sound output device on the terminal side, and causing the sound output device to output the audio data in the form of physical sound; and a device for associating the text information and the audio data and storing them in a recording medium as historical information that can be used for subsequent learning processing or improvement processing. This allows for a bottom-up, layered processing structure within the server: the bottom layer extracts temporal features from image data and generates intermediate recognition information; the middle layer interacts with generative AI models through prompts to obtain high-quality text data; and the top layer generates human-readable audio output through language processing and speech synthesis. The entire data chain is then linked and stored with user correction information. This improves the accuracy and efficiency of sign language and other gesture recognition and speech output processing at the computer technology level. Furthermore, it supports continuous optimization of the recognition and generative AI model's operating parameters through learning from historical data, thereby improving overall system performance and resource utilization.

[0074] "Image acquisition device" refers to a hardware device used to collect visual information containing scenes of body movements from physical space and generate corresponding image data, including but not limited to cameras, video capture cards, and imaging sensors.

[0075] "Image information" refers to image or video data about objects in a scene that are acquired and digitally represented by an image acquisition device, including at least recognizable image content of body parts and their movements.

[0076] "Feature extraction processing" refers to the process of applying image processing algorithms or machine learning algorithms to image information to calculate numerical features that can describe the position, shape and time changes of body parts.

[0077] "Image processing algorithm" refers to software algorithms that run on computing devices and are used to analyze, transform, segment, detect and recognize image information, including but not limited to edge detection algorithms, key point detection algorithms, target tracking algorithms and image analysis algorithms based on machine learning or deep learning.

[0078] "Body part location information" refers to numerical data used to represent the position of the human body or body parts (including the trunk, limbs, hands, etc.) in an image coordinate system or a three-dimensional spatial coordinate system, including but not limited to key point coordinates, skeleton point coordinates and their relative positional relationships.

[0079] "Shape information" refers to numerical data used to represent the geometric shape and posture characteristics of body parts at a given moment, including but not limited to outline shape, joint angles, finger opening and closing degree, and local structural features.

[0080] "Intermediate recognition information" refers to the intermediate layer data expression form generated based on feature extraction of image information, used to represent the semantic content of posture expression or gesture. This data is not directly presented to the end user, but is used for subsequent text generation or further semantic reasoning.

[0081] "Body posture" refers to a form of expression that conveys information, intentions, or semantic content through non-verbal means such as body posture, movements, gestures, or facial expressions.

[0082] "Generative artificial intelligence models" refer to artificial intelligence models that learn the mapping relationship between input and output based on training data and can automatically generate text and other content according to input conditions, including but not limited to language generation models based on deep learning.

[0083] "Prompt statements" refer to structured or semi-structured textual information used as input to generative artificial intelligence models. This information is used to explain the generation task to the model, provide context or constraints, and guide the model to generate target textual data.

[0084] "Instruction information" refers to a data set containing prompt statements and parameters or auxiliary information related to the generation task, which is used as input to a generative artificial intelligence model to trigger text generation processing.

[0085] “Text data in natural language form” refers to text data that conforms to the grammar and expression habits of natural language, including phrases, sentences or paragraphs, and is represented in a common character encoding manner.

[0086] "Language processing algorithms" refer to software algorithms that perform syntactic analysis, semantic analysis, word order adjustment, word form standardization, and text polishing on text data to improve the readability and naturalness of the language.

[0087] "Textual information organized into sentences" refers to text content that has a complete grammatical structure, reasonable word order, and conforms to semantic logic after being processed by language processing algorithms.

[0088] "Speech synthesis algorithm" refers to the computational methods and program flow that convert text information into playable audio signals, including processing steps such as text analysis, phoneme conversion, prosody generation, and vocoder synthesis.

[0089] "Audio data" refers to digital data generated by speech synthesis algorithms to represent sound waveforms, including but not limited to pulse code modulation data or compressed audio file data.

[0090] "Sound output device on the terminal side" refers to a hardware device installed on a terminal device that converts audio data into physical sounds that can be heard by the human ear, including but not limited to speakers, headphones, and headsets.

[0091] "Physical sound" refers to acoustic signals that are emitted through sound output devices and propagate in the form of sound waves through media such as air, and can be perceived by the human ear.

[0092] "Recording medium" refers to a medium capable of storing digital data in a non-transient or semi-permanent manner, including but not limited to semiconductor memory, magnetic storage medium, optical storage medium, and storage devices based on the aforementioned media.

[0093] "Historical information" refers to time-series data generated during system operation and stored in recording media, which is related to image information, text information, audio data and their corresponding relationships, and is used for subsequent analysis, learning or improvement processing.

[0094] "Learning processing" refers to the process of updating the parameters of image processing algorithms, generative artificial intelligence models, or other models based on historical information or user-corrected information in order to improve the system's performance in recognition, generation, or reasoning tasks.

[0095] "Running parameters" refer to adjustable values ​​used to control the behavior of image processing algorithms, generative artificial intelligence models, or other algorithm modules, including but not limited to network weights, thresholds, hyperparameters, and configuration items.

[0096] "Temporal features" refer to numerical data used to represent the characteristics of changes in body parts or image content over time, including but not limited to displacement, velocity, acceleration, trajectory shape, and time-related statistical features.

[0097] In one embodiment of the present invention, a server, a terminal, and a user collaboratively constitute a system for converting gesture expressions (e.g., sign language) into text information and further into audio data. The following provides a detailed description of this embodiment of the present invention from aspects such as program generation, hardware and software configuration, data structure and algorithm flow, the use of generative artificial intelligence models, and the technical effects of the system.

[0098] I. Overall Structure of the Program and System In this embodiment, the server serves as the core computing node, equipped with a processor, memory, graphics processing unit, and network interface. In one embodiment, the server employs a general-purpose computer hardware platform, and the operating system can be a server operating system. The server is equipped with a deep learning framework, such as a tensor computing framework or another deep learning framework, a computer vision library, such as a general-purpose image processing library, and a multimedia processing library, such as an audio / video codec library, for decoding video streams transmitted from the terminal. The server also installs a natural language processing toolkit, such as a language model library, and a speech synthesis engine, such as a text-to-speech engine or a local speech synthesis component.

[0099] In one embodiment, the terminal is a mobile terminal or computing terminal equipped with a camera and an audio output device, such as a smartphone, tablet device, or computing terminal equipped with a camera and a speaker. The terminal runs a client application or a web application, which calls the operating system's camera interface and audio output interface to capture image information and play audio data.

[0100] In one implementation, users include those who express themselves through gestures (e.g., people with hearing impairments) and those who understand content through hearing (e.g., people with normal hearing). Users interact with the server through a terminal and can correct the recognition results for subsequent learning and processing by the server.

[0101] II. Server-side program logic and module composition 1. Image Acquisition and Preprocessing Module In one embodiment, the server receives the video data stream uploaded by the terminal via a network protocol (e.g., a long connection based on Transmission Control Protocol, Real-Time Transport Protocol, or Real-Time Communication Protocol). The terminal uses a camera to capture continuous image frames containing the user's upper body and hand movements, encodes them (e.g., using a video compression standard), and then sends them to the server.

[0102] The server uses a multimedia processing library to decode the received video stream, obtaining frame-by-frame image data. Each frame can be represented as a two-dimensional pixel matrix, such as a three-channel integer matrix with dimensions of width × height. The server applies image processing algorithms to each frame, including operations such as resizing, color normalization, noise filtering, and contrast enhancement, to form standardized image information suitable for subsequent feature extraction. This preprocessing helps reduce the model's sensitivity to changes in illumination and resolution, thereby improving recognition robustness.

[0103] 2. Body Part Key Point and Temporal Feature Extraction Module In one implementation, the server invokes a keypoint detection model based on a deep neural network. This model can employ a convolutional neural network structure and a multi-stage regression structure to detect keypoints on the human body and hands. For each frame of image, the server calculates the image coordinates of keypoints such as the wrist, palm, and finger joints, and normalizes the coordinates (e.g., normalizes them to the range of 0 to 1) to form "body part position information".

[0104] The server performs temporal modeling on keypoint sequences across multiple consecutive frames, calculating the changes in keypoint displacement, velocity, acceleration, and joint angles between adjacent frames as "temporal features." In one embodiment, the server uses a sliding window mechanism to combine keypoints and their temporal features within a fixed time period into a high-dimensional feature tensor, which is then used as input to the subsequent semantic recognition network. Compared to directly using the original pixel sequence, this structured temporal feature extraction significantly reduces the input dimensionality and enhances the ability to abstract the essential patterns of sign language actions, thereby improving recognition accuracy and inference speed.

[0105] 3. Intermediate identification information generation and data structure In one embodiment, the server uses a sequence-modeling-based neural network architecture, such as a model composed of convolutional neural networks and recurrent neural networks or attention networks. The server inputs the aforementioned temporal feature tensor into the model, performs forward propagation on the graphics processing unit, and calculates the "semantic label probability distribution" corresponding to each time step. The server uses a connection-based temporal classification decoding or bundle search strategy to decode the probability sequence into a discrete label sequence, such as a series of identifiers representing basic sign language words or action units.

[0106] The server associates the tag sequence with the corresponding time range to form "intermediate identification information." In one data structure implementation, this intermediate identification information may include: action segment identifiers, tag ID sequences, confidence scores for each tag, start and end timestamps, and reference pointers to the original keypoint sequence. This intermediate identification information serves as a bridge layer from the visual domain to the language domain, offering advantages such as structure, compactness, and interpretability, which facilitates subsequent generative AI models in generating natural language based on higher-level semantics.

[0107] III. Generative Artificial Intelligence Models and Prompt Statement Construction 1. Construction of prompt statements and instruction messages In this embodiment, the server does not directly input the raw image information into the generative artificial intelligence model. Instead, it first converts the intermediate recognition information into textualized, abstract prompts. The server can generate descriptive text based on the tag sequence, specifying, for example, the basic meaning, order, and contextual constraints of the tags. An example of a prompt constructed by the server in one embodiment is as follows: "Please generate a coherent target language sentence based on the following action label sequence: Label sequence: [Label 1, Label 2, Label 3]. The meanings of each label are: Label 1 represents 'I', Label 2 represents 'tomorrow', and Label 3 represents 'go to the hospital'. Please output the complete sentence in natural language form." The server bundles prompts with relevant parameters (such as target language type, output length limits, etc.) to form "instruction information," which is then input into the generative AI model in text or structured data form. Through this prompt mechanism, the server can adapt the same generative AI model to different types of body language expressions and different language outputs without needing to design separate rules for each language or action pattern, thereby improving the system's flexibility and scalability.

[0108] 2. Internal Structure and Training Methods of Generative Artificial Intelligence Models In one embodiment, the server uses a sequence generation model based on a self-attention mechanism as a generative artificial intelligence model. The encoding part of this model receives the text sequence contained in the prompt statement, first mapping discrete word units into vector representations, and then extracting contextual features through a multi-layer self-attention and feedforward network. The decoding part generates output text data word by word within an autoregressive framework, making probability predictions at each step based on previously generated words and the contextual features from the encoding side.

[0109] During the model training phase, the server pre-trains using a large-scale text corpus and then fine-tunes it using labeled corpora corresponding to the intermediate recognition information. The server employs a cross-entropy loss function to measure the difference between the generated sequence and the target reference sequence, and uses gradient descent-like algorithms to update the network weight parameters. The server can also use data augmentation techniques during training, such as randomly rewriting the wording of prompts, to enhance the model's robustness to different prompt styles.

[0110] Through the above structural design, the generative AI model can generate natural and fluent text data based on its semantic patterns after receiving abstract descriptions of intermediate recognition information, rather than simply performing template replacement. This configuration enables the server to support more complex combinations of body expressions and contextual reasoning within the same framework, achieving an overall improvement in the quality of text generation.

[0111] IV. Text Information Standardization and Speech Synthesis 1. Language processing algorithms and text information standardization After the generative AI model outputs text data in natural language form, the server uses language processing algorithms to further standardize it. The server can integrate word segmentation tools, syntactic analysis tools, and language correction rules to perform lexical unification, word order fine-tuning, and spell checking on the output text. For expressions containing time, numbers, or proper nouns, the server can call rule bases or dictionary bases for standardized representation, such as parsing "tomorrow afternoon three o'clock" into a standardized time expression.

[0112] The server adds metadata to the normalized text information, including language category, timestamp, and session identifier, providing an index for subsequent audio generation and historical record storage. Through this layered language processing mechanism, the server can add another layer of rules and toolchain control on top of the generative artificial intelligence model, further reducing grammatical errors and unnatural expressions, thereby improving the understandability of downstream speech synthesis.

[0113] 2. Speech Synthesis and Audio Data Generation In one implementation, the server invokes a speech synthesis engine to convert standardized text information into audio data. The speech synthesis engine can employ a two-stage architecture: first, it uses a sequence-to-sequence acoustic model to map the text sequence into a sequence of acoustic features; then, it uses a neural vocoder model to convert the acoustic features into time-domain waveforms. The server selects parameters such as speaker type, speech rate, and volume based on system configuration to adapt to different application scenarios.

[0114] The server encodes the generated audio data into a compressed audio format, such as common audio encoding formats, to reduce network bandwidth consumption. The server also retains the original sampling rate and channel count information, enabling the terminal to correctly reproduce the audio signal. By associating and storing text information with audio data using a unified identifier, the server can achieve traceable management of the entire conversion chain, providing a data foundation for subsequent performance analysis and error diagnosis.

[0115] V. Terminal-side applications and their coupling with real-world technologies In this embodiment, the terminal is responsible for acquiring image information of the user's posture and outputting corresponding audio results. The terminal calls the operating system's camera interface to capture video frames at a fixed frame rate, and adaptively adjusts the resolution or compression ratio according to network conditions to reduce network load. The terminal parses the audio data received from the server, calls the system's audio decoding function to decode the compressed format into linear pulse code modulation data, and plays the physical sound through the audio output device.

[0116] The terminal can also simultaneously display text information on the screen, allowing users to access content visually even in noisy environments or when they have hearing impairments. In this way, the system not only completes abstract data processing but also integrates closely with physical devices such as cameras and speakers, realizing end-to-end technology applications from real-world body movements to auditory output.

[0117] VI. User Correction and Server Learning Processing In one implementation, users can view the text information generated by the server through a terminal interface and manually correct any errors found. The terminal packages the original text information, intermediate recognition information, and user correction results and sends them to the server. The server saves this data as "historical information" to a recording medium.

[0118] During offline or low-load periods, the server learns from historical information. The server can employ supervised learning, using intermediate recognition information and user-corrected text as training samples. This fine-tunes the parameters of the generative AI model, enabling it to generate outputs closer to human corrections when similar prompts are input. Simultaneously, it adjusts parameters such as thresholds and classifier weights in the image processing algorithm to reduce the occurrence of similar errors. In some embodiments, the server can also retrain the neural network of the temporal feature extraction module to more accurately distinguish similar but different body movements.

[0119] By incorporating user feedback and converting it into parameter updates within the server, this system forms a closed-loop learning mechanism, thereby achieving continuous improvement in model accuracy at the computer technology level, rather than simply simplifying manual proofreading work from a business process perspective.

[0120] VII. Technical Effects and Improvements in Computer Technology This invention achieves the following technical effects by introducing an intermediate recognition information layer, a generative artificial intelligence model driven by prompt statements, and end-to-end data association and storage: 1. Improved accuracy Instead of directly classifying the raw pixel stream, the server utilizes structured body part location information and temporal features, allowing the model to focus on key features highly correlated with posture semantics, thereby reducing interference from background noise and lighting variations. The intermediate recognition information layer makes it easier to pinpoint errors in the feature extraction or text generation stages, facilitating targeted optimization and improving overall recognition accuracy.

[0121] 2. Improved processing speed and computational efficiency The server significantly reduces the input dimensionality by compressing keypoints and temporal feature vectors, making the computational cost of subsequent deep models significantly lower than that of directly processing the original image sequences. Performing these tensor operations on the graphics processing unit further improves throughput, thereby achieving near real-time response.

[0122] 3. Improved communication load and data management The terminal sends the raw video stream at an adjustable resolution and compression ratio. The server then performs subsequent processing and storage using intermediate identification and text information, avoiding the duplication of large video files. Only raw video segments are retained when necessary, significantly reducing storage and network bandwidth requirements.

[0123] 4. Enhanced scalability and maintainability The server separates image feature extraction, semantic recognition, text generation, language processing, and speech synthesis modules, and connects them with a unified data structure through prompts. This allows for future replacement of individual modules (such as replacing a new generative artificial intelligence model or a new speech synthesis engine) without affecting the overall framework. This modular structure improves scalability and maintenance efficiency from a computer system architecture perspective.

[0124] 5. The fundamental difference from traditional manual processing methods This system internally employs high-dimensional feature modeling and end-to-end learning mechanisms without explicit rules, rather than mimicking the rule-by-rule judgments of human translators. The server's neural network automatically seeks optimal parameters through error functions and gradient updates, forming complex decision boundaries on a large amount of historical data. This decision-making process takes place in a high-dimensional space and is not simply an automation of human tasks. Rather, it leverages the advantages of computers in massively parallel computing and high-dimensional pattern recognition to improve the computer processing flow itself.

[0125] 8. Examples of prompts used in generative artificial intelligence models In one implementation, the server can interact with the generative artificial intelligence model using the following example prompts: "Please describe in detail how a system uses cameras, computer vision algorithms, and deep learning models to recognize the sign language of deaf users in real time and convert the recognition results into text and speech output. Please describe step by step what hardware and software the server uses at each stage, and what processing and operations are performed on the video data, feature data, text data, and audio data respectively." "I want to design a real-time 'sign language → text → speech' conversion system. Please explain in detail from the perspectives of the server and the terminal: how the server receives the video stream uploaded by the terminal, how it uses image processing libraries and deep learning frameworks for sign language recognition, how it uses natural language processing technology to generate text information, how it calls the speech synthesis engine to generate audio, and how it is played through the terminal's speaker. Please provide a complete processing flow." "Please use natural language to break down the processing steps of a sign language recognition and speech synthesis system. Please explain: 1) the hardware environment used by the server, such as the camera, graphics processing unit, and operating system; 2) the names of the computer vision libraries, deep learning frameworks, natural language processing tools, and speech synthesis modules called by the server; 3) the conversion process of video data, feature vectors, text data, and audio data in each step." Through the above-described embodiments, the present invention implements a hierarchical processing structure within the server, from image information to intermediate recognition information and then to natural language text and audio output. Combined with generative artificial intelligence models and user correction feedback, it achieves a systematic improvement in the process of body expression recognition and voice output processing at the computer technology level.

[0126] use Figure 11 The processing flow is explained.

[0127] Step 1: The terminal uses the camera to collect image information The terminal calls the operating system's camera interface to activate the front or rear camera and capture the user's posture and movements in real time.

[0128] Input: User's body movements in the physical space.

[0129] Output: A stream of video frame data arranged in chronological order (compressed encoding format).

[0130] The terminal encodes continuously acquired image frames into a video data stream according to the preset resolution and frame rate, and adds a timestamp and session identifier to each data packet locally, ready to send it to the server.

[0131] Step 2: The terminal sends the video data stream to the server. The terminal establishes a session with the server via a network connection and sends the encoded video data packets to the server in chronological order using a communication protocol.

[0132] Input: The video frame data stream generated in step 1 and its metadata (timestamp, session ID).

[0133] Output: Continuous video data packets transmitted over the network to the server.

[0134] The terminal dynamically adjusts video encoding parameters (such as bitrate and resolution) based on network bandwidth to reduce latency and packet loss while ensuring basic image quality.

[0135] Step 3: The server receives and decodes the video data. The server listens on the network port, buffers video data packets from the terminals in the order they are received, and verifies the integrity of the data.

[0136] Input: A sequence of video encoded data packets sent by the terminal.

[0137] Output: A sequence of original image frames (pixel matrix) arranged in chronological order.

[0138] The server uses a multimedia decoding library to decode the compressed video, restoring each data packet to a two-dimensional pixel matrix, and appending the server's receiving time and the original timestamp to each frame to form a standardized sequence of image information.

[0139] Step 4: The server preprocesses the image frames. The server performs image preprocessing on each decoded frame to ensure more stable subsequent feature extraction.

[0140] Input: Original image frame sequence (high-resolution pixel matrix).

[0141] Output: A normalized sequence of image frames after scaling, normalization, and enhancement.

[0142] The server performs image scaling (e.g., downsizing from high resolution to medium resolution), color normalization (mapping pixel values ​​to a uniform range), noise reduction, and contrast enhancement. These data processing steps reduce recognition errors caused by lighting differences and noise.

[0143] Step 5: Server detects key points of body parts The server calls a keypoint detection algorithm on each frame of standardized image to locate key points of the human body and hands.

[0144] Input: Normalized image frame sequence.

[0145] Output: A set of keypoint coordinates for each body part in each frame.

[0146] The server applies convolutional neural networks or other keypoint detection models to the image, calculates the coordinates of key points such as the wrist, palm, and finger joints, and normalizes the coordinates, representing them as proportional values ​​relative to the width and height of the image, thus forming "body part position information" of a uniform scale.

[0147] Step 6: Server construction time-series features The server combines the body part position information from multiple consecutive frames in chronological order to calculate the characteristics of the action as it changes over time.

[0148] Input: A sequence of keypoint coordinates for body parts in consecutive frames.

[0149] Output: A sequence of time-series features including displacement, velocity, acceleration, and joint angle changes.

[0150] The server performs differential operations on the coordinates of key points in adjacent frames to obtain displacement vectors; then performs time differential operations on the displacements to obtain velocity and acceleration; at the same time, it calculates geometric features such as joint angles based on the positional relationship between joints, and stacks these values ​​in time order into a high-dimensional feature tensor to describe the complete motion trajectory.

[0151] Step 7: The server generates intermediate identification information. The server inputs temporal features into the semantic recognition neural network, outputs a sequence of semantic labels corresponding to the actions, and organizes them into an intermediate recognition information structure.

[0152] Input: A sequence of time-series features (high-dimensional feature tensors).

[0153] Output: Intermediate identification information including label sequence, confidence level, and time slice information.

[0154] The server uses a pre-trained deep neural network to perform matrix multiplication and nonlinear transformation on the input features, calculates the probability distribution of each semantic tag in each time slice, and then uses a decoding algorithm to convert the continuous probability sequence into a discrete tag sequence, and records the start and end time and confidence level of each tag, thereby forming structured intermediate recognition information.

[0155] Step 8: Server construct prompt statement Based on intermediate identification information, the server converts the label sequence and its meaning into text-based prompts to drive generative artificial intelligence models.

[0156] Input: Intermediate identification information (tag ID sequence, tag meaning, time information).

[0157] Output: A prompt text describing the semantics of the action, along with indication information including parameters.

[0158] The server looks up the mapping table between tags and basic meanings, converts each tag into a corresponding basic semantic word or phrase, and combines them into explanatory text in chronological order. The server adds content such as task description, target language, and style constraints to the prompt statement to form natural language text, and then encapsulates the prompt statement and parameters together into an instruction information data structure.

[0159] Step 9: The server inputs the prompt into the generative artificial intelligence model. The server takes the pre-constructed instructions as input and calls a generative artificial intelligence model to generate text data in natural language form.

[0160] Input: Instructions containing prompts and parameters.

[0161] Output: Preliminary natural language text data.

[0162] The server performs word segmentation and embedding transformation on the prompt statement, converting discrete word units into vector representations. These vectors are then input into the sequence to generate the model. The model performs weighted summation and linear transformation on the input vectors through multi-layer attention calculation and feedforward network, and outputs the probability distribution of each word step by step at the decoding end. At each step, the server selects the word with the highest probability or uses a beam search strategy to generate the complete text sequence.

[0163] Step 10: The server performs language processing to standardize text information. The server performs language processing on the text data output by the generative artificial intelligence model, corrects word order and word usage, and outputs standardized sentences.

[0164] Input: Initially generated natural language text data.

[0165] Output: Standardized text information with complete grammar and consistent expression.

[0166] The server calls language processing algorithms to perform syntactic analysis on the text, detects unnatural word order or grammatically incomplete segments, and replaces or rearranges them according to the probability distribution given by language rules or language models. The server also performs word standardization, unifying synonym variants into standard expressions, thereby obtaining more fluent sentences.

[0167] Step 11: The server inputs text information into the speech synthesis module. The server takes standardized text information as input and calls a speech synthesis algorithm to generate audio data.

[0168] Input: Standardized text information and speech synthesis parameters (speech rate, timbre, volume).

[0169] Output: Audio data in digital format (waveform or compressed audio).

[0170] The server decomposes the text into a sequence of phonemes, uses an acoustic model to calculate the acoustic feature vector (e.g., spectral features) corresponding to each phoneme, and then calls a neural vocoder to convert the feature vector into a time-domain waveform sample. Subsequently, the server uses an audio encoding algorithm to encode the waveform into a compressed audio format to reduce the amount of data.

[0171] Step 12: The server associates and stores audio data and text information. The server associates the text information of the current session with the corresponding audio data and writes it to the recording medium.

[0172] Input: Standardized text information, generated audio data, session identifier, and timestamp.

[0173] Output: Historical information record entries stored in the recording medium.

[0174] The server creates a data record entry, which includes text content, audio file path or identifier, intermediate identification information references for backtracking, and user identification information, and writes the record to a database or file system for subsequent learning processing and performance analysis.

[0175] Step 13: The server sends audio data and text information to the terminal. The server packages the generated audio data and corresponding text information and sends it to the terminal over the network.

[0176] Input: audio data, text information, session identifier.

[0177] Output: The response data packet transmitted to the terminal over the network.

[0178] The server includes text fields, audio fields, and metadata (such as encoding format and duration) in the data packet and sends it to the corresponding terminal's session channel using an appropriate transport protocol.

[0179] Step 14: The terminal receives audio data and decodes and plays it. The terminal receives data packets returned by the server, extracts audio data and text information from them, and decodes and plays them locally.

[0180] Input: A data packet containing audio data and text information sent by the server.

[0181] Output: Physical sound played through the terminal's speakers and text displayed on the screen.

[0182] The terminal calls the audio decoding library to restore the compressed audio data into a pulse code modulation audio stream, and then sends the audio stream to the operating system's audio output interface, so that the sound output device emits the corresponding physical sound; at the same time, the terminal displays text information on the display interface for the user to visually confirm and identify the content.

[0183] Step 15: Users can proofread and correct text information (optional). Users view the text information output by the server on the terminal screen to determine whether it matches the original body posture expression.

[0184] Input: The text information displayed on the terminal.

[0185] Output: The result confirmed by the user, or text input containing corrections.

[0186] When users discover errors, they can directly modify the text content through the text editing interface provided by the terminal, such as replacing incorrect words or adding or deleting phrases. The terminal will then package the text before and after modification, along with related conversation information, into correction data.

[0187] Step 16: The terminal sends the correction data to the server. After the user completes the correction, the terminal sends the original text information, the corrected text information, and the session identifier to the server.

[0188] Input: User correction data (original text, corrected text, timestamp, session ID).

[0189] Output: Corrected data packets transmitted to the server over the network.

[0190] Before sending, the terminal formats the corrected data to ensure that the server can parse the field content and match it with the corresponding historical records.

[0191] Step 17: The server performs learning processing (offline or periodic) based on historical information. During the non-real-time phase, the server reads historical records containing intermediate recognition information, text information, and user correction information from the recording medium and performs learning processing to update the model parameters.

[0192] Input: Historical information record entries (intermediate identification information, original text, corrected text, etc.).

[0193] Output: Updated image processing algorithm parameters and generative artificial intelligence model weights.

[0194] The server constructs training samples, using intermediate recognition information as input and corrected text as the target output. It calculates the error between the model's output and the target using a loss function and updates the weight parameters of the generative AI model through backpropagation. Simultaneously, the server adjusts the threshold or classifier parameters of the image processing algorithm based on the error distribution, enabling more accurate intermediate recognition information to be generated under similar input conditions in the future. By repeatedly executing this learning process, the server continuously improves the overall recognition accuracy and generation quality of the system.

[0195] Application Example 1 The process flow corresponding to the specific processing in Use Case 1 will be described below. The various parts of the system described below are implemented by the data processing device 12 and the intelligent device 14. Furthermore, the data processing device 12 is referred to as the "server" and the intelligent device 14 is referred to as the "terminal".

[0196] In existing technologies, most solutions for automatic sign language recognition and speech output for deaf and mute individuals rely solely on a single sign language recognition model, classifying or performing simple pattern matching on a frame-by-frame basis in video captured by a camera. These solutions suffer from the following technical shortcomings: First, due to the complexity of sign language in spatial posture and continuous temporal movements, a single model has limited recognition accuracy in complex lighting, occlusion, and multi-user scenarios, easily leading to misidentification or missed recognition, resulting in insufficient overall system reliability. Second, existing systems typically directly output the recognition results as the final text, lacking adaptive verification mechanisms for uncertain results and multi-model collaborative decision-making mechanisms. They cannot automatically adjust the recognition path when model confidence is low, thus failing to meet the stability and robustness requirements of real-world service scenarios. Third, the use of generative AI models in existing technologies is mostly focused on natural language generation, lacking the integration of temporal gesture features and dedicated prompts to address the internal nuances of sign language. The existing systems lack a systematic design for optimizing recognition and speech expression. Most prompts are fixed templates and cannot be dynamically adjusted according to application scenarios, environmental conditions, and historical data, thus failing to fully realize the potential of generative AI models. Furthermore, existing systems generally treat speech synthesis as simple post-processing, lacking dedicated steps for structured rewriting and conversationalization of the synthesized input text. This makes it difficult to optimize the naturalness and comprehensibility of the speech output for specific service scenarios. Moreover, existing technologies rarely integrate and record sign language video data, temporal features, generative AI model output, and traditional recognition model output, making it difficult to form high-quality, closed-loop, iterative training data, thereby limiting the model's self-evolution capabilities in long-term operation.

[0197] Therefore, how to propose a new data processing architecture at the computer system level that can start from the original sign language video and work together through multiple stages such as hand region extraction, temporal feature calculation, multi-model collaborative recognition, dynamic generation of prompts for generative artificial intelligence models, text-to-speech expression optimization, and automatic accumulation of training data, in order to improve the accuracy, robustness, and scalability of the entire process of sign language recognition and speech output, has become an urgent technical issue to be solved in this field.

[0198] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Application Example 1 is achieved by the following means.

[0199] In this invention, the server includes: a device for acquiring time-series image information of sign language movements from an imaging device and preprocessing it to extract hand and upper body regions; a device for calculating motion feature quantities including joint and contour position information based on the extracted regions using a machine learning model and generating time-series data; a device for generating prompt statements for a generative artificial intelligence model and inputting the prompt statements and the time-series data into the generative artificial intelligence model to obtain text information corresponding to the sign language content; a device for comparing the text information with the recognition results output by an independently constructed sign language recognition model and determining the final text information to be adopted based on consistency and credibility; a device for generating prompt statements for controlling speech synthesis processing based on the final adopted text information and generating speech information using speech synthesis technology and outputting it through a speech output device; and a device for associating and storing the time-series data, the output results of the generative artificial intelligence model, and the output results of the sign language recognition model and generating training data for updating the model's learning data. This allows for the formation of a complete technical chain within the computer, encompassing multimodal temporal feature extraction, prompt-driven generative AI recognition, multi-model collaborative decision-making, speech expression optimization, and data closed-loop training. Without relying on a single recognition model, this significantly improves the overall recognition accuracy, robustness, and self-learning ability of the sign language to speech conversion process, thereby achieving substantial improvements in computer sign language processing technology and human-computer interaction performance.

[0200] A "system" refers to a collection of information processing devices consisting of multiple interconnected information processing components, used to collect, analyze, transform, and output input data to achieve specific functions.

[0201] A "server" refers to a computing device in a system that undertakes the main data processing, model inference, and control flow scheduling. It can be a physical computer, a virtual machine, or a cloud computing node.

[0202] "Imaging device" refers to a hardware device used to acquire image information of a target scene through optical means, including cameras, video recording equipment or other sensing devices that can output digital image signals.

[0203] "Time-series image information" refers to a collection of multiple frames of image data acquired continuously within a predetermined time interval, used to represent scenes or actions that change over time.

[0204] "Preprocessing" refers to data processing operations performed on the original image information before subsequent recognition or analysis, such as scaling, noise reduction, color space conversion, and region cropping, in order to improve processing efficiency and recognition accuracy.

[0205] "Hands and upper body area" refers to the image area that is identified and cropped from the whole image, containing the user's hands and upper body, and is used to focus on the main parts of the sign language movement.

[0206] "Machine learning model" refers to a mathematical model that learns parameters through training data, including neural networks, support vector machines, decision trees, etc., which are used to automatically extract features from input data and output recognition results or prediction results.

[0207] "Key points" are points in the human body or hand structure used to represent key anatomical locations, such as wrist and finger joints. Their coordinates in an image can be used to describe posture.

[0208] "Contour position information" refers to spatial position information used to represent the shape of the hand or upper body boundary, including edge lines, circumscribed rectangles, or other parameters that describe the shape of the target.

[0209] "Motion feature quantity" refers to the amount of data used to characterize the spatial position and temporal changes of sign language movements, including the coordinates, trajectory, speed, direction and other features of joints as they change over time.

[0210] "Time-series data" refers to a multidimensional data sequence arranged in chronological order, used to depict dynamic processes over time. In this invention, it is mainly used to represent the continuous changes in sign language movements.

[0211] "Generative artificial intelligence models" refer to artificial intelligence models built based on technologies such as deep learning that can automatically generate text, speech, or other content based on input data and prompts.

[0212] "Prompt statements" refer to textual information used to explain the task type, input meaning, output format, or generation requirements to generative artificial intelligence models. They are also known as prompt words, prompt text, or prompt instructions.

[0213] “Textual information” refers to language content represented in the form of symbols or characters, including words, phrases, sentences, etc., used to express the semantics corresponding to sign language.

[0214] A "sign language recognition model" is a recognition model specifically trained on sign language videos or gesture features, used to map sign language actions into text information or symbol sequences.

[0215] "Consistency" refers to the degree to which two or more identification results are identical or highly similar in content, which can be measured by similarity or matching indexes.

[0216] "Confidence" refers to the probability or degree of confidence that the identification result is considered correct, and is usually represented by the confidence score or probability output by the model.

[0217] "Finally adopted text information" refers to the text information that is determined and output based on rules such as consistency and credibility after comparing multiple candidate recognition results.

[0218] "Speech synthesis technology" refers to the technology of converting input text information into audible speech signals, including text analysis, acoustic modeling, and vocoder generation.

[0219] “Voice information” refers to voice signal data represented in digital or analog form, intended for playback through a speaker or other audio output device.

[0220] "Voice output device" refers to a hardware device that can convert speech information into sound that can be heard by the human ear, including speakers, headphones or other audio playback devices.

[0221] "Associative storage" refers to a storage method that records multiple related data items together in a searchable and traceable manner on a storage medium for subsequent joint querying or analysis.

[0222] "Training data" refers to the set of labeled data used to train or retrain machine learning models or generative artificial intelligence models, which typically includes input data and its corresponding target output.

[0223] "Learning data" refers to the set of data used or generated by a machine learning model or generative artificial intelligence model during training to update parameters, including initial training data and subsequent appended data.

[0224] In the following embodiments, the server, terminal, and user are described as the main entities in the specific implementation of the present invention. The present invention is not limited to the hardware or software examples described below; these examples are merely illustrative of a feasible technical implementation.

[0225] I. Overall System Composition As the core information processing device, the server can be a rack-mounted physical server, a cloud computing node, or an edge computing node. The server preferably runs a general-purpose operating system, such as a Linux-based distribution (e.g., Ubuntu Server, Debian), and has installed deep learning frameworks (e.g., TensorFlow, PyTorch), image processing libraries (e.g., OpenCV), network communication libraries (e.g., gRPC, HTTP client libraries), and a speech synthesis SDK.

[0226] As an interactive device close to the user, the terminal can be a smartphone, tablet computing device, or embedded terminal, running a mobile operating system (such as Android or iOS) or embedded Linux, and having applications installed to receive server data, play audio, and optionally display text.

[0227] Users, acting as both sign language providers and speech receivers, can be deaf users or service personnel. Through sign language gestures and listening to speech, users indirectly drive the processing flow of the server and terminal.

[0228] The system may further include imaging devices (such as high-definition network cameras, USB cameras), audio output devices (such as speakers, headphones), and optional display devices (such as LCD screens).

[0229] II. Image Acquisition and Preprocessing Methods on the Server Side In this implementation, the server acquires sign language video through an imaging device. The server uses the OpenCV library to drive the imaging device's driver interface (such as a USB video interface or an RTSP network stream interface), reads raw image frames at a fixed frame rate (e.g., 25 to 30 frames per second), and caches them in memory using formats such as BGR or YUV.

[0230] The server performs preprocessing operations on the images. Using OpenCV, the server performs the following data processing operations: image scaling (e.g., scaling from 1920×1080 to 640×480 to reduce the input dimension of subsequent neural networks), color space conversion (e.g., BGR to RGB to match the training configuration of deep learning models), Gaussian or median filtering to reduce sensor noise, and histogram equalization to improve visibility in low-light conditions. This preprocessing is performed on the CPU, and vectorized instructions can further improve processing speed, thereby reducing latency under the same hardware conditions.

[0231] After preprocessing, the server adds a timestamp and session identifier to each frame, organizing them into a time-series data structure within a circular buffer. This allows subsequent modules to access frame data under strict temporal relationships. Through this data structure design, the invention achieves constant-level space overhead in memory usage, avoiding memory bloat caused by long video streams and thus improving system stability.

[0232] III. Extraction patterns of the hand and upper body areas on the server side In its implementation, the server employs human detection and keypoint detection models to extract regions from preprocessed image frames. The server can use convolutional neural network-based object detection models, such as those similar to the YOLO series or single-stage detection networks, to output bounding boxes containing the upper body of the human. Within the upper body region, the server then calls hand detection or keypoint models (such as network structures like MediaPipe Hands or OpenPose) to calculate the coordinates of the keypoints on both hands.

[0233] During forward inference on each frame, the server normalizes the image tensor (subtracting the mean and dividing the variance) and performs multi-layer convolution, batch normalization, non-linear activations (such as ReLU), pooling, and fully connected layer operations on the GPU. At the network's end, the server outputs the 2D or 3D coordinates of approximately 21 joints for each hand. These operations are accelerated using a tensor computation library, significantly outperforming traditional rule-based edge detection and template matching methods, thereby improving the accuracy of hand localization in complex backgrounds.

[0234] The server eliminates short-term jitter and noise by interpolating and smoothing the keypoint coordinates along the time dimension, forming a temporal data structure that includes position, velocity, and acceleration features. The server further organizes the keypoint sequence into a matrix form, for example, [T, K×D], where T is the number of time steps, K is the number of keypoints, and D is the spatial dimension, thus providing a unified tensor input for subsequent temporal neural network processing. This structured feature representation reduces redundant pixel information, significantly lowers the data dimensionality, and consequently reduces the computational load of network inference, leading to improved processing speed and reduced communication overhead.

[0235] IV. Temporal Feature Modeling and Sign Language Recognition Patterns on the Server Side In its implementation, the server utilizes two types of models for sign language recognition: a dedicated sign language recognition model and a generative artificial intelligence model. Through this multi-model collaborative structure, the server constructs an unconventional decision-making process within the computer, avoiding reliance on a single path.

[0236] 1. Dedicated sign language recognition model The server can deploy sequence-modeling-based neural network structures, such as temporal recognition models composed of multi-layer bidirectional long short-term memory networks (Bi-LSTM) or Transformer networks based on self-attention mechanisms. The server takes the aforementioned temporal feature matrix as input and calculates the hidden states or attention weights step by step over time.

[0237] During the training phase, the server uses cross-entropy loss or connection-temporal classification loss (CTC) as the error function, and updates the network weights through backpropagation and gradient descent (such as the Adam optimizer). The server can use data augmentation strategies (such as adding small random noise to the key coordinates, random stretching and compression on the time axis, and horizontal mirroring transformation) to expand the training data, thereby improving the robustness of the model to different users and different sign language speeds.

[0238] During the inference phase, the server outputs the probability distribution of the corresponding text sequence and uses a decoding algorithm (such as beam search) to obtain the optimal text sequence and its confidence score. This model focuses on structured temporal features, has fast inference speed, and relatively controllable resource consumption.

[0239] 2. Generative Artificial Intelligence Models The server simultaneously invokes a generative artificial intelligence model, which can be a multimodal Transformer model with an encoder-decoder structure. The server encodes the temporal feature matrix into a high-dimensional representation vector and drives the decoder to generate the corresponding natural language output through prompts.

[0240] Before sending a request, the server constructs a prompt statement to clarify the task definition and output format. For example: "Please identify the user's intent based on the input sign language video frame sequence and output the corresponding simplified Chinese text." "The input is a time series of extracted hand key points. Please infer the most likely Chinese sentence and return the result with the highest confidence." The server inputs the prompts and feature data into the generative AI model. Internally, the model uses multi-layered self-attention and cross-attention mechanisms to interactively compute the prompts and feature representations, thereby understanding the sign language intentions at the semantic level. Compared to traditional models, this structure can combine richer prior knowledge and contextual semantics, thus exhibiting higher fault tolerance for complex sentences and ambiguous actions.

[0241] V. Server-Side Multi-Model Collaboration and Outcome Decision-Making Patterns In its implementation, the server does not simply adopt the output of a single model, but rather compares and fuses the outputs of a dedicated sign language recognition model and a generative artificial intelligence model. The server can employ the following decision-making strategies: The server obtains the first text result and its confidence score C1 from a dedicated sign language recognition model, and the second text result and its confidence score C2 from a generative artificial intelligence model. The server calculates the content similarity S between the two using text similarity algorithms (such as edit distance and cosine similarity of embedding vectors).

[0242] The server makes the determination based on preset rules: When S is greater than the first threshold and both C1 and C2 are higher than the second threshold, the server considers the two results to be consistent and selects the result with higher confidence as the final text information to be used.

[0243] – When S is low and one of the confidence levels is much higher than the other, the server selects the result with the higher confidence level.

[0244] When both confidence levels are low, the server can call the generative AI model again with updated prompts, such as: "Please provide two most likely Chinese sentences, label their respective confidence levels, and explain the key actions that are highly correlated with the input features." This will help obtain more reliable candidates.

[0245] By combining the structured advantages of traditional time-series models with the semantic understanding capabilities of generative artificial intelligence models through this unconventional multi-model collaboration approach, the server achieves improved recognition accuracy and the ability to automatically correct for uncertain results, thus technically surpassing the human judgment process that relies solely on visual intuition.

[0246] VI. Dynamic Generation and Optimization of Server-Side Prompt Statements In its implementation, the server does not use fixed prompts. Instead, it dynamically generates prompts based on sign language usage, environmental conditions, and historical recognition records. The server maintains a prompt template library, with each template containing placeholders such as "domain," "user type," and "complexity." The server fills in these placeholders according to the current scenario.

[0247] For example, the server can select a more suitable prompt based on parameters such as the ambient brightness of the camera, noise level, and the complexity of phrases used by the user in the past. For low-light or heavily obstructed scenes, the server can generate the following prompt: "The input data may contain partial occlusion or insufficient lighting. Please increase the tolerance for incomplete actions when generating results and return the two explanations with the highest confidence." Through this dynamic prompting mechanism, the server essentially performs "online configuration" of the generative artificial intelligence model, enabling the same model parameters to exhibit different reasoning preferences in different scenarios, thereby improving the overall recognition stability.

[0248] VII. Server-side text polishing and speech synthesis control methods In its implementation, the server calls a generative artificial intelligence model to rewrite the text information to suit speech playback before speech synthesis. The server can use prompt statements: Please rewrite the following text into a shorter, more conversational sentence suitable for voice broadcast, while maintaining the original meaning: 'Please tell me the information about this product.' "Please simplify the identified text to no more than 20 characters without changing the meaning, and avoid obscure words." After receiving the rewritten text, the server uses it as input for speech synthesis, thus avoiding speech comprehension difficulties caused by long sentences or formal written language. This step is not simply about "beautifying the language," but rather about reducing problems such as blurred speech boundaries and unnatural sentence breaks by controlling the structure of the input text, thereby lowering the probability of speech synthesis errors and auditory misunderstandings.

[0249] During the speech synthesis stage, the server can call general speech synthesis services, such as an end-to-end TTS model based on deep neural networks. The server specifies parameters such as speaker type, speech rate, and volume in the request. Internally, the speech synthesis model generates waveforms through a text front-end (word segmentation, phoneme annotation, prosody prediction) and a back-end acoustic modeling and vocoder (such as one based on convolutional or autoregressive structures). The server caches the returned audio data as MP3 or WAV files, or transmits it directly to the terminal via a streaming interface.

[0250] VIII. Server-side data recording and training data generation format In the implementation, the server associates and stores key data for each interaction, including: original image fragment identifiers, temporal feature data, output of a dedicated sign language recognition model, output of a generative artificial intelligence model, final text information, and subsequent user feedback information (such as whether it has been manually corrected).

[0251] The server stores this data as structured records, such as using a relational database or key-value store, and associates multiple results from the same session with a unique session ID. The server can periodically filter this data, selecting samples with low confidence and large discrepancies to construct a reinforcement training set for retraining a dedicated sign language recognition model. It can also be used to fine-tune the weights of generative artificial intelligence models or update their external knowledge base.

[0252] Through this data closed-loop mechanism, the server continuously optimizes model parameters within the computer, enabling the system to gradually improve its recognition accuracy and robustness over long-term operation. This technological effect is not simply due to the accumulation of human experience, but rather relies on sophisticated data structure design, sample selection strategies, and model training algorithms.

[0253] IX. Terminal-side reception, decoding, and output formats In this implementation, the terminal connects to the server via a network and uses WebSocket or HTTP long polling to receive audio data and optional text information. After receiving the data, the terminal performs integrity checks on the data packets and writes the audio data to its local cache.

[0254] The terminal decodes audio using a multimedia framework provided by the operating system (such as MediaPlayer or ExoPlayer for Android, and AVAudioPlayer for iOS). The decoding process is executed on the terminal's CPU or dedicated audio decoding hardware, including frame parsing, entropy decoding, dequantization, and PCM sample reconstruction. The terminal inputs the PCM data into the audio driver, which then drives the speaker to output sound via a digital-to-analog converter.

[0255] The terminal can also simultaneously display the text results sent by the server on the screen, allowing users to obtain information visually even in noisy environments. This processing by the terminal constitutes direct control over real-world physical devices (speakers, displays), tightly coupling the aforementioned abstract data processing with concrete physical output.

[0256] 10. User-side usage patterns and technical effects In this implementation, users simply need to use sign language naturally in front of the imaging device; they don't need to understand the internal details of the system. The server automatically converts complex sign language gestures into text suitable for voice broadcast through the aforementioned multi-stage feature extraction, multi-model collaboration, and prompt-driven generative artificial intelligence model, which is then read aloud by the terminal.

[0257] By combining keypoint temporal features and a multi-layer neural network structure, this invention significantly improves recognition accuracy compared to using only a single image classification model. Through dynamic prompts and multi-model collaborative decision-making, this invention maintains high stability even under adverse conditions such as lighting changes, occlusion, and individual user differences. Through data dimensionality compression and streaming processing, this invention achieves lower computational latency and communication load. By recording and reusing multi-source recognition results, this invention constructs an iterative training data generation process, enabling the system to automatically optimize over time, exceeding the capabilities of traditional "manual annotation + static model".

[0258] Therefore, this invention achieves substantial improvements in the computer technology itself in terms of temporal multimodal recognition, inference efficiency, and system robustness by implementing specific data structure design, deep neural network architecture, multi-model decision rules, and dynamic generation mechanism of prompt statements within the server, and completing voice and display output on the terminal and physical output device, rather than simply automating human sign language translation.

[0259] use Figure 12 The processing flow is explained.

[0260] Step 1: Users use sign language naturally in front of the imaging device.

[0261] Users express their intentions through continuous movements of their hands, fingers, and upper body, such as "Please tell me about this product."

[0262] Input: Gestures and facial expressions from the real world.

[0263] Output: Continuous sign language gestures presented in the field of view of the imaging device.

[0264] Step 2: The server acquires sign language video frames from the imaging device and performs preprocessing.

[0265] The server reads raw image frames at a fixed frame rate through the camera driver interface, saves each frame as a pixel matrix in BGR or YUV format, and performs size scaling, color space conversion, and noise reduction processing.

[0266] Input: The raw video signal output by the imaging device.

[0267] Output: Preprocessed time-series image data consisting of multiple frames.

[0268] The server performs the following actions: for each frame, it calls an image scaling function to compress the resolution from high resolution to medium resolution, calls a filtering function to remove sensor noise, and adds a timestamp and session identifier to each frame and stores them in a circular buffer.

[0269] Step 3: The server extracts the hand and upper body regions from the preprocessed image.

[0270] The server uses an object detection model to perform forward inference on each frame of the image, calculates the bounding box of the potential upper body region, detects the hand position within the region, and crops out a sub-image containing both hands and the upper body.

[0271] Input: The preprocessed frame sequence from step 2.

[0272] Output: A sequence of local images of the hand and upper body at each time step.

[0273] The server performs the following steps: it constructs a tensor input for each frame, calls a deep neural network to perform convolution, pooling, and fully connected operations to obtain the coordinates of candidate regions, then calls a cropping function to crop the original image into smaller patches, and arranges these patches in chronological order.

[0274] Step 4: The server calculates motion features of joints and contours on local images.

[0275] The server calls the keypoint detection network to calculate the coordinates and contour information of the hand joints in each cropped image, and then combines the coordinates of consecutive frames into a temporal feature matrix.

[0276] Input: The sequence of partial images of the hands and upper body from step 3.

[0277] Output: Time series data containing the coordinates and motion features of each joint at each time step.

[0278] The server performs the following steps: running multiple convolutional and fully connected layers of the keypoint network on each frame to output the two-dimensional coordinates of each joint; performing difference operations on the coordinates of consecutive frames to obtain the velocity vector; interpolating and smoothing the time axis to form a matrix of shape [time steps, feature dimension], and storing it in the feature buffer.

[0279] Step 5: The server inputs the temporal features into a dedicated sign language recognition model to generate the first text result.

[0280] The server calls a temporal neural network (such as Bi-LSTM or Transformer), takes the feature matrix as input, updates the hidden state or attention weights step by step, and outputs the corresponding text sequence and confidence score.

[0281] Input: The time series feature matrix generated in step 4.

[0282] Output: The first candidate text information and its confidence value.

[0283] The server performs the following actions: normalizes the feature matrix, inputs it into each layer of the network, performs matrix multiplication and nonlinear activation; at the output, it uses a decoding algorithm (such as beam search) to generate the most probable character sequence and calculates the probability of each sequence as the confidence level.

[0284] Step 6: The server generates prompts and organizes multimodal inputs for generative artificial intelligence models.

[0285] Based on the current session type and historical recognition data, the server selects or constructs appropriate prompt statements from the prompt statement template library, such as "The input is a time series of extracted hand key points. Please infer the most likely Chinese sentence and return the result with the highest confidence." Input: temporal feature matrix, context information of the current session, and historical identification records.

[0286] Output: A request data structure containing prompt statements and feature data.

[0287] The server performs the following actions: determining scene parameters (such as lighting and occlusion), filling in variable fields in the template, generating complete text prompts, and packaging them with the feature matrix into a unified input structure, ready to be sent to the generative artificial intelligence model.

[0288] Step 7: The server invokes a generative artificial intelligence model to generate a second textual result.

[0289] The server sends the prompts and timing features to the generative artificial intelligence model service via a network interface, and receives the text output and confidence level returned by the model.

[0290] Input: The prompt statement from step 6 and the timing feature data.

[0291] Output: The second candidate text information and its confidence level.

[0292] The server performs the following actions: constructs a request message, takes the prompt statement as text input, encodes the feature matrix into a numerical format acceptable to the model, and sends it to the remote inference service; it parses the text fields and confidence fields in the response and caches the result for subsequent decision-making.

[0293] Step 8: The server compares the two types of recognition results and determines the final text information.

[0294] The server calculates the text similarity between the first and second text results, and combines their respective confidence levels to determine the final text content to be adopted based on preset decision rules.

[0295] Input: The first text result and confidence level of step 5, and the second text result and confidence level of step 7.

[0296] Output: The final text information used.

[0297] The server performs the following actions: performs edit distance or cosine similarity calculation on the two texts to obtain similarity values; compares the similarity and confidence scores with a threshold, selects the optimal result or triggers a re-call of the generative artificial intelligence model branch; and stores the selected text in the session state.

[0298] Step 9: The server uses a generative artificial intelligence model to process the text into conversational language and optimize its broadcasting.

[0299] The server generates a new prompt based on the final text content, such as "Please rewrite the following text into a shorter, more conversational sentence suitable for voice broadcast, while keeping the original meaning unchanged: 'Please tell me the information about this product.'", and sends it along with the original text to the generative artificial intelligence model.

[0300] Input: The final text information from step 8.

[0301] Output: Rewritten, optimized text information suitable for voice playback.

[0302] The server performs the following actions: constructs a prompt statement containing the original text, inputs the combined text as a single string into the generative artificial intelligence model, receives the rewritten text output, and performs length and content checks to ensure that it does not contain any inappropriate content.

[0303] Step 10: The server will optimize the text input speech synthesis module to generate speech data.

[0304] The server calls the speech synthesis service, taking the optimized text and speech parameters (language, speaker, speech rate, etc.) as input to obtain the audio data stream.

[0305] Input: Optimized text information from step 9.

[0306] Output: Compressed audio data (such as MP3 or WAV).

[0307] The server performs the following actions: it sends text parameters via a network request, the speech synthesis system performs text analysis, acoustic modeling, and vocoder operations internally, and the server writes the audio byte stream to a cache or temporary file at the receiving end, and records the audio length and encoding format.

[0308] Step 11: The server sends the voice data and optional text results to the terminal.

[0309] The server uses an established communication channel (such as WebSocket) to push audio data and corresponding text information to the terminal.

[0310] Input: The voice data from step 10 and the text information obtained from step 8 / step 9.

[0311] Output: Audio and text data packets sent to the terminal.

[0312] The server performs the following steps: it packages the audio data into segments of a certain size, adds a session ID and sequence number; it encapsulates the text information into text messages; it sends these messages sequentially through the network interface and waits for the terminal to confirm receipt.

[0313] Step 12: The terminal receives and caches voice data and text information.

[0314] The terminal reads data packets sent by the server from the communication channel, reassembles the audio packets in sequence, writes them to local storage, and displays or caches the text information at the same time.

[0315] Input: The audio data packet and text message sent in step 11.

[0316] Output: The complete audio file stored locally, along with text content that can be displayed.

[0317] The terminal performs the following actions: verifying the integrity of the data packets (such as length verification), creating temporary files in the file system, writing the audio byte stream sequentially, storing the text content in a memory data structure, and reserving a display area on the interface.

[0318] Step 13: The terminal decodes the voice data and plays it through the audio output device.

[0319] The terminal calls the system multimedia API to open the cached audio file, performs decoding, and sends the decoded PCM data to the audio driver to drive the speaker or headphones to output sound.

[0320] Input: The audio file cached in step 12.

[0321] Output: A voice signal that the user can hear.

[0322] The terminal executes the following steps: creating a media player instance, setting the data source to a local file path, and starting the decoding and playback process; adjusting the volume according to system settings during playback, outputting continuous sound through audio hardware, and simultaneously displaying corresponding text on the screen.

[0323] Step 14: Users (sales staff or hearing people) listen to the audio and respond according to the content.

[0324] Users hear voice messages played through the terminal's speaker, such as "Please tell me the information about this product," thus understanding the intentions of deaf users and responding with verbal explanations, instructions, or by operating the terminal.

[0325] Input: The audio signal output from step 13.

[0326] Output: The user's actual service behavior or verbal response to the deaf user.

[0327] User's specific actions: Choose the response method based on the meaning of the voice, such as walking to the shelf and explaining the features of the product, or entering text information on the terminal for subsequent system processing.

[0328] Step 15: The server records relevant data for this session to generate training data.

[0329] After the session ends, the server associates and stores the temporal features, the outputs of the two types of models, the final text information, and optional user feedback information into the database for subsequent model training and performance evaluation.

[0330] Inputs: the time series features from step 4, the model outputs from steps 5 and 7, and the final text and log information from step 8.

[0331] Output: Structured storage of training data records.

[0332] The server performs the following actions: It constructs a record structure containing session ID, timestamp, feature summary, model output, and decision results, and writes it to a database or object storage for batch reading during offline training to update the parameters of the generative artificial intelligence model and sign language recognition model.

[0333] Alternatively, an emotion engine for inferring user emotions can be combined. That is, the specific processing unit 290 can also use the emotion-specific model 59 to infer user emotions and perform specific processing using user emotions.

[0334] Example 2 The flow of a specific process in Example 2 will be described. Each part of the system described below is implemented by the data processing device 12 and the smart device 14. The data processing device 12 will be referred to as the "server," and the smart device 14 as the "terminal."

[0335] In existing technologies, for communication systems that use body movements (such as sign language) as input, most solutions simply execute a fixed recognition model locally or on the server side, directly mapping video data to fixed text output, and then simply calling a speech synthesis module to generate speech. These solutions typically suffer from the following technical problems: (1) In scenarios with complex backgrounds, lighting changes, or strong continuity of actions, fixed recognition models have difficulty making full use of time series information and multi-frame motion features, resulting in insufficient accuracy of body action semantic recognition. In particular, they are prone to misidentification or omission in distinguishing continuous sentences and subtle actions, thereby reducing the overall availability of the system.

[0336] (2) Most existing systems are “one-way pipeline” structures, meaning that the coupling between image processing, recognition, text generation and speech synthesis is low. It is difficult to dynamically adjust the text output based on the dialogue context, speaker attributes and scene information, and it is impossible to flexibly customize the output style and politeness according to different users and dialogue environments.

[0337] (3) Traditional speech synthesis output is often based only on the recognized single sentence text, without considering the historical dialogue content and user attributes. It does not support flexible control of the natural language generation process through prompts, resulting in stiff output content and poor contextual coherence, making it difficult to meet the needs of natural dialogue and multi-turn communication in actual human-computer interaction scenarios.

[0338] (4) Existing systems typically do not deeply embed generative artificial intelligence models into the complete processing chain from body motion recognition to speech output. The construction of prompt information is also relatively crude, and the prompt content cannot be automatically adjusted according to dynamic elements such as motion characteristics and dialogue history, thus limiting the scope of generative artificial intelligence models in such systems.

[0339] (5) From the perspective of computer technology, the existing system is not well designed for collaborative processing of image information, time series features and natural language generation on the server side. It lacks a technical solution that can organically integrate image processing, feature extraction, generative artificial intelligence model reasoning and speech synthesis under a unified architecture, resulting in low server resource utilization efficiency, large system response delay, and difficulty in scaling up to real-time services in multiple devices and scenarios.

[0340] Therefore, it is necessary to provide a new system and its server-side processing method, which can achieve high-precision recognition of body action semantics and context-aware natural language generation by preprocessing image information in a time sequence, constructing prompt information driven by motion features, and introducing generative artificial intelligence models. This will improve the overall recognition accuracy, expression flexibility and system processing efficiency at the computer technology level.

[0341] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Embodiment 2 is achieved by the following means.

[0342] In this invention, the server includes an image preprocessing component for performing temporal segmentation and encoding of image information acquired by a camera device; a feature processing component for extracting motion features of various body parts from the preprocessed image information and generating prompt statements containing instructional content based on the motion features; a generative model invocation component for inputting the prompt statements into a generative artificial intelligence model to generate intermediate text information representing the semantic content of body movements and further generating additional prompt statements for natural language text generation based on the intermediate text information and dialogue history; a text generation component for converting the intermediate text information into natural language text information according to the additional prompt statements; and a voice output component for converting the text information into speech information through speech synthesis processing and sending the speech information to an external device. This enables the formation of a unified processing chain on the server side, encompassing multi-frame image time-series feature extraction, dynamic construction of prompt statements, generative artificial intelligence model inference, natural language text generation, and speech synthesis. This not only improves the accuracy and robustness of body motion semantic recognition but also allows for dynamic adjustment of prompt statement content based on motion characteristics, speaker attributes, and dialogue scenarios. Consequently, it generates natural language output with style and honorifics that match the context. At the computer technology level, this achieves overall optimization of the image processing flow, natural language generation flow, and speech synthesis flow, thereby enhancing system response performance and resource utilization efficiency.

[0343] "Camera equipment" refers to the general term for imaging hardware and its control components used to acquire image or video information, including body movements. It may include cameras, image sensors, lenses, and the driving circuits that work with them.

[0344] "Image information" refers to still images or continuous image frame data acquired by a camera device and represented in digital form. It is usually stored in the form of a pixel matrix and used for subsequent image processing and feature extraction.

[0345] "Video format" refers to the data representation obtained by encoding and encapsulating image frames that change over time, including but not limited to multimedia data streams and their container formats that use compression encoding.

[0346] "Preprocessing" refers to the set of processing operations performed on the original image information before feature extraction and model inference, such as format conversion, size adjustment, noise reduction, normalization, time segmentation, and encoding.

[0347] "Temporal segmentation" refers to the process of dividing consecutive image frames or video data arranged in chronological order into one or more time segments or time windows for independent or batch feature analysis and recognition processing.

[0348] "Feature information" refers to a set of numerical data extracted from image information and used to characterize the posture, position, motion trajectory and other attributes of various parts of the body. It can include key point coordinates, motion vectors, shape descriptors, etc.

[0349] "Motion feature information" refers to the feature information extracted based on the temporal changes between multiple frames of images, used to describe the continuous changes of body movements in the time dimension, including speed, acceleration, direction of motion, and trajectory pattern.

[0350] "Body movement" refers to a sequence of actions consisting of changes in posture and movements of one or more parts of the human body (such as hands, arms, head, upper body, etc.) over a certain period of time, used to express semantic content.

[0351] "Intermediate text information" refers to symbolic text data used to represent the semantic content of body movements. It is usually a basic text representation that has not yet been stylized or expanded with context, and can be used as input for subsequent natural language generation.

[0352] “Text information” refers to natural language text data that conforms to predetermined language rules and stylistic requirements, obtained by natural language generation or polishing based on intermediate text information, and is used for final output or speech synthesis.

[0353] "Dialogue history" refers to an ordered collection of text information, voice content, or related semantic information that has been previously generated or received in the same interactive session, which is used to provide contextual reference for current natural language generation and prompt statement construction.

[0354] "Prompt statements" refer to input information constructed to guide generative artificial intelligence models to produce the expected output. These include instructions, intermediate text information, dialogue history, and scene information, and are used to control the model's generative behavior.

[0355] "Generative AI models" refer to AI models that can automatically generate text, audio, or other data based on input prompts. They typically employ deep learning structures and acquire generative capabilities through training on large-scale data.

[0356] "Generative model invocation component" refers to the software and / or hardware functional module in the server that is responsible for constructing prompt statements, sending requests to generative artificial intelligence models, and receiving model output results.

[0357] "Natural language text" refers to a sequence of strings written in natural language that is intended for human reading and understanding. It has grammatical structure and semantic coherence and may include sentences, paragraphs, or dialogue content.

[0358] "Image preprocessing component" refers to a functional module in the server used to perform preprocessing operations such as format conversion, size adjustment, time segmentation, and encoding on the image information acquired by the camera device.

[0359] The “feature processing component” refers to a functional module in the server that is used to extract motion-related feature information of various parts of the body from preprocessed image information and generate prompts or intermediate text information based on the feature information.

[0360] The “text generation component” refers to a functional module in a server that generates natural language text information based on intermediate text information and dialogue history, using generative artificial intelligence models or other language generation algorithms.

[0361] "Voice output component" refers to a functional module in the server that converts text information into voice information and sends it to an external device via communication, so that the external device's voice output device can play it.

[0362] "Speech information" refers to speech signals represented in the form of digital audio data, which are obtained by speech synthesis processing of text information. It may include waveform data, encoded audio streams, etc.

[0363] "Speech synthesis processing" refers to the digital signal processing process that automatically generates corresponding speech signals based on input text information, usually implemented by speech synthesis models or algorithms.

[0364] "External device" refers to various terminal devices that are connected to the server via a communication network and are capable of receiving and playing voice information and / or displaying text information, including but not limited to mobile terminals, computing devices and dedicated playback devices.

[0365] "Display device" refers to output hardware that can present text information, intermediate text information or other visual content to users in the form of graphics or text, including displays, projection devices, etc.

[0366] The embodiments of this invention will specifically describe the system structure, data flow, algorithm composition, and technical effects by combining the collaborative work of the server, terminal, and user in a real-world operating environment. The following embodiments are merely illustrative; those skilled in the art can make various modifications and substitutions without departing from the spirit of this invention.

[0367] I. Overall System Composition The server, terminal, and user each play different functional roles in the system of this invention: 1. Terminal Composition The terminal includes: - Camera equipment: such as image sensors integrated into mobile terminals, tablet computers, and portable computing devices, or external cameras connected via wired or wireless interfaces. Camera equipment can be a standard RGB camera or a camera with depth sensing capabilities.

[0368] - Processing unit: such as a central processing unit, graphics processing unit or other general-purpose processor, used to execute local image preprocessing and network communication programs on the terminal.

[0369] - Communication unit: such as wireless communication module or wired network interface, used for bidirectional data transmission with server.

[0370] - Display devices: such as liquid crystal displays and organic light-emitting displays, used to display intermediate text information and text information.

[0371] - Audio output device: such as a speaker or headphone jack, used to play voice information received from the server.

[0372] The terminal executes local programs through its operating system (such as a general mobile operating system or desktop operating system) and its image acquisition interface and network communication interface to perform preliminary processing on the image information acquired by the camera device, and then sends the processed data to the server.

[0373] 2. Server Components The servers include: - Image preprocessing component: Deployed on the server processing unit, it is used to perform format conversion, temporal segmentation, and encoding on image information received from the terminal. This component can be implemented using image processing libraries, such as general-purpose image processing software libraries.

[0374] - Feature processing component: Used to extract motion feature information of various parts of the body from preprocessed image information. It can call pose estimation model, key point detection model or other feature extraction algorithms.

[0375] - Generative Model Invocation Component: Used to construct prompt statements and input the prompt statements and feature information into the generative artificial intelligence model to generate intermediate text information and natural language text.

[0376] - Text generation component: Used to combine intermediate text information with dialogue history for natural language processing to obtain text information.

[0377] - Voice output component: Used to call the voice synthesis module to convert text information into voice information and send it to the terminal through the communication module.

[0378] - Data storage unit: Used to store trained model parameters, prompt statement templates, dialogue history, and log information.

[0379] The server can use a general-purpose server hardware platform, such as a rack server or cloud computing platform containing a multi-core central processing unit and a graphics processing unit. The server implements the functions of the above components through a software stack including network server software, deep learning inference frameworks (such as general tensor computation frameworks), and audio processing libraries.

[0380] 3. User Roles Users include those who input through physical gestures (such as users who primarily use sign language) and those who understand through voice output (such as users with normal hearing). Users interact with the system by making physical gestures in front of the camera, observing the text on the display device, and listening to the voice output from the terminal.

[0381] II. Program Structure and General Form of Data Flow The server implements data flow between image preprocessing, feature processing, generative model invocation, text generation, and speech output components using a modular software approach. Data is transferred between components through well-defined data structures, such as: - Image frame tensor: for example, a four-dimensional array structure [time step, height, width, number of channels]; - Feature sequence: for example, a three-dimensional array structure [time steps, number of key points, dimensions]; - Intermediate textual information: such as a structure consisting of a sequence of lexical identifiers and confidence levels; - Text information: such as natural language strings and their metadata; - Prompt statements: For example, text composed of instructions, context descriptions, and intermediate text information.

[0382] With a clear data structure, the server can transfer data between internal modules in a fixed format, which facilitates parallel computing and optimized deployment on different hardware platforms, thereby improving processing speed and resource utilization efficiency.

[0383] III. Implementation Forms of Server-Side Image and Feature Processing After receiving the image information uploaded by the terminal, the server performs the following specific technical processing: 1. Implementation Forms of Image Preprocessing The server processes image information at the pixel level and over time, including: - Format conversion: The server decodes the encoded video stream from the terminal into consecutive image frames and converts the pixel format to a predefined format.

[0384] - Size normalization: The server scales each frame of image to the resolution required by the model input, thereby reducing computation and ensuring the consistency of feature distribution.

[0385] - Temporal segmentation: The server divides continuous frames into several segments based on a fixed time window or the detected action segment boundary. Each segment corresponds to one or more complete body action semantic units.

[0386] By performing these preprocessing operations uniformly on the server side, the server's powerful computing resources can be used to centrally process a large amount of data uploaded from terminals, and the overall throughput and processing speed can be improved through vectorized operations and batch processing.

[0387] 2. Implementation Forms of Feature Extraction The server utilizes a feature processing component to extract motion feature information from the preprocessed image information. Specifically, it can employ one or a combination of the following algorithm structures: - Keypoint Detection Model: The server uses a pose estimation network to detect the 2D or 3D coordinates of keypoints such as hands, fingers, wrists, and upper limbs in each frame of the image. This network can employ a convolutional neural network and a heatmap regression structure to output the probability distribution of keypoints, and obtain the keypoint coordinates through post-processing.

[0388] - Temporal feature encoding: The server combines the key point coordinates of consecutive frames in chronological order to form a time series feature, which can be encoded by a unidirectional or bidirectional recurrent neural network, temporal convolutional network or attention mechanism network to capture changes in the speed, direction and rhythm of the action.

[0389] Because the server explicitly considers the differences between multiple frames in the time direction and extracts motion feature information, rather than simply classifying single-frame images, it can significantly improve the recognition accuracy of continuous body movements and reduce the sensitivity to changes in lighting and background complexity.

[0390] IV. Implementation Forms of Generative Artificial Intelligence Models and Prompt Statements After feature extraction, the server uses a generative artificial intelligence model to generate intermediate text information and natural language text. To this end, the server constructs prompts and uses these prompts to drive the model's generation process.

[0391] 1. Implementation Forms of Prompt Statement Construction The server generates multi-level prompts based on motion feature information, intermediate recognition results, and dialogue history. For example: - First type of prompt statement: Used to generate intermediate text information from feature information. The server can construct prompt statements similar to the following: "Based on the input body movement features, the corresponding basic word sequence is output. No embellishment is needed; only the semantics directly expressed by the action are retained." - Second type of prompt statements: used to refine and expand upon intermediate text information. The server can construct the following prompt statements: "The recognized sign language text is: 'I would like to make an appointment with the doctor at 3 p.m. tomorrow.' Please generate a fluent, natural, and polite Chinese sentence without changing the original meaning, to be played to the doctor's receptionist." "The following is a sentence expressed by a deaf person in sign language: 'I would like to make an appointment with the doctor at 3 p.m. tomorrow.' Please help him generate a formal Chinese request and add a brief explanation that the reason for his appointment could be a general outpatient visit." By injecting motion features, dialogue history, and speaker attributes into the prompts, the server can control the style, politeness, and level of detail of the information generated by the generative AI model, thereby achieving fine-tuning of natural language output.

[0392] 2. Structure and Training of Generative Artificial Intelligence Models The generative artificial intelligence models used by the server can employ neural network structures based on self-attention mechanisms. The models include: - Encoding module: used to embed the text sequence in the prompt statement and capture long-distance dependencies through a multi-layer self-attention network.

[0393] - Decoding module: Used to generate the target text sequence step by step, given the encoded result.

[0394] - Structures such as positional encoding, layer normalization, and residual connections are used to stabilize training and improve expressive power.

[0395] During the training phase, the server is jointly trained using a large amount of labeled body action-text pairs and dialogue corpora. The server employs cross-entropy loss or sequence-to-sequence loss functions and iteratively updates the model parameters through backpropagation and gradient descent optimization methods (such as adaptive learning rate optimization). During training, the server can utilize data augmentation techniques, such as randomly scaling the body action timeline and performing synonym substitutions on text expressions, to improve the model's robustness to deformed inputs.

[0396] Through the above structure and training methods, the generative artificial intelligence model can not only generate natural language text from intermediate text information, but also flexibly adjust the output form according to different prompts, giving the system a significant technical advantage in the text generation process.

[0397] V. Implementation Forms of Text Generation and Speech Output After obtaining the natural language text, the server uses a text generation component and a speech output component to achieve the final output.

[0398] 1. Implementation Forms of Text Generation Components The server generates natural language text based on intermediate text information and dialogue history, utilizing generative artificial intelligence models or other language models. This component can select different generation strategies according to different application scenarios: When a strict correspondence with the original action content is required, the server can make only minor word order adjustments and grammatical corrections to keep the content basically consistent with the intermediate text information.

[0399] When more complete information is needed, the server can allow the model to generate explanatory statements or polite phrases based on the dialogue history and prompts, thereby improving comprehensibility and communication friendliness.

[0400] 2. Implementation Form of Voice Output Component The server invokes the speech synthesis module to convert text information into speech information. The speech synthesis module can employ a neural network-based acoustic model and vocoder structure. After inputting text, the server obtains the corresponding digital audio signal. Based on the speaker attributes and dialogue scenario described in the notes, the server can adjust the speech rate, tone, and emotion parameters of the generated speech to make the speech output more suitable for the specific usage scenario.

[0401] The server sends the generated voice information to the terminal via the network, and the terminal then plays it through the local audio output device, thereby converting body movements into audible speech and enabling communication in the real world.

[0402] VI. Implementation Modes of Terminal-Side Coordination Processing In this invention, the terminal not only undertakes the functions of acquiring image information and playing voice, but also performs a certain degree of local preprocessing and display control, thereby reducing the server load and improving the system response speed.

[0403] 1. Image Acquisition and Preliminary Preprocessing The terminal can perform simple compression encoding and frame rate control on the raw video stream captured by the camera device to reduce network bandwidth consumption while ensuring the integrity of motion information. The terminal can also perform simple cropping or resolution scaling locally based on device performance before sending the processing results to the server.

[0404] 2. Text and voice display After receiving the intermediate text information and final text information returned by the server, the terminal can present them on the display device in different forms. For example, the intermediate text information can be displayed in the system debugging interface, and the final text information can be displayed in a user-facing dialog box. By comparing the intermediate information with the final text, the terminal facilitates system optimization by developers and maintenance personnel.

[0405] VII. Technical Effects and Reasons Through the above structure and processing flow, the present invention possesses the following technical effects and their causal relationships at the computer technology level: 1. Improved recognition accuracy The server utilizes temporal segmentation and motion feature extraction in its feature processing component, moving beyond reliance on single-frame images to perform unified modeling of multi-frame sequences. This allows for more accurate differentiation of action boundaries and subtle changes. This time-series feature-based processing approach gives the system greater robustness and accuracy in continuous body motion recognition.

[0406] 2. Optimization of processing speed and resource utilization The server utilizes unified image preprocessing and feature processing components, employing vectorized computation and batch processing strategies to process data uploaded from multiple terminals in parallel across multi-core CPUs and graphics processing units. This improves overall throughput and reduces latency per request. Simultaneously, terminals perform appropriate preprocessing locally, reducing the decoding and scaling burden on the server and further enhancing system efficiency.

[0407] 3. Reduced communication load By compressing image data and controlling the frame rate, and by segmenting and fragmenting the data over time, the terminal can reduce the transmission of redundant image data while maintaining recognition accuracy. Furthermore, the server can trigger high-precision recognition only when valid body movements are detected, avoiding the processing of numerous invalid frames and thus reducing network and computing resource consumption.

[0408] 4. Improved text generation quality The server constructs prompts containing information such as motion characteristics, dialogue history, and speaker attributes to guide the generative AI model in generating natural language text. This makes the output text more grammatically, stylistically, and politely consistent with real-world dialogue scenarios. This prompt-based control mechanism transforms text generation from a simple replacement of fixed rules into a flexible adjustment based on learned language distribution, significantly improving the quality of the output text.

[0409] 5. Independent processing rules, different from those for human tasks. Instead of simply mimicking the human translation process, the server automatically learns the correspondence between motion features and semantic representations through the model's internal weight matrix and attention mechanism. During training, the model continuously adjusts its parameters through error functions and gradient updates, enabling the system to form a high-dimensional feature space and probability distribution rules that differ from human experience. This rule system based on statistical learning and deep network structures allows the system to adaptively process complex and variable action inputs, rather than relying on fixed and difficult-to-extend human rules.

[0410] VIII. Optional Implementation Methods and Variations This invention is not limited to the specific embodiments described above, and can be modified in various ways in the following aspects: 1. Feature extraction methods can employ different pose estimation models or motion feature extraction algorithms based on optical flow.

[0411] 2. Generative AI models can be replaced by natural language generation models of different scales or architectures, as long as they can accept prompts and output intermediate text information or text information.

[0412] 3. The speech synthesis module can adopt different types of acoustic models and vocoder structures to adapt to the needs of different languages ​​and speech styles.

[0413] 4. The terminal can be any type of computing device, including fixed devices, mobile devices, or wearable devices, as long as it can communicate with the server and has image acquisition and audio output capabilities.

[0414] Through the various implementations and modifications described above, the present invention can realize an integrated processing flow from body movements to natural language text and then to voice output in different hardware and software environments, thereby providing a technical solution with high precision, high efficiency and high scalability in the field of computer technology.

[0415] use Figure 13 The processing flow is explained.

[0416] Step 1: Users input body gestures in front of the terminal.

[0417] Input: The user's actual body movements (e.g., sign language, gestures, upper limb movements).

[0418] Output: None (the physical action is converted into a digital signal by subsequent steps).

[0419] Users face the camera device on the terminal, keep both hands in the camera area, and express semantic content by continuously swinging their arms, changing hand shapes, and changing the position of gestures, such as a sequence of actions to say "hello" or "I would like to make an appointment with a doctor at 3 pm tomorrow".

[0420] Step 2: The terminal uses camera equipment to collect image information.

[0421] Input: The user's body movements and postures in space.

[0422] Output: Raw video data or image frame sequence arranged in chronological order.

[0423] The terminal calls the operating system's camera interface to control the camera device to continuously capture images at a predetermined frame rate (e.g., 30 frames per second). The terminal writes the pixel matrix of each frame into a buffer and records it by timestamp to form the raw video stream, which is used for subsequent image processing and feature extraction.

[0424] Step 3: The terminal performs local preprocessing and encoding of the acquired image information.

[0425] Input: The raw video data or image frame sequence obtained in step 2.

[0426] Output: Video data or image data packets that have been scaled, formatted, and compressed.

[0427] The terminal performs resolution adjustment (e.g., scaling from 1920×1080 to 224×224) on each frame of the image, performs color format conversion and simple noise reduction; the terminal packages multiple frames of images in chronological order and uses compression coding algorithms to encode the data to reduce data volume, thereby reducing the bandwidth usage of subsequent network transmission.

[0428] Step 4: The terminal sends preprocessed image data to the server via the communication unit.

[0429] Input: Encoded video data or image data packets obtained in step 3, and necessary metadata (timestamp, session identifier, etc.).

[0430] Output: The data stream transmitted to the server over the network.

[0431] The terminal establishes a network connection with the server, encapsulates the encoded data into a network packet, and sends it through a specified communication protocol. When sending the packet, the terminal includes metadata such as user identifier, frame rate, and resolution so that the server can configure the decoding and processing flow based on this information.

[0432] Step 5: The server receives and decodes image data from the terminal.

[0433] Input: Encoded video data or image data packets sent by the terminal.

[0434] Output: A sequence of decoded image frames arranged in chronological order.

[0435] The server receives data streams through the network service interface, inputs the received compressed and encoded data into the multimedia decoding module, performs decoding operations (such as entropy decoding, inverse quantization, inverse transform, etc.), and restores them into continuous image frames; the server adds a unified internal time index to each frame to form an image frame sequence for feature extraction.

[0436] Step 6: The server performs image preprocessing and temporal segmentation on the image frame sequence.

[0437] Input: The sequence of decoded image frames obtained in step 5.

[0438] Output: Batch of image sequences that have been size-normalized, processed, and divided by time segments.

[0439] The server performs operations such as scaling and pixel value normalization on each frame of the image, converting the pixel matrix into a floating-point tensor. Based on a fixed time window or the detected action boundary, the server divides the continuous frames into several time segments and assigns an independent identifier to each segment for subsequent feature extraction and recognition.

[0440] Step 7: The server extracts motion feature information of various body parts from the preprocessed image fragments.

[0441] Input: Image sequence fragments obtained in step 6.

[0442] Output: A sequence of motion features containing the coordinates of key points and their temporal variations.

[0443] The server calls the feature processing component to perform key point detection operations on each frame of the image, and identifies the spatial positions of key points such as hands, fingers, wrists and upper limbs. The server combines the key point coordinates of each frame into a time series in chronological order, and calculates derived features such as inter-frame displacement, velocity and direction changes to form a motion feature sequence used to represent the dynamic pattern of body movements.

[0444] Step 8: The server generates prompts based on motion feature information for recognizing intermediate text information.

[0445] Input: The motion feature sequence and its statistical information (such as motion duration, amplitude of change, etc.) obtained in step 7.

[0446] Output: A first-class prompt statement containing instructions.

[0447] Based on the start, end, rhythm, and intensity of the action reflected in the motion features, the server generates text-based prompts, such as adding instructions like "Only the basic word sequence directly corresponding to the action needs to be output, without any embellishment." The server embeds action-related metadata (such as "continuous short actions" and "long pauses") into the prompts to guide the generative AI model to focus on basic semantic recognition.

[0448] Step 9: The server inputs motion features and the first type of prompt statement into the generative artificial intelligence model to generate intermediate text information.

[0449] Input: the motion feature sequence from step 7, and the first type of prompt statement from step 8.

[0450] Output: Intermediate textual information representing the basic semantics of body actions (e.g., lexical sequences and their confidence levels).

[0451] In the generative artificial intelligence model, the server embeds the input prompts and encodes motion features, then fuses these two types of vectors within the model. The server calculates the probability of the correspondence between features and words at each time step through a self-attention network or sequence modeling network, performs decoding operations, outputs a word sequence representing the basic semantics, and assigns a confidence value to each word to form intermediate text information.

[0452] Step 10: The server generates a second type of prompt statement based on the intermediate text information and the dialogue history.

[0453] Input: The intermediate text information obtained in step 9, and the dialogue history data stored on the server.

[0454] Output: Second-class prompts used for natural language processing and polishing.

[0455] The server extracts contextual information related to the current intermediate text from the dialogue history, such as key sentences from previous or multiple rounds of dialogue. Based on the application scenario (e.g., medical appointment, daily greeting), the server constructs instructions and embeds the intermediate text into the prompt statement. For example: "The recognized sign language text is: 'I would like to make an appointment with a doctor at 3 PM tomorrow.' Please generate a fluent, natural, and polite Chinese sentence without changing the original meaning, to be played to the doctor's receptionist." This prompt statement is then used as one of the inputs to the generative artificial intelligence model.

[0456] Step 11: The server inputs the second type of prompt statement into the generative artificial intelligence model to generate text information.

[0457] Input: The second type of prompt statement generated in step 10 and the embedded representation of the dialogue history.

[0458] Output: Text information in natural language form.

[0459] In the generative artificial intelligence model, the server encodes prompts containing intermediate text information and dialogue history, and generates natural language text word by word or phrase by phrase through multi-layer self-attention and decoding operations. During the generation process, the server adjusts the word selection and sentence structure according to the style and politeness requirements described in the prompts, and outputs text information that meets the needs of the scenario, such as "I would like to make an appointment with a doctor at 3 pm tomorrow, please help me arrange it". Step 12: The server adjusts the text information as needed based on the speaker's attributes and the context of the conversation.

[0460] Input: The text information obtained in step 11, as well as the speaker attributes and scene information inferred based on body movements or user configuration.

[0461] Output: The final text information after adjustments to style and expression.

[0462] The server processes the text information based on attributes such as the speaker's age and formality, such as replacing it with more formal or colloquial expressions. The server can also use language models or rule bases to check whether honorifics need to be added or sentences simplified, thereby obtaining the final text information that is more closely matched to the current scenario.

[0463] Step 13: The server inputs the final text information into the speech synthesis module to generate speech information.

[0464] Input: The final text information obtained in step 12.

[0465] Output: Digital audio data representing the speech signal.

[0466] The server calls the speech synthesis module to segment and convert the input text into words and phonemes. Then, it uses an acoustic model and a vocoder to calculate the speech waveform. The server adjusts the reference pitch, speech rate and intonation according to the speaker's attributes and scene information to obtain a speech data containing natural pauses and emotional changes, and outputs it in a specified format (such as sampling rate and bit depth).

[0467] Step 14: The server sends voice and text messages to the terminal through the communication unit.

[0468] Input: Voice information generated in step 13 and final text information from step 12.

[0469] Output: The response data packet transmitted to the terminal over the network.

[0470] The server encapsulates the voice information and the corresponding text information into a response structure, writes it into the network buffer, and sends it to the terminal through a pre-established connection. The server records a timestamp and session identifier in the response, which helps the terminal correctly associate the current output with the corresponding action input.

[0471] Step 15: The terminal receives and decodes the voice information sent by the server.

[0472] Input: The response data packet returned by the server, including voice and text information.

[0473] Output: A playable audio stream and a text string for display.

[0474] The terminal reads response data from the receive buffer and parses out voice and text information. The terminal performs necessary decoding or format conversion on the voice information, transforming it into an audio stream that the device can play, and stores the text information in the local cache for later presentation on the display device.

[0475] Step 16: The terminal plays voice through the audio output device and displays text information on the display device.

[0476] Input: The audio stream and text string obtained in step 15.

[0477] Output: User-audible voice output and visible text display.

[0478] The terminal sends the audio stream to the local audio driver to control the speaker or headphones to output voice; the terminal displays text information on the display device with an appropriate font size and layout, such as displaying "I would like to make an appointment with the doctor at 3 pm tomorrow, please arrange it for me" in a speech bubble, so that hearing users can understand the content expressed by body movements through both auditory and visual channels.

[0479] Step 17: The user (hearing person) receives voice and text and understands or responds to them.

[0480] Input: Voice information output by the terminal and text information displayed.

[0481] Output: The user's understanding of the semantic content of body movements and possible subsequent interactive actions.

[0482] Users listen to the voice content played by the terminal's speaker and simultaneously refer to the text information on the screen to understand the intentions of deaf users or users who primarily use body language; users can choose to respond through spoken language or by using the terminal's input interface, thereby completing a communication process based on this system.

[0483] Application Example 2 The process flow corresponding to the specific processing in Use Case 2 will be described below. The various parts of the system described below are implemented by the data processing device 12 and the intelligent device 14. In addition, the data processing device 12 is referred to as the "server" and the intelligent device 14 is referred to as the "terminal".

[0484] There are many shortcomings in the computer technology aspect of existing sign language translation systems.

[0485] First, traditional systems often use fixed rules or single-pattern recognition algorithms, which directly generate text based on hand trajectories or local image features. This lacks a comprehensive understanding of the context and dialogue, resulting in stiff and unnatural text that fails to meet the requirements of politeness, naturalness, and scene adaptability in real human-computer interaction scenarios.

[0486] Secondly, existing technologies typically process sign language content recognition and emotion recognition separately, or simply attach simple emotion tags, without jointly modeling and adjusting emotional information and parameters during the underlying text generation and speech synthesis processes. Therefore, the system struggles to accurately represent the user's true emotional state at the speech level, resulting in output speech lacking emotional color and warmth, thus limiting the high-quality communication experience between hearing-impaired and hearing individuals.

[0487] Furthermore, traditional architectures often rely on simple preprocessing on the edge and classification in the cloud, failing to fully utilize the temporal alignment and feature fusion between multimodal data (sign language gestures, facial expressions, acoustic features). They also lack an integrated pipeline design around "time series features → initial text → prompting statements → generative artificial intelligence model → corrected text → emotion-modulated speech," making it difficult for the system to maintain consistency in recognition accuracy, semantic naturalness, and emotional expression under high real-time requirements.

[0488] Furthermore, existing systems mostly directly input the recognition results into the speech synthesis module without introducing generative artificial intelligence models to participate in the optimization of intermediate text and natural language generation. They also do not conduct dedicated computer implementation design at the system level for the structure, language conditions, and candidate result selection strategies of the "prompt statements," thus failing to fully leverage the advantages of generative artificial intelligence models in natural language processing.

[0489] In summary, a technical solution is needed that improves the computer architecture and processing flow, enabling the server to: efficiently extract features from multimodal time series data; generate structured prompts for generative artificial intelligence models; automatically introduce politeness, naturalness, and scene adaptability constraints during the natural language generation stage; and tightly couple emotion recognition results with speech synthesis parameters. This would improve recognition accuracy, dialogue naturalness, and emotional expression throughout the sign language to speech conversion process, thereby achieving an overall improvement in computer processing power and human-computer interaction quality.

[0490] The specific processing performed by the specific processing unit 290 of the data processing apparatus 12 in Application Example 2 is achieved by the following means.

[0491] In this invention, the server includes: a module for acquiring visual information, including sign language gestures and facial expressions, through an imaging device, and acquiring the visual information as time-series image information; a module for extracting hand and body regions from the time-series image information using image processing computing resources, generating time-series feature quantities of sign language gestures using feature extraction computing resources, and generating initial text information representing sign language content based on the time-series feature quantities; a module for generating prompt statements for instructing natural language output to a generative artificial intelligence model based on the initial text information and dialogue scene information, and inputting the prompt statements into the generative artificial intelligence model to generate corrected text information representing sign language content; a module for estimating emotional information from the facial expressions and acoustic information acquired by an audio input device using emotion recognition computing resources, and correcting the expression content or expression mode of the corrected text information based on the emotional information; and a module for converting the corrected text information into speech data using speech synthesis computing resources while setting acoustic parameters including pitch, speech rate, and volume, and sending the speech data to a terminal device for output by an audio output device through a communication channel, based on the corrected text information and the emotional information. This allows for the construction of a unified processing pipeline on the server side, encompassing multimodal time-series feature extraction, initial text generation, prompt statement construction and generative AI model invocation, emotional information inference and text correction, and emotion-modulated speech synthesis. From a computer technology perspective, this improves the accuracy and robustness of sign language content recognition, enhances the politeness, naturalness, and scene adaptability of natural language output, and achieves more human-friendly speech output through emotion-driven acoustic parameter control, thereby comprehensively improving the performance of computer-based sign language translation and human-computer interaction.

[0492] A "system" refers to a collection of devices consisting of multiple hardware components and software modules, used to collect, transmit, process, and output processing results from input data.

[0493] A "server" refers to an information processing device in a system that is responsible for receiving data from terminals and processing it centrally, performing computational tasks such as image processing, feature extraction, emotion recognition, generative artificial intelligence model invocation, and speech synthesis.

[0494] "Terminal device" refers to an information processing device that interacts directly with the user, is equipped with an imaging device, an audio input device, and an audio output device, and exchanges data with a server through a communication channel.

[0495] An "imaging device" refers to an image acquisition device used to acquire information about the user's surroundings and physical condition, converting optical images into digital image signals.

[0496] "Visual information" refers to image data acquired by an imaging device and represented in digital form, including visual features such as the user's sign language gestures, facial expressions, and body posture.

[0497] "Time-series image information" refers to a sequence of image data composed of multiple image frames arranged in chronological order, used to represent sign language gestures and facial expressions that change over time.

[0498] "Sign language gestures" refer to continuous limb movements performed by users to express semantics, including changes in hand position, hand shape, and arm movement trajectory.

[0499] "Facial expressions" refer to information patterns formed by the movement of a user's facial muscles to express emotional states, including facial expressions such as smiling, frowning, and surprise.

[0500] "Image processing computing resources" refers to processors, memory, and related software modules in servers or other information processing devices used to perform image processing operations such as image preprocessing, region segmentation, and object detection.

[0501] "Feature extraction computing resources" refers to the set of computing hardware and software used to calculate and output key features from image information or time series data, including feature extraction algorithms and processor resources used to run the algorithms.

[0502] "Time series features" refers to a set of numerical features extracted from time series image information and arranged in chronological order to describe the dynamic changes of sign language movements.

[0503] "Initial text information" refers to the first-stage text data that roughly represents the content of sign language in text form, based on time series features.

[0504] "Dialogue context information" refers to contextual information related to the current sign language environment, including interactive roles, venue type, service purpose, and expected level of politeness, which are used to guide natural language generation.

[0505] "Generative AI models" refer to AI models that automatically generate text or other data forms based on input prompts, possessing natural language understanding and natural language generation capabilities.

[0506] "Prompt statements" refer to natural language input text used to explain task requirements, constraints, and contextual information to generative artificial intelligence models, in order to guide the generative artificial intelligence models to output results in the desired form.

[0507] "Modified text information" refers to text data that has been modified in terms of semantics, politeness, and naturalness, based on initial text information and dialogue scenario information, and generated or optimized by a generative artificial intelligence model.

[0508] "Audio input device" refers to an acoustic acquisition device used to collect user voice or ambient sound and convert it into digital audio signals.

[0509] "Acoustic information" refers to audio data acquired by an audio input device and stored in digital form, including the characteristics of speech signals such as spectrum, energy, and temporal structure.

[0510] "Emotion recognition computing resources" refers to computing modules used to analyze visual and / or acoustic information to infer a user's emotional state, including emotion recognition algorithms and computing hardware that executes the algorithms.

[0511] "Emotional information" refers to structured data output by emotion recognition computing resources, used to represent the current emotion category and intensity of a user.

[0512] "Content of expression" refers to the objective semantics contained in the textual information, that is, the written description of the facts, intentions or requests expressed by sign language.

[0513] "Mode of expression" refers to the presentation of a text in terms of politeness, tone, style, and word choice without changing its basic semantics.

[0514] "Speech synthesis computing resources" refers to the computing modules used to convert text information into speech data and the processors and storage resources on which they depend for operation, including speech synthesis algorithms and acoustic models.

[0515] "Acoustic parameters" refer to numerical parameters used to control the characteristics of synthesized speech, including pitch, speech rate, volume, timbre, and pause duration.

[0516] "Speech data" refers to audio data generated by speech synthesis computing resources, stored and transmitted in digital format, and used to be played back as sound by an audio output device.

[0517] A "communication channel" refers to a wired or wireless communication path used to transmit data between a server and a terminal device, including physical links and network protocols.

[0518] "Audio output device" refers to an output device used to convert digital voice data into audible sound and transmit it to the user, including speakers and headphones.

[0519] In this embodiment of the invention, the server, terminal, and user cooperate to achieve end-to-end conversion processing from sign language and facial expressions to emotionally charged speech. The following provides a detailed description of this embodiment in conjunction with its hardware structure, software modules, data structure, algorithm flow, and technical effects.

[0520] I. Overall System Composition In this implementation, the server acts as the central processing node, the terminal acts as the data acquisition and output node, and the user interacts with the system through sign language, facial expressions, and voice.

[0521] 1. Terminal-side hardware and software composition The terminal includes an imaging device, an audio input device, an audio output device, a display device, and a communication module.

[0522] The terminal runs a mobile operating system (such as a general mobile operating system) and the client application corresponding to this invention.

[0523] The terminal uses a built-in imaging device to capture video of the user's sign language and facial expressions, an audio input device to capture sound signals, and an audio output device to play voice data generated by the server.

[0524] 2. Server-side hardware and software composition Servers include multi-core central processing units, graphics processing units, high-speed memory, network interfaces, and large-capacity storage devices.

[0525] The server runs a general-purpose operating system and several software components, including: The server uses an image processing library to perform image preprocessing and feature extraction. The server uses a deep learning framework to perform inference for sign language recognition, pose estimation, and emotion recognition models; The server uses a speech synthesis service or a local speech synthesis module to convert text to speech. The server uses a generative artificial intelligence model service to achieve natural language generation and text correction processing.

[0526] The server stores structured data such as a sign language dictionary database, a scene configuration database, an emotion tag definition table, and model parameter files in the storage device.

[0527] II. Program Modules and Data Flow Structure 1. Terminal program module The terminal uses a video acquisition module to acquire frame-by-frame image data from the imaging device and encodes it into a compressed video stream.

[0528] The terminal uses an audio acquisition module to obtain audio samples from the audio input device and encodes them into a compressed audio stream.

[0529] The terminal uses a communication module to send compressed video and compressed audio streams to the server through a secure communication channel, and receives voice data returned by the server.

[0530] The terminal uses an audio playback module to decode the received voice data and controls the audio output device to play the voice.

[0531] 2. Server program module The server uses a receiving module to obtain multimedia data streams from multiple terminals through the communication channel and divides them into session units and time segments.

[0532] The server uses a video decoding module to decode the compressed video stream into time-series image information.

[0533] The server uses an audio decoding module to decode the compressed audio stream into time-series acoustic information.

[0534] The server uses an image preprocessing module to normalize, suppress noise, crop, and transform the color space of time-series image information.

[0535] The server uses a feature extraction module to generate intermediate feature information such as hand region, skeletal key points, and pose vectors, and organizes them into time series feature quantities.

[0536] The server uses a sign language recognition module to generate initial text information based on time-series features.

[0537] The server uses a prompt generation module to construct prompts for calling generative artificial intelligence models based on initial text information and scene information.

[0538] The server uses the generative AI model interface module to send prompts to the generative AI model and receive correction text information.

[0539] The server uses an emotion recognition module to infer emotional information from visual and acoustic information, and outputs the emotion category and intensity.

[0540] The server uses a text correction module to make secondary adjustments to the content or expression of the corrected text information based on sentiment information.

[0541] The server uses a speech synthesis module to set acoustic parameters based on corrected text information and emotional information to generate speech data.

[0542] The server uses the sending module to package the voice data and return it to the corresponding terminal through the communication channel.

[0543] III. Core Algorithms and Data Structures 1. Sign Language Feature Representation and Recognition During the image preprocessing stage, the server scales each frame of the image to a fixed size and performs normalization in the color space.

[0544] The server uses a pose estimation submodule to estimate the body skeleton for each frame of image, and outputs skeletal information containing the coordinates of multiple key points.

[0545] The server uses the hand detection submodule to locate the hand region in the image and outputs the hand bounding box and the coordinates of several hand key points.

[0546] The server assembles the key point coordinates of multiple time frames into time series features in chronological order and uses a normalization method to eliminate scale and position biases.

[0547] The server uses a deep neural network structure consisting of multi-layer convolutional networks and bidirectional recurrent networks in the sign language recognition module to encode time series features.

[0548] The server uses a sequence-to-sequence mapping or connection-based temporal classification loss decoding method at the output layer to convert the network output probability sequence into a sign language symbol sequence.

[0549] The server uses a sign language dictionary database to map the sequence of sign language symbols into initial text information, such as taking "I am looking for a product" as a rough recognition result.

[0550] By employing time-series features and deep neural network structures, the server can capture the dynamic changes in sign language movements compared to traditional recognition algorithms based solely on static frames, thereby improving the accuracy of sign language recognition and maintaining high robustness under noisy and occluded conditions.

[0551] 2. Generative Artificial Intelligence Model and Prompt Statement Design In the prompt statement generation module, the server constructs natural language prompt statements containing task instructions and constraints based on the initial text information, the current scene type, and preset language conditions.

[0552] When a customer inquires about the location of a product in a service scenario, the server can generate the following prompt: The sign language recognition output is: 'I'm looking for a product.'

[0553] Scenario: A supermarket customer asks a store clerk for help.

[0554] Task: Please rewrite this sentence in natural and polite spoken Chinese. When a user expresses gratitude with a smiley face, the server can generate the following message: The sign language message is: 'Thank you.'

[0555] Emotion: joy, high intensity.

[0556] Task: Please generate a Chinese expression that clearly conveys gratitude and a cheerful tone. When a user asks for the location of a product using sign language, the server can generate the following prompt: "When someone uses sign language to express 'Where is this product?', please convert the sign language into speech, analyze the emotions conveyed, and output the result." The server sends the above prompt to the generative AI model service interface. Based on training with a large-scale corpus, the generative AI model generates natural language text according to the task instructions and linguistic conditions in the prompt, for example: "I'm looking for an item, can you help me?" "Thank you so much for your help, I'm really glad." By explicitly designing the structure and conditions of prompt statements, the server enables generative AI models to meet the requirements of politeness, naturalness, and scene adaptability in their output, thereby improving the traditional text generation method based on fixed templates and achieving more advanced control of natural language expression within the computer.

[0557] 3. Emotion Recognition and Acoustic Parameter Adjustment In the emotion recognition module, the server inputs the facial image sequence into the expression recognition network, which can adopt a combination structure of convolutional neural network and temporal network.

[0558] In acoustic analysis, the server extracts acoustic features such as short-time energy, fundamental frequency, and Mel-frequency cepstral coefficients from the decoded audio signal and inputs them into the emotion classification network.

[0559] The server performs a weighted fusion of facial expression recognition results and acoustic emotion results, and outputs emotional information, such as emotion categories like "joy" and "sad" as well as emotion intensity values.

[0560] In the text correction module, when the emotion is happiness and the intensity is high, the server can correct "he is happy" to "he looks very happy now" so as to express the emotion more fully at the linguistic level.

[0561] In the speech synthesis module, the server sets acoustic parameters based on emotional information. For example, it raises the pitch and speeds up the speech when the speaker is happy, and lowers the pitch and speeds up the speech when the speaker is sad.

[0562] By adjusting acoustic parameters in this emotion-driven manner, the server ensures that the generated speech is acoustically consistent with the emotional information, achieving a level of nuanced emotional expression that is difficult to achieve with traditional rule-based synthesis methods, thus gaining a technological improvement in computer sound output.

[0563] 4. Model Training and Optimization During the model training phase, the server uses a large amount of labeled sign language video data and corresponding text as the training dataset.

[0564] In training the sign language recognition model, the server uses a supervised learning method, employing sequence cross-entropy or connection temporal classification loss as the error function, and updates the network weights through stochastic gradient descent or its variants.

[0565] In training the emotion recognition model, the server uses a multimodal dataset containing emotion labels and employs a multi-task learning structure to simultaneously optimize facial expression recognition and acoustic emotion recognition.

[0566] In designing the prompt message strategy, the server adjusts the prompt message template through offline simulation and manual evaluation to improve the quality of the text output by the generative artificial intelligence model.

[0567] During continuous operation, the server can collect anonymized recognition results and user feedback, and incrementally train or fine-tune the model to gradually improve recognition accuracy and generation quality, thereby achieving long-term improvement in computer learning capabilities.

[0568] IV. Technical Effects and Improvements in Computer Technology 1. Improved recognition accuracy and naturalness The server achieves higher accuracy in sign language content recognition and more natural text output compared to simple rule-based or static classification methods by combining time-series features, deep neural networks, and generative artificial intelligence models.

[0569] By embedding politeness, naturalness, and context-appropriateness constraints into the prompts, the server enables the generated text to move beyond fixed short sentences and automatically adjust its expression according to the context, thus improving the computer's natural language processing capabilities.

[0570] 2. Improvement of emotional expression and speech quality The server obtains emotional information by jointly analyzing visual and acoustic information, and uses the emotional information for text correction and voice parameter settings to achieve emotion-driven voice output, rather than just simple text broadcasting.

[0571] The server sets acoustic parameters such as pitch, speech rate, and volume during the speech synthesis stage, making the generated speech technically closer to the natural human speaking pattern, which helps reduce the misunderstanding rate and improve communication efficiency.

[0572] 3. Optimization of processing efficiency and resource utilization By performing basic compression on the terminal and centralizing heavy-load computing on the server, the server achieves terminal lightweighting and server resource centralization, which is beneficial to improving overall processing efficiency in a multi-user concurrent environment.

[0573] The server reduces redundant data transformations by using a unified data structure (time series features, intermediate feature information, sentiment information structure, etc.), thereby reducing computational complexity and memory usage.

[0574] The server transmits only the necessary compressed video and audio data and synthesized speech data at the network transport layer, avoiding repeated transmission of intermediate results and reducing communication load.

[0575] 4. The essential differences from traditional human translation In this invention, the server does not simply mimic human translation steps, but rather employs specific data structures, model architectures, and prompt statement control strategies to represent and manipulate sign language features and emotional information in a machine-processable form.

[0576] The server uses pre-trained models and online inference mechanisms to ensure that the sign language recognition and natural language generation processes follow machine learning rules and the principle of minimizing the loss function, rather than human experience rules, thus forming a processing path at the computer technology level that is independent of the human translation process.

[0577] V. Optional Implementation Methods and Variations 1. Deformation of the model structure In some implementations, the server can use an attention-based sequence-to-sequence structure instead of a recurrent network structure to achieve sign language sequence processing over longer time spans.

[0578] In other implementations, the server can share some parameters between the sign language recognition network and the emotion recognition network to reduce computational burden and improve feature representation capabilities through multi-task learning.

[0579] 2. Replacement of generative artificial intelligence models The server can select generative AI models from different vendors or of different scales based on the deployment environment, including lightweight local models and remote cloud models.

[0580] When deploying generative AI models locally on a server, a trimmed version of the model can be used to meet real-time requirements, and output quality can be maintained through prompting statements.

[0581] 3. Terminal-side preprocessing and distributed computing In some implementations, the terminal can perform partial key point detection or attitude estimation, sending only the key point coordinates to the server to further reduce bandwidth usage.

[0582] After receiving the key point sequence, the server can directly perform sign language recognition and emotion analysis in the feature space without decoding the entire video, thereby improving processing speed.

[0583] Through the various implementation forms and variations mentioned above, a concrete and complete data processing link is formed between the server, terminal, and user. This enables the generative artificial intelligence model and prompts to be closely integrated with image processing, feature extraction, emotion recognition, and speech synthesis within the computer. This achieves improvements in computer technology for sign language translation and emotional speech output, rather than remaining at the level of abstract business process automation.

[0584] use Figure 14 The processing flow is explained.

[0585] Step 1: Users perform sign language gestures using their hands and arms, and create facial expressions through changes in facial muscles, while facing the imaging device. Users keep their upper body and hands in the center of the camera frame so that the device can fully capture the sign language gestures and expressions.

[0586] Input: None (physical behavior generated by the user's natural actions).

[0587] Output: None (providing a practical scenario for subsequent data collection by the terminal).

[0588] Step 2: The terminal uses an imaging device to capture the user's sign language gestures and facial expressions at a predetermined frame rate (e.g., 30 frames per second), generating a continuous sequence of image frames. The terminal calls the system's camera interface to convert analog optical signals into digital image data, which is then cached in memory in timestamp order.

[0589] Input: Actual sign language gestures and facial expressions from the user.

[0590] Data processing / calculation: The terminal converts the optical signal into a raw pixel matrix via the image sensor and adds a timestamp to each frame.

[0591] Output: Time-series image data (multi-frame images at original resolution).

[0592] Step 3: The terminal uses an audio input device to collect sounds produced by the user (such as vocalizations, laughter, and exhalations) and generates time-series audio samples at a fixed sampling rate (e.g., 16kHz). The terminal calls the audio acquisition interface to convert analog sound waves into digital audio data and buffers it.

[0593] Input: Voice signal from the user.

[0594] Data processing / calculation: The terminal converts analog sound into PCM audio samples using an analog-to-digital converter and adds a timestamp.

[0595] Output: Time-series PCM audio data.

[0596] Step 4: The terminal compresses and preprocesses the acquired image and audio data. It scales the original image frames to a lower resolution (e.g., from 1920×1080 to 640×360) to reduce data volume and uses a video coding algorithm (e.g., H.264) to encode the image sequence into a compressed video stream. After performing noise reduction and gain control on the PCM audio, the terminal compresses it into a compressed audio stream using an audio coding algorithm (e.g., AAC).

[0597] Input: Time series raw image data, time series PCM audio data.

[0598] Data processing / computation: The terminal performs operations such as resolution scaling, color format conversion, noise reduction, and compression encoding to convert large volumes of raw data into a compressed stream with timestamps.

[0599] Output: Compressed video data stream, compressed audio data stream.

[0600] Step 5: The terminal uses its communication module to establish a secure communication connection with the server (e.g., HTTPS based on TLS or WebSocket), and encapsulates compressed video and audio data into network packets in chronological order before sending them to the server. The terminal maintains retransmission and order control in the transmission queue to reduce packet loss and out-of-order delivery.

[0601] Input: Compressed video data stream, compressed audio data stream.

[0602] Data processing / computation: The terminal segments the data stream into network data packets with session ID, frame sequence number, and timestamp, and performs encryption and verification.

[0603] Output: Multimedia data packets transmitted through the communication channel.

[0604] Step 6: The server uses a receiving module to receive multimedia data packets sent by the terminal from the communication channel and reassembles them according to the session ID and frame sequence number. The server then concatenates video and audio data packets into continuous compressed video and audio streams, respectively.

[0605] Input: Multimedia data packets from the terminal.

[0606] Data processing / calculation: The server performs packet parsing, sequence recovery, packet loss detection, and necessary retransmission requests to generate a continuous media stream.

[0607] Output: Continuously compressed video stream, continuously compressed audio stream.

[0608] Step 7: The server uses a video decoding module to decode the compressed video stream and restore it to time-series image frames; the server uses an audio decoding module to decode the compressed audio stream into a PCM audio sample sequence.

[0609] Input: Compressed video stream, compressed audio stream.

[0610] Data processing / calculation: The server calls the decoding library to perform H.264 decoding on the video to obtain a frame-by-frame image matrix, and performs AAC decoding on the audio to obtain time-series PCM data.

[0611] Output: Time-series image frames, time-series PCM audio data.

[0612] Step 8: The server uses an image preprocessing module to standardize the time-series image frames. For each frame, the server performs operations such as size unification (scaling to a fixed resolution), color space conversion (e.g., YUV to RGB), brightness normalization, and noise reduction. Based on the person detection results, the server crops the region containing the user's upper body to reduce subsequent computational load.

[0613] Input: Time series image frames.

[0614] Data processing / calculation: The server performs image scaling, cropping, denoising, and normalization to convert the original image into tensor data in a uniform format.

[0615] Output: Preprocessed time series image tensor.

[0616] Step 9: The server uses a pose estimation and hand detection module to extract human skeletons and hand keypoints from the preprocessed image tensor. The server applies a pose estimation algorithm to each frame to obtain the coordinates of multiple joints (head, shoulder, elbow, wrist, etc.), and uses a hand detection network to extract the hand region and finger joints.

[0617] Input: Preprocessed time series image tensor.

[0618] Data processing / calculation: The server inputs the image tensor into the keypoint detection network, outputs the two-dimensional coordinates of multiple keypoints in each frame, and performs coordinate normalization and visibility determination.

[0619] Output: A sequence of intermediate feature information containing key information about the hands and body.

[0620] Step 10: The server combines intermediate feature information in chronological order to form time-series feature quantities, constructing an input sequence for sign language recognition. The server performs translation and scaling normalization on key point coordinates to eliminate differences in individual body size and camera position, and calculates dynamic features such as speed and acceleration.

[0621] Input: Intermediate feature information sequence.

[0622] Data processing / calculation: The server concatenates the coordinates of key points in each frame into a multi-dimensional vector sequence, and adds time difference features and motion trajectory features to form high-dimensional time series features.

[0623] Output: Time-series features used for sign language recognition.

[0624] Step 11: The server uses a sign language recognition module to map time-series features into initial text information. The server inputs the time-series features into a deep neural network, which includes several convolutional layers for local spatial feature extraction and bidirectional recurrent layers or attention mechanisms for modeling temporal dependencies. At the output, the server uses a classification layer to output the probability distribution of sign language symbols for each time slice, and obtains the symbol sequence through a decoding algorithm (such as CTC decoding or Beam Search). Finally, the server converts the symbol sequence into coarse text, such as "I'm looking for goods," based on a sign language dictionary.

[0625] Input: Time series features.

[0626] Data processing / calculation: The server performs forward inference, calculates the class probability at each time step, performs sequence decoding to eliminate duplicate and blank symbols, and maps them to the vocabulary.

[0627] Output: Initial text information representing the sign language content.

[0628] Step 12: The server uses dialogue scenario information and a rule base to determine the current interaction scenario (e.g., supermarket consultation, hospital consultation), and associates the scenario tag with the initial text information. The server retrieves parameters such as the location type, user role, and expected politeness level corresponding to the session ID from the scenario database.

[0629] Input: Initial text information, session ID, and scene configuration data.

[0630] Data processing / calculation: The server determines scene tags based on user registration information, terminal location, and business configuration, and encapsulates them together with the initial text information into structured data.

[0631] Output: Structured input data containing initial text information and scene information.

[0632] Step 13: The server uses a prompt generation module to construct prompts suitable for generative AI models based on initial text and scene information. The server embeds task descriptions, language style requirements, and example formats into the prompts, enabling the generative AI model to output compliant natural language text.

[0633] Input: Initial text information, scene information.

[0634] Data processing / calculation: The server inserts the initial text information into a predefined prompt template and adds language constraints such as "polite, natural, and context-appropriate".

[0635] Output: Prompt statements for generative artificial intelligence models.

[0636] Step 14: The server sends prompts to the generative AI model via an interface module and receives the generated candidate text. The server specifies generation parameters such as temperature and maximum length in the interface call to control output diversity and stability. The server parses one or more candidate texts from the model's response, such as "I'm looking for an item, can you help me?".

[0637] Input: Prompt statement.

[0638] Data processing / calculation: The server packages the prompt statement into a request, sends it to the generative artificial intelligence model service, receives the response, and parses out the candidate text list.

[0639] Output: One or more candidate natural language texts.

[0640] Step 15: The server uses a candidate text filtering module to select corrected text information from multiple candidate texts based on language conditions and a preset scoring function. The server calculates scores for politeness, naturalness, and length suitability for each candidate text, and selects the text with the highest score as the corrected text information.

[0641] Input: A set of candidate natural language texts and language conditions.

[0642] Data processing / calculation: The server performs keyword detection, grammatical integrity checks, and length comparisons on the candidate texts, and calculates a comprehensive score.

[0643] Output: Corrected text information representing the sign language content.

[0644] Step 16: The server uses an emotion recognition module to jointly analyze facial expressions in time-series image frames and acoustic features in time-series PCM audio data. The server inputs facial region images into an expression classification network, outputting emotion category probabilities. It also extracts short-time energy, fundamental frequency, and MFCC features from the audio data and inputs these features into an acoustic emotion classification network. The server then fuses the two emotion results to obtain the final emotion information (e.g., emotion "joy," intensity 0.8).

[0645] Input: Time-series image frames (facial region), time-series PCM audio data.

[0646] Data processing / calculation: The server performs image cropping, acoustic feature extraction, facial expression classification, and emotion probability weighted fusion.

[0647] Output: Emotional information (emotion category and intensity).

[0648] Step 17: The server uses a text correction module to further adjust the expression of the corrected text based on sentiment information. When the server expresses strong happiness, it transforms neutral sentences into more emotionally charged expressions, such as changing "thank you" to "I am very grateful for your help, I am really happy."

[0649] Input: Correct text information and sentiment information.

[0650] Data processing / calculation: The server uses rules or reconstructs short prompts to call generative artificial intelligence models to fine-tune the wording and tone of the text.

[0651] Output: Final text information with emotional connotations.

[0652] Step 18: The server uses a speech synthesis module to generate speech data based on the final text and emotional information. The server sets acoustic parameters (such as pitch, speech rate, and volume) according to the emotion category and intensity, and inputs the final text information into the speech synthesis engine. During the synthesis process, the server adjusts the synthesis parameters to match the tone of voice with the emotion.

[0653] Input: Final text information and sentiment information.

[0654] Data processing / calculation: The server transcribes the text into a phoneme sequence, combines an acoustic model with a vocoder to generate waveforms, and adjusts acoustic parameters according to emotion to control timbre and rhythm.

[0655] Output: Voice data in digital format.

[0656] Step 19: The server uses a sending module to encapsulate the generated voice data into audio data packets and sends them to the corresponding terminals through the communication channel. The server appends a session ID and playback order identifier to the data packets to ensure that the terminals can play them in the correct order.

[0657] Input: Voice data.

[0658] Data processing / calculation: The server optionally compresses or encapsulates the voice data, segments it, encrypts it, and then sends it over the network.

[0659] Output: Voice data packets transmitted over the network.

[0660] Step 20: The terminal receives voice data packets from the server and uses the decoding module to restore them into a playable audio stream. The terminal reassembles the data packets in sequence in the buffer to ensure audio continuity. The terminal then calls the local audio playback interface to pass the audio stream to the audio output device.

[0661] Input: Voice data packet.

[0662] Data processing / calculation: The terminal performs unpacking, decryption, and decoding operations to restore the compressed audio format to a PCM stream.

[0663] Output: A playable audio stream.

[0664] Step 21: The terminal plays a voice stream through its audio output device, allowing hearing users present to hear the corresponding natural speech content. After playback is complete, the terminal can send feedback on the playback status to the server, enabling the server to perform log recording or performance statistics.

[0665] Input: A playable audio stream.

[0666] Data processing / calculation: The terminal converts digital audio signals into analog electrical signals, which are then converted into airborne sound waves by the speaker diaphragm.

[0667] Output: Real-world speech output (sound).

[0668] Step 22: The hearing user understands the meaning and emotional state of the hearing-impaired user's sign language based on the voice content output by the terminal, and responds accordingly with voice or gestures. When the user needs to continue communication, they input sign language and facial expressions into the terminal imaging device again, thus triggering a new round of the above processing.

[0669] Input: Voice prompts (sounds) from the terminal.

[0670] Data processing / calculation: Users understand information at a cognitive level and decide on subsequent communication methods (not involving computer processing).

[0671] Output: New sign language gestures and expressions provide input for the next round of system processing.

[0672] The specific processing unit 290 sends the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires sound representing user input regarding the result of the specific processing. The control unit 46A sends the sound data representing user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.

[0673] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI ​​can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI ​​including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI ​​including the generation AI.

[0674] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart device 14, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart device 14. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the smart device 14 or external devices, and the smart device 14 acquires or collects information required for processing from the data processing device 12 or external devices.

[0675] For example, the collection unit is implemented by the control unit 46A of the smart device 14 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the smart device 14 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the output device 40 of the smart device 14 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.

[0676] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the smart device 14.

[0677] Second Implementation Method Figure 3 An example of the configuration of the data processing system 210 according to the second embodiment is shown.

[0678] like Figure 3 As shown, the data processing system 210 includes a data processing device 12 and smart glasses 214. A server can be cited as an example of the data processing device 12.

[0679] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0680] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, and communication I / F 44 are also connected to the bus 52.

[0681] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.

[0682] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the user 20's surroundings (e.g., the field of view defined by an angle equivalent to the field of vision of an average healthy person).

[0683] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.

[0684] Figure 4 This illustrates an example of the main functions of the data processing device 12 and the smart glasses 214. For example... Figure 4 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.

[0685] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0686] The memory 32 stores a data generation model 58 and an emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290. The specific processing unit 290 can use the emotion-specific model 59 to infer the user's emotions and perform specific processing based on the user's emotions. In the emotion inference function (emotion-specific function) using the emotion-specific model 59, various inferences and predictions related to the user's emotions are performed, including inferences and predictions of the user's emotions, but this is not limited to this example. Furthermore, emotion inference and prediction may also include, for example, emotion analysis (parsing).

[0687] In the smart glasses 214, the processor 46 performs reception and output processing. The memory 50 stores the reception and output program 60. The processor 46 reads the reception and output program 60 from the memory 50 and executes the read reception and output program 60 on the RAM 48. The reception and output processing is implemented by the processor 46 operating as a control unit 46A according to the reception and output program 60 executed on the RAM 48. Furthermore, the smart glasses 214 has the same data generation model and emotion-specific model as the data generation model 58 and the emotion-specific model 59, and these models can also be used to perform the same processing as the specific processing unit 290.

[0688] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the smart glasses 214. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0689] Example 1 The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.

[0690] Application Example 1 The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.

[0691] Example 2 The process is the same as that of the specific process in Embodiment 2 described in the first embodiment above, so the description is omitted.

[0692] Application Example 2 The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.

[0693] The specific processing unit 290 sends the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A outputs the result of the specific processing to the speaker 240. The microphone 238 acquires sound input representing the user's input regarding the result of the specific processing. The control unit 46A sends the sound data representing the user's input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.

[0694] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI ​​can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI ​​including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI ​​including the generation AI.

[0695] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart glasses 214, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart glasses 214. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the smart glasses 214 or external devices, and the smart glasses 214 acquires or collects information required for processing from the data processing device 12 or external devices.

[0696] For example, the collection unit is implemented by the control unit 46A of the smart glasses 214 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the smart glasses 214 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the speaker 240 of the smart glasses 214 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.

[0697] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the smart glasses 214.

[0698] Third Implementation Method Figure 5 An example of the configuration of the data processing system 310 according to the third embodiment is shown.

[0699] like Figure 5 As shown, the data processing system 310 includes a data processing device 12 and a head-mounted terminal 314. A server can be cited as an example of the data processing device 12.

[0700] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0701] The head-mounted terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, display 343, and communication I / F 44 are also connected to the bus 52.

[0702] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.

[0703] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the user 20's surroundings (e.g., the field of view defined by an angle equivalent to the field of vision of an average healthy person).

[0704] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.

[0705] Figure 6 This illustrates an example of the main functions of the data processing device 12 and the head-mounted terminal 314. For example... Figure 6 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.

[0706] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0707] The memory 32 stores the data generation model 58 and the emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290.

[0708] In the head-mounted terminal 314, the processor 46 performs the acceptance / output processing. The memory 50 stores the acceptance / output program 60. The processor 46 reads the acceptance / output program 60 from the memory 50 and executes the read acceptance / output program 60 on the RAM 48. The acceptance / output processing is implemented by the processor 46 operating as a control unit 46A according to the acceptance / output program 60 executed on the RAM 48.

[0709] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the head-mounted terminal 314. In the following description, the data processing device 12 will be referred to as the "server" and the head-mounted terminal 314 will be referred to as the "terminal".

[0710] Example 1 The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.

[0711] Application Example 1 The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.

[0712] Example 2 The process is the same as that of the specific process in Embodiment 2 described in the first embodiment above, so the description is omitted.

[0713] Application Example 2 The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.

[0714] The specific processing unit 290 sends the result of the specific processing to the head-mounted terminal 314. In the head-mounted terminal 314, the control unit 46A outputs the result of the specific processing to the speaker 240 and the display 343. The microphone 238 acquires sound input representing the result of the specific processing. The control unit 46A sends the sound data representing the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.

[0715] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 includes prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI ​​can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI ​​including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI ​​including the generation AI.

[0716] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the head-mounted terminal 314, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the head-mounted terminal 314. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the head-mounted terminal 314 or external devices, and the head-mounted terminal 314 acquires or collects information required for processing from the data processing device 12 or external devices.

[0717] For example, the collection unit is implemented by the control unit 46A of the head-mounted terminal 314 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the head-mounted terminal 314 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12 to analyze the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12 to generate a menu using a generation AI. For example, the serving unit is implemented by the speaker 240 and display 343 of the head-mounted terminal 314 or the specific processing unit 290 of the data processing device 12 to provide the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.

[0718] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the head-mounted terminal 314.

[0719] Fourth Implementation Method Figure 7 An example of the configuration of the data processing system 410 according to the fourth embodiment is shown.

[0720] like Figure 7 As shown, the data processing system 410 includes a data processing device 12 and a robot 414. A server can be cited as an example of the data processing device 12.

[0721] The data processing apparatus 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" as understood in this disclosure. The computer 22 includes a processor 28, RAM 30, and memory 32. The processor 28, RAM 30, and memory 32 are connected to a bus 34. Furthermore, the database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0722] Robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and memory 50. The processor 46, RAM 48, and memory 50 are connected to a bus 52. Furthermore, the microphone 238, speaker 240, camera 42, controlled object 443, and communication I / F 44 are also connected to the bus 52.

[0723] Microphone 238 receives instructions from user 20 by receiving sounds emitted by user 20. Microphone 238 captures sounds emitted by user 20 and converts the captured sounds into sound data, which is then output to processor 46. Speaker 240 outputs sound according to instructions from processor 46.

[0724] Camera 42 is a small digital camera equipped with an optical system such as a lens, aperture and shutter, and imaging elements such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, to photograph the area around robot 414 (e.g., the field of view defined by a perspective equivalent to the field of vision of an average healthy person).

[0725] Communication I / F44 is connected to network 54. Communication I / F44 and 26 are responsible for the transmission and reception of various information between processor 46 and processor 28 via network 54. The transmission and reception of various information between processor 46 and processor 28 using communication I / F44 and 26 is performed in a secure state.

[0726] The controlled object 443 includes a display device, LEDs (light-emitting diodes) for the eyes, and motors for driving the arms, hands, and feet. The posture or movement of the robot 414 is controlled by controlling the motors in the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. In addition, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0727] Figure 8 This illustrates an example of the main functions of the data processing device 12 and the robot 414. For example... Figure 8 As shown, in the data processing device 12, specific processing is performed by the processor 28. The specific processing program 56 is stored in the memory 32.

[0728] The specific processing program 56 is an example of a "program" involved in the technology of this disclosure. The processor 28 reads the specific processing program 56 from the memory 32 and executes the read specific processing program 56 on the RAM 30. Specific processing is implemented by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0729] The memory 32 stores the data generation model 58 and the emotion-specific model 59. The data generation model 58 and the emotion-specific model 59 are used by the specific processing unit 290.

[0730] In robot 414, the processor 46 performs the acceptance and output processing. The memory 50 stores the acceptance and output program 60. The processor 46 reads the acceptance and output program 60 from the memory 50 and executes the read acceptance and output program 60 on RAM 48. The acceptance and output processing is implemented by the processor 46 acting as the control unit 46A according to the acceptance and output program 60 executed on RAM 48.

[0731] Next, the specific processing of the specific processing unit 290 of the data processing device 12 will be described. Each part of the system described below is implemented by the data processing device 12 and the robot 414. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 will be referred to as the "terminal".

[0732] Example 1 The process is the same as that of the specific process described in Embodiment 1 in the first embodiment above, so the description is omitted.

[0733] Application Example 1 The process is the same as that in the specific processing described in Application Example 1 of the first embodiment above, so the description is omitted.

[0734] Example 2 The process is the same as that of the specific process in Embodiment 2 described in the first embodiment above, so the description is omitted.

[0735] Application Example 2 The process is the same as that in the specific processing described in Application Example 2 of the first embodiment above, so the description is omitted.

[0736] The specific processing unit 290 sends the result of the specific processing to the robot 414. In the robot 414, the control unit 46A outputs the result of the specific processing to the speaker 240 and the controlled object 443. The microphone 238 acquires sound input representing the result of the specific processing. The control unit 46A sends the sound data representing the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the sound data.

[0737] Data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of data generation models 58 include ChatGPT (registered trademark) (accessible via the internet (URL: https: / / openai.com / blog / chatgpt)). Data generation model 58 is obtained through deep learning on a neural network. Input to data generation model 58 are prompt words containing instructions, and inference data such as sound data representing sound, text data representing text, and image data representing images (e.g., still image data or animation data). Data generation model 58 infers from the input inference data based on the instructions represented by the prompt words and outputs the inference result in one or more data forms, such as sound data, text data, and image data. Data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or induction. The specific processing unit 290 performs the aforementioned specific processing while using data generation model 58. The data generation model 58 can also be a model finely tuned to output inference results from prompts that do not contain instructions. In this case, the data generation model 58 can output inference results based on prompts that do not contain instructions. The data processing apparatus 12, etc., includes various data generation models 58, including AI other than the generation AI. AI other than the generation AI can be, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or Naive Bayes, and can perform various processes, but is not limited to this example. Furthermore, the AI ​​can also be an AI agent. Furthermore, when the processing of the above-mentioned parts is performed by AI, the processing can be performed partially or entirely by AI, but is not limited to this example. Furthermore, the processing performed by the AI ​​including the generation AI can be replaced by processing in the rule base, and the processing in the rule base can also be replaced by processing performed by the AI ​​including the generation AI.

[0738] Furthermore, the processing of the aforementioned data processing system 10 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the robot 414, but it can also be performed by both the specific processing unit 290 of the data processing device 12 and the control unit 46A of the robot 414. Additionally, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the robot 414 or external devices, and the robot 414 acquires or collects information required for processing from the data processing device 12 or external devices.

[0739] For example, the collection unit is implemented by the control unit 46A of the robot 414 or the specific processing unit 290 of the data processing device 12. For example, the acquisition unit uses the camera 42 or communication I / F 44 of the robot 414 to acquire step data, which is then processed by the specific processing unit 290 of the data processing device 12. For example, the analysis unit is implemented by the specific processing unit 290 of the data processing device 12, which analyzes the data from the collection unit and the acquisition unit. For example, the generation unit is implemented by the specific processing unit 290 of the data processing device 12, which uses a generation AI to generate a menu. For example, the serving unit is implemented by the speaker 240 of the robot 414 and the control object 443 or the specific processing unit 290 of the data processing device 12, which provides the generated menu to the user. The correspondence between each unit and the device or control unit is not limited to the above examples and various changes can be made.

[0740] In the above embodiments, examples of specific processing by the data processing device 12 are given, but the technology disclosed herein is not limited to this, and specific processing may also be performed by the robot 414.

[0741] Furthermore, the emotion-specific model 59, acting as an emotion engine, can determine a user's emotion based on a specific mapping. Specifically, the emotion-specific model 59 can determine a user's emotion based on an emotion graph that serves as a specific mapping (see [reference]). Figure 9 The emotion-specific model 59 can also determine the robot's emotion, and the specific processing unit 290 performs specific processing based on the robot's emotions.

[0742] Figure 9 This is a diagram representing an emotion map 400 that maps multiple emotions. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotion is. On the outer side of the concentric circles, emotions representing states or behaviors arising from mood are arranged. Emotions are concepts that include feelings and mental states. Emotions generated by reactions occurring in the brain are arranged roughly to the left of the concentric circles. Emotions derived from situational judgments are arranged roughly to the right of the concentric circles. Emotions generated by reactions occurring in the brain and derived from situational judgments are arranged roughly above and below the concentric circles. Furthermore, "pleasant" emotions are arranged above the concentric circles, and "unpleasant" emotions are arranged below them. Thus, in the emotion map 400, multiple emotions are mapped based on the structure that generates emotions, and emotions that are likely to occur simultaneously are mapped close to each other.

[0743] These emotions are distributed at the three o'clock position of the emotion map 400, typically fluctuating between peace and anxiety. In the right half of the emotion map 400, situational awareness dominates over internal sensation, thus resulting in an impression of calm.

[0744] The inner side of the emotion map 400 represents the inner state, while the outer side represents behavior. Therefore, the further outward you are from the emotion map 400, the more visible the emotion becomes (manifested in behavior).

[0745] Here, human emotions are based on various balances such as posture and blood sugar levels. When these balances deviate from an ideal state, it indicates an unpleasant state; when they approach the ideal state, it indicates a pleasant state. Emotions in robots, cars, motorcycles, etc., can also be created in the following way: based on various balances such as posture and remaining battery power, when these balances deviate from an ideal state, it indicates an unpleasant state; when they approach the ideal state, it indicates a pleasant state. Emotion maps can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a Brain Physiological Signal Analysis System for Voice Emotion Recognition and Emotion, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). In the left half of the emotion map, emotions belonging to the sensory-dominated region, called "response," are arranged. Furthermore, in the right half of the emotion map, emotions belonging to the situational cognition-dominated region, called "situation," are arranged.

[0746] In the emotion map, two types of emotions that promote learning are defined. One is a negative emotion on the situational side, in the middle or peripheral region of "repentance" or "reflection." This occurs when the robot experiences negative emotions such as "I don't want to experience this feeling again" or "I don't want to be blamed again." The other is a positive emotion on the response side, near the "desire" region. This occurs when there are positive feelings such as "wanting more" or "wanting to know more."

[0747] The emotion-specific model 59 inputs user input into a pre-trained neural network to obtain emotion values ​​representing each emotion shown in the emotion map 400, thereby determining the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network... Figure 10 As shown in the sentiment graph 900, it was trained in a way that sentiments that are configured close to each other have similar values. Figure 10 The text shows examples of emotions such as "peace of mind", "stability", and "reassurance" that have similar emotion values.

[0748] The above description focuses on the functions of the data processing device 12, but the system of this disclosure is not necessarily installed on a server. The system of this disclosure can also be installed as a general information processing system. This disclosure can also be installed, for example, as a software program running on a personal computer, an application running on a smartphone, etc. The method of this disclosure can also be provided to users in the form of SaaS (Software as a Service).

[0749] In the above embodiments, an example of a specific process being performed by a single computer 22 is given. However, the technology disclosed herein is not limited to this, and the specific process can also be distributed among multiple computers, including computer 22. For example, the data generation model 58 can be located on an external device of the data processing apparatus 12, where data is generated based on the input data.

[0750] In the above embodiments, examples of storing a specific processing program 56 in the memory 32 have been described, but the technology disclosed herein is not limited thereto. For example, the specific processing program 56 may also be stored in a portable computer-readable non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed into the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0751] Alternatively, a specific processing program 56 may be pre-stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 according to the requirements of the data processing device 12.

[0752] In addition, it is not necessary to store all the specific processing program 56 in the storage device such as the server connected to the data processing device 12 via the network 54 or in the memory 32; a portion of the specific processing program 56 may be stored in advance.

[0753] As hardware resources for performing specific processes, various processors, as shown below, can be used. For example, a CPU can be listed as a processor, which functions as a general-purpose processor that performs specific processes by executing software, i.e., a program. Furthermore, processors can be listed as special-purpose circuits such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application-Specific Integrated Circuits), which are processors with circuitry specifically designed to perform specific processes. Each processor has built-in or connected memory, and each processor executes specific processes using that memory.

[0754] The hardware resources for performing a specific process can consist of one of these various processors, or a combination of two or more processors of the same or different types (e.g., a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resources for performing a specific process can be a single processor.

[0755] As an example of a single processor, there are two approaches: First, a processor is composed of a combination of one or more CPUs and software, which functions as a hardware resource to perform a specific process; second, as represented by a SoC (System-on-a-chip), a processor is used to implement the functionality of the entire system, which includes multiple hardware resources for performing a specific process, using a single IC (Integrated Circuit) chip. In this way, the specific process is implemented by using one or more of the aforementioned processors as hardware resources.

[0756] Furthermore, the hardware architecture of these various processors, more specifically, can utilize circuits that combine semiconductor elements and other circuit components. Moreover, the specific process described above is just one example. Therefore, without departing from the main point, unnecessary steps can certainly be deleted, new steps added, or the processing order changed.

[0757] The descriptions and illustrations above are detailed explanations of a portion of the technology disclosed herein, and are merely one example of the technology disclosed herein. For example, the above descriptions of the structure, function, effect, and results are just one example of the structure, function, effect, and results of a portion of the technology disclosed herein. Therefore, without departing from the spirit of the technology disclosed herein, unnecessary parts may be deleted, new elements added, or replacements may be made to the descriptions and illustrations above. Furthermore, to avoid confusion and facilitate understanding of a portion of the technology disclosed herein, explanations of common technical knowledge that do not require special explanation under the premise of being able to implement the technology disclosed herein have been omitted from the descriptions and illustrations above.

[0758] All documents, patent applications and technical specifications set forth in this specification are incorporated herein by reference to the same extent that each document, patent application and technical specification is specifically and individually described therein and referenced by reference.

[0759] In addition, the following notes are provided in response to the above explanation.

[0760] Example 1 (Note 1) An information processing system, characterized in that it comprises: A device for acquiring image information containing body movements using an image acquisition device; A device for performing feature extraction processing on the image information, calculating temporal body part position information and shape information using image processing algorithms, and generating intermediate recognition information representing the semantic content of posture expression based on the calculation results; An apparatus for generating instruction information consisting of prompt statements that instruct a generative artificial intelligence model to generate text information based on the intermediate recognition information, and inputting the instruction information into the generative artificial intelligence model to generate text data in natural language form; A device for performing language processing algorithms on the generated text data, correcting word order and unifying expression, and outputting text information organized into sentences; A device for taking the processed text information as input to execute a speech synthesis algorithm and generate audio data; A device for sending the audio data to a sound output device on the terminal side, and for causing the sound output device to output the audio data in the form of physical sound; An apparatus for associating and storing the text information and audio data in a recording medium, so as to preserve historical information that can be used for subsequent learning or improvement processing.

[0761] (Note 2) The information processing system according to Appendix 1 is characterized in that, It also includes a learning processing device for acquiring user correction operations on the intermediate recognition information and the text information, and updating the running parameters of at least one of the image processing algorithm and the generative artificial intelligence model based on the correction operations, so as to continuously improve the recognition accuracy of posture expression and the text generation accuracy.

[0762] (Note 3) The information processing system according to Appendix 1 is characterized in that, The prompt statement includes an abstract description that summarizes the body part location information and time-series features extracted from the image information. The generative artificial intelligence model generates text data that conforms to the context of body posture expression based on the abstract description and historical information accumulated in the past.

[0763] Application Example 1 (Note 1) An information processing system, characterized in that it comprises: A device for acquiring time-series image information from an imaging device for acquiring sign language information, and for preprocessing the image information to extract the hand and upper body regions; An apparatus for calculating motion feature quantities including joint point and contour position information for the extracted hand and upper body regions using a machine learning model, and generating time series data. A device for generating prompts for a generative artificial intelligence model together with the time series data to instruct the generation of text information representing the meaning of sign language, and for obtaining text information corresponding to the sign language content from the generative artificial intelligence model based on the prompts and the time series data. A device for comparing text information obtained from the generative artificial intelligence model with the recognition results output by an independently constructed sign language recognition model, and determining the final text information to be used based on the consistency and reliability of the two. A device for generating prompt statements for controlling speech synthesis processing based on the determined text information, generating speech information using speech synthesis technology based on the prompt statements and the text information, and outputting the speech information through a speech output device; An apparatus for associating and storing the determined text information, the time series data, the generative artificial intelligence model, and the sign language recognition model's respective recognition results, and for generating training data based on the stored content to update the learning data of the generative artificial intelligence model and the sign language recognition model.

[0764] (Note 2) The information processing system according to Appendix 1 is characterized in that, The device for generating prompts for generative artificial intelligence models is configured to dynamically change the prompts based on sign language usage, environmental conditions, or historical recognition records, thereby automatically adjusting the content of the prompts so that the accuracy of the text information generated by the generative artificial intelligence model meets predetermined conditions.

[0765] (Note 3) The information processing system according to Appendix 1 is characterized in that, Before generating speech information, the system is configured to generate additional prompts that instruct the generative artificial intelligence model to convert the text information representing sign language content into a form suitable for speech prompts, and to generate speech information based on the converted text information using speech synthesis technology.

[0766] Example 2 (Note 1) An information processing system, characterized in that it comprises: Apparatus for acquiring body movements using camera equipment and preprocessing the acquired image information, segmenting it in the time direction, and encoding it into a predetermined image or video format; An apparatus for extracting feature information corresponding to the movement of various parts of the body from preprocessed image information, generating a prompt statement containing instruction content for generating intermediate text information representing the semantic content of body movements based on the feature information, and inputting the prompt statement into a generative artificial intelligence model to generate the intermediate text information. An apparatus for generating prompts for a generative artificial intelligence model based on the intermediate text information and dialogue history, and for converting the intermediate text information into natural language text to generate text information based on the prompts. A device for converting text information into speech information through speech synthesis processing, and sending the speech information to an external device through communication control, so that the speech output device of the external device outputs the speech information; A means for outputting the text information and the intermediate text information to a display device via display control.

[0767] (Note 2) According to the information processing system described in Appendix 1, the prompting statement input to the generative artificial intelligence model includes speaker attributes inferred from the body movements and additional information about the dialogue scenario, and the generative artificial intelligence model is configured to change the style and honorific expressions of the text information according to the additional information.

[0768] (Note 3) According to the information processing system described in Appendix 1, the system is characterized in that it extracts motion feature information based on the time change between multiple frames from the image information acquired by the camera device, and changes the content of the prompt statement according to the motion feature information, thereby generating multiple intermediate text information corresponding to continuous body movements.

[0769] Application Example 2 (Note 1) An information processing system, characterized in that it comprises: A device for acquiring visual information, including sign language gestures and facial expressions, through an imaging device, and for acquiring the acquired visual information as time-series image information. An apparatus for extracting hand and body regions from the image information using image processing computing resources, generating time-series features of sign language actions using feature extraction computing resources, and generating initial text information representing sign language content based on the time-series features. An apparatus for generating prompt statements for instructing natural language output to a generative artificial intelligence model based on the initial text information and dialogue scenario information, and inputting the prompt statements into the generative artificial intelligence model to generate corrected text information representing sign language content; An apparatus for inferring emotional information from the facial expressions and acoustic information acquired by the audio input device using emotion recognition computing resources, and for correcting the content or manner of expression of the corrected text information based on the emotional information; An apparatus for converting the corrected text information into speech data by using speech synthesis computing resources based on the corrected text information and the emotional information, while setting acoustic parameters including pitch, speech rate and volume. A means for transmitting the voice data to a terminal device via a communication channel and for causing an audio output device connected to the terminal device to output the voice data.

[0770] (Note 2) The information processing system according to Appendix 1 is characterized in that, The apparatus for generating prompt statements is configured to attach linguistic conditions, including politeness, naturalness, and scene adaptability, to the initial text information to form the prompt statements, and to obtain multiple candidate text information output by the generative artificial intelligence model, and to select the corrected text information from the multiple candidate text information based on the linguistic conditions and the sentiment information.

[0771] (Note 3) The information processing system according to Appendix 1 is characterized in that, The device for acquiring image information is configured to receive multiple image frames acquired at a predetermined frame rate by an imaging device mounted on a mobile terminal device through a communication channel, and to generate intermediate feature information, including hand position information and skeletal information, from the multiple image frames using the feature extraction computing resources, and to use the intermediate feature information as the time series feature quantity to generate initial text information representing the sign language content.

Claims

1. An information processing system, characterized in that, include: processor; The processor is configured to: capture sign language through a camera and analyze the captured image data to identify the content of the sign language; The processor is configured to: generate prompting information to instruct a generative artificial intelligence model to generate text data in order to convert the recognized sign language content into text data, and generate the text data based on the prompting information; The processor is configured to convert the generated text data into speech data using speech synthesis technology, and output the speech data through a speech output device.

2. The information processing system according to claim 1, characterized in that, The processor is configured to: identify a user's emotions by parsing the user's facial expressions and / or the tone of the user's voice, and generate the voice data with a tone of voice corresponding to the identified emotions.

3. The information processing system according to claim 1, characterized in that, The processor is configured to capture sign language through a camera and analyze the captured image data frame by frame to recognize sign language movements with high accuracy.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A