Display device, control method of display device and sign language interaction method

By identifying sign language content and environment information in the video, and fusion of information is carried out to identify interaction intentions, the problem of low sign language interaction performance in the prior art is solved, and more efficient user interaction and more environmentally-friendly recommended information generation is achieved.

CN120164466APending Publication Date: 2025-06-17HISENSE VISUAL TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510130793.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-05
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

In the prior art, there is a problem of low interaction performance for the mode of sign language recognition interaction, which is difficult to meet the needs of users, resulting in poor interaction performance.

Method used

By identifying the sign language content and environmental information of the collected video, corresponding text is generated, and information fusion is carried out to identify the interaction intention, and recommendation information is generated for users.

Benefits of technology

The sign language interaction performance is improved, and the recommended information generated can better match the user's needs and the interactive needs of the environment, improving the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120164466A_ABST
    Figure CN120164466A_ABST
Patent Text Reader

Abstract

The invention relates to a display device, a control method of the display device and a sign language interaction method. The display device in one embodiment comprises a display; and at least one processor configured to execute the instructions to enable the display device to: identify sign language content of the captured video, to generate an identified first text; environment information in the video is recognized, and a second text containing environment description information is generated; performing information fusion on the first text and the second text to obtain a fused text; and identifying an interaction intention of the fused text, obtaining an interaction intention identification result, determining whether an interaction intention exists or not based on the interaction intention identification result, generating recommendation information for the fused text if the interaction intention exists, and broadcasting the recommendation information. According to the scheme, the sign language interaction performance is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of sign language-based human-computer interaction, and particularly to a display device, a control method of the display device, and a sign language interaction method. Background Art

[0002] With the increasing development of display technology, display devices such as smart TVs can not only receive instructions from a remote control to complete interactions with users, but also other interaction methods with users have emerged, such as realizing voice interaction by recognizing speech, interacting by performing gesture recognition, and so on. Sign language, as a communication and language tool between deaf people and between deaf and non-deaf people, in order to meet the need to interact with deaf people, display devices have also emerged with the function of being able to recognize sign language to enable sign language interaction with deaf people.

[0003] However, in the related art, there is a problem of low interaction performance in the way of performing sign language recognition and interaction. Summary of the Invention

[0004] This application provides a display device, a control method of the display device, and a sign language interaction method to improve the interaction performance when performing human-computer interaction based on sign language.

[0005] In a first aspect, some embodiments provide a display device, the display device includes:

[0006] A display;

[0007] And at least one processor, configured to execute instructions to cause the display device to:

[0008] Recognize the sign language content of the collected video and generate a first recognized text;

[0009] Recognize the environmental information in the video and generate a second text including environmental description information;

[0010] Perform information fusion on the first text and the second text to obtain a fused text;

[0011] Recognize the interaction intention of the fused text, obtain an interaction intention recognition result, and determine whether there is an interaction intention based on the interaction intention recognition result. If there is an interaction intention, generate recommended information for the fused text and broadcast the recommended information.

[0012] In a second aspect, some embodiments provide a control method of a display device, the control method of the display device includes:

[0013] Recognize the sign language content of the collected video and generate a first recognized text;

[0014] Identify the environmental information in the video and generate a second text containing environmental description information;

[0015] Perform information fusion on the first text and the second text to obtain a fused text;

[0016] Identify the interaction intention of the fused text, obtain an interaction intention recognition result, and determine whether there is an interaction intention based on the interaction intention recognition result. If there is an interaction intention, generate recommendation information for the fused text and broadcast the recommendation information.

[0017] In a third aspect, some embodiments of the present application provide a sign language interaction method, wherein the method includes:

[0018] Identify the sign language content of the collected video and generate a recognized first text;

[0019] Identify the environmental information in the video and generate a second text containing environmental description information;

[0020] Perform information fusion on the first text and the second text to obtain a fused text;

[0021] Identify the interaction intention of the fused text, obtain an interaction intention recognition result, and determine whether there is an interaction intention based on the interaction intention recognition result. If there is an interaction intention, generate recommendation information for the fused text and broadcast the recommendation information.

[0022] Based on the display device according to the embodiments of the present application, while performing sign language content recognition on the collected video to obtain a recognized first text, it also identifies the environmental information in the video to obtain a second text of environmental description information, and performs information fusion on the first text and the second text to obtain a fused text, and performs recognition of interaction intention and generation of recommendation information for the fused text. That is, when identifying whether there is an interaction intention, it is combined with the environmental description information, so that it can identify whether there is an interaction intention in combination with the environmental information, and when generating recommendation information provided to the user in the case of an interaction intention, it is also combined with the environmental description information. That is, the generated recommendation information can not only match the needs of the sign language content, but also match the environment where the sign language content is obtained, so that the generated recommendation information can better meet the interaction needs of the environment, improve the sign language interaction performance, and on this basis, also improve the user experience. Description of the Drawings

[0023] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0024] Figure 1 Schematic diagram of an operation scenario between a display device and a control device provided by some embodiments of the present application;

[0025] Figure 2 Schematic diagram of the hardware configuration of a display device provided by some embodiments of the present application;

[0026] Figure 3 Schematic diagram of the hardware configuration of a control device provided by some embodiments of the present application;

[0027] Figure 4 Schematic diagram of the software configuration of a display device provided by some embodiments of the present application;

[0028] Figure 5 Schematic diagram of an interaction scenario of a display device provided by some embodiments of the present application;

[0029] Figure 6 Schematic flowchart of the sign language interaction process of a display device provided by some embodiments of the present application;

[0030] Figure 7 Schematic diagram of the processing timing of related models and processes in the sign language interaction of a display device provided by some embodiments of the present application;

[0031] Figure 8 Schematic flowchart of the control method of a display device in some embodiments of the present application. Detailed implementation manners

[0032] The following will explain the embodiments in detail, and the examples are shown in the drawings. When the following description involves the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following embodiments do not represent all the implementation manners consistent with the present application. They are only examples of the systems and methods consistent with some aspects of the present application detailed in the claims.

[0033] It should be noted that the brief description of the terms in the present application is only for the convenience of understanding the following described implementation manners, rather than intending to limit the implementation manners of the present application. Unless otherwise specified, these terms should be understood in their ordinary and common meanings.

[0034] In this application, terms such as "first", "second", "third", etc. in the specification, claims and the above-mentioned drawings are used to distinguish similar or like objects or entities, and do not necessarily mean to limit a specific order or sequence, unless otherwise noted. It should be understood that such terms used can be interchanged under appropriate circumstances.

[0035] The terms "comprising" and "having" and any variations thereof are intended to cover but not exclude inclusion. For example, a product or device comprising a series of components does not necessarily have to be limited to all the components clearly listed, but may include other components not clearly listed or inherent to these products or devices.

[0036] The term "module" refers to any known or later-developed hardware, software, firmware, artificial intelligence, fuzzy logic, or a combination of hardware or / and software code that can perform functions related to that element.

[0037] In the embodiments of this application, the display device 200 generally refers to a device with the capabilities of displaying images and processing data. For example, the display device 200 includes but is not limited to smart TVs, mobile terminals, computers, monitors, advertising screens, wearable devices, virtual reality devices, augmented reality devices, etc. In the following related embodiments of this application, the display device 200 is taken as a smart TV as an example for illustration.

[0038] Figure 1 It is a schematic diagram of the operation scenario between the display device and the control device provided for some embodiments of this application. As Figure 1 shown, the user can operate the display device 200 through touch operations, the mobile terminal 300, and the control device 100. For example, the control device 100 can be a remote control, a stylus, a gamepad, etc.

[0039] The mobile terminal 300 can be used as a control device for performing human-computer interaction between the user and the display device 200. The mobile terminal 300 can also be used as a communication device for establishing a communication connection with the display device 200 to perform data interaction. In some embodiments, software applications can be installed on the mobile terminal 300 and the display device 200, and they can be connected and communicate through network communication protocols to achieve the purpose of one-to-one control operations and data communication. It is also possible to transmit the audio and video content displayed on the mobile terminal 300 to the display device 200 to achieve the synchronous display function.

[0040] As Figure 1 also shown, the display device 200 also communicates with the server 400 through various communication methods. The display device 200 is allowed to establish a communication connection through a local area network (LAN), a wireless local area network (WLAN), and other networks.

[0041] The display device 200 can provide a broadcast receiving TV function, and can additionally provide an intelligent network TV function with computer support functions, including but not limited to, Internet TV, smart TV, Internet Protocol TV (IPTV), etc.

[0042] Figure 2 For some embodiments of the present application Figure 1 The hardware configuration block diagram of the display device 200 in

[0043] In some embodiments, the display device 200 may include at least one of a tuner demodulator 210, a communication device 220, a detector 230, a device interface 240, a controller 250, a display 260, an audio output device 270, a memory, a power supply, and a user input interface.

[0044] In some embodiments, the detector 230 is used to collect signals from the external environment or interact with the outside. For example, the detector 230 includes a light receiver ( Figure 2 not shown in the figure), a sensor for collecting the ambient light intensity; or, the detector 230 includes an image collector, such as a camera, which can be used to collect external environmental scenes, user attributes, or user interaction gestures. Or, the detector 230 includes a sound collector, such as a microphone, etc., for receiving external sounds.

[0045] In some embodiments, the display 260 includes a display function component for presenting a picture and a driving component for driving image display. The display 260 is used to receive the image signal output from the controller 250 for display. For example, the display 260 can be used to display video content, image content, components of a menu manipulation interface, and a user manipulation UI interface, etc.

[0046] In some embodiments, the communication device 220 is a component for communicating with external devices or the server 400 according to various communication protocol types. The display device 200 may be provided with a plurality of communication devices 220 according to different supported communication methods. For example, when the display device 200 supports wireless network communication, the display device 200 may be provided with a communication device 220 including a WiFi function. When the display device 200 supports Bluetooth connection communication, the display device 200 needs to be provided with a communication device 220 including a Bluetooth function.

[0047] The communication device 220 can enable the display device 200 to communicate with an external device or server 400 through a wireless or wired connection. Among them, the wired connection can connect the display device 200 to the external device through components such as data lines and interfaces. The wireless connection can connect the display device 200 to the external device through wireless signals or wireless networks. The display device 200 can directly establish a connection relationship with the external device or indirectly establish a connection relationship through a gateway, router, connection device, etc.

[0048] In some embodiments, the controller 250 may include at least one of a central processing unit, a video processor, an audio processor, a graphics processor, and a power processor, and first to n interfaces for input / output. The controller 250 controls the operation of the display device and responds to user operations through various software control programs stored in the memory. The controller 250 controls the overall operation of the display device 200.

[0049] In some embodiments, the controller 250 and the tuner demodulator 210 may be located in different split devices, that is, the tuner demodulator 210 may also be in an external device of the main device where the controller 250 is located, such as an external set-top box, etc.

[0050] In some embodiments, if a user inputs a user command on the graphical user interface (GUI) displayed on the display 260, the user input interface receives the user input command through the graphical user interface (GUI).

[0051] In some embodiments, the audio output device 270 may be a built-in speaker of the display device 200 or an external audio output device connected to the display device 200. Among them, for the external audio output device connected to the display device 200, the display device 200 may also be provided with an external audio output terminal, and the audio output device can be connected to the display device 200 through the external audio output terminal to output the sound of the display device 200.

[0052] In some embodiments, the user input interface 280 can be used to receive instructions from user input.

[0053] Figure 3 For some embodiments of this application Figure 1 The hardware configuration block diagram of the control device. As Figure 3 shown, the control device 100 may include: a controller 110, a communication interface 130, a user input / output interface, a memory, and a power supply.

[0054] The control device 100 is configured to control the display device 200, receive input operation instructions from the user, and convert the operation instructions into instructions recognizable and responsive by the display device 200, serving as an interaction intermediary between the user and the display device 200.

[0055] In some embodiments, the control device 100 can be an intelligent device. For example, the control device 100 can install various applications for controlling the display device 200 according to user needs.

[0056] In some embodiments, as Figure 1 shown, after installing the application for controlling the display device 200, the mobile terminal 300 or other intelligent electronic devices can perform functions similar to those of the control device 100.

[0057] The controller 110 includes a processor 112, a RAM 113, a ROM 114, a communication interface 130, and a communication bus. The controller 110 is used to control the operation and operation of the control device 100, as well as the communication and cooperation between internal components and the data processing functions between the external and the internal.

[0058] Under the control of the controller 110, the communication interface 130 realizes the communication of control signals and data signals with the display device 200. The communication interface 130 can include at least one of a WiFi chip 131, a Bluetooth module 132, an NFC module 133, and other near-field communication modules.

[0059] In some embodiments, the control device 100 includes at least one of a communication interface 130 and an input / output interface 140. When the communication interface 130 is configured in the control device 100, such as modules like WiFi, Bluetooth, and NFC, the user input instructions can be encoded through the WiFi protocol, or the Bluetooth protocol, or the NFC protocol and sent to the display device 200. Among the user input / output interfaces 140, the input interface includes at least one of a microphone 141, a touchpad 142, a sensor 143, a button 144, and other input interfaces.

[0060] The memory 190 is used to store various operating programs, data, and applications for driving and controlling the control device 100 under the control of the controller. The memory 190 can store various control signal instructions input by the user.

[0061] The power supply 180 is used to provide operating power support for each component of the control device 100 under the control of the controller.

[0062] In order to perform user interaction, in some embodiments, the display device 200 may run an operating system. The operating system is a computer program for managing and controlling hardware resources and software resources in the display device 200. The operating system may provide a user interface (control the display device), allow the user to interact with the display device 200, and support the running of various application programs.

[0063] It should be noted that the operating system may be a native operating system based on a specific operating platform, or a third-party operating system deeply customized based on a specific operating platform, or an independent operating system specially developed for the display device.

[0064] The operating system can be divided into different modules or layers according to the functions implemented, such as Figure 4 As shown, in some embodiments, the system is divided into four layers, from top to bottom, namely, the application layer (Applications) layer (referred to as "application layer"), the application framework layer (Application Framework) layer (referred to as "framework layer"), the system library layer and the kernel layer.

[0065] In some embodiments, the application layer is used to provide services and interfaces for applications so that the display device 200 can run applications and interact with users based on the applications. At least one application can be run in the application layer, and these applications can be window programs, system settings programs, clock programs, etc. that come with the operating system; they can also be applications developed by third-party developers. In specific implementations, the application packages in the application layer are not limited to the above examples.

[0066] The framework layer provides application programming interfaces (APIs) and programming frameworks for applications. The application framework layer includes some predefined functions. The application framework layer is equivalent to a processing center that determines the actions that applications in the application layer take. Through the API interface, applications can access system resources and obtain system services during execution.

[0067] like Figure 4As shown in the figure, in the embodiment of the present application, the application framework layer includes a view system, managers, content providers, etc. Among them, the view system can design and implement the interface and interaction of the application. The view system includes lists, grids, text boxes, buttons, etc. The managers include at least one of the following modules: The Activity Manager is used to interact with all the activities running in the system; the Location Manager is used to provide access to the system location service for system services or applications; the Package Manager is used to retrieve various information related to the application packages currently installed on the device; the Notification Manager is used to control the display and clearing of notification messages; the Window Manager is used to manage icons, windows, toolbars, wallpapers, and desktop widgets on the user interface.

[0068] In some embodiments, the Activity Manager is used to manage the life cycles of various applications and the general navigation back function, such as controlling the exit, opening, and backward movement of the application. The Window Manager is used to manage all window programs, such as obtaining the display screen size, determining whether there is a status bar, locking the screen, taking screenshots, and controlling the display window changes. For example, shrinking the display window, jittering the display, and distorting the display.

[0069] In some embodiments, the system runtime library layer can provide support for the framework layer. When the framework layer is used, the operating system will run the instruction libraries included in the system runtime library layer, such as C / C++ instruction libraries, to implement the functions that the framework layer needs to achieve.

[0070] In some embodiments, the kernel layer is a functional level between the hardware and software of the display device 200. The kernel layer can implement functions such as hardware abstraction, multitasking, and memory management. For example, as Figure 4 shown, hardware drivers can be configured in the kernel layer. The drivers included in the kernel layer can be at least one of the following drivers: audio driver, display driver, Bluetooth driver, camera driver, WIFI driver, USB driver, HDMI driver, sensor drivers (such as fingerprint sensors, temperature sensors, pressure sensors, etc.), and power drivers, etc.

[0071] It should be noted that the above examples are only simple classifications of the functions of the operating system, and do not limit the specific form of the operating system of the display device 200 in the embodiments of the present application. According to factors such as the functions of the display device and the type of the operating system, the number of levels and the specific level types included in the operating system may be in other forms.

[0072] With the development of technology, human-computer interaction has expanded from simple command answering or operations to the stage of free and open question answering. Taking the home scene as an example, users can use the human-computer interaction entrance (mostly voice) of the terminal (such as a TV) to realize functions such as hardware control, knowledge query, and flight reservation, which greatly facilitates the use of users and improves the efficiency and fun of life. However, for people with hearing impairments or language impairments (such as deaf-mutes), due to their natural physical defects, they cannot use the currently widely popular voice interaction method.

[0073] Sign language is a natural language expression that relies on gesture movements, facial expressions, and body postures. It mainly expresses corresponding semantics through changes in hand shapes, hand orientations, the position of the hand relative to the body, and the movement trajectory of the hand. Some sign languages also need to be supplemented by body postures and facial expressions to jointly express the meaning of sign language. Since sign language can freely express one's own views and intentions through gestures, body, and facial expressions, and the content expressed is not limited by a limited number of gestures, it is an important means for people with hearing impairments to achieve daily communication, and thus has the attributes of a natural language. Therefore, there has emerged a technology for recognizing sign language by collecting sign language pictures or videos. When interacting with a display device or other devices for sign language recognition, it is usually based on recognizing sign language from a video stream and processing the results of sign language recognition.

[0074] In the related art, when performing sign language recognition interaction, usually after recognizing sign language from a video stream to obtain the corresponding recognition text, if the recognition text includes a wake-up word, the display device is awakened and the control corresponding to the recognition text is executed.

[0075] However, in this sign language interaction process, in the case where semantic understanding is required for sign language interaction, the information feedback for the recognition text of sign language recognition usually hardly meets the needs of users, resulting in low interaction performance. In order to improve the interaction performance, it may be necessary to interact with the user multiple times before the feedback information can meet the user's needs, or when the user is expressing sign language, more information needs to be expressed in order to improve the accuracy of the feedback information, but this also places higher requirements on the user's sign language expression.

[0076] After research, it is found that in the current sign language interaction scenario, interaction is only carried out for the content expressed by the sign language itself in the video stream. During the process of sign language expression by users, personalized information of users is usually not expressed in sign language, which makes it difficult for the feedback information based on this to meet the needs of users. Taking the sign language expression text of "Recommend a movie" as an example, there may be significant differences in the types of movies that a male in his fifties and a female in her twenties hope to watch. There may also be differences in the types of movies that users expect to watch in outdoor scenarios and indoor scenarios with dim lighting. Therefore, when performing sign language recognition, the environmental information at the same time can be recognized and described, and the sign language recognition text obtained by sign language recognition and the environmental description information obtained by recognizing the environmental information are combined to comprehensively determine whether there is an interaction intention, and on this basis, determine the recommended information for users to achieve natural sign language dialogue interaction with users.

[0077] Accordingly, the display device in some embodiments of this application includes:

[0078] A display;

[0079] And at least one processor, configured to execute instructions to enable the display device to:

[0080] Recognize the sign language content of the collected video and generate a first recognized text;

[0081] Recognize the environmental information in the video and generate a second text containing environmental description information;

[0082] Perform information fusion on the first text and the second text to obtain a fused text;

[0083] Recognize the interaction intention of the fused text, obtain an interaction intention recognition result, and based on the interaction intention recognition result, determine whether there is an interaction intention. If there is an interaction intention, generate recommended information for the fused text and broadcast the recommended information.

[0084] Among them, recognizing the sign language content of the collected video means obtaining the content expressed by the user in the video through sign language by recognizing the changes in the hand shape, the orientation of the hand, the position of the hand relative to the body, and the movement trajectory of the hand of the user in the video. In some cases, the body posture and facial expression of the user in the video can also be recognized. The method for recognizing the sign language content of the video is not limited. In some embodiments, the sign language content in the video can be recognized through a sign language recognition model. Among them, the sign language recognition model can be any model that can recognize sign language content, and the type of the model is not limited. For example, it can be a neural network structure such as a 2D or 3D CNN (Convolutional Neural Networks), or a network structure such as a transformer network, but not limited thereto. In the embodiments of the present application, the model for recognizing the sign language content in the video is referred to as the first video recognition model.

[0085] Among them, the sign language recognition model (the first video recognition model) can be obtained through training. After the first video recognition model is trained by a server or other device, the trained first video recognition model can be deployed to a display device or a server for sign language content recognition.

[0086] The method for obtaining the first video recognition model through training is not limited. In some embodiments, the method for obtaining the first video recognition model through training can be as described below.

[0087] First, obtain a sample data set. Each piece of sample data in the sample data set includes a video and a label corresponding to the video, where the label is the sign language content corresponding to the video.

[0088] After obtaining the sample data set, divide the sample data set into a training set, a validation set, and a test set. For example, 70% of the sample data forms the training set, 15% of the sample data forms the validation set, and the remaining 15% of the sample data forms the test set to evaluate the generalization ability of the model.

[0089] After partitioning the sample data set, the model training process can be carried out. During the model training process, the first video recognition model is trained using the training set, and during or after the training process, the validation set is used to evaluate the performance of the first video recognition model, and the test set is used to finally test the trained first video recognition model to ensure its performance in actual applications. It can be understood that during the training of the first video recognition model, the first video recognition model performs sign language content recognition on the sample data video to obtain the recognized sign language content, and compares this sign language content with the sign language content annotated in the sample data to calculate the training loss, and adjusts the first video recognition model in combination with the calculated training loss to achieve the update of the first video recognition model. After the model training is completed, the first video recognition model updated for the last time is used as the first video recognition model obtained by training.

[0090] Identifying the environmental information in the video means identifying the information related to the environment where the user is located and the user's attributes. In some examples, the information about the environment where the user is located can be, for example, one or more of the number of people, single-person environment, multi-person environment, living room environment, outdoor environment, shopping mall environment, in-vehicle environment, dim lighting, bright environment, indoor environment shooting, indoor decoration color, other items included in the surroundings, etc., but not limited to this. The information related to the user attributes in some examples can be, for example, one or more of male, female, age group, user emotion, etc., but not limited to this.

[0091] The method of identifying the environmental information in the video is not limited. The environmental information in the video can be identified by a video recognition model (referred to as the second video recognition model in the embodiments of the present application). Among them, the second video recognition model can be obtained through training. After the second video recognition model is trained by a server or other device, the trained second video recognition model is deployed to a display device or a server to identify the environmental information in the video. The specific type of the second video recognition model is not limited. For example, it can be the LLAVA (Large Language and Vision Assistant) model.

[0092] The second video recognition model can be a large model based on the Transformer architecture, which can understand videos and pictures, and has the ability to generate text descriptions and intention speculation, or achieve the ability of intention speculation through fine-tuning. It can tokenize videos (traditional deep neural networks such as CNNs can be used), and also tokenize text, then perform token alignment operations and input them into the Transformer architecture to achieve the function of processing multiple modal signals simultaneously. Among them, Tokens are words, character sets, or segments composed of combinations of words and punctuation marks, which are used by large language models (LLMs) to decompose text. After tokenization, analyze the semantic relationships between tokens, such as the frequency of their co-occurrence or whether they are used in similar contexts. After training, generate an output token sequence based on the token relationship pattern in the input sequence.

[0093] There is no limit to the method of training the second video recognition model. In some embodiments, the method of training the second video recognition model can be as described below.

[0094] First, obtain a sample data set. Each sample data in the sample data set includes a video and the label corresponding to the video. Among them, the label is the environmental information corresponding to the video. It can be understood that the environmental information in the video can include multiple pieces of information. For example, the environmental information included in a certain sample data can include: living room, 1 person, male, about forty years old, angry, etc. After obtaining the sample data set, divide the sample data set into a training set, a validation set, and a test set. For example, 70% of the sample data forms the training set, 15% of the sample data forms the validation set, and the remaining 15% of the sample data forms the test set to evaluate the generalization ability of the model.

[0095] After dividing the sample data set, the model training process can be carried out. During the model training process, use the training set to train the second video recognition model, and during or after the training process, use the validation set to evaluate the performance of the second video recognition model, and use the test set to finally test the trained second video recognition model to ensure the performance of the model in actual applications. It can be understood that during the training of the second video recognition model, the second video recognition model performs sign language content recognition on the sample data video to obtain the recognized environmental information, and compares this environmental information with the environmental information labeled in the sample data to calculate the training loss, and adjusts the second video recognition model in combination with the calculated training loss to achieve the update of the second video recognition model. After the model training is completed, use the second video recognition model updated last time as the trained second video recognition model.

[0096] Fusing the information of the first text and the second text means fusing the text of the sign language content expressed by the user with the text of the recognized environmental information to obtain a text containing environmental description information. The method of fusing the first text and the second text is not limited. In some examples, a text processing model can be used to fuse the first text and the second text to obtain the fused text. Taking a specific example, if the first text is "movie, nice" or "recommend nice movies", and the second text containing environmental description information is "There is a man at home (in the living room), about 40 years old, and the surrounding lights are dim", then the fused text obtained after fusing the first text and the second text can be: "I am in a home scene. Please recommend nice movies that a 40-year-old man likes to watch."

[0097] By fusing the information of the first text and the second text, the limitation of a finite number of sign language instructions can be broken through, extended to a natural language expression with semantic generalization, making the interaction more natural and closer to the daily usage method, and greatly reducing the barriers for deaf-mute people and normal people in using the interaction system step by step.

[0098] Fusing the information of the first text and the second text can be to combine the text content in the first text and the text content in the second text into one text for representation, and operations such as semantic / grammatical / text correction and sentence polishing (smoothing) can be performed to further solve problems such as poor recognition accuracy of isolated words or continuous sentences caused by recognition errors in sign language recognition, which helps to further improve the success rate of interaction.

[0099] Since there may be errors in sign language recognition, the received text may be a single word, or several related words, or several words with most related and few unrelated (caused by recognition errors), or it may be a relatively vague description or a sentence that does not conform to grammar. Therefore, in the process of fusing the first text and the second text, language understanding technology can be used to complement or polish the information of these scattered, unsmooth, and unclear-intention words or sentences to make them a complete and semantically clear sentence.

[0100] And because the first text obtained by sign language recognition and the second text containing environmental description information are fused, and sentence polishing is performed during the process of fusing the second text and the second text, the accuracy of the meaning expressed by the sign language user in the current environment is improved, which also helps to further improve the success rate of interaction.

[0101] Identifying the interaction intention of the fused text means, based on the fused text, whether the user has an intention to interact with the display device, or in other words, whether the user needs to provide feedback on the information content. In some cases, the user issues specific control instructions through sign language, such as power on, power off, etc. In some cases, the user expresses the need for the display device to provide certain feedback through sign language, such as recommending movies as described above. Therefore, based on the identification of the interaction intention, it can be determined whether interactive control based on the interaction is required.

[0102] After obtaining the result of the interaction intention identification, if there is an interaction intention, recommendation information can be generated for the fused text to provide feedback to the user. The method of generating recommendation information for the fused text is not limited. For example, the fused text can be processed by a semantic understanding system to generate corresponding recommendation information. Among them, the semantic understanding system can adopt any system that can generate corresponding recommendation information for the input text, and the embodiments of the present application do not make specific limitations in this regard. Among them, in some specific examples, when the environmental information includes the user's emotion, the generated recommendation information can include recommendation information corresponding to the user's emotion. For example, if the user's emotion recognized from the video is angry, the generated recommendation information can be soothing recommendation information.

[0103] The method of broadcasting the recommendation information is not limited. For example, it can be broadcast through text display or sign language display. In some cases, it can also be broadcast through voice at the same time.

[0104] Based on the display device according to the embodiments of the present application, while recognizing the sign language content of the collected video to obtain the first recognized text, it also recognizes the environmental information in the video to obtain the second text of the environmental description information, and fuses the first text and the second text to obtain the fused text, and performs the identification of the interaction intention and the generation of the recommendation information for the fused text. That is, when identifying whether there is an interaction intention, it is combined with the environmental description information, so that it can identify whether there is an interaction intention in combination with the environmental information. And when generating the recommendation information provided to the user in the case of an interaction intention, it is also combined with the environmental description information. That is, the generated recommendation information can not only match the needs of the sign language content, but also match the environment where the sign language content is obtained. Therefore, the generated recommendation information can better meet the interaction needs of the environment, improve the interaction performance, and on this basis, improve the user efficiency. And it can achieve barrier-free sign language interaction, more accurately meet the user's needs in a specific scenario, rather than just a simple instruction, which can improve the accuracy of matching the user's needs.

[0105] In some embodiments, the processor is further configured to execute instructions to cause the display device to:

[0106] Identify the environmental information and sign language content in the video through the second video recognition model, generate a second text containing the environmental description information and sign language content, and obtain the second text probability value of the second text.

[0107] Combined with the specific examples described above, if the first text is "movie, nice" or "recommend nice movies", the second text containing environmental description information and sign language content is "There is a man at home (in the living room), about 40 years old, the surrounding lights are dim, and he is making sign language gestures. The meaning of the sign language is: want to watch nice movies", then the fused text obtained after fusing the first text and the second text can be: "I am in a home scene, please recommend nice movies that a 40-year-old man likes to watch".

[0108] Based on this embodiment, while the second video recognition model identifies the environmental information in the video, it also identifies the sign language content in the video. Thus, when fusing the first text and the second text, it can supplement the sign language content identified by the first video recognition model based on the sign language content identified by the second video recognition model, which helps to further improve the accuracy of the obtained fused text. It can be understood that in this case, during the process of training the second video recognition model, the data labels of the sample data also include the sign language content in the video.

[0109] In some embodiments, the processor is further configured to execute instructions to cause the display device to:

[0110] Identify the sign language content of the video through the first video recognition model, generate the first text, and obtain the first text probability value of the first text;

[0111] Identify the environmental information in the video through the second video recognition model, generate a second text containing the environmental description information, and obtain the second text probability value of the second text;

[0112] Based on the first text probability value and the second text probability value, determine the corresponding target text processing model from the text processing model set;

[0113] Perform information fusion on the first text and the second text through the target text processing model to obtain the fused text.

[0114] The first text probability value is the probability value of the first text obtained by the first video recognition model identifying the sign language content of the video. It represents the probability of obtaining this first text for the input video, which is usually between 0 and 1. The first text probability value can identify the credibility of obtaining the first text for the input video.

[0115] The second text probability value can identify the credibility of obtaining the second text for the input video. Since the second text output by the second video recognition model may be a text description, and this text description can be formed based on the words obtained from understanding the video, each word in the text description will have a corresponding probability. At this time, the second text probability value can be the maximum value, minimum value, weighted average value of the probabilities of each word in the text description, or a probability value calculated according to other rules. The embodiments of the present application do not make specific display in this regard.

[0116] Among them, the number of model parameters of each text processing model in the text processing model set is different, and the text processing models can all be obtained through model training. Each text processing model can perform text correction and polishing on one or more input texts to form a sentence that includes environmental information and user requirements, is grammatically smooth, and has clear semantics. The polishing of the text refers to processing such as replacing words, adjusting sentence structures, adding quotations, etc. to improve the quality of the text and perfect the expression of the text, but is not limited thereto.

[0117] The number of model parameters of a model refers to the number of variable parameters used for learning and storing knowledge in the model. Each parameter corresponds to a weight or bias term in the model network. The number of model parameters of a model characterizes the learning ability of the model, and is also related to the size of the model, the complexity and details that the model can capture from the data. Generally speaking, the larger the number of model parameters, the stronger the expression ability of the model, and theoretically it can capture more complex patterns and more detailed information, but the model is also larger, and the processing efficiency will also be affected. The smaller the number of model parameters, the weaker the expression ability of the model, but the model is also smaller, and the processing efficiency is higher.

[0118] Based on this embodiment, when recognizing the sign language content in the video, it is recognized through the first video recognition model, and the first text probability value of the obtained first text is obtained. And it is through the second video recognition model to recognize the environmental information of the video, and the second text probability value of the obtained second text is obtained, and from the text processing model set, the target text processing model matching the first text probability value and the second text probability value is determined. Thus, the credibility of recognizing the sign language content and the credibility of recognizing the environmental information of the video can be combined, and the target text processing model matching them is selected to perform information fusion on the first text and the second text, so that the credibility of the obtained fused text can be higher, and better information fusion performance can be achieved.

[0119] Among them, the method of determining the corresponding target text processing model from the text processing model set based on the first text probability value and the second text probability value is not limited. Some feasible methods are exemplified below.

[0120] In some embodiments, the set of text processing models includes a first text processing model and a second text processing model, and the number of model parameters of the first text processing model is less than the number of model parameters of the second text processing model. The specific number of model parameters of the first text processing model and the second text processing model is not limited. In some specific examples, the number of model parameters of the first text processing model may be greater than 70B (Billion), and the number of model parameters of the second text processing model may be less than or equal to 7B, but it is not limited thereto.

[0121] At this time, the processor is further configured to execute instructions to cause the display device to:

[0122] If the first text probability value is greater than or equal to the first threshold, determine the first text processing model as the target text processing model,

[0123] If the first text probability value is less than or equal to the second threshold, determine the second text processing model as the target text processing model,

[0124] If the first text probability value is greater than the second threshold and less than the first threshold, determine the corresponding target text processing model from the set of text processing models according to the second text probability value;

[0125] Wherein, the first threshold is greater than the second threshold.

[0126] Wherein, the first threshold and the second threshold are thresholds for distinguishing the credibility of the first text. When the first text probability value is greater than the first threshold, it indicates that the credibility of obtaining the first text based on the collected video is very high. When the first text probability value is less than the second threshold, it indicates that the credibility of obtaining the first text based on the collected video is very low, and the first text is not very credible. The specific setting method of the first threshold and the second threshold is not limited, as long as the feasibility can be well distinguished.

[0127] Wherein, the method of determining the corresponding target text processing model from the set of text processing models according to the second text probability value is not limited. In some embodiments, a corresponding threshold may be set for the second text probability value. If the second text probability value is greater than the threshold, the first text processing model is used as the target text processing model, otherwise the second text processing model is used as the target processing model.

[0128] Based on this embodiment, when the first text probability value is greater than the first threshold, it indicates that the credibility of the first text is relatively high and the result of sign language recognition is relatively accurate. Thus, the first text processing model with fewer model parameters can be directly used as the target text processing model, and then a lightweight text processing model with even fewer model parameters can be used to process the fused text. Therefore, when the accuracy of the first text is relatively high, the efficiency of fusing the first text and the second text can be further improved. When the first text probability value is less than the second threshold, it indicates that the credibility of the first text is very low. Thus, the second text processing model with more model parameters can be directly used as the target text processing model to process the fused text with a text processing model having more model parameters. For example, for polishing and sentence structuring, it helps to correct some recognition errors, ensure the sentences and grammar, and improve the accuracy of the fused text obtained by fusing the first text and the second text, which is beneficial to improving the accuracy of information interaction. When the first text probability value is greater than the second threshold and less than the first threshold, it indicates that the first text has a certain credibility but has not reached a relatively high level. Therefore, the corresponding target text processing model can be determined from the text processing model set according to the second text probability value, so as to be able to determine a text processing model with threshold matching by combining the credibility of the second text.

[0129] Among them, the first text processing model and the second text processing model can both be deployed on the display device side or on the cloud side. Considering that the number of model parameters of the first text processing model is small and it is a lightweight processing model, it can be deployed on the display device side to achieve fast processing speed, low power consumption, low cost, and at the same time, confidentiality can also be achieved. While the number of model parameters of the second text processing model is large, so it can be deployed on the cloud side to

[0130] It can be understood that in the above embodiment, it is described by taking two text processing models included in the text processing model set as an example. In other embodiments, the text processing model set may also include more than two text processing models. Multiple thresholds can be set for the credibility of the first text. When these multiple thresholds are arranged from small to large or from large to small, the interval between any two adjacent thresholds corresponds to one text processing model in the text processing model set. Multiple thresholds can also be set for the credibility of the second text. When these multiple thresholds are arranged from small to large or from large to small, the interval between any two adjacent thresholds corresponds to one text processing model in the text processing model set. And in the order of the thresholds from small to large, the number of model parameters of each text processing model is from large to small. Among them, it can also be that the threshold interval formed by two adjacent thresholds of the credibility of the first text and the threshold interval formed by two adjacent thresholds of the credibility of the second text respectively correspond to one text processing model in the text processing model set, but it is not limited to this.

[0131] In some embodiments, the processor is further configured to execute instructions to cause the display device to:

[0132] Perform a weighted sum of the first text probability value and the second text probability value to obtain a text probability sum value;

[0133] Determine a corresponding target text processing model from a set of text processing models according to a comparison result between the first text probability value and the text probability sum value.

[0134] The way of weighted sum in a specific example can be denoted as: α = β * a1 + (1 - β) * a2, 0 < β < 1. Where α is the text probability sum value, a1 is the first text probability value, a2 is the second text probability value, β is the weighting coefficient of the first text probability value, and 1 - β is the weighting coefficient of the second text probability value.

[0135] Among them, when performing a weighted sum of the first text probability value and the second text probability value, the weighting coefficients of the first text probability value and the second text probability value are not limited. For example, the weighting coefficients of the first text probability value and the second text probability value can both be set to 0.5. In some examples, the weighting coefficient of the first text probability value can be greater than 0.5, such as 0.7, and the weighting coefficient of the second text probability value can be less than 0.5, such as 0.3, but not limited thereto.

[0136] Based on this embodiment, after performing a weighted sum of the first text probability value and the second text probability value, the text probability sum value obtained by the weighted sum is compared with the first text probability value, so that the difference between the credibility of the first text and the comprehensive credibility of the first text and the second text can be obtained, and thus a suitable target text processing model can be selected and determined from the set of text processing models based on this difference.

[0137] Among them, the way of determining a corresponding target text processing model from a set of text processing models according to a comparison result between the first text probability value and the text probability sum value is not limited. Only some of the ways are exemplified below.

[0138] In this example, the set of text processing models includes a first text processing model and a second text processing model, and the number of model parameters of the first text processing model is less than the number of model parameters of the second text processing model.

[0139] The processor is further configured to execute instructions to cause the display device to:

[0140] If the difference between the first text probability value and the text probability sum value is greater than or equal to a preset difference threshold, determine the first text processing model as the target text processing model.

[0141] If the difference between the first text probability value and the text probability sum value is less than a preset difference threshold, determine the second text processing model as the target text processing.

[0142] The preset difference threshold is a threshold used to distinguish the difference degree between the first text probability value of the first text and the text probability sum value. When the difference between the first text probability value and the text probability sum value is greater than or equal to the preset difference threshold, it indicates that the credibility of obtaining the first text based on the collected video is very high. By performing weighted summation of the first text probability value and the second text probability, the first text probability value is greatly reduced. When the difference between the first text probability value and the text probability sum value is less than the preset difference threshold, it indicates that the influence degree of performing weighted summation of the first text probability value and the second text probability on the first text probability value is not obvious.

[0143] Based on this embodiment, when the first text probability value and the text probability sum value are greater than or equal to the preset difference threshold, it indicates that the credibility of the first text is relatively high. Thus, the first text processing model with fewer model parameters can be directly used as the target text processing model, and then a lightweight text processing model with fewer model parameters can be used to process the fused text. Thus, when the accuracy of the first text is relatively high, the efficiency of fusing the first text and the second text can be further improved. When the first text probability value and the text probability sum value are less than the preset difference threshold, it indicates that the credibility of the first text is difficult to form an obvious difference from the text probability sum value. Therefore, a text processing model with more model parameters can be used to process the fused text to improve the accuracy of the fused text obtained by fusing the information of the first text and the second text, which is beneficial to improving the accuracy of information interaction.

[0144] In some embodiments, the processor is further configured to execute instructions to cause the display device to:

[0145] Perform information fusion on the first text, the first text probability value, the second text, and the second text probability value through the target text processing model to obtain the fused text.

[0146] Thus, when performing information fusion on the first text and the second text using the target text processing model, the first text probability value and the second text probability value are combined, that is, information fusion is performed on the two based on considering the credibility of the first text and the credibility of the second text, which helps to improve the feasibility of the obtained fused text, thereby improving the accuracy of the fused text obtained by information fusion. Based on this, the generated recommendation information can not only match the needs of sign language content but also match the environment where the sign language content is obtained. Thus, the generated recommendation information can better meet the interaction needs of the environment, improving the interaction performance and, on this basis, improving the user efficiency.

[0147] Among them, when the target text processing model performs information fusion on the first text, the first text probability value, the second text, and the second text probability value, it can be combined with a prompt. The prompt can be used to clearly indicate that the results of the first text and the second text are comprehensively processed, and then a polished sentence with smooth expression, clear intention, and including the surrounding environment description is output. In the case of using the prompt, the first text, the first text probability value, the second text, the second text probability value, and the prompt can be used as the input of the target text processing model to output a polished sentence with smooth expression, clear intention, and including the surrounding environment description. The content of the prompt in a specific example can include model role information (for example: #Role You are a semantic understanding expert who can understand, synthesize, and process the content input by the user), model task information (for example: #Task You need to fully understand the user's two inputs {a1, x1} and {a2, x2}, and output a sentence with smooth expression, clear intention, and including the surrounding environment description), description information of the model input (for example: #Input Description x1 and a1 are the text and probability values recognized by the traditional sign language recognition expert model (deep learning method); x2 and a2 are the text output and probability values obtained by the video understanding large model (transformer class); the probability value can be understood as the accuracy or confidence of the text output), examples of model input and output (for example: #Example Input: {a1, x1} and {a2, x2}; Output: {xxxxx}), etc., but not limited to this.

[0148] Among them, in some examples, when the target text processing model performs information fusion on the first text and the second text to obtain the fused text, or when the target text processing model performs information fusion on the first text, the first text probability value, the second text, and the second text probability value to obtain the fused text, the stored historical interaction information can also be combined for information fusion, so as to further improve the accuracy of semantic expression and reduce semantic errors.

[0149] In some examples, in the case where the second video recognition module recognizes that there are multiple people and receives the sign language content of multiple people, information fusion can be performed by combining the recognized sign language content of multiple people. For example, the sign language content of multiple people, the environmental description information recognized by the second video recognition model, and the sign language content are repaired and polished to obtain a fused, unified, semantically clear, and grammatically compliant sentence. For example, if the sign language recognition model (the first video recognition model) can recognize and output the sign language content of multiple people, the sign language content of multiple people can be input into the target text processing model for processing. For another example, if the sign language recognition model (the first video recognition model) cannot process the recognition of multiple people at one time, the sign language content of each person can be processed sequentially, and after the recognition, the sign language content of multiple people can be uniformly input into the target text processing model for processing, but not limited to this.

[0150] In some embodiments, the processor is further configured to execute instructions to cause the display device to:

[0151] Recognize the interaction intention of the fused text, obtain an interaction intention recognition result, and obtain a first interaction intention probability value;

[0152] Monitor a wake-up instruction and obtain a second interaction intention probability value of the monitored wake-up instruction;

[0153] Perform a comprehensive calculation on the first interaction intention probability value and the second interaction intention probability value to obtain a comprehensive probability value;

[0154] Based on the comprehensive probability value, determine whether there is an interaction intention.

[0155] The manner of recognizing the interaction intention of the fused text is not limited. In some examples, an intention recognition model can be used to recognize the interaction intention of the fused text to obtain an interaction intention recognition result and a corresponding probability value, which is referred to as the first interaction intention probability value in the embodiments of the present application.

[0156] The wake-up instruction is an instruction used to indicate the wake-up of the display device, and the form of the instruction used to wake up the display device is not limited. For example, it can be an instruction expressed by a gesture or an instruction when a keyword is included in the recognized sign language content.

[0157] The manner of performing a comprehensive calculation on the first interaction intention probability value and the second interaction intention probability value is not limited.

[0158] In some examples, either the first interaction intention probability value or the second interaction intention probability value can indicate an interaction intention. For example, if the first interaction intention probability value is greater than its corresponding first probability threshold, or the second interaction intention probability value is greater than its corresponding second probability threshold, it is determined that the comprehensive probability value is 1, that is, it is considered to have an interaction intention. Taking the case where the output first interaction intention probability value is 0 or 1, and the output second interaction intention probability value is 0 or 1, and 1 indicates an interaction intention and 0 indicates no interaction intention as an example, it can be determined that the comprehensive probability value is 1, that is, it is considered to have an interaction intention when either the first interaction intention probability value or the second interaction intention probability value is 1.

[0159] In other examples, it can be the case where both the first interaction intention probability value and the second interaction intention probability value indicate an interaction intention. For example, if the first interaction intention probability value is greater than its corresponding first probability threshold, and the second interaction intention probability value is greater than its corresponding second probability threshold, it is determined that the comprehensive probability value is 1, that is, it is considered to have an interaction intention. Taking the case where the output first interaction intention probability value is 0 or 1, and the output second interaction intention probability value is 0 or 1, and 1 indicates an interaction intention and 0 indicates no interaction intention as an example, it can be determined that the comprehensive probability value is 1, that is, it is considered to have an interaction intention when both the first interaction intention probability value and the second interaction intention probability value are 1.

[0160] In other examples, it can be to perform a weighted calculation on the first interaction intention probability value and the second interaction intention probability value to obtain a comprehensive probability value. If the comprehensive probability value is greater than the corresponding probability threshold, it can be determined that there is an interaction intention. Taking the comprehensive probability value in a specific example as W = θ1*w1 + θ2*w2, where w1 and w2 are the first interaction intention probability value and the second interaction intention probability value respectively, and θ1 and θ2 are the weighting coefficients of the first interaction intention probability value and the second interaction intention probability value respectively.

[0161] Based on this embodiment, after obtaining the fused text, while obtaining the first interaction intention probability value of whether the fused text has an interaction intention, it also monitors whether a wake-up instruction is received to obtain the second interaction intention probability value of the monitored wake-up instruction. After comprehensively calculating the first interaction intention probability value and the second interaction intention probability value to obtain a comprehensive probability value, the comprehensive probability value is used to finally determine whether there is an interaction intention, enabling the determination of whether there is an interaction intention to be comprehensively made from two perspectives of the fused text and the wake-up instruction, improving the accuracy of the discrimination of the interaction intention.

[0162] Among them, the way of receiving the wake-up instruction is not limited, and several of them will be exemplified below.

[0163] In some embodiments, the processor is further configured to execute instructions to cause the display device to:

[0164] Identify gesture information in the video, determine whether a wake-up instruction is received based on the gesture information, and obtain a second interaction intention probability value of the wake-up instruction.

[0165] Gestures are pre-set fixed patterns, generally a limited number of gestures designed for high-frequency scenarios. By identifying fixed gestures in the video captured by the camera and then making a response to execute corresponding instructions, such as "increase volume", "turn off screen", "increase brightness", etc.

[0166] Gesture information for waking up the display device can be pre-set and stored in the display device or a server communicating with the display device. If the gesture information recognized from the video is the same as the stored gesture information, it can be considered that a wake-up instruction is received.

[0167] Among them, when identifying gesture information from the video and determining whether the recognized gesture information is the same as the stored gesture information, it can be performed through a gesture recognition model. The gesture recognition model can output the probability that the gesture information in the video is the same as the gesture information of the stored wake-up instruction. This probability is referred to as the second interaction intention probability value in the embodiments of the present application.

[0168] Based on this embodiment, by identifying gesture information in the video and determining whether a wake-up instruction is received through the gesture information, the user can trigger the display device in combination with the gesture information.

[0169] In some embodiments, the processor is further configured to execute instructions to cause the display device to:

[0170] The processor is further configured to execute instructions to cause the display device to:

[0171] If the first text includes a wake-up keyword, determine that a wake-up instruction is received, and determine a second interaction intention probability value of the wake-up instruction.

[0172] Keywords for waking up the display device can be pre-set and stored in the display device or a server communicating with the display device. If the recognized sign language content contains the keyword, it can be considered that a wake-up instruction is received.

[0173] Among them, when determining whether the sign language content contains the keyword corresponding to the wake-up instruction, after word segmentation of the sign language content, the similarity between the words in the sign language content and the keyword can be calculated to determine whether the keyword of the wake-up instruction is received. The maximum similarity among the similarities corresponding to these words can be used as the second interaction intention probability value, but it is not limited thereto.

[0174] Based on this embodiment, it is possible to determine whether a wake-up instruction is received by checking whether the first text contains a wake-up keyword, so that the user can express the wake-up keyword through sign language to implement the wake-up instruction.

[0175] In some embodiments, the processor is further configured to execute instructions to cause the display device to:

[0176] Perform interactive intention recognition on the video to obtain a third interactive intention probability value indicating whether the video has an interactive intention;

[0177] Perform comprehensive calculation on the first interactive intention probability value, the second interactive intention probability value, and the third interactive intention probability value to obtain a comprehensive probability value;

[0178] Based on the comprehensive probability value, determine whether there is an interactive intention.

[0179] Among them, performing interactive intention recognition on the video to obtain a third interactive intention probability value indicating whether the video has an interactive intention means directly analyzing the video to determine whether there is an interactive intention. For example, the user makes relatively many interactive actions facing the camera of the display device, or the user makes many interactive actions facing the camera of the display device and shows anxious behaviors or expressions without feedback information from the display device, etc., but not limited to this, and it can be set according to the requirements of different scenarios.

[0180] The method of performing interactive intention recognition on the video is not limited. In some examples, a third video recognition model can be used to perform interactive intention recognition on the video to obtain a recognition result and the corresponding probability value, which is called the third interactive intention probability value in the embodiments of the present application. When the third video recognition model performs interactive intention recognition on the video, other information can also be combined to perform interactive intention recognition at the same time, such as combining video, collected voice, received text, etc., and comprehensively recognizing whether there is an interactive intention, but not limited to this.

[0181] Among them, the method of performing comprehensive calculation on the first interactive intention probability value, the second interactive intention probability value, and the third interactive intention probability value is not limited.

[0182] Perform weighted sum calculation on the first interactive intention probability value, the second interactive intention probability value, and the third interactive intention probability value to obtain a comprehensive probability value. The method of performing comprehensive calculation on the first interactive intention probability value, the second interactive intention probability value, and the third interactive intention probability value is not limited.

[0183] In some examples, any one of the first interaction intention probability value, the second interaction intention probability value, and the third interaction intention probability value can indicate an interaction intention. For example, if the first interaction intention probability value is greater than its corresponding first probability threshold, or the second interaction intention probability value is greater than its corresponding second probability threshold, or the third interaction intention probability value is greater than its corresponding third probability threshold, then it is determined that the comprehensive probability value is 1, that is, it is considered to have an interaction intention. Taking the output first interaction intention probability value being 0 or 1, the output second interaction intention probability value being 0 or 1, the output third interaction intention probability value being 0 or 1, and 1 indicating an interaction intention and 0 indicating no interaction intention as an example, then when any one of the first interaction intention probability value, the second interaction intention probability value, and the third interaction intention probability value is 1, it can be determined that the comprehensive probability value is 1, that is, it is considered to have an interaction intention.

[0184] In other examples, it can be the case where at least two of the first interaction intention probability value, the second interaction intention probability value, and the third interaction intention probability value indicate an interaction intention. For example, if the first interaction intention probability value is greater than its corresponding first probability threshold, the second interaction intention probability value is greater than its corresponding second probability threshold, and the third interaction intention probability value is greater than its corresponding third probability threshold, and at least two of these three conditions are met, then it is determined that the comprehensive probability value is 1, that is, it is considered to have an interaction intention. Taking the output first interaction intention probability value being 0 or 1, the output second interaction intention probability value being 0 or 1, the output third interaction intention probability value being 0 or 1, and 1 indicating an interaction intention and 0 indicating no interaction intention as an example, then when two of the first interaction intention probability value, the second interaction intention probability value, and the third interaction intention probability value are 1, it can be determined that the comprehensive probability value is 1, that is, it is considered to have an interaction intention.

[0185] In other examples, it can be the case where the first interaction intention probability value, the second interaction intention probability value, and the third interaction intention probability value all indicate an interaction intention. For example, if the first interaction intention probability value is greater than its corresponding first probability threshold, the second interaction intention probability value is greater than its corresponding second probability threshold, and the third interaction intention probability value is greater than its corresponding third probability threshold, then it is determined that the comprehensive probability value is 1, that is, it is considered to have an interaction intention. Taking the output first interaction intention probability value being 0 or 1, the output second interaction intention probability value being 0 or 1, the output third interaction intention probability value being 0 or 1, and 1 indicating an interaction intention and 0 indicating no interaction intention as an example, then when the first interaction intention probability value, the second interaction intention probability value, and the third interaction intention probability value are all 1, it can be determined that the comprehensive probability value is 1, that is, it is considered to have an interaction intention.

[0186] In some other examples, it may be to perform a weighted calculation on the first interaction intention probability value, the second interaction intention probability value, and the third interaction intention probability value to obtain a comprehensive probability value. If the comprehensive probability value is greater than the corresponding probability threshold, it can be determined that there is an interaction intention. Taking the comprehensive probability value in a specific example as W = θ1*w1 + θ2*w2 + θ3*w3, where w1, w2, and w3 are the first interaction intention probability value, the second interaction intention probability value, and the third interaction intention probability value respectively, and θ1, θ2, and θ3 are the weighting coefficients of the first interaction intention probability value, the weighting coefficient of the second interaction intention probability value, and the weighting coefficient of the third interaction intention probability value respectively.

[0187] Based on this embodiment, it is also possible to directly perform interaction intention recognition on the video to obtain a third interaction intention probability value indicating whether there is an interaction intention in combination with the video information, so that it is possible to comprehensively determine whether there is an interaction intention from three perspectives: the fused text, the wake-up instruction, and the video image, improving the accuracy of the determination of the interaction intention.

[0188] In some embodiments, the processor is further configured to execute instructions to cause the display device to:

[0189] If there is an interaction intention, enter the wake-up mode to wake up the display device.

[0190] Wherein, in the wake-up mode, the display device is woken up, that is, the display device is in a state of displaying content and can give feedback based on received instructions or information.

[0191] Based on this embodiment, in the case of comprehensively determining that there is an interaction intention, enter the wake-up mode to wake up the display device to improve the accuracy of waking up the display device.

[0192] In some embodiments, the processor is further configured to execute instructions to cause the display device to:

[0193] In the wake-up mode, perform semantic understanding on the fused text to generate recommendation information for the fused text.

[0194] Based on this embodiment, when the display device is woken up and it is determined that there is an interaction intention, it is possible to directly perform semantic understanding on the fused text to generate recommendation information for the fused text for timely interaction feedback.

[0195] In some embodiments, the processor is further configured to execute instructions to cause the display device to:

[0196] If there is an interaction intention, enter the wake-up mode and perform semantic understanding based on the stored historical interaction information and the fused text to generate recommendation information for the fused text.

[0197] Among them, the historical interaction information refers to the information of the interaction between the display device and the user before the current time point. This information may include the first text obtained by the display device through sign language recognition, and may also include the recommended information generated when there is an interaction intention.

[0198] By combining the historical interaction information, the interaction intention can be recognized more accurately and clearly. Or in the case where there are pronouns in the fused text, the specific referents of the pronouns can be known more accurately and clearly. In a specific example, assume that the stored historical interaction information is "It's so hot today. It seems to be around 40 degrees." of User 1, and the current sign language content received from User 2 is: "Yes, I also feel very hot. Please turn on the air conditioner for us!" It can be seen that referring to the storage background of User 1, the speech of User 2 has an obvious interaction intention, which is to turn on the air conditioner; although there is only the input from User 2 and there is also a user intention, but by combining the historical interaction information of User 1, the interaction intention is more obvious.

[0199] For another example, in another example, assume that the stored historical interaction information is User 1's: "I think green apples are more nutritious than red apples" and User 2's: "I don't feel the same way." And the current sign language content received from User 2 is "Let's consult the AI intelligent assistant. Which one is more nutritious!" It can be seen that the pronoun "which" in the last sentence of User 2 has a strong association with the context, and the intention of the word it refers to is more clear when there is context.

[0200] In some embodiments, a semantic understanding model can be used to perform semantic understanding on the stored historical interaction information and the fused text to generate recommended information for the fused text. The specific type of the semantic understanding model is not limited, as long as it can combine the currently received interaction text and the historical interaction text and generate feedback information for the currently received interaction text. For example, the semantic understanding model in some examples can be an NLP (Neuro-Linguistic Programming) model.

[0201] Based on this embodiment, in the case of comprehensively determining that there is an interaction intention, enter the wake-up mode to wake up the display device, and on the basis of combining the stored historical interaction information, generate recommended information by combining the fused text, considering the historical context interaction situation, thereby improving the accuracy of the generated recommended information, and enabling sign language-based interaction to be carried out on the basis of the existing historical interaction information, which helps to realize continuous sign language conversations.

[0202] In some embodiments, the processor is further configured to execute instructions to cause the display device to:

[0203] Identify the user characteristics in the video;

[0204] Semantic understanding is performed based on the stored historical interaction information, the fused text, and the user characteristics to generate recommendation information for the fused text.

[0205] Among them, in the embodiments of the present application, user characteristics are characteristics that can identify and distinguish different users, such as the face image of the user, the iris information of the user, etc., so as to distinguish different users. For example, based on the identified user characteristics, it can be analyzed that the user in the collected video is User 1, and this User 1 may be the male or female owner of a family, etc. Without limitation, as long as it can distinguish different users.

[0206] Among them, in the case where different users can be identified, personalized portraits can be further generated for each user to obtain the portrait characteristics of each user, and the portrait characteristics of the user are stored in the historical interaction information, or the historical interaction information stores user identifiers and stores the portrait characteristics of each user identifier, so as to enable the generation of recommendation information in combination with the portrait characteristics of the user.

[0207] In the case where different users can be identified, when semantic understanding is performed based on the stored historical interaction information, the fused text, and the user characteristics to generate recommendation information for the fused text, recommendation information for this user can be generated, so as to achieve the generation of more personalized recommendation information. For example, in some examples, semantic understanding can be performed based on the stored historical interaction information of User 1 and the fused text to generate recommendation information. In some examples, recommendation information can be generated based on the stored historical interaction information, the fused text, and the portrait characteristics of User 1 (such as the preferred movie type, language style, etc., but not limited thereto). It can be understood that after the user characteristics are identified, other personalized recommendation information generation can also be performed for the identified user, and the embodiments of the present application do not make specific limitations on this.

[0208] Based on this embodiment, when generating recommendation information, on the basis of combining historical interaction information and the fused text, user characteristics are further combined, so that semantic understanding and information recommendation can be performed in combination with user characteristics, making the obtained recommendation information more matching with user characteristics, which is helpful for the recommendation of personalized information based on user characteristics. In the case where multiple people are communicating, the stored historical interaction information may include text after role recognition, making the historical memory have the ability of role recognition and making it easier to determine whether there is an interaction intention in complex scenarios.

[0209] In some embodiments, the processor is further configured to execute instructions to cause the display device:

[0210] If there is an interaction intention, supplement the fused text and the recommendation information to the historical interaction information stored in the record.

[0211] If there is no interaction intention, at least one of the first text and the second text is added to the historical interaction information stored in the record.

[0212] Based on this embodiment, when it is comprehensively determined that there is an intention to interact, after the display device wakes up, the fused text and the corresponding generated recommendation information can be added to the recorded and stored historical interaction information to serve as the historical interaction information for the next interaction, thereby assisting in the generation of recommendation information for the next interaction. When it is comprehensively determined that there is no intention to interact, one or both of the first text and the second text can be added to the recorded and stored historical interaction information to serve as the historical interaction information for the next interaction, thereby assisting in the generation of recommendation information for the next interaction.

[0213] Based on the above-mentioned embodiments, Figure 5 The schematic diagram of the interactive scene of the display device is shown as follows:

[0214] If the user expresses the intention of interaction through sign language within the display range of the camera of a display device T such as a TV or smart screen, the camera of the display device T collects the video stream and transmits the collected video stream to the sign language understanding system for processing;

[0215] The sign language understanding system performs sign language recognition on the video stream to obtain text content, such as the first text and the second text as described above, and transmits the obtained text content to the information processing system;

[0216] The information processing system processes the received text content to obtain processed text content, such as fusing the first text and the second text to obtain a fused text. In general, due to errors in sign language recognition, the received text may be a single word or multiple related words or a majority of related words and a few irrelevant words (caused by recognition errors). It may also be a vague description or an ungrammatical sentence. Therefore, through the information processing system, the language understanding technology of the large model can be used to complete or polish these scattered, incoherent, and unclearly intended words or sentences, so that they become a complete sentence with clear semantics, that is, to obtain the fused text mentioned in the above embodiment.

[0217] The fused text can be provided to the interactive intent recognition system for interactive intent recognition to identify whether there is an interactive intent. When the interactive intent recognition system performs interactive intent recognition, it is not limited to the fused text, but can also combine gesture information, wake-up words, video understanding, and other interactive intent recognition.

[0218] If an interaction intention is recognized, a human-machine dialogue can be carried out through a semantic processing system, that is, recommendation information for the fused text is generated, and the generated recommendation information can be interactively feedback through an interaction system, such as performing one or more of specific interaction actions, displaying corresponding interaction texts, sign language displaying the recommendation information, etc.

[0219] Among them, the interaction intention recognition system and the semantic processing system can jointly form a large model system, which can be deployed on the display device side or in the cloud. In some specific examples, it can be deployed in the cloud to reduce the processing power requirements for the display device. In the case of being deployed in the cloud, after the cloud processes and obtains the recommendation information, the recommendation information can be fed back to the display device for the display device to play and display.

[0220] Among them, in the training and application processes of the models (such as intention recognition models) involved in the above-mentioned embodiments, KV Cache can be combined to accelerate the model processing process. Among them, KV Cache avoids repeated calculations by caching some intermediate results, and solves the problem of slow first-word inference speed in the large model inference process by trading space for time. It mainly pre-processes the prompt words and stores their corresponding Kv Cache in the hard disk. When inference is required, the corresponding KV Cache (such as the KV Cache of the prompt words) is imported in advance, which can reduce the reading and calculation time of most prompt words. Theory and experiments prove that it can reduce the inference cost by 90%. In a specific example, the Prompt Cache / Prefix Cache system obtained by accelerating and expanding KV Cache can be used to accelerate the model training and application processes, reduce the inference time, and improve the user experience.

[0221] Based on the above examples, refer to Figure 6 、 Figure 7 As shown, the sign language interaction process of a display device in a specific example is illustrated below.

[0222] First, the camera of the display device collects video data within the video range, and determines whether there is anyone. If there is no one, continue to judge. If it is judged that there is someone, proceed to the next step to avoid unnecessary consumption of resources for sign language content recognition when there is no one in the video.

[0223] When there is someone within the video range of the camera, call the sign language recognition model (the first video recognition model) to recognize the sign language content, and output the recognized first text X1 and the recognition accuracy rate (the first text probability value) a1; at the same time, call the video understanding model (the second video recognition model) to understand the video, and output the recognized second text X2 and the second text probability value a2.

[0224] Subsequently, based on the first text probability value a1 and the second text probability value a2, a target text processing model is selected and determined from multiple text processing models. Then, the determined target text processing model is used to process the obtained first text X1, the first text probability value a1, the second text X2, and the second text probability value a2, so as to finally output a sentence that contains environmental information and user requirements, has smooth grammar, and clear semantics. For example, as mentioned in the above description: I am in a home scene. Please recommend good movies that 40-year-old men like to watch.

[0225] Based on the obtained polished text with clear semantics and correct grammar, the interaction intention can be recognized through an intention understanding model to obtain the first interaction intention probability w1 of whether there is an interaction intention in the fused text.

[0226] If the display device supports wake word recognition and gesture recognition, then based on the gesture recognition of the gesture and whether there is a wake word or gesture recognition in the first text, the second interaction intention probability w2 can be obtained.

[0227] In addition, the third video recognition model can directly recognize the interaction intention of the collected video to obtain the third interaction intention probability w3 of whether there is an interaction intention.

[0228] Subsequently, based on the first interaction intention probability w1, the second interaction intention probability w2, and the third interaction intention probability w3, comprehensively consider to identify whether the user has an interaction intention.

[0229] If there is an interaction intention, then through the semantic understanding system, the corresponding recommendation algorithm is called, and the recommendation information is generated in combination with the stored historical conversation information. The generated recommendation information can be broadcast through methods such as sign language display and content display and provided to the user. After the execution is completed, the first text and the recommendation information can be stored as new historical interaction information to assist the recognition of the next sign language interaction and continue to recognize the video content.

[0230] If there is no interaction intention, then do not enter the wake-up mode, and only store the first text obtained by sign language recognition as new historical interaction information to assist the recognition of the next sign language interaction and continue to recognize the video content.

[0231] Based on the above-mentioned example, it can be seen that the solution of the embodiment of the present application supports downward compatibility. For example, it supports the existing interaction methods based on wake words or gesture instructions, and the display device can be incrementally developed on the existing architecture without invasive architecture modification, which can greatly improve the user interaction experience without losing the original functions.

[0232] Based on the above-mentioned display device, some embodiments of the present application also provide a control method for a display device.

[0233] Reference Figure 8 As shown, the control method of the display device in some embodiments includes:

[0234] Step S801; identify the sign language content of the collected video, and generate a first recognized text;

[0235] Step S802; identify the environmental information in the video, and generate a second text containing environmental description information;

[0236] Step S803; perform information fusion on the first text and the second text to obtain a fused text;

[0237] Step S804; identify the interaction intention of the fused text, obtain an interaction intention recognition result, and determine whether there is an interaction intention based on the interaction intention recognition result. If there is an interaction intention, generate recommendation information for the fused text and broadcast the recommendation information.

[0238] Some embodiments of the present application also provide a sign language interaction method, which can be executed by a display device, a server, or any other device capable of obtaining a sign language video and generating recommendation information for sign language interaction with users. The specific sign language interaction method may include:

[0239] Identify the sign language content of the collected video, and generate a first recognized text;

[0240] Identify the environmental information in the video, and generate a second text containing environmental description information;

[0241] Perform information fusion on the first text and the second text to obtain a fused text;

[0242] Identify the interaction intention of the fused text, obtain an interaction intention recognition result, and determine whether there is an interaction intention based on the interaction intention recognition result. If there is an interaction intention, generate recommendation information for the fused text and broadcast the recommendation information.

[0243] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are shown in sequence according to the indications of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order restriction, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.

[0244] In some embodiments, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the control method of the display device in any of the above-described embodiments are implemented.

[0245] In some embodiments, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the steps of the control method of the display device in any of the above-described embodiments are implemented.

[0246] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., and are not limited thereto. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., and are not limited thereto.

[0247] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0248] The above-described embodiments merely represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.

Claims

1. A display device, characterized in that: The display device comprises: monitor; and at least one processor configured to execute instructions to cause the display device to: Recognize the sign language content of the collected video and generate a recognized first text; Identifying environmental information in the video and generating a second text containing environmental description information; Performing information fusion on the first text and the second text to obtain a fused text; Identify the interaction intent of the fused text, obtain an interaction intent identification result, and determine whether there is an interaction intent based on the interaction intent identification result. If there is an interaction intent, generate recommendation information for the fused text and broadcast the recommendation information.

2. The display device according to claim 1, characterized in that The processor is further configured to execute instructions to cause the display device to: Recognize the sign language content of the video by using a first video recognition model, generate the first text, and obtain a first text probability value of the first text; Recognize the environmental information in the video by using a second video recognition model, generate a second text containing the environmental description information, and obtain a second text probability value of the second text; Based on the first text probability value and the second text probability value, determining a corresponding target text processing model from a text processing model set; The target text processing model is used to perform information fusion on the first text and the second text to obtain the fused text.

3. The display device according to claim 2, characterized in that The display device further includes at least one of the following: Item 1: The processor is further configured to execute instructions to cause the display device to: Recognize the environmental information and the sign language content in the video by using a second video recognition model, generate a second text including the environmental description information and the sign language content, and obtain a second text probability value of the second text; Item 2: The text processing model set includes a first text processing model and a second text processing model, and the number of model parameters of the first text processing model is less than the number of model parameters of the second text processing model; The processor is further configured to execute instructions to cause the display device to: Determining a corresponding target text processing model from a text processing model set based on the first text probability value and the second text probability value includes: If the first text probability value is greater than a first threshold, the first text processing model is determined as the target text processing model, If the first text probability value is less than a second threshold, the second text processing model is determined as the target text processing model, If the first text probability value is greater than the second threshold value and less than the first threshold value, determining a corresponding target text processing model from a text processing model set according to the second text probability value; Wherein, the first threshold is greater than the second threshold; Item 3: The processor is further configured to execute instructions to cause the display device to: The target text processing model is used to perform information fusion on the first text, the first text probability value, the second text and the second text probability value to obtain the fused text.

4. The display device according to claim 1, characterized in that The processor is further configured to execute instructions to cause the display device to: Identifying the interaction intent of the fused text, obtaining an interaction intent identification result, and obtaining a first interaction intent probability value; Monitoring a wake-up instruction, and obtaining a second interaction intention probability value of the monitored wake-up instruction; Comprehensively calculating the first interaction intention probability value and the second interaction intention probability value to obtain a comprehensive probability value; Based on the comprehensive probability value, it is determined whether there is an interaction intention.

5. The display device according to claim 4, characterized in that The display device further includes at least one of the following: Item 1: The processor is further configured to execute instructions to cause the display device to: Identify gesture information in the video, determine whether a wake-up instruction is received based on the gesture information, and obtain a second interaction intention probability value of the wake-up instruction; Item 2: The processor is further configured to execute instructions to cause the display device to: The processor is further configured to execute instructions to cause the display device to: If the first text includes a wake-up keyword, it is determined that a wake-up instruction is received, and a second interaction intention probability value of the wake-up instruction is determined.

6. The display device according to claim 4, characterized in that The processor is further configured to execute instructions to cause the display device to: Performing interaction intent recognition on the video to obtain a third interaction intent probability value of whether the video has interaction intent; Comprehensively calculating the first interaction intention probability value, the second interaction intention probability value, and the third interaction intention probability value to obtain a comprehensive probability value; Based on the comprehensive probability value, it is determined whether there is an interaction intention.

7. The display device according to any one of claims 1 to 6, characterized in that: The processor is further configured to execute instructions to cause the display device to: If there is an intention to interact, the wake-up mode is entered, and semantic understanding is performed based on the stored historical interaction information and the fused text to generate recommendation information for the fused text.

8. The display device according to claim 7, characterized in that The display device further includes at least one of the following: Item 1: The processor is further configured to execute instructions to cause the display device to: identifying user characteristics in the video; Performing semantic understanding based on the stored historical interaction information, the fused text, and the user characteristics, and generating recommendation information for the fused text; Item 2: The processor is further configured to execute instructions to cause the display device to: If there is an intention to interact, the fused text and the recommendation information are added to the historical interaction information stored in the record; If there is no interaction intention, one or both of the first text and the second text are added to the historical interaction information stored in the record.

9. A method for controlling a display device, characterized in that: The method comprises: Recognize the sign language content of the collected video and generate a recognized first text; Identifying environmental information in the video and generating a second text containing environmental description information; Performing information fusion on the first text and the second text to obtain a fused text; Identify the interaction intent of the fused text, obtain an interaction intent identification result, and determine whether there is an interaction intent based on the interaction intent identification result. If there is an interaction intent, generate recommendation information for the fused text and broadcast the recommendation information.

10. A sign language interaction method, characterized in that: The method comprises: Recognize the sign language content of the collected video and generate a recognized first text; Identifying environmental information in the video and generating a second text containing environmental description information; Performing information fusion on the first text and the second text to obtain a fused text; Identify the interaction intent of the fused text, obtain an interaction intent identification result, and determine whether there is an interaction intent based on the interaction intent identification result. If there is an interaction intent, generate recommendation information for the fused text and broadcast the recommendation information.

Citation Information

Cited By

  • Hearing-impaired person communication method and system based on user instruction emphasis

    CN120412103A

  • Communication method and system for hearing-impaired people based on user instruction emphasis

    CN120412103B