Display method and device, equipment and storage medium

By acquiring audio and video information, using a deep neural network model to recognize and convert it into intelligent sign language and subtitle text, and combining it with color analysis, a TV display interface suitable for people with disabilities is generated. This solves the shortcomings of existing TV systems in terms of accessibility display and achieves an efficient and convenient accessible TV viewing experience.

CN120956972APending Publication Date: 2025-11-14SHENZHEN ZHIXIAN VISION SOFTWARE TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511155631.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-18
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Existing television systems fail to meet the needs of people with hearing, visual, and speech impairments in terms of accessible display, especially in terms of speech recognition, multilingual support, real-time subtitles, and color adjustment in complex environments. Their user interface designs are also complex and fail to meet practical requirements.

Method used

By acquiring audio and video information, a deep neural network model is used to identify audio semantics and convert them into intelligent sign language and subtitle text. Combined with color analysis and adjustment, TV display interface information suitable for people with disabilities is generated, including AI voice enhancement, sign language animation, and personalized color processing.

Benefits of technology

It significantly improves the accuracy of speech recognition, enables real-time and accurate generation of sign language and subtitles, optimizes the visual experience for colorblind people, reduces manual operation, and provides convenient and efficient multi-directional barrier-free TV display.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120956972A_ABST
    Figure CN120956972A_ABST
Patent Text Reader

Abstract

The invention discloses a display method and apparatus, a device and a storage medium. The method comprises the steps of obtaining to-be-processed audio information and to-be-adjusted picture information; identifying audio semantics based on to-be-processed audio information and a deep neural network model, and converting an intelligent sign language voice text to generate sign language animation and subtitle text information; and determining television display interface information based on the to-be-processed audio information, the sign language animation, the subtitle text information and the to-be-adjusted picture information, and controlling television barrier-free display based on the television display interface information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence television application technology, and in particular to display methods, devices, equipment and storage media. Background Technology

[0002] With the continuous development of society, improving the quality of life for people with disabilities has become an important social goal. Television, as an important tool for information dissemination and entertainment, is needed even by people with hearing, visual, and speech impairments. However, current television systems still have many shortcomings in terms of accessible display, failing to meet the needs of these special groups.

[0003] Currently, existing methods mainly rely on manual operation. Manually adjusting the audio amplification function helps people with hearing impairments. Only basic subtitle functions are provided to assist people with hearing impairments in understanding program content. For colorblind mode, users need to manually select the preset color adjustment scheme according to their colorblindness type. Sign language translation function usually requires manual recording and synchronization and cannot be generated in real time.

[0004] However, existing methods, such as simple audio amplification, cannot effectively solve speech recognition problems in complex environments. They lack support for different languages ​​and dialects, and basic subtitle functions are unsatisfactory in terms of real-time performance and accuracy, making it difficult to handle fast-paced conversations or multilingual environments. Furthermore, color enhancement technology cannot accurately adjust for different types of color blindness. In addition, the user interface design is not intuitive enough, and the operation is complex, failing to meet the actual needs of people with disabilities. Therefore, how to utilize artificial intelligence to achieve accessible television displays has become an urgent problem to be solved.

[0005] The above content is only used to help understand the technical solution of this application and does not represent an admission that the above content is prior art. Summary of the Invention

[0006] The main objective of this application is to provide a display method, apparatus, device, and storage medium, which aims to solve the technical problem of how to achieve barrier-free display of television using artificial intelligence.

[0007] To achieve the above objectives, this application proposes a display method, which includes:

[0008] Acquire the audio information to be processed and the image information to be adjusted;

[0009] Based on the audio information to be processed and a deep neural network model, the audio semantics are identified, and intelligent sign language speech text is converted to generate sign language animation and subtitle text information;

[0010] The TV display interface information is determined based on the audio information to be processed, sign language animation, subtitle text information, and screen information to be adjusted, and the TV accessibility display is controlled based on the TV display interface information.

[0011] Furthermore, to achieve the above objectives, this application also proposes a display device, which includes:

[0012] The acquisition module is used to acquire the audio information to be processed and the image information to be adjusted.

[0013] The processing module is used to identify audio semantics based on the audio information to be processed and a deep neural network model, and to convert intelligent sign language speech into text, generating sign language animation and subtitle text information;

[0014] The execution module is used to determine the TV display interface information based on the audio information to be processed, sign language animation, subtitle text information, and screen information to be adjusted, and to control the TV accessibility display based on the TV display interface information.

[0015] In addition, to achieve the above objectives, this application also proposes a display device, which includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the display method as described above.

[0016] In addition, to achieve the above objectives, this application also proposes a storage medium, which is a computer-readable storage medium, and stores a computer program on the storage medium. When the computer program is executed by a processor, it implements the steps of the display method as described above. Attached Figure Description

[0017] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0018] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 This is a flowchart illustrating the method of Embodiment 1 of this application;

[0020] Figure 2 This is a schematic diagram illustrating the method scenario shown in this application;

[0021] Figure 3 This is a schematic diagram of the processing flow of the method shown in this application;

[0022] Figure 4 This is a flowchart illustrating the second embodiment of the method shown in this application;

[0023] Figure 5This is a schematic diagram of the module structure of the display device according to an embodiment of this application;

[0024] Figure 6 This is a schematic diagram of the device structure of the hardware operating environment involved in the display method in the embodiments of this application.

[0025] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0026] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.

[0027] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.

[0028] The main solution of this application embodiment is: to acquire audio information to be processed and screen information to be adjusted; to identify audio semantics based on the audio information to be processed and a deep neural network model, and to convert intelligent sign language speech text to generate sign language animation and subtitle text information; to determine TV display interface information based on the audio information to be processed, sign language animation, subtitle text information and screen information to be adjusted, and to control TV accessibility display based on the TV display interface information.

[0029] In this embodiment, for ease of description, the following description will focus on the identification display device as the execution subject.

[0030] Because existing technologies cannot effectively solve speech recognition problems in complex environments by simply amplifying audio, lack support for different languages ​​and dialects, and perform poorly in terms of real-time performance and accuracy, making it difficult to handle fast conversations or multilingual environments, and color enhancement technology cannot accurately adjust according to different types of color blindness, while the user interface design is not intuitive enough and the operation is complicated, making it difficult to meet the actual needs of people with disabilities.

[0031] This application provides a solution for acquiring audio information to be processed and screen information to be adjusted; recognizing audio semantics based on the audio information to be processed and a deep neural network model, and converting intelligent sign language speech text to generate sign language animation and subtitle text information; determining TV display interface information based on the audio information to be processed, sign language animation, subtitle text information and screen information to be adjusted, and controlling TV accessibility display based on the TV display interface information.

[0032] As can be seen from the above embodiments, this application utilizes the audio information to be processed and a deep neural network model to identify audio semantics, converts it into intelligent sign language speech text, and generates sign language animation and subtitle text information. At the same time, it combines the screen information to be adjusted to determine the TV display interface, thereby controlling accessibility display, significantly improving the accuracy of speech recognition, realizing real-time and accurate generation of sign language and subtitles, optimizing the visual experience of colorblind people, and reducing manual operation through intelligent means, making multi-directional TV accessibility display more convenient and efficient.

[0033] Based on this, the embodiments of this application provide a display method, referring to... Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the method shown in this application.

[0034] In this embodiment, the display method includes steps S10 to S40:

[0035] Step S10: Obtain the audio information to be processed and the image information to be adjusted;

[0036] It should be noted that the audio information to be processed is pre-processed audio data, while the picture information to be adjusted is picture data after color analysis of the TV picture, which is adjusted according to the visual needs of colorblind users by adjusting the color element area information.

[0037] Understandably, the audio information to be processed can be preprocessed through adaptive noise reduction and sound quality enhancement. Adaptive noise reduction can dynamically identify and suppress background noise, such as ambient noise during TV playback, air conditioning noise, or other interference, thereby highlighting the speech signal. Sound quality enhancement, on the other hand, enhances the high-frequency and low-frequency components of speech through frequency equalization and dynamic range adjustment, making the speech fuller and easier to understand. This effectively removes background noise, highlights the speech signal, and improves the clarity and comprehensibility of the speech.

[0038] Additionally, it should be noted that the image information to be adjusted can be obtained by identifying color element areas in the image based on different types of color blindness, such as red-green color blindness and blue-yellow color blindness, and performing precise color conversion. For example, for red-green color blind users, the system will enhance the color contrast between red and green, or map red and green to other easily distinguishable colors, such as yellow and blue. The system will analyze the color distribution in the image in real time, identify the areas that need adjustment, and apply personalized color conversion schemes to optimize color contrast and hue, making it easier for color blind users to identify image information.

[0039] For ease of understanding, we will take the acquisition of audio information to be processed and image information to be adjusted as an example. The information acquisition device is the information acquisition module, and the storage device is the memory.

[0040] The information acquisition module obtains microphone audio information using array microphone technology, arranging multiple microphones to capture sound from different directions and enhance the spatial resolution of the signal. A digital signal processor (DSP) processes the audio signal in real time, performing initial filtering and gain control to preprocess the microphone audio information, resulting in the audio information to be processed. Preprocessing includes adaptive noise reduction and sound quality enhancement. The adaptive noise reduction algorithm implements a deep learning-based noise reduction model, such as a convolutional neural network or a recurrent neural network, for real-time noise suppression. The model is trained to identify common environmental noise patterns, such as wind noise and traffic noise. It collects 500 hours of audio for each type of common noise, including wind noise, traffic noise, and human voice noise, using a lightweight hybrid model architecture, a TCN temporal convolutional network, and BiGRU bidirectional gating units. The TCN temporal convolutional network has 5 layers with a dilation factor of 2. M M represents the number of layers. Short-time Fourier transform is then used to extract time-frequency features, and generative adversarial networks are used for end-to-end noise reduction. Noise reduction parameters are dynamically adjusted to adapt to different environments. Audio enhancement technology is used to process microphone audio information. In the audio enhancement technology, the frequency equalizer is used to adjust the spectrum of the audio signal and enhance the clarity of speech, especially the enhancement of high-frequency and low-frequency components. The dynamic range compressor processes the audio signal according to the set parameters. When the loudness of the audio signal exceeds the set threshold, the compressor will reduce the signal gain to reduce its loudness. The degree of compression is determined by the compression ratio. For example, a compression ratio of 4:1 means that the part of the input signal that exceeds the threshold will be compressed to one-quarter of the original volume to prevent the volume from being too loud or too soft and to ensure the audibility of the speech.

[0041] Acquiring television image information involves extracting the television image, performing color analysis on the image, identifying the color elements in the areas to be adjusted, and then performing color conversion to obtain the image information to be adjusted. This can be achieved using color analysis tools or color conversion algorithms to identify the color elements in the areas to be adjusted and perform color conversion. Color analysis tools use computer vision technology to perform real-time color analysis on the television image, identify the color areas that need adjustment, and apply image processing algorithms, such as edge detection and region segmentation, to accurately identify the color elements in the image. Color conversion algorithms, on the other hand, convert colors that the user cannot recognize into colors that the user can recognize, based on the user's color blindness type. This involves applying specific color conversion algorithms, such as color mapping and contrast enhancement. A specific color conversion algorithm can be a defined algorithm used to convert colors from one color to another. The mathematical formula for converting a color into another involves transforming the input color data into an equivalent representation of that color through calculation. For example, matrix transformations or nonlinear mappings can be used to convert the numerical representation of color from one format to another, while maintaining visual consistency as much as possible. Machine learning models are then used to predict the optimal color adjustment scheme, extracting color features from the image, such as RGB values, HSV values, brightness histograms, and color distribution, as well as contextual information, such as the scene type (nature, city, or people). Machine learning models are used to predict the optimal color adjustment scheme to ensure accurate transmission of color information. Both methods require user testing to collect feedback and continuously optimize the color adjustment algorithm to improve the user's visual experience. For example, two tests are conducted to compare the actual effects of different color adjustment schemes, and the scheme with the highest user satisfaction is selected for adjusting the image, thus obtaining the image information to be adjusted.

[0042] In one feasible implementation, step S10 may include steps A11 to A16:

[0043] Step A11: Obtain microphone audio information and television screen information;

[0044] It should be noted that microphone audio information is the raw audio signal captured by the microphone, while television picture information is the raw visual content displayed on the television screen.

[0045] Understandably, microphone audio information can be acquired using array technology. An array receives signals simultaneously from multiple sensors and utilizes signal processing techniques, such as beamforming, to combine multiple signals into a stronger one. This effectively reduces background noise interference and improves signal clarity and quality. For example, in speech recognition scenarios, microphone arrays can significantly improve the signal-to-noise ratio of the speech signal, making speech recognition more accurate. Furthermore, array technology can adjust the phase and amplitude of each sensor to form a beam in a specific direction, thereby more accurately capturing or transmitting signals and suppressing interference from other directions. For instance, microphone arrays can achieve directional sound pickup, receiving the speech signal from the direction the user is speaking while ignoring noise from other directions.

[0046] Additionally, it should be noted that television screen information can include program videos, advertisements, subtitles, and image information. Generally speaking, television screen information is intuitive and easy to understand for normal viewers, but it may be difficult for colorblind users to identify. For example, red-green colorblind users have difficulty distinguishing red and green elements in the screen, which affects their understanding of the program content. Therefore, it is necessary to perform personalized color optimization according to different types of colorblindness. By analyzing the color distribution and color element area information in the television screen information, the color contrast and hue can be precisely adjusted so that colorblind users can perceive the color information in the screen more clearly.

[0047] Step A12: Preprocess the microphone audio information to obtain the audio information to be processed. The preprocessing includes adaptive noise reduction and sound quality enhancement.

[0048] Understandably, preprocessing typically involves performing operations such as denoising, format conversion, feature enhancement, and normalization on the raw data to remove irrelevant noise and errors, converting the data into a format suitable for further analysis. For example, in audio signal processing, preprocessing removes background noise using adaptive noise reduction techniques while enhancing speech clarity through frequency equalization and dynamic range compression. In image processing, preprocessing removes noise and enhances edge information through filters, making image features more prominent, providing more accurate data, and adapting to the needs of different application scenarios.

[0049] Step A13: Use a color analysis tool to perform color analysis on the television screen information, and divide the television screen information into color areas to be adjusted according to the region segmentation image processing algorithm to determine the color element area information.

[0050] It should be noted that a color analysis tool is a software tool or algorithm used for real-time color analysis of television images. Through computer vision technology and image processing algorithms, it converts the RGB color space of the television image into a more suitable color space for analysis, statistically analyzes color distribution, and identifies color element regions and their characteristics within the image. This color analysis tool uses edge detection and region segmentation algorithms to divide the image into multiple color regions, extracting key features such as average hue, saturation, and brightness from each region. It then analyzes the importance of each region in conjunction with contextual information. Simultaneously, it calculates the contrast between adjacent color regions to determine whether they meet the visual needs of colorblind users. Using color analysis tools, various types of television content can be processed in real-time and accurately, providing colorblind users with clearer and more easily identifiable image information, thereby significantly improving their visual experience.

[0051] Understandably, color analysis uses computer vision technology to detect and analyze color information in a television image in real time. The area to be adjusted is the region of the image that is identified as causing visual impairment to colorblind users during the color analysis process. Color element region information refers to the color components in the area to be adjusted that have an important impact on the understanding of the image content. Region segmentation image processing algorithm is an algorithm used to divide an image into several non-overlapping regions. It can divide an image into several regions with similar characteristics, that is, decompose the image into multiple parts with similar characteristics. The similar characteristics can be extracted colors, textures, or brightness, ensuring that the feature differences between different regions are obvious.

[0052] Additionally, it should be noted that color analysis can identify attributes such as color distribution, hue, saturation, and brightness in different areas of the image, and extract color information related to the visual perception of colorblind users. The color areas to be adjusted may contain color combinations that are difficult to distinguish, such as red and green, blue and yellow, or areas with low color contrast. Color element area information refers to color components or areas in the television image that have a significant impact on the visual recognition and understanding of the image content by colorblind users. It can be used to identify color information that colorblind users can understand and distinguish. Color conversion uses algorithms to map colors that are difficult to distinguish to other colors that colorblind users can perceive more clearly, or to improve the visual effect by enhancing color contrast. Color element area information can include high-contrast color area elements, important information area elements, and background and foreground color combination elements. High-contrast color area elements can be color combinations of red and green or blue and yellow, which are difficult for colorblind users to distinguish. Important information area elements can be subtitle elements or icon elements, while background and foreground color combination elements can be combinations of sky and grass or buildings and natural environment, which directly affect the perception of the overall structure of the image by colorblind users.

[0053] Step A14: Use a color conversion algorithm to extract the image color features corresponding to the color element region information, and determine the color feature information and context information;

[0054] It should be noted that color feature information is quantitative data related to color extracted from the television image, while context information is the image content and scene information related to the color features.

[0055] Understandably, color feature information describes the characteristics of various color areas in an image, including hue, saturation, brightness, color distribution, and color contrast. Hue represents the basic type of color, such as red, blue, and green. A red apple in the HSV color space has a hue value close to 0°. Saturation represents the purity of a color, i.e., its vividness. The higher the saturation, the more vivid the color; the lower the saturation, the closer the color is to gray. For example, a bright red apple has high saturation, while a dull red apple has low saturation. Brightness represents the lightness or darkness of a color. The higher the brightness, the brighter the image; the lower the brightness, the darker the image. For example, a red apple photographed in sunlight has high brightness, while a red apple photographed in the shade has low brightness. Color distribution can statistically analyze the percentage and distribution area of ​​different colors in an image. For example, it can calculate the percentage of red pixels in the total number of pixels and where red pixels are mainly concentrated in the image. Color contrast can calculate the contrast between adjacent color areas. For example, the hue contrast between red and green is high, while the hue contrast between similar hues is low.

[0056] Additionally, it should be noted that contextual information refers to the image content and scene information related to color features, such as scene type, object category, color correlation, color importance, and color emotion. By combining color feature information and contextual information, color element area information can be identified more accurately, and color adjustment schemes that better meet the visual needs of colorblind users can be generated, thereby optimizing the image display effect and improving the visual experience of colorblind users.

[0057] Step A15: Concatenate the color feature information and context information into an input vector in a fixed order, and input the input vector into the machine learning model to predict the color adjustment combination to generate a set of color adjustment schemes;

[0058] It should be noted that the color adjustment scheme set is a series of color optimization strategies generated by machine learning models to provide colorblind users with more easily distinguishable and perceptible colors in the picture. User satisfaction is evaluated through version testing, and the best-performing scheme is finally selected to generate the picture information to be adjusted, thereby significantly improving the visual experience of colorblind users.

[0059] Understandably, an input vector is a data structure that combines color feature information and contextual information. It is used to input into a machine learning model to predict the probability of color adjustment combinations and generate a set of color adjustment schemes. It requires the integration of multiple feature information so that the model can learn and predict effectively. A fixed order means that the arrangement of color feature information and contextual information is fixed when constructing the input vector. For example, they are arranged in a fixed order to ensure that the vector structure input each time is consistent.

[0060] Additionally, it should be noted that color adjustment combinations are based on input color feature information and contextual information. The model calculates every possible color adjustment combination that meets the user's needs in order to predict different color adjustment schemes.

[0061] Step A16: Use the color adjustment scheme set to conduct version testing to evaluate user satisfaction, and select the target color adjustment scheme with the highest user satisfaction to perform color conversion on the color element area information to obtain the image information to be adjusted.

[0062] Understandably, color conversion involves transforming colors in an image based on the user's colorblindness type to improve their ability to distinguish colors. For example, for red-green colorblind users, the system might enhance the contrast between red and green, or map red and green to other easily distinguishable colors, such as yellow and blue. Color conversion algorithms can analyze the color distribution in the image in real time, identify areas requiring adjustment, and apply different personalized color conversions to select the most suitable solution for optimizing color contrast and hue, making it easier for colorblind users to recognize image information.

[0063] Step S20: Based on the audio information to be processed and the deep neural network model, the audio semantics are identified, and the intelligent sign language speech text is converted to generate sign language animation and subtitle text information;

[0064] It should be noted that sign language animation is an animated display that uses artificial intelligence technology to convert speech or text content into sign language actions in real time. The subtitle text information is text content extracted and optimized from intelligent sign language speech text and used to display on the TV screen.

[0065] Understandably, a deep neural network model is an optimized training model obtained by using a pre-trained deep neural network model and training it with a large amount of audio training condition information, thereby performing multilingual speech recognition. The pre-trained deep neural network model can be Transformer or BERT.

[0066] Additionally, it should be noted that sign language animation can present sign language expressions in a natural, fluent, and easy-to-understand way, ensuring that hearing-impaired individuals can intuitively obtain information. Based on a pre-trained sign language generation model, combined with speech recognition and natural language processing technologies, speech content is converted into sign language actions and displayed through virtual characters or animations, thereby accurately expressing semantics and conforming to the grammatical structure and expression habits of sign language. Subtitle text information can present audio content in a concise, accurate, and easy-to-read way, and may include auxiliary information such as dialogue content, background sound effects, speaker identity, and emotional tone, helping hearing-impaired or speech-impaired individuals to better understand the program content.

[0067] For ease of understanding, the example of generating sign language animation and subtitle text information will be used. The information acquisition device is the information acquisition module, the storage device is the memory, and the processing device is the processing module.

[0068] Based on the audio information to be processed and a deep neural network model, the audio semantics are identified, and intelligent sign language speech text is converted to generate sign language animation and subtitle text information. Specifically, the speech recognition module uses a pre-trained deep neural network model, such as Transformer or BERT, for multilingual speech recognition. The model is trained on a large-scale speech dataset, supporting the recognition of multiple languages ​​and dialects. Natural language processing techniques are used to analyze the identified text, handling homophones, grammatical structures, and contextual semantics. Named entity recognition and sentiment analysis techniques are used to ensure accurate text understanding and processing. The sign language generation module collects high-quality sign language datasets, uses MediaPipe to extract key points of hands and body, and consists of a generator and a discriminator. Realistic sign language animations are generated through adversarial training, and the naturalness and semantic consistency of the generated sign language movements are adjusted through reward parameters of a reward mechanism. These reward parameters are variables used to quantify the specific numerical value or condition of the reward in the reward mechanism. For example, they may include action accuracy parameters to measure the similarity between the sign language movement and the standard movement, and semantic accuracy parameters to measure the semantic accuracy of the sign language movement. The system considers factors such as whether the semantics expressed in the animation are consistent with the input text or speech, the smoothness of the action (measured by the continuity and naturalness of the sign language movements), user satisfaction (measured by the user's satisfaction with the sign language animation), and reward weight and threshold parameters (used to adjust the importance of different reward dimensions and define the level at which rewards are given). By setting these parameters appropriately, the reward mechanism can effectively incentivize users or the system to generate more accurate, natural, and easily understandable sign language movements and semantics, thereby improving the overall quality of the sign language animation and ensuring the naturalness and fluency of the sign language expression. The sign language database contains a rich vocabulary and sentence structure, supporting multiple sign language languages ​​such as ASL and BSL. ASL is a complete natural language, commonly used by hearing-impaired communities in the United States and Canada, expressing language through gestures, facial expressions, body postures, and spatial positioning. BSL is the main sign language used by hearing-impaired communities in the United Kingdom. BSL and ASL are two completely different languages ​​with no direct connection; however, both BSL and ASL have independent grammatical structures, vocabulary, and expressions, showing significant differences in structure and expression from spoken language, thus enabling the creation of sign language animations.

[0069] The subtitle generation system utilizes natural language processing technology to convert speech-recognized text into subtitles, supports multilingual subtitle options, and uses language models such as GPT to generate natural and fluent subtitle text, ensuring the accuracy of the subtitle content. At the same time, it uses dynamic time warping technology to align audio and text, generating accurate timestamps. Combined with NLP models, it optimizes the segmentation and display of subtitle text, ensuring that subtitles and audio content are displayed synchronously. It implements an automatic delay adjustment algorithm to dynamically adjust the subtitle display time according to the audio rhythm, providing a smooth viewing experience, thus obtaining subtitle text information.

[0070] Step S30: Determine the TV display interface information based on the audio information to be processed, sign language animation, subtitle text information, and screen information to be adjusted, and control the TV accessibility display based on the TV display interface information.

[0071] It should be noted that the information displayed on the TV screen is processed and used to control the output content of the TV's accessibility display.

[0072] It is understandable that, such as Figure 2 As shown, Figure 2 This is a schematic diagram of the display method scenario of this application, including AI voice enhancement for hearing impairment, AI subtitles for hearing impairment, AI color blindness mode for hearing impairment, and AI sign language generation for hearing impairment. The TV display interface information is an integration of multiple accessibility functions, and also includes personalized settings for users. It can provide customized visual and auditory information display according to the different needs of people with disabilities, clearly display video images and optimized colors, integrate synchronized sign language animation windows, accurate subtitle text, and user-adjustable assistive functions, such as volume enhancement and contrast adjustment, to provide people with disabilities with an accessible, easy-to-use, and efficient viewing experience.

[0073] Additionally, it should be noted that, as Figure 3 As shown, Figure 3 This is a schematic diagram of the processing flow of the display method of this application. The TV plays a TV program, outputting sound and picture. The output picture is intercepted to obtain model input data, and then fed into an AI color blindness model for picture processing to obtain picture information to be adjusted, thereby displaying it on the TV. The output sound is intercepted to obtain model input data, and then fed into an AI voice enhancement model for sound processing. The audio information to be processed is directly output for TV display. The audio information to be processed is fed into an AI subtitle model to obtain subtitle text information for TV display. The audio information to be processed is fed into an AI sign language generation model to obtain sign language animation for TV display.

[0074] For ease of understanding, we will take determining the information on the television display interface as an example. The information acquisition device is the information acquisition module, the storage device is the memory, and the execution device is the execution module.

[0075] The information collection module acquires user interaction information, including interface feedback, personalized settings feedback, and user feedback. This involves designing an intuitive user interface that supports voice control, gesture recognition, and touch operation, facilitating use by people with disabilities. The interface design adheres to accessibility principles, ensuring easy access and operation for all users. It also involves obtaining interface feedback, enabling personalized settings, providing users with options to customize system settings such as subtitle font size, sign language display position, and audio enhancement level. The user interface supports personalized configuration, allowing users to adjust the interface layout and functions according to their preferences, and obtaining personalized settings feedback. Finally, it integrates a user feedback mechanism to collect user feedback. User experience data is used to continuously optimize system performance. Data analysis tools are used to analyze user feedback, identify improvement opportunities and user needs, and ensure that the system can adapt to the ever-changing needs of users. User feedback information is obtained and categorized using data analysis tools to determine user function improvement needs. Interface feedback information, personalized setting feedback information, and user feedback information are categorized according to user improvement needs to determine adjustment information. User improvement needs include subtitle display settings, voice settings, sign language animation settings, and position settings. Based on the target TV display interface information, the TV accessibility display is controlled so that it can be used for all people with hearing and visual impairments to watch TV.

[0076] In one feasible implementation, step S30 may include steps B11 to B15:

[0077] Step B11: Obtain user interaction information, which includes interface feedback information, personalized settings feedback information, and user feedback information.

[0078] It should be noted that user interaction information is the data generated by all interactions between the user and the system when using the display system; interface feedback information is the feedback information generated by the user's intuitive feeling and operation experience of the user interface when using the TV accessibility display system; personalized setting feedback information is the feedback information generated after the user personalizes the system according to their own needs; and user feedback information is the feedback information of the user on the system's functional experience during the use of the system.

[0079] Understandably, user interaction information can include user operations on the interface, use of functions, and adjustments to system settings. Interface feedback information can represent user evaluations of the interface functions, such as user feedback that the location of function buttons is not intuitive enough or that the interface operation steps are too complicated. Improving the interface can better meet the usage habits and needs of people with disabilities. Personalized setting feedback information can represent users' personalized needs and preferences for system functions. User feedback information can also represent problems encountered by users during use and suggestions for functions that urgently need improvement, thereby better meeting the needs of people with disabilities.

[0080] Step B12: Use data analysis tools to classify user interaction information and determine user function improvement needs.

[0081] It should be noted that user feature improvement requests are suggestions and requirements made by users based on their own experience and expectations when using the product or service, to optimize, enhance, or add existing features.

[0082] Understandably, user feature improvement requests can represent the problems users encounter in actual use and the new features or improvements they hope the product can provide.

[0083] Step B13: Use data analysis tools to classify interface feedback information, personalized settings feedback information and user feedback information according to user improvement needs, and determine the adjustment information. User improvement needs include subtitle display settings, voice settings, sign language animation settings and position settings.

[0084] It should be noted that the adjustment information is the result of an analysis of the user's personalized adjustment wishes. It is used to adjust the specific instructions or parameters of the system to optimize and improve the information on the TV display interface, so that the TV accessibility display system can better meet the actual usage needs of people with disabilities.

[0085] Understandably, user improvement requests can include subtitle display settings, voice settings, sign language animation settings, and position settings. Subtitle display settings refer to user requests to adjust subtitle display, such as adjusting font size, color, and display position, or the speed and size of sign language animations. Voice settings involve user requests for new functionalities, such as adding voice control, gesture recognition, or supporting more languages ​​and dialects. These user improvement requests allow for targeted, personalized adjustments to the TV display interface. Sign language animation settings address user dissatisfaction with the performance or effects of sign language functions, such as adjusting the display method. Position settings reflect user feedback on optimized interface operations, such as adjusting interface operations, button positions, or the location of certain functional interfaces. Based on these user improvement requests, the target TV display interface information can be determined, thus better aligning with the usage habits and needs of people with disabilities, providing a more convenient, efficient, and personalized accessible TV viewing experience.

[0086] Step B14: Extract the display parameters corresponding to different user improvement requests from the adjustment information, and adjust the parameters in the TV display interface information based on the display parameters to determine the target TV display interface information.

[0087] It should be noted that the target TV display interface information is the final interface configuration information used to control the TV's accessibility display, generated by the system after processing user feedback and data analysis tools.

[0088] Understandably, the resulting target TV display interface information is an optimized interface display state that integrates user feedback and needs, providing people with disabilities with a more optimized, user-friendly, and personalized accessible display experience, ensuring that all people with disabilities can enjoy an accessible TV viewing experience.

[0089] Step B15: Based on the target TV display interface information, control the TV configuration display interface to complete the TV accessibility display.

[0090] Understandably, the system can optimize the accessibility display function based on user feedback, that is, by using the target TV display interface information to control the TV configuration display interface, so that the generated sign language animations, subtitles and colorblind images accurately meet the user's needs, are more in line with the usage habits of people with disabilities, and reduce the learning cost.

[0091] This embodiment proposes a display method that acquires audio information to be processed and screen information to be adjusted; identifies audio semantics based on the audio information to be processed and a deep neural network model, and converts it into intelligent sign language speech text to generate sign language animation and subtitle text information; determines TV display interface information based on the audio information to be processed, sign language animation, subtitle text information, and screen information to be adjusted; and controls the TV's accessible display based on the TV display interface information. This solves the technical problem of how to achieve accessible TV display using artificial intelligence. Compared with existing technologies, this application acquires optimized audio information to be processed and screen information optimized for color blindness, accurately identifies audio semantics using a deep neural network model and converts it into intelligent sign language speech text, generating sign language animation and subtitle text information, thereby obtaining TV display interface information to achieve accessible TV display control. AI voice enhancement improves audio clarity to help hearing-impaired individuals; sign language generation provides real-time translation for easier understanding; color blindness mode optimizes color display and improves visual experience; and subtitle synchronization ensures consistency between subtitles and audio, improving information acquisition efficiency and providing an accessible TV viewing experience for people with disabilities.

[0092] Based on the first embodiment of this application, in the second embodiment of this application, the same or similar content as the first embodiment can be referred to the above description, and will not be repeated hereafter.

[0093] In this embodiment, refer to Figure 4 , Figure 4 This is a flowchart illustrating the second embodiment of the method shown in this application. Step S20 specifically includes steps S21 to S24:

[0094] Step S21: Based on the audio information to be processed and the deep neural network model, the audio semantics are identified, and the human pose points are extracted using the sign language detection model to perform sign language animation conversion, thereby determining the sign language animation;

[0095] It should be noted that audio semantics is the semantic information extracted from audio signals, that is, the specific meaning expressed by the audio. Intelligent sign language voice text is the conversion of audio semantics into a text form suitable for sign language expression, which must conform to the language structure and expression habits of sign language.

[0096] Understandably, a sign language detection model is an algorithm based on computer vision and deep learning technologies used to detect in real time whether someone is using sign language in a video. It captures video frames through a camera, preprocesses the images including denoising and filtering, and uses human pose estimation technology to detect human pose points, including key points on the hands, body, and face. These human pose points are two-dimensional or three-dimensional coordinate points of various key parts of the human body detected by computer vision technology during pose estimation, such as the head, shoulders, elbows, wrists, and knees. This captures hand and body movements. By analyzing the motion trajectory and positional changes of these key points, the model can determine whether someone is using sign language. It calculates the normalized displacement features of the human pose points, extracts motion features by comparing the positional changes of key points in consecutive frames, and inputs the extracted features into temporal models such as Long Short-Term Memory networks to identify the temporal sequence features of sign language actions.

[0097] Additionally, it should be noted that using deep neural network models for semantic recognition of audio can effectively handle complex speech environments, including dialects, background noise, and multilingual mixing, more accurately converting speech into text, accurately expressing audio content, while conforming to sign language language habits, and supporting multiple languages ​​and dialects to meet the needs of different regions and groups.

[0098] For ease of understanding, we will use sign language animation as an example, where the information acquisition device is the information acquisition module, the storage device is the memory, and the processing device is the processing module.

[0099] The information acquisition module obtains audio training information, including language training information and dialect training information. This information is then input into a deep neural network model to identify the natural language in the audio information to be processed. The language recognition dataset is determined; that is, the speech recognition module uses a pre-trained deep neural network (DNN) model, such as Transformer or BERT, to perform multilingual speech recognition. The model is trained on a large-scale speech dataset, supporting the recognition of multiple languages ​​and dialects. The natural language in the language recognition dataset is extracted to analyze the corresponding audio semantics, thus determining the intelligent sign language speech text. Semantic analysis and processing are then performed, using natural language processing techniques to analyze the identified text, handling homophones, grammatical structures, and contextual semantics. Named entity recognition and sentiment analysis techniques are used to ensure accurate understanding and processing of the text. Based on the intelligent sign language speech text and the sign language database, sign language is recognized, and a sign language detection model is used to extract human pose points for sign language animation conversion, resulting in sign language animation. The sign language generation module collects high-quality sign language datasets, extracts key points of hands and body using MediaPipe, and consists of a generator and a discriminator. Through adversarial training, it generates realistic sign language animations, and adjusts the naturalness and semantic consistency of the generated sign language movements through a reward mechanism to ensure the naturalness and fluency of the sign language expression. The sign language database contains a rich vocabulary and sentence structures, supporting multiple sign language languages ​​such as ASL and BSL. ASL is a complete natural language, commonly used by hearing-impaired communities in the United States and Canada, expressing language through gestures, facial expressions, body postures, and spatial positions. BSL is the primary sign language used by hearing-impaired communities in the United Kingdom. BSL and ASL are two completely different languages ​​with no direct connection; however, both BSL and ASL have independent grammatical structures, vocabulary, and expressions, showing significant differences in structure and expression from spoken languages, thus enabling the generation of sign language animations.

[0100] In one feasible implementation, step S21 may include steps C11 to C14:

[0101] Step C11: Obtain audio training status information, which includes language training status information and dialect training status information.

[0102] It should be noted that audio training runtime information is a set of audio data used to train artificial intelligence models, language training runtime information is audio data and its annotation information for different languages, and dialect training runtime information is audio data and its annotation information for different dialects within a language.

[0103] Understandably, audio training information can include language, dialect, speech rate, intonation, and background noise. Language training information can represent the standard pronunciation, grammatical structure, and vocabulary usage of a language, enabling the model to accurately recognize and understand the speech content of different languages. Dialect training information can represent the special pronunciation, vocabulary, and intonation of a dialect, enabling the model to adapt to dialect differences in different regions and improve the model's speech recognition ability and accuracy in dialect environments.

[0104] Step C12: Input the audio training condition information into the deep neural network model to identify the natural language in the audio information to be processed, and determine the language recognition dataset;

[0105] It should be noted that a language recognition dataset is a collection of data generated during audio processing and natural language processing. It consists of audio information that has been processed and analyzed by a deep neural network model and converted into text form, including rich semantic and linguistic features, used to represent the semantic information in the audio.

[0106] Understandably, natural language is the language humans use daily, conforming to human language habits and possessing rich semantics, grammatical structures, and contextual relationships. Language recognition datasets can consist of text transcription, semantic annotation, and linguistic features. Text transcription converts audio signals into text form, obtained through deep neural network models that identify and transcribe audio, accurately reflecting the linguistic information in the audio. For example, if the audio says "The weather is nice today," then the language recognition dataset will include this text content. Semantic annotation, in addition to text transcription, adds semantic annotation to the text to better understand its meaning, including distinguishing homophones, analyzing grammatical structures, and recognizing contextual semantics. For example, the annotation for the word "bank" will indicate whether it refers to a financial institution or something else in the current context. Linguistic features are the linguistic characteristics of the audio, such as speech rate, intonation, and pauses. For instance, a faster speech rate indicates the speaker is more anxious, while a slower speech rate indicates the speaker is more relaxed.

[0107] Additionally, it should be noted that deep neural network models can handle speech recognition tasks more efficiently. By inputting diverse audio training data into the model, the model can learn a wider range of speech features and patterns, enabling it to automatically adjust its internal parameters to better adapt to the speech features of different languages ​​and dialects, thereby improving overall performance. Deep neural network models not only perform well on training data but also maintain high performance on unseen speech data, making them adaptable to various real-world application scenarios, including different speakers, speech rates, and background noise.

[0108] Step C13: Extract the audio semantics corresponding to the natural language analysis in the language recognition dataset to determine the intelligent sign language speech text;

[0109] It should be noted that audio semantics can include homonym semantics, grammatical structure semantics, and contextual semantics. Among them, homonym semantics refers to words that have the same pronunciation but different meanings. For example, "flower" can refer to flowers or expenses. It is necessary to correctly distinguish the meanings of homonyms. Grammatical structure semantics refers to the meaning expressed by the order of words and grammatical relationships in a sentence. For example, subject-verb-object structures and topic-comment structures have different expressions in different languages. Contextual semantics refers to the influence of the context on the meaning of words and sentences in a dialogue or text. For example, "bank" can refer to financial institutions, but in the context, it can refer to entities with savings capacity, and does not necessarily refer to financial institutions.

[0110] Step C14: Based on the intelligent sign language speech text and sign language database, sign language is identified, and human pose points are extracted using the sign language detection model to convert the sign language into animation, thus obtaining the sign language animation.

[0111] It should be noted that the sign language detection model analyzes key information such as hand movements, gesture shapes, and body postures in videos or images, and converts visual signals into understandable language information, thereby achieving automatic recognition and translation of sign language. It can accurately capture complex movements and subtle changes in sign language and match them with a preset sign language vocabulary database to achieve efficient and accurate communication.

[0112] Understandably, when converting sign language into animation, it is necessary to consider that sign language has a unique language structure and expression mode, which is significantly different from spoken language. For example, sign language emphasizes spatiality and the continuity of movements, and it needs to combine facial expressions and body postures to fully express semantics. Computer graphics and artificial intelligence technologies can convert intelligent sign language voice text into visual sign language actions. At this time, it is necessary to consider the grammatical structure of sign language, the continuity of gestures, and the naturalness of expression.

[0113] In one feasible implementation, step C14 may include steps D11 to D14:

[0114] Step D11: Extract sign language word and sentence features from the intelligent sign language speech text to determine the sign language vocabulary structure and sentence structure;

[0115] Understandably, sign language vocabulary is based on gestures, using hand shapes, positions, movements, and facial expressions to express different meanings. Sign language vocabulary can be categorized into different parts of speech based on their function and purpose, such as nouns, verbs, adjectives, and adverbs. Furthermore, sign language can combine multiple gestures to form more complex compound words, whose meanings are often not simply the sum of the meanings of each gesture, but rather new semantics are generated through different combinations. In contrast, sign language sentence structure first clarifies the topic of the sentence and then comments or explains it. This structure differs from the subject-verb-object structure in spoken language, emphasizing the organization and order of information, conveying emotions and intonation, expressing grammatical functions such as questions, emphasis, and negation, and can even change the semantics of a sentence.

[0116] Step D12: Match the sign language vocabulary and sentence structures with the corresponding sign language word and sentence structures in the sign language database to obtain the sign language dataset;

[0117] It should be noted that the sign language dataset is a collection of large amounts of sign language-related data used to train and optimize sign language generation models.

[0118] Understandably, sign language datasets can record video clips or still images of sign language actions, showcasing different gestures, movements, and facial expressions. They also provide detailed annotations of the content in the sign language videos or images, including the names of the gestures, the start and end times of the movements, the movement trajectories of body parts, a large number of sign language words and their corresponding Chinese or other language translations, as well as the structure and grammatical rules of sign language sentences, in order to construct a complete sign language expression.

[0119] Step D13: Extract human pose points based on the sign language dataset, and build generator and discriminator models based on the human pose points for adversarial training to generate the initial sign language animation;

[0120] It should be noted that the initial sign language animation is an animated form of sign language movements initially generated based on the sign language dataset and related algorithms during the sign language generation process.

[0121] Understandably, the initial sign language animations showcased the basic forms and processes of sign language movements, including the shape, position, trajectory of the gestures, as well as facial expressions and body postures. However, they lacked naturalness, fluency, and semantic consistency. For example, the movements might appear stiff or disjointed, or the details of some gestures might not be accurate enough to fully reflect the actual expression habits of sign language.

[0122] Additionally, it should be noted that human pose points are two-dimensional or three-dimensional coordinate points of various key parts of the human body detected by computer vision technology in human posture estimation, such as the head, shoulders, elbows, wrists, and knees. These points are used to capture hand and body movements. By analyzing the motion trajectory and positional changes of these key points and extracting human pose points, human movements and postures can be accurately captured. The positions and motion trajectories of key points in sign language movements can be tracked and recorded in real time, thereby achieving accurate analysis and understanding of sign language movements. This improves the smoothness and naturalness of sign language animation, making it closer to real sign language expression.

[0123] Step D14: Adjust the sign language motion parameters and sign language semantic parameters of the initial sign language animation using the reward parameters in the reward mechanism to obtain the sign language animation.

[0124] It should be noted that sign language animation is a type of animation generated using computer graphics and artificial intelligence technologies, used to visually translate speech or text content into sign language gestures in real time.

[0125] Understandably, sign language animation can display sign language actions through virtual characters or animations, enabling hearing-impaired individuals to intuitively obtain information. It can present sign language expressions in a natural, fluent, and easy-to-understand way, and can be generated in real time and displayed synchronously with audio or text content, providing hearing-impaired individuals with immediate information exchange. At the same time, it can be personalized according to different user needs and preferences, such as adjusting the speed, size, and position of gestures.

[0126] Additionally, it's important to note that the reward mechanism is a function that defines the ultimate optimization goal. For example, the goal might be to generate clear, natural, and grammatically correct sign language animations. Through repeated trial and error, the mechanism observes the changes in reward values ​​resulting from different parameter adjustments, learning which parameter combinations can yield higher long-term cumulative rewards. The level of the reward value directly guides the direction of parameter adjustments, favoring parameter adjustments that are expected to bring higher rewards. Reward parameters are pre-set or learnable weighting coefficients used to adjust the balance of different reward items in the reward function, thereby changing the focus of the generated animation. Sign language motion parameters are the underlying driving parameters that directly control the specific posture and movement of the virtual human body model in each frame. By evaluating the generated animation, the sign language motion parameters can be fine-tuned, thereby adjusting the virtual human's physical posture and changing the animation's appearance and trajectory. Sign language semantic parameters are parameters that describe the semantics and structure carried by the sign language actions, defining the language organization and expression methods. These can be specific sign language words or sentences to be expressed.

[0127] Step S22: Extract the semantic features of the intelligent sign language speech text to determine the information to be converted;

[0128] It should be noted that the information to be transformed is structured information extracted, analyzed, and generated from text or speech data using Natural Language Processing (NLP) technology.

[0129] Understandably, the information to be transformed is the original text converted into a format that is easier for machines to understand. It can also convert the text into a digital representation that machines can analyze and interpret, extract meaningful information from the text data, analyze the meaning behind sentences, find similar meanings in different sentences or process words with different meanings, and automatically generate natural language text based on the input information.

[0130] For ease of understanding, we will take the determination of information to be transformed as an example. The information acquisition device is the information acquisition module, the storage device is the memory, and the processing device is the processing module.

[0131] The information acquisition module acquires intelligent sign language speech text, extracts the semantic features of the intelligent sign language speech text, determines the information to be converted, that is, uses natural language processing technology to convert the speech recognition text into subtitles to support multilingual subtitle options, and obtains the information to be converted.

[0132] Step S23: Input the information to be converted into the language model for subtitle conversion and determine the converted text;

[0133] It should be noted that the converted text is the text content used to generate subtitles, which is converted from the original intelligent sign language speech text after natural language processing and language model processing.

[0134] Understandably, the converted text can accurately represent the semantics of the original audio, and through optimization and adjustment, it improves the naturalness and reading experience of the text, providing high-quality subtitle services for people with hearing impairments and significantly enhancing the accessibility viewing experience.

[0135] For ease of understanding, we will take the determination of subtitle text information as an example, where the information acquisition device is the information acquisition module, the storage device is the memory, and the processing device is the processing module.

[0136] The information to be converted is input into a language model for subtitle conversion. The converted text is determined by using a language model, such as GPT, to generate natural and fluent subtitle text, ensuring the accuracy of the subtitle content, and thus obtaining the converted text.

[0137] Step S24: Correct the subtitle display time based on the converted text, and intelligently synchronize the subtitle display to determine the subtitle text information.

[0138] Understandably, the generated subtitle text information can be used to synchronize subtitles with audio content, avoiding subtitle delays or premature appearances, making the reading more natural and smooth, and ensuring consistency between the semantics and emotional expression of the subtitles and audio content, thereby improving the accuracy of information delivery and significantly optimizing the user's visual experience.

[0139] For ease of understanding, we will take the determination of subtitle text information as an example, where the information acquisition device is the information acquisition module, the storage device is the memory, and the processing device is the processing module.

[0140] The process involves acquiring audio rhythm information, aligning the audio time series in the converted text with the subtitle text time series using a dynamic time warping algorithm, generating timestamp information, and then segmenting and adjusting the subtitle text based on the timestamp information to determine the subtitle display text information. This involves combining an NLP model to optimize the segmentation and display of the subtitle text, ensuring that the subtitles are displayed synchronously with the audio content. The subtitle display text information is then input into a natural language model to predict the segmentation and adjustment order of the target text, determining the predicted subtitle text information. Finally, the predicted subtitle text information is adapted and corrected with the audio rhythm information to obtain the subtitle text time information, providing a smooth viewing experience. The subtitle text information is obtained based on the subtitle display text information and the subtitle text time information.

[0141] In one feasible implementation, step S24 may include steps E11 to E15:

[0142] Step E11: Obtain audio rhythm information;

[0143] It should be noted that audio rhythm information is the time regularity information represented in the audio signal.

[0144] Understandably, audio rhythm information can include speech rate, pauses, and intonation variations. Speech rate refers to how fast one speaks, measured in words or syllables per minute. Pauses are brief silences during speech, used to separate sentences or express emphasis. Intonation variations are changes in pitch, used to express emotions, questions, or statements.

[0145] Step E12: Use the dynamic time warping algorithm to align the audio time series in the converted text with the subtitle text time series to generate timestamp information;

[0146] It should be noted that timestamp information is a precise time stamp used to mark the time when an event or data segment occurs, so that subtitles and audio content are displayed synchronously.

[0147] Understandably, timestamp information records the start and end times of each word or sentence in the subtitle text, ensuring that the time the subtitles are displayed on the screen matches the playback time of the audio content precisely. For example, if the audio content starts playing "The weather is nice today" at the 5th second and finishes at the 7th second, the timestamp information will record the start time of the sentence "The weather is nice today" as 5 seconds and the end time as 7 seconds.

[0148] Additionally, it should be noted that the Dynamic Time Warping algorithm is an algorithm used to measure the similarity between two time series. It is particularly suitable for handling time scaling and compression in time series, and is used to align the time series of audio and subtitle text to ensure that the subtitles are displayed synchronously with the audio content. By calculating the optimal matching path between the two time series, it aligns the time points in the audio signal with the time points in the subtitle text. For example, if a word in the audio starts at 5 seconds, but the corresponding word in the subtitle text starts at 5.2 seconds, Dynamic Time Warping will adjust the timestamp of the subtitle text to align it with the time point of the audio signal. Even when the speech rate changes or there are pauses, the subtitles can be accurately synchronized with the audio content. For example, if the speaker pauses for 1 second in the middle of a sentence, Dynamic Time Warping will adjust the timestamp of the subtitle text so that it is not displayed during the pause, ensuring the synchronization of the subtitles with the audio content.

[0149] Step E13: Based on the timestamp information, the subtitle text is segmented and adjusted to determine the subtitle display text information;

[0150] It should be noted that the subtitle text information is a textual representation of the audio content, including optimizations in semantics, syntax, and context.

[0151] Understandably, segmentation adjustment involves dividing the subtitle text into multiple logically independent paragraphs or sentence units based on the semantic structure and rhythm information of the audio content during the subtitle text generation process. The display time of each unit is then optimized to ensure that the subtitle text is semantically complete, grammatically correct, easy for users to read and understand, and consistent with the playback rhythm of the audio content. This improves the readability and synchronization of the subtitles, while maintaining the logical relationship between paragraphs or sentences to avoid semantic breaks caused by segmentation, thereby optimizing the user's visual experience.

[0152] Step E14: Input the subtitle display text information into the natural language model to predict the target text segmentation and adjustment order, and determine the subtitle prediction text information;

[0153] It should be noted that the predicted subtitle text information is the subtitle text predicted by the natural language model after analyzing and optimizing the subtitle display text information. It is a version of the subtitle text that is more in line with the user's reading habits and display needs.

[0154] Understandably, the subtitle text information can be input into a natural language model. This model could be a deep learning model, such as Transformer or BERT, trained on a large amount of text data. It would be used to understand the semantic structure, grammatical relationships, and contextual information of the text. Based on the audio rhythm information and the semantic structure of the subtitle text, it would segment and adjust the subtitle text. For example, it could break longer sentences into shorter paragraphs or merge related sentences to improve readability. Simultaneously, it would optimize the semantics of the subtitle text to ensure that each subtitle segment is semantically complete and clear. For example, it could adjust the sentence structure to better conform to the expression habits of natural language while avoiding ambiguity and redundancy. Furthermore, it could predict the optimal display order of the subtitle text. For example, if there is an important information point in the audio, it would predict to display the relevant subtitle content at that point in time to ensure that users can obtain information in a timely manner.

[0155] Step E15: Adapt and correct the subtitle display time by matching the predicted subtitle text information with the audio rhythm information to obtain the subtitle text time information.

[0156] It should be noted that the subtitle text timing information represents the specific point in time and duration at which the subtitle content is displayed on the screen.

[0157] Understandably, correcting the subtitle display time involves using technical means to adjust the display time of the subtitles so that it is precisely synchronized with the playback time of the audio content. This requires not only considering the audio rhythm, but also the speech rate and pauses, to ensure that the appearance and disappearance time of the subtitles match the audio content and avoid subtitles being delayed or displayed prematurely.

[0158] Additionally, it should be noted that the subtitle text timing information can include the start time, end time, and duration of the subtitle on the screen, ensuring precise synchronization between the subtitle and the audio content playback time, and guaranteeing that the semantics and emotional expression of the subtitle and audio content match.

[0159] This embodiment proposes a display method that identifies audio semantics based on the audio information to be processed and a deep neural network model, and uses a sign language detection model to extract human pose points for sign language animation conversion to determine the sign language animation; extracts semantic features of intelligent sign language speech text to determine the information to be converted; inputs the information to be converted into a language model for subtitle conversion to determine the converted text; corrects the subtitle display time based on the converted text, and intelligently synchronizes the display of subtitles to determine the subtitle text information. This solves the technical problem of how to achieve accessible television display using artificial intelligence. Compared with existing technologies, this application accurately identifies audio semantics, generates intelligent sign language speech text and converts it into natural and fluent sign language animation, and uses intelligent sign language speech text to recognize subtitles for intelligent synchronized display, enhancing the system's multilingual support capabilities, accurately recognizing system speech, translating sign language in real time, and synchronously displaying subtitles, providing an accessible television viewing experience for people with disabilities.

[0160] This application also provides a display device, please refer to... Figure 5 The display device includes:

[0161] The acquisition module 10 is used to acquire the audio information to be processed and the image information to be adjusted;

[0162] The processing module 20 is used to identify audio semantics based on the audio information to be processed and a deep neural network model, and to convert intelligent sign language speech text to generate sign language animation and subtitle text information;

[0163] The execution module 30 is used to determine the TV display interface information based on the audio information to be processed, sign language animation, subtitle text information and screen information to be adjusted, and to control the TV accessibility display based on the TV display interface information.

[0164] The acquisition module 10 is also used to acquire microphone audio information and television screen information;

[0165] The microphone audio information is preprocessed to obtain the audio information to be processed. The preprocessing includes adaptive noise reduction and sound quality enhancement.

[0166] Color analysis tools are used to analyze the color of the television screen information, and the television screen information is divided into color regions to be adjusted according to the region segmentation image processing algorithm to determine the color element region information.

[0167] The color conversion algorithm is used to extract the image color features corresponding to the color element region information, and to determine the color feature information and context information.

[0168] Color feature information and context information are concatenated into an input vector in a fixed order. The input vector is then fed into a machine learning model to predict color adjustment combinations and generate a set of color adjustment schemes.

[0169] Version testing was conducted on the color adjustment schemes in the set to evaluate user satisfaction. The target color adjustment scheme with the highest user satisfaction was selected to perform color conversion on the color element area information to obtain the image information to be adjusted.

[0170] The processing module 20 is also used to identify audio semantics based on the audio information to be processed and the deep neural network model, and to extract human pose points using the sign language detection model to perform sign language animation conversion and determine the sign language animation.

[0171] Extract semantic features from intelligent sign language speech text to determine the information to be converted;

[0172] Input the information to be converted into a language model for subtitle conversion and determine the converted text;

[0173] The subtitle display time is corrected based on the converted text, and the subtitles are displayed in a smart and synchronized manner to determine the subtitle text information.

[0174] The processing module 20 is also used to acquire audio training condition information, which includes language training condition information and dialect training condition information.

[0175] The audio training process information is input into the deep neural network model to identify the natural language in the audio information to be processed, and the language recognition dataset is determined.

[0176] Extract the audio semantics corresponding to the natural language analysis in the language recognition dataset to determine the intelligent sign language speech text;

[0177] Sign language is recognized based on intelligent sign language speech text and sign language database, and human pose points are extracted using sign language detection model to convert sign language into animation.

[0178] The processing module 20 is also used to extract sign language word and sentence features from intelligent sign language speech text, and to determine the sign language vocabulary structure and sentence structure;

[0179] A sign language dataset is obtained by matching the sign language vocabulary and sentence structure with the corresponding sign language vocabulary and sentence structures in the sign language database.

[0180] Human pose points are extracted from the sign language dataset, and generator and discriminator models are built based on the human pose points for adversarial training to generate the initial sign language animation.

[0181] The sign language animation is obtained by adjusting the sign language motion parameters and sign language semantic parameters of the initial sign language animation using the reward parameters in the reward mechanism.

[0182] Processing module 20 is also used to acquire audio rhythm information;

[0183] The dynamic time warping algorithm is used to align the audio time series in the converted text with the subtitle text time series to generate timestamp information.

[0184] The subtitle text is segmented and adjusted based on timestamp information to determine the subtitle display text information;

[0185] Input the subtitle display text information into the natural language model to predict the target text segmentation and adjust the order, and determine the subtitle prediction text information;

[0186] The predicted text information of the subtitles is adapted and corrected with the audio rhythm information to obtain the subtitle display time information.

[0187] The execution module 30 is also used to acquire user interaction information, including interface feedback information, personalized setting feedback information and user feedback information;

[0188] Use data analysis tools to categorize user interaction information and identify user needs for feature improvement;

[0189] Data analysis tools were used to categorize interface feedback, personalized settings feedback, and user feedback according to user improvement needs, and to determine the adjustment information. User improvement needs included subtitle display settings, voice settings, sign language animation settings, and position settings.

[0190] Extract the display parameters corresponding to different user improvement requests from the adjustment information, and adjust the parameters in the TV display interface information based on the display parameters to determine the target TV display interface information;

[0191] Control the TV display interface configuration based on the target TV display interface information to achieve barrier-free TV display.

[0192] The display device provided in this application, employing the display method described in the above embodiments, can solve the technical problem of how to achieve barrier-free television display using artificial intelligence. Compared with the prior art, the beneficial effects of the display device provided in this application are the same as those of the display method provided in the above embodiments, and other technical features in the display device are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.

[0193] This application provides a display device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, which are executed by the at least one processor to enable the at least one processor to perform the display method in the first embodiment described above.

[0194] The following is for reference. Figure 6The diagram illustrates a structural schematic of a display device suitable for implementing embodiments of this application. The display device in the embodiments of this application may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 6 The display device shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.

[0195] like Figure 6 As shown, the display device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in ROM (Read Only Memory) 1002 or a program loaded from storage device 1003 into RAM (Random Access Memory) 1004. RAM 1004 also stores various programs and data required for the operation of the display device. The processing unit 1001, ROM 1002, and RAM 1004 are interconnected via bus 1005. Input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to I / O interface 1006: input devices 1007 including, for example, touch screens, touchpads, keyboards, mice, image sensors, microphones, accelerometers, gyroscopes, etc.; output devices 1008 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 1003 including, for example, magnetic tapes, hard disks, etc.; and communication devices 1009. Communication device 1009 allows the display device to communicate wirelessly or wiredly with other devices to exchange data. Although the figures show display devices with various systems, it should be understood that it is not required to implement or possess all of the systems shown. More or fewer systems may be implemented alternatively.

[0196] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from ROM 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.

[0197] The display device provided in this application, employing the display method described in the above embodiments, can solve the technical problem of how to achieve barrier-free television display using artificial intelligence. Compared with the prior art, the beneficial effects of the display device provided in this application are the same as those of the display method provided in the above embodiments, and other technical features of the display device are the same as those disclosed in the previous embodiment method, and will not be repeated here.

[0198] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.

[0199] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0200] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, the computer-readable program instructions being used to execute the display method in the above embodiments.

[0201] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.

[0202] The aforementioned computer-readable storage medium may be included in the display device or may exist independently without being assembled into the display device.

[0203] The aforementioned computer-readable storage medium carries one or more programs that, when executed by a display device, cause the display device to: acquire audio information to be processed and screen information to be adjusted; recognize audio semantics based on the audio information to be processed and a deep neural network model, and convert intelligent sign language speech text to generate sign language animation and subtitle text information; determine television display interface information based on the audio information to be processed, sign language animation, subtitle text information, and screen information to be adjusted, and control the television for accessible display based on the television display interface information.

[0204] Computer program code for performing the operations of this application can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0205] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0206] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.

[0207] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for performing the above-described display method, and can solve the technical problem of how to achieve barrier-free display on television using artificial intelligence. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as the beneficial effects of the display method provided in the above embodiments, and will not be repeated here.

[0208] The above are only some embodiments of this application and do not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.

Claims

1. A display method, characterized in that, The method includes: Acquire the audio information to be processed and the image information to be adjusted; Based on the audio information to be processed and the deep neural network model, the audio semantics are identified, and the intelligent sign language speech text is converted to generate sign language animation and subtitle text information; Based on the audio information to be processed, the sign language animation, the subtitle text information, and the screen information to be adjusted, the TV display interface information is determined, and the TV accessibility display is controlled based on the TV display interface information.

2. The method as described in claim 1, characterized in that, The steps of acquiring the audio information to be processed and the image information to be adjusted include: Acquire microphone audio information and television image information; The microphone audio information is preprocessed to obtain audio information to be processed. The preprocessing includes adaptive noise reduction processing and sound quality enhancement processing. The television image information is analyzed using a color analysis tool, and the television image information is divided into color regions to be adjusted according to a region segmentation image processing algorithm to determine the color element region information. The color conversion algorithm is used to extract the image color features corresponding to the color element region information, and to determine the color feature information and context information. The color feature information and the context information are concatenated into an input vector in a fixed order. The input vector is then fed into a machine learning model to predict color adjustment combinations, thereby generating a set of color adjustment schemes. The user satisfaction of the proposed color adjustment schemes is evaluated by testing the schemes in the set of color adjustment schemes. The target color adjustment scheme with the highest user satisfaction is then selected to perform color conversion on the color element area information to obtain the image information to be adjusted.

3. The method as described in claim 1, characterized in that, The steps of recognizing audio semantics based on the audio information to be processed and a deep neural network model, converting intelligent sign language speech text, and generating sign language animation and subtitle text information include: Based on the audio information to be processed and the deep neural network model, the audio semantics are identified, and the human pose points are extracted using the sign language detection model to perform sign language animation conversion, thereby determining the sign language animation. Extract the semantic features of the intelligent sign language speech text to determine the information to be converted; The information to be converted is input into a language model for subtitle conversion, and the converted text is determined. The subtitle display time is corrected based on the converted text, and the subtitles are displayed in a smart and synchronized manner to determine the subtitle text information.

4. The method as described in claim 3, characterized in that, The steps of identifying audio semantics based on the audio information to be processed and the deep neural network model, and extracting human pose points using a sign language detection model to perform sign language animation conversion, and determining the sign language animation, include: Acquire audio training status information, which includes language training status information and dialect training status information. The audio training information is input into a deep neural network model to identify the natural language in the audio information to be processed, and the language recognition dataset is determined. Extract the audio semantics corresponding to the natural language analysis from the language recognition dataset to determine the intelligent sign language speech text; Based on the intelligent sign language speech text and sign language database, sign language is identified, and human pose points are extracted using a sign language detection model to convert them into sign language animation, thus obtaining sign language animation.

5. The method as described in claim 4, characterized in that, The steps of recognizing sign language based on the intelligent sign language speech text and sign language database, and extracting human pose points using a sign language detection model to convert them into sign language animation, to obtain the sign language animation, include: Extract sign language word and sentence features from the intelligent sign language speech text to determine the sign language vocabulary structure and sentence structure; A sign language dataset is obtained by matching the sign language vocabulary structure and sentence structure with the corresponding sign language vocabulary and sentence structures in the sign language database. Human pose points are extracted based on the sign language dataset, and generator and discriminator models are constructed based on the human pose points for adversarial training to generate initial sign language animation. The sign language motion parameters and sign language semantic parameters of the initial sign language animation are adjusted using the reward parameters in the reward mechanism to obtain the sign language animation.

6. The method as described in claim 3, characterized in that, The steps of correcting the subtitle display time based on the converted text and intelligently synchronizing the subtitle display to determine the subtitle text information include: Obtain audio rhythm information; The audio time series in the converted text is aligned with the subtitle text time series using a dynamic time warping algorithm to generate timestamp information. Based on the timestamp information, the subtitle text is segmented and adjusted to determine the subtitle display text information; The subtitle display text information is input into a natural language model to predict the segmentation and adjustment order of the target text, thereby determining the subtitle prediction text information; The predicted text information of the subtitles is adapted and corrected with the audio rhythm information to adjust the subtitle display time, thereby obtaining the subtitle text time information.

7. The method as described in claim 1, characterized in that, The steps of controlling the TV's accessibility display based on the TV display interface information include: Acquire user interaction information, which includes interface feedback information, personalized setting feedback information, and user feedback information; The user interaction information is categorized using data analysis tools to identify user needs for feature improvement. Using data analysis tools, the interface feedback information, the personalized settings feedback information, and the user feedback information are categorized according to user improvement needs to determine adjustment information. The user improvement needs include subtitle display settings, voice settings, sign language animation settings, and position settings. Extract the display parameters corresponding to different user improvement needs from the adjustment information, and adjust the parameters in the TV display interface information based on the display parameters to determine the target TV display interface information; Based on the target TV display interface information, the TV configuration display interface is controlled to achieve barrier-free TV display.

8. A display device, characterized in that, The device includes: The acquisition module is used to acquire the audio information to be processed and the image information to be adjusted. The processing module is used to identify audio semantics based on the audio information to be processed and the deep neural network model, and to convert intelligent sign language speech text to generate sign language animation and subtitle text information; The execution module is used to determine the TV display interface information based on the audio information to be processed, the sign language animation, the subtitle text information, and the screen information to be adjusted, and to control the TV accessibility display based on the TV display interface information.

9. A display device, characterized in that, The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the display method as described in any one of claims 1 to 7.

10. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the steps of the display method as described in any one of claims 1 to 7.