DATA STORAGE DEVICE AND METHOD FOR GESTURE GENERATION AND MANAGEMENT

The data storage device generates and manages gesture videos from audio-visual content to address the limitations of subtitles and sign language videos, offering a customizable and cost-effective immersive experience for hearing-impaired viewers.

DE102025115261A1Pending Publication Date: 2026-01-08SANDISK TECHNOLOGIES LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
DE102025115261
Authority / Receiving Office
DE · DE
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-08
Filing Date
2025-04-17
Publication Date
2026-01-08

AI Technical Summary

Technical Problem

People with hearing impairments face challenges in fully enjoying audio-visual content due to impaired audio comprehension, as subtitles often fail to convey nonverbal sounds, and sign language videos in picture-in-picture format can be distracting and costly to produce.

Method used

A data storage device generates gesture videos, such as sign language, from audio-visual content using artificial intelligence, and overlays them as a picture-in-picture window during playback, allowing users to customize their viewing experience.

Benefits of technology

Enhances the media experience for hearing-impaired individuals by providing clear non-verbal cues without additional screen space usage, reducing production costs, and optimizing resource utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A data storage device and a method for gesture generation and management are provided. In one embodiment, a data storage device is provided comprising memory and one or more processors. The one or more processors are configured individually or in combination to: extract subtitles from a video stored in memory; generate a gesture video from the subtitles; create a combined video comprising the generated gesture video combined with the video; and store the combined video in memory. Other embodiments are provided.
Need to check novelty before this filing date? Find Prior Art

Description

STATE OF THE ART

[0001] Due to impaired audio comprehension, people with hearing loss often find it difficult to enjoy audio-visual content. Without audio, the media experience may not be immersive enough. Subtitles are often provided alongside media to help viewers understand the speech in the audio. However, nonverbal sounds, such as a ringing telephone, are frequently not conveyed when using subtitles. Instead of subtitles, a person with hearing loss may prefer to see the text communication in sign language or gestures for deaf / mute individuals. Some specialized media content (e.g., certain news broadcasts) displays a video with sign language interpretation alongside the main content in a picture-in-picture (PIP) window. BRIEF DESCRIPTION OF THE DRAWINGS Fig. Figure 1A is a block diagram of a data storage device of one embodiment. Fig. Figure 1B is a block diagram illustrating a storage module of one embodiment. Fig. 1C is a block diagram illustrating a hierarchical storage system of one embodiment. Fig. 2A is a block diagram showing the components of the controller in Fig. Figure 1A illustrates a data storage device according to one embodiment. Fig. 2B is a block diagram showing components of the Fig. Figure 1A illustrates a data storage device according to one embodiment. Fig. Figure 3 is a block diagram of a host and a data storage device of an embodiment. Fig. Figure 4 is a block diagram of a system embodiment for overlaying gesture videos as picture-in-picture onto an original video. Fig. Figure 5 is a block diagram of a system of embodiment for providing gesture videos as a video track separate from an original video. Fig. 6 is a block diagram of a media container file of one embodiment. DETAILED DESCRIPTION

[0002] The following embodiments relate generally to a data storage device and a method for gesture generation and management. In one embodiment, a data storage device is provided comprising memory and one or more processors. The one or more processors are configured individually or in combination to: extract subtitles from a video stored in memory; generate a gesture video from the subtitles; create a combined video comprising the generated gesture video combined with the video; and store the combined video in memory.

[0003] In some embodiments, the combined video is created by: caching the generated gesture video in memory; decoding the generated gesture video and the video in respective frame buffers; and embedding the generated gesture video into the video.

[0004] In some embodiments, the combined video is managed as a data storage device-specific file that is abstracted to a host.

[0005] In some embodiments, the combined video is made accessible to a host and stored in a logical memory location.

[0006] In some embodiments, one or more processors are further configured individually or in combination to trim the gesture video based on a parameter.

[0007] In some embodiments, the one or more processors, individually or in combination, are further configured to store the combined video using a flash translation layer biasing scheme.

[0008] In some embodiments, the generated gesture video is combined with the video as a picture-in-picture video and overlaid on the video.

[0009] In some implementations, the subtitles are extracted from a Moving Picture Experts Group (MPEG) data stream.

[0010] In some versions, the subtitles are extracted from individual video frames.

[0011] In some versions, the gestures are generated from the subtitles using an artificial intelligence model.

[0012] In some embodiments, the data storage device includes a network-connected storage server.

[0013] In some embodiments, the memory includes a three-dimensional memory.

[0014] In a further embodiment, a method is provided that is carried out in a data storage device comprising a memory. The method comprises: extracting text information associated with a video stored in the memory; generating a gesture video from the extracted text information; and storing the gesture video as a new program in an existing transport data stream for the video.

[0015] In some versions, the gesture video is overlaid as a picture-in-picture video on top of the main video.

[0016] In some embodiments, the gestures are generated from the extracted text information using an artificial intelligence model.

[0017] In some embodiments, the transport data stream includes a Moving Picture Experts Group (MPEG) transport data stream.

[0018] In some embodiments, the data storage device includes a network-connected storage server.

[0019] In some embodiments, the data storage device includes an edge node of a content delivery network.

[0020] In some embodiments, the data storage device includes an over-the-top platform.

[0021] In a further embodiment, a data storage device is provided comprising: a memory; and means for: extracting subtitles from a video stored in the memory; generating a gesture video from the subtitles; and causing the gesture video to be overlaid on the video during playback.

[0022] Other embodiments are possible, and each embodiment can be used alone or in combination with the others. Accordingly, various embodiments will now be described with reference to the accompanying drawings. Designs

[0023] The following embodiments relate to a data storage device (DSD). As used herein, “data storage device” means a non-volatile device that stores data. Examples of DSDs include, but are not limited to, hard disk drives (HDDs), solid state drives (SSDs), tape drives, hybrid drives, etc. Details of exemplary DSDs are provided below.

[0024] Examples of data storage devices suitable for use in implementing aspects of these embodiments are given in the Fig. Examples 1A to 1C are shown. It should be noted that these are only examples and other implementations can also be used. Fig. Figure 1A is a block diagram illustrating the data storage device 100 according to one embodiment. With reference to Fig. 1A The data storage device 100 includes a controller 102 coupled to a non-volatile memory, which may consist of one or more non-volatile memory dies 104. As used herein, the term “die” refers to the collection of non-volatile memory cells and associated circuit arrangements for handling the physical operation of these non-volatile memory cells, formed on a single semiconductor substrate. The controller 102 is connected to a host system and transmits command sequences for read, program, and erase operations to the non-volatile memory die 104. As used herein, the expressions “in communication with” or “coupled with” may also mean directly in communication with / coupled with, or indirectly in communication with / coupled with, through one or more components, which may or may not be shown or described herein.Communication / coupling can be wired or wireless.

[0025] The Controller 102 (which may be a non-volatile memory controller (e.g., a flash controller, a resistive random-access memory (ReRAM) controller, a phase-change memory (PCM) controller, or a magnetoresistive random-access memory (MRAM) controller)) may include one or more components configured individually or in combination to perform specific functions, including but not limited to those described herein and illustrated in the flowcharts. As shown in Fig. As shown in Figure 2A, the controller 102 can, for example, comprise one or more processors 138 configured individually or in combination to perform functions such as, but not limited to, those described herein and illustrated in the flowcharts by executing computer-readable program code stored in one or more non-transitory memories 139 within the controller 102 and / or outside the controller 102 (e.g., in random access memory (RAM) 116 or read-only memory (ROM) 118). As another example, the one or more components can include circuit arrangements such as, but not limited to, logic gates, switches, an application-specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller.

[0026] In one embodiment, the non-volatile memory controller 102 is a device that manages data stored in non-volatile memory and communicates with a host such as a computer or electronic device running any suitable operating system. A non-volatile memory controller 102 may have various functionalities in addition to the specific functionality described herein. For example, the non-volatile memory controller may format the non-volatile memory to ensure that the memory functions properly, detect faulty non-volatile memory cells, and allocate replacement cells to replace future failed cells.A portion of the spare cells can be used to store firmware (and / or other metadata used for management and tracking) to operate the non-volatile memory controller and implement other features. When a host needs to read data from or write data to the non-volatile memory, it can communicate with the non-volatile memory controller. If the host provides a logical address for data to be read from or written to, the non-volatile memory controller can translate the logical address received from the host into a physical address within the non-volatile memory.The non-volatile memory controller can also perform various memory management functions, such as wear leveling (distributing write operations to avoid the wear and tear of certain memory blocks that would otherwise be repeatedly written to) and automatic memory cleanup (when a block is full, only the valid data pages are moved to a new block so that the full block can be deleted and reused).

[0027] The non-volatile memory die 104 can include any suitable non-volatile storage medium, including resistive random-access memory (ReRAM), magnetoresistive random-access memory (MRAM), phase-change memory (PCM), NAND flash memory cells, and / or NOR flash memory cells. The memory cells can be in the form of solid-state memory cells (e.g., flash memory cells) and can be programmable once, multiple times, or many times. The memory cells can also be single-level cells (SLC), multiple-level cells (MLC) (e.g., dual-level cells, triple-level cells (TLC), quad-level cells (QLC), etc.), or use technologies with other memory cell levels now known or later developed. Furthermore, the memory cells can be fabricated in a two-dimensional or three-dimensional manner.

[0028] The interface between controller 102 and non-volatile memory 104 can be any suitable flash interface such as Toggle-Mode 200, 400, or 800. In one embodiment, the data storage device 100 can be a card-based system such as a Secure Digital card (SD card) or a Micro Secure Digital card (Micro-SD card). In an alternative embodiment, the data storage device 100 can be part of an embedded data storage device.

[0029] Although in the Fig. As illustrated in Example 1A, the data storage device 100 (hereafter sometimes referred to as a storage module) includes a single channel between the controller 102 and the non-volatile memory die 104. However, the subject matter described here is not limited to having a single storage channel. For example, in some architectures (such as those in the Fig. 1B and Fig. Depending on the controller functions (as shown in Figure 1C), two, four, eight, or more memory channels may be present between the controller and the memory device. In all embodiments described here, more than one channel may be present between the controller and the memory die, even if only a single channel is shown in the drawings.

[0030] Fig. Figure 1B illustrates a storage module 200 that includes multiple non-volatile data storage devices 100. As such, the storage module 200 can include a storage controller 202, which is connected to a host, and to the data storage device 204, which includes a plurality of data storage devices 100. The interface between the storage controller 202 and the data storage devices 100 can be a bus interface, such as a Serial Advanced Technology Attachment (SATA) interface, a Peripheral Component Interconnect Express (PCIe) interface, a Double Data Rate (DDR) interface, or a Serial Attached Small-Scale Compute (SAS / SCSI) interface.The storage module 200 can, in one embodiment, be a solid-state storage device (SSD) or a non-volatile dual in-line storage module (NVDIMM), such as those found in server PCs or portable data processing devices like laptop computers and tablet computers.

[0031] Fig. Figure 1C is a block diagram illustrating a hierarchical storage system. A hierarchical storage system 250 includes a plurality of storage controllers 202, each of which controls a corresponding data storage device 204. Host systems 252 can access storage within the storage system 250 via a bus interface. In one embodiment, the bus interface can be a Non-Volatile Memory Express (NVMe) interface or a Fibre Channel over Ethernet (FCoE) interface. In one embodiment, the Fig. System 1C illustrated is a rack-mountable mass storage system that can be accessed by multiple host computers, such as can be found in a data center or other location where mass storage is needed.

[0032] With renewed reference to Fig. 2A In this example, the controller 102 also includes a front-end module 108, which provides an interface to a host, a back-end module 110, which provides an interface to one or more non-volatile memory die(s) 104, and various other components or modules, such as, but not limited to, a buffer manager / bus controller module, which manages buffers in RAM 116 and controls the internal bus arbitration of the controller 102. A module can include one or more processors or components, as described above. The ROM 118 can store the system startup code. Although they are in Fig. In Figure 2A, the RAM 116 and ROM 118 are shown to be arranged separately from the controller 102. In other embodiments, one or both of the RAM 116 and ROM 118 may be located inside the controller 102. In still other embodiments, sections of the RAM 116 and ROM 118 may be located both inside and outside the controller 102.

[0033] The front-end module 108 includes a host interface 120 and a physical layer interface (PHY) 122, which provide the electrical interface with the host or the next-level storage controller. The choice of host interface 120 type can depend on the type of storage used. Examples of host interfaces 120 include, but are not limited to, SATA, SATA Express, Serially Attached Small Computer System Interface (SAS), Fibre Channel, Universal Serial Bus (USB), PCIe, and NVMe. The host interface 120 typically facilitates the transmission of data, control signals, and timing signals.

[0034] The back-end module 110 includes an error correction code engine (ECC engine) 124, which encodes the data bytes received from the host and decodes the data bytes read from the non-volatile memory, correcting errors. A command sequencer 126 generates command sequences, such as program and erase command sequences, to be transferred to the non-volatile memory die 104. A RAID module (RAID = Redundant Array of Independent Drives) 128 manages the generation of RAID parity and the recovery of corrupted data. RAID parity can be used as an additional layer of integrity protection for the data written to the storage device 104. In some cases, the RAID module 128 may be part of the ECC engine 124. A memory interface 130 provides the instruction sequences to the non-volatile memory die 104 and receives status information from the non-volatile memory die 104.In one embodiment, the storage interface 130 can be a double data rate (DDR) interface such as a toggle mode 200, 400, or 800 interface. In this example, the controller 102 also includes a media management layer 137 and a flash control layer 132, which controls the overall operation of the back-end module 110.

[0035] The data storage device 100 also includes other separate components 140, such as external electrical interfaces, external RAM, resistors, capacitors, or other components that may be connected to the controller 102. In alternative embodiments, one or more of the physical layer interface 122, the RAID module 128, the media management layer 138, and the buffer management / bus controller are optional components that are not required in the controller 102.

[0036] Fig. Figure 2B is a block diagram illustrating the components of the non-volatile memory die 104 in more detail. The non-volatile memory die 104 includes a peripheral circuitry 141 and a non-volatile memory array 142. The non-volatile memory array 142 includes the non-volatile memory cells used to store data. The non-volatile memory cells can be any suitable non-volatile memory cells, including ReRAM, MRAM, PCM, NAND flash memory cells, and / or NOR flash memory cells in a two-dimensional and / or three-dimensional configuration. The non-volatile memory die 104 further includes a data buffer 156, which temporarily stores data and address decoders 148 and 150. The peripheral circuitry 141 in this example includes a state machine 152, which provides status information to the controller 102.The peripheral circuit arrangement 141 may also include one or more components configured individually or in combination to perform certain functions, including but not limited to those described herein and illustrated in the flowcharts. As in . Fig. As shown in Figure 2B, the memory die 104 can, for example, include one or more processors 168 configured individually or in combination to execute computer-readable program code stored in one or more non-transitory memories 169 located in the memory array 142 or outside the memory die 104.

[0037] As another example, the one or more components may include circuit arrangements such as, but not limited to, logic gates, switches, an application-specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller.

[0038] In addition to or instead of the one or more processors 138 (or more generally, components) in the controller 102 and the one or more processors 168 (or more generally, components) in the storage device 104, the data storage device 100 may include another set of one or more processors (or more generally, components). In general, one or more processors (or more generally, components) in the data storage device 100, regardless of their location and number, may be configured individually or in combination to perform various functions, including, but not limited to, those described herein and illustrated in the flowcharts. For example, the one or more processors (or components) may be located in the controller 102, in the storage device 104, and / or elsewhere in the data storage device 100.Furthermore, different functions can be performed using different processors (or components) or combinations of processors (or components). Additionally, means for performing a function can be implemented with a controller that includes one or more components (e.g., processors or the other components described above).

[0039] With renewed reference to Fig. 2A handles the Flash Control Layer 132 (referred to here as the Flash Translation Layer (FTL)) and flash errors, and is connected to the host. Specifically, the FTL, which may be an algorithm within the firmware, is responsible for internal memory management operations and translates write operations from the host into writes to memory 104. The FTL may be necessary because memory 104 may have a limited lifespan, be writeable only in multiples of pages, and / or be unwriteable unless erased as a block. The FTL understands these potential limitations of memory 104, which may not be visible to the host. Accordingly, the FTL attempts to translate write operations from the host into writes to memory 104.

[0040] The FTL can include a logic-to-physical address mapping (L2P mapping) (sometimes referred to herein as a table or data structure) and allocated cache memory. In this way, the FTL translates logical block addresses (“LBAs”) from the host to physical addresses in memory 104. The FTL may include other features such as, but not limited to, power-off recovery (so that the FTL's data structures can be restored in the event of a sudden power loss) and wear leveling (so that wear is even across memory blocks to prevent certain blocks from becoming excessively worn, which would lead to a higher probability of failure).

[0041] Referring again to the drawings, Fig. Figure 3 shows a block diagram of a host 300 and a data storage device 100 of an embodiment. The host 300 can take any suitable form, including, but not limited to, a computer, a mobile phone, a tablet, a wearable device, a digital video recorder, a surveillance system, etc. The host 300 in this embodiment (here a computing device) comprises one or more processors 330 and one or more memories 340. In one embodiment, computer-readable program code stored in one or more memories 340 configures the one or more processors 330 to perform, individually or in combination, the actions described herein as being performed by the host 300. Therefore, actions performed by the host 300 are sometimes considered herein to be performed by an application (computer-readable program code) running on the host 300.For example, host 300 can be configured to send data (which may initially be stored in memory 340 of the host) to data storage device 100 for storage in memory 104 of the data storage device.

[0042] As mentioned above, a lack of audio comprehension makes it difficult for people with hearing impairments to enjoy audio and video content. Without audio, the media experience may not be immersive enough. Often, subtitles (e.g., closed captions) are provided with the media so that the viewer can understand the speech in the audio. However, nonverbal sounds, such as a ringing telephone, are frequently not conveyed when using subtitles. Instead of subtitles, a viewer with a hearing impairment may prefer to see the text communication in sign language or gestures for deaf / mute individuals.

[0043] Some specialized media content (e.g., certain news broadcasts) displays a sign language video alongside the main content in a picture-in-picture (PIP) window. However, the PIP window takes up screen space and can be distracting for the viewer without obstructing the view of the main video. For this reason, viewers may be wary of this type of specialized content. Furthermore, producing the sign language video requires specialized personnel and additional effort. These issues can result in prohibitive costs for content producers.

[0044] In addition to subtitles / closed captions (which are added to the media, available in different languages, and can be turned on and off by the user) and captions for the hearing impaired (which describe nonverbal sounds in text form), a hearing aid can be used to enable hearing-impaired people to enjoy audio and video media. These devices deliver the audio directly into the person's ear, filtering out background noise and making the audio easier to understand.

[0045] These devices are useful for people who suffer from hearing loss but do not have complete hearing loss.

[0046] The following embodiments provide a data storage device and a method for gesture generation and management that can enable a hearing-impaired person to enjoy audio-video media. The data storage device can take any suitable form, such as those described in the examples above, but also others, such as a network-attached storage (NAS) device (sometimes referred to herein as a "storage server") and an edge node of a content delivery network (CDN), without limitation. It should be understood that these are merely examples and other types of data storage devices may also be used. Therefore, a specific type of data storage device should not be read into the claims unless expressly stated therein.

[0047] In general, in these embodiments, gesture videos (e.g., sign language) are automatically generated by the data storage device 100 from audio-video media content in a regular format. The gesture video can be generated, for example, using a generative artificial intelligence (AI) model. The gesture video can be displayed to the user in a picture-in-picture (PIP) window on the user's display device, at another location on the user's display device, on a secondary display device, etc. In one embodiment, the PIP window can be turned on or off by sending a request to the data storage device 100. For example, the hearing-impaired viewer can request the media content with gesture video, and the controller 102 of the data storage device 100 can generate the gesture video using the subtitle text information available in the media.In this way, a viewer can activate and deactivate the gesture video PIP window, with the data storage device 100 providing a mechanism to enable / disable the operation according to the viewer's requirements. For example, if the user accesses the media via a REST API (Representational State Transfer) or an HTTP protocol (Hypertext Transfer Protocol), the operation can be controlled using an API parameter.

[0048] It should be noted that the general concept of subtitle extraction from an MPEG data stream is well-known, as are sign language translation systems that generate gestures from text. However, in these embodiments, the generation and management of the gestures are performed in the data storage device 100, which increases the efficiency of the entire operating environment and its management based on the end-user's needs. Furthermore, data storage device-specific data storage techniques offer several advantages for the user, as explained below.

[0049] In one embodiment, the controller 102 of the data storage device 100 (e.g., a storage server system, NAS, a data storage device used in the cloud, or an over-the-top (OTT) system, etc.) has a model that extracts the text from the subtitles in an MPEG (Moving Picture Experts Group) data stream or from video frames in memory 104, uses a model to generate a gesture (e.g., sign language) from the extracted text, adds the generated gesture as another video program in a PIP format to the transport data stream (TS) (e.g., gesture image in an actual video image), manages the new data as device-specific data or as new logical data transparent to the host 300, and updates the configuration to play back the new transport data stream during video playback based on a host request with PIP.This data storage device 100 enables user-customized gesture playback as well as video playback (in PIP format). In this example, a gesture is an image sequence that is played back as a secondary video in PIP format.

[0050] An artificial intelligence (AI) model can be used to generate gestures from subtitle information. A generative AI model for video can require significant computing resources for training and inference. Inference resources can be problematic if the model is hosted on storage servers with varying capacities. Inference time is also critical, as audio / video presentations are typically time-limited. The generative AI model can be developed in-house, or an open-source model can be adapted to the specific requirements. This model can be trained by providing a large dataset of input text-to-video pairs. The input video can be low-resolution, as the output video does not require high resolution. This reduces the number of parameters in the model and shortens inference time.If subtitle information is provided with a presentation timestamp, the inference time can be within the presentation itself, allowing for output video generation during the presentation. Furthermore, the generated video can fit within the specified timestamp.

[0051] The new gesture program can be managed by the controller 102 of the data storage device 100 as an additional logical data element, which may be device-specific data and / or transparent to the host. The controller 102 of the data storage device 100 can manage additional metadata and control data as needed to support the system. In some cases, the new program can be generated during operation when an end user makes a request, or the new program can be generated and stored during the idle time of the data storage device 100. These implementations can be supported by a graphical user interface (GUI) where the end user can select an option for gesture PIP in the stored MPEG data stream.

[0052] As mentioned above, many different types of implementations can be used. In one example implementation, the controller 102 of the data storage device 100 embeds (overlays) the PIP video into the main video. In this example, the controller 102 retrieves the data from memory 104, extracts the subtitle, generates a gesture video, and temporarily stores the gesture video in memory 104 as in an SLC buffer (SLC = single-level cell). The controller 102 can decode both videos into single-frame buffers and then overlay the gesture video on top of the original video to generate a combined video. Finally, the controller 102 can store the combined video as a supporting program in memory 104.The new file can be managed as a device-specific file, abstracted by the Host 300 and used for playback when a special playback is requested due to an impairment. Alternatively, the new file can be made accessible to the Host 300 and stored in a new logical memory location marked in the file system. The Controller 102 can also use automatic NAND trimming to elapse the new content based on time or other parameters. Additionally, the Controller 102 can employ Flash Translation Layer (FTL) biasing schemes, such as protection and durability, for the newly created logical data based on the Host 300 and application requirements.

[0053] Fig. Figure 4 is a block diagram of some additional components of the controller 102 of the data storage device 100; i.e., a demultiplexer 410, a video mixer 415, and a multiplexer 420. As shown in Fig. As shown in Figure 4, the Demultiplexer 410 generates subtitles, original video, original audio, and metadata from an original media file. The Controller 102 uses a text-to-gesture model (now known or later developed) to generate a gesture video from the subtitles. The Video Mixer 415 overlays the gesture video onto the original video (e.g., as picture-in-picture video or otherwise). Finally, the Multiplexer 420 combines the output of the Video Mixer 415, the original audio, and the metadata to output media containing the gesture video.

[0054] In this embodiment, the gesture video is embedded in the main video as a small picture-in-picture (PIP) window. To achieve this, the processing application in the controller 102 can have the capability to decode the video into individual video frames. If the content is encrypted, the controller 102 can also have decryption capabilities as well as associated DRM (Digital Rights Management) security keys or other security keys. To generate PIP video, the controller 102 can decode both videos into single-frame buffers and then overlay the gesture video on top of the original video to generate a combined video. In one implementation, this is done in hardware, while in other implementations, it is done in software.

[0055] In another exemplary implementation, the controller 102 of the data storage device 100 manages a separate video track. In this example, the controller 102 retrieves the data from memory 104, extracts the subtitle, generates gesture video, and manages and stores it as a separate gesture video program in memory 104 (e.g., in an SLC buffer). Furthermore, the controller 102 adds this video program as a new program to the existing MPEG transport (TS) data stream, modifying the current data stream itself. Additionally, the controller 102 can manage two copies: a current MPEG data stream and another MPEG data stream in which the additional gesture program is embedded. The controller 102 can also apply FTL biasing schemes, such as protection and persistence, to the generated MPEG data in which a new gesture program is embedded.MPEG parameters can be modified according to the MPEG specification to accommodate the new program. The media player (e.g., 300 in the host, 100 outside the data storage device) can be responsible for decoding both video programs (the actual video and the gesture video) from the data stream and playing them back synchronously. With appropriate user settings, backward compatibility can be enabled if the playback system ignores the decoding and synchronization of the gesture program when it is not required, even though it is part of the MPEG-TS.

[0056] Fig. Figure 5 is a block diagram illustrating this embodiment. As shown in Fig. As shown in Figure 5, the Demultiplexer 410 generates subtitles, original video, original audio, and metadata from an original media file. The Controller 102 uses a text-to-gesture model (now known or later developed) to generate a gesture video from the subtitle. The Controller 102 also generates gesture video metadata. The Multiplexer 420 combines the gesture video, the original video, the original audio, and the newly created metadata and outputs media containing the gesture video. The generated gesture video metadata can be used by the Host 300 to play back the output media file.

[0057] Therefore, in this embodiment, a separate video track can be created with the output video. This track can be included in the media data stream that the data storage device 100 provides to the client. For MPEG-based media, a program mapping table can be modified to reflect the new video track. The presentation and decoding timestamps of the video frames in the gesture video can be synchronized with the original video.

[0058] The media player in the client device can be capable of displaying the picture-in-picture (PIP) video alongside the original video. The advantage of this approach is that the data storage device does not need to decrypt the original video if it is encrypted.

[0059] Many alternatives can be used in these embodiments. One alternative embodiment is extrapolated, for example, in over-the-top (OTT) platforms. In this alternative, the controller 102 of the data storage device 100 can, in response to an end-user request, perform one of the methods described above (or other methods) to dynamically add picture-in-picture (PIP) support with gesture generation to the existing content when the controller 102 determines that such video playback is required. Additionally, the gesture type (from a variety of text-to-gesture designs) can be dynamically selected based on ratings and a preference list of the end users. A method triggered in the cloud or in an edge system can be used to accomplish the task.In some cases, the controller 102 of the data storage device 100 can also convert the subtitles into a preferred audio (e.g., an audio language that is not otherwise part of the data stream) and add an additional audio program to the data stream.

[0060] Fig.Figure 6 is a block diagram of an example media container file of one embodiment. There are several methods for embedding subtitles (e.g., closed captions) in media content. For example, "SRT files" are text files bundled with the main audio-video (AV) content and may contain timestamps and speech transcripts. As another example, the text can be part of a data track, which is one way to embed the subtitles in the media package defined by a media container format. For example, in MPEG-based formats, subtitles are transmitted with their own data streams. These data streams can be identified by reading the program mapping table of the displayed program. The subtitles may be in ASCII (American Standard Code for Information Interchange) or Unicode encoding, or they may be present as a bitmap within the subtitle data stream.A subtitle in bitmap format can be converted to text using a pixel-to-text conversion method. If no subtitle information is available, a speech-to-text converter can be used to generate subtitles. The appropriate audio track in the media can be selected using audio language descriptors within the media.

[0061] These embodiments offer several advantages. For example, they can provide a richer and enhanced media experience for a hearing-impaired viewer. Furthermore, since the gesture video is generated during operation, these embodiments can reduce the need for additional storage space and avoid the costs associated with manually generating gestures. Finally, these embodiments can help a significant portion of the disadvantaged population and assist an organization in fulfilling its corporate social responsibility.

[0062] Finally, as mentioned above, any suitable type of memory can be used. Semiconductor memory devices include volatile memory devices such as dynamic random-access memory (“DRAM”) or static random-access memory (“SRAM”), non-volatile memory devices such as resistive random-access memory (“ReRAM”), electrically erasable programmable read-only memory (“EEPROM”), flash memory (which can also be considered a subset of EEPROM), ferroelectric random-access memory (“FRAM”), and magnetoresistive random-access memory (“MRAM”), as well as other semiconductor devices capable of storing information. Each type of memory device can have different configurations. For example, flash memory devices can be configured in a NAND or NOR configuration.

[0063] The memory devices can be composed of passive and / or active elements in any combination. As a non-limiting example, passive semiconductor memory elements include ReRAM device elements, which in some embodiments include a resistive switching storage element such as an antifuse, a phase-change material, etc., and optionally a control element such as a diode, etc. As another non-limiting example, active semiconductor memory elements include EEPROM and flash memory device elements, which in some embodiments include elements that comprise a charge storage region such as a floating gate, conductive nanoparticles, or a dielectric charge storage material.

[0064] Multiple memory elements can be configured to be cascaded or to allow individual access to each element. As a non-restrictive example, flash memory devices in a NAND configuration (NAND memory) typically contain cascaded memory elements. A NAND memory array can be configured to consist of multiple memory strings, where a string comprises multiple memory elements sharing a single bit line and accessed as a group. Alternatively, memory elements can be configured to allow individual access to each element, such as a NOR memory array. NAND and NOR memory configurations are examples, and memory elements can be configured in other ways as well.

[0065] The semiconductor memory elements, which are located within and / or above a substrate, can be arranged in two or three dimensions, for example as a two-dimensional memory structure or as a three-dimensional memory structure.

[0066] In a two-dimensional memory structure, the semiconductor memory elements are arranged in a single plane or a single memory device plane. Typically, in a two-dimensional memory structure, memory elements are arranged in a plane (e.g., in a plane in an xz direction) that is substantially parallel to a major surface of a substrate that supports the memory elements. The substrate can be a wafer over or in which the layer of memory elements is formed, or it can be a support substrate that is attached to the memory elements after they have been formed. As a non-restrictive example, the substrate can include a semiconductor such as silicon.

[0067] The memory elements can be arranged in an ordered array, such as in a multitude of rows and / or columns, within a single storage device level. However, the memory elements can also be arranged in irregular or non-orthogonal configurations. Each memory element can have two or more electrodes or contact lines, such as bit lines and word lines.

[0068] A three-dimensional storage array is arranged such that storage elements occupy multiple levels or multiple storage device levels, thereby forming a structure in three dimensions (i.e. in the x, y, and z directions, with the y direction being essentially perpendicular and the x and z directions being essentially parallel to the main surface of the substrate).

[0069] As a non-restrictive example, a three-dimensional memory structure can be arranged vertically as a stack of multiple two-dimensional memory device levels. As another non-restrictive example, a three-dimensional memory array can be arranged as multiple vertical columns (e.g., columns extending substantially perpendicular to the main surface of the substrate, i.e., in the y-direction), with each column containing multiple memory elements. The columns can be arranged in a two-dimensional configuration, e.g., in an xz-plane, resulting in a three-dimensional array of memory elements with elements on multiple vertically stacked memory levels. Other configurations of memory elements in three dimensions can also form a three-dimensional memory array.

[0070] As a non-restrictive example, in a three-dimensional NAND memory array, the memory elements can be coupled to form a NAND string within a single horizontal (e.g., xz) memory device layer. Alternatively, the memory elements can be coupled to form a vertical NAND string spanning multiple horizontal memory device layers. Other three-dimensional configurations are conceivable, where some NAND strings contain memory elements within a single memory layer, while other strings contain memory elements spanning multiple memory layers. Three-dimensional memory arrays can also be designed in a NOR configuration and in a ReRAM configuration.

[0071] Typically, in a monolithic three-dimensional storage array, one or more storage device layers are formed on top of a single substrate. Optionally, the monolithic three-dimensional storage array can also include one or more storage layers that are at least partially contained within the single substrate. As a non-restrictive example, the substrate can be a semiconductor such as silicon. In a monolithic three-dimensional array, the layers that form each storage device layer of the array are usually formed on top of the layers of the underlying storage device layers of the array. However, layers of adjacent storage device layers in a monolithic three-dimensional storage array can be shared, or there can be intermediate layers between the storage device layers.

[0072] On the other hand, two-dimensional arrays can be formed separately and then arranged together to create a non-monolithic, multi-layered storage device. For example, non-monolithic stacked storage devices can be constructed by forming storage layers on separate substrates and then stacking the storage layers on top of each other. The substrates can be made thinner or removed from the storage device layers before stacking, but because the storage device layers are initially formed on separate substrates, the resulting storage arrays are not monolithic three-dimensional storage arrays. Furthermore, multiple two-dimensional or three-dimensional storage arrays (monolithic or non-monolithic) can be formed on separate chips and then bundled together to form a stacked-chip storage device.

[0073] Operating and communicating with memory elements typically requires an associated circuit arrangement. Non-restrictive examples include circuit arrangements used to control and manage memory elements to perform functions such as programming and reading. This associated circuit arrangement may reside on the same substrate as the memory elements and / or on a separate substrate. For instance, a memory read / write controller may reside on a separate controller chip and / or on the same substrate as the memory elements.

[0074] The person skilled in the art will recognize that this invention is not limited to the two-dimensional and three-dimensional structures described, but covers all relevant storage structures within the basic idea and scope of protection of the invention as described here and understood by the person skilled in the art.

[0075] The foregoing detailed description is to be understood as an illustration of selected forms that the invention may take, and not as a definition of the invention. Only the following claims, including all equivalents, are intended to define the scope of protection of the claimed invention. Finally, it should be noted that any aspect of any embodiment described herein may be used alone or in combination with one another.

Claims

[1] Data storage device comprising: a storage facility; and one or more processors, individually or in combination, configured to: Extracting subtitles from a video stored in memory; Generating a gesture video from the subtitles; Creating a combined video that includes the generated gesture video combined with the video; and Saving the combined video to memory. [2] Data storage device according to claim 1, wherein the combined video is created by: Temporarily storing the generated gesture video in memory; Decoding the generated gesture video and the video in their respective frame buffers; and Embedding the generated gesture video into the video. [3] Data storage device according to claim 1, wherein the combined video is managed as a data storage device-specific file that is abstracted to a host. [4] Data storage device according to claim 1, wherein the combined video is made accessible to a host and stored in a logical storage location in memory. [5] Data storage device according to claim 1, wherein the one or more processors are further configured individually or in combination to crop the gesture video based on a parameter. [6] Data storage device according to claim 1, wherein the one or more processors are further configured individually or in combination to store the combined video using a flash translation layer biasing scheme. [7] Data storage device according to claim 1, wherein the generated gesture video is combined with the video as a picture-in-picture video and superimposed over the video. [8] Data storage device according to claim 1, wherein the subtitles are extracted from a Moving Picture Experts Group (MPEG) data stream. [9] Data storage device according to claim 1, wherein the subtitles are extracted from video frames. [10] Data storage device according to claim 1, wherein the gestures are generated from the subtitles using an artificial intelligence model. [11] Data storage device according to claim 1, wherein the data storage device comprises a network-connected storage server. [12] Data storage device according to claim 1, wherein the storage comprises a three-dimensional storage. [13] Procedures, including: Performing, in a data storage device that includes a memory: Extracting text information associated with a video stored in memory; Generating a gesture video from the extracted text information; and Saving the gesture video as a new program within an existing transport data stream for the video. [14] Method according to claim 13, wherein the gesture video is overlaid as a picture-in-picture video on top of the video. [15] Method according to claim 13, wherein the gestures are generated from the extracted text information using an artificial intelligence model. [16] Method according to claim 13, wherein the transport data stream comprises a Moving Picture Experts Group (MPEG) transport data stream. [17] Method according to claim 13, wherein the data storage device comprises a network-connected storage server. [18] Method according to claim 13, wherein the data storage device comprises an edge node of a content delivery network. [19] Method according to claim 13, wherein the data storage device comprises an over-the-top platform. [20] Data storage device comprising: a storage facility; and Means to: Extracting subtitles from a video stored in memory; Generating a gesture video from the subtitles; and To cause the gesture video to be overlaid on the video during playback.