Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

22 results about "Audio description" patented technology

Audio description, also referred to as a video description, described video, or more precisely called a visual description, is a form of narration used to provide information surrounding key visual elements in a media work (such as a film or television program, or theatrical performance) for the benefit of blind and visually impaired consumers. These narrations are typically placed during natural pauses in the audio, and sometimes during dialogue if deemed necessary.

Automated audio description system and method

An audio description system includes a memory and a processor. The memory stores source media comprising frames positioned within the source media according to a time index. The processor is configured to generate, using an image-to-text model, a textual description of each frame; identify intervals within the time index, each interval encompassing one or more positions of one or more frames; identify placement periods within the time index, each placement period being temporally proximal to an interval; generate a summary description based on at least one textual description of at least one frame positioned within a selected interval temporally proximal to a placement period; and associate the summary description with the placement period.
Owner:3PLAY MEDIA

Method for producing illustrated audiobooks for children with vision impairment

PCT designated stageWO2025254564A1Pattern printingElectrical appliancesTactile sensationTalking books
A method for producing an educational book includes creating a draft layout with tactile materials and subsequently working up and assembling same. Illustrations, scoring, perforation channels and embossing are applied to the layout. A picture has audio regions applied thereto which, when touched with a portable device, playback an audio track (audio description). All audio files are combined into a single file for recording on a portable audio playback device. The prepared material is sent for comprehensive preparation and printing to produce a finished page spread. This is followed by a decorative work-up and final manual assembly, which includes attaching the tactile materials to the working portions of the spread and folding the spread along the marked score lines. The technical result consists in the creation of a multi-sensory product that allows children with vision impairment to receive information via several sensory channels simultaneously.
Owner:REGIONALNYJ BLAGOTVORITELNYJ OBSHCHESTVENNYJ FOND ILLYUSTRIROVANNYE KNIZHKI DLYA MALENKIKH SLEPYKH DETEJ

Audio description of physical objects

In one implementation, a method of playing a sound is performed by a first device having an image sensor, one or more processors, and non-transitory memory. The method includes capturing, using the image sensor, an image of a physical environment including a physical object. The method includes determining that one or more description criteria are satisfied, wherein the description criteria includes a criterion that is satisfied when the first device detects that a gaze of the user is directed at the physical object. The method includes, in response to determining that the description criteria are satisfied, playing a sound describing the physical object.
Owner:APPLE INC

Audio processing method and device, model training method and device, storage medium and equipment

The invention relates to an audio processing method and device, a model training method and device, a storage medium and equipment, and the method achieves the distinguishing of the weights of word vectors in a first vector and the distinguishing of the weights of word vectors in a second vector through the weighting processing of the first vector of audio description and the second vector of reference description. Therefore, when the similarity between the audio description and the reference description is determined based on the weighted vector, differentiated expressions of sound in the audio description and the reference description can be captured according to the weight of the word vector, and the accuracy of similarity calculation is improved. And the audio description is measured according to the similarity between the audio description and the reference description, so that the accuracy of measurement can be improved.
Owner:BEIJING XIAOMI MOBILE SOFTWARE CO LTD +1

Automated audio description system and method

An audio description system includes a memory and a processor. The memory stores source media comprising frames positioned within the source media according to a time index. The processor is configured to generate, using an image-to-text model, a textual description of each frame; identify intervals within the time index, each interval encompassing one or more positions of one or more frames; identify placement periods within the time index, each placement period being temporally proximal to an interval; generate a summary description based on at least one textual description of at least one frame positioned within a selected interval temporally proximal to a placement period; and associate the summary description with the placement period.
Owner:3PLAY MEDIA

Comment display method, device, equipment, medium and product

This application discloses a comment display method, apparatus, device, medium, and product, applicable to the field of data processing technology. The method includes: displaying target comment data for target published content, the target comment data including at least a comment video; and displaying, in association with the comment video, at least one of the following: audio description text corresponding to the audio content in the comment video, and video description text corresponding to the comment video. Using this application can improve the display effect of comments, thereby enhancing the interactive effect of comments.
Owner:XIAOHONGSHU TECH CO LTD

Audio pitch correction method and apparatus, and electronic device

An audio pitch correction method and apparatus, and an electronic device, which relate to the technical field of audio processing. The method comprises: acquiring singing audio to be processed (S101); performing speech recognition on said singing audio, so as to obtain audio description information; and performing pitch recognition on said singing audio, so as to obtain initial pitch information (S102); on the basis of the audio description information, correcting boundary information of the initial pitch information, so as to obtain pitch information (S103); on the basis of the audio description information, determining an original song pitch template corresponding to said singing audio, and on the basis of the pitch information, correcting the original song pitch template corresponding to said singing audio, so as to obtain a reference original song pitch template (S104); and on the basis of the reference original song pitch template and the pitch information, determining adjustment pitch information, and on the basis of the adjustment pitch information, performing pitch correction on said singing audio, so as to obtain pitch-corrected audio (S105). The method combines the results of pitch recognition and speech recognition to jointly determine pitch information, and dynamically determines a reference original song pitch template on the basis of the pitch information, such that the obtained adjustment pitch information is more accurate, and pitch-corrected audio is more consistent with an original song in terms of pitch and is more aligned with the actual singing performance of a user, thereby improving the overall quality of the audio.
Owner:SHANGHAI SOULGATE TECH CO LTD

Apparatus and method for providing audio description content

A method and apparatus are described. The method includes receiving audio soundtracks including an audio soundtrack representing an audible form of written text and / or an audible description of a visual element, determining a quality level of that audio soundtrack, and modifying that audio soundtrack to include an indication of the quality level based on the determination. The apparatus includes a memory circuit that stores audio soundtracks including an audio soundtrack representing an audible form of written text and / or an audible description of a visual element. The apparatus further includes an audio processing circuit coupled to the memory circuit, the audio processing circuit configured to retrieve the set of audio soundtracks, determine a quality level of the audio soundtrack representing an audible form of written text, and modify that audio soundtrack to include an indication of the quality level based on the determination.
Owner:PARITY ENDEAVORS INC

System and method for audio guide

A method of providing audio descriptions of landmarks includes causing a user device to capture an image via an imaging sensor of the user device, comparing the captured image to a reference image to identify a landmark that appears in the captured image, providing a prompt requesting a description associated with the identified landmark to a large language model (LLM), receiving an audio file of the description associated with the identified landmark, and providing the audio file of the description associated with the identified landmark to the one or more user devices.
Owner:UNIVERSAL CITY STUDIOS LLC

Method, apparatus, device, medium and program product for processing video

The invention provides a video processing method, device and equipment, a medium and a program product. In one method, in response to receiving a generation request to generate a second video based on a first video, attribute data of the first video is acquired. And determining text data for describing the first video based on the attribute data. And generating the second video based on the first video and the audio data corresponding to the text data. In this way, the dictation image making efficiency can be effectively improved, the quality of dictation content output along with the video is guaranteed, and a user can understand the video content more easily.
Owner:BEIJING ZITIAO NETWORK TECH CO LTD

Video retrieval method based on Agent

The invention discloses an Agent-based video retrieval method, which comprises two major steps of video storage and retrieval: during video storage, an original video is subjected to time sequence segmentation to obtain video clips, a comprehensive description containing a picture description and an audio description is generated through a video description module, and then the overall description of the original video is obtained through summarization of a large language model; respectively storing the two types of descriptions into a database and generating a vector index; during retrieval, the Query processing Agent analyzes user query to clarify a retrieval intention, and the database retrieval Agent optimizes the query and obtains a retrieval result based on the vector index. In the video description module, a cue word Agent generates a targeted cue word according to picture description, an audio understanding large model is guided to extract audio information associated with a picture, and tight combination of the audio information and the picture information is achieved. The problems that in the prior art, sound and picture information association is insufficient, and user intentions are difficult to distinguish are solved, the accuracy of video retrieval is effectively improved, interaction obstacles between the user and the system are reduced, and the user experience is improved.
Owner:BESTTONE HOLDING

Systems and methods for generating audio descriptions

Systems and methods for generating audio description in a particular voice. A voice clip comprising the voice of a speaker is processed by a first transformer to generate audio tokens. An image or video to be described by the audio description is processed by a second transformer to generate visual tokens. A decoder processes the audio and visual tokens to generate acoustic tokens. At least some of the acoustic tokens are generated in a zero-shot manner. The acoustic tokens are converted to an audio waveform of the audio description in the voice of the speaker.
Owner:CNTXT FZCO

Video generation method, related device, equipment and storage medium

The invention discloses a video generation method, a related device, equipment and a storage medium. The method comprises the following steps: acquiring a target audio material; obtaining a target text material corresponding to the target audio material; performing feature extraction on the target audio material to obtain audio description information related to the target audio material; obtaining K candidate videos through a video generation model based on the target text material and the audio description information; and according to the target audio material, editing the K candidate videos to obtain a target synthetic video, the target synthetic video comprising a video clip from at least one candidate video, and the background sound of the target synthetic video being generated based on the target audio material. According to the method provided by the invention, the video editing efficiency can be improved, so that the user experience is enhanced.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Method and apparatus for processing videos, and device, medium and program product

Provided in the present disclosure are a method and apparatus for processing videos, and a device, a medium and a program product. In one method, in response to receiving a generation request for generating a second video on the basis of a first video, attribute data of the first video is acquired; on the basis of the attribute data, text data for describing the first video is determined; and the second video is generated on the basis of the first video and audio data corresponding to the text data. In this way, the efficiency of audio description production can be effectively improved, and the quality of audio content output along with videos can be ensured, thereby helping users understand the content of the videos more easily.
Owner:BEIJING ZITIAO NETWORK TECH CO LTD

Pet state recognition method and apparatus, electronic device, and computer readable medium

Embodiments of the present disclosure disclose pet state recognition methods and devices, electronic equipment and computer readable media. A specific embodiment of the method comprises: in response to receiving at least one positioning signal sent by a pet recognition device, determining the position of the signal source according to the at least one positioning signal; collecting monitoring audio and video in real time according to the position of the signal source; determining the pet type according to the position of the signal source and the monitoring video included in the monitoring audio and video; in response to the pet type and the audio description information satisfying a first condition, determining pet state description information according to the position of the signal source and the monitoring audio and video; and in response to the pet type and the audio description information not satisfying the first condition, determining pet state description information according to the position of the signal source and the monitoring video included in the monitoring audio and video. This embodiment realizes accurate and effective pet state recognition.
Owner:ADDX (BEIJING) TECH CO LTD

Audio text processing methods, devices, storage media and electronic devices

This application discloses an audio text processing method, apparatus, storage medium, and electronic device. The method includes: in response to a first operation on target audio recognition text, acquiring target structured data corresponding to the target audio recognition text, wherein the target structured data includes target audio recognition text and audio description information, wherein the audio description information is used to acquire target audio data corresponding to the target audio recognition text; storing the target structured data in a shared storage area; and in response to a second operation applied to a target location, acquiring the target structured data from the shared storage area and displaying the target audio recognition text at the target location based on the target structured data, thereby enabling copied and pasted audio text to carry audio description information, and thus enabling copied and pasted audio text to have audio playback functionality.
Owner:ANHUI IFLYREC TECH CO LTD

Gesture-Controlled Interactive Audio Adventure Application

An interactive and immersive adventure application instantiated on a user computing device is configured to receive gestures and responsively output audio descriptions. The adventure application may have pre-stored stories, maps, or virtual environments and generate stories, maps, or virtual environments on the fly using some artificial intelligence engine, such as an LLM (large language model) or a hybrid approach. The stories or maps may generally be referred to as an event structure. The adventure application can interoperate with a remote service that generates or receives the event structures, and the local adventure application can receive the event structures from the remote service. Alternatively, the user computing device's adventure application may have its own stories pre-downloaded or generated by a local LLM.
Owner:SAGASWIPE LLC

System and method for audio guide

A method of providing audio descriptions of landmarks includes causing a user device (100) to capture an image (200) via an imaging sensor (102) of the user device (100), comparing the captured image (200) to a reference image (202, 204, 206, 208) to identify a landmark that appears in the captured image (200), providing a prompt requesting a description associated with the identified landmark to a large language model (LLM) (110), receiving an audio file of the description associated with the identified landmark, and providing the audio file of the description associated with the identified landmark to the one or more user devices (100).
Owner:UNIVERSAL CITY STUDIOS LLC

Automatic mixing of audio descriptions

A computer-implemented audio processing method, the method comprising: receiving audio object data and audio description data, wherein the audio object data comprises a first plurality of audio objects; calculating a long-term loudness of the audio object data and a long-term loudness of the audio description data; calculating a plurality of short-term loudnesses of the audio object data and a plurality of short-term loudnesses of the audio description data; reading a first plurality of mixing parameters corresponding to the audio object data; generating a second plurality of mixing parameters based on the first plurality of mixing parameters, the long-term loudness of the audio object data, the long-term loudness of the audio description data, the plurality of short-term loudnesses of the audio object data, and the plurality of short-term loudnesses of the audio description data; generating a gain adjustment visualization corresponding to the second plurality of mixing parameters, the audio object data, and the audio description data; and generating mixed audio object data by mixing the audio object data and the audio description data according to the second plurality of mixing parameters, wherein the mixed audio object data comprises a second plurality of audio objects, wherein the second plurality of audio objects correspond to the first plurality of audio objects mixed with the audio description data according to the second plurality of mixing parameters.
Owner:DOLBY LABORATORIES LICENSING CORP

Audio editing method and device, electronic equipment and storage medium

PendingCN122658335AEngineeringAudio frequency
The application discloses an audio editing method and device, electronic equipment and a storage medium. The method comprises: obtaining an audio to be edited; performing perception and retrieval processing on the audio to be edited to obtain audio description information and audio tokens of the audio to be edited, and generating a prompt word based on a system instruction, the audio description information and the audio tokens; wherein the audio description information comprises: an audio content type, an audio event category, a time axis and a recommended gain value; the recommended gain value is obtained from a preset vector database according to the audio content type and an audio feature vector of the audio to be edited; the prompt word is input into a preset multi-modal large model to obtain editing tokens and acoustic tokens; and the audio to be edited, the editing tokens and the acoustic tokens are input into a preset generator to generate a mixed audio. The method realizes automatic, high-quality and scene-adaptive audio gain editing.
Owner:SAMSUNG ELECTRONICS CHINA R&D CENT +1

Audio description method, system, electronic device, and storage medium

The embodiments of the present application disclose an audio description method, system, electronic device and storage medium, wherein the method comprises: during pre-training of a contrastive language-audio pair (CLAP), jointly training an audio encoder and a text encoder to align audio-text pairs with similar semantics in a shared embedding space; during the training process, learning, using a large language model, to decode real subtitles from CLAP text embeddings generated by the text encoder to reconstruct the subtitles; during the inference process, replacing the text encoder with the audio encoder, extracting audio embeddings via the audio encoder, and finally generating final subtitles via the large language model.
Owner:AISPEECH CO LTD