Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

11 results about "Audio description" patented technology

Audio description, also referred to as a video description, described video, or more precisely called a visual description, is a form of narration used to provide information surrounding key visual elements in a media work (such as a film or television program, or theatrical performance) for the benefit of blind and visually impaired consumers. These narrations are typically placed during natural pauses in the audio, and sometimes during dialogue if deemed necessary.

Audio description of physical objects

In one implementation, a method of playing a sound is performed by a first device having an image sensor, one or more processors, and non-transitory memory. The method includes capturing, using the image sensor, an image of a physical environment including a physical object. The method includes determining that one or more description criteria are satisfied, wherein the description criteria includes a criterion that is satisfied when the first device detects that a gaze of the user is directed at the physical object. The method includes, in response to determining that the description criteria are satisfied, playing a sound describing the physical object.
Owner:APPLE INC

Comment display method, device, equipment, medium and product

This application discloses a comment display method, apparatus, device, medium, and product, applicable to the field of data processing technology. The method includes: displaying target comment data for target published content, the target comment data including at least a comment video; and displaying, in association with the comment video, at least one of the following: audio description text corresponding to the audio content in the comment video, and video description text corresponding to the comment video. Using this application can improve the display effect of comments, thereby enhancing the interactive effect of comments.
Owner:XIAOHONGSHU TECH CO LTD

Audio pitch correction method and apparatus, and electronic device

An audio pitch correction method and apparatus, and an electronic device, which relate to the technical field of audio processing. The method comprises: acquiring singing audio to be processed (S101); performing speech recognition on said singing audio, so as to obtain audio description information; and performing pitch recognition on said singing audio, so as to obtain initial pitch information (S102); on the basis of the audio description information, correcting boundary information of the initial pitch information, so as to obtain pitch information (S103); on the basis of the audio description information, determining an original song pitch template corresponding to said singing audio, and on the basis of the pitch information, correcting the original song pitch template corresponding to said singing audio, so as to obtain a reference original song pitch template (S104); and on the basis of the reference original song pitch template and the pitch information, determining adjustment pitch information, and on the basis of the adjustment pitch information, performing pitch correction on said singing audio, so as to obtain pitch-corrected audio (S105). The method combines the results of pitch recognition and speech recognition to jointly determine pitch information, and dynamically determines a reference original song pitch template on the basis of the pitch information, such that the obtained adjustment pitch information is more accurate, and pitch-corrected audio is more consistent with an original song in terms of pitch and is more aligned with the actual singing performance of a user, thereby improving the overall quality of the audio.
Owner:SHANGHAI SOULGATE TECH CO LTD

Apparatus and method for providing audio description content

A method and apparatus are described. The method includes receiving audio soundtracks including an audio soundtrack representing an audible form of written text and / or an audible description of a visual element, determining a quality level of that audio soundtrack, and modifying that audio soundtrack to include an indication of the quality level based on the determination. The apparatus includes a memory circuit that stores audio soundtracks including an audio soundtrack representing an audible form of written text and / or an audible description of a visual element. The apparatus further includes an audio processing circuit coupled to the memory circuit, the audio processing circuit configured to retrieve the set of audio soundtracks, determine a quality level of the audio soundtrack representing an audible form of written text, and modify that audio soundtrack to include an indication of the quality level based on the determination.
Owner:PARITY ENDEAVORS INC

Method, apparatus, device, medium and program product for processing video

The invention provides a video processing method, device and equipment, a medium and a program product. In one method, in response to receiving a generation request to generate a second video based on a first video, attribute data of the first video is acquired. And determining text data for describing the first video based on the attribute data. And generating the second video based on the first video and the audio data corresponding to the text data. In this way, the dictation image making efficiency can be effectively improved, the quality of dictation content output along with the video is guaranteed, and a user can understand the video content more easily.
Owner:BEIJING ZITIAO NETWORK TECH CO LTD

Video retrieval method based on Agent

The invention discloses an Agent-based video retrieval method, which comprises two major steps of video storage and retrieval: during video storage, an original video is subjected to time sequence segmentation to obtain video clips, a comprehensive description containing a picture description and an audio description is generated through a video description module, and then the overall description of the original video is obtained through summarization of a large language model; respectively storing the two types of descriptions into a database and generating a vector index; during retrieval, the Query processing Agent analyzes user query to clarify a retrieval intention, and the database retrieval Agent optimizes the query and obtains a retrieval result based on the vector index. In the video description module, a cue word Agent generates a targeted cue word according to picture description, an audio understanding large model is guided to extract audio information associated with a picture, and tight combination of the audio information and the picture information is achieved. The problems that in the prior art, sound and picture information association is insufficient, and user intentions are difficult to distinguish are solved, the accuracy of video retrieval is effectively improved, interaction obstacles between the user and the system are reduced, and the user experience is improved.
Owner:BESTTONE HOLDING

Systems and methods for generating audio descriptions

Systems and methods for generating audio description in a particular voice. A voice clip comprising the voice of a speaker is processed by a first transformer to generate audio tokens. An image or video to be described by the audio description is processed by a second transformer to generate visual tokens. A decoder processes the audio and visual tokens to generate acoustic tokens. At least some of the acoustic tokens are generated in a zero-shot manner. The acoustic tokens are converted to an audio waveform of the audio description in the voice of the speaker.
Owner:CNTXT FZCO

Video generation method, related device, equipment and storage medium

The invention discloses a video generation method, a related device, equipment and a storage medium. The method comprises the following steps: acquiring a target audio material; obtaining a target text material corresponding to the target audio material; performing feature extraction on the target audio material to obtain audio description information related to the target audio material; obtaining K candidate videos through a video generation model based on the target text material and the audio description information; and according to the target audio material, editing the K candidate videos to obtain a target synthetic video, the target synthetic video comprising a video clip from at least one candidate video, and the background sound of the target synthetic video being generated based on the target audio material. According to the method provided by the invention, the video editing efficiency can be improved, so that the user experience is enhanced.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Method and apparatus for processing videos, and device, medium and program product

Provided in the present disclosure are a method and apparatus for processing videos, and a device, a medium and a program product. In one method, in response to receiving a generation request for generating a second video on the basis of a first video, attribute data of the first video is acquired; on the basis of the attribute data, text data for describing the first video is determined; and the second video is generated on the basis of the first video and audio data corresponding to the text data. In this way, the efficiency of audio description production can be effectively improved, and the quality of audio content output along with videos can be ensured, thereby helping users understand the content of the videos more easily.
Owner:BEIJING ZITIAO NETWORK TECH CO LTD

Automatic mixing of audio descriptions

A computer-implemented audio processing method, the method comprising: receiving audio object data and audio description data, wherein the audio object data comprises a first plurality of audio objects; calculating a long-term loudness of the audio object data and a long-term loudness of the audio description data; calculating a plurality of short-term loudnesses of the audio object data and a plurality of short-term loudnesses of the audio description data; reading a first plurality of mixing parameters corresponding to the audio object data; generating a second plurality of mixing parameters based on the first plurality of mixing parameters, the long-term loudness of the audio object data, the long-term loudness of the audio description data, the plurality of short-term loudnesses of the audio object data, and the plurality of short-term loudnesses of the audio description data; generating a gain adjustment visualization corresponding to the second plurality of mixing parameters, the audio object data, and the audio description data; and generating mixed audio object data by mixing the audio object data and the audio description data according to the second plurality of mixing parameters, wherein the mixed audio object data comprises a second plurality of audio objects, wherein the second plurality of audio objects correspond to the first plurality of audio objects mixed with the audio description data according to the second plurality of mixing parameters.
Owner:DOLBY LABORATORIES LICENSING CORP