Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

371results about "Metadata video data retrieval" patented technology

Weak supervision online video moment positioning method and system based on memory perception

The invention relates to a weak supervision online video moment positioning method and system based on memory perception, and belongs to the technical field of artificial intelligence, and the method comprises the steps: carrying out the multi-modal feature fusion of a given video and a text query thereof, and obtaining the unified frame level representation at each stage; inputting the fused features into an offline module and an online module in an integral and frame-by-frame manner by using an offline guide online model architecture; in the off-line module, generating a Gaussian mask to reconstruct query of a covered part of words, and obtaining a proposal of an action starting moment; in the on-line module, the long-term historical memory in the window is used for enhancing the score, the attention weight of the score in the window is dynamically generated, and the score of the current frame is calculated in a weighted mode; taking the proposal obtained by the offline module as a pseudo tag, and providing supervision information for the score sequence of the online module; and high-performance weak supervision on-line moment positioning can be completed only by independently deducing the on-line module. The expansion capability and the application value of the model are remarkably improved.
Owner:SHANDONG UNIV

Public opinion video tag aggregation method and system based on artificial intelligence

The invention provides a public opinion video tag aggregation method and system based on artificial intelligence, and relates to the technical field of artificial intelligence. Comprising the following steps: acquiring pictures and text information in a short video, and performing semantic alignment; different large language models are adopted to generate preliminary labels for the pictures and the text information after semantic alignment; clustering the pictures and the texts after semantic alignment to obtain clusters; calculating the labeling probability of each primary label type in the current cluster by each large language model, and selecting the primary label with the highest probability sum as a clustering label of the current cluster; calculating the reliability weight of each large language model in the current cluster based on the clustering label of the current cluster; and based on the reliability weight, calculating the weighted support degree of all the large language models to different preliminary label types of each piece of data in the current cluster, calculating the weighted label of the current data, and further determining a final label. According to the method, the condition of few labels or no labels can be effectively processed, and the manual workload is greatly reduced.
Owner:SHANDONG DAZHONG INFORMATION IND CO LTD

Short video promotion optimization method and system

The invention discloses a short video promotion optimization method and system, relates to the technical field of short video promotion optimization, and aims to solve the problems that when an existing short video promotion method is used for processing subtle elements which may cause different understandings of users in video contents, intentions of merchants are difficult to accurately convey, target users are difficult to reach, and the user experience is poor. The method comprises the steps of obtaining content data of a to-be-promoted video, performing multi-meaning element identification on the content data to determine at least one potential multi-meaning element in the to-be-promoted video, and determining a situational misreading risk assessment result of the potential multi-meaning element according to the potential multi-meaning element and a preset user situation model, and according to the situational misreading risk assessment result, a pre-stored promotion constraint condition and a content modification cost model, making a decision between the content modification strategy and the audience screening strategy to generate an initial intervention strategy, executing the initial intervention strategy, and optimizing the promotion process based on user feedback data obtained after execution.
Owner:SHENZHEN FERROMAGNETIC DIGITAL TECH CO LTD

Information processing system and methods for clinical video retrieval

The present disclosure generally relates to an integrated approach for retrieving biomedical information from clinical video presentations. In particular, the present disclosure is directed to video retrieval systems and methods of text-video retrieval from clinical video presentations.
Owner:THE CURATORS OF THE UNIVERSITY OF MISSOURI

Time sequence statement positioning model training method based on proposal selection and anchor point distribution

The invention discloses a timing sequence statement positioning model training method and device based on proposal selection and anchor point distribution, and relates to the technical field of timing sequence statement positioning. The method comprises the following steps: initializing a fixed number of learnable queries according to a first static anchor point set; on the basis of the first learnable query, according to the unpruned long video and the natural language description text, performing proposal generation through a time sequence statement positioning model; based on the proposal selection module, redundant proposal filtering is carried out by using a non-maximum suppression algorithm; based on an anchor point distribution module, according to the untrimmed long video, the first static anchor point set and the candidate proposal set, performing bipartite graph matching by using a Hungary algorithm; and carrying out loss function weighting calculation according to the unpruned long video and the optimal proposal set, and carrying out parameter optimization on the time sequence statement positioning model to obtain an optimized time sequence statement positioning model. The method is an efficient and accurate time sequence statement positioning model training method combining proposal selection and anchor point distribution.
Owner:UNIV OF SCI & TECH BEIJING

Short video personalized recommendation system and method based on deep reinforcement learning

The invention discloses a short video personalized recommendation system and method based on deep reinforcement learning. Firstly, user features, video features and feedback features are converted into embedded sequence representations, then states are modeled through a self-adaptive attention mechanism, potential relationships among embedded sequences are captured, and personalized recommendation state representations reflecting user preferences are generated. An optimal recommendation strategy of deep reinforcement learning is used, a Q network comprising an LSTM layer and an MLP layer is constructed, the LSTM layer is responsible for capturing time sequence information of user behaviors, and the MLP layer combines the information with a current environment and outputs a Q value of each action. By continuously optimizing the Q network, an optimal recommendation strategy is learned, so that personalized short video recommendation is provided. Compared with a traditional short video recommendation method, the method has the advantages that the adaptive capacity is higher, the preference change of the user is quickly responded, the recommendation strategy is dynamically adjusted, and higher recommendation accuracy and diversity are achieved, so that long-term return is maximized, and the satisfaction degree of the user is improved.
Owner:HANGZHOU UNIV OF ELECTRONIC SCI & TECH PINGHU DIGITAL TECH INNOVATION RES INST CO LTD +1

An intelligent clipping application method and device through picture recognition, equipment and medium

The application relates to an intelligent clipping application method through picture recognition, which comprises the following steps: acquiring medical image video data to be clipped; performing grouped image recognition on the medical image video data to be clipped to obtain a key frame index; and performing clipping on the medical image video data to be clipped based on the key frame index. The application does not require a clipping personnel with medical knowledge to perform clipping, and the clipping personnel can complete the video clipping operation without watching the whole content of the video or repeatedly watching the video for multiple times, so that the efficiency of the video clipping is improved. The application also relates to an intelligent clipping application device through picture recognition, a storage medium and equipment.
Owner:北京泽桥数智科技有限公司

Text-video retrieval method based on multi-granularity attention

The invention discloses a text-video retrieval method based on multi-granularity attention, which comprises the following steps of: firstly, aiming at a given text-video pair, generating disturbance text characteristics through a learnable semantic preserving strategy, and constructing a text with semantic key conflicts as a difficult negative sample to carry out comparative learning so as to strengthen semantic discrimination capability; secondly, short-time actions, medium-time semantics and long-time features are extracted from the video, video content is represented in an omnibearing mode, text features serve as query signals, multi-granularity video features are dynamically weighted and fused through a multi-granularity attention mechanism, and accurate alignment of the text and the video features is achieved; and finally, calculating a similarity score of the query text and the fused video features to retrieve a video matched with the query text. According to the method, multi-granularity video features in a text-video retrieval method are considered, adaptive association of text and video semantics is established, and accurate matching is realized to improve the retrieval effect.
Owner:ZHEJIANG UNIV OF TECH

Training method of video time positioning model, video time positioning method, equipment and medium

The invention relates to the technical field of computer vision, particularly provides a training method of a video time positioning model, a video time positioning method, equipment and a medium, and aims to solve the problem of large video time positioning error. In order to achieve the purpose, the model training method comprises the steps that multiple frames of images are sampled from a training video to serve as training data, the training data and a first preset query text are coded to obtain visual features and text query features, and a video time positioning model is trained based on the visual features and the text query features, obtaining a plurality of candidate answers, respectively calculating the relative advantage value of each candidate answer based on the real answer of the first preset query text and the plurality of candidate answers, adjusting the parameters of the video time positioning model based on each relative advantage value, and continuing to execute the step of sampling multiple frames of images from the training video. Therefore, the performance and accuracy of video time positioning can be improved.
Owner:PEKING UNIV +1

Video content retrieval method, device and terminal based on voice interaction of television system

The invention discloses a video content retrieval method and device based on television system voice interaction and a terminal, and relates to the technical field of video processing, and the method comprises the steps: when a video is played for the first time, extracting a picture frame from the video at a preset frequency, converting the picture frame into a multi-dimensional image feature vector and a corresponding video timestamp, and carrying out hierarchical storage in a database, constructing vectorized data containing visual semantic information; obtaining a voice retrieval instruction, performing intention recognition and semantic understanding, extracting a detection keyword, and generating a multi-dimensional retrieval feature vector; calculating a matching degree between the multi-dimensional retrieval feature vector and a multi-dimensional image feature vector of a video picture frame stored in a database, and screening out picture frames of which the similarity is higher than a preset similarity threshold to form a retrieval candidate matching set; and determining matched picture playing. The video content retrieval method is efficient, accurate and high in interactivity, and retrieval experience and operation efficiency of the user in the video watching process are remarkably improved.
Owner:SHENZHEN COOCAA NETWORK TECH CO LTD

Dynamic news abstract generation method based on voice and visual perception

The invention discloses a dynamic news abstract generation method based on voice and visual perception, which comprises the steps of voice recognition text extraction, audio-text alignment, prompt word generation and final news abstract generation. The objective is achieved through an innovative process comprising a plurality of key optimization links. A robust long audio processing and recognition mechanism, an optimized audio-text alignment method, an elaborately designed cue word framework and an efficient fine-tuning multi-modal large language model are integrated. According to the method, deep understanding of multi-modal information and high-quality text generation are jointly realized, and the quality of four core dimensions including content coverage, semantic accuracy, logic continuity and language fluency of the abstract is comprehensively improved.
Owner:NANJING INFORMATION HIGH-SPEED RAILWAY RES INST OF SCI AND TECH

Method for identifying video author, method and device for recommending video resource

The disclosure provides a video author identification method, a video resource recommendation method and device, and relates to the technical field of computers, in particular to the technical field of video technology and intelligent recommendation. The video author identification method comprises: determining first image features of a plurality of video works of a video author according to classification results of cover image of the plurality of video works; determining first video works of a target type from the plurality of video works according to text features and the first image features of the plurality of video works; wherein the text features are obtained according to label information of the video works; determining first work features of the video author according to a proportion of the first video works in the plurality of video works; determining second work features of the video author according to a proportion of second video works containing target attributes in the plurality of video works; and determining an author type label of the video author according to the first work features and the second work features. The disclosure is helpful for mining of video authors and video resources.
Owner:BAIDU (CHINA) CO LTD

Contextual digital media processing systems and methods

Systems and methods for contextual digital media processing are disclosed herein. An example method includes receiving content from a source as digital media that are being displayed to a user, processing the digital media to determine contextual information within the content, searching at least one network for supplementary content based on the determined contextual information, and transmitting the supplementary content for use with at least one of the source or a receiving device.
Owner:SONY INTERACTIVE ENTERTAINMENT LLC

Video recommendation method and device, equipment and storage medium

The invention relates to a video recommendation method and device, equipment and a storage medium. The video recommendation method comprises the steps that historical behavior data of a target user is acquired, and the historical behavior data refers to various interaction data generated in the historical video watching process of the target user; establishing a multi-dimensional label system for carrying out structured description on the video content of the historical video, wherein each dimensional label in the multi-dimensional label system is used for representing an independent semantic classification; determining an incidence relation between the historical behavior data and each dimension label in the multi-dimension label system, and generating a multi-dimension label behavior sequence used for representing the video interest preference of the target user; and outputting a target video to be recommended to the user through a video recommendation model according to the multi-dimensional tag behavior sequence and the multi-dimensional tag features of the candidate videos. According to the method provided by the invention, the accuracy of user interest characterization is improved, the correlation of recommendation results is enhanced, and the sorting precision and the user experience are further improved.
Owner:BEIJING QIYI CENTURY SCI & TECH CO LTD

A short video recommendation method based on big data statistics

The application relates to the technical field of big data statistics, and specifically discloses a short video recommendation method based on big data statistics, which comprises the following steps: S1: obtaining the total number of short videos watched by a user in a sampling period, and dividing the same short videos according to short video labels; S2: obtaining the number of times of logging in a platform and the single online duration in the sampling period, and calculating the mean value and the dependence coefficient of the single online duration; S3: dividing the sampling period into a front interval and a rear interval, and sequentially calculating the trend coefficients of the same short videos in the front interval and the rear interval; S4: correcting the trend coefficients to obtain corrected coefficients; and S5: calculating the trend difference values of the same short videos, and analyzing the watching trend of the same short videos by the user based on the trend difference values. The application can analyze the preferences of the user and predict future trends, establish a flexible and real-time response mechanism for the change of the preferences of the user, realize the long-term development of the platform, and maintain the market competitiveness.
Owner:BEIJING ZHONGYANG TIANCHENG TECH CO LTD

A video representation method and device based on an unsupervised pre-training model

The application provides a video representation method based on an unsupervised pre-training model, comprising the following steps: obtaining a video sample set; taking video frame embedding of the video sample set and text labels of video titles as inputs, pre-training a video representation model by using a mask framework modeling and a mask language modeling respectively, and obtaining a pre-training model; reconstructing the text labels of the video titles in the video sample set by using a contrastive learning method, and obtaining a positive sample set, a medium sample set and a negative sample set; performing contrastive training on the pre-training model by using the positive sample set, the medium sample set and the negative sample set by using a dynamic queue training method, and obtaining a completed video representation model; and obtaining a video to be labeled, and performing content extraction by using the completed video representation model.
Owner:HAINAN UNIV

Capturing objects in an unstructured video stream

A method includes obtaining a first unstructured video stream that provides pixel values for a plurality of pixels and corresponds to a portion of a second unstructured video stream being displayed on a second electronic device different from the first electronic device. Obtaining the first unstructured video stream includes obtaining pass-through image data including the portion of a second unstructured video stream. The method includes generating respective pixel characterization vectors for a first portion of the plurality of pixels. Generating each of the respective pixel characterization vectors includes determining a respective instance label value. The method includes identifying a first object within the first portion of the plurality of pixels associated with a particular instance label value. The method includes generating respective semantic label values corresponding to pixels associated with the first object. The respective semantic label values are added to pixel characterization vectors associated with the first object.
Owner:APPLE INC

Transformation of database entries for improved association with related content items

A content analysis system includes processor and memory hardware storing data analyzed content items and instructions for execution by the processor hardware. The instructions include, in response to a first intermediate content item being analyzed to generate a first text description, receiving the first intermediate content item and analyzing the first text description to generate a first reduced text description. The instructions include identifying a first set of tags by applying a tag model to the first text description and generating a first analyzed content item. The instructions include adding the first analyzed content item to the analyzed content database and, in response to a displayed content item being associated with at least one tag of the first set of tags, displaying a first user-selectable link corresponding to the first analyzed content item on a portion of a user interface of a user device displaying the displayed content item.
Owner:CHARLES SCHWAB & CO INC

A robot dance automatic generation method and system based on music feature analysis

PendingCN122156404AMetadata video data retrievalBiological modelsEngineeringMulti objective model
The application provides a robot dance automatic generation method and system based on music feature analysis, and relates to the technical field of robot control. First, a music feature spectrum covering rhythm, energy and other dimensions is constructed to convert abstract music into structured information understandable by machines, and then a dance semantic label sequence bound to a time axis is generated to provide accurate basis for action selection. By combining with a forward-looking window to predict future music features, the current and future candidate actions are filtered through a multi-objective model to realize the dual matching of action and music micro-rhythm and macro-structure. The finally generated control instruction sequence directly drives the physical robot, which breaks the dependence on video materials, realizes the automatic generation of dance with creativity and artistic expression, and improves the generation efficiency and adaptive flexibility.
Owner:ZHEJIANG SILICON ARK ROBOT CO LTD

A method, device and system for storing and synchronously triggering augmented reality events in a panoramic video

This invention provides a method for storing augmented reality events in panoramic videos, a method for synchronously triggering events, and an apparatus. The method uses time-triggered special effects or footage in the panoramic video as trigger events. It uses the ID of the trigger event as the key and the content of the trigger event as the value to create an index for a linked list. Each trigger event is stored in the linked list in chronological order. Simultaneously, a balanced tree is built using the trigger time of each event as the key and the ID of each trigger event as the value. To insert or delete a trigger event, the method searches the balanced tree to obtain the IDs of time-adjacent trigger events, then finds the corresponding position in the linked list based on that ID for insertion or deletion. It also allows for timed triggering of each event at specific times. This invention enables convenient and quick insertion, deletion, and synchronous triggering of events in panoramic videos. By using a balanced tree to store events, it greatly reduces the computational requirements during event lookup and improves the synchronization during panoramic video playback.
Owner:BEIJING UNIV OF POSTS & TELECOMM

Information processing method, apparatus and electronic device

The present disclosure provides an information processing method and apparatus, and an electronic device. The method includes: displaying first information in response to a second user having generated second multimedia content using first multimedia content, where the first multimedia content is multimedia content published by a first user; and acquiring the second multimedia content in response to an operation on the first information and playing the second multimedia content, thus reducing the operation complexity of acquiring the second multimedia content.
Owner:LEMON INC(GB)

Image retrieval method, device and equipment for target image information

The embodiment of the invention provides an image retrieval method, device and equipment for target image information. The method comprises the following steps: acquiring to-be-retrieved target image information; analyzing and extracting the target image information to obtain target image feature information; matching the target image feature information in a preset image data structured index database, and retrieving to obtain a retrieval result containing matched target image data and structured description information associated with the target image data; outputting a retrieval result; wherein the image data structured index database is obtained through the following processes: obtaining multi-modal original data containing images, videos and text descriptions of the same target object; performing content analysis and feature extraction on the multi-modal original data to obtain corresponding structured description information; and storing the structured description information as an index structure supporting feature-based matching retrieval according to a preset storage format. According to the embodiment of the invention, the deep semantic content of the image can be efficiently understood and accurately matched.
Owner:CHINA AI MEDIA&ENTERTAINMENT TECH CO LTD

Imaging device, imaging method, and imaging program

To provide an imaging apparatus, an imaging method, and an imaging program, capable of suppressing an influence of a delay in output of metadata on development processing.SOLUTION: An imaging unit 119 generates image data of a video as a plurality pieces of frame data that are continuous in a temporal order. An imaging control unit 104 generates metadata indicating an imaging condition of the imaging unit in association with the frame data in a case in which the imaging condition of the imaging unit 119 is changed. An output control unit 110 and an external output I / F 111 add the metadata to the frame data and output the frame data to which the metadata is added as video data before demosaicing. In addition, in a case in which a data amount of the metadata associated with the frame data exceeds an addable data amount, the output control unit 110 and the external output I / F111 add, to the frame data, metadata selected based on a priority from among the metadata associated with the frame data.SELECTED DRAWING: Figure 4
Owner:FUJIFILM CORP

Memory perception based weakly supervised online video temporal instance localization method and system

The present application relates to a memory-aware based weakly supervised online video moment localization method and system, belonging to the field of artificial intelligence technology, comprising: multi-modal feature fusion on a given video and its text query to obtain unified frame-level representation at each stage; using an offline-guided online model architecture, the fused features are input into the offline and online modules in the form of the whole and frame by frame; in the offline module, a Gaussian mask is generated to reconstruct the hidden part of the query, obtaining the proposal of the action starting moment; in the online module, the long-term historical memory in the window is used for enhancement, and the attention weight in the window is dynamically generated, and the score of the current frame is calculated by weighting; the proposal obtained by the offline module is used as a pseudo label to provide supervision information for the score sequence of the online module; only the online module needs to be inferred separately, and the weakly supervised online moment localization with high performance is completed. The present application significantly improves the expansion capability and application value of the model.
Owner:SHANDONG UNIV

Method, electronic device and computer program product for managing large number of videos

The present disclosure relates to a method for managing a large number of videos, an electronic device and a computer program product, which are applied to a monitoring system, including detecting an event and outputting a large number of videos according to the event; labeling a plurality of tags for the large number of videos, the tags including a plurality of attributes of at least one object in the large number of videos; receiving video search information; comparing the relevance of the video search information and the tags of the large number of videos, and outputting a plurality of recommended videos in a terminal device according to the relevance, the terminal device including a display screen; detecting the resolution of the display screen, and adaptively playing the recommended videos in a user interface on the display screen according to the resolution. The method of the present disclosure uses attribute tag indexing technology closest to human search logic, effectively manages a large number of videos, automatically searches for relevant video results according to input information, greatly simplifies the search process and improves user experience.
Owner:WISTRON CORP

Real-time screenshot splicing-based training room desktop operation key step playback method

The application discloses a real-time screenshot splicing-based operation key step playback method in a practical training room, relates to the technical field of data processing, and balances integrity and storage economy by aligning desktop screenshots with multi-dimensional event timestamps and combining operation strength adaptive sampling, which not only completely captures operation details, but also avoids invalid screenshot waste; relying on screenshot splicing and incremental storage technology, a globally consistent operation context is constructed to solve the information loss problem under window switching and complex operation; through intelligent screening and atlas indexing of key steps, accurate positioning of operation steps is realized, and playback step-level time compression and event superposition presentation are matched to greatly improve the backtracking efficiency; the playback consistency verification and rollback correction mechanism guarantee the accuracy of reproduction, and finally, the operation review effect in practical training teaching is optimized, the students are helped to quickly master core skills, and the training quality is significantly improved.
Owner:SHANDONG PANLONG INFORMATION TECH CO LTD

Method for performing visual question and answer by utilizing attention mechanism from word to region

The invention belongs to the field of artificial intelligence networks, and particularly relates to a method for performing visual questions and answers by using an attention mechanism from words to regions. Comprising the following steps: (1) extracting features; and (2) generating candidate answers: firstly positioning related image areas and keywords in questions by adopting a collaborative attention mechanism, then obtaining fine-grained image features and question features, and finally fusing the two features to generate the candidate answers. According to the method, in order to generate candidate answers with higher quality, two question and answer stages are cascaded, a traditional single-stage visual model is expanded into a double-stage model, semantic information contained in the answers is fully mined, and accurate prediction of the final answers is promoted. According to the invention, image areas and keywords related to questions can be extracted and utilized, so that more accurate candidate answers can be generated.
Owner:石帅

Tag recommendation method and device, tag recommendation model training method and medium

The application relates to the technical field of artificial intelligence, in particular to a label recommendation method and device, a label recommendation model training method and a medium, wherein the method comprises the following steps: obtaining a reference label and at least two candidate labels; determining semantic similarity between the reference label and the candidate labels; determining co-occurrence probability between the reference label and the candidate labels, the co-occurrence probability being used for describing the probability that the reference label and the candidate label belong to the same label of multimedia data; and selecting a target label from the at least two candidate labels according to the semantic similarity and the co-occurrence probability. Since the semantic similarity represents the semantic similarity of the labels and the co-occurrence probability represents the probability that the labels appear in the same multimedia data, the target label obtained based on the semantic similarity and the co-occurrence probability is more consistent with the distribution assumption of the multimedia data, so that the obtained target label is more accurate.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

A video query method based on multi-modal embedding, electronic device, medium

The application discloses a video query method based on multi-modal embedding, an electronic device and a medium. The method first filters a video stream to obtain an anchor frame set, and stores anchor frame visual embedding, question text embedding and a first detection result processed by an adapter into a query cache. When responding to a user question, a candidate frame set is constructed by selecting a relevant anchor frame subset and a time neighborhood thereof, and a current frame and text embedding are distributed to a corresponding adapter through a query route to obtain a second detection result. Finally, the second detection result is subjected to deviation inspection by using adjacent anchor frame data in the cache, and the result after the deviation inspection is taken as a final result. Through anchor frame clipping of a search space, the application realizes efficient, accurate and time-consistent processing of various video query tasks in combination with an adapter and a deviation inspection mechanism.
Owner:ZHEJIANG UNIV

Short video recommendation method based on artificial intelligence

A short video recommendation method based on artificial intelligence comprises the following steps that a preset number of short videos are selected from a database according to tags, and the short videos containing the same tag are classified into one class; performing popularity calculation on the short videos of the same label; carrying out initialization processing on all short videos to be processed, carrying out normalization again, and carrying out percentage calibration on the videos again; selecting browsing records of 100 short videos with the same tag; for a single short video, querying the time of exiting the short video or switching the short video during each browsing, and recording a switching ratio; classifying the switching ratio by using a k-means algorithm to obtain a perceptual aversion point location; calculating the aversion degree of the whole audience to the video through the obtained perceptual aversion point location; and carrying out reverse sequencing through the disgust degree and recommending to audiences. According to the method, a specific switching point location is positioned by adopting a clustering method, so that the accuracy of data processing can be guaranteed, and an area which is disgusted by a user can be segmented, and therefore, the method is applied to video recommendation.
Owner:湖北雅派文化传播有限公司