Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

266results about "Video data querying" patented technology

Marketing short video generation method and device based on large language model

The embodiment of the invention provides a marketing short video generation method and device based on a large language model, and the method comprises the steps: carrying out the matching degree analysis of a hot title and product data based on the large language model, and outputting a hot topic which is most matched with a product according to the sequence of the matching degree from high to low; searching a hot video related to the most matched hot topic in a video platform, and extracting an oral copy in the hot video based on an automatic voice recognition technology; generating a creation copywriting by using a large language model based on the oral copywriting, the hot topics and the product data; and performing similarity matching on the embedded vector corresponding to the created copywriting and the semantic vector of the video slice in the database, and sorting and generating a mixed video meeting the duration requirement according to the similarity.
Owner:特赞(上海)信息科技有限公司

Multi-granularity video retrieval method and device based on multi-modal large model, computer equipment and readable storage medium

The invention discloses a multi-granularity video retrieval method and device based on a multi-modal large model, computer equipment and a readable storage medium, and the method comprises the steps: firstly obtaining video query information input by a user, carrying out the intention recognition to obtain a query field, rewriting the query information and the field to obtain a video query vector, and carrying out the retrieval of the video query vector; and retrieving in a preset retrieval video knowledge base according to the vector and the field, and finally obtaining the target retrieval video content, so as to improve the efficiency and accuracy of video retrieval and adapt to multi-field retrieval requirements.
Owner:DARK MATTER ARTIFICIAL INTELLIGENT (BEIJING) TECHNOLOGY CO LTD

Electronic device for at least one of video moment retrieval and highlight detection and operation method thereof

Proposed is an electronic device for at least one of video moment retrieval and highlight detection which includes a storage unit and a processor, wherein the processor obtains a plurality of first video features from a video, obtains a text query feature from a text query, obtains a plurality of weights from the plurality of first video features and the text query feature, obtains a plurality of second video features from the plurality of weights and the plurality of first video features, obtains a plurality of third video features from the plurality of second video features by using an encoder, obtains a plurality of fourth video features from the plurality of third video features and a time query by using a decoder, and selects at least one of time points or time periods of the video by using the plurality of fourth video features.
Owner:RES & BUSINESS FOUND SUNGKYUNKWAN UNIV +1

Multi-mode convergence media content auxiliary creation method based on AI technology

The invention discloses a multi-mode convergence media content auxiliary creation method based on an AI technology, and the method comprises the steps: extracting key features according to a creation demand text inputted by a user, carrying out the retrieval in a knowledge base through employing a knowledge graph retrieval mode according to the key features, obtaining a creation material, and generating a first draft; carrying out cross-modal feature mapping on the first draft by adopting a multi-modal alignment model, aligning text-image-video embedded vectors through contrastive learning, and dynamically adjusting the correlation of multi-modal contents by utilizing an attention mechanism to obtain a multi-modal content packet; duplicate checking is carried out on the multi-modal content packet in multiple modes, and the multi-modal content packet after duplicate checking is sent to a user for manual editing. Through multi-mode processing methods such as AI auxiliary writing, AI illustration and video generation and whole-network duplicate checking and propagation value evaluation, the content production efficiency and quality are improved, the propagation effect is optimized, and the original content is protected.
Owner:广西日报社

Video retrieval generation method and device based on sparse representation and reordering

The invention discloses a video retrieval generation method and device based on sparse representation and reordering, and relates to the technical field of cross-modal video retrieval. The method comprises the following steps: acquiring a query text, a candidate video sequence and a video retrieval generation model; obtaining a query text dense representation and a candidate video sequence dense representation according to the encoder module, the query text and the candidate video sequence; through a sparse representation module, text sparse representation is generated according to the query text dense representation, candidate video sparse representation is generated according to the candidate video sequence dense representation, and reverse index screening is performed on the candidate video sequence to obtain part of candidate videos; determining a global score of each video in the partial candidate videos through a cross attention module, and sorting the partial candidate videos; and through a generation module, generating a target text corresponding to the query text according to the sorted candidate videos and the query text. By adopting the method and the device, the retrieval efficiency is improved in a sparse representation and reordering mode.
Owner:HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN)

Character recognition model training method and apparatus, character recognition method and apparatus, device and storage medium

The present disclosure provides a character recognition model training method and apparatus, a character recognition method and apparatus, a device and a medium, relating to the technical field of artificial intelligence, and specifically to the technical fields of deep learning, image processing and computer vision, which can be applied to scenarios such as character detection and recognition technology. The specific implementing solution is: partitioning an untagged training sample into at least two sub-sample images; dividing the at least two sub-sample images into a first training set and a second training set; where the first training set includes a first sub-sample image with a visible attribute, and the second training set includes a second sub-sample image with an invisible attribute; performing self-supervised training on a to-be-trained encoder by taking the second training set as a tag of the first training set, to obtain a target encoder.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

Long video content information acquisition method based on OCR (Optical Character Recognition) and voice recognition technology

The invention discloses a long video content information acquisition method based on an OCR and voice recognition technology. The method comprises the following steps: S1, carrying out preprocessing n on input long video data to extract an image frame sequence and an audio stream; s2, inputting the image frame sequence into an OCR recognition module, inputting the audio stream into an ASR recognition module, and obtaining a preliminary recognition result; s3, constructing a multi-target fitness function, and optimizing an OCR and ASR parameter combination by using a Kanglizard optimization algorithm; s4, respectively applying the optimal parameter group to an OCR identification module and an ASR identification module to obtain an optimized identification result; s5, constructing a fusion factor graph, executing edge message passing by adopting a belief propagation algorithm, and generating a multi-modal semantic block set; and S6, processing the multi-modal semantic block set to generate a unified multi-modal content information set. According to the invention, through fusion of the horny lizard optimization algorithm and the belief propagation mechanism, high-precision recognition and multi-modal semantic consistency extraction of the image text and the voice information in the long video are realized.
Owner:华电(海西)新能源有限公司

Video space-time retrieval method and device based on grid coding

The invention discloses a grid-coding-based video space-time retrieval method and device, the grid-coding-based video space-time retrieval method is executed in computing equipment, and the method comprises the following steps: mapping a coordinate range of a target area into a corresponding grid code set according to a Beidou grid standard on the basis of the coordinate range of the target area; acquiring video data with an intersection or inclusion relationship between the spatial range of the shot object and the target area; combining the shooting time of any video data with each grid code to generate a space-time grid code set; constructing an index structure corresponding to the video data by taking each time-space grid code in the time-space grid code set as a main key; and in response to query information of the user, analyzing a space-time range in the query of the user, mapping the space-time range into a corresponding space-time grid code range, executing retrieval based on the range, and returning hit video data. According to the method, rapid positioning and accurate retrieval of large-scale aerial photography video data are realized.
Owner:BEIJING ZHIWANG YILIAN TECH CO LTD

Method and system for text search capability of live or recorded video content streamed over a distributed communication network

A server receives and rebroadcasts live streaming video content from a video capture device, such as a mobile phone or unmanned surveillance vehicle. The server includes a media server configured to stream selected video content to a client device, a video analysis system configured to analyze the live video content and generate object detection data, a storage system configured to store the generated object detection data and an identifier of the associated live video content, and a search engine configured to receive a text-based search request, search the object detection data stored in the storage system for relevant search results, and generate a list of live and stored video content associated with the relevant search results.
Owner:AERYON LABS

Shopping interface and method

A media sharing and communication system, including a recording mechanism that records a desired portion of media upon activation by a first individual user, a first user transmitter / receiver that transmits the portion of media and a message generated by the first individual user regarding the portion of media to a second individual user and is capable of transmitting a message to a second individual user, a confirmation mechanism that confirms that the second individual user is authorized to view the portion of media, a notification mechanism that notifies the first individual user if the second individual user is not authorized to receive the portion of media, a second user transmitter / receiver that receives the portion of media and voice message upon authorization of the second individual user, a search mechanism, and a video recording mechanism, an online betting module, and an online food ordering module.
Owner:TAYLOR DAVID A

Media content memory retrieval

Aspects of the subject disclosure may include, for example, a media consumption database that stores data elements describing conditions under which electronic media content is consumed by a user on an electronic device. A search of the media consumption database based on at least a portion of the conditions may result in at least a portion of the electronic media content to be re-presented to an electronic device of the user Other embodiments are disclosed.
Owner:AT&T INTELLECTUAL PROPERTY I L P

Method and system for quickly searching and positioning target in surveillance video

The invention discloses a method and a system for quickly searching and positioning a target in a surveillance video, and the method comprises the steps: carrying out the large-class marking of the target in a key frame of the surveillance video, and generating first marking data; performing coarse classification marking on the targets of each large class according to the first marking data to generate second marking data; performing fine classification marking on the roughly classified target according to the second marked data to generate third marked data; generating a hierarchical association index; calculating feature saliency; selecting the key frame with the highest feature saliency as a retrieval display picture; generating a retrieval video clip; when the retrieval input is a character, generating a first matching result; when the retrieval input is a picture, generating a second matching result; and retrieving the corresponding retrieval display pictures and the retrieval video clips from the hierarchical association index, and outputting the retrieval display pictures and the retrieval video clips according to a sequence of similarity from high to low. The method and the device are used for improving the accuracy and the efficiency of searching and positioning the target in the monitoring video.
Owner:SHENZHEN EFERCRO ELECTRONIC TECHNOLOGY CO LTD

Video editing template search method, device, electronic device and storage medium

The present disclosure relates to a video editing template search method, device, electronic device, and storage medium, in which a search keyword entered by a user is obtained, and matching is performed based on the search keyword to obtain target music that matches the search keyword and a first template video edited using the target music, and then the first template video and the target music are presented to the user in the form of a card on a search result page.
Owner:BEIJING ZITIAO NETWORK TECH CO LTD

Privacy Controls for Sharing Embeddings for Searching and Indexing Media Content

This document describes techniques and systems that enable privacy controls for sharing embeddings for searching and indexing media content. A set of images of a user's face are obtained and a machine-learned model is applied to the set of images to generate a user-specific dataset of face embeddings for the user. Media content stored in a media storage is indexed by applying the machine-learned model to the media content to provide indexed media information identifying one or more faces shown in the media content. Access to the indexed media information by another user querying the media content for images or videos depicting the user is controlled based on a digital key shared by the user with the other user, where the digital key is associated with the user-specific dataset and the user-specific dataset is usable to identify the images or videos depicting the user.
Owner:GOOGLE LLC

Internet of Things video monitoring big data privacy protection and efficient retrieval method based on artificial intelligence

The invention discloses an Internet of Things video monitoring big data privacy protection and efficient retrieval method based on artificial intelligence, belongs to the crossing field of artificial intelligence, Internet of Things and information security, and is suitable for vehicle-mounted, park, battery swap stations and other scenes. The method comprises the following steps: firstly, establishing a self-adaptive acquisition framework, and realizing multi-protocol switching and video preprocessing; secondly, extracting privacy information through an improved YOLO algorithm and a Graph-Cut technology, and combining reversible watermark embedding; constructing a multi-dimensional privacy level model for hierarchical encryption, and matching two-level storage and three-level index; then, two-factor authentication authorization is performed to extract privacy and optimize retrieval; and finally, the system state is monitored in real time and self-adaptive optimization is performed. According to the method, the problems of protocol heterogeneity, insufficient privacy protection, low retrieval efficiency and the like can be solved, and privacy security, storage overhead and retrieval efficiency are balanced.
Owner:ZHEJIANG HAISHI HUAYUE DIGITAL TECHNOLOGY CO LTD

Method and system for three-dimensional modeling of chemical industry park

The invention discloses a three-dimensional modeling method and system for a chemical industrial park, and the method comprises the steps: carrying a multi-lens inclined camera and a monitoring sensor through employing an unmanned plane, and obtaining the multi-view image data, positioning and attitude determination data of the chemical industrial park, and the real-time operation data of chemical equipment; performing data processing on the multi-view image data to generate sparse point clouds, and generating three-dimensional point clouds through dense matching and depth estimation; performing block processing on the three-dimensional point cloud, and performing splicing by adopting a point cloud registration algorithm to obtain global point cloud data; performing surface reconstruction on the global point cloud data to generate a three-dimensional grid model, and completing model texture mapping through a texture mapping algorithm; a three-dimensional visual management platform is constructed, a three-dimensional grid model and real-time monitoring data are integrated, and park global visualization, risk visualization and intelligent management of chemical engineering devices, equipment and facilities are achieved.
Owner:BEIJING UNIV OF CHEM TECH

Automatic traffic accident disposal system based on large model and block chain technology

The invention relates to a traffic accident automatic handling system based on a large model and a block chain technology. The traffic accident automatic handling system comprises a traffic accident analysis intelligent module, a traffic accident liability confirmation intelligent module, a traffic accident report auditing intelligent module and a whole process data storage and processing module based on the block chain technology. The three intelligent modules cooperate with one another by simulating the thinking modes of different roles in the real world disposal process to jointly complete automatic identification of traffic accidents. According to the invention, the automation of traffic accident disposal is realized, the accuracy and fairness of traffic accident identification are improved, and the traceability of accident original data and the transparency of the processing process are ensured.
Owner:HANGZHOU DIANZI UNIV

Systems and methods for few-shot new action recognition

A method includes: (i) receiving a query video including performance of an action; (ii) receiving a predetermined number of support videos including performance of actions, respectively, the predetermined number of support videos being less than 100 support videos; (iii) determining a similarity matrix based on a comparison of temporally ordered images of the query video with temporally ordered images of one of the support videos, respectively; (iv) determining a similarity value for the one of the support videos based on the similarity matrix; (v) repeating (iii) and (iv) for each of the support videos; (vi) identifying the highest one of the similarity values and the one of the support videos associated with the highest one of the similarity values; and (vii) setting a first indicator of the action in the query video to the same as a second indicator of the action performed in the one of the support videos.
Owner:CZECH TECH UNIV IN PRAGUE +1

Text image retrieval model training method, system and equipment and storage medium

The invention provides a text image retrieval model training method, system and device and a storage medium, and is applied to the technical field of medical image retrieval, and the method comprises the following steps: carrying out multi-level cross-modal alignment relation extraction on a to-be-processed video; gaussian noise addition is carried out on continuous picture frames in the video to be processed, and the picture frames after Gaussian noise addition are cut into a plurality of image blocks; training an image encoder by taking the image block feature corresponding to the current image block, the time sequence feature of the current image block corresponding to the previous frame of image block and the space-time position coding feature corresponding to the current image block as input features; positive and negative samples corresponding to the multi-level cross-modal alignment relationship are coded, generated text coding features and image coding features are mapped to the same semantic space, and parameters of a text and image retrieval model are optimized, so that the recall rate of the retrieval model is increased, the accuracy of a retrieval result is ensured, and the retrieval efficiency is improved. And reliable video information support is provided for medical decision and medical research.
Owner:PING AN TECH (SHENZHEN) CO LTD

Systems and methods for generating improved content based on matching mappings

Systems and methods are disclosed herein for generating content based on matching mappings by implementing deconstruction and reconstruction techniques. The system may retrieve a first content structure that includes a first object with a first mapping that includes a first list of attribute values. The system may then search content structures for a matching content structure having a second object with a second list of attributes and a second mapping including second attribute values corresponding to the second list of attributes. Upon finding a match, the system may generate a new content structure having the first object from the first content structure with the second mapping from the matching content structure. The system may then generate for output a new content segment based on the newly generated content structure.
Owner:ADEIA GUIDES INC

Editing strategy scheduling method and device, electronic device and storage medium

The invention relates to an editing strategy scheduling method and device, an electronic device and a storage medium, and the method comprises the steps: receiving an original material, a copywriting and an editing style instruction inputted by a user, and generating a corresponding dubbing audio according to the copywriting; performing multi-dimensional analysis on the original material to generate a lens-level structured index; performing semantic analysis on the copywriting to obtain copywriting semantic features; performing similarity retrieval based on the copywriting semantic features and the multi-modal semantic features of the materials to obtain a candidate shot set matched with the copywriting semantic features; generating a global style vector and a target rhythm curve through the first agent; and taking the global style vector and the target rhythm curve as control signals, driving a second agent to select and trim shots from the candidate shot set, recombining the editing sequence, and outputting a final editing sequence and an editing jump point structure. According to the style vector, the rhythm target curve and reinforcement learning, the editing style is met, and the intelligent agent editing stylized presentation is achieved.
Owner:ZHEJIANG HUAZHI WANXIANG TECHNOLOGY CO LTD

Operation process management method and system

The invention relates to an operation process management method and system. The method comprises the following steps: determining node events of different operation stages based on the number and positions of predetermined monitoring objects in an operating room; acquiring a real-time monitoring video in the operating room, identifying the number and the position of a predetermined monitoring object in the real-time monitoring video, and judging whether to trigger a node event of each operation stage; if the node event is triggered, determining starting and ending time of each operation stage based on the triggered node event; and storing the operation data according to stages and automatically generating operation record information. Existing camera equipment in the operating room is fully utilized, the image recognition technology is adopted, different stages in the operation process are automatically distinguished on the basis of the number and position relation of personnel and medical equipment in the operating room, operation records are automatically updated, time errors caused by manual operation are reduced, and the operation efficiency is improved. And the accuracy and the efficiency of the operation process recording are obviously improved.
Owner:QINGDAO WANMUCHUN MEDICAL TECH CO LTD

Ammonia water leakage detection intelligent management system and detection method

The invention provides an ammonia water leakage detection intelligent management system and a detection method, and belongs to the technical field of production safety in the polycrystalline silicon industry. The system comprises an explosion-proof law enforcement instrument which is used for shooting and uploading a telescopic rod leakage detection video in the whole process; the ammonia water leakage detection telescopic rod is used for detecting the leakage condition; the mobile phone end leakage checking module is used for generating a leakage checking task and submitting an execution condition and a leakage checking result of the leakage checking task; the video management platform is used for storing the inspection video uploaded by the explosion-proof law enforcement instrument; and the production informatization management system is used for generating a leakage checking machine account, performing intelligent re-checking on a leakage checking result, and performing leakage checking and random checking. Through the system, the efficiency and the accuracy of leak detection work are improved.
Owner:中国化学品安全协会 +1

Media content memory retrieval

Aspects of the subject disclosure may include, for example, a media consumption database that stores data elements describing conditions under which electronic media content is consumed by a user on an electronic device. A search of the media consumption database based on at least a portion of the conditions may result in at least a portion of the electronic media content to be re-presented to an electronic device of the user Other embodiments are disclosed.
Owner:AT&T INTELLECTUAL PROPERTY I L P

End-to-end multi-task video retrieval with cross attention

A method includes obtaining a video and a relational spatio-temporal query, and identifying at least one type of the relational spatio-temporal query. The at least one type of identification of the relational spatio-temporal query represents at least one of: an activity type, an object type, or a temporal type. The method further includes learning correlations between activities, objects, and time in the video using one or more cross attention models. The method further includes obtaining one or more predictions generated using one or more outputs of the one or more cross attention models based on the identified at least one type of relational spatiotemporal query. Further, the method includes generating a response to the relational spatiotemporal query based on the one or more predictions.
Owner:SAMSUNG ELECTRONICS CO LTD

system

We provide the system. [Solution] A means for automatically identifying the voice of a specific person, adding a timestamp, and recording it as audio data, A means of recognizing an object held by a specific person as image data and recording it with a timestamp, A means for editing and generating a natural language record based on the above audio data and image data, A means for automatically transmitting this record to an external communication application via a communication network, A means for performing behavior recognition within a specific environment and generating messages corresponding to specific behaviors, A means of transmitting the generated message to an information terminal, A system that includes this.
Owner:SOFTBANK GROUP CORP

Camera gun video resource positioning method and related equipment

The invention discloses a camera gun video resource positioning method and related equipment. The method comprises the following steps: acquiring camera gun information of each financial network; obtaining structured risk model data for the daily recovery high-risk model troubleshooting index; determining a target financial network point corresponding to each piece of risk model data in the structured risk model data; according to the structured risk model data, the target financial network point corresponding to each piece of risk model data, the camera gun information of each financial network point and the current popularity value of each camera gun, determining a plurality of current retrieval camera guns; according to the current user operation log, updating the current popularity value corresponding to each camera gun; and according to the updated current popularity value corresponding to each camera gun, determining a plurality of updated current retrieval camera guns. The method can achieve the precise pushing of the video resources needed by the daily redisk of the website, improves the video verification efficiency, and can be widely applied to the technical field of artificial intelligence.
Owner:GUANGDONG BRANCH OF CHINA POST GRP CO LTD

Fine-grained video clip retrieval method based on semantic concept decomposition

The invention discloses a fine-grained video clip retrieval method based on semantic concept decomposition, which comprises the following steps of: firstly, constructing a training sample set, constructing a video clip retrieval model, respectively extracting text concept representation and video concept representation for a query and a video, generating a query correlation matrix according to the query, respectively carrying out sparse concept merging on the text concept representation and the video concept representation by utilizing the query correlation matrix, then fusing, and decoding by a moment decoding module to obtain the starting time and the ending time of a video clip; and training the video clip retrieval model by adopting the training sample set, and subsequently performing video clip retrieval by adopting the trained video clip retrieval model. According to the method, the text concept representation and the video concept representation are extracted on the basis of semantic concept decomposition, and fine-grained concepts in semantics can be more effectively captured and recognized, so that the performance and efficiency of video clip retrieval are improved.
Owner:UNIV OF ELECTRONICS SCI & TECH OF CHINA