Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

460results about "Video data clustering/classification" patented technology

Model deployment method, end-side device, and storage medium

The present disclosure relates to the technical field of target detection, and particularly relates to a model deployment method, an end-side device and a storage medium, which are used for solving the problem in the related art of the accuracy of a deployed model being low. The method comprises: performing target detection on a video frame image input into a first model, and acquiring a first target detection result and a first confidence; if the first confidence is greater than or equal to a first confidence threshold value, recording the video frame image and the first target detection result as samples in a training set; if the first confidence is less than the first confidence threshold value, performing target detection on the video frame image on the basis of a second model, and recording the video frame image and an acquired second target detection result as samples in the training set; and training the first model on the basis of the training set, and replacing the current first model with a trained first model for subsequent target detection. In this way, the accuracy and model generalization capability of a first model are improved.
Owner:HISENSE GRP HLDG CO LTD

Semantic-tree-based ai content management platform

A data processing system implements receiving a call requesting a generative model to generate a semantic tree for a source content; constructing a first prompt including the source content and instructions to the model to analyze a semantic structure of the source content and to generate a semantic outline and content chunks of the source content, the semantic outline including one or more topics each connected with one or more of the content chunks, to compute one summary for each of the content chunks, to apply indices to reference each topic node of the semantic tree to one of the topics, and to apply indices to reference each leaf node of the semantic tree to one of the content chunks and the respective summary; providing the first prompt to the model and receiving the semantic tree of the source content; and storing the semantic tree in a database.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Systems and methods for multimodal indexing of video using machine learning

Systems, methods, and computer-readable media are disclosed for systems and methods multimodal indexing of video using machine learning. An example method may include deceiving, by a video encoder of an audio-video transformer neural network comprising one or more computer processors coupled to memory, a first frame and a second frame associated with a first segment of a video. The example method may also include receiving, by an audio encoder of the audio-video transformer neural network, an audio spectrogram comprising first audio data associated with the first segment of the video. generating, by the video encoder, a first video embedding. The example method may also include generating, by the audio encoder, a first audio embedding. The example method may also include determining a fusion of the first video embedding and the first audio embedding using a multimodal bottleneck token. The example method may also include determining an output including the first video embedding and the first audio embedding. The example method may also include determining a classification of the first portion of the video based on the output.
Owner:AMAZON TECH INC

Visual display method and system applied to security and protection and storage medium

The invention relates to the technical field of security and protection visualization, in particular to a security and protection visualization display method and system and a storage medium. The method comprises the following steps: acquiring a data stream acquired by security and protection equipment, and compressing the data stream into a unified acquisition frame set; identifying a road boundary point set in the unified collection frame set, and detecting a dynamic target and a static target at each moment in the unified collection frame set; performing semantic annotation according to the motion state change condition of the dynamic target at each moment in the unified collection frame set to obtain a road boundary semantic distribution map; according to the road boundary semantic distribution map, identifying and predicting the motion state change condition of the static target; and predicting an overlapping time point of the motion state of the target based on the motion state change condition of the static target and the motion state change condition of the dynamic target, and transmitting the overlapping time point to communication equipment of the static target through the Internet of Things so as to execute a voice prompt task. According to the invention, the display precision and response efficiency of security and protection visualization can be improved.
Owner:JINGGANGSHAN YUJIE FIRE SCI & TECH

Semantic-tree-based ai content management platform

A data processing system implements receiving a call requesting a generative model to generate a semantic tree for a source content; constructing a first prompt including the source content and instructions to the model to analyze a semantic structure of the source content and to generate a semantic outline and content chunks of the source content, the semantic outline including one or more topics each connected with one or more of the content chunks, to compute one summary for each of the content chunks, to apply indices to reference each topic node of the semantic tree to one of the topics, and to apply indices to reference each leaf node of the semantic tree to one of the content chunks and the respective summary; providing the first prompt to the model and receiving the semantic tree of the source content; and storing the semantic tree in a database.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Film and television supervision analysis method and system based on data mining and storage medium

The invention relates to the technical field of data processing, and discloses a film and television supervision analysis method and system based on data mining and a storage medium. The method comprises the following steps: collecting film and television data in a panoramic manner, and carrying out cross-media transcoding processing to obtain a multi-dimensional data set; performing video semantic extraction, audio emotion recognition and text tendency analysis on the data set to generate a time sequence mark list; scoring the key elements according to social influence, propagation depth and audience acceptability to form a dynamic score table; analyzing the propagation path based on the score table, and constructing an influence map; sorting supervision points according to the atlas, and making an intelligent supervision scheme; the execution deviation is analyzed through effect feedback, and the supervision parameters are optimized. According to the method, multi-dimensional content feature extraction, social influence evaluation and propagation path analysis of the film and television works are realized in a mass multimedia content environment, a precise and differentiated intelligent supervision scheme is formed, and the method has self-optimization and adjustment capabilities at the same time.
Owner:HANGZHOU RUNHAO CULTURE MEDIA CO LTD

Adaptive sample selection for data item processing

Methods, systems, and apparatuses, including computer programs encoded on computer storage media, for receiving a query relating to a data item that includes multiple data item samples and processing the query and the data item to generate a response to the query. In particular, the described techniques include adaptively selecting a subset of the data item samples using a selection neural network conditioned on features of the data item samples and the query. Then processing the subset and query using a downstream task neural network to generate a response to the query. By adaptively selecting the subset of data item samples according to the query, the described techniques generate responses to queries that are more accurate and require less computation resources than would be the case using other techniques.
Owner:GOOGLE LLC

Systems and methods for trigger-based updates to camograms for autonomous checkout in a cashier-less shopping

Systems and methods for tracking inventory items in an area of real space are disclosed. The method includes receiving a signal generated in dependence on sensors. The signal indicates a change to a portion of an image of an area of real space. The method includes, in response to receiving the signal, implementing a trained location detection model to determine, based on inputs, whether an inventory item identified in the portion of the image has changed a position in the area of real space. The method includes implementing a trained item classification model to determine a classification of the inventory item. The method includes updating an inventory database with inventory item data determined in dependence on the classification of the inventory item to provide an updated map of the area of real space as a result of the received signal indicating the change to the portion of the image.
Owner:STANDARD COGNITION CORP

Media information processing method and device, storage medium and electronic equipment

The invention discloses a media information processing method and device, a storage medium and electronic equipment. The method comprises the steps that to-be-retrieved source media information and prompt template information are obtained, the source media information and the prompt template information are input into a target retrieval model, target media information is determined, and the target retrieval model represents a model obtained through joint training according to at least two retrieval task types; a loss function used in the training process of the target retrieval model is a joint loss function, the joint loss function is determined by at least two sub-loss functions in one-to-one correspondence with the at least two retrieval task types, and the target media information represents media information obtained by retrieving the source media information according to the prompt template information. According to the method and the device, the technical problem of relatively low retrieval efficiency of the media information caused by single multi-modal retrieval task is solved.
Owner:TENCENT TECH (BEIJING) CO LTD

Multi-granularity video retrieval method and device based on multi-modal large model, computer equipment and readable storage medium

The invention discloses a multi-granularity video retrieval method and device based on a multi-modal large model, computer equipment and a readable storage medium, and the method comprises the steps: firstly obtaining video query information input by a user, carrying out the intention recognition to obtain a query field, rewriting the query information and the field to obtain a video query vector, and carrying out the retrieval of the video query vector; and retrieving in a preset retrieval video knowledge base according to the vector and the field, and finally obtaining the target retrieval video content, so as to improve the efficiency and accuracy of video retrieval and adapt to multi-field retrieval requirements.
Owner:DARK MATTER ARTIFICIAL INTELLIGENT (BEIJING) TECHNOLOGY CO LTD

Intelligent conference recording and recording method and system

The invention discloses an intelligent conference recording and recording method and system. According to the invention, cloud service equipment receives multi-modal information sent by edge equipment; processing the multi-modal information, and generating conference record information in combination with context sensing error correction, content segmentation, a dynamic segmentation strategy and a preset storage strategy; carrying out encryption processing on the conference record information based on the participant role information, and generating target format conference information; processing the target format conference information to generate initial conference scene mode information including a conference outline, a PPT and a mind map; processing the target format conference information and the initial conference scene mode information to generate conference information data after priority ranking and strategy processing; performing semantic analysis, feature fusion and cross-modal association processing on the conference information data to generate target conference scene mode information; and determining a target application program based on the target conference scene mode information, and sending the target format instruction information to the target application program.
Owner:XIAMEN RGBLINK SCI & TECH CO LTD

Storing feature vectors in one or more memory processing units

Disclosed embodiments include a computational memory system. The computational memory system includes at least one computational memory chip including one or more processor subunits and one or more memory banks formed on a common substrate. The at least one computational memory chip is configured to store one or more portions of an embedding table in the one or more memory banks, the embedding table including one or more feature vectors. The one or more processor subunits are configured to receive a sparse vector indicator from a host external to the at least one computational memory chip and, based on the received sparse vector indicator and the one or more portions of the embedding table, generate one or more vector sums.
Owner:NEUROBLADE LTD

Methods and systems for segmenting video content based on speech data and for retreiving video segments to generate videos

A method includes receiving a series of video segments and providing the series of video segments as input to a first machine learning model to produce text data. The text data is provided as input to a second machine learning model to produce categorized text data that includes a classification indication. The classification indication is added to metadata of the video segment, and the categorized text data is provided as input to a third machine learning model to produce a semantic vector. The method also includes causing the video segment and the metadata that includes the classification indication to be stored at a location of a database based on the semantic vector, the database being configured to be searched based on a search query associated with the semantic vector.
Owner:VIDEOFORCEAI INC

Multi-modal accident scene library construction method for end-to-end automatic driving test

The invention belongs to the technical field of automatic driving test, and particularly relates to a multi-mode accident scene library construction method for end-to-end automatic driving test. The method specifically comprises the following steps: step 1, extracting text information based on an MCAT-BiLSTM-CRF algorithm; 2, designing the ontology architecture, and storing accident scene information by using a knowledge graph to obtain an accident text knowledge base; step 3, derivative expansion is carried out on the accident scene element combination based on a HyCon-Sg-Net algorithm; step 4, performing enhanced fine tuning on the Open-Sora model to realize modal conversion from the text knowledge base to the accident video database so as to construct a multi-modal accident scene library; the method can be used for testing the performance of an end-to-end automatic driving automobile in an extreme accident scene, and the adaptive capacity of an automatic driving algorithm to the extreme accident scene is remarkably improved.
Owner:JILIN UNIVERSITY

Video segmentation method, server, storage medium, and program product

The present application provides a video segmentation method, a server, a storage medium, and a program product. In the method of the present application, video data to be segmented is segmented into multiple data segments, unimodal features of the data segments, including text features of a text modality and visual features of a visual modality, are respectively extracted by means of a video topics segmentation model, and then the text features and visual features of the data segments are fused, so that the fusion of multimodal information can be performed at the intermediate representation level, the relationship and interaction between different modalities can be better captured, and higher-quality multimodal fusion features of the data segments are obtained. Furthermore, on the basis of the multimodal fusion features of the data segments, whether the data segments are topic boundaries is predicted, so that the topic boundaries of the video data can be accurately predicted, improving the accuracy of topic boundary recognition, thereby improving the accuracy and quality of video topics segmentation results.
Owner:ALIBABA (CHINA) CO LTD

Video retrieval method and device, electronic equipment and storage medium

The invention discloses a video retrieval method and device, electronic equipment and a storage medium. The method comprises the following steps: acquiring each current to-be-retrieved video; dividing each video to be retrieved into a plurality of split videos; for each target retrieval dimension, obtaining a text description of the target retrieval dimension of each split video; wherein the target retrieval dimension comprises a subtitle dimension, an audio dimension and a visual dimension; the text description of the visual dimension of one split video is combined by the context text description of the split video and the text description of each key frame of the split video; combining the text description of the target retrieval dimension of each split video with a retrieval problem, and performing retrieval enhancement to obtain a plurality of retrieval videos based on the target retrieval dimension; weighting the text description of each target retrieval dimension of each retrieval video to obtain a plurality of sub-mirror multi-dimensional text descriptions; and combining the multi-dimensional text description of each sub-mirror with a retrieval problem to carry out retrieval enhancement to obtain a final retrieval result.
Owner:HUNAN HAPPLY SUNSHINE INTERACTIVE ENTERTAINMENT MEDIA CO LTD

Public opinion video tag aggregation method and system based on artificial intelligence

The invention provides a public opinion video tag aggregation method and system based on artificial intelligence, and relates to the technical field of artificial intelligence. Comprising the following steps: acquiring pictures and text information in a short video, and performing semantic alignment; different large language models are adopted to generate preliminary labels for the pictures and the text information after semantic alignment; clustering the pictures and the texts after semantic alignment to obtain clusters; calculating the labeling probability of each primary label type in the current cluster by each large language model, and selecting the primary label with the highest probability sum as a clustering label of the current cluster; calculating the reliability weight of each large language model in the current cluster based on the clustering label of the current cluster; and based on the reliability weight, calculating the weighted support degree of all the large language models to different preliminary label types of each piece of data in the current cluster, calculating the weighted label of the current data, and further determining a final label. According to the method, the condition of few labels or no labels can be effectively processed, and the manual workload is greatly reduced.
Owner:SHANDONG DAZHONG INFORMATION IND CO LTD

Video tag processing method and apparatus, and computer device and storage medium

A video tag processing method, comprising: determining a video to be processed, and performing extraction on the video to obtain modality information of at least two modalities (202); on the basis of the modality information of the at least two modalities, respectively performing cross retrieval according to feature dimensions of the at least two modalities, so as to obtain respective multi-modal retrieval tags of the modality information of the at least two modalities (204); on the basis of the respective multi-modal retrieval tags of the modality information of the at least two modalities, determining at least one candidate tag for the video (206); and on the basis of the at least one candidate tag and the modality information of the at least two modalities, performing tag prediction, so as to obtain a video tag for the video (208).
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Video scene content label determination method based on knowledge graph

The invention relates to the technical field of video content analysis, and discloses a video scene content label determination method based on a knowledge graph. The method comprises the following steps: inputting a multi-modal feature sequence formed by visual, audio and text features extracted from a video stream into a knowledge graph inference engine comprising an entity relationship network and a semantic association rule base; performing node mapping through an engine to generate an initial scene entity set; performing hierarchical reasoning on the set based on a semantic association rule base to obtain a scene semantic topological structure; screening core entities according to entity weight distribution in the topological structure, and generating a candidate tag set; performing time sequence consistency verification on the candidate tag set, and correcting tag time sequence offset in combination with timestamp information; performing cross-modal disambiguation on the corrected label set by using an entity relationship network to eliminate semantic conflicts; and according to a disambiguation result, constructing a scene label knowledge sub-graph containing entity attributes and relation path constraints.
Owner:GUANGZHOU JUNHE INFORMATION TECH CO LTD

Multi-modal content based automated feature recognition

A system includes a computing platform having processing hardware, and a memory storing software code and a machine learning (ML) model-based feature classifier. When executed, the software code receives media content including a first media component corresponding to a first media mode and a second media component corresponding to a second media mode, encodes the first media component using a first encoder to generate multiple first embedding vectors, and encodes the second media component using a second encoder to generate multiple second embedding vectors. The software code further combines the first embedding vectors and the second embedding vectors to provide an input data structure for a neural network mixer, process, using the neural network mixer, the input data structure to provide feature data corresponding to a feature of the media content, and predict, using the ML model-based feature classifier and the feature data, a classification of the feature.
Owner:DISNEY ENTERPRISES INC

Semantic-tree-based AI content management platform

A data processing system implements receiving a call requesting a generative model to generate a semantic tree for a source content; constructing a first prompt including the source content and instructions to the model to analyze a semantic structure of the source content and to generate a semantic outline and content chunks of the source content, the semantic outline including one or more topics each connected with one or more of the content chunks, to compute one summary for each of the content chunks, to apply indices to reference each topic node of the semantic tree to one of the topics, and to apply indices to reference each leaf node of the semantic tree to one of the content chunks and the respective summary; providing the first prompt to the model and receiving the semantic tree of the source content; and storing the semantic tree in a database.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Computer-implemented system and method for providing a stream or playback of a live performance recording

The present disclosure relates to a computer-implemented system and method for providing streams and playbacks of video and audio recordings in respect of artist performances before live audiences. In particular, the disclosure relates to a system and method of providing a stream or playback of video and audio recordings captured in respect of an artist's live performance to users located around the world in exchange for a fee, with particular limitations placed upon access to the recording until the fee is paid.
Owner:SLOSS GLENN

Clustering algorithm-based education platform course optimization recommendation system and method

The invention discloses an education platform course optimization recommendation system and method based on a clustering algorithm, and relates to the technical field of course recommendation. All reserved similar courses are divided into a plurality of display clusters through the clustering algorithm, and after historical behavior data of a user is analyzed, the plurality of display clusters are subjected to display priority ranking; the display clusters can be sorted from top to bottom on a platform display page and can also be sorted in a sliding superposition mode, and after historical use data of similar courses in each display cluster are obtained, a sorting index of each similar course is generated through a regression analysis algorithm; and dynamically adjusting the recommendation sequence of all courses in the display cluster according to the sorting index. According to the recommendation system, the similar courses are subjected to clustering analysis to obtain the display cluster, all the similar courses in the display cluster are sorted and then recommended to the user, the course recommendation effect is improved, and the learning experience of the user is effectively optimized.
Owner:HUNAN ANNA INTELLIGENT TECH CO LTD

Time sequence statement positioning model training method based on proposal selection and anchor point distribution

The invention discloses a timing sequence statement positioning model training method and device based on proposal selection and anchor point distribution, and relates to the technical field of timing sequence statement positioning. The method comprises the following steps: initializing a fixed number of learnable queries according to a first static anchor point set; on the basis of the first learnable query, according to the unpruned long video and the natural language description text, performing proposal generation through a time sequence statement positioning model; based on the proposal selection module, redundant proposal filtering is carried out by using a non-maximum suppression algorithm; based on an anchor point distribution module, according to the untrimmed long video, the first static anchor point set and the candidate proposal set, performing bipartite graph matching by using a Hungary algorithm; and carrying out loss function weighting calculation according to the unpruned long video and the optimal proposal set, and carrying out parameter optimization on the time sequence statement positioning model to obtain an optimized time sequence statement positioning model. The method is an efficient and accurate time sequence statement positioning model training method combining proposal selection and anchor point distribution.
Owner:UNIV OF SCI & TECH BEIJING

Interrogation record and audio and video recording linkage method and system

The invention provides an interrogation record and audio and video recording linkage method and system, and belongs to the technical field of information processing, and the method comprises the steps: obtaining interrogation record text data and segmented video recording data, and respectively labeling category labels; performing semantic extraction on the record text data to obtain semantic feature information, and converting the feature information into text semantic feature vectors; frame extraction is carried out on the video data, image features of each frame of the video data are extracted, similarity screening is carried out, and the image features are converted into key frame semantic feature vectors; constructing a cross-modal embedding space model, respectively encoding the text semantic feature vector and the key frame semantic feature vector of the category label through two encoders, then splicing and fusing to obtain a fused feature vector, converting and mapping the fused feature vector into a sharing layer, completing feature association of the record text data and the segmented video recording data, and obtaining the segmented video recording data. And subsequently, cross-modal search is carried out. According to the method, cross-modal information association is realized through association of the semantic feature vectors of the record and video recording information, and a basis is provided for subsequent rapid calling, research and judgment.
Owner:SHAANXI POLICE VOCATIONAL COLLEGE (SHAANXI POLITICAL & LEGAL MANAGEMENT CADRE COLLEGE)

Data processing method and system for online intelligent infringement comparison

The invention discloses a data processing method and system for online intelligent infringement comparison, and relates to the technical field of electronic digital data processing.The method comprises the steps that firstly, short video data are obtained, text analysis and audio feature analysis are conducted on the short video data, and therefore a first similarity index of each short video is calculated; then, pre-analyzed short videos are screened out according to the first similarity indexes, the image content of the pre-analyzed short videos is further analyzed, finally, an infringement comparison judgment result is obtained, a user is prompted, text, audio and image features are integrated, the comprehensiveness and accuracy of infringement judgment are ensured, and the user experience is improved. The infringement comparison efficiency and effect are greatly improved, potential infringement behaviors can be quickly and accurately recognized, and the rights and interests of content creators are protected.
Owner:ORIGINAL GUARD (WUXI) TECH CO LTD

Method to generate a database for synchronization of a text, a video and / or audio media

The present document discloses a method to generate a database for synchronization of a text, a video and / or audio media, namely a computer-implemented method to generate a database for synchronization of a text, a video and / or audio media with a video stream to be generated using said database, comprising the steps of: receiving an input text; splitting the input text into text segments; clustering the text segments into at least one text group; labelling each text segment with a sequential timestamp; and, storing each text segment with the label on a data record wherein the text group comprises a time interval label which corresponds to the duration of at least a portion of a video and / or audio media. It is further disclosed a method for retrieving information from said database, a system, and a computer program thereof.
Owner:UNIVERSITY OF MINHO

Intelligent event video retrieval data storage and display method

The invention relates to an intelligent event video retrieval data storage and display method, which comprises the following steps: acquiring video data of each video source, and inserting a track area number of a track area where a target appears in a video frame in which the target is identified to exist; according to the target identifier of the target, the time information and the track area number, track data are obtained and stored; based on the track area number and the retrieval condition of the time information, retrieving from the track data to obtain a target identifier of the corresponding target; corresponding target video data are obtained through integration based on the target identifier; and playing the target video data corresponding to the target identifier in a display window, and displaying the corresponding target attribute. By means of the video retrieval method and device, the track area number of the target can be inserted into the video frame with the target, the track data are formed and stored, the target video data of the needed target can be rapidly obtained and played based on the track area number and the time information during retrieval, and the problem that the video retrieval efficiency is low is solved.
Owner:ZHEJIANG DAHUA TECH CO LTD +1

Intelligent film and television content recommendation method and system based on artificial intelligence

The invention discloses a film and television content intelligent recommendation method and system based on artificial intelligence. The method comprises the steps of multi-source data collection, data optimization, group film and television content recommendation, personalized recommendation optimization sorting and film and television content intelligent recommendation. The invention relates to the technical field of data processing, in particular to an intelligent film and television content recommendation method and system based on artificial intelligence. According to the scheme, group recommendation and personalized optimization are innovatively combined, and dynamic balance between group representativeness and individual difference is achieved; a density correction mechanism, weighted distance calculation, a representative point merging mechanism and a fuzzy weighted distribution strategy are innovatively proposed to improve the clustering algorithm, so that the clustering accuracy is improved, and the group recommendation precision and the user satisfaction are improved; and an adaptive reference individual selection strategy and an iteration-based global exploration weight improvement optimization algorithm are adopted, so that the accuracy of a model output result is remarkably improved, and the stability of personalized film and television content recommendation optimization sorting is realized.
Owner:CHINA UNICOM VIDEO TECH CO LTD

Method for classifying and controlling transmission of a file

A method and system for classifying a video file within an environment in which the file is located and when a file is classified as sensitive, controlling transmission of the file outside the environment. Classifying the video comprises analysing, using at least one machine learning model, the video to recognise any individuals in the video; obtaining a transcript of any speech in the video and generating, using the analysis, obtained transcript and a database of individuals linked to the environment, a labelled transcript which identifies each individual linked to the environment that is in the video. Information about each identified individual may be obtained from a connected database. A first generative AI model generates a text-based summary of the video by using the labelled transcript and information about identified individuals as prompts. A second generative AI model then determines a sensitivity classification of the video using the generated text-based summary.
Owner:VARONIS SYSTEMS INC