Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

41results about "Video data clustering/classification" patented technology

Capturing objects in an unstructured video stream

ActiveUS12651455B2Image analysisVideo data clustering/classificationPattern recognitionComputer graphics (images)
A method includes obtaining a first unstructured video stream that provides pixel values for a plurality of pixels and corresponds to a portion of a second unstructured video stream being displayed on a second electronic device different from the first electronic device. Obtaining the first unstructured video stream includes obtaining pass-through image data including the portion of a second unstructured video stream. The method includes generating respective pixel characterization vectors for a first portion of the plurality of pixels. Generating each of the respective pixel characterization vectors includes determining a respective instance label value. The method includes identifying a first object within the first portion of the plurality of pixels associated with a particular instance label value. The method includes generating respective semantic label values corresponding to pixels associated with the first object. The respective semantic label values are added to pixel characterization vectors associated with the first object.
Owner:APPLE INC

Coordinated control method and system of fitness videos and training equipment

The present application relates to the fields of sensor technology and Internet of Things, and provides a fitness video and training equipment cooperative control method and system, which comprises: collecting a cloud-based fitness teaching video stream to extract a sequence of key points of a skeleton of a demonstration action, an action intensity level and a standard motion trajectory, collecting real-time physiological data and real-time motion posture of a user; identifying a target action of each stage in the fitness teaching video stream, generating a target control parameter of a training equipment corresponding to the target action to construct a video-parameter mapping library of the training equipment; determining a physical fitness adaptation coefficient of the user, calculating an action completion degree and a motion risk coefficient of the user, and generating a safety control threshold of the training equipment; and generating a cooperative control instruction of the training equipment to execute cooperative control of the training equipment. The present application can improve the adaptability and safety of the training equipment cooperative control.
Owner:DONGGUAN BOQUN ELECTRONIC SCI & TECH CO LTD

Method, electronic device and computer program product for managing large number of videos

PendingCN122112304AVideo data clustering/classificationMetadata video data retrievalMonitoring systemTerminal equipment
The present disclosure relates to a method for managing a large number of videos, an electronic device and a computer program product, which are applied to a monitoring system, including detecting an event and outputting a large number of videos according to the event; labeling a plurality of tags for the large number of videos, the tags including a plurality of attributes of at least one object in the large number of videos; receiving video search information; comparing the relevance of the video search information and the tags of the large number of videos, and outputting a plurality of recommended videos in a terminal device according to the relevance, the terminal device including a display screen; detecting the resolution of the display screen, and adaptively playing the recommended videos in a user interface on the display screen according to the resolution. The method of the present disclosure uses attribute tag indexing technology closest to human search logic, effectively manages a large number of videos, automatically searches for relevant video results according to input information, greatly simplifies the search process and improves user experience.
Owner:WISTRON CORP

Image recognition-based leather product comparison identification method and system

PendingCN122265812AVideo data indexingVideo data clustering/classificationReference sampleEngineering
The application relates to the technical field of image recognition, and particularly discloses a leather product comparison and identification method and system based on image recognition. The application extracts a multi-dimensional standardized feature vector reflecting inherent physical and mechanical properties by collecting dynamic deformation videos of a to-be-inspected leather and a reference sample under controlled micro force, receives identification task description parameters, maps core discriminant features from a physical and mechanical feature knowledge base, dynamically instantiates a self-adaptive identification model through a meta-learning model, generates a feature weighting scheme and a dynamic decision threshold, calculates a mechanical feature matching degree and generates a visual preliminary report, adjusts the scheme in combination with user interactive correction instructions, updates the result and feeds back data optimization meta-models. The application upgrades the identification basis to essential mechanical properties, realizes task self-adaptive decision and man-machine collaborative optimization, can improve identification reliability and scene adaptability, makes the decision process transparent and interpretable, has a continuous optimization capability, and is suitable for multi-class requirements such as authenticity identification and traceability.
Owner:海宁中国皮革城网络科技有限公司

Video frame processing method and system based on semantic guidance, and storage medium

PendingCN122112302AOvercome the shortcoming of easily selecting redundant informationFilter out background noiseSemantic analysisVideo data clustering/classificationPattern recognitionVideo retrieval
The application discloses a kind of based on semantic guide's video frame processing method, system and storage medium, comprising: obtaining original video, obtains semantic guide text;The time position information of each frame is encoded after fusion with the visual features of this frame, form time sequence enhancement features, according to time sequence enhancement features to video frame grouping, select multiple representative and diversity key frames from pre-sampling frame;Multiple candidate local cropping regions are generated for selected key frame, the semantic similarity between each candidate region and semantic guide text is calculated respectively, and the candidate region with the highest similarity to semantic guide text is cropped out. Through the semantic perception of key frame selection in time dimension, and adaptive frame cropping in spatial dimension, the fine-grained refinement and alignment of video content are realized, thereby the performance of cross-modal text-video retrieval is significantly improved.
Owner:XIANGTAN UNIV

An operating room medical staff behavior management system

PendingCN122091126AVideo data clustering/classificationBiological modelsOperating theatresReoperative surgery
This invention provides an operating room medical staff behavior management system, belonging to the field of information technology management consulting technology. It includes: a process configuration module for determining the corresponding surgical list for each medical staff member based on their job permissions and configuring the behavior monitoring process; a process parsing module for setting monitoring targets for medical staff based on the parsing results of each monitoring step; a personnel analysis module for identifying non-standard behaviors of medical staff based on the behavior monitoring results of each monitoring step and analyzing all non-standard behaviors of each medical staff member involved in the current surgery to obtain comprehensive behavioral norms; and a behavior management module for determining the list of permitted dispatch personnel and behavioral norms for all hospital areas based on the comprehensive behavioral norms and the analysis results of medical staff in the corresponding surgical scenarios, storing this list on a cloud platform for managers of different hospital areas to view. This effectively achieves efficient, standardized, and managed operating rooms.
Owner:BEIJING DEREKANG INTELLIGENT EQUIP CO LTD

Method for presenting video recording apparatus, electronic device, medium and program product

The embodiment of the disclosure discloses a method for presenting a video recording, an apparatus for presenting a video recording, a device, and a storage medium. The method includes: receiving a video recording viewing operation; in response to the video recording viewing operation, displaying a personal recording page of a current user, and presenting video information of a plurality of first video recordings posted by the current user in the personal recording page according to a recording shooting theme.
Owner:BEIJING ZITIAO NETWORK TECH CO LTD

Target retrieval method and device, and storage medium

PendingEP4475007A4Video data indexingVideo data clustering/classificationComputer hardware
Provided are a target retrieval method and device, and a storage medium. The target retrieval method includes acquiring structured feature information that is input when information retrieval is performed on a preset information retrieval database; acquiring the totality of semi-structured feature information within a first preset period and a first preset range from the information retrieval database according to time information and range information contained in the input structured feature information and using the semi-structured feature information as to-be-retrieved thermal data; acquiring real-time video streams of the complete set of cameras within a second preset period and a second preset range; acquiring semi-structured feature information of a potential target in the real-time video streams; and comparing the semi-structured feature information of the potential target with the semi-structured feature information in the thermal data and determining whether the potential target is a retrieval target according to the comparison result.
Owner:ZHEJIANG UNIVIEW TECH CO LTD

Method and system for searching for event in captured image

PCT designated stageWO2026127182A1Video data indexingVideo data clustering/classification
A method and a system for searching for an event in a captured image are provided. According to some embodiments, a method for searching for an event in a captured image, which is performed by a computing system, may comprise the steps of: acquiring a real-time captured image; acquiring an image sequence for a predefined key event from the real-time captured image; generating an event embedding vector representing data for the image sequence by using a multimodal model, and storing the event embedding vector in an event database; receiving an event search query in the form of a natural language from a user terminal; performing preprocessing on the event search query, and acquiring a search target in a standardized format; generating a search embedding vector for the search target by using the multimodal model; performing a search for a first embedding vector included in the event database by using the search embedding vector; and as a result of performing the search, transmitting, to the user terminal, data on an event corresponding to a second embedding vector having a similarity to the search embedding vector equal to or greater than a reference value.
Owner:MOTOV CO LTD

Methods and systems for segmenting video content based on speech data and for retreiving video segments to generate videos

A method includes receiving a series of video segments and providing the series of video segments as input to a first machine learning model to produce text data. The text data is provided as input to a second machine learning model to produce categorized text data that includes a classification indication. The classification indication is added to metadata of the video segment, and the categorized text data is provided as input to a third machine learning model to produce a semantic vector. The method also includes causing the video segment and the metadata that includes the classification indication to be stored at a location of a database based on the semantic vector, the database being configured to be searched based on a search query associated with the semantic vector.
Owner:CIPIO INC

A method and system for automatically generating sales text videos

ActiveCN121071181BHigh originalityIncrease randomnessMetadata audio data retrievalTelevision system detailsEngineeringAudio frequency
This invention relates to the field of video generation technology, specifically disclosing a method and system for self-generating sales text videos. The method includes receiving sales text uploaded by a user; performing semantic analysis on the sales text to segment it into sentences and words in each sentence; creating a sales video based on the words; recognizing the sales video using AI to obtain recognized text; comparing the recognized text with the sales text to calculate text similarity; and reading audio from a preset audio library based on the text similarity and inserting it into the sales video. This invention performs semantic analysis on the sales text, segments it into words, creates a sales video based on the words, recognizes the sales video, compares the recognition results with the sales text, and adjusts the audio insertion part according to the comparison results, thereby optimizing the originality and randomness of the sales video.
Owner:TAIDOU TECH GRP CO LTD

Computer-implemented system and method for providing a stream or playback of a live performance recording

ActiveUS12647626B2Video data clustering/classificationRecord information storageEngineeringAudio frequency
The present disclosure relates to a computer-implemented system and method for providing streams and playbacks of video and audio recordings in respect of artist performances before live audiences. In particular, the disclosure relates to a system and method of providing a stream or playback of video and audio recordings captured in respect of an artist's live performance to users located around the world in exchange for a fee, with particular limitations placed upon access to the recording until the fee is paid.
Owner:SLOSS GLENN

Video clip retrieval model generation method and device, equipment and storage medium

ActiveCN119202312BVideo data clustering/classificationCharacter and pattern recognition
Embodiments of the present application provide a video segment retrieval model generation method, device and equipment, and a storage medium, relating to the field of artificial intelligence. The method comprises: obtaining a plurality of video blocks and their corresponding video text information and prior boxes by performing segmentation processing on a video segment; performing feature coding processing on the plurality of video blocks and their corresponding video text information to generate video coding information and text coding information corresponding to the plurality of video blocks; constructing a first graph network structure; updating nodes in the first graph network structure according to the video coding information and the text coding information to obtain a second graph network structure; determining a loss function of a video segment retrieval model according to labels corresponding to the prior boxes; and training the video segment retrieval model based on the second graph network structure and the loss function to obtain a target video segment retrieval model. The present application aims to solve the problem that complete videos cannot be accurately retrieved through video segments in related technologies.
Owner:PING AN TECH (SHENZHEN) CO LTD

Video search method, apparatus, medium, and computing device

ActiveCN117112839BVideo data clustering/classificationCharacter and pattern recognition
This disclosure provides a video search method, apparatus, medium, and electronic device, relating to the field of image technology. The data processing method includes: acquiring target parameters for each frame of a first image in a first video, the target parameters indicating at least one of the image category, application scenario, and purpose of the first image; determining a second image among the various first images based on the target parameters, wherein the image category of the second image is a preset category, the application scenario of the second image is a preset application scenario, and / or the purpose of the second image is a preset purpose; determining a first similarity between the third image and each frame of the second image based on the image features of the second images and the image features of a third image to be searched; and determining a video similar to the third image among a plurality of second videos based on the first similarity, wherein the first video is any one of the second videos. This disclosure improves the efficiency of finding similar videos of an image.
Owner:HANGZHOU NETZHIYI INNOVATION TECH CO LTD

Systems and methods for automated transformation of sports data

PCT designated stageWO2026117502A1Video data clustering/classificationCharacter and pattern recognitionData packEngineering
A method including receiving video data including a plurality of video frames captured during a sporting occasion, wherein each frame of the plurality of video frames includes data corresponding to one or more agents. The method including receiving event data associated with the video data. The method including processing the plurality of video frames and the event data to generate imputed tracking data, wherein the imputed tracking data includes positional data and movement data for each of the one or more agents. The method including determining one or more metrics based on the imputed tracking data, wherein the one or more metrics includes at least one of a pass option, a pressure, a line detection, and a marking. The method including determining an output based on the one or more metrics for the one or more agents.
Owner:STATS LLC

A false video source tracking method and system for hash data lookup and a medium

PendingCN122113065AVideo data clustering/classificationDigital data protectionFeature extractionData set
The application discloses a false video source tracking method and system based on hash data search and a medium, and relates to the technical field of data processing. The method comprises the following steps: uniformly preprocessing a video dataset to be tracked, and constructing a video sample set with consistent structures; constructing a triple sample set comprising original video samples, corresponding fake video samples and different source comparison video samples based on the video sample set; inputting the triple sample set into a hash feature extraction network to generate video hash codes, and dynamically training a plurality of hash centers based on a hash triple loss function; generating a temporary hash center through a voting mechanism and iteratively optimizing the distribution distance between the hash centers during the training process; extracting a video hash code of a video to be detected, performing hash data search in the hash centers, and determining a target hash center; outputting a source tracking result of the video to be detected based on a real video associated with the target hash center, and realizing stable positioning of a unique source video.
Owner:ZHONGKE TIANWANG (GUANGDONG) TECH CO LTD +1

An audio and video parsing method based on noise label learning

ActiveCN121682445BVideo data clustering/classificationSpeech analysisNoise (video)Noise
The application belongs to the technical field of deep learning, and relates to an audio and video parsing method based on noise label learning, which comprises the following steps: preprocessing original audio and video to obtain a segment-level input sequence; constructing a mutual learning noise-resistant double-flow network; training the mutual learning noise-resistant double-flow network according to a training set; comparing the validation set indicators of two sub-networks in the trained mutual learning noise-resistant double-flow network, and taking the sub-network with the larger validation set indicator as an audio and video parsing model; and parsing through the audio and video parsing model according to a test set to obtain a video prediction result. The mutual learning noise-resistant double-flow network is composed of two sub-networks with the same structure but different initializations, a cross filtering mechanism is executed according to the clean masks generated by the two sub-networks during the training of the mutual learning noise-resistant double-flow network, and the dynamic confidence ratio is gradually reduced through a cosine strategy, so that the problems of high pseudo-label noise rate and easy overfitting noise in the existing audio and video parsing task are solved.
Owner:UNIV OF ELECTRONIC SCI & TECH OF CHINA CHENGDU COLLEGE

Generating video from a static image using interpolation

A media application selects pairs of candidate images from a set of images associated with a user account, where each pair includes a first static image and a second static image from the user account. The media application applies a filter to select a particular pair of images from the pairs of candidate images. The media application generates one or more intermediate images based on the particular pair of images using an image interpolator. The media application generates a video that includes three or more frames arranged in a sequence, where a first frame of the sequence is the first static image, a last frame of the sequence is the second static image, and each of the one or more intermediate images is a corresponding intermediate frame in the sequence between the first frame and the last frame.
Owner:GOOGLE LLC

Video title generation method and apparatus, electronic device, storage medium, and product thereof

ActiveCN116431859BVideo data clustering/classificationSpecial data processing applications
This disclosure provides a method, apparatus, electronic device, storage medium, and products for generating video titles, relating to the field of big data processing technology, and particularly to the field of video data processing technology. The specific implementation of the video title generation method is as follows: extracting main words and feature words from the video information of the video to be processed; generating initial tags based on the main words and feature words; the initial tags containing at least one main word and at least one feature word; filtering target tags from the initial tags based on the search question of the video to be processed; and generating the title of the video to be processed based on the target tags.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

Search method and search device for video material segments, electronic device

PendingCN122173678AVideo data clustering/classificationMetadata video data retrievalComputer graphics (images)Data profiling
This application relates to the field of multimedia data analysis technology, and discloses a method, retrieval device, and electronic device for retrieving video clips. The retrieval method includes: extracting a first semantic feature sequence with semantic invariance from the query video; obtaining a second semantic feature sequence of candidate video clips to be retrieved; matching the first and second semantic feature sequences to determine multiple candidate matching frame pairs between the query video and the candidate video clips; performing temporal alignment and deduplication processing on the multiple candidate matching frame pairs to determine the matching result between the query video and the candidate video clips. This application can improve the accuracy of retrieving video clips.
Owner:MIAOZHEN INFORMATION TECHNOLOGY (ZIYANG) CO LTD +1

A knowledge-enhanced iterative self-optimizing text-to-video method and system

ActiveCN122138025AVideo data clustering/classificationBiological modelsAlgorithmText entry
This invention provides an iterative self-optimizing text-based video generation method and system based on knowledge enhancement, relating to the field of video generation technology. The method includes: acquiring the user's original text input; matching physical rule sets from a static physical knowledge base and recalling historical constraint sets from a dynamic constraint memory base; performing physical rule constraints and memory fusion to generate knowledge-enhanced video generation prompts; inputting the generated prompts into a pre-trained diffusion-based T2V model to generate the target video; performing static appearance verification and dynamic physical verification on the target video, and calculating semantic consistency scores and physical common sense scores; based on the verification results and scores, determining whether the iteration termination condition is met; if so, outputting the final video and optimized prompts. This invention forms a complete closed loop of "knowledge retrieval - planning generation - verification feedback - memory accumulation," significantly improving the generalization ability in out-of-distribution scenarios and achieving continuous iterative improvement in the system's generation quality.
Owner:SHANDONG UNIV

Systems and methods for automated transformation of sports data

PendingUS20260147832A1Database management systemsVideo data clustering/classificationData packEngineering
A method including receiving video data including a plurality of video frames captured during a sporting occasion, wherein each frame of the plurality of video frames includes data corresponding to one or more agents. The method including receiving event data associated with the video data. The method including processing the plurality of video frames and the event data to generate imputed tracking data, wherein the imputed tracking data includes positional data and movement data for each of the one or more agents. The method including determining one or more metrics based on the imputed tracking data, wherein the one or more metrics includes at least one of a pass option, a pressure, a line detection, and a marking. The method including determining an output based on the one or more metrics for the one or more agents.
Owner:STATS LLC

Device and method for occurrence frequency statistics for each CCTV event

PCT designated stageWO2026116577A1Video data browsing/visualisationVideo data clustering/classificationCCTV - Closed circuit televisionAlgorithm
A CCTV event type-specific statistical device, according to an embodiment disclosed in the present document, may comprise: a display device; and a processor functionally connected to the display device, wherein the processor may acquire event analysis data detected from at least some CCTV channels among a plurality of CCTV channels, calculate, on the basis of the event analysis data, the frequency of occurrence for each CCTV event type according to an expression time, configure, as a three-dimensional graph, the calculated frequency of occurrence for each CCTV event type according to the expression time, and display the configured three-dimensional graph on the display device.
Owner:KOREA ELECTRONICS TECH INST

DISPLAY DEVICE AND METHOD FOR CONTROLLING THE SAME

ActiveDE112018007902B4Input/output for user-computer interactionTelevision system detailsInternal memoryDisplay design
Display device including: a tuner that is equipped to receive a radio signal; a communication module that is set up to perform communication with an external server and / or an external remote control; a display designed to show content contained in the received broadcast signal, the content being received from the external server or stored in internal memory; and a controller set up to control the tuner, the communication module and / or the display, the controller is further configured to capture a screen image displaying the content, characterized by the fact that the controller is further configured for: to extract a first keyword from the recorded screen image and information contained in the broadcast signal of an electronic program guide, to generate a confidence value corresponding to the first keyword, to cause the display to show a first menu containing the extracted first keyword, to transmit the first keyword, first feedback information corresponding to a selection of the first keyword, and / or the confidence value of the first keyword to the external server based on the fact that an input selecting the first keyword is received from the external remote control. to receive a second keyword, a corrected confidence value, and second feedback information from the external server, to cause the display to show a second menu containing the received second keyword, and to cause the display to show a first screen corresponding to the second keyword, based on the fact that an input to select the second keyword is received from the external remote control.
Owner:LG ELECTRONICS INC

A multi-modal based video structured label construction method and device and medium

PendingCN122112306Aimprove accuracyeasy to identifyVideo data clustering/classificationBiological modelsPattern recognitionFrame sequence
The application discloses a multi-modal-based video structured label construction method and device and medium, and relates to the technical field of computer vision and artificial intelligence. The method comprises the following steps: acquiring a visual frame sequence, an audio stream and text information in a video to be labeled, and extracting features of the visual frame sequence, the audio stream and the text information to obtain multi-modal video features; performing cross-modal time sequence alignment on the multi-modal video features to obtain standard multi-modal video features, and performing deep semantic fusion on the standard multi-modal video features to generate a unified joint semantic representation; constructing a structured label knowledge graph with hierarchical semantic relationships by using unsupervised clustering and relationship mining, and matching the joint semantic representation corresponding to the video to be labeled with the structured label knowledge graph to output a pathized label set reflecting semantic levels. The application significantly improves the label construction accuracy by cross-modal time sequence alignment, deep semantic fusion and construction of a label knowledge graph.
Owner:QINGDAO QIANRUI DIGITAL INFORMATION TECHNOLOGY CO LTD