Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

59 results about "Video content analysis" patented technology

Video content analysis (also video content analytics, VCA) is the capability of automatically analyzing video to detect and determine temporal and spatial events. This technical capability is used in a wide range of domains including entertainment, health-care, retail, automotive, transport, home automation, flame and smoke detection, safety and security. The algorithms can be implemented as software on general purpose machines, or as hardware in specialized video processing units.

Al-based video content analysis method and system

The invention relates to the technical field of content recognition, in particular to an Al-based video content analysis method and system, and the method comprises the following steps: carrying out the image segmentation learning based on a video sequence through employing a U-Net convolutional neural network, analyzing scenes and elements in a video frame, recognizing and isolating key visual elements in a video through network learning, and carrying out the recognition of the key visual elements in the video. Comprising objects and figures. According to the method, image segmentation learning is carried out by adopting the U-Net convolutional neural network, key visual elements in the video can be identified and isolated more accurately, a clearer basis is provided for follow-up scene change and key event tracking, video content analysis is carried out by applying the graph neural network, the identification capability of dynamic scenes and events is enhanced, and the identification efficiency is improved. Deeper structured understanding is provided for video content indexing, the video quality is analyzed through a structural similarity index evaluation method, and the reason of visual quality reduction can be accurately recognized and improved.
Owner:CHENDA (GUANGZHOU) NETWORK TECH CO LTD +1

Video generation method, system and device based on large language model and medium

The invention discloses a video generation method, system and device based on a large language model and a medium, and the method comprises the steps: obtaining product information inputted by a user, the product information comprising the name, description and selling point of a product; preprocessing the product information, and carrying out semantic information split processing on the preprocessed product information through a large language model to obtain split description information corresponding to the product information; the method comprises the following steps: acquiring an original video, and performing video clip segmentation on the original video to generate a plurality of video clips and picture description information corresponding to the video clips; semantic matching processing is carried out on the sub-mirror description information and the picture description information, and a video clip with the highest similarity with each piece of sub-mirror description information is obtained through matching; and splicing the video clips according to the sequence of the sub-mirror description information to generate a complete video. Compared with the prior art, the product video can be automatically and efficiently generated by integrating natural language processing, video content analysis and an intelligent matching algorithm.
Owner:GUANGZHOU TAIDONG TECH CO LTD

Video content analysis and post-production optimization method and system based on cloud computing

The invention discloses a video content analysis and post-production optimization method and system based on cloud computing, and relates to the technical field of video processing, and the method comprises the following steps: in a video content analysis stage, obtaining a video through a data interface, dividing the video into frames, analyzing the character action change, forming an action event stream, and storing the action event stream in a database; extracting video content dimension features at the same time; constructing a video content comprehensive feature scoring model, and obtaining a comprehensive feature score through weighted summation in combination with the dimension features of the video content; in the later production and adjustment stage, the performance of the video in each dimension is evaluated based on the comprehensive scoring model, and the video content is adjusted according to the evaluation result and the video clip; according to the method, the frame number meeting the set matching degree threshold condition in the video frame is calculated, the proportion of the frame number higher than the threshold value and the frame number lower than the threshold value is analyzed, when the proportion meets the preset condition, whether the video content needs to be further adjusted in the post-production stage or not is marked, and video content analysis and post-production optimization are achieved.
Owner:NANJING SHUHAI XINGCHEN VISION TECHNOLOGY CO LTD

Methods, systems, and devices for capturing video content associated with performing an athletic skill and determining biomechanic adjustments for performing the athletic skill

Aspects of the subject disclosure may include, for example, obtaining current video content of a player repeatedly performing a physical skill, analyzing the current video content based on previous video content, the previous video content comprises other video content of the player repeatedly performing the physical skill, determining biomechanic metrics of the player performing the physical skill based on the analysis, and determining each biomechanic metric of a portion of the biomechanic metrics does not satisfy a respective biomechanic metric success rate. Further embodiments include generating a first image of the player performing the physical skill from the current video content, generating a second image of the player performing the physical skill from the previous video content, and presenting the first image and the second image simultaneously and indicating the portion of the biomechanics that did not satisfy the respective biomechanic metric success rate. Other embodiments are disclosed.
Owner:ATHLETIQ LLC

Video scene content label determination method based on knowledge graph

The invention relates to the technical field of video content analysis, and discloses a video scene content label determination method based on a knowledge graph. The method comprises the following steps: inputting a multi-modal feature sequence formed by visual, audio and text features extracted from a video stream into a knowledge graph inference engine comprising an entity relationship network and a semantic association rule base; performing node mapping through an engine to generate an initial scene entity set; performing hierarchical reasoning on the set based on a semantic association rule base to obtain a scene semantic topological structure; screening core entities according to entity weight distribution in the topological structure, and generating a candidate tag set; performing time sequence consistency verification on the candidate tag set, and correcting tag time sequence offset in combination with timestamp information; performing cross-modal disambiguation on the corrected label set by using an entity relationship network to eliminate semantic conflicts; and according to a disambiguation result, constructing a scene label knowledge sub-graph containing entity attributes and relation path constraints.
Owner:GUANGZHOU JUNHE INFORMATION TECH CO LTD

Illegal capital clue identification and early warning method and device

The invention relates to the technical field of big data analysis, and particularly provides an illegal capital clue identification and early warning method and device, and the method comprises the following steps: S1, data collection and preprocessing; s2, text feature extraction and semantic modeling; s3, video content analysis and behavior recognition; s4, performing multi-modal fusion and risk scoring; and S5, early warning output and visual display are carried out. Compared with the prior art, the method has the advantages that potential illegal information can be automatically extracted from massive news reports and short video data, and efficient and accurate risk early warning is realized.
Owner:天元大数据信用管理有限公司

Automatic movie and television video script extraction method based on multi-modal large model

The invention relates to an automatic movie and television video script extraction method based on a multi-modal large model, and belongs to the field of artificial intelligence and video content analysis. Aiming at the problem that an existing automatic speech recognition tool cannot generate a structured script, the method comprises the following steps: firstly, carrying out multi-mode decomposition on an input video, extracting frames by adopting a scene self-adaptive strategy, and establishing sound and picture timestamp alignment mapping; then extracting features through a CLIP-ViT visual feature encoder and a Whisper audio feature encoder, and associating semantics by using a cross-modal attention mechanism; role identity recognition, scene type recognition and action character description are achieved, and finally a script file conforming to the standardization specification is generated and comprises a scene title, a time code mark and a special narrative mark. Compared with a mode of manually dictating marks and an automatic voice recognition tool, the method effectively improves the accuracy and efficiency of movie and television video script extraction.
Owner:BEIJING INST OF COMP TECH & APPL

Language-driven hour-level traffic video analysis method and system

The invention discloses a language-driven hour-level traffic video analysis method and system, relates to the technical field of video content analysis and retrieval, can generate a video-specific open world detector, can accurately identify and filter fine-grained and composite semantic targets, and improves the accuracy of video analysis. And an hour-level video analysis service driven by open vocabularies and natural languages is really realized. In order to achieve the purpose, the technical scheme of the invention comprises the following steps of: 1) receiving an hour-level traffic video, and positioning a video clip related to natural language query in the video; and 2) automatically constructing an open world target detector in the video clip positioned in the step 1). And step 3) finishing cross-time-period target trajectory generation on the video clip positioned in the step 1). The invention also provides a language-driven hour-level traffic video analysis system for executing the method. The system comprises a video clip positioning module, an open world target detector construction module and a target trajectory extraction module.
Owner:BEIJING INST OF TECH

Video content automatic auditing method and system based on AI drive

The invention discloses a video content automatic auditing method and system based on AI drive, and relates to the technical field of video content analysis, and the method comprises the steps: obtaining video information to be audited, and carrying out the multi-level segmentation of the video information; in the slices, node simulation is carried out according to the content independence of the video information; analyzing associated branches and historical features of each node to generate attribute features of the nodes; content auditing is carried out through an auditing strategy of multiple virtual agents, and illegal space analysis is carried out according to an unstable part in an auditing result; and generating a final video content auditing report according to the auditing result and the analysis result of the illegal space. According to the method, the structured recognition and intelligent auditing efficiency of the multi-modal video content is improved, accurate positioning, label attribution and automatic early warning of complex violation behaviors are realized, and the content security management capability is enhanced.
Owner:JIANGSU BROADCASTING CORPORATION

Video encoding method and apparatus, computer device, and storage medium

A method of video encoding is described. The method includes segmenting original video data to obtain an original video segment including multiple video images. Video content analysis is performed on the original video segment to obtain a video image processing parameter corresponding to the original video segment. Image processing is performed on a video image in the multiple video images in the original video segment based on the video image processing parameter to obtain a processed video segment. An encoding parameter of the processed video segment can be determined based on image feature data of the processed video segment. The processed video segment can be encoded based on the encoding parameter to obtain an encoded video segment.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Mine monitoring video key frame extraction method

The invention relates to the technical field of video content analysis, in particular to a mine surveillance video key frame extraction method, which comprises the following steps: extracting multi-dimensional features of each frame of a candidate frame set, obtaining a fusion value, and obtaining a candidate key frame set; based on an Euclidean distance method, calculating an Euclidean distance between continuous frames of the candidate key frame set, and determining a total clustering number according to a relationship between the Euclidean distance and a preset clustering threshold value; and clustering the frames in the candidate key frame set by using a preset mixed GWO-FCM clustering algorithm and the total clustering number to obtain a key frame set. According to the method, a hybrid clustering algorithm GWO-FCM combining grey wolf optimization and fuzzy C-means is adopted, global search and local optimization capabilities are considered, and the accuracy and representativeness of key frame clustering are ensured. The high-quality extraction of the key frames is realized, the number of redundant frames is obviously reduced, and the compactness and readability of the video abstract are improved.
Owner:XJ GRP CORP

Video content intelligent analysis and statistical method under support of large model

The invention belongs to the technical field of video content analysis, and more specifically relates to a video content intelligent analysis and statistical method under the support of a large model. According to the method, rich visual features are extracted through a Swin Transform model, and feature weighting is dynamically adjusted according to an application scene; organizing the extracted features into a time sequence matrix so as to facilitate subsequent time sequence information fusion; through a time sequence modeling network, a time dependency relationship between video frames is captured, a new feature matrix containing dynamic information is generated, then the accuracy and robustness of behavior recognition are improved, finally, a space-time correlation modeling technology is utilized, accurate tracking of a target in a video is achieved, and the behavior recognition accuracy and robustness are improved in combination with a behavior recognition result. Comprehensively analyzing and counting the video content; the accuracy of video analysis is effectively improved, and the problems of dynamic change and insufficient continuous capture of behavior recognition in the prior art are solved.
Owner:CHENGDU SHUXI TECH CO LTD

Systems and methods for interactive content viewing and discovery

Disclosed are computerized systems and methods for a decision intelligence (DI)-based framework that automatically and / or dynamically provides an interactive content viewing and content discovery experience to users. The framework includes functionality for real-time video content analysis and interactive entity identification that, inter alia, provides novel capabilities to viewing users related to the extraction, analysis and subsequent interaction with content depicted within video frames during playback of such content. The framework implements artificial intelligence / machine learning (AI / ML) approaches for video processing, user interaction and information delivery through a series of interconnected processes and subsystems. In some implementations, rendered content can be parsed and mined for real-world and / or digital content depicted therein that relate to real-world entities and / or digital resources, whereby interaction with such entities is provided in the form of a provided interface, electronic message and / or recommendations for further information discovery, or some combination thereof.
Owner:RIGH INC

Intelligent equipment linkage control method and system based on video content analysis

The invention discloses an intelligent equipment linkage control method and system based on video content analysis, and relates to the field of video content analysis, and the intelligent equipment linkage control method based on video content analysis comprises the following steps: S1, obtaining video parameters, and carrying out the preprocessing; s2, pixel motion vectors are extracted, and a time sequence motion field is formed; s3, extracting a pixel dominant motion information set, calculating an average speed value, and generating a time sequence speed parameter set; s4, performing zero crossing point identification on the time sequence speed parameter set, and generating a time sequence displacement control signal; and S5, sending the time sequence displacement control signal to the external intelligent equipment, and driving the external intelligent equipment to generate physical motion with the video content by adopting the displacement signal. According to the method, by introducing preprocessing means such as multi-scale image analysis, resolution compression and graying conversion, key information is efficiently and accurately extracted from the video, and the precision and efficiency of subsequent analysis are improved.
Owner:SHENZHEN WEIAI TECHNOLOGY CO LTD

A video intelligent slicing method based on multi-modal information fusion and semantic constraint

The application discloses a kind of video intelligent slice method based on multi-modal information fusion and semantic constraint, it is related to video content analysis and intelligent editing technical field, it is characterized in that it proposes semantic priority audiovisual collaborative slice scheme, obtains shot interval by shot boundary detection, constructs voice interval and effective word information by voice recognition, adopts voice interval as semantic mask to remove visual cut point in the process of voice expression, and based on fragment length and effective word, content density determination is carried out to long fragment, and further adopts visual re-cutting, audio energy driving and uniform cutting Secondary slicing strategy of division, the method provided by the application solves the problem of semantic fragmentation and content perception loss in the prior art, can obtain video fragment with complete semantics, high content density and appropriate length, and is suitable for various application scenarios such as automatic editing, short video generation and video content retrieval.
Owner:YUELAI INTELLIGENT MEDIA (TIANJIN) TECHNOLOGY CO LTD

Video texture mapping and real-time rendering method and system based on digital twin scene

The application provides a video texture mapping and real-time rendering method and system based on a digital twin scene, relates to the technical field of digital twin rendering, and comprises the following steps: acquiring a real scene monitoring video stream and performing real-time image content analysis; dividing the texture details of different regions of the video stream according to the features and numerical values of the image analysis and generating region detail identifiers associated with the video content; identifying the dynamic features of the video picture based on the identifiers and generating scene state codes representing the comprehensive picture conditions; dynamically selecting a target engine and corresponding rendering quality parameters from preset heterogeneous rendering engines in combination with the current load state of a graphics processor; and mapping the video stream image data to the surface of a corresponding model of the digital twin scene as dynamic texture to realize efficient rendering of the picture, so that dynamic texture mapping and adaptive real-time rendering based on video content analysis and hardware load in the digital twin scene can be realized.
Owner:BEIJING ZHIHUI YUNZHOU TECH CO LTD

Audio and video content analysis method and device

The embodiment of the present invention proposes a method and device for parsing audio and video content, which belongs to the field of deep learning. The audio and video to be analyzed are split, feature extracted and feature fused to obtain comprehensive visual features and auditory features. The auditory features and comprehensive visual features are input into a parsing algorithm obtained by optimization training using weakly supervised learning. Through the parsing algorithm, the auditory features and comprehensive visual features are modeled and perceptually predicted with respect to modality to obtain the action events contained in the audio and video to be analyzed and the category and modality to which each action event belongs. The parsing algorithm of the present application proposes a modality perception module and a timing perception module, which can coordinate modality and timing for evidence mining, thereby greatly reducing the model's sensitivity to pseudo-label noise generated during modal classification, improving the robustness of modality dependency judgment and the accuracy of timing annotation, and overcoming the uncertainty problem caused by the lack of timing annotation under weak supervision settings.
Owner:UNIV OF ELECTRONICS SCI & TECH OF CHINA +1

Systems and methods for interactive content viewing and discovery

Disclosed are computerized systems and methods for a decision intelligence (Dl)-based framework that automatically and / or dynamically provides an interactive content viewing and content discovery experience to users. The framework includes functionality for real-time video content analysis and interactive entity identification that, inter alia, provides novel capabilities to viewing users related to the extraction, analysis and subsequent interaction with content depicted within video frames during playback of such content. The framework implements artificial intelligence / machine learning (AI / ML) approaches for video processing, user interaction and information delivery through a series of interconnected processes and subsystems. In some implementations, rendered content can be parsed and mined for real -world and / or digital content depicted therein that relate to real-world entities and / or digital resources, whereby interaction with such entities is provided in the form of a provided interface, electronic message and / or recommendations for further information discovery, or some combination thereof.
Owner:RIGH INC

Video processing collaboration method and device, equipment and storage medium

The application provides a video processing cooperation method and device, equipment and a storage medium. The method comprises the following steps: sending a video rendering capability request and a video content analysis capability request to a terminal device; receiving a video rendering capability response and a video analysis capability response of the terminal device, wherein the video rendering capability response comprises a video rendering capability of the terminal device, and the video analysis capability response comprises a video analysis capability of the terminal device; determining an optimal video rendering cooperation configuration according to the video rendering capability of the terminal device, and determining an optimal video analysis cooperation configuration according to the video analysis capability of the terminal device. Therefore, the idle computing resources of the terminal device can be fully utilized under limited cloud server computing resources, so that a user can have a better cloud game quality experience.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

A video generation method, system, device and medium based on a large language model

The application discloses a video generation method, system, device and medium based on a large language model. The method obtains product information input by a user, the product information including the name, description and selling point of the product. The product information is preprocessed, and the preprocessed product information is subjected to semantic information split-screen processing through a large language model to obtain split-screen description information corresponding to the product information. An original video is obtained, and the original video is subjected to video segment splitting to generate a plurality of video segments and picture description information corresponding to the video segments. The split-screen description information and the picture description information are subjected to semantic matching processing, and the video segments with the highest similarity to each split-screen description information are matched. The video segments are spliced according to the sequence of the split-screen description information to generate a complete video. Compared with the related art, the application can automatically and efficiently generate a product video by integrating natural language processing, video content analysis and intelligent matching algorithms.
Owner:GUANGZHOU TAIDONG TECH CO LTD

A network public opinion video content analysis method based on multi-modal fusion

The application provides a network public opinion video content analysis method based on multi-modal fusion, and relates to the technical field of network information processing.The method comprises the following steps: by splitting the public opinion video into an audio stream and an image frame sequence, structurally analyzing picture text and speech content respectively, and completing semantic embedding and cross-modal alignment under a unified time axis, introducing a bidirectional cross-modal attention and a gating fusion mechanism to realize deep fusion of multi-modal information, and further combining multi-role consistency, semantic superposition effect and emotion and fact collaborative features for joint modeling at an event level scale, so as to realize overall understanding of network public opinion video content and accurate determination of public opinion risk level.The application can fuse multi-modal information as a whole and realize joint modeling to analyze network public opinion video content.
Owner:WUHAN FIBERHOME PUTIAN INFORMATION TECH CO LTD

Video content analysis method, server, product, equipment and storage medium

The invention discloses a video content analysis method, a server, a product, equipment and a storage medium, and relates to the field of video analysis, and the method comprises the steps: calling a preset large language model which cannot be subjected to subjective reasoning, carrying out the information extraction of the video display content of a video file at each moment, and obtaining an information extraction result; and performing risk identification on the extracted structured event data based on a risk identification rule, extracting verification evidence of the risk assessment result of each event entity from the video display content, and outputting a content analysis result based on each risk assessment result and the verification evidence. According to the method and the device, on the basis of the preset large language model which cannot be subjectively reasoned, the disassembly of the complex description logic is realized, the accuracy of each event entity obtained by disassembly is improved, the risk assessment result of each event entity is verified by using the verification evidence, and the content analysis precision of the finally output video file is improved.
Owner:BEIJING QIYI CENTURY SCI & TECH CO LTD

Multi-mode-based video advertisement generation method, system, equipment and medium

The invention discloses a multi-mode-based video advertisement generation method, system and device and a medium, and the method comprises the steps: obtaining a video material of a product, and carrying out the preprocessing of the video material; performing video content analysis and video clip scoring on the preprocessed video material data to obtain a corresponding video content analysis result and a video clip score; generating a mixed cutting video corresponding to the product according to the video clip score; according to the analysis result of the video content, the name of the product and the selling point, generating an advertisement copywriting corresponding to the product; forming audio and subtitle data according to the content of the advertisement copywriting; and integrating the mixed video, the audio and the subtitles to generate a video advertisement of the product. According to the method, technologies such as video content analysis, propaganda copywriting automatic generation and an intelligent matching algorithm are fused, and an efficient automatic process from user input to finished video is constructed, so that the problems that traditional video advertisement generation needs a large amount of manual intervention and the generation efficiency is relatively low are solved.
Owner:GUANGZHOU TAIDONG TECH CO LTD

Systems and methods for interactive content viewing and discovery

Disclosed are computerized systems and methods for a decision intelligence (DI)-based framework that automatically and / or dynamically provides an interactive content viewing and content discovery experience to users. The framework includes functionality for real-time video content analysis and interactive entity identification that, inter alia, provides novel capabilities to viewing users related to the extraction, analysis and subsequent interaction with content depicted within video frames during playback of such content. The framework implements artificial intelligence / machine learning (AI / ML) approaches for video processing, user interaction and information delivery through a series of interconnected processes and subsystems. In some implementations, rendered content can be parsed and mined for real-world and / or digital content depicted therein that relate to real-world entities and / or digital resources, whereby interaction with such entities is provided in the form of a provided interface, electronic message and / or recommendations for further information discovery, or some combination thereof.
Owner:RIGH INC

Video texture mapping and real-time rendering method and system based on digital twin scene

The invention provides a video texture mapping and real-time rendering method and system based on a digital twin scene, and relates to the technical field of digital twin rendering. The method comprises the following steps: dynamically dividing texture detail levels of different regions of a video stream according to features and numerical values of image analysis, generating region detail identifiers associated with video contents, identifying dynamic features of a video picture based on the identifiers, and generating a scene state code representing a comprehensive picture condition; a target engine and corresponding rendering quality parameters are dynamically selected from preset heterogeneous rendering engines in combination with the current load state of a graphics processor, and video stream image data are mapped to the surface of a model corresponding to a digital twin scene as dynamic textures, so that efficient rendering of pictures is realized. Dynamic texture mapping and self-adaptive real-time rendering based on video content analysis and hardware load in a digital twinning scene can be realized.
Owner:BEIJING ZHIHUI YUNZHOU TECH CO LTD

Video abstract extraction method based on industry large model

The invention provides a video abstraction extraction method based on an industry large model, and belongs to the field of artificial intelligence and machine learning. The large model is finely adjusted and optimized on specific industry data, so that the large model can deeply understand and adapt to specific video contents of the industry, the quality and the industry adaptability of the video abstraction are remarkably improved, and the video abstraction extraction efficiency is improved. And a more accurate, efficient and reliable solution is provided for video content analysis, retrieval and management of each industry.
Owner:SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD

Multimodal model driven video recognition method and apparatus

This application provides a multimodal model-driven video recognition method and apparatus, relating to the field of video content analysis technology. The method includes: acquiring a target short video and performing frame segmentation, audio separation, and text content extraction to obtain video frame sequences, audio data, and text data; extracting visual features, speech features, and text features based on the video frame sequences, audio data, and text data, respectively; inputting these multimodal features into a multimodal fusion network based on an attention mechanism, and performing weighted fusion by calculating cross-modal interaction weights between the features to generate a fused feature representation; inputting the fused feature representation into a business order recognition classifier to obtain the business order recognition result of the target short video. The technical solution of this application can improve the accuracy of short video business order recognition and is applicable to scenarios such as short video platform content review, advertising monitoring, and business data analysis.
Owner:BEIJING STAR RIVER EXCELLENCE TECH CO LTD

Systems and methods for interactive content viewing and discovery

Disclosed are computerized systems and methods for a decision intelligence (DI)-based framework that automatically and / or dynamically provides an interactive content viewing and content discovery experience to users. The framework includes functionality for real-time video content analysis and interactive entity identification that, inter alia, provides novel capabilities to viewing users related to the extraction, analysis and subsequent interaction with content depicted within video frames during playback of such content. The framework implements artificial intelligence / machine learning (AI / ML) approaches for video processing, user interaction and information delivery through a series of interconnected processes and subsystems. In some implementations, rendered content can be parsed and mined for real-world and / or digital content depicted therein that relate to real-world entities and / or digital resources, whereby interaction with such entities is provided in the form of a provided interface, electronic message and / or recommendations for further information discovery, or some combination thereof.
Owner:RIGH INC

Systems and techniques for retraining models for video quality assessment and for transcoding using the retrained models

A trained model is retrained for video quality assessment and used to identify sets of adaptive compression parameters for transcoding user generated video content. Using transfer learning, the model, which is initially trained for image object detection, is retrained for technical content assessment and then again retrained for video quality assessment. The model is then deployed into a transcoding pipeline and used for transcoding an input video stream of user generated content. The transcoding pipeline may be structured in one of several ways. In one example, a secondary pathway for video content analysis using the model is introduced into the pipeline, which does not interfere with the ultimate output of the transcoding should there be a network or other issue. In another example, the model is introduced as a library within the existing pipeline, which would maintain a single pathway, but ultimately is not expected to introduce significant latency.
Owner:GOOGLE LLC

Chinese zither performance video content intelligent analysis and retrieval system

The invention discloses a zither performance video content intelligent analysis and retrieval system, and belongs to the technical field of video content analysis and retrieval, the system comprises a multi-modal feature extraction module, a manifold topology transformation module, a semantic scene analysis module and a self-adaptive retrieval module, deep fusion of visual and audio features is realized through manifold topology transformation, and the video content is analyzed and retrieved. A topological structure is established based on Riemannian metrics, and precise recognition and time sequence segmentation of scenes such as technique display, track playing and teaching explanation are achieved. The adaptive retrieval module constructs a closed-loop feedback mechanism, dynamically adjusts feature extraction weights and measurement parameters according to retrieval confidence, and realizes continuous optimization of system performance, the problems of insufficient multi-modal information fusion, inaccurate scene recognition, low retrieval efficiency and the like in the prior art are solved, the scene recognition accuracy is improved by more than 12%, and the system performance is improved by more than 12%. And the retrieval precision is improved by more than 30%.
Owner:JIANGXI UNIVERSITY OF FINANCE AND ECONOMICS