Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

30813results about "Selective content distribution" patented technology

Method and apparatus for reducing the number of control messages transmitted by a set top terminal in an SDV system

A method is provided by which a subscriber accesses an SDV channel using a set top terminal. The method begins when the set top terminal receives a user request to tune to a first SDV channel. An active services list is also received over an access network. The active services list includes an entry for each currently available SDV program and a time-to-live (TTL) associated therewith. Tuning information is identified for the first SDV channel from its entry in the active services list. The set top terminal tunes to the first SDV channel using the identified tuning information. The channel change information associated with the user request is locally stored in set top terminal for transmission over the access network at a later time.
Owner:GENERAL INSTR CORP

Cloud storage backend

A video streaming system including a video cache system having a cache database and a video storage device, and including a video server configured to communicate with a gateway system configured to store multiple video streams, communicate with a cloud backup system configured to store video in a file system, and communicate with a cloud server to receive a stream request, the stream request indicating requested video that comprises multiple video segments. The video server is further configured to receive the stream request and determine a playlist that indicates where on the video cache system, the cloud backup system, and the gateway system the video segments that are needed to fulfill the stream request are stored, and to communicate needed video segments to the video cache system in preparation for providing streaming video to fulfill the stream request.
Owner:SAMSARA INC

Marketing video auditing method based on AI

The invention provides an AI-based marketing video auditing method, and relates to the technical field of AI marketing video auditing, and the method comprises the steps: obtaining a multi-modal data original structure set, and extracting image semantic features, voice expression features, text semantic features and scene label information, and obtaining an image semantic feature set, a visual rhythm feature set, a voice expression feature set, a voice and picture synchronous association vector structure, a text semantic feature set and a subtitle semantic and image main body linkage relation graph. By constructing an image semantic feature set, a voice expression feature set, a text semantic feature set and a visual rhythm feature set and fusing the image semantic feature set, the voice expression feature set, the text semantic feature set and the visual rhythm feature set into a multi-modal content fusion feature tensor, unified modeling of an AI marketing video at visual, auditory and semantic levels can be realized; and subsequent microscopic consistency detection, compliance knowledge graph and emotion semantic conflict identification are effectively performed, so that full-link risk perception and accurate auditing of video contents are realized.
Owner:SHANGHAI WANGMAI INFORMATION TECH GRP CO LTD

AI intelligent short video generation method and system based on multi-agent collaboration

The invention relates to an AI intelligent short video generation method and system based on multi-agent collaboration. The method comprises the following steps: S1, input analysis and sub-shot script generation; s2, obtaining materials; s3, sub-shot video generation; step S4: editing and synthesizing; wherein the plurality of sub-shot video clips are edited and synthesized into a complete video according to a script sequence, and style fusion processing is performed on the whole video by using a picture style unification algorithm; step S5, background music matching; and S6, outputting the slices. Through the structured sub-shot script generation, intelligent material completion, multi-shot style fusion and music synchronization technology, a professional short video of multi-scene coherent narrative can be generated, the structured sub-shot script is automatically generated, and intelligent material completion, picture style unification and music rhythm and emotion expression synchronization are realized.
Owner:苏州日报社

Video editing method and related equipment

The invention discloses a video editing method and related equipment, and the method comprises the steps: receiving an editing intention of a user, carrying out the semantic analysis of an intention analysis model, and generating an editing copywriting, a style and a structured instruction; the edited copywriting is split according to the edited copywriting; performing cross-modal analysis on video materials by using a video understanding model trained based on an open-source multi-modal model, and accurately positioning sub-lens segments; secondly, editing the sub-lens segments through an editing model to generate a preliminary sheet, performing quantitative scoring through a video scoring model, and obtaining a scoring result according to a multi-dimensional index; and finally, the result is fed back to the editing model, and iterative adjustment is carried out until a final slice is produced. According to the method, the editing efficiency is improved, the manual operation time is shortened, the material adaptability is enhanced, the personalized requirement is met, the editing effect consistency is improved, the video content understanding is deepened, the editing effect is optimized, the technical threshold is reduced, the existing technical problems are solved comprehensively, and an efficient, intelligent and personalized editing scheme is provided.
Owner:GUANGZHOU QUYAN NETWORK TECH CO LTD

Avatar JSON interchange file format

Some embodiments of a method may include: obtaining Avatar JSON Interchange File (AJIF) data; decoding the AJIF data to generate avatar data; obtaining animation parameters from the avatar data; obtaining user movement data corresponding to a movement of a user; generating updated avatar data based on the user movement data; and sending the updated avatar data to a client device. For some embodiments, a method may include: generating Avatar JSON Interchange File (AJIF) data corresponding to an avatar; sending, to an application server, the AJIF data; obtaining user movement data corresponding to a movement of a user; sending, to the application server, the user movement data corresponding to the movement of the user; receiving, from the application server, scene update data, wherein the scene update data includes updated avatar data corresponding to the movement of the user; and rendering a scene update based on the scene update data.
Owner:INTERDIGITAL CE PATENT HOLDINGS SAS

3D gaussians splatting in scene description

Some embodiments of a method may include: obtaining information for a three-dimensional (3D) Gaussian model corresponding to a 3D scene, wherein the information comprises a set of attributes of the 3D Gaussian model; parsing the information for a first attribute of the set of attributes, wherein the first attribute corresponds to a position of the 3D Gaussian model; parsing the information for a second attribute of the set of attributes, wherein the second attribute corresponds to a color of the 3D Gaussian model; parsing the information for a third attribute of the set of attributes, wherein the third attribute corresponds to a covariance of the 3D Gaussian model; and rendering the 3D scene using the parsed attributes of the 3D Gaussian model.
Owner:INTERDIGITAL CE PATENT HOLDINGS SAS

3D gausians splatting in scene description

Some embodiments of a method may include: obtaining information for a three-dimensional (3D) Gaussian model corresponding to a 3D scene, wherein the information comprises a set of attributes of the 3D Gaussian model; parsing the information for a first attribute of the set of attributes, wherein the first attribute corresponds to a position of the 3D Gaussian model; parsing the information for a second attribute of the set of attributes, wherein the second attribute corresponds to a covariance of the 3D Gaussian model; parsing the information for third, fourth, and fifth attributes, wherein the third, fourth, and fifth attributes correspond to first, second, and third sets of spherical harmonics coefficients associated with the 3D Gaussian model; parsing the information for a sixth attribute of the set of attributes, wherein the sixth attribute is an alpha coefficient for the 3D Gaussian model; and rendering the 3D scene using the parsed attributes.
Owner:INTERDIGITAL CE PATENT HOLDINGS SAS

Methods For Generating Advertisement Videos Consistent With The Context And Storyline Of A Primary Video Stream

Embodiments include methods for generating advertisement videos for insertion into a video stream to promote a product, service, or brand in a manner that is consistent with the context and storyline of the video stream before and at the time of ad insertion. Methods may include capturing an image from the video stream and generating caption text using an image-to-text description model. A product, service, or brand that is consistent with the context and storyline of the captured image is selected and ad video sequence description text is generated that includes descriptions and a storyline blending descriptions of the selected product, service, or brand with the context and storyline of the primary video stream. The ad video sequence description text is used to prompt a text-to-video generation model that generates a new advertisement video clip, which is inserted into the primary video stream before distribution to video content rendering devices.
Owner:CHARTER COMM OPERATING LLC

Cross-modal knowledge optimization system for improving localization adaptability of large language model

PendingCN120542521ASemantic analysis2D-image generationDeep knowledgeEngineering
The invention relates to the field of natural language processing. The invention discloses a cross-modal knowledge optimization system for improving localization adaptability of a large language model, and the system comprises a modal data collection module which collects cross-modal localization data of characters, audios and videos; the data preprocessing module is connected with the data acquisition module and preprocesses acquired data; the cross-modal knowledge fusion module is connected with the preprocessing module, fuses data and large language model general knowledge, and performs mining association by means of a cross-modal learning algorithm to form a localized cross-modal knowledge graph; the model fine-tuning module is connected with the fusion module and uses the map to finely tune the large language model; the evaluation feedback module is connected with the fine tuning module and evaluates the localization performance of the model. Through multi-modal data integration, dynamic preprocessing, deep knowledge fusion, efficient fine adjustment and intelligent feedback, various problems in the large language model localization process are systematically solved, and the expression of the model in dialect understanding, cultural questions and answers and cross-modal tasks is remarkably improved.
Owner:HANGZHOU LANGSHI VIDEO TECH CO LTD

Video stream dynamic fragment encryption and block chain evidence storage method

The invention discloses a video stream dynamic fragmentation encryption and block chain evidence storage method, and relates to the technical field of video content security, and the method comprises the steps: calculating a color histogram difference value and an optical flow vector change rate between adjacent frames of an input video, marking the difference value as a scene switching point when the difference value exceeds a preset threshold value, and storing the scene switching point; the method comprises the following steps: preliminarily dividing a video into a plurality of scene segments according to scene switching points, performing content complexity evaluation on the scene segments, calculating gray level co-occurrence matrix characteristics of each frame of image through texture density analysis, calculating edge complexity to extract the number and distribution of Canny edges, and performing motion vector statistics to analyze the size and direction of inter-frame object displacement. The change rate between adjacent pixels in the color space is measured according to the color change gradient; the video stream dynamic fragment encryption and block chain evidence storage method is suitable for video contents of different types and complexities, and has relatively high detection accuracy and robustness.
Owner:HANGZHOU MEICHANG IOT TECH CO LTD

Multimode Heterogeneous Cellular Communication with Integrated Multi-Channel Link Diversity, Physical Layer Optimization, and Radio Access Handover

The invention overcomes the constraints imposed by conventional wireless communication standards by introducing a physical-layer optimized architecture that decouples transmission control from any single protocol. Rather than being confined by the predefined behaviors of LTE, 5G NR, Wi-Fi, or NB-IoT, the system implements a unified control system that dynamically manages radio parameters—such as modulation, coding, and power—based on real-time link quality and network context. This cross-standard, multimode capability effectively supersedes traditional standard-driven implementations, enabling adaptive, low-latency, and spectrum-efficient communication in complex heterogeneous environments.The system dynamically controls radio access network (RAN) and physical layer parameters, including carrier aggregation, dynamic spectrum allocation, modulation and coding scheme (MCS) adaptation, beamforming configuration, channel coding, transmit power control, and frequency selection. The system architecture includes multi-mode base stations, relay nodes, and edge access points that support inter-RAT handover, fast radio link recovery, and seamless mobility across diverse wireless technologies.
Owner:CHINA ENTROPY CO LTD (AIOT ENTROPY CO LTD)

Method for providing video and electronic device supporting the same

An electronic device is provided. The electronic device includes a memory, and at least one processor electrically connected to the memory, wherein the at least one processor is configured to obtain a video including an image and an audio, obtain information on at least one object included in the image from the image, obtain a visual feature of the at least one object, based on the image and the information on the at least one object, obtain a spectrogram of the audio, obtain an audio feature of the at least one object from the spectrogram of the audio, combine the visual feature and the audio feature, obtain, based on the combined visual feature and audio feature, information on a position of the at least one object the information indicating the position of the at least one object in the image, obtain an audio part corresponding to the at least one object in the audio, based on the combined visual feature and audio feature, and store, in the memory, the information on the position of the at least one object and the audio part corresponding to the at least one object.
Owner:SAMSUNG ELECTRONICS CO LTD

Multi-user data storage docking and secure transmission method based on AI

The invention provides an AI-based multi-user data storage docking and secure transmission method, which comprises the following steps of: acquiring a propagation path of a hot topic according to a multi-modal content association graph, performing fragmentation storage on the propagation path by adopting a clustering algorithm, and generating a fragmentation storage index table; when the relevance of different modal contents is higher than a preset threshold value, determining that the sensitivity grading result is consistent with the multi-modal relevance, performing dynamic rule adjustment by adopting a deep learning model, generating a dynamic auditing rule set, and obtaining boundary judgment parameters of the rule set; and distributing the dynamic auditing rule set and the decision transparency index to each service line by adopting a cross-platform data synchronization protocol according to the auditing decision log, generating a cross-platform consistency auditing standard, and obtaining a synchronization log of standard execution.
Owner:SHENZHEN MINGHUI INTELLIGENT TECH CO LTD

Audio and video monitoring and early warning method and system based on multi-modal model driving

The invention relates to the technical field of big data, and discloses an audio and video monitoring and early warning method and system based on multi-modal model driving, and the method comprises the steps: S1, collecting audio and video signals in a monitoring scene, and extracting the energy distribution characteristics of the audio and video signals, wherein the energy distribution characteristics at least comprise the frequency energy distribution of the audio signal and the brightness and color change rate of the video signal; s2, multi-modal feature extraction of energy distribution features is carried out based on audio and video signals, and environment background feature vectors are constructed in combination with environment background information; s3, dynamically adjusting anomaly detection thresholds of the audio signal and the video signal based on the environmental background feature vector; s4, comparing the audio and video signal energy distribution characteristics extracted in real time with an anomaly detection threshold value, and when the audio and video signal energy distribution characteristics exceed the range of the anomaly detection threshold value, determining an abnormal event and obtaining early warning information; and S5, sending the early warning information from the edge equipment to a monitoring center through a low-bandwidth communication protocol.
Owner:TIANHE COLLEGE GUANGDONG POLYTECHNIC NORMAL UNIV +2

Live broadcast bullet screen real-time feedback method and system based on interactive semantic matching

The invention relates to the technical field of user interaction, in particular to a live broadcast bullet screen real-time feedback method and system based on interaction semantic matching, and the method comprises the steps: receiving and caching user bullet screens in real time, carrying out the density analysis and priority evaluation of bullet screen flows through the combination of filtering and resource optimization rules, dynamically adjusting the response frequency, and screening key bullet screens; a distillation processing mechanism intention unit is introduced, a user intention label in a key bullet screen is extracted, a vector is generated, historical behavior data of a sending user is obtained at the same time, the intention label, the vector and current real-time scene data of a live broadcast room are fused, and a bullet screen vector is constructed; performing similarity calculation in combination with the live broadcast content and user historical behaviors to obtain first feedback content; based on a reinforcement learning sentiment analysis strategy, combining user preference memory to generate second feedback content; and embedding the second feedback content into the current live broadcast interface through the client. According to the invention, real-time forward feedback and intelligent guidance of the live broadcast bullet screen are realized through interactive semantic matching.
Owner:HANGZHOU XINGMAI YUNSHANG TECHNOLOGY CO LTD

Full-process AI image creation method and system based on diffusion type image generation

The invention discloses a full-process AI image creation method and system based on diffusion type image generation, and relates to the field of AI image creation, and the method comprises the steps: obtaining a natural language creation instruction inputted by a user, and analyzing and extracting a main body element data set and a visual style label data set; generating a sequential task sequence including script generation, split mirror design, picture generation and video synthesis; generating a script frame and a split script data set through the large model; calculating the semantic similarity between the split content and the main body element, and performing supplementary optimization; generating a diffusion model cue word and obtaining rendering parameters to generate a rendered picture; the video is synthesized after the visual coverage is verified, and final output is achieved through feedback adjustment; according to the method, a complete control link from creation guiding to generation output can be constructed, and the customized generation requirement in a complex creation scene is met.
Owner:王永泉

Video stream processing method for dynamic Gaussian compression and adaptive code rate regulation

The invention discloses a video stream processing method for dynamic Gaussian compression and adaptive code rate regulation, which is suitable for scenes such as virtual reality, augmented reality and three-dimensional video, and comprises the following steps: S1, Gaussian attribute modeling and initialization; s2, constructing a binary hash grid; s3, constructing a deformation prediction network; s4, designing a mask pruning mechanism; s5, entropy modeling and arithmetic coding and decoding module design; s6, model training; and S7, video stream transmission under multiple code rates. According to the method, a unified scheme combining Gaussian volume cloud coding and adaptive video transmission is proposed for the first time, the video data storage and transmission cost is remarkably reduced, and the comprehensive performance superior to that of an existing method is obtained on multiple real and synthetic data sets.
Owner:THE CHINESE UNIV OF HONG KONG (SHENZHEN) FUTURE NETWORK OF INTELLIGENCE INST +1

Controllers in MPEG avatar representation format

Some embodiments of a method may include: obtaining a MPEG Avatar Representation Format (MARF)-based file; parsing the MARF-based file into a data structure; selecting an asset from the data structure; verifying the asset comprises a controller set property; obtaining controller set data corresponding to the controller set property; obtaining container data using the controller set data; determining a mime type using the controller set data; decoding the container data using the mime type; obtaining a mesh and a skeleton using the container data; and animating an avatar corresponding to the mesh and the skeleton.
Owner:INTERDIGITAL CE PATENT HOLDINGS SAS

Grouping intruder video events from smart security cameras into priority groups for efficient video monitoring

A system comprising security devices and a computing device. The security devices may each be configured to detect an event in response to video frames and audio and communicate the video frames, the audio and a notification of the event. The computing device may be configured to execute one or more steps, the steps comprising performing a classification of the notifications of the events into a plurality of event levels and displaying a list of the notifications of the events according to the classification. Each of the plurality of event levels determined to be high priority events are added to a top of the list as the high priority events are detected. Each of the plurality of event levels determined to be a lower priority than the high priority events are added to the list below the high priority events as the events with the lower priority are detected.
Owner:KUNA SYSTEMS CORP

Audio and video player control method based on voice instruction

The invention relates to the technical field of audio and video control, and discloses an audio and video player control method based on a voice instruction. The method comprises the steps that an original voice instruction stream of a user is collected, the instruction stream comprises a time domain audio signal sequence, an environment noise spectrum and user pronunciation characteristic parameters, and voice information can be comprehensively captured; multi-modal instruction analysis processing is carried out on the original voice instruction stream, a structured control instruction set containing acoustic control intention identification, semantic operation object description and context correlation parameters is generated, and the analysis precision is improved; then executing player state adaptation based on the set, generating a dynamic control response sequence containing an equipment state adjustment command, a media content positioning parameter and an interface interaction logic identifier, driving a player to execute a multi-dimensional control operation and generating real-time play control effect feedback data; and finally, multi-modal analysis parameters are optimized according to feedback data, a self-adaptive instruction analysis strategy is generated, and the control experience of a user on the audio and video player is optimized.
Owner:ONWAY TECH LTD

Multi-terminal scene synchronization control method and device, equipment and medium

The invention relates to the technical field of equipment operation and maintenance, can be applied to business scenes of financial science and technology, medical health and the like, and discloses a multi-terminal scene synchronization control method, which comprises the following steps: a control terminal receives current virtual scene position and current virtual scene view angle information sent when a user terminal starts a virtual scene, and generates a synchronization preview picture based on the information; the control terminal triggers an anti-control mode and sends an operation authority forbidding instruction to the user terminal, and the user terminal forbids an operation input signal acquisition function after receiving the instruction; and the control terminal sets a target virtual scene position and a target virtual scene view angle parameter of the user terminal through the interactive interface, generates a rendering execution instruction based on the target parameter, and controls the user terminal to load the instruction to complete virtual scene rendering. The target virtual scene position and the view angle parameter of the user terminal are set at the control end, dynamic guidance and picture control of the control end on the user terminal are achieved, and view angle conflicts are avoided in cooperation with an operation authority forbidding mechanism.
Owner:CHINA PING AN LIFE INSURANCE CO LTD

Short video intelligent editing method and system based on multi-modal analysis

The invention discloses a short video intelligent editing method and system based on multi-modal analysis, and relates to the technical field of video editing. The method is used for improving editing efficiency and visual experience and comprises the following steps: extracting lip motion features of a character, visual saliency features of a commodity and a voice emotion intensity value from a target short video stream to form multi-modal time sequence data; afterwards, the voice stream is recorded, a product keyword timestamp is extracted, the alignment degree is calculated through dynamic time warping in combination with a visual saliency peak value, and a preliminary editing point set is generated through weighted evaluation in combination with an emotional intensity value; constructing an editing decision optimization model based on deep reinforcement learning, taking the multi-modal features as state input, adjusting the retention probability of editing points through a joint reward function, and selecting an optimal transition mode; and the lip movement and voice synchronization error before and after the editing point and the emotional and visual continuity of the transition section are analyzed, the discontinuous region is smoothed, and the edited finished product is output, so that precise short video intelligent editing is realized.
Owner:ANHUI XINGBANG DIGITAL TECHNOLOGY GROUP CO LTD

Video display and storage method and system based on wayland protocol and storage medium

The invention belongs to the technical field of computers, and provides a video display and storage method and system based on a wayland protocol and a storage medium in order to solve the problems that an existing wayland video system is high in rendering delay, high in CPU load, low in storage efficiency and poor in cross-platform compatibility. Video frames are directly written into an Overlay layer through DRM-KMS and mixed by bypassing a synthesizer, meanwhile, CPU memory amplitude overhead is eliminated through zero copy transmission, an ISP-GPU-VPU direct transmission assembly line is constructed, full-link DMA-BUF direct transmission is achieved, and delay is greatly reduced; iSP preprocessing, GPU shader scaling and VPU coding are all executed by special hardware, so that the load is reduced, and the multi-path processing capability is improved; vPU dynamic code rate compression is utilized, segmented storage is carried out according to events / time, metadata is synchronized to a database, redundant storage is avoided, and space is saved.
Owner:CHANGSHA YINGBEIDI ELECTRONIC TECH CO LTD

Multi-modal diffusion-based long video role scene decoupling generation method and system

The invention discloses a long video role scene decoupling generation method and system based on multi-modal diffusion, and relates to the technical field of image processing, and the method comprises the steps: S1, synthesizing the advanced features of a role and a scene through a SigLIP encoder and a DINOv2 encoder; s2, performing cross-modal feature fusion on the advanced features to obtain joint features, and compressing the joint features to obtain compact vectors; s3, generating text features according to the text prompt; s4, potential codes are generated from an input video through a causal 3D convolution encoder, the potential codes pass through a linear projection matrix and then are spliced with a memory state for dimension reduction, and a segmented potential vector sequence is obtained; s5, performing decoupling perception generation on the segmented potential vector sequence through an improved 3D-UNet, and performing deconvolution up-sampling reconstruction after deterministic sampling to obtain an RGB video segmented sequence; according to the method, the key problems of rough dynamic control, limited generation length and over-high resource consumption in long video generation are solved, and the quality and efficiency of the generated video are remarkably improved.
Owner:湖南马栏山视频先进技术研究院有限公司

Video stream processing method and device, equipment and medium

The invention relates to the technical field of artificial intelligence, can be applied to business scenes of medical health, financial science and technology and the like, and discloses a video stream processing method which comprises the steps of collecting current environment parameters, generating a mode switching instruction and determining a target processing mode; obtaining multi-dimensional context awareness data, and selecting a target detection model; key area coordinate parameters in the video frame sequence are extracted, and grading resolution parameters are determined; and based on the target processing mode, the target detection model and the grading resolution parameter, constructing a video processing strategy matrix, executing the video processing strategy matrix to perform coding processing on the video stream, and generating a target coding video stream. According to the invention, through intelligent mode switching based on the current environment parameters, dynamic adjustment of the video processing mode is realized, and the adaptive capacity of the system in a complex environment is improved; through target detection model selection in combination with multi-dimensional context awareness data, the detection precision is optimized, and the reliability of visual analysis is improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Event information guided image deblurring and high frame rate reconstruction method

The invention discloses an event information guided image deblurring and high-frame-rate reconstruction method. The method comprises the following steps: step 1, obtaining a blurred image and an event stream corresponding to the blurred image to construct a training data set; 2, constructing a deblurring and frame insertion combined reconstruction network based on a dynamic cross-modal fusion module; step 3, training to obtain a trained joint reconstruction network based on the cross-modal attention mechanism and the UNet architecture; and step 4, inputting a to-be-reconstructed blurred image and a corresponding event stream into the trained deblurring and frame insertion combined reconstruction network based on the dynamic cross-modal fusion module to obtain a final multi-frame reconstruction image. According to the method, a higher weight dynamic state is given to an event mode in a high-speed motion scene, an image data weight is added in a static scene, and a scale feature fusion strategy is adopted for a seriously fuzzy scene, so that the problems of semantic difference and noise interference are reduced, and the image recovery quality of a fuzzy region is improved.
Owner:HUNAN UNIV

Short video active defense encryption system based on device fingerprint and dynamic confusion field

The invention relates to the technical field of short video encryption, and discloses a short video active defense encryption system based on a device fingerprint and a dynamic confusion field, and the key point of the technical scheme is that the system comprises a device fingerprint generation module, a mother video encryption module, a slice encryption module, a behavior recognition and defense module and an encryption logic update regulation and control module. The system generates a unique device fingerprint hash value through a multi-modal feature, generates a dynamic confusion field in combination with a chaotic system and a quantum random number, realizes differential encryption of a mother video and slices, and is embedded with zero-knowledge consanguinity proof to support traceability verification. The behavior recognition and defense module monitors user behaviors in real time and dynamically adjusts a confusion strategy or triggers an active defense mechanism, and the encryption logic updating regulation and control module optimizes the updating frequency according to playing data and reduces the batch crawling risk. According to the invention, the security and anti-attack capability of the short video content can be effectively improved.
Owner:HANGZHOU POPCORN EAGLE EYE TECH CO LTD

Video stream real-time coding and decoding transmission method under cluster

The invention relates to the technical field of cluster video stream processing, and discloses a video stream real-time coding and decoding transmission method under a cluster. The method comprises the following steps: firstly, acquiring video stream coding parameter text data, link state time sequence data and equipment performance index data of multiple nodes of a target cluster to form a transmission link data set; semantic analysis is carried out on the coding parameter text data to obtain a coding semantic feature vector, dynamic fluctuation features are extracted from the link state time sequence data to obtain a link fluctuation feature vector, and cross-modal fusion is carried out to generate a fusion transmission feature set; generating an abnormal association degree score set by using a pre-trained multi-layer sensing network model, and obtaining an abnormal source node and an equipment defect type by combining root cause tracing according to the abnormal association degree score set; and finally, generating a dynamic optimization strategy and feeding back to the transmission control system to trigger parameter calibration. According to the method, the abnormal root cause can be accurately traced, the transmission parameters are optimized, and the cluster video stream transmission quality is improved.
Owner:ZHEJIANG VERSATILE MEDIA

Multi-modal time sequence alignment AI video translation method and system

The invention relates to the technical field of subtitle translation, in particular to a multi-modal time sequence alignment AI video translation method and system, and the method comprises the steps: 1, carrying out the multi-modal analysis of a to-be-translated video, and obtaining audio separation data, voiceprint feature data and visual time sequence data; 2, performing cross-language translation and context optimization on the basis of the voice of the audio separation data to generate a target language text, and synthesizing target language voice retaining the original voice color in combination with the voiceprint feature data and the target language text; generating a mouth shape animation matched with the target language voice based on the lip key point data and the limb action time sequence data; and step 3, performing four-dimensional alignment on the target language voice, the translated text, the mouth shape animation and the limb action sequence through a cross-modal time sequence encoder, and dynamically adjusting the layout of the bilingual subtitles to adapt to a video picture. According to the method and the device, multi-mode synchronization can be taken into consideration during video translation, so that the body actions such as voice, subtitles and mouth shapes are kept aligned.
Owner:HANGZHOU BAOMIHUA TECH CO LTD