Intelligent pricing matching system and method based on video data asset transaction

CN122736672APending Publication Date: 2026-09-11XIAN WENYU NETWORK INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610896678.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-22
Publication Date
2026-09-11

AI Technical Summary

Technical Problem

[0002]随着数字经济与人工智能产业快速发展,视频数据已成为重要的数据资产,市场中视频数据资产的交易需求持续增长,但当前行业内普遍缺乏专业化、智能化的交易支撑体系,传统交易模式多依靠人工核验、经验估价与线下对接,不仅存在数据资产标识不统一、行为数据处理滞后、资产维度刻画片面等问题,还难以精准评估视频内容质量、热度、稀缺度及 AI 训练价值,定价方式主观随意、供需撮合效率低下,同时视频原始文件未做精细化切片与多模态特征挖掘,数据资产利用率不足,且无法动态迭代评估与定价规则,制约了视频数据资产交易的市场化发展

Benefits of technology

本申请所提供的基于视频数据资产交易的智能定价撮合系统及方法,通过全链路模块化设计实现视频数据资产从采集、标准化处理、特征挖掘到交易撮合、定价优化的一体化运转,统一资产标识完成数据资产规范化管理,依托流式计算实时构建多维度视频资产画像,结合视频切片与多模态特征提取深度挖掘数据价值;同时依托智能匹配、动态定价竞价机制大幅提升供需撮合效率与定价合理性,搭配训练反馈模块形成闭环迭代体系,持续优化资产评估、推荐及定价参数,有效解决传统交易模式效率低、评估片面、定价主观、资产利用率不足等痛点,推动视频数据资产交易的智能化发展。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122736672A_ABST
    Figure CN122736672A_ABST
Patent Text Reader

Abstract

The application discloses an intelligent pricing and matching system and method based on video data asset transaction, comprising: a data acquisition module for collecting large-scale behavior data of a video file to form tradable video data assets; an event access module for encapsulating the behavior data into standardized event messages; a streaming computing module for performing big data analysis on the standardized event messages, including heat, supply and demand relationship analysis, user demand portrait and video scarcity evaluation, and dynamically generating a video asset portrait; a video analysis and slicing module for decoding and scene slicing of the video file; a multi-modal feature extraction module for extracting modal features and correlating them to the video data assets; an intelligent recommendation matching module for matching constraint conditions with the video asset portrait to calculate matching degrees and generate matching results; and an intelligent pricing and bidding module for generating reference prices and bidding strategies. The application improves transaction efficiency and data utilization value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to an intelligent pricing and matching system for video data asset transactions, belonging to the field of asset transaction pricing and matching technology. Background Technology

[0002] With the rapid development of the digital economy and artificial intelligence industry, video data has become an important data asset. The market demand for video data asset transactions continues to grow. However, the industry currently lacks a professional and intelligent transaction support system. Traditional transaction models rely heavily on manual verification, experience-based valuation, and offline connections. This not only results in problems such as inconsistent data asset identification, lagging behavioral data processing, and one-sided asset dimension characterization, but also makes it difficult to accurately assess the quality, popularity, scarcity, and AI training value of video content. Pricing methods are subjective and arbitrary, and the efficiency of supply and demand matching is low. At the same time, the original video files have not been finely sliced ​​and multimodal feature mined, resulting in insufficient utilization of data assets and the inability to dynamically iterate evaluation and pricing rules, which restricts the market-oriented development of video data asset transactions. Summary of the Invention

[0003] According to one aspect of this application, an intelligent pricing and matching system for video data asset transactions is provided, which improves transaction efficiency and data utilization value.

[0004] The intelligent pricing and matching system for video data asset transactions is characterized by including: The data acquisition module is used to collect video files and large-scale behavioral data generated around the video files, and to generate a unified asset identifier for each video file, forming tradable video data assets. The event access module is used to encapsulate the behavioral data into standardized event messages and transmit them in real time through a message queue; The streaming computing module is used to perform real-time big data analysis on the event stream composed of the standardized event messages and dynamically generate video asset profiles. The big data analysis includes popularity analysis, supply and demand analysis, real-time user demand profiles, and video scarcity assessment. The video asset profile includes at least several dimensions from the following: content tags, quality scores, popularity scores, scarcity scores, transaction conversion scores, training value scores, and price reference factors. The video parsing and slicing module is used to decode the video file and segment the scene, dividing the complete video into multiple video segments and generating an independent segment identifier for each video segment; A multimodal feature extraction module is used to extract modal features from the video file, wherein the modal features include at least two of visual features, audio features, text features and temporal features, and associate the extracted features with the corresponding video data assets; The intelligent recommendation and matching module is used to calculate the matching degree between the video asset profile and the constraints input by the demander, and generate recommendation or matching results. The intelligent pricing and bidding module is used to dynamically generate reference prices and bidding strategies based on the multi-dimensional attributes of video assets, historical transaction information, supply and demand relationships, and training value scores. The training feedback module is used to select samples that meet the conditions from video assets and output them to the model for training. Based on the performance indicators obtained from training, it updates the training value score, recommendation weight, or pricing parameters of the video assets in reverse.

[0005] Furthermore, the video parsing and slicing module divides the video file into multiple video segments according to the segmentation rules, and binds a corresponding start time and end time to each video segment, and obtains the audio features, image tags and independent quality scores of the segment from the results of the multimodal feature extraction module. The classification rules include at least one of the following: camera changes, voice pauses, subtitle timelines, action changes, or scene transitions.

[0006] Furthermore, the visual features include at least one of keyframe features, scene features, character features, object features, motion features, and image sharpness; The audio features include at least one of the following: speech content, background music, speech rate, and emotional features; The text features include at least one of title, description, caption, comment, and OCR text; The temporal features include at least one of shot sequence, action change, or scene change.

[0007] Furthermore, the streaming computing module dynamically updates the popularity score, transaction conversion score, or training value score in the video asset profile based on the judgment indicators. The evaluation metrics include at least one of the following: video playback completion rate, repeat viewing rate, user dwell time, sharing rate, purchase conversion rate, and training call frequency.

[0008] Furthermore, the intelligent recommendation matching module performs semantic parsing and vectorization modeling on the text input by the demander, extracts character attributes, scene features, style tendencies, emotional atmosphere, and time environment constraints, and performs cross-modal alignment calculations with the visual feature vectors, text semantic vectors, or style vectors of the video data assets. Combining text semantic similarity, visual content similarity, style matching degree, video quality indicators, and historical interaction data, a multi-objective fusion ranking model is constructed.

[0009] Furthermore, the execution of the intelligent pricing and bidding module includes: For newly uploaded video data assets, price factors will not be considered in the initial stage, and they will be pushed to the first batch of demand pools in a priority and indiscriminate manner. The average transaction price of the first N completed orders is used as the base price for the video data asset, where N is a preset positive integer; Once the base pricing is established, the video data asset will no longer be pushed to demanders with budgets lower than the base pricing; Specifically, whenever the number of new transactions for the video data asset reaches a predetermined quantity, the unit price is automatically increased by a first percentage, and the recommended exposure is increased by a second percentage. If no new transactions occur for the video data assets within a consecutive number of days, the price will automatically decrease by a third percentage, the recommended exposure will decrease by a fourth percentage, and the market demand trend will be reassessed. The price of the video data asset shall not be lower than the base price.

[0010] Furthermore, the training metrics acquired by the training feedback module include model accuracy, recall, video understanding score, action recognition accuracy, scene recognition accuracy, text-video matching degree, generated video quality score, and sample contribution. The training value score, recommendation ranking, and pricing parameters of the video data asset are updated based on the training metrics.

[0011] According to another aspect of this application, a smart pricing matching method for video data asset transactions is provided, comprising the following steps: S1. Collect video files and large-scale behavioral data generated around the video files, and generate a unified asset identifier for each video file to form tradable video data assets; S2. Encapsulate the behavioral data into standardized event messages and transmit them in real time through a message queue; S3. Perform real-time big data analysis on the event stream composed of the standardized event messages to dynamically generate video asset profiles. The big data analysis includes popularity analysis, supply and demand relationship analysis, real-time user demand profiles, and video scarcity assessment. The video asset profile includes at least several dimensions from the following: content tags, quality scores, popularity scores, scarcity scores, transaction conversion scores, training value scores, and price reference factors. S4. Decode and segment the video file, divide the complete video into multiple video segments, and generate an independent segment identifier for each video segment; S5. Extract modal features from the video file, the modal features including visual features, audio features, text features and temporal features, and associate the extracted features with the corresponding video data assets; S6. Based on the constraints input by the demand side, calculate the matching degree with the video asset profile and generate recommendation or matching results; S7. Based on the multidimensional attributes of the video data assets, historical transaction information, supply and demand relationship, and the training value score, dynamically generate reference prices and bidding strategies; S8. Select samples that meet the conditions from the video data assets and output them to the model for training. Based on the performance indicators obtained from the training, update the training value score, recommendation weight, or pricing parameters of the video data assets in reverse.

[0012] Furthermore, in step S4, the video file is divided into multiple video segments according to the segmentation rules, and a corresponding start time and end time are bound to each video segment. The audio features, image tags and independent quality scores of the segment are obtained from the results of the multimodal feature extraction module. The classification rules include at least one of the following: camera changes, voice pauses, subtitle timelines, action changes, or scene transitions; In step S6, the text input by the demander is semantically parsed and vectorized, and at least one of the following is extracted: character attributes, scene features, style tendencies, emotional atmosphere, or time environment constraints. Cross-modal alignment calculation is performed with the visual feature vector, text semantic vector, or style vector of the video data asset. Combining text semantic similarity, visual content similarity, style matching degree, video quality indicators, and historical interaction data, a multi-objective fusion ranking model is constructed.

[0013] Furthermore, S7 includes: For newly uploaded video data assets, price factors will not be considered in the initial stage, and they will be pushed to the first batch of demand pools in a priority and indiscriminate manner. The average transaction price of the first N completed orders is used as the base price for the video data asset, where N is a preset positive integer; Once the base pricing is established, the video data asset will no longer be pushed to demanders with budgets lower than the base pricing; Specifically, whenever the number of new transactions for the video data asset reaches a predetermined quantity, the unit price is automatically increased by a first percentage, and the recommended exposure is increased by a second percentage. If no new transactions occur for the video data assets within a consecutive number of days, the price will automatically decrease by a third percentage, the recommended exposure will decrease by a fourth percentage, and the market demand trend will be reassessed. A price protection mechanism is set up to ensure that the price of the video data asset is not lower than the base price.

[0014] The beneficial effects that this application can produce include: The intelligent pricing and matching system and method for video data asset transactions provided in this application achieve integrated operation of video data assets from acquisition, standardized processing, feature mining to transaction matching and pricing optimization through a modular design across the entire chain. It unifies asset identification to complete the standardized management of data assets, builds multi-dimensional video asset profiles in real time based on streaming computing, and deeply mines data value by combining video slicing and multimodal feature extraction. At the same time, it greatly improves the efficiency of supply and demand matching and the rationality of pricing by relying on intelligent matching and dynamic pricing bidding mechanisms. Combined with a training feedback module, it forms a closed-loop iterative system to continuously optimize asset evaluation, recommendation and pricing parameters, effectively solving the pain points of traditional transaction models such as low efficiency, one-sided evaluation, subjective pricing and insufficient asset utilization, and promoting the intelligent development of video data asset transactions. Attached Figure Description

[0015] Figure 1 This is a schematic diagram of a smart pricing matching system for video data asset transactions according to one embodiment of this application. Figure 2 This is the overall system logic framework of an intelligent pricing matching system based on video data asset transactions in one embodiment of this application; Figure 3 This is a subsystem architecture diagram of the intelligent pricing and bidding module in one embodiment of this application; Figure 4 This is a sequence diagram of message interactions between the buyer, seller, and bidding matching engine in one embodiment of this application; Figure 5 This is a flowchart of an intelligent pricing and matching method for video data asset transactions in one embodiment of this application. Detailed Implementation

[0016] The present application is described in detail below with reference to the embodiments, but the present application is not limited to these embodiments.

[0017] See Figure 1-4 A smart pricing and matching system for video data asset transactions includes: The data acquisition module is used to collect video files and large-scale behavioral data generated around the video files, and to generate a unified asset identifier for each video file, forming tradable video data assets. The event access module is used to encapsulate the behavioral data into standardized event messages and transmit them in real time through a message queue; The streaming computing module is used to perform real-time big data analysis on the event stream composed of the standardized event messages and dynamically generate video asset profiles. The big data analysis includes popularity analysis, supply and demand analysis, real-time user demand profiles, and video scarcity assessment. The video asset profile includes at least several dimensions from the following: content tags, quality scores, popularity scores, scarcity scores, transaction conversion scores, training value scores, and price reference factors. Preferably, when updating multiple profile dimensions of the same video asset, the streaming computing module executes the updates in the order of dependency: first, it updates the content tags and quality scores; then, it updates the popularity scores and transaction conversion scores; next, it updates the scarcity scores and training value scores; and finally, it calculates the price reference factor based on the latest values ​​of all the above dimensions to ensure the consistency and timeliness of the profile data.

[0018] The video parsing and slicing module is used to decode the video file and segment the scene, dividing the complete video into multiple video segments and generating an independent segment identifier for each video segment; A multimodal feature extraction module is used to extract modal features from the video file, wherein the modal features include at least two of visual features, audio features, text features and temporal features, and associate the extracted features with the corresponding video data assets; The intelligent recommendation and matching module is used to calculate the matching degree between the video asset profile and the constraints input by the demander, and generate recommendation or matching results. The intelligent pricing and bidding module is used to dynamically generate reference prices and bidding strategies based on the multi-dimensional attributes of video assets, historical transaction information, supply and demand relationships, and training value scores. The training feedback module is used to select samples that meet the conditions from video assets and output them to the model for training. Based on the performance indicators obtained from training, it updates the training value score, recommendation weight, or pricing parameters of the video assets in reverse.

[0019] Specifically, the data acquisition module completes the assetization and unique ownership confirmation of unstructured video data, breaking the traditional problem of video files lacking asset attributes and being unable to be traded in compliance with regulations. The acquired data includes original video files from various sources and full-link behavioral trajectory data generated around the videos. The original video files include all types of video resources such as surveillance videos, self-media videos, industry scene videos, sports event videos, and government videos. The full-link behavioral trajectory data includes user browsing, clicking, collecting, downloading, searching, sharing, transaction inquiries, transaction records, and secondary reuse information. Furthermore, through a combination of hash algorithms, timestamps, and source traceability coding, a globally unique and tamper-proof digital asset ID is generated for each original video file. This identifier will be bound to the video's source information, ownership information, original file, and behavioral logs, completing the standardized digital asset upgrade of video data ownership confirmation, traceability, transaction, and statistics, providing a unique identity certificate for subsequent use. At the same time, the module has data cleaning, deduplication, and compliance screening capabilities, automatically filtering damaged videos, illegal videos, and duplicate resources to ensure the compliance and validity of the assets entering the database.

[0020] The event access module addresses the issues of fragmented and scattered raw behavioral data formats, high transmission latency, and the inability to perform unified calculations. User- and platform-generated video behavioral data is scattered, inconsistently formatted, and exists in various forms, including structured, semi-structured, and unstructured, making it unsuitable for direct intelligent analysis and computation. The event access module is responsible for standardized data encapsulation and efficient transmission. It unifies the format of various types of raw behavioral data, completes field completion, de-identifies data, and performs compliance processing, encapsulating fragmented and diverse behavioral data into standardized event messages. Each message includes asset ID, behavior type, behavior time, user entity, behavior parameters, and scene information. Simultaneously, the module incorporates a high-performance message queue mechanism to achieve millisecond-level real-time transmission, asynchronous push, traffic shaping, and data caching of standardized event messages, preventing data loss and congestion in high-concurrency scenarios. This module supports real-time access to massive amounts of high-frequency user behavior, ensuring that the subsequent streaming computing module can continuously and in real-time acquire the latest behavioral data, enabling dynamic updates of video asset status.

[0021] The streaming computing module is based on standardized event messages transmitted in real time and relies on a big data streaming computing engine to realize dynamic iteration and real-time updates of video asset profiles. Unlike traditional static profiles, the profiles generated by this module can be adjusted in real time according to market behavior and changes in asset status. It abandons the offline batch computing mode and adopts real-time incremental computing logic to continuously consume event data in the message queue and quantify and dynamically update the various value dimensions of video assets.

[0022] Specifically, the video asset profile includes: content tags, which automatically generate precise tags for video scenes, themes, content, industries, and applicable scenarios through AI semantics and scene recognition, enabling accurate asset classification and retrieval; quality scores, which are quantitatively scored based on objective indicators such as video resolution, frame rate, clarity, image stability, absence of watermarks, absence of blur, and audio fidelity; popularity scores, which dynamically assess the market attention of the asset based on market behavior data such as real-time views, interactions, searches, and dissemination; scarcity scores, which combine platform asset inventory, the number of similar resources in the industry, and scene uniqueness to assess the exclusivity and scarcity value of the video asset; transaction conversion scores, which assess the market transaction potential of the asset based on the historical consultation rate, transaction rate, and repurchase rate of the video and similar videos; training value scores, which assess the suitability and effectiveness of video data for artificial intelligence model training, algorithm iteration, and scene fitting; and price reference factors, which integrate all dimensions of data to generate quantitative factors that can be directly used for intelligent pricing, providing a core basis for dynamic pricing.

[0023] It is worth noting that the quality rating includes:

[0024] in, For resolution mapping function, ; The frame rate normalization function, ; This is the bit rate normalization function. ; Image sharpness is scored using Laplace variance. , The maximum variance threshold; , , and The first, second, third, and fourth preset weights are respectively 0.3, 0.2, 0.2, and 0.3; the final quality score range is normalized to [0,1], and can be multiplied by 100 to convert to a percentage. The popularity rating includes: ; Where t is the current time window; ; ; The calculation of the scarcity score includes:

[0025] in, This represents the total number of video assets on the platform. The number of assets with the same content tag as this asset; The initial value of the training value score The calculations include:

[0026] in, To provide rich visual features, To provide rich audio features, Enrich text features with indicators; For visual feature weights, For audio feature weights, Text feature weights.

[0027] Furthermore, traditional video transactions are mostly conducted as whole segments, which suffers from poor flexibility, resource waste, and limited adaptability. The video parsing and slicing module enables the fine-grained splitting of video assets and the construction of the smallest transaction units. The module supports intelligent decoding, format compatibility, and frame parsing of complete original video files, adapting to video resources of various encoding formats and completing standardized parsing processing of video files. On this basis, through scene recognition, shot segmentation, and content breakpoint detection technologies, long and complex complete videos are automatically divided into multiple fine-grained video segments with independent content, single scenes, and standardized durations. This avoids the problems of low efficiency and inconsistent standards in manual slicing. At the same time, each segmented video segment generates an independent and unique segment identifier, which is associated and bound to the original video asset identifier, realizing two-level asset ownership of the original video and sub-segments. The split video segments can independently participate in retrieval, matching, pricing, and trading, meeting the needs of demanders for on-demand procurement, segment procurement, and precise procurement, and improving the utilization rate and transaction flexibility of video data assets.

[0028] The multimodal feature extraction module breaks through the limitations of traditional methods that rely solely on surface information in videos to assess value. By using multimodal AI technology to deeply mine the underlying features of videos, it achieves comprehensive quantification of asset value. Videos are typical multimodal data, including visual, audio, text, and temporal information. The module can extract features from at least two modalities simultaneously and accurately associate all feature data with the corresponding original video assets and video clip assets.

[0029] The intelligent recommendation and matching module breaks down information barriers between demanders and asset providers, solving the problems of inefficiency, low matching accuracy, and supply-demand mismatch in traditional transaction retrieval. Demanders can input personalized constraints through the platform, including custom filtering conditions such as the content scene, quality parameters, duration, modal features, scarcity level, price range, training purpose, and popularity requirements of the required videos. Based on machine learning matching algorithms, the system performs comprehensive matching calculations between the user-input constraints and the multi-dimensional profiles and multi-modal features of video assets already constructed by the system. Through multi-layered logic such as weighted scoring, similarity ranking, and scene suitability verification, the system filters out video assets and slice resources that meet the requirements, generating an asset recommendation list and supply-demand matching results. It also supports priority ranking, prioritizing the push of assets with the highest suitability, best cost performance, and strongest value matching. This not only meets the content procurement needs of ordinary users but also accurately matches the model training data procurement needs of AI companies and research institutions, improving the efficiency of transaction matching.

[0030] The intelligent pricing and bidding module replaces the traditional manual and fixed pricing model, realizing market-oriented, intelligent, and dynamic pricing of video data assets. The core pricing criteria include: multi-dimensional attributes of the video asset such as quality, popularity, scarcity, and feature richness; historical transaction information such as the transaction price, frequency, and price fluctuation patterns of similar assets on the platform; current market supply and demand relationships such as platform asset inventory, user demand, and supply-demand gap; and model training value scoring. Through big data algorithms, market price fluctuations are calculated in real time, dynamically generating a real-time reference price for each video asset and each slice of asset. Simultaneously, it incorporates an intelligent bidding strategy system, supporting various transaction modes such as listing pricing, dynamic negotiation, and auction. It can automatically adjust bidding rules based on asset scarcity level and market popularity. A bidding premium mechanism is activated for scarce, high-value assets, while a parity-optimal matching mechanism is implemented for ordinary assets, ensuring both maximum asset value and fair and market-oriented transaction prices.

[0031] After model training is completed, the training feedback module collects performance metrics in real time during the training process and feeds them back to the system. Based on real training effect data, it dynamically updates core parameters and corrects the training value score of video assets, ensuring that the value score matches the real AI training effect and avoiding subjective scoring bias. It also optimizes the recommendation weight of the intelligent recommendation matching module, increasing the matching priority of high-value training data. Furthermore, it iterates the pricing parameters of the intelligent pricing module, making the pricing more in line with the real training value and market reuse value of the assets. Through continuous training feedback iteration, the system's profile accuracy, matching accuracy, and pricing rationality will continuously improve with the accumulation of transaction and training data.

[0032] It is worth noting that the update of the training value score includes:

[0033] in, For the relative increase of the core indicator, if ,but ; The marginal contribution of the sample in the margin is calculated using the Shapley value.

[0034] In one embodiment, a video sample V is used to train an action recognition model. Before training, the model's accuracy was 82%. After adding a batch containing V, the accuracy increased to 82.5%, an improvement of ΔM = 0.005. Shapley value calculations show that V's contribution ranks in the top 10% among all samples, with C = 0.8. The original training value score for this video was 65 points, so the new score is 65 + 0.1 × (0.005 × 0.8 × 100) = 65.04. Although the improvement is small, the cumulative effect is significant. Simultaneously, the recommendation weight increases by 5%, and the weight of the training value factor in the pricing parameters is increased from 0.1 to 0.11.

[0035] The video parsing and slicing module divides the video file into multiple video segments according to the segmentation rules, and binds a corresponding start time and end time to each video segment. It also obtains the audio features, image tags and independent quality scores of the segment from the results of the multimodal feature extraction module. The classification rules include at least one of the following: camera changes, voice pauses, subtitle timelines, action changes, or scene transitions.

[0036] Specifically, the system intelligently divides a complete video file into multiple video segments with independent content, complete scenes, and coherent semantics according to preset segmentation rules. These segmentation rules include at least one of the following: shot changes, audio pauses, subtitle timelines, action changes, or scene transitions. Single-rule or multi-rule fusion judgments can be made based on video content characteristics to avoid defects such as scene fragmentation, semantic truncation, and loss of key information caused by fixed-length slicing. After video slicing, a unique segment identifier is generated for each video segment, achieving dual-layer asset ownership management of a single complete video and multiple subdivided video segments. Simultaneously, each segmented video segment undergoes comprehensive attribute binding, automatically assigning each segment a corresponding start time, end time, representative keyframes, corresponding time-segment subtitle text, segment audio features, image tags, and an independent quality score, ensuring each video segment possesses standardized asset attributes. The split video segments can independently participate in the platform's asset retrieval, feature matching, intelligent pricing, supply and demand matching, and transaction flow, supporting demanders to accurately purchase partial video content on demand, effectively improving the utilization and market-based trading of video data assets.

[0037] It is worth noting that the detection and determination of the lens change includes: For adjacent frames and Calculate the normalized pixel difference and color histogram chi-square distance If ΔP > 0.3 and ΔH > 0.5, then it is determined to be a lens shear point; The detection of speech pauses includes: calling an open-source speech activity detection library to divide the audio stream into frames; when non-speech frames exceeding 500ms are continuously detected, they are determined as a potential semantic segmentation point, and the nearest shot cut point near that point is preferred as the final slice boundary. In one embodiment of this application, for a video segment, if a pause in speech is detected first, followed by a camera switch within 0.8 seconds, the system will determine this point as a forced slicing point. If only a camera change is detected but the speech is continuous, it will be considered an in-scene transition and no slicing will be performed.

[0038] The visual features include at least one of keyframe features, scene features, character features, object features, motion features, and image sharpness; The audio features include at least one of the following: speech content, background music, speech rate, and emotional features; The text features include at least one of title, description, caption, comment, and OCR text; The temporal features include at least one of shot sequence, action change, or scene change.

[0039] Specifically, visual features extract visual dimensions such as objects, scenes, colors, actions, people, environment, and composition, suitable for AI model training scenarios such as image recognition and visual inspection, and used to characterize the content composition, visual quality, and scene performance of video images; audio features extract audio features such as human voice, ambient sound, sound effects, timbre, volume, and speech content, adapted to scenarios such as speech recognition, voiceprint detection, and audio classification, and used to characterize the content information and auditory emotional attributes of the audio accompanying the video; text features extract text information, keywords, semantic content, and other text features from the video through OCR recognition, subtitle extraction, and on-screen text parsing, used to collect the semantic information of the text inside and outside the video; temporal features extract temporal dynamic features such as the changing patterns of video frames, the rhythm of scene switching, the temporal logic of actions, and the development of events, adapted to advanced AI training scenarios such as temporal prediction, behavior analysis, and dynamic recognition, and used to depict the dynamic evolution of the video over time.

[0040] It is worth noting that the object features in the visual features are extracted by the YOLOv8 model and output as an 85-dimensional vector, including 80 COCO class probabilities, 4 bounding box coordinates and 1 confidence score. The object detection results of the keyframes of the whole video are pooled to finally obtain a fixed-length object existence probability vector. The title description semantics in the text features are encoded into a 384-dimensional dense vector using the Sentence-BERT (all-MiniLM-L6-v2) model and stored in the FAISS vector index for subsequent cross-modal alignment calculation with the requester's text. The temporal features are extracted using a pre-trained I3D model to create a 1024-dimensional temporal feature vector for each video segment. This vector comprehensively represents the evolution of actions within the segment.

[0041] The streaming computing module dynamically updates the popularity score, transaction conversion score, or training value score in the video asset profile based on the judgment indicators. The evaluation metrics include at least one of the following: video playback completion rate, repeat viewing rate, user dwell time, sharing rate, purchase conversion rate, and training call frequency.

[0042] Specifically, the popularity score is mainly calculated based on user dwell time, video completion rate, repeat viewing rate, and sharing rate. The higher the values ​​of these four indicators, the higher the market exposure and user attention of the video asset, and the popularity score is adjusted upward in real time. Conversely, the score is dynamically adjusted downward as traffic decreases and user interaction declines. The transaction conversion score uses the purchase conversion rate as the criterion. It combines data on pre-conversion behaviors such as playback, consultation, and sharing to calculate the asset's transaction potential in real time. When user purchase behavior increases and the conversion rate improves, the transaction conversion score is automatically increased incrementally. If there is no conversion behavior for a long period of time, the score is gradually adjusted downward. The training value score is dynamically updated based on the training call frequency. It counts the frequency with which each video asset is called, trained, and tested by the AI ​​model. The higher the frequency of asset training reuse and the stronger the reuse stability, the higher its data adaptability and training value. The corresponding training value score is adjusted upward in real time. Assets that have not been called for a long time and have poor training effects are automatically downgraded and have their scores deducted.

[0043] The intelligent recommendation matching module performs semantic parsing and vectorization modeling on the text input by the demander, extracts character attributes, scene features, style tendencies, emotional atmosphere, and time environment constraints, and performs cross-modal alignment calculations with the visual feature vectors, text semantic vectors, or style vectors of the video data assets. Combining text semantic similarity, visual content similarity, style matching degree, video quality indicators, and historical interaction data, a multi-objective fusion ranking model is constructed.

[0044] Specifically, FFmpeg / ffprobe is used to perform structured parsing of the video, obtaining basic information such as video resolution, duration, frame rate, bit rate, and encoding format. Keyframes are automatically extracted based on time interval and shot change detection algorithms. A visual content understanding model is used to analyze the keyframes for characters, scenes, actions, objects, text, and visual style, generating structured visual semantic information. To address the problem that traditional video retrieval relies solely on single-frame analysis, a temporal context-based video semantic fusion mechanism is introduced to model the semantic relationships between consecutive keyframes, using BERT... The system generates video-level content summaries using the Decoder multimodal generation model, enabling dynamic understanding of video event chains, shot logic, and overall atmosphere. During the demand matching phase, the system performs semantic parsing and vectorization modeling of user demand text, automatically extracting multidimensional constraints such as character attributes, scene features, style tendencies, emotional atmosphere, and temporal environment. These constraints are then cross-modal aligned with visual feature vectors, text semantic vectors, and style vectors from the video material side. Combining text semantic similarity, visual content similarity, style matching degree, video quality indicators, and historical interaction data, a multi-objective fusion ranking model is constructed to achieve accurate recommendation and intelligent transaction matching of video materials, thereby improving video material retrieval efficiency, content reuse rate, and supply-demand matching accuracy.

[0045] It is worth noting that the multi-objective fusion ranking model is implemented using a lightweight gradient booster, and the input feature vector of the model is the concatenated form: ,in, Embedded for users, Embedded in video, Scoring is based on bilinear alignment. and These are text cosine similarity and visual cosine similarity, respectively. To rate the quality, Historical click-through rate; In model training, the three optimization objectives are whether clicks, inquiries, and transactions occurred in historical transactions. A weighted Pairwise Ranking Loss is used for training. Specifically, for a single search by the same customer, the loss function for positive samples (videos with completed transactions) and negative samples (videos with exposure but no clicks) is as follows:

[0046] in, These are marginal parameters; and These are the model's prediction scores for positive and negative samples, respectively. In one embodiment of this application, the requester inputs a video of a rainy nighttime city road for autonomous driving training; the system parses the constraints: {Scene: City road, Weather: Rain, Lighting: Night, Purpose: Training}; the constraints are vectorized, and cosine similarity is calculated with the multimodal features of all videos to select Top-100 candidates. Then, using the LightGBM model, combined with the quality score and historical popularity of the candidate videos, the system sorts and outputs the final recommendation list; for example, a video that meets the criteria of nighttime, rainy day, high definition, and high training frequency is ranked first.

[0047] The execution of the intelligent pricing and bidding module includes: For newly uploaded video data assets, price factors will not be considered in the initial stage, and they will be pushed to the first batch of demand pools in a priority and indiscriminate manner. The average transaction price of the first N completed orders is used as the base price for the video data asset, where N is a preset positive integer; Once the base pricing is established, the video data asset will no longer be pushed to demanders with budgets lower than the base pricing; Specifically, whenever the number of new transactions for the video data asset reaches a predetermined quantity, the unit price is automatically increased by a first percentage, and the recommended exposure is increased by a second percentage. If no new transactions occur for the video data assets within a consecutive number of days, the price will automatically decrease by a third percentage, the recommended exposure will decrease by a fourth percentage, and the market demand trend will be reassessed. The price of the video data asset shall not be lower than the base price.

[0048] Specifically, the intelligent pricing and bidding module is used to achieve cold start pricing for video data assets without prior benchmarks, dynamic price adjustment based on market supply and demand, and intelligent exposure matching control. This solves the technical problems of new video assets having no historical transaction data, no basis for pricing, prices deviating from the market, and supply and demand imbalance. For newly uploaded video data assets with no transaction records, a cold start mechanism is adopted. In the initial stage of asset listing, no price constraints are introduced, and no price screening is performed on demanders. New video data assets are pushed to the platform's first batch of demand pools without discrimination, ensuring that newly listed assets can obtain initial exposure and initial transaction opportunities, and quickly accumulate the first batch of market transaction sample data. Based on the transaction data accumulated in the cold start stage, the first N valid transaction orders of the video data asset are counted, and the average of the actual transaction prices of the N orders is used as the benchmark price of the video data asset, where N is a positive integer preset by the system. This forms an initial price benchmark that conforms to the actual market acceptance, avoiding the subjective bias of manual pricing and the rigidity of fixed initial pricing. Once the base price of the video data asset is officially generated, the module activates a price threshold filtering push mechanism, preventing the video data asset from being pushed to users with budgets below the base price. This achieves precise tiered matching of asset value and user spending power, improving transaction efficiency. Simultaneously, a dynamic adaptive pricing and exposure linkage mechanism based on transaction activity adjusts price and exposure weights in real time according to changes in asset transaction activity and market demand: when the cumulative number of new transactions for a video data asset reaches the system's preset quantity, it is determined that the asset has strong market demand and scarcity, automatically increasing the asset's unit price by a first percentage and simultaneously increasing the asset's platform recommendation exposure by a second percentage, achieving premium value and traffic support for high-demand assets; when no new transactions occur for a preset number of consecutive days, it is determined that the asset's market activity has declined and demand has decreased. When demand weakens, the system automatically lowers the unit price of the asset by a third percentage, simultaneously reducing the platform's recommended exposure by a fourth percentage. It also re-collects market supply and demand data, user search behavior, and price trends of similar assets to iterate and reassess market demand trends, adapting to the latest market changes. Furthermore, a price floor protection mechanism is set up to ensure that the real-time transaction price of video data assets never falls below the previously generated base price during all dynamic price adjustments. This guarantees the stability of the underlying value of video data assets and avoids asset value distortion and market disorder caused by unlimited price declines. Therefore, through cold start foundation building, benchmark pricing, threshold push, popularity premium, cold volume price reduction, and floor price protection, the system achieves intelligent, adaptive, and dynamic iteration of video data asset prices based on supply and demand, transaction volume, and market trends, constructing a market-oriented, traceable, and dynamically iterative video asset pricing system.

[0049] In one embodiment of this application, in order to incentivize content providers to continuously upload high-quality and massive amounts of video content, the platform constructs an intelligent bidding and dynamic pricing mechanism based on real market feedback to achieve a dynamic balance between content value, traffic distribution, and market demand.

[0050] Based on a video content understanding model, multi-dimensional feature analysis is performed on the materials, including: video content semantics, scene type, character and action features, video quality, resolution and duration, style and atmosphere, material scarcity, historical transaction performance, user clicks, collections and conversion behavior; combined with content similarity calculation results and user demand vectors, cold start market validation is performed on newly uploaded materials. Initially, the platform does not consider price factors but prioritizes pushing materials indiscriminately to the first batch of target demand pools that meet the needs of the scenarios. This is dynamically determined based on asset type, such as 100 demanders. A market value model for the materials is quickly established through real transaction behavior. The average transaction price of the first 10 completed orders is used as the base price for the materials, forming the initial market value anchor. Once the base price is established, the system will no longer push materials to demanders with budgets lower than the base price to prevent low-price transactions from damaging the value system of high-quality materials. In the subsequent circulation stage, a time-series-based Hybrid Boosting Regression Model is introduced to dynamically predict and adjust material prices and traffic. The model comprehensively analyzes user interest trends, video popularity changes, transaction speed, market supply and demand, time decay factors, material competition, creator historical performance, user budget distribution, and click-through rate and conversion rate. Through a multi-condition joint decision-making mechanism, price and exposure resources are dynamically and collaboratively controlled. For every 10 new completed orders, the unit price of the materials will automatically increase by 10%. We also recommend increasing exposure by 100% and expanding the reach of our content. As the conversion rate of creative materials continues to improve, the system further enhances the recommendation weight and bidding priority; When a highly relevant user has insufficient budget, the system automatically recommends similar alternative materials or adjusts the push strategy.

[0051] If no new transactions are generated for a particular stock for three consecutive days, the system will automatically trigger the failed bid and market feedback correction mechanism. Material prices will automatically decrease by 10%; Recommended exposure decreased by 100% simultaneously; The system reassesses the popularity of materials and market demand trends; Automatically analyze the reasons for failed bids, including factors such as high prices, insufficient content matching, and decreased demand; Dynamically adjust subsequent bidding ranges and recommendation strategies.

[0052] To prevent vicious price competition, a price protection mechanism is set up, and the price of the materials must not be lower than the base price. At the same time, when the materials lack market feedback for a long period of time, the system can gradually reduce the recommended resources, and the push volume can be reduced to 0 at the lowest point, so as to optimize the overall traffic utilization efficiency and resource allocation efficiency of the platform.

[0053] Compared to traditional fixed-price and static recommendation models, this application integrates multimodal content analysis, time series prediction, dynamic bidding mechanisms, and multi-condition correlation decision-making to achieve intelligent and adaptive adjustment of the video material pricing system and traffic distribution system, thereby improving the efficiency of material transactions, the revenue capacity of high-quality content, and the overall commercial conversion level of the platform.

[0054] It is worth noting that the hybrid Boosting regression model is implemented using XGBoost, and its pricing function is: ; in, The feature vector includes: asset quality score, popularity score, scarcity score, training value score, average transaction price of similar assets on the platform, number of transactions in the past 7 days, inventory, and user demand index. Output the base forecast price; yes The model's time-series correction term for the residuals.

[0055] As a preferred implementation, when the number of new transactions for video assets reaches a preset quantity M=10, a price adjustment is triggered; the unit price is automatically increased by the first percentage, which is set to 5%~15%, and the system automatically selects the percentage based on the scarcity of the asset. If the scarcity is >80 points, 15% is applied, otherwise 5% is applied; the recommended exposure is increased by the second percentage, which is set to 100%, that is, the exposure is doubled; if no transaction occurs within the consecutive reservation period T=3 days, the price is automatically decreased by the third percentage, which is 10%, and the exposure is decreased by the fourth percentage, which is 50%.

[0056] A newly uploaded drone aerial video has no historical data. The system indiscriminately pushes it to 100 relevant clients; the first N=10 orders that are successfully completed have transaction prices of 48, 52, 50, 49, 51, 53, 47, 50, 52, and 48 yuan respectively. Therefore, the base price is... After that, customers with a budget of less than 50 yuan will not see this video. When 10 new orders are placed and the 20th order is completed, if the scarcity is high, the price will be automatically increased by 10% to 55 yuan, and the recommended exposure will be doubled.

[0057] The training metrics obtained by the training feedback module include model accuracy, recall, video understanding score, action recognition accuracy, scene recognition accuracy, text-video matching degree, generated video quality score, and sample contribution. The training value score, recommendation ranking, and pricing parameters of the video data asset are updated based on the training metrics.

[0058] Specifically, the platform automatically selects video data assets that meet the sample quality criteria as training samples and inputs them into the training task to complete data learning, performance testing, and parameter iteration optimization. Simultaneously, it collects multi-dimensional performance data throughout the training process in real time, constructing a standardized training indicator system. These indicators include model accuracy, recall, video understanding score, action recognition accuracy, scene recognition accuracy, text-video matching accuracy, generated video quality score, and sample contribution, comprehensively quantifying the actual empowering effect and data application value of the video samples. Accuracy and recall are used to evaluate the overall training fit; video understanding score and text-video matching accuracy are used to measure the training gain effect of multimodal semantic matching; action recognition accuracy and scene recognition accuracy are used to characterize the adaptability of video content to visual perception tasks; and generated video... The quality score is used to evaluate the application effect of video materials in content generation tasks; the sample contribution is used to quantify the contribution weight of a single video asset to the overall accuracy improvement and generalization ability optimization; the above multi-dimensional training indicators are weighted and fused to output a quantitative comprehensive training effect score, and based on this score, multi-dimensional parameters are updated in reverse iteration to accurately correct the training value score of the corresponding video data assets, so that the score results fully match the real training gain effect and avoid the subjective bias of traditional static scoring; at the same time, the recommendation ranking weight of the intelligent recommendation matching module is updated, and the ranking priority of video assets with high training contribution and excellent application effect is increased to strengthen the accurate matching of high-quality data; simultaneously, the core pricing parameters of the intelligent pricing and bidding module are updated in conjunction with the update, and the sample training contribution and optimization gain are included in the pricing factor to realize the market-based pricing of the training value of video data assets.

[0059] See Figure 5 The intelligent pricing and matching method for video data asset transactions includes the following steps: S1. Collect video files and large-scale behavioral data generated around the video files, and generate a unified asset identifier for each video file to form tradable video data assets; S2. Encapsulate the behavioral data into standardized event messages and transmit them in real time through a message queue; S3. Perform real-time big data analysis on the event stream composed of the standardized event messages to dynamically generate video asset profiles. The big data analysis includes popularity analysis, supply and demand relationship analysis, real-time user demand profiles, and video scarcity assessment. The video asset profile includes at least several dimensions from the following: content tags, quality scores, popularity scores, scarcity scores, transaction conversion scores, training value scores, and price reference factors. S4. Decode and segment the video file, divide the complete video into multiple video segments, and generate an independent segment identifier for each video segment; S5. Extract modal features from the video file, the modal features including visual features, audio features, text features and temporal features, and associate the extracted features with the corresponding video data assets; S6. Based on the constraints input by the demand side, calculate the matching degree with the video asset profile and generate recommendation or matching results; S7. Based on the multidimensional attributes of the video data assets, historical transaction information, supply and demand relationship, and the training value score, dynamically generate reference prices and bidding strategies; S8. Select samples that meet the conditions from the video data assets and output them to the model for training. Based on the performance indicators obtained from the training, update the training value score, recommendation weight, or pricing parameters of the video data assets in reverse.

[0060] Specifically, the system collects various raw video files and all user behavior data generated around these files. The collected data undergoes cleaning, deduplication, and compliance filtering. A globally unified and unique asset identifier is generated for each video file, completing the assetization and ownership confirmation of the video data. This constructs a standardized video data asset that is traceable, quantifiable, and tradable. The system also uniformly encapsulates and processes the collected fragmented and multi-formatted video behavior data, converting it into standardized event messages with standardized fields and structures. Asynchronous caching, traffic shaping, and millisecond-level real-time transmission of these event messages are achieved through message queues, ensuring the integrity, real-time performance, and orderliness of the behavior data. Furthermore, the system enhances the real-time event stream composed of continuously transmitted standardized event messages. Streamlined computing and dynamic analysis, relying on real-time judgment indicators such as video playback completion rate, repeat viewing rate, user dwell time, sharing rate, purchase conversion rate, and training call frequency, dynamically update and generate video asset profiles. These video asset profiles are multi-dimensional quantitative profiles, including at least several dimensions such as content tags, quality scores, popularity scores, scarcity scores, transaction conversion scores, training value scores, and price reference factors, enabling real-time iterative updates to the value status of video assets. The original video files undergo unified decoding, format adaptation, and frame parsing preprocessing. Based on at least one segmentation rule among shot changes, voice pauses, subtitle timelines, action changes, or scene transitions, the complete video file is intelligently segmented into multiple semantically complete and scene-based segments. Independent video segments; each video segment is assigned a unique segment identifier, which is then bound to the segment's start and end times, keyframes, subtitle text, audio features, image tags, and independent quality scores, enabling refined segmentation and two-layer ownership management of video assets; multi-dimensional feature mining is performed on video files and subdivided video segments to extract multimodal features, including visual features, audio features, text features, and temporal features, and all features are structurally associated with the corresponding video data assets; visual features include at least one of keyframe features, scene features, character features, object features, motion features, and image clarity; audio features include at least one of speech content, background music, speech rate, and emotional features; text features include at least one of... Features include at least one of title, description, subtitle, comment, and OCR text; temporal features include at least one of shot sequence, action change, or scene change, constructing a complete multi-dimensional feature system for video assets; semantic parsing and vectorization modeling of the text requirements input by the demand side, extracting constraints such as character attributes, scene features, style tendencies, emotional atmosphere, and time environment, and performing cross-modal alignment calculations between the demand vector and the visual feature vector, text semantic vector, and style vector of the video data assets; combining text semantic similarity, visual content similarity, style matching degree, video quality indicators, and historical interaction data to construct a multi-objective fusion ranking model, generating the optimal asset recommendation result and supply-demand matching result through comprehensive scoring and ranking;Based on the multidimensional attributes of video data assets, historical transaction information, real-time market supply and demand, and training value scores, the system dynamically calculates asset reference prices and generates appropriate bidding strategies. Newly uploaded video assets employ a cold-start mechanism, initially being pushed indiscriminately to the demand pool, with the average price of the first N completed orders serving as the base price. After the base price is generated, low-budget demand is blocked, and prices are adaptively adjusted based on asset transaction volume. Increased transaction volume leads to a premium, while decreased sales result in a price reduction, ensuring the asset price never falls below the base price, achieving dynamic market-based pricing control. Samples meeting the quality criteria are selected from compliant and high-quality video data assets for training. Real-time data collection during training includes multidimensional performance metrics such as model accuracy, recall, video understanding score, action recognition accuracy, scene recognition accuracy, text-video matching degree, generated video quality score, and sample contribution. A comprehensive training effect score is calculated through weighted fusion of these metrics, iteratively updating the training value score of the corresponding video data asset, platform recommendation ranking weight, and intelligent pricing core parameters, thus optimizing data transactions, application training, and parameter self-calibration.

[0061] In step S4, the video file is divided into multiple video segments according to the segmentation rules, and a corresponding start time and end time are bound to each video segment. The audio features, image tags and independent quality scores of the segment are obtained from the results of the multimodal feature extraction module. The classification rules include at least one of the following: camera changes, voice pauses, subtitle timelines, action changes, or scene transitions; In step S6, the text input by the demander is semantically parsed and vectorized, and at least one of the following is extracted: character attributes, scene features, style tendencies, emotional atmosphere, or time environment constraints. Cross-modal alignment calculation is performed with the visual feature vector, text semantic vector, or style vector of the video data asset. Combining text semantic similarity, visual content similarity, style matching degree, video quality indicators, and historical interaction data, a multi-objective fusion ranking model is constructed.

[0062] Specifically, the video slicing process employs diversified intelligent segmentation rules for segmentation, including at least one of shot changes, voice pauses, subtitle timelines, action changes, or scene transitions. After slicing is completed, each video segment is precisely bound to a corresponding start and end time, and the audio features, image tags, and independent quality scores corresponding to the video segment are obtained from the output of the multimodal feature extraction module, so that each subdivided video segment has complete, searchable, and tradable standardized attributes. The system performs semantic parsing and vectorization modeling on the text input by the demand side, extracting at least one of the following: character attributes, scene features, style tendencies, emotional atmosphere, or time environment constraints. The extracted demand features are then aligned across modalities with the visual feature vectors, text semantic vectors, or style vectors of the video data assets. Combined with text semantic similarity, visual content similarity, style matching degree, video quality indicators, and historical interaction data, a multi-objective fusion ranking model is constructed to achieve accurate matching and intelligent ordering of video assets.

[0063] S7 includes: For newly uploaded video data assets, price factors will not be considered in the initial stage, and they will be pushed to the first batch of demand pools in a priority and indiscriminate manner. The average transaction price of the first N completed orders is used as the base price for the video data asset, where N is a preset positive integer; Once the base pricing is established, the video data asset will no longer be pushed to demanders with budgets lower than the base pricing; Specifically, whenever the number of new transactions for the video data asset reaches a predetermined quantity, the unit price is automatically increased by a first percentage, and the recommended exposure is increased by a second percentage. If no new transactions occur for the video data assets within a consecutive number of days, the price will automatically decrease by a third percentage, the recommended exposure will decrease by a fourth percentage, and the market demand trend will be reassessed. A price protection mechanism is set up to ensure that the price of the video data asset is not lower than the base price.

[0064] Specifically, for newly uploaded video data assets, price is not considered in the initial listing phase. They are prioritized and pushed to the platform's first batch of demand pools to quickly accumulate initial transaction data. The transaction prices of the first N orders completed for each video data asset are collected, and the base price of the video data asset is calculated by averaging these prices, where N is a preset positive integer. After the base price is generated, a price threshold push mechanism is activated, preventing the video data asset from being pushed to demanders with budgets lower than the base price, thus achieving precise tiered matching of supply and demand. At the same time, the price and exposure weight are dynamically adjusted in conjunction with transaction popularity: when a new transaction order for a video data asset is placed... When a preset quantity is reached, the market demand for the asset is deemed strong. The asset's unit price is automatically increased by the first percentage, and the platform's recommended exposure is simultaneously increased by the second percentage, strengthening the trading and circulation capabilities of high-quality and popular assets. When no new transactions occur for a preset number of consecutive days, the market demand is deemed to have declined. The asset's unit price is automatically reduced by the third percentage, and the recommended exposure is simultaneously reduced by the fourth percentage. The market demand trend is reassessed to adapt to the latest market conditions. A price protection mechanism is set up throughout the process to ensure that the real-time transaction price of video data assets is never lower than the base price, guaranteeing the stability of the underlying value of video data assets.

[0065] Example An autonomous driving company had a requirement for "nighttime, rainy days, and urban expressways." The system extracted constraints, matched them with a video library, and recommended a 15-second high-definition dashcam video clip that met the criteria. The initial price for this video was 80 yuan. After three transactions, the price increased to 96 yuan. The requesting company purchased the video for 96 yuan. Subsequently, 10 autonomous driving companies used the video for training. The training feedback module increased its training value score from 70 to 92, doubled the recommendation weight, updated the pricing factor, and the price seen by new users was updated to 120 yuan.

[0066] The above description is merely a few embodiments of this application and is not intended to limit this application in any way. Although this application discloses preferred embodiments as described above, it is not intended to limit this application. Any changes or modifications made by those skilled in the art without departing from the scope of the technical solution of this application using the disclosed technical content are equivalent to equivalent implementation cases and fall within the scope of the technical solution.

Claims

1. An intelligent pricing and matching system for video data asset transactions, characterized in that: include: The data acquisition module is used to collect video files and large-scale behavioral data generated around the video files, and to generate a unified asset identifier for each video file, forming tradable video data assets. The event access module is used to encapsulate the behavioral data into standardized event messages and transmit them in real time through a message queue; The streaming computing module is used to perform real-time big data analysis on the event stream composed of the standardized event messages, and dynamically generate video asset profiles. The big data analysis includes popularity analysis, supply and demand analysis, real-time user demand profiles, and video scarcity assessment. The video asset profile includes at least several dimensions from the following: content tags, quality scores, popularity scores, scarcity scores, transaction conversion scores, training value scores, and price reference factors.

2. A video parsing and slicing module, used to decode and segment the video file, divide the complete video into multiple video segments, and generate an independent segment identifier for each video segment; A multimodal feature extraction module is used to extract modal features from the video file, wherein the modal features include at least two of visual features, audio features, text features and temporal features, and associate the extracted features with the corresponding video data assets; The intelligent recommendation and matching module is used to calculate the matching degree between the video asset profile and the constraints input by the demander, and generate recommendation or matching results. The intelligent pricing and bidding module is used to dynamically generate reference prices and bidding strategies based on the multi-dimensional attributes of video assets, historical transaction information, supply and demand relationships, and training value scores. The training feedback module is used to select samples that meet the conditions from video assets and output them to the model for training. Based on the performance indicators obtained from training, it updates the training value score, recommendation weight, or pricing parameters of the video assets in reverse.

3. The intelligent pricing and matching system for video data asset transactions according to claim 1, characterized in that, The video parsing and slicing module divides the video file into multiple video segments according to the segmentation rules, binds a corresponding start time and end time to each video segment, and obtains the audio features, image tags and independent quality scores of the video segment from the results of the multimodal feature extraction module. The classification rules include at least one of the following: camera changes, voice pauses, subtitle timelines, action changes, or scene transitions.

4. The intelligent pricing and matching system for video data asset transactions according to claim 1, characterized in that, The visual features include at least one of keyframe features, scene features, character features, object features, motion features, and image sharpness; The audio features include at least one of the following: speech content, background music, speech rate, and emotional features; The text features include at least one of title, description, caption, comment, and OCR text; The temporal features include at least one of shot sequence, action change, or scene change.

5. The intelligent pricing and matching system for video data asset transactions according to claim 1, characterized in that, The streaming computing module dynamically updates the popularity score, transaction conversion score, or training value score in the video asset profile based on the judgment indicators. The evaluation metrics include at least one of the following: video playback completion rate, repeat viewing rate, user dwell time, sharing rate, purchase conversion rate, and training call frequency.

6. The intelligent pricing and matching system for video data asset transactions according to claim 1, characterized in that, The intelligent recommendation matching module performs semantic parsing and vectorization modeling on the text input by the demander, extracts character attributes, scene features, style tendencies, emotional atmosphere, and time environment constraints, and performs cross-modal alignment calculations with the visual feature vectors, text semantic vectors, or style vectors of the video data assets. Combining text semantic similarity, visual content similarity, style matching degree, video quality indicators, and historical interaction data, a multi-objective fusion ranking model is constructed.

7. The intelligent pricing and matching system for video data asset transactions according to claim 1, characterized in that, The execution of the intelligent pricing and bidding module includes: For newly uploaded video data assets, price factors will not be considered in the initial stage, and they will be pushed to the first batch of demand pools in a priority and indiscriminate manner. The average transaction price of the first N completed orders is used as the base price for the video data asset, where N is a preset positive integer; Once the base pricing is established, the video data asset will no longer be pushed to demanders with budgets lower than the base pricing; Specifically, whenever the number of new transactions for the video data asset reaches a predetermined quantity, the unit price is automatically increased by a first percentage, and the recommended exposure is increased by a second percentage. If no new transactions occur for the video data assets within a consecutive number of days, the price will automatically decrease by a third percentage, the recommended exposure will decrease by a fourth percentage, and the market demand trend will be reassessed. The price of the video data asset shall not be lower than the base price.

8. The intelligent pricing and matching system for video data asset transactions according to claim 1, characterized in that, The training metrics obtained by the training feedback module include model accuracy, recall, video understanding score, action recognition accuracy, scene recognition accuracy, text-video matching degree, generated video quality score, and sample contribution. The training value score, recommendation ranking, and pricing parameters of the video data asset are updated based on the training metrics.

9. A smart pricing and matching method for video data asset transactions, characterized in that, Includes the following steps: S1. Collect video files and large-scale behavioral data generated around the video files, and generate a unified asset identifier for each video file to form tradable video data assets; S2. Encapsulate the behavioral data into standardized event messages and transmit them in real time through a message queue; S3. Perform real-time big data analysis on the event stream composed of the standardized event messages to dynamically generate video asset profiles. The big data analysis includes popularity analysis, supply and demand relationship analysis, real-time user demand profiles, and video scarcity assessment. The video asset profile includes at least several dimensions from the following: content tags, quality scores, popularity scores, scarcity scores, transaction conversion scores, training value scores, and price reference factors. S4. Decode and segment the video file, divide the complete video into multiple video segments, and generate an independent segment identifier for each video segment; S5. Extract modal features from the video file, the modal features including visual features, audio features, text features and temporal features, and associate the extracted features with the corresponding video data assets; S6. Based on the constraints input by the demand side, calculate the matching degree with the video asset profile and generate recommendation or matching results; S7. Based on the multidimensional attributes of the video data assets, historical transaction information, supply and demand relationship, and the training value score, dynamically generate reference prices and bidding strategies; S8. Select samples that meet the conditions from the video data assets and output them to the model for training. Based on the performance indicators obtained from the training, update the training value score, recommendation weight, or pricing parameters of the video data assets in reverse.

10. The intelligent pricing and matching method for video data asset transactions according to claim 8, characterized in that, In step S4, the video file is divided into multiple video segments according to the segmentation rules, and a corresponding start time and end time are bound to each video segment. The audio features, image tags and independent quality scores of the segment are obtained from the results of the multimodal feature extraction module. The classification rules include at least one of the following: camera changes, voice pauses, subtitle timelines, action changes, or scene transitions; In step S6, the text input by the demander is semantically parsed and vectorized, and at least one of the following is extracted: character attributes, scene features, style tendencies, emotional atmosphere, or time environment constraints. Cross-modal alignment calculation is performed with the visual feature vector, text semantic vector, or style vector of the video data asset. Combining text semantic similarity, visual content similarity, style matching degree, video quality indicators, and historical interaction data, a multi-objective fusion ranking model is constructed.

11. The intelligent pricing and matching method for video data asset transactions according to claim 8, characterized in that, S7 includes: For newly uploaded video data assets, price factors will not be considered in the initial stage, and they will be pushed to the first batch of demand pools in a priority and indiscriminate manner. The average transaction price of the first N completed orders is used as the base price for the video data asset, where N is a preset positive integer; Once the base pricing is established, the video data asset will no longer be pushed to demanders with budgets lower than the base pricing; Specifically, whenever the number of new transactions for the video data asset reaches a predetermined quantity, the unit price is automatically increased by a first percentage, and the recommended exposure is increased by a second percentage. If no new transactions occur for the video data assets within a consecutive number of days, the price will automatically decrease by a third percentage, the recommended exposure will decrease by a fourth percentage, and the market demand trend will be reassessed. A price protection mechanism is set up to ensure that the price of the video data asset is not lower than the base price.