Method for ott content operation optimization based on multi-modal perception and agent decision
By employing multimodal perception and intelligent agent decision-making methods, the baseband video signal of OTT network TV clients is sampled and digitized, solving the technical problem of the inability to automatically analyze and understand interface data in existing technologies. This enables automated and real-time scanning of OTT content operation, improves the accuracy and coverage of analysis results, supports data-driven decision-making, and enhances operational efficiency.
Patent Information
- Application Number
- CN202511156433.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-19
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2045-08-19
AI Technical Summary
Existing technologies cannot achieve automated and large-scale interface data exploration and content operation optimization in OTT content analysis. They are difficult to understand the complex layout and operational meaning of TV screens and lack the ability to make horizontal comparisons and differentiated analyses.
By employing multimodal perception and intelligent agent decision-making methods, the baseband video signal of the OTT network TV client is sampled and digitized, parsed into a set of screen content, and an interaction sequence is generated to construct a TV service model. In-depth analysis and temporal differential comparison are then performed to output an evaluation report.
It has enabled automated and real-time scanning of OTT content operations, improved the accuracy and coverage of analysis results, generated content arrangement suggestions, supported data-driven decision-making, and improved operational efficiency and accuracy.
Smart Images

Figure CN120730095B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of image communication, in particular to an OTT content operation optimization method based on multi-modal perception and agent decision. BACKGROUND
[0002] In an interactive video system, the system architecture usually contains a head-end system responsible for managing and distributing content, and a large number of client devices for receiving and presenting content. On the client device, the system presents a graphical user interface (GUI) through which the user interacts with the video service. The core of this GUI is usually an interactive program guide (IPG), which shows the user available video resources in the form of hierarchical menus, classified lists and content recommendations. Therefore, the interface layout, content organization and its dynamic changes of the IPG are not only the key to the core user experience of the service, but also directly reflect the content supply state at a certain time point. When it is necessary to systematically catalog or analyze the content presentation of a video service, the existing technology mainly relies on manual operation. The operator needs to simulate the end user, manually navigate the GUI interface of the client device using a remote control or the like, and record its content through screenshots or video recordings for subsequent analysis.
[0003] The existing technology faces technical bottlenecks in realizing the automatic analysis of the content of the video service client: first, the information black box problem of the client interface: from the technical process, the structured content information in the service head-end, after being encoded, transmitted and finally rendered by the client device as visual images on the screen, its original data structure has been lost. For the outside, only the final pixel stream can be perceived, and the data structure behind it cannot be directly accessed or queried, which makes the client interface a "black box" that cannot be directly probed. Moreover, this analysis method lacks the ability to scale and systematically compare, and cannot compare and analyze the differences of the operation content, column hotspots and main resources of different OTT platforms or television manufacturers. Further, the existing technology cannot accurately and comprehensively understand the complex layout, text and image mixed layout content on the television screen and the operation implications implied behind them.
[0004] Therefore, an OTT content operation optimization method based on multi-modal perception and agent decision is proposed. SUMMARY
[0005] The present application aims to provide an OTT content operation optimization method based on multi-modal perception and agent decision to solve the problems raised in the background.
[0006] To achieve the above-mentioned purpose, the present application provides the following technical solution: an OTT content operation optimization method based on multi-modal perception and agent decision, the method comprising:
[0007] The baseband video signal output by the OTT network TV client is sampled and digitized, and the baseband video signal is used to present the visual content of the interactive program guide;
[0008] The two-dimensional pixel information in the baseband video signal is parsed into a set of screen content containing business information and content metadata through a multimodal model; navigation command signals are generated and sent to the OTT network TV client through an automated exploration engine, and the subsequently returned baseband video signal is synchronously sampled to construct a complete interaction sequence.
[0009] Multiple interaction sequences obtained through iterative exploration are associated and aggregated to generate a TV service model of service access points and their transition relationships.
[0010] The internal structure of the television service model is analyzed in depth and compared with the historical television service model in the time domain. The decision analysis agent generates the strategy and compiles all the analysis, comparison and generation results to output an evaluation report of model data and strategy.
[0011] Preferably, the specific implementation process of sampling and digitizing the baseband video signal of the OTT network TV client includes:
[0012] The OTT network TV client establishes a physical communication link with the external video acquisition device through a high-definition multimedia interface; the OTT network TV client generates a baseband video signal, which is then digitally converted by the video acquisition device to generate static video frames in real time that represent the interactive program guide screen displayed on the terminal.
[0013] Preferably, the specific implementation process of parsing the baseband video signal into a set of screen content using a multimodal model includes:
[0014] The static video frame is input into the multimodal model, which recovers and extracts the image components encoded as visual elements during the client rendering process from the pixel domain of the TV screen, and identifies and classifies the graphic elements carrying service information in the screen; for the identified graphic elements, the position information in the screen coordinate system is calculated and output; optical character recognition is performed on text elements to extract the string content; and the classification, string content and position information of all elements are compiled into a set of screen content describing the content presented by the static video frame.
[0015] Preferably, the process of constructing a complete interaction sequence through an automated exploration engine includes:
[0016] The automated exploration engine performs the exploration process, parsing the screen content set, identifying the graphic elements that switch pages, and adding them to the queue to be explored. Based on a predefined TV service navigation strategy, the exploration engine generates interactive instructions and sends them to the control actuator to simulate the navigation instructions issued by the physical remote control device to change the state of the TV client. It also synchronizes with subsequent video frame sampling to complete the systematic data acquisition and construct a complete interaction sequence.
[0017] Preferably, the specific implementation process of the television service model that generates service access points and their transition relationships through a loop includes:
[0018] Following the exploration process, an analysis process is performed. First, a state signature is generated based on the interaction sequence, and the corresponding interactive program guide is identified as a service access point using this state signature. When a navigation command triggers a change in the presented content from the starting service access point to the target service access point, the directional transition relationship between the two is established and recorded. The exploration and analysis process is repeated to construct a visualized television service model from all the discovered service access points and their transition relationships.
[0019] Preferably, the process of performing in-depth analysis of the television service model includes:
[0020] The shortest path algorithm is applied to calculate the minimum number of interaction steps from the entry service access point to key TV content assets and service access points, evaluating the discoverability and navigation efficiency of paid content and value-added services. A centrality metric is used to score and rank the importance of all service access points, identifying core content aggregation pages and navigation hubs. A community discovery algorithm is run to divide the TV business model into different business logic domains, reflecting the service packaging and bundling strategies of the TV business. Finally, text content associated with each service access point is aggregated.
[0021] Preferably, the process of performing time-domain difference comparison between the television service model and multiple historical models includes:
[0022] The current period's television service model and multiple historical period's television service models are selected as inputs for differential analysis. A service access point correlation comparison process is executed, establishing a one-to-one correspondence between the current and historical models based on each status signature. Matched service access point pairs are compared one by one, identifying and recording changes in internal attributes. Unmatched service access points are identified as newly added or deleted. Based on the established service access point correspondence, navigation paths resulting from changes in television service interaction logic are identified and recorded for additions, deletions, and redirections. Differences in service menu structure, content entry locations, and channel arrangement order between the two maps are calculated and quantified to assess the periodic adjustments to content arrangement strategies. All changes are compiled into a time-series record of content and navigation changes.
[0023] Preferably, the specific implementation process of outputting an evaluation report of model data and strategy includes:
[0024] The various TV program performance indicators obtained from the deep analysis are deeply integrated with the content and navigation change records obtained from the time-domain difference comparison to form an aggregated record representing the platform's operational dynamics. A content strategy analysis agent, assigned preset market goals and content arrangement guidelines, is activated to comprehensively evaluate the aggregated record, identifying the fit and differences between the current platform content strategy and market operation hotspots. Based on the fit and differences, content arrangement suggestions for the next cycle of content operation are automatically generated, including recommendations for key content types, optimization of promotional resource positions, and responses to emerging trends. The deep analysis, time-domain comparison, and content arrangement suggestions are logically assembled into an explanatory text description. A graphical engine is invoked to generate trend charts showing changes in content distribution efficiency performance indicators. The completed content is output as a static report file, which integrates content operation status analysis and future strategy guidance.
[0025] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0026] 1. In terms of automation and scalability, controlling HID devices through multimodal models transforms a method reliant on manual browsing and monitoring into an automated and near-real-time method for acquiring operational content. This enables comprehensive and synchronous scanning and cataloging of the content operation status of OTT platforms within a short timeframe, significantly improving efficiency and coverage compared to existing technologies, thus laying the foundation for systematic industry competitive analysis.
[0027] 2. In terms of semantic understanding, by applying a multimodal model to perform deep semantic analysis on video frames and generating service state signatures robust to dynamic visual content, the system accurately restores screen pixels to a structured representation containing business logic. It can also reliably identify unique interface states even under interference from dynamic content such as advertisements and badges. This significantly improves the accuracy, consistency, and depth of the analysis results, enabling it to penetrate the visual surface and effectively understand the content arrangement and navigation logic behind the interface.
[0028] 3. During the decision-making phase, deep topology analysis is conducted on the constructed business navigation network model, and a strategy-making agent with specific business objectives is introduced to perform comprehensive judgment and forward-looking planning. This elevates the analysis from describing the objective state to guiding future strategies. It can automatically evaluate the effectiveness of content arrangement and generate optimization guidance including content type recommendations and promotion plans, providing data-driven and quantifiable decision support capabilities for content operations. Attached Figure Description
[0029] Figure 1 This is a flowchart of an OTT content operation optimization method based on multimodal perception and agent decision-making proposed in an embodiment of this invention application;
[0030] Figure 2 This is a schematic diagram of the structure of multimodal analysis proposed in an embodiment of this invention application;
[0031] Figure 3 This is a diagram illustrating the interactive process for generating an evaluation report as proposed in an embodiment of this invention. Detailed Implementation
[0032] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0033] Please see Figures 1-3 The present invention relates to an OTT content operation optimization method based on multimodal perception and intelligent agent decision-making, the specific implementation steps of which are as follows:
[0034] The baseband video signal output by the OTT network TV client, which presents interactive program guide visual content, is sampled and digitized;
[0035] The two-dimensional pixel information in the baseband video signal is parsed into a set of screen content containing business information and content metadata through a multimodal model; navigation command signals are generated and sent to the OTT network TV client through an automated exploration engine, and the subsequently returned baseband video signal is synchronously sampled to construct a complete interaction sequence.
[0036] By iteratively associating and aggregating the multiple parsed interaction sequences, a TV service model of the service access point and its transition relationship is generated.
[0037] The internal structure of the television service model is analyzed in depth; a time-domain differential comparison is performed with the topology model stored in historical cycles; a decision analysis agent generates a strategy; all analysis, comparison and generation results are finally compiled to output an evaluation report on content navigation efficiency and presentation strategy.
[0038] The technical solution of the present invention will be further described in detail below with reference to specific embodiments.
[0039] Example 1
[0040] This application discloses an OTT content operation optimization method based on multimodal perception and intelligent agent decision-making, including the optimization process of the central analysis server identifying the content operated by the manufacturer. (See attached document for details.) Figure 1 The specific implementation steps of the method proposed in this invention include: S1, sampling and digitizing the baseband video signal output by the OTT network TV client, which presents interactive program guide visual content; S2, parsing the two-dimensional pixel information in the baseband video signal into a set of screen content containing service information and content metadata through a multimodal model; S3, generating and sending navigation command signals to the OTT network TV client through an automated exploration engine, and synchronously sampling the subsequently returned baseband video signal to construct a complete interaction sequence; S4, associating and aggregating multiple interaction sequences obtained through iteration to generate a TV service model of service access points and their transition relationships; S5, performing in-depth analysis of the internal structure of the TV service model and comparing it with historical TV service models using temporal differential analysis; S6, generating strategies through a decision analysis agent, finally compiling all analysis, comparison, and generation results, and outputting an evaluation report of model data and strategies.
[0041] Furthermore, the baseband video signal output by the OTT network TV client, which presents the interactive program guide visual content, is sampled and digitized; corresponding to step S1 above; the specific implementation process includes:
[0042] Specifically, a physical communication link is first established between the HDMI output port of the OTT client and the HDMI input port of the video capture card or a device with equivalent specifications in the video capture subsystem via a high-speed data cable conforming to the HDMI 2.0 standard. When the OTT client is running, it generates a standard, unencrypted baseband video signal, accurately representing any interactive program guides or other user interfaces displayed on its screen. The video capture card intercepts this signal. The chipset inside the capture card performs analog-to-digital conversion and hardware compression, digitizing the continuous video signal into a series of discrete static video frames in real time. For example, for a 4K signal source, the capture card can capture it at 3840x2160 resolution and 60fps, or downsample it to 1920x1080 resolution and 60fps to reduce the subsequent processing load. These digitized video frame data are transmitted at high speed to the central analysis server through a high-bandwidth interface.
[0043] By offloading computationally intensive video encoding and decoding tasks from the central analysis server through the video acquisition subsystem, valuable computing resources can be focused on performing core image analysis tasks. This design not only achieves platform independence but also optimizes the resource allocation and processing efficiency of the entire system.
[0044] Furthermore, the two-dimensional pixel information in the baseband video signal is parsed into a set of screen content containing service information and content metadata using a multimodal model; corresponding to step S2 above, the specific implementation process includes:
[0045] Each frame of video image is input into the central analysis server. First, the zero-shot object detection model GroundingDINO identifies the bounding boxes of all potential UI components in the video frame. These components include video thumbnails, text blocks, buttons, icons, etc. The Segment Anything Model generates a pixel-accurate segmentation mask for each detected component, segmenting the image of each component and feeding it into the Multimodal Large Language Model (MLLM). Leveraging its understanding of UI design patterns, the model classifies the components, identifying them as "buttons," "icons," "text labels," etc., and infers their potential functions. For components containing text information, the OCR engine is invoked to extract their string content. Finally, the classification of all components, their text content, and their precise position information in the screen coordinate system—the bounding box coordinates—are stored in the screen content set.
[0046] By employing a multimodal model, unstructured, purely visual data from video frames was successfully converted into a structured, semantic, and machine-readable format. This effectively reverse-engineered the UI, enabling subsequent automated systems to understand and manipulate the interface in a meaningful way, resulting in greater robustness and versatility. Furthermore, regardless of changes in the application's underlying code, as long as the visual presentation remains consistent, stable recognition and interaction are possible.
[0047] Furthermore, through an automated exploration engine, navigation command signals are generated and sent to the OTT network TV client, and the subsequently returned baseband video signals are synchronously sampled to construct a complete interaction sequence; corresponding to step S3 above, the specific implementation process includes:
[0048] The exploration engine parses the screen content collection, identifies all elements categorized as "activatable" or "interactive," and adds them to an "exploration queue." It then selects a target element from the queue using a predefined breadth-first search. The engine generates a navigation instruction based on the selected target element. For example, for a button located at coordinates, the instruction might be "perform a click operation at coordinates." This instruction is sent to a control actuator, which either simulates the infrared and Bluetooth signals of a physical remote control or directly drives the app's interface elements at the software level using a UI automation framework. This allows direct interaction with UI controls, unaffected by changes in resolution, layout, or coordinates, and more accurately obtains the attributes of any UI element for result verification. The central analysis server synchronizes this operation with the video capture subsystem, recording the screen state before the operation, the operation itself, and the new state appearing on the screen after the operation. This group of operations constitutes a complete interaction. By linking these interactions chronologically, a complete interaction sequence is constructed.
[0049] The exploration engine fully automates the extremely time-consuming and labor-intensive process of manually surveying application UI maps. It can explore all possible user journeys in a detailed and repeatable manner, far exceeding the speed and breadth of human testers. The resulting large-scale UI interaction dataset is comprehensive and systematic in a way that is unmatched by manual testing.
[0050] Furthermore, a television service model is generated by iteratively generating service access points and their transition relationships; corresponding to step S4 above, the specific implementation process includes:
[0051] For each unique screen interface encountered during the exploration process, the central analysis server generates a unique "state signature." Specifically, all elements in the screen content set are deterministically sorted according to their screen positions. The sorted element type, text content, and coordinates are then concatenated into a unique, standardized string. This string itself serves as the signature for the service state. The central analysis server identifies each unique service state signature as a service access point. For each recorded interaction, the central analysis server creates a transition relationship between the service access point representing state A and the service access point representing state B. This transition relationship is weighted by the time spent transitioning from state A to state B during synchronous sampling. This transforms the vague subjective experience of "feeling fast" or "feeling lag" into an objective performance metric that can be precisely measured, compared, and tracked, making subsequent maintenance more convenient and the optimization process more intuitive. This process is repeated until the queue to be explored is empty, and all discovered service access points and their transition relationships are constructed into a complete and visualized television service model. This model is specifically a directed graph G=(V,E), where V is the set of service access points and E is the set of transition relationships.
[0052] This step transforms thousands of linear, independent sequences of interactions into a single, holistic, structured model. This enables the analysis of complex, non-linear user actions and reveals deep relationships between all screen interfaces within the service, allowing the entire team to collaborate based on the same model.
[0053] Furthermore, through in-depth analysis of the inherent structure of the television service model, the specific implementation process corresponding to step S5 above includes:
[0054] The central analytics server applies Dijkstra's algorithm to the model to calculate the minimum number of interaction steps required to reach key service access points from a specified entry point, such as the application homepage, to various critical service access points, such as a specific movie category, paid content, or a user subscription page. The Betweenness Centrality algorithm is then performed on all service access points in the graph, identifying those located on the shortest path between the largest number of other service access point pairs. Service access points with high centrality scores are key navigation hubs or potential flow bottlenecks in the application. For example, a search results page or main menu page typically has a high centrality score.
[0055] The Louvain algorithm runs on the model, dividing it into multiple service access point clusters. Connections between service access points within a cluster are much stronger than those between clusters. These automatically discovered communities often directly correspond to logically independent business domains within OTT services, such as "live TV," "video-on-demand," "kids' zone," or "account management." A central analytics server aggregates all text content extracted from each service access point in the graph via OCR. By performing frequency analysis on this text, content "hotspots" on the platform are identified—the programs, genres, actors, or services most frequently promoted throughout the service.
[0056] By analyzing the internal structure, abstract models are transformed into a set of objective, quantifiable performance metrics. Precise data replaces subjective evaluation. This enables operators to make data-driven decisions, accurately identify pain points in user experience, and gain a deep understanding of the actual functional architecture and layout of their applications.
[0057] Furthermore, the specific implementation process of performing time-domain difference comparison with historical cycle models includes:
[0058] Using the unique state signature of each service access point, a one-to-one service access point match is performed between the current model and the historical model. Service access points that exist in the current model but not in the historical model are marked as "new screens".
[0059] Service access points that exist in the historical model but not in the current model are marked as "deleted screen".
[0060] For service access point pairs that initially match successfully, the central analytics server compares their internal attributes and detects "content changes" by extracting text content via OCR, such as changing movie thumbnails on the same page layout. Based on the established service access point mappings, the central analytics server further compares the transition relationships between models. If a transition relationship exists between two matching service access points in the new model but not in the old model, this is identified as a new path. If a transition relationship exists in the old model but not in the new model, this is identified as a removal path. If the target service access point of a transition relationship originating from a matching service access point changes, this indicates a redirection path. The central analytics server quantifies these differences and compiles them into a time-series change log.
[0061] The comparison provides an automated, longitudinal audit trail of platform evolution. It replaces the tedious manual tracking of UI changes and provides a powerful tool for decision analysis. Operators can use this trail to correlate specific UI modifications with changes in key performance indicators.
[0062] Furthermore, the decision analysis agent generates strategies, compiles all analysis, comparison, and generation results, and outputs an evaluation report on content navigation efficiency and presentation strategies; corresponding to step S6 above, the specific implementation process includes:
[0063] A decision analytics agent is launched to integrate the outputs from deep analytics and temporal analysis. It correlates a decrease in navigation efficiency with a specific path redirection event. This agent is assigned specific business objectives and content orchestration guidelines, such as "maximizing user engagement with original content" or "shortening the purchase path for video-on-demand." Based on these objectives, it analyzes the fused data to identify the alignment and deviation between the current platform strategy and the established goals. Based on its analysis, the agent automatically generates specific, actionable content orchestration suggestions for the next operational cycle. These suggestions might include: "Promote 'new sci-fi series' to the homepage carousel to respond to current trends," or "Add a direct link to 'My Subscriptions' to the user avatar dropdown menu, reducing the navigation path from four clicks to two." A natural language generation module receives structured analytics data such as metrics, logs, and suggestions, and transforms it into coherent, human-readable narrative text. This process follows multiple stages: content planning, sentence aggregation, and grammatical structuring. Simultaneously, a visualization engine generates charts from the quantified data, such as trend graphs for navigation efficiency and pie charts for community size. Finally, the text and visual charts are compiled and formatted into a static report or interactive dashboard. This report and dashboard integrate in-depth analysis of the current operational status of the OTT platform and clear guidance for future strategies. On the dashboard, users can dynamically filter data and compare models across different time periods, even simulating path changes resulting from UI modifications. Compared to static files, this provides a more intuitive understanding of the current operational status and facilitates faster comprehension of decision-making recommendations.
[0064] This step fully automates the entire process from analysis to strategy, not only presenting complex analytical results to decision-makers in an easy-to-understand way but also proactively generating actionable recommendations, thus bridging the gap between raw data and business decisions. This significantly reduces the cognitive burden on operators and dramatically accelerates the iteration cycle of data-driven platform optimization.
[0065] The true value of this entire process lies in building a complete, automated decision-making loop for OTT operations: "observation" through video capture, "location" through modeling and analysis, "decision-making" through intelligent agents generating suggestions, and finally, "action" by the operator based on the report. This invention automates the first three stages, greatly improving the deployment agility of OTT service providers.
[0066] Example 2
[0067] In the process of converting visual signals into structured data, reference Figure 2 The specific implementation method is as follows:
[0068] The video acquisition device receives the signal, and its internal analog-to-digital converter and field-programmable gate array sample the video signal 60 times per second, converting each frame of analog signal into digital pixel data. This data is encoded in real time to generate a series of high-definition still video frames representing the instantaneous image on the terminal screen.
[0069] The central analysis server provides Grounding DINO with a set of text prompts: "a button," "an image," "a piece of text," "an icon," and "an input box." Grounding DINO simultaneously understands the meaning of this text and the content of a complete video frame. It then uses bounding boxes to draw a bounding box around the "Subscribe Now" button, another around the poster of the popular TV series, and yet another around the words "Action & Adventure." Each bounding box generated by Grounding DINO is used as a prompt and sequentially fed to SAM. For the rough bounding box of the "Subscribe Now" button, SAM precisely identifies the button's rounded corners and edges, generating a segmentation mask that only contains the button's pixels. The mask only segments the button itself from the original image. Similarly, it generates a precise rectangular mask for the poster of the popular TV series and a mask that tightly wraps around the text. Each segmented UI element is sent to MLLM for analysis. When MLLM sees the "Subscribe Now" image block, it analyzes its visual features (rectangle, colored fill, shadow) and internal text patterns, and combines this with its vast UI knowledge base to assign a semantic tag to each UI element. For example, a button is categorized as a button that triggers a subscription action; a poster image of a popular TV series is categorized as an image thumbnail; and the text "Action & Adventure" is categorized as a text tag. The "Subscribe Now" and "Action & Adventure" elements identified by MLLM as containing text are sent to the OCR engine to recognize the characters. The central analysis server outputs information for compilation. For the "Subscribe Now" button, a data record {id: 3, type: button, text: Subscribe Now, bounding_box: {x1: 450, y1: 820, x2: 630, y2: 880}, mask} is generated, becoming an entry in the screen content set.
[0070] To explore the entire application, the automated exploration engine first analyzes the screen content of the "Application Homepage," identifying all clickable elements such as the "TV Series Channel" button, the "Movie Channel" button, and the "Search" icon. These clickable elements are placed into an "Exploration Queue" ["TV Series Channel" button, "Movie Channel" button, "Search" icon], and a breadth-first search strategy is applied. The homepage is divided into layer 0, the TV Series Channel page, the Movie Channel page, and the Search page into layer 1, and all pages accessed from these layers into layer 2. Exploration begins from layer 0; all pages in layer 1 must be explored before any page in layer 2 can be explored. The exploration engine first retrieves the "TV Series Channel" button from the head of the "Exploration Queue." For Android-based boxes, a shell command sent via the Android Debug Bridge (ADB) is transmitted; for other systems, it might be a hexadecimal infrared code. Through a USB infrared transmitter connected to the analysis server, a precise infrared light signal, simulating a remote control, is emitted and sent to the OTT client. At the moment the navigation command is sent, the video capture device begins continuously capturing video frames output by the OTT client at 60fps. The central analysis server continuously compares adjacent frames. When it finds 30 consecutive frames with identical content, it determines that the screen has completed all rendering and loading and entered a stable state. The content of the last stable frame is taken as the final result of this navigation command execution. The central analysis server gathers all the key information of this interaction to form an interaction sequence {starting state signature, navigation command, target state signature, time}. After completing and storing this interaction sequence, the central analysis server returns to the first step, retrieves the next target "movie channel" button from the "exploration queue," and repeats the entire process, continuously generating interaction sequences.
[0071] Example 3
[0072] The specific implementation method during the business model construction process is as follows:
[0073] The automated exploration engine enters the application's homepage and waits for all animations and main content to load. Then, it obtains a structured "screen content set" using a multimodal visual model. Element A: {Type: Carousel, Content: Popular TV Series Promotion.jpg, Coordinates: [100, 100, 1820, 600]}, Element B: {Type: Icon Button, Content: Search, Coordinates: [50, 950, 150, 1050]}, Element C: {Type: Icon Button, Content: My Account, Coordinates: [1770, 950, 1870, 1050]}. The central analysis server sorts all elements in the screen content set according to a "vertical first, horizontal second" rule. Following the sorted order, it concatenates the key information (type, text content, coordinates) of each element into a long string, separating elements with semicolons. This string is the homepage's state signature in a specific state. The carousel itself is marked as a dynamic area to ensure the relative stability of the homepage's state signature.
[0074] This unique signature is defined as a service access point, labeled "Application Homepage" in the business model. The exploration engine interacts with each interactive element on the homepage, recording each complete "interaction sequence." The central analytics server, on the homepage, simulates clicking element B. After 420 milliseconds, the interface stabilizes on the "Search Page," creating a transition relationship from the homepage to the search page with the attributes {Homepage signature, Clicked element B, Search page signature, 420ms}. The central analytics server then returns to the homepage, clicks element C, and after 680 milliseconds reaches the "User Center Page," creating a transition relationship from the homepage to the user center with the attributes {Homepage signature, Clicked element C, User Center Page signature, 680ms}. Through this process, multiple transition relationships with different time weights radiate from the single service access point "Application Homepage," pointing to different functional pages, thus initially forming a star-shaped network structure centered on the "Application Homepage" SAP. The automated exploration engine continues to sequentially visit the "Search Page" and "User Center Page," generating state signatures on these pages to define new service access points. It then explores all links on these pages, creating new transition relationships. As the exploration process deepens, the original star-shaped structure centered on the homepage evolves into a complex, interwoven, and vast network. This network is the final television business model.
[0075] Example 4
[0076] In the process of analyzing and comparing television service models, reference was made to Figure 3 The specific implementation method is as follows:
[0077] Dijkstra's algorithm creates a distance table recording the distances from the homepage to all other service access points. Initially, the distance from the homepage to itself is set to 0, and the distances to all other service access points are set to infinity. Starting with the homepage, it examines all directly connected service access points, updating the distances to these points to 1. Next, from all service access points with known distances, the "Movie Channel" with the shortest distance and not yet fully processed is selected. Starting from the "Movie Channel," all its service access points are examined. If it connects to the VIP purchase page, the distance to the purchase page is updated to 1. The algorithm continuously repeats this process, expanding outwards layer by layer, constantly updating the distance table until the shortest distance to the target VIP purchase page is calculated, and then that number is output.
[0078] The Betweenness Centrality algorithm theoretically traverses all possible service access point pairs in the graph, calculates the shortest path for each pair, and finds all possible shortest paths between them. Then, the algorithm examines the search results page we are currently analyzing to see how many times it appears in these shortest paths, and records the sum of the probabilities of it appearing on the shortest paths between all service access point pairs as the centrality score.
[0079] Initially, each service access point belongs to its own independent community. The Louvain algorithm iterates through each service access point, attempting to move it from its current community to the community of its neighbors, calculating the gain in "modularity" that this action brings. Then, following a greedy approach, it moves the service access point to the neighboring community that provides the maximum gain in modularity. This process is repeated several times until moving any service access point can no longer improve the overall modularity. Each stable community formed in this stage is aggregated into a new "super service access point." Based on these super service access points, a completely new, coarser-grained network is built. Then, the movement is repeated again on this new network. The algorithm eventually outputs a stable community partitioning scheme, assigning all service access points to different clusters. These automatically discovered communities often closely match the actual logical partitioning of OTT services. For example, the algorithm might automatically identify: Community A (Video on Demand Domain): including the homepage, movie / TV series channels, details page, playback page, etc. Community B (Account Management Domain): including the "My" page, login page, order history, settings page, etc. Community C (Children's Zone): Contains all pages related to children's content.
[0080] The central analysis server iterates through every service access point in the business model, extracting all text from its screen content set using OCR, such as program names, actors, introductions, button text, and promotional text, compiling them into a massive text corpus. The corpus is preprocessed to remove meaningless stop words and segment sentences into individual words or phrases. Finally, the total frequency of each word or phrase in the entire corpus is calculated. A list of "hot words" is output, sorted by frequency from highest to lowest.
[0081] During the comparison process, the central analysis server traverses every service access point in the current model and attempts to find service access points with the exact same state signature in the historical model. If a service access point's state signature exists in the current model but not in the historical model, it will be marked as a "new screen." Conversely, if a service access point's state signature exists in the historical model but not in the current model, it will be marked as a "deleted screen." Because state signatures are highly sensitive to content, even if the overall layout of a page remains unchanged, its state signature will change if the promotional content changes. In the historical model, the movie channel page's state signature contains the text "Hot Recommendation: Popular Movie 1" extracted via OCR. In the current model, the movie channel page's status signature now contains the text "Hot Recommendation: Popular Movie 2" because the operations team changed the recommended movies to new releases this week. The central analysis server would initially determine that the movie channel page in the historical model is a "deleted screen," while the movie channel page in the current model is a "new screen." However, further analysis of the similarity between these two signatures reveals that they are highly similar. Therefore, in the final change log, the central analytics server will record it as a "content change": the old movie channel page is replaced by a new, updated version.
[0082] If a transition relationship exists between two matching service access points in the current model, but not in the historical model, this is identified as a "new path," meaning a new navigation entry or shortcut has been added. Conversely, if a transition relationship existed in the historical model, but the corresponding relationship has disappeared in the current model, it is identified as a "removed path." In the historical model, a transition relationship starting from the "My Account" page pointed to the old order history page. In the current model, a transition relationship starting from the same "My Account" page now points to a completely new member subscription management page. Comparing the results, the central analytics server determines this to be a "redirect path," precisely describing how the user's behavior hasn't changed, but the result of that behavior has fundamentally changed.
[0083] The central analytics server first performs quantitative statistics on all changes, forming a high-level summary. In addition to macro-level statistics, the central analytics server also generates a detailed, item-by-item, time-series change log. Each record includes the type of change, the signature of the service access point or transition relationship involved, and a possible contextual description.
[0084] Example 5
[0085] The specific implementation method for agent decision-making and report generation is as follows:
[0086] After analyzing and comparing all the data from the TV business model, the agent discovered that the navigation depth of the VIP purchase page had worsened from 3 steps to 5 steps this week. In the change log, the agent found a record describing a button on the "Movie Channel" page that originally pointed to the VIP purchase page but was now redirecting to a generic "Member Activity Center" page. The agent correlated these two independent findings, forming a strong causal hypothesis: "The decrease in navigation efficiency is highly likely caused by a specific path redirection event."
[0087] When the agent is instantiated, it is given specific business objectives: "shortening the purchase path for video-on-demand" and content arrangement principles: "maximizing user engagement with original content." The agent compares the obtained "causal relationships" with these preset principles, comparing the inferred fact that the VIP purchase page path is longer with the business objectives. A serious deviation is identified: "The current platform strategy runs counter to the established business objectives." In response to this deviation, the agent automatically generates a structured suggestion: "It is recommended to restore the direct navigation link to the 'VIP purchase page' on the 'Movie Channel page,' with the goal of shortening the navigation path from the current 5 clicks to an optimized 3 or less." Text analysis shows that "new sci-fi dramas" is a hot topic across the platform, but navigation analysis shows that its entry point is buried deep. This aligns with the content arrangement principles and represents an opportunity to amplify. Therefore, the agent generates a suggestion: "It is recommended to elevate the 'new sci-fi dramas' topic entry point to the carousel or prime position on the 'application homepage' to respond to current hot trends and maximize user engagement."
[0088] A natural language generation module receives all metrics, logs, and suggestions from the preceding stages, generating data in the following order: summarizing key changes, analyzing specific impacts, and finally providing optimization recommendations. Simultaneously, a visualization engine utilizes a chart library to generate trend charts from historical navigation efficiency data, visually displaying the changing curves of VIP purchase path depth; it also generates pie charts of community size based on the Louvain algorithm's community segmentation results and word clouds from the hot keyword list. All generated narrative text and visualizations are then fed into a layout engine. Following a report template, the engine professionally lays out and formats the content, adding brand elements such as the company logo.
[0089] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. An OTT content operation optimization method based on multimodal perception and agent decision-making, characterized in that, include: The baseband video signal output by the OTT network TV client is sampled and digitized, and the baseband video signal is used to present the visual content of the interactive program guide; The two-dimensional pixel information in the baseband video signal is parsed into a set of screen content containing business information and content metadata through a multimodal model. Through the automated exploration engine, navigation command signals are generated and sent to the OTT network TV client. The interface elements of the App are driven directly at the software level through infrared and Bluetooth signals, and the baseband video signals returned subsequently are sampled synchronously to construct a complete interaction sequence. Multiple interaction sequences obtained through iterative exploration are associated and aggregated to generate a state signature, which is then used to identify the interactive program guide as a service access point. When navigation instructions trigger changes in the presented content from the starting service access point to the target service access point, the directional transition relationship between the two is established and recorded. Through iterative exploration and analysis, a television service model of the service access point and its transition relationship is generated. The internal structure of the TV business model is analyzed in depth. The shortest path algorithm is applied to calculate the minimum number of interaction steps from the entry service access point to the key TV content assets and service access points, and to evaluate the discoverability and navigation efficiency of paid content and value-added services. A centrality metric is used to score and rank the importance of service access points, identifying core content aggregation pages and navigation hubs; a community discovery algorithm is run to divide the TV business model into different business logic domains. And perform time-domain difference comparison with historical television service models; The decision analysis agent generates the strategy, compiles all the analysis, comparison and generation results, and outputs an evaluation report of the model data and strategy.
2. The OTT content operation optimization method based on multimodal perception and agent decision-making according to claim 1, characterized in that, The process of sampling and digitizing the baseband video signal of the OTT network TV client includes: establishing a physical communication link between the OTT network TV client and an external video acquisition device through a high-definition multimedia interface; the OTT network TV client generating a baseband video signal, which is then digitized by the video acquisition device to generate static video frames in real time that represent the interactive program guide screen displayed on the terminal.
3. The OTT content operation optimization method based on multimodal perception and agent decision-making according to claim 1, characterized in that, The specific implementation process of parsing the baseband video signal into a set of screen content using a multimodal model includes: inputting a static video frame into the multimodal model; recovering and extracting image components encoded as visual elements during client rendering from the pixel domain of the television screen; identifying and classifying graphic elements carrying service information in the screen; calculating and outputting the position information of the identified graphic elements in the screen coordinate system; performing optical character recognition on text elements to extract string content; and compiling the classification, string content, and position information of all elements into a set of screen content describing the content presented by the static video frame.
4. The OTT content operation optimization method based on multimodal perception and agent decision-making according to claim 1, characterized in that, The specific implementation process of constructing a complete interaction sequence through the automated exploration engine includes: the automated exploration engine performs an exploration process, during which it parses the set of screen content, identifies the graphic elements that switch pages, adds them to the queue to be explored, and generates interaction commands based on a predefined TV service navigation strategy. These commands are then sent to the control actuator to simulate the navigation commands issued by the physical remote control device to change the state of the TV client. The process is synchronized with subsequent video frame sampling to complete the systematic data acquisition and construct a complete interaction sequence.
5. The OTT content operation optimization method based on multimodal perception and agent decision-making according to claim 1, characterized in that, The specific implementation process of performing temporal differential comparison between the television service model and multiple historical models includes: selecting the television service model of the current period and the television service models of multiple historical periods as input for differential analysis; executing the service access point correlation comparison process, establishing a one-to-one correspondence between service access points between the current and historical models based on each state signature; comparing each matched service access point pair one by one, identifying and recording changes in internal attributes; identifying unmatched service access points as newly added or deleted; based on the established service access point correspondence, identifying and recording navigation paths added, deleted, and redirected due to changes in television service interaction logic; calculating and quantifying the differences in service menu structure, content entry location, and channel arrangement order in the two maps, and evaluating the periodic adjustment of content arrangement strategies; and compiling all changes into a temporal content and navigation change record.
6. The OTT content operation optimization method based on multimodal perception and agent decision-making according to claim 1, characterized in that, The specific implementation process of outputting a model data and strategy evaluation report includes: deeply integrating various TV program performance indicators obtained from the deep analysis with content and navigation change records obtained from the time-domain difference comparison to form an aggregated record representing the platform's operational dynamics; activating a content strategy analysis agent assigned preset market goals and content arrangement guidelines to comprehensively evaluate the aggregated record and identify the fit and differences between the current platform content strategy and market operation hotspots; automatically generating content arrangement suggestions for the next cycle of content operation based on the fit and differences, including recommendations for key content types, optimization of promotional resource positions, and responses to emerging trends; logically assembling the deep analysis, time-domain comparison, and content arrangement suggestions into an explanatory text description; calling a graphical engine to generate trend charts of changes in content distribution efficiency performance indicators; and outputting the completed content layout as a static report file, which integrates content operation status analysis and future strategy guidance.
Citation Information
Patent Citations
Automatic UI (User Interface) interactive exploration method based on multi-modal large model
CN118467032A
Personalized operation strategy generation method and system based on data processing
CN119939118A