Remote diagnosis method for internet television and internet television system

CN122802671APending Publication Date: 2026-09-22CHINA MOBILE GRP GANSU CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611154821.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-31
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

[0004]本发明提供了一种互联网电视的远程诊断方法及互联网电视系统,以解决相关技术仅能通过平台下发任务或用户报修人工启动诊断等被动式诊断,无法对机顶盒状态实现持续秒级检测与故障即时感知的问题

Benefits of technology

[0011]本发明实施例的技术方案,首先,通过所述远程诊断平台基于所述连接获得所述机顶盒的设备在线状态,并基于所述设备在线状态生成检测任务,并将所述检测任务下发至所述软探针;以设备在线状态触发自动生成检测任务,实现了诊断流程的自动化与按需启动,避免了盲目下发指令带来的资源浪费,确保了诊断工作的精准性和高效性;接着,通过所述软探针按所述检测任务持续采集秒级检测指标并回传至所述远程诊断平台,所述远程诊断平台根据所述秒级检测指标生成故障信息,并根据所述故障信息装配故障采集任务下发至所述软探针;通过持续采集秒级检测指标,能够捕捉到设备运行过程中的微小波动和瞬时异常,极大提升了故障发现的灵敏度和时效性;接着,通过所述软探针根据所述故障采集任务进行数据采集以生成诊断数据集,并将所述诊断数据集上传至所述远程诊断平台,所述诊断数据集包括所述机顶盒的业务流量的数据包文件、设备诊断结果和所述机顶盒的实时画面流;在确认故障后,自动装配并下发针对性的数据采集任务,且采集内容涵盖多维度信息构建立体化的诊断数据集,使得故障分析更加全面、直观且具备可追溯性;最后,通过所述远程诊断平台根据所述诊断数据集提取关键路径内容,并基于所述关键路径内容判定故障判定结果,根据所述故障判定结果下发处置任务至所述软探针,并根据所述软探针对所述处置任务的执行结果更新任务状态;实现了从数据提取、故障判定到处置任务下发及状态更新的完整闭环,不仅提高了故障处理的自动化水平,还通过执行结果的反馈机制确保了任务状态的实时准确,有效缩短了故障恢复时间,保证了故障发现的及时性与诊断的全面性,实现了处置流程的自动化与可追踪,显著提升了故障检测和处置的效率与智能化水平。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122802671A_ABST
    Figure CN122802671A_ABST
Patent Text Reader

Abstract

The application discloses a remote diagnosis method of an Internet TV and an Internet TV system, and the Internet TV system comprises a set top box and a remote diagnosis platform, a soft probe is arranged in the set top box, the set top box and the remote diagnosis platform are connected through the soft probe to establish a message queue telemetry transfer protocol connection, and the method comprises the following steps: the remote diagnosis platform obtains the equipment online state of the set top box based on the connection and generates a detection task to be sent to the soft probe; the soft probe collects second-level detection indexes according to the detection task and returns the second-level detection indexes, the remote diagnosis platform generates fault information according to the second-level detection indexes, assembles a fault collection task and sends the fault collection task to the soft probe; the soft probe generates a diagnosis data set according to the fault collection task and uploads the diagnosis data set to the remote diagnosis platform; the remote diagnosis platform extracts key path content according to the diagnosis data set, judges a fault judgment result, sends a disposal task to the soft probe, and updates a task state according to an execution result of the soft probe on the disposal task; and the efficiency and intelligent level of fault detection and disposal are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of personal and home service technology, and in particular to a remote diagnostic method and system for Internet television. Background Technology

[0002] Internet TVs have fully penetrated modern family life. As the core smart audio-visual entertainment terminal device in the home, its operational stability directly affects the audio-visual experience of end users.

[0003] Current technologies for diagnosing internet TVs largely rely on the platform proactively issuing diagnostic tasks or users manually initiating the diagnostic process after reporting a fault. This passive approach cannot provide continuous, second-level monitoring of the set-top box's operational status. Consequently, faults are often only discovered some time after they occur, or even after a user complaint, hindering immediate fault detection and rapid response. Therefore, a remote diagnostic method for internet TVs is urgently needed to improve the efficiency of fault diagnosis and handling. Summary of the Invention

[0004] This invention provides a remote diagnostic method and system for Internet TVs, which solves the problem that related technologies can only perform passive diagnostics through platform-issued tasks or user-reported repairs, and cannot achieve continuous, second-level detection of set-top box status and real-time fault perception.

[0005] According to one aspect of the present invention, a remote diagnostic method for Internet television is provided, applied to an Internet television system, the Internet television system including a set-top box and a remote diagnostic platform, wherein a soft probe is deployed in the set-top box, and the set-top box and the remote diagnostic platform establish a message queue telemetry transmission protocol connection through the soft probe, wherein the method includes: The remote diagnostic platform obtains the online status of the set-top box based on the connection, generates a detection task based on the online status of the device, and sends the detection task to the soft probe. The soft probe continuously collects second-level detection indicators according to the detection task and sends them back to the remote diagnostic platform. The remote diagnostic platform generates fault information based on the second-level detection indicators and assembles a fault collection task based on the fault information and sends it to the soft probe. The soft probe collects data according to the fault collection task to generate a diagnostic dataset, and uploads the diagnostic dataset to the remote diagnostic platform. The diagnostic dataset includes data packet files of the set-top box's service traffic, device diagnostic results, and the set-top box's real-time video stream. The remote diagnostic platform extracts critical path content from the diagnostic dataset, determines the fault determination result based on the critical path content, issues a handling task to the soft probe based on the fault determination result, and updates the task status based on the execution result of the handling task by the soft probe.

[0006] According to another aspect of the present invention, an Internet television system is provided, including a set-top box and a remote diagnostic platform. The set-top box is equipped with a soft probe, and the set-top box and the remote diagnostic platform establish a message queue telemetry transmission protocol connection via the soft probe. The remote diagnostic platform obtains the online status of the set-top box based on the connection, generates a detection task based on the online status of the device, and sends the detection task to the soft probe. The soft probe continuously collects second-level detection indicators according to the detection task and sends them back to the remote diagnostic platform. The remote diagnostic platform generates fault information based on the second-level detection indicators and assembles a fault collection task based on the fault information and sends it to the soft probe. The soft probe collects data according to the fault collection task to generate a diagnostic dataset, and uploads the diagnostic dataset to the remote diagnostic platform. The diagnostic dataset includes data packet files of the set-top box's service traffic, device diagnostic results, and the set-top box's real-time video stream. The remote diagnostic platform extracts critical path content from the diagnostic dataset, determines the fault determination result based on the critical path content, issues a handling task to the soft probe based on the fault determination result, and updates the task status based on the execution result of the handling task by the soft probe.

[0007] According to another aspect of the present invention, a remote diagnostic device for Internet television is provided, applied to an Internet television system, the Internet television system including a set-top box and a remote diagnostic platform, wherein a soft probe is deployed in the set-top box, and the set-top box and the remote diagnostic platform establish a message queue telemetry transmission protocol connection through the soft probe, wherein the device includes: The detection task generation module is used by the remote diagnostic platform to obtain the device online status of the set-top box based on the connection, generate a detection task based on the device online status, and send the detection task to the soft probe. The fault information generation module is used for the soft probe to continuously collect second-level detection indicators according to the detection task and send them back to the remote diagnostic platform. The remote diagnostic platform generates fault information based on the second-level detection indicators and assembles a fault collection task based on the fault information and sends it to the soft probe. A diagnostic dataset generation module is used by the soft probe to collect data according to the fault collection task to generate a diagnostic dataset, and upload the diagnostic dataset to the remote diagnostic platform. The diagnostic dataset includes data packet files of the set-top box's service traffic, device diagnostic results, and the set-top box's real-time video stream. The fault determination module is used by the remote diagnostic platform to extract critical path content based on the diagnostic dataset, determine the fault determination result based on the critical path content, issue a handling task to the soft probe based on the fault determination result, and update the task status based on the execution result of the handling task by the soft probe.

[0008] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the remote diagnostic method for Internet television according to any embodiment of the present invention.

[0009] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the remote diagnostic method for Internet television according to any embodiment of the present invention.

[0010] According to another aspect of the present invention, embodiments of this disclosure also provide a computer program product, including a computer program that, when executed by a processor, implements the remote diagnostic method for Internet television as described in any of the embodiments of this disclosure.

[0011] The technical solution of this invention embodiment firstly involves obtaining the online status of the set-top box through the remote diagnostic platform based on the connection, generating a detection task based on the online status, and sending the detection task to the soft probe. The automatic generation of the detection task triggered by the online status of the device automates the diagnostic process and enables on-demand startup, avoiding resource waste caused by blindly issuing instructions and ensuring the accuracy and efficiency of the diagnostic work. Next, the soft probe continuously collects second-level detection indicators according to the detection task and transmits them back to the remote diagnostic platform. The remote diagnostic platform generates fault information based on the second-level detection indicators and assembles a fault acquisition task based on the fault information, sending it to the soft probe. Continuous collection of second-level detection indicators can capture minute fluctuations and instantaneous anomalies during device operation, greatly improving the sensitivity and timeliness of fault detection. Finally, the soft probe collects data according to the fault acquisition task to generate a diagnostic dataset, and uploads the diagnostic dataset to the remote diagnostic platform. The diagnostic dataset includes data packet files of the set-top box's service traffic, device diagnostic results, and the set-top box's real-time video stream. After confirming a fault, a targeted data collection task is automatically assembled and issued, and the collected content covers multi-dimensional information to construct a three-dimensional diagnostic dataset, making fault analysis more comprehensive, intuitive, and traceable. Finally, the remote diagnostic platform extracts key path content from the diagnostic dataset, determines the fault judgment result based on the key path content, issues a handling task to the soft probe based on the fault judgment result, and updates the task status based on the soft probe's execution result of the handling task. This achieves a complete closed loop from data extraction and fault judgment to handling task issuance and status update, which not only improves the automation level of fault handling but also ensures the real-time accuracy of task status through the execution result feedback mechanism, effectively shortens fault recovery time, ensures the timeliness of fault detection and the comprehensiveness of diagnosis, realizes the automation and traceability of the handling process, and significantly improves the efficiency and intelligence level of fault detection and handling.

[0012] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0013] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0014] Figure 1 This is a flowchart of a remote diagnostic method for Internet TV provided according to Embodiment 1 of the present invention; Figure 2 This is a flowchart of a remote diagnostic method for Internet TV provided according to Embodiment 2 of the present invention; Figure 3 This is a schematic diagram of the structure of an Internet television system according to Embodiment 3 of the present invention; Figure 4 This is a schematic diagram of the structure of a remote diagnostic device for Internet TV provided in Embodiment 4 of the present invention; Figure 5 This is a schematic diagram of the structure of an electronic device that implements the remote diagnostic method for Internet TV according to an embodiment of the present invention. Detailed Implementation

[0015] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0016] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0017] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0018] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0019] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.

[0020] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.

[0021] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0022] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.

[0023] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.

[0024] Example 1 Figure 1 This is a flowchart of a remote diagnostic method for an Internet TV provided in Embodiment 1 of the present invention. Applied to an Internet TV system, the system includes a set-top box and a remote diagnostic platform. The set-top box is equipped with a soft probe. The set-top box and the remote diagnostic platform establish a message queue telemetry transmission protocol connection through the soft probe. This embodiment is applicable to remote diagnostics of Internet TVs. The method can be executed by a remote diagnostic device for the Internet TV, which can be implemented in hardware and / or software. Optionally, it can be implemented through an electronic device, such as a mobile terminal, a PC, or a server. Figure 1 As shown, the method may specifically include: S110. The remote diagnostic platform obtains the online status of the set-top box based on the connection, generates a detection task based on the online status of the device, and sends the detection task to the soft probe.

[0025] In this embodiment of the invention, the Internet TV system can be a TV system capable of remote diagnosis. The Internet TV system may include a set-top box and a remote diagnostic platform. The remote diagnostic platform can be a platform-side system that receives data reported from the set-top box, issues detection tasks, and processes diagnostic results. The set-top box (STB) can be a terminal carrier used to carry Internet TV service processes, network interfaces, storage directories, and local running permissions. The software probe (SP) can be a terminal-side running component deployed in the set-top box, used to perform second-level detection indicator collection, fault detection, data acquisition operations, real-time video stream transmission from the set-top box, and subsequent task reception. The Message Queuing Telemetry Transport (MQTT) protocol can be used to establish a long-term connection channel between the set-top box-side software probe and the remote diagnostic platform. The device online status can be understood as the set-top box terminal's network connectivity and communication capability status determined by the remote diagnostic platform based on the communication link and heartbeat signal. The device online status may include current online status, status update time, connection maintenance status, allowed distribution flags, and terminal range information corresponding to the device identification registration result. The detection task can be a set of tasks generated by the remote diagnostic platform for set-top boxes and soft probes. The detection task can include information such as task number, device identifier, collection item configuration and reporting rhythm configuration.

[0026] Before obtaining the device's online status, soft probe deployment, MQTT protocol connection establishment, and device identification registration are required to obtain real-time connection information. Then, based on this real-time connection information, heartbeat detection, device online status checks, and connection maintenance are performed to generate the device's online status. Specifically, heartbeat detection continuously confirms whether the long connection between the soft probe and the remote diagnostic platform remains valid; device online status checks determine whether the set-top box is currently online based on heartbeat detection results, session response records, and reporting interval records; and connection maintenance performs operations such as reconnection, resubscription, and status correction in case of connection interruption, topic anomalies, or reporting delays.

[0027] Optionally, the software probe's running directory, startup directory, and data cache directory can be established in the set-top box, and the software probe's process name, startup method, and reconnection method can be registered in the set-top box's process table. The software probe deployment can include multiple steps such as installation package verification, version number reading, running permission binding, network access permission binding, and automatic startup registration. Installation package verification can be used to verify the integrity of the software probe's installation package. Version number reading can be used to record the current software version of the software probe. Running permission binding can be used to enable the software probe to access the second-level detection indicator acquisition interface, log directory, data acquisition call interface, and video stream return interface. Network access permission binding can be used to enable the software probe to connect to the access address of the remote diagnostic platform. Automatic startup registration can be used to restart the software probe after the set-top box restarts.

[0028] For set-top boxes that have already deployed software probes, the deployment status can be verified. The verification may include, but is not limited to, whether the software probe process exists, whether the version number matches, whether the running directory is complete, and whether the last exit record is abnormal. If a version mismatch or a corrupted running directory is found, an overwrite installation operation can be performed. If the process does not exist and the installation status is normal, a restart operation can be performed. If an abnormal exit record is found, the abnormal exit record can be written to the local log and the subsequent connection establishment process can continue.

[0029] After the soft probe is deployed, an MQTT protocol connection establishment operation can be performed. Specifically, the soft probe can read the access address, access port, topic name, authentication information, and reconnection parameters of the remote diagnostic platform, generate a connection request message, and send a connection request to the remote diagnostic platform. After receiving the connection request, the remote diagnostic platform verifies the authentication information, device source, and access topic, and sends back a connection confirmation result. The connection request message may contain multiple pieces of information such as device identifier, soft probe version number, set-top box model, network address, and current time record. The connection confirmation result may contain multiple pieces of information such as connection status, platform-side session identifier, and subsequent heartbeat detection parameters. The topic name of the remote diagnostic platform is used to distinguish between connection topics, reporting topics, and task topics. Connection topics can be used for initial access, task topics can be used for subsequent detection task distribution, and reporting topics can be used for subsequent second-level data return, collection result upload, and execution result return.

[0030] If the initial MQTT connection establishment fails, the software probe performs an interval reconnection operation based on the reconnection parameters. If multiple connection failures occur consecutively, the reason for failure, the time of failure, and the number of failures are written locally, and the set-top box is placed in a pending access state. The connection is then initiated again once the network is restored. It is important to emphasize that the MQTT connection establishment in this application is not simply a network connection establishment, but rather a simultaneous confirmation of the access relationship between the set-top box, the software probe, and the remote diagnostic platform. This ensures that subsequent online device status checks, detection task distribution, and second-level detection indicator reporting are all conducted on the same real-time connection link.

[0031] Based on the above scheme, optionally, the remote diagnostic platform obtains the set-top box's online status based on the connection, including: after establishing a message queue telemetry transmission protocol connection between the set-top box and the remote diagnostic platform based on the soft probe, sending a device identification registration message of the set-top box to the remote diagnostic platform, the device identification registration message carrying the set-top box's basic device fields, the soft probe version number, and the network address; after completing the device identification registration based on the device identification registration message, the remote diagnostic platform sends the set-top box's registration result back to the soft probe; the soft probe organizes the connection status, platform session identifier, and the registration result into real-time connection information, and periodically sends heartbeat signals to the remote diagnostic platform based on the real-time connection information, so that the remote diagnostic platform generates the set-top box's online status based on the reception status of the heartbeat signals.

[0032] The device identification registration message can be the device registration message sent to the remote diagnostic platform after the set-top box soft probe establishes its first protocol connection. The device basic fields can be understood as the set-top box's basic identity information fields. The soft probe version number can be the version code of the soft probe software deployed on the set-top box, used by the platform to adapt to the corresponding version's task distribution and data parsing rules. The network address can be the network identifier currently connected to by the set-top box. The registration result can be the feedback information sent back to the soft probe by the platform after completing device registration. The real-time connection information can be the current communication link status information compiled by the soft probe. The platform session identifier can be the unique session code for a single communication connection between the platform and the set-top box, used to bind all tasks, data, and interaction records of this connection. The heartbeat signal can be a status probe message sent periodically by the soft probe to the platform, used to proactively inform the platform that the terminal is online and the link is normal. The heartbeat signal reception status can be the platform's reception result of the heartbeat signal.

[0033] Optionally, after establishing a message queue telemetry transmission protocol connection between the set-top box and the remote diagnostic platform based on the soft probe, a device identification registration process can be performed. The device identification can be the registration content used by the remote diagnostic platform to distinguish different set-top boxes and their soft probes. The device identification may include various identification information such as set-top box serial number, soft probe instance number, set-top box model, network address, version number, terminal range, and first access time.

[0034] Specifically, the software probe can read the device's basic fields from the set-top box system information, the software probe version number and installation time from the local deployment information, and the network address from the network interface information. These fields are then assembled into a device identification registration message and sent to the remote diagnostic platform. After receiving the device identification registration message, the remote diagnostic platform can compare the device's basic fields with the existing registration content on the platform side. If the same device already has old registration content, the corresponding version number, network address, and last access time are updated. When a new device appears for the first time, a new device registration entry is generated. If the device changes its network address, upgrades the software probe version number, or is redeployed after a factory reset, the old registration record is retained and a currently valid registration record is generated. This ensures that subsequent execution result feedback and cross-step tracing can all be mapped to the same set-top box. After completing the device identification registration, the set-top box's registration result can be sent back to the software probe.

[0035] Furthermore, information such as connection status, platform session identifier, device identifier registration result, and platform-side confirmation time can be compiled into real-time connection information using a soft probe, and this real-time connection information can be written into a local cache and a remote diagnostic platform session table. The real-time connection information includes at least several pieces of information such as connection status, platform-side session identifier, device identifier registration result, access topic, and reconnection parameters.

[0036] Specifically, an online device session table can be established through the remote diagnostic platform according to the platform-side session identifier in the real-time connection information. The soft probe enters a working state of periodically sending heartbeats, i.e., heartbeat signals, according to the access topic and reporting topic. The heartbeat signal may contain information such as device identifier, current connection status, the time of the last data transmission, and the current soft probe process status. After receiving the heartbeat signal, the remote diagnostic platform compares the received time with the time of the previous cycle and registers three types of records: normal heartbeat, delayed heartbeat, or missing heartbeat. Among them, a normal heartbeat indicates that the connection link and the soft probe process are both active, a delayed heartbeat indicates that the link may fluctuate or the soft probe's operating pressure is increasing, and a missing heartbeat indicates that connection maintenance processing is required.

[0037] Optionally, an online status check can be performed after heartbeat detection. Specifically, the remote diagnostic platform does not directly determine the online status based on a single heartbeat result. Instead, it generates the device's online status by combining multiple recent heartbeat cycles, the most recent reporting time, and session response records. The device's online status includes at least three status items: online, fluctuating online, and offline. Specifically, the online status corresponds to continuous heartbeat arrivals and normal session responses; the fluctuating online status corresponds to heartbeat delays, short-term reconnections, or short-term reporting interruptions, but the platform-side session identifier remains valid; and the offline status corresponds to multiple consecutive missing heartbeats, disappearance of session responses, or active cancellation of the soft probe.

[0038] To ensure stable operation of online device status checks, a soft probe can read the local process table, network interface status, and the most recent transmission result before each heartbeat transmission. If the local process table is normal but the network interface is abnormal, a network abnormality flag is written into the heartbeat content; if the network interface is normal but the previous transmission failed, a retransmission flag is written. At the same time, when the remote diagnostic platform parses the heartbeat content, it can write the network abnormality flag and the retransmission flag into the session table to distinguish between two scenarios: network interruption and abnormal exit of the soft probe. This prevents subsequent detection tasks from directly classifying fluctuating online devices as offline devices.

[0039] Based on the above scheme, optionally, the device online status includes fluctuating online status; the method further includes: when the soft probe detects a connection confirmation timeout, topic subscription failure, or heartbeat transmission failure, the soft probe performs a local connection maintenance operation. If no confirmation is received from the remote diagnostic platform after the local connection maintenance operation is completed, the soft probe enters an interval reconnection according to the reconnection parameters. The local connection maintenance operation includes resending the connection request, resubscribing to the task topic, and restoring the reporting topic; when the device online status is fluctuating online status, the remote diagnostic platform retains the platform session identifier and marks the last valid time, waits for the remote diagnostic platform to re-establish a connection with the soft probe, and then re-issues the detection task that was not completed due to the connection interruption.

[0040] Fluctuating online status can be understood as the set-top box terminal being connected to the network, but the communication link is unstable, resulting in alternating periods of disconnection and reconnection, which is not considered completely offline. Connection confirmation timeout can be understood as the state where, after the soft probe initiates a connection request, it does not receive a connection confirmation response from the platform within a preset time. Topic subscription failure can be understood as the state where the platform task topics or data reporting topics subscribed to by the soft probe in the MQTT protocol become invalid or unbound, making it unable to receive platform messages. Heartbeat sending failure can be understood as the state where, after the soft probe triggers the heartbeat sending logic, the heartbeat message fails to be sent, is lost, or is rejected by the platform. Local connection persistence operation can be a keep-alive operation performed locally by the soft probe when the link is abnormal, such as resending connection requests, resubscribing to task topics, or restoring reporting topics. Reconnection parameters can be the soft probe's preset reconnection configuration parameters. Last valid time can be the time node recorded by the platform before the set-top box link is interrupted, the last time a normal communication and heartbeat was valid.

[0041] Optionally, connection maintenance processing can be performed after the device's online status check. Specifically, when the soft probe detects a connection confirmation timeout, topic subscription failure, or heartbeat transmission failure, a local connection maintenance operation can be performed first. The local connection maintenance operation may include, but is not limited to, resending the connection request, resubscribing to the task topic, and restoring the reporting topic. If no confirmation is received from the remote diagnostic platform after the local connection maintenance operation is completed, an interval reconnection will be initiated according to the reconnection parameters.

[0042] When the device's online status is fluctuating, the remote diagnostic platform can perform a connection maintenance operation on the platform side for fluctuating online or offline devices. The platform side connection maintenance operation may include, but is not limited to, retaining the original platform side session identifier, marking the last valid time, and resetting the task distribution status, thereby avoiding the repeated generation of multiple test tasks during short-term fluctuations. After the remote diagnostic platform and the soft probe re-establish a connection, the test tasks that were not completed due to the connection interruption will be re-distributed.

[0043] Furthermore, information such as each heartbeat detection result, device online status change record, connection hold trigger reason, and connection recovery result can be written into the device online status log. The device online status log can be used to determine whether the detection task should be issued, whether the second-level data return frequency should be reduced, and to trace the connection link status when the execution result return fails.

[0044] After completing the above processing, the current online status, status update time, connection hold status, and allow-to-issue flag of the device can be organized into the device online status through the remote diagnostic platform. Among them, the online, fluctuating online, and offline records in the device online status can also affect the way detection tasks are received. For example, online devices directly enter the second-level detection indicator configuration, fluctuating online devices re-receive detection tasks according to the recovered session, and offline devices retain the pending recovery access flag and are processed after the real-time connection information is re-established.

[0045] By employing sophisticated heartbeat detection and connection maintenance processing, the online status of devices is categorized into online, fluctuating online, and offline. Differentiated detection task management strategies are then implemented based on these different statuses, ensuring reliable execution of detection tasks and accurate recording of statuses. This enhances the reliability and continuity of detection tasks in complex network environments.

[0046] Optionally, based on the above scheme, generating a detection task based on the online status of the device includes: reading the allow-to-distribute flag in the online status of the device; when the allow-to-distribute flag is allowed, the remote diagnostic platform generates a detection task and distributes it to the soft probe via a task topic through a message queue telemetry transmission protocol; when the allow-to-distribute flag is pending recovery, the remote diagnostic platform generates a detection task but does not send it immediately, and distributes it after the soft probe reconnects; when the allow-to-distribute flag is prohibited, the remote diagnostic platform registers the detection task as pending distribution without generating task content and does not perform the distribution operation.

[0047] The "Allow Distribution" flag can be a task distribution permission identifier configured by the remote diagnostic platform for each set-top box, used to control the generation and distribution logic of detection tasks. The "Allow" state can be understood as the normal state of the distribution flag, indicating that the terminal link is normal and can directly receive and execute detection tasks. The "Pending Recovery" state can be understood as the abnormal transitional state of the distribution flag, indicating that the terminal link is fluctuating, the task is temporarily stored and not immediately distributed, waiting to be resent after reconnection. The "Prohibited" state can be understood as the locked state of the distribution flag, indicating that the terminal is abnormally offline or malfunctioning. The "Pending Distribution" state can be understood as the storage state of the platform's task ledger, only registering task information without performing distribution operations. The task topic can be understood as the message topic channel in the MQTT protocol used by the platform to distribute various tasks to the terminal.

[0048] Specifically, the remote diagnostic platform can read the "allow delivery" flag in the device's online status to generate formal testing tasks for online devices, generate recovery testing tasks for fluctuating online devices, and register offline devices as pending tasks without sending them immediately. The testing task generation process is not simply writing the task name; instead, it configures specific task content based on information such as the device identification registration result, set-top box model, soft probe version number, and current device online status.

[0049] The detection task includes at least the following configuration items: second-level detection metrics, detection task status, terminal range, reporting rhythm, fault detection items, and reserved items for subsequent automatic data collection conditions. Second-level detection metrics may include multiple indicators such as player running status, network reception status, screen display status, storage usage status, and process activity status. The detection task status can be understood as the current state of the detection task between waiting to receive, received, executing, and abnormal interruption. The detection task status may include, but is not limited to, multiple status types such as waiting to be sent, sent, received, executing, and abnormal interruption. The terminal range can be understood as the single device range or group device range corresponding to the current detection task. The terminal range may include the single device range and the group device range formed by platform registration.

[0050] For a single device, the detection task is directly bound to the current device identifier; for a group of devices, the remote diagnostic platform generates batch detection tasks for the same group when the devices are online in the same state, and attaches an independent task instance to each device, so that when the detection task is obtained later, it can read the unified configuration and the independent receiving record of each device.

[0051] Furthermore, after a testing task is generated, a testing task status registration process can be performed. Specifically, the remote diagnostic platform can generate a corresponding task registration record for each testing task. This record can include information such as task number, device identifier, testing task status, generation time, distribution time, receipt status, and last modification time. The testing task status registration is not statically written but continuously updated throughout the entire process.

[0052] For example, a detection task can be marked as pending when it is first generated, as sent to the task topic, as sent, as received after the soft probe returns a confirmation, as being executed once the soft probe starts collecting second-level detection metrics, and as an abnormal interruption if the device's online status fluctuates and the task delivery is interrupted. For devices with fluctuating online records during the previous connection maintenance process, the remote diagnostic platform can add a recovery source marker to the task registration record, enabling subsequent second-level detection metric configurations to read the differences between the tasks before and after recovery. This prevents fault detection from misjudging connection anomalies as lag or screen glitches.

[0053] Optionally, a real-time modification entry for detection tasks can be set up. When the remote diagnostic platform needs to update the second-level detection indicators based on changes in the online status of the device, adjustments to the platform-side management strategy, or subsequent returned execution results, the modification time and content can be written into the original task registration record through the remote diagnostic platform, while keeping the task number unchanged and only updating the detection task status and task content fields, thereby forming a continuous record for the same device and the same task chain.

[0054] After task registration is completed, the detection task can be issued. Specifically, the detection task content can be sent to the soft probe based on the task topic through the remote diagnostic platform. After receiving the detection task, the soft probe first verifies the task number, device identifier, and version number, then writes the detection task to the local task cache and sends back a receipt confirmation. If the soft probe finds that the task number is duplicated and the most recent task content has not changed, it only sends back a duplicate receipt flag and does not re-initialize; if it finds that the task number is the same but the second-level detection indicators or reporting rhythm have changed, it overwrites the local task cache with the modified task content and sends back a modified receipt flag.

[0055] For online devices with fluctuating activity, if the session is interrupted again during the delivery process, the remote diagnostic platform will retain the detection task status as delivered but not received, and wait for the updated online status of the device to trigger the delivery again; for offline devices, the detection task registration record will remain in the pending delivery status, and will not be sent into the task topic.

[0056] Optionally, information such as task number, second-level detection indicators, detection task status, terminal range, reporting rhythm, and reception confirmation status can be organized into detection tasks through a remote diagnostic platform. By establishing a continuously invoked detection task chain through the access relationship between the set-top box, soft probe, and remote diagnostic platform, a three-layer basic data structure is established, consisting of real-time connection information, device online status, and detection tasks. This forms a front-end link that fixes subsequent second-level detection indicator collection, fault detection, and automatic data collection condition generation under the same task number and device identifier, making subsequent data sources more centralized and the invocation relationship clearer.

[0057] S120. The soft probe continuously collects second-level detection indicators according to the detection task and sends them back to the remote diagnostic platform. The remote diagnostic platform generates fault information based on the second-level detection indicators and assembles a fault collection task based on the fault information and sends it to the soft probe.

[0058] The second-level detection metrics can be understood as operational fields continuously collected by the soft probe on the set-top box side and transmitted back to the remote diagnostic platform. These fields may include information such as player operating status, network reception status, screen display status, storage usage status, and process activity status. Fault information can be fault record data generated by the soft probe based on the second-level detection metrics, containing information such as fault type, fault period, and related metrics. Fault collection tasks can be data collection tasks specifically issued by the platform based on fault information, used to obtain complete data for fault tracing.

[0059] Optionally, the soft probe process in the set-top box can read the detection task in the local task cache, and verify the task number, device identifier and receiving time. After verification, it can continuously collect second-level detection indicators according to the detection task.

[0060] Based on the above scheme, optionally, the detection task includes a task number, a device identifier, and a collection item configuration, wherein the collection item configuration is used to indicate multiple target collection items to be collected; the soft probe continuously collects second-level detection indicators according to the detection task and sends them back to the remote diagnostic platform, including: the soft probe maps the collection item configuration in the detection task to the data acquisition entry corresponding to the set-top box, and performs second-level detection indicator collection on multiple target collection items at a fixed period; after each collection is completed, the soft probe assembles the attribute values ​​of multiple target collection items collected in the current period into a second-level reporting record according to the task number, device identifier, and collection time of the detection task, and sends the second-level reporting record to the remote diagnostic platform through a message queue telemetry transmission protocol.

[0061] The task number can be a unique code for each detection task, used to bind task data, reporting records, fault information, diagnostic data, and other information throughout the entire process. The collection item configuration can be understood as the preset collection rule configuration in the detection task, used to specify all terminal indicator items that the soft probe needs to collect. The target collection item can be the indicator item defined in the collection item configuration that needs to be collected at the second level. The data acquisition entry can be understood as the reading interface corresponding to various indicators within the set-top box system, which may include status reading interface, network interface, display interface, storage directory, and process table. The second-level detection indicator can be the specific value read by the soft probe from the corresponding data acquisition entry according to the indicator name, which may include player running status, network receiving status, screen display status, storage occupancy status, and process activity status. The target collection item attribute value can be understood as the quantified attribute parameters of each indicator obtained in a single collection. The second-level reporting record can be the standardized reporting data assembled by the soft probe periodically. The collection time can be the timestamp of the completion of a single second-level indicator collection.

[0062] Optionally, the collection item configuration is used to indicate at least several target collection items to be collected. Specifically, the collection items can first be loaded by mapping the collection item configuration in the detection task to the corresponding data acquisition entry of the set-top box through the soft probe. For example, the player running status, network receiving status, screen display status, storage usage status, and process activity status in the detection task can be mapped to the corresponding status reading interface, network interface, display interface, storage directory, and process table of the set-top box. At the same time, the collection order can also be loaded by writing the multiple target collection items to be collected into the collection loop of the soft probe at a fixed period, and the return rhythm can be loaded by writing the second-level data return cycle into the reporting loop of the soft probe and keeping it associated with the device online status. When the device online status is online, the return is carried out at a normal rhythm; when the device online status is fluctuating online, the return is carried out at a recovery rhythm; when the device online status update is abnormal, the data is temporarily stored locally and waits for the connection to be restored.

[0063] Furthermore, a persistent soft probe process can perform second-level detection indicator collection, with the execution location located locally on the set-top box. Optionally, the player's running status can be obtained by reading the current playback process, channel switching records, and playback session records. Network reception status can be obtained by reading network interface transmission and reception records, cache dwell records, and session connection records. Screen display status can be obtained by reading the current display frame status, black screen markers, and abnormal rendering markers. Storage occupancy status can be obtained by reading the occupancy records of system partitions, service directories, and cache directories. Process activity status can be obtained by scanning the activity records of current foreground and background processes.

[0064] During the data acquisition process, if a target acquisition item fails to be read, a failure record, including the failure time, failed acquisition item, and number of failures, is written locally on the set-top box using a soft probe. Other acquisition items are then processed, and the data is collected again in the next cycle. If the same acquisition item fails multiple times consecutively, it is marked as an abnormal acquisition item and sent to the remote diagnostic platform along with the next second-level data transmission.

[0065] Optionally, after each acquisition, a second-level data backhaul process is performed. The attribute values ​​of multiple target acquisition items, such as player running status, network receiving status, screen display status, storage usage status, and process activity status obtained in the current period, are assembled into a second-level reporting record by the task number, device identifier, and acquisition time of the detection task. The second-level reporting record is then sent to the remote diagnostic platform via the reporting topic of the MQTT protocol.

[0066] Furthermore, after receiving the second-level reporting records, the remote diagnostic platform can register information such as the return time, return status, and missing item markers according to the task number. In addition, for grouped device range scenarios, multiple set-top boxes can execute the same detection task separately, but each generates an independent second-level reporting record. For example, if multiple set-top boxes in a certain cell are within the same grouped device range, the remote diagnostic platform can uniformly issue detection tasks to that grouped device range. The soft probes in each set-top box collect the local player's operating status, network reception status, and screen display status, and then transmit them back separately. The platform still registers each device individually according to its device identifier, preventing field misuse.

[0067] Finally, the running class fields corresponding to the current collection cycle can be organized into second-level detection indicators through soft probes, and the second-level detection indicators can be written into the platform-side second-level detection table and local cache.

[0068] Based on the above scheme, optionally, the remote diagnostic platform generates fault information according to the second-level detection indicators, including: the soft probe performs stutter detection on the second-level detection indicators, wherein the stutter detection is a joint check of the player running status, network receiving status, and process activity status in the second-level detection indicators, extracting continuous playback session records, data receiving pause records, and process scheduling stagnation records from the second-level detection indicators and aligning them within the same period; if playback session pause, data receiving pause, and no foreground process progress record occur simultaneously within the same period, the period is marked as a suspected stutter period; if multiple adjacent periods continuously exhibit suspected stutter periods, the time period corresponding to these multiple periods is registered as a stutter fault time period; the soft probe performs screen flickering detection on the second-level detection indicators, wherein the screen flickering detection is a joint check of the screen display status and player running status in the second-level detection indicators, extracting abnormal rendering markers and black screen markers from the second-level detection indicators. The system records playback session switching records and display frame anomaly records and matches them according to playback time periods. If abnormal rendering markers and display frame anomaly records appear consecutively within the same playback time period, they are marked as suspected screen distortion periods. If the suspected screen distortion period does not disappear across adjacent acquisition cycles, the period is registered as a screen distortion fault period. After the stuttering detection and screen distortion detection are completed, the soft probe determines the fault type based on the presence and start and end times of the stuttering fault period and the screen distortion fault period. The fault types include stuttering faults, screen distortion faults, and concurrent faults. Concurrent faults are when stuttering fault periods and screen distortion fault periods exist simultaneously within the same time period. The soft probe assembles fault records according to the task number, device identifier, fault type, fault start time, fault end time, associated second-level detection indicators, and current device online status. The fault records are written as fault information into the local fault cache and uploaded to the fault registration table of the remote diagnostic platform.

[0069] The player's running status can be understood as the running status parameters of the set-top box's video playback program. Network reception status refers to the real-time status of the set-top box's network data reception. Process activity status can be understood as the running and scheduling status of the set-top box's foreground playback process and background service processes. Continuous playback session records can be time-series records of the player's continuous playback, pause, and interruption. Data reception pause records can be time records of brief pauses in network data reception and rates returning to zero. Process scheduling pause records can be state records of playback processes not being scheduled and being paused. Suspected stuttering periods can be abnormal periods within a single acquisition period where playback pauses and process pauses occur simultaneously. Stuttering fault periods can be time periods consisting of multiple consecutive suspected stuttering periods, confirming the existence of playback stuttering. Screen display status refers to the real-time display status of the TV screen. Abnormal rendering markers can be status markers indicating failed or distorted image rendering when the set-top box is rendering images. Black screen markers can be status markers indicating no image output and a completely black screen. Playback session switching records can be understood as records of the player switching playback sources and restarting playback sessions. Display frame anomaly records can include anomalies such as lost video playback frames, frame errors, frame distortion, and image misalignment. Suspected screen flickering periods can be suspected fault periods within the same playback time segment where rendering anomalies and display frame anomalies occur simultaneously. Fault types can be classifications of fault conditions identified by the soft probe, including but not limited to stuttering, screen flickering, and concurrent faults. Concurrent faults can be anomalies where the set-top box simultaneously experiences playback stuttering and screen flickering within the same time segment. The fault start time is the start time of the stuttering fault period or the screen flickering fault period, and the fault end time is the end time of the stuttering fault period or the screen flickering fault period. When the fault type is a concurrent fault, the fault start time is the earlier of the stuttering fault period and the screen flickering fault period, and the fault end time is the later of the stuttering fault period and the screen flickering fault period. The fault cache can be a temporary data storage area on the soft probe's local machine, used to temporarily store unreported fault record data. The fault registration form can be a data form used by the remote diagnostic platform backend to uniformly store and manage all fault records. Fault logs can be standardized data records generated by integrating all fault information.

[0070] Optionally, based on the second-level detection metrics, stuttering detection, screen flickering detection, and fault registration processing can be performed to generate fault information. The second-level detection metrics may include various information such as player running status, network reception status, screen display status, storage usage status, and process activity status.

[0071] Specifically, fault detection can be performed through a fault detection unit. The fault detection unit can be a processing unit inside the soft probe that receives second-level detection indicators and outputs fault information. It is deployed locally on the set-top box and is executed after each second-level data back transmission is completed.

[0072] Optionally, stuttering detection can be performed on the second-level detection indicators using a soft probe. Stuttering detection can be a process of jointly checking the player's running status, network reception status, and process activity status. Specifically, the soft probe can extract continuous playback session records, data reception pause records, and process scheduling pause records from the second-level detection indicators, and align these three types of records within the same period. If playback session pauses, data reception pauses, or no progress records of the foreground process occur simultaneously within the same period, then that period is marked as a suspected stuttering period. If multiple adjacent periods continuously show suspected stuttering periods, then that time period is registered as a stuttering fault period, where the stuttering fault period includes the start and end times. Here, alignment within the same period means checking records from different acquisition items side by side according to the same acquisition time window, thereby avoiding misjudgments caused by a single network jitter or a single process switch.

[0073] Furthermore, screen distortion detection can be performed on second-level detection indicators using a soft probe. This screen distortion detection can be a process of jointly checking the screen display status and the player's operating status. Specifically, the soft probe can extract abnormal rendering markers, black screen markers, playback session switching records, and abnormal display frame records from the second-level detection indicators. Time periods with switching markers are first registered as observation periods, and these records are matched according to playback periods. If abnormal rendering markers and abnormal display frame records appear consecutively within the same playback period, they are marked as suspected screen distortion periods. If the suspected screen distortion period does not disappear across adjacent acquisition cycles, the period is registered as a screen distortion fault period, which includes the start and end times.

[0074] Furthermore, boundary constraint processing can be performed. For short-term screen changes occurring during channel switching, boot loading, and application switching, a software probe can first read the switching markers in the player's running status to determine whether to enter the screen distortion detection phase; time periods with switching markers are first registered as observation periods and not directly written into the screen distortion fault period. For cases of momentary network interruption but uninterrupted playback session, the software probe can register them as network fluctuation periods and not directly write them into the stuttering fault period.

[0075] Optionally, after the lag detection and screen distortion detection are completed by the soft probe, the fault type can be determined based on the presence and start and end times of the lag fault period and the screen distortion fault period. The fault type may include lag fault, screen distortion fault, and concurrent fault, where concurrent fault refers to the situation where lag fault period and screen distortion fault period exist simultaneously within the same time period. If the fault occurs for the first time in the current cycle, it is recorded as a new fault. If the fault continues across cycles, it is recorded as a persistent fault. If the fault record disappears in a subsequent cycle, it is recorded as an ended fault.

[0076] Optionally, the soft probe determines the fault type and fault period based on the detection results of the stuttering fault period and the screen distortion fault period: if only the stuttering fault period exists, the fault type is stuttering fault, and the fault start time and fault end time are taken as the start time and end time of the stuttering fault period; if only the screen distortion fault period exists, the fault type is screen distortion fault, and the fault start time and fault end time are taken as the start time and end time of the screen distortion fault period; if both exist simultaneously, the fault type is concurrent fault, and the fault start time is taken as the earlier of the two start times, and the fault end time is taken as the later of the two end times.

[0077] Optionally, after the stuttering and screen flickering detection, a fault registration process can also be performed. Specifically, a fault record can be assembled using a soft probe according to multiple pieces of information, such as task number, device identifier, fault type, fault start time, fault end time, associated second-level detection indicators, and current device online status. The fault record is then written as the fault information into the local fault cache and uploaded to the fault registration form of the remote diagnostic platform.

[0078] For example, suppose in a certain Internet TV playback scenario, a user is watching a live channel. The soft probe collects playback session pause records, data reception pause records, and abnormal rendering markers for several consecutive cycles. The fault detection unit first registers the stuttering fault period, then registers the screen flickering fault period, and generates concurrent fault cases under the same task number, thus sending the fault records of the same period to the subsequent automatic data acquisition condition generation process all at once. After completing the above processing, the soft probe organizes the fault type, fault period, associated second-level detection indicators, and current device online status into a fault case, and records the fault case in the local fault cache.

[0079] Based on the collection of second-level detection indicators, fault conditions are generated through lag detection and screen flickering detection. These fault conditions then directly trigger the automatic generation of data collection conditions, registration of collection range, and assembly of data collection tasks, forming a fault collection task. This task can include the data collection range reserved before and after the fault period, the collection object interface, and linkage information bound to the fault type. This realizes the transformation from passive diagnosis to proactive intelligent diagnosis, ensuring the precise synchronization of data collection and evidence gathering with the occurrence of the fault.

[0080] Based on the above scheme, optionally, the remote diagnostic platform assembles a fault acquisition task according to the fault information and sends it to the soft probe. Specifically, the soft probe generates data acquisition conditions according to the fault information. The data acquisition conditions include time period conditions and traffic threshold conditions. The time period conditions define the calculation rules for the start time and end time of data acquisition, and the traffic threshold conditions define the network traffic determination rules for initiating data acquisition. The soft probe calculates the start time and end time of data acquisition according to the time period conditions and registers the acquisition object interface and acquisition object session. The remote diagnostic platform receives the data acquisition conditions, the start time of data acquisition, the end time of data acquisition, the acquisition object interface, and the acquisition object session, and assembles them into the fault acquisition task by combining the task number, device identifier, and current device online status, and sends it to the soft probe.

[0081] The data acquisition conditions can be understood as the triggering and scope conditions for deep fault data collection. Data acquisition conditions include at least time period conditions and traffic threshold conditions. Time period conditions can be conditions that limit the time range of fault data collection, used to define the start and end times of collection. Traffic threshold conditions can be judgment threshold conditions that limit the initiation of network traffic data packet collection; precise collection is only triggered when a preset abnormal traffic value is reached. The data acquisition start time and end time are the start and end time nodes of deep fault data collection, respectively. The collection object interface can include object interfaces such as the set-top box system interface, network interface, and playback service interface that need to collect data. The collection object session can include the corresponding video playback session, network connection session, and business interaction session within the fault period.

[0082] Optionally, data acquisition conditions can be generated, acquisition range registered, and acquisition task assembly processed based on fault information to generate a fault acquisition task and send it to the soft probe. This can be achieved through the collaborative execution of the acquisition control unit in the soft probe and the task assembly unit in the remote diagnostic platform. The acquisition control unit is located locally on the set-top box and is used to generate local automatic data acquisition conditions based on fault information; the task assembly unit is located on the remote diagnostic platform and is used to form an executable fault acquisition task based on fault information and platform-side task policies.

[0083] Specifically, automatic data acquisition conditions can be generated first. These automatic data acquisition conditions refer to the set of conditions that trigger the tcpdump acquisition operation, which may include fault type conditions, time period conditions, and traffic threshold conditions. The fault type condition limits acquisition to cases where the current fault is a lag fault, a screen flickering fault, or a concurrent fault. The time period condition limits the start and end times of data acquisition; the start time can be a reserved period before the fault start time, and the end time can be a reserved period after the fault end time, thus covering network interaction records before and after the fault. The traffic threshold condition limits acquisition to when the network interface experiences an abnormal surge, abnormal drop, or prolonged lack of traffic change during the fault period, thereby avoiding frequent acquisition during fault-free or abnormal traffic periods.

[0084] Furthermore, when generating automatic data acquisition conditions, the acquisition control unit can also read the device's online status. If the current device is online, the conditions to be executed are written directly; if the current device is fluctuating online, the conditions to be restored are written and resent after the connection is restored; if the current device is offline, only the fault information and automatic data acquisition condition records are retained, and local acquisition is not started.

[0085] Optionally, after the automatic data acquisition conditions are generated, a data acquisition range registration process can be performed. Specifically, the data acquisition start time, data acquisition end time, acquisition target interface, and acquisition target session can be registered by the acquisition control unit based on the fault start time, fault end time, and current network interface status. The acquisition target interface refers to the network interface currently carrying Internet TV service traffic on the set-top box. The acquisition target session refers to the network connection record associated with the current player's operating status. For a single device range, the data acquisition range registration can be directly bound to the current device identifier; for a grouped device range, multiple set-top boxes with the same fault type and the same service time period can be registered as a grouped acquisition range through a remote diagnostic platform, while still maintaining the independent data acquisition start time and data acquisition end time for each set-top box.

[0086] Specifically, when the device is online or fluctuating online, the data acquisition control unit can, according to the calculation rules defined in the time period conditions, use the reserved period before the fault start time as the data acquisition start time and the reserved period after the fault end time as the data acquisition end time, and register the acquisition object interface and acquisition object session according to the current network interface status.

[0087] Furthermore, to achieve continuous operation and version evolution, task version management can be implemented. Specifically, before assembling a collection task, the task assembly unit can read the detection task version number and collection strategy version number corresponding to the current task number. If the detection task version number has been updated but the collection strategy version number has not, the original collection strategy version number is used with a version difference marker added; if the collection strategy version number has been updated, the new version number and effective time are written into the fault collection task. Version management ensures that subsequent TCP packet capture tool (tcpdump) collection operations can be executed according to the corresponding version. The platform can also trace the collection scope and triggering conditions used at the time when parsing the traffic packets of the faulty node.

[0088] Optionally, after the collection range is registered, a collection task assembly process can be performed. Specifically, the automatic collection conditions and collection range registration content uploaded by the set-top box can be received through the task assembly unit in the remote diagnostic platform, and a fault collection task can be assembled by combining the task number, device identifier, terminal range, and current online status of the device. The fault collection task includes at least the following information: automatic data collection conditions, data collection start time, data collection end time, collection object interface, task version number, upload method, and linked video stream return flag. The linked video stream return flag can be used to notify the set-top box to start real-time video stream return simultaneously with the tcpdump collection operation.

[0089] Optionally, for platform-side batch acquisition scenarios, multiple fault acquisition tasks can be generated by the task assembly unit according to the grouped device range, and then sent to the corresponding soft probes one by one. For scenarios where there are already incomplete acquisitions on the set-top boxes, the acquisition control unit does not directly overwrite the old tasks, but writes queuing marks and priority order into the new tasks. The priority order can be determined by the fault type and the recentity of the fault time period. For example, if a live broadcast channel triggers stuttering faults on multiple set-top boxes simultaneously during the evening peak period, the platform-side task assembly unit can assemble multiple fault acquisition tasks according to the grouped device range, and write the fault type, acquisition range, and linkage video stream return mark for the same channel and the same time period into their respective tasks. Subsequently, each set-top box soft probe executes them separately, avoiding the mixing of acquisition files.

[0090] After completing the above processing, the information such as automatic data acquisition conditions, data acquisition start time, data acquisition end time, acquisition object interface, task version number, upload method, and linkage video stream return mark can be organized into a fault acquisition task through the remote diagnostic platform, and the fault acquisition task can be written into the platform-side task table and the set-top box local task cache.

[0091] By directly converting fault information into automatic data acquisition conditions and fault acquisition tasks, instead of using the platform to initiate ordinary diagnostic tasks and then select the tcpdump path, the system automatically forms the acquisition range and task version when stuttering, screen flickering, or concurrent faults occur. This ensures that the obtained fault acquisition tasks, second-level detection indicators, fault information, and subsequent video stream feedback maintain the same task number and the same fault time period. When the platform parses the traffic data packets of the fault node later, it can directly correspond to the fault occurrence link, resulting in a more concentrated time period and more complete fields.

[0092] S130. The soft probe collects data according to the fault collection task to generate a diagnostic dataset, and uploads the diagnostic dataset to the remote diagnostic platform. The diagnostic dataset includes data packet files of the set-top box's service traffic, device diagnostic results, and the set-top box's real-time video stream.

[0093] The diagnostic dataset can be a collection of fault tracing data generated after the soft probe performs fault collection tasks. This data includes, but is not limited to, set-top box service traffic data packet files, device diagnostic results, and set-top box real-time video streams. The service traffic data packet files can be archived files of all network data packets generated during set-top box playback and network interaction, used to trace network anomalies. Device diagnostic results can be quantitative diagnostic data and detection results generated after detecting the set-top box's operating status. The real-time video stream can be the live video stream of the television playback during the period of set-top box failure, which can be used to visually locate visual faults such as screen stuttering and distortion.

[0094] Optionally, a fault acquisition task can be obtained through a soft probe, performing tcpdump acquisition operations, uploading data acquisition results, and processing the real-time video stream from the set-top box to obtain a diagnostic dataset. The diagnostic dataset is then uploaded to the remote diagnostic platform. The diagnostic dataset includes at least the data packet file of the set-top box's service traffic, device diagnostic results, and the set-top box's real-time video stream.

[0095] S140. The remote diagnostic platform extracts critical path content based on the diagnostic dataset, determines the fault determination result based on the critical path content, issues a handling task to the soft probe based on the fault determination result, and updates the task status based on the execution result of the handling task by the soft probe.

[0096] The critical path content can be extracted from the diagnostic dataset, directly corresponding to the link data and event records of fault occurrence, persistence, and recovery. The fault determination result can be the final fault result output by the remote diagnostic platform after analyzing the critical path content, which may include, but is not limited to, fault type, fault cause, and fault node information. The handling task can be a targeted remote repair or anomaly handling task issued by the platform based on the fault determination result, used to resolve set-top box terminal faults. The execution result can be the terminal execution status data, such as task success, failure, or anomaly, fed back by the soft probe after receiving and running the handling task.

[0097] Optionally, the remote diagnostic platform can automatically parse the traffic data packets of the faulty node, extract the critical path content, and present the critical path content based on the diagnostic dataset to obtain the critical path content. Based on the critical path content, the platform can determine the fault determination result, issue a handling task to the soft probe based on the fault determination result, and update the task status based on the execution result of the handling task by the soft probe.

[0098] Based on the above scheme, optionally, the remote diagnostic platform extracts critical path content according to the diagnostic dataset, including: the parsing unit in the remote diagnostic platform reads network interaction records from the data packet files in the diagnostic dataset, aligns them with the network diagnostic results in the diagnostic dataset, and obtains fault link records arranged in chronological order; the parsing unit filters out the connection segments, dwell segments, and recovery segments corresponding to the fault conditions from the fault link records, binds the start time, end time, session status, and corresponding device diagnostic records of each segment, and generates critical path content.

[0099] The parsing unit can be a platform-side processing unit that receives the diagnostic dataset and outputs the critical path content. Network interaction records can be understood as the original records of all network connections, data transmissions, and session interactions contained in the data packet file. Fault link records can be time-series data records arranged chronologically to reconstruct the entire process of fault occurrence, development, and recovery. Connection segments can be the stage data of network session establishment and service connection activation in the fault link. Stasis segments can be the stage data of network stagnation, data freezes, and persistent abnormal video playback in the fault link. Recovery segments can be the final stage data of the fault link where the abnormality disappears, the network recovers, and playback returns to normal.

[0100] Based on the above scheme, optionally, the parsing unit reads network interaction records from the data packet file in the diagnostic dataset and aligns them with the network diagnostic results in the diagnostic dataset, including: the parsing unit calls the data packet file in the diagnostic dataset according to the task number, and establishes a data reading window for the same fault period according to the data acquisition start time, data acquisition end time, and corresponding screen time period; the parsing unit reads service connection records, session establishment records, traffic dwell records, and abnormal interruption records from the data packet file, and aligns them with the network interface connectivity status, session connection status, and service traffic arrival status in the network diagnostic results in the diagnostic dataset; the parsing unit extracts network interaction records corresponding to the player's running status segment by segment according to the same task number, the same device identifier, and the same fault period, and superimposes the background process scan results, central processing unit (CPU) information diagnostic records, and streaming service test diagnostic results in the diagnostic dataset onto the same time period to obtain the fault link records arranged in chronological order.

[0101] The data reading window can be a data reading interval defined by the platform based on the fault period, and only valid fault data within the window is read. Service connection records can be records of the establishment, maintenance, and disconnection of connections between the set-top box and the streaming media server and service server. Session establishment records can be records of the creation, initialization, and binding of video playback sessions and network communication sessions. Traffic dwell records can be time-series records of service traffic interruptions and transmission pauses. Abnormal interruption records can be records of abnormal disconnections of network sessions and playback sessions. Network interface connectivity status can be the status of the set-top box's network interface, including connected, disconnected, and abnormal states. Service traffic arrival status can be a description of whether the service data stream has arrived at the terminal normally.

[0102] Specifically, the parsing unit can retrieve the fault node data packet file by task number, and then establish a data reading window for the same fault period based on the data acquisition start time, data acquisition end time, and corresponding screen time. Next, it reads service connection records, session establishment records, traffic dwell records, and abnormal interruption records from the fault node data packet file, and aligns them with the network interface connectivity status, session connection status, and service traffic arrival status in the network diagnostic results. The automated parsing does not merely display the original collected files; instead, the parsing unit extracts network interaction records corresponding to the player's operating status segment by segment according to the same task number, the same device identifier, and the same fault period. Then, it overlays the background process scan results, CPU information diagnostic records, and streaming service test diagnostic results onto the same time period to obtain fault link records arranged in chronological order.

[0103] Furthermore, connection segments, dwell segments, and recovery segments directly corresponding to fault conditions such as lag, screen flickering, or concurrent failures can be filtered from the fault link records. The start time, end time, session state, and corresponding device diagnostic records of each segment are then bound together to extract the critical path content. These bound records can be written as time-segment link content accessible to the platform, enabling the presentation of critical path content. The presented content includes at least records of the fault initiation segment, fault duration segment, and fault recovery segment, with each record corresponding to the device diagnostic results and network diagnostic results in the diagnostic dataset.

[0104] For example, in a scenario where a live streaming channel experiences buffering and screen tearing, the parsing unit can first read the data packets of the faulty node, then extract the session connection status of the same time period from the network diagnostic results, and extract the stream service recovery record from the stream service test diagnostic results. Finally, the abnormal interruption record, recovery delay record, and background process continuous resident record are written side by side in the same fault link record in chronological order, and the critical path content is extracted from them.

[0105] Optionally, the parsing unit can first locate the faulty node's data packet file by task number, and then extract the time series of network interface connectivity status, session connection status, and service traffic arrival status from the network diagnostic results. Next, the parsing unit can then analyze the data according to the data collection start time. Data collection end time and corresponding screen time period Establish a unified data reading window for the fault period, and use deep packet inspection technology to parse the data packet files of the faulty nodes packet by packet to obtain service connection records. (This may include information such as connection establishment time, source / destination IP, port, and protocol type), session establishment records. (This may include information such as session ID, establishment time, and status code), traffic dwell time records. (This may include information such as packet arrival interval, cumulative bytes within the window, etc.) and abnormal interruption records. (This may include information such as interruption time and cause code). Finally, the parsing unit can align the timestamps of these read records with the corresponding fields in the network diagnostic results to ensure that each record can be mapped to a second-level timeline of the same fault period.

[0106] To quantify the degree of anomaly in network interactions within each time segment, a parsing unit of length can be used for each segment. Time slices (e.g.) Define a session anomaly metric (seconds). The session anomaly index is calculated as follows: the session anomaly index is obtained by multiplying the interruption indication function value by a first weight, subtracting the difference between the current time slice's cumulative received bytes and the maximum reference bytes during the peak period of the entire network by a second weight, multiplying the quotient of the root mean square interval of the current time slice's data packet arrival by the jitter normalization constant by a third weight, and then dividing by the sum of the three weights. The first weight is set to 0.5, the second weight to 0.3, and the third weight to 0.2. The session anomaly index is between 0 and 1.

[0107] The session anomaly index can be calculated from the original packet records based on the following formula: ; in, Indicates the first The session anomaly score for each time slice, with a value range of [value missing]. ; This represents a time-slice index, corresponding to a data collection period with a fixed step size. A continuous time segment divided into seconds; This is a weighting coefficient, dimensionless, and can be set to values ​​such as 0.5, 0.3, 0.2, etc., to balance the contributions of interruption, traffic, and jitter to the anomaly score. This indicates an indicator function, i.e., if the first... If there is an exception interrupt record in the time slice memory, then The value is 1 if it is not 1, otherwise it is 0. Indicates the first The cumulative number of bytes received within a time slice can be determined based on the traffic statistics in the network diagnostic results, and the unit is bytes. This indicates the maximum reference number of bytes for this interface during the peak period across the entire network. It can be obtained from historical statistics of the platform and is measured in bytes. Indicates the first The root mean square of the arrival interval of data packets within a time slice can be used to reflect jitter. It can be determined based on network diagnostic results and is measured in milliseconds (ms). This represents the jitter normalization constant, which can be taken as the average interval during normal playback (e.g., 10ms), used to normalize jitter values ​​to a dimensionless range; the denominator is... The sum of weights is used to ensure Normalize after weighted summation.

[0108] For example, suppose a certain time slice Interruption records exist. , , And then calculate Since it exceeds 1, it is truncated to 1, indicating a serious anomaly.

[0109] Next, the background process scan results, CPU information diagnostic records, and streaming service test diagnostic results can be overlaid onto the same fault period using the parsing unit. To evaluate the correlation between each diagnostic record and the player's operating status, a fault link record can be constructed. The fault link record is a vector sequence arranged in chronological order. Each vector element includes a timestamp, the session anomaly level, the average CPU utilization, the memory utilization, the background process anomaly flag, and the streaming service recovery latency. The average CPU utilization and memory utilization range from zero to one, and the background process anomaly flag has a value of zero or one, so as to merge the anomaly levels of multiple dimensions into a link vector ordered by time. ; in, This represents a fault link record, which is a vector sequence arranged in chronological order, with each element corresponding to multi-dimensional information of a time slice; Indicates the first The timestamp of each time slice is measured in seconds. Indicates the first Session anomaly level for each time slice; Indicates the first The average CPU utilization within a time slice can be determined based on CPU information diagnostic records, with a value range of [value missing]. (Percentage normalization); Indicates the first The memory usage rate within a time slice can be determined based on memory information diagnostic records, with a value range of [value missing]. ; Indicates the first An abnormal flag for a background process within a time slice can be determined based on the background process scan results, with a value of 0 (normal) or 1 (abnormal, such as continuous unresponsiveness). Indicates the first The recovery delay of streaming services within a time slice can be determined based on the results of streaming service test and diagnosis, and is measured in seconds. Indices representing the start and end time slices within a data collection period can be derived from the data collection start time. and data collection time according to The result is obtained through conversion.

[0110] Optionally, connection segments, pause segments, and recovery segments directly corresponding to stuttering, screen flickering, or concurrent failures can be filtered from the fault link records to extract critical path content. Specifically, a sliding window detection algorithm can be used in the parsing unit to find continuous anomalies exceeding a threshold. The longest interval (e.g., 0.7) is taken as the fault duration segment, and its preceding interval is... One time slice is used as the fault initiation segment, then... Each time slice is designated as a fault recovery segment. Next, the start time, end time, session status, and corresponding device diagnostic records of these three segments are recorded. , , ) and network diagnostic results Bind together to generate critical path content This may include the fault initiation segment. Fault duration Fault recovery section The three-segment record, each containing a time interval, average anomaly rate, and associated diagnostic data, enables automated analysis and intelligent judgment of the critical path of the faulty link, replacing the inefficient mode of manual data collection and interpretation.

[0111] By automatically parsing the fault node data packet files in the diagnostic dataset and combining them with multi-source diagnostic records, a fault link record organized by time axis is constructed. Then, the critical path content is extracted, so that the network interaction, equipment status and service quality information during the fault period form a structured and traceable link data, which provides a unified input for subsequent intelligent judgment and avoids the tedious process of manually splicing multiple diagnostic records.

[0112] Based on the above scheme, optionally, the remote diagnostic platform determines the fault determination result based on the critical path content, including: the determination unit in the remote diagnostic platform calculates the device-side anomaly dominance, network-side anomaly dominance, and flow service anomaly dominance respectively, obtains the probability of each cause through normalization, takes the cause corresponding to the highest probability as the main cause, determines the fault cause category, and locates the fault node according to the main cause and the specific record in the critical path content.

[0113] The judgment unit is a platform-side processing unit that categorizes the critical path content and outputs the fault judgment result. Device-side anomaly dominance can be understood as the weighted score of the impact of set-top box terminal hardware, system, and process anomalies on the fault. Network-side anomaly dominance can be understood as the weighted score of the impact of network link, session, and traffic anomalies on the fault. Streaming service anomaly dominance can be understood as the weighted score of the impact of cloud streaming media service and playback service anomalies on the fault. The primary cause can be the core fault cause with the highest anomaly probability value, playing a dominant role in the fault. The fault cause category can be the fault attribution classification, including but not limited to various fault cause types such as device-side faults, network-side faults, streaming service faults, and concurrent faults. The fault node can be the specific location node where the fault occurs.

[0114] Specifically, the judgment unit can first read the fault initiation segment, fault duration segment, and fault recovery segment in the critical path content segment by segment, and then compare the CPU information diagnostic record, memory information diagnostic record, storage information diagnostic record, background process scan results, network diagnostic results, and stream service test diagnostic results, etc., to determine whether the current fault is dominated by device-side records, network-side records, stream service records, or by multiple types of records.

[0115] Fault cause determination can be a process of sorting out the types of fault sources. The determination content can include device-side abnormal causes, network-side abnormal causes, streaming service abnormal causes, and concurrency abnormal causes. Fault node determination is a process of sorting out the location of the fault. The determination content can include set-top box local nodes, network interface nodes, and streaming service connection nodes.

[0116] Furthermore, within the same fault period, if the CPU information diagnostic record, memory information diagnostic record, and background process scan results all show continuous anomalies, while the network diagnostic results show no obvious interruption records, then the judgment unit registers this period as a device-side anomaly cause and identifies the fault node as the set-top box local node. If the network diagnostic results and the session interruption records in the fault node's data packets appear simultaneously, while the device diagnostic results show no continuous anomalies, then the period is registered as a network-side anomaly cause and the fault node is identified as the network interface node. If the streaming service test diagnostic results show a streaming service recovery delay record, and the fault recovery segment in the critical path content is significantly delayed, then the period is registered as a streaming service anomaly cause and the fault node is identified as the streaming service connection node. For cases where multiple types of records appear simultaneously within the same period, the judgment unit can prioritize writing the record that appears first in the fault initiation segment into the fault judgment result, and then write the remaining records as related causes, thereby preserving the correspondence between the main cause and related causes.

[0117] For example, if a set-top box experiences a lag during peak evening hours, the critical path content shows that the network interface record is interrupted first, followed by a delay in the stream service recovery record, while the CPU information diagnostic record remains stable. Based on this, the judgment unit can write the main cause as a network-side abnormality, the related cause as a stream service abnormality, and determine the fault node as the network interface node.

[0118] Based on the above scheme, optionally, the determination unit calculates the device-side anomaly dominance, network-side anomaly dominance, and flow service anomaly dominance respectively, obtains the probability of each cause through normalization, takes the cause corresponding to the highest probability as the main cause, and determines the fault cause category, including: the determination unit reads the fault initiation segment, fault duration segment, and fault recovery segment in the critical path content segment by segment, and compares them with the CPU information diagnostic records, memory information diagnostic records, storage information diagnostic records, background process scan results, network diagnostic results, and flow service test diagnostic results in the diagnostic dataset respectively; for the fault duration segment, the device-side anomaly dominance, network-side anomaly dominance, and flow service anomaly dominance are calculated respectively: the device-side anomaly dominance is the sum of the average CPU utilization rate of each time slice in the fault duration segment multiplied by the first device weight, the memory utilization rate multiplied by the second device weight, and the background process anomaly mark multiplied by the third device weight. The average value is calculated as follows: the weight of the first device is 0.4, the weight of the second device is 0.3, and the weight of the third device is 0.3. The network-side anomaly dominance is the average value of the session anomaly degree of each time slice within the fault duration. The streaming service anomaly dominance is the average value of the sum of the indicator function values ​​of the streaming service recovery delay exceeding a preset threshold within the fault duration. The device-side probability, network-side probability, and streaming service probability are obtained through exponential normalization. The cause corresponding to the highest probability is taken as the primary cause. If the difference between the highest probability and the second highest probability is less than 0.1, it is marked as a concurrent anomaly cause. The determination unit locates the fault node based on the primary cause and the specific records in the critical path content, including: if the primary cause is a device-side anomaly, the fault node is determined to be the set-top box local node; if the primary cause is a network-side anomaly, the fault node is determined to be the network interface node; if the primary cause is a streaming service anomaly, the fault node is determined to be the streaming service connection node.

[0119] Among these, CPU average utilization can be understood as the average resource usage ratio of the set-top box's central processing unit within a single time slice. Memory utilization can be understood as the resource usage ratio of the set-top box's running memory within a single time slice. Background process anomaly flags can be understood as flags indicating abnormal states such as background processes freezing, crashing, or excessive resource usage. Session anomaly score can be a quantitative score of the severity of network session anomalies within a single time slice. Streaming service recovery latency can be the time taken to restore normal transmission and playback after a streaming media service anomaly. Indicator function value can be a binary function used to determine whether an anomaly has been triggered. Concurrency anomaly cause can be understood as an anomaly caused by multiple factors when the difference between the highest probability and the second highest probability is less than 0.1. Set-top box local node can be understood as a fault occurring in the set-top box device itself, mainly caused by hardware, system, process, or resource anomalies. Network interface node can be understood as a fault occurring in the terminal network link, network card, or routing transmission link. Streaming service connection node can be understood as a fault occurring in the cloud streaming media service, playback source service, or service connection link link.

[0120] Optionally, the fault initiation segment in the critical path content can be determined by the determination unit. Fault duration Fault recovery section Read segment by segment and compare the diagnostic records of the central processing unit, memory, storage, background process scans, network and streaming services in the diagnostic dataset.

[0121] For each segment, calculate the equipment-side anomaly dominance. Network-side anomaly dominance HeLiu service abnormal dominance Taking the fault duration as an example, the confidence score for each candidate cause category can be calculated based on the following formula: ; ; ; in, These are the confidence scores for the causes of anomalies on the device side, network side, and flow service during the fault duration, respectively, corresponding to the device-side anomaly dominance, network-side anomaly dominance, and flow service anomaly dominance, and are dimensionless. Indicates the duration of the fault The total number of time slices included; This represents a set of time slice indexes indicating the duration of a fault, which can be found in the critical path content. Extracted from; The device weights for the device-side sub-indicators are 0.4, 0.3, and 0.3 respectively, which are the device weights corresponding to the average CPU utilization, memory utilization, and background process abnormality flags. These weights are used to comprehensively consider the impact of CPU, memory, and process abnormalities. This indicates an indicator function, if the streaming service recovery is delayed. Exceeding the threshold The value is 1 if it is 1, otherwise it is 0. This indicates the threshold for the recovery delay of the streaming service, which is preset to 2 seconds and can be determined according to the platform's business needs.

[0122] For example, suppose the fault duration consists of 5 time slices, and the duration of each time slice is... They are respectively , for , for ,but If at the same time for ,but If the streaming service recovery delay exceeds the threshold only in the first and second time slices, then .

[0123] Furthermore, the device-side probability, network-side probability, and flow service probability can be obtained by normalization based on the following formula: ; ; ; in, These represent the probabilities of the causes of anomalies on the device side, network side, and flow service, respectively, i.e. device side probability, network side probability, and flow service probability, with values ​​between 0 and 1, and the sum of the three is 1; is a natural constant, approximately 2.71828, used to map the score to non-negative values ​​using an exponential function and perform softmax normalization.

[0124] Furthermore, the cause corresponding to the highest probability can be taken as the primary cause. If the difference between the highest probability and the second highest probability is less than 0.1, it is marked as a concurrent anomaly cause. The fault node is determined based on the specific records in the primary cause and the critical path content: if the primary cause is an equipment-side anomaly cause and the duration is within the specified period... or If the value remains consistently high, the node is identified as a local node of the set-top box; if the main cause is a network-side anomaly and the session abnormality is high... If the value remains consistently high and retransmissions or packet loss are detected in the data packet file of the faulty node, the node is identified as a network interface node; if the main cause is an abnormality in the streaming service and the recovery delay frequently exceeds the threshold, the node is identified as a streaming service connection node.

[0125] Continuing with the above numerical example, if the calculation yields... , , If the main cause is an abnormality on the device side and the CPU and memory usage remains high, then the faulty node is identified as the local node of the set-top box.

[0126] Finally, the determination unit can organize the main cause, related causes, fault node determination results, and corresponding task numbers and equipment identifiers into a fault determination result. Its structure may include information such as task number, device identifier, main cause, list of related causes, fault node identifier, device-side anomaly dominance, network-side anomaly dominance, and flow service anomaly dominance.

[0127] Based on the critical path content, the system automatically identifies the causes and nodes of failures by quantifying and normalizing the confidence levels of each cause category. This avoids the inefficiency of manually analyzing multi-source data and outputs the results in a structured manner, providing a clear basis for subsequent closed-loop management.

[0128] Based on the above scheme, optionally, the step of issuing a handling task to the soft probe according to the fault determination result and updating the task status according to the execution result of the handling task by the soft probe includes: the handling unit in the remote diagnostic platform selects a handling path according to the fault cause and fault node in the fault determination result, generates a corresponding restart task or cleanup process, and issues the task topic to the soft probe through the message queue telemetry transmission protocol; after receiving and executing the restart task or the cleanup process, the soft probe assembles the execution result into a return message and sends it to the remote diagnostic platform; after receiving it, the remote diagnostic platform updates the task registration record and binds the current execution record with the fault determination result.

[0129] The system includes a handling unit for generating and issuing restart tasks or cleanup processes, with the software probe responsible for receiving and executing them. The handling path can be a remote repair solution path corresponding to different fault causes and fault nodes. Restart tasks can be platform-issued restart handling tasks for terminal processes, players, or network services. Cleanup processes can be platform-issued cleanup handling processes for terminal cache cleanup, junk process cleanup, and network cache reset. Feedback messages can be feedback messages assembled by the software probe after executing the handling tasks, including execution results and terminal status. Task registration records can be handling task ledger records stored in the platform's backend, containing information on the entire process of task issuance, execution, and completion.

[0130] Specifically, the handling unit first reads the fault determination result, and then selects a handling path based on the fault cause and fault node in the fault determination result. When the fault determination result points to the local node of the set-top box, and the device-side abnormality is dominant, the handling unit generates a cleanup process. When the fault determination result points to the local node of the set-top box, and the critical path content shows that the fault duration cannot be recovered for a long time, the handling unit generates a restart task. When the fault determination result contains both device-side and network-side abnormalities, the handling unit first generates a cleanup process, and then decides whether to continue generating a restart task based on the receipt status after the cleanup process is executed. The restart task includes at least a task number, device identifier, distribution time, execution method, and receipt requirements. The cleanup process includes at least background process processing records, cache directory processing records, installation package directory processing records, and junk file directory processing records.

[0131] Furthermore, a restart task or cleanup process can be sent to the soft probe via a task topic using the MQTT protocol from the remote diagnostic platform. After receiving the task, the soft probe first verifies the task number and device identifier, and then enters the execution phase according to the task content. Specifically, when executing a restart task, the soft probe first writes a local execution flag, then calls the set-top box restart interface, and re-enters the real-time connection and online status management link after restarting. When executing a cleanup process, the soft probe cleans up background processes, cache directories, installation package directories, and junk file directories item by item according to the directory and process records in the cleanup process. Records that cannot be processed are written to the local exception log, and the execution results are sent back for processing after the set-top box completes the restart task or cleanup process.

[0132] Specifically, the execution results, such as task number, device identifier, execution start time, execution end time, execution status, and abnormal log location, can be assembled into a feedback message by a soft probe and sent to the remote diagnostic platform. After receiving the message, the remote diagnostic platform updates the task registration record, binds the current execution record with the aforementioned fault determination result, and sends the execution status back to the device online status update link and the detection task status registration link of the soft probe.

[0133] For example, in a screen flickering scenario caused by an anomaly on a certain device side, the remote diagnostic platform first issues a cleanup process based on the fault determination result. After the soft probe completes the background process processing and cache directory processing, it sends back the execution status. If the sent-back status shows that the player's running status has recovered after the cleanup process is completed, the platform registers the execution result and ends the task. If the sent-back status shows that the fault continues, the platform continues to issue a restart task under the same task number link.

[0134] By directly converting critical path content and fault determination results into restart tasks or cleanup processes, the existing technology no longer requires manual review of data collection results before deciding whether to remotely control the system. This ensures that the resulting handling chain corresponds one-to-one with the task number, fault node determination results, and execution result feedback. The platform can complete the parsing, determination, and handling according to the same fault chain. When the system subsequently re-enters real-time connection and online status management, the detection task status can continue to be updated along the original chain.

[0135] Optionally, it can be read by the processing unit. The main causes and fault nodes are identified, and then the corresponding handling path is selected according to the preset handling strategy table. To quantify the expected cost of different handling paths, the cleanup process cost can be calculated based on the following formula. and the cost of restarting the task : ; ; in, These represent the costs of the cleanup process and the task restart, respectively, with values ​​ranging from [value range missing]. The smaller the value, the better the handling path; These represent the weights of time cost and success rate cost, respectively, and can be dimensionless, for example, 0.6 and 0.4 respectively. These represent the expected time for the cleanup process and the restart task, respectively, which can be determined based on historical statistics (e.g., 5s and 45s), and are measured in seconds. These represent the maximum allowed cleanup time and restart time, respectively, and can be preset to 30s and 120s, respectively, in seconds; These represent the success rates of the historical cleanup process and the restart task, respectively, which can be obtained through platform statistics (e.g., 0.95 and 0.98), and are dimensionless.

[0136] For example, assuming a set-top box has a history cleanup success rate of 0.95% and a restart success rate of 0.98%, then the following can be calculated: , If the comparison shows that the cleanup process is less costly, then the cleanup process should be selected first.

[0137] Optionally, the handling unit can further refine the process based on the primary cause and the fault node: when the fault determination result points to the local node of the set-top box and the device-side anomaly is dominant, the cleaning process should be prioritized even if the cleaning cost is slightly higher; when the result points to the local node of the set-top box and the critical path content shows that the fault duration cannot be recovered for a long time, for example, the fault duration exceeds a preset threshold. If the fault diagnosis result includes both device-side and network-side anomalies, a cleanup process will be generated first, and the decision on whether to continue generating a restart task will be based on the receipt status after the cleanup execution. The final generated restart task or cleanup process will include a task number, device identifier, issuance time, execution method (such as "restart immediately" or "clean up background processes + cache"), and receipt requirements.

[0138] Optionally, the remote diagnostic platform can issue the handling task via a task topic using the MQTT protocol. After receiving the task, the soft probe verifies the task number and device identifier, then proceeds to the execution phase. When executing the restart task, the soft probe first writes a local execution flag and then calls the set-top box restart interface. During the cleanup process, it cleans up background processes, cache directories, installation package directories, and junk file directories item by item according to the process. After execution, the soft probe can assemble information such as the task number, device identifier, execution start time, execution end time, execution status (success / failure / partial success), and exception log location into a feedback message, which is then sent to the remote diagnostic platform via an MQTT reporting topic. Upon receiving the message, the platform updates the task registration record and binds this execution record with the fault determination result to form the execution result. If the status update after the cleanup process shows that the player's running status has been restored, the platform will terminate the task; if the fault persists, the platform will issue a restart task under the same task number.

[0139] By directly converting fault diagnosis results into executable handling tasks, and automatically selecting the optimal handling path through cost calculation and strategy rules, and ensuring that the handling process is bound to the original fault task number, a closed-loop linkage between diagnosis and handling is achieved. This seamlessly connects data parsing, intelligent judgment, and remote handling, automating and managing the diagnostic process, improving handling efficiency and accuracy, and greatly enhancing the automation level and traceability of remote fault handling.

[0140] The technical solution of this invention embodiment firstly involves obtaining the online status of the set-top box through the remote diagnostic platform based on the connection, generating a detection task based on the online status, and sending the detection task to the soft probe. The automatic generation of the detection task triggered by the online status of the device automates the diagnostic process and enables on-demand startup, avoiding resource waste caused by blindly issuing instructions and ensuring the accuracy and efficiency of the diagnostic work. Next, the soft probe continuously collects second-level detection indicators according to the detection task and transmits them back to the remote diagnostic platform. The remote diagnostic platform generates fault information based on the second-level detection indicators and assembles a fault acquisition task based on the fault information, sending it to the soft probe. Continuous collection of second-level detection indicators can capture minute fluctuations and instantaneous anomalies during device operation, greatly improving the sensitivity and timeliness of fault detection. Finally, the soft probe collects data according to the fault acquisition task to generate a diagnostic dataset, and uploads the diagnostic dataset to the remote diagnostic platform. The diagnostic dataset includes data packet files of the set-top box's service traffic, device diagnostic results, and the set-top box's real-time video stream. After confirming a fault, a targeted data collection task is automatically assembled and issued, and the collected content covers multi-dimensional information to construct a three-dimensional diagnostic dataset, making fault analysis more comprehensive, intuitive, and traceable. Finally, the remote diagnostic platform extracts key path content from the diagnostic dataset, determines the fault judgment result based on the key path content, issues a handling task to the soft probe based on the fault judgment result, and updates the task status based on the soft probe's execution result of the handling task. This achieves a complete closed loop from data extraction and fault judgment to handling task issuance and status update, which not only improves the automation level of fault handling but also ensures the real-time accuracy of task status through the execution result feedback mechanism, effectively shortens fault recovery time, ensures the timeliness of fault detection and the comprehensiveness of diagnosis, realizes the automation and traceability of the handling process, and significantly improves the efficiency and intelligence level of fault detection and handling.

[0141] Example 2 Figure 2 This is a flowchart of a remote diagnostic method for Internet TV provided in Embodiment 2 of the present invention. It further describes the process of a soft probe collecting data according to the fault acquisition task to generate a diagnostic dataset, and uploading the diagnostic dataset to the remote diagnostic platform. Specific implementation details can be found in the description of this embodiment. Technical features that are the same as or similar to those in the foregoing embodiments will not be repeated here. Figure 2 As shown, the method may specifically include: S210. The remote diagnostic platform obtains the online status of the set-top box based on the connection, generates a detection task based on the online status of the device, and sends the detection task to the soft probe.

[0142] S220. The soft probe continuously collects second-level detection indicators according to the detection task and sends them back to the remote diagnostic platform. The remote diagnostic platform generates fault information based on the second-level detection indicators and assembles a fault collection task based on the fault information and sends it to the soft probe.

[0143] S230. The soft probe performs data acquisition operations based on the data acquisition start time, data acquisition end time, and acquisition object interface in the fault acquisition task to obtain a data packet file.

[0144] Among them, the data acquisition operation can be a full-dimensional data acquisition operation performed by the soft probe on network traffic, device status, and video stream.

[0145] Optionally, a data acquisition operation can be performed using a soft probe based on the data acquisition start time, data acquisition end time, and acquisition target interface in the fault acquisition task to obtain a data packet file.

[0146] Based on the above scheme, optionally, the soft probe performs data acquisition operations according to the data acquisition start time, data acquisition end time, and acquisition object interface in the fault acquisition task, including: the data acquisition control unit in the soft probe reads the fault acquisition task, calls the data acquisition tool to perform data acquisition operations when the data acquisition start time arrives, the calling object is the acquisition object interface, and the range is the business traffic period corresponding to the data acquisition start time to the data acquisition end time, to obtain the data packet file; during the data acquisition, the video stream return unit in the soft probe starts the return according to the linkage video stream return mark in the fault acquisition task, reads the real-time picture of the set-top box and returns it to the remote diagnostic platform, the return period of the real-time picture corresponds to the data acquisition period; after the data acquisition is completed, the data acquisition control unit registers the file name, binds the task number and binds the device identifier of the data packet file, and uploads the data packet file to the remote diagnostic platform according to the upload method in the fault acquisition task.

[0147] The data acquisition control unit can be an internal unit of the software probe used to manage acquisition tasks, trigger acquisition operations, and manage data archiving. The data acquisition tool can be a built-in acquisition tool of the software probe used to capture network data packets, system operation data, and other information. The service traffic period can be the time interval corresponding to video playback or network service interaction during a fault occurrence. The video stream return unit can be a functional unit within the software probe responsible for real-time image capture, encoding, and cloud return. The linked video stream return flag is used to instruct the software probe to simultaneously start real-time video stream return from the set-top box when performing data acquisition.

[0148] Specifically, the fault acquisition task can be read first through the data acquisition control unit in the soft probe, and then the task number, device identifier, data acquisition object interface and task version number can be verified. After the verification is successful, the data acquisition tool is called to perform the data acquisition operation when the data acquisition start time arrives. The calling object is the acquisition object interface, the calling range is the business traffic period corresponding to the data acquisition start time to the data acquisition end time, and the calling result is the original data packet file.

[0149] Furthermore, before initiating the data acquisition control unit, it can first check the device's online status, network interface activity records, and remaining space in the local storage directory. When the device is online and there are service traffic records on the network interface, data acquisition is initiated. When the device's online status is fluctuating, it first enters a short-term waiting period, then reads the network interface activity records again. After reading the activity records, data acquisition is resumed. When the device's online status is abnormal and the waiting period has expired, the current fault acquisition task is written into the pending acquisition record.

[0150] During data acquisition, the video stream backhaul unit in the soft probe synchronously reads the set-top box's real-time screen. The set-top box's real-time screen is a continuous screen record corresponding to the current playback session, which may include screen switching records, black screen records, abnormal rendering records, and current channel screen records. The video stream backhaul unit starts backhaul according to the linked video stream backhaul marker, reserving a short-term screen before the data acquisition start time and retaining a short-term screen after the data acquisition end time, so that the screen time period maintains a correspondence with the original data packet file.

[0151] After the data acquisition is completed, the original data packet file can be registered by the data acquisition control unit, the task number can be bound, and the device identifier can be bound. Then, the acquisition result upload message can be sent to the remote diagnostic platform according to the upload method in the fault acquisition task. The acquisition result upload message can include information such as task number, device identifier, acquisition file location, data acquisition start time, data acquisition end time, and upload time.

[0152] For example, during live broadcast, once stuttering and screen flickering issues have been registered and fault collection tasks have been generated, the soft probe initiates tcpdump data collection according to the task number and simultaneously transmits the set-top box's real-time video, thus obtaining both network-side and video-side records during the same fault period. After completing the above processing, the soft probe organizes the collected file location, the bound task number, the device identifier, and the corresponding video time period into a fault node data packet file.

[0153] S240. The soft probe synchronously transmits the real-time image of the set-top box during the data acquisition process, and collects the device operating status of the set-top box within the corresponding data acquisition time period.

[0154] Among them, the equipment operating status can be understood as the real-time operating parameters and status of the set-top box during the fault period.

[0155] Optionally, a soft probe can be used to synchronously transmit the real-time image from the set-top box during data acquisition, and to collect the device operating status of the set-top box during the corresponding data acquisition period.

[0156] Based on the above scheme, optionally, the soft probe collects the device operating status of the set-top box during the corresponding data acquisition period, including: the device diagnostic unit in the soft probe reads the operating status of the set-top box during the corresponding period from the start time of data acquisition to the end time of data acquisition, and obtains device diagnostic results, the device diagnostic results including central processing unit information diagnostic records, memory information diagnostic records and storage information diagnostic records.

[0157] The device diagnostic unit is a status reading module located locally on the set-top box, used to read the operating status of the set-top box during the period from the start time of data collection to the end time of data collection, and output the device diagnostic results.

[0158] Specifically, the device diagnostic unit can read the set-top box's operating status during the period corresponding to the start and end times of data collection. First, it can perform a Central Processing Unit (CPU) information diagnostic, which may include records of CPU usage, foreground process usage, and background process usage. The device diagnostic unit reads these records segment by segment according to the collection period, mapping the current playback process, system process, and background process to their respective time periods. Next, it performs a memory information diagnostic. The memory information diagnostic record may include records of total memory usage, service process memory usage, and cache usage. After reading these records, the device diagnostic unit aligns them with the collection period in the fault node's data packet file to obtain the memory activity for each time period. Finally, it performs a storage information diagnostic, which may include records of system partition usage, service directory usage, and cache directory usage, checking for any abnormal storage directory writes, temporary file accumulation, or sudden increases in cache directory usage during the collection period.

[0159] Furthermore, when reading CPU usage records, memory usage records, and storage usage records, the device diagnostic unit can also read local time records and task numbers, binding status items within the same acquisition period to the same task number; when a status item fails to be read, the failed item, failure time, and failure reason can be written into the status supplementary acquisition record, and the remaining diagnostic items can continue to be executed in the current cycle, pending supplementary recording in the next local diagnostic cycle.

[0160] To facilitate subsequent retrieval, all diagnostic records can be uniformly organized. Specifically, the device diagnostic unit can merge CPU usage records, foreground process usage records, background process usage records, total memory usage records, business process memory usage records, cache usage records, system partition usage records, business directory usage records, and cache directory usage records into a single diagnostic record based on task number, device identifier, and collection period, and then write it to the local diagnostic cache and the platform-side diagnostic receiving area.

[0161] For example, after a set-top box experiences a screen flickering fault, a soft probe can be used to obtain the data packets of the faulty node. Then, during the corresponding collection period, CPU usage records, service process memory usage records, and cache directory usage records are read. Assuming the reading results show that the player process usage is consistently high and the cache directory shows a short-term surge, the device diagnostic unit can bind these records to the same task number to form the local status data required for subsequent diagnosis. Finally, the device diagnostic unit can organize the bound CPU information diagnostic records, memory information diagnostic records, and storage information diagnostic records into a device diagnostic result.

[0162] S250, the soft probe merges the data packet file, the real-time video, and the device operating status according to the task number, generates a diagnostic dataset, and uploads it to the remote diagnostic platform. The diagnostic dataset includes the data packet file of the set-top box's service traffic, the device diagnostic results, and the set-top box's real-time video stream.

[0163] Optionally, background process scanning, network diagnostics, and streaming service test diagnostics can be performed based on the device diagnostic results to generate a diagnostic dataset.

[0164] Based on the above scheme, optionally, the soft probe merges the data packet file, the real-time image, and the device operating status according to the task number to generate a diagnostic dataset, including: the diagnostic linkage unit in the soft probe performs background process scanning, network diagnosis, and streaming service test diagnosis under the same task number to obtain background process scanning results, network diagnosis results, and streaming service test diagnosis results; the diagnostic linkage unit merges the data packet file, the real-time image, the device diagnosis results, the background process scanning results, the network diagnosis results, and the streaming service test diagnosis results according to the task number, according to the device identifier, according to the collection time period, and according to the image time period to generate the diagnostic dataset and upload it to the remote diagnostic platform.

[0165] The diagnostic linkage unit is used to continue performing background process scanning, network diagnostics, and flow service test diagnostics under the same task number, and to merge device-side records and network-side records to generate a diagnostic dataset.

[0166] Specifically, the diagnostic linkage unit can first perform a background process scan, which may include reading the current background process table of the set-top box, comparing the background process activity time period, registering the background process continuous resident record, and registering the background process abnormal restart record. The background process activity time period can be consistent with the collection time period and CPU information diagnostic record, so that the background process scan result can be directly mapped to the fault node data packet file. Next, network diagnosis is performed, which may include reading the network interface connectivity status, reading the session connection status, and reading the service traffic arrival status. When performing network diagnosis, the diagnostic linkage unit does not regenerate a new task number, but directly uses the aforementioned task number, so that the network diagnosis, device diagnosis result, and fault node data packet are on the same link. Finally, streaming service test diagnosis can be performed, which can be a diagnostic operation to locally verify the streaming service access process corresponding to the current playback session. The diagnostic content may include whether there is an interruption record for streaming service access, whether there is a pause record for streaming service switching, and whether there is an abnormal delay record for streaming service recovery.

[0167] Furthermore, after the diagnostic linkage unit completes the background process scanning, network diagnostics, and streaming service test diagnostics, all records can be merged. The input for merging can be the device diagnostic results, background process scanning results, network diagnostic results, and streaming service test diagnostic results. The merging actions can include merging by task number, by device identifier, by collection time period, and by screen time period. For multiple records under the same task number, the diagnostic linkage unit can rearrange them according to the order from the collection start time to the collection end time, and then write the corresponding background process records, network interface records, and streaming service records into a unified data structure to generate a diagnostic dataset. The diagnostic dataset includes at least the task number, device identifier, fault node data packet location, CPU information diagnostic records, memory information diagnostic records, storage information diagnostic records, background process scanning results, network diagnostic results, and streaming service test diagnostic results.

[0168] For example, in a live streaming buffering scenario, after generating fault node data packets and device diagnostic results, the diagnostic linkage unit can continue to read the set-top box's background process table, find that a certain background process is continuously resident during the fault period, and read the network interface connectivity status and streaming service access records, find that there are streaming service recovery delay records in the same period. Then, all these records are merged into the same task number to form a diagnostic dataset that can be directly called by the platform.

[0169] Based on the fault node data packet file, not only are CPU information diagnosis, memory information diagnosis, and storage information diagnosis added, but background process scanning, network diagnosis, and flow service test diagnosis are also merged into the same task number and the same collection period, forming a diagnostic dataset that can be directly sent to the platform for parsing. Compared with the existing technology of storing collection records, device status records, and business side records in a scattered manner, this makes the subsequent extraction of critical path content have a correspondence with the same source time period, the same source task, and the same source device. The platform does not need to splice multiple diagnostic records from inconsistent sources.

[0170] Driven by fault acquisition tasks, heterogeneous data from multiple sources, such as tcpdump data acquisition results, real-time video streams from set-top boxes, CPU / memory / storage diagnostic information, background process scans, network diagnostics, and streaming service test diagnostic results, are uniformly organized under the same task number to form a structured diagnostic dataset. This aligns fault clues from the device side, network side, and business side in time and space, solving the pain point of scattered and difficult-to-correlate diagnostic data in existing technologies, and providing a complete and consistent data foundation for subsequent automated analysis and judgment.

[0171] S260. The remote diagnostic platform extracts critical path content based on the diagnostic dataset, determines the fault determination result based on the critical path content, issues a handling task to the soft probe based on the fault determination result, and updates the task status based on the execution result of the handling task by the soft probe.

[0172] The technical solution of this invention firstly involves the soft probe performing data acquisition operations based on the start and end times of the data acquisition in the fault acquisition task, along with the acquisition target interface, to obtain a data packet file. By clearly defining the start and end times of data acquisition and the acquisition target interface, precise targeting and on-demand execution of the fault acquisition operation are achieved, avoiding redundant capture of invalid data and ensuring that the acquired data packet file has high relevance and effectiveness. Next, the soft probe synchronously transmits the real-time image from the set-top box during data acquisition, collecting the device operating status of the set-top box within the corresponding data acquisition period. Simultaneously, real-time images are transmitted and the device operating status is recorded while data is being acquired. This approach enables multi-dimensional, three-dimensional data acquisition. By employing a spatiotemporal synchronization mechanism for multi-source data, it overcomes the limitations of single-dimensional data, providing rich and intuitive contextual information for subsequent cross-validation of fault causes. Finally, the soft probe merges the data packet file, real-time video feed, and device operating status according to task numbers, generating a diagnostic dataset which is then uploaded to the remote diagnostic platform. By merging and packaging the data packet, real-time video feed, and device operating status according to unified task numbers, a highly structured and ordered management of heterogeneous data is achieved. This facilitates parsing and correlation analysis by the remote diagnostic platform, ensuring the integrity and auditability of the fault tracing chain and laying a solid data foundation for intelligent diagnosis and rapid decision-making by the remote diagnostic platform.

[0173] Example 3 Figure 3 This is a schematic diagram of an Internet television system according to Embodiment 3 of the present invention, including a set-top box and a remote diagnostic platform. The set-top box is equipped with a soft probe, and the set-top box and the remote diagnostic platform establish a message queue telemetry transmission protocol connection through the soft probe. Figure 3 As shown, where, The remote diagnostic platform obtains the online status of the set-top box based on the connection, generates a detection task based on the online status of the device, and sends the detection task to the soft probe. The soft probe continuously collects second-level detection indicators according to the detection task and sends them back to the remote diagnostic platform. The remote diagnostic platform generates fault information based on the second-level detection indicators and assembles a fault collection task based on the fault information and sends it to the soft probe. The soft probe collects data according to the fault collection task to generate a diagnostic dataset, and uploads the diagnostic dataset to the remote diagnostic platform. The diagnostic dataset includes data packet files of the set-top box's service traffic, device diagnostic results, and the set-top box's real-time video stream. The remote diagnostic platform extracts critical path content from the diagnostic dataset, determines the fault determination result based on the critical path content, issues a handling task to the soft probe based on the fault determination result, and updates the task status based on the execution result of the handling task by the soft probe.

[0174] Based on the above solution, optionally, the remote diagnostic platform obtains the online status of the set-top box based on the connection, including: After establishing a message queue telemetry transmission protocol connection between the set-top box and the remote diagnostic platform based on the soft probe, the set-top box sends a device identification registration message to the remote diagnostic platform. The device identification registration message carries the set-top box's basic device fields, the soft probe version number, and the network address. After completing the device identification registration based on the device identification registration message, the remote diagnostic platform sends the registration result of the set-top box back to the soft probe; The soft probe compiles the connection status, platform session identifier, and registration result into real-time connection information, and periodically sends heartbeat signals to the remote diagnostic platform based on the real-time connection information, so that the remote diagnostic platform can generate the device online status of the set-top box according to the reception status of the heartbeat signals.

[0175] Based on the above solution, optionally, the online status of the device includes fluctuating online status; the method further includes: When the soft probe detects a connection confirmation timeout, topic subscription failure, or heartbeat sending failure, the soft probe performs a local connection maintenance operation. If no confirmation is received from the remote diagnostic platform after the local connection maintenance operation is completed, it enters an interval reconnection according to the reconnection parameters. The local connection maintenance operation includes resending the connection request, resubscribing to the task topic, and restoring the reporting topic. When the device's online status is fluctuating, the remote diagnostic platform retains the platform session identifier and marks the last valid time. After the remote diagnostic platform re-establishes a connection with the soft probe, it re-issues the detection tasks that were not completed due to the connection interruption.

[0176] Based on the above solution, optionally, the step of generating a detection task based on the online status of the device includes: Read the permission flag in the online status of the device; When the permission to send is set to allowed, the remote diagnostic platform generates a detection task and sends it to the soft probe via a task topic using the message queue telemetry transmission protocol. When the permission to send is marked as pending recovery, the remote diagnostic platform generates a detection task but does not send it immediately. It sends the task only after the soft probe is reconnected. When the allowed distribution flag is set to prohibited, the remote diagnostic platform will register the detection task as pending distribution without generating task content or performing the distribution operation.

[0177] Based on the above scheme, optionally, the detection task includes a task number, a device identifier, and a collection item configuration, wherein the collection item configuration is used to indicate multiple target collection items to be collected; the soft probe continuously collects second-level detection indicators according to the detection task and transmits them back to the remote diagnostic platform, including: The soft probe maps the collection items in the detection task to the data acquisition entry corresponding to the set-top box, and performs second-level detection index collection on multiple target collection items at a fixed period. After each acquisition is completed, the soft probe assembles the attribute values ​​of multiple target acquisition items acquired in the current cycle into a second-level reporting record according to the task number, device identifier and acquisition time of the detection task, and sends the second-level reporting record to the remote diagnostic platform through the message queue telemetry transmission protocol.

[0178] Based on the above solution, optionally, the remote diagnostic platform generates fault information according to the second-level detection indicators, including: The soft probe performs stutter detection on the second-level detection indicators. The stutter detection involves a joint check of the player's running status, network reception status, and process activity status in the second-level detection indicators. It extracts continuous playback session records, data reception pause records, and process scheduling stagnation records from the second-level detection indicators and aligns them within the same period. If playback session pause, data reception pause, and no progress record of the foreground process occur simultaneously within the same period, the period is marked as a suspected stutter period. If multiple adjacent periods continuously show suspected stutter periods, the time period corresponding to these multiple periods is registered as a stutter fault time period. The soft probe performs screen flickering detection on the second-level detection indicators. The screen flickering detection involves a joint check of the screen display status and player operation status in the second-level detection indicators. Abnormal rendering markers, black screen markers, playback session switching records, and abnormal display frame records are extracted from the second-level detection indicators and matched according to the playback period. If abnormal rendering markers and abnormal display frame records appear consecutively in the same playback period, they are marked as suspected screen flickering periods. If the suspected screen flickering periods do not disappear across adjacent collection cycles, the period is registered as a screen flickering fault period. After the lag detection and the screen distortion detection are completed, the soft probe determines the fault type based on the presence and start and end times of the lag fault period and the screen distortion fault period. The fault type includes lag fault, screen distortion fault and concurrent fault. The concurrent fault is when lag fault period and screen distortion fault period exist at the same time. The soft probe assembles a fault record according to the task number, device identifier, fault type, fault start time, fault end time, associated second-level detection indicators, and current device online status. The fault record is written as the fault information into the local fault cache and uploaded to the fault registration table of the remote diagnostic platform.

[0179] Based on the above solution, optionally, the remote diagnostic platform assembles a fault acquisition task according to the fault information and sends it to the soft probe, specifically including: The soft probe generates data acquisition conditions based on the fault information. The data acquisition conditions include time period conditions and traffic threshold conditions. The time period conditions define the calculation rules for the start time and end time of data acquisition, and the traffic threshold conditions define the network traffic determination rules for initiating data acquisition. The soft probe calculates the data acquisition start time and data acquisition end time based on the time period conditions, and registers the acquisition object interface and acquisition object session; The remote diagnostic platform receives the data acquisition conditions, the data acquisition start time, the data acquisition end time, the acquisition object interface, and the acquisition object session, and combines them with the task number, device identifier, and current device online status to assemble the fault acquisition task and send it to the soft probe.

[0180] Based on the above scheme, optionally, the soft probe collects data according to the fault acquisition task to generate a diagnostic dataset, and uploads the diagnostic dataset to the remote diagnostic platform, including: The soft probe performs data acquisition operations based on the data acquisition start time, data acquisition end time, and acquisition target interface in the fault acquisition task to obtain a data packet file; The soft probe synchronously transmits the real-time image of the set-top box during the data acquisition process, and collects the device operating status of the set-top box during the corresponding data acquisition period. The soft probe merges the data packet file, the real-time video, and the device operating status according to the task number, generates a diagnostic dataset, and uploads it to the remote diagnostic platform.

[0181] Based on the above scheme, optionally, the soft probe performs data acquisition operations according to the data acquisition start time, data acquisition end time, and acquisition object interface in the fault acquisition task, including: The data acquisition control unit in the soft probe reads the fault acquisition task, and when the data acquisition start time arrives, it calls the data acquisition tool to perform the data acquisition operation. The calling object is the acquisition object interface, and the range is the business traffic period corresponding to the data acquisition start time to the data acquisition end time, and the data packet file is obtained. During data acquisition, the video stream backhaul unit in the soft probe starts backhaul according to the linkage video stream backhaul mark in the fault acquisition task, reads the real-time image of the set-top box and backhauls it to the remote diagnostic platform. The backhaul time of the real-time image corresponds to the data acquisition time. After data acquisition is completed, the data acquisition control unit registers the file name, binds the task number and device identifier of the data packet file, and uploads the data packet file to the remote diagnostic platform according to the upload method in the fault acquisition task.

[0182] Based on the above scheme, optionally, the soft probe collects the device operating status of the set-top box during the corresponding data acquisition period, and merges the data packet file, the real-time screen, and the device operating status according to the task number to generate a diagnostic dataset, including: The device diagnostic unit in the soft probe reads the operating status of the set-top box during the period from the start time of data acquisition to the end time of data acquisition, and obtains the device diagnostic results, which include central processing unit information diagnostic records, memory information diagnostic records, and storage information diagnostic records. The diagnostic linkage unit in the soft probe performs background process scanning, network diagnosis, and streaming service test diagnosis under the same task number, and obtains background process scanning results, network diagnosis results, and streaming service test diagnosis results. The diagnostic linkage unit merges the data packet file, the real-time screen, the device diagnostic results, the background process scan results, the network diagnostic results, and the streaming service test diagnostic results by task number, by device identifier, by collection time period, and by screen time period to generate the diagnostic dataset and uploads it to the remote diagnostic platform.

[0183] Based on the above scheme, optionally, the remote diagnostic platform extracts key path content according to the diagnostic dataset, including: The parsing unit in the remote diagnostic platform reads network interaction records from the data packet file in the diagnostic dataset, aligns them with the network diagnostic results in the diagnostic dataset, and obtains fault link records arranged in chronological order. The parsing unit filters out the connection segment, dwell segment, and recovery segment corresponding to the fault condition from the fault link record, and binds the start time, end time, session status, and corresponding device diagnostic record of each segment to generate critical path content.

[0184] Based on the above scheme, optionally, the parsing unit reads network interaction records from the data packet file in the diagnostic dataset and aligns them with the network diagnostic results in the diagnostic dataset, including: The parsing unit calls the data packet file in the diagnostic dataset according to the task number, and establishes a data reading window for the same fault period according to the data acquisition start time, data acquisition end time and corresponding screen time period; The parsing unit reads service connection records, session establishment records, traffic dwell records and abnormal interruption records from the data packet file, and aligns them with the network interface connectivity status, session connection status and service traffic arrival status in the network diagnostic results of the diagnostic dataset. The parsing unit extracts network interaction records corresponding to the player's running status segment by segment according to the same task number, the same device identifier, and the same fault time period. It then overlays the background process scan results, central processing unit information diagnostic records, and streaming service test diagnostic results in the diagnostic dataset onto the same time period to obtain the fault link records arranged in chronological order.

[0185] Based on the above solution, optionally, the remote diagnostic platform determines the fault determination result based on the critical path content, including: The judgment unit in the remote diagnostic platform calculates the device-side anomaly dominance, network-side anomaly dominance, and flow service anomaly dominance, respectively. It obtains the probability of each cause through normalization, takes the cause corresponding to the highest probability as the primary cause, determines the fault cause category, and locates the fault node based on the primary cause and the specific records in the critical path content.

[0186] Based on the above scheme, optionally, the determination unit calculates the device-side anomaly dominance, network-side anomaly dominance, and flow service anomaly dominance respectively, obtains the probability of each cause through normalization, takes the cause corresponding to the highest probability as the primary cause, and determines the fault cause category, including: The determination unit reads the fault initiation segment, fault duration segment, and fault recovery segment in the critical path content segment by segment, and compares them with the central processing unit information diagnostic record, memory information diagnostic record, storage information diagnostic record, background process scan results, network diagnostic results, and streaming service test diagnostic results in the diagnostic dataset. For the fault duration, the device-side anomaly dominance, network-side anomaly dominance, and streaming service anomaly dominance are calculated respectively: the device-side anomaly dominance is the average of the sum of the CPU average utilization rate of each time slice within the fault duration multiplied by a first device weight, the memory utilization rate multiplied by a second device weight, and the background process anomaly flag multiplied by a third device weight, wherein the first device weight is 0.4, the second device weight is 0.3, and the third device weight is 0.3; the network-side anomaly dominance is the average of the session anomaly degree of each time slice within the fault duration; the streaming service anomaly dominance is the average of the sum of the indicator function values ​​of the streaming service recovery delay exceeding a preset threshold within the fault duration. The device-side probability, network-side probability, and flow service probability are obtained by exponential normalization. The cause corresponding to the highest probability is taken as the main cause. If the difference between the highest probability and the second highest probability is less than 0.1, it is marked as a cause of concurrency anomaly. The determination unit locates the fault node based on the main cause and the specific records in the critical path content, including: if the main cause is a device-side abnormality, the fault node is determined to be the local node of the set-top box; if the main cause is a network-side abnormality, the fault node is determined to be the network interface node; if the main cause is a streaming service abnormality, the fault node is determined to be the streaming service connection node.

[0187] Based on the above scheme, optionally, the step of issuing a handling task to the soft probe according to the fault determination result, and updating the task status according to the execution result of the handling task by the soft probe, includes: The handling unit in the remote diagnostic platform selects a handling path based on the fault cause and fault node in the fault determination result, generates a corresponding restart task or cleanup process, and sends the task topic to the soft probe through the message queue telemetry transmission protocol. After receiving and executing the restart task or the cleanup process, the soft probe assembles the execution result into a feedback message and sends it to the remote diagnostic platform. Upon receiving the message, the remote diagnostic platform updates the task registration record and binds the execution record with the fault determination result.

[0188] The technical solution of this invention embodiment firstly involves obtaining the online status of the set-top box through the remote diagnostic platform based on the connection, generating a detection task based on the online status, and sending the detection task to the soft probe. The automatic generation of the detection task triggered by the online status of the device automates the diagnostic process and enables on-demand startup, avoiding resource waste caused by blindly issuing instructions and ensuring the accuracy and efficiency of the diagnostic work. Next, the soft probe continuously collects second-level detection indicators according to the detection task and transmits them back to the remote diagnostic platform. The remote diagnostic platform generates fault information based on the second-level detection indicators and assembles a fault acquisition task based on the fault information, sending it to the soft probe. Continuous collection of second-level detection indicators can capture minute fluctuations and instantaneous anomalies during device operation, greatly improving the sensitivity and timeliness of fault detection. Finally, the soft probe collects data according to the fault acquisition task to generate a diagnostic dataset, and uploads the diagnostic dataset to the remote diagnostic platform. The diagnostic dataset includes data packet files of the set-top box's service traffic, device diagnostic results, and the set-top box's real-time video stream. After confirming a fault, a targeted data collection task is automatically assembled and issued, and the collected content covers multi-dimensional information to construct a three-dimensional diagnostic dataset, making fault analysis more comprehensive, intuitive, and traceable. Finally, the remote diagnostic platform extracts key path content from the diagnostic dataset, determines the fault judgment result based on the key path content, issues a handling task to the soft probe based on the fault judgment result, and updates the task status based on the soft probe's execution result of the handling task. This achieves a complete closed loop from data extraction and fault judgment to handling task issuance and status update, which not only improves the automation level of fault handling but also ensures the real-time accuracy of task status through the execution result feedback mechanism, effectively shortens fault recovery time, ensures the timeliness of fault detection and the comprehensiveness of diagnosis, realizes the automation and traceability of the handling process, and significantly improves the efficiency and intelligence level of fault detection and handling.

[0189] Example 4 Figure 4 This is a schematic diagram of a remote diagnostic device for Internet TV provided in Embodiment 4 of the present invention. It is applied to an Internet TV system, which includes a set-top box and a remote diagnostic platform. A soft probe is deployed in the set-top box, and the set-top box and the remote diagnostic platform establish a message queue telemetry transmission protocol connection through the soft probe. This device is used to execute the remote diagnostic method for Internet TV provided in any of the above embodiments. This device and the remote diagnostic methods for Internet TV in the above embodiments belong to the same inventive concept. Details not described in detail in the embodiments of the remote diagnostic device for Internet TV can be referred to the embodiments of the remote diagnostic methods for Internet TV described above. Figure 4 As shown, the device includes: a detection task generation module 410, a fault information generation module 420, a diagnostic dataset generation module 430, and a fault determination module 440.

[0190] The system includes: a detection task generation module 410, used by the remote diagnostic platform to obtain the online status of the set-top box based on the connection, generate a detection task based on the online status, and send the detection task to the soft probe; a fault information generation module 420, used by the soft probe to continuously collect second-level detection indicators according to the detection task and send them back to the remote diagnostic platform, the remote diagnostic platform to generate fault information based on the second-level detection indicators, and assemble a fault collection task based on the fault information and send it to the soft probe; a diagnostic dataset generation module 430, used by the soft probe to collect data according to the fault collection task to generate a diagnostic dataset, and upload the diagnostic dataset to the remote diagnostic platform, the diagnostic dataset including data packet files of the set-top box's service traffic, device diagnostic results, and the set-top box's real-time video stream; and a fault determination module 440, used by the remote diagnostic platform to extract critical path content based on the diagnostic dataset, determine the fault determination result based on the critical path content, send a handling task to the soft probe based on the fault determination result, and update the task status based on the soft probe's execution result of the handling task.

[0191] The technical solution of this invention embodiment is as follows: First, the detection task generation module 410 obtains the online status of the set-top box through the remote diagnostic platform based on the connection, generates a detection task based on the online status, and sends the detection task to the soft probe. The automatic generation of the detection task is triggered by the online status of the device, realizing the automation and on-demand start of the diagnostic process, avoiding the waste of resources caused by blindly issuing instructions, and ensuring the accuracy and efficiency of the diagnostic work. Next, the fault information generation module 420 continuously collects second-level detection indicators through the soft probe according to the detection task and sends them back to the remote diagnostic platform. The remote diagnostic platform generates fault information based on the second-level detection indicators and assembles a fault acquisition task based on the fault information and sends it to the soft probe. By continuously collecting second-level detection indicators, minute fluctuations and instantaneous anomalies during device operation can be captured, greatly improving the sensitivity and timeliness of fault detection. Then, the diagnostic dataset generation module 430 collects data through the soft probe according to the fault acquisition task to generate a diagnostic dataset, and uploads the diagnostic dataset to... The diagnostic dataset, transmitted to the remote diagnostic platform, includes data packet files of the set-top box's service traffic, device diagnostic results, and the set-top box's real-time video stream. Upon fault confirmation, a targeted data collection task is automatically assembled and issued, with the collected content covering multi-dimensional information to construct a comprehensive diagnostic dataset, making fault analysis more comprehensive, intuitive, and traceable. Finally, the fault determination module 440 extracts key path content from the diagnostic dataset through the remote diagnostic platform, determines the fault determination result based on the key path content, issues a handling task to the soft probe based on the fault determination result, and updates the task status based on the soft probe's execution result of the handling task. This achieves a complete closed loop from data extraction and fault determination to handling task issuance and status update, not only improving the automation level of fault handling but also ensuring the real-time accuracy of the task status through the execution result feedback mechanism, effectively shortening fault recovery time, ensuring the timeliness of fault detection and the comprehensiveness of diagnosis, realizing the automation and traceability of the handling process, and significantly improving the efficiency and intelligence level of fault detection and handling.

[0192] Based on the above scheme, optionally, the detection task generation module 410 includes: a message sending submodule, a registration result feedback submodule, and a status generation submodule. The message sending submodule is used to send a device identification registration message of the set-top box to the remote diagnostic platform after establishing a message queue telemetry transmission protocol connection between the set-top box and the remote diagnostic platform based on the soft probe. The device identification registration message carries the device basic fields of the set-top box, the soft probe version number, and the network address. The registration result feedback submodule is used for the remote diagnostic platform to send the registration result of the set-top box back to the soft probe after completing the device identification registration based on the device identification registration message. The status generation submodule is used for the soft probe to organize the connection status, platform session identifier, and registration result into real-time connection information, and periodically send heartbeat signals to the remote diagnostic platform based on the real-time connection information, so that the remote diagnostic platform can generate the device online status of the set-top box according to the reception status of the heartbeat signals.

[0193] Based on the above scheme, optionally, the online status of the device includes fluctuating online status; the device further includes: a connection maintenance module and an interruption delivery module. The connection maintenance module is used to perform a local connection maintenance operation when the soft probe detects a connection confirmation timeout, topic subscription failure, or heartbeat transmission failure. If no confirmation is received from the remote diagnostic platform after the local connection maintenance operation is completed, an interval reconnection is initiated according to reconnection parameters. The local connection maintenance operation includes resending the connection request, resubscribing to the task topic, and restoring the reporting topic. The interruption delivery module is used to, when the device's online status is fluctuating online status, have the remote diagnostic platform retain the platform session identifier and mark the last valid time, wait for the remote diagnostic platform and the soft probe to re-establish a connection, and then re-deliver the detection tasks that were not completed due to the connection interruption.

[0194] Based on the above scheme, optionally, the detection task generation module 410 includes: a distribution tag reading submodule, a first distribution submodule, a second distribution submodule, and a third distribution submodule. The distribution tag reading submodule is used to read the allow distribution tag in the online status of the device; the first distribution submodule is used to generate a detection task and distribute it to the soft probe via a task topic through a message queue telemetry transmission protocol when the allow distribution tag is allowed; the second distribution submodule is used to generate a detection task but not send it immediately when the allow distribution tag is pending recovery, waiting for the soft probe to reconnect before distribution; the third distribution submodule is used to register the detection task as pending distribution without generating task content and without performing a distribution operation when the allow distribution tag is prohibited.

[0195] Based on the above scheme, optionally, the detection task includes a task number, device identifier, and collection item configuration, wherein the collection item configuration is used to indicate multiple target collection items to be collected; the fault information generation module 420 includes: an indicator collection submodule and a reporting record assembly submodule. The indicator collection submodule is used by the soft probe to map the collection item configuration in the detection task to the data acquisition entry corresponding to the set-top box, and to perform second-level detection indicator collection on multiple target collection items at a fixed period; the reporting record assembly submodule is executed after each collection, wherein the soft probe assembles the attribute values ​​of the multiple target collection items collected in the current period into a second-level reporting record according to the task number, device identifier, and collection time of the detection task, and sends the second-level reporting record to the remote diagnostic platform through a message queue telemetry transmission protocol.

[0196] Based on the above scheme, optionally, the fault information generation module 420 includes: a stuttering detection submodule, a screen flickering detection submodule, a fault type determination submodule, and a fault record assembly submodule. The stuttering detection submodule is used by the soft probe to perform stuttering detection on the second-level detection indicators. The stuttering detection involves a joint check of the player's running status, network reception status, and process activity status in the second-level detection indicators. It extracts continuous playback session records, data reception pause records, and process scheduling pause records from the second-level detection indicators and aligns them within the same period. If playback session pauses, data reception pauses, and no foreground process progress records occur simultaneously within the same period, that period is marked as a suspected stuttering period. If multiple adjacent periods continuously exhibit suspected stuttering periods, the time periods corresponding to these multiple periods are registered as stuttering fault time periods. The screen flickering detection submodule is used by the soft probe to perform screen flickering detection on the second-level detection indicators. The screen flickering detection involves a joint check of the screen display status and player running status in the second-level detection indicators. It extracts abnormal rendering markers, black screen markers, playback session switching records, and display frame abnormal records from the second-level detection indicators and aligns them within the same period. The system matches playback time periods. If abnormal rendering markers and abnormal display frame records appear consecutively within the same playback time period, it is marked as a suspected screen distortion period. If the suspected screen distortion period persists across adjacent acquisition cycles, it is registered as a screen distortion fault period. The fault type determination submodule is used by the soft probe to determine the fault type based on the presence and start and end times of the stuttering fault period and the screen distortion fault period after the stuttering detection and the screen distortion detection are completed. The fault types include stuttering fault, screen distortion fault, and concurrent fault. The concurrent fault is when stuttering fault and screen distortion fault periods exist simultaneously within the same time period. The fault record assembly submodule is used by the soft probe to assemble fault records according to the task number, device identifier, fault type, fault start time, fault end time, associated second-level detection indicators, and current device online status. The fault records are written as fault information into the local fault cache and uploaded to the fault registration table of the remote diagnostic platform.

[0197] Based on the above scheme, optionally, the fault information generation module 420 specifically includes: a data acquisition condition generation submodule, a time calculation submodule, and a data acquisition task assembly submodule. The data acquisition condition generation submodule is used by the soft probe to generate data acquisition conditions based on the fault information. These data acquisition conditions include time period conditions and traffic threshold conditions. The time period conditions define the calculation rules for the data acquisition start time and data acquisition end time, and the traffic threshold conditions define the network traffic determination rules for initiating data acquisition. The time calculation submodule is used by the soft probe to calculate the data acquisition start time and data acquisition end time based on the time period conditions, and to register the acquisition object interface and acquisition object session. The data acquisition task assembly submodule is used by the remote diagnostic platform to receive the data acquisition conditions, the data acquisition start time, the data acquisition end time, the acquisition object interface, and the acquisition object session, and to assemble them into the fault acquisition task by combining the task number, device identifier, and current device online status, and then send it to the soft probe.

[0198] Based on the above scheme, optionally, the diagnostic dataset generation module 430 includes: a data acquisition submodule, a running status acquisition submodule, and a dataset generation submodule. The data acquisition submodule is used by the soft probe to perform data acquisition operations according to the data acquisition start time, data acquisition end time, and acquisition object interface in the fault acquisition task, obtaining a data packet file; the running status acquisition submodule is used by the soft probe to synchronously transmit the real-time image of the set-top box during data acquisition, and to collect the device running status of the set-top box within the corresponding data acquisition time period; the dataset generation submodule is used by the soft probe to merge the data packet file, the real-time image, and the device running status according to the task number, generate a diagnostic dataset, and upload it to the remote diagnostic platform.

[0199] Based on the above scheme, optionally, the data acquisition submodule includes: a data acquisition unit, a return transmission unit, and a data packet file upload unit. The data acquisition unit is used by the data acquisition control unit in the soft probe to read the fault acquisition task, and when the data acquisition start time arrives, to call the data acquisition tool to perform data acquisition operations. The calling object is the acquisition object interface, and the range is the business traffic period corresponding to the data acquisition start time to the data acquisition end time, to obtain the data packet file. The return transmission unit is used by the video stream return transmission unit in the soft probe to start return transmission according to the linkage video stream return transmission mark in the fault acquisition task during data acquisition, and to read the real-time image of the set-top box and transmit it back to the remote diagnostic platform. The return transmission period of the real-time image corresponds to the data acquisition period. The data packet file upload unit is used by the data acquisition control unit to register the filename, bind the task number, and bind the device identifier of the data packet file after data acquisition is completed, and to upload the data packet file to the remote diagnostic platform according to the upload method in the fault acquisition task.

[0200] Based on the above solution, optionally, the operating status acquisition submodule includes: an operating status acquisition unit. The operating status acquisition unit is used by the device diagnostic unit in the soft probe to read the operating status of the set-top box during the time period corresponding to the data acquisition start time to the data acquisition end time, and obtain device diagnostic results. The device diagnostic results include central processing unit information diagnostic records, memory information diagnostic records, and storage information diagnostic records. Based on the above scheme, optionally, the dataset generation submodule includes: a diagnostic unit and a merging unit. The diagnostic unit is used by the diagnostic linkage unit in the soft probe to perform background process scanning, network diagnostics, and streaming service test diagnostics under the same task number, obtaining background process scanning results, network diagnostic results, and streaming service test diagnostic results. The merging unit is used by the diagnostic linkage unit to merge the data packet file, the real-time screen, the device diagnostic results, the background process scanning results, the network diagnostic results, and the streaming service test diagnostic results by task number, by device identifier, by collection time period, and by screen time period, generating the diagnostic dataset and uploading it to the remote diagnostic platform.

[0201] Based on the above scheme, optionally, the fault determination module 440 includes: an alignment submodule and a critical path content generation submodule. The alignment submodule is used by the parsing unit in the remote diagnostic platform to read network interaction records from the data packet files in the diagnostic dataset, align them with the network diagnostic results in the diagnostic dataset, and obtain fault link records arranged in chronological order. The critical path content generation submodule is used by the parsing unit to filter out connection segments, dwell segments, and recovery segments corresponding to the fault conditions from the fault link records, and bind the start time, end time, session state, and corresponding device diagnostic records of each segment to generate critical path content.

[0202] Based on the above scheme, optionally, the alignment submodule includes: a reading window establishment unit, an alignment unit, and a fault link record generation unit. The reading window establishment unit is used by the parsing unit to call the data packet file in the diagnostic dataset according to the task number, and to establish a data reading window for the same fault period according to the data acquisition start time, data acquisition end time, and corresponding screen time period. The alignment unit is used by the parsing unit to read service connection records, session establishment records, traffic dwell records, and abnormal interruption records from the data packet file, and to align them with the network interface connectivity status, session connection status, and service traffic arrival status in the network diagnostic results of the diagnostic dataset. The fault link record generation unit is used by the parsing unit to extract network interaction records corresponding to the player's running status segment by segment according to the same task number, the same device identifier, and the same fault period, and to superimpose the background process scan results, CPU information diagnostic records, and streaming service test diagnostic results in the diagnostic dataset onto the same time period to obtain the fault link records arranged in chronological order.

[0203] Based on the above scheme, optionally, the fault determination module 440 includes a fault node location submodule. The fault node location submodule is used by the determination unit in the remote diagnostic platform to calculate the device-side anomaly dominance, network-side anomaly dominance, and flow service anomaly dominance, respectively. It then obtains the probability of each cause through normalization, selects the cause with the highest probability as the primary cause, determines the fault cause category, and locates the fault node based on the primary cause and the specific records in the critical path content.

[0204] Based on the above scheme, optionally, the fault node location submodule includes: a comparison unit, a calculation unit, a cause marking unit, and a node determination unit. The comparison unit is used by the determination unit to read the fault initiation segment, fault duration segment, and fault recovery segment in the critical path content segment by segment, and compare them respectively with the CPU information diagnostic records, memory information diagnostic records, storage information diagnostic records, background process scan results, network diagnostic results, and streaming service test diagnostic results in the diagnostic dataset; the calculation unit is used to calculate the device-side anomaly dominance, network-side anomaly dominance, and streaming service anomaly dominance for the fault duration segment: the device-side anomaly dominance is the average of the sum of the CPU average utilization rate of each time slice within the fault duration segment multiplied by a first device weight, the memory utilization rate multiplied by a second device weight, and the background process anomaly mark multiplied by a third device weight, wherein the first device weight is 0.4, the second device weight is 0.3, and the third device weight is 0.3; The network-side anomaly dominance is the average of the session anomalies in each time slice within the fault duration; the streaming service anomaly dominance is the average of the sum of the indicator function values ​​for streaming service recovery delays exceeding a preset threshold within the fault duration; the cause marking unit is used to obtain the device-side probability, network-side probability, and streaming service probability through exponential normalization, and take the cause corresponding to the highest probability as the primary cause; if the difference between the highest probability and the second highest probability is less than 0.1, it is marked as a concurrent anomaly cause; the node determination unit is used to locate the fault node based on the primary cause and the specific records in the critical path content, including: if the primary cause is a device-side anomaly, the fault node is determined to be the set-top box local node; if the primary cause is a network-side anomaly, the fault node is determined to be the network interface node; if the primary cause is a streaming service anomaly, the fault node is determined to be the streaming service connection node.

[0205] Based on the above scheme, optionally, the fault determination module 440 includes: a fault handling submodule and a handling execution submodule. The fault handling submodule is used by the handling unit in the remote diagnostic platform to select a handling path based on the fault cause and fault node in the fault determination result, generate a corresponding restart task or cleanup process, and send the task topic to the soft probe via a message queue telemetry transmission protocol. The handling execution submodule is used by the soft probe to receive and execute the restart task or the cleanup process, assemble the execution result into a return message, and send it to the remote diagnostic platform. Upon receiving the message, the remote diagnostic platform updates the task registration record and binds the current execution record to the fault determination result.

[0206] The remote diagnostic device for Internet TV provided in this embodiment of the invention can execute the remote diagnostic method for Internet TV provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the method.

[0207] Example 5 Figure 5 A schematic diagram of an electronic device 10 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0208] like Figure 5 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0209] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0210] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as remote diagnostic methods for Internet television.

[0211] In particular, according to embodiments of the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present invention include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication unit 19, or installed from storage unit 18, or installed from ROM 12. When the computer program is executed by processor 11, it performs the functions defined in the methods of the embodiments of the present invention.

[0212] In some embodiments, the remote diagnostic method for Internet TV can be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the remote diagnostic method for Internet TV described above can be performed. Alternatively, in other embodiments, processor 11 can be configured to perform the remote diagnostic method for Internet TV by any other suitable means (e.g., by means of firmware).

[0213] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0214] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0215] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0216] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0217] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0218] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0219] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0220] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A remote diagnostic method for Internet television, characterized in that, An application is made to an internet television system, the internet television system including a set-top box and a remote diagnostic platform, wherein the set-top box is equipped with a soft probe, and the set-top box and the remote diagnostic platform establish a message queue telemetry transmission protocol connection through the soft probe, wherein the method includes: The remote diagnostic platform obtains the online status of the set-top box based on the connection, generates a detection task based on the online status of the device, and sends the detection task to the soft probe. The soft probe continuously collects second-level detection indicators according to the detection task and sends them back to the remote diagnostic platform. The remote diagnostic platform generates fault information based on the second-level detection indicators and assembles a fault collection task based on the fault information and sends it to the soft probe. The soft probe collects data according to the fault collection task to generate a diagnostic dataset, and uploads the diagnostic dataset to the remote diagnostic platform. The diagnostic dataset includes data packet files of the set-top box's service traffic, device diagnostic results, and the set-top box's real-time video stream. The remote diagnostic platform extracts critical path content from the diagnostic dataset, determines the fault determination result based on the critical path content, issues a handling task to the soft probe based on the fault determination result, and updates the task status based on the execution result of the handling task by the soft probe.

2. The method according to claim 1, characterized in that, The remote diagnostic platform obtains the online status of the set-top box based on the connection, including: After establishing a message queue telemetry transmission protocol connection between the set-top box and the remote diagnostic platform based on the soft probe, the set-top box sends a device identification registration message to the remote diagnostic platform. The device identification registration message carries the set-top box's basic device fields, the soft probe version number, and the network address. After completing the device identification registration based on the device identification registration message, the remote diagnostic platform sends the registration result of the set-top box back to the soft probe; The soft probe compiles the connection status, platform session identifier, and registration result into real-time connection information, and periodically sends heartbeat signals to the remote diagnostic platform based on the real-time connection information, so that the remote diagnostic platform can generate the device online status of the set-top box according to the reception status of the heartbeat signals.

3. The method according to claim 2, characterized in that, The online status of the device includes fluctuating online status; the method further includes: When the soft probe detects a connection confirmation timeout, topic subscription failure, or heartbeat sending failure, the soft probe performs a local connection maintenance operation. If no confirmation is received from the remote diagnostic platform after the local connection maintenance operation is completed, it enters an interval reconnection according to the reconnection parameters. The local connection maintenance operation includes resending the connection request, resubscribing to the task topic, and restoring the reporting topic. When the device's online status is fluctuating, the remote diagnostic platform retains the platform session identifier and marks the last valid time. After the remote diagnostic platform re-establishes a connection with the soft probe, it re-issues the detection tasks that were not completed due to the connection interruption.

4. The method according to claim 1, characterized in that, The generation of detection tasks based on the online status of the device includes: Read the permission flag in the online status of the device; When the permission to send is set to allowed, the remote diagnostic platform generates a detection task and sends it to the soft probe via a task topic using the message queue telemetry transmission protocol. When the permission to send is marked as pending recovery, the remote diagnostic platform generates a detection task but does not send it immediately. It sends the task only after the soft probe is reconnected. When the allowed distribution flag is set to prohibited, the remote diagnostic platform will register the detection task as pending distribution without generating task content or performing the distribution operation.

5. The method according to claim 1, characterized in that, The detection task includes a task number, a device identifier, and a collection item configuration, wherein the collection item configuration is used to indicate multiple target collection items to be collected; The soft probe continuously collects second-level detection indicators according to the detection task and transmits them back to the remote diagnostic platform, including: The soft probe maps the collection items in the detection task to the data acquisition entry corresponding to the set-top box, and performs second-level detection index collection on multiple target collection items at a fixed period. After each acquisition is completed, the soft probe assembles the attribute values ​​of multiple target acquisition items acquired in the current cycle into a second-level reporting record according to the task number, device identifier and acquisition time of the detection task, and sends the second-level reporting record to the remote diagnostic platform through the message queue telemetry transmission protocol.

6. The method according to claim 5, characterized in that, The remote diagnostic platform generates fault information based on the second-level detection indicators, including: The soft probe performs stutter detection on the second-level detection indicators. The stutter detection involves a joint check of the player's running status, network reception status, and process activity status in the second-level detection indicators. It extracts continuous playback session records, data reception pause records, and process scheduling stagnation records from the second-level detection indicators and aligns them within the same period. If playback session pause, data reception pause, and no progress record of the foreground process occur simultaneously within the same period, the period is marked as a suspected stutter period. If multiple adjacent periods continuously show suspected stutter periods, the time period corresponding to these multiple periods is registered as a stutter fault time period. The soft probe performs screen flickering detection on the second-level detection indicators. The screen flickering detection involves a joint check of the screen display status and player operation status in the second-level detection indicators. Abnormal rendering markers, black screen markers, playback session switching records, and abnormal display frame records are extracted from the second-level detection indicators and matched according to the playback period. If abnormal rendering markers and abnormal display frame records appear consecutively in the same playback period, they are marked as suspected screen flickering periods. If the suspected screen flickering periods do not disappear across adjacent collection cycles, the period is registered as a screen flickering fault period. After the lag detection and the screen distortion detection are completed, the soft probe determines the fault type based on the presence and start and end times of the lag fault period and the screen distortion fault period. The fault type includes lag fault, screen distortion fault and concurrent fault. The concurrent fault is when lag fault period and screen distortion fault period exist at the same time. The soft probe assembles a fault record according to the task number, device identifier, fault type, fault start time, fault end time, associated second-level detection indicators, and current device online status. The fault record is written as the fault information into the local fault cache and uploaded to the fault registration table of the remote diagnostic platform.

7. The method according to claim 6, characterized in that, The remote diagnostic platform assembles a fault acquisition task based on the fault information and sends it to the soft probe, specifically including: The soft probe generates data acquisition conditions based on the fault information. The data acquisition conditions include time period conditions and traffic threshold conditions. The time period conditions define the calculation rules for the start time and end time of data acquisition, and the traffic threshold conditions define the network traffic determination rules for initiating data acquisition. The soft probe calculates the data acquisition start time and data acquisition end time based on the time period conditions, and registers the acquisition object interface and acquisition object session; The remote diagnostic platform receives the data acquisition conditions, the data acquisition start time, the data acquisition end time, the acquisition object interface, and the acquisition object session, and combines them with the task number, device identifier, and current device online status to assemble the fault acquisition task and send it to the soft probe.

8. The method according to claim 1, characterized in that, The soft probe collects data according to the fault acquisition task to generate a diagnostic dataset, and uploads the diagnostic dataset to the remote diagnostic platform, including: The soft probe performs data acquisition operations based on the data acquisition start time, data acquisition end time, and acquisition target interface in the fault acquisition task to obtain a data packet file; The soft probe synchronously transmits the real-time image of the set-top box during the data acquisition process, and collects the device operating status of the set-top box during the corresponding data acquisition period. The soft probe merges the data packet file, the real-time video, and the device operating status according to the task number, generates a diagnostic dataset, and uploads it to the remote diagnostic platform.

9. The method according to claim 8, characterized in that, The soft probe performs data acquisition operations based on the data acquisition start time, data acquisition end time, and acquisition target interface in the fault acquisition task, including: The data acquisition control unit in the soft probe reads the fault acquisition task, and when the data acquisition start time arrives, it calls the data acquisition tool to perform the data acquisition operation. The calling object is the acquisition object interface, and the range is the business traffic period corresponding to the data acquisition start time to the data acquisition end time, and the data packet file is obtained. During data acquisition, the video stream backhaul unit in the soft probe starts backhaul according to the linkage video stream backhaul mark in the fault acquisition task, reads the real-time image of the set-top box and backhauls it to the remote diagnostic platform. The backhaul time of the real-time image corresponds to the data acquisition time. After data acquisition is completed, the data acquisition control unit registers the file name, binds the task number and device identifier of the data packet file, and uploads the data packet file to the remote diagnostic platform according to the upload method in the fault acquisition task.

10. The method according to claim 8, characterized in that, The soft probe collects the device operating status of the set-top box during the corresponding data acquisition period, and merges the data packet file, the real-time screen, and the device operating status according to the task number to generate a diagnostic dataset, including: The device diagnostic unit in the soft probe reads the operating status of the set-top box during the period from the start time of data acquisition to the end time of data acquisition, and obtains the device diagnostic results, which include central processing unit information diagnostic records, memory information diagnostic records, and storage information diagnostic records. The diagnostic linkage unit in the soft probe performs background process scanning, network diagnosis, and streaming service test diagnosis under the same task number, and obtains background process scanning results, network diagnosis results, and streaming service test diagnosis results. The diagnostic linkage unit merges the data packet file, the real-time screen, the device diagnostic results, the background process scan results, the network diagnostic results, and the streaming service test diagnostic results by task number, by device identifier, by collection time period, and by screen time period to generate the diagnostic dataset and uploads it to the remote diagnostic platform.

11. The method according to claim 1, characterized in that, The remote diagnostic platform extracts key path content based on the diagnostic dataset, including: The parsing unit in the remote diagnostic platform reads network interaction records from the data packet file in the diagnostic dataset, aligns them with the network diagnostic results in the diagnostic dataset, and obtains fault link records arranged in chronological order. The parsing unit filters out the connection segment, dwell segment, and recovery segment corresponding to the fault condition from the fault link record, and binds the start time, end time, session status, and corresponding device diagnostic record of each segment to generate critical path content.

12. The method according to claim 11, characterized in that, The parsing unit reads network interaction records from the data packet files in the diagnostic dataset and aligns them with the network diagnostic results in the diagnostic dataset, including: The parsing unit calls the data packet file in the diagnostic dataset according to the task number, and establishes a data reading window for the same fault period according to the data acquisition start time, data acquisition end time and corresponding screen time period; The parsing unit reads service connection records, session establishment records, traffic dwell records and abnormal interruption records from the data packet file, and aligns them with the network interface connectivity status, session connection status and service traffic arrival status in the network diagnostic results of the diagnostic dataset. The parsing unit extracts network interaction records corresponding to the player's running status segment by segment according to the same task number, the same device identifier, and the same fault time period. It then overlays the background process scan results, central processing unit information diagnostic records, and streaming service test diagnostic results in the diagnostic dataset onto the same time period to obtain the fault link records arranged in chronological order.

13. The method according to claim 1, characterized in that, The remote diagnostic platform determines the fault determination result based on the critical path content, including: The judgment unit in the remote diagnostic platform calculates the device-side anomaly dominance, network-side anomaly dominance, and flow service anomaly dominance, respectively. It obtains the probability of each cause through normalization, takes the cause corresponding to the highest probability as the primary cause, determines the fault cause category, and locates the fault node based on the primary cause and the specific records in the critical path content.

14. The method according to claim 13, characterized in that, The determination unit calculates the device-side anomaly dominance, network-side anomaly dominance, and flow service anomaly dominance, respectively. It then normalizes the probabilities of each cause, selects the cause with the highest probability as the primary cause, and determines the fault cause category, including: The determination unit reads the fault initiation segment, fault duration segment, and fault recovery segment in the critical path content segment by segment, and compares them with the central processing unit information diagnostic record, memory information diagnostic record, storage information diagnostic record, background process scan results, network diagnostic results, and streaming service test diagnostic results in the diagnostic dataset. For the fault duration, the device-side anomaly dominance, network-side anomaly dominance, and streaming service anomaly dominance are calculated respectively: the device-side anomaly dominance is the average of the sum of the CPU average utilization rate of each time slice within the fault duration multiplied by a first device weight, the memory utilization rate multiplied by a second device weight, and the background process anomaly flag multiplied by a third device weight, wherein the first device weight is 0.4, the second device weight is 0.3, and the third device weight is 0.3; the network-side anomaly dominance is the average of the session anomaly degree of each time slice within the fault duration; the streaming service anomaly dominance is the average of the sum of the indicator function values ​​of the streaming service recovery delay exceeding a preset threshold within the fault duration. The device-side probability, network-side probability, and flow service probability are obtained by exponential normalization. The cause corresponding to the highest probability is taken as the main cause. If the difference between the highest probability and the second highest probability is less than 0.1, it is marked as a cause of concurrency anomaly. The determination unit locates the fault node based on the main cause and the specific records in the critical path content, including: If the primary cause is a device-side anomaly, the fault node will be identified as the local node of the set-top box. If the primary cause is a network-side anomaly, the fault node will be identified as the network interface node. If the primary cause is a streaming service anomaly, the fault node will be identified as the streaming service connection node.

15. The method according to claim 1, characterized in that, The step of issuing a handling task to the soft probe based on the fault determination result, and updating the task status based on the execution result of the handling task by the soft probe, includes: The handling unit in the remote diagnostic platform selects a handling path based on the fault cause and fault node in the fault determination result, generates a corresponding restart task or cleanup process, and sends the task topic to the soft probe through the message queue telemetry transmission protocol. After receiving and executing the restart task or the cleanup process, the soft probe assembles the execution result into a feedback message and sends it to the remote diagnostic platform. Upon receiving the message, the remote diagnostic platform updates the task registration record and binds the execution record with the fault determination result.

16. An Internet television system, characterized in that, The system includes a set-top box and a remote diagnostic platform. The set-top box is equipped with a soft probe, and the set-top box and the remote diagnostic platform establish a message queue telemetry transmission protocol connection via the soft probe. The remote diagnostic platform obtains the online status of the set-top box based on the connection, generates a detection task based on the online status of the device, and sends the detection task to the soft probe. The soft probe continuously collects second-level detection indicators according to the detection task and sends them back to the remote diagnostic platform. The remote diagnostic platform generates fault information based on the second-level detection indicators and assembles a fault collection task based on the fault information and sends it to the soft probe. The soft probe collects data according to the fault collection task to generate a diagnostic dataset, and uploads the diagnostic dataset to the remote diagnostic platform. The diagnostic dataset includes data packet files of the set-top box's service traffic, device diagnostic results, and the set-top box's real-time video stream. The remote diagnostic platform extracts critical path content from the diagnostic dataset, determines the fault determination result based on the critical path content, issues a handling task to the soft probe based on the fault determination result, and updates the task status based on the execution result of the handling task by the soft probe.