Massive network live broadcast batch data acquisition method and system

Through the combination of the group control system, Appium framework and Scrapy-Redis framework, multi-platform synchronous data collection and multimodal analysis are achieved, which solves the problem of efficient collection and intelligent supervision of massive online live broadcast data and improves supervision efficiency and accuracy.

CN120769077APending Publication Date: 2025-10-10SHANDONG PROVINCIAL MARKET SUPERVISION & MONITORING CENT
View PDF 16 Cites 0 Cited by

Patent Information

Application Number
CN202511269383.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-08
Publication Date
2025-10-10

AI Technical Summary

Technical Problem

Existing technologies are unable to efficiently collect massive amounts of online live broadcast data. They have problems such as narrow coverage, insufficient analysis capabilities, and limitations in anti-crawl mechanisms, and are unable to meet the needs of multimodal analysis and intelligent supervision.

Method used

The group control system module is used to centrally manage multiple mobile terminal devices, and the Appium framework is combined to realize automated collection and simulate user interaction. The Scrapy-Redis framework is used to build a distributed crawler engine, and a multimodal large model is used for video understanding and semantic analysis to generate violation analysis reports.

Benefits of technology

It significantly improves data collection efficiency and analysis accuracy, can cover multiple live broadcast platforms at the same time, identify complex violations, generate structured analysis reports, and support large-scale live broadcast content supervision.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120769077A_ABST
    Figure CN120769077A_ABST
Patent Text Reader

Abstract

The invention provides a mass network live broadcast batch data acquisition method and system, and belongs to the field of data processing and information. The method comprises the steps that a plurality of mobile terminal devices are managed and controlled in a centralized mode through a group control system module, and synchronous operation of multi-platform live broadcast APPs is achieved; an automatic acquisition module is constructed based on an Appium framework, real user interaction behaviors are simulated, and metadata of a live broadcast room is captured; the method comprises the following steps: constructing a distributed crawler engine by using a Scrapy-Redis framework, analyzing a live streaming media source address in real time, and performing block storage and format conversion on a live video stream; and carrying out video understanding and semantic analysis on the live broadcast content by adopting a multi-modal large model, identifying violation behaviors, and generating a violation analysis report. According to the method, the problems of low efficiency, narrow coverage and limited analysis capability of a traditional live broadcast supervision technology are solved, and the automation level and accuracy of large-scale live broadcast content supervision are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of data processing and information technology, and in particular relates to a method and system for collecting batch data of massive network live broadcasts. Background Art

[0002] The statements in this section merely provide background information related to the present invention and do not necessarily constitute prior art.

[0003] With the rapid development of internet technology, live streaming has become an important means of communication in e-commerce, entertainment, education and other fields. Especially in the live streaming sales scenario, the massive amount of live streaming content poses a huge challenge to market supervision. Traditional live streaming content supervision technology has the following main problems: (1) Single-channel acquisition is inefficient. Traditional audio and video acquisition technologies (such as TV commercial monitoring) typically use a "one acquisition card for one signal" approach, which is unable to cope with the massive concurrency of live streaming. In a live streaming environment where everyone can become a host, traditional methods are difficult to achieve effective coverage in terms of manpower, equipment, and cost, resulting in a large number of blind spots in supervision.

[0004] (2) Anti-crawling mechanisms limit data acquisition. Existing web crawler technologies have low collection success rates and poor stability when facing anti-crawling mechanisms such as dynamic encryption, IP blocking, and verification codes on live streaming platforms. Some technologies attempt to achieve multi-channel collection by opening multiple simulators, but due to platform security policies, the simulator environment is often identified and access is restricted, resulting in data collection failures.

[0005] (3) Insufficient content analysis capabilities. Current live broadcast supervision relies mainly on manual review or single-modal analysis (such as keyword matching and image recognition), which makes it difficult to effectively identify complex violations (such as false propaganda and obscure advertising). In addition, traditional AI models lack multimodal fusion capabilities and are unable to perform joint semantic understanding of video, audio, and text, resulting in a high misjudgment rate.

[0006] In summary, existing technologies struggle to meet the demands for efficient collection, multimodal analysis, and intelligent supervision of massive amounts of live streaming data. Therefore, a comprehensive solution is urgently needed that integrates distributed data collection, multi-device collaboration, and multimodal AI analysis to enhance the breadth of coverage, depth of analysis, and efficiency of live streaming supervision. Summary of the Invention

[0007] To overcome the shortcomings of the aforementioned prior art, the present invention provides a method and system for collecting batch data from massive live webcasts. This system utilizes a cluster of group-controlled devices to achieve simultaneous data collection across multiple platforms, employs a distributed crawler architecture to process highly concurrent live streams, and integrates a multimodal AI model to enable in-depth analysis of video content. This system addresses the low efficiency, narrow coverage, and limited analytical capabilities of traditional live broadcast monitoring technologies, significantly improving the automation and accuracy of large-scale live broadcast content monitoring. It is particularly suitable for scenarios such as market regulation and content review that require real-time processing of massive amounts of live broadcast data.

[0008] To achieve the above objectives, one or more embodiments of the present invention provide the following technical solutions: The first aspect of the present invention provides a method for collecting batch data of massive network live broadcasts; A method for collecting batch data of massive network live broadcasts, comprising: Centrally control multiple mobile terminal devices through the group control system module to achieve synchronous operation of multi-platform live broadcast APP; Build an automated collection module based on the Appium framework to simulate real user interaction behaviors and capture live broadcast room metadata; Use the Scrapy-Redis framework to build a distributed crawler engine, parse the source address of live streaming media in real time, and perform block storage and format conversion on live video streams; A multimodal large model is used to perform video understanding and semantic analysis on live broadcast content, identify violations, and generate violation analysis reports.

[0009] As a further technical solution, the group control system module is used to centrally control multiple mobile terminal devices to achieve the synchronous operation of multi-platform live broadcast APP, including: Start multiple mobile devices, connect them to the test host via USB debugging cables and connect them to the network cables to ensure they are in the same local area network; Install and run the group control software on the test host to check its ability to identify and remotely operate mobile devices; Configure the network port and IP segment so that the mobile device can automatically obtain an IP address and connect to the network normally; Install the same version of the target application APK package on all mobile devices.

[0010] As a further technical solution, the automated collection module is built based on the Appium framework to simulate real user interaction behaviors and capture live broadcast room metadata, including: Use Docker container to deploy Appium, achieving integrated packaging of Appium and related dependencies; Use ADB commands to check the connectivity between the server and the mobile device; Assign a separate port number to each mobile device, start the corresponding Appium Server container and establish a connection; By simulating the user operation process, the page location and operation path of the required collected data are clarified, and an automated collection logic is formed. Based on the obtained automated collection logic, feature extraction is performed on the new application state after user interaction, and data classification and structured storage are performed according to predefined rules; by setting scheduled tasks or trigger mechanisms, collection tasks can be started on demand.

[0011] As a further technical solution, we use the Scrapy-Redis framework to build a distributed crawler engine, which can parse the source address of live streaming media in real time, store live video streams in blocks, and convert their formats, including: Deploy the Scrapy framework and FFmpeg tools; Analyze the network requests of the live broadcast platform to determine the URL structure and request parameters of the live video stream; Based on the Scrapy framework, FFmpeg tools and asynchronous programming technology, combined with the Scrapy-Redis middleware to build a distributed crawler engine; The distributed crawler engine is deployed to the server to achieve continuous monitoring and collection of live streaming media data.

[0012] As a further technical solution, the analysis of the network request of the live broadcast platform to determine the URL structure and request parameters of the live video stream includes: Use browser developer tools to capture network requests related to live streaming; Trigger the real-time network request of the live broadcast platform and record the relevant request information; Parse the structure of the live stream URL and determine the variable parameters and their generation rules; Reverse engineer request parameters containing encrypted fields.

[0013] As a further technical solution, a multimodal large model is used to perform video understanding and semantic analysis on live broadcast content, identify violations, and generate violation analysis reports, including: Build a large model prompt word template; Read the live video file to be analyzed from the mass storage module; Use a large multimodal model to analyze the information of each video segment and extract key features; The text results output by the multimodal large model are imported into the LLM large model and analyzed in combination with the domain knowledge base to determine whether there is any violation; if the violation judgment is met, an alarm message is generated and relevant information is stored for evidence preservation.

[0014] As a further technical solution, the multimodal model includes: a graphic and text multimodal service model and a language multimodal service model; wherein the graphic and text multimodal service model is used to analyze the content of live video images; the language multimodal service model is used to identify the content of live audio; The output results of the graphic and text multimodal service model and the language multimodal service model are input into the large language model for semantic-level violation judgment.

[0015] A second aspect of the present invention provides a system for collecting batch data of massive network live broadcasts.

[0016] A massive network live broadcast batch data collection system, comprising: The terminal device control module is configured to: centrally control multiple mobile terminal devices through the group control system module to achieve the synchronous operation of multi-platform live broadcast APP; The data capture module is configured to: build an automated collection module based on the Appium framework, simulate real user interaction behaviors, and capture live broadcast room metadata; The streaming media processing module is configured to: build a distributed crawler engine using the Scrapy-Redis framework, parse the source address of live streaming media in real time, and perform block storage and format conversion on the live video stream; The intelligent analysis module is configured to use a multimodal large model to perform video understanding and semantic analysis on live broadcast content, identify violations, and generate violation analysis reports.

[0017] The third aspect of the present invention provides a computer-readable storage medium having a program stored thereon, which, when executed by a processor, implements the steps of a method for collecting batch data of massive network live broadcasts as described in the first aspect of the present invention.

[0018] The fourth aspect of the present invention provides an electronic device, including a memory, a processor, and a program stored in the memory and runnable on the processor. When the processor executes the program, it implements the steps in the method for collecting batch data of massive network live broadcasts as described in the first aspect of the present invention.

[0019] One or more of the above technical solutions have the following beneficial effects: (1) The present invention realizes the centralized control and synchronous operation of multiple devices and platforms through the group control system module, which can cover the mainstream live broadcast platforms at the same time and significantly improve the efficiency of data collection. The automated collection module based on the Appium framework is used to simulate real user interaction behaviors (such as searching, clicking, sliding, etc.) to effectively circumvent anti-crawler restrictions and ensure the stability and continuity of data collection. The streaming media processing module is used to build a distributed crawler engine based on the Scrapy framework to achieve high concurrent requests and real-time analysis of live streaming media. Compared with traditional technologies, the system of the present invention has horizontal expansion capabilities, can dynamically allocate resources, support the synchronous collection and processing of thousands of live streams, and meet the timeliness requirements of large-scale data collection.

[0020] (2) The intelligent analysis module of the present invention builds an end-to-end video understanding solution based on an open-source pre-trained multimodal large model. By extracting key sentences from live audio through ASR, combining it with VLM to analyze the text and image information in the video, and then using LLM to interpret the content at a semantic level, it can accurately identify illegal advertising, false propaganda, and other behaviors, and generate structured analysis reports (such as violation type, risk level, and evidence fragments). The application of LLM significantly improves the depth and accuracy of content analysis, providing a scientific basis for decision-making for regulatory authorities.

[0021] Advantages of additional aspects of the present invention will be given in part in the following description and in part will be obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] The accompanying drawings, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.

[0023] Figure 1 This is a flow chart of the method of the first embodiment.

[0024] Figure 2 Schematic diagram of the overall method framework in the first embodiment.

[0025] Figure 3 This is a system structure diagram of the second embodiment. DETAILED DESCRIPTION

[0026] It should be noted that the following detailed descriptions are exemplary and intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the present invention belongs.

[0027] It should be noted that the terms used herein are for describing particular embodiments only and are not intended to limit the exemplary embodiments according to the present invention.

[0028] In the absence of conflict, the embodiments of the present invention and the features thereof may be combined with each other.

[0029] Example 1 This embodiment discloses a method for collecting batch data of massive network live broadcasts; like Figure 1 and Figure 2 As shown, a method for collecting batch data of massive network live broadcasts includes: Step S1: Centrally control multiple mobile terminal devices through the group control system module to achieve synchronous operation of multi-platform live broadcast APPs; Step S2: Build an automated collection module based on the Appium framework to simulate real user interaction behaviors and capture live broadcast room metadata; Step S3, using the Scrapy-Redis framework to build a distributed crawler engine, parse the source address of the live streaming media in real time, and perform block storage and format conversion on the live video stream; In step S4, a multimodal large model is used to perform video understanding and semantic analysis on the live broadcast content, identify violations, and generate a violation analysis report.

[0030] Specifically, the above steps also include the following: Step S1: Centrally control multiple mobile terminal devices through the group control system module to achieve synchronous operation of multi-platform live broadcast APPs.

[0031] Step S11, turn on the power and trigger the power-on command to start multiple mobile terminal devices required for group control; connect each mobile terminal to the test host via a USB debugging cable, and connect the network cable to ensure that all devices are in the same local area network environment.

[0032] Step S12, connect the other end of the USB debugging cable to the test host, install and run the group control software on the host; detect whether the software can normally identify and display the system interface of all connected mobile devices, and verify its function of remotely executing operations.

[0033] Step S13: Enable the network port through the group control software, configure the IP network segment, and enable each mobile terminal to automatically obtain an IP address within the local area network to achieve normal network connection of multiple mobile terminal systems and complete initialization configuration and debugging.

[0034] Step S14: To ensure UI consistency and interaction stability, the same version of the target live broadcast application APK package is installed on all mobile terminals.

[0035] Through the group control system module, multiple mobile terminal devices are centrally controlled. A technical solution combining software and hardware is adopted to simulate the real mobile system environment, avoiding network access obstacles caused by security restrictions of some applications in the simulator environment, thereby realizing the synchronous operation of multi-platform live broadcast APP.

[0036] Step S2: Build an automated collection module based on the Appium framework to simulate real user interaction behaviors and capture live broadcast room metadata.

[0037] In Appium's environment deployment, considering its reliance on components like the Android SDK, JDK, and Node.js, traditional manual deployment methods present complex configuration, frequent version conflicts, and difficulty migrating environments. In this implementation, Appium is deployed using a Docker container. The "docker pull appium / appium" command is used to download an image file from the Docker repository. This image file integrates Appium and its dependent components, such as the Android SDK, JDK, and Node.js. Containerized deployment improves environment configuration efficiency and migration flexibility, ensuring consistency and reusability across different deployment scenarios.

[0038] Use the ADB command to detect the connection status between the Linux server and all target mobile devices. If the connection is successful, the IP and port of the connected device will be returned and the next step will be entered; if the connection is unsuccessful, the IP and port will not be returned.

[0039] Assign an independent port number to each mobile device, start the corresponding Appium Server container, and establish a connection channel with the device through the IP:Port method.

[0040] Furthermore, in the process of configuring the automated collection function, first configure the required dependent libraries for operation, such as by building a Python development environment and installing necessary dependent libraries such as Appium-Python-Client. Secondly, by analyzing the user-defined data targets and simulating the user operation process, the page location and operation path of the required collected data are determined, and the automated collection logic is formed. Based on the UI structure characteristics of the target platform, the data targets are decomposed into specific UI element positioning requirements. Specifically: First, the top-level abstract "live data collection" business goal (such as "collecting live room user nickname, fan number, live goods selling status") is broken down into a series of executable and concrete interaction scenarios, each of which corresponds to a complete user operation process on the target platform, and each process points to a sub-collection requirement of the data target, ensuring that the decomposed scenarios can cover the full link of data collection. For example, if the data target is "collect the fan number of a specific live room on a certain live platform", it is broken down into the following interaction scenarios: (1) Start the target live APP and load to the application main page; (2) Enter the live room identification information (such as live room ID, anchor name) in the main page search box and perform the search operation; (3) Locate the target live room in the search result list and click to enter the live room detail page; (4) Locate the UI area that displays the fan number in the live room detail page.

[0041] For each decomposed interaction scenario, combined with the UI hierarchy of the target platform, extract all UI components related to data collection in this scenario, the UI component is a visual unit that can be perceived and interacted with by users in the target platform interface, and is directly or indirectly associated with the storage and display of the data target.

[0042] Continuing the above example, for the interaction scenario "locate the UI area that displays the fan number in the live room detail page", the decomposed UI components include: Fan number display text box (UI unit for directly displaying fan number value); Fan number label text box (UI unit for labeling "fans", "follow", etc. semantics, to assist in confirming the accuracy of the fan number display text box); Live room detail page top navigation bar (used to confirm that the current page is the live room detail page, to avoid positioning to similar UI components on other pages).

[0043] Based on the target platform UI structure characteristics (such as the uniqueness, stability, and accessibility of UI component attributes), develop positioning strategies for each UI component that adapt to its attribute characteristics, the positioning strategy needs to meet the element recognition specifications of the Appium framework, and preferentially select UI component attributes with uniqueness and low change frequency to ensure the stability and anti-interference of the positioning logic. Convert the above determined positioning strategy into executable code instructions that meet the syntax specifications of Appium-Python-Client, and configure positioning timeout time, retry mechanism and other auxiliary logic to ensure that the code instructions can be called by the automated collection module, achieving accurate identification and data extraction of UI components.

[0044] Finally, based on the Appium-Python-Client interface and the above-mentioned automated collection logic, a collection script is constructed that supports parallel operation of multiple devices and has status monitoring and fault tolerance capabilities; according to the constructed script, starting from the initial state of the application, user interaction operations are simulated or recorded, and data is extracted from the page elements after the operation. After the data is extracted, the data is preprocessed and structured according to predefined rules. In this embodiment, the extracted live room structured data can be stored in a MySQL database. Among them, the predefined rules refer to the rules for data processing and data storage, such as processing and storing the extracted "number of followers and number of fans" data. The predefined rules are to store the number of followers and the number of fans as two fields in the database as digital types. The elements extracted from the page are such as the text of "13 followers and 27,000 fans". Data processing requires extracting the number of followers and fans in the text separately, and converting them into digital types for storage, for example, 27,000 is converted to .

[0045] By utilizing thread pool or process pool management, synchronous and stable access and data extraction of multiple live broadcast pages can be achieved, and metadata such as the live broadcast room nickname, number of fans, live broadcast room access link, and whether the live broadcast is for selling goods can be captured.

[0046] Then, the constructed automated collection module is deployed on the server, and the collection task is started on demand by setting a scheduled task or trigger mechanism.

[0047] Therefore, in step S2, multiple Appium services are started through containerization technology, and Appium services can be connected to mobile devices through different ports without complex coding, which is convenient and easy.

[0048] Step S3, using the Scrapy-Redis framework to build a distributed crawler engine, parse the source address of the live streaming media in real time, and perform block storage and format conversion on the live video stream.

[0049] Step S31: Install the Scrapy framework and its dependent libraries, and deploy the FFmpeg tool to support real-time video stream processing. FFmpeg is a complete cross-platform audio and video solution that provides powerful audio and video processing capabilities. By deploying the FFmpeg tool, real-time video stream processing can be achieved.

[0050] Step S32: First, access the target live broadcast room address through the browser, and use the Network panel of the browser's built-in developer tools to set filtering conditions to capture network requests related to the live stream; trigger the real-time network request of the live broadcast platform through user operations, and record the request method, request header, request parameters, response status code and response content; secondly, perform regular expression matching on the captured live stream URL, extract fixed prefixes and dynamic parameters, and determine the variable parameters and their generation rules by comparing the request differences under different operations. If the request parameters contain encrypted fields, simulate the request and use the debugging tool to locate the parameter generation logic in the front-end JavaScript to reverse the decryption algorithm. At the same time, analyze the relevant fields in the request header to verify whether it is necessary to obtain temporary credentials through the interface and design a simulated request process.

[0051] Step S33, based on the Scrapy framework and FFmpeg tools, introduce the Scrapy-Redis middleware to realize cross-node task scheduling and deduplication management, and build a distributed crawler system; wherein, asynchronous programming is used to use the built-in Twisted asynchronous network framework in the Scrapy framework to implement non-blocking I / O operations, and the FFmpeg tool is used to asynchronously process video stream data in the video stream acquisition scenario, and the FFmpeg command is registered as an asynchronous task through the asyncio interface to realize parallel processing of video stream downloading and transcoding.

[0052] Furthermore, the task scheduling of the middleware is implemented with the help of the Redis-based scheduler and deduplication components provided by Scrapy-Redis. The Redis database is used as the central scheduler of the distributed system, and the queue of URLs to be crawled is stored in the Redis database, so that each crawler node can asynchronously pull tasks from the queue to achieve cross-node collaborative scheduling and load balancing. The Scrapy-Redis deduplication mechanism ensures the accuracy of cross-node data deduplication.

[0053] In the above process, crawling tasks (URL lists) are centrally stored in a Redis database and retrieved and executed by multiple independent crawler nodes from the Redis queue. Computationally intensive tasks such as data collection and parsing are processed in parallel by multiple crawler nodes rather than being concentrated on a single node. By having multiple nodes working in parallel, the system can simultaneously initiate and process a number of requests far exceeding the capabilities of a single machine, significantly improving the overall speed and efficiency of data collection. At the same time, since the failure of a single crawler node will not cause the entire collection task to fail, other healthy nodes can continue to work. Even if some nodes temporarily fail due to network problems or restrictions on the target website, the task queue still exists and can be continued after the node is restored or when a new node is added.

[0054] Step S34: deploying the crawler program to the server to achieve continuous monitoring and real-time collection of live streaming data, so as to parse the source address of the live streaming in real time.

[0055] In step S4, a multimodal large model is used to perform video understanding and semantic analysis on the live broadcast content, identify violations, and generate a violation analysis report.

[0056] Step S41 constructs a large, industry-specific model prompt word template. This template includes analysis steps for determining the compliance of the content to be inspected (such as identifying absolute terms, checking for false advertising, and determining whether specific industry restrictions are violated), and outputs a compliance conclusion. The live video file to be analyzed is loaded into the analysis system, which supports simultaneous acquisition of multiple video data sets, with the amount of video data configurable based on actual needs.

[0057] Step S42, multimodal large model analysis: use the pre-trained multimodal large model to analyze each video segment, and use custom prompt words to guide the model to focus on and extract target features in the video (such as illegal words, product information, behavioral characteristics, etc.) to achieve video screening; wherein, build a language multimodal service model to directly perform semantic analysis and speech recognition on the live audio segment, and identify illegal speech in the audio. wherein, the language multimodal service can be built based on the qwen2-audio-instruc-7b large model, which is used to directly perform semantic analysis and speech recognition on the audio segment; build A multimodal image-text service model is built to understand long videos, locate events within seconds, and efficiently extract key information. It can perform in-depth analysis of dynamic video content and single-frame images, identifying exaggerated promotional terms, inappropriate product descriptions, product types, product information, and other content in the video. The image-text multimodal service can be built based on the qwen2.5-vl-instruct-72b large model to understand long videos, locate events within seconds, and efficiently extract key information. This allows for in-depth analysis of videos and images, not only analyzing dynamic video content but also extracting content from single-frame images.

[0058] In step S43, the text output by the multimodal large model is imported into the LLM large model and analyzed in combination with the domain knowledge base to determine whether there is any violation. If at least one of the analysis results of the multimodal large model and the LLM large model meets the violation judgment conditions, an alarm message is generated, and the original video, analysis report, screenshots of the violation points, and other materials are stored to achieve evidence fixation.

[0059] Among them, a "baseline model + domain knowledge enhancement" technical framework is established. By building a multimodal model and a dynamic RAG architecture, professional corpus of the market supervision industry is imported in stages. Through the knowledge base, the judgment basis and similar cases related to violations can be associated and brought out.

[0060] We selected a general-purpose, large-scale multimodal model as a baseline (e.g., Qwen2-Audio-Instruct-7B for audio analysis and Qwen2.5-VL-Instruct-72B for video and image analysis) and initially adapted it for market regulation scenarios. This adaptation involved using prompt word engineering to guide the model's focus on features relevant to violations (e.g., absolute terms, false advertising statements), and constraining the model's output format to ensure it generates structured text for easier processing.

[0061] Collect professional corpus in the field of market supervision, clean and deduplicate these corpuses, and build a vectorized knowledge base to support efficient semantic retrieval.

[0062] The text output from the multimodal model (such as identified product information and promotional language) is fed into the RAG module. First, the knowledge base is searched for relevant regulatory provisions and similar cases. The search results are then combined with the multimodal output and fed into the LLM for comprehensive analysis, generating a violation determination and supporting evidence. Through efficient knowledge base retrieval and multimodal analysis, the violation is accurately linked to the legal basis and similar cases, generating a traceable basis for determination, improving the accuracy and interpretability of the system's compliance analysis.

[0063] The framework effect is tested through actual live broadcast data. The pertinence of prompt word templates, coverage of knowledge base and retrieval accuracy are improved according to the false alarm / missing alarm situation; the logical consistency of LLM violation judgment is continuously optimized.

[0064] Example 2 This embodiment discloses a system for collecting batch data of massive network live broadcasts; like Figure 3 As shown, a massive network live broadcast batch data collection system includes: The terminal device control module is configured to: centrally control multiple mobile terminal devices through the group control system module to achieve the synchronous operation of multi-platform live broadcast APP; The data capture module is configured to: build an automated collection module based on the Appium framework, simulate real user interaction behaviors, and capture live broadcast room metadata; The streaming media processing module is configured to: build a distributed crawler engine using the Scrapy-Redis framework, parse the source address of live streaming media in real time, and perform block storage and format conversion on the live video stream; The intelligent analysis module is configured to use a multimodal large model to perform video understanding and semantic analysis on live broadcast content, identify violations, and generate violation analysis reports.

[0065] Example 3 The embodiment aims to provide a computer readable storage medium.

[0066] A computer readable storage medium, having stored thereon a computer program, which, when executed by a processor, implements the steps of the mass network live batch data acquisition method according to the embodiment 1.

[0067] Embodiment four The embodiment aims to provide an electronic device.

[0068] An electronic device, comprising a memory, a processor, and a program stored in the memory and executable on the processor, wherein the processor implements the steps of the mass network live batch data acquisition method according to the embodiment 1 when executing the program.

[0069] The steps and methods involved in the devices of the above embodiments two, three and four correspond to the embodiment one, and the specific implementation can refer to the related description part of the embodiment one. The term "computer readable storage medium" should be understood as including a single medium or multiple media of one or more instruction sets; and should also be understood as including any medium capable of storing, encoding or carrying instruction sets for execution by a processor and causing the processor to perform any method in the present application.

[0070] Those skilled in the art should understand that each module or step of the above-mentioned present application can be realized by a general computer device, alternatively, they can be realized by program codes executable by a computing device, so that they can be stored in a storage device for execution by a computing device, or they can be respectively made into each integrated circuit module, or a plurality of modules or steps among them can be made into a single integrated circuit module to realize. The present application is not limited to any specific combination of hardware and software.

[0071] The above describes the specific embodiments of the present application in combination with the accompanying drawings, but is not a limitation on the protection scope of the present application, and those skilled in the art should understand that various modifications or changes made by those skilled in the art on the basis of the technical solutions of the present application without creative labor are still within the protection scope of the present application.

Claims

1. A method for collecting batch data of massive network live broadcasts, characterized in that: include: Centrally control multiple mobile terminal devices through the group control system module to achieve synchronous operation of multi-platform live broadcast APP; Build an automated collection module based on the Appium framework to simulate real user interaction behaviors and capture live broadcast room metadata; Use the Scrapy-Redis framework to build a distributed crawler engine, parse the source address of live streaming media in real time, and perform block storage and format conversion on live video streams; A multimodal large model is used to perform video understanding and semantic analysis on live broadcast content, identify violations, and generate violation analysis reports.

2. A method for collecting batch data of massive network live broadcasts according to claim 1, characterized in that: The group control system module is used to centrally control multiple mobile terminal devices to achieve the synchronous operation of multi-platform live broadcast APP, including: Start multiple mobile devices, connect them to the test host via USB debugging cables and connect them to the network cables to ensure they are in the same local area network; Install and run the group control software on the test host to check its ability to identify and remotely operate mobile devices; Configure the network port and IP segment so that the mobile device can automatically obtain an IP address and connect to the network normally; Install the same version of the target application APK package on all mobile devices.

3. A method for collecting batch data of massive network live broadcasts according to claim 1, characterized in that: The automated collection module is built based on the Appium framework to simulate real user interaction behaviors and capture live broadcast room metadata, including: Use Docker container to deploy Appium, achieving integrated packaging of Appium and related dependencies; Use ADB commands to check the connectivity between the server and the mobile device; Assign a separate port number to each mobile device, start the corresponding Appium Server container and establish a connection; By simulating the user operation process, the page location and operation path of the required collected data are clarified, and an automated collection logic is formed. Based on the obtained automated collection logic, feature extraction is performed on the new application state after user interaction, and data classification and structured storage are performed according to predefined rules; by setting scheduled tasks or trigger mechanisms, collection tasks can be started on demand.

4. A method for collecting batch data of massive network live broadcasts according to claim 1, characterized in that: Use the Scrapy-Redis framework to build a distributed crawler engine, parse the source address of live streaming media in real time, and perform block storage and format conversion on live video streams, including: Deploy the Scrapy framework and FFmpeg tools; Analyze the network requests of the live broadcast platform to determine the URL structure and request parameters of the live video stream; Based on the Scrapy framework, FFmpeg tools and asynchronous programming technology, combined with the Scrapy-Redis middleware to build a distributed crawler engine; The distributed crawler engine is deployed to the server to achieve continuous monitoring and collection of live streaming media data.

5. A method for collecting batch data of massive network live broadcasts according to claim 4, characterized in that: Analyzing the network request of the live broadcast platform to determine the URL structure and request parameters of the live video stream includes: Use browser developer tools to capture network requests related to live streaming; Trigger the real-time network request of the live broadcast platform and record the relevant request information; Parse the structure of the live stream URL and determine the variable parameters and their generation rules; Reverse engineer request parameters containing encrypted fields.

6. A method for collecting batch data of massive network live broadcasts according to claim 1, characterized in that: Use a multimodal large model to understand and analyze live content, identify violations, and generate violation analysis reports, including: Build a large model prompt word template; Read the live video file to be analyzed from the mass storage module; Use a large multimodal model to analyze the information of each video segment and extract key features; The text results output by the multimodal large model are imported into the LLM large model and analyzed in combination with the domain knowledge base to determine whether there is any violation; if the violation judgment is met, an alarm message is generated and relevant information is stored for evidence preservation.

7. A method for collecting batch data of massive network live broadcasts according to claim 1, characterized in that: The multimodal model includes: a graphic and text multimodal service model and a language multimodal service model; wherein the graphic and text multimodal service model is used to analyze the content of live video images; the language multimodal service model is used to identify the content of live audio; The output results of the graphic and text multimodal service model and the language multimodal service model are input into the large language model for semantic-level violation judgment.

8. A massive network live broadcast batch data collection system, characterized by: include: The terminal device control module is configured to: centrally control multiple mobile terminal devices through the group control system module to achieve the synchronous operation of multi-platform live broadcast APP; The data capture module is configured to: build an automated collection module based on the Appium framework, simulate real user interaction behaviors, and capture live broadcast room metadata; The streaming media processing module is configured to: build a distributed crawler engine using the Scrapy-Redis framework, parse the source address of live streaming media in real time, and perform block storage and format conversion on the live video stream; The intelligent analysis module is configured to use a multimodal large model to perform video understanding and semantic analysis on live broadcast content, identify violations, and generate violation analysis reports.

9. A computer-readable storage medium having a program stored thereon, characterized in that: When the program is executed by a processor, the steps of a method for collecting batch data of massive network live broadcasts as described in any one of claims 1 to 7 are implemented.

10. An electronic device comprising a memory, a processor, and a program stored in the memory and executable on the processor, wherein: When the processor executes the program, the steps in the method for collecting batch data of massive network live broadcasts as described in any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Crawler analysis platform based on government affair big data

    CN110175280A

  • Mobile phone group control system combining software with hardware equipment

    CN111726706A

  • Financial live broadcast violation detection method, device and equipment and readable storage medium

    CN113038153A

  • Live broadcast method, device, equipment, medium and program product

    CN118433430A

  • Big data processing method and system based on classification algorithm

    CN118585688A