News client agent cooperation service method, device, medium and product

CN122433783BActive Publication Date: 2026-09-29ZHEJIANG BAORONG MEDIA TECH (ZHEJIANG) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202610865449.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-16
Publication Date
2026-09-29
Estimated Expiration
2046-06-16

AI Technical Summary

Technical Problem

整个交互过程由单一模型闭环完成,未引入多处理单元的动态协作机制,也缺乏对任务语义进行分解与分发的能力

Benefits of technology

[0021]本发明实施例的技术方案,通过获取目标用户在使用新闻客户端的过程中产生的目标用户交互信息;基于所述目标用户交互信息生成目标任务指令,并根据所述目标任务指令的任务类型,确定至少一个子任务;分别确定与各所述子任务匹配的目标智能体,并基于所述目标智能体执行各所述子任务,得到各中间结果;对各所述中间结果进行语义对齐与多模态融合,生成面向所述目标用户的协同服务内容,可以显著提升复杂新闻交互场景下的服务准确性、内容完整性与上下文一致性,并增强新闻客户端在复杂交互场景下的服务稳定性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122433783B_ABST
    Figure CN122433783B_ABST
Patent Text Reader

Abstract

The application discloses a news client intelligent body cooperative service method and device, medium and product, and relates to the technical field of artificial intelligence. The method comprises the following steps: obtaining target user interaction information generated by a target user in the process of using a news client; wherein the target user interaction information comprises operation behavior data of the target user or submitted multi-modal input content; generating a target task instruction based on the target user interaction information, and determining at least one subtask according to the task type of the target task instruction; determining a target intelligent body matched with each subtask respectively, and executing each subtask based on the target intelligent body to obtain each intermediate result; and performing semantic alignment and multi-modal fusion on each intermediate result to generate cooperative service content for the target user. The scheme of the application significantly improves the service accuracy, content integrity and context consistency in a complex news interaction scene, and enhances the service stability of the news client in the complex interaction scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method, device, medium, and product for collaborative services of intelligent agents in a news client. Background Technology

[0002] With the rapid development of big data models and artificial intelligence technologies, news clients are evolving from one-way content distribution to intelligent, multimodal, and interactive services, and users are placing higher demands on news experiences that are contextually coherent, domain-specific, and responsive.

[0003] Currently, mainstream news apps' intelligent interaction systems generally adopt a processing architecture centered on a single, general-purpose large language model. Under this architecture, all user input is uniformly received and responded to by the same large model. For example, when a user is reading a medical news article about influenza and further inputs "What should I do if I have similar symptoms?", the system will still directly generate an answer from the general-purpose large model, without calling external professional resources or switching to a dedicated processing unit with medical knowledge. The entire interaction process is completed in a closed loop by a single model, lacking a dynamic collaboration mechanism among multiple processing units and the ability to decompose and distribute task semantics. This inability to dynamically schedule professional processing resources for complex user requests leads to inaccurate responses and service fragmentation in cross-domain interactions. Summary of the Invention

[0004] This invention provides a method, device, medium, and product for intelligent agent collaborative services in news clients, which significantly improves service accuracy, content integrity, and contextual consistency in complex news interaction scenarios, and enhances the service stability of news clients in complex interaction scenarios.

[0005] According to one aspect of the present invention, a method for collaborative services of intelligent agents in a news client is provided, the method comprising:

[0006] Acquire target user interaction information generated by the target user during the use of the news client; wherein, the target user interaction information includes the target user's operation behavior data or submitted multimodal input content;

[0007] Based on the target user interaction information, a target task instruction is generated, and at least one sub-task is determined according to the task type of the target task instruction;

[0008] Each target agent matching the sub-task is determined, and each sub-task is executed based on the target agent to obtain intermediate results.

[0009] Semantic alignment and multimodal fusion are performed on each of the intermediate results to generate collaborative service content for the target user.

[0010] According to another aspect of the present invention, a news client intelligent agent collaborative service device is provided, the device comprising:

[0011] The acquisition module is used to acquire target user interaction information generated by the target user during the use of the news client; wherein, the target user interaction information includes the target user's operation behavior data or submitted multimodal input content;

[0012] The subtask determination module is used to generate a target task instruction based on the target user interaction information, and determine at least one subtask according to the task type of the target task instruction;

[0013] The agent determination module is used to determine the target agent that matches each of the sub-tasks, and to execute each of the sub-tasks based on the target agent to obtain intermediate results.

[0014] The collaborative service content generation module is used to perform semantic alignment and multimodal fusion on the intermediate results to generate collaborative service content for the target user.

[0015] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:

[0016] At least one processor;

[0017] and a memory communicatively connected to the at least one processor;

[0018] The memory stores a computer program that can be executed by the at least one processor, which is then executed by the at least one processor to enable the at least one processor to perform the news client intelligent agent collaborative service method according to any embodiment of the present invention.

[0019] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the news client intelligent agent collaborative service method according to any embodiment of the present invention.

[0020] According to another aspect of the present invention, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the news client intelligent agent collaborative service method described in any embodiment of the present invention.

[0021] The technical solution of this invention involves acquiring target user interaction information generated by a target user during the use of a news client; generating target task instructions based on the target user interaction information; determining at least one sub-task according to the task type of the target task instructions; determining target intelligent agents matching each sub-task; executing each sub-task based on the target intelligent agents to obtain intermediate results; and performing semantic alignment and multimodal fusion on each intermediate result to generate collaborative service content for the target user. This significantly improves service accuracy, content completeness, and contextual consistency in complex news interaction scenarios and enhances the service stability of news clients in complex interaction scenarios.

[0022] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0023] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0024] Figure 1 This is a flowchart of a news client intelligent agent collaborative service method according to Embodiment 1 of the present invention;

[0025] Figure 2 This is a flowchart of a news client intelligent agent collaborative service method according to Embodiment 2 of the present invention;

[0026] Figure 3 This is a schematic diagram of the structure of a news client intelligent agent collaborative service device according to Embodiment 3 of the present invention;

[0027] Figure 4 This is a schematic diagram of the structure of an electronic device that implements the news client intelligent agent collaborative service method according to an embodiment of the present invention. Detailed Implementation

[0028] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0029] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0030] Example 1

[0031] Figure 1 This is a flowchart of a news client intelligent agent collaborative service method according to Embodiment 1 of the present invention. This embodiment is applicable to situations where multiple intelligent agents are deployed in a news client and collaboratively provide services to users. The method can be executed by a news client intelligent agent collaborative service device, which can be implemented in hardware and / or software and can be configured in electronic devices such as computers, servers, or tablet computers. Figure 1 As shown, the method includes:

[0032] Step 110: Obtain target user interaction information generated during the target user's use of the news client.

[0033] The target user interaction information includes the target user's operational behavior data or submitted multimodal input content.

[0034] In this embodiment, the target user can be any user who is using a news client, such as a freelancer, staff member, doctor, or student. This embodiment does not limit the target user.

[0035] Target user interaction information refers to a set of data generated actively or passively by target users during the use of a news client, which can be used to characterize their current intent or contextual state. This includes, but is not limited to, operational behavior data such as page browsing history, clickstream, dwell time, swiping behavior, favorites / shares / comments, and search keyword input; multimodal input content refers to information explicitly submitted by users that includes two or more modalities, such as news tips containing both voice and images, text comments with images, or uploaded video clips with accompanying text descriptions.

[0036] Optionally, in this embodiment, during the operation of the news client, when the target user performs page navigation, content clicks, scrolling, liking, favorites, forwarding, or enters search content in the search box, a behavior log record containing user identifier, device information, timestamp, current news item ID, operation type, and operation parameters is generated and cached in the local database by session dimension. Simultaneously, the client provides a multimodal input component on the news details page, the tip-off entry, or the comment section; when a user activates this component and submits data containing at least two modalities of images, audio, video, or text, the corresponding interface can be called to obtain the original media file and its metadata can be extracted synchronously, including but not limited to the image's shooting time, geographical location, shooting device model, and image size information; the audio's duration and number of channels; the video's frame rate and encoding format; and user-added text descriptions.

[0037] In this embodiment, all modal data and their metadata are bound to the same interaction session identifier, forming a structured multimodal input package. After the user completes the interaction, the client merges the accumulated operation behavior logs within this session with the multimodal input package into a single interaction context data body, which is then serialized and uploaded to the collaborative service processing platform via an encrypted channel for subsequent steps.

[0038] For example, when a target user is reading a news report about urban flooding, they stay on the page for 45 seconds and scroll down to read the full text. They then click the "Report a Report" button, activate their camera via the client to record a 10-second video of the flooded area, and enter "The floodwaters at the intersection of XX Road and YY Street have submerged the car wheels" in the accompanying text box. The user's actions—staying on the page and clicking—are recorded as user behavior data, while the submitted video and text constitute multimodal input content. Both are associated with the same session identifier and uploaded to the server as target user interaction information.

[0039] Step 120: Generate target task instructions based on target user interaction information, and determine at least one subtask according to the task type of the target task instructions.

[0040] Among them, the target task instruction refers to a structured task description generated after semantic parsing of the target user's interaction information, used to drive the execution of subsequent services. It may include elements such as task intent, target object, and contextual constraints. The task type is a classification of the target task instructions, used to distinguish service requirements of different natures, such as event reporting, knowledge Q&A, opinion commenting, or service request.

[0041] In this embodiment, determining at least one subtask based on the task type means breaking down complex or composite target task instructions into several atomic operation units that can be executed independently by a single intelligent agent. For example, reporting a fire and inquiring about personnel information can be broken down into two subtasks: extracting event elements and querying reports.

[0042] Optionally, in this embodiment, after receiving the target user's interaction information, keyword extraction, named entity recognition, and dependency parsing are performed on the text content in the target user's interaction information. A contextual feature vector is constructed by combining the operation behavior data (e.g., clicking the "report" button, uploading pictures, etc.). Subsequently, the feature vector is input into a pre-trained lightweight classifier (e.g., SVM or decision tree) and matched to a preset set of task types. Finally, according to the matched task type, the corresponding subtask generation template is loaded from the rule base, filled with specific parameters, and one or more subtasks are output.

[0043] In one optional implementation of this embodiment, after receiving the target user's interaction information, natural language preprocessing can be performed on the text content, including word segmentation, part-of-speech tagging, and named entity recognition, to identify elements such as time, location, people, and event keywords. Simultaneously, key signals in the operation behavior data are analyzed, such as whether the "I want to report a news item" entry is triggered, whether it contains image or video upload records, and the currently viewed news category. The aforementioned text semantic elements and behavioral features are combined into a structured feature vector, which is input to a task type classifier. This classifier is a decision tree model trained based on historical user interaction logs, and its output is a predefined task type identifier. The task type can include: event reporting, fact-checking, opinion solicitation, service consultation, etc., but this embodiment does not limit it.

[0044] Furthermore, based on the classification results, the corresponding decomposition strategy is retrieved from the subtask rule base. For example, if the task type is event reporting, the event reporting subtask template is invoked. This template defines two subtasks: Subtask A: extracts the five elements of an event (time, location, people, event, and result) from multimodal input; Subtask B: queries the local news database to see if there are relevant reports based on the extracted location and time. Specific values ​​extracted from the interaction information (e.g., XX subway station, 2026-01-14:13:20) are filled into the template parameters to generate two subtask objects with clear input references, processing logic identifiers, and output format requirements, and these objects are written to the task scheduling queue.

[0045] Step 130: Determine the target agent that matches each subtask, and execute each subtask based on the target agent to obtain each intermediate result.

[0046] In this embodiment, the target intelligent agent refers to a functional unit with specific domain processing capabilities already deployed in the news client, such as "Chao Xiaobang" for event analysis, "Chao Xiaokang" for medical Q&A, or "Chao Xiaopai" for image and text generation. This embodiment does not limit the specific target intelligent agent to these components. The intermediate results are structured outputs returned by each target intelligent agent after completing its sub-task, such as a tip-off form, health advice text, or generated images. This embodiment also does not limit the intermediate results to these outputs.

[0047] Optionally, in this embodiment, after receiving the subtask list generated in the above steps, the agent matching and invocation process is executed sequentially for each subtask. First, the capability requirement identifier is extracted from the subtask object, including the required processing modality (e.g., text, image, speech), task semantic keywords (e.g., event extraction, medical question answering, or text-to-image processing), and output format requirements. Further, the agent registry is queried to obtain metadata for all currently online agents. This metadata includes: agent tags, capability tag sets, domain semantic vectors, current memory availability indicators, and task success rates over the past 24 hours. The scheduling engine calculates the cosine similarity between the semantic vector of the subtask and the domain vector of each agent, combines this with their availability and historical success rates, and calculates a matching score.

[0048] Furthermore, the agent with the highest score can be selected as the target agent; then, subtask instructions and input data references are sent to the target agent; the target agent executes local processing logic, generates intermediate results, and sends them back to the collaborative service bus for use in subsequent fusion steps.

[0049] For example, when a user is reading a health news article about a new dengue fever vaccine, they might ask via voice, "I had dengue fever last year, can I get this vaccine now? Will there be any side effects?" This request is broken down into two sub-tasks: Sub-task one is to determine whether the user's past infection history affects their eligibility for vaccination, and sub-task two is to obtain information on the known side effects of the vaccine. Different target agents are matched for each of these sub-tasks. Specifically, sub-task one is matched with an agent possessing medical knowledge reasoning capabilities, and sub-task two is matched with an agent connected to a drug regulatory database. The two target agents independently execute their assigned sub-tasks; the former generates a medical conclusion including vaccination recommendations and risk assessments, while the latter returns a structured list of side effects and contraindications, forming two independent intermediate results.

[0050] Step 140: Perform semantic alignment and multimodal fusion on each intermediate result to generate collaborative service content for the target user.

[0051] Semantic alignment refers to mapping intermediate results from different agents to a unified semantic space, eliminating semantic conflicts or redundancy caused by differences in expression, modal heterogeneity, or lack of context. For example, high fever in text can be associated with body temperature ≥39℃ in medical knowledge base as the same clinical manifestation.

[0052] Optionally, in this embodiment, after generating each intermediate result, entity standardization can be further performed on the text-based intermediate results; for example, location names can be unified into administrative division codes, time expressions can be converted into a unified time format, and medical terms can be mapped to a standard dictionary; for image-based intermediate results, their visual feature vectors can be extracted and combined with the semantic vectors of the generated prompt text to construct a joint image-text representation, etc. Furthermore, a cross-result semantic graph is constructed, where nodes are key elements in each intermediate result (e.g., event subject, time, location, conclusion), and edges are logical relationships (e.g., causality, comparison, or supplementation, etc.). A graph neural network is used to calculate the consistency score between nodes to identify and mark potential conflicts. In the absence of conflicts, the output is organized according to a preset template: if an image is included, it is embedded in a response card, and an explanatory paragraph generated from the text intermediate results is attached below it. This paragraph is ensured to be consistent with the image content through referential resolution; if only text is included, a fluent paragraph is generated by sequence splicing and conjunction insertion, and style adaptation is performed through a language model to match the overall tone of the news client. The final generated collaborative service content is encapsulated in a preset format, including text blocks, media resource locators, structured metadata, and rendering instructions, and is pushed to the client front end for display.

[0053] For example, in the above example, if the intermediate results returned by the agent are: those with a history of infection can be vaccinated to prevent reinfection with other serotypes, but immunosuppression must be ruled out; common side effects include injection site pain and mild fever; contraindications include allergy to the vaccine components or pregnancy. Further, during the fusion phase, it is identified that the vaccination recommendation and side effects belong to the same medical decision-making context. The semantics of the two are aligned to the topic of vaccine use assessment, eliminating discrepancies in expression, and integrated into a coherent answer through connecting sentences: Based on your past infection history, you belong to the eligible population for vaccination, but it is recommended to confirm the absence of immunosuppression before vaccination; common side effects of this vaccine include injection site pain and mild fever; vaccination is not recommended if you are allergic to the vaccine components or are pregnant. Simultaneously, a diagram illustrating the vaccine's mechanism of action generated by the news semantic parsing module can be used as an illustration to ensure semantic consistency between text and image. Finally, the text and image are combined into a collaborative service content pushed to the user.

[0054] The technical solution of this embodiment obtains target user interaction information generated by the target user during the use of the news client; generates target task instructions based on the target user interaction information, and determines at least one sub-task according to the task type of the target task instructions; determines the target intelligent agent matching each sub-task, and executes each sub-task based on the target intelligent agent to obtain each intermediate result; performs semantic alignment and multimodal fusion on each intermediate result to generate collaborative service content for the target user, which can significantly improve the service accuracy, content completeness and context consistency in complex news interaction scenarios, and enhance the service stability of the news client in complex interaction scenarios.

[0055] Example 2

[0056] Figure 2 This is a flowchart of a news client intelligent agent collaborative service method according to Embodiment 2 of the present invention. This embodiment is a further refinement of the above technical solution, and the technical solution in this embodiment can be combined with various optional solutions in one or more of the above embodiments. Figure 2 As shown, the method includes:

[0057] Step 210: Obtain target user interaction information generated during the target user's use of the news client.

[0058] Optionally, in this embodiment, obtaining target user interaction information generated during the use of the news client may include: collecting target user operation behavior data generated on the news browsing interface in real time through the client's data tracking module, including page scrolling speed, content area dwell time, scroll position sequence, click coordinate sequence, and column switching events; and / or receiving multimodal input content actively submitted by the target user, including text data, voice audio data, or image data; aligning the operation behavior data and / or multimodal input content according to timestamps and encapsulating them into an interaction event sequence, and determining the interaction event sequence as target user interaction information.

[0059] Among them, operational behavior data refers to the unconscious or semi-conscious interaction signals generated by users in the news client interface, which are used to reflect their reading interests, attention distribution, and intention tendencies. For example, page scrolling speed can distinguish between quick skimming and in-depth reading; the duration of dwell time in content areas can indicate the degree of attention paid to a certain text or image; scroll position sequence can record the vertical trajectory of the user's browsing, etc.

[0060] Multimodal input content can refer to information input initiated by the user that includes two or more perceptual modalities, such as simultaneously entering text and uploading a screenshot in the comment section (text + image), or verbally expressing opinions via a voice button (audio).

[0061] In one optional implementation of this embodiment, the news client initializes the data collection module upon startup, continuously capturing user actions on the news details page, list page, and comment section. When the user swipes the screen, the current scroll offset is sampled at a first preset time interval (e.g., 50 seconds), the displacement difference between two adjacent reads is calculated and divided by the time interval to obtain the swipe speed; when a news paragraph area remains within the visible range of the screen for more than a second preset time interval (e.g., 800 milliseconds), the unique identifier of the paragraph, the timestamp of entering the visible area, and the timestamp of leaving the visible area are recorded; when the user touches the screen, the screen coordinates (x, y) of the touch event and the identifier of the clicked control are captured, and the timestamp of the event is recorded; when it is detected that the user switches from the current news section to another section, the section name before the switch, the section name after the switch, and the timestamp of the switch completion are recorded.

[0062] When a user uses the input function, if the voice input button is triggered, the client starts audio recording, saves the collected audio data as a target format file, and associates a timestamp with the file when the recording ends; if the user chooses to upload an image, the client calls the system image selection interface to obtain the image file and store it in the application's private directory, while recording the image file path and the operation completion timestamp; if the user inputs text, the client caches the text content and marks the input completion timestamp.

[0063] It should be noted that, in this embodiment, all the above-mentioned behavioral data entries and multimodal input entries carry their own timestamps; the client sorts these entries by timestamp from smallest to largest to form an ordered event sequence; each entry contains an event type field, data content or data reference, and a timestamp; finally, the client packages the ordered event sequence into a structured data packet, attaches the current session identifier, and sends it to the server as target user interaction information.

[0064] For example, a user opens a tech news article about AI chips, slowly scrolls the page (average speed 120 pixels / second), pauses for 4.2 seconds on a paragraph containing a chip architecture diagram, and clicks the "favorite" button in the lower right corner (coordinates x=320, y=680). The user then enters the comments section, presses the voice input button, and verbally states: "This energy efficiency improvement is crucial," attaching a comparison chart of XX chip parameters taken from another article. The client accurately timestamps the scrolling speed, pause duration, click coordinates, audio file, and image file, organizing them chronologically into an interaction event sequence: [Scrolling event → Pause event → Click event → Voice input event → Image upload event], and uploads this sequence as the target user's interaction information.

[0065] The solution in this embodiment synchronously collects fine-grained operational behavior data and user-submitted multimodal input content, and performs strict alignment and serialization encapsulation based on high-precision timestamps to construct a high-fidelity, multi-dimensional user interaction context representation. The generated interaction event sequence is characterized by clear structure, complete temporal sequence, and modal traceability, providing a comprehensive and consistent input foundation for the generation of subsequent task instructions, significantly enhancing the system's ability to understand and its robustness against complex, ambiguous, or cross-modal user intentions.

[0066] Step 220: Generate target task instructions based on target user interaction information, and determine at least one subtask according to the task type of the target task instructions.

[0067] Optionally, in this embodiment, generating target task instructions based on target user interaction information and determining at least one subtask according to the task type of the target task instructions may include: performing natural language understanding and intent parsing on the target user interaction information based on a pre-fine-tuned large model to generate a task semantic vector representing user needs, and constructing a structured target task instruction based on the task semantic vector; and decomposing the target task instruction into at least one subtask based on the task semantic tags in the target task instruction.

[0068] The sub-tasks include content creation sub-tasks, event reporting and analysis sub-tasks, domain knowledge Q&A sub-tasks, or news recommendation sub-tasks.

[0069] In this embodiment, the pre-fine-tuned large model refers to a specialized language model obtained by further supervised or instruction-based fine-tuning using news-related corpus data on top of a general large-scale language model. The training data used in this fine-tuning process includes, but is not limited to: news text, user comments, reporter / editor instructions, question-and-answer pairs, event reporting text, multi-turn dialogue logs, etc., covering multiple news verticals such as content creation, fact-checking, healthcare, and technology interpretation. Through this fine-tuning, the large model, while retaining its general language understanding capabilities, significantly enhances its ability to recognize and generate news contexts, user interaction intentions, and task instructions. It can accurately parse complex interactive information containing operational behavior context and multimodal input references, and output structured task instructions that conform to system scheduling specifications.

[0070] The task semantic vector is a dense vector representation generated by a large model after performing deep semantic encoding on the target user's interaction information. It can capture the core semantic features of the user's intent, such as asking about vaccine applicability or requesting the generation of news images. The target task instruction is a task description object generated by further structuring based on this vector, which includes fields such as task type label, target object, and constraints.

[0071] In one optional implementation of this embodiment, the target user interaction information can be input into a pre-tuned large model. Further, the large model performs joint semantic analysis on the text content, operational behavior context, and multimodal input references in the interaction information to identify the user's true intent in the current news reading scenario and outputs a task semantic vector. This vector can represent the user's core needs in a dense numerical form. Further, based on this task semantic vector, the large model generates a target task instruction according to a preset structured template. This instruction includes task semantic tags, target objects, contextual constraints, and reference information pointing to the original input data.

[0072] Furthermore, the task semantic tags in the target task instruction are parsed, and predefined subtask decomposition rules are matched based on these tags to break down the target task instruction into at least one atomic subtask. In this embodiment, each subtask explicitly specifies its processing type, which is limited to one of the following: content creation subtask, event reporting parsing subtask, domain knowledge question answering subtask, or news recommendation subtask. It also includes the required input data path, expected output format, and processing objective description. All generated subtasks are organized into a task list and written to a scheduling queue for use in subsequent agent matching and execution phases.

[0073] This embodiment's solution automatically decomposes target task instructions into standardized sub-task types according to predefined task semantic tags, achieving atomic decomposition of complex user needs and ensuring that each sub-task has clear processing boundaries, input dependencies, and output specifications. This not only supports diverse news interaction scenarios but also provides a structurally consistent and semantically complete task input foundation for subsequent accurate matching and parallel execution of multiple agents, significantly improving the system's depth of understanding of complex intentions and the level of automation in task scheduling.

[0074] Step 230: Determine the target agent that matches each subtask, and execute each subtask based on the target agent to obtain each intermediate result.

[0075] Optionally, in this embodiment, determining the target agent matching each subtask and executing each subtask based on the target agent to obtain intermediate results may include: obtaining the task semantic vector of the target subtask and obtaining the agent adaptation features of each vertical agent from the registered set of vertical agents; determining the adaptation score of each vertical agent based on the task semantic vector and the agent adaptation features of each vertical agent; taking the vertical agent with the highest adaptation score as the target agent of the target subtask, calling the target agent to execute the target subtask, and outputting intermediate results matching the type of the target subtask.

[0076] The agent's adaptation features include at least one of the following: domain representation, current computing resource status, and historical success rate of similar subtasks; wherein, domain representation refers to the business scope that the agent is good at handling, such as medical health, social events, or image and text generation; current computing resource status should reflect a normalized index of CPU, memory, or GPU load, with a higher value indicating more idle time; historical success rate of similar subtasks refers to the proportion of successes that the agent has achieved in the past when handling semantically similar subtasks.

[0077] In an optional implementation of this embodiment, after obtaining each subtask, a target subtask to be processed can be retrieved from the task queue, and its attached task semantic vector can be read. At the same time, a query is initiated to the agent registration center to obtain the adaptation feature data of all currently registered vertical agents. The adaptation features include at least one of the following: the domain representation vector obtained by semantically encoding the capability description text submitted during registration, the current computing resource status value collected in real time and normalized by the runtime monitoring module, and the historical success rate of the agent in processing semantically similar subtasks in the past thirty days, which is statistically obtained by the task log system.

[0078] Furthermore, an adaptation score is calculated for each vertical agent. This involves calculating the cosine similarity between the task semantic vector and the agent's domain representation vector to obtain a semantic matching score. This matching score, the current computing resource status value, and the historical success rate are then weighted and summed according to preset weights, with the sum of weights being 1 and defined by the system configuration file. The final value obtained is the adaptation score for that agent. Further, the adaptation scores of all vertical agents are iterated, and the agent with the highest score is selected as the target agent. An execution request is sent through its registered unified calling address, containing the subtask identifier, input data reference path, and expected output format. Upon receiving the request, the target agent loads the corresponding business logic module, executes the processing flow, and generates an intermediate result matching the subtask type. This intermediate result is encapsulated and returned to the scheduling service, verified, and stored in the collaborative service bus, completing the execution of this subtask.

[0079] Optionally, in this embodiment, the adaptation score of each vertical intelligent agent is determined based on the task semantic vector and the agent adaptation features of each vertical intelligent agent, including: calculating the domain matching score of each vertical intelligent agent based on the task semantic vector and the domain representation of each vertical intelligent agent; calculating the resource availability score of each vertical intelligent agent based on the current computing resource status of each vertical intelligent agent; calculating the historical performance score of each vertical intelligent agent based on the historical success rate of the same type of subtasks of each vertical intelligent agent; and weighting and summing the domain matching score, resource availability score and historical performance score of each vertical intelligent agent according to preset weights to obtain the adaptation score of each vertical intelligent agent.

[0080] In an optional implementation of this embodiment, after obtaining each subtask, the task semantic vector of the target subtask (any subtask) can be further obtained, and the adaptation feature data of all registered vertical intelligent agents can be synchronized from the intelligent agent registration center. Further, the following calculations are performed sequentially for each vertical intelligent agent: First, the cosine similarity between the task semantic vector and the agent's domain representation vector is calculated as its domain matching score; second, its current computing resource status value is directly used as its resource availability score; third, its historical success rate is used as its historical performance score; finally, the above three scores are weighted and summed according to the preset weight coefficients in the configuration file to obtain the agent's adaptation score. After the adaptation scores of all vertical intelligent agents are calculated, the system sorts them from high to low scores, selects the top-ranked intelligent agent as the target intelligent agent, and triggers its execution process.

[0081] In an optional implementation of this embodiment, the matching score can be calculated using the following weighted formula:

[0082] ;

[0083] in, Indicates the first Sub-tasks and the first The matching score between agents is higher, indicating a better fit. , , These are adjustable weighting coefficients, and ; Indicates the first The task semantic vectors of each subtask; For the first The domain vector of each agent; For the first The current computational availability of each agent; Indicates the first The historical success rate of an agent when handling subtasks similar to the current task type in the past, that is, the proportion of successful completions out of the total number of attempts.

[0084] In one example of this embodiment, after reading a news article about "increasing the coverage of community elderly canteens," a user leaves a comment: "Are there any such canteens in our district? How do I apply?" The news client breaks down this request into a sub-task: querying the elderly meal assistance service outlets and application methods in the user's administrative district. The task semantic vector of this sub-task focuses on local public service retrieval. Adaptive features of all vertical agents are obtained. One agent's domain representation vector is highly correlated with community elderly care, its current computing resource status is 0.92 (indicating ample resources), and its historical success rate for handling similar public service query sub-tasks in the past month is 94%. Another general question-answering agent is online but has low domain matching and a resource status of only 0.45. After weighted scoring, the former has the highest score and is identified as the target agent. This agent calls the local service public interface to obtain the list of meal assistance points and application process in the user's district, returning an intermediate result: "There are currently 12 elderly canteens in your XX district, covering 8 streets. Residents over 60 years old can register and apply at the community service center with their ID card. See the relevant website for details."

[0085] In one optional implementation of this embodiment, registration requests from various vertical-category intelligent agents are received through a unified API protocol. These intelligent agents include: the "Chao Xiaotu" intelligent agent, which provides functions such as text-to-image generation, image-to-text generation, summary expansion, topic-based image matching, and lead-in polishing based on content creation scenarios; the "Chao Xiaosu" intelligent agent, which automatically parses elements such as time, location, theme, and people in natural language and generates structured reporting forms based on user reports and social event feedback information; and the "Chao Xiaokang" intelligent agent, which is geared towards the medical and health news field and has the ability to answer medical knowledge questions, generate health advice, and recommend related news. Each intelligent agent submits its capability description and service interface during registration and connects to a unified context data bus to share semantic information, achieve parallel tasks, result aggregation, and context inheritance.

[0086] In the multi-agent collaborative processing, the system receives the result sequence returned by each agent. These results are then fed into the fusion module, which performs optimization calculations with the goal of minimizing the squared error between all results and the fused result. ;in, For the first Intermediate results returned by each agent The final output after fusion is obtained by iteratively solving the objective function to obtain a unified output after weighted integration of the results of each agent, thus achieving semantic consistency and information complementarity of multi-source results.

[0087] Further deployment of a news semantic parsing and content summarization module. This module, based on a Transformer-structured semantic architecture model, performs topic clustering, opinion identification, and event chain extraction on the news text. Specifically, the news text is input into a Transformer encoder to obtain the hidden state corresponding to each token. ; Event chains are represented in the form of quadruples, i.e. ,in Indicates an action, Represents an object, Indicates time, Representing locations; all event chains form a directed graph. ,in For a set of events, A set of relationships between events, where a relationship is defined as follows: This forms a cause-and-effect graph structure.

[0088] It also deploys an intelligent comment interaction module to perform semantic clustering and sentiment recognition on user comments in the comment section; based on the semantic clustering results, it automatically generates semantically relevant and naturally linguistically natural replies, supporting personalized expressions under different personality settings; sentiment analysis calculates the sentiment score of the comment text through a sentiment quantification function. ;in, and These represent the probabilities of positive and negative sentiment, respectively. and The weights are used as coefficients; simultaneously, the comment text is classified into intents using semantic vector clustering, with the objective function being: ,in, For the semantic vector of the comment text, It is the cluster center to which it belongs.

[0089] In another optional implementation of this embodiment, a neural network model can be trained by collecting user swiping speed, dwell time, scrolling range, and click hotspots to predict when a user will refresh or switch sections. A news API is requested in advance in the background to preload images and summaries, achieving seamless refresh. If an increase in user interest in a certain type of news is detected, a corresponding topic page is proactively pushed, achieving intent-driven intelligent push. The refresh probability is calculated using the predictive model based on the user behavior sequence. ;

[0090] in, For the model at time The output vector, It is the Sigmoid activation function. , These are model parameters; when the predicted probability exceeds a set threshold, a pre-refresh process is triggered.

[0091] This embodiment achieves dynamic and precise allocation of subtasks to execution units by integrating task semantic vectors with multi-dimensional adaptation features of vertical-category agents. Domain representation ensures the professionalism of capability matching, current computing resource status guarantees the real-time performance and stability of system response, and historical success rates introduce an experience feedback mechanism to improve service reliability. This significantly enhances the task scheduling accuracy, resource utilization efficiency, and overall service robustness of the multi-agent collaborative system.

[0092] Step 240: Perform semantic alignment and multimodal fusion on each intermediate result to generate collaborative service content for the target user.

[0093] Optionally, in this embodiment, semantic alignment and multimodal fusion of the intermediate results to generate collaborative service content for the target user may include: extracting the contextual semantic elements of each intermediate result, clustering the intermediate results based on the contextual semantic elements to generate at least one event context group; and integrating the cross-modal features of the intermediate results in each event context group through a pre-trained multimodal fusion model to generate collaborative service content containing text, images, and structured information.

[0094] Contextual semantic elements refer to key semantic units extracted from each intermediate result, used to characterize the core information of the event to which they belong, and may include time, location, subject, event type, state, etc. Event context groups are sets formed by clustering the contextual semantic elements of all intermediate results after measuring their similarity. In this embodiment, intermediate results within the same group describe the same event or highly related sub-events.

[0095] In one optional implementation of this embodiment, all intermediate results that have been completed can be read from the collaborative service bus, and semantic elements can be extracted from each intermediate result: if it is a text-based result, time, location, person, event action and status can be extracted through named entity recognition and dependency parsing; if it is an image-based result, the prompt text or image metadata used when it was generated can be parsed to extract the corresponding semantic elements; if it is structured data, the predefined field values ​​can be directly read as semantic elements.

[0096] In this embodiment, the semantic elements of all intermediate results are converted into element vectors in a unified format, and the pairwise association scores are calculated based on the entity overlap and semantic similarity between elements. Furthermore, a hierarchical clustering algorithm is used to group intermediate results with association scores higher than a threshold into the same event context group, ensuring that results within each group point to the same news event or user focus. For each event context group, its contained text, images, and structured data are input into a pre-trained multimodal fusion model. This model first encodes the content of each modality separately, then aligns key semantic nodes through a cross-modal attention mechanism, suppresses conflicting information, strengthens complementary details, and finally outputs a coherent natural language description, embedding resource identifiers of relevant images and attaching structured data cards. This output is the collaborative service content for the target user, which can be packaged and pushed to the client for display.

[0097] In one optional implementation of this embodiment, text data and image data input by the user are received and fed into a text encoder and an image encoder, respectively, to obtain corresponding text semantic vectors and image semantic vectors. The text semantic vectors and image semantic vectors are then input into a multimodal alignment module, which calculates the similarity between the two in a shared semantic space to complete the semantic alignment of the text and the image, and outputs the aligned multimodal joint representation. Further, the multimodal joint representation is input into a cross-modal generation module, which generates target content based on the alignment result, including: generating matching news images based on text, generating corresponding descriptive text based on images, or generating news titles and summaries by fusing multimodal information. Finally, the generated text, images, and associated metadata are combined into collaborative service content and output to the client for presentation.

[0098] A multimodal alignment mechanism can unify the semantic space of text and images, supporting user interaction with the AI ​​system through text, voice, and images. After receiving user input, the system uses a cross-modal generation mechanism to convert text into images, images into text, and voice into text, assisting journalists and users in creating news content. Furthermore, the system supports AI in automatically matching images, generating titles and summaries for news content, reducing the workload of manual editors. The multimodal alignment mechanism calculates the semantic distance between text and images and introduces a maximum value constraint term to construct a joint optimization objective function, achieving semantic consistency between text and images and improving the accuracy and consistency of cross-modal content generation.

[0099] The solution in this embodiment effectively solves the problem of event fragmentation that may exist due to different sources of multi-source heterogeneous results by extracting the contextual semantic elements of intermediate results and performing clustering, ensuring that subsequent fusion operations are performed within semantically consistent event units.

[0100] Step 250: Perform keyword matching and semantic feature extraction on the text portion of the collaborative service content, and compare the extraction results with a preset content filtering rule base to obtain text detection results; perform image classification and region feature detection on the image portion of the collaborative service content, and compare the detection results with a preset content filtering rule base to obtain image detection results; convert the audio portion of the collaborative service content into corresponding text through speech recognition, perform keyword matching and semantic feature extraction on the corresponding text, and compare the extraction results with a preset content filtering rule base to obtain audio detection results; based on the text detection results, image detection results, and audio detection results, perform filtering operations on the collaborative service content to generate filtered collaborative service content, and return the filtered collaborative service content to the target user.

[0101] Keyword matching involves precisely or fuzzily comparing text against a pre-defined sensitive word library to determine if it contains prohibited words. Semantic feature extraction uses an embedding model to obtain the deep semantic vector of the text and calculates its similarity to a prohibited semantic template. Image classification in the image section is used to identify the overall content category; region feature detection uses an object detection model to locate specific sensitive regions in the image. Audio data can be first converted into text through speech recognition and then reused in the text detection process.

[0102] In one optional implementation of this embodiment, after obtaining the collaborative service content, text, image, and audio components can be separated from it. For the text component, keyword matching is first performed, segmenting the text and comparing it with a keyword set in a preset content filtering rule base, recording the matched items. Simultaneously, the text is input into a semantic encoding model to generate a text semantic feature vector, and the cosine similarity between this vector and each violating semantic template vector in the rule base is calculated. If the similarity exceeds a preset threshold, it is marked as a semantic violation. For the image component, it is input into a multi-task image analysis model, which simultaneously performs image classification. The system performs region feature detection, outputting an overall image category label and bounding boxes and category labels for several sensitive regions. It then matches the overall and region category labels against a set of image violation categories in a rule base; if an intersection exists, the image is marked as violating the rule. For the audio portion, it first converts the audio into corresponding text using a speech recognition engine. Then, it performs the same keyword matching and semantic feature extraction process as the text portion and compares the results with the rule base to obtain the audio detection results. Finally, it aggregates the text, image, and audio detection results. If any detection result indicates the presence of violating content, the entire collaborative service content is marked as pending filtering. Based on the filtering strategy configured in the rule base, the system removes the violating content or replaces it with compliant placeholders, generating filtered collaborative service content, which is then returned to the target user's terminal for display.

[0103] The solution in this embodiment achieves full-dimensional compliance review of collaborative service content by constructing a unified content security detection mechanism covering text, images, and audio modalities. While ensuring the richness of content generation, it significantly improves the system's accuracy and consistency in identifying complex, cross-modal violations, meeting the high compliance requirements for content security in news applications.

[0104] This invention, through the deep integration of a multi-agent collaborative mechanism and a semantic understanding model, significantly improves the intelligence level of interaction between users and artificial intelligence, enabling multi-role, contextualized dialogues around news semantics, increasing the average user conversation duration by approximately 40%. Compared to traditional keyword extraction methods, the semantic understanding module improves the accuracy of key point extraction and stance judgment by approximately 30%, and the generated reading reports are closer to news logic. The system preloads content based on user behavior prediction, reducing the average refresh latency to 40% of the original, effectively improving client response performance. The intelligent agent in the comment section automatically generates personalized replies based on semantic clustering. The system saw a 35% increase in interaction volume and a reduction in cold-start comments. Reporters and editors utilized vertical AI agents for automated image matching, lead generation, and manuscript polishing, resulting in approximately a 45% increase in news production efficiency and a significant reduction in manual editing workload. The platform supports rapid integration of various types of AI agents, reducing the deployment time for new vertical models to one-third of the original time, and possesses the capability to expand into fields such as education, tourism, and finance. Generated content undergoes dual review by AI agents and human review to ensure that outputs comply with news and public opinion guidance and regulatory requirements. The system reduces human customer service and editing costs by over 50%, while simultaneously improving user retention and usage time, enhancing the platform's commercial conversion potential.

[0105] Example 3

[0106] Figure 3 This is a schematic diagram of the structure of a news client intelligent agent collaborative service device according to Embodiment 3 of the present invention. Figure 3 As shown, the device includes: an acquisition module 310, a subtask determination module 320, an agent determination module 330, and a collaborative service content generation module 340.

[0107] The acquisition module 310 is used to acquire target user interaction information generated by the target user during the use of the news client; wherein, the target user interaction information includes the target user's operation behavior data or submitted multimodal input content;

[0108] The subtask determination module 320 is used to generate a target task instruction based on the target user interaction information, and determine at least one subtask according to the task type of the target task instruction;

[0109] The agent determination module 330 is used to determine the target agent that matches each of the sub-tasks, and to execute each of the sub-tasks based on the target agent to obtain intermediate results.

[0110] The collaborative service content generation module 340 is used to perform semantic alignment and multimodal fusion on the intermediate results to generate collaborative service content for the target user.

[0111] In an optional implementation of this embodiment, the news client intelligent agent collaborative service device further includes: a filtering module, used to perform keyword matching and semantic feature extraction on the text portion of the collaborative service content, and compare the extraction results with a preset content filtering rule base to obtain text detection results;

[0112] Image classification and region feature detection are performed on the image portion of the collaborative service content, and the detection results are compared with the preset content filtering rule base to obtain the image detection results;

[0113] For the audio portion of the collaborative service content, speech recognition is used to convert it into corresponding text. Keyword matching and semantic feature extraction are performed on the corresponding text, and the extraction results are compared with the preset content filtering rule base to obtain the audio detection results.

[0114] Based on the text detection results, image detection results, and audio detection results, a filtering operation is performed on the collaborative service content to generate filtered collaborative service content, and the filtered collaborative service content is returned to the target user.

[0115] In an optional implementation of this embodiment, the acquisition module 310 is specifically used to collect the operation behavior data generated by the target user on the news browsing interface in real time through the client-side tracking module. The operation behavior data includes page scrolling speed, content area dwell time, scroll position sequence, click coordinate sequence, and column switching events.

[0116] And / or, receive multimodal input content actively submitted by the target user, the multimodal input content including text data, voice / audio data, or image data;

[0117] The operation behavior data and / or the multimodal input content are aligned according to timestamps and encapsulated into an interaction event sequence, which is then identified as the target user interaction information.

[0118] In an optional implementation of this embodiment, the subtask determination module 320 is specifically used to perform natural language understanding and intent parsing on the target user interaction information based on the pre-fine-tuned large model, generate a task semantic vector representing user needs, and construct a structured target task instruction based on the task semantic vector.

[0119] Based on the task semantic tags in the target task instruction, the target task instruction is decomposed into at least one subtask;

[0120] The subtasks include content creation subtasks, event reporting and analysis subtasks, domain knowledge Q&A subtasks, or news recommendation subtasks.

[0121] In an optional implementation of this embodiment, the agent determination module 330 is specifically used to obtain the task semantic vector of the target subtask and obtain the agent adaptation features of each vertical agent from the registered set of vertical agents.

[0122] Based on the task semantic vector and the agent adaptation features of each vertical agent, the adaptation score of each vertical agent is determined.

[0123] The vertical category agent with the highest adaptation score is selected as the target agent for the target subtask, and the target agent is invoked to execute the target subtask, outputting intermediate results that match the type of the target subtask.

[0124] In an optional implementation of this embodiment, the agent adaptation features include at least one of the following: domain representation, current computing resource status, and historical success rate of similar subtasks;

[0125] The agent determination module 330 is also specifically used to calculate the domain matching score of each vertical agent based on the task semantic vector and the domain representation of each vertical agent;

[0126] Based on the current computing resource status of each vertical agent, calculate the available resource score for each vertical agent;

[0127] Based on the historical success rate of the same sub-tasks of each vertical agent, calculate the historical performance score of each vertical agent.

[0128] The domain matching score, resource availability score, and historical performance score of each vertical agent are weighted and summed according to preset weights to obtain the adaptation score of each vertical agent.

[0129] In an optional implementation of this embodiment, the collaborative service content generation module 340 is specifically used to extract the contextual semantic elements of each intermediate result, and to cluster the intermediate results based on each contextual semantic element to generate at least one event context group.

[0130] By using a pre-trained multimodal fusion model, cross-modal features are integrated into the intermediate results of each event context group to generate collaborative service content containing text, images, and structured information.

[0131] The news client intelligent agent collaborative service device provided in the embodiments of the present invention can execute the news client intelligent agent collaborative service method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of executing the method.

[0132] In the technical solutions of the embodiments of the present invention, the collection, storage, use, processing, transmission, provision and disclosure of user interaction information (such as operation behavior data, multimodal input content, etc.) all comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0133] Example 4

[0134] Figure 4 This is a schematic diagram of the structure of an electronic device implementing the news client intelligent agent collaborative service method of this invention. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (such as helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0135] like Figure 4 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded into the RAM 13 from the storage unit 18. The RAM 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0136] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0137] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, digital signal processors (DSPs), and any suitable processor, controller, microcontroller, etc. Processor 11 executes the various methods described above, such as the news client intelligent agent collaborative service method.

[0138] In some embodiments, the news client intelligent agent collaborative service method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the news client intelligent agent collaborative service method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to execute the news client intelligent agent collaborative service method by any other suitable means (e.g., by means of firmware).

[0139] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0140] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0141] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, RAM, ROM, erasable programmable read-only memory (EPROM), optical fibers, compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0142] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device for displaying information to the user (e.g., a cathode ray tube (CRT) or liquid crystal display (LCD) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0143] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0144] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system. It addresses the shortcomings of traditional physical hosts and Virtual Private Servers (VPS) in terms of management difficulty and weak business scalability.

[0145] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0146] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

[0147] This invention also provides a computer program product, including a computer program that, when executed by a processor, implements a database detection method as provided in any embodiment of this application.

[0148] In implementing the computer program product, computer program code for performing the operations of this invention can be written in one or more programming languages ​​or a combination thereof. Programming languages ​​include object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including LANs or WANs—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0149] It should be noted that in the embodiments of this application, certain software, components, models and other existing solutions in the industry may be mentioned. These should be regarded as exemplary and are only intended to illustrate the feasibility of implementing the technical solution of this application. However, they do not mean that the solution has been or necessarily used.

[0150] Note that the above description is merely a preferred embodiment of the present invention and the technical principles employed. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein, and various obvious changes, readjustments, and substitutions can be made without departing from the scope of protection of the present invention. Therefore, although the present invention has been described in detail through the above embodiments, the present invention is not limited to the above embodiments, and may include many other equivalent embodiments without departing from the concept of the present invention, the scope of which is determined by the scope of the appended claims.

Claims

1. A method for collaborative services of intelligent agents in a news client, characterized in that, The method includes: Acquire target user interaction information generated by the target user during the use of the news client; wherein, the target user interaction information includes the target user's operation behavior data or submitted multimodal input content; Based on the pre-fine-tuned large model, natural language understanding and intent parsing are performed on the target user interaction information to generate task semantic vectors representing user needs. Based on the task semantic vectors, structured target task instructions are constructed. Based on the task semantic tags in the target task instructions, the target task instructions are decomposed into at least one sub-task. The sub-tasks include content creation sub-tasks, event reporting and parsing sub-tasks, domain knowledge question answering sub-tasks, or news recommendation sub-tasks. Obtain the task semantic vector of the target subtask, and obtain the agent adaptation features of each vertical agent from the set of registered vertical agents; determine the adaptation score of each vertical agent based on the task semantic vector and the agent adaptation features of each vertical agent; select the vertical agent with the highest adaptation score as the target agent of the target subtask, call the target agent to execute the target subtask, and output the intermediate result that matches the type of the target subtask; A cross-result semantic graph is constructed, where nodes in the semantic graph are key elements in each intermediate result, and edges represent the logical relationships between key elements. A graph neural network is used to calculate consistency scores between nodes, identify and mark potential conflicts, and perform semantic alignment on each intermediate result. Multimodal fusion is performed on the semantically aligned intermediate results to generate collaborative service content for the target user.

2. The news client intelligent agent collaborative service method according to claim 1, characterized in that, The method further includes: Keyword matching and semantic feature extraction are performed on the text portion of the collaborative service content, and the extraction results are compared with a preset content filtering rule base to obtain text detection results; Image classification and region feature detection are performed on the image portion of the collaborative service content, and the detection results are compared with the preset content filtering rule base to obtain the image detection results; For the audio portion of the collaborative service content, speech recognition is used to convert it into corresponding text. Keyword matching and semantic feature extraction are performed on the corresponding text, and the extraction results are compared with the preset content filtering rule base to obtain the audio detection results. Based on the text detection results, image detection results, and audio detection results, a filtering operation is performed on the collaborative service content to generate filtered collaborative service content, and the filtered collaborative service content is returned to the target user.

3. The news client intelligent agent collaborative service method according to claim 1, characterized in that, The acquisition of target user interaction information generated during the use of the news client includes: The client-side tracking module collects real-time operational behavior data of the target user on the news browsing interface. The operational behavior data includes page scrolling speed, duration of stay in content area, scroll position sequence, click coordinate sequence, and column switching events. And / or, receive multimodal input content actively submitted by the target user, the multimodal input content including text data, voice / audio data, or image data; The operation behavior data and / or the multimodal input content are aligned according to timestamps and encapsulated into an interaction event sequence, which is then identified as the target user interaction information.

4. The news client intelligent agent collaborative service method according to claim 1, characterized in that, The agent adaptation features include at least one of the following: domain representation, current computing resource status, and historical success rate of similar subtasks; The process of determining the adaptation score of each vertical agent based on the task semantic vector and the agent adaptation features of each vertical agent includes: Based on the task semantic vector and the domain representation of each vertical agent, the domain matching score of each vertical agent is calculated; Based on the current computing resource status of each vertical agent, calculate the available resource score for each vertical agent; Based on the historical success rate of the same sub-tasks of each vertical agent, calculate the historical performance score of each vertical agent. The domain matching score, resource availability score, and historical performance score of each vertical agent are weighted and summed according to preset weights to obtain the adaptation score of each vertical agent.

5. The news client intelligent agent collaborative service method according to claim 1, characterized in that, The step of semantically aligning and multimodal fusion of the intermediate results to generate collaborative service content for the target user includes: Extract the contextual semantic elements of each intermediate result, and cluster the intermediate results based on the contextual semantic elements to generate at least one event context group; By using a pre-trained multimodal fusion model, cross-modal features are integrated into the intermediate results of each event context group to generate collaborative service content containing text, images, and structured information.

6. An electronic device, characterized in that, The electronic device includes: At least one processor; and a memory communicatively connected to the at least one processor; The memory stores a computer program that can be executed by the at least one processor, which is then executed by the at least one processor to enable the at least one processor to perform the news client intelligent agent collaborative service method according to any one of claims 1-5.

7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that are used to cause a processor to execute the news client intelligent agent collaborative service method according to any one of claims 1-5.

8. A computer program product comprising a computer program that, when executed by a processor, implements the news client intelligent agent collaborative service method according to any one of claims 1-5.

Citation Information

Patent Citations

  • Multi-agent collaborative task planning method, related device, equipment and storage medium

    CN120723402A