Task dynamic processing method and system based on large language model, medium and equipment

Through a dynamic task processing method based on a large language model, combined with multimodal wake-up signals and knowledge graphs, the limitations of AI agents in dynamic response and cross-platform execution are solved, and full-scene adaptive triggering and in-depth scenario understanding are realized, and the adaptive capabilities and execution efficiency of AI systems in complex scenarios are improved.

CN120066730APending Publication Date: 2025-05-30ZHONG FU TONG CO LTD
View PDF 0 Cites 6 Cited by

Patent Information

Application Number
CN202510225542.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing AI agents have significant limitations in dynamic response and cross-platform execution, and lack the ability to perceive multiple signals such as user behavior patterns and environmental states, resulting in insufficient adaptability to trigger scenarios.

Method used

The task dynamic processing method based on the large language model is adopted, and the task semantic vector containing application scenario classification is generated by receiving multi-modal wake-up signals (including voice input, text input, device behavior mode input and environment state automatic trigger signals), and the user needs are analyzed through the large language model to generate a task semantic vector containing application scenario classification. Scenario correlation analysis is carried out based on the knowledge graph, a task processing process framework is generated, and executable tool chains are dynamically matched to complete cross-platform task execution.

Benefits of technology

It realizes adaptive triggering of all scenarios, in-depth scenario understanding, supports elastic configuration of the optimal execution path, realizes seamless connection of cross-platform operations, reduces user interaction steps, and demonstrates efficiency advantages and fault tolerance in complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120066730A_ABST
    Figure CN120066730A_ABST
Patent Text Reader

Abstract

The invention discloses a task dynamic processing method and system based on a large language model, a medium and equipment, which adopt a multi-mode fusion wake-up mechanism, integrate voice, text, equipment behavior modes and environment sensor data, and realize full-scene self-adaptive triggering. On the basis of a task semantic vector generation technology of a large language model, in combination with the association reasoning capability of a knowledge graph, deep scene understanding is realized. The dynamic tool chain matching mechanism supports the elastic configuration of the optimal execution path based on the real-time updating capability of the process database. Through the third-party application interface, seamless connection of cross-platform operation is achieved, a user does not need to pay attention to the difference of bottom-layer applications, user interaction steps are reduced, and the efficiency advantage and the fault-tolerant capability are shown in a complex scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computers, and in particular, to a method, system, medium and device for dynamically processing tasks based on a large language model. Background Art

[0002] AI agents have been widely used in actual applications. For example, a variety of AI agents such as Xiaoai Tongxue, Doubao, and Youyou are pre-configured in the terminal devices used by users to help users solve daily problem requirements. However, existing AI agents have significant limitations in dynamic response and cross-platform execution. Traditional wake-up mechanisms mostly rely on single-modal input (such as pure voice or text instructions), lacking the ability to perceive multiple signals such as user behavior patterns and environmental states, resulting in insufficient adaptability to trigger scenarios. For example, when the user is in a noisy environment, voice wake-up is likely to fail; in specific device operation scenarios, simply relying on text input cannot capture the user's potential intentions. Summary of the Invention

[0003] In view of this, the purpose of the present invention is to propose a method, system, medium and device for dynamically processing tasks based on a large language model.

[0004] In order to achieve the above technical objectives, the technical solutions adopted by the present invention are as follows:

[0005] In the first aspect, the present invention provides a method for dynamically processing tasks based on a large language model, including:

[0006] Receiving a multi-modal wake-up signal, the multi-modal wake-up signal including at least one of voice input, text input, device behavior pattern input, and environmental state automatic trigger signal;

[0007] Converting the multi-modal wake-up signal into language information, parsing the user's needs through the large language model with the language information, and generating a task semantic vector including application scenario classification;

[0008] Performing scenario association analysis on the task semantic vector based on the knowledge graph to obtain the target application scenario;

[0009] Generating a task processing flow framework according to the target application scenario;

[0010] Dynamically matching an executable tool chain corresponding to the task processing flow framework according to a pre-constructed process database;

[0011] Invoking the function modules in the executable tool chain through a third-party application interface to complete cross-platform task execution.

[0012] In some embodiments, the device behavior pattern input includes at least one of a specific shaking pattern, a specific key pattern, and a specific gesture; the environmental state automatic trigger signal includes a trigger condition for reaching a preset area based on the matching of a geographical location sensor and a timestamp.

[0013] In some embodiments, the environmental state automatic trigger signal is configured to be generated through the following steps:

[0014] Real-time monitor the movement trajectory of the user's terminal device and the system time;

[0015] When it is detected that the user leaves the work location coordinates and the moving direction points to the home address coordinates, predict the user's needs through the historical behavior pattern database, and generate a home service type wake-up instruction according to the user's needs.

[0016] In some embodiments, the knowledge graph includes a multi-layer semantic network structure, and the multi-layer semantic grid result includes a first layer, a second layer, and a third layer; the first layer is configured to divide the online service scenario and the offline entity operation scenario; the second layer is configured to divide the commodity procurement module, the housekeeping service module, and the smart home control module according to the service field; the third layer is configured to define the association rules and execution priorities between the commodity procurement module, the housekeeping service module, and the smart home control module.

[0017] In some embodiments, generating a task processing flow framework according to the target application scenario includes:

[0018] Extract the intent entity and constraint conditions in the language information through a large language model;

[0019] Match the intent entity and constraint conditions with the knowledge graph to generate a process topology structure, and the process topology structure includes a subtask sequence, an execution parameter transfer path, and an exception handling strategy.

[0020] In some embodiments, the subtask sequence is configured to be constructed using a dynamic sharding mechanism;

[0021] The construction steps of the subtask sequence include:

[0022] Split the complex task output by the knowledge graph into independently executable sub-operation units;

[0023] Establish a data dependency relationship graph between multiple sub-operation units;

[0024] Through the context awareness module, automatically fill in the cross-unit parameters in the data dependency relationship graph, and the cross-unit parameters are configured as unknown parameters in the data dependency relationship graph.

[0025] In some embodiments, the method further includes:

[0026] After the cross-platform task execution is completed, the execution result is returned to the user;

[0027] The execution result includes at least one of an execution report in natural language form, a physical state feedback signal of the smart home device, and a confirmation message sent by the third-party service platform.

[0028] In a second aspect, the present invention further provides a task dynamic processing system based on a large language model, which is applicable to the processing method described in the first aspect. The system includes a signal acquisition module, a semantic parsing module, a scenario processing module, a process execution module, and a third-party application module; the signal acquisition module is used to receive multi-modal wake-up signals, and the multi-modal wake-up signals include at least one of voice input, text input, device behavior pattern input, and environmental state automatic trigger signals; the semantic parsing module is used to convert the multi-modal wake-up signals into language information, parse the user requirements through the large language model for the language information, and generate a task semantic vector including application scenario classification; the scenario processing module is used to perform scenario association analysis on the task semantic vector based on the knowledge graph to obtain the target application scenario, and the scenario processing module is further used to generate a task processing flow framework according to the target application scenario; the process execution module is used to dynamically match the executable tool chain corresponding to the task processing flow framework according to the pre-constructed process database; and, the process execution module is used to call the function modules in the executable tool chain through the third-party application interface to complete the cross-platform task execution; the third-party application module is configured to establish a communication connection with the process execution module.

[0029] In a third aspect, the present invention further provides a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described in the first aspect is implemented.

[0030] In a fourth aspect, the present invention further provides an electronic device, including a memory and a processor, where the memory is used to store one or more computer program instructions, and wherein the one or more computer program instructions are executed by the processor to implement the method described in the first aspect.

[0031] Adopting the above technical solutions, compared with the prior art, the present invention has the following beneficial effects:

[0032] This technical solution adopts a multi-modal fusion wake-up mechanism, integrates voice, text, device behavior patterns, and environmental sensor data, and realizes full-scenario adaptive triggering. Based on the task semantic vector generation technology of large language models and combined with the associative reasoning ability of knowledge graphs, it realizes in-depth scenario understanding. The dynamic toolchain matching mechanism relies on the real-time update ability of the process database and supports the flexible configuration of the optimal execution path. Through the third-party application interface, seamless connection of cross-platform operations is achieved, and users do not need to pay attention to the underlying application differences, reducing the user interaction steps and showing efficiency advantages and fault tolerance capabilities in complex scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0034] Figure 1 It is a step diagram of the task dynamic processing method. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0035] The following will further describe the present invention in detail with reference to the drawings and embodiments. It should be specifically noted that the following embodiments are only used to illustrate the present invention, but do not limit the scope of the present invention. Similarly, the following embodiments are only partial embodiments of the present invention rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.

[0036] Please refer to Figure 1 , in the first aspect, this embodiment provides a task dynamic processing method based on a large language model, including:

[0037] S101. Receive a multi-modal wake-up signal, where the multi-modal wake-up signal includes at least one of voice input, text input, device behavior pattern input, and environmental status automatic trigger signal;

[0038] S102. Convert the multi-modal wake-up signal into language information, parse the user's needs through the large language model with the language information, and generate a task semantic vector including application scenario classification;

[0039] S103. Conduct scenario association analysis on the task semantic vector based on the knowledge graph to obtain the target application scenario;

[0040] S104. Generate a task processing flow framework according to the target application scenario;

[0041] S105. Dynamically match the executable tool chain corresponding to the task processing flow framework according to the pre-built process database;

[0042] S106. Invoke the function modules in the executable tool chain through a third-party application interface to complete cross-platform task execution.

[0043] The steps shown in this embodiment can be understood as the application of an AI agent. The current wake-up method of the AI agent is relatively single, while the AI agent shown in this embodiment supports the triggering of multi-modal wake-up signals. Specifically, the multi-modal wake-up signal refers to a composite trigger signal collected through multiple sensing channels (such as voice, text, device sensors, environmental sensors). For example, the user inputs their voice data in a way such as using a microphone, and the user inputs information using the text input component on the terminal device. The device behavior mode input specifically includes: the combination of the power-off key and the volume key, the finger touch signal on the touch screen (such as three-finger swipe, etc.), and the instantaneous acceleration detected by the acceleration sensor integrated in the terminal device when the terminal device shakes as a trigger signal, and so on. By analogy, the environmental state automatic trigger can be in multiple modes. For example, the change in the user's location when getting off work enables the AI agent to perceive the time when the user arrives home, and then make a home furnishing plan in advance. Another example is that when the temperature and humidity sensor detects that the indoor temperature exceeds 30°C, the air conditioner adjustment task is automatically started. The setting of the multi-modal wake-up signal can break through the limitations of a single interaction method and achieve full-scenario perception triggering.

[0044] Language information conversion is also the process of converting unstructured multi-modal signals into structured text descriptions. For example, the voice signal is converted into "Please send the meeting minutes to the team" through ASR (Automatic Speech Recognition); the device behavior signal (such as continuously shaking the mobile phone) is mapped to "The user needs to initiate an emergency contact call"; the environmental signal (such as PM2.5 exceeding the standard) is converted into "The current air quality is poor, it is recommended to turn on the air purifier". This method can achieve unified input of heterogeneous data and provide standardized data for subsequent semantic parsing.

[0045] The large language model can be understood as a pre-trained deep learning model with language processing capabilities (such as GPT-4, PaLM, Claude). The large language model can decompose the language information containing the user's instructions into multiple phrases (such as parsing "Send an email to Zhang San" as {Recipient: Zhang San, Action: Email Sending}); identify the domain to which the language information belongs (such as classifying "Reserve a meeting room" as the "Office Collaboration Scenario"); generate a task semantic vector: output a vector representation containing scene labels (such as 0.87 Office + 0.12 Schedule) and operation intentions. Specifically, for example, when inputting "What should I bring for the business trip tomorrow morning", the output vector may be encoded as [Scene: Business Trip, Action: Generate a packing list, Associated Entities: Flight Schedule / Weather Data].

[0046] The scene association analysis of the knowledge graph can be understood as logical reasoning based on a domain knowledge base (such as an enterprise office knowledge graph, a smart home device graph). Specifically, keywords in the task vector (such as "projector") are linked to device nodes in the knowledge graph; implicit requirements are discovered through graph paths (such as "preparing presentation materials" → associated with "checking printer ink level" and "file format conversion"); the target scene is derived by integrating the context (such as for the "working from home scene", the local device interface should be preferentially called instead of the cloud service). The similarity between the task vector and the graph nodes can be calculated using a graph neural network (GNN).

[0047] The task processing flow framework can be understood as a standardized processing template generated according to the target scene. For example, the meeting arrangement scene framework includes: extracting the list of participants, querying available calendar slots, booking a meeting room, and sending invitations. Preferably, it can be realized by mining the process patterns based on historical task logs.

[0048] The pre-built process database is a database of pre-stored structured process templates, including operation libraries for basic functions such as {file conversion, API call, data query} and other basic functions; a mapping table between scenes and processes, as well as toolchain metadata. Optionally, an update mechanism can be set, such as dynamically expanding new processes using online learning.

[0049] Dynamic matching means selecting the optimal tool combination according to the real-time context (such as device status, network latency). For example, when it is detected that the user is offline, the local OCR tool is automatically selected instead of the cloud API; when the target service is detected to be busy, a load balancing strategy is enabled to switch to the backup interface. The executable toolchain is a set of function modules arranged in sequence to complete the entire process.

[0050] The third-party application interface call integrates cross-platform services through a standardized adaptation layer (such as REST API, gRPC). For example, the local printing service of the Windows system and the Reminders application interface of iOS are simultaneously called to achieve seamless operation of "printing a file and adding a reminder".

[0051] In this embodiment, multi-modal signals such as voice, text, device behavior, and environmental sensors are received (such as voice commands, device combination operations, or environmental threshold triggers), and the heterogeneous inputs are unified into structured language descriptions through a signal conversion module; subsequently, a large language model is used to parse the user's intention and generate a semantic vector containing scene tags (such as identifying the "business trip preparation" requirement and encoding it as a business travel scene vector), and context reasoning is performed in combination with a knowledge graph to determine the target scene (such as associating flight information and weather data); a task process framework is automatically generated based on scene features (such as the "itinerary planning - reminder setting - file sorting" logic chain for the business travel scene), and the optimal tool chain is dynamically matched through a pre-built process database (such as preferentially calling local services when offline); finally, cross-platform function modules are scheduled through a standardized interface (such as operating the email system and smart home devices simultaneously), realizing the automated execution of multi-terminal collaborative tasks. This method significantly improves the adaptive ability and execution efficiency of the AI system in complex scenarios through a technical closed-loop of multi-modal perception, semantic understanding, and dynamic orchestration.

[0052] Correspondingly, in the second aspect, this embodiment also provides a task dynamic processing system based on a large language model, which is applicable to the processing method described in the first aspect. The system includes a signal acquisition module, a semantic parsing module, a scene processing module, a process execution module, and a third-party application module; the signal acquisition module is used to receive multi-modal wake-up signals, and the multi-modal wake-up signals include at least one of voice input, text input, device behavior pattern input, and environmental state automatic trigger signals; the semantic parsing module is used to convert the multi-modal wake-up signals into language information, parse the user's requirements through the large language model, and generate a task semantic vector containing application scene classification; the scene processing module is used to perform scene association analysis on the task semantic vector based on the knowledge graph to obtain the target application scene, and the scene processing module is also used to generate a task processing process framework according to the target application scene; the process execution module is used to dynamically match an executable tool chain corresponding to the task processing process framework according to a pre-built process database; and, the process execution module is used to call function modules in the executable tool chain through a third-party application interface to complete cross-platform task execution; the third-party application module is configured to establish a communication connection with the process execution module.

[0053] This embodiment adopts a multi-modal fusion wake-up mechanism, integrating voice, text, device behavior patterns, and environmental sensor data to achieve full-scenario adaptive triggering. For example, in a driving scenario, the system can prioritize parsing the steering wheel press signal over voice input to ensure operation safety; in a smart home environment, temperature and humidity sensor data can automatically trigger air conditioner adjustment tasks without explicit instructions. Based on the task semantic vector generation technology of large language models, combined with the associative reasoning ability of knowledge graphs, it realizes in-depth scenario understanding. For example, when the user says "Prepare business trip materials for tomorrow morning", the system can associate multi-source knowledge nodes such as flight information, hotel reservations, and weather data, and automatically generate a composite task flow including file printing, itinerary reminders, and clothing suggestions. The dynamic toolchain matching mechanism relies on the real-time update ability of the process database to support the flexible configuration of the optimal execution path. For example, when it detects that the target cloud storage service is down, the system can automatically switch to a backup platform and maintain task continuity. Through a standardized third-party interface abstraction layer, seamless cross-platform operation is achieved, and users do not need to concern themselves with underlying application differences. For example, when calling the document editing function, it can automatically adapt to different ecological interfaces such as Office 365 and Google Docs. This embodiment significantly improves the intelligence and robustness of task processing through multi-dimensional technological innovations.

[0054] In some embodiments, the device behavior pattern input includes at least one of a specific shaking pattern, a specific key pattern, and a specific gesture; the environmental state automatic trigger signal includes a trigger condition for reaching a preset area based on the matching of a geographical location sensor and a timestamp.

[0055] In this embodiment, the device behavior pattern input refers to triggering instructions by detecting specific physical operations of the user on the device, including a preset shaking pattern (such as shaking the mobile phone twice continuously to initiate an emergency call), a specific key combination (such as long-pressing the volume key and the power key simultaneously to activate the voice assistant), and a custom gesture (such as drawing a "C" shape on the screen to directly open the camera); the environmental state automatic trigger signal refers to a composite condition formed by a geographical location sensor (such as GPS / Beidou positioning) and a timestamp matching logic (such as 18:00 - 19:00 on weekdays). For example, when the system detects that the user's mobile phone enters the geographical fence range of the residential community at 18:30, it automatically executes the preset task of "turning on the home air conditioner and starting the floor cleaning robot", realizing scenario-based intelligent response without active operation.

[0056] In this embodiment, by integrating a dual-trigger mechanism of physical interaction and environmental perception, the scene adaptability and operation convenience of intelligent devices are significantly improved: on the one hand, diverse device behavior mode inputs (such as specific shakes / button presses / gestures) allow users to quickly trigger functions in emergency or inconvenient voice operation scenarios, enhancing the flexibility and reliability of human-computer interaction; on the other hand, intelligent environment judgment based on geographical location and timestamp (such as automatically turning on devices when arriving home) realizes seamless task execution, reducing the manual operation steps of users, especially showing efficient automatic response capabilities in multi-task concurrent scenarios.

[0057] In some embodiments, the environmental state automatic trigger signal is configured to be generated through the following steps:

[0058] Continuously monitor the movement trajectory and system time of the user's terminal device in real time;

[0059] When it is detected that the user leaves the work location coordinates and the moving direction points to the home address coordinates, predict the user's needs through the historical behavior pattern database, and generate a home service type wake-up instruction according to the user's needs.

[0060] In this embodiment, the system continuously collects the user's location coordinates through the positioning modules such as GPS and Beidou of the terminal device, records the movement trajectory at a preset sampling frequency (such as once per second), and synchronously binds the system timestamp accurate to milliseconds; uses trajectory smoothing algorithms such as Kalman filtering to eliminate positioning jitter errors, and combines with the geofence technology to judge in real time whether the user crosses the preset regional boundary (such as within a range of 50 meters from the work location coordinates). When it is detected that the user leaves the work area and the deviation between the moving direction and the azimuth of the home address is less than the threshold, trigger the environmental state analysis process. For example, when the user leaves the company fence range at 18:15 and the movement trajectory shows that they are moving towards the home coordinates at a speed of 30 km / h, after the system determines that it meets the "commuting period" condition, it enters the demand prediction stage.

[0061] Based on the user's periodic behavior data (such as arrival time at home, device operation records) stored in the historical behavior pattern database, an LSTM time series prediction model can be used to analyze the similarity between the current moving speed, direction and historical trajectory, and calculate the time window for the expected arrival at the target area. Combining the user's preferences (such as turning on the air conditioner 10 minutes earlier in summer) and external environmental parameters (such as real-time weather), generate a home service queue containing device control instructions and execution times, and send it to the home Internet of Things gateway through message protocols such as MQTT. For example, if it is predicted that the user will arrive at the residence at 18:50, the system automatically generates an instruction of "start the living room air conditioner to 24°C at 18:45" and sends it to the smart home central control.

[0062] In this embodiment, through the fusion of spatio-temporal data and personalized prediction models, precise pre-execution of home services is achieved. Commands are automatically triggered based on movement trajectories and time patterns, reducing manual operations by users; the startup time of devices is dynamically calculated (such as calculating the air conditioner startup time by backtracking according to the arrival time), reducing ineffective energy consumption; control parameters are adjusted in combination with variables such as season and weather (such as switching to floor heating priority in winter), enhancing the user experience. Breaking through the traditional passive response mode of smart homes, a "perception-prediction-execution" closed-loop is constructed, significantly improving the system's intelligence level and user experience.

[0063] In some embodiments, the knowledge graph includes a multi-layer semantic network structure. The multi-layer semantic grid structure includes a first layer, a second layer, and a third layer; the first layer is configured to divide online service scenarios and offline entity operation scenarios; the second layer is configured to divide into a commodity procurement module, a housekeeping service module, and a smart home control module according to service fields; the third layer is configured to define the association rules and execution priorities between the commodity procurement module, the housekeeping service module, and the smart home control module.

[0064] In this embodiment, the knowledge graph uses a three-layer semantic network structure to achieve refined scenario modeling. The first layer (scenario classification layer) divides online service scenarios (such as e-commerce shopping, online reservation) and offline entity operation scenarios (such as logistics distribution, device control) through ontology, and establishes mutually exclusive and inclusive relationships between scenarios; the second layer (domain module layer) is further divided into vertical modules such as commodity procurement (SKU management, price comparison strategy), housekeeping service (cleaning schedule, personnel scheduling), and smart home control (device linkage, energy consumption optimization) according to service types on the basis of scenario classification; the third layer (rule logic layer) defines cross-module association rules (such as automatically triggering the call of the logistics interface after commodity procurement) and execution priorities (urgent housekeeping needs > regular procurement tasks), and represents the dependency relationship between modules through weighted directed edges.

[0065] For example, when the user triggers the "preparation for a family gathering" task, the system first executes the smart home control module (lighting / air conditioner adjustment), and then concurrently calls the commodity procurement module (fresh food ordering) and the housekeeping service module (kitchen cleaning), and dynamically adjusts the resource allocation ratio of each subtask according to historical data.

[0066] In this embodiment, through the gradual refinement of scenario-domain-rules, efficient decomposition and collaborative scheduling of complex tasks are achieved. The modular design improves the task response speed and reduces the resource conflict rate; based on the hierarchical association rules, logical contradictions are automatically avoided (such as not starting the cooking equipment before the goods arrive); each layer supports independent updates (such as only expanding the second layer when adding a medical and health module), reducing the system transformation cost. This embodiment provides an interpretable and evolvable decision-making framework for multi-modal task processing.

[0067] In some embodiments, generating a task processing flow framework according to a target application scenario includes:

[0068] Extracting intent entities and constraint conditions in language information through a large language model;

[0069] Matching the intent entities and constraint conditions with a knowledge graph to generate a process topology structure, which includes a subtask sequence, an execution parameter transfer path, and an exception handling strategy.

[0070] In this embodiment, through a large language model, in-depth semantic parsing is performed on user input to identify core intent entities (such as the "restaurant" entity and the "reservation" action in "reserving a restaurant") and constraint conditions (such as "within 200 yuan per person" and "need a window seat"). The large language model uses joint training of named entity recognition (NER) and dependency syntactic analysis to ensure accurate extraction of key parameters in complex sentence patterns (such as negation and parallel structures), and at the same time, through the attention mechanism, explicit requirements and implicit preferences are distinguished (such as the implicit "preference for private rooms" constraint in a "quiet environment").

[0071] In this embodiment, the extracted entities and constraint conditions are mapped to the nodes of the knowledge graph. Specifically, the subtask sequence is an ordered operation chain generated based on the service link in the knowledge graph, such as "restaurant reservation → route planning → reminder setting"; the parameter transfer path can establish cross-task parameter dependencies, such as automatically transferring the address of the successfully reserved restaurant to the navigation module; according to the fault tolerance rules in the graph (such as "enable an alternative list when the restaurant is full"), branch logic is preset, which is the exception handling strategy. The execution path is optimized through a graph traversal algorithm to ensure a balance between resource consumption and success rate.

[0072] For example, when the user inputs "Reserve a hot pot restaurant suitable for 6 people for dinner this Friday at 7 pm, which should be near the subway station and support private rooms", the system executes the following process:

[0073] Entity extraction: Identify intent entities {action: reserve, object: hot pot restaurant}, constraint conditions {time: Friday 19:00, number of people: 6, requirement: private room, location: near the subway station};

[0074] Graph matching: Retrieve nodes that meet the conditions in the catering service sub-graph and associate the link "hot pot restaurant reservation → private room availability query → subway station distance calculation";

[0075] Topology generation:

[0076] Subtask sequence: [restaurant search → private room confirmation → online reservation → generate navigation link]

[0077] Parameter transfer: Transfer the longitude and latitude of the selected restaurant to the Gaode Map API to generate a navigation path;

[0078] Exception strategy: If there is no private room in the preferences, try alternative restaurants in descending order of rating and trigger user confirmation.

[0079] In this embodiment, through semantic deep parsing and graph-based process orchestration, efficient automated processing of complex tasks is achieved, supporting real-time updating of the knowledge graph to cope with service changes (such as the entry of new restaurants), resulting in an increase in task success rate; reducing repeated queries through parameter passing (such as address reuse); and presetting exception strategies to reduce the process interruption rate.

[0080] In some embodiments, the subtask sequence is configured to be constructed using a dynamic sharding mechanism;

[0081] The steps for constructing the subtask sequence include:

[0082] Split the complex task output by the knowledge graph into independently executable sub-operation units;

[0083] Establish a data dependency graph among multiple sub-operation units;

[0084] Through the automatic filling of cross-unit parameters in the data dependency graph by the context-aware module, the cross-unit parameters are configured as unknown parameters in the data dependency graph.

[0085] In this embodiment, splitting the complex task into sub-operation units can be understood as follows: Based on the modular decomposition strategy of the knowledge graph, a composite task (such as "organizing a team dinner") is decomposed into sub-operation units that can be executed in parallel or serially. Each sub-operation unit (such as "restaurant screening", "budget checking", "personnel notification") corresponds to the smallest executable node in the knowledge graph, ensuring that it does not depend on external states when running independently; further, independent computing resources (such as dedicated API call quotas, memory buffers) are allocated to each sub-operation unit to avoid interference between tasks. Through the definition of service boundaries between knowledge graph nodes (such as payment services and schedule services belonging to different subgraphs), cutting is performed in combination with the principle of business logic integrity.

[0086] Modeling the data dependency graph among sub-operation units using a graph structure can be understood as follows: The data dependency graph includes explicit dependencies and implicit dependencies. Explicit dependencies are direct parameter passes (such as restaurant address as the input of the navigation service); implicit dependencies are temporal constraints (such as payment must be completed first to trigger order confirmation). By extracting interface definitions through static code analysis and mining historical dependency patterns from runtime logs, a weighted directed graph (edge weights represent dependency strength) is generated, thereby obtaining a complete data dependency graph.

[0087] For the unknown parameter nodes in the data dependency graph, a context-aware module can be used for filling, specifically including historical pattern inference, environmental state completion, preference model prediction, etc. Specifically, historical pattern inference is to retrieve similar task records (e.g., 80% of the user's dinner arrangements are between 18:00 and 19:30); environmental state completion is to utilize device sensor data (e.g., if currently located in the office area, it is default to execute after work); preference model prediction is to recommend matching parameters based on the user profile (e.g., vegetarian preference). Deploy a lightweight decision tree model to generate candidate parameters in real time and filter out invalid options through a confidence threshold.

[0088] Specific examples are as follows:

[0089] When the user initiates the task of "preparing for a product launch event":

[0090] Task splitting: Decompose it into sub-units such as [venue rental → equipment debugging → invitation sending → media docking], etc.;

[0091] Dependency graph: Media docking depends on the completion of invitation sending (obtaining the list of participants), and equipment debugging depends on the venue coordinates (logistics path planning);

[0092] Parameter filling: When the venue area is not specified, it is automatically filled according to the historical ratio of the number of participants (e.g., 100 people ≈ 200㎡), and the user is triggered to confirm.

[0093] In this embodiment, through the dynamic sharding and intelligent parameter completion mechanism, the efficient and reliable execution of complex tasks is realized. For adding new sub-operation units, only the dependency graph needs to be updated, and the system reconstruction cost is reduced; the execution efficiency is optimized, and the parallelization of independent sub-tasks shortens the overall processing time; the real-time monitoring of the dependency graph supports local retry (e.g., a single API call failure does not affect the whole), and the fault tolerance ability is improved; the user experience is enhanced, and the automatic parameter filling reduces the user input steps, especially remarkable in multi-constraint tasks.

[0094] In some embodiments, the method further includes:

[0095] After the cross-platform task is executed, return the execution result to the user;

[0096] The execution result includes at least one of an execution report in natural language form, a physical state feedback signal of a smart home device, and a confirmation message sent by a third-party service platform.

[0097] In this embodiment, the execution report in natural language form can be: Using a natural language generation (NLG) engine (such as a GPT-3 fine-tuning model), convert the structured task log (such as API response codes, device status codes) into a user-readable text summary.

[0098] The text summary can be a success indicator, an anomaly explanation, or an operation suggestion. The success indicator clearly marks the task completion rate (such as "3 / 4 sub-tasks completed"); the anomaly explanation provides a layman's explanation for a failed operation (such as "Failed to turn on the air conditioner: The filter needs to be replaced is detected"); the operation suggestion provides a solution for unfinished items (such as "Click to retry" or "Contact the property management to check the circuit").

[0099] The steps for generating the physical state feedback signal of the smart home device can be: polling the device status in real time through the IoT protocol (such as MQTT, CoAP); or, in the cases of direct feedback, indirect verification, and anomaly detection, directly feedback by reading the device sensor data (such as the current temperature of the air conditioner, the brightness value of the light); indirectly verify the physical changes through the camera (such as whether the curtain is closed). Anomaly detection compares the difference between the expected state and the actual state (such as triggering an alarm when the set temperature is 24°C but the actual temperature is 28°C).

[0100] The third-party service platform sends a confirmation message, which can be API callback listening, including subscribing to the event notification interfaces of the third-party platform (such as the Alipay payment success callback, the SF Express logistics status push);

[0101] In this embodiment, through the multi-modal feedback fusion mechanism, the transparency and user trust of task execution are significantly improved. By cross-verifying the structured log and the physical signal, the root cause of the problem can be quickly located (such as distinguishing between network failures and device failures); the natural language report reduces the technical understanding threshold, enabling non-professional users to clearly master the details of task execution; automatically aggregating messages from multiple platforms reduces the time for users to manually verify, which is particularly valuable in cross-border logistics and cross-brand smart home scenarios.

[0102] In a third aspect, this embodiment also provides a computer-readable storage medium, on which computer program instructions are stored, and the computer program instructions, when executed by a processor, implement the method described in the first aspect.

[0103] The computer program involved in this embodiment can be stored in a computer-readable storage medium, which includes but is not limited to magnetic disks, magnetic tapes, magnetic cards, floppy disks, flash memories, optical discs, optical cards, read-only memories (ROMs), random access memories (RAMs), erasable programmable ROMs (EPROMs), and electrically erasable programmable ROMs (EEPROMs), etc., and also includes other biological, physical, or chemical structures that can achieve functions similar to or equivalent to the above-listed storage media, such as units with information storage capabilities like DNA, RNA, proteins, etc. In a specific embodiment, the storage medium involved can be one of the above medium types or a combination of the above medium types. In different embodiments, the computer program involved in the embodiment can be centrally stored in a single medium or distributedly stored in multiple media. The memory containing the computer-readable storage medium can be a non-volatile memory or a random access memory. These computer-readable storage media can be built into the device or can be an external device or a part of an external device connected to the device involved in the embodiment. In some embodiments, the memory with the computer-readable storage medium is deployed locally; in other embodiments, a scheme of deploying the memory away from the processor can also be adopted, such as a network-attached memory accessed via an RF circuit or an external port and a communication network, where the communication network can be the Internet, one or more intranets, local area networks (LANs), wide area wireless networks (WLANs), storage area networks (SANs), etc., or an appropriate combination thereof, as long as the computer device can access the memory. In addition, the computer program involved in the embodiment can be stored in plaintext / ciphertext form or can be designed as training data and be integrally reorganized and implicitly stored in the parameter states of a deep neural network or other machine learning models through model training.

[0104] In a fourth aspect, this embodiment further provides an electronic device, including a memory and a processor, where the memory is used to store one or more computer program instructions, and wherein the one or more computer program instructions are executed by the processor to implement the method described in the first aspect.

[0105] The processor described in this embodiment can be implemented by hardware, firmware, software, or a combination thereof. It can use circuits, one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), central processing units (CPUs), controllers, microcontrollers, microprocessors, or at least one of the above, and also includes other physical, biological, or chemical structures that can achieve functions similar to or equivalent to those of the above-listed processors, such as biological neurons, quantum computing units, DNA computing units, etc., so that the processor can execute some steps, all steps, or any combination of the steps mentioned in the computer programs or methods involved in the various embodiments of the present application.

[0106] This technical solution adopts a multi-modal fusion wake-up mechanism, integrates voice, text, device behavior patterns, and environmental sensor data, and realizes full-scenario adaptive triggering. Based on the task semantic vector generation technology of large language models, combined with the associated reasoning ability of knowledge graphs, it realizes in-depth scene understanding. The dynamic tool chain matching mechanism relies on the real-time update ability of the process database to support the flexible configuration of the optimal execution path. Through the third-party application interface, seamless connection of cross-platform operations is achieved, and users do not need to pay attention to the differences in underlying applications, reducing the user interaction steps and showing efficiency advantages and fault tolerance capabilities in complex scenarios.

[0107] The above are only some embodiments of the present invention, and thus do not limit the protection scope of the present invention. Any equivalent device or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present invention.

Claims

1. A task dynamic processing method based on a large language model, characterized in that: include: Receiving a multimodal wake-up signal, wherein the multimodal wake-up signal includes at least one of a voice input, a text input, a device behavior mode input, and an automatic trigger signal of an environmental state; Convert the multimodal wake-up signal into language information, parse the language information into user needs through a large language model, and generate a task semantic vector including application scenario classification; Performing scenario association analysis on the task semantic vector based on the knowledge graph to obtain the target application scenario; Generate a task processing flow framework according to the target application scenario; Dynamically matching an executable tool chain corresponding to the task processing process framework according to a pre-built process database; The function modules in the executable tool chain are called through a third-party application interface to complete cross-platform task execution.

2. The task dynamic processing method based on a large language model according to claim 1 is characterized in that: The device behavior mode input includes at least one of a specific shaking mode, a specific key mode, and a specific gesture; The automatic environmental status trigger signal includes a trigger condition of reaching a preset area based on a match between a geographic location sensor and a timestamp.

3. The task dynamic processing method based on a large language model according to claim 1 is characterized in that: The environmental status automatic trigger signal is configured to be generated through the following steps: Real-time monitoring of the movement trajectory and system time of the user's terminal device; When it is detected that the user leaves the workplace coordinates and moves towards the home address coordinates, the user needs are predicted through the historical behavior pattern database, and a home service wake-up instruction is generated according to the user needs.

4. The task dynamic processing method based on a large language model according to claim 1 is characterized in that: The knowledge graph includes a multi-layer semantic network structure, and the multi-layer semantic grid result includes a first level, a second level, and a third level; The first level is configured to divide online service scenarios and offline entity operation scenarios; The second level is configured to divide the commodity purchasing module, the housekeeping service module, and the smart home control module according to the service field; The third level is configured to define association rules and execution priorities among the commodity procurement module, the housekeeping service module, and the smart home control module.

5. The task dynamic processing method based on a large language model according to claim 1, characterized in that: Generating a task processing flow framework according to the target application scenario includes: Extracting the intent entities and constraints in the language information through a large language model; The intention entity and constraint conditions are matched with the knowledge graph to generate a process topology structure, which includes a subtask sequence, an execution parameter transfer path, and an exception handling strategy.

6. The task dynamic processing method based on a large language model according to claim 5 is characterized in that: The subtask sequence is configured to be constructed using a dynamic sharding mechanism; The steps of constructing the subtask sequence include: Splitting the complex task output by the knowledge graph into independently executable sub-operation units; Establish a data dependency graph between multiple sub-operation units; The cross-unit parameters in the data dependency graph are automatically filled in by the context-aware module, and the cross-unit parameters are configured as unknown parameters in the data dependency graph.

7. The task dynamic processing method based on a large language model according to claim 1 is characterized in that: The method further comprises: After the cross-platform task is executed, the execution result is returned to the user; The execution result includes at least one of an execution report in natural language form, a physical state feedback signal of the smart home device, and a confirmation message sent by a third-party service platform.

8. A task dynamic processing system based on a large language model, characterized in that: The method according to any one of claims 1 to 7, wherein the system comprises a signal acquisition module, a semantic analysis module, a scene processing module, a process execution module and a third-party application module; The signal acquisition module is used to receive a multimodal wake-up signal, wherein the multimodal wake-up signal includes at least one of a voice input, a text input, a device behavior mode input, and an automatic trigger signal of an environmental state; The semantic parsing module is used to convert the multimodal wake-up signal into language information, parse the language information into user needs through a large language model, and generate a task semantic vector containing application scenario classification; The scenario processing module is used to perform scenario association analysis on the task semantic vector based on the knowledge graph to obtain a target application scenario, and the scenario processing module is also used to generate a task processing flow framework according to the target application scenario; The process execution module is used to dynamically match the executable tool chain corresponding to the task processing process framework according to the pre-built process database; and the process execution module is used to call the functional module in the executable tool chain through the third-party application interface to complete the cross-platform task execution; The third-party application module is configured to establish a communication connection with the process execution module.

9. A computer-readable storage medium storing computer program instructions, characterized in that: The computer program instructions implement the method according to any one of claims 1 to 7 when executed by a processor.

10. An electronic device comprising a memory and a processor, characterized in that: The memory is used to store one or more computer program instructions, wherein the one or more computer program instructions are executed by the processor to implement the method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Cross-platform task intelligent collaboration method and device based on natural language and storage medium

    CN120335970A

  • A natural language-based cross-platform task intelligent collaboration method, device and storage medium

    CN120335970B

  • Intelligent customer service method, device and system

    CN120353906A

  • Intelligent customer service method, device and system

    CN120353906B

  • Intelligent home recommendation method and system based on cross-modal fusion

    CN120929675A