Task processing method, device and equipment

By deploying intelligent agents in local electronic devices, cross-device aggregation of multimodal data and personalized model training are achieved, solving the problems of data fragmentation across multiple devices and slow training response. This improves the real-time performance and intelligence of task processing, and reduces data latency and privacy risks.

CN121523819APending Publication Date: 2026-02-13LENOVO (BEIJING) LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511554970.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-28
Publication Date
2026-02-13

AI Technical Summary

Technical Problem

In existing technologies, data fragmentation across multiple devices, slow training response, and low personalization result in insufficient real-time performance and intelligence in task processing. Furthermore, there is a lack of dynamic adaptation capabilities to users' long-term goals, leading to data transmission delays and privacy and security risks.

Method used

By deploying intelligent agents in local electronic devices, multimodal data is aggregated across devices. Based on preset parameter management and timestamp management, real-time data synchronization and personalized model training are achieved. Using preset parameter and timeline management mechanisms, combined with user-defined target tasks, personalized model training and inference are driven to form a closed-loop learning mechanism.

Benefits of technology

It achieves efficient integration of data from multiple devices and local real-time feedback, improves the real-time and personalized level of task processing, reduces data latency and privacy risks, and ensures the model's adaptability and security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121523819A_ABST
    Figure CN121523819A_ABST
Patent Text Reader

Abstract

The invention provides a task processing method, device and equipment, and the method comprises the steps: obtaining first multi-modal data sent by at least one second electronic equipment through an intelligent agent in first electronic equipment; managing the first multi-modal data and pre-stored second multi-modal data based on preset parameters; in response to an instruction of selecting a target task, determining target data matched with the target task in the first multi-modal data and the second multi-modal data; and executing the target task based on the target data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence, and includes, but is not limited to, a task processing method, apparatus, and device. Background Technology

[0002] With the widespread use of smart devices, the amount of multimodal data generated by users across various terminals is increasing daily. This data includes various forms such as text, images, and voice, and is characterized by real-time and personalization. To achieve efficient task processing, such as health monitoring, personalized recommendations, or intelligent interaction scenarios, it is necessary to fuse and analyze data from different devices and respond quickly based on the current context.

[0003] In related technologies, data from multiple devices is typically processed and models trained using a centralized server. This method relies on cloud computing resources, extracting relevant features from historical data and performing inference operations upon receiving a user's task request. However, due to data transmission latency, long training cycles, and a lack of dynamic adaptation to long-term user goals, task response is untimely and results lack generalization. It struggles to address issues such as fragmented data from multiple devices, slow training response, low personalization, and separation of training and inference, limiting the real-time performance and intelligence of task processing. Therefore, there is an urgent need for a task processing method that can integrate local multi-source data, support immediate feedback, and possess continuous learning capabilities. Summary of the Invention

[0004] This application provides a task processing method, apparatus, and device.

[0005] The technical solution of this application embodiment is implemented as follows: In a first aspect, embodiments of this application provide a task processing method, including: An intelligent agent configured in a first electronic device acquires first multimodal data sent by at least one second electronic device; The first multimodal data and the pre-stored second multimodal data are managed based on preset parameters; In response to the instruction to select a target task, target data matching the target task is determined from the first multimodal data and the second multimodal data; Execute the target task based on the target data.

[0006] Secondly, embodiments of this application provide a task processing apparatus, including: An acquisition module is used to acquire first multimodal data sent by at least one second electronic device using an intelligent agent configured in a first electronic device; The management module is used to manage the first multimodal data and the pre-stored second multimodal data based on preset parameters; The determination module is used to determine the target data that matches the target task in response to the instruction to select the target task from the first multimodal data and the second multimodal data; The task execution module is used to execute target tasks based on target data.

[0007] Thirdly, embodiments of this application provide an electronic device, including a memory and a processor. The memory stores a computer program that can run on the processor, and the processor executes the program to implement the above-described method.

[0008] Fourthly, embodiments of this application provide a storage medium storing executable instructions for implementing the above-described method when executed by a processor.

[0009] Fifthly, embodiments of this application provide a computer program product, including a computer program or instructions, which, when executed by a processor, implement the steps in the above-described method. Attached Figure Description

[0010] Figure 1 A schematic diagram illustrating the implementation flow of a task processing method provided in an embodiment of this application; Figure 2 This is a schematic diagram illustrating an implementation process for determining target data, provided in an embodiment of this application. Figure 3 A schematic diagram of a display interface provided for an embodiment of this application; Figure 4 A schematic diagram illustrating the implementation flow of a task processing method provided in an embodiment of this application; Figure 5 This is a schematic diagram of the composition structure of a task processing device provided in an embodiment of this application; Figure 6 This is a schematic diagram of a hardware entity of an electronic device provided in an embodiment of this application. Detailed Implementation

[0011] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the specific technical solutions of the embodiments will be further described in detail below with reference to the accompanying drawings. The following embodiments are used to illustrate this application, but are not intended to limit the scope of this application.

[0012] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0013] In the following description, the terms "first, second, third" are used merely to distinguish similar objects and do not represent a specific ordering of objects. It is understood that "first, second, third" may be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.

[0014] This application provides a task processing method, such as... Figure 1 As shown, the method includes: Step S110: The intelligent agent set in the first electronic device acquires first multimodal data sent by at least one second electronic device; Here, the first electronic device is a high-performance local computing device that deploys intelligent agents, such as a personal computer, workstation, or edge server; the second electronic device refers to other terminal devices, such as mobile phones, wearable devices, sensors, cameras, etc. The data generated by the second electronic device is diverse, including but not limited to images, video, audio, text, and physiological signals, collectively referred to as multimodal data. Because the data generated by the second electronic device originates from different terminal devices and is heterogeneous, a unified management mechanism can be established to support the fusion and analysis of this multimodal data.

[0015] The intelligent agent, installed in the first electronic device, acts as an intelligent proxy module and operates on the first electronic device. It has the ability to receive external input, coordinate data flow, perform task matching, and model inference. The intelligent agent installed in the first electronic device can actively listen for data transmission requests from multiple second electronic devices and collect and cache first multimodal data according to preset communication protocols (such as HTTP, MQTT, WebSocket, etc.).

[0016] For example, when a user wears a smartwatch and takes a photo with their phone, the smartwatch can record the user's physiological indicators (such as heart rate and steps), while the phone can upload image data. At this time, the intelligent agent installed in the first electronic device can obtain the corresponding first multimodal data from both the smartwatch and the phone (two second electronic devices), and integrate this data into timeline management for subsequent processing.

[0017] During implementation, the intelligent agent installed in the first electronic device ensures real-time performance and consistency through a cross-device data synchronization mechanism, thereby preventing the loss or delay of the first multimodal data. Furthermore, with all first multimodal data processed locally, the risk of data leakage can be significantly reduced, and privacy and security can be effectively improved.

[0018] Step S120: Manage the first multimodal data and the pre-stored second multimodal data based on preset parameters; Here, preset parameters refer to a pre-configured set of rules or algorithms used to classify, filter, weight, and store newly accessed first multimodal data and existing second multimodal data. These parameters may include: data type identification rules, time window control strategies, feature extraction algorithms, weight allocation logic, etc. They determine which first multimodal data will be retained, how they will be organized, and whether they will participate in subsequent target task matching and model training.

[0019] The pre-stored second multimodal data refers to historical data that has been previously collected and stored in the first electronic device. This type of data is usually stored in a local database or file system. The second multimodal data has also undergone multimodal processing and may exist in the form of images, voice, text, etc., and is arranged in an orderly manner according to the timeline.

[0020] For example, suppose a user records a voice diary entry and takes a photo of their work scene every day. This second-modal data can be stored as pre-stored second-modal data during the first run. Subsequently, whenever new voice or images are received, the agent in the first electronic device can determine, based on preset parameters, whether these new voices or images belong to the same target task (such as content creation optimization) and decide whether to add them to the current dataset.

[0021] During implementation, dynamic management of the first multimodal data and pre-stored second multimodal data enables the continuous construction of high-quality training sets, improving the model's adaptability and accuracy. Simultaneously, preset parameters can be adjusted based on user feedback or the progress of the target task, giving the model stronger adaptive capabilities.

[0022] Step S130: In response to the instruction to select a target task, determine the target data that matches the target task from the first multimodal data and the second multimodal data; Here, the target task is a long-term or phased task direction explicitly set by the user, such as health optimization, assisted diagnosis, or design creation. Users can input the target task into the device through interface operation, voice commands, or API calls, and can filter data related to the target task accordingly.

[0023] In some embodiments, the user can also set up multiple tasks and trigger an instruction to select a target task to be executed.

[0024] Target data refers to content highly relevant to the target task extracted from both the first and second multimodal data sets. Target data may include sensor readings within a specific time period, specific types of images, audio clips, etc. To find target data, techniques such as keyword matching, semantic analysis, timestamp comparison, and pattern recognition can be employed.

[0025] For example, if a user's goal is to optimize their fitness and health, the system can automatically filter out all exercise-related data, such as running records, heart rate changes, and fitness videos. This exercise-related data will then be used for subsequent personalized model training to generate health recommendations that better suit the user's needs.

[0026] During implementation, by accurately matching target data, the focus can be placed on the core issues that users care about, avoiding interference from irrelevant information and improving model training efficiency and result quality. Furthermore, the extraction process of target data can incorporate a Human-in-the-Loop mechanism, allowing users to confirm or correct preliminary screening results, enhancing controllability and transparency.

[0027] Step S140: Execute the target task based on the target data.

[0028] After extracting the target data, task processing will be executed based on this data, such as generating personalized recommendations, performing health diagnoses, optimizing design schemes, and generating videos. Execution methods can include model inference, rule engine triggering, and automated script execution. The specific task processing method depends on the nature of the target task and the characteristics of its architecture.

[0029] For example, for a health diagnosis task, a predictive model might be trained using the target data to predict a user's health status for the coming week and provide optimization suggestions regarding diet, exercise, and sleep. Conversely, for a design and creation task (generating videos), an image style transfer model might be generated based on the target data to help users quickly express their creative ideas.

[0030] During implementation, by executing target tasks based on target data, a learning-by-doing intelligent experience can be achieved, allowing users to continuously optimize their behavior and results in actual operation. Furthermore, since all processing is completed locally, it also offers advantages such as low latency, high security, and strong privacy protection.

[0031] In this embodiment, an intelligent agent is deployed within a first electronic device to achieve cross-device aggregation and management of the first multimodal data. This, combined with user-defined target tasks, drives personalized model training and inference, forming a closed-loop learning mechanism. Throughout the process, attention is paid not only to the diversity and richness of the first multimodal data and pre-stored second multimodal data, but also to the relevance of the target task and user goal orientation, thereby providing more accurate and efficient intelligent services.

[0032] In some embodiments, the preset parameter is a time parameter, and the above step S120 "managing the first multimodal data and the pre-stored second multimodal data based on the preset parameter" can be implemented through the following steps: Step 121: Determine the corresponding timestamp based on the time of acquiring each of the first multimodal data; Here, a timestamp refers to the precise time information recording the moment an event occurs, usually represented in numerical form. Timestamps are used to identify the specific acquisition time of each piece of primary multimodal data, determining its exact position on the timeline. Through timestamps, heterogeneous primary multimodal data generated in real time from different devices can be arranged in an orderly manner, establishing a temporal order, which facilitates subsequent data fusion and model training.

[0033] During implementation, when users record videos with their mobile phones, wear wearable devices to record heart rate data, or type text on their computers, these primary multimodal data points can be automatically timestamped. The timestamps are not only used for sorting but also serve as part of a triggering mechanism to determine whether each piece of primary multimodal data is temporally relevant to the current task objective, thereby deciding whether to include it in the training set.

[0034] By adding a timestamp to each piece of first multimodal data, unified time management of real-time data across terminals can be achieved, ensuring data consistency and timeliness, thereby improving the accuracy and efficiency of subsequent data processing.

[0035] Step 122: Based on the timestamp, connect the first multimodal data to the timeline management of the second multimodal data.

[0036] Here, timeline management refers to integrating newly added first-modal data with existing second-modal data into a unified time-series structure according to timestamps. The timeline management module is not only a data storage container but also a key mechanism for dynamically triggering model updates and data analysis. With the support of timeline management, key time nodes can be identified, data pattern changes can be detected, and model training strategies can be adjusted based on the detected changes in data patterns.

[0037] During implementation, the timeline management module continuously maintains a dynamic timeline structure, inserting first-mode multimodal data streams from different devices into their corresponding timestamps. For example, in a sports and health scenario, heart rate sensor data, step counter data, and user-inputted subjective feelings (such as fatigue or ease) can be uniformly incorporated into the timeline management to form a complete activity trajectory. Subsequently, updates to the personalized training module can be triggered based on key events recorded by the timeline management module (such as sudden increases in heart rate or continuous high-intensity exercise).

[0038] Timeline management enables unified organization and dynamic response to multimodal data across terminals, allowing for timely adjustments to the model training process based on changes in user behavior.

[0039] In this embodiment, the first multimodal data is labeled based on timestamps and then connected to the timeline management module. This operation enables the orderly integration of data across terminals, thereby enhancing data relevance and consistency, and further promoting efficient training and real-time feedback of personalized models.

[0040] In some embodiments, the step S130 above, "determining target data matching the target task from the first multimodal data and the second multimodal data," can be achieved through the following steps: Step S131: Obtain the preset parameters of the target task; Here, preset parameters refer to a set of key conditions defined by the user to complete a specific task, used to filter and match relevant multimodal data. Preset parameters can include specific information such as time, location, and people, or abstract labels or semantic features. For example, in a medical auxiliary diagnosis scenario, preset parameters might include patient ID, medical record time range, and image type; in a sports and health management scenario, preset parameters might include target training intensity, execution time period, and participants. By clearly defining preset parameters, artificial intelligence can accurately locate the dataset most relevant to the task, avoiding invalid data from interfering with model training and inference.

[0041] By acquiring preset parameters, data can be filtered according to the user's long-term task goals, thereby improving the efficiency and accuracy of subsequent processing while reducing the consumption of unnecessary computing resources.

[0042] Step S132: If the preset parameter is determined to be a time parameter, based on the time points corresponding to the first multimodal data and the second multimodal data on the time axis, determine the multimodal data that satisfies the preset matching parameter; Here, the time parameter refers to preset conditions related to the time dimension, such as data from 9 AM to 11 AM, or all data prior to a certain day. The timeline is a dynamically maintained data stream management structure that arranges and updates multimodal data from different devices in chronological order. When the preset parameter is a time parameter, the data on the timeline can be scanned, and multimodal data that fits the time range can be extracted as candidate data. For example, in a health monitoring task, if the user sets the preset parameter to exercise data within the past week, sensor data and video clips within the past week can be extracted from the timeline.

[0043] Filtering data based on the time axis ensures that the selected data is highly relevant to the task objective in the time dimension, improving the relevance of model training and supporting the modeling and prediction of historical behavioral trends.

[0044] Step S133: If the preset parameter is determined to be a location parameter, the multimodal data in the first multimodal data and the second multimodal data that matches the location parameter is determined as the multimodal data that satisfies the preset matching parameter; Here, location parameters refer to preset conditions related to spatial location, such as Chaoyang District in Beijing, company meeting room, gym, etc. The matching of the preset location parameters to the multimodal data can be determined based on geographical location information (such as GPS coordinates, Wi-Fi signal fingerprints, Bluetooth beacons, etc.). For example, in a smart home optimization task, if the user sets the preset parameters as temperature and humidity data for the living room area, then sensor data collected in the living room area can be selected as valid input.

[0045] By filtering data using location parameters, it can be ensured that the spatial environmental information upon which the mission objective depends is accurately captured. This data processing method is suitable for applications such as indoor and outdoor navigation and environmental perception, and can improve spatial perception capabilities.

[0046] Step S134: If the preset matching parameter is determined to be a character parameter, the multimodal data that matches the character parameter in the first multimodal data and the second multimodal data is determined as the multimodal data that satisfies the preset matching parameter. Here, "personal parameters" refer to preset conditions related to user identity or role, such as user A, doctor B, coach C, etc. The identities of individuals in multimodal data are identified through methods such as facial recognition, voice recognition, and device binding, and multimodal data related to the target individual's identity is then filtered out. For example, in a personalized learning task, if the preset parameter is student D's learning records, data such as video recordings, notes, and answer records related to student D can be extracted.

[0047] By filtering data based on user parameters, highly personalized data processing can be achieved. This ensures that model training is based on the behavioral trajectories of specific users. Personalized data processing improves the accuracy of recommendations and predictions. The system enhances the user experience by optimizing recommendation and prediction functions.

[0048] Step S135: Match the multimodal data that meets the preset parameters with the target task to determine the target data.

[0049] Here, target data refers to the dataset ultimately identified as highly relevant to the target task. After completing the above steps, candidate data can be weighted and evaluated by considering multiple parameters such as time, location, and people, and the optimal data combination can be selected as the target data. This process considers not only the matching degree of individual parameters but also factors such as task priority, data quality, and historical performance for comprehensive decision-making. For example, in intelligent creation tasks, high-resolution images, high-quality audio, and text content highly relevant to the task theme can be prioritized to form a complete material library.

[0050] By matching comprehensive parameters, the data that best fits the task objective can be efficiently identified, providing high-quality input for subsequent model training and inference, thereby significantly improving response speed and intelligence level.

[0051] In this embodiment, by introducing preset parameters across multiple dimensions such as time, location, and people, refined filtering and matching of multimodal data is achieved. This ensures that the selected data is highly relevant to the task objective, thereby improving the accuracy and efficiency of model training and ultimately enabling a more personalized and intelligent interactive experience.

[0052] In some embodiments, step S135 above, "matching the multimodal data that satisfies the preset parameters with the target task to determine the target data," is as follows: Figure 2 As shown, this can be achieved through the following steps: Step S210: Determine the task keywords for the target task; Here, task keywords refer to semantic elements or behavioral descriptions that users define and use in a long-term task, possessing core meaning, to characterize the main objectives and focus of the target task. For example, in a health management task, users can use sleep quality, heart rate fluctuations, and exercise intensity as task keywords; in a creative design task, users can set image style, color scheme, and composition layout as task keywords. Task keywords not only help understand the intent of the target task but also serve as the basis for filtering and associating multimodal data.

[0053] By extracting task keywords, we can focus on data segments most relevant to users' long-term goals, avoiding interference from a large amount of irrelevant information, thereby improving the relevance and efficiency of model training. Furthermore, task keywords can serve as a basis for subsequent feature weighting and data prioritization, ensuring that the model continuously aligns with the user's actual needs.

[0054] Step S220: Determine the relevance between the multimodal data that satisfies the preset parameters and the task keywords; Here, relevance refers to the degree of semantic or behavioral matching between the content of multimodal data and the task keywords in a given set of multimodal data. Relevance calculation can be achieved through natural language processing techniques (such as word vector matching), image recognition algorithms (such as object detection), or sensor data analysis. For example, for a health management task, if the task keyword is sleep quality, we can assess indicators such as the presence of snoring, stable heart rate, and suitable ambient lighting in each audio segment, and then comprehensively determine the relevance between the multimodal data and the task keywords.

[0055] Quantifying relevance helps distinguish which data is more likely to contribute to the target task. Through this process, key information can be quickly located within massive datasets, reducing unnecessary computational overhead and improving the effectiveness of personalized training. Furthermore, dynamic changes in relevance reflect the user's task progress and provide real-time feedback for model optimization.

[0056] There is a close relationship between task keywords and relevance. Task keywords define the core focus, while relevance measures whether multimodal data aligns with the core focus defined by the task keywords. Therefore, only when multimodal data has sufficiently high relevance can it be considered a valuable data source and further involved in the subsequent model training process.

[0057] Step S230: If the correlation meets the correlation threshold, the multimodal data that meets the preset parameters is determined as the target data.

[0058] Here, the relevance threshold is a preset numerical standard used to determine whether a segment of multimodal data is relevant enough to be included in the training set. The relevance threshold can be dynamically adjusted based on task type, user preferences, or historical performance. For example, in medical auxiliary diagnostic tasks, due to the sensitivity and importance of the data, the relevance threshold may be set higher to ensure that only high-quality data participates in training; while in creative design tasks, the relevance threshold can be appropriately lowered to capture more sources of inspiration.

[0059] When the relevance of multimodal data exceeds a relevance threshold, the multimodal data is marked as target data and included in the personalized training set. This relevance threshold-based filtering method ensures the quality of training data while avoiding the omission of potentially valuable information. The target data filtering mechanism allows the model to focus on data that truly drives the target task, thereby improving the model's accuracy and generalization ability.

[0060] In this embodiment, the core focus of the target task is first identified using task keywords; second, relevance analysis is performed to determine the degree of matching between multimodal data and the target task; finally, a relevance threshold is set and data is filtered to ensure that the data used for training is sufficiently representative and effective. This hierarchical processing approach, consisting of task keywords, relevance analysis, and relevance threshold filtering, improves the level of intelligence and enhances the accuracy and adaptability of personalized training. Effectively focusing on the core needs of users allows for precise filtering of high-value data, thereby significantly improving personalized training effectiveness and task completion efficiency.

[0061] In some embodiments, step S140 "execute the target task based on the target data" can be achieved through the following steps: Step 141: Obtain the feature weights and prompt words of the target task; Here, feature weights refer to the dynamically assigned values ​​based on the long-term task objective, assigning importance to each feature in the input data during task execution. Feature weights determine the influence of different features on the training and inference of the target model. For example, in a video generation task, feature weights might refer to the user-preset video style: realistic, cartoonish, retro, etc. In a health diagnosis task, heart rate might have a higher feature weight, while step count might have a lower one. These feature weights can be adjusted during subsequent model training. By dynamically adjusting feature weights, the key information of the user's current task can be more accurately focused on, thereby improving personalized processing results.

[0062] Cue words are a set of keywords or phrases used to guide the target model in understanding the task intent. They are typically generated from task descriptions, user needs, and historical behaviors. The purpose of cue words is to help the target model quickly identify the task type and invoke the corresponding processing logic. In creative design tasks, cue words might be in the style of creative illustrations or have a hand-drawn feel.

[0063] There is a synergistic relationship between feature weights and cue words: cue words provide task direction for the target model, while feature weights determine which data elements are more critical in that direction. This mechanism, formed by cue words and feature weights, enables the target model to quickly locate the most valuable information for the task in multimodal data and make more accurate judgments accordingly.

[0064] Step 142: Determine the target model in the intelligent agent that corresponds to the target task; Here, the target model refers to a machine learning model pre-set or dynamically generated within the agent for a specific task, used to perceive and analyze target data and perform the task. The target model can vary depending on the task type; for example, convolutional neural networks (CNNs) are used to extract image features in health diagnosis tasks, while generative adversarial networks (GANs) are used to generate image content in creative design tasks. The choice of target model directly affects the accuracy and efficiency of the task.

[0065] The process of determining the target model depends on the matching degree between task feature weights and prompt words. It can automatically select the most suitable model structure and parameter configuration according to the needs of the current task. For example, when the prompt words indicate that the task is biased towards visual analysis, an image recognition target model can be selected; when the task is biased towards text analysis, a natural language processing target model can be selected.

[0066] During implementation, the target model not only needs to possess good generalization ability but also should be able to flexibly adapt to changes in feature weights and prompts. The requirement for good generalization ability and adaptability in the target model indicates the need for a certain adaptive mechanism, allowing it to quickly adjust its structure and parameters when facing new tasks to maintain optimal performance.

[0067] Step 143: Utilize the target model to perform perceptual analysis on the target data based on the feature weights and the prompt words, in order to execute the target task.

[0068] Here, perceptual analysis refers to the process by which the target model, after receiving target data, performs multi-dimensional analysis of the data by combining feature weights and prompts. The core objective of perceptual analysis is to extract task-relevant features from complex data and map these features into the output of the target model. For example, in medical auxiliary diagnostic tasks, feature weights can be used to emphasize image clarity and lesion areas, and prompts can be used to identify suspected lesion sites.

[0069] The results of perceptual analysis directly determine the effectiveness of task execution. By dynamically adjusting feature weights and prompts, the accuracy of perceptual analysis is optimized, enabling the target model to adapt to different task scenarios and user needs. Furthermore, perceptual analysis supports a real-time feedback mechanism, allowing users to view the decision-making basis of the target model during the analysis process and further optimize task parameters based on feedback information.

[0070] In this embodiment, the direction and focus of the task are first clarified by acquiring feature weights and prompt words; secondly, a suitable target model is selected based on the acquired feature weights and prompt words; finally, the selected target model, combined with feature weights and prompt words, is used to perform perceptual analysis on the target data, thereby completing the specific task. This better meets the user's goal requirements in long-term tasks, further realizing a highly efficient and intelligent learning-by-doing experience. It ensures that efficient and accurate execution capabilities are maintained even when facing diverse tasks.

[0071] In some embodiments, this application also provides a method for training a target model, which can be implemented through the following steps: Step S150: In response to the first instruction to train the target model, the target data is used as training data to train the target model; During implementation, once the electronic device receives the instruction to train the target model, it can input the collected target data as training data into the target model. Target data refers to a filtered and weighted dataset that is relevant to the user's long-term task. Target data may originate from real-time acquisition by multiple terminal devices (such as PCs, mobile phones, wearable devices, etc.), which collect data from various modalities such as images, audio, and sensors.

[0072] A target model is a personalized model that is dynamically built and optimized based on a user's long-term tasks, such as AI models for content creation, health diagnosis, or smart home control. The target model can adjust its parameters based on continuously updated training data to better match the user's task requirements.

[0073] Training models on local high-performance computing devices can avoid uploading sensitive target data to the cloud, thus ensuring privacy and security, while reducing training latency and enabling a real-time interactive experience of learning and using simultaneously.

[0074] Step S160: During the training of the target model, at least one of the following information obtained from training the target model is displayed in a visualization interface in the form of a cloud map: the task keywords of the target task, the feature weights of the target task, and the sample contribution of each target data.

[0075] During the training of the target model, the electronic device provides a cloud-map-style visualization interface to display key information about the training process. This cloud-map visualization is a graphical representation, typically using word size, color, and position to reflect the importance and influence of the displayed information. For example, larger task keywords indicate greater importance of those keywords in the target model training; color changes can represent trends in feature weights; and sample contribution can be reflected by node distribution density, demonstrating the degree of influence of different target data on the target model.

[0076] The information displayed in the cloud-map-style visualization interface includes, but is not limited to: Task keywords for the target task refer to words or concepts highly relevant to the long-term task set by the user. For example, in a creative design task, composition, color scheme, and materials might be the main task keywords. These task keywords reflect the core elements that the target model is currently focusing on and help users understand the direction the target model is learning in.

[0077] Feature weights refer to the relative importance of each feature in the target model in prediction or decision-making. Higher feature weights indicate a greater influence of the feature on the output of the target model. For example, in health diagnosis tasks, vital signs such as heart rate and blood pressure may have higher feature weights.

[0078] Sample contribution refers to the degree to which each piece of target data contributes to the final performance of the target model during training. Target data with high sample contribution means that it plays a crucial role in improving the accuracy and generalization ability of the target model. For example, certain particularly complex or representative case data may have a significant impact on the optimization of the target model.

[0079] The purpose of the cloud-map-style visualization interface is to allow users to intuitively understand the training status and progress of the target model, enhancing interpretability and user engagement. Users can adjust task settings, choose whether to continue using a certain type of target data to train the target model, and even manually modify some feature weights based on the information provided by the cloud-map-style visualization interface, thereby forming a human-machine collaborative training mechanism.

[0080] For example, in medical auxiliary diagnosis scenarios, doctors can view a visualization interface in the form of a cloud map during the training of the target model to observe which symptoms or examination indicators are assigned higher feature weights, and determine whether the target model is focusing on the correct medical features. If it is found that the target model relies too much on certain unimportant features, doctors can intervene in a timely manner, adjust the training strategy, and ensure that the target model conforms to clinical standards.

[0081] In addition, cloud-based visualization interfaces can also be used in education and creative fields. For example, in design tools, designers can use cloud-based visualization interfaces to see which elements (such as lines, colors, and fonts) are frequently used in generating target models, thereby optimizing the designer's creative direction.

[0082] In this embodiment, a cloud-map-style visualization interface is introduced during the training of the target model. This interface displays key information such as the task keywords, feature weights, and sample contribution of each data point. This approach improves the transparency and controllability of the target model training, thereby enhancing user understanding and trust in the model and further promoting the widespread application of personalized intelligence.

[0083] In some embodiments, training the target model includes the following steps: Step S170: In response to the second instruction to adjust task keywords, adjust the task keywords of the target task; Here, the second instruction to adjust task keywords refers to a user-issued command to modify the set of keywords relevant to the current target task. This second instruction can be a manually entered list of new task keywords or keyword update suggestions recommended by the electronic device based on contextual semantics. For example, in a health monitoring task, the initial task keywords might be heart rate, sleep quality, etc. If the user wishes to expand the scope of focus, they can add new task keywords such as respiratory rate, body movement, etc., using the second instruction to adjust task keywords.

[0084] Task keywords determine the priority of content identification and learning by the target model during training. Allowing users to actively adjust these task keywords more precisely guides the target model to focus on the user's long-term goal domain. This approach differs from traditional methods that passively adapt models to historical data; instead, it employs dynamic optimization based on the user's long-term goals. The target model adjusts according to the user's long-term objectives, thereby enhancing the level of personalization.

[0085] Step S180: In response to the third instruction to adjust feature weights, adjust the feature weights of the target task.

[0086] Here, the third instruction to adjust feature weights refers to the instruction issued by the user or electronic device to modify the degree of importance assigned to each feature by the target model. Feature weights reflect the relative influence of different features in the target model's decision-making. These weights are usually automatically generated during the training process, but the current method allows users to intervene according to their actual needs. For example, in creative design tasks, users may want to increase the weight of image style features while reducing the influence of color saturation.

[0087] By allowing users to participate in adjusting feature weights, the interpretability and controllability of the target model can be enhanced. Users can subjectively judge which factors are more important based on their own needs, and this subjective judgment can be directly fed back into the target model training, thereby achieving personalized model training that more closely reflects user intent. Furthermore, the user-participatory feature weight adjustment mechanism aligns with the Human-in-the-Loop design philosophy, improving the efficiency and satisfaction of human-computer collaboration.

[0088] In this embodiment, by introducing user-adjustable task keywords and feature weights, the training of the target model can be more aligned with the user's long-term goals and real-time needs. Introducing user-adjustable task keywords and feature weights improves the relevance and personalization of the target model, thereby more effectively guiding the user towards their set goals and ultimately building a learning-by-doing, continuously optimized intelligent interactive experience.

[0089] In some embodiments, the above step S120, "managing the first multimodal data and the pre-stored second multimodal data based on preset parameters," can be implemented through the following steps: Step 121: Set at least one parameter module on the visual interface; Here, a visual interface refers to a graphical interactive interface where users can directly operate and view the running status of a method. It typically includes elements such as buttons, sliders, and charts to display parameter configurations and data status. Visual interfaces not only improve human-computer interaction efficiency but also enhance the transparency and controllability of methods.

[0090] A parameter module is a set of parameters used to adjust or control specific functions, such as data fusion methods, weight allocation strategies, and model training frequencies. Each parameter module can be set independently or linked with other parameter modules to form a flexible configuration system. By presenting these parameters as modules, users can quickly switch or adjust the relevant settings of parameter modules according to their actual needs.

[0091] For example, when using a timeline to manage data, multimodal data can be displayed based on different time period modules.

[0092] In the personalized health management system, users can select different parameter modules on the visual interface to set goals (such as weight loss rate, sleep quality optimization, etc.). The personalized health management system adjusts the processing logic of multimodal data in real time according to the target parameters set by the user, thereby achieving more accurate health advice.

[0093] Step 122: Display the first multimodal data and the second multimodal data corresponding to each parameter module.

[0094] Here, the first multimodal data refers to the new data currently collected by the device and input into the method. It typically includes various types of information such as images, audio, and sensor data, and is dynamic and real-time. The second multimodal data is the existing historical data in the method, which may come from previous task execution processes or structured data after preliminary processing.

[0095] Displaying the first and second multimodal data for each parameter module means that the method can show the comparison between the input data and existing data under different parameter configurations. Displaying the comparison between the first and second multimodal data for each parameter module helps users intuitively understand the impact of different parameter settings on the data processing results.

[0096] In this embodiment of the application, by setting parameter modules on the visual interface and displaying the first type of multimodal data and the second type of multimodal data corresponding to each module, users can more clearly understand and control the data processing flow, thereby improving the flexibility and accuracy of personalized model training, and ultimately achieving a more efficient intelligent interactive experience of learning and using simultaneously.

[0097] In the future of smart living and work, users will increasingly use multiple devices (PCs, mobile phones, wearables, sensors, etc.) to complete tasks. These devices will generate massive amounts of multimodal, real-time data. Certain tasks (such as professional creation, health diagnosis, real-time interactive learning, and smart home optimization) not only require real-time fusion of the massive amounts of multimodal, real-time data generated by multiple devices, but also require high-precision model training and inference locally to quickly output personalized results. Typical high-computing-power applications include creative design, medical assisted diagnosis, and sports and health management. These tasks involve large amounts of data, high computational complexity, and frequent model updates, thus requiring high-performance local computing power to meet the requirements of real-time performance, privacy, and accuracy.

[0098] The main problems currently are: existing systems lead to data fragmentation, with a lack of unified aggregation and correlation analysis of user behavior data across different devices; cloud-based models have long training or update cycles, making it difficult to respond promptly to changes in user behavior; most systems passively adapt based on historical data without adjusting data weights according to users' long-term goals; this centralized processing approach poses a high privacy risk, especially when sensitive behavioral data is involved; furthermore, the inference and training processes are often fragmented, lacking an immediate feedback mechanism.

[0099] To address the aforementioned issues, this application proposes a task-driven personalized data training and interaction mechanism based on high-performance local computing devices. Through cross-terminal real-time data fusion, task target relevance perception, and timeline management mechanisms, it solves the problems of scattered data sources, lack of task-driven approaches, high response latency, and significant privacy risks in existing personalized training systems. This achieves a low-latency, privacy-secure, and explainable personalized intelligent learning experience.

[0100] Figure 3 A schematic diagram of a display interface provided for an embodiment of this application, such as... Figure 3 As shown, the interface includes: model cloud visualization 31, implementation timeline data flow 32, task relevance self-awareness and analysis 33, user-selectable parameters and personalized training mechanism 34, user task and model task orchestration, among which, Model cloud visualization presentation 31: The personalized trained model will be presented in the form of a cloud map, showing the feature words, weight changes and data distribution related to the task, so that users can intuitively understand how the model is approaching the task goal.

[0101] Real-time Timeline Data Stream 32: Continuously maintains a dynamic timeline, integrating data streams (video, audio, sensors, etc.) from different devices and arranging them in chronological order. The dynamic timeline is not only a data display tool but also a real-time trigger.

[0102] Automatic Task Relevance Sensing and Analysis 33: When new data enters the timeline, its relevance is calculated based on long-term task objectives. Once the relevance of the new data exceeds a threshold, the sensing and analysis module is triggered to perform rapid parsing (e.g., extracting tags for people and events).

[0103] User-selectable personalized training mechanism 34: Analysis results are presented to the user, who can actively or according to their needs confirm whether to add the data to the personalized training set. It supports both automatic addition and semi-automatic confirmation modes, balancing efficiency and user control (Human-in-the-Loop).

[0104] User Task 35: Display information about the user's currently executing task.

[0105] Model Task Orchestration 36: Displays the execution information of tasks during the current model execution process.

[0106] To achieve the above functions, this application provides a task processing implementation method, such as... Figure 4 As shown, this can be achieved through the following steps: Step S410: Data stream access to the timeline; Electronic devices aggregate real-time data streams from multiple terminals such as PCs, mobile phones, and wearables, and manage the data streams according to a timeline to ensure the correlation and timeliness of the data.

[0107] During implementation, user-assigned tasks can be obtained in advance, such as "editing the most timely player videos before the match." Then, when data is triggered to access the timeline, for example, when a new clip enters, it includes at least one of the following information: player name, event ID, timestamp, popularity, etc. All of the above data is then entered into the timeline manager (merged chronologically, deduplicated, and fingerprinted).

[0108] Step S420: Determine if it is related to the task; Based on the task keywords, the relevance between multimodal data and task keywords is determined.

[0109] If multimodal data with a relevance less than the relevance threshold is determined to be relevant to the task, proceed to step S430. If multimodal data with a relevance greater than or equal to the relevance threshold is determined to be relevant to the task, proceed to step S440.

[0110] Step S430: Automatically add to storage; Store multimodal data that is not relevant to the task.

[0111] Step S440: Enter perception analysis; Initiate and execute task-related automatic sensing operations and perform model analysis; Based on the user-defined long-term tasks, the system calculates relevance weights for new data and automatically triggers data analysis and the generation of interim results. It displays the user's task area and the running status of its corresponding specific model tasks.

[0112] For example, during the perception analysis process, people / events in the video can be detected, shot data can be generated, and rapid verification can be performed.

[0113] Step S450: Automatic or manual addition to training.

[0114] During implementation, users can participate in personalized training mechanisms (human-machine collaboration).

[0115] Users can choose whether to add relevant data to the training set in the analysis results. This allows for controllable and interpretable optimization of the model, and further improves the effectiveness of personalized training.

[0116] Step S440: Present the model cloud map; During the model training process, model cloud map visualization feedback can be achieved.

[0117] The results of personalized model optimization are presented in the form of cloud maps, showing feature words, weight changes and data distribution, helping users to intuitively understand the training progress and effect.

[0118] During the learning process, different keywords are displayed, and the latest keywords in the learning process are displayed in a carousel.

[0119] In some embodiments, a "Cancel" button can also be set, which can terminate the current learning session if the user triggers the button.

[0120] Upon completion of training, the changes in task keywords, feature weights, and sample contribution are presented.

[0121] Step S460, Task completed.

[0122] For example, the generated video can be automatically published after the task is confirmed to be completed.

[0123] During implementation, local high-performance computing is performed and privacy protection measures are implemented. All data processing, model training, and inference are completed on local computing devices to prevent data leakage and reduce latency, thereby ensuring user privacy and response speed.

[0124] In this embodiment, for the first time, a long-term task explicitly set by the user is used as the core driving factor for data filtering and weighting. Combined with a cross-device real-time multimodal data fusion and time axis management mechanism, a closed-loop system that can dynamically and adaptively continuously optimize the model is formed.

[0125] The cloud map is not a purely visual interface; it is deeply integrated with the training mechanism. The cloud map is generated directly from the changes in the model's feature weights and the calculation of data contribution. Users can adjust task parameters or data filtering strategies based on the cloud map results, which can in turn affect the model's training dataset and training weights. Therefore, it is part of the system interaction and training logic, rather than an independent user interface.

[0126] A task-driven triggering mechanism and timeline event awareness are introduced to ensure that model updates are not only periodic but also automatically triggered by key data events, which is superior to traditional online learning in terms of responsiveness and targeting.

[0127] Based on the foregoing embodiments, this application provides a task processing device, which includes various modules, each module including sub-modules, and each sub-module including units. It can be implemented by a processor in an electronic device; of course, it can also be implemented by specific logic circuits. In the implementation process, the processor can be a central processing unit (CPU), a microprocessor unit (MPU), a digital signal processor (DSP), or a field programmable gate array (FPGA), etc.

[0128] Figure 5 This is a schematic diagram of the composition structure of the task processing device provided in the embodiments of this application, as shown below. Figure 5 As shown, the device 500 includes: The acquisition module 510 acquires first multimodal data sent by at least one second electronic device using an intelligent agent set in the first electronic device; Management module 520 is used to manage the first multimodal data and the pre-stored second multimodal data based on preset parameters; The determination module 530 is configured to, in response to an instruction to select a target task, determine target data that matches the target task from the first multimodal data and the second multimodal data; The task execution module 540 is used to execute the target task based on the target data.

[0129] In some embodiments, the preset parameter is a time parameter. The management module 520 includes a first determining submodule and an access submodule. The first determining submodule is used to determine the corresponding timestamp based on the time of acquiring each of the first multimodal data. The access submodule is used to access the first multimodal data into the time axis management of the second multimodal data based on the timestamp.

[0130] In some embodiments, the determining module 530 includes a first acquisition submodule, a second determining submodule, a third determining submodule, a fourth determining submodule, and a matching submodule. The first acquisition submodule is used to acquire preset parameters of the target task. The second determining submodule is used to, when the preset parameters are determined to be time parameters, determine multimodal data that satisfies the preset matching parameters based on the corresponding time points on the time axis of the first multimodal data and the second multimodal data. The third determining submodule is used to, when the preset parameters are determined to be location parameters, determine the multimodal data in the first and second multimodal data that matches the location parameter as the multimodal data that satisfies the preset matching parameters. The fourth determining submodule is used to, when the preset matching parameters are determined to be person parameters, determine the multimodal data in the first and second multimodal data that matches the person parameter as the multimodal data that satisfies the preset matching parameters. The matching submodule is used to match the multimodal data that satisfies the preset parameters with the target task to determine the target data.

[0131] In some embodiments, the matching submodule includes a first determining unit, a second determining unit, and a third determining unit, wherein the first determining unit is used to determine the task keywords of the target task; the second determining unit is used to determine the relevance between the multimodal data that satisfies the preset parameters and the task keywords; and the third determining unit is used to determine the multimodal data that satisfies the preset parameters as the target data if the relevance satisfies a relevance threshold.

[0132] In some embodiments, the task execution module 540 includes a second acquisition submodule, a fifth determination submodule, and a perception analysis submodule, wherein the second acquisition submodule is used to acquire the feature weights and prompt words of the target task; the fifth determination submodule is used to determine the target model in the agent corresponding to the target task; and the perception analysis submodule is used to perform perception analysis on the target data based on the feature weights and prompt words using the target model to execute the target task.

[0133] In some embodiments, the task processing apparatus further includes a training module and a display module, wherein the training module is configured to train the target model using the target data as training data in response to a first instruction to train the target model; and the display module is configured to display, in the form of a cloud map, at least one of the following information obtained from training the target model during the training process: the task keywords of the target task, the feature weights of the target task, and the sample contribution of each piece of target data.

[0134] In some embodiments, the task processing device further includes a first adjustment module and a second adjustment module, wherein the first adjustment module is configured to adjust the task keywords of the target task in response to a second instruction to adjust task keywords; and the second adjustment module is configured to adjust the feature weights of the target task in response to a third instruction to adjust feature weights.

[0135] In some embodiments, the management module 520 includes a setting submodule and a display submodule, wherein the setting submodule is used to set at least one parameter module on a visual interface; and the display submodule is used to display the first multimodal data and the second multimodal data corresponding to each parameter module.

[0136] The descriptions of the above device embodiments are similar to those of the above method embodiments, and have similar beneficial effects. For technical details not disclosed in the device embodiments of this application, please refer to the descriptions of the method embodiments of this application for understanding.

[0137] It should be noted that, in the embodiments of this application, if the above methods are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, or the parts that contribute to related technologies, can be embodied in the form of software products. These computer software products are stored in a storage medium and include several instructions to cause electronic devices (such as mobile phones, tablets, laptops, desktop computers, etc.) to execute all or part of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), magnetic disks, or optical disks. Thus, the embodiments of this application are not limited to any specific hardware and software combination.

[0138] Correspondingly, embodiments of this application provide a storage medium storing a computer program thereon, which, when executed by a processor, implements the steps in the task processing method provided in the above embodiments.

[0139] Correspondingly, embodiments of this application provide an electronic device, Figure 6 A schematic diagram of a hardware entity of an electronic device provided in an embodiment of this application, such as... Figure 6 As shown, the hardware entity of the device 600 includes a memory 601 and a processor 602. The memory 601 stores a computer program that can run on the processor 602. When the processor 602 executes the program, it implements the steps in the task processing method provided in the above embodiments.

[0140] The memory 601 is configured to store instructions and applications executable by the processor 602, and can also cache data to be processed or already processed by the processor 602 and the various modules in the electronic device 600 (e.g., image data, audio data, voice communication data and video communication data), which can be implemented by flash memory or random access memory (RAM).

[0141] It should be noted that the descriptions of the storage medium and device embodiments above are similar to the descriptions of the method embodiments above, and have similar beneficial effects. For technical details not disclosed in the storage medium and device embodiments of this application, please refer to the descriptions of the method embodiments of this application for understanding.

[0142] It should be understood that the phrase "one embodiment" or "an embodiment" throughout the specification means that a specific feature, structure, or characteristic related to the embodiment is included in at least one embodiment of this application. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. Furthermore, these specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. It should be understood that in the various embodiments of this application, the sequence numbers of the above-described processes do not imply a sequential order of execution; the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application. The sequence numbers of the above-described embodiments are merely descriptive and do not represent the superiority or inferiority of the embodiments.

[0143] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0144] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods, such as: multiple units or components can be combined, or integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the various components shown or discussed can be through some interfaces, and the indirect coupling or communication connection between devices or units can be electrical, mechanical, or other forms.

[0145] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units. They may be located in one place or distributed across multiple network units. Some or all of the units may be selected to achieve the purpose of this embodiment according to actual needs.

[0146] In addition, each functional unit in the various embodiments of this application can be integrated into one processing unit, or each unit can be a separate unit, or two or more units can be integrated into one unit; the integrated unit can be implemented in hardware or in the form of hardware plus software functional units.

[0147] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium includes various media that can store program code, such as mobile storage devices, read-only memory (ROM), magnetic disks, or optical disks.

[0148] Alternatively, if the integrated units described above are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, or the parts that contribute to related technologies, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause an electronic device (which may be a mobile phone, tablet computer, laptop computer, desktop computer, etc.) to execute all or part of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, ROM, magnetic disks, or optical disks.

[0149] The methods disclosed in the several method embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments.

[0150] The features disclosed in the several product embodiments provided in this application can be arbitrarily combined without conflict to obtain new product embodiments.

[0151] The features disclosed in the several method or device embodiments provided in this application can be arbitrarily combined without conflict to obtain new method or device embodiments.

[0152] The above description is merely an embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A task processing method applied to a first electronic device, the method comprising: An intelligent agent configured in the first electronic device acquires first multimodal data sent by at least one second electronic device; The first multimodal data and the pre-stored second multimodal data are managed based on preset parameters; In response to an instruction to select a target task, target data matching the target task is determined from the first multimodal data and the second multimodal data; The target task is executed based on the target data.

2. The method as described in claim 1, wherein the preset parameter is a time parameter, and the step of managing the first multimodal data and the pre-stored second multimodal data based on the preset parameter includes: The corresponding timestamp is determined based on the time when each of the first multimodal data is acquired; Based on the timestamp, the first multimodal data is integrated into the timeline management of the second multimodal data.

3. The method of claim 2, wherein determining the target data matching the target task from the first multimodal data and the second multimodal data comprises: Obtain the preset parameters of the target task; When the preset parameter is determined to be a time parameter, multimodal data that satisfies the preset matching parameter is determined based on the time points corresponding to the first multimodal data and the second multimodal data on the time axis. If the preset parameter is determined to be a location parameter, the multimodal data that matches the location parameter in the first multimodal data and the second multimodal data is determined as the multimodal data that satisfies the preset matching parameter; If the preset matching parameter is determined to be a character parameter, the multimodal data that matches the character parameter in the first multimodal data and the second multimodal data is determined as the multimodal data that satisfies the preset matching parameter; The multimodal data that meets the preset parameters is matched with the target task to determine the target data.

4. The method of claim 3, wherein matching the multimodal data satisfying the preset parameters with the target task to determine the target data includes: Determine the task keywords for the target task; Determine the relevance between the multimodal data that satisfies the preset parameters and the task keywords; If the correlation is determined to meet the correlation threshold, the multimodal data that meets the preset parameters is determined as the target data.

5. The method of claim 1, wherein performing the target task based on the target data comprises: Obtain the feature weights and prompt words for the target task; Determine the target model in the intelligent agent that corresponds to the target task; The target model is used to perform perceptual analysis on the target data based on the feature weights and the prompt words in order to execute the target task.

6. The method of claim 5, further comprising: In response to a first instruction to train the target model, the target data is used as training data to train the target model; During the training of the target model, at least one of the following information obtained from training the target model is displayed in a visualization interface in the form of a cloud map: the task keywords of the target task, the feature weights of the target task, and the sample contribution of each of the target data.

7. The method of claim 6, further comprising: In response to the second instruction to adjust the task keywords, the task keywords of the target task are adjusted; In response to a third instruction to adjust feature weights, the feature weights of the target task are adjusted.

8. The method according to any one of claims 1 to 7, wherein managing the first multimodal data and the pre-stored second multimodal data based on preset parameters comprises: Set at least one parameter module on the visual interface; Display the first multimodal data and the second multimodal data corresponding to each of the parameter modules.

9. A task processing apparatus, the apparatus comprising: The acquisition module uses an intelligent agent set in the first electronic device to acquire first multimodal data sent by at least one second electronic device; The management module is used to manage the first multimodal data and the pre-stored second multimodal data based on preset parameters; A determination module is configured to, in response to an instruction to select a target task, determine target data that matches the target task from the first multimodal data and the second multimodal data; The task execution module is used to execute the target task based on the target data.

10. A first electronic device, comprising a memory and a processor, the memory storing a computer program executable on the processor, the processor executing the program to acquire first multimodal data sent by at least one second electronic device using an agent disposed in the first electronic device; The first multimodal data and the pre-stored second multimodal data are managed based on preset parameters; In response to an instruction to select a target task, target data matching the target task is determined from the first multimodal data and the second multimodal data; The target task is executed based on the target data.