Version self-updating method, computing device, storage medium and computer program product
Patent Information
- Application Number
- CN202610747680.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-28
- Publication Date
- 2026-09-11
- Estimated Expiration
- 2046-05-28
AI Technical Summary
[0003]随着智能处理单元在代码辅助、知识问答、内容创作和数据分析等场景的广泛应用,用户对智能处理单元的期望已从通用能力演进为个性化处理能力,然而,不同用户的输入习惯和服务偏好均不同,为了能够使智能处理单元更加通用的服务大量用户,通常采用统一的系统提示指令和固定的交互策略,忽略了不同用户的个性化需求,影响用户体验
[0011]上述方法中,可以获取目标对象在第一时间区间内与智能处理单元在多个交互维度的交互行为数据,并根据多个交互维度的交互行为数据,确定目标对象的目标行为画像信息,并根据目标行为画像信息,抽象出目标对象在和智能处理单元交互过程中的行为规则信息,在根据行为规则信息确定触发预设更新条件的情况下,对智能处理单元的当前版本进行更新,获得智能处理单元的目标更新版本,使得智能处理单元的目标更新版本能够学习目标对象的交互行为习惯和服务偏好,使得智能处理单元能够满足用户的个性化需求,进一步提升用户体验。
Smart Images

Figure CN122285048B_ABST
Abstract
Description
Technical Field
[0001] The embodiments in this specification relate to the field of artificial intelligence technology, and in particular to self-updating methods, computing devices, storage media, and computer program products. Background Technology
[0002] Currently, in order to handle a large number of model inference tasks, multiple artificial intelligence models can usually be deployed on the server to process model inference tasks. At the same time, multiple intelligent processing units can be run. These intelligent processing units can be used to autonomously access the artificial intelligence models to execute model inference tasks after receiving model inference tasks.
[0003] With the widespread application of intelligent processing units in scenarios such as code assistance, Q&A, content creation, and data analysis, users' expectations for these units have evolved from general-purpose capabilities to personalized processing capabilities. However, different users have different input habits and service preferences. In order to enable intelligent processing units to serve a large number of users more universally, uniform system prompts and fixed interaction strategies are usually adopted, ignoring the personalized needs of different users and affecting the user experience. Therefore, an effective technical solution is urgently needed to solve the above problems. Summary of the Invention
[0004] In view of this, embodiments of this specification provide a version self-updating method. One or more embodiments of this specification also relate to a version self-updating apparatus, a computing device, a computer-readable storage medium, and a computer program product, to address the technical deficiencies existing in the prior art.
[0005] According to a first aspect of the embodiments of this specification, a version self-update method is provided, applied to an intelligent processing unit, comprising: The system acquires interaction behavior data of the target object with the intelligent processing unit in multiple interaction dimensions within a first time interval, and determines the target behavior profile information of the target object based on the interaction behavior data. Based on the target behavior profile information, determine the behavior rule information corresponding to the target object; When the preset update conditions are determined based on the behavior rule information, the current version of the intelligent processing unit is updated to obtain the target update version of the intelligent processing unit. The target update version is a version obtained by performing a self-evolution update at the code logic level on the current version and passing the update test.
[0006] According to a second aspect of the embodiments of this specification, a version self-updating device is provided, applied to an intelligent processing unit, comprising: The acquisition module is configured to acquire interaction behavior data of the target object with the intelligent processing unit in multiple interaction dimensions within a first time interval, and determine the target behavior profile information of the target object based on the interaction behavior data. The determination module is configured to determine the behavior rule information corresponding to the target object based on the target behavior profile information; The update module is configured to update the current version of the intelligent processing unit when a preset update condition is determined based on the behavior rule information, so as to obtain a target update version of the intelligent processing unit, wherein the target update version is a version obtained by performing a self-evolution update at the code logic level on the current version and passing the update test.
[0007] According to a third aspect of the embodiments of this specification, a computing device is provided, comprising: Memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions, which, when executed by the processor, implement the steps of the above method.
[0008] According to a fourth aspect of the embodiments of this specification, a computer-readable storage medium is provided that stores a computer program / instructions that, when executed by a processor, implement the steps of the above-described method.
[0009] According to a fifth aspect of the embodiments of this specification, a computer program product is provided, including a computer program / instructions that, when executed by a processor, implement the steps of the above-described method.
[0010] This specification provides an embodiment of a version self-update method applied to an intelligent processing unit, comprising: acquiring interaction behavior data of a target object with the intelligent processing unit in multiple interaction dimensions within a first time interval, and determining target behavior profile information of the target object based on the interaction behavior data; determining behavior rule information corresponding to the target object based on the target behavior profile information; and updating the current version of the intelligent processing unit when a preset update condition is determined based on the behavior rule information to obtain a target updated version of the intelligent processing unit, wherein the target updated version is a version obtained by performing a self-evolution update at the code logic level on the current version and passing an update test.
[0011] The above method can acquire interaction behavior data of the target object with the intelligent processing unit in multiple interaction dimensions within a first time interval. Based on the interaction behavior data in multiple interaction dimensions, the target behavior profile information of the target object is determined. Based on the target behavior profile information, the behavior rule information of the target object in the interaction process with the intelligent processing unit is abstracted. When the preset update conditions are determined based on the behavior rule information, the current version of the intelligent processing unit is updated to obtain the target update version of the intelligent processing unit. This allows the target update version of the intelligent processing unit to learn the interaction behavior habits and service preferences of the target object, enabling the intelligent processing unit to meet the personalized needs of users and further improve the user experience. Attached Figure Description
[0012] Figure 1 This is a flowchart illustrating a version self-update method provided in one embodiment of this specification; Figure 2 This is a schematic diagram of the structure of the intelligent processing unit in a version self-update method provided in one embodiment of this specification; Figure 3 This is a schematic diagram of a three-level evolutionary scheduling in a version self-update method provided in one embodiment of this specification; Figure 4 This is a schematic diagram illustrating the construction and updating of behavioral profile information in a version self-updating method provided in one embodiment of this specification; Figure 5 This is a schematic diagram illustrating the construction of a behavior rule database in a version self-update method provided in one embodiment of this specification; Figure 6 This is a schematic diagram of negative feedback attribution analysis in a version self-update method provided in one embodiment of this specification; Figure 7 This is a schematic diagram illustrating a simulation test in a version self-update method provided in one embodiment of this specification; Figure 8 This is a schematic diagram of shadow parallel execution verification in a version self-update method provided in one embodiment of this specification; Figure 9 This is a schematic diagram of the structure of a version self-updating device provided in one embodiment of this specification; Figure 10 This is an architecture diagram of a version self-updating system provided in one embodiment of this specification; Figure 11 This is a structural block diagram of a computing device provided in one embodiment of this specification. Detailed Implementation
[0013] Many specific details are set forth in the following description to provide a full understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar extensions without departing from the spirit of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.
[0014] The terminology used in one or more embodiments of this specification is for the purpose of describing particular embodiments only and is not intended to be limiting of the one or more embodiments of this specification. The singular forms “a,” “described,” and “the” as used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.
[0015] It should be understood that although the terms first, second, etc., may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first may also be referred to as second without departing from the scope of one or more embodiments of this specification, and similarly, second may also be referred to as first. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to a determination."
[0016] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this manual are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data comply with the relevant laws, regulations and standards of the relevant regions, and corresponding operation portals are provided for users to choose to authorize or refuse.
[0017] In one or more embodiments of this specification, a large model refers to a deep learning model with a large number of model parameters, typically containing hundreds of millions, tens of billions, hundreds of billions, trillions, or even tens of trillions of model parameters. A large model can also be called a foundation model. It is pre-trained using large-scale unlabeled corpora to produce a pre-trained model with hundreds of millions of parameters. Such models can adapt to a wide range of downstream tasks and have good generalization ability. Examples include Large Language Models (LLMs) and multi-modal pre-training models.
[0018] In practical applications, large models only require a small number of samples to fine-tune the pre-trained model before they can be applied to different tasks. Large models can be widely used in fields such as Natural Language Processing (NLP) and Computer Vision. Specifically, they can be applied to computer vision tasks such as Visual Question Answering (VQA), Image Captioning (IC), and Image Generation, as well as NLP tasks such as text-based sentiment classification, text summarization, and machine translation. The main application scenarios for large models include digital assistants, intelligent robots, search, online education, office software, e-commerce, and intelligent design.
[0019] First, the terms and concepts used in one or more embodiments of this specification will be explained.
[0020] Intelligent Processing Unit: An automated processing system with a Large Language Model (LLM) as the core reasoning engine, combined with memory mechanisms, planning capabilities, and tool usage capabilities.
[0021] Meta-self-evolutionary control layer: A higher-order control layer located above the core capabilities of the agent, responsible for observing, analyzing, and driving the agent's autonomous evolution in multiple orthogonal dimensions.
[0022] Agent: An autonomous task-executing intelligent agent, a software entity based on a large language model that possesses autonomous decision-making and execution capabilities.
[0023] Behavioral profile information: a structured vector representation of multi-dimensional features such as user interaction behavior, preferences, and knowledge domain, serving as the core input for evolutionary decision-making.
[0024] Evolutionary dimensions: orthogonal evolutionary dimension, interaction dimension, independent ability dimension of agent self-evolution, each dimension can evolve independently but there are cross-dimensional relationships.
[0025] With the widespread application of autonomous task-execution agents driven by large language models in scenarios such as code assistance, knowledge-based question answering, content creation, and data analysis, users' expectations for these agents have evolved from "general capabilities" to "personalized assistants that understand me." However, current agents still have the following problems in terms of personalized adaptation.
[0026] First, current AI agents use uniform system prompts and fixed interaction strategies for all users. However, different users vary significantly in terms of response detail, coding style, tool preferences, and communication styles, making it difficult for the AI agent to perceive and adapt to these differences. For example, a programming assistance AI agent provides programming assistance services to multiple users, where 40% of users prefer detailed code comments, while 60% prefer concise code generation. With a uniform configuration for the AI agent, regardless of whether a strategy of generating detailed code comments or concise code is adopted, most users will need to frequently manually correct the AI agent's output style, wasting a significant amount of time.
[0027] Secondly, when the agent makes an error and it is corrected by the user, the correct answer is only remembered in the current session. In the next session or when encountering similar but not identical scenarios with other users, the agent may still make the same type of error. The negative feedback from the user in a single session only exists in the agent's short-term contextual memory and cannot be translated into a lasting performance improvement.
[0028] Furthermore, the current skill set of the intelligent agent (including callable tools, types of operations that can be performed, etc.) is determined when the intelligent agent is deployed. During operation, it will provide services to users based on the deployed skill set. However, as the user's work scenario continues to evolve, the capability boundary of the intelligent agent remains static and cannot expand other skills to meet the user's needs.
[0029] Current AI agents attempt to achieve personalization through global prompts, typically compressing all preference information into a single, flat prompt text. Adjusting one preference can affect the normal execution of other aspects. For example, changing the preference to "more detailed comments in the generated code" via the prompt text might cause the agent's output to become verbose.
[0030] In order to optimize the agent, the current agent may over-correct its own logic based on limited negative feedback information, or accidentally break the correct behavior of another module while correcting one module, causing the agent to malfunction or even become completely unusable.
[0031] Therefore, there is an urgent need for an effective technical solution to address the above problems.
[0032] This specification provides a version self-updating method, and also relates to a version self-updating device, a computing device, a computer-readable storage medium, and a computer program product, which will be described in detail in the following embodiments.
[0033] See Figure 1 , Figure 1A flowchart of a version self-update method according to an embodiment of this specification is shown, applied to an intelligent processing unit, and specifically includes the following steps.
[0034] Step 102: Obtain the interaction behavior data of the target object with the intelligent processing unit in multiple interaction dimensions within the first time interval, and determine the target behavior profile information of the target object based on the interaction behavior data.
[0035] Specifically, the version self-update method provided in the embodiments of this specification can be applied to the update of intelligent processing units.
[0036] The target object can be understood as the object using the intelligent processing unit, such as a user. The first time interval can be understood as a preset period of time after the target object first uses the intelligent processing unit, such as 30 days after the user's first use, or the period of 100 uses. This first time interval can be used to construct the target object's target behavior profile information. The target behavior profile information describes the target object's behavioral preferences and habits when using the intelligent processing unit. The process of the target object using the intelligent processing unit can be understood as the process of the target object and the intelligent processing unit interacting through text or voice to enable the intelligent processing unit to complete tasks. For example, if the intelligent processing unit is used to provide programming assistance services, the target object can input text commands into the intelligent processing unit, and the intelligent processing unit can generate code for the target object based on the text commands. Multiple interaction dimensions can be understood as the interaction dimensions between the target object and the intelligent processing unit. Interaction behavior data can be understood as the interaction behavior data between the target object and the intelligent processing unit, such as the target object's question-and-answer behavior, retrieval behavior, tool call behavior, error correction behavior, etc.
[0037] In practical applications, target behavior profile information can be stored in a storage structure based on an interaction dimension tree and sub-feature information, a storage structure based on a unified high-dimensional vector space representation, or a storage structure based on a knowledge graph association network representation.
[0038] In addition, the intelligent processing unit can also be used to provide other services, such as text generation services, image processing services, etc., but this specification does not limit the embodiments thereto.
[0039] For ease of understanding, this specification uses the user as the target object and the intelligent processing unit as an example to illustrate the embodiments. In addition, the target object can also be other objects that can call the intelligent processing unit, and this specification does not limit this.
[0040] Based on this, within a preset time period after the user first uses the intelligent processing unit, interaction behavior data of the user and the intelligent processing unit in multiple interaction dimensions can be obtained, and target behavior profile information of the target object can be determined based on the interaction behavior data.
[0041] It is understandable that the interaction behavior data acquired within the first time interval can be multiple interaction behavior data of multiple interactions between the target object and the intelligent processing unit, and multiple interaction behavior data can also be acquired for each interaction process between the target object and the intelligent processing unit.
[0042] In practical applications, see Figure 2 , Figure 2 A schematic diagram of the intelligent processing unit is shown in one embodiment of a version self-update method provided in this specification. Figure 2 As shown, the intelligent processing unit can be an agent. Above the agent's basic capability layer, a meta-self-evolutionary control layer can be built. This basic capability layer can be used for task execution, while the meta-self-evolutionary control layer can execute the aforementioned version self-update method to update the intelligent processing unit. The meta-self-evolutionary control layer can continuously observe the interaction behavior between the user and the agent, maintain structured target behavior profile information, and drive the agent to autonomously evolve across multiple interaction dimensions through a three-level evolutionary triggering mechanism. These multiple interaction dimensions can be understood as multiple orthogonal evolutionary dimensions used to update the agent.
[0043] Further, the step of acquiring interaction behavior data between the target object and the intelligent processing unit in multiple interaction dimensions within a first time interval, and determining the target behavior profile information of the target object based on the interaction behavior data, includes: In response to the target object's request to invoke the intelligent processing unit, the intelligent processing unit processes the request and constructs initial behavioral profile information of the target object. The system acquires interactive behavior data across multiple dimensions of the target object during the process of calling the intelligent processing unit within the first time interval, and updates the initial behavior profile information based on the interactive behavior data to obtain the target behavior profile information of the target object, so that the intelligent processing unit can process the scheduling request of the target object based on the target behavior profile information.
[0044] The initial behavioral profile information can be understood as the initial behavioral profile information constructed in response to the first call request of the target object to the intelligent processing unit. The initial behavioral profile information may include initial feature information of multiple interaction dimensions. The target behavioral profile information can be understood as the behavioral profile information obtained by continuously updating the initial behavioral profile information based on the interaction behavior data of the target object and the intelligent processing unit within the first time interval. Accordingly, the target behavioral profile information may include updated feature information of multiple interaction dimensions.
[0045] Understandably, the target behavior profile information is also dynamically and continuously updated during the process of the target object using the intelligent processing unit. That is to say, after the first time interval, the target behavior profile information can be updated based on the interactive behavior data of the target object each time it uses the intelligent processing unit, so that the behavioral habit information and behavioral preference information described by the target behavior profile information can be more accurate. This allows the intelligent processing unit to process the target object's subsequent scheduling requests in accordance with user habits and user preferences based on the target behavior profile information.
[0046] Based on this, in response to the target object's initial call request to the intelligent processing unit, the interaction behavior data of the first interaction between the target object and the intelligent processing unit can be determined based on the processing procedure of the intelligent processing unit in handling the initial call request. Based on this interaction behavior data, the initial behavioral profile information of the target object is initialized and constructed. In subsequent first time intervals, in response to each subsequent call request from the target object to the intelligent processing unit, the interaction behavior data of each interaction between the target object and the intelligent processing unit is determined based on the processing procedure of the intelligent processing unit in handling each call request. This interaction behavior data may include interaction behavior data from multiple interaction dimensions. Furthermore, the initial behavioral profile information of the target object can be updated based on the interaction behavior data of each interaction. Further, the initialization feature information of each interaction dimension in the initial behavioral profile information of the target object can be updated to obtain the target behavioral profile information of the target object.
[0047] In summary, by initializing and continuously updating the behavioral profile information of the target object, the target behavioral profile information can quickly respond to changes in the target object's behavioral habits and preferences, thereby learning the target object's behavioral habits and preferences. This allows the intelligent processing unit to process subsequent call requests of the target object based on the target behavioral profile information in a way that better matches the target object's preferences and habits, achieving personalized adaptation for the target object.
[0048] Furthermore, the step of acquiring interactive behavior data across multiple interaction dimensions of the target object during the process of calling the intelligent processing unit within the first time interval includes: According to the preset update rules, the interactive behavior data of the target object in calling the intelligent processing unit within the first time interval is obtained in multiple interactive dimensions.
[0049] Among them, the preset update rules can be understood as the rules for obtaining and updating interactive behavior data according to the update trigger conditions. The preset update rules can include instant updates, periodic updates, and event-driven targeted updates.
[0050] Specifically, the interaction behavior data between the target object and the intelligent processing unit can be obtained according to the preset update rules, so as to update the initial behavior profile information based on the interaction behavior data.
[0051] In specific implementation, the step of acquiring interactive behavior data across multiple interactive dimensions of the target object during the process of calling the intelligent processing unit within the first time interval, according to a preset update rule, includes: If the target object completes the invocation request by calling the intelligent processing unit within the first time interval, the interactive behavior data corresponding to the invocation request is obtained; and / or According to a preset time interval, acquire interaction data between the target object and the intelligent processing unit during the process of the target object calling the intelligent processing unit within the preset time interval, wherein the preset time interval is the time interval for triggering the update, and the preset time interval is shorter than the first time interval; and / or The system acquires interactive behavior data of the target object during the process of calling the intelligent processing unit within the first time interval, which meets preset event conditions.
[0052] The preset time interval can be understood as a pre-defined update time interval, which can be used to trigger periodic updates. The preset event condition can be understood as the condition that triggers a certain event, which can be used for event-driven targeted updates.
[0053] Specifically, when the preset update rule is real-time update, interaction behavior data corresponding to each call request can be collected based on the processing procedure of the intelligent processing unit after each call request is completed by the target object. When the preset update rule is periodic update, interaction behavior data of each interaction between the target object and the intelligent processing unit within a preset time interval can be obtained. When the preset update rule is event-driven targeted update, interaction behavior data corresponding to a specific event is collected in each interaction between the target object and the intelligent processing unit, provided that a specific event is identified.
[0054] In practical applications, in combination with the above Figure 2 The meta-evolutionary control layer can include a three-level evolution scheduler, which can be used to uniformly manage evolutionary tasks with three triggering modes, namely, the aforementioned instant update, periodic update and event-driven directional update, to ensure the orderliness and resource efficiency of update operations.
[0055] Real-time updates can be triggered after each interaction between the target object and the intelligent processing unit, specifically after the target object completes each call to the intelligent processing unit to process a request. After acquiring the interaction behavior data for each interaction, the sub-feature vectors of the corresponding interaction dimension can be slightly updated. For example, the feature weights of the interaction dimension corresponding to this interaction in the initial behavior profile information can be updated. Furthermore, if an explicit preference signal is detected during this interaction (such as the target object's corrective behavior), the configuration of the interaction dimension corresponding to that explicit preference signal can be fine-tuned. Then, in the next interaction, processing can be directly based on the fine-tuned configuration. Since this real-time update only involves incremental updates of feature vectors and a small number of conditional judgments, resource consumption is extremely low, and the configuration change range of a single interaction's real-time micro-evolution is limited, thus preventing drastic behavior changes caused by a single interaction.
[0056] Periodic updates can be triggered according to a preset time interval (i.e., a preset period), such as every day or every week. This can be triggered by a timer, for example, by setting it to trigger a periodic update every Monday. After an update is triggered according to the preset time interval, a global comprehensive analysis can be performed on all interaction behavior data of the target object and the intelligent processing unit within that preset time interval. This analysis can cover all interaction dimensions corresponding to all interaction behavior data. Specifically, in-depth statistical analysis can be performed on all accumulated interaction behavior data within the preset time interval to identify long-term trends across interactions. The dimensional correlations between various interaction dimensions corresponding to all interaction behavior data can be detected; these correlations may include collaborative and conflict modes. The confidence level of all behavior rule information can be reassessed, and behavior rule information with a confidence level below a preset confidence threshold or that is invalid can be deleted. Skill usage frequency can be analyzed to identify capability gaps in the intelligent processing unit (i.e., areas where users have needs but the intelligent processing unit currently does not adequately cover), triggering a skill self-discovery and learning process. Furthermore, domain knowledge dimensions can be organized and the knowledge graph updated. Although periodic updates consume more resources, they are less frequent, and changes from periodic deep updates can only take full effect after subsequent gray-scale verification.
[0057] Event-driven targeted updates can be triggered when a specific event is detected in the interaction between the target object and the intelligent processing unit. This specific event can include: negative feedback use cases accumulating to exceed a preset feedback threshold in a certain interaction dimension (e.g., receiving negative feedback from the target object 3 times in the same interaction dimension); a task failing consecutively N times; a user explicitly requesting a change in behavior (e.g., a user instructing the user to stop using this style); or a significant change in the user's work context (e.g., a user's call request switching from a front-end project to a back-end project). When a specific event is detected, the corresponding interaction behavior data can be obtained and precisely targeted to the interaction dimension attributable to triggering the event. Significant configuration adjustments can be made to that interaction dimension, and a large language model can be used to analyze the negative feedback attribution records related to the specific event, generating a comprehensive dimension evolution plan, which can then be validated through a phased rollout. Because event-driven targeted updates only perform in-depth analysis on the specific interaction dimension corresponding to a specific event, resource consumption is moderate, and event-driven targeted updates require phased rollout before execution, while also retaining rollback points.
[0058] See Figure 3 , Figure 3 A schematic diagram of a three-level evolutionary scheduling in a version self-update method according to an embodiment of this specification is shown. Figure 3 As shown, event-driven targeted update (i.e., event-driven targeted evolution), as the third level of evolution, can perform in-depth analysis on a specific interaction dimension corresponding to a specific event when a specific event is detected during the interaction between the target object and the intelligent processing unit. It then generates an evolutionary plan for that specific interaction dimension, performs gray-scale verification on the evolutionary plan, and updates the specific interaction dimension based on the successful gray-scale verification, while abstracting the corresponding behavioral rule information. Detected specific events may include, for example, negative feedback exceeding a threshold, continuous task processing failures, or explicit user requests. Subsequently, real-time update, as the first level of evolution, incrementally updates the weight information of the interaction dimension corresponding to each interaction between the target object and the intelligent processing unit, with limitations on the update magnitude. Periodic updates, as the second level of evolution, can accumulate interaction behavior data of each interaction between the target object and the intelligent processing unit in the first level of evolution. They are triggered at preset time intervals, and statistical analysis is performed on the interaction dimensions corresponding to the accumulated interaction behavior data. Correlation analysis is also performed on the dimensional relationships between interaction dimensions. In addition, capability gap identification, skill self-discovery and self-learning, and re-evaluation of behavioral rule information are carried out. If anomalies are detected, the system can switch to the third level of evolution for event-driven targeted updates.
[0059] In summary, the three-level scheduling architecture of instant evolution, periodic deep evolution, and event-driven targeted evolution takes into account the timeliness (real-time response to changes in user preferences), depth (periodic global re-analysis), and targeting (precise response to negative events) of evolution.
[0060] Further, updating the initial behavioral profile information based on the interaction behavior data to obtain the target behavioral profile information of the target object includes: Interaction behavior signals are collected from the interaction behavior data, wherein the interaction behavior signals include a first type of behavior signal and a second type of behavior signal, the first type of behavior signal is used to characterize the explicit preference information of the target object, and the second type of behavior signal is used to characterize the behavior pattern information of the target object; The interaction behavior signal is mapped to the corresponding target interaction dimension, and the interaction behavior signal is encoded to obtain the updated feature information corresponding to the target interaction dimension, wherein the target interaction dimension is any one of the plurality of interaction dimensions; The initial behavioral profile information is updated based on the updated feature information corresponding to the target interaction dimension to obtain the target behavioral profile information of the target object.
[0061] Interactive behavior signals can be understood as interactive behaviors or behavior patterns collected in interactive behavior data. The first type of behavior signal can be understood as explicit behavior signals, which can be understood as behavior signals of the target object actively operating. The first type of behavior signal can be, for example, the user correcting the output behavior of the intelligent processing unit (e.g., do not use this way of writing, use the functional way of writing), the user evaluation feedback behavior (e.g., like or dislike), the configuration change behavior, and the explicit instruction behavior (e.g., sending a message instruction to the intelligent processing unit "keep the reply more concise in the future").
[0062] The second type of behavioral signal can be understood as implicit behavioral signal, which can be understood as a signal inferring the behavioral pattern of the target object. The second type of behavioral signal may include the user's editing distance to the output of the intelligent processing unit (significantly modifying the output or directly adopting it), the task abandonment rate, the response acceptance / rejection ratio, the frequency of repeated requests for the same type of task, the spectrum of tool / function usage, etc.
[0063] The target interaction dimension can be understood as the interaction dimension corresponding to the interaction behavior signal among multiple interaction dimensions.
[0064] Based on this, multiple interaction behavior signals can be collected from the interaction behavior data. For each interaction behavior signal, it can be mapped to the corresponding target interaction dimension, and the interaction behavior signal can be encoded to obtain the updated feature information corresponding to the target interaction dimension. Then, based on the updated feature information corresponding to the target interaction dimension, the sub-feature information corresponding to the target interaction dimension in the initial behavior profile information is updated to obtain the target behavior profile information. Here, the updated feature information can be understood as the updated feature increment, and the sub-feature information can be understood as the sub-feature vector corresponding to the target interaction dimension.
[0065] In practical applications, in combination with the above Figure 2 The meta-evolutionary control layer may include an interactive behavior sensor, which can be used to collect explicit and implicit behavioral signals during the interaction between the target object and the intelligent processing unit in real time, and encode the explicit and implicit behavioral signals into a structured behavioral event stream.
[0066] The interactive behavior sensor can collect explicit and implicit behavioral signals from the interactive behavior data of each interaction between the target object and the intelligent processing unit. It can collect explicit behavioral signals based on user-initiated actions and implicit behavioral signals based on behavioral pattern inference. Each collected interactive behavior signal is mapped to a corresponding target interaction dimension and encoded as an updated feature vector of the sub-feature vector of that target interaction dimension.
[0067] In one embodiment of this specification, at least two interaction dimensions may include an interaction performance dimension, a tool invocation dimension, a code generation dimension, a context management dimension, and a domain knowledge dimension. Each interaction dimension may include multiple sub-feature information. The sub-feature information corresponding to the interaction performance dimension may include the following indicators: response detail preference, response format preference (code block, table, or text), language style (technical language / plain natural language), and interface interaction mode preference. For example, if the interaction behavior signal is "the user compresses a long response three times consecutively," the interaction behavior signal is mapped to the interaction performance dimension, and the encoded update feature information is "reduce the detail weight."
[0068] The sub-feature information corresponding to the tool invocation dimension can include the following indicators: the frequency distribution of each tool, the tool combination pattern, and the mapping preference between invocation tasks and tools. For example, if the interaction behavior signal is "the user always uses the search tool first when doing coding tasks", this interaction behavior signal can be mapped to the tool invocation dimension, and the encoded update feature information can be "adjust tool invocation priority".
[0069] The sub-feature information corresponding to the code generation dimension can include the following metrics: programming language preference, naming style, architectural pattern preference, code comment density preference, etc. For example, if the interaction behavior signal is "the user changed the variable name from camelCase to serpentine", this interaction behavior signal can be mapped to the code generation dimension, and the encoded update feature information is "update naming style feature".
[0070] The context management dimension can be understood as the memory management dimension. The sub-feature information corresponding to this context management dimension can include the following indicators: memory item retrieval hit rate, memory expiration speed preference, and context length sensitivity. For example, if the interaction behavior signal is "the user frequently mentions details of a conversation from 3 days ago", this interaction behavior signal can be mapped to this context management dimension, and the encoded update feature information can be "increase the weight of long-term memory retention".
[0071] The sub-feature information corresponding to the domain knowledge dimension can include the following indicators: active domain distribution, domain depth level, cross-domain association pattern, etc. For example, if the interaction behavior signal is "80% of the user's tasks in the last 2 weeks involve microservice architecture", this interaction behavior signal can be mapped to the domain knowledge dimension, and the encoded update feature information is "strengthen the domain knowledge weight of microservice".
[0072] Based on the updated feature information of the above encoding, the sub-feature information corresponding to the target interaction dimension is updated to realize the update of the initial behavior profile information and obtain the target behavior profile information.
[0073] In addition, at least two interaction dimensions may also include security compliance dimensions, collaboration preference dimensions, etc. The interaction dimensions can be determined according to actual needs, and the embodiments in this specification do not limit this.
[0074] In summary, by decomposing the behavioral characteristics of the target object into the aforementioned multiple interaction dimensions, and independently modeling each interaction dimension as sub-feature information and updating it independently, cross-interference between multiple interaction dimensions is avoided.
[0075] Further, updating the initial behavioral profile information based on the updated feature information corresponding to the target interaction dimension to obtain the target behavioral profile information of the target object includes: Determine the confidence weight information of the updated feature information corresponding to the target interaction dimension; Based on the time decay function, the updated feature information, and the confidence weight information, the sub-behavioral profile information corresponding to the target interaction dimension in the initial behavioral profile information is updated to obtain the target behavioral profile information of the target object.
[0076] Among them, the confidence weight information of the updated feature information corresponding to the target interaction dimension can be understood as the confidence weight information attached to the sub-feature information of the target interaction dimension. The confidence weight information can be used to represent the importance of the sub-feature information, and the sub-behavior profile information can be understood as the sub-feature information.
[0077] Based on this, the confidence weight information attached to the sub-feature information of the target interaction dimension can be determined. Based on the time decay function, the updated feature information and the confidence weight information, the sub-feature information corresponding to the target interaction dimension in the initial behavior profile information is updated to obtain the target behavior profile information of the target object.
[0078] In practical applications, in combination with the above Figure 2 The meta-evolutionary control layer also includes a behavior profile information manager, which is used to maintain multi-dimensional structured profile information for each user. The behavior profile information of the target object is organized into a tree structure of orthogonal evolutionary dimensions. An orthogonal evolutionary dimension can be understood as an interaction dimension. Each interaction dimension contains sub-feature information, confidence weight information, and time decay coefficient.
[0079] In multiple interaction dimensions, each sub-feature of each interaction dimension is accompanied by a confidence weight w, with the value of w ranging from 0 to 1. In the initial behavioral profile information, the confidence weight of each sub-feature of each interaction dimension is initialized to 0.3, gradually increasing as consistency signals accumulate. Furthermore, by introducing a time decay function d(t) = e^(-λt), the influence of long-term interaction behavior signals naturally decays over time, ensuring that the target behavioral profile information reflects the user's recent state. For explicit behavior signals, a higher single-instance weight increment can be assigned; for implicit behavior signals, a lower but cumulatively significant weight increment is assigned. The weight increment is the weight added to the confidence weight information. When the confidence weight of a certain sub-feature is lower than a preset weight threshold (e.g., lower than 0.2), the sub-feature can be marked as an uncertain feature. This marked sub-feature does not participate in evolutionary decision-making, i.e., it is not updated.
[0080] See Figure 4 , Figure 4 This diagram illustrates the construction and updating of behavioral profile information in a version self-updating method according to an embodiment of this specification. Figure 4As shown, for each interaction event between the target object and the intelligent processing unit (i.e., the interaction event between the user and the intelligent agent), the interaction behavior perceiver collects interaction behavior signals and determines the signal type of the interaction behavior signals. If the interaction behavior signal is determined to be an explicit behavior signal, explicit preference instructions are extracted. If the interaction behavior signal is determined to be an implicit behavior signal, behavior pattern features are extracted. The explicit preference instructions and behavior pattern features are mapped from the interaction behavior signals to the interaction dimensions. The sub-feature information (i.e., sub-feature vectors) of the interaction dimensions corresponding to each interaction behavior signal in multiple interaction dimensions can be updated. For example, the sub-feature information of the interaction performance dimension, the tool call dimension, the code generation dimension, the context management dimension, and / or the domain knowledge dimension can be updated. Combined with the weighting of confidence weight information and time decay, the persistent update of the behavior profile information (initial behavior profile information or target behavior profile information) is realized.
[0081] In addition, when updating the confidence weight information, it can also be updated based on sliding window statistics or Bayesian posterior probability, but the embodiments in this specification do not limit this.
[0082] In summary, by introducing confidence weight information and time decay function of sub-feature information, behavioral profile information can quickly respond to recent changes in the target object's behavior, while avoiding noise interference from the target object's occasional behavior.
[0083] In practical applications, the multiple interaction dimensions include interaction performance dimension, tool invocation dimension, code generation dimension, context management dimension, and domain knowledge dimension; The step of updating the initial behavior profile information based on the interaction behavior data to obtain the target behavior profile information of the target object includes: Based on the interaction behavior data of the target object and the intelligent processing unit in the interaction performance dimension, the response format information, response detail information, interface layout information, and / or recommendation strategy information in the initial behavior profile information are updated to obtain the target behavior profile information of the target object; and / or Based on the interaction behavior data of the target object and the intelligent processing unit in the tool invocation dimension, the tool invocation strategy information, tool parameter information, and / or tool information to be updated in the initial behavior profile information are updated to obtain the target behavior profile information of the target object; and / or Based on the interaction behavior data between the target object and the intelligent processing unit in the code generation dimension, the coding style information, architecture pattern information, technology stack adaptation information, and / or code quality information in the initial behavior profile information are updated to obtain the target behavior profile information of the target object; and / or Based on the interaction behavior data between the target object and the intelligent processing unit in the context management dimension, the storage granularity information, retrieval strategy information, context weight information, and / or context structure information in the initial behavior profile information are updated to obtain the target behavior profile information of the target object; and / or Based on the interaction behavior data between the target object and the intelligent processing unit in the domain knowledge dimension, the domain knowledge experience information, domain knowledge example information and / or object decision information in the initial behavior profile information are updated to obtain the target behavior profile information of the target object.
[0084] The sub-features of the interaction performance dimension can include response format information, response detail information, interface layout information, and / or recommendation strategy information. The sub-features of the tool invocation dimension can include tool invocation strategy information, tool parameter information, and / or information about tools to be updated. The sub-features of the code generation dimension can include coding style information, architectural pattern information, technology stack adaptation information, and / or code quality information. The sub-features of the context management dimension can include storage granularity information, retrieval strategy information, context weight information, and / or context structure information. The sub-features of the domain knowledge dimension can include domain knowledge experience information, domain knowledge example information, and / or object decision information.
[0085] Understandably, sub-feature information from multiple interaction dimensions can be represented in the behavioral profile information of the target object, and can be used to characterize the target object's behavioral habits and preferences.
[0086] In practical applications, each interaction dimension can update its sub-feature information based on its own update strategy and the collected interaction behavior signals.
[0087] When updating response format information in the interaction performance dimension, the proportion and nesting depth of code blocks, tables, lists, and text in the intelligent processing unit's output can be adjusted based on the user's receiving / modification mode. When updating response detail information, a mapping relationship between task type and response detail can be established; for example, users prefer detailed responses in debugging tasks and concise responses in quick modification tasks, and continuous optimization can be performed. When updating recommendation strategy information, personalized proactive recommendation strategies can evolve based on the user's task mode; for example, some users prefer the intelligent processing unit to proactively propose optimization suggestions, while others prefer the intelligent processing unit to passively process instructions. When updating interface layout information, if the intelligent processing unit has a configurable front-end interface, the layout and priority of buttons, panels, or shortcut controls in the front-end interface can be automatically adjusted based on the user's function using a heatmap.
[0088] When updating tool invocation strategy information at the tool invocation dimension, optimized tool invocation order and combination patterns can be learned from user interaction behavior. For example, if a user's workflow of first searching relevant documentation, then analyzing code structure, and finally writing the implementation is discovered, this scheduling strategy can be adopted in similar tasks. When updating tool parameter information, the parameter selection for each tool invocation can be optimized, such as search scope, code analysis depth, and test coverage granularity. Updates to be updated tool information can include new skill self-discovery and new skill self-learning. New skill self-discovery can be understood as identifying capability gaps in the intelligent processing unit by analyzing user interaction behavior outside the capability boundaries of the intelligent processing unit, such as frequently mentioning operations that the intelligent processing unit does not support or repeatedly manually completing a certain type of task. New skill self-learning can be understood as automatically discovering and learning suitable new tools or services from the available external tool registry by analyzing the large language model for the identified capability gaps, and then expanding them to the skill set of the intelligent processing unit after verification.
[0089] When updating coding style information in the code generation dimension, a coding style map can be constructed. This allows for the continuous extraction of coding style features (including indentation, naming conventions, comment density, function granularity, and error handling patterns) from user code modification behavior, building a coding style map specific to the target object. When updating architectural pattern information, the system can learn user architectural preferences in different scenarios, such as a preference for dependency injection over singletons, composition over inheritance, and functional pipelines over object-oriented programming. When updating technology stack adaptation information, the system can automatically detect changes in the user's project's technology stack and adjust library selection, API call methods, and configuration formats in code generation. When updating code quality information, the system can update user-specific code quality evaluation standards based on the user's code inspection patterns.
[0090] When updating the storage granularity information of the context management dimension, the storage granularity of memory entries can be dynamically adjusted based on the frequency of user citations of context information. For example, domains frequently cited by users are stored with finer granularity, while domains cited less frequently are stored with coarse-grained summaries. When updating the retrieval strategy information, the relevance ranking algorithm parameters for memory retrieval can be optimized to make the retrieval results more closely match the actual needs of users. When updating the context weight information, the forgetting weight of context memories can be adjusted. Different users have different timeliness requirements for different types of memories. For example, some users frequently cite architectural decisions from months ago, while others only care about content from the past week. The forgetting decay curves for various types of memories are automatically adjusted for different users. When updating the context structure information, the classification hierarchy and association methods of context memories can be automatically adjusted based on the user's knowledge structure and work patterns.
[0091] When updating domain knowledge experience information, domain knowledge entries can be extracted from each successful task execution. For example, in this user's microservice project, inter-service communication uniformly uses an event-driven pattern, and the domain-related parts of negative experiences are recorded as pitfall knowledge, such as whether a certain function can be used in a user's project. This knowledge is then proactively avoided in similar scenarios. When updating domain knowledge example information, repeatedly verified effective operation patterns can be extracted as domain knowledge example entries and automatically updated as the domain changes. When updating object decision information, important architectural decisions made by the user and their rationale can be recorded, serving as constraint references in subsequent related scenarios.
[0092] It is understandable that the user operations described above can be interpreted as collected interactive behavior data or interactive behavior signals. The above five interaction dimensions are based on the example of the intelligent processing unit provided in this embodiment of the specification for programming assistance services. The intelligent processing unit can also be used to provide other services; therefore, the interaction dimensions can be set according to the services provided by the intelligent processing unit, and this embodiment of the specification does not limit this.
[0093] In summary, by updating each interaction dimension independently, the independent evolution of each interaction dimension is achieved, thus preventing the update of one interaction dimension from affecting the normal operation of other interaction dimensions.
[0094] Furthermore, the step of updating the initial behavioral profile information based on the interaction behavior data to obtain the target behavioral profile information of the target object further includes: If it is determined that the interactive behavior data corresponds to at least two interactive dimensions, the dimensional association relationship between the at least two interactive dimensions is determined. Based on the dimensional relationships and the interaction behavior data, the sub-behavior profile information corresponding to the at least two interaction dimensions in the initial behavior profile information is updated to obtain the target behavior profile information of the target object.
[0095] Among them, the dimensional relationships between at least two interaction dimensions can include collaborative relationships, conflict relationships, and transmission relationships.
[0096] In practical applications, in combination with the above Figure 2 The meta-evolutionary control layer can include a multi-dimensional evolutionary executor, which can be used to perform specific update operations such as configuration adjustments, skill optimizations, code pattern updates, and memory strategy reorganization across multiple interaction dimensions. Each interaction dimension has an independent evolutionary strategy engine, and also includes a cross-dimensional correlation analysis module to detect and process the dimensional relationships between interaction dimensions.
[0097] Although each interaction dimension is orthogonal in the configuration space (and can be adjusted independently), there are cross-dimensional correlations in users' actual behavior patterns. For example, under a specific task type, users may have collaborative preferences in both interaction style and coding style. Cross-dimensional correlation analysis can capture and utilize dimensional correlations while maintaining orthogonality and independence.
[0098] There must be a synergistic relationship between at least two interaction dimensions. This can be understood as a positive correlation between the evolutionary directions of at least two interaction dimensions. For example, when the user instruction intelligent processing unit executes a code refactoring task, the interaction performance dimension prefers detailed responses, while the code generation dimension prefers rigorous generation. When at least two interaction dimensions with a synergistic relationship are detected, updating one of the interaction dimensions can automatically suggest the synergistic adjustment direction of other interaction dimensions associated with that dimension.
[0099] A conflicting relationship exists between at least two interaction dimensions. This can be understood as the evolutionary directions of at least two interaction dimensions producing contradictory effects. For example, the evolutionary direction of the interaction performance dimension requires concise responses, but the evolutionary direction of the code generation dimension requires detailed comments, leading to a longer overall output. When the cross-dimensional correlation analysis module detects a conflicting relationship between at least two interaction dimensions, it can resolve the conflict through priority arbitration or scenario-based isolation (e.g., prioritizing the code generation dimension when performing code tasks, and prioritizing the interaction performance dimension after completing question-and-answer tasks).
[0100] At least two interaction dimensions have a transmission relationship, which can be understood as the evolution of one interaction dimension triggering a chain reaction of evolutionary needs in another. For example, if the domain knowledge dimension discovers that a user is switching to a new technology stack, this can be transmitted to the code generation dimension, requiring an update to the coding style, and to the tool usage dimension, requiring the learning of new tools. Based on this, transmission routing rules can be established between at least two interaction dimensions. These rules describe the transmission route from the source interaction dimension to the target interaction dimension. When the source interaction dimension undergoes significant evolution, it automatically sends an evolutionary suggestion signal to the target interaction dimension. In essence, the evolution of an interaction dimension can be understood as the updating of its sub-feature information.
[0101] Based on this, given that at least two interaction dimensions are determined to correspond to the interaction behavior data and that there is a dimensional relationship between the at least two interaction dimensions, the sub-feature information corresponding to the at least two interaction dimensions in the initial behavior profile information can be updated by associating the dimensional relationship between the at least two interaction dimensions and the interaction behavior data to obtain the target behavior profile information of the target object.
[0102] In summary, by analyzing the dimensional relationships between at least two interaction dimensions, co-evolution among at least two interaction dimensions can be achieved.
[0103] Step 104: Determine the behavior rule information corresponding to the target object based on the target behavior profile information.
[0104] Among them, behavioral rule information can be understood as behavioral pattern rules of the target object abstracted from the target object's behavioral preferences and habits.
[0105] Based on this, behavioral rule information corresponding to the target object can be abstracted and extracted from the target object's target behavior profile information.
[0106] Further, determining the behavior rule information corresponding to the target object based on the target behavior profile information includes: Based on the target behavior profile information, determine the object feedback information corresponding to the target interaction behavior data, wherein the object feedback information includes positive feedback information and negative feedback information, and the target interaction behavior data is any one of the interaction behavior data; If the object feedback information is determined to be positive feedback information, rules are extracted based on the target interaction behavior data to obtain positive behavior rule information; If the object feedback information is determined to be negative feedback information, rules are extracted based on the target interaction behavior data to obtain negative behavior rule information, wherein the behavior rule information includes the positive behavior rule information and the negative behavior rule information.
[0107] Among them, object feedback information can be understood as the feedback information of the target object to the output information and processing information of the intelligent processing unit on the target interaction behavior data. Positive feedback information indicates that the target object recognizes the output information and processing information of the intelligent processing unit, while negative feedback information indicates that the target object does not recognize the output information and processing information of the intelligent processing unit. Positive behavior rule information can be understood as the rule information abstracted from the target interaction behavior data recognized by the target object, while negative behavior rule information can be understood as the rule information abstracted from the target interaction behavior data not recognized by the target object.
[0108] Specifically, based on the target behavior profile information, the output information of the target object to the intelligent processing unit regarding the target interaction behavior data, as well as the object feedback information of the processed information, can be determined. If the object feedback information is determined to be positive feedback information, rules can be extracted from the target interaction behavior data to obtain positive behavior rule information. If the object feedback information is determined to be negative feedback information, rules can be extracted from the target interaction behavior data to obtain negative behavior rule information.
[0109] Furthermore, after determining that the object feedback information is the positive feedback information, and extracting rules based on the target interaction behavior data to obtain positive behavior rule information, the method further includes: The positive behavior rule information is graded to obtain the graded positive behavior rule information; The marked positive behavior rule information is stored in the behavior rule database.
[0110] The process of assigning a level label to positive behavior rule information can be understood as assigning a criticality level label to the positive behavior rule information. The higher the criticality of the positive behavior rule information, the more important it is for the update of the intelligent processing unit. The behavior rule database can be used to store positive behavior rule information, which is then used for subsequent testing of behavior rule information replay for candidate update versions of the intelligent processing unit.
[0111] In practical applications, the "correct behavior" of the intelligent processing unit is not determined by the static test cases predefined by the developers, but is learned from the successful interactions between the intelligent processing unit and the target object in history. The interaction behavior data of these successful interactions can be extracted into replayable behavior contracts (i.e. positive behavior rule information) as verification test cases, thereby forming a set of behavior contract verification test cases (i.e. behavior rule database) that grows synchronously with the evolution of the intelligent processing unit.
[0112] See Figure 5 , Figure 5A schematic diagram illustrating the construction of a behavior rule database in a version self-update method according to an embodiment of this specification is shown. Figure 5 As shown, after each interaction between the user and the intelligent processing unit, the system evaluates whether the interaction was a successful interaction recognized by the user (i.e., whether the feedback information corresponding to the target interaction behavior data is positive feedback). The criteria include: explicit positive feedback from the user (likes, acceptance), the user's edit distance to the output being below a threshold (close to zero modification for direct use), and successful task execution without subsequent correction. For interaction behavior data that meets the above criteria, structured contract records (i.e., positive behavior rule information) can be extracted, including: input summary, expected behavior pattern, output feature constraints, and the corresponding interaction dimension.
[0113] Furthermore, to avoid test case explosion, instead of recording the specific input and output of each interaction, the interaction behavior data of each interaction is abstracted into a behavior contract at the behavior pattern level. For example, instead of recording the input as writing a response component and the output as functional component code, it is abstracted as "front-end component generation task, which must use the functional paradigm and include the target type definition." The abstraction level is determined according to the interaction dimension corresponding to the target interaction behavior data. For example, for the interaction performance dimension, positive behavior rule information can be abstracted into format / style features; for the code generation dimension, positive behavior rule information can be abstracted into architecture / style pattern features; and for the tool call dimension, positive behavior rule information can be abstracted into tool call sequence patterns.
[0114] Furthermore, after extracting multiple positive behavior rule information, similarity calculation can be performed on the multiple positive behavior rule information, and the positive behavior rule information with similarity higher than the similarity threshold can be merged and deduplicated, retaining the generalized version with a wider coverage.
[0115] Furthermore, after extracting multiple positive behavioral rule information, and determining that each positive behavioral rule information represents a new behavioral pattern, a criticality level can be assigned and marked for each positive behavioral rule information. The criticality level can include a first level, a second level, and a third level. The first level is the core level, involving the core functions of the intelligent processing unit, such as the basic correctness of code generation and the security of tool calls. When testing candidate update versions of the intelligent processing unit using the behavioral rule database, if a positive behavioral rule information at the first level fails, the replacement of that candidate update version is blocked. The second level of positive behavioral rule information involves user preferences (such as behaviors that users have explicitly corrected). When testing candidate update versions of the intelligent processing unit using the behavioral rule database, if the pass rate of the second level of positive behavioral rule information is less than 95%, the replacement of that candidate update version is blocked. The third level of positive behavioral rule information involves users' implicit habit preferences. When testing candidate update versions of the intelligent processing unit using the behavioral rule database, if the pass rate of the second level of positive behavioral rule information is less than 85%, a warning can be issued for that candidate update version. Understandably, the level label of positive behavior rule information can be dynamically adjusted based on the number of times verification is triggered and user feedback. If the positive behavior rule information obtained after deduplication is not a new behavior pattern, the confidence weight information and coverage count of the current positive behavior rule information can be updated to cover the existing positive behavior rule information in the existing contract library (i.e., behavior rule database).
[0116] Understandably, each interaction between the user and the intelligent processing unit may abstract new positive behavioral rule information, and the behavioral rule database continuously enriches as the user uses the intelligent processing unit. When a user's preferences change, the positive behavioral rule information corresponding to the previous preferences is marked as "evolved" through evolutionary rules and no longer participates in subsequent verification. Alternatively, the positive behavioral rule information corresponding to the previous preferences can be directly deleted from the behavioral rule database. Furthermore, in the behavioral rule database, for each piece of positive behavioral rule information, the version information of the intelligent processing unit at the time of its generation is recorded for easy traceability later.
[0117] In another embodiment of this specification, for interactive behavior data where the object's feedback information is negative, structured attribution analysis and extraction of negative behavior rule information can be performed. Specifically, the interaction dimension corresponding to the interactive behavior data can be determined, and pattern abstraction and rule generalization can be performed to transform a single negative feedback event into negative behavior rule information that can be generalized across scenarios, achieving efficient evolution by learning a class of rules from each mistake.
[0118] For target interaction behavior data, when a negative feedback event is detected, it indicates that the object's feedback information is negative, triggering the attribution analysis process for the target interaction behavior data. A negative feedback event can be understood as a negative feedback interaction behavior within the target interaction behavior data. Negative feedback events include: explicit user correction of the intelligent processing unit's output, user rejection or withdrawal of the intelligent processing unit's operation, task execution failure (detected by the intelligent processing unit's self-check or user feedback), and the user manually completing the same type of task N times consecutively without using the intelligent processing unit (implicit abandonment signal). Based on this, the meta-self-evolutionary control layer of the intelligent processing unit can save the complete context of the negative feedback event: the original user request, the intelligent processing unit's output, the user's correction / expected output, a snapshot of the target behavior profile information at the time of interaction, and the version of the intelligent processing unit used at that time.
[0119] In combination with the above Figure 2 The meta-evolutionary control layer also includes a negative feedback attribution analysis engine, which is used to perform structured root cause analysis on negative feedback events, determine the interaction dimension corresponding to the negative feedback event, extract generalizable evolutionary rules (i.e. negative behavior rule information), and mark the applicable scenario boundaries.
[0120] The negative feedback attribution analysis engine can perform structured analysis of the context of negative feedback events to determine the attribution dimension and root cause category of the error. When the attribution dimension is determined to be the interaction performance dimension, the root cause category could be, for example, an overly lengthy / brief response, format mismatch, or inappropriate tone. Attribution analysis can be performed by comparing the differences in output format before and after the user's correction (e.g., length, structure, tone words). When the attribution dimension is determined to be the tool invocation dimension, the root cause category could be, for example, the selection of an inappropriate tool, an inappropriate tool invocation order, or the omission of a necessary tool. Attribution analysis can be performed by analyzing whether the user's correction involved tool-related instructions or whether the user used tools not invoked by the intelligent processing unit. When the attribution dimension is determined to be the code generation dimension, the root cause category could be, for example, incompatible coding style, inappropriate architecture selection, mismatched naming conventions, or over / under-design. Attribution analysis can be performed by comparing the differences in style and structure between the user's modified code and the original code output by the intelligent processing unit. When the attribution dimension is determined to be the context management dimension, root cause categories could include, for example, failing to remember important context, referencing outdated information, or forgetting the user's previous explicit requests. Attribution analysis can be conducted by determining whether the error is related to missing or incorrect contextual information. When the attribution dimension is determined to be the domain knowledge dimension, root cause categories could include, for example, lack of domain-specific expertise or using the opposite pattern of that domain. Attribution analysis can be conducted by determining whether the error is caused by insufficient domain knowledge.
[0121] Furthermore, when performing attribution analysis, it is possible to perform attribution analysis based on multiple interaction dimensions driven by a large language model, or based on rule-based pattern matching, or based on statistical causal inference. The embodiments in this specification do not limit this approach.
[0122] See Figure 6 , Figure 6 This diagram illustrates a negative feedback attribution analysis in a version self-update method according to an embodiment of this specification. The specific scenario characteristics of negative feedback events (i.e., negative interaction events) can be abstracted into scenario patterns. For example, the target interaction behavior data corresponding to negative feedback information, "class components were used instead of functional components in the user's React project," can be abstracted as "in the user's front-end project, the component writing preference is functional paradigm." Structured evolutionary rules (i.e., negative behavior rule information) can be generated, in the following format: triggering condition: scenario pattern, evolutionary action: dimension, direction, and magnitude, applicable boundary: effective scenario range. The newly generated negative behavior rule information (i.e., evolutionary rules) is initialized with a low initial confidence level, such as 0.4. If it is continuously verified in subsequent interactions (i.e., no more similar negative feedback information is triggered), the confidence level is gradually increased. If it is overturned by counterexamples, the confidence level of the negative behavior rule information is reduced or discarded. Specifically, after context restoration and snapshot saving of negative interaction events, attribution analysis can be performed on the negative interaction events across multiple interaction dimensions. The attribution dimensions can be determined to determine whether they are attributions based on interaction performance, tool invocation, code generation, context management, or domain knowledge. After pattern abstraction and rule generalization, structured evolutionary rules are abstracted.
[0123] In summary, by extracting positive and negative behavioral rule information, the intelligent processing unit learns positive behavioral patterns and avoids the recurrence of erroneous behavioral patterns, thereby improving the processing quality and efficiency of the intelligent processing unit and enabling it to meet the personalized needs of users that change over time.
[0124] Step 106: When the preset update condition is determined based on the behavior rule information, the current version of the intelligent processing unit is updated to obtain the target update version of the intelligent processing unit, wherein the target update version is the version obtained by performing a self-evolution update at the code logic level on the current version and passing the update test.
[0125] The target update version can be understood as the version obtained by self-evolving the current version of the intelligent processing unit and passing the update test. Self-evolving the current version can be understood as the intelligent processing unit performing self-evolution at the code logic level.
[0126] Specifically, when the preset update conditions are determined based on behavioral rule information, the current version of the intelligent processing unit can be updated to obtain the target updated version of the intelligent processing unit. Furthermore, during the self-evolutionary update at the code logic level of the current version, the code corresponding to the current version can be updated. For example, the code corresponding to the current version can be modified, or code can be added to the code corresponding to the current version to implement new functions, thereby obtaining the target updated version of the intelligent processing unit.
[0127] In practical applications, the intelligent processing unit can be implemented through source code. The source code of the intelligent processing unit can include runtime source code and functional source code. The runtime source code can be used to support the basic operation of the intelligent processing unit, while the functional source code can be used to implement the functions of the intelligent processing unit, such as encoding and proofreading functions. The functional source code can be updated or deleted during the user's use of the intelligent processing unit, or code can be added to implement a new function during the user's use of the intelligent processing unit, so as to achieve self-evolution and update of the intelligent processing unit's code logic level.
[0128] Further, the step of updating the current version of the intelligent processing unit to obtain a target updated version of the intelligent processing unit when a preset update condition is determined based on the behavior rule information includes: When a preset update condition is determined based on the behavior rule information, the current version of the intelligent processing unit is updated to obtain a candidate update version of the intelligent processing unit. The candidate update version is tested to obtain test results. If the test results are determined to be successful, the candidate update version is determined as the target update version and the version is replaced.
[0129] Among them, the candidate update version can be understood as the updated version obtained by self-evolution of the current version of the intelligent processing unit. Self-evolution of the current version can be understood as self-evolution of the intelligent processing unit at the code logic level, which may include rewriting skill scripts, modifying architectural components, adding new functional processing logic, etc.
[0130] The triggering preset conditions are determined based on the behavior rule information, which can be that the number of times the negative behavior corresponding to the negative behavior rule information occurs reaches a preset number threshold.
[0131] Based on this, when the preset update conditions are determined according to the behavior rule information, the intelligent processing unit can be self-evolved at the code logic level to obtain candidate update versions of the intelligent processing unit. The candidate update versions can then be tested to obtain test results. If the test results are successful, the candidate update version can be determined as the target update version and the version can be replaced.
[0132] In practical applications, in combination with the above Figure 2 The meta-evolutionary control layer also includes an evolution safety and verification mechanism, which can be used to perform grayscale verification of each evolutionary change. The new configuration is applied in some interactions in an A / B comparison manner (i.e., normal operation and simulated operation). The full effect is applied after the effect is confirmed. At the same time, it provides evolution rollback capability to ensure that any evolutionary change can be safely revoked.
[0133] In addition, gray-scale verification of each evolutionary change can be achieved by comparing and verifying before and after the time series or by offline evaluation and verification based on the prediction model.
[0134] Further, the step of testing the candidate update version, obtaining test results, and, if the test result is determined to be a successful test, identifying the candidate update version as the target update version and performing version replacement includes: Retrieve the behavior rule information to be tested from the behavior rule database; In a simulated execution environment, the candidate update versions of the intelligent processing unit are simulated and tested according to the test behavior rule information to obtain the simulation test results; If the simulation test result is determined to be a pass, in response to the target object's call request to the intelligent processing unit within a second preset time interval, the call request is processed based on the candidate update version and the current version to obtain the candidate processing result output by the candidate update version and the current processing result output by the current version. The candidate processing results and the current processing results are compared and analyzed to obtain the comparison and analysis results; If the candidate update version meets the test pass conditions based on the comparative analysis results, the candidate update version is determined as the target update version and the version is replaced.
[0135] The behavior rule information to be tested can be understood as any behavior rule information included in the behavior rule database, including positive and negative behavior rule information. The simulated execution environment can be understood as a sandbox environment isolated from the normal operating environment of the intelligent processing unit, used for simulated testing. The second preset time interval can be understood as the time interval following the first preset time interval, which can be a future time period after the simulated test of the candidate update version of the intelligent processing unit. The test passing condition can be that the comparative analysis result meets a preset similarity threshold, or it can be that the user satisfaction index of the candidate processing result is greater than the user satisfaction index of the current processing result; this embodiment does not limit this.
[0136] Based on this, any behavioral rule information in the behavioral rule database can be identified as the behavioral rule information to be tested. In a simulated execution environment, candidate update versions of the intelligent processing unit are simulated and tested based on the behavioral rule information to be tested, and the simulation test results are obtained. If the simulation test result is determined to be a pass, the actual call request to the intelligent processing unit by the target object within a second preset time interval can be responded to. The actual call request is processed based on the current version of the intelligent processing unit, the current processing result is obtained, and the current processing result is returned to the user who initiated the actual call request. Simultaneously, the actual call request is processed based on the candidate update version of the intelligent processing unit, and candidate processing results are obtained. If the comparison analysis between the current processing result and the candidate processing result determines that the candidate update version meets the test pass conditions, the candidate update version is determined as the target update version and version replacement is performed.
[0137] In practical applications, see Figure 7 , Figure 7 The diagram illustrates a simulated test in a version self-update method according to one embodiment of this specification. Any code-level self-evolutionary change (rewriting skill scripts, modifying processing logic, adjusting architectural components) must pass full verification of the behavior contract verification test case set (i.e., the behavior rule database) before replacing the current version of the intelligent processing unit. Replacement is only allowed if all test cases pass; failure of any first-level positive behavior rule information test results in a blocking process.
[0138] When the meta-evolutionary control layer decides to execute code-level self-evolution (e.g., deciding to rewrite a skill script based on accumulated negative feedback attribution), it can build candidate updated versions of the intelligent processing unit in an isolated sandbox execution environment (i.e., a simulated testing environment). These candidate updated versions include: modified code + other unmodified modules + a snapshot of the current user's evolutionary configuration. The sandbox execution environment is completely isolated from the production environment, and any behavior of the candidate updated versions does not affect the user's actual usage.
[0139] Load all valid behavioral contracts from the behavioral contract verification test case set, sorting them from highest to lowest criticality. For each behavioral contract, simulate execution in a sandbox execution environment: take the input summary of the behavioral contract as input, execute the candidate updated version, and compare whether the output satisfies the behavioral pattern and output feature constraints defined in the behavioral contract. For Level 1 behavioral contracts, the test pass rate must be greater than 100%. Verification will be immediately halted and replacement blocked if any behavioral contract fails. For Level 2 behavioral contracts, the test pass rate must be greater than or equal to 95%. If it is lower, the contract will be blocked. For Level 3 behavioral contracts, the test pass rate must be greater than or equal to 80%. If it is lower, a warning will be issued, but the contract can be allowed to proceed with higher-privilege operations.
[0140] After the simulation test, a structured simulation test report can be generated as the simulation test result. The simulation test report can include: the pass rate statistics of behavioral contracts at each criticality level, the specific failure reasons of behavioral contracts that failed the test, and the assessment of the scope of behavioral changes of candidate update versions relative to the current version. If the simulation test fails, the failure information can be fed back to the evolutionary safety and validator, which will drive the code to be corrected and the test to be carried out again. See Figure 8 , Figure 8 This diagram illustrates a shadow parallel execution verification process in a version self-update method according to an embodiment of this specification. The simulated test process described above verifies the behavior contract replay, but behavior contract replay verification is based on historical behavior contract verification cases and cannot fully cover future real-world interaction scenarios. Therefore, after the simulated test passes, shadow execution verification can be performed: the candidate update version processes the same real user requests in parallel with the current version in an isolated environment (i.e., the simulated test environment), the output differences are compared, and the processing result of the current version is returned to the user.
[0141] Each user request to the intelligent processing unit is sent to both the current version and a candidate update version. The output of the current version is returned to the user normally, while the output of the candidate update version is not returned to the user and is only used for comparative analysis. The comparative analysis dimensions include: output semantic consistency, behavioral pattern matching degree, differences in tool call sequence, and differences in execution success rate.
[0142] The shadow run of the candidate update version can continue for N interactions (the size of the shadow window). The overall performance score of the candidate update version relative to the current version is calculated. If the performance of the candidate update version within the shadow window is not worse than that of the current version, then the shadow run is verified.
[0143] After the shadow run verification passes, the current version snapshot can be taken as the rollback baseline version and stored in the version history stack. During the request intervals of the intelligent processing unit (i.e., the inactive execution window), the current version of the intelligent processing unit is switched to the candidate update version that has passed the test (i.e., the target update version) to achieve safe hot replacement. The first M interactions after the switch can enter the observation period of the target update version, continuously monitoring user satisfaction indicators. If anomalies occur during the observation period (such as an increase in task failure rate or an abnormal increase in user correction frequency), the version rollback conditions can be triggered to roll back the intelligent processing unit to the rollback baseline version.
[0144] Furthermore, after determining the candidate update version as the target update version and performing version replacement, the method further includes: Store the historical versions of the intelligent processing unit in the historical version database; If the conditions for triggering a version rollback are determined, the intelligent processing unit performs a version rollback.
[0145] The conditions for triggering version rollback can include rollback based on time, rollback based on interaction dimensions, and rollback triggered by the user. The historical version database can be used to store historical versions of the intelligent processing unit.
[0146] Based on this, after the candidate update version is determined as the target update version and the version is replaced, the current version of the intelligent processing unit is stored as a historical version in the historical version database. When the conditions for triggering version rollback are determined, the intelligent processing unit can be rolled back, and the historical version corresponding to the rollback condition can be selected for rollback.
[0147] The version history stack (i.e., the historical version database) retains the K most recent stable versions (K is configurable, and the default value can be set to 5). When the version rollback conditions are triggered, it can roll back to a specified historical stable version, or roll back only a specific skill script / code module while retaining the evolution results of other modules, or the user can actively trigger the rollback.
[0148] In combination with the above Figure 2 The persistent storage layer can include a behavioral profile information database, a negative experience attribution record database, an evolutionary rule knowledge base, a skill registration and ability graph, a domain knowledge self-growing library, and an evolutionary history and rollback log to achieve persistent storage of this data.
[0149] Furthermore, during the gray-scale verification of the intelligent processing unit, in the subsequent N interactions, a candidate update version is randomly applied with a 50% probability, and the current version is applied with a 50% probability. The performance of the two versions in user satisfaction metrics, such as output acceptance rate, edit distance, and task success rate, is compared. If the user satisfaction metric of the candidate update version is greater than that of the current version and the improvement exceeds a significance threshold, the candidate update version is fully implemented. If the user satisfaction metric of the candidate update version is lower than that of the current version, the candidate update version is discarded, and the evolution rule corresponding to the candidate update version is marked as invalid.
[0150] Before each evolution change, the intelligent processing unit stores the current configuration version in the evolution history log. It supports rollback by interaction dimension (only rollback of changes in a certain interaction dimension without affecting other interaction dimensions), rollback by time (rollback to the configuration version at a certain historical point in time), and users can actively trigger rollback (such as user command to restore to the configuration of last week).
[0151] For example, taking a user as a full-stack developer using a programming aid agent as an example, on the first day of using the agent, as a cold start period, the user starts using the agent with the initial configuration set to default. In the code task, the user rewrites the class components output by the agent into function components three times in a row, and updates them in real time. The "component paradigm preference" sub-feature information of the code generation dimension is slightly adjusted from the default value towards the "functional" direction. In two conversations, the user asks the agent to search for documents before coding, triggering an instant update: the confidence weight information of the "pre-search of code task" sub-feature information of the tool call dimension gradually increases.
[0152] On the third day of the user's use of the agent (rapid adaptation period), the user submitted code using Vitest instead of Jest. The agent still used Jest to generate tests. The user corrected the user by saying "I used Vitest," triggering an event-driven targeted update. Negative feedback attribution analysis attributed the issue to "code generation dimension - test framework selection." Rule generalization generated evolutionary rules (positive behavior rule information) "[the user's front-end project], [test framework preference Vitest]." After gray-scale verification, the rules took full effect, and the agent automatically used Vitest in all front-end projects thereafter.
[0153] On the 7th day of the user's use of the agent (deep evolution cycle), a periodic update is triggered. Analyzing the cumulative interaction data of the week, it is found that the user accepts detailed replies in debugging tasks (average edit distance 12%) and prefers concise replies in quick modifications (edit distance for long replies 45%). Cross-dimensional correlation analysis reveals that the user prefers rigorous naming (snake_case) in backend tasks and camelCase in frontend tasks. Capability gap identification: the user manually performed database migration operations 4 times this week, and the agent did not provide this capability, triggering optimization of the skill self-discovery process memory strategy: the user frequently references architectural decisions from 5 days ago, extending the retention period of architecture-related memories.
[0154] On day 30 of the user's use of the agent (maturity stage), the agent has formed a complete behavioral profile: Interaction performance: detailed during debugging, concise during modification; preference for a mixed table + code block format; Tool usage: searching documentation, analyzing existing code, and writing implementations before coding tasks; automatically learned database migration tools; Code generation: camelCase, functional, and Vitest for front-end, snakeCase, DDD, and PyTest for back-end; Context management: architectural decisions are retained for 30 days, and specific implementation details are retained for 7 days; Domain knowledge: accumulated 47 architectural rules and 12 pitfalls for the user's project; Behavioral contract verification test case set: accumulated 312 behavioral contracts (28 at level 1, 97 at level 2, and 187 at level 3). The user's correction frequency for the agent has decreased from 10+ times per day on day 1 to 1-2 times per day. On the 45th day of user interaction with the agent (code self-evolution event), the agent accumulated 5 negative feedback attributions regarding the "file search" skill: The user repeatedly manually corrected the agent's search strategy in a large project (the agent always performed global searches, leading to excessive noise; the user preferred to limit the directory scope before searching). An event-driven targeted update triggered code-level self-evolution. The meta-self-evolution control layer decided to rewrite the file search skill script, adding "intelligent directory scope narrowing" logic, resulting in a candidate update version. The behavior contract was fully validated: the candidate update version replayed 312 behavior contracts in an isolated sandbox. Level 1 (28 contracts): all passed (core code generation, secure tool calls, and other basic capabilities remained unaffected); Level 2 (97 contracts): 96 passed (98.9% pass rate, exceeding the 95% threshold); Level 3 (187 contracts): 171 passed (91.4% pass rate, exceeding the 80% threshold). Shadow parallel execution: the candidate update version and the current version processed 20 real search requests in parallel. Search result accuracy: 87% for the candidate update version and 62% for the current version, a significant improvement, with no degradation in other capabilities. Safe hot-swap: All verifications passed. The file search skill script was replaced, and the metrics remained normal within the 10-interaction observation period. User experience: Search results are significantly more accurate, eliminating the need to manually specify directory ranges.
[0155] In summary, by leveraging behavioral profile information across multiple interaction dimensions and a three-tiered evolutionary scheduling system, the system continuously learns from explicit and implicit signals in every interaction between the user and the intelligent processing unit. This causes the deviation between the behavioral profile information and the user's true preferences to monotonically decrease with the number of interactions, achieving a qualitative leap from requiring correction for most interactions to almost eliminating the need for correction. Through dimensional attribution, pattern abstraction, and rule generalization, specific corrections are elevated to evolutionary rules at the interaction dimension level. These rules are effective for all variant scenarios within the same interaction dimension, achieving a fundamental breakthrough from remembering one instance to learning an entire class. Through capability gap analysis in periodic deep evolution, the system identifies which task types users frequently abandon using the intelligent processing unit or which require manual completion. This automatically identifies the intelligent processing unit's skill blind spots, and then acquires new skills through a tool registry matching and self-learning process, ensuring that the intelligent processing unit's capability boundaries expand synchronously with the expansion of user needs. By orthogonally decomposing multiple interaction dimensions, each dimension is isolated in the configuration space. This means that a configuration change in one interaction dimension will not affect the parameters of another. Furthermore, cross-dimensional correlation analysis can proactively detect and handle relationships between multiple interaction dimensions when necessary, enabling synergistic enhancement, conflict arbitration, and transmission notification among multiple interaction dimensions. This transforms cross-interference between interaction dimensions from uncontrollable side effects into conscious collaborative decision-making. An automatic behavioral contract accumulation mechanism continuously extracts verification test cases from historical successful interactions. The content of these test cases reflects users' real personalized behavioral patterns rather than generic behaviors. When code self-evolution occurs, full contract verification can detect whether user-approved behaviors have been violated. Overlaying shadow parallel execution can detect adaptability to real-world scenarios, and setting up multiple version rollbacks provides a backup for failures, achieving a complete security closed loop from offline verification and online verification to fault recovery.
[0156] The above method can acquire interaction behavior data of the target object with the intelligent processing unit in multiple interaction dimensions within a first time interval. Based on the interaction behavior data in multiple interaction dimensions, the target behavior profile information of the target object is determined. Based on the target behavior profile information, the behavior rule information of the target object in the interaction process with the intelligent processing unit is abstracted. When the preset update conditions are determined based on the behavior rule information, the current version of the intelligent processing unit is updated to obtain the target update version of the intelligent processing unit. This allows the target update version of the intelligent processing unit to learn the interaction behavior habits and service preferences of the target object, enabling the intelligent processing unit to meet the personalized needs of users and further improve the user experience.
[0157] Corresponding to the above method embodiments, this specification also provides embodiments of a version self-updating device applied to an intelligent processing unit. Figure 9A schematic diagram of a version self-updating device according to one embodiment of this specification is shown. Figure 9 As shown, the device includes: The acquisition module 902 is configured to acquire the interaction behavior data of the target object with the intelligent processing unit in multiple interaction dimensions within a first time interval, and determine the target behavior profile information of the target object based on the interaction behavior data. The determination module 904 is configured to determine the behavior rule information corresponding to the target object based on the target behavior profile information; The update module 906 is configured to update the current version of the intelligent processing unit when a preset update condition is determined based on the behavior rule information, so as to obtain a target update version of the intelligent processing unit, wherein the target update version is a version obtained by performing a self-evolution update at the code logic level on the current version and passing the update test.
[0158] In an optional embodiment, the acquisition module 902 is further configured to: In response to the target object's request to invoke the intelligent processing unit, the intelligent processing unit processes the request and constructs initial behavioral profile information of the target object. The system acquires interactive behavior data across multiple dimensions of the target object during the process of calling the intelligent processing unit within the first time interval, and updates the initial behavior profile information based on the interactive behavior data to obtain the target behavior profile information of the target object, so that the intelligent processing unit can process the scheduling request of the target object based on the target behavior profile information.
[0159] In an optional embodiment, the acquisition module 902 is further configured to: According to the preset update rules, the interactive behavior data of the target object in calling the intelligent processing unit within the first time interval is obtained in multiple interactive dimensions.
[0160] In an optional embodiment, the acquisition module 902 is further configured to: If the target object completes the invocation request by calling the intelligent processing unit within the first time interval, the interactive behavior data corresponding to the invocation request is obtained; and / or According to a preset time interval, acquire interaction data between the target object and the intelligent processing unit during the process of the target object calling the intelligent processing unit within the preset time interval, wherein the preset time interval is the time interval for triggering the update, and the preset time interval is shorter than the first time interval; and / or The system acquires interactive behavior data of the target object during the process of calling the intelligent processing unit within the first time interval, which meets preset event conditions.
[0161] In an optional embodiment, the acquisition module 902 is further configured to: Interaction behavior signals are collected from the interaction behavior data, wherein the interaction behavior signals include a first type of behavior signal and a second type of behavior signal, the first type of behavior signal is used to characterize the explicit preference information of the target object, and the second type of behavior signal is used to characterize the behavior pattern information of the target object; The interaction behavior signal is mapped to the corresponding target interaction dimension, and the interaction behavior signal is encoded to obtain the updated feature information corresponding to the target interaction dimension, wherein the target interaction dimension is any one of the plurality of interaction dimensions; The initial behavioral profile information is updated based on the updated feature information corresponding to the target interaction dimension to obtain the target behavioral profile information of the target object.
[0162] In an optional embodiment, the acquisition module 902 is further configured to: Determine the confidence weight information of the updated feature information corresponding to the target interaction dimension; Based on the time decay function, the updated feature information, and the confidence weight information, the sub-behavioral profile information corresponding to the target interaction dimension in the initial behavioral profile information is updated to obtain the target behavioral profile information of the target object.
[0163] In one optional embodiment, the multiple interaction dimensions include interaction performance dimension, tool invocation dimension, code generation dimension, context management dimension, and domain knowledge dimension; The acquisition module 902 is further configured to: Based on the interaction behavior data of the target object and the intelligent processing unit in the interaction performance dimension, the response format information, response detail information, interface layout information, and / or recommendation strategy information in the initial behavior profile information are updated to obtain the target behavior profile information of the target object; and / or Based on the interaction behavior data of the target object and the intelligent processing unit in the tool invocation dimension, the tool invocation strategy information, tool parameter information, and / or tool information to be updated in the initial behavior profile information are updated to obtain the target behavior profile information of the target object; and / or Based on the interaction behavior data between the target object and the intelligent processing unit in the code generation dimension, the coding style information, architecture pattern information, technology stack adaptation information, and / or code quality information in the initial behavior profile information are updated to obtain the target behavior profile information of the target object; and / or Based on the interaction behavior data between the target object and the intelligent processing unit in the context management dimension, the storage granularity information, retrieval strategy information, context weight information, and / or context structure information in the initial behavior profile information are updated to obtain the target behavior profile information of the target object; and / or Based on the interaction behavior data between the target object and the intelligent processing unit in the domain knowledge dimension, the domain knowledge experience information, domain knowledge example information and / or object decision information in the initial behavior profile information are updated to obtain the target behavior profile information of the target object.
[0164] In an optional embodiment, the acquisition module 902 is further configured to: If it is determined that the interactive behavior data corresponds to at least two interactive dimensions, the dimensional association relationship between the at least two interactive dimensions is determined. Based on the dimensional relationships and the interaction behavior data, the sub-behavior profile information corresponding to the at least two interaction dimensions in the initial behavior profile information is updated to obtain the target behavior profile information of the target object.
[0165] In an optional embodiment, the determining module 904 is further configured to: Based on the target behavior profile information, determine the object feedback information corresponding to the target interaction behavior data, wherein the object feedback information includes positive feedback information and negative feedback information, and the target interaction behavior data is any one of the interaction behavior data; If the object feedback information is determined to be positive feedback information, rules are extracted based on the target interaction behavior data to obtain positive behavior rule information; If the object feedback information is determined to be negative feedback information, rules are extracted based on the target interaction behavior data to obtain negative behavior rule information, wherein the behavior rule information includes the positive behavior rule information and the negative behavior rule information.
[0166] In an optional embodiment, the determining module 904 is further configured to: The positive behavior rule information is graded to obtain the graded positive behavior rule information; The marked positive behavior rule information is stored in the behavior rule database.
[0167] In an optional embodiment, the update module 906 is further configured to: When a preset update condition is determined based on the behavior rule information, the current version of the intelligent processing unit is updated to obtain a candidate update version of the intelligent processing unit. The candidate update version is tested to obtain test results. If the test results are determined to be successful, the candidate update version is determined as the target update version and the version is replaced.
[0168] In an optional embodiment, the update module 906 is further configured to: Retrieve the behavior rule information to be tested from the behavior rule database; In a simulated execution environment, the candidate update versions of the intelligent processing unit are simulated and tested according to the test behavior rule information to obtain the simulation test results; If the simulation test result is determined to be a pass, in response to the target object's call request to the intelligent processing unit within a second preset time interval, the call request is processed based on the candidate update version and the current version to obtain the candidate processing result output by the candidate update version and the current processing result output by the current version. The candidate processing results and the current processing results are compared and analyzed to obtain the comparison and analysis results; If the candidate update version meets the test pass conditions based on the comparative analysis results, the candidate update version is determined as the target update version and the version is replaced.
[0169] In an optional embodiment, the update module 906 is further configured to: Store the historical versions of the intelligent processing unit in the historical version database; If the conditions for triggering a version rollback are determined, the intelligent processing unit performs a version rollback.
[0170] The aforementioned device can acquire interaction behavior data of the target object with the intelligent processing unit in multiple interaction dimensions within a first time interval. Based on the interaction behavior data in multiple interaction dimensions, it can determine the target object's target behavior profile information. Based on the target behavior profile information, it can abstract the behavior rule information of the target object during the interaction with the intelligent processing unit. When the preset update conditions are determined based on the behavior rule information, the current version of the intelligent processing unit is updated to obtain the target update version of the intelligent processing unit. This allows the target update version of the intelligent processing unit to learn the target object's interaction behavior habits and service preferences, enabling the intelligent processing unit to meet the user's personalized needs and further improve the user experience.
[0171] The above is an illustrative scheme of a version self-update device according to this embodiment. It should be noted that the technical solution of this version self-update device and the technical solution of the version self-update method described above belong to the same concept. For details not described in detail in the technical solution of the version self-update device, please refer to the description of the technical solution of the version self-update method described above.
[0172] See Figure 10 , Figure 10 This specification illustrates an architecture diagram of a version self-updating system according to an embodiment of the present specification. The version self-updating system may include a client 100 and a server 200. Client 100 is used to send a request to the server 200 to invoke the target object; Server 200 is used to acquire interaction behavior data of the target object with the intelligent processing unit in multiple interaction dimensions within a first time interval, and determine the target behavior profile information of the target object based on the interaction behavior data; determine the corresponding behavior rule information of the target object based on the target behavior profile information; and update the current version of the intelligent processing unit when a preset update condition is triggered based on the behavior rule information to obtain the target updated version of the intelligent processing unit, wherein the target updated version is a version obtained by performing a self-evolution update at the code logic level on the current version and passing update testing. It is also used to process call requests and obtain processing results.
[0173] Client 100 is also used to receive processing results sent by server 200.
[0174] The self-updating system may include multiple clients 100 and a server 200. Clients 100 can be referred to as end-side devices, and the server 200 can be referred to as cloud-side devices. Multiple clients 100 can establish communication connections through the server 200. Users can interact with the server 200 through clients 100 to receive data sent by other clients 100, or to send data to other clients 100, etc.
[0175] In this system, client 100 and server 200 establish a connection via a network. The network provides the medium for communication between client 100 and server 200. The network can include various connection types, such as wired or wireless communication links or fiber optic cables. Data transmitted by client 100 may need to undergo encoding, transcoding, compression, or other processing before being published to server 200.
[0176] Client 100 can be a browser, an app (application), a web application such as an H5 (HyperText Markup Language 5) application, a lightweight application (also known as a mini-program), or a cloud application. Client 100 can be developed based on the software development kit (SDK) of the corresponding service provided by server 200, such as a real-time communication (RTC) SDK. Client 100 can be deployed on a computing device and depends on the device or certain apps on the device to run. The computing device may have a display screen and support information browsing, such as a personal mobile terminal like a mobile phone, tablet, or personal computer. Various other types of applications can also be configured on the computing device, such as human-computer interaction applications, model training applications, text processing applications, web browser applications, shopping applications, search applications, instant messaging tools, email clients, and social media platform software.
[0177] Server 200 may include servers providing various services, such as servers providing communication services to multiple clients, servers supporting backend training of models used on clients, and servers processing data sent by clients. It should be noted that server 200 can be implemented as a distributed server cluster composed of multiple servers, or as a single server. The server can also be a server in a distributed system, or a server integrated with blockchain. The server can also be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology.
[0178] It is worth noting that the version self-update method provided in the embodiments of this specification is generally executed by the server. However, in other embodiments of this specification, the client may also have similar functionality to the server, thereby executing the version self-update method provided in the embodiments of this specification. In other embodiments, the version self-update method provided in the embodiments of this specification may also be executed jointly by the client and the server.
[0179] Figure 11 A structural block diagram of a computing device 1100 according to an embodiment of this application is shown. The components of the computing device 1100 include, but are not limited to, a memory 1110 and a processor 1120. The processor 1120 is connected to the memory 1110 via a bus 430, and a database 1150 is used to store data.
[0180] The computing device 1100 also includes an access device 1140, which enables the computing device 1100 to communicate via one or more networks 1160. Examples of these networks include Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or combinations of communication networks such as the Internet. The access device 1140 may include one or more of any type of wired or wireless network interface (e.g., a network interface card (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, a Wi-MAX (Worldwide Interoperability for Microwave Access) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, a Near Field Communication (NFC) interface, and so on.
[0181] In one embodiment of this application, the aforementioned components of the computing device 1100 and Figure 11 Other components, not shown, can also be connected to each other, for example, via a bus. It should be understood that... Figure 11 The block diagram of the computing device shown is for illustrative purposes only and is not intended to limit the scope of this application. Those skilled in the art can add or replace other components as needed.
[0182] The computing device 1100 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or personal computers (PCs). The computing device 1100 can also be a mobile or stationary server.
[0183] The processor 1120 is used to execute the following computer program / instructions, which, when executed by the processor, implement the steps of the above-described version self-update method.
[0184] The above is an illustrative scheme of a computing device according to this embodiment. It should be noted that the technical solution of this computing device and the technical solution of the above-described version self-update method belong to the same concept. For details not described in detail in the technical solution of the computing device, please refer to the description of the technical solution of the above-described version self-update method.
[0185] An embodiment of this specification also provides a computer-readable storage medium storing a computer program / instructions that, when executed by a processor, implement the steps of the above-described version self-update method.
[0186] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on its differences from other embodiments. In particular, the computer-readable storage medium embodiments are described simply because they are substantially similar to the version self-update method embodiments; relevant parts can be referred to in the description of the version self-update method embodiments.
[0187] An embodiment of this specification also provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of the above-described version self-update method.
[0188] The above is an illustrative scheme of a computer program product according to this embodiment. It should be noted that the technical solution of this computer program product and the technical solution of the above-described version self-update method belong to the same concept. For details not described in detail in the technical solution of the computer program product, please refer to the description of the technical solution of the above-described version self-update method.
[0189] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.
[0190] The computer instructions include computer program code, which may be in the form of source code, object code, executable file, or certain intermediate forms. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium may be appropriately added or removed according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media may not include electrical carrier signals and telecommunication signals.
[0191] It should be noted that the above description describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in a different order than that shown in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous. Secondly, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments of this specification.
[0192] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0193] The preferred embodiments disclosed above are merely illustrative of this specification. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the embodiments described herein. These embodiments are selected and specifically described in this specification to better explain the principles and practical applications of the embodiments, thereby enabling those skilled in the art to better understand and utilize this specification. This specification is limited only by the claims and their full scope and equivalents.
Claims
1. A version self-update method, applied to an intelligent processing unit, comprising: The system acquires interaction behavior data of the target object with the intelligent processing unit in multiple interaction dimensions within a first time interval, and determines the target behavior profile information of the target object based on the interaction behavior data. Based on the target behavior profile information, the behavior rule information corresponding to the target object is determined. The behavior rule information is the behavior pattern rule of the target object abstracted from the target object's behavior preferences and habits. The behavior rule information includes positive behavior rule information and negative behavior rule information. The positive behavior rule information is the rule information abstracted from the target interaction behavior data that the target object approves of. The negative behavior rule information is the rule information abstracted from the target interaction behavior data that the target object does not approve of. When the preset update conditions are determined based on the behavior rule information, the current version of the intelligent processing unit is updated to obtain the target update version of the intelligent processing unit. The target update version is a version obtained by performing a self-evolution update at the code logic level on the current version and passing the update test.
2. The method according to claim 1, wherein acquiring the interaction behavior data of the target object with the intelligent processing unit in multiple interaction dimensions within a first time interval, and determining the target behavior profile information of the target object based on the interaction behavior data, includes: In response to the target object's request to invoke the intelligent processing unit, the intelligent processing unit processes the request and constructs initial behavioral profile information of the target object. The system acquires interactive behavior data across multiple dimensions of the target object during the process of calling the intelligent processing unit within the first time interval, and updates the initial behavior profile information based on the interactive behavior data to obtain the target behavior profile information of the target object, so that the intelligent processing unit can process the scheduling request of the target object based on the target behavior profile information.
3. The method according to claim 2, wherein obtaining the interactive behavior data of the target object in calling the intelligent processing unit within the first time interval, comprising: According to the preset update rules, the interactive behavior data of the target object in calling the intelligent processing unit within the first time interval is obtained in multiple interactive dimensions.
4. The method according to claim 3, wherein obtaining interactive behavior data of multiple interactive dimensions of the target object during the process of calling the intelligent processing unit within the first time interval according to a preset update rule includes: If the target object completes the call request processing by the intelligent processing unit within the first time interval, the interactive behavior data corresponding to the call request is obtained. and / or According to a preset time interval, the interaction behavior data between the target object and the intelligent processing unit during the process of the target object calling the intelligent processing unit within the preset time interval is obtained, wherein the preset time interval is the time interval for triggering the update, and the preset time interval is shorter than the first time interval; and / or The system acquires interactive behavior data of the target object during the process of calling the intelligent processing unit within the first time interval, which meets preset event conditions.
5. The method according to claim 2, wherein updating the initial behavioral profile information based on the interaction behavior data to obtain the target behavioral profile information of the target object includes: Interaction behavior signals are collected from the interaction behavior data, wherein the interaction behavior signals include a first type of behavior signal and a second type of behavior signal, the first type of behavior signal is used to characterize the explicit preference information of the target object, and the second type of behavior signal is used to characterize the behavior pattern information of the target object; The interaction behavior signal is mapped to the corresponding target interaction dimension, and the interaction behavior signal is encoded to obtain the updated feature information corresponding to the target interaction dimension, wherein the target interaction dimension is any one of the plurality of interaction dimensions; The initial behavioral profile information is updated based on the updated feature information corresponding to the target interaction dimension to obtain the target behavioral profile information of the target object.
6. The method according to claim 5, wherein updating the initial behavioral profile information based on the updated feature information corresponding to the target interaction dimension to obtain the target behavioral profile information of the target object includes: Determine the confidence weight information of the updated feature information corresponding to the target interaction dimension; Based on the time decay function, the updated feature information, and the confidence weight information, the sub-behavioral profile information corresponding to the target interaction dimension in the initial behavioral profile information is updated to obtain the target behavioral profile information of the target object.
7. The method according to any one of claims 2-6, wherein the plurality of interaction dimensions includes an interaction performance dimension, a tool invocation dimension, a code generation dimension, a context management dimension, and a domain knowledge dimension; The step of updating the initial behavior profile information based on the interaction behavior data to obtain the target behavior profile information of the target object includes: Based on the interaction behavior data of the target object and the intelligent processing unit in the interaction performance dimension, the response format information, response detail information, interface layout information, and / or recommendation strategy information in the initial behavior profile information are updated to obtain the target behavior profile information of the target object; and / or Based on the interaction behavior data of the target object and the intelligent processing unit in the tool invocation dimension, the tool invocation strategy information, tool parameter information, and / or tool information to be updated in the initial behavior profile information are updated to obtain the target behavior profile information of the target object; and / or Based on the interaction behavior data between the target object and the intelligent processing unit in the code generation dimension, the coding style information, architecture pattern information, technology stack adaptation information, and / or code quality information in the initial behavior profile information are updated to obtain the target behavior profile information of the target object; and / or Based on the interaction behavior data between the target object and the intelligent processing unit in the context management dimension, the storage granularity information, retrieval strategy information, context weight information, and / or context structure information in the initial behavior profile information are updated to obtain the target behavior profile information of the target object; and / or Based on the interaction behavior data between the target object and the intelligent processing unit in the domain knowledge dimension, the domain knowledge experience information, domain knowledge example information and / or object decision information in the initial behavior profile information are updated to obtain the target behavior profile information of the target object.
8. The method according to any one of claims 2-6, wherein updating the initial behavioral profile information based on the interaction behavior data to obtain the target behavioral profile information of the target object further includes: If it is determined that the interactive behavior data corresponds to at least two interactive dimensions, the dimensional association relationship between the at least two interactive dimensions is determined. Based on the dimensional relationships and the interaction behavior data, the sub-behavior profile information corresponding to the at least two interaction dimensions in the initial behavior profile information is updated to obtain the target behavior profile information of the target object.
9. The method according to any one of claims 1-6, wherein determining the behavior rule information corresponding to the target object based on the target behavior profile information includes: Based on the target behavior profile information, determine the object feedback information corresponding to the target interaction behavior data, wherein the object feedback information includes positive feedback information and negative feedback information, and the target interaction behavior data is any one of the interaction behavior data; If the object feedback information is determined to be positive feedback information, rules are extracted based on the target interaction behavior data to obtain positive behavior rule information; If the object feedback information is determined to be negative feedback information, rules are extracted based on the target interaction behavior data to obtain negative behavior rule information, wherein the behavior rule information includes the positive behavior rule information and the negative behavior rule information.
10. The method according to claim 9, further comprising, after determining that the object feedback information is the positive feedback information and extracting rules based on the target interaction behavior data to obtain positive behavior rule information, the method further comprises: The positive behavior rule information is graded to obtain the graded positive behavior rule information; The marked positive behavior rule information is stored in the behavior rule database.
11. The method according to any one of claims 1-6, wherein updating the current version of the intelligent processing unit to obtain a target updated version of the intelligent processing unit when a preset update condition is determined based on the behavior rule information includes: When a preset update condition is determined based on the behavior rule information, the current version of the intelligent processing unit is updated to obtain a candidate update version of the intelligent processing unit. The candidate update version is tested to obtain test results. If the test results are determined to be successful, the candidate update version is determined as the target update version and the version is replaced.
12. The method according to claim 11, wherein testing the candidate update version, obtaining test results, and determining the candidate update version as the target update version and performing version replacement upon determining that the test result is a successful test, comprises: Retrieve the behavioral rule information to be tested from the behavioral rule database; In a simulated execution environment, the candidate update versions of the intelligent processing unit are simulated and tested according to the test behavior rule information to obtain the simulation test results; If the simulation test result is determined to be a pass, in response to the target object's call request to the intelligent processing unit within a second preset time interval, the call request is processed based on the candidate update version and the current version to obtain the candidate processing result output by the candidate update version and the current processing result output by the current version. The candidate processing results and the current processing results are compared and analyzed to obtain the comparison and analysis results; If the candidate update version meets the test pass conditions based on the comparative analysis results, the candidate update version is determined as the target update version and the version is replaced.
13. The method according to claim 11, further comprising, after determining the candidate update version as the target update version and performing version replacement: Store the historical versions of the intelligent processing unit in the historical version database; If the conditions for triggering a version rollback are determined, the intelligent processing unit performs a version rollback.
14. A computing device, comprising: Memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions, which, when executed by the processor, implement the steps of the method according to any one of claims 1 to 13.
15. A computer-readable storage medium storing a computer program / instructions that, when executed by a processor, implement the steps of the method according to any one of claims 1 to 13.
16. A computer program product comprising a computer program / instructions that, when executed by a processor, implement the steps of the method according to any one of claims 1 to 13.
Citation Information
Patent Citations
Large model application publishing method and device, equipment and medium
CN119473344A
Method and device for intelligent computing cloud platform to realize agent self-evolution through computing power
CN121614114A