A model inference method and program product
Patent Information
- Application Number
- CN202611062905.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-17
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2046-07-17
AI Technical Summary
但现有编码模型完全无法获取上述运行时动态数据,只能基于源代码等静态数据进行推断,导致编码模型定位问题缺陷的准确率和效率较低,需要开发人员花费大量时间反复进行手动调试和验证
[0019]在上述的实现过程中,通过在弹窗显示事件触发后捕获当前线程的完整调用栈并进行过滤,保留业务代码层的调用栈帧,使得弹窗触发的来源路径更清晰简洁,减少大量栈帧对后续定位的干扰。将过滤后的调用栈与弹窗的标识信息、时间戳和所属父容器一并存入环形缓冲区,使得系统能够持久记录最近发生的弹窗事件及其触发来源。在目标上下文中包含弹窗状态信息时,从环形缓冲区中查询弹窗记录,能够快速识别异常弹窗,并通过过滤后的调用栈精准定位到触发该弹窗的业务方法,编码模型可以直接对该业务方法进行分析和修复,缩短了定位时间,提高了代码修复的效率。
Smart Images

Figure CN122596261B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and more specifically, to a model reasoning method and program product. Background Technology
[0002] In actual development, code defect fixing and feature debugging account for about half of the total development time. Especially in the development scenario of mobile terminal applications, the location and fixing of defects are highly dependent on the dynamic data generated by the application at runtime. However, the existing coding model cannot obtain the above-mentioned runtime dynamic data at all, and can only make inferences based on static data such as source code. This results in low accuracy and efficiency of the coding model in locating problems and defects, requiring developers to spend a lot of time repeatedly performing manual debugging and verification. Summary of the Invention
[0003] The purpose of this application is to provide a model reasoning method, program product, electronic device and storage medium to improve the above-mentioned technical problems.
[0004] In a first aspect, embodiments of this application provide a model inference method, comprising: collecting dynamic data during the runtime of a terminal application; receiving a user's inference request and determining the task type corresponding to the inference request based on the inference request; determining a corresponding main chain and auxiliary chain based on the task type, and determining a target context from the dynamic data and static code context based on the main chain and auxiliary chain; wherein the main chain represents an information source branch that is preferentially selected into the target context, and the auxiliary chain represents an information source branch that is supplementarily selected into the target context; the information source branch includes a dynamic branch and a static branch, the dynamic branch pointing to the dynamic data, and the static branch pointing to the static code context; inputting the target context into an encoding model, and having the encoding model perform model inference based on the target context.
[0005] In the above implementation process, dynamic data from the terminal application runtime is collected, and the target context is determined from the dynamic data and static code context. This target context is then input into the coding model, enabling the coding model to utilize both static and dynamic data during inference. Defect localization shifts from a speculation-driven approach based on static code to an evidence-driven approach based on runtime data, thereby improving the accuracy of model inference. The task type is determined based on the inference request, and the target context is determined from the dynamic data and static code context according to the data priority corresponding to the task type. This makes the information input into the model more targeted, reduces interference from irrelevant information, and reduces the time and resources spent on exploration during model inference. Inputting the priority-selected target context into the coding model for model inference improves the efficiency and success rate of model inference.
[0006] Optionally, in this embodiment of the application, collecting dynamic data during the runtime of the terminal application includes: collecting dynamic data through a runtime data collection module deployed on the terminal application, wherein the runtime data collection module includes multiple data processors divided according to data categories, the multiple data processors independently collect and cache various types of dynamic data during the runtime of the terminal application, and provide a data query interface within the terminal application; the dynamic data includes at least one of the following: network communication data, page navigation status data, pop-up lifecycle data, view inspection data, key-value pair storage data, and event tracking data.
[0007] In the above implementation process, by deploying a runtime data acquisition module containing multiple independent data processors, different types of dynamic data can be acquired and cached separately, with each processor operating independently without interference. This reduces the coupling within the module and facilitates the subsequent expansion to new data categories. Each processor caches the acquired data within the terminal application, reducing the performance overhead of repeated acquisition. By providing a unified data query interface within the terminal application, external systems can obtain the latest runtime status during the normal operation of the terminal application.
[0008] Optionally, in this embodiment of the application, after collecting dynamic data during the runtime of the terminal application, the method further includes: obtaining dynamic data from the data processor through a protocol translation intermediate layer; the protocol translation intermediate layer establishes communication connections with the encoding model and the terminal application respectively, and the protocol translation intermediate layer is used to convert the tool call request of the encoding model into a message format recognizable by the terminal application, and convert the dynamic data returned by the terminal application into structured text corresponding to the encoding model.
[0009] In the above implementation process, a standardized data interaction channel is established between the encoding model and the terminal application by deploying a protocol translation middleware layer. The protocol translation middleware layer uses an inter-process communication protocol to connect to the encoding model, allowing the encoding model to initiate data acquisition requests without needing to know the communication details of the terminal application. The protocol translation middleware layer converts the standardized tool call requests of the encoding model into message formats that the terminal application can recognize, enabling communication between different protocol systems. The protocol translation middleware layer converts the dynamic data returned by the terminal application into structured text that the encoding model can consume.
[0010] Optionally, in this embodiment of the application, determining the corresponding main chain and auxiliary chain according to the task type, and determining the target context from the dynamic data and static code context based on the main chain and auxiliary chain, includes: determining the main chain and auxiliary chain corresponding to the task type according to the pre-established mapping relationship between task types and information source branches; obtaining the priority of dynamic data when the dynamic branch is used as the main chain for each task type; and determining the target context from the dynamic data and static code context according to the priority of the main chain, auxiliary chain and dynamic data corresponding to the task type.
[0011] In the above implementation process, by determining the corresponding main chain and auxiliary chains based on the task type, different types of development tasks can automatically match the most suitable context source for their characteristics. The main chain prioritizes acquiring the most relevant information, while the auxiliary chain supplements it when the main chain is insufficient, improving the relevance of the target context to the current task. When a dynamic branch serves as the main chain, by setting the priority order of dynamic data for each task type, the system can acquire dynamic data in order of its importance to the current task, avoiding wasting time on acquiring secondary data, thereby improving the targeting and efficiency of context acquisition. Ultimately, this allows the target context to convey more effective information with less data, reducing the interference of irrelevant information on the inference of the coding model.
[0012] Optionally, in this embodiment of the application, receiving a user's reasoning request and determining the task type corresponding to the reasoning request based on the reasoning request includes receiving a natural language description corresponding to the reasoning request input by the user, using keyword rules or natural language intent recognition, and determining the task type based on the natural language description; the task type includes at least one of: adding new features, fixing defects, optimizing performance, and user interface issues.
[0013] In the above implementation process, the natural language descriptions input by users are classified through keyword rules or natural language intent recognition. This automatically identifies whether the user's task type is adding new features, fixing defects, optimizing performance, or addressing user interface issues, eliminating the need for users to manually specify the task type and simplifying the user operation process. Keyword rule matching is simple to implement and has a fast response time, quickly handling most common expressions; natural language intent recognition has stronger semantic understanding capabilities, handling more diverse and colloquial user input. The two methods can be used individually or in combination, flexibly adapting to different application scenarios.
[0014] Optionally, in this embodiment of the application, after receiving the natural language description corresponding to the inference request input by the user, the method further includes: if the inference request is a multi-turn dialogue, correcting at least one of the following based on the context data access information obtained by the encoding model in subsequent rounds: the priority relationship of the task type, the priority order of the information source branches, and the priority order of the dynamic data; wherein, the access information includes the access information of the encoding model to other context data during the inference process after obtaining the target context.
[0015] In the above implementation process, in multi-turn dialogue scenarios, the invocation of other contextual data by the encoding model during inference serves as a feedback signal, enabling the identification and timely correction of deficiencies in the initial configuration. When information is insufficient, the encoding model proactively requests additional data. These requests reflect the actual data requirements of the current task. Adjusting the priority relationships of task types, information source branches, or dynamic data based on these requirements allows for more accurate context provision in subsequent turns. As the number of dialogue turns increases and the context provisioning strategy is continuously optimized, the system can increasingly accurately understand user intent and provide efficient contextual support, improving the model inference efficiency in multi-turn interaction scenarios.
[0016] Optionally, in this embodiment, the dynamic data includes pop-up lifecycle data; the method for collecting pop-up lifecycle data includes: registering pop-up lifecycle callback hooks; for dialog fragment type pop-ups, listening to the display and close events of dialog fragments through the lifecycle callback of the fragment manager; for traditional non-fragment type dialogs, wrapping the display and close methods of traditional dialogs through the delegate pattern to listen to display and close events; determining pop-up lifecycle data; the pop-up lifecycle data is used as pop-up state information in the target context when the task type is a user interface problem.
[0017] In the above implementation process, by registering pop-up lifecycle callback hooks, differentiated listening schemes are adopted for different types of pop-ups: pop-ups of dialog fragment classes utilize the lifecycle callbacks of the fragment manager provided by the operating system for listening, while traditional dialogs use a delegate pattern to wrap their display and close methods for listening. This collected pop-up lifecycle data is incorporated into the target context as pop-up state information when the task type is a user interface problem. During model inference, the coding model can directly obtain information such as the class name, duration, and parent container of all active pop-ups in the current terminal application, reducing manual operations and improving the efficiency of handling user interface problems.
[0018] Optionally, in this embodiment, the target context is input into the encoding model, and the encoding model performs model reasoning based on the target context, including: after the pop-up display event is triggered, capturing the call stack of the current thread and filtering the call stack, removing the stack frames of the framework layer and the third-party library layer to obtain the filtered call stack; the filtered call stack includes the call stack frames of the business code layer; storing the association information of the filtered call stack and the pop-up in a circular buffer; the association information of the pop-up includes the pop-up's identification information, timestamp, and parent container; if the target context contains pop-up state information, querying the pop-up record from the circular buffer, if there is an abnormal pop-up, locating the business method that triggered the pop-up from the filtered call stack, and the encoding model generating repair code based on the business method.
[0019] In the above implementation, by capturing and filtering the complete call stack of the current thread after the pop-up display event is triggered, and retaining the call stack frames of the business code layer, the source path of the pop-up trigger is made clearer and simpler, reducing the interference of a large number of stack frames on subsequent location. The filtered call stack, along with the pop-up's identification information, timestamp, and parent container, is stored in a circular buffer, enabling the system to persistently record the most recently occurring pop-up events and their triggering sources. When the target context contains pop-up state information, querying the pop-up record from the circular buffer can quickly identify abnormal pop-ups, and the filtered call stack can accurately locate the business method that triggered the pop-up. The coding model can directly analyze and repair this business method, shortening the location time and improving the efficiency of code repair.
[0020] Optionally, in this embodiment of the application, the target context is input into the encoding model, and the encoding model performs model inference based on the target context. This includes: inputting the target context into the encoding model; if the capacity of the context window of the encoding model is insufficient to accommodate all the target contexts, then the target contexts are input into the context window of the encoding model in descending order of priority of dynamic data, until the remaining capacity of the context window is insufficient to accommodate the input of context data of the next priority, and the encoding model performs model inference based on the input target contexts.
[0021] In the above implementation process, considering the limited capacity of the encoding model's context window, when the total amount of target context data exceeds the window capacity, dynamic data is input sequentially in descending order of priority. This ensures that the highest priority data enters the encoding model's context window first, thus preserving more information most relevant to the current task. The encoding model performs inference based on the input high-priority context data, reducing inference failures caused by exceeding the window capacity with a single input, and also reducing information loss due to data truncation. This improves the localization success rate and generation quality of model inference under limited context window conditions.
[0022] Secondly, embodiments of this application also provide a model inference apparatus, comprising: a dynamic data acquisition module for acquiring dynamic data during the runtime of a terminal application; a request receiving module for receiving a user's inference request and determining the task type corresponding to the inference request based on the inference request; a context determination module for determining a target context from the dynamic data and static code context based on the task type and the priority of the data corresponding to the task type; and an inference module for inputting the target context into an encoding model, and having the encoding model perform model inference based on the target context.
[0023] Thirdly, embodiments of this application also provide a computer program product, including computer program instructions, which are executed by a processor to perform the method provided in the first aspect or any implementation thereof.
[0024] Fourthly, embodiments of this application also provide an electronic device, including: a processor and a memory, the memory storing computer program instructions, which are executed by the processor to perform the method provided in the first aspect or any implementation thereof.
[0025] Fifthly, embodiments of this application also provide a computer-readable storage medium storing computer program instructions, which, when executed by a processor, perform the method provided in the first aspect or any implementation thereof.
[0026] This application provides a model inference method, program product, electronic device, and storage medium. By collecting dynamic data during the runtime of a terminal application and determining the target context from the dynamic data and static code context, this target context is input into the coding model. This allows the coding model to utilize both static and dynamic data during inference, transforming defect localization from a speculation-driven approach based on static code to an evidence-driven approach based on runtime data, thereby improving the accuracy of model inference. The task type is determined based on the inference request, and the target context is determined from the dynamic data and static code context according to the data priority corresponding to the task type. This makes the information input into the model more targeted, reduces interference from irrelevant information, and reduces the time and resources spent on exploration during model inference. Inputting the priority-selected target context into the coding model for model inference improves the efficiency and success rate of model inference. Attached Figure Description
[0027] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments of this application will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0028] Figure 1 A flowchart illustrating a model reasoning method provided in an embodiment of this application; Figure 2 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0029] The embodiments of the technical solution of this application will now be described in detail with reference to the accompanying drawings. These embodiments are only used to more clearly illustrate the technical solution of this application and are therefore merely examples, and should not be used to limit the scope of protection of this application.
[0030] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs; the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit this application.
[0031] In the description of the embodiments of this application, technical terms such as "first" and "second" are used only to distinguish different objects and should not be construed as indicating or implying relative importance or implicitly specifying the number, specific order, or primary and secondary relationship of the indicated technical features. In the description of the embodiments of this application, "multiple" means two or more, unless otherwise explicitly defined.
[0032] Currently, coding models typically assist developers in completing code writing by reading static data such as source code, Git history, and project documentation. Static data can reflect fixed information about the software during the design and coding phases.
[0033] However, in actual development, code defect fixing and feature debugging consume a significant amount of developers' time. This is especially true in terminal application development scenarios, where locating and fixing many defects heavily relies on dynamic information generated during application runtime. This includes network requests and server responses, the current page's navigation status, whether user interface components are displaying correctly, locally cached data, and the hardware and system information of the terminal device. Existing coding models are completely unable to access this dynamic runtime data and can only infer from static data such as source code, resulting in low accuracy and efficiency. This necessitates developers spending considerable time repeatedly performing manual debugging and verification.
[0034] This application provides a model inference method that collects dynamic data from the runtime of a terminal application and determines the target context from the dynamic data and static code context. This target context is then input into the coding model, enabling the coding model to utilize both static and dynamic data during the inference process. Defect localization shifts from a speculation-driven approach based on static code to an evidence-driven approach based on runtime data, thereby improving the accuracy of model inference. The method determines the task type based on the inference request and identifies the target context from the dynamic data and static code context according to the data priority corresponding to the task type. This makes the information input into the model more targeted, reduces interference from irrelevant information, and decreases the time and resources spent on exploration during model inference. Inputting the priority-selected target context into the coding model for model inference improves the efficiency and success rate of model inference.
[0035] Please see Figure 1 The illustrated diagram shows a flowchart of a model inference method provided in an embodiment of this application. The model inference method provided in this application can be applied to electronic devices, which may include physical devices such as servers, PCs, tablets, or smartphones, or virtual devices such as virtual machines or containers. The electronic device can be a single device, a combination of multiple devices, or a cluster of a large number of devices. The model inference method may include: Step S110: Collect dynamic data during the runtime of the terminal application.
[0036] Step S120: Receive the user's inference request and determine the task type corresponding to the inference request based on the inference request.
[0037] Step S130: Determine the corresponding main chain and auxiliary chain according to the task type, and determine the target context from dynamic data and static code context based on the main chain and auxiliary chain; wherein, the main chain represents the information source branch that is preferentially selected into the target context, and the auxiliary chain represents the information source branch that is supplemented into the target context; the information source branch includes dynamic branch and static branch, the dynamic branch points to dynamic data, and the static branch points to static code context.
[0038] Step S140: Input the target context into the encoding model, and the encoding model performs model inference based on the target context.
[0039] In step S110, the terminal application can be any program running on a terminal device, such as a mobile application, desktop application, or embedded system application. Dynamic data refers to the state information of the terminal application that changes over time during its operation. This information is not statically stored in code or documentation beforehand, but is determined in real time by factors such as the application's actual execution environment, user operations, and network interactions. Dynamic data can include various types of data such as communication content between the application and the server, the currently displayed page or interface hierarchy, the attribute status of each visual component in the interface, data persistently stored locally by the application, hardware and system information of the terminal device, and system operation logs.
[0040] There are several ways to collect dynamic data: one is to embed dedicated collection logic into the source code of the terminal application, that is, to insert data logging points on critical execution paths, and automatically record relevant data when the application reaches that location. Another way is to obtain the application's runtime information through the debugging or performance monitoring interfaces provided by the terminal operating system. External tools can also be used to capture and analyze the application's runtime memory state, network packets, etc. The collected data can be cached in the terminal application's local storage or uploaded to a remote server over the network. The collection process can be continuous or triggered as needed.
[0041] In step S120, the user's inference request refers to the problem description or task instruction entered by the developer when using the coding model to assist in development, repair, and debugging. It is usually presented in natural language text, such as "Fix login timeout issue" or "Optimize homepage loading speed." After receiving the inference request, the request content needs to be analyzed to determine its corresponding task type. The task type characterizes the type of development, repair, and debugging work that the user needs to complete. Its classification dimensions can be flexibly defined according to the actual application scenario. For example, task types may include: fixing errors or exceptions based on existing code (i.e., defect repair), adding unimplemented features (i.e., feature addition), improving code execution efficiency or resource consumption (i.e., performance optimization), adjusting user interface display or interaction effects (i.e., interface adjustment), etc. More sub-types can be extended in different scenarios, such as security hardening, code refactoring, and compatibility adaptation.
[0042] Determining task types can be based on pre-defined rules, such as building a keyword lexicon to map specific words in user descriptions to corresponding task types; alternatively, a trained machine learning classification model can be used to automatically identify the semantics of user input; or the user's description can be matched with historical requests for similarity, and the type of the historical requests can be used for inference. When the conclusions of multiple determination methods are inconsistent, a fusion decision can be made according to preset priorities or confidence levels. If the type still cannot be determined, it can be set as a default type or the user can be prompted for confirmation. Through these methods, the user's natural language request can be transformed into a structured task type label, providing a basis for subsequent targeted selection of context.
[0043] In step S130, the static code context refers to non-runtime information sources related to the current development task, such as code files, project documents, and version commit records. The target context refers to the information ultimately selected and prepared for input into the coding model. Since the dynamic data collected in step S110 covers multiple categories, and the static code context may also contain a large number of code files, and the coding model has limited processing capacity, it is necessary to filter out the most relevant content. The filtering is based on the task type and the priority of the data corresponding to each task type. Data priority is used to measure the importance of different categories of data for solving the task under a specific task type; the higher the importance, the higher the priority, and the more likely it should be selected into the target context.
[0044] When determining the target context, the corresponding main chain and auxiliary chains are first determined based on the task type. Information source branches include dynamic branches and static branches, where dynamic branches point to the dynamic data collected in step S110, and static branches point to the static code context. The main chain represents the information source branch that is preferentially selected into the target context, i.e., the data source most relied upon by the current task; the auxiliary chain represents the information source branch that is supplemented into the target context when the main chain information is insufficient to support model inference. Different task types may correspond to different main chains and auxiliary chains. For example, a defect repair task can set the dynamic branch as the main chain and the static branch as the auxiliary chain, while a new feature task can set the static branch as the main chain and the dynamic branch as the auxiliary chain.
[0045] After determining the main chain and auxiliary chains, data is selected from dynamic data and static code contexts in the order of main chain priority and auxiliary chain supplementation. If the main chain is a dynamic branch, dynamic data of the corresponding category is retrieved from the data processor in descending order of priority according to the task type and selected into the target context. If the main chain is a static branch, data is directly selected from the static code context. After the main chain data has been selected into the target context, if it is determined that the selected data is insufficient to support model inference, the data pointed to by the auxiliary chains is then added into the target context to supplement the selection.
[0046] There are several ways to determine priorities: One approach is to set a static priority order for each task type. For example, for defect repair tasks, network communication data can be prioritized over UI status data, and UI status data can be prioritized over cached data. Each time a request is processed, data is selected sequentially according to this order until the preset conditions are met. Another approach is to dynamically calculate priorities. For example, priorities can be sorted in real-time based on the frequency or weight of keywords mentioned in the user description across different data categories, with data categories appearing more frequently in the description receiving higher priority. A hybrid approach can also be used, where a basic priority order is first determined based on task type, and then fuzzy matching is performed across different data categories based on specific entities in the user description, such as an interface name or a page name. Data categories with higher matching scores are given an improvement in the basic priority order.
[0047] After prioritizing, data is selected from both dynamic and static code contexts in descending order of priority, and combined to form the target context. The selection can stop when a preset limit is reached for the number of data entries, a preset limit for the number of tokens, or when the currently selected data can constitute a complete problem context. This method of determining the target context reflects the differences in task types and flexibly adapts to the specific content of user requests, improving the problem of diluting or even overwhelming effective information caused by indiscriminately inputting all data into the encoding model.
[0048] To illustrate the above process, consider a specific scenario: A user inputs "Login interface error," and the task classifier identifies it as a "defect repair" task based on the keyword "error." The system determines the dynamic branch as the main chain and the static branch as the auxiliary chain. When determining the target context, it prioritizes querying network communication data (filter condition filter="login") according to the priority of the defect repair task. After obtaining the HTTP 422 error and information indicating missing fields in the response body, the system locates the corresponding login interface source code from the auxiliary chain (static branch). In contrast, if a fixed strategy (without priority distinction) is used, the encoding model will first read all login-related source code (e.g., 2000 lines of source code) and then attempt to obtain runtime clues. Using this solution, the encoding model first obtains runtime data (e.g., 20 lines of structured results), directly locating the specific error, significantly improving the localization efficiency of the encoding model during inference.
[0049] In step S140, the encoding model refers to an artificial intelligence model capable of understanding and performing model inference, typically a large language model trained on a large-scale code corpus. After determining the target context in step S130, the context needs to be organized according to the input format required by the encoding model. Different encoding models may have different input format requirements; for example, some models require the input content to be wrapped in specific prompt word templates, while others require different types of context to be placed in different input segments. The data in the target context can be formatted according to the interface specifications of the adopted encoding model. If the total amount of data in the target context exceeds the upper limit of the encoding model's single processing capacity, a truncation strategy can be adopted, prioritizing the retention of high-priority data and discarding or compressing low-priority data; alternatively, a block strategy can be adopted, splitting the target context into multiple batches and inputting them into the model sequentially, then combining the outputs from each batch. After the formatted target context is input into the encoding model, the model performs inference based on this context.
[0050] During inference, the coding model combines the actual runtime state reflected in dynamic data with the program structure in the static code context to understand the current problem and then outputs model inference data. The output data can be descriptive content for locating defects, such as marking the function and line number where the problem occurs, describing the phenomenon and analyzing the cause of the problem; the output data can also be directly executable code snippets for fixing the issue, optimized and improved versions of the code, or layout descriptions of the interface adjustments, etc., depending on the task type and the content requested by the user. The model inference output data can be directly displayed to the developers or automatically saved or applied by development tools. During inference, the coding model no longer relies solely on static code for speculation but also uses the actual runtime state of the terminal application as the basis for inference, thereby improving the accuracy of model inference.
[0051] In the implementation of the above embodiments: By collecting dynamic data during the runtime of the terminal application and determining the target context from the dynamic data and static code context, this target context is input into the coding model. This allows the coding model to utilize both static and dynamic data during inference, transforming defect localization from a speculation-driven approach based on static code to an evidence-driven approach based on runtime data, thereby improving the accuracy of model inference. The task type is determined according to the inference request, and the target context is determined from the dynamic data and static code context according to the data priority corresponding to the task type. This makes the information input into the model more targeted, reduces interference from irrelevant information, and also reduces the time and resources spent on exploration during model inference. Inputting the target context selected according to priority into the coding model for model inference improves the efficiency and success rate of model inference.
[0052] Optionally, in this embodiment, collecting dynamic data during the runtime of the terminal application includes: collecting dynamic data through a runtime data collection module deployed on the terminal application, wherein the runtime data collection module includes multiple data processors divided according to data categories, the multiple data processors independently collect and cache various types of dynamic data during the runtime of the terminal application, and provide a data query interface within the terminal application; the dynamic data includes at least one of the following: network communication data, page navigation status data, pop-up lifecycle data, view inspection data, key-value pair storage data, and event tracking data.
[0053] The data acquisition module consists of multiple data processors, each corresponding to a specific data category. For example, the network communication processor is responsible for recording network requests and responses sent by the application, the page navigation processor tracks the user's navigation paths between different interfaces, and the pop-up lifecycle processor monitors the pop-up and closing of various pop-ups. Each data processor is logically independent; that is, each processor is only responsible for acquiring data of its own category without interfering with others. This independent design reduces coupling between processors, ensuring that adding new data categories or modifying the acquisition logic of one processor will not affect the normal operation of other processors.
[0054] After collecting data, each processor caches the data in the terminal application's memory cache or local storage file, reducing the need for re-collection on each query and improving response speed. Furthermore, the data acquisition module internally launches a data query interface service exposed to the outside world to receive data query requests. When an external system sends a query request through this interface, the data acquisition module reads data from the corresponding data processor cache based on the data category specified in the request and returns it. Without interrupting the normal operation of the terminal application, external systems can obtain the latest running status of the terminal application as needed.
[0055] The specific methods for collecting the above-mentioned types of dynamic data are as follows: For network communication data, the data processor can intercept network requests from terminal applications, record the request URL, request header, and request body when the request is sent, record the response status code, response header, and response body when the response is returned, and can also record the time consumed for each request.
[0056] For page navigation state data, the data processor can listen to the page switching events of the terminal application's route manager and record the current page identifier, source page, entry time, exit time, and parameters passed between pages.
[0057] For pop-up lifecycle data, the data processor can use various methods to listen: for example, for pop-ups based on Fragments, lifecycle callbacks can be registered through FragmentManager to listen for their display and close events; for traditional Dialogs, their show and dismiss methods can be wrapped using the proxy pattern, and the recording logic can be triggered when the methods are called, thereby capturing the display and close times of all pop-ups.
[0058] For view inspection data, the data processor can traverse the view tree structure of the current window of the terminal application to obtain information such as the type, position coordinates, size, visibility status, text content, and enabled / disabled status of each view control. For key-value pair storage data, the data processor can proxy the local storage interface of the terminal application to record the key name, key value, and write time each time a key-value pair is written.
[0059] For event tracking data, the data processor can register event listeners with the terminal application to capture user interaction events such as clicks, swipes, and inputs, and record the event type, trigger location, trigger time, and target control information. The data processor can operate asynchronously, meaning it does not block the main thread of the terminal application while collecting data, thus ensuring the normal and smooth operation of the application.
[0060] In the implementation of the above embodiments: by deploying a running data acquisition module containing multiple independent data processors, different types of dynamic data can be acquired and cached separately, with each processor operating independently without interference. This reduces the coupling within the module and facilitates the subsequent expansion of new data categories. Each processor caches the acquired data within the terminal application, reducing the performance overhead caused by repeated acquisition. By providing a unified data query interface within the terminal application, external systems can obtain the latest running status during the normal operation of the terminal application.
[0061] Optionally, in this embodiment of the application, after collecting dynamic data during the runtime of the terminal application, the method further includes: obtaining dynamic data from the data processor through a protocol translation intermediate layer; the protocol translation intermediate layer establishes communication connections with the encoding model and the terminal application respectively, and the protocol translation intermediate layer is used to convert the tool call request of the encoding model into a message format recognizable by the terminal application, and convert the dynamic data returned by the terminal application into structured text corresponding to the encoding model.
[0062] The role of the protocol translation middleware is to establish a bidirectional communication bridge between the encoding model and the terminal application, which can be implemented through a lightweight service process. For example, the protocol translation middleware connects to the encoding model via an inter-process communication protocol (such as stdio), connects to the mobile terminal application via the WebSocket protocol, and obtains system-level data through a system-level debugging channel (such as ADB, Android Debug Bridge), forming a complementary data acquisition architecture with WebSocket and ADB dual channels. Specifically, the WebSocket protocol is used to obtain structured data within the mobile terminal application, such as network request content, page navigation status, and pop-up lifecycle data; the system-level debugging channel is used to obtain system-level data external to the terminal application, such as screenshots, logs, and device information. The two channels have distinct functions and complement each other, jointly covering all runtime information requirements. Structured data includes network request content, page navigation status, and pop-up lifecycle data.
[0063] The protocol translation middleware layer converts tool call requests from the encoding model into message formats recognizable by mobile terminal applications and distributes these messages to corresponding data processors using an action prefix routing mechanism. This mechanism uses the prefix of the action field in the message as the routing key, directly mapping the message to the corresponding data processor through string prefix matching, without requiring additional deserialization or content parsing of the message body. Compared to common routing methods such as message type fields, reflection / annotations, conditional branching, or content matching, the action prefix routing mechanism offers the following advantages: zero parsing overhead; decoupling from the underlying transport protocol, naturally supporting streaming data processing; dynamic addition or removal of data processors at runtime without system downtime; natural support for wildcard and layered action extensions; language and platform independence, suitable for message distribution between heterogeneous systems. It also converts dynamic data returned by mobile terminal applications into structured text corresponding to the encoding model.
[0064] Tool invocation requests from the encoding model follow standardized protocol formats, such as the MCP protocol (Model Context Protocol). These standardized protocols define specifications for the encoding model to call external tools and obtain additional context through a unified interface. Upon receiving this standardized request, the protocol translation middleware parses and converts it into a message format recognizable by the end application. For example, it converts it into a message object containing action and parameter fields, in the format {action:"<prefix>_<operation>",...parameters}. After conversion, the protocol translation middleware uses an action prefix routing mechanism to distribute the message to the corresponding data processor. This routing mechanism uses the prefix portion of the action field in the message as the routing key. For example, prefixes like createUser: and updateOrder: can directly map the message to the corresponding data processor, completing the routing without additional parsing of the message content.
[0065] This routing method boasts advantages such as zero parsing overhead and extremely high routing speed, while being decoupled from the underlying transport protocol and allowing for dynamic addition or removal of processors without downtime. Upon receiving a message, the corresponding data processor in the terminal application reads the relevant dynamic data from the cache based on the operations and parameters in the message and returns the data to the protocol translation middleware layer. After receiving the dynamic data returned by the terminal application, the protocol translation middleware layer converts this data into a structured text format that the encoding model can consume, and then passes it to the encoding model as input for subsequent inference. Through this method, cross-protocol and cross-process data interaction between the encoding model and the terminal application is achieved, enabling the encoding model to obtain dynamic runtime data from the terminal application in real time.
[0066] In the implementation of the above embodiments: a standardized data interaction channel is established between the encoding model and the terminal application by deploying a protocol translation middleware layer. The protocol translation middleware layer uses an inter-process communication protocol to connect the encoding model, allowing the encoding model to initiate data acquisition requests without needing to know the communication details of the terminal application. The protocol translation middleware layer converts the standardized tool call requests of the encoding model into message formats recognizable by the terminal application, enabling communication between different protocol systems. The protocol translation middleware layer converts the dynamic data returned by the terminal application into structured text that the encoding model can consume.
[0067] Optionally, in this embodiment, the target context is determined from dynamic data and static code context based on the task type and the priority of the data corresponding to the task type, including: Based on the pre-established mapping relationship between task types and information source branches, the main chain and auxiliary chain corresponding to the task type are determined; where the main chain represents the information source branch that is preferentially selected into the target context, and the auxiliary chain represents the information source branch that is supplementarily selected into the target context.
[0068] Information source branches refer to the source channels of the target context, used to determine which type of data source the target context input to the encoding model comes from. Pre-establishing a mapping relationship between task types and information source branches means specifying the corresponding main chain and auxiliary chains for each task type through configuration or coding. This mapping relationship can be stored in configuration files, database tables, or memory-mapped tables and loaded at system startup. The main chain represents the information source branch that is preferentially selected when assembling the target context, i.e., the data source that the encoding model relies on more; the auxiliary chain represents the information source branch that is supplemented when the main chain information is insufficient to support model inference. The main chain and auxiliary chain together constitute the complete context supply strategy for this task type.
[0069] The main and secondary chains can be determined as follows: After receiving a task type tag, the system uses that tag as the key to query a pre-established mapping table and directly reads the corresponding main and secondary chain identifiers. For example, in the mapping table, the main chain for defect repair is a dynamic branch, and the secondary chain is a static branch. This method allows for a quick and accurate determination of the priority order for context selection in the current task, avoiding the need for re-decision-making each time a request is processed.
[0070] The information source branches include dynamic branches and static branches. Dynamic branches point to dynamic data, while static branches point to static code context.
[0071] Dynamic branches are the information source channels for collected dynamic data; they are paths to obtain context through runtime data such as network communication data, page navigation state data, pop-up lifecycle data, view inspection data, key-value pair storage data, and event tracking data. Static branches are the information source channels pointing to the static code context, which includes, but is not limited to, the application's source code files, project documentation, commit history in version control systems, code comments, and dependency library documentation. The context provided by dynamic branches reflects the real-time state and behavior of the application during runtime, while the context provided by static branches reflects the structural information of the application during its design and development process.
[0072] In practical implementation, the system can distinguish between these two branches through different data acquisition interfaces. Data for the dynamic branch is acquired in real-time or near real-time from the data processor of the terminal application through a protocol translation middleware layer; data for the static branch is acquired from local or remote code repositories through methods such as file system reading, code parsing tool extraction, or version control system API calls. By explicitly defining these two branches in the system design and specifying an independent data acquisition path for each branch, the context can be flexibly selected from one or both branches according to the task type.
[0073] For each task type, prioritize dynamic data when retrieving dynamic branches as the main chain.
[0074] Priority refers to the order in which different categories of dynamic data are selected when a dynamic branch is determined as the main branch. Different task types have different requirements for different categories of dynamic data, so priorities need to be configured for each task type. For example, for defect repair tasks, network communication data usually best reflects the cause of the error, so it has the highest priority, followed by page navigation status data and cache data; for performance optimization tasks, system log data contains clues about performance bottlenecks, so it has the highest priority, followed by network latency data and device information data; for user interface issues, layout check data directly reflects the interface rendering status, so it has the highest priority, followed by screenshot data and pop-up status data.
[0075] The priority of dynamic data can be pre-stored in the system as a configuration table. Each row in the configuration table contains three fields: task type, dynamic data category, and priority level. When the system determines that the main chain corresponding to a certain task type is a dynamic branch, it uses that task type as the query condition to read the corresponding list of dynamic data categories from the configuration table and sorts them from highest to lowest priority. If a dynamic data category cannot be retrieved or retrieval fails, the system can automatically skip that category and continue to retrieve data of the next higher priority.
[0076] The target context is determined from the dynamic data and static code context based on the priority of the main chain, auxiliary chain, and dynamic data corresponding to the task type.
[0077] Determining the target context refers to the process of selecting data in an orderly manner from dynamic data sources and static code context sources, based on the main chain, auxiliary chains, and priority information determined in previous steps, and assembling them into the information set for the final input encoding model. For example, the system executes the selection process in the following order: First, the main chain branch is determined according to the task type. If the main chain is a dynamic branch, the dynamic data corresponding to that task type is retrieved in order of priority, from high priority to low priority, and the corresponding category of dynamic data is queried from the data processor and selected into the target context. If the main chain is a static branch, the static code context is read from the code repository or file system and selected into the target context. Second, it is determined whether the amount of data or the completeness of information selected into the target context meets preset conditions. These preset conditions may be whether the total number of tokens in the target context has reached a preset threshold of the encoding model context window, or whether the selected data covers key entities in the problem description, such as interface names, page names, and error codes. If the conditions are met, the selection process stops. If the conditions are not met, the selection of auxiliary chains is initiated: data corresponding to the auxiliary chain branches are selected into the target context, and the order of supplementation also follows the priority rules of the branch itself: if the auxiliary chain is a dynamic branch, it is selected according to the priority order of dynamic data; if the auxiliary chain is a static branch, it is selected from the static code context. Through the above ordered selection process, the final target context highlights the information most relevant to the current task while retaining the ability to supplement when core information is insufficient. In addition, during multi-round dialogue, if the main chain information still fails to support the solution of the problem after multiple rounds of reasoning, the system automatically upgrades the priority of the auxiliary chain, realizing the dynamic switching between the main and auxiliary chains. That is, the original auxiliary chain becomes the new main chain, and the original main chain is demoted to an auxiliary chain. Subsequent rounds reassemble the target context according to the new main and auxiliary chain relationship.
[0078] In the implementation of the above embodiments: by determining the corresponding main chain and auxiliary chain according to the task type, different types of development tasks can automatically match the most suitable context source for their characteristics. The main chain prioritizes acquiring the most relevant information, and the auxiliary chain provides supplementary information when the main chain is insufficient, thereby improving the relevance of the target context to the current task. When a dynamic branch serves as the main chain, by setting the priority order of dynamic data for each task type, the system can acquire various types of dynamic data in order of their importance to the current task, avoiding wasting time on acquiring secondary data, thus improving the targeting and efficiency of context acquisition. Ultimately, this allows the target context to convey more effective information with less data, reducing the interference of irrelevant information on the inference of the coding model.
[0079] Optionally, in this embodiment of the application, receiving a user's inference request and determining the task type corresponding to the inference request based on the inference request includes: The system receives a natural language description corresponding to a user's input inference request, uses keyword rules or natural language intent recognition to determine the task type based on the natural language description, and the task type includes at least one of the following: adding new features, fixing defects, optimizing performance, and addressing user interface issues.
[0080] Inference requests refer to the problem descriptions or task instructions that developers submit to the coding model in natural language when they encounter specific problems or need to complete specific development tasks during the coding process.
[0081] One approach to determining task types is based on keyword rule matching. The system pre-builds a corresponding keyword lexicon for each task type. For example, defect repair tasks might include words like "error," "crash," "exception," "unable," and "failure"; performance optimization tasks might include words like "lag," "slow," "power consumption," and "high memory usage"; user interface issues might include words like "not displaying," "obstructed," "misaligned," and "unclickable"; and adding new features might include words like "add," "implement," and "increase." Upon receiving a natural language description from the user, the system calculates the number of matched words or a weighted score for each task type, and identifies the task type with the highest score as the task type corresponding to the current request. When multiple task types have similar scores or none of them match, meaning the task category cannot be accurately determined based on the natural language description input by the user, the system can default to determining the task type as defect repair. This is because defect repair is a common task scenario in software development practice and has the strongest dependence on runtime dynamic data. If other types are used as the default value (such as "adding a new feature"), the system may prioritize obtaining the static code context and ignore the key runtime data, thus missing the best opportunity to locate the root cause of the defect. The "defect repair" type uses dynamic branches as the main chain, which enables the coding model to obtain runtime data in the initial stage. Even if the task type is corrected later through dynamic adjustment strategies, the initially obtained runtime data can still be used as an auxiliary reference.
[0082] Another approach to determining task types is based on a natural language intent recognition model. First, a large number of user request samples labeled with task types are pre-collected, and these samples are used to train a text classification model. During training, the model learns the mapping relationship between natural language text and task type labels. In actual use, the system inputs the user's natural language description into the model, and the model outputs the probability distribution of the description belonging to various task types. The system selects the task type with the highest probability as the determination result.
[0083] In the implementation of the above embodiments: by classifying the natural language descriptions input by users through keyword rules or natural language intent recognition, the system can automatically identify whether the user's task type is to add new features, fix defects, optimize performance, or address user interface issues, without requiring the user to manually specify the task type, thus simplifying the user operation process. Keyword rule matching is simple to implement and has a fast response time, capable of quickly processing most common expressions; natural language intent recognition has stronger semantic understanding capabilities, able to handle more diverse and colloquial user input. Both methods can be used individually or in combination, flexibly adapting to different application scenarios.
[0084] Optionally, in this embodiment of the application, after receiving the natural language description corresponding to the inference request input by the user, the method further includes: If the inference request is a multi-turn dialogue, based on the context data retrieval status of the encoding model in subsequent rounds, at least one of the following should be corrected: the priority relationship of the task type, the priority relationship of the information source branches, and the priority order of dynamic data. The retrieval status includes the inference status of the encoding model for other context data after obtaining the target context.
[0085] Multi-turn dialogue refers to a process of multiple interactions between a user and a coding model. After the user asks a question initially, the coding model provides an initial response. The user then asks follow-up questions, provides additional information, or provides feedback, leading to further responses from the coding model, and so on, in a series of dialogues. In a single-turn dialogue, the system determines the task type and priority based solely on the user's initial description and executes it. However, in a multi-turn dialogue scenario, as the dialogue progresses, more information about the current problem can be obtained and adjustments can be made.
[0086] The subsequent calls for retrieving context data refer to the record of requests initiated by the encoding model to retrieve other context data after it has obtained the determined target context and begun inference, in order to complete the inference task. During inference, the encoding model evaluates whether the currently acquired information is sufficient to support its model inference. If it finds that the information is insufficient, it will call interfaces through tools such as the MCP protocol to request additional data. The data categories, frequency, and timing of these requests together constitute the call details.
[0087] During implementation, every tool call request from the encoding model can be continuously monitored and recorded. When a multi-turn dialogue is detected as an inference request, the list of context data categories requested by the encoding model in the most recent round or multiple rounds of inference is extracted from the monitoring records. If the system finds that the encoding model frequently requests a dynamic data category that is lower in the current priority order, or requests a data category under the current task type that is not included in the main chain, it indicates that the current priority order or main chain settings may need to be adjusted and cannot fully meet the actual needs of model inference.
[0088] The configuration is modified based on these call patterns. The modifications target at least one of the following: task type, priority relationships of information source branches, and priority order of dynamic data. The modification can be automated according to preset rules, such as counting the number of requests for various data types and rearranging their priority order from highest to lowest request frequency; or it can be semi-automatic, where the system generates modification suggestions for user confirmation before application. For example, if the task type classification initially categorizes the user description "not displayed" as "defect repair," but the coding model still requires more "layout check, screenshot, and pop-up status" data after repeatedly obtaining high-priority dynamic branch data (network communication, navigation status, cache) during subsequent inference, the system determines that the actual needs of the current task better match the data pattern of the "UI problem" type. In subsequent classifications, "not displayed" is changed from "defect repair" to "UI problem," and the main and auxiliary chains and data acquisition priorities are adjusted accordingly. After the modification is complete, the new configuration will take effect in subsequent rounds. Through this dynamic modification mechanism, the system can gradually optimize the context supply strategy to better suit the actual needs of the current problem.
[0089] In the implementation of the above embodiments: In multi-turn dialogue scenarios, the invocation of other context data by the encoding model during inference serves as a feedback signal, enabling the identification and timely correction of deficiencies in the initial configuration. When information is insufficient, the encoding model proactively requests additional data. These requests reflect the actual data requirements of the current task. Adjusting the priority relationships of task types, information source branches, or dynamic data based on these requirements allows for more accurate context provision in subsequent rounds. As the number of dialogue rounds increases, the context provision strategy is continuously optimized, enabling the system to understand user intent more accurately and provide efficient context support, thereby improving the model inference efficiency in multi-turn interaction scenarios.
[0090] Optionally, in this embodiment, the dynamic data includes pop-up lifecycle data; the methods for collecting pop-up lifecycle data include: Register pop-up lifecycle callback hooks. For dialog fragment type pop-ups, listen to the display and close events of the dialog fragment through the lifecycle callback of the fragment manager. For traditional dialogs that are not fragment types, wrap the display and close methods of the traditional dialogs through the delegate pattern to listen to the display and close events; determine the pop-up lifecycle data.
[0091] A pop-up is a visual component that appears as an independent window on a terminal application interface, used to display prompts, obtain user input, or show loading progress, etc. Pop-ups in terminal applications can be divided into two categories: one is pop-ups implemented based on dialog fragments, and the other is traditional dialog boxes. Dialog fragments are standard components provided by mobile operating systems such as Android. The fragment manager is a component provided by the operating system to manage the lifecycle and transactions of all fragments.
[0092] In mobile application UI issues, pop-up occlusion is a common interaction problem. When a pop-up is not properly closed, it covers the top of the page, intercepting all touch events and preventing the user from interacting with the underlying UI. The difficulty in locating this type of problem lies in the fact that the display and closure of pop-ups are usually controlled by asynchronous callbacks (such as network response callbacks or animation end callbacks). Scenarios where closure is missed often occur in abnormal branches (error callbacks, timeout callbacks), rather than the normal execution path. Developers need to trace back "who popped up the pop-up and when," but the trigger sources for pop-ups are scattered across multiple places in the code, and the runtime call chain is not visible. Existing layout inspectors can only detect the existence of pop-ups but cannot provide information about the pop-up's trigger call stack. Without information about the pop-up's source, the coding model can only search through all `.show()` call points in the source code one by one, resulting in inefficient reasoning.
[0093] For dialog fragment pop-ups, by registering lifecycle callbacks with the FragmentManager, when any dialog fragment executes lifecycle events such as onStart (display) or onDismiss (close), the FragmentManager will automatically trigger the registered callback methods. Inside the callback methods, information such as the class name, display time, and close time of the pop-up can be obtained, thereby determining the lifecycle data of the pop-up type.
[0094] For traditional dialog boxes that are not fragment classes, the proxy pattern is used. The proxy pattern is a software design pattern that wraps the original object with a proxy object, inserting additional logic before or after calling the original object's methods. In practice, when the terminal application starts, the system replaces the `show` and `dismiss` methods of the `Dialog` with proxy methods through bytecode manipulation or runtime dynamic proxy. When the developer calls `Dialog.show()`, the proxy method is actually executed. The proxy method first records the current timestamp and the dialog box class name as a display event, and then calls the original `show` method to actually display the dialog box. Similarly, when `Dialog.dismiss()` is called, the proxy method first records the close time, and then calls the original `dismiss` method. Essentially, it adds pluggable listening capabilities to these two behaviors without modifying the original `Dialog` class. Through these two methods, the system can capture all dialog box display and close behaviors and determine the dialog box lifecycle data.
[0095] Pop-up lifecycle data is used as pop-up state information in the target context when the task type is a user interface issue.
[0096] Pop-up lifecycle data refers to the collected records of each pop-up, including the pop-up's class name, timestamp of the display event, timestamp of the close event, its parent container, and the parameters passed when the pop-up was displayed. Pop-up state information refers to the set of pop-ups currently displayed in the terminal application at a specific moment, and the duration of each pop-up from its initial display to the current moment.
[0097] When the current task type is determined to be a user interface issue, the problem described by the user is likely related to abnormal pop-up behavior, such as the pop-up obscuring the underlying interface and preventing user operation, the pop-up not disappearing for an extended period, or the pop-up displaying incorrect content. In these scenarios, the current state of the pop-up can pinpoint the root cause of the problem. Therefore, when determining the target context, if the task type is detected as a user interface issue, pop-up lifecycle data is included as pop-up state information in the target context. For example, the pop-up lifecycle data processor in the runtime data acquisition module queries the list of currently active pop-ups to obtain the class name, display timestamp, duration, and call stack information when the pop-up was triggered for each active pop-up. This information is then assembled into a structured pop-up state description text and added to the target context's dataset for use by the encoding model during subsequent inference.
[0098] In the implementation of the above embodiments: by registering pop-up lifecycle callback hooks, differentiated listening schemes are adopted for different types of pop-ups: pop-up fragments utilize the lifecycle callbacks of the fragment manager provided by the operating system for listening, while traditional dialogs use a delegate pattern to wrap their display and close methods for listening. This collected pop-up lifecycle data is incorporated into the target context as pop-up state information when the task type is a user interface problem. During model inference, the coding model can directly obtain information such as the class name, duration, and parent container of all active pop-ups in the current terminal application, reducing manual operations and improving the efficiency of handling user interface problems.
[0099] Optionally, in this embodiment of the application, the target context is input into the encoding model, and the encoding model performs model inference based on the target context, including: After the pop-up display event is triggered, the call stack of the current thread is captured and filtered to remove the stack frames of the framework layer and third-party library layer, thus obtaining the filtered call stack; the filtered call stack includes the call stack of the business code layer.
[0100] The call stack refers to the sequence of all method calls that have not yet been completed in the current thread at a certain point in program execution. Each method call corresponds to a stack frame in the call stack, which contains information such as the class name, method name, source code file name, and line number of the method. When a pop-up display event is triggered, for example, by capturing the display event through a listener mechanism, the complete call stack of the current thread can be obtained.
[0101] Because the call stack of a terminal application may contain a large number of method calls from the operating system framework layer and third-party dependency libraries, these stack frames are not helpful in locating the pop-up trigger point in the business logic code; instead, they increase the burden of subsequent processing. Therefore, it is necessary to filter the call stack and remove stack frames from the framework layer and third-party library layer. The framework layer refers to the core framework code provided by the operating system running the terminal application, such as classes with package name prefixes like android.* and androidx.* on the Android platform; the third-party library layer refers to the code of external dependency libraries referenced by the terminal application, such as classes with package name prefixes like com.google.* or other third-party SDK code. The filtering rule matches based on the application package name prefix: the system pre-configures the terminal application's own package name prefix (e.g., com.example.myapp), traverses each stack frame in the call stack, retains only stack frames whose class names begin with the application package name prefix, and removes all mismatched stack frames. Through this package name prefix-based filtering method, the dozens of stack frames that may be contained in the original call stack are compressed to contain only a few keyframes of business logic code, and the coding model can directly locate the trigger source of the pop-up from these keyframes.
[0102] The filtering can also be implemented by: pre-configuring a list of filtering rules, which records the package name prefixes that need to be removed; after capturing the complete call stack, the system iterates through each stack frame, checking if the class name of the stack frame begins with any prefix in the list; if so, the stack frame is removed; otherwise, it is retained. After the iteration is complete, the remaining stack frames that were not removed constitute the filtered call stack. These stack frames mainly come from the business code layer of the terminal application itself. The filtered call stack significantly reduces the number of stack frames, making the trigger source of the pop-up window clearer and more identifiable.
[0103] The filtered call stack and pop-up association information are stored in a circular buffer; the pop-up association information includes the pop-up's identifier, timestamp, and parent container.
[0104] A circular buffer is a fixed-size data storage structure. When the circular buffer is full, newly written data overwrites the oldest written data. Circular buffers are suitable for storing records generated within a recent period, ensuring access to the latest data while preventing memory overflow due to infinite data growth. The associated information of a pop-up includes at least one of the following: the pop-up's identifier, a timestamp, and its parent container. The pop-up's identifier uniquely distinguishes different pop-ups; for example, it can be generated using the pop-up's class name combined with its memory address or an auto-incrementing sequence number. The timestamp records the moment the pop-up's display event occurred, using the system clock to obtain a time value accurate to milliseconds. The parent container records the Activity or Fragment to which the pop-up is attached, determined by obtaining the container context when the pop-up is displayed.
[0105] After obtaining the filtered call stack and the associated information of the pop-up window, the system combines these two pieces of data into a single pop-up window record and writes it to a circular buffer. Each pop-up window record in the circular buffer fully stores the pop-up window's identifier, timestamp, parent container, and filtered call stack. The circular buffer can always retain the most recent pop-up window records for subsequent queries.
[0106] If the target context contains pop-up status information, the pop-up record is queried from the circular buffer. If an abnormal pop-up exists, the business method that triggered the pop-up is located from the filtered call stack, and the coding model generates repair code based on the business method.
[0107] When the target context contains pop-up status information, it is necessary to further determine whether there are any abnormal pop-ups. Abnormal pop-ups are those that remain open for an extended period after being displayed, such as loading prompt pop-ups that are not closed properly after data is returned, or error prompt pop-ups that are not released after user confirmation. The currently active pop-up records are read from the circular buffer, and the duration from the timestamp in each record to the current moment is calculated. If the duration of a pop-up exceeds a preset threshold, it is determined to be an abnormal pop-up. The threshold can be flexibly configured according to the pop-up type or business requirements.
[0108] For abnormal pop-ups, the system extracts a filtered call stack from the pop-up record. This filtered call stack contains the method call sequence that triggered the pop-up's display, with the topmost stack frame being the business method that directly triggered it. This business method is a function within the static code. After locating the business method, the system can locate the complete source code corresponding to that method within the static code context based on the class name, method name, and line number recorded in the stack frame of the filtered call stack. This source code is then input into the coding model. The coding model combines the source code of the business method with the pop-up state information to infer whether there are any logical vulnerabilities in the method and generates corrective code containing the correct closing logic.
[0109] In the implementation of the above embodiments: by capturing and filtering the complete call stack of the current thread after the pop-up display event is triggered, and retaining the call stack frames of the business code layer, the source path of the pop-up trigger is clearer and simpler, reducing the interference of a large number of stack frames on subsequent location. The filtered call stack, along with the pop-up's identification information, timestamp, and parent container, is stored in a circular buffer, enabling the system to persistently record the most recently occurring pop-up events and their triggering sources. When the target context contains pop-up state information, querying the pop-up record from the circular buffer can quickly identify abnormal pop-ups, and the business method that triggered the pop-up can be accurately located through the filtered call stack. The coding model can directly analyze and repair this business method, shortening the location time and improving the efficiency of code repair.
[0110] Taking a specific scenario as an example: a user describes "buttons cannot be clicked," the system determines the task type as a UI issue. After obtaining layout inspection data, it finds a pop-up window covering the top of the page, intercepting all touch events. The system checks the pop-up's status and finds a LoadingDialogFragment that has been present for 15 seconds without closing, classifying it as an abnormal pop-up. The system locates the trigger source in the filtered call stack of this pop-up to the ProfileViewModel.submitProfile() method. Reading the source code of this method reveals that the dismiss() method was correctly called to close the pop-up in the successful callback of the network request, but the dismiss() call was omitted in the error callback branch, causing the pop-up to fail to close when the network request fails. Based on this, the coding model generates corrective code to add the dismiss() call in the error callback. Through this method, the time spent locating the issue can be reduced from an average of 15 minutes for manual troubleshooting to approximately 30 seconds.
[0111] Optionally, in this embodiment of the application, the target context is input into the encoding model, and the encoding model performs model inference based on the target context, including: The target context is input into the encoding model. If the capacity of the encoding model's context window is insufficient to accommodate all the target contexts, the target contexts are input into the encoding model's context window in descending order of priority of dynamic data, until the remaining capacity of the context window is insufficient to accommodate the input of context data of the next priority. The encoding model then performs model inference based on the input target contexts.
[0112] The context window refers to the maximum input length that an encoding model can process in a single inference process, and it can be measured in units of tokens. Different types of encoding models have different context window capacities. When the total amount of data contained in the determined target context exceeds the upper limit of the encoding model's context window capacity, it is impossible to input all the target context into the model at once, otherwise it will lead to input truncation or model errors. Therefore, it is necessary to input the target context in batches or truncation in a certain order.
[0113] This application embodiment uses the priority of dynamic data as the basis for the input order. Based on the priority order of different categories of dynamic data established for the current task type (e.g., in a defect repair task, network communication data has a higher priority than navigation status data, and navigation status data has a higher priority than cached data), the system inputs the dynamic data in the target context into the encoding model's context window in descending order of priority. For static code context, it can be used as the lowest priority data, inputting it only when there is still remaining capacity after all dynamic data has been input, or it can be interspersed among the dynamic data depending on the task type. During the input process, the system continuously monitors the used and remaining capacity of the context window. When the remaining capacity is sufficient to accommodate all context data of the next priority, that data is input completely; when the remaining capacity is insufficient to accommodate all context data of the next priority, input stops. Therefore, the context window is fully utilized at this point, and incomplete data fragments are no longer forcibly inserted. The encoding model uses all the data currently input into the context window as the basis for inference. Through this method, even with a limited context window, the model can still obtain key information about the current task.
[0114] In the implementation of the above embodiments: considering the limited capacity of the encoding model's context window, when the total amount of data in the target context exceeds the window capacity, the dynamic data is input sequentially in descending order of priority. This ensures that the highest priority data enters the encoding model's context window first, thereby preserving as much information as possible that is most relevant to the current task. The encoding model performs inference and model reasoning based on the input high-priority context data, reducing inference failures caused by exceeding the window capacity with a single input, and also reducing information loss caused by data truncation, thus improving the success rate of model inference and the quality of generation under limited context window conditions.
[0115] This application provides a model inference apparatus, including: The dynamic data acquisition module is used to collect dynamic data during the runtime of terminal applications; The request receiving module is used to receive inference requests from users and determine the task type corresponding to the inference request based on the inference request. The context determination module is used to determine the target context from dynamic data and static code context based on the task type and the priority of the data corresponding to the task type. The inference module is used to input the target context into the encoding model, which then performs model inference based on the target context.
[0116] Optionally, in this embodiment of the application, the model inference device includes a dynamic data acquisition module, which is used to acquire dynamic data through a runtime data acquisition module deployed on the terminal application. The runtime data acquisition module includes multiple data processors divided according to data categories. The multiple data processors independently acquire and cache various types of dynamic data during the runtime of the terminal application and provide a data query interface within the terminal application. The dynamic data includes at least one of the following: network communication data, page navigation status data, pop-up lifecycle data, view inspection data, key-value pair storage data, and event tracking data.
[0117] Optionally, in this embodiment of the application, the model inference device further includes a protocol translation intermediate layer, used to obtain dynamic data from the data processor through the protocol translation intermediate layer; the protocol translation intermediate layer establishes communication connections with the encoding model and the terminal application respectively, and the protocol translation intermediate layer is used to convert the tool call request of the encoding model into a message format that can be recognized by the terminal application, and convert the dynamic data returned by the terminal application into structured text corresponding to the encoding model.
[0118] Optionally, in this embodiment of the application, the model inference device includes a context determination module, which is used to determine the main chain and auxiliary chain corresponding to the task type based on a pre-established mapping relationship between task types and information source branches; wherein the main chain represents the information source branch that is preferentially selected into the target context, and the auxiliary chain represents the information source branch that is supplementarily selected into the target context; the information source branches include dynamic branches and static branches, with dynamic branches pointing to dynamic data and static branches pointing to static code context; for each task type, the priority of dynamic data when the dynamic branch is used as the main chain is obtained; and the target context is determined from the dynamic data and static code context based on the priority of the main chain, auxiliary chain, and dynamic data corresponding to the task type.
[0119] Optionally, in this embodiment of the application, the model inference device includes a request receiving module, which is used to receive a natural language description corresponding to a user-input inference request, and determine the task type based on the natural language description using keyword rules or natural language intent recognition; the task type includes at least one of: adding new features, fixing defects, optimizing performance, and user interface issues.
[0120] Optionally, in this embodiment of the application, the model inference apparatus further includes an adjustment module, used to, if the inference request is a multi-turn dialogue, correct at least one of the priority relationship of the task type, the information source branch, and the priority order of dynamic data according to the calling situation of the encoding model in subsequent rounds of context data acquisition; wherein, the calling situation includes the calling situation of the encoding model for other context data during the inference process after acquiring the target context.
[0121] Optionally, in this embodiment of the application, the model inference device includes dynamic data including pop-up lifecycle data; the dynamic data acquisition module is further used to register pop-up lifecycle callback hooks, for dialog fragment type pop-ups, to listen to the display and close events of dialog fragments through the lifecycle callback of the fragment manager, and for traditional non-fragment type dialogs, to wrap the display and close methods of traditional dialogs through the proxy pattern to listen to the display and close events; determine pop-up lifecycle data; the pop-up lifecycle data is used as pop-up state information in the target context when the task type is a user interface problem.
[0122] Optionally, in this embodiment of the application, the model inference device includes a context determination module, which is used to capture the call stack of the current thread after the pop-up display event is triggered, and filter the call stack, removing the stack frames of the framework layer and the third-party library layer to obtain a filtered call stack; the filtered call stack includes the call stack frames of the business code layer; the association information between the filtered call stack and the pop-up is stored in a circular buffer; the association information of the pop-up includes the pop-up's identification information, timestamp, and parent container; if the target context contains pop-up state information, the pop-up record is queried from the circular buffer; if an abnormal pop-up exists, the business method that triggered the pop-up is located from the filtered call stack, and the coding model generates repair code based on the business method.
[0123] Optionally, in this embodiment of the application, the model inference device includes a context determination module, which is used to input the target context into the encoding model. If the capacity of the context window of the encoding model is insufficient to accommodate all the target contexts, the target contexts are sequentially input into the context window of the encoding model in descending order of priority of dynamic data, until the remaining capacity of the context window is insufficient to accommodate the input of context data of the next priority. The encoding model then performs model inference based on the input target contexts.
[0124] It should be understood that this device corresponds to the model reasoning method embodiment described above and is capable of performing the various steps involved in the above method embodiment. The specific functions of this device can be found in the description above, and detailed descriptions are omitted here to avoid repetition. The device includes at least one software functional module that can be stored in memory or embedded in the device's operating system (OS) in the form of software or firmware.
[0125] Please see Figure 2The diagram shows a structural schematic of an electronic device provided in an embodiment of this application. An electronic device 300 provided in this application includes a processor 310 and a memory 320. The memory 320 stores machine-readable instructions executable by the processor 310. When the machine-readable instructions are executed by the processor 310, the method described above is performed.
[0126] Figure 2 The components shown can be implemented using hardware, software, or a combination thereof. Electronic device 300 may be a physical device, such as a server or PC, or a virtual device, such as a virtual machine or virtualization container. Furthermore, electronic device 300 is not limited to a single device; it can be a combination of multiple devices or a cluster of numerous devices.
[0127] This application also provides a storage medium storing a computer program, which is executed by a processor to perform the above-described method.
[0128] The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0129] This application also provides a computer program product, including computer program instructions, which are executed by a processor to perform the method described above.
[0130] It should be understood that the disclosed apparatus and methods can also be implemented in other ways, given the several embodiments provided in this application. The apparatus embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of apparatus, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code, which contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, or they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram and / or flowchart, and combinations of blocks in block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0131] In addition, the functional modules in the various embodiments of this application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.
[0132] The above description is only an optional implementation of the embodiments of this application, but the protection scope of the embodiments of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the embodiments of this application should be covered within the protection scope of the embodiments of this application.
Claims
1. A model reasoning method, characterized in that, include: Collect dynamic data during the runtime of terminal applications; Receive the user's reasoning request and determine the task type corresponding to the reasoning request based on the reasoning request; Based on the pre-established mapping relationship between task types and information source branches, the main chain and auxiliary chain corresponding to the task type are determined, and based on the main chain and auxiliary chain, the target context is determined from the dynamic data and static code context; wherein, the main chain represents the information source branch that is preferentially selected into the target context, and the auxiliary chain represents the information source branch that is supplementarily selected into the target context; the information source branch includes dynamic branches and static branches, the dynamic branch points to the dynamic data, and the static branch points to the static code context; The target context is input into the encoding model, which then performs model inference based on the target context. Collect dynamic data during the runtime of the terminal application, including: The dynamic data is collected through a runtime data collection module deployed on the terminal application. This runtime data collection module includes multiple data processors categorized by data type. Each data processor independently collects and caches various types of dynamic data generated during the terminal application's runtime and provides a data query interface within the terminal application. The dynamic data includes at least one of the following: network communication data, page navigation status data, pop-up lifecycle data, view inspection data, key-value pair storage data, and event tracking data. Receive a user's inference request, and determine the task type corresponding to the inference request based on the inference request, including: The system receives a natural language description corresponding to a user-input reasoning request, utilizes keyword rules or natural language intent recognition, and determines the task type based on the natural language description. The task type includes at least one of the following: adding new features, fixing defects, optimizing performance, and addressing user interface issues.
2. The method according to claim 1, characterized in that, After collecting dynamic data during the runtime of the terminal application, the method further includes: The protocol translation intermediate layer obtains the dynamic data from the data processor; the protocol translation intermediate layer establishes communication connections with the encoding model and the terminal application respectively; the protocol translation intermediate layer is used to convert the tool call request of the encoding model into a message format that the terminal application can recognize, and convert the dynamic data returned by the terminal application into structured text corresponding to the encoding model.
3. The method according to claim 1, characterized in that, Based on the task type, determine the corresponding main chain and auxiliary chain, and based on the main chain and auxiliary chain, determine the target context from the dynamic data and static code context, including: For each task type, the priority of the dynamic data when the dynamic branch is used as the main chain is determined; The target context is determined from the dynamic data and static code context based on the priority of the main chain, the auxiliary chain, and the dynamic data corresponding to the task type.
4. The method according to claim 1, characterized in that, After receiving the natural language description corresponding to the inference request input by the user, the method further includes: If the inference request is a multi-turn dialogue, at least one of the following is corrected based on the context data retrieval status of the encoding model in subsequent rounds: the priority relationship of the task type, the priority order of the information source branches, and the priority order of the dynamic data; wherein, the retrieval status includes the retrieval status of other context data by the encoding model during the inference process after obtaining the target context.
5. The method according to claim 1, characterized in that, The dynamic data includes pop-up window lifecycle data; the methods for collecting the pop-up window lifecycle data include: Register pop-up lifecycle callback hooks. For dialog fragment type pop-ups, listen to the display and close events of the dialog fragment through the lifecycle callback of the fragment manager. For traditional non-fragment type dialogs, wrap the display and close methods of the traditional dialogs through the delegate pattern to listen to the display and close events; determine the pop-up lifecycle data. The pop-up lifecycle data is used as pop-up state information in the target context when the task type is a user interface problem.
6. The method according to claim 5, characterized in that, The target context is input into the encoding model, and the encoding model performs model inference based on the target context, including: After the pop-up display event is triggered, the call stack of the current thread is captured and filtered to remove the stack frames of the framework layer and the third-party library layer, thus obtaining the filtered call stack; the filtered call stack includes the call stack frames of the business code layer. The filtered call stack and pop-up association information are stored in a circular buffer; the pop-up association information includes the pop-up's identifier, timestamp, and parent container. If the target context contains pop-up status information, the pop-up record is queried from the circular buffer. If an abnormal pop-up exists, the business method that triggered the pop-up is located from the filtered call stack, and the coding model generates repair code based on the business method.
7. The method according to claim 1, characterized in that, The target context is input into the encoding model, and the encoding model performs model inference based on the target context, including: The target context is input into the encoding model. If the capacity of the encoding model's context window is insufficient to accommodate all the target contexts, the target contexts are input into the encoding model's context window sequentially according to the priority of the dynamic data from high to low, until the remaining capacity of the context window is insufficient to accommodate the input of context data of the next priority. The encoding model then performs model inference based on the input target contexts.
8. A computer program product, characterized in that, It includes computer program instructions that, when executed by a processor, perform the method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Supply chain low-code platform service management method, equipment and medium
CN121301049A
Pose measuring device for accuracy improvements in kinematics
DE102024138646A1