Data reasoning method, device, equipment and computing medium
By retrieving similarity data and historical demonstration examples from the database, a preset number of demonstration examples are obtained to train the adapter, which solves the problem of inaccurate output of large language models in dynamic web page environments and improves the inference accuracy of the model.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- MIAOZHEN INFORMATION TECH CO LTD
- Filing Date
- 2025-12-18
- Publication Date
- 2026-04-21
AI Technical Summary
Existing intelligent proxies based on large language models struggle to accurately execute complex processes when web page styles and DOM structures change frequently, resulting in inaccurate output.
By obtaining the retrieval results of multimodal data in the database, if no process template with a similarity threshold above the threshold is found or there are insufficient historical demonstration examples, a preset number of demonstration examples are obtained to train the adapter, and the inference result is determined by using the trained adapter and the inference model.
The model's inference accuracy when faced with unseen webpage structures has been improved, and the operational accuracy in dynamic webpage environments has been enhanced by fine-tuning the model parameters through adapters.
Smart Images

Figure CN121328749B_ABST
Abstract
Claims
1. A data reasoning method, characterized in that, include: Acquire multimodal data as input to the trained inference model; Determine the retrieval results of the multimodal data in the database, which stores process templates corresponding to the multimodal sample data input during the training of the inference model, as well as historical demonstration examples corresponding to each process template; If the search results indicate that no process template with a similarity greater than a preset similarity threshold to the multimodal data was found, or if the at least one process template was found and the number of historical demonstration examples corresponding to each of the at least one process templates is less than a preset number of example examples, then a preset number of demonstration examples based on the multimodal data are obtained. The adapter is trained based on the preset number of demonstration examples to obtain a trained adapter; The trained adapter and the trained inference model are used to determine the inference result corresponding to the multimodal data; Determining the retrieval results of the multimodal data in the database includes: Generate a query vector corresponding to the multimodal data; Obtain the vector index of each process template in the database; The similarity between the multimodal data and each process template is determined based on the similarity between the query vector and each of the vector indices; Determine whether there exists at least one process template whose similarity to the multimodal data is greater than a preset similarity threshold; If present, determine the historical demonstration example corresponding to the at least one process model; The at least one process template and the corresponding historical demonstration examples of each of the at least one process template are determined as search results; If the query does not exist, an empty string will be considered as a search result. The method further includes: If the search result indicates that at least one process template with a similarity greater than a preset similarity threshold to the multimodal data has been retrieved, or if the at least one process template has been retrieved and the number of historical demonstration examples corresponding to the process template in the at least one process template is greater than or equal to the preset example number threshold, then a prompt message is generated based on the search result and the multimodal data. The prompt message is used to assist the trained inference model in determining the semantic information corresponding to the multimodal data. The prompt information and the multimodal data are input into the trained inference model to obtain the inference result.
2. The method according to claim 1, characterized in that, The step of determining the inference result corresponding to the multimodal data using the trained adapter and the trained inference model includes: Multiple micro-rank matrices are determined based on the trained adapter; The multiple micro-rank matrices are respectively inserted into their respective self-attention layers in the trained inference model to obtain the target inference model; The multimodal data is input into the target inference model to obtain the inference result.
3. The method according to claim 1, characterized in that, The generation of prompt information based on the search results and the multimodal data includes: Obtain interface element information from the multimodal data; The target process template is determined from the process templates in the search results; Based on the interface element information and the target process template, a prompt message is generated.
4. The method according to claim 3, characterized in that, The database also includes semantic element mapping information, and the generation of prompt information based on the interface element information and the target process template includes: Obtain the key semantic information corresponding to each placeholder in the target process template; The semantic element mapping information determines the key element identifier corresponding to each key semantic information. Search for the key element information corresponding to each key element identifier in the interface element information; The prompt information is obtained by replacing the corresponding placeholders with the information of each key element.
5. The method according to any one of claims 1-4, characterized in that, The method further includes: The action sequence and confidence level in the inference result are analyzed to obtain the analysis result; The parsing results are executed under a preset environment to obtain the execution results. The execution results are used to compare with the task requirement information in the multimodal data to verify the inference results.
6. A data inference device, characterized in that, Applied to the server side, including: The acquisition module is used to acquire multimodal data that is input into the trained inference model; The determination module is used to determine the retrieval results of the multimodal data in the database, wherein the database stores the process templates corresponding to the multimodal sample data input during the training of the inference model and the historical demonstration examples corresponding to each process template; The determining module is further configured to, if the search result indicates that no process template with a similarity greater than a preset similarity threshold to the multimodal data was found, or if the at least one process template was found and the number of historical demonstration examples corresponding to each of the at least one process templates is less than a preset example number threshold, then obtain a preset number of demonstration examples based on the multimodal data. The training module is used to train the adapter based on the preset number of demonstration examples to obtain a trained adapter; The determining module is further configured to use the trained adapter and the trained inference model to determine the inference result corresponding to the multimodal data; The determining module is specifically used for: Generate a query vector corresponding to the multimodal data; Obtain the vector index of each process template in the database; The similarity between the multimodal data and each process template is determined based on the similarity between the query vector and each of the vector indices; Determine whether there exists at least one process template whose similarity to the multimodal data is greater than a preset similarity threshold; If present, determine the historical demonstration example corresponding to the at least one process model; The at least one process template and the corresponding historical demonstration examples of each of the at least one process template are determined as search results; If the query does not exist, an empty string will be considered as a search result. The generation module is configured to generate prompt information based on the search results and the multimodal data if the search results indicate that at least one process template with a similarity greater than a preset similarity threshold has been retrieved, or if at least one process template has been retrieved and the number of historical demonstration examples corresponding to the process template in the at least one process template is greater than or equal to the preset example number threshold, and the prompt information is used to assist the trained inference model in determining the semantic information corresponding to the multimodal data; The input module is used to input the prompt information and the multimodal data into the trained inference model to obtain the inference result.
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, The processor executes the computer program to implement the method as described in any one of claims 1-5.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, The program is executed by a processor to implement the method as described in any one of claims 1-5.
Citation Information
Patent Citations
Recommendation method and device based on large language model, equipment and storage medium
CN117973545A
Image searching method and device
CN118503473A