Automatic transaction operation method, system and hardware based on AI agent and multi-modal training execution framework

Through the automated transaction operation method based on the AI ​​agent and multimodal training execution framework, the existing AI systems have solved the problems of high manual participation, insufficient intelligence level, difficulty in automation adaptation and maintenance in the field of automated transaction processing, and efficient automation operation and multi-field applications have been achieved, improving operational efficiency and user experience.

CN120010964AInactive Publication Date: 2025-05-16SHENZHEN HUOLING JIANMU INTELLIGENT CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510156486.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-12
Publication Date
2025-05-16
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing AI systems have problems such as high manual participation, insufficient intelligence, difficulty in automation adaptation and difficulty in automation maintenance in the field of automated transaction processing.

Method used

Through automated transaction operation methods based on AI agents and multimodal training execution framework, remote wireless control, user information collection and preprocessing, the use of encrypted network protocols, the definition of automated operation specifications, the implementation of programmable training model logic control rules, the interpretation and execution of logical execution units, the script merging and updating mechanism, the construction of crowdsourcing data acquisition systems, the application of artificial correction mechanisms, and the use of multi-model AI inference decision-making mechanisms.

Benefits of technology

It reduces manual participation, improves intelligence level, improves automation adaptation and maintenance capabilities, realizes natural human-computer interaction, supports automation applications in multiple fields, reduces labor costs and resource waste, and improves operational efficiency and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120010964A_ABST
    Figure CN120010964A_ABST
Patent Text Reader

Abstract

The invention discloses an automatic transaction operation method, system and hardware based on an AI (artificial intelligence) agent and a multi-modal training execution framework, and relates to the fields of artificial intelligence application, automation technology and the like. The problems are solved through combination of software and hardware, the hardware AI intelligent agent customizes an operating system, has the capabilities of remote wireless control and the like, simulates user operation by adopting screenshot and click, and adapts to various operating systems; a multi-modal AI training reasoning execution framework of software maintains reasoning execution scripts for service providers, business automation is achieved by automatically generating and updating the scripts, training models, logic control rules and the like are covered, and a manual correction mechanism is provided. Data is collected for analysis by means of internet community crowdsourcing. By means of AI deep learning and other technologies, the AI agent is endowed with strong capacity, automatic scripts are inferred and executed, stable operation of the automatic scripts is guaranteed, and natural interaction is achieved by supporting multiple technologies. The method is rich in application scene, has high economic and social values, and promotes industry transformation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of artificial intelligence applications, automation technology and big data analysis, and specifically to an automated transaction operation method, system and hardware based on an AI agent and a multimodal training execution framework. Background Art

[0002] As the core driving force of the new generation of information technology, artificial intelligence has demonstrated transformative potential in many fields by simulating the perception, reasoning and decision-making capabilities of human intelligence. Since the breakthroughs in deep learning and neural network technology, AI has gradually moved from theoretical research to practical applications, especially in the fields of computer vision, natural language processing, and automated control. The core value of AI technology lies in replacing or assisting humans in completing repetitive and highly complex tasks through algorithm optimization and computing power enhancement, thereby improving efficiency and reducing labor costs, becoming an important technical foundation for promoting the digital transformation of the industry.

[0003] Although large models have outstanding performance in general fields such as text generation and knowledge question answering, their implementation in specific industry scenarios still faces multiple bottlenecks. First, the training and reasoning of large models rely on massive computing resources, and it is difficult to adapt to automated operation scenarios with high real-time requirements; secondly, general models lack a deep understanding of the business logic of vertical fields and cannot directly meet the needs of precise operating rules in scenarios such as smart travel and e-commerce operations. In addition, existing AI systems are mostly limited to a single modality, and it is difficult to effectively integrate cross-device sensor data and dynamic environmental feedback, resulting in insufficient adaptability for cross-platform task execution. Traditional automation scripts are limited by technology, require repeated development, have high maintenance costs, and rely on manual programming for updates. These problems have prevented large model technology from being applied on a large scale in the field of automated transaction processing.

[0004] In summary, it is necessary to propose automated business operation methods, systems and hardware based on AI agents and multimodal training execution framework to solve the above problems. Summary of the invention

[0005] The present invention aims to provide an automated transaction operation method, system and hardware based on AI agent and multimodal training execution framework to solve the problems of excessive manual participation, insufficient intelligence level, difficulty in automated adaptation and difficulty in automated maintenance in existing digital services.

[0006] To achieve the above object, the present invention provides the following method:

[0007] Through the hardware operating system, the AI ​​agent can have remote wireless control capabilities, thereby realizing remote operation of the equipment.

[0008] User information is collected through the AI ​​intelligent agent and the training data is pre-processed.

[0009] Based on information collection security specifications, encrypted network protocols and dedicated communication channels are used in the process of collecting information to strengthen user information management and privacy protection and ensure user data security.

[0010] A set of automated operation specifications are defined and implemented, and the main operation objects are apps, applets, web pages and other applications provided by third-party service providers running on customized systems. The specifications only use two operation methods, namely screenshots and screen clicks, to completely simulate the interaction between normal users and devices. They do not rely on the API interface and network packet capture technology of service providers, and can be adapted to most hardware operating systems on the market. Operations can even be achieved through external cameras and external robotic arms, further reducing restrictions on the underlying operating system.

[0011] A set of programmable logic control rules for training models are defined and implemented. The rules include logic control (judgment, branching, loop) and logic operation (screen judgment / status judgment of hardware devices, operation of hardware devices). They can be used to build training models for any business. By manually summarizing and concluding each business scenario and describing it using rules, a general business training model is formed.

[0012] Through the logic execution unit, the logical control rules of the training model are interpreted and executed to generate the inference execution path.

[0013] Through the script merging and updating mechanism, multiple inference execution paths are merged into one script, and updated and iterated when the path changes.

[0014] By building a crowdsourcing data collection system, we collect training scenarios and data. Based on the Internet social model, we allow users across the country to upload sample data on their own and collect trigger conditions for various business scenarios to provide rich data support for training.

[0015] Through the manual correction mechanism, during the training and reasoning process, if new scenarios not covered by the training model or incomplete processing logic arise, manual intervention will be conducted to review, add corresponding scenarios or enhance the robustness of the rules to ensure the completeness and robustness of the training model.

[0016] Through the multi-model AI reasoning decision-making mechanism, AI is allowed to reason in limited scenarios during model training. Through constrained reasoning, the number of scenarios at each stage is informed to AI, allowing it to choose from given scenarios; at the same time, prior knowledge is provided, and the characteristics of each scenario are listed in detail as a basis for judgment, thereby improving the accuracy of AI output.

[0017] Through the multimodal AI reasoning execution framework, script task issuance, execution, report collection and result analysis can be realized to ensure the smooth progress of automated transaction operations.

[0018] An algorithm for calculating the usage of public resources based on big data, which calculates the probability of public resource availability through historical data and current user information.

[0019] Based on the above method, the present invention realizes the following system:

[0020] The Internet community operation module and the data crowdsourcing collection module are used to publish training data collection tasks, collect public resource information, user operations, and build an AI application ecosystem.

[0021] Multimodal AI training module, used to automatically train AI models and generate and maintain inference execution scripts.

[0022] Multimodal AI reasoning execution module, used to execute reasoning execution scripts.

[0023] The hardware includes: a customized operating system that provides remote control capabilities, obtains advanced system permissions, and opens up the underlying operating system commands.

[0024] Compared with the prior art, the present invention has the following beneficial effects:

[0025] Significant technical advantages: Deep learning and reinforcement learning technologies are used to give AI agents powerful complex decision-making processing capabilities, and machine learning algorithms are used to improve data processing efficiency. Automated scripts distinguish between exploration and execution stages, reducing script learning and maintenance costs. Manual correction mechanisms ensure stable script operation, while supporting natural language processing, speech recognition and other technologies to achieve natural human-computer interaction.

[0026] Rich application scenarios: In the field of travel, it can analyze data such as parking lots and gas stations in real time, plan travel routes, avoid congestion, and improve travel efficiency; accurately locate targets, monitor vehicle status and driving environment, and ensure driving safety. It can also be expanded to e-commerce, finance, security and other fields, such as automatic e-commerce order processing, financial risk assessment, and security surveillance video analysis.

[0027] High economic and social value: reduce labor costs and resource waste, improve operational efficiency in various industries; improve user experience and save time for users; promote digital transformation in various industries and promote social and economic development, while ensuring data security and stable system operation. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0029] In order to more clearly show the situation of the embodiments of the present application in practical applications, the key drawings involved in the embodiments will be described in detail below. These drawings intuitively present the core technical content of the invention, including system architecture, method flow, etc. For ordinary technicians in this field, by studying these drawings, they can more clearly understand the technical solution of the present application, and can quickly grasp the technical points based on these drawings without conducting complex experiments or in-depth research. At the same time, in the subsequent implementation of the technical solution of the present application, it can also be flexibly adjusted and optimized according to these drawings in combination with actual scenarios.

[0030] FIG1 is a flow chart of the business execution of the present invention.

[0031] FIG2 is a flow chart of the multimodal AI training phase and the reasoning execution phase of the present invention.

[0032] FIG3 is a multimodal AI training framework component and execution flow chart of the present invention.

[0033] Figure 4 It is a schematic diagram of the interaction and operation flow chart of the main modules of the present invention. Specific implementation plan

[0034] The technical solutions in the embodiments of the present invention will be described clearly and completely below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0035] In the description of this specification, "include", "including", "have", "contain" and the like are open terms, which means including but not limited to. The description with reference to terms such as "one embodiment", "a specific embodiment", "some embodiments", "for example", etc. means that the specific features, structures or characteristics described in combination with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures or characteristics described may be combined in a suitable manner in any one or more embodiments or examples. The order of steps involved in each embodiment is used to schematically illustrate the implementation process of the present application, wherein the order of steps is not fixed and may be appropriately adjusted according to actual needs.

[0036] Figure 1 shows the overall business process, Figure 2 shows the technical details of the AI ​​model training execution framework, and Figure 3 shows the components and execution process of the AI ​​model training framework. Taking the parking fee payment business of a parking lot as an example, the specific steps are as follows:

[0037] The specific process includes:

[0038] S1, crowdsourcing data collection. Through Internet community operations, guide users to upload business data for training models. These data include parking lot app information used by users, parking payment QR codes, license plate numbers of vehicles in parking lots, etc.

[0039] S2, parking lot coding. The parking lot is coded in a specific coding format, which is PKXXXXXXXXXXXXXXXX0001. Among them, PK is the business code, representing the parking lot; the middle 16 bits are the geographic location information code related to longitude and latitude. The specific coding rule is to divide the earth into 4^16 = 4.2 billion blocks according to the Mercator projection, and the size of each block is about 0.5×0.5 square kilometers. Then, the coordinates of the blocks are converted into quaternary codes through quadtree key-value coding; the last four bits are counting codes, which start from 1 and increase in sequence according to the order of parking lot entry.

[0040] S3, business data classification and training data preprocessing, including extracting the longitude and latitude of the parking lot for geographic fence division, identifying the URL character features corresponding to the parking payment QR code, analyzing the domain name to determine whether it corresponds to the same service provider, and performing OCR recognition on the license plate number.

[0041] S4, builds and manages hardware clusters, and connects customized hardware to the device management system in a clustered manner. The device management system uniquely identifies the hardware through customized hardware coding, and uses remote control functions to achieve device online and offline management, work status management, training task issuance, and training log recovery. At the same time, the system monitors the hardware operating load and CPU / memory consumption in real time to dynamically allocate training tasks.

[0042] S5, AI training model construction. The AI ​​training model is the core module of the multimodal AI training model framework, which is built based on custom logic control rules. The logic control rules are a set of coding language-like functions that provide two types of programming keywords. One type is the logic control keyword, which has process control functions such as judgment, fork, loop, jump, exception handling, and end, which are used to organize the business logic of the model; the other type is the logic operation keyword, which gives the ability to operate the hardware device, including hardware interaction functions such as click, slide, virtual button, screenshot, etc., which is analogous to the process of users operating hardware. It also has screen judgment functions such as text recognition, image classification, image comparison, image target recognition, image feature extraction, and image understanding, which can simulate the process of users observing the screen and making judgments. Among them, the hardware device operation capability obtains the operation authority of the hardware from the bottom layer by customizing the hardware system; the screen judgment capability integrates image processing related technologies, among which AI image recognition capability is particularly important. The overall judgment ability of the image is enhanced through AI technology, which is used to identify screen scenes and context logic.

[0043] S6, AI training stage. In the parking payment business scenario, the AI ​​training process needs to exhaust all the pages of the current parking lot digital application (usually WeChat / Alipay applet). This process needs to be triggered by different license plate numbers, and the covered states include payment orders, insufficient parking time without payment, monthly card users without payment, paid without payment, exited without payment, no parking information found, etc.

[0044] S7, constrain the number of AI judgment scenarios in the AI ​​training model and provide detailed scenario descriptions. The processes of different parking lot payment applications are different. For example, after entering the homepage, you may need to enter the license plate number before querying, need to change the historical license plate before querying, an input box appears and the first two digits have defaulted to the current provincial and municipal license plate number, need to select the vehicle type before continuing the operation, pop-up ads, user agreements, login registration, and follow-up. Even for the same scenario, the page presentation of different parking lot payment applications may be very different. For example, in the license plate number input scenario, there may be multiple independent input boxes, long strip input boxes, or a button needs to be triggered to pop up the input interface. By enumerating all scenarios in the AI ​​training model, AI is prompted to choose from limited scenarios, which greatly ensures the accuracy of AI judgment; at the same time, for each scenario, various possible forms are described in detail to provide sufficient basis for AI to make correct judgments.

[0045] S8, get the AI ​​training reasoning execution path. The AI ​​training model defines the processing flow of all scenarios, but for specific parking lot applications, the specific operations need to be determined based on the page layout and operation logic of the application. For example, the AI ​​training model defines the business process node of "clicking the payment button", but the shape and position of the payment button may be different under different digital implementation methods, and some digital implementations may even require 2-3 steps to find the final payment button. The key to the training process is to find the payment operation path of the parking lot.

[0046] S9, manual labeling when training fails. The AI ​​training model cannot cover all situations of all parking lot payment applications in the initial stage. When the AI ​​training model is not defined during the training process, manual intervention is required to review whether the rules defined by the model are accurate and complete. There are two types of errors:

[0047] 1. New scenarios that are not covered by the training model. In this case, you need to add the corresponding scenarios to the training model and retrain it.

[0048] 2. The training model has covered the scenario, but the processing logic is not complete. At this time, the robustness of the rules needs to be enhanced. For example, a business node is defined in the training model as "find the close button", but the close button of the current parking lot is an X icon instead of pure text. We need to add image recognition capabilities based on the original OCR recognition.

[0049] S10, the inference execution paths are merged into inference execution scripts. When the trained inference execution paths cover all the scenarios of parking lot service payment, multiple paths need to be merged into a directed graph, namely the inference execution script. The key here is to find the common nodes of multiple paths for merging. These common nodes are usually nodes that will generate logical bifurcations for scene judgment, page element search, page feature search, etc.

[0050] In S11, the new parking lot automatic payment function was launched. When the automatic script of a parking lot is available, targeted message push and global announcement will be made through the operating Internet community, so that users whose activity tracks match the parking lot and have relevant search and location records will be informed of the launch first. At the same time, internal test invitations will be issued through the community to encourage users to conduct offline verification.

[0051] S12, user use. When users use the parking payment function through the hardware terminal customized by the system, the system will obtain user-related information for big data analysis of the availability of parking spaces, specifically obtaining the user's location, parking start and end time, parking duration, payment amount, etc. We make predictions by establishing a regression model, which covers the current number of parking spaces, the future number of parking spaces, and the parking lot tiered charging standard. This provides a decision-making basis for users to choose parking lots.

[0052] S13, user triggers automatic payment. When the user triggers the automatic payment function through the system-customized hardware, the hardware terminal needs to collect the user's current latitude and longitude information, compare it with the geographic fences of each parking lot in the background, find the parking lot where the user is located, and then find the inference execution script corresponding to the parking lot. Combined with the license plate number and payment password provided by the user, the inference execution script is executed to complete the payment.

[0053] S14: Manual labeling is performed when execution fails. Execution failure includes two situations:

[0054] 1. The corresponding parking lot is not included. In this case, an upload activity will be posted to users through the Internet community to guide users to upload the parking lot information for training.

[0055] 2. The inference execution script is not perfect. In this case, you need to analyze the execution log, find the missing points in the script, improve the AI ​​model, and retrain a new inference execution script.

[0056] The above description is merely a preferred embodiment of one or more embodiments of this specification and is not intended to limit one or more embodiments of this specification. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of one or more embodiments of this specification shall be included in the scope of protection of one or more embodiments of this specification.

[0057] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.

[0058] The user information involved in this manual (including but not limited to user device information, user personal information, etc.) will be information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

Claims

1. An automated transaction operation method based on AI agent and multimodal training execution framework, characterized in that: The following steps are involved: S1. Give AI agents remote wireless control capabilities through customized hardware operating systems to achieve remote operation of target devices; S2. Collect user behavior data through encrypted network protocols and dedicated communication channels, and perform desensitization and secure storage; S3. Based on automated operation specifications, simulate the user's operation on the third-party application running on the device. The operation method only includes screenshots and screen clicks, and does not rely on the service provider's API interface or network packet capture technology; S4. Construct programmable training model logic control rules, including logic control modules (judgment, branching, looping) and logic operation modules (judgment of device screen status and generation of operation instructions), and generate general training models for business scenarios; S5. The logic execution unit interprets and executes the logic control rules to generate an inference execution path; S6. Through the script merging and updating mechanism, multiple inference execution paths are merged into an executable script, and dynamically updated according to path changes; S7. Collect trigger conditions and sample data of multiple business scenarios based on crowdsourcing data collection system for training model optimization; S8. Intervene and correct rules for scenarios or logical errors not covered during training or reasoning through manual correction mechanisms; S9. Improve the accuracy of AI output through multi-model AI reasoning decision-making mechanism, combined with scene quantity constraints and prior knowledge description; S10. Implement the issuance, execution, report collection and result analysis of script tasks through a multimodal AI reasoning execution framework; S11. By defining a set of user information collection and preprocessing mechanisms on the AI ​​agent side, combining user information and historical data, the probability of public resource use is predicted through big data computing algorithms.

2. The method according to claim 1, characterized in that The S3 further includes: completely simulating user behavior through screenshots and screen click operations, and adapting to various operating systems.

3. The method according to claim 1, characterized in that The S4 further includes: defining a set of programmable languages ​​for describing training models.

4. The method according to claim 1, characterized in that: In S7, the crowdsourcing data collection system is based on the Internet social model, where users actively upload business scenario data, and compliance is ensured through data desensitization and privacy protection mechanisms.

5. The method according to claim 1, characterized in that The S8 also includes a manual correction module, which manually annotates the interface element features and updates the training model when the inference execution script does not cover the new scene or a recognition error occurs.

6. The method according to claim 1, characterized in that The multimodal AI training reasoning execution framework further includes: Multimodal AI training models for logical modeling of common processes in real-world business scenarios; The logic execution unit,records the operation paths by traversing the digital interfaces of different,service providers and generates differentiated reasoning execution scripts.

7. The method according to claim 1, characterized in that The multimodal AI training reasoning execution framework is provided with an AI reasoning decision-making mechanism, which constrains the AI ​​to select an operation path within a limited range by providing preset scenario options and judgment basis to the AI.

8. The method according to claim 1, characterized in that: The task dispatching module in the inference execution phase is executed through a distributed architecture scheduling script, and the result analysis module optimizes the business strategy by counting execution logs.

9. An automated transaction operating system based on AI agent and multimodal training execution framework, characterized in that: For implementing the automated transaction operation method according to any one of claims 1 to 8, the system comprises: The Internet community operation module and the data crowdsourcing collection module are used to publish training data collection tasks, collect public resource information, user operations, and build an AI application ecosystem; Multimodal AI training module, used to automatically generate and maintain inference execution scripts; Multimodal AI reasoning execution module, used to execute scripts and provide feedback on operation results; The manual correction module is used to manually intervene in and correct rules of the training model and execution process.

10. The system according to claim 9, characterized in that Further including: Natural language processing and speech recognition modules are used to enable natural interaction between users and AI agents.

11. The system according to claim 9, characterized in that The multimodal AI reasoning execution module further includes: The task scheduling unit is used to dynamically assign script execution priorities based on resource usage probabilities.

12. An automated transaction operation hardware device for implementing the method according to any one of claims 1 to 8, characterized in that: include: Customized operating system, which can be used for remote control, obtaining advanced system permissions, and operating system low-level commands; Used to implement automated operations of AI reasoning execution scripts to implement an automated transaction operation method based on an AI agent and a multimodal training execution framework as described in any one of claims 1 to 8.