Large model calling method and device based on operation scene, equipment and storage medium
By implementing scenario-based adaptation and dynamic key management, the issues of data security and flexibility in large model invocation are resolved, inference accuracy and development efficiency are improved, and precise adaptation and secure isolation of large model services are achieved.
Patent Information
- Application Number
- CN202511153339.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-18
- Publication Date
- 2025-12-02
AI Technical Summary
Existing large model invocation schemes have data security risks, lack flexibility and optimization space, and cannot automatically switch to the most suitable model according to the nature of the task, resulting in low inference accuracy and low development efficiency.
The context detection module collects operation data, determines the target operation scenario based on scene recognition rules, uses the key selection module to find the target API key and service address, constructs an API request to call the large model service, and realizes scenario-based adaptation and dynamic key management.
It improved the inference accuracy of large model services, reduced the risk of sensitive information leakage, and achieved a significant improvement in development efficiency and security isolation.
Smart Images

Figure CN121050802A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of large model technology, and in particular to a method, apparatus, device and storage medium for calling large models based on operational scenarios. Background Technology
[0002] Large language models have demonstrated powerful capabilities in general task processing. Currently, AI code assistant plugins are widely used in various integrated development environments (IDEs). These plugins greatly improve developers' coding efficiency.
[0003] However, existing implementations typically suffer from the following issues:
[0004] Single API Key Configuration: Most plugins allow users to configure only one fixed API key in their settings. This key is bound to a unique cloud-based large language model service, typically specified by the plugin provider. All types of requests, whether involving the completion of sensitive business code or general programming inquiries, are sent to this external, single cloud server for processing.
[0005] Data security risks: For enterprises, sending internal source code containing trade secrets, algorithms, and architecture to third-party cloud services poses a significant risk of data leakage. This is the core reason why many enterprises hesitate to introduce such AI assistants.
[0006] Lack of flexibility and optimization space: Different tasks have different requirements for large models. Internal models may have an advantage in understanding internal codebases, while external general-purpose models are more powerful in handling open-domain knowledge question answering. Existing technologies often adopt a "one-size-fits-all" approach, failing to call the most suitable model based on the nature of the task, which is neither efficient nor economical.
[0007] Manual switching is cumbersome: Even if some plugins or IDEs allow configuration of different proxies or server endpoints, this switching still requires developers to manually enter the settings page to make changes. The process is cumbersome and disrupts the workflow, making it impossible to achieve seamless and automatic switching based on the current task scenario. Summary of the Invention
[0008] This invention provides a method, apparatus, device, and storage medium for invoking large models based on operational scenarios, in order to solve the technical problem of low inference accuracy of traditional large models in vertical industry applications.
[0009] According to one aspect of the present invention, a method for invoking a large model based on an operational scenario is provided, the method comprising:
[0010] The context detection module collects user operation data in the integrated development environment, and determines the target scene identifier of the target operation scene corresponding to the operation data based on the preset scene recognition rules and the operation data.
[0011] The key selection module searches for the target API key and target service address corresponding to the target scenario identifier in the configuration storage module based on the target scenario identifier.
[0012] An API request is constructed based on the target API key and the target service address to invoke the large model service.
[0013] According to another aspect of the present invention, a large model invocation device based on an operational scenario is provided, the device comprising:
[0014] The scene determination module is used to collect user operation data in the integrated development environment through the context detection module, and determine the target scene identifier of the target operation scene corresponding to the operation data based on the preset scene recognition rules and the operation data.
[0015] The key lookup module is used to search for the target API key and target service address corresponding to the target scenario identifier in the configuration storage module based on the target scenario identifier, through the key selection module.
[0016] The model invocation module is used to construct an API request based on the target API key and the target service address to invoke the large model service.
[0017] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:
[0018] At least one processor; and
[0019] A memory communicatively connected to the at least one processor; wherein,
[0020] The memory stores a computer program that can be executed by the at least one processor, which enables the at least one processor to execute the large model invocation method based on the operation scenario as described in any embodiment of the present invention.
[0021] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the large model invocation method based on an operation scenario as described in any embodiment of the present invention.
[0022] The technical solution of this invention collects user operation data in the integrated development environment through a context detection module. Based on preset scene recognition rules and the operation data, it determines the target scene identifier corresponding to the operation data. This enables scenario-based adaptation of large model service calls, avoiding insufficient generalization of a single model for multiple scenario tasks and improving the fit between output results and actual business needs. Then, the key selection module searches for the target API key and target service address corresponding to the target scene identifier in the configuration storage module based on the target scene identifier, reducing the risk of sensitive information leakage. Finally, an API request is constructed based on the target API key and target service address to call the large model service. This solves the problem of low inference accuracy of traditional large models in vertical industry applications, achieving the beneficial effects of scenario-based dynamic key management and automated request construction, resulting in accurate adaptation, secure isolation, and significantly improved development efficiency for large model service calls.
[0023] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0024] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0025] Figure 1 This is a flowchart of a large model invocation method based on an operation scenario provided in Embodiment 1 of the present invention;
[0026] Figure 2 This is a flowchart of a large model invocation method based on an operation scenario provided by Embodiment 2 of the present invention;
[0027] Figure 3 This is a schematic diagram of a large model calling device based on an operation scenario according to Embodiment 3 of the present invention;
[0028] Figure 4 This is a schematic diagram of the structure of an electronic device that implements the large model invocation method based on the operation scenario in this embodiment of the invention. Detailed Implementation
[0029] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0030] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0031] Example 1
[0032] Figure 1 This is a flowchart of a large model invocation method based on an operation scenario, provided in Embodiment 1 of the present invention. This embodiment is applicable to large model invocation scenarios based on operation scenarios. The method can be executed by a large model invocation device based on an operation scenario. This device can be implemented in hardware and / or software and can be configured in an electronic device. Figure 1 As shown, the method includes:
[0033] S110. Collect user operation data in the integrated development environment through the context detection module, and determine the target scene identifier of the target operation scene corresponding to the operation data based on the preset scene recognition rules and the operation data.
[0034] The context detection module can be understood as, within an Integrated Development Environment (IDE), dynamically sensing the user's current working state by monitoring user actions (such as code editing, debugging, and running) and code context information (such as code structure, variable status, and project dependencies). Operation data can be understood as the specific records of user actions within the IDE, and scene recognition rules can be understood as predefined logical patterns used to map operation data to specific scenes. The target operation scene can be understood as the user's current working stage matched according to the rules. Scene identifiers can be understood as unique identifiers that identify the target scene; for example: explicit identifier: the interface prompts "Currently in debug mode"; implicit identifier: the metadata tag DEBUGGING_SCENE (used for subsequent analysis).
[0035] Specifically, user input events (such as key presses and mouse clicks) are captured through IDE plugins or hook functions. Code status (such as current file, cursor position, and variable values) and project configuration (such as dependencies and compilation options) are simultaneously acquired. For example, when a user presses F5 to run code, the operation time, code version, and execution log are recorded. A mapping relationship between scenarios and operation modes is defined (such as regular expressions and decision trees). The collected operation data is compared with a rule base to trigger scenarios that meet the conditions. For example, five consecutive failed runs plus adding a breakpoint activates the "Debug Scenario." If the operation data matches multiple rules, the most relevant scenario is selected based on scenario weight (such as frequency and time consumption). Scenario judgments are adjusted based on historical operation trends (such as switching from "Development" to "Test"). The scenario name is displayed in the IDE interface (such as a "Test Mode" notification in the status bar). Scenario identifiers are embedded in code metadata or log files for subsequent analysis, such as generating development behavior reports.
[0036] Optionally, the operation data includes at least one of the following: focus window, input command, and file type; the target operation scenario includes at least one of the following: code writing and chat consultation.
[0037] In this context, "Focused Window" can be understood as the window or panel currently active in the IDE (such as a code editor, terminal, or debugging console). "Input Command" refers to the specific instructions triggered by the user within the IDE via command line, menu, or keyboard shortcuts. "File Type" refers to the format of the file currently being manipulated. "Code Writing" refers to the scenario where the user focuses on writing or modifying source code. "Chat Consultation" refers to the scenario where the user consults with the IDE's built-in chat tools (such as an AI assistant).
[0038] Specifically, the operational data may include, but is not limited to: keyboard input (code writing, shortcut key usage), mouse operations (clicking, dragging, selecting), IDE function calls (debugging, refactoring, version control), and code change history (saving, undoing, branch switching). Target operational scenarios may also include code debugging, feature development, and bug fixing.
[0039] Optionally, the preset scene recognition rules include: when the focus window is a code editor and the input command is to generate code, the target operation scene is determined to be a code writing scene; or when the focus window is a chat window and the input command includes an explanation, the target operation scene is determined to be a chat consultation scene.
[0040] Here, "code generation" can be understood as the operation of automatically generating code snippets by the user entering specific commands (such as IDE shortcuts or AI assistant commands). "Including explanation" can be understood as the user's input commands containing explanatory text about the code's function, logic, or problem, for consultation or clarification.
[0041] Specifically, the system records the commands entered by the user in the IDE and identifies the user's current operating window (code editor or chat tool). If the focused window is the code editor and the entered command is "gen:loop", it matches the "code writing scenario". If the focused window is the chat window and the entered command contains "explanation", it matches the "chat consultation scenario".
[0042] S120. The key selection module searches for the target API key and target service address corresponding to the target scene identifier in the configuration storage module based on the target scene identifier.
[0043] The key selection module can be understood as the component responsible for dynamically selecting the corresponding API key based on the target scenario identifier. The configuration storage module can be understood as a database or configuration file that centrally stores the mapping relationship between scenarios, keys, and service addresses. The target API key can be understood as an authentication credential bound to the target scenario, used to call external services (such as AI code generation APIs). The target service address can be understood as the network access address of the external service (such as the API URL), determining the target location for sending requests.
[0044] Specifically, the key selection module obtains the target scenario identifier (such as a code writing scenario or a chat consultation scenario) passed from the context detection module. It then queries the configuration storage module using the scenario identifier as the key to find the corresponding configuration item, verifying the key's validity and checking if the API key has expired or been revoked to ensure call security. Finally, it returns the matched target API key and target service address to the caller (such as an IDE's code generation function or a chat tool).
[0045] Optionally, the configuration storage module adopts an encrypted storage mechanism, and the configuration storage module stores at least two mapping relationships between scenario identifiers and API keys and service addresses.
[0046] The encrypted storage mechanism can be understood as a measure to encrypt and protect sensitive data (such as API keys) in the configuration storage module using cryptographic techniques (such as AES symmetric encryption and RSA asymmetric encryption). The mapping relationship can be understood as the association rules between scenario identifiers and API keys / service addresses, defining the different service resources that need to be invoked in different scenarios.
[0047] Specifically, during initialization, the configuration storage module encrypts the stored API keys and service addresses (e.g., using the AES-256 encryption algorithm). The encryption key is managed by a security module (e.g., a hardware security module, HSM) to ensure that only authorized services can decrypt it. Administrators define mapping relationships for at least two scenarios (e.g., code writing, chat consultation) in the configuration storage module. Each mapping includes a scenario identifier, an encrypted API key, and an encrypted service address. When the key selection module queries the configuration, the system retrieves the encrypted mapping data based on the scenario identifier. The API key and service address are decrypted by the security module, ensuring that sensitive information only exists briefly in memory. Encryption keys and API keys are rotated periodically to avoid long-term exposure risks. Audit logs record all access and modification operations of scenario mappings and track abnormal behavior.
[0048] In this embodiment of the invention, by using encrypted storage and dynamic decryption, the system can ensure the security of sensitive data while achieving flexible adaptation to services in multiple scenarios.
[0049] Optionally, the step of using the key selection module to search for the target API key and target service address corresponding to the target scenario identifier in the configuration storage module based on the target scenario identifier includes:
[0050] If the target scenario identifier is a code writing identifier, then search the configuration storage module for the first API key and the first service address corresponding to the code writing identifier; or
[0051] If the target scenario is identified as a chat consultation identifier, the second API key and the second service address corresponding to the chat consultation identifier are searched in the configuration storage module.
[0052] The code writing identifier can be understood as an identifier marking the "code writing scenario," used to accurately match the API key and service address required for code generation in the configuration storage module. The chat consultation identifier can be understood as an identifier marking the "chat consultation scenario," corresponding to the low-privilege API key and service address required for natural language interaction. The first API key can be understood as a high-privilege key, used for scenarios requiring read / write operations such as code generation. The second API key can be understood as a low-privilege key, supporting only query or interpretation operations, enhancing security. The first service address can be understood as the interface of the code generation service. The second service address can be understood as the interface of the chat consultation service.
[0053] Specifically, the context detection module passes the determined scenario identifier to the key selection module. The key selection module uses `CODE_WRITING_SCENE` as the key to search for the corresponding first API key and first service address in the configuration storage module. Alternatively, it uses `CHAT_CONSULT_SCENE` as the key to search for the second API key and second service address. The keys and service addresses in the configuration storage module are stored in encrypted form and dynamically decrypted during queries, ensuring that sensitive information is only briefly exposed in memory. The decrypted API key and service address are then returned to the caller (such as the code generation function of an IDE or a chat tool), completing the service call preparation.
[0054] S130. Construct an API request based on the target API key and the target service address to invoke the large model service.
[0055] In this context, an API request can be understood as a structured data packet sent by a client (such as an IDE) to a server (large model service). This packet contains authentication information, operation instructions, and input content, used to trigger specific services (such as code generation or problem solving). A large model service can be understood as an AI service built upon a large language model, capable of handling natural language tasks (such as text generation, translation, and interpretation), and requires API calls.
[0056] Specifically, this can be understood as formatting the target API key according to the service requirements and encapsulating it in the request header. Based on the documentation requirements of the large model service, organize the input data (such as user-provided prompts) and parameters (such as generation length and temperature), send the request to the target service address using the POST method, and send the constructed request packet to the server over the network. After the server processes the request, it returns the result. Parse the JSON data returned by the server, extract the required content (such as the model-generated answer) or handle errors (such as invalid key or service timeout).
[0057] Optionally, constructing the API request based on the target API key and the target service address includes:
[0058] Convert the user-input operation data into a target format that conforms to the target API specification;
[0059] Embed the target API key in the request header, specify the target service address of the request as the target service address, and dynamically concatenate path parameters and query parameters;
[0060] Select the target request method according to the API specification, and bind the transformed data to the request body of the target request method to construct the API request.
[0061] The target API specification can be understood as the input and output standards defined by the large model service, including data format (e.g., JSON / XML), required parameters, and parameter types. Path parameters can be understood as the dynamic part of the URL used to locate resources. Query parameters can be understood as key-value pairs in the URL, used for additional conditions such as filtering, sorting, or pagination. The target request method can be understood as the operation type defined by the HTTP protocol. The request header can be understood as the metadata part of the HTTP request, containing authentication information (e.g., API key), data format (e.g., JSON), ensuring the server correctly parses the request. The request body can be understood as the core content part of the HTTP request, carrying the actual data to be transmitted (e.g., user-input text, model parameters). The HTTP method can be understood as the operation instruction defining the request type. Methods include POST, which submits data to the server (e.g., calling a large model to generate text), and GET, which retrieves data from the server (e.g., querying model information).
[0062] Specifically, the process involves converting user-input data (such as natural language questions) into the format required by the API (such as JSON objects) and filling in the required parameters. It embeds the target API key and declares the data format, concatenates the base address and path parameters, and appends query parameters. Based on the API documentation, it selects the appropriate method (e.g., POST for generating text, GET for querying model information). The POST method places the converted data in the request body, while the GET method appends parameters to the URL query string. The constructed request is sent over the network, and the server returns JSON data. The response body is parsed to extract the required content (such as generated text) or handle errors.
[0063] The technical solution of this invention collects user operation data in the integrated development environment through a context detection module. Based on preset scene recognition rules and the operation data, it determines the target scene identifier corresponding to the operation data. This enables scenario-based adaptation of large model service calls, avoiding insufficient generalization of a single model for multiple scenario tasks and improving the fit between output results and actual business needs. Then, the key selection module searches for the target API key and target service address corresponding to the target scene identifier in the configuration storage module based on the target scene identifier, reducing the risk of sensitive information leakage. Finally, an API request is constructed based on the target API key and target service address to call the large model service. This solves the problem of low inference accuracy of traditional large models in vertical industry applications, achieving the beneficial effects of scenario-based dynamic key management and automated request construction, resulting in accurate adaptation, secure isolation, and significantly improved development efficiency for large model service calls.
[0064] Example 2
[0065] Figure 2 This is a flowchart of a large model invocation method based on an operation scenario, provided in Embodiment 2 of the present invention. This embodiment is a further refinement of the above embodiment. Optionally, after invoking the large model service, the method further includes: receiving the response result data returned by the large model and returning the response result data to the user interface.
[0066] like Figure 2 As shown, the method includes:
[0067] S210. Collect user operation data in the integrated development environment through the context detection module, and determine the target scene identifier of the target operation scene corresponding to the operation data based on the preset scene recognition rules and the operation data.
[0068] S220. The key selection module searches for the target API key and target service address corresponding to the target scene identifier in the configuration storage module based on the target scene identifier.
[0069] S230. Construct an API request based on the target API key and the target service address to invoke the large model service.
[0070] S240: Receive the response result data returned by the large model and return the response result data to the user interface.
[0071] The response result data can be understood as the structured information returned by the large model after processing the user request, typically including generated text, answers, suggestions, or error messages. The user interface can be understood as the visual interface through which the user interacts with the system (such as a webpage, app interface, or IDE plugin window).
[0072] Specifically, the system continuously monitors the return data from the large model service via API until a complete response is received or a timeout occurs. The raw response (e.g., JSON format) is parsed into a program-readable data structure, key fields are extracted, and the data format is adjusted according to user interface requirements, such as adding line breaks, highlighting keywords, and displaying long text in blocks. If the response data contains errors (e.g., API rate limiting, invalid keys), a prompt is displayed in the user interface (e.g., "Service temporarily unavailable, please try again").
[0073] The technical solution of this invention receives the response result data returned by the large model and returns the response result data to the user interface, significantly improving the user interaction experience and task execution efficiency.
[0074] Example 3
[0075] Figure 3 This is a schematic diagram of a large model invocation device based on an operational scenario, provided in Embodiment 3 of the present invention. Figure 3 As shown, the device includes: a scene determination module 310, a key lookup module 320, and a model invocation module 330.
[0076] The scenario determination module 310 is used to collect user operation data in the integrated development environment through the context detection module, and determine the target scenario identifier corresponding to the operation data based on the preset scenario recognition rules and the operation data; the key lookup module 320 is used to look up the target API key and target service address corresponding to the target scenario identifier in the configuration storage module based on the target scenario identifier through the key selection module; and the model invocation module 330 is used to construct an API request based on the target API key and target service address to invoke the large model service.
[0077] The technical solution of this invention collects user operation data in the integrated development environment through a context detection module. Based on preset scene recognition rules and the operation data, it determines the target scene identifier corresponding to the operation data. This enables scenario-based adaptation of large model service calls, avoiding insufficient generalization of a single model for multiple scenario tasks and improving the fit between output results and actual business needs. Then, the key selection module searches for the target API key and target service address corresponding to the target scene identifier in the configuration storage module based on the target scene identifier, reducing the risk of sensitive information leakage. Finally, an API request is constructed based on the target API key and target service address to call the large model service. This solves the problem of low inference accuracy of traditional large models in vertical industry applications, achieving the beneficial effects of scenario-based dynamic key management and automated request construction, resulting in accurate adaptation, secure isolation, and significantly improved development efficiency for large model service calls.
[0078] Optionally, the operation data includes at least one of the following: focus window, input command, and file type; the target operation scenario includes at least one of the following: code writing and chat consultation.
[0079] Optionally, the preset scene recognition rules include:
[0080] When the focus window is the code editor and the input command is to generate code, the target operation scenario is determined to be a code writing scenario; or
[0081] If the focus window is a chat window and the entered command includes an explanation, the target operation scenario is determined to be a chat consultation scenario.
[0082] Optionally, the configuration storage module adopts an encrypted storage mechanism, and the configuration storage module stores at least two mapping relationships between scenario identifiers and API keys and service addresses.
[0083] Optionally, the model invocation module includes:
[0084] The data conversion unit is used to convert user-input operation data into a target format that conforms to the target API specification.
[0085] The parameter specification unit is used to embed the target API key in the request header, specify the target service address of the request as the target service address, and dynamically concatenate path parameters and query parameters;
[0086] The request construction unit is used to select the target request method according to the API specification and bind the transformed data to the request body of the target request method to construct the API request.
[0087] Optionally, the key lookup module includes:
[0088] The first data lookup unit is configured to, when the target scenario identifier is a code writing identifier, search in the configuration storage module for the first API key and the first service address corresponding to the code writing identifier; or
[0089] The second data lookup unit is used to look up the second API key and the second service address corresponding to the chat consultation identifier in the configuration storage module when the target scenario identifier is a chat consultation identifier.
[0090] Optionally, the device further includes:
[0091] The result feedback module is used to receive the response result data returned by the large model after calling the large model service, and return the response result data to the user interface.
[0092] The large model invocation device based on operation scenario provided in the embodiments of the present invention can execute the large model invocation method based on operation scenario provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0093] Example 4
[0094] Figure 4 A schematic diagram of an electronic device 10 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0095] like Figure 4 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 may also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0096] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0097] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 executes the various methods and processes described above, such as methods based on large model calls of operational scenarios.
[0098] In some embodiments, the method-based large model invocation based on the operational scenario can be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or mounted on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the method-based large model invocation based on the operational scenario described above can be performed. Alternatively, in other embodiments, processor 11 can be configured to perform the method-based large model invocation based on the operational scenario by any other suitable means (e.g., by means of firmware).
[0099] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0100] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0101] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0102] To provide interaction with a service recipient, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the service recipient; and a keyboard and pointing device (e.g., a mouse or trackball) through which the service recipient can provide input to the electronic device. Other types of devices can also be used to provide interaction with the service recipient; for example, feedback provided to the service recipient can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the service recipient can be received in any form (including voice input, speech input, or tactile input).
[0103] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., service-seeking computers with a graphical service-seeking interface or a web browser through which the service-seeking party can interact with the implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0104] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0105] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0106] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A method for invoking a large model based on an operational scenario, characterized in that, include: The context detection module collects user operation data in the integrated development environment, and determines the target scene identifier of the target operation scene corresponding to the operation data based on the preset scene recognition rules and the operation data. The key selection module searches for the target API key and target service address corresponding to the target scenario identifier in the configuration storage module based on the target scenario identifier. An API request is constructed based on the target API key and the target service address to invoke the large model service.
2. The method according to claim 1, characterized in that, The operation data includes at least one of the following: focus window, input command, and file type; the target operation scenario includes at least one of the following: code writing and chat consultation.
3. The method according to claim 2, characterized in that, The preset scene recognition rules include: When the focus window is the code editor and the input command is to generate code, the target operation scenario is determined to be a code writing scenario; or If the focus window is a chat window and the entered command includes an explanation, the target operation scenario is determined to be a chat consultation scenario.
4. The method according to claim 1, characterized in that, The configuration storage module adopts an encrypted storage mechanism and stores at least two mapping relationships between scenario identifiers, API keys, and service addresses.
5. The method according to claim 1, characterized in that, The step of constructing an API request based on the target API key and the target service address includes: Convert the user-input operation data into a target format that conforms to the target API specification; Embed the target API key in the request header, specify the target service address of the request as the target service address, and dynamically concatenate path parameters and query parameters; Select the target request method according to the API specification, and bind the transformed data to the request body of the target request method to construct the API request.
6. The method according to claim 1, characterized in that, The step of searching for the target API key and target service address corresponding to the target scenario identifier in the configuration storage module based on the target scenario identifier by the key selection module includes: If the target scenario identifier is a code writing identifier, then search the configuration storage module for the first API key and the first service address corresponding to the code writing identifier; or If the target scenario is identified as a chat consultation identifier, the second API key and the second service address corresponding to the chat consultation identifier are searched in the configuration storage module.
7. The method according to claim 1, characterized in that, After calling the large model service, the following is also included: Receive the response data returned by the large model and return the response data to the user interface.
8. A large model invocation device based on an operational scenario, characterized in that, include: The scene determination module is used to collect user operation data in the integrated development environment through the context detection module, and determine the target scene identifier of the target operation scene corresponding to the operation data based on the preset scene recognition rules and the operation data. The key lookup module is used to search for the target API key and target service address corresponding to the target scenario identifier in the configuration storage module based on the target scenario identifier, through the key selection module. The model invocation module is used to construct an API request based on the target API key and the target service address to invoke the large model service.
9. An electronic device, characterized in that, include: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the large model invocation method based on any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the large model invocation method based on any one of claims 1-7.